Safe Invest AI Lead Automation
Multilingual LLM lead pipeline
Messages prospects send on WhatsApp, Instagram or Messenger, in French, English or Darija, are turned into structured CRM leads automatically.
- Organisation
- Safe Invest Property · Marrakech
- Role
- AI Engineering Intern · Sole developer
- Timeline
- July – August 2026 · 9 weeks
- Languages
- French · English · Darija
Built during an internship. The agency’s code, data and conversations are confidential and not shown here.
- Extraction accuracy
- 98%
- How often the extracted lead data matched the annotated reference, measured on 96 messages in French, English and Darija.
- Fewer LLM calls
- ≈4×
- The drop in LLM call volume once messages sent in quick succession were grouped before extraction.
- End-to-end scenarios
- 13/14
- Full runs from incoming message to CRM lead, validated after voice-note transcription and lead↔property matching were added.
- CRM fields per lead
- 9
- The fields in each CRM record built from a conversation. 3 of them are computed in code rather than by the model.
01Context
Safe Invest Property is a real-estate agency in Marrakech whose prospects write in through WhatsApp, Instagram and Messenger, in French, English and Moroccan Darija. During a 9-week internship I was the sole developer of the system that turns those conversations into CRM leads.
02Problem
Inquiries arrive as free-form conversations: several short messages in a row, voice notes, three languages. Each one has to become a complete, consistent lead record, without creating duplicates when a platform delivers the same webhook twice.
03Architecture
01Channels
WhatsApp · Instagram · Messenger
Inbound webhooks
02Intake
Idempotent ingestion
Duplicate webhooks absorbed by DB constraints
02Buffer
Burst debouncing
One LLM call per burst
03Voice
Voice-note transcription
Audio to text
04LLM
Structured extraction
Claude API · forced tool use
98% accuracy
04Code
Deterministic fields
Computed, not generated
05Data
PostgreSQL
5-table schema
06Match
Lead ↔ property matching
06Output
Structured CRM lead
04My contribution
Sole developer, end to end: pipeline architecture, the PostgreSQL data model, LLM extraction and its evaluation, voice-note transcription, the lead-to-property matching engine and the production deployment.
05Technical decisions
- 01
Structured output, not free text
The LLM returns each lead through a forced tool call with a fixed schema, so every response maps straight onto CRM fields instead of being parsed out of prose.
- 02
Deterministic logic stays in code
Fields that follow fixed rules are computed in code rather than generated by the model, which removes a whole class of errors by design.
- 03
Idempotency in the database
Constraints in the 5-table PostgreSQL schema absorb duplicate webhook deliveries, so a repeated delivery doesn’t create a second lead.
- 04
Debounce message bursts
Prospects often send several short messages in a row. They are grouped before extraction, so a burst costs one LLM call instead of one per message.
06Results
- Deployed to production: conversations from three messaging channels become CRM leads automatically
- One pipeline for French, English and Moroccan Darija
- Voice notes transcribed and leads matched to the agency’s properties
- Duplicate webhook deliveries absorbed in production by idempotency constraints
07Stack
- Claude API
- Structured outputs
- PostgreSQL
- Webhooks
- WhatsApp · Instagram · Messenger
08Source code
The source code isn’t public: it was written during an internship and stays private. Other projects are on GitHub.
Building something similar?
Start a project