Osama Moftah
Applied AI engineer · document intelligence, agent reliability · Berlin
I build reliable AI and data systems that turn uncertain inputs into decisions people can trust: evaluation gates, retrieval checks, entity resolution, and drift monitoring that make outputs safe to use.
My work spans document intelligence, agents, and production data systems: turning fragmented, changing information into traceable workflows with clear schemas, measurable evaluation, and evidence behind every result.
I work across extraction, retrieval, agents, and production data systems using Python, PyTorch,LangGraph, and RAGAS-style evaluation. I define failure states and evaluation metrics before optimizing the model, with the goal of building systems that can explain what they did and fail safely when evidence is weak.
What I do
Document intelligence
Turn fragmented public documents into traceable records with extraction, entity resolution, and provenance behind every result.
Prospecto · HF Agentic SearchEvaluation and governance
Make model quality a testable, monitored claim with golden sets, adversarial cases, citations, and regression gates.
Compliance Agent · Eval-Gated Deploy PipelineReliable agent workflows
Design multi-step systems where state, evidence, and human review survive every handoff.
HF Agentic Search · DiffDDLSide projects
Small end-to-end systems I build to find out where an idea breaks in practice.
see the deskProofall work →
| Project | Shipped | Tested | Public |
|---|---|---|---|
| HF Agentic Search: Evidence-Led Dataset DiscoveryAgent orchestration · 2 test modules · live Gradio Space you can run · 24 commits over a month Shipped Tested Public | |||
| Compliance Agent: Explainable Policy CopilotRetrieval & explainable screening · 9 test modules · GitHub Actions · golden-sample test Shipped Tested Public | |||
| LegalDrift: Semantic Version Control for LawStatistical drift detection · 9 test modules · GitHub Actions · published to PyPI 0.1.0 Shipped Tested Public | |||
| DiffDDL: Semantic Schema MigrationsSchema safety for agents · 1 test module · GitHub Actions Shipped Tested Public | |||
| Procurement Intelligence at National ScaleEntity resolution at scale · In production, no public repo Shipped Tested Public |
Every row links to the thing that backs it. A blank cell is a blank cell — Prospecto is shipped and running, but has no public repo, so “public” stays empty rather than implied.
Open to the right Applied AI problem
I’m open to full-time Applied AI/ML roles and a small number of technical advisory engagements around evaluation, document intelligence, and agent reliability. I also run AI architecture and alignment workshops for teams that need to surface assumptions before implementation.