Osama Moftah

Applied AI engineer · document intelligence, agent reliability · Berlin

I build reliable AI and data systems that turn uncertain inputs into decisions people can trust: evaluation gates, retrieval checks, entity resolution, and drift monitoring that make outputs safe to use.

My work spans document intelligence, agents, and production data systems: turning fragmented, changing information into traceable workflows with clear schemas, measurable evaluation, and evidence behind every result.

I work across extraction, retrieval, agents, and production data systems using Python, PyTorch,LangGraph, and RAGAS-style evaluation. I define failure states and evaluation metrics before optimizing the model, with the goal of building systems that can explain what they did and fail safely when evidence is weak.

What I do

Document intelligence

Turn fragmented public documents into traceable records with extraction, entity resolution, and provenance behind every result.

Prospecto · HF Agentic Search

Evaluation and governance

Make model quality a testable, monitored claim with golden sets, adversarial cases, citations, and regression gates.

Compliance Agent · Eval-Gated Deploy Pipeline

Reliable agent workflows

Design multi-step systems where state, evidence, and human review survive every handoff.

HF Agentic Search · DiffDDL

Side projects

Small end-to-end systems I build to find out where an idea breaks in practice.

see the desk

Proofall work →

ProjectShippedTestedPublic
HF Agentic Search: Evidence-Led Dataset DiscoveryAgent orchestration · 2 test modules · live Gradio Space you can run · 24 commits over a month Shipped Tested Public
Compliance Agent: Explainable Policy CopilotRetrieval & explainable screening · 9 test modules · GitHub Actions · golden-sample test Shipped Tested Public
LegalDrift: Semantic Version Control for LawStatistical drift detection · 9 test modules · GitHub Actions · published to PyPI 0.1.0 Shipped Tested Public
DiffDDL: Semantic Schema MigrationsSchema safety for agents · 1 test module · GitHub Actions Shipped Tested Public
Procurement Intelligence at National ScaleEntity resolution at scale · In production, no public repo Shipped Tested Public

Every row links to the thing that backs it. A blank cell is a blank cell — Prospecto is shipped and running, but has no public repo, so “public” stays empty rather than implied.

Open to the right Applied AI problem

I’m open to full-time Applied AI/ML roles and a small number of technical advisory engagements around evaluation, document intelligence, and agent reliability. I also run AI architecture and alignment workshops for teams that need to surface assumptions before implementation.