OCR, VLMs, RAG & agents
Software Engineer, AI Data Systems
Build reliable pipelines and agentic systems that turn difficult source material into traceable, production-grade intelligence.
What you will do
- Build and improve OCR and vision-language-model pipelines for laboratory reports, pricelists, screenshots, PDFs, and semi-structured documents.
- Design extraction, canonicalization, entity-resolution, validation, provenance, and human-review workflows.
- Scale ingestion with queues, batching, retries, idempotency, caching, observability, and explicit failure handling.
- Develop retrieval and RAG systems that return relevant evidence with precise citations and visible limitations.
- Optimize LLM workflows across quality, latency, cost, model choice, prompting, context construction, and structured outputs.
- Build agent-driven automation with constrained tools, durable state, evaluations, approval boundaries, and audit trails.
- Create reusable scaffolding for services, jobs, evaluations, migrations, monitoring, and operational runbooks.
- Ship safely in a fast-paced environment through focused tests, staged rollouts, and reversible changes.
What we are looking for
- Strong production experience with Python, SQL, APIs, backend services, and data pipelines.
- Hands-on experience with OCR, document AI, multimodal or VLM extraction, or comparable unstructured-data systems.
- Understanding of scalable systems including queues, concurrency, batching, caches, retries, idempotency, and observability.
- Experience building retrieval, search, vector, or RAG systems and evaluating answer quality.
- Practical experience optimizing LLM systems for reliability, latency, and cost rather than demos alone.
- Ability to design agentic workflows with explicit permissions, human review, and failure recovery.
- Strong product judgment and comfort working from incomplete information without hiding uncertainty.
- No formal degree is required; relevant shipped systems and applied judgment matter more.
Useful additions
- Experience with document provenance, entity resolution, knowledge graphs, or temporal data.
- Familiarity with browser automation, resilient web ingestion, PDF processing, or image-quality diagnostics.
- Experience with security, privacy, infrastructure operations, and production incident response.
- Peptide, laboratory-data, life-sciences, marketplace, or regulated-data domain experience.
What early progress looks like
- Establish an evaluated benchmark for a high-value OCR or VLM extraction workflow.
- Improve extraction coverage and accuracy while preserving source traceability.
- Productionize an agentic workflow with observable state, safe approval boundaries, and measurable time savings.
- Reduce a documented reliability, latency, or operating-cost bottleneck in ingestion or retrieval.
- Leave reusable scaffolding and operational documentation behind each shipped system.