OCR, VLMs, RAG & agents

Software Engineer, AI Data Systems

Build reliable pipelines and agentic systems that turn difficult source material into traceable, production-grade intelligence.

Engagement Part-time contractor Location Remote worldwide Schedule Scope agreed together

What you will do

  • Build and improve OCR and vision-language-model pipelines for laboratory reports, pricelists, screenshots, PDFs, and semi-structured documents.
  • Design extraction, canonicalization, entity-resolution, validation, provenance, and human-review workflows.
  • Scale ingestion with queues, batching, retries, idempotency, caching, observability, and explicit failure handling.
  • Develop retrieval and RAG systems that return relevant evidence with precise citations and visible limitations.
  • Optimize LLM workflows across quality, latency, cost, model choice, prompting, context construction, and structured outputs.
  • Build agent-driven automation with constrained tools, durable state, evaluations, approval boundaries, and audit trails.
  • Create reusable scaffolding for services, jobs, evaluations, migrations, monitoring, and operational runbooks.
  • Ship safely in a fast-paced environment through focused tests, staged rollouts, and reversible changes.

What we are looking for

  • Strong production experience with Python, SQL, APIs, backend services, and data pipelines.
  • Hands-on experience with OCR, document AI, multimodal or VLM extraction, or comparable unstructured-data systems.
  • Understanding of scalable systems including queues, concurrency, batching, caches, retries, idempotency, and observability.
  • Experience building retrieval, search, vector, or RAG systems and evaluating answer quality.
  • Practical experience optimizing LLM systems for reliability, latency, and cost rather than demos alone.
  • Ability to design agentic workflows with explicit permissions, human review, and failure recovery.
  • Strong product judgment and comfort working from incomplete information without hiding uncertainty.
  • No formal degree is required; relevant shipped systems and applied judgment matter more.

Useful additions

  • Experience with document provenance, entity resolution, knowledge graphs, or temporal data.
  • Familiarity with browser automation, resilient web ingestion, PDF processing, or image-quality diagnostics.
  • Experience with security, privacy, infrastructure operations, and production incident response.
  • Peptide, laboratory-data, life-sciences, marketplace, or regulated-data domain experience.

What early progress looks like

  • Establish an evaluated benchmark for a high-value OCR or VLM extraction workflow.
  • Improve extraction coverage and accuracy while preserving source traceability.
  • Productionize an agentic workflow with observable state, safe approval boundaries, and measurable time savings.
  • Reduce a documented reliability, latency, or operating-cost bottleneck in ingestion or retrieval.
  • Leave reusable scaffolding and operational documentation behind each shipped system.