RemoteProject-based / Advisory
AI Engineer
Production LLM/RAG pipelines, prompting, retrieval, eval frameworks. Fix wrong answers and hallucinations with measurable benchmarks.
About the role
We're looking for an AI Engineer to join our team on production AI recovery engagements. You'll work on real LLM/RAG systems—wrong answers, hallucinations, high cost, slow latency—and ship fixes with before/after benchmarks. Audit-first, evidence-driven.
Responsibilities
- •Troubleshoot production LLM/RAG pipelines end-to-end—prompting, retrieval, reranking, tool calls
- •Build and maintain eval frameworks, golden sets, and groundedness checks
- •Ship PRs with measurable impact on cost/conversation, quality score, TTFT
- •Collaborate with clients on audit baselines and recovery roadmaps
Requirements
- •Hands-on experience with production LLM/RAG systems
- •Comfortable with Python, eval frameworks (e.g. promptfoo, ragas), and observability
- •Evidence-based mindset—prove improvements with metrics
- •Strong written communication for async collaboration
Nice to have
- •Experience with retrieval optimization (chunking, embeddings, reranking)
- •Familiarity with cost optimization (routing, caching, context trimming)
What we offer
- •Remote-first, async-friendly
- •Work on cutting-edge AI—new models, new techniques every day
- •Project-based—no long-term lock-in, join what excites you
Ready to apply?
No formal application. Send us a short intro, relevant work, and how you'd like to collaborate.