Back to Careers
RemoteProject-based / Advisory

AI Performance & Cost Engineer

Inference cost optimization, TTFT/P95 latency, routing, caching. Token economics and pipeline optimization.

About the role

Our clients struggle with inference cost spikes and slow P95 latency. We need an engineer who can optimize token spend, routing, caching, context trimming, and pipeline bottlenecks. You'll prove every improvement with cost/conversation and TTFT metrics.

Responsibilities

  • Optimize inference cost—routing, caching, context budgets, model selection
  • Reduce TTFT and P95 latency—pipeline tuning, timeouts, throughput
  • Build cost and latency baselines, then ship fixes with before/after proof
  • Collaborate on audit reports and ROI roadmaps

Requirements

  • Experience with LLM serving, token economics, and cost optimization
  • Comfortable with latency analysis—P50/P95/P99, timeouts, bottlenecks
  • Evidence-based—measure before and after, no promises without proof
  • Python and observability tooling (metrics, traces)

Nice to have

  • Experience with model routing, speculative decoding, or batching
  • Familiarity with cloud cost attribution for AI workloads

What we offer

  • Remote-first, async-friendly
  • Work on high-impact optimizations—50%+ cost cuts are common
  • Project-based—flexible, join projects that fit

Ready to apply?

No formal application. Send us a short intro, relevant work, and how you'd like to collaborate.