View all jobs

AI Engineer / LLM Systems Engineer

  • Remotely, Anywhere

About the Role

We are looking for an AI Engineer to lead the development of AI-first products and solutions. This role combines advanced LLM engineering, system architecture, and high-performance inference optimization.

The ideal candidate has strong software engineering fundamentals and hands-on experience building production-grade AI systems, including advanced RAG pipelines, multi-agentic workflows, and LLM inference infrastructure. This person will take ownership of complex technical initiatives, work directly with senior stakeholders and strategic partners, and drive AI solutions from concept to production.

Responsibilities

  • Lead the development and delivery of high-priority AI-first features and products, often working under tight deadlines.
  • Drive technical implementation of AI-first solutions in collaboration with major partners and clients.
  • Design system architecture for AI-powered products, including advanced LLM applications, RAG pipelines, and multi-agentic workflows.
  • Validate and improve existing AI products based on user feedback and evolving business requirements.
  • Define and manage technical scope for AI initiatives in collaboration with C-level executives, marketing, directors, and other senior stakeholders.
  • Validate, benchmark, and optimize LLM inference infrastructure.
  • Conduct load and performance testing, identify bottlenecks, and implement inference optimization improvements.
  • Optimize model serving performance through batching, quantization, and inference engine configuration.
  • Monitor and analyze inference performance using metrics such as TTFT, TPOT, and ITL.

Requirements

  • Strong software engineering background with substantial hands-on experience building production AI/LLM systems.
  • Python programming skills.
  • Practical experience with LLM engineering, RAG, vector search, and multi-agentic workflows.
  • Experience designing and implementing scalable AI system architectures.
  • Experience with LLM inference frameworks such as vLLM, TensorRT-LLM, and/or Triton Inference Server.
  • Understanding of LLM inference optimization, including batching, quantization, and performance tuning.
  • Experience benchmarking and troubleshooting LLM inference performance.
  • Understanding of inference performance metrics, including TTFT, TPOT, and ITL.
  • Strong system design and problem-solving skills.
  • Ability to independently drive complex technical initiatives and communicate effectively with senior stakeholders.

Nice to Have

  • Experience with Golang and Bash scripting.
  • Experience with GPU-based deployments and cloud infrastructure.
  • Experience working with self-hosted and open-weight models.
  • Experience optimizing inference infrastructure for high-throughput and low-latency workloads.
  • Previous experience leading AI-first product development or strategic technical collaborations.

Engagement Details

  • Remote
  • Contract / Staff Augmentation