Skip to main content
Remote Atlas

Job description

Role Overview

We are seeking an AI Researcher with deep experience in inference optimization to design, evaluate, and deploy high-performance inference systems for large-scale machine learning models. You will work at the intersection of model architecture, systems engineering, and hardware-aware optimization, improving latency, throughput, and cost efficiency across real-world production environments.

Key Responsibilities

  • Research and develop techniques to optimize inference performance for large neural networks.

  • Improve latency, throughput, memory efficiency, and cost per inference.

  • Design and evaluate model-level optimizations (quantization, pruning, KV-cache optimization, architecture-aware simplifications).

  • Implement systems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization).

  • Benchmark inference workloads across hardware accelerators.

  • Collaborate with engineering teams to deploy optimized inference pipelines.

  • Translate research insights into production-ready improvements.

Required Qualifications

  • Strong background in machine learning, deep learning, or AI systems.

  • Hands-on experience optimizing inference for large-scale models.

  • Proficiency in Python and modern ML frameworks (e.g., PyTorch).

  • Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime).

  • Ability to design experiments and communicate results clearly.

Preferred / Nice-to-Have Qualifications

  • Experience deploying production inference systems at scale.

  • Familiarity with distributed and multi-GPU inference.

  • Experience contributing to open-source ML or inference frameworks.

  • Authorship or co-authorship of peer-reviewed research papers in machine learning, systems, or related fields.

  • Experience working close to hardware (CUDA, ROCm, profiling tools).

What Success Looks Like

  • Measurable gains in latency, throughput, and cost efficiency.

  • Optimized inference systems running reliably in production.

  • Research ideas successfully translated into deployable systems.

  • Clear benchmarks and documentation that inform product decisions.

Relevant Research Areas (Bonus)

  • Long-context inference optimization

  • Speculative decoding

  • KV-cache compression and paging

  • Efficient decoding strategies

  • Hardware-aware inference design

Originally posted on Himalayas

Apply kit

Sign in to copy a field card for the employer’s ATS. We never submit applications for you.

arbeitnowCurated job boardRemoteSeniority not stated

AI Scientist

Mistral.ai

Zurich

Atlas fit 6

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tacklin…

  • ai/ml
  • devops
  • go
  • java
  • python

Sign in to track applications

Details
arbeitnowCurated job boardOnsiteSeniority not stated

AI Research Engineer

Sonarsource

Geneva

Atlas fit 5

Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code verification and governance, we solve a critical prob…

  • ai/ml
  • cloud
  • data & agentic
  • java
  • python

Sign in to track applications

Details
arbeitnowCurated job boardHybridSenior

Senior Software Engineer (Java) - Remediation Agent

Sonarsource

Geneva

Atlas fit 4

Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code review and verification, we solve a critical problem:…

  • agentic sdlc
  • ai/ml
  • cloud
  • java
  • python

Sign in to track applications

Details
arbeitnowCurated job boardOnsiteSeniority not stated

AI Researcher - Post-Training

Sonarsource

Geneva

Atlas fit 3

Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code verification and governance, we solve a critical prob…

  • ai/ml
  • data & agentic
  • java
  • javascript
  • python

Sign in to track applications

Details