Software Engineer (Model Performance Systems)

Auto Import

  • We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering
  • You will not just be building the automated “speedometer and diagnostic” suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work
  • Benchmarking: Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific worklo (e.g., long-context window, KV cache reuse, disaggregated serving)
  • DevEx Improvement: Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance “dev machines” optimized for model experimentation
  • Tool Development: Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis
  • System Profiling: Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack
  • Monitoring & Observability: Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance
  • Continuous Integration: Automate performance testing via CI/CD pipelines to catch regressions and build release workflow automation for the model runtimes stack
  • Optimization Automation: Build tools to find the “Pareto frontier”—identifying the absolute best configuration (latency vs. cost vs. quality) for a given model and workload

Benefits

  • Remote-first work environment. The Baseten team is welcome to work from wherever they want; fully remote, in our San Francisco office, or a mix of both. Today, our team (including our founding team) is spread across the United States, Canada, and Armenia. We provide a $1,000 stipend for you to make your home-office comfortable and productive
  • Regular in-person team summits. We get together as a team three times a year to plan, workshop, and most importantly, get to know each other better
  • Unlimited PTO. We ask that everyone take at least 4 weeks of vacation. And we have a company-wide break between Christmas and New Year’s Day
  • Full healthcare coverage. Medical, dental and vision insurance for you and your family
  • Paid parental leave. 16-weeks fully paid parental leave (adoptive and non-birth parents included) and flexibility with schedules while returning to work
  • Company-sponsored 401(k) for you to contribute to
  • Learning and development budget. We encourage you to take classes, attend conferences, and invest in your craft and we’ll cover expenses to make it happen
Apply Now →