At a glance
Build and optimize the infrastructure that runs large-scale model evaluations, from LLM inference and GPU scheduling to orchestration and internal tooling for next-generation models.
Summarized by AI from the original posting
What you'll do
- Build reliable, efficient, user-friendly infrastructure for large-scale evaluations
- Collaborate with modeling and evaluation teams to develop evals and maintain evaluation signal quality
- Build internal tooling to support training next-generation models
- Identify and resolve performance bottlenecks across inference, capacity fleet management, and asynchronous evaluation orchestration
Requirements
- 01Experience building, debugging, and optimizing large-scale distributed systems
- 02Willingness to solve complex problems across the technology stack
- 03Experience in LLM inference
- 04Experience in GPU compute management, workload scheduling, and dynamic resource optimization
- 05Experience developing interfaces for comparing models and evaluations
Perks
Skills
Full description
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
The RL infrastructure team is looking for an engineer to help develop our evaluation infrastructure.
RESPONSIBILITIES:
- Build highly reliable, efficient, and easy to use infrastructure that runs all our evaluations at scale
- Collaborate closely with modeling & evaluation teams on developing new evals, maintaining evaluation signal quality, and building internal tooling to support training our next-generation of models
- Identify and resolve performance bottlenecks in all layers of the eval infrastructure stack, including inference, capacity fleet management, asynchronous eval orchestration, and more
BASIC QUALIFICATIONS:
- Experience in building, debugging, and optimizing efficiency of large-scale distributed systems
- Willingness to dive deep and solve hard problems at all levels of the stack
PREFERRED SKILLS AND EXPERIENCE:
- Experience in LLM inference
- Experience in GPU compute management, workload scheduling, and dynamic resource optimization
- Experience in developing interfaces for comparing models and evals
COMPENSATION AND BENEFITS:
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.