Skip to main content
Remote Atlas
ashbyOfficial ATS boardOnsiteSeniorFullTime

Staff Network Production Engineer, Deployment

Crusoe

San Francisco, CA - US

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About This Role:

Crusoe Cloud is seeking a high-energy, detail-oriented Senior Network Production Engineer to support the physical and logical implementation of our global network. As we rapidly expand our footprint of high-performance compute (HPC) and GPU-based AI infrastructure, this role plays a key part in bringing new data centers and edge sites online, directly enabling the scale and reliability our AI cloud customers depend on.

This is a technical, hands-on, full-time role that sits at the intersection of Network Engineering, Data Center Operations, and Project Management. The ideal candidate is comfortable both on-site and remote, follows and helps refine deployment standards, and takes ownership of ensuring every switch, router, and fiber optic link is deployed to spec, validated, and seamlessly handed off to our Operations team.

 

What You'll Be Working On:

  • Network Deployment Execution: Execute on-site and remote network deployments for new data center and edge site builds, from rack-and-stack through final cutover.

  • Implementation Planning: Implement architecture designs into concrete deployment steps: cable maps, port assignments, device configs, and site-specific runbooks.

  • Automation Development: Write and maintain Python/Ansible automation and ZTP workflows to stage, configure, and validate batches of switches and routers without manual touch.

  • Acceptance Testing: Support burn-in and site acceptance testing (SAT) on new clusters, chase down failures, and help drive clean handoffs to Operations.

  • Fabric Configuration & Troubleshooting: Configure and troubleshoot Arista, Juniper, and NVIDIA/Mellanox gear in leaf-spine fabrics, including BGP, EVPN-VXLAN, and LLDP issues at the link and fabric level.

  • Vendor & Physical Layer Coordination: Work directly with structured cabling vendors, remote hands, and data center providers on-site to resolve physical layer issues — bad fiber runs, power/cooling deviations, mislabeled patch panels — before they block turn-up.

  • Physical/Link-Layer Diagnostics: Diagnose physical and link-layer problems using OTDRs, light meters, and packet captures when something doesn't come up clean.

  • Inventory & Capacity Tracking: Track hardware inventory and support turn-up of new backbone and edge interconnect capacity against deployment schedules.

  • Cross-Team Collaboration: Collaborate with other engineers on tricky sites or configs, and flag recurring issues back to the team so they get addressed at the process or automation level.

  • On-Call Support: Participate in an on-call rotation, responding to and troubleshooting network incidents affecting production infrastructure.

 

What You'll Bring to the Team:

  • 5+ years of experience in network engineering with a focus on large-scale data center deployments and infrastructure projects.

  • Strong knowledge of physical layer standards: solid experience with structured cabling (SMF/MMF, MPO/MTP), optical transceivers (400G/800G), and data center power/cooling requirements.

  • Solid routing and switching knowledge: hands-on experience configuring Arista (EOS), Juniper (Junos), and NVIDIA/Mellanox platforms in a leaf-spine architecture.

  • Protocol familiarity: working understanding of BGP, EVPN-VXLAN, and LLDP as they relate to large-scale fabric provisioning.

  • Automation-minded: proficiency in Python and Ansible for automating repetitive deployment tasks and validating configuration state.

  • Logistical skills: ability to manage multiple projects simultaneously across different time zones and physical locations.

  • Troubleshooting skills: ability to diagnose physical layer and link-layer issues using OTDRs, light meters, and packet captures.

  • Education: Bachelor's degree in a technical field or equivalent practical experience in hyperscale or ISP environments.

 

Bonus Points:

  • Experience working in hyperscale, ISP, or large multi-tenant data center environments.

  • Familiarity with GPU cluster networking (e.g., RDMA/RoCE, InfiniBand, or NVIDIA NCCL-aware fabric design).

  • Exposure to network monitoring and observability tooling (e.g., Prometheus, Grafana, telemetry-based fabric health checks).

  • Vendor certifications such as CCNP, JNCIP, or Arista ACE.

  • Prior experience mentoring junior engineers or contributing to deployment standards/documentation.

 

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

 

Compensation Range

Compensation will be paid in the range of up to $165,000 - $200,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Apply kit

Sign in to copy a field card for the employer’s ATS. We never submit applications for you.

arbeitnowCurated job boardUnknownSenior

Member of Technical Staff – AI Inference platform

Lyceum

Zürich

Atlas fit 6

Your mission You will make Lyceum's AI inference platform reliable, secure, and scalable - ensuring it performs under pressure as we grow to thousands of concu…

  • ai/ml
  • devops
  • engineering
  • go
  • python

Sign in to track applications

Details
arbeitnowCurated job boardRemoteSenior

Engineering Manager, Platform

Firmus

Sydney, Australia

Atlas fit 5

Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure from model to grid. Founded…

  • ai/ml
  • cloud
  • devops
  • go
  • platform engineering

Sign in to track applications

Details
arbeitnowCurated job boardUnknownSenior

Senior iOS Developer - Berlin

Resident Advisor

Berlin, Germany

Atlas fit 4

About RA Resident Advisor has been a home for electronic music culture since 2001. Event discovery, ticketing, news and long-form editorial, artist profiles, o…

  • ai/ml
  • android
  • engineering
  • epd
  • graphql

Sign in to track applications

Details
ashbyOfficial ATS boardRemoteSenior

Staff Software Engineer - Postgres Control Plane

Snowflake

US-WA-Bellevue

Atlas fit 3

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized b…

  • ai/ml
  • cloud
  • go
  • java
  • postgresql

Sign in to track applications

Details