- Salary: $400k - $600k
- Locations: Austin, Chicago, New York
- Job Type: Full Time
- Job Category: Networking
Description
We are seeking a senior Network Engineer to design and operate ultra-high-performance network fabrics for large-scale GPU clusters used in trading research, model training, and inference. This role is ideal for network engineers from leading technology companies or hyperscalers who have scaled GPU/AI infrastructure and are ready to apply that expertise to one of the most performance-critical environments in finance.
Key Responsibilities
- Design and deploy high-performance network architectures for GPU clusters, including InfiniBand (HDR/NDR), RoCE, and high-speed Ethernet (100/200/400GbE).
- Optimize network topology, routing, and congestion control to maximize GPU utilization and minimize job completion times across thousands of accelerators.
- Ensure low-latency, high-throughput connectivity between GPU clusters, storage systems, and trading infrastructure across multiple data centers.
- Build comprehensive observability and telemetry for GPU network performance: bandwidth utilization, latency, packet loss, ECN marks, queue depths, and buffer utilization.
- Automate network provisioning, configuration management, validation, and rollback using Python, Go, or similar, integrated with CI/CD pipelines and orchestration platforms (Kubernetes, Slurm, or equivalent).
Required Skills
- 5+ years of experience in network engineering, with at least 2+ years focused on HPC, AI/ML, or large-scale GPU infrastructure at hyperscalers, cloud providers, or major technology firms.
- Deep expertise in high-performance networking technologies: InfiniBand (HDR/NDR), RoCE, GPUDirect, NCCL tuning, and high-speed Ethernet.
- Strong understanding of distributed training workloads, collective communication patterns (all-reduce, all-gather), and their network requirements.
- Proficiency with Linux networking and performance tooling: tcpdump, perf, eBPF, ethtool, traffic shaping, and performance debugging.
- Experience with congestion control, adaptive routing, and QoS in GPU/HPC environments.
- Solid scripting and automation capabilities in Python, Go, or equivalent;
Apply Today
Thank you for your interest in this opportunity. Please complete the form below and upload any relevant documents. A member of our team will review your application and be in touch soon.