Hyphen Connect Limited
LLM Pre-training & Distributed Engineer (AI Infrastructure)
Be an Early Applicant
Lead orchestration and optimization of large-scale LLM pretraining across 1,000+ GPUs. Manage distributed training with PyTorch/DeepSpeed/Megatron-LM, tune networking and memory (InfiniBand/RDMA), and implement checkpointing and robust failure recovery for long-running jobs.
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar Jobs
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Build and own shared Go libraries and platform capabilities for data access, messaging, service communication, observability, resilience, and multi-cloud portability. Design adoptable APIs, implement resilience primitives, ensure security of dependencies, participate in architecture governance, consult with teams, operate libraries (on-call), and partner with Data Services, Infrastructure, SRE, and Observability teams.
Top Skills:
Circuit BreakersData StoresDistributed TracingFeature ManagementGoLoad SheddingMessage BrokersMetrics PipelinesMulti-CloudRate LimitingResilience PatternsSdksService CommunicationStructured Logging
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Build and maintain shared Go libraries and platform capabilities for data access, messaging, service communication, observability, and resilience. Design adoptable APIs, multi-cloud abstractions, and resilience primitives. Own security posture, operate libraries (on-call), consult engineering teams, participate in architectural governance, and collaborate with Data Services, Infrastructure, SRE, and Observability teams.
Top Skills:
Data StoresDistributed TracingGoMessage BrokersMetrics PipelinesMulti-CloudStructured Logging
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Deliver technical product presentations and demos, configure installations during POVs, liaise with Product Management for account requirements, train and support Sales on technical and competitive messaging, act as a trusted advisor to prospects/customers to ensure project success, and travel as required.
Top Skills:
AIAvAWSAzureBashCloud SecurityCrowdstrikeEdrFirewallForensicsGCPHipsIdentity PreventionIdsIncident ResponseLinuxmacOSMdrOsi ModelPowershellPythonSIEMVdiVirtualizationWindowsXdrZero-Trust
What you need to know about the Sydney Tech Scene
From opera to comedy shows, the Sydney Opera House hosts more than 1,600 performances a year, yet its entertainment sector isn't the only one taking center stage. The city's tech sector has earned a reputation as one of the fastest-growing in the region. More specifically, its IT sector stands out as the country's third-largest, growing at twice the rate of overall employment in the past decade as businesses continue to digitize their operations to stay competitive.

