Hyphen Connect Limited
LLM Pre-training & Distributed Engineer (AI Infrastructure)
Be an Early Applicant
Lead orchestration and optimization of large-scale LLM pretraining across 1,000+ GPUs. Manage distributed training with PyTorch/DeepSpeed/Megatron-LM, tune networking and memory (InfiniBand/RDMA), and implement checkpointing and robust failure recovery for long-running jobs.
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar Jobs
Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Database • Analytics
Leads InterSystems’ Australia and New Zealand business, owning regional sales strategy, revenue growth, customer relationships, partnerships, and P&L accountability. Manages sales and sales engineering teams, develops business plans and KPIs, engages senior executives, coordinates cross-functional resources, expands the customer base, and ensures customer success. The role requires extensive regional travel, enterprise software expertise, strong technical selling skills, and the ability to lead through influence across a multinational organization.
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Build and refine core AI and machine learning infrastructure supporting model development, training, evaluation, deployment, and monitoring. Design scalable distributed ML systems, collaborate with product and engineering teams, lead projects from technical design through launch, review code, document solutions, resolve infrastructure challenges, and mentor junior engineers.
Top Skills:
AWSAzureCi/CdDaskGCPGoGpusJavaJaxLarge Language Models (Llms)MlopsPythonPyTorchRayRetrieval-Augmented Generation (Rag)ScalaSparkTensorFlow
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Lead regional security incident response for high-severity enterprise incidents, directing investigations, communications, coordination, and remediation. Build incident response tools, systems, playbooks, and programs; conduct exercises and post-incident reviews; mentor team members; collaborate across security teams; engage law enforcement and industry bodies; and produce threat intelligence. The role requires expertise across multiple security domains, cloud and network technologies, scripting, and incident response leadership.
Top Skills:
FirewallsGCPGoIntrusion Detection Systems (Ids)Network ProtocolsPython
What you need to know about the Sydney Tech Scene
From opera to comedy shows, the Sydney Opera House hosts more than 1,600 performances a year, yet its entertainment sector isn't the only one taking center stage. The city's tech sector has earned a reputation as one of the fastest-growing in the region. More specifically, its IT sector stands out as the country's third-largest, growing at twice the rate of overall employment in the past decade as businesses continue to digitize their operations to stay competitive.


