Cloud Destinations
Sr Ethernet AI Network Engineer
Apply for this role
Share this job
Job Description
Sr Ethernet AI Network Engineer
to provide senior network engineering services for Ethernet based AI data center fabrics supporting GPU compute clusters, storage connectivity, and management plane services. This role requires strong operational judgment, advanced troubleshooting capabilities, and the ability to resolve complex network issues within large scale production environments.
Note
: This is a W2 only role — C2C, C2H will not be considered This is a 12 month remote contract opportunity within the United States.
Responsibilities
- Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments.
- Troubleshoot Layer 1 through Layer 4 issues involving optics, transceivers, cabling, link bring up, VLANs, MLAG, ECMP, BGP, underlay and overlay reachability, congestion, and packet loss.
- Validate network readiness for distributed training workloads and large east west traffic patterns.
- Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, Quality of Service policy, MTU mismatches, routing instability, and oversubscription.
- Partner with Linux, storage, and cluster deployment teams to isolate host versus network fault domains.
- Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and maintenance activities.
- Capture packet level and counter based evidence to support root cause analysis.
- Develop operational standards for cable mapping, port policy consistency, and fabric health validation.
Qualifications
Required Qualifications:
- 7 or more years of experience in large scale data center networking, including high bandwidth Ethernet fabrics.
- Strong experience with spine leaf architectures, routing, switching, and production troubleshooting.
- Hands on experience with BGP, EVPN, VXLAN, MLAG, ECMP, Quality of Service, Priority Flow Control, RoCE considerations, and telemetry analysis.
- Experience validating optics, breakout configurations, cable plant integrity, and port level consistency.
- Proven ability to troubleshoot distributed application impacts caused by network behavior.
- Experience using switch command line interfaces, automation tools, and packet and counter analysis workflows.
- Strong documentation skills for topology diagrams, incident timelines, and remediation planning.
Preferred Qualifications
- Direct experience supporting AI fabrics carrying large scale GPU collective traffic.
- Familiarity with SONiC, Cumulus Linux, or similar network operating systems used in AI data centers.
- Experience with streaming telemetry, Prometheus, Grafana, and network site reliability engineering operating models.
Tools and Technologies
- Ethernet Data Center Fabrics
- Spine Leaf Network Architectures
- BGP
- EVPN
- VXLAN
- MLAG
- ECMP
- Quality of Service
- Priority Flow Control
- RoCE
- SONiC
- Cumulus Linux
- Prometheus
- Grafana
- Network Telemetry Platforms
- Linux
- Packet Analysis Tools
- Network Automation Tooling
Pay: $80.00 - $85.00 per hour
Expected hours: 40.0 per week
Benefits
* Health insurance
Work Location: Remote
Keep looking
Similar Remote Jobs That Pay Well
Cloud Destinations
Senior Network Security Engineer
ServiceNow
Staff Voice AI Engineer – SRE/DevOps
Booz Allen Hamilton
Enterprise Cybersecurity Artificial Intelligence Engineer, Lead
Booz Allen Hamilton
Enterprise Cybersecurity Artificial Intelligence Engineer, Lead
Teichert Construction