Joyent
Infrastructure Software Engineer - Topology and Observability
Share this job
Job Description
Mountain View, CA or Remote
Joyent powers the global cloud infrastructure and developer platform providing back-end services for Samsung's billions of devices. Joyent's data center footprint is within 100ms latency to 70% of the world's population, while our multi-cloud, Kubernetes-based developer platform extends our reach to additional resource regions. We're operating at hyperscale to power workloads that bring capability and delight to Samsung's employees and customers.
Job Summary
---------------
Our infrastructure spans multiple layers — physical network, compute hosts, virtualization, managed services, and the tenant workloads running on top of them. Each layer has its own configuration management, inventory, and monitoring systems. What we lack is a tool that cuts across those layers and answers, in one view, "which hosts, network paths, and services does this workload depend on?"
In this role you will design and build a system that unifies scattered configuration, inventory, and observability data into a single infrastructure knowledge graph, and delivers multi-layer topology maps and service dependency maps on top of it. Users can overlay physical, logical, and service layers selectively, drill down hierarchically from a region to an individual node, and jump from any component to its related alerts, tickets, runbooks, and dashboards. The system serves as a shared operational view for the network, compute, SRE, and security teams, and becomes the foundation for fast root cause analysis, failure-domain analysis, and pre-change impact assessment.
### Job Responsibilities
- Design and maintain an infrastructure topology data model spanning physical, logical, and service layers — schemas, node and edge metadata, and cross-layer relationships
- Build read-only, low-impact collectors and parsers that gather data from configuration management tools, network devices, hypervisors, and metrics/log systems
- Build pipelines that reconcile declared configuration against observed state to detect and report drift
- Develop the interactive web frontend and backend services providing multi-layer, hierarchical drill-down navigation, directional traffic and dependency flow visualization, and deep links into related systems
- Design dependency-based blast radius calculation, failure-domain analysis, and change scenario simulation
- Provide topology and dependency data as the foundation for alert correlation and automated root cause analysis (RCA), and integrate with incident response, alerting, and ticketing systems and AIOps pipelines — including human-in-the-loop review and approval workflows
- Define data accuracy and freshness metrics; own the reliability and performance of the system itself
- Gather requirements from consuming teams and expand adoption in stages (static artifacts hosted service)
### Skills & Competencies
- Understanding of hypervisor-based virtualization and compute host operations
- Understanding of the data models and limitations of observability stacks (Prometheus, OpenTelemetry, logging and tracing systems)
- Sound judgment in designing data collection that minimizes impact on production systems
- Ownership -Take ownership of projects, ensuring excellence in execution and accountability for results. Foster a sense of responsibility and pride in delivering high-quality work
- Innovation - Drive innovation by proposing and implementing creative solutions to challenges. Stay abreast of industry trends and technologies, bringing fresh ideas to the table
- Customer focus - Understand and prioritize customer needs, striving to exceed expectations in every interaction. Collaborate with cross-functional teams to ensure the delivery of customer-centric solutions
- Teamwork - Embrace a collaborative and inclusive approach, working seamlessly with colleagues to achieve common goals
### Education & Experience
- 5+ years in infrastructure, platform, SRE, or network automation, with hands-on experience across both cloud and on-premises environments
- Working knowledge of L2/L3 networking: routing protocols (e.g. BGP), overlay networks (e.g. VXLAN), and the Linux networking stack
- Experience building Python data pipelines, including schema definition and validation tooling
- Experience modeling and processing data in analytical stores (columnar databases, graph databases, or similar)
- Experience visualizing graph/topology data on the web (D3, Cytoscape, or similar) and handling the performance and readability challenges of large graphs
### Preferred qualifications
- Experience operating large-scale multi-region cloud or IaaS environments
- Experience building or integrating CMDBs, inventories, or network sources of truth (NetBox, Nautobot, or similar)
- Experience with service dependency mapping, failure-domain analysis, or what-if / digital-twin style simulation
- AIOps experience: alert correlation and noise reduction, anomaly detection, automated root cause analysis, and building incident response automation pipelines
- Experience applying LLM-based agents to operational workflows — particularly using LLMs to parse and summarize unstructured sources (documents, free-form configuration, alert text) wrapped in deterministic validation and human review
- Experience with automated diagram generation (draw.io, Graphviz, or similar)
- Experience collaborating with security teams on exposure surface or attack path visualization
### What success looks like (first 12 months)
- 3 months: Data model and collectors for the core layers are operational, and a validated physical and logical topology exists for at least one environment
- 6 months: Two or more teams use the system in real incident analysis, and configuration-vs-observed drift is reported on a regular cadence
- 12 months: The system runs as a hosted service, and impact assessment and change what-if analysis are part of standard procedures
### What this role is not
- A monitoring operations role focused on dashboard maintenance
- A pure frontend or pure network engineering role — we are looking for someone who combines infrastructure understanding with data and visualization skills
Compensation and Benefits
-----------------------------
Compensation for this position will vary among specific regions due to geographical differentials in the labor market, and actual pay will be determined considering factors such as relevant skills, experience, and comparison to other employees in the role. Therefore, the annual base compensation range for this role (depending on the geographical location) is expected to be between $135000 and $190000.
Regular full-time employees (salaried or hourly) have access to benefits including Medical, Dental, Vision, Life Insurance, 401(k), Employee Purchase Program, Vacation and Sick leave, electronic reimbursement and many more. In addition, regular full-time employees (salaried or hourly) are eligible for bonus compensation based on individual, department, and company performance.
About Joyent
----------------
Joyent, a wholly-owned subsidiary of Samsung, is the open cloud company. Joyent builds technology, at the pinnacle of scale, performance, stability, and security to accelerate the transformation toward the mobile and cloud-centric world. Joyent designs, builds and manages market competitive cloud computing solutions and services for Samsung Electronics and its partners at global scale.
How To Apply
----------------
To apply, please submit a brief introduction, a copy of your resume, and a link to your Github or LinkedIn profile to jobs@joyent.com with Infrastructure Software Engineer - Topology and Observability in the subject. We are an equal-opportunity employer, building a diverse and inclusive team. Qualified applicants with criminal histories will be considered for the position in a manner consistent with the Fair Chance Ordinance.
Joyent is committed to employing a diverse workforce and providing Equal Employment Opportunities for all individuals regardless of race, color, religion, gender, age, national origin, marital status, sexual orientation, gender identity, status as a protected veteran, genetic information, status as a qualified individual with a disability, or any other characteristic protected by law.
*Disclaimer:* *This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities and activities may change or new ones may be assigned at any time with or without notice.*
Keep looking
Similar Remote Jobs That Pay Well
Pear Suite
Staff/Software Software Engineer
Quantum Sky