ORAC

Principal Core Infrastructure Engineer

Oracle · Santa Clara, CA, United States / United States · Software Engineering

Cloud, Infra & DevToolsEntwicklungStaff / Principal115–235 k$/Jahr

Bewerben — 2 Links

Melden Sie sich an, um diese Bewerbung zu verfolgen

Anmelden →

Erkannter Tech-Stack

Beschreibung

In this high-impact role, you will design and build distributed systems and algorithms that measure, analyze, and explain network behavior across OCI’s data centers and WAN. You will combine active measurements, network topology, routing state, flow signals, and switch telemetry to identify where latency, congestion, packet loss, and failures originate.

The broader vision is to make observability an active part of the network control loop. The platform will transform trusted, confidence-scored insights into signals that automated controllers can use to steer traffic, mitigate congestion, isolate faults, and restore network health safely. These capabilities will provide the intelligence foundation for increasingly autonomous, self-healing cloud networks while maintaining explainability, policy controls, and operational safeguards.

This position requires deep systems expertise and hands-on experience with networking, algorithms, telemetry, and distributed computing. You will develop topology-aware observability capabilities, congestion-triangulation and queue-wait reconstruction algorithms, confidence-scored fault localization, and production services that operate reliably at cloud scale.

Responsibilities:

  • Design, implement, and maintain critical components of OCI’s Next Gen Network Observability platform.
  • Build scalable, topology-aware measurement, analytics, and inference systems for data center and WAN networks.
  • Develop algorithms for path reconstruction, congestion triangulation, queue-wait estimation, fault localization, root cause analysis, and customer-impact attribution.
  • Ingest and correlate signals from active probes, routing and topology systems, switches, ASICs, queues, flows, logs, metrics, events, and streaming telemetry.
  • Produce ranked, confidence-scored diagnoses that localize problems to the responsible path, link, queue, device, or shared network resource.
  • Convert observability insights into safe, explainable, and controller-consumable signals for automated traffic steering, congestion mitigation, fault isolation, and service recovery.
  • Design control-loop safeguards that account for uncertainty, stale or conflicting evidence, policy constraints, and unintended interactions between automated actions.
  • Design systems that remain accurate and available despite missing or delayed telemetry, topology changes, clock uncertainty, and partial network failures.
  • Build validation frameworks that measure detection and localization time, accuracy, false-positive rates, confidence calibration, corrective-action effectiveness, and resilience across complex failure scenarios.
  • Apply AI-assisted software development lifecycle practices across design, implementation, testing, debugging, code review, documentation, and operational analysis while maintaining engineering, security, and quality standards.
  • Write robust, well-tested production code in languages such as Java, Go, C++, Rust, or Python.
  • Own and resolve complex production issues involving distributed services, large-scale telemetry pipelines, real-time network state, and automated control workflows.
  • Lead design, and code reviews while maintaining high standards for correctness, scalability, reliability, safety, and operational readiness.
  • Mentor engineers and help establish strong engineering practices for network measurement, analytics, inference, and automated remediation systems.
  • Partner with network controller, SRE, hardware, network operations, security, AI/ML, and product teams to deliver closed-loop capabilities end to end.

Preferred Qualifications:

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent professional experience.
  • 6-10+ years of experience building large-scale distributed systems, network modeling software, network measurement platforms, or telemetry and observability systems.
  • Strong knowledge of networking fundamentals, including L2/L3 forwarding, routing, switching, BGP, ECMP, data center fabrics, queues and buffers, congestion, and packet loss.
  • Advanced programming experience in at least one language such as Java, Go, C++, Rust, or Python.
  • Strong understanding of algorithms, distributed systems, data structures, system reliability, and performance engineering.
  • Experience with one or more of the following: active network measurement, streaming telemetry, time-series or stream processing, graph algorithms, statistical inference, probabilistic modeling, or multi-source data fusion.
  • Familiarity with ML pipelines and models, including data preparation, training and evaluation, deployment, versioning, monitoring, and integration into production systems.
  • Hands-on experience using AI-assisted development tools and workflows across the software development lifecycle.
  • Experience building systems that process high-volume, time-sensitive data and operate reliably under incomplete, delayed, or contradictory inputs.
  • Demonstrated ability to validate complex algorithms and models against ground truth and translate experimental results into production-quality systems.
  • Experience designing automated or closed-loop systems with appropriate safety controls, rollback mechanisms, and human oversight.
  • Track record of technical leadership, mentorship, and delivery across complex, cross-functional projects.
  • Experience operating production systems in cloud-scale or large enterprise environments.

Join the OCI Networking – Next Gen Network Observability team to build the intelligence and automation foundation for self-healing networks—where measurement, inference, and automated control work together to detect problems, select safe corrective actions, and restore network health at cloud scale —helping Oracle’s cloud networks operate with greater visibility, resilience, and efficiency.

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.


Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC4


Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Ähnliche Angebote

VERÖFFENTLICHT AM 15. AUGUST 2026 · ERSTMALS GESEHEN AM 10. SEPTEMBER 2026