DevOps Manager, Infrastructure

Job Locations US-Remote
ID
2026-4571
Category
IT Operations
Type
Full Time

Overview

A great infrastructure team is defined not by how fast it responds to incidents, but by how rarely incidents occur, because the team built the systems, automation, and AI-augmented operations that prevent them. The DevOps Manager, Infrastructure is the people leader and hands-on technical operator who makes that true at Origami Risk. 

This manager leads the team that builds, runs, and continuously improves the cloud infrastructure that Engineering teams and clients depend on every day that spans AWS, self-hosted Windows and Active Directory, network and web-tier appliances, identity and secrets, and the automation platform that ties it all together. This is a hands-on technical leader who coaches Engineers on infrastructure engineering craft, sets a high bar for reliability and operational discipline, drives AI adoption into daily DevOps workflows, and partners with Engineering, Security, and Product so that infrastructure is a source of competitive advantage and not a constraint on what Engineering can deliver. On this team, manual work is a problem to be solved, automation is the default, and every Engineer is expected to use AI to work faster and diagnose more accurately. 

 

Starting base pay for this role is between $145,000 and $185,000. The actual base pay is dependent upon many factors, such as transferable skills, work experience, business needs, training, location, and market demands. The base pay range is subject to change and may be modified in the future. This role will be eligible for a bonus as well as competitive medical, dental, and vision benefits, wellness reimbursement, life insurance, and a 401(k) with company match. We offer vacation and sick leave benefits (under a flexible time off policy in most states).

Responsibilities

What this team runs:

This is a broad estate spanning cloud-native and traditional infrastructure, operated to one modern standard. The successful candidate is energized by that breadth, not surprised by it. The team owns: 

  • Cloud (AWS): the core AWS estate — compute, container platforms, networking, and the supporting services Engineering and clients depend on, provisioned as code. 
  • Windows & Active Directory: the self-hosted Active Directory environment and Windows Server fleet, including the application servers that host the platform. 
  • Network & Web Tier: the firewall estate and web-tier proxy and load-balancing layer that route, secure, and front application traffic. 
  • Automation Platform: the automation and configuration-management platform that drives consistent, repeatable change across the estate. 
  • Identity & Secrets: the shared secrets-management platform, consumed self-service by teams across Engineering. 
  • DNS, Domains & Certificates: the DNS, domain, and certificate lifecycle that keeps services reachable and trusted. 
  • Storage, File & Access Services: the storage, file, backup, and remote-access services that support the business. 

 

Key responsibilities:

People Leadership & Team Development 

  • Leads, develops, and retains a team of Cloud Infrastructure Engineers building a high-performance culture defined by technical excellence, an automation mindset, and operational ownership. 
  • Sets clear goals, role expectations, and success criteria for each engineer; conducts regular 1:1s and gives direct, constructive feedback with a visible path for growth. 
  • Conducts structured performance reviews that recognize strong operational contributions, address delivery or quality gaps specifically, and grow both AI fluency and foundational infrastructure depth. 
  • Identifies capability gaps across the team encompassing IaC, Windows/AD administration, networking, container orchestration, security controls, and AI-assisted operations, while building targeted hiring, upskilling, and cross-training plans. 
  • Fosters a culture where every manual process is a temporary state, automation is the permanent solution, and AI tooling is an expected part of how each engineer works. 
  • Supports Engineering hiring across the organization by participating in technical interviews and helping calibrate the infrastructure engineering bar. 

Cloud Infrastructure Operations 

  • Owns the operational health of the organization's AWS infrastructure setting team standards for provisioning, configuration management, environment consistency, patching cadences, capacity planning, and cost optimization across all environments. 
  • Leads the team's infrastructure-as-code practice ensuring resources are defined, versioned, peer-reviewed, and deployed through Terraform, and that manual changes are treated as exceptions to be immediately codified. 
  • Oversees the team's operation of the Ansible Automation Platform encompassing roles, inventory, and configurations as the primary configuration-management system across the estate. 
  • Ensures reliable operation of the self-hosted Active Directory forests, Windows Server fleets, and Windows RDS, including Group Policy, OU structure, service-account and gMSA management, and identity hygiene. 
  • Owns certificate, DNS, and domain lifecycle management across, ensuring renewals and changes are reliable and never a source of outages. 
  • Ensures our Vault secrets service is operated to a high standard as a shared, self-service secrets platform for multiple teams. 
  • Drives cloud cost efficiency as a continuous discipline, by right-sizing, reserved capacity, auto-scaling, and resource-cleanup automation that control spend as the platform scales. 
  • Oversees core AWS services, EC2, ECS, EKS, Lambda, RDS, S3, VPC, Route 53, IAM, CloudFront, and Auto Scaling, while ensuring configurations are standardized, documented, and aligned with architectural and security standards. 

Network, Firewall & Web-Tier Operations 

  • Owns operation of the Fortinet firewall estate and NGINX Plus web tier encompassing IPS, routing, firewall policy, object management, SSL inspection, and VIP configuration that proxies traffic to NGINX that is deployed and configured through Terraform and Ansible. 
  • Ensures the infrastructure architecture reflects network security controls, traffic-routing policies, and connectivity requirements — operating infrastructure that is reliable and secure at every layer. 
  • Holds the team accountable for firmware and security-update discipline across firewalls, NGINX, FTP, and other network-facing systems. 

Delivery Automation & Release Support 

  • Ensures the infrastructure and automation the team owns integrate cleanly with the delivery pipeline — enabling automated provisioning and deployment workflows and holding the team accountable for the infrastructure reliability that automated releases depend on. 
  • Partners with the Release Automation and Engineering Systems teams so that pipeline tooling has the environment provisioning, configuration, and deployment automation it needs — with appropriate access controls, standards alignment, and change tracking. 
  • Oversees deployment automation for containerized services, Lambda functions, ECS tasks, and EKS workloads — ensuring patterns are consistent, traceable, and recoverable across environments. 
  • Coaches Engineers to diagnose and eliminate the root causes of flaky, intermittent infrastructure failures rather than treating them as acceptable operational noise. 

AI-Augmented DevOps Operations 

  • Drives adoption of AI development tools — including Claude Code and Cursor — across the team's daily workflow, establishing norms for AI-assisted IaC authoring, automation scripting, and operational documentation. 
  • Champions AI-assisted operations as a management discipline — integrating AI-powered log analysis, anomaly detection, runbook generation, and LLM-augmented incident analysis into workflows that reduce mean time to diagnosis. 
  • Uses AI to accelerate the team's most demanding work — large-scale infrastructure audit and compliance analysis, drift detection, capacity planning, and intelligent alerting — and shares effective patterns with engineering leadership. 
  • Evaluates AI-native infrastructure tools and makes evidence-based recommendations to the Director of DevOps on where AI investment will deliver the highest operational return. 
  • Holds team members accountable for AI tool adoption and proficiency — incorporating AI fluency into performance conversations, development plans, and hiring criteria. 

Observability, Monitoring & Incident Response 

  • Leads the team's ownership of infrastructure and application observability — monitoring dashboards, alerting policies, log aggregation, and anomaly detection across the stack using New Relic, Datadog, CloudWatch, and related tooling. 
  • Holds the team accountable for alert quality — building monitoring that surfaces actionable signals rather than noise that erodes on-call confidence. 
  • Owns the team's on-call model — rotation schedules, escalation paths, and readiness standards — ensuring every engineer has the runbooks, tooling access, and AI-assisted investigation capabilities to resolve incidents effectively. 
  • Leads blameless post-incident reviews, produces structured root cause analyses, and owns follow-through on corrective actions that prevent recurrence. 
  • Tracks infrastructure reliability metrics — availability, MTTR, change failure rate, and deployment frequency — reporting trends and improvement actions to the Director of DevOps. 

Security, Compliance & Access Management 

  • Holds the team accountable for infrastructure security controls across all environments — IAM policy governance, security-group management, secrets management, encryption standards, vulnerability patching, and compliance with baselines aligned to NIST 800-53, ISO/IEC 27001, and SOC 2. 
  • Ensures the team conducts regular infrastructure security assessments — reviewing IAM permissions, security-group configurations, exposed endpoints, and compliance configurations — and drives timely remediation of identified gaps. 
  • Owns infrastructure access governance — provisioning and de-provisioning that are timely, least-privilege, regularly reviewed, and audit-ready at all times. 
  • Supports audit readiness and evidence collection for SOC 2, ISO/IEC 27001, and NIST 800-53 — maintaining accurate records of infrastructure configurations, access grants, and change history. 
  • Partners with Application Security Engineers to integrate infrastructure-level controls — container scanning, secrets management, and network-policy enforcement — into the automation workflows the team owns. 

Cross-Functional Collaboration & Stakeholder Management 

  • Serves as the primary operational partner to Engineering delivery teams — providing infrastructure consultation during system design and ensuring services are operationally sound and reliably deployable. 
  • Collaborates with the Engineering Systems and Release Workflow teams on pipeline governance, release planning, and deployment coordination — ensuring required infrastructure capacity is available, validated, and ready before promotion. 
  • Represents the DevOps team in Engineering leadership forums — communicating infrastructure health, operational risk, team capacity, and improvement roadmap with data-backed specificity. 
  • Manages vendor and tooling relationships relevant to the infrastructure domain — cloud provider support, monitoring platforms, firewall/appliance vendors, and licensing — ensuring SLAs are met and escalations are handled with urgency. 

Qualifications

  • Bachelor's Degree in Computer Science, Software Engineering, Information Systems, or a related field. 
  • 6+ years of cloud infrastructure engineering experience, including 2+ years in a management role with direct people-leadership accountability for a DevOps or Infrastructure Engineering team. 
  • Demonstrated track record of leading and developing technical teams — hiring, coaching, performance management, and building a high-performance culture. 
  • Deep hands-on proficiency with AWS cloud services — EC2, ECS, EKS, Lambda, RDS, S3, VPC, IAM, Route 53, CloudFront, and Auto Scaling — designing and operating production AWS environments at scale. 
  • Proficiency with infrastructure-as-code tooling — Terraform required (AWS CDK a plus) — with the ability to review, guide, and hold the team accountable for IaC quality and consistency. 
  • Hands-on experience administering Windows Server and Active Directory in production — Group Policy, service accounts/gMSA, and identity management — with the depth to lead a team operating self-hosted AD. 
  • Working knowledge of network and web-tier infrastructure — firewalls (Fortinet or equivalent NGFW), reverse proxy/load balancing (NGINX or equivalent), DNS, and certificate lifecycle management. 
  • Proficiency with configuration management and automation — Ansible / Ansible Automation Platform — and scripting in Python, PowerShell, and Bash, with the depth to coach on and review automation code. 
  • Experience with container technologies including Docker, ECS, and EKS — image build pipelines, registry management, and containerized deployment automation. 
  • Experience with observability and monitoring platforms such as New Relic, Datadog, CloudWatch, or SumoLogic — at a depth sufficient to set team standards and review team-produced monitoring. 
  • Demonstrated experience driving AI tool adoption within a DevOps or infrastructure engineering team — AI-assisted IaC authoring, AI-augmented log analysis, or intelligent alerting and automation. 
  • Strong troubleshooting, incident management, and on-call leadership skills — able to lead the team through high-pressure production incidents with structure, calm, and effective coordination. 

Preferred Qualifications 

  • AWS certification at the associate or professional level (e.g., AWS Solutions Architect, AWS DevOps Engineer – Professional). 
  • Experience operating HashiCorp Vault or an equivalent secrets-management platform as a shared service. 
  • Experience with CI/CD pipeline platforms — Azure DevOps Pipelines and/or GitHub Actions — sufficient to partner effectively with release and pipeline teams. 
  • Experience implementing SRE practices — SLOs, error budgets, and chaos engineering — within a managed infrastructure team. 
  • Familiarity with Temporal or AWS Step Functions for workflow orchestration in infrastructure automation or deployment pipelines. 
  • Experience with log aggregation tooling including Elasticsearch, Logstash, Kibana (ELK), Graylog, or Fluent. 
  • Experience supporting compliance frameworks — SOC 2, NIST 800-53, or ISO/IEC 27001 — including managing team accountability for control evidence and audit readiness. 
  • Experience evaluating and adopting AI-native infrastructure intelligence tools — AI-assisted deployment risk scoring, predictive scaling, or LLM-powered operational runbook generation. 

Benefits

  • Medical and Dental coverage available for employees, dependents, domestic partners, and spouses
  • Paid Time Off – Flexible options plus 10 paid company holidays where available**
  • All full-time positions are hybrid, with many eligible to be completely remote
  • Fully Paid by Origami Risk – Vision insurance, Short & Long-Term Disability Insurance, and Basic Life Insurance
  • Generous family leave options—including adoption and foster care placements
  • Pre-Tax Savings Accounts – Flexible Spending Account, Health Savings Account, Commuter Benefits, Dependent Care Savings Account
  • Retirement Savings – 401(k) with company match up to 4%
  • Employee Assistance Program (EAP) – Confidential & Free support offered to colleagues facing personal or work-related complications
  • Education Assistance Program – to help colleagues pursue industry/role-specific certifications
  • Wellness Benefits – reimbursement program to invest in healthy habits as well as support better colleague productivity and stress management
  • Additional coverages available – Pet Insurance, Critical Illness Insurance, and Voluntary Life & AD&D coverage
**Flexible PTO not available in California or the UK

Who We Are

Origami Risk delivers single-platform SaaS solutions that help organizations best navigate the complexities of risk, insurance, compliance, and safety management.

 

Founded by industry veterans who recognized the need for risk management technology that was more configurable, intuitive, and scalable, Origami continues to add to its innovative product offerings for managing both insurable and uninsurable risk; facilitating compliance; improving safety; and helping insurers, MGAs, TPAs, and brokers provide enhanced services that drive results.

 

A singular focus on client success underlies Origami’s approach to developing, implementing, and supporting our award-winning software solutions. 

 

Origami Risk is proud to be an equal opportunity employer. We thrive and benefit from diversity and are committed to creating an inclusive and equitable environment for all employees. We do not discriminate against any individual based upon race, religion, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, color, sex, national origin, age, marital status, military or veteran status, disability, or any other characteristic protected by applicable law.

 

Caution: Be alert to recruiting scams. We have received reports of individuals impersonating Origami Risk recruiters to deceive candidates into disclosing personal information. These impostors use fake Origami Risk domain names and email addresses. Please double-check that any email address from an Origami Risk recruiter ends with origamirisk.com or talent.icims.com. And to confirm the legitimacy of any recruiting communication, feel free to email transparencycheck@origamirisk.com.

Options

Sorry the Share function is not working properly at this moment. Please refresh the page and try again later.
Share on your newsfeed