Posted today · Greenhouse · scaleai✓ Direct employer / ATS application
Senior Cloud Infrastructure Engineer
Scale AI
Role details
What you’ll be doing
Scale is powering this generative AI wave by providing the data and infrastructure for companies to build large-scale foundation models. AI is rapidly changing the world, and Scale is growing to meet that rapid demand. Our customers include OpenAI, Microsoft, Adept, Stability AI and many more major players in this space! The Platform team is responsible for building the core abstractions and infrastructure on which the products can be built and iterated rapidly. We are looking for an Operations-focused AWS Engineer with deep experience in infrastructure management, lifecycle maintenance, and security hardening. We are looking for Ops specialists who thrive on optimizing, securing, and maintaining the health of large-scale AWS environments. You have a growth mindset and are comfortable learning new technologies. You Will: Manage Infrastructure Lifecycle : Lead the end-to-end lifecycle of our AWS infrastructure, including routine patching, version upgrades, and system maintenance to ensure high availability. Vulnerability Management : Work closely with the Security Team to identify, prioritize, and remediate vulnerabilities across our cloud footprint. AWS Optimization : Continuously monitor and optimize AWS resource utilization for performance, reliability, and cost-efficiency. Automate Operations : Use scripting and automation tools to streamline repetitive operational tasks such as fleet-wide patching and configuration audits. Incident Response & Troubleshooting : Provide technical expertise for troubleshooting infrastructure-level issues and participate in operational health reviews. Security Compliance : Build and maintain systems that adhere to strict security standards, ensuring our environment remains compliant through response and proactive mitigations. Qualifications: Extensive AWS Experience : 6+ years of experience with core AWS services (EC2, VPC, IAM, S3, RDS, EKS) and experience managing multi-account environments. Patching & Upgrades : Proven track record of managing large-scale patching programs and performing major version upgrades for OS and middleware with minimal downtime. Operational Tooling : Proficiency with AWS Systems Manager (SSM), Terraform, Atlantis, and Kubernetes for configuration management and automation. Scripting : Familiarity with writing functional scripts (Bash, Python, etc.) to automate operational workflows - focused on system management rather than application development. Security Focus : Experience with vulnerability scanning tools and a strong understanding of how to harden cloud infrastructure against common threats. Monitoring & Metrics : Experience using Datadog, CloudWatch, or similar tools to monitor system health and drive optimization efforts. Multi Cloud Experience: Experience with Azure, Google Cloud Platform is a bonus. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is: $143,000 — $178,000 USD PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants. About Us: At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies