Recently posted · Greenhouse · cloudflare✓ Direct employer / ATS application
Senior Infrastructure Engineer, Storage Platform
Cloudflare
Role details
What you’ll be doing
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations | Austin, US; Seattle, US or London, UK About the Role Emerging Technologies & Incubation (ETI) builds and launches new products on Cloudflare’s global network. Within ETI, you’ll join the Storage Infrastructure team to build and operate a shared storage platform for the teams behind stateful products such as R2, Workers KV, and Durable Objects. We manage the underlying storage hardware, distributed databases, and object storage clusters, including fleet lifecycle automation, capacity, hardware validation, failure-domain planning, observability, and production operations. In this role, you will build platform software and automation that makes these fleet operations safe, repeatable, and scalable. Responsibilities Design and build automation and operator tooling for provisioning, configuring, expanding, upgrading, and decommissioning storage hardware, distributed databases, and object storage clusters. Engineer globally distributed storage fleets to tolerate hardware failures, network disruption, capacity pressure, and recovery events across multiple failure domains. Develop observability, alerting, and safety controls that help engineers understand fleet health and make auditable production changes. Diagnose incidents across Linux, storage, networking, and distributed systems; participate in on-call and implement durable fixes that reduce operational toil. Use AI tools throughout development and operations to accelerate debugging, incident triage, root-cause analysis, and toil reduction, while applying sound engineering judgment and verifying outputs before taking production action. Validate new storage hardware and characterize its performance during normal operation, failures, rebuilds, and other recovery conditions. Partner with R2, Workers KV, Durable Objects, Network, Capacity Planning, Performance Engineering, and Infrastructure Operations to turn service requirements into infrastructure capabilities. Desirable Skills, Knowledge, and Experience Experience designing, building, and operating infrastructure platforms or large-scale production fleets, including automation of lifecycle operations. Strong Linux systems knowledge and experience troubleshooting across compute, storage, and networking layers. Experience operating distributed systems in production, including observability, incident response, reliability improvements, and safe change management. Ability to write maintainable software in at least one programming language, such as Go, Rust, or Python. Strong written and verbal communication skills, with experience collaborating across engineering teams and technical stakeholders. Bonus Points Experience operating distributed databases or object storage systems, with knowledge of storage fundamentals such as filesystems, SSD behavior, replication, and rebuild dynamics. Experience with bare-metal provisioning, hardware lifecycle management, or storage hardware validation. Exp