CloudLinux is a global remote-first company. We are driven by our principles: do the right thing, employees first, we are remote first, and we deliver high-volume, low-cost Linux infrastructure and security products that help companies to increase the efficiency of their operations. Every person on our team supports each other and does what we can to ensure we all are successful.
Check out our website for more information: cloudlinux.com
Imunify360 is a multi-layer Linux server security suite — WAF, IDS/IPS, malware scanning and cleanup, proactive defence, patch management, reputation — running as an agent on hundreds of thousands of customer servers, backed by a cloud estate of scanning, correlation and signature-delivery services on our own bare metal.
Roughly 70 components currently ship without a defined service level indicator. Some are internal services we can scrape. Many are agent-side subsystems running on machines we do not own, reporting through a heartbeat we designed for something else. There is monitoring, and there are dashboards, and there is no coherent answer to the question “is this component doing its job right now, and how would we know if it stopped?”.
We know the cost of that gap precisely, because we recently paid it: a security control was silently disabled across a large fraction of the fleet for 61 days. Every dashboard was green. The telemetry reported a ruleset version but not whether the control that consumed it was switched on, so a configuration change was indistinguishable from a broken updater. Three independent safety mechanisms existed and all three were gated behind the same condition that caused the failure.
Your job is to make that class of failure detectable in hours instead of months, across the whole product line, and to build the system that keeps it detectable as the product changes.
This is a greenfield charter inside a brownfield estate. You are not inheriting an SRE team, an SLO framework or a paging culture. You are defining them, with the engineering leads, and then making them stick.
What you’ll do
What you’ll bring
First year, in outcomes
30 days: Component inventory with named owners. SLI taxonomy and tiering agreed. 3 pilot components fully instrumented end to end as the reference implementation.
90 days: Collection pipeline in production. Tier-1 components (the ones whose failure is a customer security exposure) carry SLO, alert, runbook, owner. Escalation routing live for tier-1.
180 days: All ~70 components have a defined SLI and an owner. Alert taxonomy enforced; page volume and actionable-rate measured and published. Squad on-call operating.
365 days: Mean time to detect a silent control-degradation is under 24 hours, measured, against a 61-day baseline. Error-budget policy influences release decisions. The function is documented well enough that hire #2 and #3 are additive, not archaeological.
How we work
Remote-first and async across nine time zones. Weekly PO sync and architecture sync; monthly demo and OKR review; quarterly architecture summit. Decisions land as ADRs. Every output carries an owner and a due date. Postmortems are blameless and published, and we correct ourselves on the record when we get something wrong.
By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which provides detailed information on how we maintain and handle your data.