Cloud & DevOps
Cloud architecture, Kubernetes, Terraform, and zero-downtime CI/CD — engineered by senior SREs who have run platforms processing tens of billions of requests a year.
01 — Approach
How we engage.
- Greenfield cloud architecture — multi-account AWS / GCP / Azure landing zones, VPC design, identity, secrets management, and a written rollout plan that survives the first audit.
- Kubernetes platforms — production EKS, GKE, and AKS clusters with GitOps deploys, service mesh where it earns its keep, autoscaling that does not lie to you, and runbooks your on-call team will actually read.
- Infrastructure as code — Terraform modules, Pulumi, and policy-as-code (OPA, Checkov, tflint) that pass review on day one and stay reviewable as the team scales.
- CI/CD pipelines — GitHub Actions, GitLab CI, Buildkite — with progressive delivery, blue/green and canary deploys, signed artifacts, and rollbacks that take seconds, not pages.
- Observability platforms — OpenTelemetry collection pipelines, Prometheus, Grafana, Loki, Tempo, and the dashboards that actually get opened during an incident.
- SRE practices & on-call — SLOs, error budgets, incident response playbooks, post-mortems, and an on-call rotation your engineers do not quietly resent.
- Cloud cost optimisation — FinOps audits that find the 30–60% of cloud spend most platforms leak: idle GPUs, oversized nodes, egress, and the slow drift of "temporary" resources.
- Cloud migrations — data centre to AWS, AWS to GCP, monolith to microservices — done in cuts you can roll back, not big-bang launches.
- Senior-only delivery. A staff-level SRE owns the engagement — not a delivery manager with a reference architecture and three offshore juniors.
- Boring infrastructure is good infrastructure. We pick the smallest set of moving parts that solves the problem. Kafka only when you need Kafka. Service mesh only when you need a service mesh.
- Cost-aware by default. Every architecture proposal includes a monthly cost estimate, scaling assumptions, and the line items most likely to surprise you.
- Zero-downtime as a contract, not a wish. Blue/green or canary deploys, database migrations done online, feature flags for risky rollouts, rollback playbooks for everything.
- Secure by construction. Least-privilege IAM, secrets in Vault or SSM, signed container images, SBOMs, and a security review before launch — not after the breach.
- Stay-on retainers. Most clients keep us on a monthly SRE retainer covering on-call rotation help, drift remediation, cost reviews, and the slow steady work of keeping a platform reliable as the product changes.
- Cloud: AWS, GCP, Azure — depending on the workload and the team's existing footprint.
- Orchestration: Kubernetes (EKS, GKE, AKS), ECS, Cloud Run, Lambda, App Runner.
- Infrastructure as code: Terraform, Terragrunt, Pulumi, OpenTofu, Crossplane.
- CI/CD: GitHub Actions, GitLab CI, Buildkite, Argo CD, Flux, Spinnaker.
- Observability: OpenTelemetry, Prometheus, Grafana, Loki, Tempo, Datadog, Honeycomb.
- Security & policy: OPA / Gatekeeper, Checkov, tflint, Trivy, Cosign, HashiCorp Vault.
- Data: PostgreSQL, ClickHouse, Redis, Kafka, ScyllaDB — operated on Kubernetes or as managed services.
- Cloud architecture & audit — $25k–$75k, two to six weeks. Written architecture, risk register, FinOps report, prioritised remediation plan.
- Build & migration — $90k–$400k+, three to six months. Fixed scope, milestone-billed, with a documented cut-over plan.
- Embedded SRE squad — monthly retainer, one to three engineers. Best for ongoing platform evolution alongside your internal team.
- On-call augmentation — monthly retainer. Senior on-call coverage for teams that cannot yet sustain a 24/7 rotation.
02 — What's included
Every engagement ships with.
Senior lead
A 10+-year practitioner who stays on the work, end-to-end.
Design system
A scalable foundation, not screen-by-screen one-offs.
Production deploys
Fortnightly increments to a staging URL.
Documentation
Runbooks, ADRs, and onboarding materials.
03 — Process
Four phases. Always.
Discovery
1–2 weeks. Audit, listen, scope.
Design
2–4 weeks. Prototypes you can click.
Build
6–16 weeks. Two-week cadences.
Stewardship
Ongoing. Continuity beats handoff.
04 — Common questions
Frequently Asked Questions
How much do Cloud & DevOps services cost?
A focused cloud architecture and audit engagement starts at $25,000 (two to six weeks). A full build or cloud migration typically lands between $90,000 and $400,000 depending on scope, regulatory requirements, and the size of the existing footprint. Monthly SRE retainers start at $15,000 and scale with on-call coverage and engineer count.
How long does a cloud migration take?
A small monolith-to-cloud cut takes 8–12 weeks. A large multi-service migration with stateful workloads runs four to nine months. We migrate in cuts you can roll back — not big-bang launches — so production stays usable through the entire process.
AWS, GCP, or Azure — which should we pick?
We are cloud-agnostic and will tell you honestly. AWS still wins on breadth of services and the largest ecosystem of vendors. GCP wins on data and ML workloads, networking quality, and Kubernetes ergonomics. Azure wins inside Microsoft-heavy enterprises and for clients with Office 365 / Entra integration requirements. We have shipped production work on all three and will recommend by workload, team experience, and existing commitments.
Do we actually need Kubernetes?
Probably not at first. Most teams under twenty engineers are better served by ECS, Cloud Run, App Runner, or Fly — they get 80 percent of the value without the operational burden. We move teams to Kubernetes when service count, traffic control, or multi-region requirements demand it — and we tell teams when it is the wrong answer.
Can you reduce our cloud bill?
Yes. A FinOps audit typically finds 30–60 percent of monthly cloud spend in oversized nodes, idle GPUs, untargeted egress, orphaned resources, and unused reserved capacity. We have repeatedly cut six-figure monthly bills by a third in the first quarter — without sacrificing performance or reliability.
Do you handle on-call and incident response?
Yes. Our SRE retainer can include senior on-call augmentation, incident response playbook authoring, blameless post-mortem facilitation, and a rotation alongside your internal engineers. We do not replace your on-call permanently — but we help teams that cannot yet sustain a 24/7 rotation bridge the gap.
How do you ensure our infrastructure is secure?
Security is built in from architecture, not bolted on. Every engagement includes least-privilege IAM, secrets management with Vault or SSM, signed container images, SBOMs, OPA / Gatekeeper policy, and an end-of-build security review by our cybersecurity practice. Compliance frameworks (SOC 2, ISO 27001, HIPAA, PCI) are scoped at audit time.
Do you offer ongoing infrastructure maintenance?
Yes. Most clients keep us on a monthly SRE retainer covering drift remediation, Terraform and Kubernetes version upgrades, cost reviews, library churn, on-call augmentation, and the slow steady work of keeping a platform reliable as the product evolves.
05 — Selected work
Related projects.
— From the journal