Cloud & DevOps

Infrastructure and release decisions built around the product.

Start with the boundary

Choose Cloud & DevOps when infrastructure or release responsibility
defines the work

Not every problem is an infrastructure problem. We start by establishing where the actual limitation sits.

TIER01

Application-level fix

TIER02

Configuration change

TIER03

Bounded infrastructure change

TIER04

Infrastructure & release ownership

Current operating state

Establish the current environment and who
controls each boundary

Change is only safe once the current state, its risks, and its owners are actually known.

01/

Current runtime and environments

Map providers, environments, access, and who currently controls each boundary.

02/

Release path

Trace build, test, approval, deployment, and rollback as they actually happen today, not as documented.

03/

Data and recovery posture

Establish what is backed up, how often, and whether restore has ever actually been tested.

04/

Operating signals

Identify which logs, metrics, and alerts are useful now versus which exist but nobody looks at.

Shared responsibility

Make cloud security
a shared-responsibility boundary

Infrastructure change is product-consequential. It deserves the same review rigor as a feature change.

(001)

Treat infrastructure change as product-consequential

A deployment, migration, or scaling change is reviewed with the same rigor as a feature change, because it can affect users just as directly.

(002)

Shared responsibility, not full takeover

Provider, platform, and team responsibilities are named explicitly; nothing is assumed to be handled automatically.

(003)

Coexistence before cutover

Migrations run old and new systems in parallel with validation before the old path is retired.

(004)

Cost as a visible input

Capacity and spend are tracked against actual usage, not assumed to shrink from a single optimisation pass.

Technology decisions

Cloud & DevOps
platform stack

The stack follows the requirement, not the other way round. These are the categories this engagement typically has to decide.

01

Terraform & IaC

02

Docker & Kubernetes

03

AWS & Google Cloud

04

GitHub Actions CI/CD

05

Datadog & Prometheus

06

Linux & Cloudflare

Possible scope

Make build, test, approval, release, observation, and
rollback one path

The exact scope depends on the current state and agreed boundary. None of the following should be assumed automatically.

Environment, access, and dependency inventoryRelease-path walkthroughBackup and recovery verificationRisk and responsibility record
CI/CD pipeline design or repairRollback and roll-forward pathsBackup, restore, and retention policyIncident and escalation runbook
Log, metric, and trace coverage for the actual riskAlert thresholds tied to real signalsDashboards scoped to owners, not vanity metricsCost and capacity visibility
Coexistence and validation planCutover and rollback criteriaDocumentation of the resulting stateOngoing coverage decision
Infrastructure decisions

Plan coexistence, cutover, validation, and
reversal before migration

Every infrastructure decision trades something specific for something else. The useful choice is the one whose reason is understood.

01

Managed platform or operated infrastructure?

A managed service reduces operational surface at the cost of control; the right choice depends on team capacity and constraint tolerance.

02

How many environments earn their keep?

Every extra environment is a maintenance and consistency cost; add one only when it catches something staging alone would miss.

03

Automatic or reviewed recovery?

Automatic rollback is safer for well-understood failures; some changes need a human decision before reverting.

04

Optimise now or keep capacity margin?

Aggressive cost-cutting can remove the buffer that absorbs real traffic spikes.

Inspectable outputs

Leave infrastructure evidence, access, ownership,
and handoff explicit

Deliverables should make the current and resulting infrastructure state visible and testable.

01/

Current-state and risk record

What exists today, what is fragile, and what depends on what.

02/

Release and recovery evidence

A pipeline and rollback path that has actually been exercised, not just documented.

03/

Observability baseline

Logs, metrics, and alerts that map to real failure modes.

04/

Handoff or coverage agreement

Runbooks and access transferred, or continuing operational responsibility explicitly agreed.

Delivery and engagement

Choose Cloud & DevOps when infrastructure and operation
define the need

The starting engagement should reduce the operational risk that matters before committing to broader change.

01
BOUNDED DISCOVERY

Infrastructure assessment

A bounded review of the current runtime, release path, and recovery posture.

02
FIXED-SCOPE MILESTONE

Defined infrastructure change

A scoped engagement to implement one release, recovery, or observability improvement.

03
STAGED RELEASES

Staged migration

Coexistence, validation, and cutover run as a sequence of reviewable stages.

04
CONTINUOUS CADENCE

Ongoing operational coverage

Continued release, monitoring, and incident responsibility as part of an active roadmap.

Fit and routing

Use the service
that owns the defining condition.

Cloud & DevOps is the right starting point when infrastructure and operation define the need. A different service may own the wider product need.

01/ ALTERNATIVE

Custom Software Development

Choose this when the defining need is the application itself, not its infrastructure.

02/ ALTERNATIVE

Product Modernisation & Improvement

Choose this when infrastructure change is one part of a broader existing-product transition.

03/ ALTERNATIVE

Ongoing Product Engineering

Choose this when continuing operational responsibility is the actual ask, not a bounded change.

04/ ALTERNATIVE

AI-Enabled Software & Automation

Choose this when the infrastructure question is really about hosting or serving an AI workload.

Frequently Asked Questions

Cloud & DevOps questions,
answered clearly.

Technical delivery and architecture details. Key decisions regarding code quality, integrations, and handoff rituals.

No. Provider and tooling decisions follow workload, state, release, recovery, security, portability, cost, and team constraints.

Build the product.
Start with context.

Start a conversation