cloudrop.it

Infrastructure

A private compute fabric.

Four regions, one control plane, one identity plane, and a restore rehearsed every week. Every dependency in the request path is something the fabric can drain, move or replace on its own.

Recovery point objective

15 min

Worst case exposure on a full region loss. Recovery time objective is 40 minutes, measured on the weekly rehearsal rather than estimated.

The map

Four regions. One address.

  • WAW1 Operational

    Warsaw

    Europe

    8ms

    p50 in-region

  • FRA1 Operational

    Frankfurt

    Europe

    11ms

    p50 in-region

  • SIN1 Operational

    Singapore

    Asia

    24ms

    p50 in-region

  • TYO1 Operational

    Tokyo

    Asia

    19ms

    p50 in-region

Every region carries the full stack: edge, identity, application, state. A region is drained and refilled without a ticket, because there is no ticket. Traffic follows the catalogue, so an Asian lane is served from Asia.

The numbers

Regions
4

Warsaw, Frankfurt, Singapore, Tokyo.

Availability
99.99%

Platform target, contracted on the top tier.

Events / day
8.4M

Ingested, deduplicated, reconciled.

Geographic vaults
3

Encrypted replication targets, all off-region.

The stack

Five planes, one spine.

  1. Edge

    TLS terminates at the edge of every region. What sits behind it answers to the fabric and to nothing else.

    Regional entry

    anycast address · automated certificates · per-host routers

    Rate control

    per-tenant budgets · burst absorption · request shedding

  2. Identity

    One identity plane in front of every request, in every region. Hand-rolled auth has never shipped here.

    Single issuer

    OIDC · authorization code · hardware keys supported

    Session fabric

    host-bound · re-checked on every mutation, not on entry

  3. Application

    Thirty-two modules inside one process boundary, by decision. One deployable, one version, one place a bug can hide.

    Application plane

    typed end to end · schema-validated on every input

    Scheduler

    idempotent jobs · 90-second push cadence · replayable

    Model gateway

    every AI call routed · budget-gated · fully logged

  4. State

    State is regional. Reads stay local to the region that serves them, writes stay ordered.

    Relational core

    one database per platform · forward-only migrations

    Event spine

    streams · consumer groups · at-least-once per module

    Object store

    S3 API · presigned URLs · buckets pinned to a region

  5. Durability

    A backup nobody has restored is a rumour. This plane exists to keep it from becoming one.

    Vault replication

    three geographic vaults · encrypted before it moves

    Key custody

    keys held inside the fabric · never travel with the ciphertext

Why

Owning the fabric is what makes the guarantees real.

Every promise on this page reduces to one question: who can move the machine. A platform assembled from managed parts inherits its ceiling from whoever runs those parts. Its recovery time is a support queue. Its egress cost is somebody else's pricing page. Its worst day is scheduled by a vendor.

Cloudrop runs on capacity it controls, in four regions, behind one identity plane. Recovery time is a measured number because we own the restore. Model spend has a ceiling because the gateway is ours to compile the ceiling into. A region drains and refills on our schedule, at our hour, for our reasons.

A dependency you cannot restart is not a dependency. It is a landlord.

The honest cost: the pager is ours. Certificates, kernel updates, disk headroom, vault verification, the weekly restore. That is priced in and written down as runbooks, which is the part most infrastructure stories quietly skip.

Durability

The restore, with times in it.

  1. Every 15 min

    Snapshot taken

    A consistent snapshot of the relational core plus an object manifest, checksummed inside the region that owns it. This cadence is the recovery point objective: fifteen minutes, worst case.

  2. T+90s

    Encrypted, then replicated

    Encryption happens inside the fabric. What crosses a region boundary is ciphertext and a checksum, fanned out to three geographic vaults on three separate failure domains.

  3. T+6m

    Verified against the manifest

    Every vault compares object count, size and checksum against the source manifest. A mismatch pages. Silence pages louder.

  4. Weekly

    Restore rehearsed

    A vault copy is restored into a throwaway region and a fixed set of row counts is asserted end to end. Forty minutes, measured, is where the recovery time objective comes from.

Recovery point · 15 min

Fifteen minutes of exposure, worst case. The push workers are idempotent, so a replay after a restore converges on the truth instead of duplicating it.

Recovery time · 40 min

Forty minutes, end to end, taken from the weekly rehearsal. The number is measured because an unrehearsed figure is a guess with a decimal point on it.

Discipline

Rules that outrank convenience.

  • Migrations run forward. Every schema change runs once, in order, against a ledger. Forward is the only direction the runner knows, which keeps a bad deploy a one-step problem instead of a second untested deploy at the worst hour.
  • State has an address. Every byte of state resolves to a path a human can find at three in the morning, without knowing a container runtime's internals.
  • Every boundary validates. Request, response, environment, config. A schema at the door turns what would have been an incident into a rejected payload and a log line.
  • Cutovers run dual-stack. Old and new hostnames serve in parallel, certificates are issued ahead of the flip, and the rollback is one environment variable.
  • Production is the proof. Every change is confirmed on the running platform before it counts as shipped. That is the whole bar, and it is the one that holds.