ibee
Origin Architecture for Cloudflare: Where App, Storage, and APIs Should Run

Origin Architecture for Cloudflare: Where App, Storage, and APIs Should Run

SabariDevOps Engineer
September 18, 20269 min read

Most teams treat Cloudflare like a magic front door. Then they discover the real problem: their “origin” design quietly forces every request to bounce across regions, every upload to miss caching, and every authenticated API call to depend on session state that nobody actually plans for.

The migration took six weeks. Three of them were spent undoing a decision made in the first hour. Not because Cloudflare failed. Because your origin architecture did not match how your app reads and writes data.

You do not get to design this by vibes. You design it by access patterns, state management, and failure behavior.

The failure mode: you chose origins before you chose data behavior

Here is what nobody tells you about Cloudflare origin design: Cloudflare can accelerate and protect traffic, but it cannot replace your need for correct data placement, state management, and origin capacity planning. If your app depends on fast reads and low-latency writes, your origin placement has to honor that. If your API is authenticated and personalized, caching rules and session strategy matter more than CDN speed.

Most teams get this wrong in one of two ways.

First, they centralize everything “for simplicity” in one region, then wonder why global users feel laggy. You can hide some latency with caching for static assets, but you cannot hide it for every authenticated API call and every database write.

Second, they copy a multi-region pattern from another company without mapping it to their own state model. They replicate compute, then forget sessions. They replicate storage, then forget consistency. They add failover, then discover their writes are not idempotent, so retries create duplicates.

Your origin architecture is not a diagram. It is a set of decisions about where your application code runs, where your storage lives, and how your data access path behaves under load and failure.

Required visual: where app, API, and storage should actually land

Cloudflare edge routing to app/API servers, object storage, and database with cacheable and non-cacheable paths

Shows how cacheable requests can stay at the edge while dynamic requests reach the correct regional origins.

Start with three questions, not a region count

You only need three answers to place your app, API, and storage correctly.

The first question: what requests are cacheable, and what headers and semantics do you control? If your static assets have stable URLs and correct cache headers, you can push a lot of traffic to the edge and reduce origin load. If your “GET /api/profile” returns user-specific data with no-store semantics, caching will not save you. Treat it as dynamic and plan for origin latency and capacity.

The second question: where does state live? Your state is not “the database.” Your state is also sessions, tokens, job queues, and any in-memory assumptions you made to get through early development. If you use server-side sessions stored in memory, you have already committed to session affinity or shared session storage. If you rely on background jobs, your queue and worker placement becomes part of origin architecture too.

The third question: what is your failure behavior? When an origin region degrades, do you want to fail open with stale reads, fail closed with errors, or reroute to a secondary origin? Cloudflare can route traffic to multiple origins, but your app must tolerate partial failures. That means timeouts, retries with idempotency keys, and “withdraw routes instead of blackholing traffic” behavior at the infrastructure level.

Once you answer these, the “single region vs multi region” debate gets simpler. You are not choosing a topology for its own sake. You are choosing it to satisfy your access patterns.

The three origin patterns that actually work

You can map most real systems into three origin patterns. Each has clear trade-offs.

Pattern 1: Single-region origin (baseline, but only if your users are not global pain)

You run app, API, and database in one primary region. Cloudflare handles global entry, security, and caching for cacheable content.

This pattern wins when your product does not require tight global latency for authenticated flows, and when your consistency model stays simple. Debugging is straightforward because you have one place to look.

The downside is predictable: users far from the region pay latency tax on every dynamic request, and you have one region that must carry the write load.

Use this when you are still stabilizing your data model and you want fewer moving parts. Do not use it if your core user journeys require sub-100ms time-to-response worldwide. You will end up building multi-region later under pressure, with more migration risk.

Pattern 2: Multi-region compute with a deliberate storage strategy

You run app and API in multiple regions. Storage can be either replicated or centralized, depending on your consistency needs and what you can tolerate during failover.

This pattern wins when you need global latency for dynamic traffic. It also improves availability if you design failover correctly.

The downside is the one most teams underestimate: state. If your API writes to a centralized database, you still pay cross-region latency for writes. If you replicate storage, you must handle conflicts and consistency. If you use sessions, you must decide between session affinity, shared session storage, or token-based stateless auth.

A good multi-region design is not “active-active everywhere.” It is “active where it matters, and centralized where it reduces complexity.” You pick the trade-off knowingly, not accidentally.

Pattern 3: Edge-heavy with origin specialization (the pragmatic split)

You push static and cacheable content to the edge and keep dynamic compute limited to a smaller set of regional origins. Storage gets split by access pattern: object storage for media and documents, databases for transactional data, and caches for hot reads.

This is often the best starting point for teams that have mixed workloads. Static assets get fast delivery without touching your app. Dynamic APIs hit your compute. Media uploads and downloads go to object storage optimized for that workload.

The trade-off is operational clarity. You must keep your boundaries clean: do not route everything through the app “because it’s easier.” Define which services own which data paths, then enforce it with routing and deployment rules.

This pattern also helps cost, because origin capacity becomes proportional to dynamic traffic, not total traffic.

Where to run each component: app, API, storage

App compute and dynamic APIs

Your app servers and API servers should run in the region(s) that minimize end-to-end latency to the data they actually touch. That includes database reads and writes, cache lookups, and any object storage calls you do during request handling.

If your API is authenticated and personalized, assume it will be non-cacheable. That means you should focus on connection reuse, autoscaling behavior, and predictable tail latency. Also plan for retry safety. Retries happen. Your origin must not turn a retry into a second charge, a duplicate order, or a double write.

If you use server-side sessions, decide now how sessions survive failover. Most teams either force sticky sessions at the load balancer layer or they migrate to token-based stateless auth. Both are valid. Neither is “we’ll figure it out later.”

Object storage for user uploads and media

Object storage usually belongs in a place optimized for throughput and durability, not in the same region as your database. The key is that object storage requests are typically independent from your transactional state, and they can be served efficiently with caching at the edge.

If your users upload images, videos, documents, or exports, treat object storage as a first-class origin. Then let Cloudflare cache downloads aggressively when it is safe.

Also, stop thinking of “disk” and start thinking of “objects.” If you store user uploads on VM disks, you couple your storage lifecycle to your compute lifecycle. That is how you end up with brittle scaling and painful migrations later.

Databases and the “single source of truth”

Your database (or whatever holds the authoritative state) needs the simplest consistency story you can manage while meeting your latency SLOs.

Single-region databases are simplest. Multi-region databases can be necessary, but they come with a cost in complexity. If you do multi-region writes, you must design conflict handling and idempotency. If you do multi-region reads with a single primary writer, you must design for replication lag and stale reads.

Most teams get this wrong by treating replication as a performance feature. It is a correctness feature too. Decide what correctness guarantees you need for each entity: orders, balances, profiles, audit logs, and so on.

A practical routing and placement checklist

You can turn the above into a buildable plan without over-engineering.

First, tag every endpoint and workflow by cacheability and statefulness. GET endpoints that return public content can be cacheable. Authenticated endpoints that return user-specific data usually are not. Upload endpoints are special: they should go directly to object storage when possible, or through a very thin upload service that writes to storage and returns an object key.

Second, decide your session strategy. If you rely on server-side sessions, plan for shared storage or sticky routing. If you use stateless tokens, make sure your token validation does not require round trips to a centralized service that becomes a bottleneck.

Third, design idempotency for every write path that can be retried. That includes “create” endpoints and background job triggers. If you cannot make writes idempotent, retries will corrupt your data the moment you have a transient origin failure.

Fourth, keep your origin capacity model honest. Your app autoscaling can handle spikes, but your database and storage cannot be treated as infinite. If you scale compute without scaling storage throughput and IOPS, you will just move the bottleneck downstream.

Finally, test failover with real traffic patterns. Do not just simulate 500 errors. Simulate region degradation, increased latency, and partial outages. Then verify your timeouts and retry behavior match your correctness expectations.

If you are using IBEE for object storage, you can keep uploads and downloads S3-compatible so you do not rewrite app code while you adjust routing and caching. That matters because origin architecture changes are already stressful; you do not want to add a storage migration and a code migration at the same time.

One concrete action you can take today

Write a one-page “origin map” for your app that lists every endpoint and workflow with three labels: cacheable or not, stateful or stateless, and which data store it touches. Then pick one workflow that is currently painful (usually authenticated API calls or uploads) and redesign its origin path first. Do not start with region count. Start with the access pattern that hurts you.

Related articles