I/O and Network Timing: "When and How Long Should You Hold a Lock and Wait?" — A Trade-off Framework for High-Traffic Sy
대용량 트래픽을 대응하고 비용을 아끼면서 트레이드오프를 감소시키는 방향으로 볼수있을까?
If you've ever stared at a pull request and wondered, "Will this actually hold up under real traffic?" — you're asking exactly the right question. Most performance disasters don't happen in local development. They happen when 10,000 users hit your system simultaneously, and suddenly that innocent-looking loop is firing 10,001 database queries instead of one.
This guide is built for engineers who care about trade-offs: frontend developers worried about bundle size and render waterfalls, backend developers dealing with query performance and connection limits, fullstack and product engineers balancing UX with system constraints, and platform or infrastructure engineers designing the safety nets that prevent cascading failure. The core shift is learning to read code not just for correctness, but for actual cost — in time, money, and system stability — at every layer of the stack.
Let's walk through the framework, clear up the most common misconceptions, and build a mental model you can apply in your next code review.
The Core Mindset Shift: Visualizing Real Cost Per Line of Code
Before diving into specific patterns, there's one fundamental reframe that changes how you evaluate any piece of code:
"When this line executes, what is the actual cost incurred in the browser, across the network, and on the server or database?"
Most engineers are trained to think about correctness and readability. Senior engineers add performance. But engineers who can operate at scale think in cost per execution path — CPU cycles, network round trips, memory allocation, and dollar spend on cloud infrastructure.
The four-lens framework below gives you a structured way to apply that thinking consistently, whether you're reviewing a React component, a REST endpoint, a database migration, or an infrastructure config.
Lens 1: I/O and Network Timing — "When and How Long Are You Blocking?"
The majority of bottlenecks in production systems are not CPU-bound. They are I/O-bound: your code is sitting idle, waiting for a network response, a disk read, or a database query to come back. Identifying and eliminating unnecessary waiting is the single highest-leverage optimization most teams can make.
Non-Blocking and Delegation: Stop Waiting for Work That Doesn't Need a Response
The question to ask: "Does this task need to complete before we respond to the user?"
If the answer is no — think analytics events, audit logs, welcome emails — then holding the API response open while you process it is pure waste. It increases latency for the user and keeps server resources occupied longer than necessary.
Real-world example: An e-commerce platform tracks "add to cart" events for analytics. A junior implementation awaits the analytics write inside the cart API handler. During a flash sale, 5,000 concurrent users hit the endpoint. The analytics write — which adds 80ms on average — now blocks 5,000 concurrent response threads. The fix is straightforward: push the event to a background queue (AWS SQS, Redis, or even a simple sendBeacon call on the frontend) and respond immediately. The analytics write happens asynchronously, user-facing latency drops, and the system handles 3x more concurrent requests without adding infrastructure.
Code patterns to look for in review:
navigator.sendBeacon()orfetchwithkeepalive: trueon the frontend (fire-and-forget telemetry)- SQS, BullMQ, or Redis-based queue delegation on the backend
- Background workers consuming from a queue rather than blocking API handlers
Batching and the N+1 Query Problem: Stop Making 101 Trips When One Will Do
The question to ask: "Am I making N network calls where I could make one?"
This is one of the most expensive anti-patterns in production systems, and it's remarkably easy to write without noticing.
N+1 Problem Example — The Classic Case
Imagine a dashboard that shows 10 recent orders, each with the associated product name:
// ❌ N+1 Anti-pattern: 1 query for orders + 10 queries for items = 11 DB round trips
const orders = await db.query('SELECT * FROM orders LIMIT 10');
for (const order of orders) {
const item = await db.query(
'SELECT * FROM items WHERE order_id = $1',
[order.id]
);
order.item = item;
}
This fires 11 separate database round trips. At 5ms per query (optimistic), that's 55ms of pure network overhead — before any actual computation. Under load, this becomes catastrophic.
// ✅ Batched approach: 2 queries total, regardless of how many orders
const orders = await db.query('SELECT * FROM orders LIMIT 10');
const orderIds = orders.map(o => o.id);
const items = await db.query(
'SELECT * FROM items WHERE order_id = ANY($1)',
[orderIds]
);
const itemMap = Object.fromEntries(items.map(i => [i.order_id, i]));
orders.forEach(order => { order.item = itemMap[order.id]; });
Two round trips. Always. Whether you have 10 orders or 10,000.
Where this pattern shows up across the stack:
- Frontend: Fetching user data for each item in a list via separate API calls inside a
useEffectloop - Backend (GraphQL): The classic DataLoader pattern solves this at the resolver level by batching and caching within a single request
- Backend (REST): Nested
forloops withawaitinside hitting any external service — database, third-party API, internal microservice - ORM-level: Sequelize or Prisma without
include/JOINwill silently generate N+1 queries
Payload size also matters here. SELECT * when you need three columns is a hidden cost — you're paying for network transmission and DB serialization of data you'll throw away. Always extract only the columns you need.
Lens 2: Caching and Lazy Execution — "Are You Repeating Expensive Work Unnecessarily?"
Caching: Stop Paying for the Same Answer Twice
The question to ask: "How often does this data actually change? Could I cache the result?"
Caching is layered, and every layer has a different cost-performance profile. The goal is to intercept the request at the cheapest possible layer:
- Browser cache (free, zero server cost) — HTTP
Cache-Control,ETag - CDN edge cache (Cloudflare, CloudFront) — eliminates origin server load entirely
- Application-level cache (Redis, Memcached) — eliminates DB queries
- Database query cache — eliminates disk I/O
A word on React 19 and Next.js fetch caching — because this is genuinely confusing:
React 19 introduces Request Memoization via the cache() function. Within a single server-side render pass, duplicate fetch calls to the same URL are deduplicated — only one network request fires. This is helpful for shared data across multiple Server Components in one render.
However, Next.js 15 changed the default: fetch is now uncached by default. The aggressive automatic caching from Next.js 13–14 caused too many stale-data bugs in production. As of Next.js 15, you must explicitly opt into caching:
// Cached for 60 seconds
const data = await fetch('/api/products', { next: { revalidate: 60 } });
// Permanently cached until manually invalidated
const data = await fetch('/api/config', { cache: 'force-cache' });
// Explicitly uncached (default in Next.js 15)
const data = await fetch('/api/live-prices', { cache: 'no-store' });
The trade-off to understand: Caching always introduces a consistency vs. freshness trade-off. Cached product prices might be 60 seconds stale. For a product catalog, that's fine. For a live stock ticker or a payment status, it's unacceptable. Always ask: "What is the cost of showing stale data here?" before caching.
Lazy Execution and Dynamic Import: Stop Paying Upfront for Work You Might Not Need
The question to ask: "Am I loading or computing something the user hasn't asked for yet?"
There are two distinct flavors of lazy execution that often get conflated:
1. Code-level laziness (Dynamic Import / Code Splitting)
// ❌ Loads the entire chart library in the initial bundle — even on pages with no chart
import { HeavyChartLibrary } from 'heavy-chart-lib';
// ✅ Loads only when the user actually navigates to the analytics page
const HeavyChartLibrary = React.lazy(() => import('heavy-chart-lib'));
Dynamic Import doesn't delay the execution of a computation — it delays the download of the code itself. If your initial JS bundle includes a PDF generation library that only 2% of users ever trigger, you're forcing 100% of users to download and parse code they'll never use. That's pure tax on Time-to-Interactive (TTI).
2. Data-level laziness (Lazy Loading / On-demand fetching)
This means deferring database queries or API calls until the data is actually needed, rather than eagerly fetching everything upfront. A classic ORM-level example is only loading related entities when you traverse the relationship, not on initial load.
The shared philosophy: "Costs that aren't needed right now should be deferred." Both patterns reduce initial load — one on the network/parse dimension, one on the DB/API dimension.
Lens 3: Serverless Economics — "Will This Architecture Explode Under Load?"
The Serverless Misconception: "More Traffic = Less Cost"
This is one of the most dangerous misunderstandings in modern cloud architecture, and it's particularly common among engineers who are new to serverless.
The question to ask: "If traffic spikes 100x, does my cost spike linearly — or exponentially?"
Fixed infrastructure (EC2, ECS) has a ceiling: traffic above capacity degrades or fails, but your monthly bill stays predictable. Serverless (AWS Lambda, Vercel Functions) scales up gracefully — but billing is per invocation × execution duration × memory. A DDoS attack, a buggy infinite-retry loop, or a viral moment can turn a $200/month function into a $40,000 surprise invoice in 24 hours.
The DB connection problem is the bigger architectural risk:
Traditional servers maintain a connection pool — say, 20 persistent connections to PostgreSQL. Under load, requests queue behind those 20 connections. Manageable.
Serverless scales differently. When 2,000 concurrent Lambda instances spin up, each one tries to open its own DB connection. PostgreSQL has a hard connection limit (typically 100–500 depending on instance size). At 2,000 simultaneous connection attempts, the database falls over — not because of query load, but because of connection exhaustion.
❌ Without a pooler:
[2,000 Lambda instances] → [2,000 DB connections] → DB crashes
✅ With a connection pooler:
[2,000 Lambda instances] → [RDS Proxy / PgBouncer] → [20–50 DB connections] → DB healthy
Solutions: AWS RDS Proxy, PgBouncer, Prisma Data Proxy, PlanetScale, or Neon's serverless driver all solve this by sitting between your serverless functions and the database and managing connection multiplexing.
The Right Architecture: Serverless as the Front, Queues as the Buffer
Serverless is genuinely the right tool for absorbing unpredictable traffic spikes. The key insight is that serverless should absorb the traffic without directly transmitting that load to downstream systems:
[Explosive user traffic]
│
▼
[Serverless — Lambda / Vercel] ──→ Returns "Request received" immediately
│
▼ (enqueue)
[AWS SQS / Redis / Kafka] ← Buffer that absorbs the spike
│
▼ (consume at safe rate)
[Worker / DB] ← Processes at a rate it can handle
This architecture decouples ingestion rate from processing rate. The serverless layer scales infinitely to accept requests. The queue smooths out the burst. The worker processes at a sustainable pace. The user gets an immediate acknowledgment. Nobody waits, nothing falls over.
Hybrid architecture for mature systems: Baseline traffic (the steady stream that's always there) runs on containerized infrastructure (ECS, EKS) because it's far more cost-efficient per request. Serverless handles only the burst capacity — the flash sale, the press mention, the viral spike. This ratio optimization is where significant infrastructure cost savings come from at scale.
Mandatory safety valves when using serverless:
- Reserved Concurrency: Cap Lambda at 500 concurrent instances to protect downstream systems
- Timeout configuration: If a Lambda waits indefinitely for a slow DB, you pay for every second. Set aggressive timeouts (2–3 seconds) and fail fast
- Circuit Breaker pattern: Stop sending traffic to a failing dependency rather than letting retry storms amplify the failure
Comprehensive Trade-off Checklist: What to Look for in Every Code Review
Pulling all of the above together into a practical tool you can apply immediately:
The three questions to ask on every PR:
- "If this function fails or slows down, does the entire system or UX halt?" → Single point of failure — candidate for async decoupling
- "If this hits 10,000 TPS, what breaks first — the DB, the server, or the external API?" → Identify the bottleneck and add caching or
Comments
Loading comments...