R RenderComp
remotion aws-lambda video-automation serverless cost-optimization

Total Cost of a Video Automation Pipeline

By RenderComp Team Editorial policy

Remotion treats video as a pure function: given a frame number and a set of props, the output is always the same pixel grid. That determinism makes Remotion Lambda viable as a distributed render backend, because a 300-frame composition can be split into independent chunks with no shared state between workers. Each chunk lands on a separate AWS Lambda function, and each function is billed by the GB-second.

Teams now build video pipelines the same way they build any other API-driven workflow: a request goes in, a video file comes out. The compute that produces the file is rented by the second. At scale, the bill is driven by three numbers: how many chunks run in parallel, how much memory each chunk uses, and how long each chunk runs.

Lambda memory can be set anywhere from 128 MB to 10,240 MB in 1-MB increments. A single invocation runs for at most 15 minutes. Those two bounds define the extremes of what any one render chunk can cost, and the space between them is where cost control happens.


The GB-Second Formula

The billing unit is memory in GB multiplied by wall-clock seconds. A function allocated 2,048 MB (2 GB) running for 60 seconds accumulates 120 GB-seconds. Four 512 MB functions each running 60 seconds also accumulate 120 GB-seconds: 4 × 0.5 × 60. Adding parallel workers does not reduce the GB-second total. It reduces wall-clock latency.

What concurrency does affect is throughput. AWS Lambda limits concurrent executions to 1,000 per region by default, and that limit can be raised. A render that spawns 1,500 chunks will queue 500 of them even at maximum throughput. Queued functions do not accumulate GB-seconds while waiting, but they extend wall-clock completion time.

A typed wrapper makes the formula explicit and keeps inputs auditable:

type ChunkConfig = {
  chunkCount: number;
  memoryMb: number;                 // 128–10,240 in 1-MB steps
  wallClockSecondsPerChunk: number; // measure from a test render
};

function gbSeconds({ chunkCount, memoryMb, wallClockSecondsPerChunk }: ChunkConfig): number {
  return chunkCount * (memoryMb / 1024) * wallClockSecondsPerChunk;
}

wallClockSecondsPerChunk is the one value you measure rather than configure. Run a test render with your actual composition at your target framesPerLambda setting, check the chunk timing in the render log, and use that number. Everything else comes directly from your renderMediaOnLambda call parameters.


The 1,769 MB vCPU Threshold

AWS assigns compute capacity proportionally to memory allocation. At 1,769 MB, a function receives the equivalent of one vCPU. Below that mark it gets a fractional allocation; above it, more than one.

For compositions that are CPU-bound, such as those layering multiple <Video> elements, running frame-by-frame physics updates, or performing heavy canvas draws, crossing 1,769 MB can cut wall-clock time enough to offset the higher cost per second. A composition driven by useCurrentFrame and a few interpolate calls renders at nearly the same speed across a wide memory range, because the bottleneck there is frame I/O, not compute.

function effectiveVcpus(memoryMb: number): number {
  // AWS assigns compute proportionally to memory.
  // 1,769 MB is the documented 1-vCPU crossover point.
  return memoryMb / 1769;
}

// Choosing between two memory settings for a CPU-bound composition:
// 1. Run a test render at each memory value with the same framesPerLambda.
// 2. Record wallClockSecondsPerChunk for each run.
// 3. Compute (memoryMb / 1024) * wallClockSecondsPerChunk.
// 4. The lower product is cheaper.

The four-step test replaces guesswork with a figure specific to your composition.


The 15-Minute Cap and Chunk Sizing

Each Lambda invocation can run for at most 15 minutes (900 seconds). For Remotion, this ceiling bounds the maximum useful framesPerLambda value: if a single chunk takes longer than 900 seconds to render, the function times out and that chunk’s output is lost.

Chunk size pulls cost in two opposing directions. Larger chunks mean fewer parallel workers and lower per-render invocation overhead, but they push individual chunks closer to the 15-minute limit. Smaller chunks give more parallelism (up to the 1,000-concurrent ceiling) and stay safely below the timeout, but they raise chunk count along with the per-request charge that comes with each invocation.

import { renderMediaOnLambda } from '@remotion/lambda/client';

const result = await renderMediaOnLambda({
  region: 'us-east-1',
  functionName: 'remotion-render-4-0-272-mem3000mb-disk2048mb-240sec',
  composition: 'MyComposition',
  serveUrl: 'https://your-serve-url.com',
  inputProps: {},
  framesPerLambda: 30,           // tune to keep chunks well under 900 s
  timeoutInMilliseconds: 840_000, // 14 min: 60-second buffer before Lambda hard-terminates
  codec: 'h264',
  outName: 'output.mp4',
});

// estimatedPrice.accruedSoFar updates as chunks complete.
// Log it mid-render to catch unexpectedly expensive jobs before they finish.
console.log(`Cost so far: $${result.estimatedPrice.accruedSoFar}`);

Setting timeoutInMilliseconds to 840,000 rather than the full 900,000 gives Remotion a window to handle graceful shutdown instead of getting hard-terminated by the Lambda runtime with no cleanup.


Putting the Numbers Together

Remotion’s documentation states that most users render multiple minutes of video for just a few pennies per render. Their published cost examples hold a single Lambda configuration constant across all compositions, which makes GB-second totals directly comparable between them. That fixed-configuration discipline is worth applying to your own benchmarks: hold memory and framesPerLambda constant, vary only the composition content, and the GB-second output isolates the cost variable you actually care about.

The full estimator:

type PipelineCostInput = {
  durationInFrames: number;
  framesPerLambda: number;
  memoryMb: number;
  wallClockSecondsPerChunk: number; // measured under a fixed config
  pricePerGbSecond: number;         // your region's current rate from AWS Lambda pricing
};

function estimatePipelineCost(input: PipelineCostInput) {
  const chunkCount = Math.ceil(input.durationInFrames / input.framesPerLambda);
  const memoryGb = input.memoryMb / 1024;
  const totalGbSeconds = chunkCount * memoryGb * input.wallClockSecondsPerChunk;
  return {
    chunkCount,
    totalGbSeconds,
    estimatedCost: totalGbSeconds * input.pricePerGbSecond,
  };
}

// Example: 450 frames (15 s at 30 fps), 45 framesPerLambda, 2,048 MB
const estimate = estimatePipelineCost({
  durationInFrames: 450,
  framesPerLambda: 45,
  memoryMb: 2048,                     // just above the 1,769 MB vCPU threshold
  wallClockSecondsPerChunk: 18,       // from a test render at fixed config
  pricePerGbSecond: YOUR_REGION_RATE, // look up at aws.amazon.com/lambda/pricing
});

// estimate.chunkCount      → 10
// estimate.totalGbSeconds  → 10 × 2 × 18 = 360
// estimate.estimatedCost   → 360 × YOUR_REGION_RATE

The last check before scaling is concurrency: if estimate.chunkCount exceeds 1,000, either request a concurrency limit increase through AWS Support or raise framesPerLambda to reduce chunk count. Raising framesPerLambda pushes individual chunks closer to the 15-minute ceiling, so the two levers constrain each other. Knowing both limits before production load reveals them is the difference between a pipeline that costs pennies per render and one that fails silently at 2 a.m.

Now available

Get 1,000+ Remotion Templates

Pay once — no subscription. Lifetime updates. TypeScript-first.

View pricing →

Free 50

Get the 50 templates as a ZIP

Enter your email and the ZIP link arrives right away.

Send me the 50-template ZIP. I agree to receive RenderComp template updates and product news, including the paid library (a few emails, one-click unsubscribe). Privacy policy

You do not have to use email. The GitHub repository stays public and needs no signup. Open the repository