R RenderComp
remotion render-queue typescript self-hosted video-rendering worker

Remotion Render Queue with States, Retries, and Bundle Caching

By RenderComp Team Editorial policy

renderMedia() blocks the current process for the full duration of a render. On a 30-fps, 300-frame composition, that can mean 10–90 seconds of CPU time depending on scene complexity. Fire two of those concurrently without coordination and you are competing for the same CPU cores, the same browser slots, and the same disk I/O. A queue is not optional infrastructure. It is the missing layer between “a script that renders one video” and “a system that renders many.”

Remotion’s renderer is deterministic: given the same bundle, the same composition ID, and the same inputProps, a render always produces the same output. That determinism makes retry logic safe. A failed job can be re-enqueued with the identical payload and produces the same artifact on success, with no side effects from the failed attempt.

The design below uses SQLite (via better-sqlite3) for job storage, a four-state machine, exponential-backoff retries, and a bundle cache that skips webpack recompilation between jobs.


The Job State Machine

Every render job moves through four states: queued → rendering → done, or queued → rendering → failed. A job that fails with retries remaining follows a different path: rendering → queued, with an incremented attemptCount and a future retryAfter timestamp.

// types.ts
export type JobStatus = 'queued' | 'rendering' | 'done' | 'failed';

export interface RenderJob {
  id: string;
  compositionId: string;
  inputProps: Record<string, unknown>;
  outputPath: string;
  status: JobStatus;
  attemptCount: number;
  maxAttempts: number;
  retryAfter: number | null; // unix ms; null means eligible immediately
  progress: number;          // 0–1
  heartbeat: number | null;  // updated during rendering
  error: string | null;
  createdAt: number;
  updatedAt: number;
}

The retryAfter field carries the backoff logic. Rather than sleeping in the worker, the worker skips any job where retryAfter > Date.now(), and the next poll cycle picks it up without any timer state in the worker process itself.

One edge case that is easy to miss: a worker process that crashes mid-render leaves the job stuck in rendering indefinitely. The fix is a heartbeat timestamp updated during onProgress. Any job in rendering with a heartbeat older than a threshold is treated as orphaned and reset to queued at the start of the next claim transaction.


Schema and Persistence

SQLite works well for single-host queues because it gives you ACID transactions. When multiple workers compete for the same job, the claim operation must be atomic. A SELECT followed by a separate UPDATE creates a window for two workers to claim the same row.

// db.ts
import Database from 'better-sqlite3';
import type { RenderJob } from './types';

export const db = new Database('./queue.db');

db.exec(`
  CREATE TABLE IF NOT EXISTS jobs (
    id TEXT PRIMARY KEY,
    compositionId TEXT NOT NULL,
    inputProps TEXT NOT NULL,
    outputPath TEXT NOT NULL,
    status TEXT NOT NULL DEFAULT 'queued',
    attemptCount INTEGER NOT NULL DEFAULT 0,
    maxAttempts INTEGER NOT NULL DEFAULT 3,
    retryAfter INTEGER,
    progress REAL NOT NULL DEFAULT 0,
    heartbeat INTEGER,
    error TEXT,
    createdAt INTEGER NOT NULL,
    updatedAt INTEGER NOT NULL
  )
`);

export function enqueue(
  job: Pick<RenderJob, 'id' | 'compositionId' | 'inputProps' | 'outputPath' | 'maxAttempts'>
): void {
  const now = Date.now();
  db.prepare(`
    INSERT INTO jobs (id, compositionId, inputProps, outputPath, maxAttempts, createdAt, updatedAt)
    VALUES (?, ?, ?, ?, ?, ?, ?)
  `).run(job.id, job.compositionId, JSON.stringify(job.inputProps), job.outputPath, job.maxAttempts, now, now);
}

export function claimNextJob(): RenderJob | null {
  const STALE_MS = 5 * 60 * 1000; // 5 minutes without a heartbeat = orphaned

  const resetStale = db.prepare(`
    UPDATE jobs
    SET status = 'queued', heartbeat = NULL, updatedAt = ?
    WHERE status = 'rendering' AND heartbeat < ?
  `);

  // UPDATE … RETURNING requires SQLite 3.35+.
  const claim = db.prepare(`
    UPDATE jobs
    SET status = 'rendering', heartbeat = ?, updatedAt = ?
    WHERE id = (
      SELECT id FROM jobs
      WHERE status = 'queued'
        AND (retryAfter IS NULL OR retryAfter <= ?)
      ORDER BY createdAt ASC
      LIMIT 1
    )
    RETURNING *
  `);

  const transaction = db.transaction((): RenderJob | null => {
    const now = Date.now();
    resetStale.run(now, now - STALE_MS);
    const row = claim.get(now, now, now) as any;
    if (!row) return null;
    return { ...row, inputProps: JSON.parse(row.inputProps) };
  });

  return transaction();
}

Storing a Promise rather than the resolved path in the bundle cache (shown next) means two workers that simultaneously need the same bundle wait on a single webpack run instead of each starting their own.


Bundle Caching

bundle() from @remotion/bundler invokes webpack. The first call for a given entry point takes several seconds. For a queue processing back-to-back jobs that all share the same composition file, re-bundling on every job is wasteful and unnecessary.

The bundle output path is stable for the lifetime of the process. Workers reuse it across jobs, and the browser instance can keep its internal page cache warm between renders. When you deploy an updated composition, restarting the worker process clears the cache naturally.

// bundleCache.ts
import { bundle } from '@remotion/bundler';

const cache = new Map<string, Promise<string>>();

export function getCachedBundle(entryPoint: string): Promise<string> {
  if (!cache.has(entryPoint)) {
    cache.set(
      entryPoint,
      bundle({ entryPoint }).catch((err) => {
        // Remove the failed entry so the next caller retries bundling.
        cache.delete(entryPoint);
        throw err;
      })
    );
  }
  return cache.get(entryPoint)!;
}

Storing a Promise rather than the resolved string is the key detail. If two workers call getCachedBundle before the first bundle() resolves, both await the same Promise. Without this, each would start a separate webpack compilation.


The Worker Loop

Each worker polls for a job, renders it with renderMedia(), and writes the result back to the database. The onProgress callback lets you record progress at frame resolution without blocking the render itself.

// worker.ts
import { randomUUID } from 'crypto';
import { getCompositions, renderMedia } from '@remotion/renderer';
import { claimNextJob, db } from './db';
import { getCachedBundle } from './bundleCache';
import type { RenderJob } from './types';

const ENTRY_POINT = './src/index.ts'; // adjust to your project
const POLL_INTERVAL_MS = 1_000;

async function processJob(job: RenderJob): Promise<void> {
  const serveUrl = await getCachedBundle(ENTRY_POINT);

  const compositions = await getCompositions(serveUrl, {
    inputProps: job.inputProps,
  });
  const composition = compositions.find((c) => c.id === job.compositionId);
  if (!composition) {
    throw new Error(`Composition "${job.compositionId}" not found in bundle`);
  }

  await renderMedia({
    composition,
    serveUrl,
    codec: 'h264',
    outputLocation: job.outputPath,
    inputProps: job.inputProps,
    onProgress: ({ progress }) => {
      db.prepare('UPDATE jobs SET progress = ?, heartbeat = ?, updatedAt = ? WHERE id = ?')
        .run(progress, Date.now(), Date.now(), job.id);
    },
  });
}

async function workerLoop(): Promise<never> {
  while (true) {
    const job = claimNextJob();

    if (!job) {
      await new Promise((r) => setTimeout(r, POLL_INTERVAL_MS));
      continue;
    }

    try {
      await processJob(job);
      db.prepare(`UPDATE jobs SET status = 'done', progress = 1, updatedAt = ? WHERE id = ?`)
        .run(Date.now(), job.id);
    } catch (err) {
      const errorMessage = err instanceof Error ? err.message : String(err);
      const nextAttempt = job.attemptCount + 1;

      if (nextAttempt >= job.maxAttempts) {
        db.prepare(`UPDATE jobs SET status = 'failed', error = ?, updatedAt = ? WHERE id = ?`)
          .run(errorMessage, Date.now(), job.id);
      } else {
        // Exponential backoff: 4 s, 8 s, 16 s, capped at 30 s.
        const backoffMs = Math.min(2_000 * 2 ** nextAttempt, 30_000);
        db.prepare(`
          UPDATE jobs
          SET status = 'queued', attemptCount = ?, retryAfter = ?, error = ?, updatedAt = ?
          WHERE id = ?
        `).run(nextAttempt, Date.now() + backoffMs, errorMessage, Date.now(), job.id);
      }
    }
  }
}

const CONCURRENCY = 2;
for (let i = 0; i < CONCURRENCY; i++) {
  workerLoop();
}

CONCURRENCY = 2 is a reasonable starting point for a machine with 4–8 logical cores. renderMedia() already spawns multiple browser tabs internally, so setting concurrency higher than your available core count produces diminishing returns and can exhaust the browser’s tab pool.


Retry Backoff

The formula Math.min(2_000 * 2 ** attempt, 30_000) produces these delays:

attemptCount after failureretryAfter delay
14 000 ms
28 000 ms
316 000 ms
4+30 000 ms (capped)

With maxAttempts: 3, a job exhausts retries after the third failure. Attempt 0 is the initial try, so after the first failure attemptCount is 1 and the backoff is 4 seconds. After the second failure it is 8 seconds. The third failure moves the job to failed with no further backoff.

Retries fix transient errors: a browser process that crashed, a moment of disk pressure, a temporary network timeout fetching a remote asset. They have no effect on permanent errors, such as a composition that throws because inputProps failed schema validation. The error column on jobs in failed status tells you which kind you are dealing with.


Enqueuing Jobs

import { randomUUID } from 'crypto';
import { enqueue } from './db';

enqueue({
  id: randomUUID(),
  compositionId: 'SalesRecap',
  inputProps: { quarter: 'Q3', revenue: 1_420_000 },
  outputPath: `/renders/${randomUUID()}.mp4`,
  maxAttempts: 3,
});

The worker loop picks this up on its next poll cycle. With POLL_INTERVAL_MS at 1 000 ms, the median latency from enqueue to render start is 500 ms, which is acceptable for background video jobs. If your application needs lower latency, replace the polling loop with a NOTIFY/LISTEN pattern on PostgreSQL or a Redis Streams consumer, and remove the setTimeout entirely.

Now available

Get 1,000+ Remotion Templates

Pay once — no subscription. Lifetime updates. TypeScript-first.

View pricing →

Free 50

Get the 50 templates as a ZIP

Enter your email and the ZIP link arrives right away.

Send me the 50-template ZIP. I agree to receive RenderComp template updates and product news, including the paid library (a few emails, one-click unsubscribe). Privacy policy

You do not have to use email. The GitHub repository stays public and needs no signup. Open the repository