seifashraf.tech
01work02notes03blog04contact
CV
open to work
All writing

Published 8 October 2026· 5 min read

BullMQ jobs that run once: the outbox pattern in NestJS

Calling queue.add() next to a database write loses jobs or invents them. Here is the outbox setup I run in production with NestJS, Postgres, and BullMQ.

Written by Seif Ashrafnestjsbullmqpostgresqlredisarchitecture

Most NestJS and BullMQ examples enqueue a job right after saving a row. That works until the process dies between the two calls, and then you either lose a job or run one for data that never existed.

On Lov, a speech-therapy platform, jobs send booking emails and react to payments. Neither can be lost and neither can happen twice. This is the setup that got me there.

1. The problem with queue.add() in a request

There are two obvious places to enqueue, and both are wrong.

Where you enqueueWhat a crash does
After the transaction commitsThe row is saved, the job is never added. Lost work.
Inside the transactionThe job is in Redis, then the transaction rolls back. A phantom job.

The second one is worse. A handler emails a patient about a booking that does not exist.

Postgres and Redis cannot share a transaction. So the fix is to stop writing to Redis in the request at all.

2. Write the event in the same transaction

Every state change that needs a side effect also inserts a row into an outbox table. Same transaction, same commit.

await db.transaction(async (tx) => {
  const booking = await bookings.confirm(tx, bookingId);
  await tx.insert(outbox).values({
    eventType: "booking.confirmed",
    aggregateId: booking.id,
    payload: { bookingId: booking.id },
  });
});

If the booking rolls back, the event rolls back with it. There is nothing in Redis to clean up, because nothing was sent there.

3. A relay that stamps first, then publishes

A relay runs as a continuous loop in the worker process. It is not a queued job. Queueing the thing that feeds the queue puts it behind its own output.

Each pass claims a batch of unpublished rows, stamps them, commits, and only then publishes.

const claimed = await db.transaction(async (tx) => {
  const rows = await tx.execute(sql`
    SELECT id, event_type, aggregate_id, payload, occurred_at
    FROM outbox
    WHERE relayed_at IS NULL
    ORDER BY occurred_at
    LIMIT ${batchSize}
    FOR NO KEY UPDATE SKIP LOCKED
  `);
  if (rows.length) await stampRelayed(tx, rows.map((r) => r.id));
  return rows;
});
 
await publish(claimed); // after the commit, never inside it

SKIP LOCKED lets several relays run at once without claiming the same rows.

The order matters. Publishing inside the claiming transaction brings back the phantom job from section 1. Committing the stamp first can never invent an event. It does open a different gap, a row that is stamped but never published, and section 5 closes it.

4. Use the outbox id as the job id

await queue.add(event.eventType, toWire(event), {
  jobId: event.id,
  attempts: 3,
  backoff: { type: "exponential", delay: 5000 },
});

BullMQ refuses to add a job whose id is still in the queue. So publishing the same event twice is harmless while the first copy is waiting or running.

Set attempts yourself. BullMQ gives a job one attempt by default, so a single transient error would drop the event.

That only covers jobs still in Redis. Completed jobs get trimmed, so the queue forgets them. The real guarantee has to live in Postgres.

5. Dedupe and effects in one transaction

The consumer records each event id in a processed_events table. The insert and the handler's work share a transaction.

await db.transaction(async (tx) => {
  const inserted = await tx
    .insert(processedEvents)
    .values({ eventId: event.id })
    .onConflictDoNothing()
    .returning({ eventId: processedEvents.eventId });
 
  if (inserted.length === 0) return; // seen before, skip
 
  await handler(event, tx); // handlers get the transaction, they never open their own
});

That pairing is the whole guarantee. Mark first in a separate transaction, and a crash mid-handler records the event as done with nothing done. Mark afterwards, and every redelivery repeats the effect.

Handlers must do their database work on the transaction they are given. A handler that opens its own connection steps outside the guarantee.

With that in place, the gap from section 3 is easy to close. A recovery pass re-publishes rows that were stamped a while ago and have no processed row. If the original job is still queued, BullMQ rejects the duplicate. If it already ran, the consumer skips it.

6. Decisions that kept it calm

  • Two queues, split by guarantee. events carries outbox rows, where every job id is an outbox id. maintenance carries scheduled sweeps with ids BullMQ generates. They never mix, so a sweep can never hit the event dedupe.
  • Emails are not jobs. The handler writes a pending delivery row in the event's transaction, and a maintenance job sends it. A rolled-back event never emails, and no email text sits in Redis.
  • Unknown event types are marked processed and logged. They must not block the queue or be re-published forever.
  • Malformed jobs are dropped with an error log. A retry cannot fix a bad shape.
  • Failed jobs are kept. removeOnComplete: 1000, removeOnFail: 5000. The failed set is the thing I alert on.

The relay and the consumer each expose a single-pass method, relayOnce() and process(job). Integration tests call those directly against a real Postgres and Redis instead of waiting on timers.

What I'd do again

All of it, in this order: outbox row in the transaction, stamp then publish, job id from the outbox id, dedupe inside the handler's transaction, a recovery pass for the gap.

It is more code than queue.add(). In exchange, "did that email send twice?" has an answer I can read from two tables.

If you are building something where a job touches money or a customer's inbox, tell me about it. The live product is in the project viewer.

Share Email

See it live

Lovlov.build8.dev

Speech-therapy assessment platform, France

Building something like this?

I reply within 24 hours.

Tell me about your project
NextStripe webhooks in NestJS: raw body, signature, safe retries

© 2026 Seif Ashraf · Cairo · +20 100 700 4828

GitHubLinkedInInstagramFacebookblog· no cookies
Let's talk