Skip to main content
Back to Blog

Mastering the SNS + SQS Fan-Out Pattern in Event-Driven Systems

10 min read
1 → N, one event, many queues
In this article
  1. 01Why Does the Fan-Out Pattern Matter?
  2. 02How Does SNS and SQS Fan-Out Work?
  3. 03How Does Message Filtering Work?
  4. 04How Do You Handle Failures and Retries?
  5. 05Do You Need Strict Ordering and Deduplication?
  6. 06How Do You Monitor This Setup? (The Operational Playbook)
  7. 07How Do You Scale Your Consumers?
  8. 08When Should You Actually Use This Pattern?
  9. 09When Should You Avoid It?
  10. 10Key Takeaways

When I need asynchronous processing, the first tool I reach for is usually a message queue. But what happens when a single event needs to trigger several independent workflows? Standard queues are strictly point-to-point. In these cases, you actually need a fan-out approach. If you are building on AWS (Amazon Web Services), the SNS (Simple Notification Service) + SQS (Simple Queue Service) pattern is the absolute standard solution.

Why Does the Fan-Out Pattern Matter?

Consider a standard e-commerce order placement. When a customer completes a checkout, you usually need to do a few things:

  • Send a confirmation email to the customer
  • Update the inventory
  • Trigger fraud detection
  • Push analytics events
  • Notify the warehouse If each of these tasks belongs to a separate service, a single queue just won’t cut it. You would need your publishing service to know about every single consumer, and it would have to send messages to each queue individually. This creates tight coupling, which is exactly what event-driven architecture is meant to eliminate. Instead, fan-out completely decouples the publisher from the consumers. The publisher simply sends one message to a single topic. From there, the topic distributes copies to all of its subscribers. Then, every subscriber processes its copy independently. If you ever need to add a new consumer, you just create a new SQS queue and subscribe it to the topic. The publisher doesn’t have to change at all.

How Does SNS and SQS Fan-Out Work?

The underlying pattern is actually pretty straightforward:

  • SNS Topic: This receives the initial event directly from the publisher.
  • SQS Queues: Each subscriber has its own dedicated queue, which is subscribed to the topic.
  • Consumers: Every service polls its own queue independently.
{
  "TopicArn": "arn:aws:sns:us-east-1:123456789:order-placed",
  "Subscriptions": [
    { "Protocol": "sqs", "Endpoint": "arn:aws:sqs:us-east-1:123456789:email-queue" },
    { "Protocol": "sqs", "Endpoint": "arn:aws:sqs:us-east-1:123456789:inventory-queue" },
    { "Protocol": "sqs", "Endpoint": "arn:aws:sqs:us-east-1:123456789:analytics-queue" }
  ]
}
JSON

In this setup, each queue receives its very own copy of the original message. If the email service becomes slow or goes down completely, this does not affect the inventory updates at all. Every consumer scales independently of the others. Also, you can use dead letter queues to catch failures on a strict per-consumer basis.

How Should You Structure Your Messages?

You should always design your SNS messages with a clear, stable schema. Make sure to include enough context for all consumers, so they don’t have to make callback queries back to the publisher. Here’s an example.

{
  "eventType": "ORDER_PLACED",
  "version": "1.0",
  "timestamp": "2026-03-05T14:30:00Z",
  "correlationId": "abc-123-def-456",
  "payload": {
    "orderId": "ord-789",
    "customerId": "cust-456",
    "items": [{"sku": "SKU-001", "quantity": 2, "price": 29.99}],
    "totalAmount": 59.98,
    "currency": "USD"
  }
}
JSON

Notice the correlationId field. This ID is essential for tracing a single business event across all of your consumers. When the email service processes this message, it then logs the correlation ID. And when the inventory service processes its own copy, it logs the exact same ID. Later, during an incident, you can query your log aggregator using that specific correlation ID. This gives you a clear view of the full processing chain across all of your services.

How Does Message Filtering Work?

Not every subscriber needs to receive every single message. Luckily, SNS supports filter policies that let you route messages selectively.

{
  "FilterPolicy": {
    "order_type": ["premium", "enterprise"],
    "region": ["us-east-1"]
  }
}
JSON

Using a policy like this keeps your queues lean. It also prevents you from wasting computing resources on messages that a specific service simply does not care about. SNS evaluates these filter policies before delivery, and because of this, filtered-out messages never reach the queue, which means you don’t incur any SQS costs for them. SNS supports two different filter policy scopes. The default is MessageAttributes, which filters based on metadata. The alternative is MessageBody, which filters directly on the payload content itself.

Body-based filtering gives you more flexibility, but it is slightly more expensive to evaluate. For most of my use cases, I find that message attributes are perfectly sufficient and actually more performant.

How Do You Handle Failures and Retries?

One of the biggest benefits of this pattern is that each SQS queue provides independent retry mechanisms and failure isolation. Here is how I configure these settings:

  • Visibility timeout: This controls how long a message is hidden after a consumer picks it up. You should set this to at least 6x your expected processing time. This gives the system enough buffer to handle retries gracefully.
  • Dead letter queues (DLQ): These catch messages that fail after a certain number of retries. You must always configure a DLQ because without one, poison messages will cycle through your queue forever.
  • Redrive policies: These policies allow you to replay DLQ messages once you have fixed the underlying bugs. Because of this setup, a bug in your analytics consumer will never block your order confirmations. Each failure domain will be completely isolated.

What Is the Best Way to Handle Poison Messages?

A poison message is a message that will never process successfully, no matter how many times you retry it. It might have an unexpected schema, reference a missing resource, or trigger a specific bug in your consumer. Without a DLQ, this poison message will either block the queue entirely or waste compute power by retrying indefinitely. I recommend configuring your redrive policy to move messages to the DLQ after 3 to 5 failed attempts. Once they are there, you need to monitor that DLQ closely. Here’s how:

  • Set up alerts for when the DLQ depth exceeds zero. Any message in a DLQ means that something is actively wrong.
  • Build a dashboard that shows the DLQ depth per queue over time. A growing DLQ points to a systemic issue, rather than just a transient failure.
  • After you fix the bug, use the SQS DLQ redrive feature. This allows you to safely replay the messages back to their source queue.

Do You Need Strict Ordering and Deduplication?

Standard SNS and SQS provide at-least-once delivery with best-effort ordering. If you need strict ordering, you have to use SNS FIFO (First-In-First-Out) with SQS FIFO queues. Going the FIFO route guarantees two things:

  • Messages within a specific message group are processed in exact order.
  • You get exactly-once processing (as long as it falls within the 5-minute deduplication window). The main tradeoff here is throughput. FIFO topics only support 300 publishes per second, or up to 3,000 if you use batching. In practice, I find that most fan-out use cases simply do not need strict ordering. For instance, the email service does not care if it processes Order A before Order B. The analytics service doesn’t care either. If you do need ordering for a specific consumer, like inventory updates that must be ordered by SKU, you can use FIFO for that specific flow and standard queues for the rest. To achieve this, you will need to use separate topics for your ordered and unordered flows.

Why Must Your Consumers Be Idempotent?

Because standard SQS uses at-least-once delivery, every single consumer must be able to handle duplicate messages. SQS can deliver the exact same message more than once, and SNS can deliver that message to the same queue multiple times. To handle this, you must design your consumers to be idempotent. Here is my standard approach:

  • Use the MessageId or a specific business key as a deduplication key.
  • Always check if the message has already been processed before executing the business logic.
  • Use database-level constraints, such as unique indexes or conditional updates, as a final safety net.

How Do You Monitor This Setup? (The Operational Playbook)

To keep everything running smoothly, I track a few specific metrics for every single queue:

  • ApproximateNumberOfMessagesVisible: This shows the number of messages waiting to be processed. If this metric is growing, it means your consumers are not keeping up.
  • ApproximateAgeOfOldestMessage: This tells you how long the oldest message has been waiting. I always set an alarm if this exceeds my SLA, such as 60 seconds for order confirmations.
  • NumberOfMessagesSent / NumberOfMessagesReceived: This measures your overall throughput. A sudden drop in sent messages usually indicates an issue on the publisher side.
  • DLQ depth: As mentioned earlier, any non-zero value here requires immediate investigation.

How Do You Scale Your Consumers?

SQS consumers scale horizontally by design. Using AWS Lambda is the simplest option here. You just need to configure an event source mapping with a specific batch size and concurrency limit. If you are running ECS or EKS workloads, you should use SQS-based autoscaling. The most reliable way to scale these is based on this: ApproximateNumberOfMessagesVisible / numberOfConsumers

When Should You Actually Use This Pattern?

I highly recommend reaching for the SNS and SQS fan-out pattern when:

  • Multiple independent consumers need to react to the exact same event.
  • You want strict, loose coupling between your publisher and your subscribers.
  • Your consumers have completely different scaling, latency, or reliability requirements.
  • You need per-consumer failure isolation and custom retry policies.

When Should You Avoid It?

Despite how powerful it is, this pattern is not always the right choice. You should look for alternatives in these specific scenarios:

  • Single consumer: If you only have one consumer, just use SQS directly. There is absolutely no need to add the SNS layer.
  • Complex routing logic: If you need content-based routing with rich rule expressions, you should consider AWS EventBridge instead.
  • Cross-account or cross-region setups: While this is possible with SNS and SQS, it adds significant operational complexity with IAM policies and VPC configuration.
  • Very high throughput with strict ordering: Remember that FIFO topics have strict throughput limits. If you need high-throughput ordered streams, you should consider a tool like Kafka instead.

Key Takeaways

To wrap things up, here are the most important points to remember when building event-driven systems with this pattern:

  • The SNS and SQS fan-out pattern perfectly decouples event publishers from consumers.
  • Every consumer gets its own independent scaling, retries, and failure isolation.
  • You must include correlation IDs in your messages to guarantee end-to-end traceability across all consumers.
  • Use filter policies to prevent unnecessary message processing and save compute costs.
  • Design every single consumer to be idempotent, because at-least-once delivery is the default behavior.
  • Always configure DLQs for every queue and set up alerts for any non-zero depth.
  • Only use FIFO variants when ordering truly matters, and always keep their throughput limits in mind.
  • Keep your messages small (under 256KB), or use S3 for large payloads alongside the extended client library.

Related Articles

Browse All Articles