← Back to Questions
Microservices - Scenario based questions

Kafka broker goes down during message processing. What happens and how will you recover?

Learn Kafka broker goes down during message processing. What happens and how will you recover? with simple explanations, real-time examples, interview tips and practical use cases.

Kafka Broker Goes Down During Message Processing. What Happens And How Will You Recover?

In enterprise event-driven microservices architecture, broker failures are common distributed system problems. A Kafka broker may go down because of hardware crashes, network failures, disk corruption, cloud infrastructure issues, Kubernetes pod failures, or maintenance operations. Apache Kafka is designed as a distributed fault-tolerant system, so broker failures do not necessarily cause complete system outages if the cluster is configured correctly.


Main Goal

Prevent Data Loss
Maintain Availability
And Recover Automatically
During Broker Failures

What Is A Kafka Broker?

A broker is a Kafka server responsible for:

  • Storing messages
  • Serving producers
  • Serving consumers
  • Managing partitions

Kafka Cluster Example

Broker-1
Broker-2
Broker-3

Topics Are Split Into Partitions

Order Topic

Partition-0
Partition-1
Partition-2

Replication Example

Partition-0 Leader → Broker-1
Replica → Broker-2
Replica → Broker-3

Why Replication Is Important?

Replication prevents message loss during broker failures.


What Happens When Broker Goes Down?

Behavior depends on:

  • Replication factor
  • Leader election
  • Producer configuration
  • Consumer configuration

Scenario 1: Broker Without Replication

Architecture

Partition-0 Stored Only In Broker-1

Broker-1 Crashes

Broker-1 DOWN

Result

  • Partition unavailable
  • Messages inaccessible
  • Possible permanent data loss
  • Consumers stop reading
  • Producers fail

Production Verdict

Never Use Single Replica
In Production

Scenario 2: Broker With Replication

Architecture

Partition-0 Leader → Broker-1
Replica → Broker-2
Replica → Broker-3

Broker-1 Crashes

Broker-1 DOWN

Kafka Recovery Process

Replica Promotion Happens
Automatically

New Leader Election

Broker-2 Becomes New Leader

Result

  • Cluster continues working
  • Minimal downtime
  • No data loss
  • Consumers reconnect automatically

Production Principle

Replication
Provides Fault Tolerance

1. Leader Election

Kafka partitions always have one leader.


Example

Partition-0 Leader → Broker-1
Followers → Broker-2, Broker-3

If Leader Fails

Follower Replica
Becomes New Leader

Benefits

  • Automatic recovery
  • High availability

2. In-Sync Replicas (ISR)

Kafka maintains ISR replicas.


Definition

Replicas Fully Synced
With Leader

Example

ISR = Broker-1, Broker-2

If Leader Crashes

Only ISR replicas can become leaders.


Benefits

  • Prevent stale data promotion
  • Improve consistency

3. Producer Behavior During Broker Failure

Producers may experience temporary failures.


Possible Errors

  • Leader not available
  • Network exception
  • Timeout exception

Recovery Strategy

  • Automatic retries
  • Reconnect to new leader
  • Metadata refresh

Important Producer Configuration

acks=all
retries=10
enable.idempotence=true

Why acks=all?

Producer waits until all replicas acknowledge.


Benefits

  • Prevent data loss
  • Strong durability

4. Consumer Behavior During Broker Failure

Consumers may temporarily stop consuming.


What Happens?

Leader Changes
Consumer Rebalances
Consumption Resumes

Possible Temporary Effects

  • Small delays
  • Consumer rebalance
  • Offset reassignments

Benefits Of Consumer Groups

  • Automatic recovery
  • Load balancing

5. Replication Factor

Replication factor determines fault tolerance.


Production Recommendation

Replication Factor = 3

Meaning

1 Leader
2 Replicas

Benefits

  • Tolerate broker failures
  • Improve availability

6. min.insync.replicas

Critical production configuration.


Example

min.insync.replicas=2

Meaning

At least 2 replicas must acknowledge writes.


Benefits

  • Prevent data loss
  • Improve consistency

7. Idempotent Producers

Broker failures may cause retries.


Risk

Duplicate Messages

Solution

enable.idempotence=true

Benefits

  • Prevent duplicates
  • Exactly-once producer guarantees

8. Kafka Transactions

Distributed processing may require transactions.


Example

Consume Message
      ↓
Update Database
      ↓
Publish New Event

Problem

Broker crashes during processing.


Solution

Kafka Transactions

Configuration

transactional.id=payment-txn

Benefits

  • Exactly-once processing
  • Atomic operations

9. Retry Mechanisms

Broker failures require retry handling.


Example

Producer Send Fails
      ↓
Retry Automatically

Important

Retries must use backoff.


Correct Configuration

retry.backoff.ms=1000

Benefits

  • Avoid retry storms
  • Improve cluster stability

10. Dead Letter Queue Handling

Some messages may still fail permanently.


Example

Message Retry Failed 5 Times
      ↓
Move To DLQ

Benefits

  • Prevent message loss
  • Enable manual recovery

11. Monitoring During Broker Failure

Monitoring is critical.


Monitor

  • Broker health
  • ISR shrinkage
  • Consumer lag
  • Partition availability
  • Under replicated partitions

Popular Monitoring Stack

  • :contentReference[oaicite:0]{index=0}
  • :contentReference[oaicite:1]{index=1}

Benefits

  • Early failure detection
  • Faster recovery

12. Kubernetes Recovery

Kafka commonly runs on Kubernetes.


Platform

  • :contentReference[oaicite:2]{index=2}

If Broker Pod Crashes

Kubernetes Restarts Pod
Automatically

Benefits

  • Self-healing infrastructure
  • Automatic recovery

13. Disaster Recovery

Multi-data-center replication improves resiliency.


Example

Primary Kafka Cluster
        ↓
Replica Cluster

Popular Tool

  • :contentReference[oaicite:3]{index=3}

Benefits

  • Regional disaster recovery
  • Business continuity

14. Banking Example

Digital Banking System

Microservices:

  • Payment Service
  • Fraud Detection Service
  • Notification Service
  • Transaction Ledger Service

Flow

Payment Initiated
      ↓
Kafka Event
      ↓
Fraud Validation
      ↓
Ledger Update

Problem

Broker-1 crashes during payment processing.


Kafka Recovery

Broker-2 Promoted As Leader

Producer Behavior

Retries Triggered
Metadata Refreshed

Consumer Behavior

Consumer Rebalance Happens
Consumption Resumes

Result

  • No payment loss
  • No duplicate processing
  • Minimal downtime
  • Automatic recovery

15. Common Problems

Problem Cause
Data Loss No replication
Duplicate Messages Retries
Consumer Lag Leader election delays
Partition Unavailability Insufficient ISR replicas
Retry Storms Aggressive retries

Solutions

Problem Solution
Data Loss Replication factor
Duplicates Idempotent producers
Availability Leader election
Recovery Automatic retries
Disaster Recovery Cross-cluster replication

16. Production Best Practices

  • Use replication factor 3
  • Enable acks=all
  • Configure min.insync.replicas properly
  • Enable idempotent producers
  • Use Kafka transactions
  • Monitor ISR shrinkage
  • Implement DLQ handling
  • Use retry backoff
  • Deploy brokers across multiple nodes
  • Use Kubernetes self-healing

17. Exactly Once Flow

Producer Transaction
        ↓
Kafka Topic
        ↓
Consumer Transaction
        ↓
Database Commit

Benefits

  • No duplicates
  • No data loss
  • Reliable distributed processing

Final Interview Answer

When a Kafka broker goes down during message processing, the behavior depends mainly on Kafka replication configuration and fault-tolerance settings. In enterprise production environments, Kafka topics are configured with multiple replicas, typically a replication factor of 3, where one broker acts as leader and others act as follower replicas. If the leader broker fails, Kafka automatically performs leader election and promotes one of the in-sync replicas as the new leader, ensuring high availability and minimal downtime. Producers may temporarily receive exceptions such as leader not available or timeout errors, but they automatically recover using retries and metadata refresh mechanisms. Consumers may experience short pauses during consumer group rebalancing but automatically reconnect to the new leader and resume consumption. To prevent data loss, enterprises configure producers with settings like acks=all, enable.idempotence=true, and proper retry policies with exponential backoff. Critical systems also use Kafka transactions for exactly-once processing semantics. Monitoring tools such as :contentReference[oaicite:4]{index=4} and :contentReference[oaicite:5]{index=5} are used to monitor broker health, ISR shrinkage, and consumer lag. In cloud-native environments, :contentReference[oaicite:6]{index=6} automatically restarts failed Kafka broker pods, improving resiliency and recovery speed. For disaster recovery, enterprises implement cross-cluster replication using :contentReference[oaicite:7]{index=7}. Overall, proper replication, leader election, retries, idempotency, transactions, and observability ensure that Kafka clusters recover automatically from broker failures with minimal downtime and without losing business-critical messages.

Why this Microservices - Scenario based questions question is important?

This interview question helps candidates understand real-time backend development concepts, practical problem solving, coding fundamentals, system design basics and production-ready application behavior.

Practice this question carefully for Java backend roles, Spring Boot developer interviews, microservices interviews, company interviews and full-stack developer preparation.

About the Author

Naresh Kumar is a Senior Java Backend Engineer with experience building enterprise applications using Java, Spring Boot, Microservices, Docker, Kubernetes and Cloud technologies.