vs

In the Apache Kafka distributed stream processing architecture, Consumer Groups serve as the core unit for coordinating consumption logic. When a consumer retrieves and successfully processes a message, it must inform the Broker that the message has been consumed so that subsequent reads can resume from the same position. This process is achieved through "Offsets." Kafka provides two primary offset submission strategies: manual submission and automatic submission. Understanding the differences, applicable scenarios, and potential risks between the two is crucial for building high-availability, low-latency data consumption systems.
Manual submission involves the consumer explicitly invoking commitSync() (synchronous) or commitAsync() (asynchronous) interfaces to update offsets after completing business logic processing. In this mode, the Broker does not actively intervene in the consumer's submission behavior; the timing and frequency of submissions are entirely controlled by the application.

The core advantage of this mechanism lies in accuracy. Since the submission action follows immediately after business logic, developers can ensure that offsets are only updated to Kafka after data has been truly processed, verified, and even written to downstream systems. If an exception occurs during processing (such as data format errors or unsatisfied business rules), the consumer can choose to rollback the offset or mark the message as failed, thereby avoiding duplicate consumption or loss of critical data.

Typical application scenarios for manual submission include:

  • Scenarios requiring strict At-Least-Once delivery guarantees: For example, financial transaction records or order processing, where any data loss is unacceptable.
  • Complex data cleaning and transformation workflows: When messages require multi-step logical validation, offsets must be submitted only after every step succeeds.
  • Scenarios requiring consumption rate adjustment based on business state: Dynamically controlling the consumption speed to match downstream processing capabilities.

Mechanism and Limitations of Automatic Offset Submission

In contrast, automatic submission is actively managed by the Kafka Broker. Consumers simply need to periodically invoke the poll() method to fetch messages; the Broker will persistently update the current fetched but unsubmitted offsets based on the configured auto.commit.interval.ms (defaulting typically to 5 seconds). Additionally, in newer versions of the Kafka client, enable.auto.commit can be set to false in conjunction with manual submission, or enabled to simplify development.

The primary characteristic of automatic submission is simplicity. It reduces code implementation complexity by eliminating the need to write submission logic after every block of business logic, making it ideal for rapid prototyping or scenarios where consumption logic is extremely simple and data accuracy requirements are low (such as simple log collection).

However, the automatic submission mechanism carries significant risks:

  1. Risk of Duplicate Consumption: If the consumer crashes or shuts down after fetching a message but before submitting the offset, the Broker will resume reading from the last successfully submitted offset upon the next startup. This means the same batch of messages may be fetched and processed again. While Kafka itself supports idempotent design to mitigate this issue, it increases system complexity and resource consumption.
  2. Processing Latency: Due to the fixed submission frequency (e.g., every 5 seconds), if business processing takes a long time, the actual consumption progress may lag behind the Broker's perceived progress.
  3. Inability to Rollback: Once automatic submission occurs, it cannot be easily rolled back to an earlier position even if data errors are discovered later, unless manual intervention is performed.

Key Configuration Parameter Comparison and Best Practices

In actual deployment, developers must weigh and select strategies based on business requirements and correctly configure relevant parameters. Below are the key configuration differences between the two modes:

  • Automatic Submission Configuration:

    • auto.commit.interval.ms: Controls how often the Broker checks for unsubmitted offsets (default 5000ms).
    • enable.auto.commit: A boolean value controlling whether automatic submission is enabled (default true).
    • max.poll.interval.ms: Defines the maximum time window for a single message fetch. If processing time exceeds this value, the Broker assumes the consumer has crashed and reassigns partitions.
  • Manual Submission Configuration:

    • Typically requires no specific Broker-side automatic submission parameters, but attention must be paid to setting the client's enable.auto.commit to false.
    • Relies heavily on commitSync() or commitAsync() calls within the code.
    • If using asynchronous submission, note that max.poll.records should not be set too high to avoid excessively long processing times per batch affecting overall throughput.

Best Practice Recommendations:
For core business data streams in production environments, it is strongly recommended to use manual submission. Although the code volume is slightly higher, it maximizes the guarantee of eventual data consistency. During implementation, offset submission should be encapsulated within transactions or ensure rollback logic is executed within exception handling blocks. For example, when processing orders, update the database status first; only after confirming success should commitSync() be called.

For non-critical data streams such as log analysis or monitoring metric collection, if the system possesses idempotent consumption capabilities, automatic submission may be considered to simplify architecture, provided that duplicate message rates are monitored closely and max.poll.interval.ms is set reasonably to prevent consumers from being mistakenly judged as crashed, triggering rebalancing.

Summary and Outlook

Manual and automatic submission are two complementary mechanisms within Kafka consumers. Manual submission sacrifices some development convenience for extremely high data reliability, suitable for scenarios with strict accuracy requirements; automatic submission trades simplified code for the risk of duplicate consumption, suitable for fault-tolerant log-based scenarios. As the Kafka ecosystem evolves, future client libraries may offer finer-grained control options, such as transaction-based offset submission or smarter rebalancing strategies. However, the core logic—that clearly determines who decides when to mark a message as "consumed"—will remain the key factor in building reliable stream processing systems.