How Batch Ingestion Works
Batch ingestion extracts data from source systems at scheduled intervals, transforms it in a staging area, and loads the results into warehouse tables. The classic ETL (extract, transform, load) pipeline is a batch process: a scheduler triggers extraction at midnight, transformations run for two hours, and analysts see refreshed data by morning.
ELT (extract, load, transform) follows the same scheduling pattern but defers transformation to the warehouse engine. Raw data lands in staging tables, and SQL-based transforms run inside the warehouse where columnar compute handles aggregations efficiently.
Batch ingestion is simple to reason about. Each run processes a bounded dataset. If a run fails, you reprocess the same window. Idempotent pipelines (running the same batch twice produces the same result) are straightforward to build because the input set is fixed.
The trade-off is latency. Data is only as fresh as the last completed batch. An hourly batch means analysts work with data that is between zero and sixty minutes stale. For financial reporting, inventory reconciliation, and most analytical dashboards, this staleness is acceptable and often invisible.
Cost is predictable with batch pipelines. Compute spins up for the batch window, processes the data, and shuts down. You pay for defined processing time rather than continuous resource allocation.
How Real-Time Ingestion Works
Real-time ingestion (also called streaming ingestion) processes records as they arrive, typically through a message broker like Apache Kafka, Amazon Kinesis, or Google Pub/Sub. Producers write events to topics; consumers read those events and write them into warehouse tables continuously.
True row-level real-time insert into a data warehouse has historically been expensive because warehouses optimise for bulk operations, not single-row writes. Modern platforms address this with micro-batch streaming: Snowpipe (Snowflake) loads files from cloud storage as they appear, typically achieving one-to-two-minute latency. Redshift streaming ingestion from Kinesis materialises data through views with near-real-time freshness. BigQuery supports streaming inserts at row level with a per-row cost premium.
The operational complexity is higher than batch. Streaming pipelines must handle out-of-order events, late-arriving data, schema changes mid-stream, backpressure when consumers fall behind producers, and exactly-once or at-least-once delivery semantics. Monitoring must track consumer lag, partition health, and throughput continuously rather than checking batch-run success once per cycle.
Real-time ingestion suits use cases where stale data has a measurable cost: fraud detection, live dashboards for operations teams, dynamic pricing, and alerting systems. If nobody looks at the data between batch windows, streaming adds cost and complexity without delivering value.
Cost and Complexity Comparison
Batch pipelines cost less to build and operate for the same data volume. A scheduled job using a managed orchestrator (Airflow, dbt Cloud, or a cloud-native scheduler) requires minimal infrastructure. Failures are retried at the next window. On-call burden is low because issues surface during predictable processing windows.
Streaming pipelines require always-on infrastructure: message brokers, consumer applications, and monitoring dashboards. Kafka clusters or managed equivalents carry a base cost whether throughput is high or low. Consumer applications must scale with source-system event rates, which can spike unpredictably. The engineering team needs streaming expertise (partitioning strategy, offset management, serialisation formats) on top of warehouse skills.
A useful cost heuristic: if your batch window can run in under an hour and stakeholders do not need sub-hour freshness, batch ingestion will cost a fraction of an equivalent streaming pipeline. If your data sources already emit events to a broker (common in microservice architectures), adding a warehouse consumer is incremental rather than greenfield.
Many teams overshoot on real-time. Before committing to streaming, ask the consuming team a direct question: what decision changes if data is 15 minutes old instead of 15 seconds old? If the answer is nothing, batch wins on cost and simplicity.
Hybrid architectures are the norm in mature data platforms. Transactional data that feeds operational dashboards streams in near-real-time. Reference data, historical backfills, and large analytical aggregations run as scheduled batches. The two patterns coexist in the same warehouse with clear ownership boundaries.
Choosing the Right Pattern for Each Workload
Map each data source and consumer to a freshness requirement before choosing an ingestion pattern. Create a simple table with four columns: source system, data volume per day, required freshness, and consuming application. This exercise usually reveals that most workloads are fine with hourly or daily batches, and only a handful need sub-minute latency.
For sources that need real-time ingestion, evaluate whether micro-batch (one-to-five-minute latency) satisfies the requirement. Micro-batch streaming (Snowpipe, Fivetran streaming, Airbyte incremental) is significantly simpler than true event-at-a-time processing and covers the majority of near-real-time use cases.
Reserve true event-stream processing (Apache Flink, Spark Structured Streaming, or ksqlDB) for workloads that genuinely need sub-second latency or complex event processing (windowed aggregations, sessionisation, pattern detection) before data reaches the warehouse.
Consider the downstream consumption pattern as well. If a dashboard refreshes every five minutes, there is no benefit to ingesting data every five seconds. Align ingestion frequency with the slowest link in the consumption chain to avoid paying for freshness that nobody uses.
This content is general information about data ingestion patterns and does not constitute professional or financial advice.
Operational Considerations and Failure Handling
Batch pipelines fail discretely. A nightly job either succeeds or fails, and the failure is visible in the orchestrator's run history. Recovery means fixing the issue and rerunning the batch for the affected window. Data consistency is straightforward: either the batch loaded or it did not.
Streaming pipelines fail continuously. A consumer might fall behind (consumer lag), skip messages (offset mismanagement), or process duplicates (at-least-once delivery without deduplication). Detecting these failures requires real-time monitoring of consumer lag metrics, dead-letter queues for poison messages, and reconciliation jobs that compare source counts to warehouse counts.
Schema evolution is harder in streaming. A batch pipeline can detect schema changes at extraction time and halt the run for review. A streaming consumer processing events in real time must handle old-schema and new-schema events simultaneously during the transition window. Schema registries (Confluent Schema Registry, AWS Glue Schema Registry) help by enforcing compatibility rules before producers publish new schemas.
Backfill and replay are essential capabilities for both patterns. Batch pipelines backfill by rerunning historical windows. Streaming pipelines replay by resetting consumer offsets to an earlier position in the message log, which requires that the broker retains messages long enough for replay (Kafka's retention policy controls this). Design your retention and replay strategy before going to production; retrofitting it after a data-loss incident is stressful and error-prone.
This content is general information about data ingestion patterns and does not constitute professional or financial advice.