See where CDP is headed with AI — Agentic World 2026, Oct 5–7, Miami →
日本語で読む
Glossary

Data Ingestion

Data ingestion connects multiple sources and transports raw data into a central repository for analysis. Learn batch vs. real-time methods and CDP benefits.

CDP.com Staff CDP.com Staff 13 min read

Data ingestion is the process of connecting to multiple data sources and transporting the data from each source into a single repository, typically a database, data warehouse, or data lake. Once the data is in the central repository, it can be accessed and analyzed by anyone in the organization with access rights. Data ingestion can occur in batches on a schedule, or it can occur in real-time with a steady flow of data from the source system into the central repository.

Although data ingestion is often used interchangeably with data integration, the two are not the same. Data ingestion imports the data in the new repository in its raw form. With data integration, the data is transformed as part of the process of moving it from the source system through an ETL (Extract, Transform, Load) process. In addition, in some architectures, integrating data means the data stays in the source systems but is accessible through a centralized application, like a search engine.

Benefits of Data Ingestion

The most significant benefit of data ingestion is speed: data reaches a central repository quickly because no transformation is required during the move. Once ingested, data can be cleaned for consistency and correctness, then go through transformation processes as part of a broader data pipeline.

Centralizing data is also essential for analytics systems that derive common themes and insights across the full dataset.

Data Ingestion as a Core CDP Capability

Data ingestion is the foundational layer of every customer data platform (CDP). A CDP must ingest data from dozens or hundreds of source systems, including marketing automation, CRM, ERP, web analytics, e-commerce, mobile apps, social media, point-of-sale, and IoT devices. This breadth of ingestion is what differentiates CDPs from single-purpose tools.

CDPs must handle both batch and streaming ingestion simultaneously. Batch ingestion processes historical data loads on a schedule, while real-time data processing captures events as they occur, enabling immediate personalization and triggered messaging. Agentic CDPs depend on streaming ingestion to keep customer profiles current so AI agents can make sub-second decisions.

Once ingested, CDP data goes through identity resolution to match records across sources, deduplication to merge profiles, and cleansing to resolve discrepancies. The unified data is then available to analytics engines, machine learning models, and data activation pipelines that deliver audiences to external systems for campaigns and programs.

Schema flexibility is another critical requirement. CDPs must accept structured data (transaction records, form submissions), semi-structured data (JSON event logs, API responses), and unstructured data (customer service transcripts, social media posts) without requiring rigid schemas to be defined in advance.

Challenges with Data Ingestion

Ensuring that data ingestion is performed securely is critical, especially with customer data and proprietary information. Proper data governance policies are essential for managing both the transport and storage of ingested data. Only authorized analytics tools, systems, and people should have access to the ingested repository.

How data ingestion works, step by step

Ingestion looks like a single action from the outside, but every ingestion job runs the same five-step cycle, and each step has a characteristic failure that only shows up downstream.

1. Connect. The job authenticates to the source with an API key, OAuth token, service account, or read-only database credential. Credentials expire, and expired credentials are the most common silent stop: the job keeps running on its schedule while the source rejects every request. Rotating credentials on a fixed calendar and alerting on authentication failures — not just job failures — closes most of that gap.

2. Extract. The job reads data out of the source: polling an API page by page, receiving pushed events from an SDK or webhook, downloading files from object storage, or tailing a database’s change log. Extraction is where pagination and rate limits bite. A poller that assumes the API’s default page size imports the first page and quietly drops everything after it; a poller that ignores rate-limit headers gets throttled and produces a half-read that still reports success.

3. Transform in flight. Ingestion applies only light, per-record reshaping during transport: renaming fields to a common schema, casting types, normalizing timestamps to a single reference timezone, and filtering out obvious noise such as test events. Heavy reshaping — joining sources, deriving metrics, harmonizing categories — belongs to integration downstream. When teams push heavy transformation into the ingestion step, every source change breaks the transport job instead of only the modeling that depends on it.

4. Load. The job writes records into the repository with a deduplication strategy: a primary key, an upsert, or an event identifier that makes the write idempotent. Without idempotent loads, any retry — a webhook delivered twice, a nightly job re-run after a failure — double-counts records, and inflated counts are hard to spot weeks after the fact.

5. Monitor. Someone compares row counts, freshness lag, and error rates against expectations. Ingestion failures are quiet by nature: an empty table is easier to miss than a broken page, and the first visible symptom is often a dashboard showing yesterday’s numbers or a segment that stopped growing. The consumers that inherit the gap include dashboards, activation jobs, and any AI agent acting on customer records without a human reviewing the input.

One distinction runs through all five steps: pull versus push. Pulling — polling an API, reading a file drop — means the ingestion side controls timing and can batch, retry, and back off on its own schedule. Pushing — webhooks, SDK event streams — means the source controls timing, so the ingestion side must be continuously available and absorb bursts. Most CDPs run both, because sources differ, and the mix is a design decision rather than a default.

The pattern across all five steps is the same: ingestion rarely fails loudly. It fails by delivering less than expected while reporting success. That is why ingestion is treated as the first leg of a broader data pipeline — the leg whose correctness decides whether everything after it starts from honest data.

Matching the ingestion method to the source

There is no universally correct ingestion method. The source’s shape — how it exposes data, how often it changes, and how quickly decisions consume it — decides the method, and a mismatch surfaces as either stale profiles or throttled, failing syncs.

Three dimensions drive the choice. Volume: how many records the source produces per day. Volatility: how often existing records change, not just how often new ones appear. Latency: how stale a record may be when a decision uses it. A source with high volume, low volatility, and relaxed latency — a transaction archive, for example — suits batch. A source with low volume but high volatility — an active shopping cart, for example — rewards streaming. Change data capture exists for the awkward middle: high volume and high volatility, where polling either costs too much or falls behind.

Source situationIngestion methodWhat to establish firstWhy the method fitsFailure mode when mismatched
Cloud SaaS tool with a REST API (CRM, email platform, support desk)Scheduled API pollingRate limits, auth scope, pagination depth, and whether the API supports “changed since” filtersIncremental pulls move volume proportional to change instead of table sizePolling too aggressively hits rate limits and the source throttles or cuts off the sync; polling too rarely leaves profiles weeks stale
Website and app eventsStreaming push from an SDK or webhookAn event schema contract, payload size limits, and tolerance for out-of-order arrivalEvents arrive while intent is live, so the profile reflects behavior with minimal delayWebhook retries without idempotency keys turn one purchase into three and inflate conversion counts
Transactional database (orders, subscriptions)Log-based change data capturePrimary keys, the source’s log retention window, and a policy for schema changesCaptures inserts, updates, and deletes continuously without adding query load to a production systemA log retention window shorter than the sync interval forces a full reload after any outage
Files dropped by a partner or legacy systemScheduled batch file loadDelimiter and encoding consistency, stable column order, and an arrival time you can alert onThe cheapest method for high volumes no decision consumes within minutesThe partner adds a column and positional parsing shifts every field, silently corrupting loads
Historical archive or one-time migrationBulk one-time loadWhich system is the source of truth for identity fields, and a planned deduplication passVolume justifies a one-off job tuned for throughput rather than latencyRunning a bulk load through tooling sized for steady-state syncs takes days or stalls on rate limits

Two habits keep the choice honest. First, start from the decision the data feeds: if nothing consumes a source within an hour of an event, batch is the right answer regardless of how current streaming looks. Second, re-check the match when the source grows: a source that moves from thousands to millions of records turns an incremental poller into a throttled one, and the fix is a method change, not a longer schedule. In practice one source often justifies two methods — a bulk load for history, then a steady-state incremental or streaming sync from the point of cutover — and treating the two as separate jobs with separate monitoring is what keeps the cutover clean.

Who owns data ingestion?

Data ingestion sits between teams, which is exactly why it breaks. The data engineering side owns whether records arrive: connectors, schedules, credentials, and load correctness. The marketing and analytics side owns whether the right records arrive at all: which sources matter, which fields a segment or model will actually use, and how fresh the data must be for the decision it feeds. When one side assumes the other defined the contract, the failure is predictable — a source gets connected because it was easy to connect, fields sync that no one uses, and a field that segmentation depends on was never mapped at all.

The practical fix is an explicit schema contract per source: which fields are expected, their types, which field serves as the deduplication key, and what freshness commitment applies. A contract makes drift visible — when the source adds or renames a field, someone knows which downstream consumer to warn. Without one, drift surfaces as a silently empty column, usually discovered by whoever built the campaign on top of it.

Ownership also decides incident response. When a sync fails overnight, someone has to know whether that is a broken connector (data engineering), a changed source (the team that owns the relationship), or an expired credential (platform operations). Teams that write this down recover in hours; teams that improvise ownership recover in days.

Common ingestion failure modes and how to fix them

Most ingestion problems are not dramatic outages. They are quiet distortions that pass the job’s own success checks and surface weeks later as wrong numbers. The six below cover most of what goes wrong in transport; each has a fix that is cheaper to apply at ingestion time than to repair downstream.

Failure modeHow it shows upWhat to do about it
Schema drift — a source adds, renames, or removes a fieldMapped fields arrive empty; reports show gaps instead of errorsAlert on unexpected null rates per field, map against a pinned schema version, and fail the load on unrecognized structure rather than passing it through
Duplicate records from retriesCounts inflate; the same event appears two or three timesRequire idempotency keys and deduplicate on load, not during analysis
Rate limiting and throttlingSyncs slow down or fail intermittently, worse at peak hoursPull incrementally, respect the source’s throttling headers, and stagger schedules across sources
Timestamp and timezone inconsistencyEvents land on the wrong day; recency-based segments misfireNormalize to a single reference timezone during transport and keep the original offset in a separate field
Partial loads after a crashHalf a batch lands; totals look plausible because the missing rows are spread randomlyWrap loads transactionally or checkpoint progress, and alert on row-count deltas between runs
Late-arriving dataYesterday’s events appear today, after yesterday’s reports were already sentDefine a late-arrival window and a reprocessing policy instead of assuming data is final at load time

Fixing these at the transport layer matters because every downstream consumer — identity resolution, segmentation, activation — inherits the same defect and would otherwise have to detect and repair it separately, if it notices at all. Automation raises the stakes: an agentic data platform can act on ingested records without a person reviewing the gap first, so a silent ingestion defect becomes an automated decision made on incomplete data. Checks and rules belong to the validation discipline; the fixes above decide whether that discipline receives honest data in the first place.

FAQ

What is the difference between data ingestion and data integration?

Data ingestion imports raw data from source systems into a central repository without transforming it, preserving the original format and structure. Data integration, on the other hand, involves transforming and harmonizing data as part of the transfer process through ETL (Extract, Transform, Load) pipelines. Ingestion prioritizes speed and simplicity, while integration focuses on making data immediately usable by standardizing formats and resolving inconsistencies during the move.

What is the difference between batch and real-time data ingestion?

Batch data ingestion collects and transfers data in scheduled intervals—hourly, daily, or weekly—making it efficient for large volumes of data that do not require immediate processing. Real-time data ingestion streams data continuously from source systems as events occur, enabling near-instant availability for analytics and activation. The choice between batch and real-time depends on your use case: real-time is essential for personalization and time-sensitive customer interactions (see real-time CDP), while batch is sufficient for reporting and historical analysis.

What types of data sources can be ingested into a CDP?

A CDP can ingest data from virtually any customer-facing system, including CRM platforms, marketing automation tools, web and mobile analytics, e-commerce platforms, point-of-sale systems, social media, customer support software, and third-party data providers. CDPs support both structured data (like transaction records and form submissions) and semi-structured or unstructured data (like JSON event logs and customer service transcripts). This broad ingestion capability is what enables CDPs to create comprehensive, unified customer profiles.

How often should data be ingested into a CDP?

Cadence should follow how quickly decisions consume the data, not how often the source can send it. Daily batch ingestion is enough for reporting and long-cycle campaigns. Hourly or 15-minute syncs suit teams that refresh segments between sends. Streaming earns its extra cost only when a decision window is measured in minutes, such as abandoned-cart follow-up or in-session personalization. Choosing the slowest cadence that still meets the deadline keeps API load, cost, and failure surface down.

What is an ingestion connector?

An ingestion connector is a prebuilt integration that handles the mechanics of moving data from one specific source: authentication, pagination, rate limits, and field mapping. Connectors replace the custom scripts every team would otherwise write and maintain against the same APIs. The trade-off is coupling. When a source changes its API, the connector must be updated before syncs recover, so a connector’s maintenance record matters as much as its feature list.

This article is also available in: データ取り込み(データインジェスト)とは?CDPの基盤機能 · Ingestão de dados: o que é, em lote e em streaming

CDP.com Staff
Written by

The CDP.com staff has collaborated to deliver the latest information and insights on the customer data platform industry.