Back to cases

One pipeline, many file formats: Claude-assisted schema resolution at ingestion

EffectiveSoft built an AI-assisted pipeline that sorts it out on the way from Amazon S3 to Amazon Redshift.

enterprise ai data ingestion
enterprise ai data ingestion

    Fifteen years of data from over 100 sources, and every one of them structures files differently. EffectiveSoft built an AI-assisted pipeline that sorts it out on the way from Amazon S3 to Amazon Redshift. Rule-based matching handles what it can, Anthropic’s Claude resolves the column names it can’t, and a final compatibility check against the live warehouse schema gates every load. Anything still ambiguous becomes a one-reply Slack approval that teaches the system for next time. Manual mapping effort fell 60%, onboarding went from weeks to days, and throughput scales 3.5x horizontally.

    • Client

    • Country

    • Service

    Client context

    EffectiveSoft client, Webbula, is a data quality and technology company founded in 2009. It helps brands and marketers clean, verify, and enrich customer data. The company focuses on identifying email-based threats and other risky or low-quality signals and providing curated datasets to enable marketing campaigns targeted at audiences based on real consumer data.

    The company is undergoing a broad modernization initiative to consolidate multiple mature products into a unified platform while simultaneously enhancing the architecture of their core data platform.

    EffectiveSoft serves as the delivery partner driving both initiatives.

    Challenge

    Over 15 years of operation, Webbula had accumulated terabytes of valuable data from more than 100 external sources. These data sources are frequently updated and refreshed, requiring regular processing and maintenance. The datasets are stored within the company’s AWS ecosystem, with incoming files arriving in Amazon S3. Each source follows its own data model: some files include clearly defined headers, while others contain no metadata at all. Many also use different naming conventions or structures for the same attributes. This forces human intervention to properly process the data to update their data models.

    To optimize the process and keep pace with growing data volumes and an increasing number of sources, and to provide downstream data processing a consistent foundation, Webbula needed to modernize how incoming data is standardized at scale.

    Fully deterministic ingestion pipelines have a structural weakness: they only work for inputs they were explicitly programmed to handle. In Webbula’s environment:

    • Provider files arrived as CSV, TSV, pipe-delimited, Excel, and inside GZIP, ZIP, RAR, 7z, tar, bzip2, and standard containers—some far too large to buffer in memory.
    • The same logical column appeared under dozens of raw names (e.g., email_addr, EmailAddress, e-mail), and some files carried no header row at all.
    • When a deterministic pipeline met a file it didn’t recognize, it either failed the whole load or silently skipped the file, leaving it for manual review, mapping, and re-ingestion by a human engineer.
    • Every layout change meant hand-writing new schema mappings—engineering work on the critical path of data delivery.

    The goal of the engagement was to remove the human bottleneck: reduce manual schema-mapping work to the small set of genuinely ambiguous cases, and make everything else flow through automatically.

    Solution

    AI-powered data ingestion pipeline architecture
    AI-powered data ingestion pipeline architecture
    AI-powered data ingestion pipeline architecture

    Business impact

    The intelligent ingestion layer transformed one of the most labor-intensive stages of Webbula’s data pipeline. By combining deterministic schema resolution with Claude-assisted mapping, normalization, and classification for novel or ambiguous fields, we significantly reduced the manual effort required to onboard new data sources and prepare datasets for analytics. Files with known or inferable schemas flow from S3 to Redshift with zero human involvement; only genuinely ambiguous files reach a person as a one-reply Slack approval.

    Different file formats through one entry point. The platform can now ingest a wider variety of file formats without requiring custom mapping rules for each source. This accelerates the onboarding of new datasets while reducing operational overhead and improving scalability. Seven compression/archive families and four tabular formats are handled by a single pipeline invocation against an S3 prefix, replacing per-format handling logic.

    Auditable by construction. Deterministic routing, confidence thresholds in config, per-run result manifests in S3, and source-file lineage on every loaded row make the pipeline’s behavior inspectable end to end.

    AI-powered workflow automation

    See what we offer

    Key outcomes

    • Reduced manual data preparation and schema-mapping effort by 60%.
    • Accelerated onboarding of new data sources and file formats from weeks to days.
    • Eliminated a recurring bottleneck in the Amazon S3-to-Amazon Redshift ingestion workflow.
    • Created a self-improving metadata repository that increases automation accuracy over time.
    • Maintained data quality through a human-in-the-loop review process for ambiguous files.
    • 3.5x faster processing through horizontal scaling. Benchmarked at four parallel workers versus one, the deadlock-free load design lets throughput scale by adding containers, not engineering effort.

    Contact us

    Our team would love to hear from you.

      Let’s connect

      Fill out the form, and we’ve got you covered.

      What happens next?

      • Our expert will follow up after reviewing your needs.
      • If required, we’ll sign an NDA to ensure privacy.
      • Our Pre-Sales Manager will send you a proposal.
      • Then, we get started on your project.

      Our locations

      Say hello to our friendly team at one of these locations.

      • San Diego, California

        4445 Eastgate Mall, Suite 200
        92121, 1-800-288-9659

      • San Francisco, California

        50 California St #1500
        94111, 1-800-288-9659

      • Pittsburgh, Pennsylvania

        One Oxford Centre, 500 Grant St Suite 2900
        15219, 1-800-288-9659

      • Durham, North Carolina

        RTP Meridian, 2530 Meridian Pkwy Suite 300
        27713, 1-800-288-9659

      • San Jose, Costa Rica

        C. 118B, Trejos Montealegre
        10203, 1-800-288-9659

      Join our newsletter

      Stay up to date with the latest news, announcements, and articles.

        Error text
        error message
        You must accept the terms and conditions to continue.
        title
        content
        View project