AI Data Preparation Platform
← All projects

AI / Data · Belgium · 2025

AI Data Preparation Platform

A collaborative annotation and data cleaning platform that accelerated an AI research team's model training pipeline by 3×.

SectorAI / Data
Client locationBelgium
Duration5 months
Year2025
Overview

A Belgian AI research and consulting firm was spending more time cleaning and labelling data than training models. Their annotation workflow involved emailing CSV files between team members, manual deduplication in Excel, and no version control over dataset iterations. AstraWorks built a purpose-built data preparation platform that brought the entire pipeline — ingestion, cleaning, annotation, versioning and export — into a single collaborative workspace.

The challenge

Data annotation tools on the market were either too generic (built for image labelling), too expensive for a boutique firm, or couldn't handle the unstructured tabular and text data the team worked with. The client also needed the platform to integrate with their existing Python-based model training pipelines via an API — not just export CSV files.

Our solution

We designed a platform with four core modules: (1) a data ingestion pipeline supporting CSV, JSON, Parquet and database connections; (2) a cleaning workspace with configurable rule-based transformations and outlier detection; (3) a collaborative annotation interface for text classification, entity labelling, and structured data tagging with multi-reviewer support and consensus scoring; (4) a dataset versioning system with an API endpoint that training pipelines can query directly to pull the latest approved dataset version.

Key features

  • Multi-format data ingestion (CSV, JSON, Parquet, SQL)
  • Configurable cleaning rules and outlier detection
  • Text classification and entity annotation interface
  • Multi-reviewer annotation with consensus scoring
  • Dataset versioning and change history
  • REST API for direct pipeline integration
  • Python SDK for programmatic access
  • Audit trail for compliance and reproducibility

Technology stack

React (annotation UI)FastAPIPostgreSQLCelery + Redis (pipeline jobs)MinIO (data storage)Python SDKDocker / Kubernetes

// outcome

What changed after launch

The team's average time from raw data ingestion to training-ready dataset dropped from 3 weeks to under 5 days. Inter-annotator agreement improved measurably with the consensus scoring system. The platform is now used across 4 active research projects simultaneously.

Have a similar challenge?

We'd like to hear about it. No pitch — just a conversation about what you're building.

Start a conversation →