Product vision
OneOpen ML Studio is an open-source, local-first and self-hosted platform for building training-ready machine-learning datasets and managing the full data → model lifecycle.
Definition
A collaborative, framework-neutral ML data platform that converts raw data into validated, versioned and reproducible training-ready datasets — then trains, compares and registers models.
Who it is for
Individuals and teams who need to:
Import raw data
Explore and profile datasets
Annotate multiple modalities
Review and approve annotations
Clean and transform data
Detect quality problems
Version datasets
Generate reproducible training splits
Train models
Compare experiments
Register trained models
Run active-learning cycles
Collaborate with multi-user workflows
Deploy locally, on private servers, or on private cloud
Platform scope
Annotation
+ Data preparation
+ Quality validation
+ Dataset versioning
+ Training readiness
+ Experiment tracking
+ Model training
+ Active learning
+ Collaboration
+ Local / team / enterprise deployment
Modalities
Computer vision, text, tabular, audio, video, and LLM fine-tune datasets — each with a task-specific label UI. In-process training is deepest for CV (and sklearn for text/tabular); audio, video, and LLM emphasize import → curate → package for external trainers.
What we are not
OneOpen ML Studio is not trying to be a full cloud data warehouse, feature store, or production inference platform in early releases. Focus stays on high-quality training data and the reproducible path from data to trained model.