Skip to main content
Datasets in Pioneer are collections of labeled examples used to train and evaluate models. Each dataset has a name you define, and Pioneer versions it automatically as you add or regenerate data. You reference a dataset by name when starting a training job or running an evaluation — so the name you choose is the stable identifier you’ll use throughout your workflow.

How datasets are created

You create datasets in two ways: Synthetic data generation — Use POST /generate to have Pioneer produce labeled examples from a description of your domain and the labels you care about. This is the fastest way to bootstrap a dataset without any existing labeled data.
POST /generate also supports task types beyond NER — pass a different task_type and its required field: The endpoint returns 202 immediately with a job_id; generation itself runs asynchronously. Once generated or labeled, examples are stored in your dataset automatically. Poll GET /generate/jobs/:job_id until status is ready (or failed, in which case check the error field) before starting training.

Uploading your own dataset

Uploading your own data: Use POST/felix/datasets/upload/url if you already have labeled data. This is a three-step process:

Step 1. Get a presigned upload URL

Only dataset_name is required — dataset_type defaults to "ner" if omitted, and accepts "ner", "classification", "custom", or "decoder" (the type field is "training" by default; "benchmark" is rejected here since benchmark datasets are system-managed). The response includes ‘presigned_url’, ‘dataset_id’, and ‘version_number’.

Step 2. Upload the file directly to S3

This is a direct HTTP PUT to S3. Do not include your API key here.

Step 3. Trigger processing

This call returns immediately (202) with status uploading; the dataset then moves through the remaining statuses in the background: initialized → uploading → converting → validating → ready Poll GET /felix/datasets/{name}/{version} until status is ready (or failed, in which case check processing_error) before starting a training job. You can also pass latest in place of a version number to always fetch the newest version.

Listing your datasets

Retrieve all datasets in your account:
The response lists each dataset by name along with metadata such as creation time and version count. Datasets with status failed are excluded by default — pass include_failed=true to see them too.

Inspecting a dataset

To see the versions and details of a specific dataset, pass its name:
This returns version history and example counts, which is useful for confirming the dataset is ready before training.

Deleting a dataset

Deleting a dataset by name soft-deletes it and all its versions — S3 data is preserved in case you need to restore it. If any evaluation or training job is currently pending/running against the dataset, the delete is rejected with 409 Conflict until that job finishes. Once deletion succeeds, already-completed training jobs and evaluations remain queryable, but you can no longer start new jobs referencing it.
Dataset storage is free. You are not charged for storing datasets in Pioneer, regardless of size or number of versions.

Dataset endpoints summary

Data Privacy: If you would like to opt out of having your data used in Fastino’s model training, please email support@fastino.ai and we will ensure your data is excluded from our training pipelines.