Open chat
24,110 records · 18.1m tokens
Review dataset clusters, compare their scale and fingerprints, and tag the slices that need curation or expansion.
Training clusters
4 clusters · 58.5k records
24,110 records · 18.1m tokens
18,240 records · 12.8m tokens
9,420 records · 9.4m tokens
6,710 records · 5.7m tokens
Clusters turn a flat record list into an actionable map. Sort by size or recency, identify dominant pockets, and keep review tags beside each group.
Group semantically related records across the full dataset.
Compare record counts, token volume, fingerprints, and existing tags.
Mark clusters for review, curation, generation, or export.
See where repeated themes may overpower the intended training mix.
Identify small or missing clusters that need additional records.
Apply shared tags without maintaining a separate tracking sheet.
Join the waitlist for early access to training-data clustering.