Data Architecture & Provenance
PRIME DATA MELBOURNE AUSTRALIA
For Developers and Advanced
HOW IT WORKS
## Dataset Architecture: The 17k-to-9 Mapping Matrix
We do not sell this dataset as a tool for "creative text generation." It is engineered specifically as a deterministic routing and classification matrix.
Here is exactly how the dataset is structured:
*The Input Volume:** 17,524 unique, message-derived analysis records extracted from real-world SMS communication.
*The Classification Bottleneck:** These 17,524 chaotic records are aggressively mapped down into nine (9) controlled intent and response patterns (`prompt_input` / `ideal_agent_response` pairs).
Why 9 patterns? Because in high-risk environments, you do not want your AI generating 17,000 unique ways to handle a safety violation. You want the AI to recognize thousands of chaotic user variations and strictly route them into one of nine legally approved, compliance-safe workflows. This dataset trains your agent to recognize the noise and enforce the guardrail.
2. Add this right next to the "Download 50-Record Sample" button:
### Download the Evaluation Sample (50 Records)
We provide a highly curated 50-record mini-dataset so your ML team can test our deterministic classification mapping before purchasing a Tier 1 or Tier 2 license.
To ensure complete transparency and reproducible evaluation, this sample is statically generated.
Sample Dataset Provenance:
*Total Size:** 50 distinct structured JSONL objects
*Coverage:** Covers 7 of the 9 master response templates
*Extraction Seed:** Fixed random seed (`20260713`)
*Data Hygiene:** 0 malformed lines; exact duplicate objects excluded prior to sampling
*File Integrity (SHA-256):** `a20b4fc1d52852dc20c651674e6553970208077caad452edf146d0e1d8b0f444`
Sample Stratification:
* [10] Safety & Guardrail records
* [15] High-Urgency / Negative escalation records
* [25] Standard Workflow records
[ Download sample_50_records.jsonl ]