Data Architecture & Provenance

PRIME DATA MELBOURNE AUSTRALIA

For Developers and Advanced

HOW IT WORKS

## Dataset Architecture: The 17k-to-9 Mapping Matrix

We do not sell this dataset as a tool for "creative text generation." It is engineered specifically as a deterministic routing and classification matrix.

Here is exactly how the dataset is structured:

*The Input Volume:** 17,524 unique, message-derived analysis records extracted from real-world SMS communication.

*The Classification Bottleneck:** These 17,524 chaotic records are aggressively mapped down into nine (9) controlled intent and response patterns (`prompt_input` / `ideal_agent_response` pairs).

Why 9 patterns? Because in high-risk environments, you do not want your AI generating 17,000 unique ways to handle a safety violation. You want the AI to recognize thousands of chaotic user variations and strictly route them into one of nine legally approved, compliance-safe workflows. This dataset trains your agent to recognize the noise and enforce the guardrail.

2. Add this right next to the "Download 50-Record Sample" button:

### Download the Evaluation Sample (50 Records)

We provide a highly curated 50-record mini-dataset so your ML team can test our deterministic classification mapping before purchasing a Tier 1 or Tier 2 license.

To ensure complete transparency and reproducible evaluation, this sample is statically generated.

Sample Dataset Provenance:

*Total Size:** 50 distinct structured JSONL objects

*Coverage:** Covers 7 of the 9 master response templates

*Extraction Seed:** Fixed random seed (`20260713`)

*Data Hygiene:** 0 malformed lines; exact duplicate objects excluded prior to sampling

*File Integrity (SHA-256):** `a20b4fc1d52852dc20c651674e6553970208077caad452edf146d0e1d8b0f444`

Sample Stratification:

* [10] Safety & Guardrail records

* [15] High-Urgency / Negative escalation records

* [25] Standard Workflow records

[ Download sample_50_records.jsonl ]