Audit-grade provider data with field-level provenance. Survives compliance review.
Most provider-data vendors say "we're accurate" and offer an NDA. We ship a per-dataset, per-field, versioned Audit Pack a buyer can attach to their internal audit response. Methodology is public. Reproducibility is documented. Limitations are explicit. PDF + JSON downloads are free.
The compliance artifact your buyer's audit team needs
Named Fonteum datasets publish an Audit Pack when the required methodology, source inputs, limitations, and downloads are available. Pack contents and history depth vary by dataset; a pack does not imply that every field has provenance or that every past row can be replayed. Health-tech buyers can retain the published artifact with their own audit record.
What changes for your buyer:the data conversation moves from "take our word for it" to "cite our methodology version in your audit." That changes who buys (data team → compliance owner), changes ACV (3-10× typical), and changes renewal dynamics (compliance budgets stick).
- Fonteum Pilot: contracted-dataset exports on the source-specific delivery schedule stated in the agreement.
- Fonteum Standard: customer-scoped packs and exports for the contracted datasets; no universal weekly refresh is claimed.
- Fonteum Enterprise: custom delivery schedules and methodology references for the contracted scope, with source dates disclosed per output.
Browse published Audit Packs → · Sample: Dermatology Audit Pack →
The link your buyer's product points at
When a named dataset has a published Audit Pack, a compliance officer can retain that artifact with an internal audit response. The corresponding methodology page is the public URL their product can cite. Examples include /methodology/dermatology-supply, /methodology/cardiology-supply, and /methodology/nursing-home-quality.
Published methodology pages identify the version, available source and observation dates, reproducibility inputs, limitations, field schema, version notes, and citation formats that the particular dataset supplies. Missing dates, history, provenance, or downloads remain explicit rather than being inferred from a platform-wide cadence.
What changes for your buyer's buyer: the auditor stops asking your compliance officer to defend Fonteum. They click the link. Fonteum defends itself. The compliance review closes faster; the integration sticks longer.
Sample: Dermatology methodology page → · Global methodology + changelog →
Generic NLQ returns a number. We return the evidence trail.
Natural-language questions about U.S. provider supply, answered with the source URL, last-checked timestamp, methodology version, and per-claim limitations the auditor needs to verify each claim independently. Try the public demo →
What changes for your buyer's buyer: the response can expose the available source, date, methodology, and limitation fields behind a supported number. Fields can be null, and consequential claims still need confirmation at the named authority. Out-of-scope questions use the documented refusal code rather than inventing an answer.
How it's Ribbon-proof: the contract only works because the underlying data was built provenance-first (§192 per-field provenance, §194 Audit Pack, §195 refresh tracking, §196 versioned methodology pages). A vendor retrofitting LLM features on a non-provenanced data graph cannot ship this — the citations would have nothing to point at.
Public demo (free, 10/day per IP) → · Authenticated API contract →
The 4th differentiator. Designed for buyers building copilot, RAG, or patient-navigation LLM products.
Generic CSV/JSON exports are designed for ETL pipelines. AI-native exports are designed for LLM grounding: pre-chunked text representations with methodology embedded for grounding, citations embedded for defensibility, and source provenance per chunk so the buyer's LLM cites real Fonteum data — not hallucinated facts.
Three formats per dataset at GET /api/v1/exports/[dataset]/llm-ready:
- NDJSON — newline-delimited, streaming-friendly. Pipe directly into an embedding worker without buffering.
- JSON — single-document RAG envelope. Best for batch ingestion + indexing into a vector DB.
- text-blocks — plain text, paste-into-LLM-prompt-ready,
--- CHUNK ---delimiters.
Embedding-free is intentional. Buyers run their own embedding model (OpenAI / Voyage / Cohere / whichever fits their stack) and own their vector storage. We provide LLM-ready text + provenance + methodology metadata; they retain full control of their AI infrastructure. A 1-2 paragraph synthetic narrative is generated via Claude Sonnet on first request and cached per methodology version, so cache hits are free.
Why this is Ribbon-proof: retrofitting LLM-ready exports onto a non-provenanced data graph means inventing citations that don't trace anywhere. Our chunks cite real audit-pack methodology pages (example) backed by real public sources (CMS NPPES, U.S. Census PEP V2025). The buyer's compliance team sees citations in every LLM-generated response. Their auditor verifies via the public methodology page. The whole audit-grade trust stack works inside their LLM product.
Endpoint contract + integration examples (Pinecone / Chroma / PGVector) →
Provider-discovery data layer
When your patient-facing product needs a defensible, source-cited list of providers in a specialty + geography. NPPES Type-1 enumeration with state-level license matching where applicable.
Example: Dermatology results in Wyoming with available last-checked dates and source URLs; missing metadata remains null.
Network-adequacy analysis
When a payer-network or care-coordination product needs to model provider-to-population ratios at state or county granularity. Per-100k density figures derived from the same per-source provenance contract.
Example: Per-state OBGYN supply at 2.52 / 100k, with the limitation footnotes a diligence team needs.
Access-gap visualizations
When you ship maps or dashboards that need defensible underlying counts. The downloadable CSV/JSON behind every research study is the same data we'd license — with source-row traceability for any state.
Example: Gastroenterology supply: 6 jurisdictions covering 14.7M residents below the AGA-aligned 4/100k threshold.
Sub-specialty enrichment
When your CRM or product database has provider names but lacks the NUCC taxonomy specificity to filter by sub-specialty. We provide the sub-specialty layer with explicit Type-1 / Type-2 disclosure where it matters.
Example: OBGYN sub-specialty breakouts (MFM, REI, GynOnc, Urogyn) with the §185 Type-1 NPPES disclosure pattern.
Every research study ships with a downloadable CSV
You can sample the data shape before scheduling a call. Each shipped study includes a downloadable per-state CSV with the same provenance shape we license to paying customers. Three concrete examples:
- Dermatology supply by state (NPPES 2026) — active dermatologists, per-state per-100k figures, downloadable CSV/JSON
- Cardiology supply by state (NPPES 2026) — active cardiologists, per-state per-100k figures, downloadable CSV/JSON
- Gastroenterology supply by state (NPPES 2026) — active gastroenterologists, per-state per-100k figures, downloadable CSV/JSON
- Discovery (30 min, free). We learn what specialties + geographies your product needs. You see the methodology, the limitations, and a sample CSV.
- Pilot scope (1 week). We agree on 1–3 specialty datasets, refresh cadence, and integration shape. Pricing range: $2,500–$5,000/mo.
- Pilot agreement (signed, 1 week). 90-day term, 30-day no-penalty termination, mutual NDA, internal-use license. Net 30 invoice.
- Pilot delivery (Day 1 onward). CSV / JSON exports on the agreed cadence. Email support, business-hours response within 1 business day.
- Day-75 review. Joint check-in to assess fit. If pilot continues to Standard tier, the agreement converts. If not, the pilot ends cleanly at Day 90.
The pilot is designed so a procurement team can de-risk the buy. No long-term lock-in, no per-seat sprawl, no legal sprawl. Pilot agreement and security/SLA pages are linked from the discovery email.
Everything a procurement team typically asks for, already published:
- Security posture — data scope, infrastructure, attestation roadmap (no overstatement)
- SLA — uptime targets and support-response targets per tier
- Refresh cadence — per-source data freshness, transparently stated
- Methodology — how every published figure was computed
- Data provenance — the per-record source/date/confidence contract
- Standard B2B terms — pilot agreements supersede where they conflict
Ready to discuss a pilot?
30-minute discovery call. We learn your scope, you see the data. No commitment.