# Datagoat > Datagoat is a governed decision engine for agents: it answers typed questions about cases from what happened to cases like them. Send a record (past cases and their yes/no outcomes), questions keyed by id (yesno, score, choice, rank) and the cases to answer; get structured answers keyed by the same ids: per case a chance, the columns that moved it (each with the band of values behind it), and a signed Verdict, plus the pattern where the outcome is most common, or an honest refusal. Deterministic and glass box; a refusal bills no answered cases. Same shape as Jev (typed questions in, typed answers out); Jev reads what a case says, Datagoat reads recorded outcomes. Use it for: will this customer churn, which contract keeps them, which leads to call first, which machines will fault, which agent runs will fail — decisions you make again and again and find out about later. Do not use it for: reading free text, counting, arithmetic, predicting a number, or questions with no recorded history. ## Which journey Five journeys organise the docs, the agent skills and the MCP prompts; `dg_describe` lists them with a ready-to-run call and `start`, the journey for the caller. - [Try it](https://datagoat.io/docs/quickstart): when the user has no data yet, or wants to see an answer and a refusal before using their own. Not when they already have a table of past cases (Ask your data). First call: dg_describe, then dg_ask (a sample's ready-to-run ask). - [Ask your data](https://datagoat.io/docs/ask-your-data): when the user has a table of past cases with a yes/no outcome (churned, converted, faulted) and a question about it. Not when they will score cases for many customers of their own product (Ship a product). First call: dg_add_dataset (upload: true, one per source), then dg_map (returns the ask; dg_backtest and dg_ask follow; a schedule needs fetch_url sources). - [Ship a product](https://datagoat.io/docs/build): when the user will score many customers' cases repeatedly, on a schedule, inside their own product. Not when it is a one-off question about one table (Ask your data). First call: dg_ask (with namespace and model_ttl_days), then dg_ask (by model_ref, no fit). - [Run it](https://datagoat.io/docs/run): when the user already has a model_ref in use and is learning what happened to the cases it scored. Not when no model has been fitted yet (Ask your data). First call: dg_report_outcomes, then dg_schedule (or refit_of + dg_drift by hand). - [Prove it](https://datagoat.io/docs/prove): when the user must show someone the calls were right, or measure whether acting on them worked. Not when they only need the answer (Ask your data). First call: dg_verify, then dg_track_record (or dg_evidence, whether acting on the calls worked). ## Notes for agents - Put every question about one record in one `dg_ask` call; questions on the same outcome and the same `outcome_is_desirable` share one fit. - Each answer has a `state`: `answered`, `refused` (the record holds no reliable pattern; the same call returns the same refusal) or `not_yet` (too few labeled rows or positives; `needs` says how many more). In Datagoat `state` is the answer's status, not the input. - Every chance, level, reason, range, cut point and direction comes from the answer; reasons and patterns are associations, not causes. - Explain one case with its `reasons` (each has `range.text`, e.g. "tenure_months 17 to under 26"); describe the at-risk group with the answer's `pattern` (e.g. "support_tickets over 3") and each case's `pattern_match` ("meets 3 of 4"). The pattern is a few conditions; the chance uses more columns. - `choice` needs `outcome_is_desirable`; the other types take it when you know it. - A record that is not one row per case takes `time_column` and a `shape`; name the columns yourself. - Verify a Verdict with `dg_verify` (no key needed) before anyone relies on it; the SDKs' `verify_offline` / `verifyOffline` check its Ed25519 signature with no call to Datagoat. - A person's own file reaches Datagoat from a chat through `dg_add_dataset` with `upload: true`: it returns `upload_page`, a link they open in a browser to upload the CSV (up to 250 MB, link valid one hour), and `upload_pieces_url`, where code sends the same file in pieces of up to 4 MB on `api.datagoat.io` alone. - Building a product for your own customers: `dg_suggest` lists the yes/no questions a table could be asked (`worth_asking` is a screen; `defined_by` names columns that restate the outcome); `namespace` on `dg_ask` keeps each customer's models and usage apart; `model_ttl_days` (1 to 365) sets how long a fitted model answers, `dg_extend_model` renews it, `dg_delete_model` removes it; a question with `model_ref` instead of `outcome_column` scores new rows, or a newer record's ids, from the model with no fit. Several rows per case (a deal at each stage): `group_column`, or the `snapshots` shape for an event log read as of each moment. Answers list `model_columns` (what a model reads) and, when whole cases were held out, `quality.held_out`; an events or snapshots record asked about by `model_ref` is read with the fit's own activity types, so a small daily batch works; `dg_drift` answers `no_check_yet` until a refit has been compared. - Over MCP, an answer larger than one tool result (about 60,000 characters, roughly 60 cases) returns its first page; `dg_page` with `page.next_cursor` returns the rest in order, and `page.answer_url` is the whole answer as JSON. `export: "csv"` on `dg_ask` adds a CSV of every case. REST returns whole answers. ## Start - [Introduction](https://datagoat.io/docs/introduction): what Datagoat is, the four question types, the five journeys, what it is and isn't - [Quick start](https://datagoat.io/docs/quickstart): try it: a free test key, a first answer and a verified Verdict on the samples - [Connect](https://datagoat.io/docs/install): MCP in Claude, ChatGPT, Cursor, VS Code, Codex; the Claude Code plugin and skills; the SDKs; AgentFactory ## Journeys - [Ask your data](https://datagoat.io/docs/ask-your-data): upload a table or logs, map them, find the questions worth trying, check it, ask, read the answer - [Ship a product](https://datagoat.io/docs/build): one namespace per customer, fit once, score from the model, model lifetime, cases seen at several moments - [Put it in an agent](https://datagoat.io/docs/agents): gate an action on state, band and your threshold, with a person before the side effect - [Run it](https://datagoat.io/docs/run): report outcomes, refit with refit_of, act on dg_drift, renew or retire models - [Prove it](https://datagoat.io/docs/prove): keep signed calls, verify them offline, grade them in a walk-forward replay, and measure whether acting worked ## Concepts - [The record](https://datagoat.io/docs/record): what Datagoat learns from, how to send it, cases, the free samples - [Questions](https://datagoat.io/docs/questions): the four types, question ids, asking several at once - [Answers](https://datagoat.io/docs/answers): answer state, quality, the band, reasons, the pattern, Verdicts and offline verification, levers, closing the loop, cost - [Watching the work](https://datagoat.io/docs/watching): a pending ask's stages as they happen (the event stream, MCP progress, the watch page) and a model's calibration page, all verbatim - [Profiles](https://datagoat.io/docs/profiles): per customer namespace: the words answers use, columns never used, chance or bands for end users - [Shapes](https://datagoat.io/docs/shapes): event logs, series, panels, sensor signals, agent traces, and snapshots (an event log read as of moments you choose) ## Reference - [The ask call](https://datagoat.io/docs/ask): every field of dg_ask and its errors - [API reference](https://datagoat.io/docs/api): the twenty-one operations over REST and MCP, authentication, the OpenAPI file - [SDKs](https://datagoat.io/docs/sdks): Python and TypeScript: every method, offline verification, retries, volume - [Errors](https://datagoat.io/docs/errors): every error code Datagoat returns, what it means and what to do (generated) - [Known limits](https://datagoat.io/docs/limits): what Datagoat is not good at, and the numbers: rows, sizes, rates - [Glossary](https://datagoat.io/docs/glossary): every term once, with its Jev counterpart ## Operations - Try it: dg_describe, dg_ask, dg_verify, dg_map, dg_backtest. Skills: [datagoat-first-run](https://datagoat.io/skills/datagoat-first-run/SKILL.md) - Ask your data: dg_add_dataset, dg_map, dg_backtest, dg_suggest, dg_preflight (deprecated), dg_ask, dg_poll, dg_page. Skills: [datagoat-ask](https://datagoat.io/skills/datagoat-ask/SKILL.md) - Ship a product: dg_ask, dg_profile, dg_extend_model, dg_delete_model, dg_map. Skills: [datagoat-product](https://datagoat.io/skills/datagoat-product/SKILL.md), [datagoat-gate](https://datagoat.io/skills/datagoat-gate/SKILL.md) - Run it: dg_report_outcomes, dg_schedule, dg_delete_schedule, dg_drift, dg_extend_model, dg_delete_model. Skills: [datagoat-product](https://datagoat.io/skills/datagoat-product/SKILL.md) - Prove it: dg_verify, dg_backtest, dg_attest, dg_report_outcomes, dg_evidence, dg_track_record. Skills: [datagoat-prove](https://datagoat.io/skills/datagoat-prove/SKILL.md) - All 5 skills as a Claude Code plugin: `/plugin marketplace add dmilstein-match/datagoat-skills`; listed at https://datagoat.io/.well-known/agent-skills/index.json - MCP: https://api.datagoat.io/mcp (one prompt per journey); OpenAPI: https://api.datagoat.io/openapi.json - The journeys as Arazzo 1.1.0 workflows (the first calls of each, over the OpenAPI operations): https://api.datagoat.io/journeys.arazzo.yaml ## Cost - $0.00002 per answered case ($20 per million); 1,000 fits free each month, then $0.01 each - Refusals and not_yet bill no answered cases (a refusal's fit counts like any fit); verification, sample records, storage, preflight, suggest, drift, outcome reports, attestations, evidence, the track record, profiles and renewing or deleting models are free; a `model_ref` question runs no fit - A card on file (https://datagoat.io/billing) is needed before asking about your own data; the samples need none. Without one, own-data asks return 402 payment_required ## Optional - [Map your sources](https://datagoat.io/docs/map): a table, or event logs with a table, into a record: the proposal, your answers to its questions, the record, replays - [Backtest a record](https://datagoat.io/docs/backtest): day 0: what a mapped record's own past supported, cutoff by cutoff, against a naive rule, with a signed Verdict - [yesno](https://datagoat.io/docs/questions/yesno): the chance the outcome happens (Jev's Noul, learned) - [score](https://datagoat.io/docs/questions/score): that chance as a named level, cut where you choose - [choice](https://datagoat.io/docs/questions/choice): the best option for each case, with each option's chance - [rank](https://datagoat.io/docs/questions/rank): a whole record in order of the chance - [Worked examples](https://datagoat.io/docs/examples): deal risk refused then answered, machine faults, agent runs, churn and the best offer, each with what it does not claim - [Patterns](https://datagoat.io/docs/patterns): fan-out, gate, grade, route, rerank, explain, check, prove, watch, read-then-learn - [Datagoat and Jev](https://datagoat.io/docs/jev): how Datagoat relates to TypeSafe's Jev, and using both - [How answers are checked](https://datagoat.io/docs/checked): the rules every answer passes, what you can check, and what Datagoat does not claim - [Your data](https://datagoat.io/docs/data): what is stored, for how long, and how to delete it - Every docs page in one file: https://datagoat.io/llms-full.txt The free sample records, each asked about with no data of your own (`dg_describe` has a ready-to-run ask for each): - `sample:saas_churn` (table): 800 synthetic SaaS accounts: tenure, charges, support tickets, logins, plan, seats. Which will churn? - `sample:b2b_leads` (table): 800 synthetic B2B leads: pages viewed, demo requests, company size, touch latency. Which will convert? - `sample:telco_churn` (table): 800 synthetic telecom accounts: contract, tech support, tenure, charges. Which will churn? - `sample:customer_events` (events): An event log: a year of visits, purchases, refunds and support tickets from 800 synthetic shop customers. Which customers have gone quiet? - `sample:store_weekly` (series): A time series: 60 synthetic stores over 60 weeks, with sales, footfall, stock cover and late deliveries. Which stores will run out of stock next week? - `sample:usage_panel` (panel): A panel: 800 synthetic SaaS accounts over 12 monthly periods. The outcome is read from the trend in usage. Which accounts are declining? - `sample:sensor_stream` (signals): Signals: four sensors on 40 synthetic machines every six hours for 30 days, with fault intervals in the same table. Which machines will fault in the next three days? - `sample:agent_traces` (traces): Traces: 800 synthetic AI-agent runs with the agent, task and tool of each, one tool degrading mid-month. Which runs will fail? - `sample:deal_activity` (snapshots): Snapshots: sales activity (calls, emails, meetings) on 1,200 synthetic B2B deals, read as of the moment each deal entered each of three stages (the snapshot table sample:deal_stages). Which deals will be won? - `sample:deal_stages` (part of sample:deal_activity): The snapshot table of sample:deal_activity: one row per deal per stage entered (3,600 rows), with the deal's amount, when it closed and whether it was won. Asked alone it is refused: the stage and amount do not predict the outcome; the activity before each stage does. - `sample:parley_billing` (table): Parley's billing export: 811 synthetic small-company accounts of an AI support assistant, one row each (plan, billing cycle, helpdesk, seats, mrr, status, signed_up_at, cancelled_at). Raw, as exported: no outcome column. Map it with its two logs (dg_map, `map`) into a record of who cancels. - `sample:parley_events` (a log, part of sample:parley_billing): Parley's product log: 32,170 timed events (ai_resolution, login, handoff) from the accounts of sample:parley_billing, 2024-01-01 to 2025-11-30. - `sample:parley_tickets` (a log, part of sample:parley_billing): Parley's support tickets: 4,543 tickets (how_to, bug, billing) opened by the accounts of sample:parley_billing. - `sample:parley_record` (built by dg_map from `sample:parley_billing`, `sample:parley_events`, `sample:parley_tickets`): The Parley record: the documented mapping of sample:parley_events and sample:parley_tickets (logs) with sample:parley_billing (a table), each account read every 4 weeks for whether it cancels within 90 days. Built by dg_map, not a file: ask about today's open accounts, or backtest it.