type -> format id -> kvscience/tasks def -> the KVMACHINES telemetry build and the Claude Science page tests (group claude-science: one task per docs page, one MLflow run per page in MLflow experiment 2, kvmachines/claude-science, written straight to the sqlite store at /Users/zhouk/workspace/claudia/var/mlflow/mlflow.db -- not through a server, which a Claude Science sandbox cannot reach), one record per task, rendered by site/index.html at https://kvscience.com; state is done | running | waiting | building | open | blocked; `needs` names who or what the task waits on; `evidence` is the command, file or page that proves the state; deployed on push to main keys -> type, id, then alphabetical line -> key -> value; parse by splitting on the first " -> " updated -> 2026-10-10T13:22:59Z type -> task id -> honeycomb-organization evidence -> claude.ai admin settings, Claude Code, managed settings v12; data/admin-settings/server-managed-settings.json in the claudia repository needs -> the operator pastes the Honeycomb ingest key (environment test, name mac-collector) into the managed settings editor state -> done title -> Honeycomb for every agent in the organization what -> every Claude Code session of every member, on any host, exports traces, metrics and logs straight to api.honeycomb.io (team kvmachines.com, environment test, dataset claude-code) through the organization's server-managed settings type -> task id -> mlflow-managed evidence -> enabledPlugins mlflow-tracing@mlflow-plugins and MLFLOW_CLAUDE_TRACING_ENABLED in managed settings v12; the Mac's tracking URI in its user and machine settings needs -> nothing state -> done title -> MLflow tracing in managed settings what -> the mlflow-tracing plugin is enabled for the whole organization; its Stop hook posts each session as a trace into experiment kvmachines/claude-code; only the tracking URI is per host type -> task id -> mlflow-cloud evidence -> the tunnel is up and the REST API answers through it: GET /api/2.0/mlflow/experiments/get returns kvmachines/claude-science and runs/search returns all 4 runs, while the same call without the service token gets 302 to the Access login. Tunnel kv-mlflow (ca49064a-06c1-4bab-a5db-612c546337a6) -> http://127.0.0.1:5000, proxied CNAME mlflow.kv-keys.com, Access app kv-mlflow with an Operator policy and the non-identity token kv-mlflow agent gotcha -> MLflow's own middleware rejected the proxied request with "Invalid Host header - possible DNS rebinding attack detected" because mlflow.kv-keys.com is not in the server's --allowed-hosts; fixed in the tunnel rather than the server by setting originRequest.httpHostHeader to 127.0.0.1:5000, the same rewrite kv-mirror uses needs -> nothing for the Mac; a cloud VM still needs CF-Access-Client-Id and CF-Access-Client-Secret in its environment alongside MLFLOW_TRACKING_URI state -> done title -> MLflow tracing in cloud sessions what -> a cloud VM gets the plugin and flags from the same managed settings; its setup script writes MLFLOW_TRACKING_URI https://mlflow.kv-keys.com into the VM's user settings type -> task id -> mac-experiment evidence -> MLflow experiment 1 holds trace tr-934e895af1037aa7b0d94f19c46c99fa (2026-10-10T11:10:41Z, state OK, claude_code_version 2.1.296), found with POST /api/3.0/mlflow/traces/search needs -> nothing state -> done title -> One traced session on the Mac what -> one headless prompt on claudia's account ran with the plugin and the local tracking URI; MLflow recorded the trace with its cost type -> task id -> linux-headless evidence -> claudia data/linux/Dockerfile and scripts/claudia linux-test build|run; build log var/log/linux-test-build.log needs -> a long-lived token the operator creates with claude setup-token and stores in Proton Pass as kvmachines-claude-oauth state -> building title -> Linux headless test what -> an Ubuntu 24.04 image with Claude Code, uv and MLflow builds under apple/container on the Mac; one headless prompt in it is traced the same way type -> task id -> honeycomb-research evidence -> issues 38 to 43 in kvkeys/kvkeys-in-chrome, label telemetry, board github.com/orgs/kvkeys/projects/5; issue 38 runs on the finance seat needs -> five more cloud sessions on Premium seats (design, legal, engineering) state -> running title -> Honeycomb research, six issues what -> each issue is one research task a cloud session works from its own issue: send-data and limits, the honeycombio GitHub org, the MCP server and agent skill, resources as code, the Claude Code signal mapping, the collector type -> task id -> kvscience-site evidence -> https://kvscience.com returns 200 and serves this file; worker kvscience in account 634ff28065e23b4f4a0b6200d62bc97b, custom domains kvscience.com and www.kvscience.com attached to zone bb7e8bef7641bf01f0e6eb48fd204311. GitHub Actions had never run (total_count 0) because deploy.yml triggers on push to main while the only branch was master; master is now renamed to main and set as the default how -> the first deploy was made through the Cloudflare connector (an authenticated MCP session on the account), not wrangler, because the sandbox has no Cloudflare API token. It is a bootstrap worker that inlines site/index.html and site/tasks.kv rather than the wrangler [assets] build; CI replaces it with the real assets build on the first successful run needs -> the repository secret CLOUDFLARE_API_TOKEN, which does not exist yet (actions/secrets total 0); until it is added in Settings -> Secrets and variables -> Actions the deploy job fails at the wrangler step, and every publish has to go through a Claude Science session state -> done title -> kvscience.com progress site what -> this public page, behind Cloudflare, renders tasks.kv so anyone can follow the build without a login type -> task id -> claude-science evidence -> Claude Science 0.1.60 serves on the Mac (port 8000), signed in as KVMACHINES from the scientist seat; MLflow experiment 2 kvmachines/claude-science holds the page tests, one run per page needs -> the next page test, run by a Claude Science session on the scientist seat, never by claudia (the operator, 2026-10-10: "please use not claudia usage", "use scientist@kvmachines.com") state -> running title -> Claude Science experiment what -> a Claude Science session runs one experiment that tests the build: reads its telemetry settings, sends one request through them, reports whether Honeycomb and MLflow received it type -> task id -> motion-videos evidence -> none yet needs -> renders in 16:9 and 9:16, named yyyymmdd-- state -> open title -> Motion videos, one per MLflow implementation what -> three thirty-second Motion videos: managed settings, cloud sessions, the Mac and Linux experiments type -> task id -> claude-science-01-overview evidence -> MLflow experiment 2 run 46bcaac6ea124e27b25a66fea40226ed, attempt 2, pass 1: all five claims of the page hold on this install — version (Settings > General > About reads 0.1.60), kernels (Python 3.11.16, R 4.5.3, bash 5.3.20), sandbox (bind EPERM, no DNS, proxy-only egress, private IPs hard-refused), artifacts (versioned, each with a latest_version_id and lineage), reviewer (host.findings() answers, count 0); attempt 1 run 33be6d00bdcd4f4d81b48f6d391796f0 scored pass 0 only because it read the version from a stale 0.1.55 bundle group -> claude-science corrected -> there is only ONE install; the earlier claim that /Applications/Claude Science.app was "a separate stale install at 0.1.55" was wrong. That bundle is the launcher shell and its CFBundleShortVersionString is the launcher's own version; the daemon that serves the app runs from ~/.claude-science/runtime/0.1.60-release and About reports that. The Mac session trashed the bundle on this agent's advice, found the running app (pid 16393) executing from ~/.Trash, and moved it back. Read the version from `ls ~/.claude-science/runtime` or the daemon log, never from the plist, and never delete that bundle needs -> nothing page -> https://claude.com/docs/claude-science/overview run -> 46bcaac6ea124e27b25a66fea40226ed state -> done title -> Claude Science · Claude Science what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-02-get-started evidence -> MLflow experiment 2 (kvmachines/claude-science) run 88a253340f2e4880b5a1a6f0ac52719a, attempt 1: installed (0.1.60) and serving on port 8000, but the app was signed into the retired AKWLABS account, so pass 0; attempt 2 follows the sign-in fix; attempt 2 run d20d535575ec426b9c0a5353dda97a65: signed in as KVMACHINES from the scientist seat, environments ready, pass 1 group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/get-started state -> done title -> Claude Science · Get started what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-03-run-on-remote-linux-server evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/run-on-remote-linux-server state -> open title -> Claude Science · Run on a remote server what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-04-core-concepts evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/core-concepts state -> open title -> Claude Science · Core concepts what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-05-safeguards evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/safeguards state -> open title -> Claude Science · Safeguards what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-06-multiple-computers evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/multiple-computers state -> open title -> Claude Science · Multiple computers what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-07-artifacts evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/artifacts state -> open title -> Claude Science · Artifacts what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-08-comments evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/comments state -> open title -> Claude Science · Comments what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-09-the-reviewer evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/the-reviewer state -> open title -> Claude Science · The reviewer what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-10-tools-and-environments evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/tools-and-environments state -> open title -> Claude Science · Tools and environments what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-11-troubleshooting evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/troubleshooting state -> open title -> Claude Science · Troubleshooting what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-12-connectors-and-skills evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/connectors-and-skills state -> open title -> Claude Science · Connectors and skills what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-13-remote-compute-clusters evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/remote-compute-clusters state -> open title -> Claude Science · Remote compute clusters what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-14-compute-providers evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/compute-providers state -> open title -> Claude Science · Compute providers what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-15-cloud-storage evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/cloud-storage state -> open title -> Claude Science · Cloud storage what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-16-literature-access evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/literature-access state -> open title -> Claude Science · Literature access what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-17-custom-connectors evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/custom-connectors state -> open title -> Claude Science · Custom connectors what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-18-enable-claude-science evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/enable-claude-science state -> open title -> Claude Science · Enable Claude Science what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-19-how-claude-science-works-with-your-data evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/how-claude-science-works-with-your-data state -> open title -> Claude Science · How Claude Science works with your data what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-20-admin-controls evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/admin-controls state -> open title -> Claude Science · Admin controls what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-21-manage-on-devices evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/manage-on-devices state -> open title -> Claude Science · Manage on devices what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-22-corporate-networks evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/corporate-networks state -> open title -> Claude Science · Corporate networks what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-23-network-requirements evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/network-requirements state -> open title -> Claude Science · Network requirements what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-24-monitor-usage evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/monitor-usage state -> open title -> Claude Science · Monitor usage what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-25-glossary evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/glossary state -> open title -> Claude Science · Glossary what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-26-command-line-settings evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/command-line-settings state -> open title -> Claude Science · Command line settings what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-27-configuration-file-reference evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/configuration-file-reference state -> open title -> Claude Science · Configuration reference what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-28-changelog evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/changelog state -> open title -> Claude Science · Changelog what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page type -> task id -> claude-science-29-legal-and-compliance evidence -> none yet group -> claude-science needs -> one MLflow run in experiment kvmachines/claude-science that records the test of this page (pass 1 or 0, the evidence) after the page was followed on the operator's Mac page -> https://claude.com/docs/claude-science/legal-and-compliance state -> open title -> Claude Science · Legal and compliance what -> follow the page as written, one page per run; record what worked and what did not as the run's params, metrics and tags; fix before moving to the next page