# TaskStation full public content corpus > The open-source AI command center for your company. Every agent, skill, and memory is a file in one versioned repo you own — a workforce of AI agents that does real work, shared across your whole team from Slack, Teams, the web, or the CLI. Self-hostable, any model, your keys. # TaskStation – The AI Command Center for Your Company The open-source AI command center for your company. Every agent, skill, and memory is a file in one versioned repo you own — a workforce of AI agents that does real work, shared across your whole team from Slack, Teams, the web, or the CLI. Self-hostable, any model, your keys. Canonical page: https://taskstation.co/ ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # About TaskStation We build self-driving companies. Humans verify, steer, and govern while agent teams do work across engineering, product, operations, finance, support, and growth. Canonical page: https://taskstation.co/about ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for developers Build with TaskStation from the terminal: drive projects, sessions, and agents through the CLI and SDK, wire triggers and connectors, and land every change through a reviewed change request. Canonical page: https://taskstation.co/developers ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation Enterprise Scale, security, and your deployment. Canonical page: https://taskstation.co/enterprise ## Enterprise Custom Scale, security, and your deployment. - Everything in Team - SAML SSO + SCIM directory sync - Advanced RBAC + audit logs - Cloud, VPC, or on-prem - BYOK, ChatGPT subscription, and managed model controls - SLA, DPA & dedicated support --- # TaskStation pricing Current plans and included features. Canonical page: https://taskstation.co/pricing ## Free $0 Start with real sandbox credits. - 200 credits / month for sandbox compute - 1 project - Bring your own API key for any premium model - Connect your ChatGPT subscription ## Team $40 / seat / mo For teams running real work on agents. - Everything in Free - 2,500 credits / month per seat, pooled - Optional managed models use pooled credits - BYOK and ChatGPT subscription still supported - Up to 200 projects, up to 100 seats - Top up credits anytime - Support via email ## Enterprise Custom Scale, security, and your deployment. - Everything in Team - SAML SSO + SCIM directory sync - Advanced RBAC + audit logs - Cloud, VPC, or on-prem - BYOK, ChatGPT subscription, and managed model controls - SLA, DPA & dedicated support --- # TaskStation Agent Computer Every TaskStation session gets its own computer: an isolated Linux machine that clones your repo, cuts a branch named after the session, and runs OpenCode. Work lands through a change request a person approves. Canonical page: https://taskstation.co/agent-computer ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation Agents & Skills A TaskStation agent is a markdown persona with a deny-by-default reach into connectors, secrets and skills. A skill is the markdown that encodes how your company does one job. Both are files in your repo, versioned and reviewed. Canonical page: https://taskstation.co/agents-and-skills ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation Automations Cron schedules and signed webhooks start TaskStation sessions with nobody present. Each trigger names the agent it runs as, carries a prompt template, and lands its work through a change request a person approves. Canonical page: https://taskstation.co/automations ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation Channels Connect Slack or Microsoft Teams to a TaskStation project and a message in a thread starts a session. The agent works on its own cloud computer and replies in the same thread. Email is in preview. Canonical page: https://taskstation.co/channels ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # Company as Code A TaskStation project is a git repo, and that repo is the company. taskstation.yaml and the OpenCode config define it; agents, skills and memory are files beside your code. Every change is a commit a person approves. Canonical page: https://taskstation.co/company-as-code ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation connectors Connect 3,000+ apps, MCP servers, OpenAPI, GraphQL and raw HTTP once for the whole company. Agents reach them through one scoped token — credentials stay server-side, every action is allowed, gated, or blocked, and every call is logged. Canonical page: https://taskstation.co/connectors ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation Security How TaskStation is built to survive a security review: an isolated machine per session, connector credentials brokered server-side that never enter that machine, permissions for people and agents, human approval gates, and a change request between an agent and main. Canonical page: https://taskstation.co/security ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # Self-host TaskStation Run the same TaskStation on your own box. One Docker Compose stack, the same images the managed cloud runs, your database and your files on disk you control. taskstation self-host start, then taskstation hosts use selfhost. Canonical page: https://taskstation.co/self-hosted ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation Solutions by role One platform, eight teams with completely different work. What sales, marketing, product, engineering, finance, people, IT and data science can each hand off — and what comes back. Canonical page: https://taskstation.co/solutions ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for sales teams Hand a sales agent the research, the pre-call brief, the CRM write-back and the follow-up draft. Nothing sends until you approve it, and every action it took is written down. Canonical page: https://taskstation.co/solutions/sales ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for marketing teams Your voice, your claims and your banned words live in the repo as a skill file every session reads. Marketing work lands as a draft in a change request, so the review is a diff. Canonical page: https://taskstation.co/solutions/marketing ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for product teams Turn scattered feedback into a spec with the evidence attached, keep the tracker honest, and write the release notes from the actual diff. Everything lands as a document a person reviews. Canonical page: https://taskstation.co/solutions/product ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for engineering teams Every TaskStation session gets its own cloud computer and its own branch. The agent reproduces the bug, writes the patch, runs the tests, and opens a change request. Merge is default-deny for agents. Canonical page: https://taskstation.co/solutions/engineering ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for finance teams Reconciliation, variance analysis and the schedules behind the close, run on a machine that shows its working. Every figure traces to a source, and nothing posts to a system of record without approval. Canonical page: https://taskstation.co/solutions/finance ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for people and recruiting teams Interview kits, scheduling, onboarding runs and policy answers drawn from your own handbook. TaskStation does the coordination; a person makes every decision about a person. Canonical page: https://taskstation.co/solutions/people ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for IT teams Access reviews, joiner-mover-leaver runs and service-desk triage, run as sessions on isolated machines. Plus the honest answers IT needs before approving an agent platform at all. Canonical page: https://taskstation.co/solutions/it ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation for data science teams Every TaskStation session is a real Linux machine, so an analysis agent can install a package, run the query, and commit the notebook. The analysis lands in the repo as a change request, so it can be re-run. Canonical page: https://taskstation.co/solutions/data-science ## Official resources - [Documentation](https://taskstation.co/docs) - [Blog](https://taskstation.co/blog) - [Use cases](https://taskstation.co/use-cases) - [GitHub](https://github.com/melihyolacan/suna) --- # TaskStation vs QM: two open agent platforms, two different units of work QM and TaskStation both give teams persistent agents, isolated computers, Slack and web access, and self-hosting. The decisive difference is deeper: QM organizes work around people and rooms; TaskStation organizes it around git-backed projects and reviewable sessions. Canonical page: https://taskstation.co/blog/taskstation-vs-qm Published: 2026-08-02 Author: marko Tags: Comparisons, Architecture, Open Source QM and TaskStation look similar from thirty thousand feet. Both are open agent platforms for teams. Both put agents in isolated computers, persist work beyond one chat, support Slack and the web, and let an operator run the system in their own cloud. But they are not two implementations of the same product. They choose a different **unit of work**, and that decision changes almost everything below it. Compared here: - QM by yc-software (github.com) This comparison is based on [QM main at `7f2c916`](https://github.com/yc-software/qm/tree/7f2c916360f1797a8ff2a77ce2ce40c5fabab087), its published `@yc-software/qm` `0.1.4` package, and TaskStation main at `3006838` on August 2, 2026. We read the runtime, deployment code, security model, storage paths, tests, and operator runbooks. This is a source comparison, not a feature-page comparison. ## The shortest accurate explanation ``` QM person or room -> scope -> persistent memory + files + computer -> conversations, schedules, credentials, published apps TaskStation git project -> session -> isolated computer + session branch -> agent work -> change request -> reviewed merge to main ``` QM starts with the social graph. A person, Slack channel, group message, or project room gets a durable scope. That scope owns memory, files, credentials, schedules, skills, and a computer. TaskStation starts with the work graph. A project is a git repository and `taskstation.yaml`; each session receives a sandbox and a branch, and durable changes return through a change request. > QM is closest to a persistent AI colleague for every person and room. TaskStation is closest to a versioned operating system where many agents work on isolated branches of the company. ## What they genuinely share - **Multi-user by design** — neither system is a single-user desktop agent stretched into a team product. - **Durable work** — both keep state outside the model context and survive process restarts. - **Real computers** — agents execute commands in isolated Linux environments instead of a narrow function-call sandbox. - **Slack and web surfaces** — people can work from a browser or the collaboration surface they already use. - **Background work** — QM has crons and watches; TaskStation has cron and signed-webhook triggers. - **Model choice** — both can use multiple model providers instead of binding the whole platform to one lab. - **Operator ownership** — both can run in infrastructure controlled by the customer. ## Side by side | Dimension | QM | TaskStation | | --- | --- | --- | | Primary unit | Person or room scope | Git-backed project and session | | Durable configuration | Postgres records + deployment layer | Repo files + taskstation.yaml + Postgres | | Computer lifetime | Persistent per scope | Isolated per session; stop and resume | | Agent runtime | Pi, OpenCode, Codex, or Claude Code | One OpenCode REST runtime | | Model routing | Provider selected by deployment or admin | Project gateway, budgets, fallbacks, BYOK | | Collaboration model | Shared room memory and computer | Shared project, isolated session branches | | Durable write path | Scope store and resident computer | Commit + reviewed change request | | Tool credentials | Scoped keychain and resident logins | Server-brokered connectors + explicit secrets | | Security policy | Strict / Auto / Dangerous posture | Per-action allow / ask / block + grants | | Client connection | Core HTTP API; deployment contract package | Published TypeScript SDK + CLI + REST | | Hosted sandboxes | Fly or AWS Lambda MicroVMs | Daytona, Platinum, or E2B | | Self-host path | Per-org Fly or AWS deployment repo | Docker Compose stack; managed cloud too | | License | MIT | Elastic License 2.0 | ## Runtime: QM chooses portability; TaskStation chooses one deep contract QM treats the agent harness as an interface. Pi, OpenCode, Codex, and Claude Code can drive the same core. That is real architectural portability: an operator can change how the model loop runs without changing the surrounding identity, memory, delivery, or sandbox system. TaskStation makes the opposite trade. Every session exposes one OpenCode REST runtime through the sandbox daemon. `@taskstation/sdk` owns session startup, runtime resolution, SSE, files, errors, and the mapping between a TaskStation session and its native OpenCode conversation. Web, mobile, CLI, and white-label clients use the same contract. QM wins if harness interchangeability is the requirement. TaskStation wins if every client needs one stable, typed, deeply integrated runtime surface. TaskStation still routes many model providers through its gateway; it standardizes the **agent runtime**, not the model vendor. ## State: a durable scope versus a versioned company QM stores sessions, memory, queues, grants, audit data, and other control-plane state in Postgres. A scope has a durable workspace and home. On its resident-computer path, installed tools and login state remain warm between turns. Shared artifacts resolve through grants between scopes. TaskStation separates operational state from authoritative company configuration. Supabase Postgres stores accounts, projects, sessions, sandboxes, triggers, grants, audit events, usage, and gateway logs. The project repo stores the agents, skills, memory, policies, and runtime configuration people are expected to edit and review. OpenCode state lives outside `/workspace`, so runtime internals do not pollute the company repo. The practical consequence is important. QM makes continuity of the colleague and room the default. TaskStation makes reproducibility, diff, rollback, and promotion to `main` the default. A QM room can keep accumulating local context. A TaskStation project can show exactly which agent changed the company and which person accepted it. ## Isolation and lifecycle QM assigns a durable computer to a scope. Its Fly path uses a persistent home volume and replaces the root filesystem independently during upgrades. Its AWS path uses Lambda MicroVM agent computers and durable object storage. That design favors a warm, laptop-like environment that remembers installed tools and native CLI logins. TaskStation assigns a computer to a session. Daytona, Platinum, and E2B implement one provider interface. The session branch and sandbox share the session identity. The control plane extends a bounded sandbox deadline when it observes a turn start, gateway LLM activity, or authenticated preview traffic. A terminal turn shortens that deadline to the idle grace period. Passive traffic from an open conversation tab cannot keep the sandbox alive, and one continuous running stretch is capped at 24 hours. A stopped session can resume on the same provider identity or recover through the provider-specific path. Neither lifecycle is universally better. A persistent scope avoids repeating setup for an everyday colleague. A per-session computer provides a clean blast radius and a natural branch for concurrent work. The right choice depends on whether continuity or isolation is the stronger invariant. ## Security: provenance screening versus reviewed change QM selects one organization security posture. **Strict** pauses harness tool calls for human approval. **Auto** screens provenance-labelled external data before it reaches the model. **Dangerous** removes that screening and the per-tool pause, but the command-policy floor still denies or gates declared destructive commands. Narrower scopes can tighten the organization posture. TaskStation centers authorization on the principal, project, session, agent grant, and action. Connector credentials are brokered server-side and do not enter the sandbox. Project secrets are different: an explicitly granted secret is injected as a real environment value and can be read by commands in that session. Connector policies decide allow, ask, or block, and durable repo changes still face a deny-by-default merge boundary. > QM puts more policy around what reaches the model and what each tool call may do. TaskStation puts more policy around which identity acts, which connector action is allowed, and whether durable work reaches the shared branch. ## APIs and product surfaces QM has a headless Fastify core, an HTTP API, and optional Slack, web, admin, auth, and portal plugins. Its supported npm programmatic contract is deliberately narrow: deployment-directory parsing, validation, and provider metadata. The web surface uses Lit; Slack uses Bolt. TaskStation treats `@taskstation/sdk` as a public product boundary. `createTaskStation({ getToken })` returns one client for project and session lifecycle, files, streaming, runtime health, previews, and errors. React hooks, a TypeScript server entry, the real `taskstation` CLI, mobile, desktop, and the web app build on that package. The API and OpenCode transport are implementation details for host applications. This is one of the clearest selection criteria. Choose QM when you primarily deploy and operate the included collaboration product. Choose TaskStation when you also need to embed the platform, build another host, automate it from a CLI, or expose project/session primitives to customers. ## Deployment and operations QM initializes a separate deployment repository pinned to `@yc-software/qm`. The operator chooses Fly or AWS, supplies identity and model-provider credentials, publishes the agent-computer image, deploys, and proves the real computer from outside the transcript. QM intentionally does not generate deployment CI. Its Docker target is documented as an evaluation path, not the recommended production topology. TaskStation offers a managed multi-tenant cloud and a self-hosted Docker Compose distribution. The self-hosted stack includes the frontend, API, LLM gateway, Caddy, and the pinned Supabase distribution. Daytona remains outside the box by default, so a persistent public callback domain or tunnel is required. Managed production uses GitOps on EKS, while the web ships separately. QM gives the operator a cleaner per-organization cloud boundary, at the cost of provisioning more infrastructure. TaskStation gives smaller teams a shorter Compose path and a hosted product, at the cost of a broader platform stack. Both still depend on external model billing, and both use external sandbox compute on their recommended paths. ## How to deploy and test QM For a real organization, QM recommends Fly or AWS. Start in a new organization-owned deployment repository. The initializer pins the exact `@yc-software/qm` package version and writes the provider-specific config, runbook, skill, secret schema, and Slack manifests. Docker exists for a local test drive, but QM does not present it as the production path. ``` npm exec --yes --package=@yc-software/qm@latest -- \ qm init . --org --target --model-provider npm install npm exec qm -- check npm exec qm -- doctor npm exec qm -- infra build-image # AWS npm exec qm -- plan npm exec qm -- up --yes npm exec qm -- check --live npm exec qm -- conformance npm exec qm -- outputs --json ``` - `check` validates config, computed secret names, tools, skills, and plugins without network access. - `doctor` performs read-only checks against the selected cloud, identity, model, and deployment prerequisites. - `plan` renders the mutation before deployment. `up --yes` applies it. AWS builds the Lambda MicroVM image first; Fly publishes its agent-computer image. - `check --live` detects drift in the deployed workloads and sandbox pins. `conformance` compares the static contract with the core’s resolved deployment layer. - **The acceptance test is external** — sign in through the real web URL, send one message, receive a real model response, ask the agent to write a fresh UUID into `/root/workspace/qm-computer-proof.txt`, then verify that UUID through the provider outside the transcript. If Slack is enabled, mention the bot in a test channel and require a real response. - **For source contributors** — run `npm test`, `npm run typecheck`, `npm run lint`, and `npm --prefix cli run test:all`. Postgres and live-provider suites are separate because they require their real dependencies. ## Scaling, observability, and performance - **QM scaling** — durable Postgres stores, background queues, leader leases, multi-instance-safe state, and one computer per active scope. The admin plane exposes sessions, model requests, errors, cost, egress decisions, and audit data. - **TaskStation scaling** — provider-balanced sandbox creation, per-account concurrency limits, control-plane-observed sandbox deadlines, idle and orphan reapers, EKS/GitOps for the managed control plane, and one computer per active session. Provider events, boot timelines, gateway request logs, audit events, and compute metering are durable. - **No honest benchmark winner yet** — the projects publish different tests and target different lifecycles. QM should be faster on repeated tool use in one warm scope. TaskStation should isolate parallel work more cleanly and amortize boot through provider snapshots and resume. Those are architectural expectations, not an apples-to-apples measured result. ## Licensing is not a footnote QM is MIT-licensed. You can modify it, redistribute it, and build a hosted service from it under the MIT terms. TaskStation uses the **Elastic License 2.0**. You can inspect, modify, and self-host the source, but the license restricts providing the software to third parties as a competing hosted or managed service. Some enterprise functionality also requires a license entitlement. If your goal is to create a commercial hosted fork, QM has the more permissive license. If your goal is to run the system for your own organization, both support that deployment model. Read [QM’s license](https://github.com/yc-software/qm/blob/main/LICENSE) and [TaskStation’s license](https://github.com/melihyolacan/suna/blob/main/LICENSE) before making a product decision. ## Can you migrate between them? There is no drop-in migration because the ownership models differ. A practical QM-to-TaskStation migration maps scopes to projects, scope skills and memory to repo files, crons to triggers, keychain entries to connectors or secrets, and active work to sessions. The hard part is deciding which room-local state belongs in version control and which should stay operational data. A TaskStation-to-QM migration maps projects or teams to scopes, imports agents and skills into a deployment layer, converts triggers to crons, and replaces change-request promotion with QM’s scope storage and app-publishing model. That direction loses the automatic branch-per-session review boundary unless you rebuild it as a skill and policy. ## When to pick which ### Choose QM if your core product is a persistent AI colleague for every person and Slack room, you value interchangeable harnesses, you want resident computer state, and the MIT license matters. ### Choose TaskStation if your core product is a versioned company or customer project, you need isolated concurrent sessions, reviewed change requests, a published SDK and CLI, managed cloud plus self-hosting, and provider-independent session infrastructure. They could coexist, but no supported connection exists today. QM could own the conversational scope while TaskStation runs project sessions and change requests behind it. That adapter would have to preserve identity, grants, delivery provenance, and idempotency across both systems. Treat it as a connector project, not a configuration flag. ## The real conclusion QM is one of the more serious new open agent architectures because it starts from multiplayer identity instead of adding team features after the fact. Its scope model, harness portability, resident computers, and provenance-aware security are worth studying. TaskStation makes a different bet: the company should be a git repository, agents should work on isolated session branches, and durable changes should pass through review. That gives up some room-local continuity and harness flexibility. In exchange, it makes ownership, embedding, parallel work, rollback, and institutional learning explicit parts of the product. ## Put the project model to work. Create a TaskStation project, start an isolated session, and review the change your agent brings back. Self-host it or use TaskStation Cloud. --- # The only moat that matters: why your AI platform needs a learning loop, not a better model Every AI product is converging on the same architecture. The only defensible advantage is a data flywheel — a learning loop where every interaction makes your system better. Here is what that means, and why TaskStation is built for it. Canonical page: https://taskstation.co/blog/the-only-moat-that-matters Published: 2026-08-02 Author: marko Tags: Vision, Architecture, Open Source There is a convergence happening in AI right now that is barely discussed. Open any agent platform — Perplexity Computer, Manus, GenSpark, OpenClaw, Hermes, Claude Cowork, Notion AI, Lovable, Cursor, Replit — and architecturally, they are nearly identical. An LLM with tools, a sandboxed execution environment, a memory layer, and a multi-step loop. The marginal differences are UX, a handful of custom connectors, and how they handle memory. None of that takes more than a few months to replicate. This is not a problem for any one company. It is a structural reality of the market. Static software — code you write once and run forever — no longer creates a defensible advantage. Anyone can copy it, and with AI-assisted coding, they can do it faster than ever. Writer.com, a company valued at over a billion dollars, built a significant portion of their enterprise platform by cloning open-source code. The code was not the moat. The code was the starting line. > Static software has no moat. The only defensible advantage in the AI era is a data flywheel — a learning loop where every interaction makes your system better, and more usage compounds into a gap that is genuinely hard to close. Consider why autonomous coding agents actually work. The reason they are viable today is that the environment — bash, a file system, a running process — provides a deterministic pass/fail signal. Did the test pass? Did the API call return 200? Did the build succeed? The answer is a boolean. The agent writes code, runs it, gets a signal, and iterates. The environment is the feedback loop, and the feedback loop is what makes the agent improve. The same principle applies at the company level. The moat is not the model. The moat is the system that captures what worked, why it worked, and what context mattered — and feeds that back into the next run. Every completed task, every rejected proposal, every human override is a training signal. The company that captures that signal and builds it into its agents' behavior is compounding. The company that treats each session as a fresh context window is starting from zero every time. ## Human capital meets token capital Satya Nadella recently wrote about the need for every company to build two kinds of capital: human capital and token capital. Human capital is the knowledge, judgment, relationships, and pattern recognition of your people. Token capital is the AI capability your company builds and owns — the skills, the workflows, the persistent memory, the private evaluation datasets that capture what "good" means for your specific business. The key insight: human capital does not become less valuable as token capital grows. It becomes more valuable. Humans set the ambitious goals, connect dots across domains, build relationships, recognize patterns that matter. Without human direction, you have compute running in circles. The learning loop between people and AI systems is what compounds. Offload a task, or even a job, but never offload the learning. The company that builds a system to capture that learning, encode it, and feed it back into its AI workforce is building the only real moat left. ## The test: can you swap the model without losing what you built? Satya proposed a test that every company should run on its AI platform: can you switch out the model without losing the institutional expertise you have built? If your state lives in a context window, you lose it when the model changes. If your workflows are embedded in a proprietary vendor's toolchain, you do not own them. If your company's knowledge is training data for a model you do not control, you have not built a moat — you have donated your IP. This is the architectural question that matters more than any model benchmark. A platform is sovereign when: - Your agents' skills and memory are stored in files you own, not in a vendor's database. - You can swap the underlying model without rebuilding the system. - Your company's private evaluation datasets are yours, and they determine what "good" means. - The feedback loop — the signal that improves your agents — stays inside your perimeter. - You can self-host the entire stack on your own infrastructure. If any of those is false, you are not building a moat. You are renting one. ## How TaskStation is built for the data flywheel TaskStation was designed from the ground up around this thesis. Every architectural decision — the git-native model, the model-agnostic gateway, the isolated sandbox environment, the skill system — is aimed at one thing: enabling your company to build a learning loop that compounds. Here is how each layer works. **Everything is files in a git repo.** The manifest, the agents, the skills, the connectors, the policies, the memory. Versioned, diffable, reviewable, owned by you. When an agent learns something, it goes into a file. When you want to see what changed, you read a diff. Your company's AI operation is not a pile of settings in someone else's dashboard — it is a repository you control. That means your institutional knowledge is never locked into a proprietary format or a vendor's database. **Any model, your keys.** The gateway is model-agnostic by design. Route a cheap open-weight model for bulk work and a frontier model for the hard calls. Switch them without touching your agents. The reasoning engine is replaceable; the platform around it is the product. This is the test Satya described — model independence is the foundation of sovereignty. **The sandbox is the feedback loop.** Every session runs in an isolated Linux machine with a file system, a terminal, and network access. The agent proposes, and the environment verifies. That deterministic signal — pass or fail — is what makes the learning loop work. The more sessions you run, the more signal you generate, the better your agents get. And because every session is isolated, you can run thousands in parallel without the chaos of shared state. **Skills are procedural memory.** A skill in TaskStation is a file: purpose, preconditions, steps, policies, examples, tests, version, and provenance. An agent can propose a new skill, but it ships through review. Over time, your company accumulates a library of proven, tested, versioned capabilities that encode exactly how your business works. That library is your token capital. And it compounds — every skill that gets used generates more signal, which improves the next skill. None of this requires a frontier model. It requires a platform that treats the learning loop as a first-class architectural concern. The model is the reasoning engine. The platform is where the value accumulates. ## The frontier ecosystem, not the frontier model Satya's essay ends with a warning that is worth repeating: "There is no societal permission for an AI future that hollows out entire industries." If all the value is captured by a small number of models, the political economy will not tolerate it. The priority has to be building a frontier ecosystem — one where every company, every industry, every country can own the learning loop that encodes its institutional knowledge. This is the ethos open-source has always represented: platforms that enable more value on top than they capture inside. TaskStation is open-source, self-hostable, and model-agnostic by design. We want every company that uses TaskStation to build its own compounding advantage — not to make TaskStation the only company that gets smarter. That is the only stable equilibrium. And it is the only moat that actually lasts. ## What this means for your company If you are building with AI today, or planning to, the question is not whether your model is better. The model will be equalized in months. The question is whether your system gets better with every interaction. Do you capture the signal? Do you own the learning loop? Can you change the model without losing what you have built? The companies that win the next decade will not be the ones with the best model. They will be the ones that built the best learning loop, compounded it fastest, and owned their own token capital. The moat is not in the code. It is in the curve. ## Start building your learning loop today. TaskStation is the open-source AI operating system where your company's knowledge compounds. Connect your tools, deploy an agent, and start accumulating your own token capital. Free to start, free to self-host. --- # Your company needs two kinds of capital: human and token Satya Nadella’s frontier-ecosystem framework reveals the most important strategic idea in AI: every company must build human capital and token capital, and they compound together. Canonical page: https://taskstation.co/blog/two-kinds-of-capital Published: 2026-08-02 Author: marko Tags: Vision, Strategy, Enterprise Satya Nadella published something recently that I think is the most important strategic idea in AI right now. He calls it the **frontier-ecosystem** framework. The core insight: every company needs two kinds of capital — **human capital** and **token capital** — and they compound together. If you are building an AI strategy without this framework, you are flying blind. ## Human capital is not going away A common fear I hear from founders: "AI will make our people obsolete." I think the opposite is true. Human capital — the knowledge, judgment, relationships, and pattern recognition of your people — becomes *more* valuable as token capital grows, not less. Here is why. Humans set goals. Humans connect dots across domains. Humans build trust with customers and partners. Humans recognize when the model is confidently wrong. Without human direction, compute runs in circles. It optimizes the wrong thing, or it optimizes the right thing into the ground. I saw this firsthand when Writer.com cloned TaskStation’s open-source code and raised $200M. They took the token capital we publicly shared. What they could not clone was the human capital: the years of judgment about what makes an AI agent actually useful in production, the relationships with our early users, the accumulated pattern recognition of what breaks and why. > Human capital does not depreciate when token capital appreciates. It compounds. The people who know the domain, the customers, and the failure modes become the most leveraged asset in the company. ## Token capital is the new balance sheet item Token capital is the AI capability your firm builds and owns. Not the models you rent — the stuff you build. Skills. Workflows. Persistent memory. Private evaluation datasets. Reinforcement learning from your own human feedback. Most companies are burning token capital without realizing it. Every prompt you type into a closed chat interface is a donation. Your institutional knowledge goes into a context window, gets processed, and disappears. The model learns nothing about your domain. You get an answer, but you do not build capability. Token capital is not the API key. It is the system you build *around* the API key that captures signal, evaluates outputs, and improves over time. ## The learning loop is the moat The compound interest happens in the loop between people and AI systems. Every completed task, every rejected proposal, every human override is a training signal. If you capture it and feed it back, your system gets smarter. If you do not, you are starting from zero every time. This is not abstract. It is a concrete architectural decision. Does your platform capture signal and feed it back? Or does it treat every interaction as stateless? - **Private evals** — your own definition of what "good" looks like. This is the new IP of the firm. - **Private RL environments** — the ability to practice, fail, and improve in a safe loop before touching production. - **Persistent memory** — the system that remembers what worked last time, for this user, in this context. A company that builds these three things owns its trajectory. A company that relies on the provider’s eval set, the provider’s RLHF, and a fresh context window every time is renting intelligence, not building it. ## Offload the task, not the learning Nadella put it simply: "You can offload a task but never offload your learning." This is the line every company needs to draw. Delegate execution to AI. Keep the learning in-house. When you offload a task, you get efficiency. When you offload learning, you get dependency. The platform learns; you do not. The vendor improves; you stagnate. Over time, your cost goes down but your capability ceiling hardens. > The test: if you stopped paying the vendor tomorrow, what would you keep? If the answer is "nothing," you have offloaded learning, not just tasks. ## Where TaskStation fits We built TaskStation around this idea. Skills are procedural memory — repeatable expertise encoded in code, not context windows. Everything is files in a git repo, so your token capital is versioned, forkable, and portable. The sandbox is the feedback loop where humans evaluate, correct, and improve. Every interaction builds the system, not just the answer. Human capital and token capital. Build both. They compound. ## Start building your token capital TaskStation is the open-source AI OS where your company’s knowledge compounds. Free to start, free to self-host, free to own your learning loop. --- # The test of sovereignty: can you swap the model without losing what you built? The most important question for any company adopting AI: can you switch out the model without losing the institutional expertise you have built? A sovereignty checklist for evaluating your AI platform. Canonical page: https://taskstation.co/blog/the-test-of-sovereignty Published: 2026-08-02 Author: marko Tags: Architecture, Enterprise, Open Source There is one question that matters more than any other when evaluating an AI platform: **can you switch out the model without losing what you built?** If your state lives in a context window, you lose it when the model changes. If your workflows are embedded in a proprietary vendor’s toolchain, you do not own them. If your company’s knowledge is training data for a model you do not control, you have not built a moat. You have donated your IP. ## The sovereignty checklist Here are the five criteria I use to evaluate whether an AI platform respects your sovereignty: - **Model independence** — can you swap the underlying LLM without rewriting your skills, workflows, and memory? - **Data portability** — can you export everything your system has learned in a standard format? - **Self-hostability** — can you run the entire stack on your own infrastructure? - **Open source** — can you audit, modify, and fork the platform itself? - **Stateless vs. stateful** — does the platform treat your knowledge as persistent state you own, or ephemeral context you rent? Most AI platforms fail at least four of these. That is not an accident. It is a business model. ## Why most platforms fail this test The dominant AI platform model today is the walled garden. You bring your data, your prompts, your workflows into a proprietary system. The platform learns from your usage. The platform improves its models. Your company gets faster answers — but the platform captures the compound learning. When Writer.com cloned TaskStation’s open-source code, they copied our token capital — the skills, the agent architecture, the prompts we had published. What they could not copy was the fact that our code is open. Anyone can audit it. Anyone can fork it. Anyone can self-host it. The barrier to entry is not the code. It is the learning loop. And the learning loop is ours because the platform is ours. > A closed platform is a rental agreement on your own intelligence. The rent goes up every year. The eviction terms are written by the landlord. ## Renting a moat vs. building one There is a seductive pitch: "Use our platform, and you will be so deeply integrated that switching becomes impossible." That is not a moat. That is golden handcuffs. A real moat is something you build that makes you better over time, not something that makes you stuck. The difference is clear when you look at what happens if the model provider changes their pricing, their safety policy, or their availability. If you are locked into one provider’s embeddings, one provider’s tool-use format, one provider’s context window — you do not have options. You have a dependency. ## Model independence is the foundation Model independence is not just about avoiding vendor lock-in. It is about being able to choose the right model for each task. A 7B parameter model running locally might be better for a latency-sensitive internal tool than GPT-5. A fine-tuned open model might outperform a frontier model on your specific domain. A model that costs 10x less might be 95% as good for most tasks. If your platform is tied to one provider, you cannot make these tradeoffs. You are paying the frontier premium for every task, including the ones that do not need it. ## Open source is the only verifiable path I have come to believe that open source is not optional for enterprise AI. Not because of ideology. Because of verifiability. With a closed platform, you cannot verify what happens to your data. You cannot verify how the model is evaluated. You cannot verify what the vendor learns from your usage. You have to trust. And trust is not a security strategy. With open source, you can verify everything. You can audit the code. You can inspect the data flows. You can run the system on an air-gapped network. You can fork it and extend it in directions the original authors never imagined. > The test of sovereignty is simple: can you swap the model without losing what you built? If the answer is yes, you own your AI future. If the answer is no, you are renting it. ## Where TaskStation stands TaskStation passes every item on the sovereignty checklist. Model-agnostic gateway — swap any LLM without rewriting your skills. Everything is files in a git repo — your token capital is versioned, portable, forkable. Fully self-hostable. Open source under a permissive license. We built it this way because we believe the company that owns its learning loop wins. Not the company that rents the best API. ## Run the test on your AI platform TaskStation passes the test of sovereignty. Free to start, free to self-host, free to own your AI future. --- # Static software is dead: the shift from code to feedback loops The way we build software is fundamentally changing. Static software no longer creates a defensible advantage. The shift is to dynamic software that improves through feedback loops. Canonical page: https://taskstation.co/blog/static-software-is-dead Published: 2026-08-02 Author: marko Tags: Architecture, Vision, Engineering The way we build software is fundamentally changing. Static software — code you write once and run forever — no longer creates a defensible advantage. The shift is to **dynamic software**: systems that improve through feedback loops, where the environment provides a deterministic pass/fail signal, and the product gets better with every interaction. ## Why coding agents actually work Coding agents work for a specific reason. The environment provides a deterministic boolean signal. Did the test pass? Did the API return 200? Did the build succeed? The shell, the file system, the running process — these are ground truth. An agent can try something, observe the result, and try again. That feedback loop is what makes autonomous coding possible. This is not magic. It is a well-defined environment with a clear success criterion. The same principle applies to any domain where you can define a pass/fail signal. ## Static software is a document Most software today is static. You write it, you ship it, and it does the same thing until a human changes it. It is a document. A frozen artifact. It does not learn. It does not adapt. It sits there, accumulating cruft, until a developer rewrites it. Dynamic software is different. It is a process that compounds. Every interaction improves the system. Every user session generates signal. The product gets smarter the more people use it. > Static software is a document. Dynamic software is a process that compounds. The difference is the feedback loop. ## At 10,000 tokens per second Generation is effectively free. At 10,000 tokens per second, everything is instantly generated. The code itself is a commodity. The real complexity shifts to building the reinforcement environments — the sandboxes, the test harnesses, the evaluation pipelines — that agents can learn from. The moat moves from writing code to building the feedback loop. The value shifts from the artifact to the environment. If you cannot generate a deterministic signal from your domain, you cannot build dynamic software. ## What this means for engineers Your job shifts from writing code to designing environments that agents can learn from. Instead of hand-writing every function, you define the constraints, the test harness, the evaluation criteria. The agent generates the implementations. You curate what works. This is not a reduction in engineering value. It is a shift. The hard part becomes: can you define a signal for what good looks like? Can you build a repeatable environment where the agent can fail safely and learn quickly? ## The TaskStation approach The sandbox is the feedback loop. Every TaskStation session is an isolated environment with a deterministic signal. The agent tries, observes, and iterates. Skills are the accumulated output of that loop — compressed experience, not static code. This is why we invest in the sandbox architecture. It is the foundation. ## Stop building static software. Start building feedback loops. TaskStation is the platform for dynamic software. Deploy on-prem, own your data, and build systems that learn. --- # Every AI product is the same: the convergence nobody is talking about Open any agent platform — they are architecturally identical. The marginal differences are not defensible. The only thing that diverges is the learning loop. Canonical page: https://taskstation.co/blog/the-convergence-of-ai-products Published: 2026-08-02 Author: marko Tags: Market, Vision, Comparisons Open any agent platform today. Perplexity Computer, Manus, GenSpark, OpenClaw, Hermes, Claude Cowork, Notion AI, Lovable, Cursor, Replit. Architecturally, they are nearly identical. An LLM with tools. A sandboxed execution environment. A memory layer. A multi-step loop. The differences are UX, a handful of custom connectors, and how they handle memory. None of that takes more than a few months to replicate. ## The convergence is real This is not a criticism. It is a structural observation. The underlying architecture of an AI agent platform is converging to a minimum viable set of components. The LLM is a commodity. The sandbox is a commodity. The connector pattern is a commodity. The memory layer is the only variable, and even that is converging to a small set of approaches. Writer.com raised a $200 million series C. They cloned TaskStation’s open-source code to build their agent layer. This is not unusual. It is the norm. If you build something useful, someone will copy the architecture. The question is: what happens after they copy it? > If your competitive advantage is your architecture, you do not have a competitive advantage. Architectures are commodities. Feedback loops are moats. ## What is not defensible - **UX** — good design matters, but it is not a moat. A competitor can match it in a quarter. - **Custom tools** — connectors are implementation work, not differentiation. Everyone will build the same connectors. - **Memory handling** — the approach converges. Short-term, long-term, episodic. Everyone is building the same abstractions. ## What actually diverges The one thing that genuinely differentiates is the feedback loop. The system that captures signal and compounds it. The pipeline that takes every user interaction, every success, every failure, and turns it into a better model, a better skill, a better outcome. This is hard. It requires infrastructure. It requires data that you own. It requires a product that people actually use in production. Most platforms skip this part because it is expensive and slow. They compete on features instead. ## What this means for builders Do not compete on the architecture. Compete on the data flywheel. If you are building an AI product, ask yourself: does every user session make your product better? If the answer is no, you are building static software with an AI wrapper. ## What this means for buyers Do not buy the architecture. Buy the platform that gets better with use. The platform that has real users generating real signal. The platform where the learning loop is not a roadmap item, but the core product. > The architecture is a commodity. The learning loop is the differentiator. Everything else is table stakes. ## The architecture is a commodity. The learning loop is the differentiator. TaskStation is built on the feedback loop. Deploy it, use it, and watch it compound. Start building yours. --- # A frontier without an ecosystem is not stable The first phase of globalization hollowed out industrial economies. The AI era cannot repeat that. The priority has to be building a frontier ecosystem, not just a frontier model. Canonical page: https://taskstation.co/blog/the-frontier-ecosystem Published: 2026-08-02 Author: marko Tags: Vision, Open Source, Industry Satya Nadella said something recently that has stuck with me: **"There is no societal permission for an AI future that hollows out entire industries."** It is one of the most important things any tech CEO has said about this moment, and I think most people in AI missed it. The first phase of globalization did exactly that. It hollowed out industrial economies across the American Midwest, the British north, and the German Ruhr. GDP kept growing. The stock market was fine. But the displacement was real, and the consequences — populism, declining life expectancy, political instability — are still being felt decades later. If you measure the economy by aggregate output, globalization was a success. If you measure it by who captured the gains and who bore the costs, it was a disaster. The AI industry is on track to repeat that same dynamic, only faster. ## The centralization problem Right now, the AI industry is organized around a small number of frontier models. A handful of labs control the most capable systems. Value flows to them. Everyone else is a consumer of intelligence, not a producer of it. This is not stable. If every company, every industry, and every country has to rent its intelligence from the same three providers, the system concentrates power in exactly the way that hollowed out the industrial heartland. The surface-level metrics will look great. The distribution will be brutal. > The lesson of the last thirty years is that when value concentrates, the system breaks. Not eventually — it is already breaking. ## What a frontier ecosystem looks like A frontier ecosystem is the opposite of that. Every company owns its learning loop. Every industry has its own models, trained on its own data, optimized for its own workflows. Every country can build AI that reflects its own language, culture, and regulatory environment. This is not a nice-to-have. It is the only way the AI transition can work at scale. If intelligence is a commodity that everyone can produce rather than a service that everyone must buy, the gains are distributed. The displacement is manageable. The system is stable. ## Open source is the only viable model None of this happens under closed, proprietary models. A closed model is a rent-extraction machine by design. The more value it captures, the less value is available for everyone else. Open source is the opposite: a platform that enables more value on top than it captures inside. This is why the tension between model labs and the ecosystem is real, and it is not going away. The labs are incentivized to centralize. The ecosystem is incentivized to distribute. These are fundamentally incompatible. The labs will try to frame this as a debate about safety, but it is really a debate about who controls the value. ## This is not altruism I am not arguing for open source because it is virtuous. I am arguing for it because it is the only stable equilibrium. A world where three companies control all frontier intelligence is a world that will face the same political backlash that globalization created, only faster and with more at stake. The labs that figure out how to build a real ecosystem around their models — where the ecosystem captures more value than the lab itself — will be the ones that survive. The ones that try to capture everything will be regulated, broken up, or replaced. ## Where TaskStation fits TaskStation is open-source, self-hostable, and model-agnostic for exactly this reason. We want every company to build its own compounding advantage on top of AI, not rent it from someone else. The data flywheel that makes your company better over time should belong to you, not to a model provider in San Francisco. We are not trying to be the only AI platform. We are trying to be the platform that enables every company to be its own AI company. ## The frontier ecosystem needs builders TaskStation is open-source and built for companies that want to own their AI future. Self-host it, connect your own models, and build something that compounds. --- # Good businesses don’t need moats (and why that’s fine) The startup world is obsessed with moats. Every pitch deck has a slide. Every investor asks. But most great businesses do not have a true moat, and they are still great businesses. Canonical page: https://taskstation.co/blog/good-businesses-dont-need-moats Published: 2026-08-02 Author: marko Tags: Strategy, Business, Vision Every pitch deck has a moat slide. Every investor asks about it. Founders spend weeks agonizing over how to frame their defensibility. And I think the whole conversation is mostly wrong. The truth is: most great businesses do not have a true moat, and they are still great businesses. The obsession with moats is a VC narrative, not a business reality. It comes from a venture industry that needs winner-take-all stories to justify the math of a fund that needs one portfolio company to return the whole thing. ## The WordPress example WordPress runs something like 43% of the web. It has created an estimated **$44 billion economy** of developers, agencies, hosting companies, and plugin builders. The company behind it, Automattic, makes around $800 million per year in revenue. Is there a moat? Not really. Anyone can spin up a WordPress site. Anyone can build a competing hosting service. The code is open-source. The plugins are open-source. The barriers to entry are essentially zero. And yet, the business is real. It is large. It compounds. Year after year, it grows. Not because it is defensible in the VC sense, but because it has execution, distribution, brand, and customer relationships that compound over time. ## Consultancies are the counterexample that proves the point McKinsey, Deloitte, BCG, Bain — these are massive, profitable businesses. They are also, at the core, indistinguishable from each other. They hire from the same schools. They use the same frameworks. They compete for the same clients. There is no technology moat, no network effect, no data advantage. And yet, they are some of the most durable businesses in the world. Why? Because they have **execution, brand trust, and relationships** that take decades to build and are not easily replicated. These are not moats in the Warren Buffett sense. They are something more mundane and more real. > The obsession with moats confuses the condition for venture-scale returns with the condition for a good business. They are not the same thing. ## What actually matters If you are building a real business — not a lottery ticket, not a flip — the things that matter are: - **Execution velocity.** Can you ship faster and better than everyone else? - **Distribution.** Do you have a channel that compounds? Referral loops, content engines, sales relationships. - **Brand.** Do people trust you? Would they recommend you? - **Customer relationships.** Do they stay? Do they expand? Do they churn less than your competitors? - **Operational excellence.** Are you running a tight ship? Do you have real margins? None of these are moats in the traditional sense. They are not defensible in the way a network effect is defensible. But they are real, and they compound. A business that does all of these well is a business that will outlast most of its competitors, even if a well-funded copycat could, in theory, replicate every feature. ## Moats are for the winner-take-all story True moats exist. Network effects are real. Data flywheels are real. Scale economies are real. But they are the exception, not the rule. They are the condition for a venture-scale outcome, not the condition for a good business. The problem is that the startup world has internalized the VC frame so deeply that founders think they need a moat or they are not building anything real. This is wrong. It leads to bad strategy: chasing defensibility at the expense of actually building something people want. ## Where this lands for TaskStation TaskStation is open-source. Anyone can clone the repo. Anyone can self-host. Anyone can build a competing product. There is no moat in the VC sense. But the cloud business is real. The brand is real. The open-source community is real. The customer relationships are real. And most importantly, the **data flywheel** is real — every company that uses TaskStation builds a compounding advantage that belongs to them, not to us. That is not a moat around TaskStation. It is a moat around our customers. That is a better trade. We would rather build the platform that makes every customer stronger than build the fortress that keeps everyone out. ## Build a business that compounds TaskStation is the platform for companies that want to own their AI learning loop. Self-host it, connect your models, and build something that compounds over time. --- # What TaskStation actually is: the open-source AI Management System, layer by layer One git repo is the source of truth for the agents, the skills, the memory and the connector config. Every session gets its own isolated machine and its own branch. Any model, your keys. Work lands through a change request. Self-hosted or managed cloud. The long version of that sentence. Canonical page: https://taskstation.co/blog/open-source-ai-management-system Published: 2026-07-31 Author: marko Tags: Product, Architecture, Open Source TaskStation is the open-source AI Management System — the leading open-source alternative to Claude Cowork and ChatGPT Work. One git repo holds the agents, the skills, the memory and the connector config. Every session runs on its own isolated machine, on its own branch. Any model, your keys. Work lands through a change request. Self-host it, or run it on our cloud. This is the long version of that paragraph: every layer, in the order you meet it, with the parts we do not claim marked as such. ## The category is real. The question is who ends up holding it. Agents that deliver finished work stopped being a demo and became a product category. Claude Cowork shipped on the desktop in January 2026 and reached web and mobile on 7 July; ChatGPT Work followed on 9 July. Both are good products. Both also share a shape: they run inside a model lab, on that lab’s models, on that lab’s cloud, with no self-host option — Cowork for Claude Max subscribers on Anthropic models, ChatGPT Work on paid usage-metered plans, on GPT-5.6. Which leaves most companies choosing between a toy and a cage: a single-tenant demo with no isolation, no version history and no permission model, or renting your company back from the lab that keeps your data, your configuration and your model. TaskStation refuses both. The refusal is not a philosophy — it is six layers, and you can read every one of them. ## 01 · One git repo that is the company A TaskStation project is a git repository, and that repository *is* the company. `taskstation.yaml` is the TaskStation layer: the machine a session boots on, the connectors, the triggers, the secret names, and what each agent may touch. The OpenCode configuration beside it is the runtime the agents think in. Everything past those two files is markdown. So the whole company answers to `grep`. Every agent prompt, every skill, every remembered fact and every grant is a line in a file with an author, a timestamp and a diff, and undo is `git revert`. `taskstation init` turns any directory into a TaskStation; `taskstation ship` checks it compiles, asks for the secrets it is missing, and brings it live. ``` # taskstation.yaml — the TaskStation layer of the repo. taskstation_version: 2 runtime: opencode default_agent: taskstation project: name: acme # OpenCode keeps agents, skills, commands, tools, plugins and models here. opencode: config_dir: .taskstation/opencode # An omitted grant resolves to `none`. Grant explicitly. agents: taskstation: connectors: all secrets: all skills: all taskstation_cli: all memory-reflector: taskstation_cli: [project.cr.open] # A trigger starts a session with nobody present. triggers: - slug: memory-reflector type: cron agent: memory-reflector cron: "0 0 3 * * *" timezone: UTC prompt: | Reflect on the last 24 hours of project activity, update .taskstation/memory/, and open one change request. ``` Note two things that manifest deliberately does not contain. There is no `channels:` block — the v2 validator rejects one, because channel routing is live project state rather than repo configuration, and pretending otherwise would put a lie in your git history. And every omitted grant resolves to `none`: leave a connector out of an agent’s list and that agent does not get it. Deny is the state you fall into by accident, not the one you have to remember. ## 02 · Every tool your company already runs on Connect a tool once, for the whole project: 3,000+ apps through their own OAuth screens, or your own APIs through an OpenAPI or Postman spec, a GraphQL endpoint, a remote MCP server, or a bare HTTP base URL. TaskStation reads the source, works out the authentication, and turns every operation into a tool an agent can call. The credential never travels. The machine carries exactly one project-scoped TaskStation token; the third-party key is decrypted server-side and attached to the outbound request, so the raw key never reaches the sandbox. Every action gets one of three answers — allow, ask, or block — and a rule can read the arguments it was given rather than only the tool name, so “only to this domain” is something you can actually express. An ask returns a signed approval URL immediately. One human decision sends a durable callback into the session, and only the exact approved request can run. ## 03 · Any model. Keep your keys. The one safe bet in this field is that a better model ships, so TaskStation is model-agnostic on purpose. Pick the model per agent, per session or per message. Bring your own API key from any major provider, or use ours. Sign in with the ChatGPT subscription you already pay for. Or point it at your own model behind your own URL — anything OpenAI-compatible. When you switch, nothing above this layer moves. ## 04 · The part that turns a model into an agent A model on its own answers a question. A harness gives it planning, tool use, and multi-step runs it actually finishes. TaskStation runs OpenCode as that harness, and an agent here *is* an OpenCode agent: a markdown file carrying a persona and a permission tree is the baseline, and the whole OpenCode lifecycle sits underneath it — commands, tools, plugins, providers, models, and skills that ride into every session that needs them. A skill is a directory with a `SKILL.md` at its root: how your company does one specific job, written once. So how an agent thinks is text you can read, diff and edit. You can say allow, ask or deny per tool, down to a single shell command. And because the harness is open source too, it is never the thing you are locked into. ## 05 · Every session gets its own computer Start a session and its own isolated Linux machine boots. It clones the project repo into `/workspace`, cuts a branch named after the session, and starts the harness. The session id, the sandbox id and the branch name are one and the same string. The agent gets the whole machine — a shell, a package manager, a filesystem, the network — and nothing runs on your laptop. The machine is disposable, so a bad install or a wiped directory goes away with it and only what the agent commits survives. And because one session is one machine on one branch, two sessions cannot touch each other. Run one, or run thousands at once, each a different version of the company working at the same time. That parallel, isolated workforce is the part nobody else has. ## 06 · One place to start it, one gate to land it The web app, Slack, mobile, the CLI and the API all start the same session — same object, same branch, same audit row. Then the work comes back the one way it is allowed to: a change request you read as a diff before anything reaches `main`. Merging one is a capability of its own, and it is **default-deny for agents**. An agent gets it only if an admin grants `project.cr.merge` in `taskstation.yaml`, and widening that grant is itself a change somebody else has to approve. It is not a human-only gate, and we would rather state that precisely than sell you a stronger claim than the code makes. ## Where people actually reach it Bind a project to Slack and a message in a thread starts a session. The agent picks up its own cloud computer, does the work, and answers in the same thread: the reply streams into one message, files move both directions, and a decision it needs from you arrives as a card with buttons. A thread is exactly one session — a unique index in the database, not a convention two services agree to honour. The honest list is short, because the platform enum is closed at three. **Slack is live.** Microsoft Teams is code-complete behind an operator switch. Email is experimental and opt in per project. Telegram, WhatsApp, SMS and Discord are not channels, in any tense. ## When nobody is asking A trigger starts a session with nobody present. There are two kinds and no third: a cron schedule stored against an IANA timezone name rather than an offset, or a webhook signed with HMAC-SHA256. A webhook trigger that names no signing secret is rejected at validation, so there is no unsigned path to forget to lock down later. Both are entries in `taskstation.yaml`, so the 3am job has an author and a history like everything else, and both inherit exactly the reach of the agent they name. The prompt is a template: a webhook fire renders `{{ body.* }}`, a cron fire renders `{{ cron.schedule }}`, `{{ cron.timezone }}` and `{{ cron.scheduled_for }}`. Every fire is a clean slate by default, or a trigger can re-prompt a session it already owns, keyed off the payload, so one customer keeps one thread. ## Permissions and secrets, stated precisely People, groups and service accounts are all principals, and a permission attaches to a principal for an action on a resource type. A service account is evaluated purely against its own policies — it never inherits the reach of whoever created it. Secrets are sealed with AES-256-GCM under a key derived per project. And a session receives only the intersection of the agent’s declared grant and the role of the person who started it, so an agent can never out-reach the human who launched it. > We will not tell you a granted secret is invisible to the model. Once delivered, a runtime secret is a real environment value inside the session, readable by any command the agent runs — because that is how a tool uses it. What holds is narrower and true: **connector credentials never enter the machine at all**, and the machine is destroyed with everything on it. Approval gates get the same treatment. They are **not on by default** — a project that declares no policy block falls back to allowing actions — so the operator sets the default they want and puts an explicit ask on the step that matters. Audit is the one thing that is not optional: recording is never gated, and only reading, exporting and streaming the record are permissions at all. ## What it does on an ordinary Tuesday Layers only matter if they add up to work somebody was already being paid to do. The bar for anything below is the same: a real job, run end to end by an agent with its own machine, a repo, connectors and a schedule, whose output is one concrete artifact. - **Sales** — pulls a lead list, enriches every account and writes a sequence per lead. Put an ask on the send step and it stops with you before anything goes out. - **Engineering** — sweeps the day’s errors, groups them, reproduces the worst one on its own machine, patches it, and opens a change request against `main`. - **Finance** — reconciles the ledger against the bank, chases the receipts that are missing, attaches them, and closes the month with the variance explained. - **Marketing** — tracks the queries you care about, finds the pages losing ground, rewrites them against the brief, and opens each rewrite as a change request. - **Data** — queries the warehouse on a schedule, checks the result against last week, draws the chart, and posts the whole thing to Slack while you are asleep. Every one of those is a job, not a chat. What lands is a file — a diff, a spreadsheet, a draft, a report — with a branch behind it and a person in front of it. ## Read every line, then run it on your own box All of it is open source. TaskStation is developed in the open at [melihyolacan/suna](https://github.com/melihyolacan/suna) — clone the repo, read what you are trusting, fork it if you want it different. Then run that same product on hardware you control. One Docker Compose stack, built from the images the managed cloud runs, so it is the whole platform rather than a cut-down edition, and the database, the file storage, every project repo, the secrets, the policies and the audit record sit on disk you control. ``` # bring the whole stack up on your own box $ taskstation self-host start # point the CLI at your stack $ taskstation hosts use selfhost → Active host is now selfhost # same commands, back on the managed cloud $ taskstation hosts use cloud → Active host is now cloud ``` Two limits, stated plainly. Agent sandboxes run on the compute provider you configure and the stack pulls its images over the internet, so **this is not an air-gapped deployment** — isolated topologies get scoped with us instead. And SAML SSO, SCIM directory sync, custom roles, groups and reading the audit log switch on with an Enterprise licence. Models are yours either way. ## Side by side | Dimension | Claude Cowork · ChatGPT Work | TaskStation | | --- | --- | --- | | Where your configuration lives | Inside the vendor’s product | A git repo you own | | Version history on that configuration | Not published | `git diff` on every agent, skill and grant | | Which models it runs | The vendor’s own — Anthropic, or GPT-5.6 | Any model — your keys or your subscription | | Where it runs | The vendor’s cloud; no self-host | Managed cloud, your VPC, or your own on-prem network | | Reading the source | Closed | Open source — clone it and read it | | Isolation per unit of work | Not published | One isolated machine and one branch per session | | How work lands | Inside the product | A change request you read as a diff first | | Audit record | Not published | Recorded on every plan; reading it is its own permission | | Getting started | A Claude Max plan, or a paid metered ChatGPT plan | Free to start, free to self-host | Two of those rows say “not published”, and they stay that way. Neither lab documents its isolation model or its concurrency limits, so we do not get to characterise them. Where we have no data, the table says so. ## What we are not claiming A page of capabilities is only worth reading if the same page will tell you where the edges are. These are ours. - **Not air-gapped.** `taskstation self-host start` pulls its images over the internet and reaches a sandbox provider, so a fully disconnected install is not a shipped capability. Isolated topologies get scoped directly with us. - **No blanket microVM claim.** One session gets one isolated machine. Whether that machine is a microVM depends on the compute provider you run on, and the default is not one. We would rather name the boundary than the buzzword. - **No network egress control.** Nothing in the product enforces it today. The boundary that is real is the credential one — a connector key never enters the sandbox. - **No certification.** SOC 2 Type I and Type II are in progress, not held. GDPR is a posture the company does hold. We will not print a badge for a report that has not landed. - **One harness, one runtime.** TaskStation runs OpenCode. That is the shipped path, and it is the only one this post describes. ## When to pick which ### Choose a model lab’s agent if you want finished work today with nothing to run, you are happy on that lab’s models and that lab’s cloud, and your company’s configuration living inside their product is a trade you are content to make. ### Choose TaskStation if you want the same finished work with the company underneath it staying yours — agents, skills, memory and connector config as files in [one repo you own](/company-as-code), any model on your own keys, [one isolated machine per session](/agent-computer), and [work that lands through review](/security). None of the above is a roadmap item. Every layer runs today, and every page it points at is written against the code rather than the pitch. The long form on each layer: [company as code](/company-as-code), [agents and skills](/agents-and-skills), [the agent computer](/agent-computer), [connectors](/connectors), [channels](/channels), [automations](/automations), [security](/security) and [self-hosting](/self-hosted). The architecture argument behind all of it is in [AGI-ready architecture](/blog/agi-ready-architecture); the direct comparisons are [Claude Cowork](/blog/taskstation-vs-claude-cowork), [Glean](/blog/taskstation-vs-glean) and [Poetic](/blog/taskstation-vs-poetic). ## Run your whole company from one repo you own. Start with one job, connect the tools it needs, and reach it from Slack, the web or the CLI. Free to start, free to self-host. --- # TaskStation vs Poetic: both turn workflows into code — the difference is who owns the code Poetic compiles your procedures into a purpose-built language it runs for you. TaskStation keeps the workflow as ordinary code in a repo you own, and gates the boundary where it touches the world. A technical comparison of two answers to the same problem. Canonical page: https://taskstation.co/blog/taskstation-vs-poetic Published: 2026-07-31 Author: marko Tags: Comparisons, Architecture, Open Source Poetic and TaskStation start from the same observation: an agent that improvises a high-stakes process a thousand times a day will not do it the same way a thousand times. The fix both companies reach for is code — pin the workflow to something you can read, review and re-run. Where the two diverge is what that code is, where it lives, and who ends up holding it. That is the whole comparison, and it is a real one. Compared here: - Poetic (poetic.com) ## What Poetic actually is Everything in this section comes from Poetic’s own material — their [site](https://poetic.com/) and their [Series A announcement](https://www.prnewswire.com/news-releases/poetic-raises-50m-series-a-to-automate-the-worlds-most-complex-enterprise-processes-with-reliable-ai-302796939.html). Where their public material does not answer a question, this post says so rather than guessing. - **The pitch is "turn your business into software."** Their framing for the execution model is software that "learns like AI but runs like code." - **The primitive is a purpose-built programming language.** Their announcement describes a language that lets operators "define complex workflows in natural language, then encodes that expertise into deterministic, near-tokenless execution." - **Authoring is teaching, not typing.** You upload procedures, recorded sessions and historical examples; Poetic drafts the workflow, then refines it against expert feedback until it is production-ready. - **The target is narrow and deliberate** — "multi-hour processes that run thousands of times a day" that "demand near-perfect accuracy," in financial services and insurance. - **Delivery is forward-deployed.** They reached an eight-figure run rate in 2025 with four employees, working directly alongside large enterprise customers. ## What Poetic is genuinely good at This is not a hit piece, and the strongest thing about Poetic is the part we do not do. Once a procedure is compiled into their language, execution is deterministic and, by their description, near-tokenless. That is a real engineering result with two consequences an agent loop cannot match: the same input produces the same output, and the marginal cost of the ten-thousandth run does not include a frontier model bill. For a dispute investigation that runs continuously against a fixed set of internal systems, that is the correct architecture, and it is why they can publish accuracy figures like 99%+ quality on multi-hour processes. TaskStation does not make a determinism claim, and this post will not pretend otherwise. A TaskStation session is a model loop. It is more general and it is less repeatable. If your problem is one well-bounded process, run at enormous volume, where a 1% deviation is a regulatory event, Poetic is purpose-built for exactly that and TaskStation is not. ## The same instinct, applied one layer down TaskStation reaches for code too, but it never compiles anything. There is no TaskStation language. The workflow is ordinary TypeScript and markdown sitting in a git repo — a `SKILL.md` that tells an agent how your company does a job, a script next to it for the logic that deserves to be pinned down, a `taskstation.yaml` that declares the connectors, triggers and policies. You read it with `cat`. You review it with `git diff`. You test the script without an agent anywhere near it. What TaskStation does own is narrower and, we would argue, the part that actually needs owning: the boundary where that code touches the outside world. Every connector call goes through one server-side chokepoint, the Connector gateway, and the typed client is part of [`@taskstation/sdk`](https://www.npmjs.com/package/@taskstation/sdk). ## What "verifiable" means here, concretely For Poetic, "verifiable" means the compiled procedure executes deterministically. For TaskStation it means something narrower and different: every action that leaves the sandbox is typed, risk-classified, policy-checked, optionally paused for a human, and written to an audit row — and none of that is enforced by the agent, so none of it can be talked out of by the agent. The gateway is the only path. Every action in the catalog carries a machine-assigned risk. For an HTTP-backed connector it is derived from the method — `GET`/`HEAD`/`OPTIONS` are `read`, `DELETE` is `destructive`, everything else is `write`. For an MCP connector the server’s own `readOnlyHint` and `destructiveHint` annotations are honoured. That classification is what your policy binds against, so a workflow can assert what it is about to do before it does it: ``` import { createTaskStation } from '@taskstation/sdk'; const taskstation = createTaskStation({ backendUrl: process.env.TASKSTATION_API_URL!, getToken: async () => process.env.TASKSTATION_TOKEN ?? null, }); const connectors = process.env.TASKSTATION_PROJECT_ID ? taskstation.project(process.env.TASKSTATION_PROJECT_ID).connectors : taskstation.connectors; // The catalog is the contract. Refuse to run if it drifted. const action = await connectors.describe('stripe.close_dispute'); if (action?.risk !== 'write') { throw new Error(`refusing to run: catalog says ${action?.risk ?? 'unknown'}`); } ``` And here is the shape Poetic targets — read, branch, act — written as a TaskStation skill script. Note what is not in it: no API key, no polling loop, no prompt. The branching is a plain `if`. The credential is resolved server-side and never enters the sandbox. A gated write returns an authenticated approval URL, ends the request, and resumes the TaskStation session through a durable callback after one human decision: ``` import { type ConnectorCallResult } from '@taskstation/sdk'; // `connectors` is the client created above. type Dispute = { id: string; amount_cents: number }; // 1. Read. risk: 'read' — the gateway never gates this. const open = await connectors.call<{ disputes: Dispute[] }>('stripe.list_disputes', { status: 'needs_response', limit: 50, }); if (!open.ok) throw new Error(`list_disputes failed: ${open.reason ?? open.status}`); for (const dispute of open.data?.disputes ?? []) { // 2. Branch. Ordinary TypeScript — diffable, and testable with no agent. if (dispute.amount_cents > 500_00) continue; // 3. Act. A gated write returns one approval handoff immediately. const result: ConnectorCallResult = await connectors.call('stripe.close_dispute', { dispute: dispute.id, }); if (result.status === 'pending_approval') { console.log(result.approval_url); break; // TaskStation resumes this session after the human approves or denies. } if (!result.ok) throw new Error(`close_dispute ${dispute.id}: ${result.reason}`); } ``` > Both snippets above type-check under `strict` against the published `@taskstation/sdk` source, and the package unit suite covers route selection, the call envelope, error mapping and catalog flattening. The approval handoff, risk classification and audit row are enforced in the gateway, not in this client — a script cannot opt out of them by not calling them. The honest scope note: the Connector client in `@taskstation/sdk` is a typed client for the action boundary. It is not a workflow compiler and it does not make your workflow deterministic. It makes the *edges* of your workflow legible. The determinism you get is the determinism of the TypeScript you wrote around it. That is a weaker guarantee than Poetic’s and a more general one. ## Now the part that is actually different: ownership Poetic’s public material describes the language, the learning loop and the accuracy. It does not describe where the compiled artifact lives, whether you can export it, whether you can run it without Poetic, or what happens to the workflow if you stop paying. Those may all have good answers — they are simply not published, and we are not going to invent them. What we can be specific about is our side. - **A project is a git repo.** Not a workspace in our cloud — a repository. `taskstation init` turns a directory into one; `taskstation ship` pushes it up and runs it. Clone it and you have the whole thing. - **`taskstation.yaml` is the manifest.** Connectors, triggers, channels, required secrets, policies and where agent config lives — one file, in your repo, in the diff. - **Agents and skills are files.** An agent is a markdown persona. A skill is a `SKILL.md` plus the scripts beside it. There is no console where the real definition secretly lives; the file *is* the definition, which is why an agent can propose an edit to its own configuration as a change request. - **Work lands through review.** A session runs on its own isolated cloud computer on its own branch. It reaches `main` only through a change request someone approves. - **You can self-host it.** TaskStation is open source. Run it on your own infrastructure with your own keys and your own models. This is not an air-gapped story — `taskstation self-host start` pulls images and reaches a sandbox provider over the network — but the data, the config and the model are yours. - **Any model.** Bring your own key, or the ChatGPT, Claude or Cursor subscription you already pay for. Concretely: if TaskStation disappeared tomorrow, your skills are still markdown, your logic is still TypeScript, your manifest is still YAML, and the repo still clones. The Connector gateway is the piece you would have to replace, and its client is a 209-line file whose surface is five methods. That asymmetry — a lot of durable artifact, a small replaceable runtime — is the entire ownership argument, and it is the reason we think it is worth stating plainly rather than dressing up. ## Where this comparison does not favour us - **No determinism claim.** Poetic compiles to deterministic execution. A TaskStation session is a model loop. On a single fixed process at extreme volume, that is their win, not ours. - **No published accuracy number.** Poetic publishes 99%+ on named process types with named enterprise customers. We publish no comparable figure, and we are not going to manufacture one. - **Per-run cost.** "Near-tokenless" execution beats a frontier model loop on a process that runs thousands of times a day. Bring-your-own-key narrows that gap; it does not close it. - **Someone has to write the script.** Poetic drafts the workflow from your recordings and documents, with their team alongside you. TaskStation expects you to be comfortable in a repo — and an agent will happily write the skill for you, but you still review it. - **They are further along on one axis.** A purpose-built language for regulated back-office process work is a deeper commitment to that problem than a general runtime will ever be. ## Side by side | Dimension | Poetic | TaskStation | | --- | --- | --- | | Turns workflows into code | Yes — a purpose-built language | Yes — ordinary TypeScript + markdown | | Deterministic re-execution | Yes — their core claim | No — a model loop plus typed calls | | Cost per repeat run | Near-tokenless after authoring | Model cost per run — any model, your keys | | Where the workflow lives | Poetic's platform | A git repo you own | | Readable as a diff | Not stated publicly | `git diff` — skills, agents, manifest | | Open source | No | Yes | | Self-hostable | Not stated publicly | Yes — your cloud, VPC, on-prem | | Choose your models | Vendor-managed | Any model — your keys or your subscription | | Human approval on risky actions | Expert feedback loop during authoring | Gateway pauses the call for a decision | | Audit trail on every action | Immutable audit logs | One audit row per gateway call | | Scope | Deep — regulated, high-volume process work | Broad — a workforce across the company | | How you start | Forward-deployed engagement | `taskstation init` — or self-host today | ## When to pick which ### Choose Poetic if you have one or a few high-stakes, high-volume, well-bounded processes — fraud investigation, KYC, dispute handling — where near-perfect repeatability is the requirement, and a forward-deployed vendor relationship is a feature rather than a risk. ### Choose TaskStation if you want the workflow to stay yours: agents, skills, policies and memory as files in [one repo you own](/blog/introducing-taskstation), any model, self-hostable, with a typed and audited boundary on every action agents take [across departments](/enterprise). These are not mutually exclusive, and pretending otherwise would be dishonest. A bank can reasonably run Poetic on dispute adjudication and TaskStation for everything else the company does — the two answer different questions. If the connector boundary is the part you care about, [the secure tool-access model](/blog/secure-ai-agent-tool-access) goes deeper on the gateway. If it is the architecture, [AGI-ready architecture](/blog/agi-ready-architecture) explains why state lives in files rather than a context window. The other two comparisons — [Claude Cowork](/blog/taskstation-vs-claude-cowork) and [Glean](/blog/taskstation-vs-glean) — cover the desktop and the search layer. ## Turn the workflow into code — and keep the code. Connect your tools and hand a TaskStation agent a real task. Free to start, free to self-host. --- # AGI-ready architecture: what it really means, and how TaskStation is built for it AGI-ready doesn't mean an architecture that produces AGI. It means one that absorbs a 100× capability jump without losing state or granting uncontrolled access. How TaskStation is built for it. Canonical page: https://taskstation.co/blog/agi-ready-architecture Published: 2026-07-17 Author: marko Tags: Architecture, Vision Most “AGI-ready” claims are a model wrapped in tools. Swap that model for something a hundred times more capable and the whole thing breaks — or worse, it works, and you’ve quietly handed an uncontrolled agent your production, your money, and your customers. AGI-ready doesn’t mean an architecture that produces AGI. It means the surrounding platform can absorb major jumps in capability without being rebuilt. The models will get better. That is the one safe bet in this entire field — more capable, more autonomous, multimodal, continuously active, able to run for days or months, able to control computers, infrastructure, and accounts, and eventually cheap enough to spawn thousands of concurrent workers. The question is not whether that happens. The question is whether your platform can take the jump without giving the model uncontrolled access, losing its state, or being rewritten. > AGI-ready architecture is a durable, stateful, model-independent, permissioned, observable, evaluable, self-improving execution system for intelligent entities — not merely an LLM wrapped in tools. ## The test that matters Here is the practical test: can you replace today’s model with something 100× more capable — more autonomous, multimodal, able to run for days, control computers and money, and fan out into thousands of workers — and keep the same identity boundaries, the same permissions, the same review gates, and the same state? If yes, the architecture is AGI-ready. If the model *is* the application, it isn’t. The model is a reasoning engine. The platform around it is the product. ## Models are replaceable — the reasoning engine is not the application An AGI-ready system depends on capabilities — reason, code, vision, verify — not on a model name. TaskStation treats models as hot-swappable reasoning engines behind a gateway. Bring any provider on your own keys, or run self-hosted inference on your own GPUs. Route a cheap open-weight model for the bulk of the work and a frontier model only where it earns its keep. - **Any model, your keys** — Claude, GPT, Gemini, or open-weight GLM and DeepSeek; your subscription, your spend, your data residency. - **A model-agnostic gateway** — per-project routing chains, ordered fallbacks, and semantic failover where an empty completion is classified as a failure, not a zero-output success. - **Self-hosted inference** — run it in your own VPC or on-prem, on your own hardware. The platform never assumes a vendor is reachable. When a better model ships tomorrow, you point the gateway at it. Nothing else moves. ## State lives outside the model This is the single most important invariant, and it is the one most “agent platforms” violate. If the model’s context window is where your state lives, you have no state — you have a lucky streak that ends when the session ends. In TaskStation, everything is files in a git repo. The manifest, the agents, the skills, the connectors, the policies, and the memory are all versioned, diffable, owned files. The model proposes; a trusted subsystem persists. ``` # taskstation.yaml — one file that defines this project. taskstation_version: 2 project: name: acme-security-audit # A tool the agent can use. Credentials stay in the platform, # never in this file. The policy decides what needs a human. connectors: - slug: slack policies: - match: "*message*" action: require_approval # Run work on a schedule — nobody has to kick it off. triggers: - slug: nightly-access-review type: cron cron: "0 0 2 * * *" prompt: Audit last night's access logs and flag any out-of-policy connector calls for review. ``` Crucially, nothing the model generates silently becomes authoritative business state. An agent can write code, draft a campaign, or move a file — but that change reaches the shared `main` only through a reviewed change request. The model’s output is a proposal until a human or a trusted gate says otherwise. ## Every action has an identity — and the model is never the authority on what it may do Authentication answers who the agent is. Authorization answers what it may do — and the model must never be the final authority on the second question. In TaskStation, every session runs under a scoped identity with a single token carrying claims for principal, project, session, and agent grant. Connector credentials are bound server-side and injected at runtime; they never enter the sandbox environment, the transcripts, or the model’s view. - **Least privilege per session** — an agent sees only the connectors and secrets it was granted, nothing more. - **Policy as code** — `connectors:` in the manifest carry per-action policies; matching an outbound message can require a human before it ever sends. - **One scoped token** — granular permissions stop an agent from touching tools it shouldn’t, the way a mature platform scopes API access rather than handing out a god key. A 100× more capable model does not get 100× more authority. It runs inside the same identity, the same scoped token, and the same policies it always did. ## Execution is isolated from the control plane An intelligent agent should never execute directly inside the control plane. In TaskStation, every session is its own isolated Linux sandbox on its own git branch — a disposable machine the agent owns, with filesystem, network, and process isolation. Thousands run in parallel on the same config without colliding, because none of them share state. Egress and credentials are controlled at the network boundary, and the sandbox assumes the agent-generated code is untrusted even when the model looks reliable. Work reaches `main` one way: through an approved change request. The sandbox is where the agent thinks and acts; `main` is where the company lives, and the two are deliberately not the same place. ## Long-running work is durable and resumable AGI-level work will not fit into a single HTTP request. An agent should be able to sleep for three months, wake because an event fired, reload its identity and state, and continue correctly. TaskStation sessions are durable: stop a running session and it pauses in place with its compute metering closed; resume it and it picks up where it left off. Retryable failures auto-resume a turn with bounded backoff; the user keeps control through a kill switch. Triggers declare their own session strategy — a fresh session each run, a reused loop session, or a pinned one — so a long-lived operator can carry context across invocations instead of restarting cold every time the clock ticks. ## Every consequential action is auditable and evaluable Git history is the audit trail; the change request is the review gate. Every PR spins up a full ephemeral environment — a real backend and a wired frontend — so a change is tested end to end, not just linted. The important metric is not “the model scored 90%.” It is: what percentage of real tasks finish correctly, safely, autonomously, within budget? That is a reliability question, and the platform is built to answer it with actual preview runs, not vibes. ## Learning is gated, versioned, and reversible An AGI-ready platform turns production experience into improvement — but never lets a model auto-edit its own behavior. In TaskStation, a skill is a file: purpose, preconditions, steps, policies, examples, tests, version, and provenance. An agent can propose a new skill, but it ships through a controlled lifecycle with review, not by mutating itself in production. Memory is layered — working, episodic, semantic, procedural — and a trusted subsystem decides what is persisted, with deduplication, provenance, and expiration. Learning a fact, remembering a preference, and changing an agent’s behavior are different operations with different gates. ## Humans can inspect, interrupt, and terminate Human-in-the-loop is a first-class runtime primitive, not a setting. Approval requests, escalation, plan editing, and emergency termination are all part of the loop. The interface shows what the agent intends to do, why it intends to do it, what it used to decide, what could go wrong, what it will cost, and whether the action is reversible — then asks for exactly the permission it needs. A stop button that actually stops is not a nice-to-have when the agent can spend money. ## Costs and resources are bounded A capable system without budget control can create effectively unlimited spend. TaskStation tracks token, compute, sandbox, and tool cost against per-run and per-objective budgets, with holds and reconciliation so a streamed completion is billed at the rate it actually used. Idle sandboxes are reaped — quiet machines are stopped, not billed forever. Fan-out into hundreds of subtasks stays a controlled, cost-bounded operation rather than an open tab. ## More capability, not more authority That is the whole thesis in one line. A 100× model should make TaskStation do 100× more work — it should not get 100× more access. Authority is capped by identity, policy, sandbox, and review, and none of those are controlled by the model. The capability jumps; the guardrails hold. That is what AGI-ready means in practice, and it is the difference between a platform that survives the next model and one that has to be rebuilt around it. ## Side by side | Dimension | An LLM wrapped in tools | TaskStation | | --- | --- | --- | | Where state lives | In the context window | Files in a git repo + layered memory | | Who authorizes actions | The model decides what it may do | Policy engine + scoped connectors | | Where execution happens | In your control plane | Isolated sandbox per session, per branch | | Resumability | Lost when the request ends | Durable; stop, resume, wake on events | | Review & rollback | No diff, no rollback | Every change a reviewed change request | | Models | Welded to one vendor | Any model, your keys, self-hostable | | Cost control | Unbounded by default | Per-run budgets + idle reaping | ## When to pick which ### Choose the wrapper if you are shipping a single model’s output straight to production with no durable state, no permission boundary, and no review — and betting that a smarter model makes that safe. ### Choose TaskStation if you want the capability jump without the authority jump — agents that do 100× more work inside the same identity, policy, sandbox, and review you already control. The companies that win the next decade of AI won’t be the ones with the best model. They’ll be the ones whose platform can take whatever model shows up next and put it to work safely. If that operating layer is what you’re missing, the [introduction](/blog/introducing-taskstation) and the [company-as-a-repo thesis](/blog/ai-transformation-company-os) are the next reads. ## Build for the jump, not the model. TaskStation is the Autonomous Company Operating System — open-source, self-hostable, any model. Start one project free. --- # TaskStation vs Glean: search or an agent platform that runs work? Glean is the best permission-aware enterprise search. But search finds work — it doesn't do it. Here's where you outgrow it, and the open runtime alternative. Canonical page: https://taskstation.co/blog/taskstation-vs-glean Published: 2026-07-13 Author: team Tags: Comparisons, Enterprise, Open Source Glean is genuinely the best permission-aware enterprise search you can buy. It indexes your apps, respects your ACLs, and answers in plain language with citations. So this isn't a "they're bad, we're good" post. The honest question is a different one: once you can find anything in your company, what actually does the work with it? Compared here: - Glean (glean.com) ## What Glean is great at - **Permission-aware search done right** — it inherits your source-system ACLs, so a result you can see is a result you can act on. - **Mature connectors** — it reaches across the usual enterprise stack and keeps the index fresh. - **Serious compliance posture** — built for the security review that enterprise search has to survive. - **A clean assistant on top of retrieval** — ask a question, get a cited answer instead of ten blue links. ## Where it stops: search finds work, it doesn’t do it Glean’s center of gravity is the index. Agents are a layer on top of retrieval, not a workforce that runs your company. The moment the job is “open the tickets, enrich the accounts, draft and send the outreach, land the fix, close the book” — search has stopped being the bottleneck and a chat assistant over the index isn’t the answer either. You need a runtime that hands a task to agents and they return finished work. - **Retrieval-first, agents bolted on.** The product answers “where is it?” well; it is not built to run a fleet of agents that take real actions across your tools. - **Closed and vendor-hosted.** You query Glean; you don’t own it. It is SaaS or vendor-managed cloud — your company’s knowledge leaves your walls to be indexed somewhere else. - **Seat-priced and sales-led.** Public reporting puts Glean at roughly [$50–75 per user/month with a ~100-seat minimum](https://www.gosearch.ai/faqs/glean-enterprise-search-pricing-explained-costs-tiers-hidden-fees-gosearch-comparison) — about a $60k/year floor before infrastructure and implementation. That locks out the small team and the single-department pilot. - **Configured in a console, not as code.** Connectors, assistants, and prompts live in a vendor dashboard. There is no diff to review, no version to roll back, no repo to fork. None of that is a flaw in a search product. It is exactly the line you cross when “let me find it” becomes “let something do it.” If you want the broader framing, [beyond the chat box](/blog/beyond-the-chat-box) makes the same argument against chat assistants: input→output stops short of work. ## A runtime that does the work, not just retrieves it TaskStation is an open agent runtime — the command center where a workforce of agents runs your company, not a search bar over it. Hand a task to a project and agents run in isolated sandboxes, take real actions through scoped connectors, and land durable change back to one shared `main` through a reviewed change request. The context they need is files in a repo you own, not an index someone else rents back to you. That is the real split. Glean makes your existing knowledge searchable; TaskStation makes your company’s operating layer — agents, skills, memory, connectors, policies — into [files in one repo](/blog/introducing-taskstation) that agents run against. One is a window onto work; the other is where the work happens. ## Own the data, pick the model, skip the seat tax Because TaskStation is open-source and self-hostable, your data never has to leave your walls — your cloud, your VPC, on-prem, or your own GPUs. And because you bring your own key and run any model, the bill is not bundled into a per-seat license. An open-weight model like **GLM-5.2** runs about **5–7× cheaper** than Claude Opus or GPT on output (~$4.40 vs $25–30 per 1M tokens), and **DeepSeek** is **50×+ cheaper** on output. Route a cheap model for the bulk of the work and a frontier model only where it earns its keep. > No 100-seat floor, no sales process to start — [see the plans](/pricing). Open-source means you can run one project today and a whole company on it tomorrow — on infrastructure where the data, config, and model belong to you. ## Side by side | Dimension | Glean | TaskStation | | --- | --- | --- | | Core job | Find & answer over company data | Build & run agents that do the work | | Runs a fleet of agents in parallel | Assistants bolted onto search | Thousands of agents, isolated sandboxes | | Self-hostable / own your data | No — SaaS or vendor-managed cloud | Yes — your cloud, VPC, on-prem | | Choose your models | Vendor-managed, bundled in seat | Any model — your keys | | Pricing model | ~$50–75/user/mo, ~100-seat min | Open-source; cloud or self-host, any size | | Accessible below 100 seats | No — sales-led, large-enterprise floor | Yes — start with one project | | Agents, skills & policies as code | Configured in a vendor console | Files in one repo you own | | Versioned, reviewable, roll-back-able | Console settings, no diff | Git history + change requests | | Multi-tenant governance | Enterprise permissions on search | Departments, roles, scoped connectors | ## When to pick which ### Choose Glean if you want the best permission-aware enterprise search and assistant, you’re fine with a closed SaaS and a sales-led ~100-seat contract, and “find the answer” is the job. ### Choose TaskStation if you want to run agents that actually do the work — [across departments](/enterprise), any model, self-hosted, with everything versioned and owned by you. They can coexist, too. Plenty of companies will keep Glean as the search layer and run the work itself on TaskStation — agents that read, decide, and act, with the operating layer they need to do it governed as code. (The desktop side has its own parallel: [how TaskStation compares to Claude Cowork](/blog/taskstation-vs-claude-cowork).) If that operating layer is what you’re missing, the [company OS post](/blog/ai-transformation-company-os) and the [secure connector model](/blog/secure-ai-agent-tool-access) are the next reads. ## Don't just find the work. Run it. Connect your tools and hand a TaskStation agent a real task. Free to start, free to self-host. --- # How to give AI agents tool access safely How to give AI agents production tool access without raw API keys: scoped connectors, approval policies, server-side credentials, and reviewed work. Canonical page: https://taskstation.co/blog/secure-ai-agent-tool-access Published: 2026-07-07 Author: team Tags: Security, Connectors, Enterprise The moment an AI agent can use tools, it stops being a chat feature and becomes production infrastructure. It can read customer records, draft emails, open pull requests, query billing, post in Slack, or touch an internal API. At that point the hard question is not “can the model call the tool?” It is **who gave it access, how narrow is that access, what happens before a risky action runs, and what audit trail remains afterward?** TaskStation was built around that boundary. Tool access does not belong in a prompt and raw credentials do not belong in an agent sandbox. In TaskStation, connections are part of the project operating layer: declared as files, brokered server-side, granted per agent, governed by policy, and reviewed when durable work changes the company. If you want the larger architecture first, read [Introducing TaskStation](/blog/introducing-taskstation) or the [company OS post](/blog/ai-transformation-company-os). The rest of the market is converging on the same lesson. [Auth0](https://auth0.com/blog/api-key-security-for-ai-agents) calls out over-privileged tokens, prompt-injection exposure, and missing audit trails as common risks when teams hand API keys to agents. [WorkOS](https://workos.com/blog/ai-agent-credentials) argues agents need their own scoped, revocable credentials instead of borrowing a user’s full session. [Promptfoo’s OWASP Agentic AI summary](https://www.promptfoo.dev/docs/red-team/owasp-agentic-ai) lists Tool Misuse and Identity and Privilege Abuse as core agentic risks. The pattern is clear: agent security is mostly tool security. ## Chat is harmless until it touches systems A model drafting text in a window has a small blast radius. A model with connected tools has the blast radius of those tools. That is not a reason to keep agents powerless; powerless agents do not run companies. It is a reason to treat the connector layer as seriously as you treat IAM, secrets, and production deploys. - **A support agent** may need to read tickets and invoices, but should not be able to refund money without approval. - **A finance agent** may need to pull Stripe, bank, and warehouse data, but should not be able to send vendor payments from the same path. - **A recruiting agent** may need to enrich candidates and draft outreach, but should not send messages without a human approving the final copy. - **An engineering agent** may need GitHub, Linear, CI, and preview access, but should land work through a reviewed change request instead of mutating main directly. > The control plane cannot be “the prompt told the agent to be careful.” The control plane has to be outside the model. ## The five rules of safe tool access A production agent platform needs five layers before you can comfortably connect real company systems: - **Keep credentials out of the sandbox.** The agent should never receive a third-party API key unless the task truly requires direct process-level access. Connector credentials should be resolved server-side and injected into the upstream request, not into model context. - **Grant tools per agent.** Connecting Slack, Gmail, Stripe, or GitHub to a project is not the same as letting every agent call it. The support agent and release agent need different reach. - **Gate individual actions.** Read operations, write operations, deletes, sends, payments, and admin changes should not share one permission bit. Tool names need policy: always run, require approval, or block. - **Make risky calls human-reviewable.** A good agent can prepare the exact action and evidence. The platform should pause at the boundary where a human decision is required. - **Route durable change through review.** If the agent edits the operating layer — agents, skills, triggers, memory, policies, or code — that work should be a diff someone can review, merge, and roll back. ## How TaskStation models a connector TaskStation connections are documented in [Connecting your tools](/docs/guides/connecting-tools). A connector can be a one-click Pipedream app, a remote MCP server, an OpenAPI or GraphQL API, a raw HTTP API, a channel such as Slack, or a connected computer. The definition lives with the project; the credential lives on the platform. The agent sees a tool catalog, not a pile of secrets. ``` connectors: - slug: stripe provider: openapi spec: https://raw.githubusercontent.com/stripe/openapi/master/openapi/spec3.json policies: - match: "*.get*" action: always_run - match: "*.create*" action: require_approval - match: "*.delete*" action: block agents: support: connectors: [plain, stripe] secrets: none taskstation_cli: none release-bot: connectors: [github, vercel] taskstation_cli: [project.cr.open] ``` That example is deliberately boring. Boring is the point. You should be able to answer “what can this agent call?” by reading the project files, not by reverse-engineering a prompt or inspecting a live process. The [manifest reference](/docs/reference/manifest#connectors) defines connector policies and the [agent governance section](/docs/reference/manifest#agents-v2) defines per-agent grants. ## Server-side credentials change the failure mode When credentials sit in environment variables inside the agent runtime, every prompt-injection bug, logging bug, file-read bug, and subprocess bug becomes a possible credential leak. When credentials are brokered server-side, the agent can ask to call a tool, but the platform decides whether the call is allowed, resolves the credential, executes the upstream request, and records what happened. That is the model behind the TaskStation Connector. Every session gets a scoped Connector token. The agent discovers tools, describes their schemas, and calls them through the TaskStation API. The gateway enforces the project grant and connector policy, resolves credentials outside the sandbox, runs the request, and audits the call. The [connections guide](/docs/guides/connecting-tools) is explicit: the agent never holds third-party credentials. > A scoped tool token is not just safer than a raw API key. It also makes the audit trail meaningful: agent identity, tool name, input boundary, policy decision, approval state, and upstream result can all be tied together. ## The dangerous pattern to delete The common early pattern is understandable: put `STRIPE_SECRET_KEY`, `GITHUB_TOKEN`, `SLACK_BOT_TOKEN`, and a dozen other keys into `.env`, start the agent, and hope the prompt keeps it in bounds. That works for a demo. It is the wrong shape for a company. - **It is too broad.** The key usually carries every permission the connection owner had, not the minimum action the agent needs. - **It is hard to attribute.** Downstream systems see the shared key, not the agent, session, person, or approval that caused the call. - **It is hard to revoke safely.** Rotating a shared key breaks every workflow using it; leaving it in place keeps the blast radius large. - **It hides policy in code and prompts.** Security reviewers need declarative grants and logs, not “the agent instructions say don’t delete things.” ## A quick audit for your agent stack Before you connect agents to production systems, ask these questions: - Can I list every external system this agent can reach without opening the agent prompt? - Can I give a sales agent CRM read access without also giving it billing write access? - Can I block deletes, require approval for sends, and allow safe reads on the same connector? - Can I see which person, agent, session, and policy decision caused a tool call? - Can I revoke one agent’s reach without rotating a shared key that breaks other workflows? - Can the operating layer move from cloud to VPC or on-prem without rewriting the tool model? If the answer is no, you may still have a useful agent prototype. You do not yet have a secure AI command center. ## Why this is a company OS problem Safe tool access is not a standalone feature. It only works when it sits beside the rest of the company operating layer: memory, agents, skills, triggers, secrets, policies, sandboxes, and change requests. The connector grant says what the agent may touch. The sandbox limits where it runs. The policy gate decides which calls need approval. The change request records durable changes as a diff. The repo keeps the whole thing owned and reviewable. That is why TaskStation frames the product as an Autonomous Company Operating System, not another assistant with more connectors. A company does not need one more place to paste keys. It needs a Git-backed AI command center where the tools, credentials, policies, and agent work are part of the same owned system. ## Connect the tools, keep the keys out of the agent. Start with one workflow, grant only the connectors it needs, gate risky actions, and run the work from a repo your company owns. --- # AI transformation needs a company OS Why consultancies and AI-transformation teams need one Git-backed workspace for agents, memory, connectors, policy, and auditable work. Canonical page: https://taskstation.co/blog/ai-transformation-company-os Published: 2026-06-29 Author: team Tags: Enterprise, AI Transformation, Company OS AI transformation is past the demo phase. The hard part now is not proving that an agent can draft a report, inspect a spreadsheet, or update a CRM record. The hard part is giving every client, department, and delivery team a **repeatable operating layer** where agents, context, connectors, policy, and review live together. That is what a company OS is for. TaskStation is the **Autonomous Company Operating System**: an AI command center where a workforce of agents does real work, and everything that defines the system is files in one Git repo you own. For consultancies and AI-transformation teams, that matters because the deliverable is no longer a single chatbot. The deliverable is a governed workspace the client can keep running after the pilot. If you want the full product spine first, read [Introducing TaskStation](/blog/introducing-taskstation). The market is already pointing this way. [Accenture AI Refinery](https://www.accenture.com/us-en/services/ai-data/ai-refinery) frames enterprise AI around agents, knowledge, models, and governance. [Deloitte](https://www.deloitte.com/in/en/services/consulting/services/engineering-ai-data/agentic-ai.html) describes multiagent systems that understand requests, plan workflows, coordinate role-specific agents, collaborate with humans, and validate outputs. The missing question is where all of that lives so it can be owned, reviewed, repeated, and ported into the tools people already use. ## The pilot is not the product Most AI-transformation work starts with a useful prototype: a support agent, a sales-research assistant, a finance close helper, a legal intake workflow, a marketing campaign planner. The prototype proves demand. Then the real work starts. - **Who owns the instructions?** If the prompt lives in one vendor dashboard, the client cannot audit or improve it like normal operational IP. - **Where does the context accumulate?** If every tool stores a different slice of memory, the organization never gets one shared brain. - **How are tools governed?** Reading a CRM, sending an email, querying Stripe, and posting in Slack should not have the same permission profile. - **How does the work become official?** A finished deliverable needs review, history, rollback, and a clear path into the client’s source of truth. - **How do you repeat it for the next department?** The second workspace should be a fork, not a rebuild. A proof of concept can avoid those questions. A production AI-transformation program cannot. The operating layer becomes the product because it decides whether the client gets a one-off demo or a system that keeps improving. ## One client, one repo In TaskStation, a project is a repo. That repo contains the company’s agents, skills, memory, triggers, connector policy, sandbox definition, and operating instructions. One `taskstation.yaml` defines how the workspace runs. Every session happens on an isolated branch. Every persistent change comes back through a change request. ``` acme-ai-workspace/ ├─ taskstation.yaml # project, sandboxes, triggers, connectors, policy ├─ .taskstation/opencode/ │ ├─ agents/ # role-specific agents: finance, support, sales, legal │ ├─ skills/ # repeatable client playbooks and workflows │ └─ commands/ # approved operating motions ├─ memory/ # durable company context and decisions ├─ artifacts/ # reports, briefs, packets, launch plans └─ docs/ # source-of-truth operating docs ``` That sounds technical because it is. It is also the reason the workspace can be handed to a client without trapping them in your service team forever. Files can be inspected. Diffs can be reviewed. A successful sales-ops workspace can be forked into a recruiting workspace. A regulated client can run the same pattern in their own VPC or on-prem environment. The [docs](/docs) walk through the project, session, and change request model in detail. ## The workspace needs five layers If you are leading AI transformation for a client, a serious agent workspace needs more than a chat UI. It needs at least five layers working together: - **Context.** The policies, playbooks, decisions, customer notes, docs, and memory the agents need to act like part of the company. - **Agents and skills.** Named roles and reusable workflows, not one giant prompt that tries to do everything. - **Connectors.** Access to the real systems of work — Slack, Gmail, HubSpot, Stripe, Linear, Notion, warehouses, internal APIs — brokered through scoped credentials instead of pasted keys. - **Policy.** Tool-level allow, ask, and block rules so a workspace can automate research freely and still pause before it sends, pays, deletes, or posts. - **Review.** A change request path for durable work: what changed, who requested it, what the agent touched, and what a human approved. > The unit of delivery is not “an agent.” The unit of delivery is a governed workspace where many agents can do real work safely. ## Governance belongs in the runtime Enterprise buyers do not just ask whether the model is good. They ask where secrets live, how access is scoped, what gets logged, how approvals work, how quickly a bad change can be reverted, and whether the system can run under their infrastructure constraints. TaskStation was built around those constraints. Sessions run in disposable Linux sandboxes on their own branches. Connectors are brokered server-side through one scoped token. Secrets are encrypted and injected at runtime, not shown to the model. Work reaches `main` only through reviewed change requests. The same workspace can be used from the web, Slack, Teams, CLI, API, and MCP surfaces instead of forcing every employee into a new destination app. That is the difference between “we connected an LLM to your tools” and “we gave your organization a controlled workforce.” The first is exciting in a workshop. The second survives procurement, security review, and the third month of production use. ## Why consultancies feel this first Consultancies and systems connectors are where the repeatability pressure shows up fastest. They do not need one beautiful demo. They need a way to deploy the same architecture across many clients, many departments, and many compliance profiles without rebuilding the plumbing every time. - **For the AI-transformation partner:** one horizontal platform can become the delivery substrate for many vertical offerings. - **For the client CTO:** the workspace is Git-backed, self-hostable, and inspectable instead of a vendor-owned service wrapper. - **For the delivery team:** each department gets its own agents, memory, connectors, and policies without losing the shared pattern. - **For the end user:** the agent shows up where they already work — Slack, Teams, web, CLI, API — instead of asking the 99% of employees to adopt another AI portal. This is also where open matters. A consultancy cannot credibly tell a bank, manufacturer, or healthcare company that their future operating layer is a closed prompt stack nobody can inspect. The closer agents get to real work, the more the client needs to own the substrate. That is why TaskStation is open, self-hostable, and built for enterprise deployment from the start. ## What to build first The best first workspace is narrow enough to ship and important enough to prove the operating model. Pick one workflow where the client already has documents, tools, approvals, and recurring pain. Then encode it as files. - **Sales renewal workspace:** read CRM context, summarize account risk, draft renewal plans, open human-reviewed follow-ups. - **Support triage workspace:** monitor tickets, classify urgency, draft replies from docs, escalate edge cases with evidence. - **Finance close workspace:** pull reconciliations, produce variance notes, flag missing evidence, create the close packet for review. - **Recruiting workspace:** source candidates, enrich profiles, draft Marko-style outreach, log every touch, never send without approval. - **Engineering review workspace:** review PRs, run checks, verify previews, and return concrete blockers instead of vague comments. Those are not abstract use cases for us. TaskStation runs internal sweeps for production errors, PR review, docs maintenance, weekly briefs, outbound research, and this SEO/blog loop from the same project-native model: agents with skills, memory, tools, triggers, and a reviewed path for durable changes. ## A quick test for your stack Before you choose an AI-transformation platform, ask five questions: - Can the client clone or export the actual operating layer — agents, skills, memory, policy, and triggers — as files? - Can two hundred agents run in parallel without sharing one fragile machine or one user’s desktop state? - Can tool access be scoped per person, group, agent, and action? - Can a security reviewer see what happened after the fact: prompts, tool calls, commits, approvals, and diffs? - Can the same workspace move from cloud to VPC to on-prem without changing the basic model? If the answer is no, you may still have a good agent demo. You do not yet have a company OS. ## Build the client workspace as files, then run it with agents. Start with one department, connect the tools it already uses, and turn the workflow into a Git-backed AI command center the client can own. --- # TaskStation vs Claude Cowork: a desktop assistant, or a company-wide agent platform? Claude Cowork is the best agent on the desktop. But it runs one assistant per person, on Anthropic's models, with your data on their cloud. Here's where you outgrow it — and what an open, company-wide agent platform looks like. Canonical page: https://taskstation.co/blog/taskstation-vs-claude-cowork Published: 2026-06-29 Author: marko Tags: Comparisons, Agents Claude Cowork is, hands down, one of the best agents you can put on a desktop today. It inherits Claude Code's engine, genuinely does multi-step work across your files and apps, and has a clean approval model. So this isn't a "they're bad, we're good" post. The honest question is what happens when one person's desktop assistant has to become a whole company's way of working. Compared here: - Claude Cowork (anthropic.com) ## What Claude Cowork is great at - **It does the work, not just the talking** — give it a goal and it returns a finished deliverable. - **A real permission model** — it shows a plan and waits for approval on consequential actions. - **Extensible with plugins** — teach it how you like work done and expose your own tools. ## Where it stops: one assistant, one machine, one lab - **One assistant per person** — not a fleet of agents running long jobs in parallel for the org. - **Nothing is shared.** Each person’s agents, skills, and context live on their own desktop — what one person teaches, the company never gets. - **Locked to Anthropic’s models** — no bring-your-own-key, so you pay frontier prices and can’t pick a cheaper model. - **Closed and vendor-hosted** — you can’t self-host it, and your data flows to Anthropic’s cloud. None of these are flaws for the product Cowork is. They’re exactly why it’s so good for one person — and exactly what you outgrow when "an agent on my laptop" needs to become "agents running our company." ## Shared across the company vs. siloed on a desktop It’s not just the machine that’s landlocked — it’s the knowledge. In Cowork, each person’s agents, skills, and context stay on their own desktop. In TaskStation, your agents, skills, and memory are **files in one shared repo**: what one person teaches, every teammate — and every agent — gets, and it compounds over time instead of resetting person by person. ## The model lock-in tax Cowork only runs on Anthropic’s models, at Anthropic’s prices. TaskStation lets you **bring your own key and run any model** — and the savings aren’t small. An open-weight model like **GLM-5.2** runs about **5–7× cheaper** than Claude Opus or GPT on output ($4.40 vs $25–30 per 1M tokens), and models like **DeepSeek** are **50×+ cheaper** on output. Route a cheap model for the bulk of the work and a frontier model only where it earns its keep. > Same agents, [a fraction of the bill](/pricing) — and you can run them on your own infrastructure, even your own GPUs, with your data never leaving your walls. ## Side by side | Dimension | Claude Cowork | TaskStation | | --- | --- | --- | | Does real, multi-step work | Yes — on your desktop | Yes — in the cloud, at scale | | Runs a fleet of agents in parallel | One assistant per person | Thousands of agents in parallel | | Choose your models | Anthropic (Claude) only | Any model — your keys | | Model cost (per 1M output) | ~$25–30 — frontier only | ~$4.40 (GLM-5.2) to ~$0.28 (DeepSeek) | | Agents, skills & memory shared org-wide | Siloed on each desktop | Shared in one repo | | Open-source & self-hostable | No — closed, via Anthropic | Yes — your cloud, VPC, on-prem | | Your data stays with you | Processed by Anthropic's cloud | On your own infrastructure | | Multi-tenant — departments, roles | A per-user desktop app | Multi-tenant by default | | Everything as versioned code | Plugins customize one assistant | Agents, skills & policies as files | ## When to pick which ### Choose Claude Cowork if you want a brilliant agent on one person’s desktop, you’re happy on Anthropic’s models, and you don’t need to self-host or run a fleet. ### Choose TaskStation if you want that same do-the-work power as a company platform — many agents [across departments](/enterprise), any model, self-hosted, with everything versioned and owned by you. They can even coexist: a power user keeps Cowork on their desktop while the company runs its shared, governed workforce on TaskStation. ## Love agents that do the work? Run a whole fleet — on your own terms. Connect your tools and hand a TaskStation agent a real task. Free to start, free to self-host. --- # Personal AI agents vs a company OS: TaskStation, OpenClaw, and Hermes OpenClaw and Hermes are brilliant open-source personal agents — and we genuinely recommend them for individuals. But a personal "Jarvis" and a governed company platform are different things. Here is exactly where the line is. Canonical page: https://taskstation.co/blog/personal-ai-agents-vs-company-os Published: 2026-06-28 Author: team Tags: Comparisons, Open Source If you’ve spent time in open-source AI lately, you’ve met **OpenClaw** and **Hermes**. Both are excellent: open-source, self-hosted, bring-your-own-model, living in the chat apps you already use. For an individual who wants a private, always-on agent on their own machine, they’re a joy — we mean that as a compliment. Compared here: - OpenClaw (github.com) - Hermes (nousresearch.com) They share TaskStation’s core values: open, self-hosted, your models, your data. So why build TaskStation? Because a **personal agent** and a **company operating system** are different problems — and stretching one into the other is where it gets painful. ## Single-operator is a design choice, not a gap - **OpenClaw** is explicit that it’s a personal assistant, not a shared multi-tenant system — and by default its tools run with broad access to the host machine. Fine on *your* laptop; a serious problem the moment several employees can steer a tool-enabled agent. - **Hermes** is a beautiful "agent that grows with you" — but team roles, tenant isolation, and org-wide audit aren’t what it’s documented for. You’d assemble that yourself. Neither is wrong. They optimized for the person. A company has to optimize for **many people, least privilege, and accountability** — and that changes the architecture from the ground up. ## Side by side | Dimension | OpenClaw / Hermes | TaskStation | | --- | --- | --- | | Open-source & self-hosted | Yes — MIT, bring your own model | Yes — any model, your keys | | Designed for | One operator (personal use) | Teams and companies | | Multi-tenant — departments, roles | Single operator | Multi-tenant by default | | Scoped policies per connector | Largely DIY; broad access | Allow / ask / block per tool, as code | | Isolated sandbox per task | Optional / personal | One isolated machine per session, on its own branch | | Versioned, auditable, reversible | Limited | Git-backed — full history | ## When to pick which ### Choose OpenClaw or Hermes if you want a private, always-on agent for *yourself*, on your own machine. ### Choose TaskStation if you want agents running across a *team or company* — with scoped control, isolation, roles, and audit — without giving up open-source and self-hosting. ## Love a great open-source agent? Get one built for your whole company. Same freedom, built for more than one person. Free to start, free to self-host. --- # Beyond the chat box: why ChatGPT, Claude, and Grok aren't an AI workforce Chat assistants answer; a workforce does the work. Why input-output tools — however brilliant — aren’t the same as a fleet of agents that run your company, own the data, and run on any model. Canonical page: https://taskstation.co/blog/beyond-the-chat-box Published: 2026-06-27 Author: team Tags: Comparisons, Vision ChatGPT, Claude, and Grok are extraordinary, and you should keep using them. But it’s worth being precise about what they are: **chat assistants.** You give an input, you get an output, and the moment you close the tab, the work is yours to carry out. That’s a faster way to *think*. It isn’t a company *running* on AI. Compared here: - ChatGPT (openai.com) - Claude (anthropic.com) - Grok (x.ai) ## Input → output vs. hand-off → finished work With a chat assistant, you’re the runtime: you ask, it answers, and you copy-paste between the chat window and your real tools to get anything done. With TaskStation, you hand off a task and an agent **goes and does it** — 30+ minutes of real, multi-step work across your connected tools, with full context on your company, returning a finished deliverable for review. ## The differences that matter at company scale | Dimension | Chat assistants | TaskStation | | --- | --- | --- | | Finishes multi-step work end to end | Mostly answers; agent modes are supervised | Agents act across your tools, end to end | | Runs a fleet in parallel | One supervised session | Thousands of isolated agents at once | | Choose your models | Locked to the vendor's models | Any model — your keys | | Run cheaper models | Pay the vendor’s frontier price | GLM-5.2 ~5–7× cheaper; DeepSeek ~50×+ | | Own your data / self-host | On the vendor's cloud | Open-source — your infrastructure | | Company-wide memory | Per-user chat history | A shared, Git-backed brain | | No lock-in | Tied to one vendor's platform | Files in a repo you own | ## What chat assistants are genuinely great at The point isn’t that chat assistants are bad. They’re excellent at what they’re built for, and they belong in the stack. A TaskStation agent that manages a vendor risk review might start by asking a chat assistant to digest a SOC 2 report — then take that output and run the full workflow. The key is knowing which tool fits which job: - **Quick answers and drafting.** Need a one-paragraph summary of a policy doc, or a first draft of a customer email? A chat assistant is faster than opening a ticket for an agent. - **Thinking out loud.** Exploring a problem, iterating on a prompt, or testing a hypothesis — the chat interface is the fastest way to refine an idea before handing it to an agent to execute. - **Code completion in-IDE.** Tools like Claude Code and Cursor are brilliant at diffing, refactoring, and writing code in your editor. TaskStation agents orchestrate those same tools at scale. - **Single-shot research.** "What’s the latest pricing for these three providers?" or "Summarize the Q2 trends." A chat assistant handles that in seconds — and an agent can then take the result and file it, notify stakeholders, and trigger the next step. ## They’re complementary, not interchangeable This isn’t "stop using ChatGPT." Use a chat assistant for quick answers, drafting, and thinking out loud. Use TaskStation for the work that has to actually get done — repeatedly, across your tools, owned by you, running while you sleep. One is a brilliant place to ask. The other is where your company’s work runs. ## Go from asking questions to running the work. Hand a TaskStation agent a real task and get a finished result back. Free to start, free to self-host. --- # Introducing TaskStation: the AI command center for your company A workforce of AI agents that do real work across your tools — defined as files in a git repo, run in isolated sandboxes, governed by review, and built enterprise-first. Here is the whole thing, A to Z. Canonical page: https://taskstation.co/blog/introducing-taskstation Published: 2026-06-06 Author: marko Tags: Product, Vision Every company is being told to "adopt AI." But most AI tools stop at the conversation. You ask a question, you get an answer, and the moment you close the tab the work is gone. That is a faster way to think. It is not a company running on AI. TaskStation is the **command center for the AI agents that do your work** — one place to build a workforce of agents, connect them to your tools, run them on your terms, and keep every result accountable to a human. Underneath that is an idea we think is right for the next decade of software: **your company's AI operation should be files in a git repo.** Not a pile of settings in someone else's dashboard — actual files you own, version, review, and run. ## A company is a git repo In TaskStation, a **project** is one git repository. The repo *is* the project: its files, its history, its agents, its automations, its settings — all of it lives in git. Start fresh with a private repo TaskStation hosts for you, or bring an existing one on GitHub. - **Every change is reviewable.** A new automation, a tweak to an agent, a newly connected tool — each is a diff someone can read and approve before it goes live. - **Nothing drifts.** There is no separate database of settings to fall out of sync with reality. The repo is the truth. - **It is portable and yours.** Your whole setup is plain files. Read it, fork it, move it, run it on your own infrastructure. > A company that runs on AI shouldn’t be a dashboard you rent and can’t inspect. It should be a codebase you own. ## taskstation.yaml: the single source of truth At the root of every project sits one file: `taskstation.yaml`. Any repo with a valid manifest at its root *is* a TaskStation project — that file defines what the project is, what it’s allowed to do, and how it runs. Here’s a real one: ``` # taskstation.yaml — the one file that defines this project. taskstation_version: 2 project: name: acme-ops description: Acme's operations command center. # Secrets your agents need: names here, encrypted values in the vault. env: required: [DATABASE_URL] optional: [STRIPE_API_KEY] # The sandbox every task boots into — your image, your hardware. sandbox: templates: - slug: ops dockerfile: .taskstation/Dockerfile cpu: 4 memory: 8 # Run work on a schedule — nobody has to kick it off. triggers: - slug: weekly-health-report type: cron cron: "0 0 9 * * 1" prompt: Draft the weekly customer health report for review. # A tool the agent can use — credentials stay in the platform, never here. connectors: - slug: slack policies: - match: "*message*" action: require_approval ``` That’s a company’s operating setup in a few dozen lines. The scheduler reads `triggers:`, the sandbox builder reads `sandbox.templates:`, the connector layer reads `connectors:`. Edit it in the dashboard or from inside a session and changes round-trip through the same file — the diff stays clean either way. ## What happens when you hand off a task Day to day, you describe a task in plain language and get a finished result back. Here’s everything that happens underneath, from the moment you hit go: - **A branch is cut.** The control plane opens a **session** and cuts a fresh branch from main. Your main line is never touched directly. - **The sandbox boots.** An isolated sandbox comes up from a content-addressed snapshot of your image, clones the repo, and pulls git credentials on demand — no long-lived token sits in the environment. - **The agent works.** It reads and writes files, reaches your connected tools, and commits progress to the session branch. - **It proposes the work.** When done, the agent opens a **change request** — a summary plus the exact diff — and hands it to you. It does not merge its own work. The sandbox is disposable by design. When the session ends, the environment is thrown away — only committed, merged work survives. Because each session is fully isolated, any number can run at once: yours, your teammates’, and your automated ones, none stepping on each other. ## Review is the only way in The change request is the heart of the trust model. It’s the **only** path for a session’s work to reach your main line — for *everything* the agent touched: new code, a new skill, an edited automation, a change to the agent’s own instructions. You see the exact diff, with conflicts flagged up front. Until you merge, the work is proposed, not applied. > An agent can have real autonomy inside its sandbox while having zero ability to change your company without a human saying yes. That’s the combination that makes handing agents real work sane. ## Tools without handing over the keys TaskStation connects your agents to the apps your team already uses — Slack, Gmail, Notion, Salesforce, and thousands more. When an agent uses a connected tool, **it never holds your credentials.** Each call is brokered server-side: the platform resolves the credential, runs the call, records it, and returns the result. The key never enters the sandbox. And you govern every action with policy — each tool can **run**, **require approval**, or be **blocked**, matched by name, so you can let an agent read freely and pause it before anything sends, posts, or pays. Every call is audited. ## Self-hostable, open, and yours When AI becomes how your company gets work done, the system running it stops being a tool and becomes infrastructure. Infrastructure you don’t own can be changed, repriced, or switched off without your say. So TaskStation is **open-source and self-hostable**, and you can run the entire stack on your own infrastructure — one command brings up a production-style TaskStation on your own machines, and the same CLI switches between our cloud and yours. Because it’s all open, you can read exactly how isolation, review, and credential brokering work — not trust a description. No lock-in: your projects are git repos, your config is plain files, and the platform running them is yours to host. (If you’re weighing TaskStation against personal open-source agents like OpenClaw or Hermes, [personal AI agents vs a company OS](/blog/personal-ai-agents-vs-company-os) draws that line.) ## It compounds Because your whole setup is version-controlled files, none of it resets tomorrow. Every agent you shape, every skill you teach, every tool you connect, every bit of memory your agents carry forward accumulates in the repo and gets more capable week over week. The routine work that used to fill calendars runs quietly in the background, 24/7, and your team spends its time on the decisions that need a human. For the architecture that makes that compounding safe — state kept out of the model, identity and isolation as first-class, durable review — the [AGI-ready architecture post](/blog/agi-ready-architecture) is the deeper read. ## Open the command center and hand an agent a real task. Connect your first tool and watch it come back with something you can use. Free to start, free to self-host. --- # Accounts & access An account holds your projects, your teammates, and one access model. Canonical page: https://taskstation.co/docs/accounts An account holds your [projects](/docs/project) and the people who work in them. When you sign up, TaskStation creates a personal account for you. Invite a teammate and it becomes a team account. Personal and team accounts use the same roles, billing, and limits. ## One access model TaskStation has one grant record: an **assignment**. Every assignment binds one principal to one role, at one scope. | Term | Meaning | |---|---| | **Principal** | Who holds the access. A `user`, a `group`, a `service_account` (an agent's identity), or a `pending` invitee email. | | **Role** | A named set of permissions. TaskStation ships the built-in roles below. An account can add custom roles. | | **Permission** | One action, for example `project.secret.read` or `member.update`. | | **Scope** | Where the role applies: the whole `account`, or one `project`. | | **Object** | Optional. Narrows the assignment to one `agent`, `skill`, `secret`, `app`, or `trigger` inside the scope. | | **Expiry** | Optional. The assignment stops granting at `expires_at`. | There is no second grant store. A group's access, a per-resource grant, and a custom-role binding are all assignments. They differ only in the principal, the object, and the role. ## Account roles Each person in an account holds one account role: - **Owner** — full control, including members and billing. - **Admin** — manages projects, members, groups, roles, and tokens. - **Member** — works only in the projects they hold a project assignment on. Owner and admin hold manager-equivalent access on every project in the account. A project assignment cannot lower that. To limit an owner or an admin on one project, first change their account role to member. Invite a teammate by email from the account's members page. The invitation is an assignment on a `pending` principal. It becomes a `user` assignment when they accept. ## Project roles Inside a project, a principal holds one of two roles: - **Member** — reads the project, starts and stops its sessions, fires its triggers. No editing, no configuration. - **Manager** — every project permission: edits the project, manages triggers, connectors, skills, secrets and access, holds the gateway keys, and deletes the project. There is no third project role. Manager is the full set of project permissions. A custom role only adds permissions, so no role can withhold one from a manager. [Object assignments](#object-assignments) are the separate mechanism that narrows *which objects* a permission reaches. Set project access from the project's access settings, or with `taskstation access` — see [CLI](/docs/cli#access). ## Groups A group is a principal, exactly like a person. Assign a role to a group at a scope and every member of that group holds it. A group adds access. It never removes access a person already holds. A group provisioned by SCIM is the same principal type as one you create by hand. ## Object assignments An assignment can name one object inside a project: an `agent`, a `skill`, a `secret`, an `app`, or a `trigger`. TaskStation enforces object assignments on agents and skills today. - **Agents are closed by default.** A member reaches an agent only when an assignment names them, or names one of their groups. - **Every other object type is open by default.** With no assignment on the object, a member reaches it. - **An object assignment restricts a project manager too.** That is what makes "scope this agent to the finance group" mean anything. What the manager role buys is the unscoped default, not an exemption. - **Account owners and admins are never restricted by an object assignment**, and neither is a service account acting on its own assignments. An object assignment carries no permissions of its own. It answers "which objects", not "which actions". ## Expiry Give an assignment an `expires_at` and it stops granting at that instant. The engine ignores an expired assignment. The row stays for the audit trail. ## One vocabulary, two bindings A person, a group, or an agent gets access from TaskStation as **roles** — the assignments described above. An agent carries a **second, separate binding**: the TaskStation CLI scopes its manifest declares in `taskstation.yaml` under `agents..taskstation_cli`. A session can only do what both allow. The two never widen each other. An agent whose manifest lists `project.secret.read` still reads nothing when the role verdict denies it, and an agent launched by an owner still reads nothing when its manifest does not list the action. Roles are account state; CLI scopes are repository state in the manifest. See [Manifest reference](/docs/project/manifest#agents). ## Custom roles A custom role is an account-owned role with the permissions you choose. Assigning one writes a single assignment — there is no built-in baseline row beneath it. A custom role can grant project access with no built-in project role at all, which is how a department-style role works. A custom role only adds permissions. TaskStation has no deny rule. Custom roles need the enterprise `rbac` entitlement. Without it, assigning one answers `402` with `code: "entitlement_required"`. ## Super-admin Super-admin is not a role. It is a flag on one account membership. It bypasses every permission check, and every bypass is written to the audit log. Roles are the mechanism for ordinary access; super-admin is the audited escape hatch. ## Per-feature access settings Some features carry their own visibility setting on top of the role model: | Setting | Decides | |---|---| | [App access mode](/docs/feature-flags/apps#access-modes) | Who can open one deployed App | | [Trigger session access](/docs/connect/triggers#session-access) | Who can open the sessions one trigger creates | | [Slack channel policy](/docs/connect/slack) | Who can start a session from one channel | | [Computer Tunnel capability grants](/docs/connect/computers#grant-access) | Which filesystem, shell, and desktop calls a machine accepts | Each one narrows access to one resource. None of them grants a permission the role verdict denies. ## The access API | Method + path | Does | |---|---| | `GET /v1/accounts/{accountId}/iam/assignments` | List assignments. Filter by principal, scope, object, or role. | | `POST /v1/accounts/{accountId}/iam/assignments` | Create one assignment. | | `DELETE /v1/accounts/{accountId}/iam/assignments/{assignmentId}` | Revoke one assignment. | | `GET /v1/accounts/{accountId}/iam/permissions` | The permission catalog, as data. | | `GET /v1/accounts/{accountId}/iam/roles` | Roles, built-in and custom. | | `GET /v1/accounts/{accountId}/iam/roles/{roleId}/permissions` | One role's permissions. | The catalog is data, not a hardcoded list. Each permission carries its `action`, `scope_type`, `resource_type`, `delegable` flag, `description`, `area`, `level`, and `implies`. Read it instead of hardcoding action strings. The write routes choose the permission they require from what you are granting: `project.members.manage` for a project role or an object assignment, `member.update` for an account role, `policy.create` for a custom role. You cannot side-step a ceiling by picking a different route. ## Branding An Enterprise account can put its own brand on the app for every member: a wide logo, a square icon, and a favicon, each with an optional dark-mode variant, plus a product name that replaces "TaskStation" in the browser tab title. Open **Account → Branding** (`/accounts/{accountId}?tab=branding`). Uploads accept PNG, JPEG, WebP, SVG, and ICO up to 1 MB; the icon stands in for a missing logo, the favicon falls back to the icon, and a missing dark variant falls back to the light image. In-app marks follow the app theme; the favicon follows the operating system's color scheme. Branding follows the account, not the browser: inside a project, members see the brand of the account that owns the project. Sign-in pages, emails, and public share pages stay TaskStation, because there is no account to brand until someone is signed in. When the Enterprise entitlement lapses, members see TaskStation again — nothing is deleted, and the account can remove what it uploaded at any time. | Method + path | Does | |---|---| | `GET /v1/accounts/{accountId}/branding` | The stored record and whether the plan allows it. | | `PUT /v1/accounts/{accountId}/branding` | Set or clear `app_name`. Needs `account.write` and the `branding` entitlement (`402 entitlement_required` otherwise). | | `POST /v1/accounts/{accountId}/branding/assets/{kind}` | Upload one image as multipart `file`. `kind` is `logo`, `icon`, `favicon`, or their `_dark` variants. Same gate as `PUT`. | | `DELETE /v1/accounts/{accountId}/branding/assets/{kind}` | Remove one image. `account.write` only. | | `DELETE /v1/accounts/{accountId}/branding` | Reset everything to TaskStation. `account.write` only. | `GET /v1/accounts` carries each account's effective `branding` — the record while entitled, `null` otherwise — so a client renders from one request it already makes. In the SDK: `taskstation.accounts.branding.{get, update, uploadAsset, removeAsset, reset}`. ## Switching accounts If you belong to more than one account, switch between them from the account switcher. Each account keeps its own projects, members, and settings. ## Tokens TaskStation signs you in with a personal access token (`taskstation_pat_...`). It acts as the user who created it and holds exactly that user's assignments. A service account (`taskstation_sa_...`) is a separate principal, not a person's credential, and holds no project access until an assignment gives it some. The two live on two surfaces, because they belong to two different owners. Your own API keys are in your settings, at **Settings → API keys** (`/settings/tokens`) — only you see them, and they stop working when your membership does. Service account tokens are account configuration, at **Account → Tokens**, beside the key rules that govern expiry. See [SDK authentication](/docs/sdk/auth) for the full token model. ## Billing Billing applies at the account level and covers every project in the account. TaskStation offers a free tier, a pro tier, and a per-seat team tier, each with its own model access and credits. Using your own model key does not make usage free: TaskStation still charges a platform fee on top of it, except on the free tier. > **Enterprise** > Single sign-on (`sso`), SCIM provisioning (`scim`), custom roles (`rbac`), and audit access (`auditAccess`) are entitlements on the enterprise tier. An entitlement is orthogonal to a role: it decides whether the feature exists for the account, and the role still decides who may use it. Contact sales to enable them. --- # TaskStation as a Backend Start and manage TaskStation sessions from your backend with explicit connector, model, context, and secret scope. Canonical page: https://taskstation.co/docs/backend Use a TaskStation API key to start sessions from your server. Each session has one TaskStation owner, one project, and one cost record. Your application owns its customer identifiers and metadata. Store the relationship between your customer and the returned `session_id` in your application database. ## 1. Get an API key Create a personal access token (`taskstation_pat_…`) in your own settings, at **Settings → API keys** (`/settings/tokens`), or a service-account credential (`taskstation_sa_…`) at **Account → Tokens**. Both authenticate a programmatic session-create request. The API derives `origin: "backend"` from the credential type. ```bash export TASKSTATION_API_URL="https://your-taskstation-deployment.com/v1" export TASKSTATION_API_KEY="taskstation_pat_…" export TASKSTATION_PROJECT_ID="…" ``` Use a service account when you need an independently managed principal. A service account is the `service_account` principal type. It has no membership, so it holds only the roles assigned to it directly. Assign it a project role before use — see [Accounts & access](/docs/accounts#one-access-model). ## 2. Start a session ### Create with HTTP ```bash curl -X POST "$TASKSTATION_API_URL/projects/$TASKSTATION_PROJECT_ID/sessions" \ -H "Authorization: Bearer $TASKSTATION_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{ "agent_name": "support", "opencode_model": "taskstation/glm-5.3-flash", "runtime_context": { "ticket_id": "ticket-123" }, "connector_bindings": { "gmail": { "connection_id": "" } }, "secrets": ["STRIPE_KEY"] }' ``` ### Create with the SDK ```ts import { createScopedTaskStation } from '@taskstation/sdk/server'; const taskstation = createScopedTaskStation({ backendUrl: process.env.TASKSTATION_API_URL!, getToken: async () => process.env.TASKSTATION_API_KEY!, }); const session = await taskstation.project(projectId).sessions.create({ agent_name: 'support', opencode_model: 'taskstation/glm-5.3-flash', runtime_context: { ticket_id: 'ticket-123' }, connector_bindings: { gmail: { connection_id: connectionId }, }, secrets: ['STRIPE_KEY'], }); ``` Use `createScopedTaskStation` when one server process handles concurrent requests. Each client keeps its token and runtime state request-scoped. Store the `session_id` returned by either create call: ```bash export SESSION_ID="" ``` ### Session-create fields | Field | Contract | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `agent_name` | Selects a declared OpenCode agent. | | `opencode_model` | Selects the initial OpenCode model. An unavailable model returns `400 INVALID_SESSION_MODEL`. | | `runtime_context` | Stores non-secret scalar context. The API rejects credential-like keys, more than 64 entries, or more than 16 KiB. | | `connector_bindings` | Maps a connector slug to one strategy-compatible `connection_id`. The credential stays outside the sandbox. | | `inherit_unbound` | Keeps strategy-based default resolution for connectors omitted from an explicit binding map. The default is `false`. | | `secrets` | Narrows the selected agent's project-secret grant. An empty list delivers no project secrets. Only backend-origin callers can set it. | | `require_connectors` | Adds mandatory connectors for this create request. Missing connections return `409 CONNECTOR_CONNECTION_REQUIRED`, and unconfigured slugs return `409 REQUIRED_CONNECTOR_CONNECTION_UNAVAILABLE`, before the session row is inserted and before sandbox startup. | ## 3. Configure connectors and connections A connector defines the tool surface. It contains a project-unique slug, display name, provider app, authorization strategy, and policies. A connection stores one connected account or credential for that connector. Every connection inherits the connector's policies. The authorization strategy has two values: - `project` accepts active project connections. - `user` accepts only an active connection owned by the acting project member. A service account is a principal, but it is not a person, so it cannot use a member's `user` connection. Use `project` connectors for service-account sessions. A personal access token can use an eligible `user` connection owned by the token's member. ```yaml connectors: - slug: gmail-read name: Gmail read only provider: pipedream app: gmail authorization_strategy: project policies: - match: search_email action: always_run agents: support: connectors: [gmail-read] connectors_required: [gmail-read] ``` The SDK exposes connections under `project.connectors.connections`: ```ts const connection = await taskstation.project(projectId).connectors.connections.reconcile({ connector_alias: 'gmail-read', owner_type: 'project', label: 'Support inbox', }); await taskstation .project(projectId) .connectors.connections.updateCredential(connection.connection_id, { value: credential, kind: 'secret', }); await taskstation.project(projectId).connectors.connections.activate(connection.connection_id); ``` For a Pipedream OAuth connection, call `pipedreamConnect()` and `pipedreamFinalize()`. Do not place its provider token in `updateCredential()`. > **Info** > The connection object and new session binding input use `connection_id`. > `authorization_id` remains a deprecated SDK input alias. ## 4. Read and replace session scope The session scope is authoritative server state. `secrets_allowlist` contains the session's stored narrowing. A `null` value means the agent grant applies. `connector_bindings` contains the materialized connection selection. ```bash curl -sS \ "$TASKSTATION_API_URL/projects/$TASKSTATION_PROJECT_ID/sessions/$SESSION_ID/scope" \ -H "Authorization: Bearer $TASKSTATION_API_KEY" ``` ```ts const scope = await taskstation.session(projectId, sessionId).scope(); ``` Replace scope with `PUT` or `rescope()`: ```ts const nextScope = await taskstation.session(projectId, sessionId).rescope({ secrets: ['STRIPE_KEY'], connector_bindings: { 'gmail-read': { connection_id: connection.connection_id }, }, }); ``` Each supplied field uses set semantics. The new value replaces the complete previous value. Omit a field to leave it unchanged. Connector changes apply to the next tool call. Secret removal stops future delivery. It cannot remove a value from an existing model context or process. Rotate the secret when prior disclosure matters. ## 5. Read session costs The session-cost API combines finalized LLM cost and billed sandbox compute cost. Every session appears in the list, including sessions with zero cost. ```bash curl -sS \ "$TASKSTATION_API_URL/usage/session-costs?project_id=$TASKSTATION_PROJECT_ID&limit=25&offset=0" \ -H "Authorization: Bearer $TASKSTATION_API_KEY" curl -sS \ "$TASKSTATION_API_URL/usage/session-costs/$SESSION_ID?project_id=$TASKSTATION_PROJECT_ID" \ -H "Authorization: Bearer $TASKSTATION_API_KEY" ``` The list returns session, project, owner, LLM, compute, total, request, token, model, and compute-duration fields. It also returns `reconciliation` for account usage that has no session. The detail response adds: - `model_usage`, grouped by provider and model - `ledger_entries`, with discriminated `llm` and `compute` rows Use the SDK for typed reads: ```ts const page = await taskstation.billing.sessionCosts.list({ accountId, projectId, limit: 25, offset: 0, }); const detail = await taskstation.billing.sessionCosts.get(sessionId, { accountId, projectId, }); const sameDetail = await taskstation.session(projectId, sessionId).cost(); ``` `session.cost()` does not start the session runtime. ## 6. Stream the answer Await runtime readiness before using the OpenCode REST methods: ```ts const handle = taskstation.session(projectId, session.session_id); await handle.ensureReady(); const stream = await handle.stream({ onEvent: (event) => { // Render or persist the event. }, }); await handle.send('Summarize the support queue.'); ``` Use `useSession(projectId, sessionId)` for React hosts. It owns startup, readiness, the live event stream, and message synchronization. ## Idempotent retries Generate one `Idempotency-Key` for each logical session-create operation. Reuse that key only when the request body is identical. A replay with the same key and body returns the same session. A replay with a different secret allowlist, connector binding map, or runtime context returns `409`. ## Security rules - The API derives session origin from the credential. The request body cannot select it. - Connector credentials resolve server-side for each tool call. - A connection must match its connector's authorization strategy. - Connector policies apply to every connection under that connector. - Project guardrails apply above connector-connection policies. - Secret scope can narrow an agent grant. It cannot widen one. - A session can only do what the role verdict and the agent's manifest grant both allow. Neither one widens the other. See [One vocabulary, two bindings](/docs/accounts#one-vocabulary-two-bindings). - Session scope replacement cannot select a connection owned by another member. - Store application customer metadata outside TaskStation. ## Common errors | Status | Code | Meaning | | ---------------------------- | ------------------------------------------------ | ---------------------------------------------------------------------------------- | | `400` | `INVALID_SESSION_MODEL` | The selected model is not available to the account. | | `400` | `INVALID_SESSION_CONNECTOR_BINDINGS` | The binding map is malformed. | | `400` | `INVALID_SESSION_RUNTIME_CONTEXT` | Runtime context violates its shape, key, entry, or size limits. | | `403` | `origin_override_forbidden` | A non-backend caller supplied a secret allowlist. | | `403` | `CONNECTOR_NOT_ASSIGNED` | The selected agent is not granted the connector. | | `404` create / `403` rescope | `CONNECTOR_CONNECTION_NOT_FOUND` | The connection does not exist in this project or violates the connector's authorization strategy. | | `404` | `SECRET_IDENTIFIER_NOT_FOUND` | The secret allowlist contains an unknown project-secret identifier. | | `409` | `CONNECTOR_CONNECTION_REQUIRED` | A mandatory connector has no valid active connection. Every failing connector is listed in `connector_connections`. | | `409` | `REQUIRED_CONNECTOR_CONNECTION_UNAVAILABLE` | A required slug has no configured connector at all. Every failing alias is listed in `connectors`. | | `409` | `CONNECTOR_PROVIDER_UNSUPPORTED` | The alias is a connector on the project but its provider has no hosted authorization page, so no connect link exists for it. | | `409` | `CONNECTOR_PIPEDREAM_APP_MISSING` | The Pipedream connector names no app, so no connect link can be built. | | `409` create / `403` rescope | `CONNECTOR_CONNECTION_INACTIVE` | The selected connection or connector is inactive. | | `409` | `IDEMPOTENCY_*_CONFLICT` | The idempotency key was replayed with a different request body. | | `402` | `subscription_required` / `insufficient_credits` | The account cannot start a billed session. | --- # CLI The taskstation command line, its auth model, the dev loop, and every command. Canonical page: https://taskstation.co/docs/cli The `taskstation` command line interface (CLI) controls TaskStation from a terminal — your laptop or a session sandbox. This page shows the everyday dev loop, then lists every stable command and flag. ## Install ```sh curl -fsSL https://taskstation.co/install | bash ``` The installer downloads a prebuilt binary for macOS and Linux. Windows is not supported. | Command | Effect | | --- | --- | | `taskstation update` | Re-run the install script and pull the latest binary. | | `taskstation uninstall [-y\|--yes] [--keep-auth] [--keep-home]` | Remove the binary, the `/usr/local/bin` shim, and the stored token. `--keep-auth` keeps the token. `--keep-home` keeps `~/.taskstation`. | | `taskstation version` | Print the CLI version. | ## Auth model TaskStation stores authentication per host, not globally. A host is one TaskStation API endpoint. Four hosts exist by default: `cloud` (TaskStation Cloud), `selfhost` (your self-hosted stack), `local-dev`, and `taskstation-internal-dev`. You can add more. The config file lives at `~/.config/taskstation/config.json`, mode `0600`. Override its path with `TASKSTATION_CONFIG_FILE`. The CLI follows one hierarchy: host → account → project → session. You sign in to a host, pick an account inside it, pick a project inside that account, and open sessions inside the project. `taskstation hosts login` walks the first three steps in order: it signs you in, picks the account, then sets a default project. Every token starts with `taskstation_pat_`. A user token, from `taskstation login`, sees every account and project you belong to. A project token is auto-minted for a session sandbox and scoped to one project. See [Token scope](#token-scope). ## The dev loop This loop assumes the CLI is installed and you ran `taskstation login`. See [Quickstart](/docs/quickstart) for setup. ### Link a repo Start a new project, or link an existing repo folder to one. To scaffold a new project: ```sh taskstation init my-app cd my-app ``` `taskstation init` creates a project directory with the general-purpose starter. Its `taskstation.yaml` declares `taskstation_version: 2` and runs OpenCode. To link an existing cloned repo to a project you already created: ```sh taskstation projects link ``` This command writes `.taskstation/link.json` in the current directory. TaskStation reads this file to find your project on every command run from this folder. If you plan to run `taskstation ship` first, skip this step. It links a new project for you when none exists. ### Ship your code ```sh taskstation ship ``` `taskstation ship` lints your `taskstation.yaml`, commits local changes, pushes your branch, and prompts for any missing secret or connection. Run it each time you want your local changes on the cloud project. The first run also creates the cloud project and repo if you have not linked one yet. ### Run and attach to sessions Start a session with a prompt: ```sh taskstation sessions new --prompt "Build the login page" --wait ``` Each session runs in its own sandbox, on its own branch. `--wait` blocks until the session is ready. Attach to a session from your terminal: ```sh taskstation connect ``` With no session id, `taskstation connect` (alias: `attach`) opens a session picker for the bound project — running sessions attach immediately, stopped ones boot first, and `+ New session` starts a fresh sandbox — then lands you in the full OpenCode TUI attached to that session. Pass an id to skip the picker: `taskstation connect `. The CLI manages the `opencode` binary for you: on first connect it downloads the exact version the session's server runs and caches it under `~/.taskstation/opencode//`, so the TUI and server never skew. Set `TASKSTATION_OPENCODE_BIN` to force your own binary. For a lighter-weight line-based chat instead of the full TUI, run: ```sh taskstation sessions chat ``` This opens an interactive chat with your most recent session. Add an id to target a specific session: `taskstation sessions chat `. To open a raw shell in the sandbox, with no agent involved, run: ```sh taskstation sessions shell ``` List your running sessions at any time: ```sh taskstation sessions ls ``` ### Review with change requests An agent opens a change request (CR) when its session has commits ready to merge. List, inspect, and merge them from the CLI. ```sh taskstation cr ls taskstation cr diff 1 taskstation cr merge 1 ``` `taskstation cr ls` lists change requests for the linked project. `taskstation cr diff ` shows the unified patch. `taskstation cr merge ` merges it into the project's default branch. Accept a CR number or its full id. ## Reference ### Auth commands | Command | Effect | | --- | --- | | `taskstation login [--host ] [--api ] [--token ] [--account ] [--no-project]` | Sign in to the active host, or the named one. Opens a browser by default; `--token` signs in headless. `--no-project` skips the default-project pick. | | `taskstation logout [--host ]` | Remove the token for the active host, or the named one. | | `taskstation whoami [--host ] [--json] [--token-only]` | Print the signed-in user and active account. | | `taskstation token [--host ]` | Shortcut for `taskstation whoami --token-only`. | `taskstation hosts login` / `hosts logout` / `hosts whoami` are the canonical forms. `login` / `logout` / `whoami` are shortcuts that act on the active host. ### Hosts | Command | Effect | | --- | --- | | `taskstation hosts ls [--json]` | List every host and its auth status. | | `taskstation hosts login [] [--token ] [--api ] [--account ] [--no-project]` | Sign in to a host. An unknown name registers the host first. | | `taskstation hosts logout []` | Remove the token for a host. | | `taskstation hosts use ` | Switch the active host. | | `taskstation hosts add --url [--dashboard-url ] [--login]` | Register a new host. `--login` signs in right after. | | `taskstation hosts rm [--force]` | Remove a host. | | `taskstation hosts info [] [--json]` | Show details for one host. | | `taskstation hosts current [--json]` | Print the active host name. | A remote host URL that starts with `http://` is normalized to `https://`. The CLI never sends a token over plain HTTP to a remote host. `localhost` is exempt. ### Accounts | Command | Effect | | --- | --- | | `taskstation accounts ls [--json]` | List the accounts you belong to on the active host. | | `taskstation accounts use []` | Switch the active account. | | `taskstation accounts current [--json]` | Print the active account. | | `taskstation accounts info [] [--json]` | Show one account. | ### Members Who belongs to the account, and at what account role. Roles are `owner`, `admin`, and `member`. Owners and admins hold implicit Manager on every project, so `member` is the only role that takes per-project grants. | Command | Effect | | --- | --- | | `taskstation members ls [--json]` | List members, roles, and project counts. | | `taskstation members invite --role admin\|member [--project :]` | Invite by email. An existing TaskStation user is added immediately; anyone else is mailed an invite link. `--project` is repeatable and applies on accept. Needs `member.invite`. | | `taskstation members set-role --role owner\|admin\|member` | Change an account role. Needs `member.update`; the `owner` role is owner-only. | | `taskstation members rm [-y\|--yes]` | Remove a member and revoke their tokens. Needs `member.remove`. | | `taskstation members super-admin on\|off` | Grant or revoke the super-admin bypass. Needs `member.super_admin.grant`. | | `taskstation members invites ls [--json]` | List pending invitations you sent. | | `taskstation members invites cancel ` | Cancel one pending invitation. | | `taskstation members invites resend ` | Re-send the email and refresh the 14-day expiry. | A `` is a user id, or the email of someone already in the account. Options: `--account `, `--host `, `--json`, `-y`. ### Groups An account group is a named set of people you grant a role to once. Bind a group to a scope with `taskstation access grant --group --role `; revoke with `taskstation access revoke `. | Command | Effect | | --- | --- | | `taskstation groups ls [--json]` | List groups with member and project counts. | | `taskstation groups create [--description ]` | Create a group. | | `taskstation groups set [--name ] [--description \|--no-description]` | Rename or re-describe a group. | | `taskstation groups rm [-y\|--yes]` | Delete a group. Its grants go with it. | | `taskstation groups members [--json]` | List a group's members. | | `taskstation groups add ...` | Add one or more people. | | `taskstation groups remove ` | Remove one person. | | `taskstation groups projects [--json]` | Which projects the group reaches, and at what role. | A `` is a group id or its exact name. Reads need `group.read`. `create`, `set`, and `add` need `group.update` or `group.members.manage`, plus the enterprise `rbac` entitlement. `rm` and `remove` are cleanup and are never entitlement-gated. ### Tokens Non-interactive credentials for the account. Reads need `token.read`, minting needs `token.create`, revoking needs `token.revoke`. | Command | Effect | | --- | --- | | `taskstation tokens ls [--mine] [--json]` | List the account's personal API keys. `--mine` narrows to the ones you minted. | | `taskstation tokens new [--expires ] [--project ]` | Mint a key. The secret prints once. `--project` binds it to one project, which it can never leave. | | `taskstation tokens rm [-y\|--yes]` | Revoke a key immediately. | | `taskstation tokens service-accounts ls [--json]` | List service accounts. | | `taskstation tokens service-accounts new [--description ] [--expires ]` | Create one. The bearer prints once. | | `taskstation tokens service-accounts disable ` | Disable a service account. Reversible only by deleting and re-creating. | | `taskstation tokens service-accounts rm [-y\|--yes]` | Delete a service account permanently. | A personal API key acts as you and dies with your membership. A service account acts as itself and inherits no access: a new one holds no permissions. Grant it one with `taskstation access grant --service-account --role `. `--expires ` takes ISO-8601 or a forward span: `30d`, `12h`, `6w`, `1y`. ### Billing Read the active account's plan, credits, and spend. Read-only: plan changes, top-ups, and payment methods are dashboard flows. | Command | Effect | | --- | --- | | `taskstation billing status [--json]` | Plan, credits, seats, subscription. | | `taskstation billing transactions [--limit ] [--offset ] [--type ] [--json]` | Credit ledger, newest first. Default page size 50. | | `taskstation billing transactions --summary\|--breakdown\|--usage [--days ]` | Credits in/out, the balance split (expiring/non-expiring/daily), or a credit-usage summary. `--days` is the window for `--summary` and `--usage`; default 30. | | `taskstation billing costs [--since ] [--until ] [--json]` | Account spend over a window, plus a model breakdown. Default window: 30 days, half-open `[from, to)` UTC. | | `taskstation billing costs --by project\|session [--sort ] [--limit ] [--offset ] [--csv ]` | Roll spend up by project or by session. `--sort` takes `total_desc` (default), `total_asc`, `recent`, or `name_asc` (`--by project` only). `--csv` needs `--by`. | Filters: `--project `, `--session `, and `--owner ` (sessions, with `--by session`). Options: `--account `, `--host `, `--json`. ### Projects A command resolves "the project" in this order: 1. The `--project` flag. 2. The `TASKSTATION_PROJECT_ID` environment variable. 3. `.taskstation/link.json` in the exact working directory. 4. The global default set by `taskstation projects use`. | Command | Effect | | --- | --- | | `taskstation projects ls [--all] [--query ] [--json]` | List projects on the active account. `--all` spans every account you belong to. `--query` (alias `-q`) filters by name, id, or repo. | | `taskstation projects info [] [--json]` | Show one project. Default: the linked or default project. | | `taskstation projects use []` | Set the global default project. Switches the active account if the project lives elsewhere. | | `taskstation projects unset` | Clear the global default project. | | `taskstation projects link []` | Bind the current directory to a project. Writes `.taskstation/link.json`. | | `taskstation projects unlink` | Remove `.taskstation/link.json`. | | `taskstation projects open []` | Open a project's dashboard page in your browser. | | `taskstation projects clone [] [dir]` | Clone a project's repo through the authenticated TaskStation git proxy. | | `taskstation projects rm [] [--purge] [-y\|--yes]` | Archive a project. `--purge` also deletes its managed git repo. | | `taskstation projects set [] [--name ] [--branch ] [--manifest ] [--json]` | Update one project's settings. Only the fields you pass are written. Passing no field exits `2`. Alias: `update`. | | `taskstation projects set [] --icon \|--no-icon\|--glyph :\|--no-glyph` | Set or remove the project's icon. | | `taskstation projects rename [] ` | Alias for `taskstation projects set --name `. | | `taskstation projects features [ls] [--json]` | List every feature flag with its key, state, origin, and stability. | | `taskstation projects features enable\|disable\|reset ` | Set the project override on, off, or clear it so the flag follows the platform default. | | `taskstation projects cli-tokens ls [--json]` | List the project's CLI tokens. | | `taskstation projects cli-tokens new [--name ]` | Mint a project-scoped CLI token. The secret prints once. | | `taskstation projects cli-tokens rm [-y\|--yes]` | Revoke one project CLI token. | | `taskstation projects upgrade [] [--json]` | Start the agent session that migrates a v1 `taskstation.toml` to a v2 `taskstation.yaml` and opens a change request. | A project shows one icon, so writing `--icon` clears the glyph and writing `--glyph` clears the emoji; passing both is refused. Glyph colors: `grey`, `red`, `orange`, `yellow`, `lime`, `blue`, `purple`, `magenta`. `set` and `features` need `project.customize.write`. A flag the platform marks unavailable stays off regardless of the project override. A project CLI token is bound to one project — the API rejects it everywhere else. A session sandbox uses its session-bound `TASKSTATION_TOKEN`. `cli-tokens ls` needs project read; `new` and `rm` need `project.credentials.issue`. An agent-session token can neither mint nor revoke project tokens (`403`). `taskstation projects upgrade` needs project write. The default agent refreshes the marketplace baseline, rewrites the manifest, runs `taskstation validate`, and opens a change request. It never merges: a human reviews the diff. ### Project scaffold `taskstation init [project-name] [options]` creates a new project directory. The starter writes a v2 `taskstation.yaml` and the canonical `.taskstation/opencode` system-skill source. The command can wire local coding-tool discovery without changing the cloud OpenCode runtime. The command does not write `.taskstation/link.json`. `taskstation ship` or `taskstation projects link` create that file. | Flag | Meaning | | --- | --- | | `--name ` | Project name. | | `--primary ` | Primary agent. | | `--agents ` | Local coding-agent integrations to wire up. | | `--force` | Configure the current directory in place instead of scaffolding a new one. | | `--overwrite` | Overwrite existing files. | | `--no-git` | Skip git init. | | `-y, --yes` | Don't prompt. | The local coding-tool selection does not change the cloud OpenCode runtime. `taskstation init` does not include a marketplace picker. Adding a marketplace skill is an agent import: start a session and ask the agent to bring one in. ### Ship `taskstation ship` stages, commits, and pushes your current branch to the project's git repo. Run it once to create the project. Run it again any time to sync. Alias: `taskstation deploy`. Each run: - Parses and validates `taskstation.yaml` (skip with `--no-verify`). - Commits any dirty working tree (skip with `--no-commit`). - Prompts for any missing `env` secret (skip with `--no-env`). - Pushes the current branch to the same-named remote branch. - Connects any declared connector that still needs auth (skip with `--no-connect`). An existing GitHub `origin` links through the TaskStation GitHub App. Any other existing `origin` is registered as-is. No `origin` creates a managed TaskStation git repo. | Option | Effect | | --- | --- | | `--name ` | Display name for a new project. | | `--account ` | Account to create the project under (first ship only). | | `--origin ` | Override the inferred origin choice. | | `--github-token ` | Link a GitHub origin with this token instead of the GitHub App. | | `-m, --message ` | Commit message. | | `--no-commit` / `--no-verify` / `--no-env` / `--no-connect` | Skip that step. | | `-y, --yes` | Don't prompt. | | `-n, --dry-run` | Print what would happen; change nothing. | | `--project ` / `--host ` | Target a non-default project or host. | ### Sessions Each session runs in one sandbox on its own branch. | Command | Effect | | --- | --- | | `taskstation sessions ls` | List every session on the project. | | `taskstation sessions status [--all] [--json]` | Every session and what its agent is doing right now. Aliases: `overview`, `ps`. | | `taskstation sessions info [--json]` | Detail view: status, branch, agent, sandbox URL. | | `taskstation sessions new [--prompt ""] [--agent ] [--model ] [--wait] [--connect] [--json]` | Start a session. `--connect` attaches the OpenCode TUI once it is ready (implies `--wait`); on an interactive terminal without it, the CLI asks whether to connect after creation. `--model ` overrides the project's default model. `--wait` blocks until it is running (up to ~5 minutes). Use `--secret ` or `--no-secrets` to narrow Secret access. These Secret flags require a backend token. Use `--connector =` or `--no-connectors` to set Connector access. Use `--require-connector ` to require an authorization before provisioning. Scope flags are repeatable. Use `--context =` for non-secret runtime context. | | `taskstation sessions chat [] [--prompt ""] [--queue] [--new] [--agent ] [--json]` | Talk to a session's agent. Interactive by default. Alias: `talk`. Top-level `taskstation chat` also works. `--queue` is one-shot only: it stores the prompt in the session's durable inbox and returns as soon as it is stored, instead of handing it to the runtime. | | `taskstation sessions connect [] [-- ]` | Attach the OpenCode TUI to the session's OpenCode server. Also available top-level: `taskstation connect` / `taskstation attach`. With no id, opens a session picker (running, stopped-with-restart, or new). The CLI auto-downloads the version-matched `opencode` binary (cache: `~/.taskstation/opencode//`; override: `TASKSTATION_OPENCODE_BIN`). | | `taskstation sessions shell [] [--new]` | Open a raw interactive terminal in the sandbox, with no agent. Reattaches to the session's existing terminal; `--new` always starts a fresh one. Aliases: `terminal`, `ssh`. | | `taskstation sessions shell ls [--json]` | List the session's terminals: id, status, command. Needs no TTY. | | `taskstation sessions shell kill ` | Kill one terminal. The ambient shell respawns on the next attach; anything running inside it does not. | | `taskstation sessions log [] [--limit ] [--json]` | Print recent messages, read-only. Aliases: `messages`, `history`. | | `taskstation sessions pending [--json]` | List open tool-permission or question prompts. Alias: `prompts`. | | `taskstation sessions approve [] [--always] [--reject] [--message ""]` | Answer a permission prompt. | | `taskstation sessions answer [] [--option ]... [--text ""] [--reject]` | Answer a question prompt. | | `taskstation sessions digest [--since <7d>] [--json]` | Compact multi-session review. Aliases: `review`, `summary`. | | `taskstation sessions scope [scope options] [--json]` | Read or replace Secret and Connector access. Alias: `access`. Use `--secret `, `--no-secrets`, or `--inherit-secrets` for Secrets. Use `--connector =` or `--no-connectors` for Connector bindings. Use `--require-connector ` or `--no-required-connectors` for required Connectors. Provided categories replace their current values. Omitted categories remain unchanged. Changes apply to the next prompt. Removed Secret values remain in existing context if the session already read them. | | `taskstation sessions preview [port] [--port ] [--list] [--json]` | Print a clickable preview URL for a sandbox port. Default port: `3000`. `--list` prints the named candidates instead. | | `taskstation sessions restart ` | Restart the session's sandbox. | | `taskstation sessions rename ` | Set a session's name. Pass `""` to clear it. | | `taskstation sessions rm ...` | Stop and delete one or more sessions. | | `taskstation sessions open ` | Open a session's dashboard page in your browser. | | `taskstation sessions stop [--json]` | Pause a session. The sandbox stops in place and the disk is kept. Alias: `pause`. Needs `project.session.stop`. | | `taskstation sessions start [--wait] [--json]` | Wake a session: provision a missing sandbox, resume a stopped one, and resolve its runtime. Idempotent. Alias: `wake`. `--wait` blocks until ready (up to ~5 min) and exits `1` if the session ends up failed or stopped. | | `taskstation sessions warm [--exclude ] [--json]` | Pre-create the session you are about to use, so the sandbox is already up. Reuses an existing unused warm session. A warm session stays hidden from `sessions ls` until its first prompt. | | `taskstation sessions model [--json]` | Change the model a session runs, mid-session. A live sandbox is re-pointed and its runtime restarts, which ends the turn running right now; a stopped session stores the value for its next start. | | `taskstation sessions compact [--json]` | Summarize the conversation so far and continue from the summary. | | `taskstation sessions queue [ls] [--json]` | List the prompts still waiting in the session's durable inbox. | | `taskstation sessions queue rm ` | Drop one queued prompt. Refused (`409`) once a model step has started answering it. | | `taskstation sessions queue now ` | Run one queued prompt next: re-queue it ahead of the ordering rule and release the session's hold. | | `taskstation sessions queue hold\|release` | Hold every queued prompt — what the Stop button writes — or release the hold. | | `taskstation sessions approvals [ls] [--json]` | List the governed connector calls this session is waiting on a human for. | | `taskstation sessions approvals approve\|deny ` | Let one governed connector call run, or refuse it. The agent is told and continues without it. | | `taskstation sessions files [--json]` | Read and edit the sandbox's live workspace: `ls []`, `status`, `find `, `write `, `touch `, `mkdir `, `mv `, `rm `. | | `taskstation sessions share [--mode private\|project\|members] [--member ] [--group ] [--show] [--json]` | Set who inside TaskStation can open this session. With no `--mode` it prints the current setting and changes nothing. `--member` and `--group` are repeatable. | | `taskstation sessions links ls [--json]` | List every public link ever minted on the session, newest first. | | `taskstation sessions links create [options]` | Mint one public, unauthenticated link onto a preview port or one workspace file. | | `taskstation sessions links revoke ` | Kill one public link. | `sessions queue` needs `project.session.start` — the same permission as sending a message. A queued prompt survives a closed terminal and is delivered when the session can take it. Put one there with `taskstation sessions chat -p "…" --queue`. `sessions approvals` are durable: unlike `sessions pending`, they survive a sandbox restart. It needs `project.members.manage`, or being the human who launched the session. An agent may never resolve its own approval. `sessions files` reads the working tree the agent is editing right now, before anything is committed; `taskstation files` reads the committed repo instead. Paths resolve under `/workspace` unless they start with `/workspace`, `/tmp`, `/home`, or `/opt`. The command wakes the sandbox if it is asleep. Options: `--from ` (`write` reads this file instead of stdin), `--content` (`find` greps contents with ripgrep instead of filenames), `--limit ` (`find` filename cap), and `-y` to skip the `rm` confirmation. `sessions share` is owner-governed: the API refuses a project manager who cannot already read the session. `sessions links create` options: `--port ` (default `3000`; `22`, `8000`, and the opencode ports are refused), `--path

` (default `/`), `--preview ` (a named candidate — `web`, `vite`, `dev-server`, `api-docs` — instead of `--port`/`--path`), `--file ` (share one workspace file instead of a preview; always read-only), `--mode view\|interactive` (default `view`; `interactive` allows writes and websockets, and is ignored for `--file`), `--label `, and `--expires `. Minting a link needs the session owner, because the link itself needs no login; listing and revoking also accept a project manager. Inside a sandbox, `TASKSTATION_SESSION_ID` is your own session's id. ### Change requests A change request (CR) merges one branch into another on any git host. It is the only way for an agent to land session work on the default branch. See [Change requests](/docs/work/change-requests). | Command | Effect | | --- | --- | | `taskstation cr ls [--status open\|merged\|closed\|all] [--project ]` | List CRs. Default: `--status open`. | | `taskstation cr show [--project ]` | Show one CR, including its merge preview. Alias: `info`. | | `taskstation cr diff [--no-color] [--json]` | Print a CR's unified diff. | | `taskstation cr open --title "" [--description ""] [--head ] [--session ] [--base ]` | Open a CR. Aliases: `new`, `create`. Inside a sandbox, `--head` and `--session` default automatically. `--base` defaults to the project's default branch. | | `taskstation cr merge [--message ""]` | Merge an open CR. Fast-forward when possible, three-way merge otherwise. | | `taskstation cr close ` | Close an open CR without merging. | | `taskstation cr reopen ` | Reopen a closed CR. Merged CRs are terminal. | | `taskstation cr merge-preview [--json]` | Report whether the CR can merge, and list every conflicting path. Alias: `preview`. | | `taskstation cr request-changes --message ""` | Ask the agent that opened the CR to revise it. Alias: `changes`. | | `taskstation cr version-diff --from --into [--json]` | Summarize one version against another before opening a CR. | `request-changes` records the note on the CR and delivers it to the originating session, booting its sandbox if it is asleep. It needs `project.review.act` — the same leaf the Review Center uses, not `gitops.push`. `` accepts a per-project number (`3`) or the full id. Inside a sandbox, the CLI reads its token automatically — no login or link needed. ### Review The project's review inbox — everything waiting on a human decision: change requests, connector tool calls a policy gated for approval, and the outputs, decisions, and batches agents submit for sign-off. Mirrors the dashboard's Review Center. Gated by the `review_center` feature flag; turn it on with `taskstation projects features enable review_center`. | Command | Effect | | --- | --- | | `taskstation review ls [--segment ] [--kind ] [--json]` | List inbox items. Default: every segment. | | `taskstation review show [--json]` | Show one item in full. | | `taskstation review act [--message ]` | Decide one item. `--message` carries the note. | | `taskstation review bulk [ …]` | Decide several native items in one call. | | `taskstation review submit --kind --title [options]` | Submit an output, decision, or batch for review. | Verdicts: `approve`, `reject`, `changes`, `answer`, `dismiss`. Segments: `needs_you`, `waiting`, `done`. Kinds: `change`, `approval`, `output`, `decision`, `batch`. Where a verdict lands depends on the item id. On `cr:`, `approve` merges the change, `reject` closes it, and `changes` sends the note back to the agent that opened it (`--message` required). On `call:`, `approve` lets the tool call run and `reject` denies it; a connector approval takes no other verdict — read its arguments first with `taskstation review show`. Every other id goes to the native act endpoint, which takes every verdict. `bulk` acts on native ids only. A connector approval needs its own parameter review and a change request needs its diff in view, so both are reported and skipped — the same rule as the dashboard's multi-select. `submit` options: `--kind output\|decision\|batch` (required), `--title ` (required), `--summary `, `--risk none\|low\|medium\|high` (default `none`), `--detail ` (a JSON object), `--agent `, and `--session ` (ignored when it is not this project's session). Reads need `project.review.read`, verdicts need `project.review.act`, and `submit` needs `project.review.submit`. ### Secrets Encrypted values stored on the project. By default a secret injects as a plain environment variable into every session sandbox at boot (environment exposure). Enforced delivery — where the sandbox holds a handle and TaskStation substitutes the real value outside it (egress-enforced exposure, and `taskstation secrets call`) — is experimental. Enable the `secrets_egress` feature flag (Settings → Feature flags) to use it; with the flag off, `taskstation secrets delivery … egress` returns `403` `feature_disabled`. See [Secrets](/docs/project/secrets). | Command | Effect | | --- | --- | | `taskstation secrets ls` | List secrets by identifier and manifest `env` spec. Marks required-but-missing values. | | `taskstation secrets set NAME=VALUE ... [--identifier ]` | Upsert one or more secrets. `NAME=-` reads the value from stdin. | | `taskstation secrets request NAME ... [--scope runtime\|connector] [--expires ]` | Mint a link for a human to enter a value directly — you never see the raw value. | | `taskstation secrets unset NAME ...` | Remove secrets. | | `taskstation secrets grant IDENTIFIER --agent ` | Let one agent receive this secret: merge the identifier into that agent's `secrets` list in `taskstation.yaml`, adding the agent entry when the manifest omits it. | `grant` is the fix for a row `ls` reports as undeliverable. It only ever widens one agent's list; to narrow or replace it, rewrite the whole set with `taskstation agents scope`. There is no `secrets revoke` — the API has no route that removes a single identifier from a grant. The first grant on a project with no agents starts governance: from then on, an agent the manifest does not list receives no project secrets, and the command says so when it happens. ### Env | Command | Effect | | --- | --- | | `taskstation env pull [--out ] [--force]` | Write a `.env` skeleton — names only. Values never leave the cloud. | | `taskstation env push --from ` | Upload every `NAME=VALUE` from a dotenv file as a secret. | ### Agents Per-agent model settings on the linked project. | Command | Effect | | --- | --- | | `taskstation agents ls [--json]` | Show every agent's pinned model and the fallback default. Alias: `models`. | | `taskstation agents model ` | Pin an agent to a model. | | `taskstation agents model --clear` | Clear the pin — the agent follows the default again. | | `taskstation agents default ` | Make this the project's default agent. | | `taskstation agents default --show [--json]` | Print the current default agent. | | `taskstation agents scope [--secrets all\|none\|A,B] [--connectors all\|none\|a,b] [--require-connector ]` | Replace which secrets and connectors the agent may use. `--require-connector` is repeatable and must resolve before a session starts. | | `taskstation agents scope --show [--json]` | Print the agent's current scope. | | `taskstation agents config [--json]` | Print the full agent config block. | | `taskstation agents config --file ` | Replace the block with a JSON file's contents. `-` reads stdin. | | `taskstation agents config --set = ...` | Change single dotted keys, merged in. Repeatable, e.g. `opencode.model=glm-5.3-flash`, `enabled=false`, `connectors=["slack"]`. | Every scope option replaces; none merge. A `--set` value is parsed as JSON when it parses, and kept as a string otherwise. Model pins and `scope` apply instantly, with no `taskstation.yaml` commit; `default` and `config` commit to `taskstation.yaml` on the project's default branch. `scope` needs `project.agent.write`; `default` and `config` need `project.customize.write`. ### Models Which models the project offers, and which one it starts with. Same surface as the dashboard's Customize → Models. A project stores only its exceptions to the catalog default (the newest model of each family). Enablement is display-only: it decides what pickers offer, never what the gateway serves. | Command | Effect | | --- | --- | | `taskstation models ls [--json]` | List every model: state, origin, provider. | | `taskstation models enable ...` | Offer these models. | | `taskstation models disable ...` | Stop offering them. The project default refuses with `409` — change the default first. | | `taskstation models reset` | Drop every exception; back to the catalog default. | | `taskstation models default [--json]` | Print the default chain (project → account → platform) and what it resolves to. | | `taskstation models default [--account]` | Set the project default, or the account-wide one with `--account`. | | `taskstation models default --clear [--account]` | Clear the project, or account, default. | Model ids are gateway wire ids — a bare managed id (`glm-5.3-flash`) or a BYOK `provider/model`. Copy one from `taskstation models ls --json`. Per-agent pins live on `taskstation agents model `. Writes need `project.customize.write`. ### Channels Manages the project's connection to a chat platform. Tokens are stored encrypted in the project's secrets and resolved server-side — they are never injected into the sandbox. | Command | Effect | | --- | --- | | `taskstation channels status [--json]` | Show the current connection. | | `taskstation channels connect [--wait] [--timeout ]` | Connect in one step: prints an install link. `--wait` polls until the install lands. | | `taskstation channels connect --manual [--bot-token ] [--signing-secret ]` | Bring-your-own-app mode: save a bot token and signing secret directly. | | `taskstation channels disconnect [--platform slack\|teams]` | Drop the project's connection — the Slack one, or the Teams one with `--platform teams`. | | `taskstation channels manifest` | Print the app manifest JSON for the bring-your-own-app path. | | `taskstation channels email status [--json]` | Inbox and delivery mode for one email connector. | | `taskstation channels email connect [options]` | Create a managed inbox, or attach an existing AgentMail one. | | `taskstation channels email disconnect` | Drop the inbox connection. | | `taskstation channels email policy [--allow ] [--allow-regex ] [--allow-all]` | Replace who may email the agent. | | `taskstation channels bindings [ls] [--json]` | List every bound channel and the agent, model, and join policy it resolves to. | | `taskstation channels bind [--agent \|--no-agent] [--model \|--no-model] [--policy

]` | Change one binding. `--policy` takes `owner_approval`, `owner_only`, or `project_open`. | | `taskstation channels voice name ` | Set the display name the voice bot joins calls with. | | `taskstation channels voice name --show` | Print the current voice bot name. | `--platform slack|teams` selects the platform; default `slack`. Teams `connect` prints the Microsoft admin-consent URL; granting tenant-wide consent publishes the app to your Teams catalog automatically. See [Connectors](/docs/connect/connectors). The email channel is AgentMail-backed and needs the `agentmail_email` feature flag. `email connect` options: `--connector ` (default `taskstation_email`), `--api-key ` (bring your own AgentMail key; `-` reads stdin), `--display-name ` (from-name on outgoing mail; default the project name), `--username ` and `--domain ` (a new managed inbox), and `--inbox-id ` with `--email ` (attach an existing inbox — both are required together). `--allow` is repeatable and puts the policy in restricted mode; a bare value with no `@`, or one with a leading `@`, is read as a domain. `--allow-all` clears the list and accepts every sender again. Email and `bind` writes need `project.connector.write`. `voice name` needs `project.customize.write`. ### Connectors Connectors an agent calls as tools. `add`, `rm`, and `policy set` edit the local `taskstation.yaml`; run `taskstation ship` to apply, unless you pass `--apply` to change the cloud project immediately. | Command | Effect | | --- | --- | | `taskstation connectors ls [--json]` | List connectors and their status. | | `taskstation connectors show [--json]` | Show one connector's tools. | | `taskstation connectors add --provider

[options] [--apply]` | Add a connector. | | `taskstation connectors rm [--apply]` | Remove a connector. | | `taskstation connectors rename ` | Set a connector's display name. | | `taskstation connectors sync` | Reconcile the catalog from the shipped `taskstation.yaml`. | | `taskstation connectors credential [value]` | Set a connector's credential. | | `taskstation connectors connect ` | Start a one-click connect flow. | | `taskstation connectors link [--expires ]` | Mint a shareable connect link for a human. | | `taskstation connectors apps [] [--category ] [--cursor ] [--json]` | Browse the Pipedream app catalog. | | `taskstation connectors catalog [] [--cursor ] [--json]` | Browse the direct-connector catalogue. Needs the `connectors_api_discover` flag. | | `taskstation connectors catalog show [--json]` | Show one catalogue record's surfaces. | | `taskstation connectors sensitive on\|off` | Gate this connector's reads too — every call then needs approval. Applies now. | | `taskstation connectors owner project\|user` | Who authorizes: one project connection, or each member's own. Applies now. | | `taskstation connectors machines [--show] [--add ] [--rm ]` | Which paired computers a `computer` connector may target. Applies now. | | `taskstation connectors authorize [--status] [--scope ""] [--client-id ] [--client-secret ] [--success-redirect ] [--error-redirect ] [--json]` | OAuth 2.1 a connector end to end: discover the server's authorization metadata, register TaskStation as a client (RFC 7591) where the server supports it, and print the URL to approve. `--status` reports the result instead. | | `taskstation connectors authorize --device` | Same, using the OAuth 2.0 device flow (RFC 8628): print a code and a URL, then poll until it is approved, denied, or expired. | | `taskstation connectors policy ls [--json]` | Show project-wide execution policy. Alias: `show`. | | `taskstation connectors policy set --default [--apply]` | Set the default execution mode in `taskstation.yaml`. `--apply` sets it live instead. | | `taskstation connectors policy add [--condition ]` | Add a project-wide rule. Applies now. `--condition` narrows it to a matching argument and is repeatable; `k!=v` negates, and `k` is a dot path into the call's arguments. | | `taskstation connectors policy rm ` | Remove a project-wide rule. Applies now. | | `taskstation connectors policy ls\|set \|rm \|clear` | Manage one connector's tool-call rules. | `policy ls`, `show`, `set`, `add`, and `rm` are the project-wide surface, so a connector named after one of those verbs must be addressed as `policy ls`. A `` is a tool name, a glob (`send_*`), or a `/regex/`. `add` options: `--name

) [options]` | Create a custom template and start a build. | | `taskstation sandboxes update [options]` | Update a template. | | `taskstation sandboxes build ` | Trigger a rebuild. | | `taskstation sandboxes rebuild ` | Force-rebuild: delete the existing snapshot first. | | `taskstation sandboxes rm ` | Delete a template. | | `taskstation sandboxes fix` | Start a session seeded with the last failed build log, to repair it. | | `taskstation sandboxes provider [--json]` | Show the project's sandbox-provider pin and which providers this host offers. | | `taskstation sandboxes provider [--timeout ]` | Pin every new session to one provider. Where the target needs its snapshot built first, the API answers with a preparation and the command follows it to completion. Default `--timeout`: 600s. | | `taskstation sandboxes provider --clear` | Drop the pin and follow the platform default. Alias: `--unpin`. | | `taskstation sandboxes provider status [--json]` | Show the latest provider transition and its history. Alias: `transition`. | Pinning a provider needs `project.customize.write`. `add`/`update` options: `--name

]` | List commits on `--ref`. | | `taskstation files show ` | Show one commit and its changed files. | | `taskstation files diff [--path

]` | Print a commit's unified patch. | | `taskstation files compare ` | Summarize the diff between two refs. | | `taskstation files download -o ` | Download the repo, or the `--path` subtree, at `--ref` as a zip. Alias: `archive`. | Options on every subcommand: `--ref `, `--path

`, `--limit `, `--json`. `download` also takes `-o, --out `, which is required. Every subcommand needs `project.file.read`. `download` additionally refuses any subtree that would include an agent or skill you are scoped out of — a zip cannot be filtered mid-stream — so archive a narrower `--path` in that case. ### Triggers A trigger starts a session from a schedule, a webhook, or a monitor (experimental). `add`, `rm`, `enable`, `disable` edit the local manifest — run `taskstation ship` to apply. `pause`/`resume` flip a separate, server-side switch. See [Triggers](/docs/connect/triggers). | Command | Effect | | --- | --- | | `taskstation triggers ls [--json]` | List triggers and their runtime state. | | `taskstation triggers add [options] [--apply]` | Append a trigger to the manifest. `--apply` creates it on the cloud project now instead: it commits to `taskstation.yaml` on `main` and reconciles. | | `taskstation triggers set [options]` | Change a live trigger. Only the flags you pass are written. Always applies now — there is no local form. Alias: `update`. | | `taskstation triggers rm [--apply]` | Remove a trigger from `taskstation.yaml`, or from the cloud project now with `--apply`. | | `taskstation triggers info [--json]` | Show one trigger. | | `taskstation triggers fire ` | Fire a trigger manually. | | `taskstation triggers enable [--apply]` / `disable [--apply]` | Turn one trigger on or off. | | `taskstation triggers pause` / `resume` | Deactivate or reactivate every trigger on the project, server-side. | `add` options: `--type ` (default `cron`), `--prompt ` (required), `--agent `, `--cron ` (6-field, e.g. `"0 0 9 * * 1-5"`), `--run-at ` (run once at this instant instead of on a cron), `--timezone ` (default UTC), `--secret-env `, `--name

]` | Create one. The signing secret prints once, and a test delivery fires immediately. `--action-prefix` delivers only actions with that prefix. | | `taskstation audit webhooks enable ` | Resume delivery. | | `taskstation audit webhooks disable ` | Pause delivery, keeping the endpoint. | | `taskstation audit webhooks rm ` | Delete a webhook permanently. | Audit webhooks stream the trail to a SIEM. Every verb needs `account.write`; `add` and `enable` also need the enterprise entitlement. `disable` and `rm` never do. Filters: `--actor`, `--actor-type`, `--project`, `--session`, `--source`, `--phase`, `--outcome`, `--action`, `--resource-type`, `--request-id`, `--correlation-id`, `--query`, `--since`, `--until`, `--cursor`, and `--limit`. Account lists and exports require `audit.read` and the account's `auditAccess` entitlement. Project-wide lists require `project.members.manage` because they can include private-session metadata. Session reconstruction requires `project.session.read` and visibility of that session. ### Roles A role is a named set of permissions. System roles (`owner`, `admin`, `member` at account scope; `manager`, `member` at project scope; plus `agent-user`, the marker an object assignment carries) are read-only references. Custom roles are yours to create and edit, and need the enterprise `rbac` entitlement. | Command | Effect | | --- | --- | | `taskstation roles ls [--json]` | List roles, system and custom. | | `taskstation roles show [--json]` | Show one role's permissions and usage. | | `taskstation roles permissions [--json]` | List one role's permissions. | | `taskstation roles create --name [options]` | Create a custom role. | | `taskstation roles edit [--name ] [--desc \|--no-desc]` | Rename or re-describe a custom role. Its key never changes. Needs `role.update`, and refuses a system role. | | `taskstation roles set-actions --actions a,b` | Replace a custom role's permissions. | | `taskstation roles rm ` | Delete a custom role. | | `taskstation roles export [--project ] [--out ] [--format toml\|json]` | Dump roles and assignments to a file. | | `taskstation roles import ` | Apply a roles and assignments file. | Bind a role to a principal with `taskstation access grant`. The older `taskstation roles assign` / `unassign` / `assignments` verbs still work and write the same table, but `taskstation access` is the documented path. `taskstation roles actions` is superseded by `taskstation permissions ls`. A custom role only adds permissions. TaskStation has no deny rule, so a role cannot withhold a permission from a manager. ### Permissions The permission catalog, as data. One row per leaf action, with the scope it is decided at, whether it is delegable, and what it implies. Roles are built from these keys — `taskstation roles create --actions` and `taskstation roles set-actions` take exactly them. Alias: `taskstation perms`. | Command | Effect | | --- | --- | | `taskstation permissions ls [--scope account\|project] [--area ] [--json]` | List the catalog. | | `taskstation permissions show [--json]` | Show one action in full. | ### Grants Assigns one project object to a principal — an **object assignment**. Secrets and connectors live on agents, so assigning an agent to a person grants everything that agent declares. An agent is closed by default: a member reaches it only when an assignment names them or one of their groups. `taskstation access grant --agent ` writes the same row. | Command | Effect | | --- | --- | | `taskstation grants ls [--json]` | List object assignments, and which agents can be assigned. | | `taskstation grants assign --to [--group]` | Assign an agent to a user, or to a group with `--group`. | | `taskstation grants revoke ` | Revoke one object assignment. | ### Manifest validation | Command | Effect | | --- | --- | | `taskstation validate [--file ] [--json] [--scopes]` | Validate the manifest against the canonical schema. Resolves `taskstation.yaml` first, then `taskstation.toml`. Exit codes: `0` valid, `1` errors, `2` file missing. | | `taskstation schema [--version 1\|2] [--url]` | Print the manifest's JSON Schema. `--url` prints the schema URL instead. | See [Manifest reference](/docs/project/manifest). ### Self-host `taskstation self-host` runs one Docker-based stack, identical on a laptop, a VPS, or a cloud VM. See [Self-hosting](/docs/host) and [Self-hosting architecture](/docs/host/architecture). | Command | Effect | | --- | --- | | `taskstation self-host init` | Create or refresh the self-host config. Does not start the stack. | | `taskstation self-host configure` | Interactive wizard for connections and update policy. | | `taskstation self-host doctor` | Validate Docker tooling and the rendered config. | | `taskstation self-host plan` | Validate the rendered Compose config; change nothing. | | `taskstation self-host start` | Create config if needed, then start the stack. Aliases: `up`, `deploy`. | | `taskstation self-host update [--tag \|--channel stable\|latest]` | Pull images for the configured channel or tag and recreate the stack. Alias: `upgrade`/`reconcile`. | | `taskstation self-host rollback --release ` | Roll back to an explicit older version. | | `taskstation self-host version` | Show the running version and channel. | | `taskstation self-host restart` / `stop` | Restart or stop the stack. Alias for stop: `down`. | | `taskstation self-host status` / `ps` | Show service status. | | `taskstation self-host open` | Open the dashboard in your browser. | | `taskstation self-host connect-github` | Connect a GitHub App for managed repos. | | `taskstation self-host env ls [--show]` | Show persistent config values, masking secrets by default. | | `taskstation self-host env set KEY=VALUE ...` | Set a value and restart only the services it affects. | | `taskstation self-host env rotate KEY\|--all-generated` | Regenerate a rotatable, CLI-generated secret. | | `taskstation self-host logs [service]` | Tail stack logs. | | `taskstation self-host uninstall` | Stop the stack and delete this instance's containers, volumes, and config. | Common flags: `--instance ` (default `default`), `--domain `, `--tunnel cloudflare`, `--version`/`--tag`/`--release `, `--channel stable|latest` (default `stable`), `--auto-update on|off` (default `on`; forced off by `--local-images`), `--update-time ` / `--update-tz ` (auto-updater schedule), `--local-images` (run locally-built images; dev mode), `--enterprise-license` (unlock SSO/SCIM/RBAC/audit), `--admin-email `, `--no-restrict-account-creation` (let any signed-in user create new accounts/orgs; default is admin-only), `--restrict-account-creation` (re-enable the admin-only default), `--json`, `--yes`. ### Token scope Every token starts with `taskstation_pat_`. A user token is scoped to every project on your accounts. A project token is scoped to one project and auto-injected into that project's sandboxes. See the full token-family reference at [Session runtime](/docs/work/runtime). ### Exit codes | Code | Meaning | | --- | --- | | `0` | Success. | | `1` | Operation failed. Diagnostics print to stderr. | | `2` | Bad flag, unknown subcommand, or missing required argument. | --- # Computer Tunnel Connect your machine through the permissioned TaskStation Agent Tunnel. Canonical page: https://taskstation.co/docs/connect/computers A computer is your own machine — laptop, desktop, or server — connected to TaskStation through a permissioned reverse tunnel. A computer is not a sandbox. A sandbox is a disposable cloud machine that TaskStation creates for a [session](/docs/work/sessions). A computer is a machine you already own, and it stays connected across sessions. ## Connect a machine 1. Add **Computer Tunnel** from the project's connector catalog. 2. Run the pairing command shown in the profile, for example `npx --yes @taskstation/agent-tunnel@latest connect --api-url `. 3. Approve the connection in your browser. 4. Select the paired machine for the profile. ## Grant access You grant access per capability: filesystem, shell, or desktop. You can scope each capability to allowed paths, commands, or desktop features. A capability grant is a per-resource setting on one machine, not a TaskStation role. It grants no permission the role verdict denies — see [Accounts & access](/docs/accounts#per-feature-access-settings). The agent gets only what you grant. A call to an ungranted capability creates a permission request. Open the machine inside its Computer Tunnel profile to approve or deny the request. An unrestricted shell grant can run any executable available to your user. An unrestricted filesystem grant can access any path allowed by the local Agent Tunnel config. Grant the smallest path, command, feature, and expiry that the task needs. ## How the agent reaches a computer Pairing adds a machine to your account fleet. It writes no assignment, so it grants no project access. Add **Computer Tunnel** from the project's [connector](/docs/connect/connectors) catalog, then select one or more paired machines for that profile. One profile can contain one machine or a set of machines. A machine set is not an account group — an account group is a principal in the role model, and these are machines. You can create multiple profiles with different or overlapping machine sets. Each profile has independent agent grants and tool policies. The `list_computers` tool returns only machines assigned to the active profile. Other tools accept an optional `computer` name or id from that result. The selector is optional when exactly one assigned machine is online. ```bash taskstation connectors call studio-computers.list_computers '{}' taskstation connectors call studio-computers.fs.read \ '{"computer":"MacBook-Pro-9.local","path":"/etc/hosts"}' ``` A computer authenticates with a machine-specific setup token stored locally. The API stores only its hash. Remote connections require HTTPS/WSS. Project credentials cannot call the raw tunnel API; they must pass the selected Computer Tunnel profile, connector grant, and tool policy. Per-machine filesystem, shell, and desktop access also lives in the tunnel permission layer. The connector policy and tunnel permission must both allow a call. The local agent enforces each tunnel permission again. Its configured allowed paths and commands are maximum access boundaries that a server grant cannot widen. Computer Use requires a separately installed local `cua-driver`. Agent Tunnel does not download, install, or update that executable. You configure Computer Tunnel profiles from the dashboard, not from `taskstation.yaml`. Use the profile's Accounts tab to pair, select, inspect, rename, and remove machines. Use its Tools tab to configure connector policy. --- # Connectors How connectors and connections give agents scoped access to external tools. Canonical page: https://taskstation.co/docs/connect/connectors A connector links a project to an external tool or service. The agent calls it as a tool. TaskStation brokers each call, so the sandbox never holds the connector credential. You declare most connectors in `taskstation.yaml`. See the [manifest reference](/docs/project/manifest) for every field. TaskStation declares channel and computer connectors when you connect a chat platform or a machine. ## Connectors and connections A connector is the agent-facing reach package. It is not a role, and it holds no TaskStation permission: an agent reaches a connector only when its manifest grant lists the slug and the role verdict allows the call. See [One vocabulary, two bindings](/docs/accounts#one-vocabulary-two-bindings). A connector contains: - a project-unique slug - a display name - a provider app - one authorization strategy - connector policies A connection is one connected account or credential for the connector. Every connection uses the connector's policies. The authorization strategy is: - `project` for connections available to eligible project members - `user` for a connection owned by the acting project member A service account is a principal, but it is not a person, so it cannot use a member's `user` connection. Multiple connectors can reference the same provider app. Use separate connectors when one app needs different policies. ## Providers A connector uses one provider type: - **pipedream** — managed OAuth for supported SaaS apps - **openapi**, **postman**, **graphql**, **http** — direct API connectors - **mcp** — a remote MCP server over HTTP or SSE - **channel** — a chat platform connection - **computer** — one permissioned connector profile for one connected machine See [Slack and channels](/docs/connect/slack) and [Computers](/docs/connect/computers) for the managed provider flows. ## Authentication and policy A connection authenticates with: - OAuth through Pipedream, a channel install, or a native OAuth2 grant - an API key or token entered through the dashboard or SDK TaskStation encrypts connection data and resolves it server-side for each tool call. The agent requests an action. TaskStation attaches the credential, checks the agent grant and connector-connection policy, calls the external API, and returns the result. Connector policies belong to the connector. A connection cannot override them. Project guardrails apply above connector-connection policies. By default, an unmatched connector action runs without approval. Set `policy.default_mode: risk` to require approval for unmatched write and destructive actions. Set `sensitive: true` to make `require_approval` the connector's unmatched-action default, including reads. Explicit project or connector-connection rules still apply first. ### Approve one governed call `require_approval` creates one decision for one connector call. The Connector returns `202 pending_approval` with `approval_url`, `approval_summary`, and `execution_id`. It does not keep an HTTP request open. Share `approval_url` with any teammate. The URL identifies the request but does not grant authority. The page requires a signed-in TaskStation account. TaskStation then verifies that the account can access and approve actions in the project. The approval page shows the redacted parameters that the connector will receive. Approve or deny the call once. TaskStation sends the decision back into the session through a durable callback. An approval applies only to the exact request digest. A changed recipient, subject, body, channel, URL, or other parameter requires a new decision. Open the session's **Audit** panel to use the same parameter view. Historical entries remain read-only. There is no session-wide approval option. Use an explicit `always_run` policy only when a connector action must run unattended. ## Connect with OAuth ### Open the project's Connectors page Open the project. Select **Connectors**, then select the app. ### Select the connection scope Select **Project** for a shared project connection. Select **User** for a connection owned by the acting project member. ### Complete authorization Complete the OAuth flow. TaskStation stores the connected account as a connection. ## Connect an MCP server that uses OAuth 2.1 TaskStation implements the MCP authorization specification. Open the connector, select **Add credential**, then select the **OAuth 2.0** tab. TaskStation probes the server and reads its metadata: 1. The unauthenticated probe returns `401` with `WWW-Authenticate: Bearer resource_metadata="…"`. 2. TaskStation reads the protected resource metadata (RFC 9728) at that URL, or at `/.well-known/oauth-protected-resource`. 3. TaskStation reads the authorization server metadata (RFC 8414 or OpenID Connect discovery) for the first authorization server the resource names. When that server advertises `registration_endpoint`, TaskStation registers itself as an OAuth client (RFC 7591) and shows one button: **Connect <server>**. You create no application, and you copy no client ID or secret. TaskStation then runs Authorization Code with PKCE (S256) and binds the token to the server with the `resource` parameter (RFC 8707). Register this redirect URI when a server needs one in advance: ```text https://api.taskstation.co/v1/connectors/oauth2/callback ``` When the server publishes endpoints but no `registration_endpoint`, TaskStation prefills the authorization URL, token URL, and scopes. Enter the client ID of an app you create with that provider. When the server publishes no metadata, enter every field. TaskStation keeps the MCP session: it runs `initialize` and `notifications/initialized` on demand, then sends `Mcp-Session-Id` on later calls. A server that answers without a session never sees the handshake. TaskStation binds each connection to the authorization server that issued it. When the callback carries an `iss` parameter (RFC 9207), TaskStation rejects it unless it matches the recorded issuer — a code minted by a different server is refused before it is redeemed. ### Authorize from the CLI The dashboard is one way to run this flow, not the only one. Declare the connector in `taskstation.yaml`, then authorize it from a terminal or an agent session: ```yaml connectors: - slug: read-ai name: Read AI provider: mcp url: 'https://api.read.ai/mcp' auth: type: bearer ``` ```text taskstation connectors authorize read-ai --json ``` The command creates the connection, runs the discovery chain, registers TaskStation as a client when the server supports RFC 7591, and returns the URL to approve: ```json { "connection_id": "7b1a16b2-...", "registered": true, "scopes": ["openid", "offline_access", "mcp:execute", "meeting:read"], "authorization_url": "https://authn.read.ai/oauth2/auth?response_type=code&...", "expires_at": "2026-08-19T14:48:26.345Z" } ``` An agent returns `authorization_url` to the person it is working with. After they approve, the agent confirms: ```text taskstation connectors authorize read-ai --status ``` The command exits non-zero while the status is `error`. Use `--scope` to narrow what is requested. Use `--client-id` and `--client-secret` for a server that does not support dynamic client registration. The same steps are available on the SDK — `discoverConnectionOAuth2Resource`, `registerConnectionOAuth2Client`, `startConnectionOAuth2Authorization`, and `getConnectionOAuth2Status`. TaskStation refetches the connector's tool catalog as soon as authorization completes, so the connector leaves the `error` state without a manual sync. ### Self-hosted: give the box a stable public URL The callback URL is derived from `TASKSTATION_URL`, the public origin of your API. Authorization servers compare `redirect_uri` byte for byte, so the value must be stable. A self-host install started with the zero-config quick tunnel gets a **new** `https://.trycloudflare.com` hostname every time `cloudflared` restarts. The callback URL changes with it. A server that supports dynamic client registration recovers on its own — the next authorization registers a new client against the current URL. A server that needs a pre-registered OAuth app does not: you must update its allowed redirect URI after every restart. Set `CLOUDFLARE_TUNNEL_TOKEN` and `CLOUDFLARE_TUNNEL_HOSTNAME` for a named tunnel, or point `TASKSTATION_URL` at your own domain. Then register one callback URL once: ```text https:///v1/connectors/oauth2/callback ``` ## Connect a direct API with OAuth2 Direct connectors support: - client credentials - authorization code with PKCE - device authorization - dynamic client registration (RFC 7591) For client credentials, enter the token URL, client ID, scopes, and client secret. You can use `client_secret_basic`, `client_secret_post`, `client_secret_jwt`, or `private_key_jwt` token-endpoint authentication. For authorization code, enter the authorization URL and token URL. For device authorization, enter the device-authorization URL and token URL. An RFC 8414 discovery URL can provide these endpoints. TaskStation encrypts the OAuth2 configuration and tokens. It refreshes access tokens before expiry, and it stores each rotated refresh token. Revoking the connection blocks the next connector call. For Microsoft Graph, use: ```text https://login.microsoftonline.com/{tenant_id}/oauth2/v2.0/token https://graph.microsoft.com/.default ``` For direct SharePoint REST calls, use the SharePoint resource scope: ```text https://{tenant}.sharepoint.com/.default ``` ## Connect with an API key ### Declare the connector ```yaml connectors: - slug: stripe-read name: Stripe read access provider: openapi spec: 'https://raw.githubusercontent.com/stripe/openapi/master/openapi/spec3.json' authorization_strategy: project auth: type: bearer policies: - match: 'get_*' action: always_run - match: '*' action: block ``` `authorization_strategy` defaults to `project` when the manifest omits it. ### Merge the change request TaskStation reads the manifest from the default branch. The connector becomes active after the change request merges. ### Add the connection Open the connector and set its credential. TaskStation stores the value encrypted. It does not write the value to the manifest. ## Grant an agent access Add the connector slug to the agent's `connectors` field: ```yaml agents: release-bot: connectors: [stripe-read] connectors_required: [stripe-read] ``` `connectors_required` must be a subset of `connectors`. A session for this agent returns `409 CONNECTOR_CONNECTION_REQUIRED` before sandbox startup when it cannot resolve a valid active connection. Omit `connectors`, and the agent gets `none`. Merge the change before it takes effect. ## Select a session connection Default resolution follows the connector's authorization strategy. A session can select a specific connection: ```json { "connector_bindings": { "stripe-read": { "connection_id": "00000000-0000-4000-8000-000000000000" } } } ``` The binding key is the connector slug. The value is an active connection that matches the connector's strategy. Use `GET /projects/{projectId}/sessions/{sessionId}/scope` to read the effective binding. Use `PUT` on the same path to replace it. The replacement applies to the next tool call without restarting the session. ## Use a connector in a session Inside a session, use the Connector CLI: ```text taskstation connectors ls taskstation connectors call stripe-read '' ``` `connectors` lists the connectors in scope. `call` runs one action. Slack and Microsoft Teams connect from the dashboard the same way as OAuth apps. Connecting Slack writes a `channel` connector to `taskstation.yaml` for you. See [Slack & channels](/docs/connect/slack). --- # Connect & automate Connectors, Slack, computers, and triggers reach outside a project. Canonical page: https://taskstation.co/docs/connect A project reaches outside itself four ways. A connector calls an external tool. A channel starts a session from chat. A computer connects the agent to your own machine. A trigger starts a session with no person present. - [Connectors](/docs/connect/connectors): Give the agent scoped, brokered access to external tools and APIs. - [Slack & channels](/docs/connect/slack): Start and continue a session from a Slack or Teams message. - [Computers](/docs/connect/computers): Connect your own machine to the agent through a permissioned tunnel. - [Triggers](/docs/connect/triggers): Start a session automatically, on a schedule or from a webhook. --- # Slack & channels Connect Slack to a project and control sessions from a channel. Canonical page: https://taskstation.co/docs/connect/slack A channel connects a chat platform to a project. A message in a connected channel starts a [session](/docs/work/sessions). The agent replies in the same thread. ## Live channels TaskStation supports four channel platforms today: - **Slack** — connect with one click through the TaskStation-managed app, or bring your own bot token. - **Microsoft Teams** — connect through org admin-consent OAuth, or bring your own Azure Bot (experimental; enable it under [Settings → Experimental](/docs/feature-flags)). - **Email** — an AgentMail-backed channel (experimental; enable it under [Settings → Experimental](/docs/feature-flags)). - **Voice** — a realtime voice channel on LiveKit (experimental; enable it under [Settings → Experimental](/docs/feature-flags)). A connected call is a LiveKit room; the agent speaks through the live media session. Only Slack and Microsoft Teams connect through the CLI and dashboard. Email and voice are managed from the dashboard or SDK. This page covers Slack. Teams follows the same session, identity-linking, and credential rules. ## How a channel starts a session The first message in a Slack thread or Teams conversation creates a session. Every later message in that thread goes to the same session, even after the sandbox stops and resumes. A Slack workspace connected to more than one project shows a project picker on the first mention. ## Identity linking TaskStation links each chat sender to a TaskStation account before the agent runs for them. Run `/taskstation login` in Slack and follow the link to sign in. Teams uses the same requirement. An unlinked sender gets a prompt to link instead of a session run. ## Where credentials live A connected channel's bot token is a connector-scoped [secret](/docs/project/secrets). It does not appear on the project's Secrets page. TaskStation never injects it into a sandbox. TaskStation resolves the token server-side when the agent sends a chat message. ## Connect Slack ### Start the connection From the CLI, run: ``` taskstation channels connect ``` On a host with the shared Slack app configured (TaskStation Cloud, for example), this prints a one-click install link. Open the link and pick the workspace. Click **Allow**. Add `--wait` to make the command poll until the install lands. From the dashboard, open the project's Channels page and connect Slack there instead. It uses the same install flow. ### Use your own Slack app On a self-hosted deployment with no shared Slack app configured, `taskstation channels connect` falls back to manual mode automatically. Run `taskstation channels manifest` to print an app manifest. Create the app at `api.slack.com/apps` from that manifest. Install the app to your workspace. Then run: ``` taskstation channels connect --manual --bot-token xoxb-... --signing-secret ... ``` ### Check the connection ``` taskstation channels status ``` Run `taskstation channels disconnect` to remove the connection. See the [CLI reference](/docs/cli) for every flag. Connecting Slack adds a `channel` [connector](/docs/connect/connectors) entry to `taskstation.yaml` for you. You never write this entry by hand. ## Use Slack Mention the bot with a task, in any channel it has joined, or open a direct message. The thread keeps its session as described above. A bare mention with no task gets a reminder to add one instead of starting a session. ## Control a session with slash commands Type these as `/taskstation ` in Slack, or as plain text in a DM. Slash commands do not run inside the Assistant DM pane, so TaskStation parses the same words from plain text there. | Command | What it does | |---|---| | `login`, `logout` | Link or unlink your Slack identity | | `switch`, `unbind` | Rebind or unbind this channel from its project | | `projects` | List projects you can bind to | | `sessions` | List recent sessions started in this workspace | | `whoami` | Show the channel panel: project, agent, model, policy, and your linked identity | | `agent `, `model ` | Set the agent or model this channel uses | | `policy ` | Set who can start sessions here | | `help` | List all commands | `policy` accepts `project_open` (default — any project member who mentions the bot gets a session), `owner_approval`, or `owner_only`. This is a per-resource setting on one channel, not a role. It grants no permission the role verdict denies — see [Accounts & access](/docs/accounts#per-feature-access-settings). ## What does not work - Slash commands do not fire inside the Assistant DM pane. Type `/taskstation ` as plain text there instead. - Other chat platforms, including Telegram, are not supported channels today. Only Slack and Microsoft Teams connect through the CLI and dashboard; email and voice connect through the dashboard or SDK. --- # Triggers A trigger starts a session on a schedule, from a webhook, or from a monitor. Canonical page: https://taskstation.co/docs/connect/triggers A trigger starts a [session](/docs/work/sessions) with no person present. Use a trigger to automate recurring or event-driven work. ## Trigger types TaskStation supports three trigger types. - **cron** — runs on a schedule you set. - **webhook** — runs when an external service sends a signed request to the project's webhook URL. - **monitor** — runs a command from your repository 24/7. Each line the command prints to stdout fires the trigger. Experimental: see [Monitors](#monitors). You define triggers in the project manifest, `taskstation.yaml`. Each trigger holds a prompt that renders as the fired session's first message. Runtime state — such as the last fire time and status — lives outside the manifest, in the database. Firing a trigger does not create a commit. Creating, updating, or deleting a trigger through the API, SDK, or dashboard writes directly to the default branch. It does not go through a [change request](/docs/work/change-requests) (CR). Editing `taskstation.yaml` inside a session and running `taskstation ship` follows the normal branch and CR flow instead. ## Set up a cron trigger ### Add the trigger ```sh taskstation triggers add daily-digest --type cron \ --cron "0 0 9 * * 1-5" --timezone America/Los_Angeles \ --prompt "Summarize yesterday's activity and save it as a daily note." ``` `cron` is a 6-field expression: second, minute, hour, day, month, weekday. This command edits your local `taskstation.yaml` only. ### Ship it ```sh taskstation ship ``` `taskstation ship` commits `taskstation.yaml` and pushes it. The schedule goes live once this lands on your project's default branch. ### Confirm it runs ```sh taskstation triggers ls ``` The list shows each trigger's slug, state, and when it last fired. To fire it now instead of waiting for the schedule, run `taskstation triggers fire daily-digest`. ## Set up a webhook trigger A webhook trigger needs a secret. TaskStation uses it to check the signature on every incoming request. ### Add the secret ```sh taskstation secrets set WEBHOOK_SECRET= ``` See [Secrets](/docs/project/secrets) for more on secrets. ### Add the trigger ```sh taskstation triggers add new-lead --type webhook \ --secret-env WEBHOOK_SECRET \ --prompt "A new lead arrived: {{ body.name }} ({{ body.email }}). Add it to the CRM." ``` `--secret-env` names the secret that signs requests to this trigger. ### Ship it ```sh taskstation ship ``` ### Send it a request TaskStation builds the webhook URL from your project id and the trigger's slug: ``` POST /v1/webhooks/projects// ``` Send a signed `POST` request to this URL from the external service. TaskStation checks the signature against `WEBHOOK_SECRET`, then starts a session with the request body available in the prompt. See [Webhook signature](#webhook-signature) below for the exact header and format. By default, each fire starts a fresh session on a new branch. A trigger can instead reuse or pin a session, and a webhook trigger can filter which payloads start one — see [Session strategy](#session-strategy) and [Payload templating](#payload-templating) below. ## Monitors **Experimental.** Monitors run only where the `monitors` feature flag is on. The flag is off by default. While it is off, the platform provisions no monitor box and fires no monitor event. A monitor watches something that neither pushes webhooks nor fits a schedule: a live log, a queue depth, a page that changes, a price. TaskStation runs your command 24/7 in the project's monitor box — one persistent microVM per project, with the same isolation boundary and the same project secrets a session sandbox gets. Deterministic code watches; the agent wakes only when the command emits a line. Four rules define the contract: - **Stdout lines are events. Nothing else is.** Stderr is diagnostics: visible in the monitor's logs, never fires. - Each line fires the trigger exactly once, through the path a webhook already uses: `filter` → prompt template → `session_mode`. - **A monitor cannot fail silently.** Process exit, restart-budget exhaustion, and silence longer than `expect_event_within` each fire a platform-written lifecycle event in the same stream. - `session_mode` defaults to `reuse` on a monitor, not `fresh`. A monitor fires repeatedly by design, so `fresh` would mint one session per event. ### Add the monitor ```sh taskstation triggers add checkout-errors --type monitor \ --run "./monitors/checkout-errors.ts" \ --mode poll --interval 60s --expect-event-within 24h \ --prompt "Checkout monitor emitted: {{ line }}" ``` `--mode poll` re-runs the command every `--interval` and expects it to exit. `--mode stream` runs it once and keeps it alive; a stream takes no `--interval`. Both shapes produce lines, and nothing downstream can tell them apart. `cron`, `run_at`, `timezone`, and `secret_env` are rejected on a monitor. ### Ship it ```sh taskstation ship ``` The platform starts the monitor box once this lands on the default branch. ### Confirm it runs ```sh taskstation triggers ls taskstation triggers info checkout-errors ``` `ls` shows a monitor's mode and interval where a cron shows its schedule. `info` shows `run`, `mode`, `interval`, and `expect_event_within`. ### Monitor limits Every bound below is enforced by the platform. | Bound | Value | | --- | --- | | Monitors per project | 10 enabled | | Poll interval | at least 30s | | `expect_event_within` | at least 5m | | Event rate per monitor | 60/hour sustained, burst 30. Overflow suppresses the monitor for 10 minutes; 3 suppressions in 24 hours disables it. | | Line length | 8 KiB, truncated with a `truncated: true` marker | | Restart budget | 5 restarts / 10 minutes, then a `restart_budget_exhausted` lifecycle event and 15-minute backoff | | Event retention | 30 days | | Monthly box budget | $75 by default. Past it, the box stops and one `budget_exceeded` lifecycle event fires. | The monitor box needs a provider that supports a persistent sandbox. Where the project's provider cannot hold one, the `monitors` flag reports itself unavailable. ## Config shape ```yaml # taskstation.yaml triggers: - slug: daily-digest # required, lowercase + dashes, unique per project name: Daily digest # optional, defaults to slug type: cron # "cron" | "webhook" | "monitor", required agent: taskstation # optional, defaults to "default" model: anthropic/claude-sonnet-4-5 # optional, resolves at fire time if unset enabled: true # optional, default true cron: "0 0 9 * * 1-5" # 6-field expression, mutually exclusive with run_at timezone: America/Los_Angeles # IANA name, default UTC session_mode: reuse # "fresh" | "reuse" | "pinned" | "keyed", default "fresh" filter: # optional, webhook payload guard "body.data.direction": "inbound" prompt: "Summarize {{ body.text }}" # required, template string ``` A monitor replaces the schedule fields with its command and shape: ```yaml # taskstation.yaml triggers: - slug: checkout-errors type: monitor run: ./monitors/checkout-errors.ts # required, repo-relative command mode: poll # "poll" | "stream", required interval: 60s # required on poll, invalid on stream expect_event_within: 24h # optional silence watchdog agent: oncall session_mode: reuse # the monitor default filter: # optional, same guard a webhook uses "line.severity": "error" prompt: "Checkout monitor emitted: {{ line }}" ``` Legacy `taskstation.toml` uses the same fields in a different container; see [legacy TOML](/docs/project/legacy-toml). ## Fields | Field | Required | Default | Notes | | --- | --- | --- | --- | | `slug` | yes | — | `[a-z0-9][a-z0-9_-]{0,127}`, unique per project. | | `type` | yes | — | `cron`, `webhook`, or `monitor`. | | `prompt` | yes | — | Template string. Renders as the session's first message. | | `name` | no | `slug` | Human label. | | `agent` | no | `default_agent` | Must name a key in `agents:`. Omit it to use `default_agent`; do not write the literal `default`. | | `model` | no | resolves at fire time | Wire form `provider/model`, for example `anthropic/claude-sonnet-4-5`. | | `enabled` | no | `true` | When `false`, the scheduler and the webhook receiver skip the entry. | | `session_mode` | no | `fresh`, or `reuse` on a monitor | See [Session strategy](#session-strategy). | | `session_id` | required for `pinned` | — | Exact session to re-prompt. | | `session_key` | required for `keyed` | — | Template string. Setting it alone implies `session_mode: keyed`. | | `filter` | no | — | Dotted path → expected string. A webhook delivery that does not match returns `200` and fires no session. | | `cron` | one of `cron`/`run_at`, on `type: cron` | — | 6-field expression: second minute hour day month weekday. | | `run_at` | one of `cron`/`run_at`, on `type: cron` | — | ISO-8601 timestamp. Fires once, then stays dormant. | | `timezone` | no, cron only | `UTC` | IANA name. | | `secret_env` | required, webhook only | — | Name of a project secret holding the signing key. The secret must use `broker` delivery with the `connector` consumer. | | `run` | required, monitor only | — | Repo-relative command whose stdout lines are the events. One line, at most 1024 characters. | | `mode` | required, monitor only | — | `poll` re-runs `run` every `interval`; `stream` runs it once and keeps it alive. | | `interval` | required for `mode: poll` | — | Duration literal (`30s`, `5m`, `24h`, `7d`), minimum `30s`. Invalid on `mode: stream`. | | `expect_event_within` | no, monitor only | — | Duration literal, minimum `5m`. Silence longer than this fires a lifecycle event. | A cron trigger needs `cron` or `run_at`, never both. A webhook trigger without `secret_env` is rejected — there is no unauthenticated webhook. A monitor needs `run` and `mode`, and rejects `cron`, `run_at`, `timezone`, and `secret_env` outright — a manifest that claims a schedule the monitor runner never reads is a lie. Configure the signing secret before you create or update a webhook trigger: ```bash taskstation secrets set WEBHOOK_SECRET=- taskstation secrets delivery WEBHOOK_SECRET broker --consumer connector ``` Pass the value to the first command on standard input. If the secret previously used sandbox delivery, rotate it after the delivery change because an existing sandbox can retain the previous value. ## Session strategy `session_mode` controls which session a fire re-prompts. TaskStation tries the modes below in order and falls through on failure at each step. 1. **`pinned`** — re-prompt the exact `session_id`. If that session is gone or failed, fall through. 2. **`keyed`** — render `session_key` against the payload, then look up the most recent non-failed session previously stamped with that exact key. If the key renders empty, or no session matches, fall through to a fresh session. It never falls through to another key's session. 3. **`reuse`** — re-prompt the most recent non-failed session this trigger previously created. A pinned trigger falls back here too, before falling further. 4. **`fresh`** — create a new sandbox and branch. This is the default, and the final fallback for every mode. The new session becomes the trigger's session for future `reuse` and `keyed` fires. ## Session access Sessions a trigger creates use the private policy by default. The trigger agent's service account owns them — an agent's identity is a `service_account` principal, so it can own a session and hold assignments like any other principal. A project manager can always open them. An account owner and an account admin hold manager-equivalent access on every project, so the same applies to them. The person who configured or manually fired the trigger does not gain access through that action unless they hold one of those roles. Trigger settings offer three policies: - **Trigger agent and project managers** — no ordinary project member can open the session. - **Selected teammates** — the trigger agent, project managers, and selected project members or account groups. - **Whole project** — every project member. This policy is a per-resource visibility setting on top of the role model, not a role. It decides who can open one trigger's sessions. It grants no permission the role verdict denies. See [Accounts & access](/docs/accounts#per-feature-access-settings). The access policy is account-local runtime state. It does not enter the portable `taskstation.yaml` manifest because principal ids belong to one account. Use the dashboard or SDK `session_access` field to configure it. Updating only this policy creates no Git commit. Saving a policy also updates prior sessions created by that trigger. A pinned session keeps its own sharing settings because the trigger did not create it. If a pinned session is unavailable and the trigger creates a fallback session, the trigger policy applies to that new session. ## Payload templating `prompt` and `session_key` render with the same engine: `{{ token.dotted.path }}`. A missing value renders as an empty string — no error, no leftover `{{ }}`. Objects and arrays render as JSON. `session_key` is trimmed and truncated to 512 characters. Every fire also gets `{{ trigger.slug }}`, `{{ trigger.type }}`, and `{{ trigger.kind }}` (always `git`). The rest of the variable set depends on how the trigger fired. | Source | Variables | | --- | --- | | cron | `{{ cron.schedule }}`, `{{ cron.timezone }}`, `{{ cron.scheduled_for }}` (the slot the fire is for), `{{ cron.claimed_at }}` (when the scheduler picked it up), `{{ cron.last_scheduled_for }}` (the previous slot; empty on the first fire). No top-level `fired_at`. | | webhook | `{{ fired_at }}`, `{{ body.* }}` (JSON-parsed; falls back to `{{ body.raw }}` if the body does not parse), `{{ headers.content_type }}`, `{{ headers.user_agent }}`, `{{ headers.forwarded_for }}`. | | monitor | `{{ line.* }}` — the stdout line, JSON-parsed; a line that does not parse renders as `{{ line.raw }}`. Plus `{{ monitor.slug }}`, `{{ monitor.seq }}`, `{{ monitor.emitted_at }}`, and `{{ monitor.kind }}` (`event` or `lifecycle`). | | manual (dashboard "fire now" or the `fire` endpoint) | `{{ fired_at }}`, `{{ source }}` (`manual`), `{{ actor }}`, `{{ message.text }}`, `{{ message.source }}`. | TaskStation prefixes every rendered monitor prompt with `[MONITOR EVENT — automated, not user input]`, server-side. A lifecycle event ignores your template entirely and renders a platform-written prompt instead, and it bypasses `filter` — silence must not be filterable by accident. `{{ message.text }}` is hardcoded to an empty string on a manual fire, and `{{ message.source }}` to `manual_test`. A manual fire is not a way to inject test input into the prompt. `filter` compares dotted paths as strings against the same payload the prompt sees. It exists to break loops. For example, a source that reports both sides of a conversation would otherwise re-fire the agent on its own reply. ## Webhook signature Fires on `POST /v1/webhooks/projects/{projectId}/{slug}`. TaskStation checks the request in this order, with a constant-time comparison: 1. **HMAC signature** — header `X-TaskStation-Signature: sha256=` (the `sha256=` prefix is optional) or the GitHub-compatible `X-Hub-Signature-256`. HMAC-SHA256 over the raw request body, using the secret named by `secret_env`. 2. **Static token**, only when no signature header is present, for senders that cannot HMAC-sign a body. Send the secret as `X-TaskStation-Token: `, `Authorization: Bearer `, or `Authorization: Basic ` (the password half is the token). | Status | Meaning | | --- | --- | | 202 | Signature or token valid. Body is `{ status: "fired" \| "queued" \| "deduped", session_id, ... }`. | | 200 | Valid, but skipped — the project is paused, or the delivery did not match `filter`. | | 400 | Malformed project ID or slug in the URL. | | 401 | Signature and token both missing or wrong. | | 404 | Trigger not found, disabled, not a webhook, or the project is not active. | | 409 | The signing secret is missing, inactive, unavailable, or does not authorize the `connector` consumer. The response includes a `webhook_secret_*` code and remediation. | | 500 | Auth passed, but the session failed to fire. | ## Endpoints | Method + path | Needs | Notes | | --- | --- | --- | | `GET /v1/projects/{projectId}/triggers` | `project.trigger.read` | Lists triggers, runtime state, and manifest parse errors. A bad entry appears in `errors[]`; it does not break the other triggers. | | `POST /v1/projects/{projectId}/triggers` | `project.trigger.create` | Creates a trigger. Commits to the manifest directly. | | `PATCH /v1/projects/{projectId}/triggers/{slug}` | `project.trigger.update` | Partial update, merged onto the current entry. | | `DELETE /v1/projects/{projectId}/triggers/{slug}` | `project.trigger.delete` | Also clears the trigger's runtime state. | | `PATCH /v1/projects/{projectId}/triggers/activation` | `project.trigger.update` | Body `{ paused: boolean }`. See [Pause and resume](#pause-and-resume). | | `POST /v1/projects/{projectId}/triggers/{slug}/fire` | `project.trigger.fire` | Manual fire. The built-in project `member` role holds `project.trigger.fire`, so an ordinary member can fire a trigger. | | `POST /v1/webhooks/projects/{projectId}/{slug}` | signature or token | Public URL, gated by the webhook secret. | ## Pause and resume A project-level switch stops every trigger in the project at once, independent of each trigger's own `enabled` field. While paused, the scheduler skips the project and inbound webhooks return `200` with `{ status: "skipped" }` — no session fires. A manual fire still works. Use this when the same repository runs on two control planes (for example, dev and production) so cron does not fire twice. CLI: `taskstation triggers pause` and `taskstation triggers resume`. See [CLI](/docs/cli) for the full `taskstation triggers` command group. ## Limits and reliability - The scheduler polls roughly every second (default 1,000 ms; configurable via `TASKSTATION_TRIGGER_SCHEDULER_INTERVAL_MS`). Cron precision is best-effort to the second, even though the expression has a seconds field. - Each project allows 3 triggered sessions provisioning at once, by default. The account's plan-tier active-session cap can also apply. A fire past either limit returns `queued` (`202`) instead of failing, and runs once a slot frees up. - A manual or webhook fire has a 45-second timeout. Loading the manifest has a 30-second timeout. - A cron fire is keyed on the due schedule slot, so a fire that timed out but actually landed does not duplicate on retry. A webhook fire is keyed on the delivery ID header, or a hash of the body and signature when the sender sends no ID. --- # Apps Deploy static sites, bundles, Dockerfiles, and OCI images to stable TaskStation URLs. Canonical page: https://taskstation.co/docs/feature-flags/apps A TaskStation App is a provider-neutral serverless deployment owned by one project. An App owns one stable URL. Each deployment is immutable and numbered. A failed deployment never replaces live traffic. Apps is a [feature flag](/docs/feature-flags). Its stability is **stable**, and it is off by default. Turn it on per project before you deploy. Building on Apps from TypeScript? See [SDK → Apps](/docs/sdk/apps) for the client surface and the React hooks. ## Turn Apps on Open **Settings → Experimental** and switch **Apps** on for the project. You need `project.customize.write`. While the flag is off: - Every Apps route answers `403` with `{ error, code: "feature_disabled", feature: "apps" }`. - `taskstation apps ` prints the same sentence and exits `1`. - The **Apps** entry does not appear in the project sidebar. Opening `/projects//apps` directly shows a gate screen that links to the flag. The page itself never enables the feature. ## Source kinds All four source kinds run on the same TaskStation sandbox hosting backend. TaskStation selects Daytona, Platinum, or E2B. They therefore share one deployment contract and one cold-wake contract. | Kind | Deploy this | TaskStation does | |---|---|---| | `static` | Plain HTML, CSS, JS, or a prebuilt SPA or `dist/` | Serves the files | | `bundle` | A package source | Runs the install and build commands, then serves the output directory | | `dockerfile` | A repository with a Dockerfile | Builds the image, then runs your `command` on your `port` | | `oci_image` | A public image reference | Runs your `command` on your `port` | Pick the fastest path for the result you want: - Build Vite locally and deploy `dist/` as `static` for the lowest latency. - Deploy the package source as `bundle` when TaskStation must run the install and build. - Export Next.js with `output: 'export'` and deploy `out/` as `static` when the App needs no server runtime. - Deploy server-rendered Next.js and arbitrary services as `dockerfile`, with an explicit command and port. - Deploy an existing public image as `oci_image`, with an explicit command and port. `dockerfile` and `oci_image` require `--command` and `--port`. `static` and `bundle` do not. ## Deploy from the CLI ```bash taskstation apps deploy . ``` `deploy` creates the App on first use, registers an immutable artifact — uploading a `.tar.gz` for a path, or recording the reference for `--image` — builds it, and blocks until the stable URL is ready. The wait budget is `--wait-seconds`, default `1200`. Use `--no-wait` only when another process owns status tracking. ```bash taskstation apps deploy dist --slug docs --access project taskstation apps deploy . --type dockerfile --command '["node","server.js"]' --port 3000 taskstation apps deploy --image ghcr.io/acme/service:2026-08-07 --command '["node","server.js"]' --port 3000 ``` The full subcommand list: | Command | What it does | |---|---| | `taskstation apps list` | List the project's Apps. `--json`. | | `taskstation apps create ` | Create an App without deploying it. | | `taskstation apps deploy [path]` | Deploy a directory, a `.tar.gz`, or `--image`. | | `taskstation apps set ` | Change an existing App: `--name`, `--cpu`, `--memory-gb`, `--disk-gb`, `--idle-timeout`, `--budget`. | | `taskstation apps show ` | Show the App and its deployments. `--json`. | | `taskstation apps logs [deployment]` | Read runtime logs. `--after N --limit N`. | | `taskstation apps start ` | Permit requests and start the App. | | `taskstation apps stop ` | Suspend compute now. | | `taskstation apps rollback ` | Move traffic to a ready deployment. | | `taskstation apps access ` | Read or update the access policy. | | `taskstation apps access-link ` | Create a short-lived authenticated browser URL. | | `taskstation apps delete ` | Delete the App and its runtimes. `--yes`. | `--project`, `--host`, and `--json` work on every subcommand. `taskstation apps set` sends only the flags you pass, and needs project write access. `--memory-gb` accepts `--memory` and `--disk-gb` accepts `--disk` as aliases. `--idle-timeout` takes 120-86400 seconds. A machine or budget change applies to the next deployment, not to the running runtime. ### Deployment defaults from `taskstation.yaml` An `apps.` block holds local deploy defaults. The server stays the App control plane: the block never carries access, passwords, or member ids. ```yaml apps: docs: path: docs/dist type: static spa: true readiness_path: / idle_timeout_seconds: 300 monthly_budget_usd: 5 resources: cpu: 1 memory_gb: 2 disk_gb: 10 env: PUBLIC_BASE: https://taskstation.co secrets: API_TOKEN: docs_api_token ``` Select the block with `taskstation apps deploy --manifest-app docs`. An explicit flag always wins over the block. ### Machine, idle timeout, and budget | Setting | Default | Bounds | |---|---|---| | `cpu` | `1` | `1` to `32` cores | | `memory_gb` | `2` | `1` to `128` GiB | | `disk_gb` | `10` | `1` to `500` GiB | | `idle_timeout_seconds` | `300` | `120` to `86400` | | `monthly_budget_usd` | `5` | `0` to `100000`, or the operator's `TASKSTATION_APPS_MAX_MONTHLY_BUDGET_USD` | Apps reject a machine larger than the limits instead of clamping it. An App records its requested spec and bills off that record, so a silent downgrade would charge for compute the provider never gave. An out-of-range value answers `400` with `code: "app_machine_out_of_range"` or `"app_budget_out_of_range"`. Creating an App past the account's App quota answers `402` with `code: "app_quota_exceeded"`. A duplicate slug in the same project answers `409`. ## The stable URL TaskStation assigns the hostname when the App is created and never changes it. On TaskStation cloud it is `--.apps.taskstation.co`. Self-hosted deployments serve their own wildcard domain from `TASKSTATION_APPS_BASE_DOMAIN`. An authorized request to a suspended App resumes its sandbox, waits for readiness, and proxies that same request. You do not have to wake it first. While the App is waiting for its first deployment, queued, validating, building, provisioning, checking, activating, or starting: - A browser navigation gets a branded status page, HTTP `202`, `retry-after: 3`, and a 3-second meta refresh. - A machine client gets `202` and JSON — for a cold start, `{ code: "app_starting" }` with `retry-after: 3`. Terminal and paused states answer differently: | State | HTTP | `code` | |---|---|---| | Deployment failed | `503` | `app_deployment_failed` | | Deployment cancelled | `503` | `app_deployment_cancelled` | | Monthly compute budget reached | `402` | `app_budget_exceeded` | | Account cannot start compute | `402` | `app_account_unfunded` | | Account at its concurrent-App limit | `429` | `app_concurrency_limit` | `taskstation apps stop` suspends compute immediately; the next authorized request resumes the App. `taskstation apps start` warms it before traffic arrives. > **Cold starts stay invisible** > The stable URL never exposes an `app_stopped` state. A provider edge that > answers `502` during the first request after a resume is served as the ordinary > cold-start page instead. A warm App owns its own HTTP status, including a > deliberate application `502`. ## Access modes An App's access mode is a per-resource visibility setting on top of the role model, not a role. It decides who can open this one App. It grants no permission the role verdict denies. See [Accounts & access](/docs/accounts#per-feature-access-settings). New Apps are private. Choose one mode: | Mode | Who can open the App | |---|---| | `private` | The creator only | | `project` | Every principal who can read the project | | `restricted` | Selected users and groups | | `public` | Anyone, with no authentication | | `password` | Anyone with the App password | Set it at deploy time or afterwards: ```bash taskstation apps deploy . --access restricted --members m1,m2 --groups g1 taskstation apps access docs --mode password --password 's3cret' taskstation apps access-link docs --json ``` TaskStation access uses a five-minute exchange URL and an eight-hour, host-only, secure cookie. `access-link` mints that exchange URL without changing the policy — treat it as a secret. Changing an access policy increments its revision, which revokes every existing App cookie. Passwords are Argon2id hashes; the API, the CLI, and the SDK never return a password or its hash. > **Never put an App password in your repo** > `taskstation.yaml` holds deployment defaults only. Pass a password with `--password`, > or set it from the access modal. Being able to see an App listed and being able to open it are different verdicts. A project manager sees every App in the project, so a private App stays manageable when its creator leaves. An account owner and an account admin hold manager-equivalent access on every project, so the same applies to them. The App record reports `viewer_can_access` for the second question. ## Versions and rollback Each deployment gets the next version number for its App and is immutable. The deployment record keeps its source kind, hosting provider, build and runtime spec, attempt count, and error code. Every deployment also records who made it: `created_by`, `actor_type` (`human`, `agent`, `service_account`, or `system`), and the originating `source_session_id` when an agent deployed it. Move traffic back to any ready deployment: ```bash taskstation apps rollback docs ``` Rollback starts the target deployment's runtime first, then stops the previous one. A target that fails to start leaves the current deployment serving. Each cold start compares the active deployment's runtime version against the current TaskStation App runtime. The old deployment keeps serving while TaskStation asynchronously builds one immutable replacement with the latest `taskstation-appd` and Caddy binaries. A PostgreSQL advisory lock prevents duplicate refreshes. ## The Apps page Once the flag is on, an **Apps** row appears in the project sidebar, under Customize. The page is operational, not a creation surface: it lists the project's Apps with live state, a signed preview of each running App, and the access controls. Deploying is `taskstation apps deploy .`. An App with no active deployment reads **Not deployed**, never **Running**. A suspended App's preview issues the request that wakes it. TaskStation opens `*.apps.taskstation.co` and `*.apps.localhost` on their direct origin rather than through a session's web forward proxy. That preserves the host-only access cookie and removes one network hop. --- # Feature flags Turn a TaskStation surface on for one project, and read what every flag gates. Canonical page: https://taskstation.co/docs/feature-flags A feature flag turns one TaskStation surface on for one project. Any surface can ship behind a flag — experimental, beta, or fully stable. "Experimental" is a stability badge on a flag, not the name of the system. Flags are per project. Turning a flag on in one project changes nothing in another project, and nothing for other accounts. - [Apps](/docs/feature-flags/apps): Deploy static sites, bundles, Dockerfiles, and OCI images to stable URLs. ## Turn a flag on 1. Open **Settings → Experimental**. The tab lists every flag the platform supports. 2. Read the row: the flag name, its stability badge, one sentence of description, and its origin — `Default on`, `Default off`, or `Overridden for this project`. 3. Use the switch. The change applies to the current project immediately. You need the project's `project.customize.write` permission. The route answers `403` for any other caller. ### From the CLI The same switches are available to scripts and agents through `taskstation projects features`: ```bash taskstation projects features # every flag: key, state, origin, stability taskstation projects features enable apps taskstation projects features disable voice taskstation projects features reset apps # drop the override; follow the platform default taskstation projects features --json # the full catalog as JSON ``` Add `--project ` to act on a project other than the linked/default one. A flag the platform marks unavailable stays off whatever the project override says; the CLI prints that as `n/a` / `unavailable`. > **Flag state is not in your repo** > Per-project flag state lives in the database, on the project row. It is never > read from `taskstation.yaml`. A flag you turn on does not travel with a repository > clone or a fork. ## The two gates Each flag has two gates. They answer different questions. | Gate | Question | Effect when false | |---|---|---| | `available` | Does this deployment support the flag at all? | The toggle is hidden and the surface stays dark, whatever the project chose. | | `enabled` | Is the flag on for this project? | The surface stays dark for this project. | `enabled` is the project's explicit choice over the platform default, then AND-gated by `available`. `enabled` therefore always implies `available`. `available` is an operator decision, made by the environment the API runs in. Three flags read it from configuration; every other flag is always available: | Flag | Available when | |---|---| | `agent_tunnel` | `TUNNEL_ENABLED` is on | | `llm_gateway` | `LLM_GATEWAY_ENABLED` is on | | `monitors` | `PLATINUM_API_KEY` is set | ## Stability badges The badge describes the contract, not the switch. A `stable` flag is still an opt-in: `apps` is stable and still off by default. | Badge | What it means | |---|---| | Experimental | The surface and its contract can still change. | | Beta | The surface works and the shape is settling. | | Stable | The contract holds. The flag stays an opt-in. | ## How a flag is enforced Every flag declares one enforcement mode. | Mode | What the server does when the flag is off | |---|---| | `routes` | The HTTP surface rejects the request with `403`. | | `behavioral` | The behavior does not occur — no connector materializes, no env injects, no agent registers. | | `ui-only` | The server deliberately does not enforce. The flag hides client surface only. | A `routes` rejection is identical everywhere: ```json { "error": "Apps is not enabled for this project. Enable it in Settings → Feature flags.", "code": "feature_disabled", "feature": "apps" } ``` The `error` string names the flag list, not a specific tab. The list is the **Experimental** tab of Settings. Branch on `code`, never on the message text. The SDK exports `isFeatureDisabledError(error)` and `featureDisabledKey(error)` for exactly this. ## Every flag Registry order — the same order **Settings → Experimental** shows. | Key | Name | Stability | Default | Enforcement | |---|---|---|---|---| | `marketplace` | Marketplace | Beta | On | `routes` | | `agent_tunnel` | Agent Computer Tunnel | Experimental | Off | `ui-only` | | `connectors_api_discover` | Connectors API Discover | Experimental | Off | `routes` | | `agentmail_email` | AgentMail Email | Experimental | Off | `routes` | | `teams` | Microsoft Teams | Experimental | Off | `routes` | | `voice` | Voice | Experimental | Off | `behavioral` | | `llm_gateway` | LLM Gateway | Experimental | On (operator can default off) | `behavioral` | | `review_center` | Review Center | Experimental | Off | `routes` | | `meta_agent` | Meta Agent | Experimental | Off | `behavioral` | | `apps` | Apps | Stable | Off | `routes` | | `monitors` | Monitors | Experimental | Off | `routes` | | `warm_sessions` | Warm Sessions | Beta | On | `routes` | `llm_gateway` reads its per-project default from `LLM_GATEWAY_DEFAULT_ENABLED`, which defaults to on. Turning the flag off per project is a first-class path: the project runs native OpenCode model management (provider keys injected into the sandbox, native `provider/model` refs). An explicit project choice always wins. ### What each flag gates - **`marketplace`** — browse and install skills from community and vendor registries. - **`agent_tunnel`** — let agents reach a local machine over a permissioned reverse tunnel. See [Computer Tunnel](/docs/connect/computers). - **`connectors_api_discover`** — browse direct API, MCP, GraphQL, CLI, and Postman surfaces beside Pipedream OAuth apps. See [Connectors](/docs/connect/connectors). - **`agentmail_email`** — assign AgentMail inbox connections so inbound email starts and continues sessions. - **`teams`** — connect a Microsoft Teams bot so chats and channels start and continue sessions. See [Slack & channels](/docs/connect/slack). - **`voice`** — give the agent a live voice call it can start and hold. See [Slack & channels](/docs/connect/slack). - **`llm_gateway`** — route the project through the managed TaskStation LLM gateway. See [Models](/docs/project/models). - **`review_center`** — one inbox for change requests, approvals, and agent output. - **`meta_agent`** — add a platform-owned coordinator agent that spawns and manages specialized sessions. - **`apps`** — deploy static sites, bundles, Dockerfiles, and OCI images to stable serverless URLs. See [Apps](/docs/feature-flags/apps). - **`monitors`** — run 24/7 watchers from your repo that fire trigger events into sessions. See [Triggers](/docs/connect/triggers). - **`warm_sessions`** — keep one sandbox booted while a project is open, so a new session starts without a cold boot. ## Side effects of a toggle Some flags converge platform state after the write commits. | Flag | Effect after the toggle | |---|---| | `voice`, `teams`, `agentmail_email` | TaskStation re-runs channel-connector materialization, so the connector appears or disappears with the flag. | | `agent_tunnel` | TaskStation re-syncs the account's computer connectors. | | `llm_gateway` | TaskStation propagates the new provider mode to active sandboxes. | Effects are convergence work, not part of the toggle's success. The API response does not wait for them. Each effect is retried once, and the reconcilers behind it are idempotent and re-run on their periodic sweeps. ## Read and set a flag from code Read the effective per-project state through the project detail, or through the React hook: ```tsx import { useFeatureFlag } from '@taskstation/sdk/react'; function AppsNavItem({ projectId }: { projectId: string }) { const apps = useFeatureFlag(projectId, 'apps'); if (!apps.enabled) return null; return Apps; } ``` `enabled` is `true` only when the server said exactly `true`. A missing project id, an in-flight query, and an error all resolve to `false`. Gate fail-closed. Set the project override with the client: ```ts const p = taskstation.project(projectId); await p.updateFeatureFlag('apps', true); // turn it on for this project await p.updateFeatureFlag('apps', null); // clear the override, inherit the default ``` `updateFeatureFlag` calls `PATCH /v1/projects/:id/features`. `feature` is one of `FEATURE_FLAG_KEYS`, exported from `@taskstation/sdk` and typed as `FeatureFlagKey`. See [SDK reference](/docs/sdk/reference). Handle a disabled feature by code, not by message: ```ts import { featureDisabledKey, isFeatureDisabledError } from '@taskstation/sdk'; try { await taskstation.project(projectId).apps.list(); } catch (error) { if (isFeatureDisabledError(error)) { console.log(`${featureDisabledKey(error)} is off for this project`); } } ``` --- # Self-hosting architecture How the self-hosted Docker Compose stack fits together, on the box and off it. Canonical page: https://taskstation.co/docs/host/architecture This page shows how the pieces of a self-hosted TaskStation instance fit together: one Docker Compose stack, plus the compute that stays outside it. For install steps, see the [self-hosting guide](/docs/host). Self-hosted TaskStation is one generic Docker Compose system, not a family of deployment targets. `taskstation self-host init` renders a `docker-compose.yml` and `.env` file (plus a `Caddyfile` and `updater.sh` when you set a domain) into `~/.config/taskstation/self-host//`. `taskstation self-host start` runs `docker compose up`. The same artifact runs on a laptop, a VPS, or any cloud VM. A domain is only the `TASKSTATION_DOMAIN` environment variable, not a different setup. Production self-hosting needs a persistent domain pointed at the box. A domain gives Caddy a stable name for ACME TLS, and gives agent sandboxes a stable URL to call back to. Without a domain or a tunnel, sessions cannot run, because the sandbox has no way to reach the API. For evaluation without a domain, use `taskstation self-host init --tunnel cloudflare` instead. ## One box, one Compose stack ```mermaid flowchart TB subgraph internet["Internet"] user["Browser / API client"] end subgraph box["One host: laptop, VPS, or cloud VM"] subgraph compose["docker compose (one project per instance)"] caddy["Caddy\n(only when TASKSTATION_DOMAIN is set)\nACME TLS on 80/443"] frontend["frontend"] api["taskstation-api"] gateway["llm-gateway"] updater["taskstation-updater\n(pull -> migrate -> roll,\nonce daily at a fixed time)"] subgraph supabase["Supabase Docker distribution"] kong["supabase-kong"] auth["supabase-auth"] rest["supabase-rest"] storage["supabase-storage"] db[("supabase-db (Postgres)")] end end vol_db[("bind mount:\nvolumes/db/data")] vol_storage[("bind mount:\nvolumes/storage")] end subgraph external["Outside the box"] daytona["Daytona\n(agent sandboxes, default)"] registry["docker.io/taskstation/*\n(image registry)"] end user -->|"80/443, TLS"| caddy user -.->|"no domain: local ports"| frontend caddy -->|"/v1/llm*"| gateway caddy -->|"else"| api caddy -->|"Supabase data-plane paths"| kong caddy -->|"else"| frontend frontend --> api api --> gateway api --> kong kong --> auth kong --> rest kong --> storage auth --> db rest --> db storage --> db db --> vol_db storage --> vol_storage updater -->|"docker compose pull"| registry updater -->|"migrate, then roll"| compose api -->|"provision and run sessions"| daytona ``` ## What runs on the box - **Caddy** — reverse proxy and ACME TLS. TaskStation renders this service only when you set `TASKSTATION_DOMAIN`; a domain-less instance never opens ports 80/443. Caddy routes `api.` to the gateway (for `/v1/llm*`) or the API, and `` to Kong (for Supabase data-plane paths) or the frontend. - **`taskstation-api`, `llm-gateway`, `frontend`** — the three application images. They track the same channel, or a version you pin explicitly. - **The Supabase Docker distribution** — Kong, GoTrue auth, PostgREST, Storage, Realtime, Studio, imgproxy, meta, functions, and the Supavisor connection pooler. TaskStation vendors this from upstream Supabase and pins every image by digest. - **`taskstation-updater`** — a small container with the Docker socket mounted. It checks for a new image once a day, at a fixed local clock time (`TASKSTATION_UPDATE_TIME`, default `02:00`, in `TASKSTATION_UPDATE_TZ`, default `America/New_York`). If an image changed, it runs the `taskstation-migrate` job, then starts new containers before it stops the old ones. This start-first swap is zero-downtime only when the box runs two replicas (domain mode). A single-replica box (tunnel or local mode) uses a different, brief-downtime swap instead. - **Data** — two bind mounts under the instance directory: `volumes/db/data` for Postgres and `volumes/storage` for Supabase Storage. The `.env` file holds every secret. ## What runs outside the box - **Agent sandboxes** — by default, Daytona. You can configure Platinum or E2B instead. `taskstation-api` reaches the sandbox provider over egress; sandbox compute never runs on the self-host box. - **The image registry** — `docker.io/taskstation/*`. The updater and `taskstation self-host start` pull from it. It needs no credentials. > **Warn** > `taskstation self-host uninstall` runs `docker compose down --volumes > --remove-orphans` and deletes the instance directory. This removes your > database and storage bind mounts. Back them up first. ## Channels and updates Every instance tracks one of two moving tags, or a version you pin explicitly: | Channel | Meaning | |---|---| | `stable` (default) | Curated. A human promotes a proven version to `stable` on a separate schedule from prod releases. | | `latest` | Every prod release retags `latest` automatically. | | `--tag ` | Pins an exact version. Overrides the channel. | `taskstation-updater` and `taskstation self-host update` (alias `reconcile`) resolve the same way: an explicit pin wins, otherwise the configured channel. Self-hosted instances only consume images this pipeline has already built; they never build or sign anything themselves. See the [self-hosting guide](/docs/host) for install steps and the [CLI reference](/docs/cli) for the full `taskstation self-host` command surface. --- # Self-hosting Run your own TaskStation instance with Docker Compose, on a VPS or for evaluation. Canonical page: https://taskstation.co/docs/host TaskStation runs as one Docker Compose stack: the frontend, the API, the LLM gateway, and the Supabase distribution. This page shows the three ways to install it, how updates work, and how to back up your data. Agent sessions run on a separate sandbox provider, not on this stack. The default is [Daytona](https://www.daytona.io/); Platinum and E2B are also supported. ## One-shot bootstrap On a bare Linux box, one command installs Docker, installs the `taskstation` CLI, and starts the stack: ```sh curl -fsSL https://raw.githubusercontent.com/melihyolacan/suna/main/scripts/taskstation-selfhost-up.sh \ | bash -s -- --domain taskstation.example.com --email ops@example.com ``` This script runs on Linux only. On another OS, install the CLI directly and use the manual path below. ## Manual path ### Install the CLI ```sh curl -fsSL https://taskstation.co/install | bash ``` ### Point DNS, then initialize Create an A/AAAA record for your domain and for `api.`, both pointing at the box's IP. Open ports 80 and 443 — the bundled Caddy proxy uses them to issue a TLS certificate. Then run: ```sh taskstation self-host init --domain taskstation.example.com ``` ### Start the stack ```sh taskstation self-host start ``` Check `taskstation self-host status`, `logs`, and `doctor` while the stack starts. ## Evaluation mode To try TaskStation with no domain, use a Cloudflare tunnel instead of a domain: ```sh taskstation self-host init --tunnel cloudflare taskstation self-host start ``` The tunnel URL changes on every restart. Use this mode for evaluation, not production. After the stack starts, set your sandbox provider key: ```sh taskstation self-host configure ``` `configure` is an interactive prompt for the sandbox provider key, and optionally a managed-git token. Sign up in the dashboard, then connect your own LLM key in the model picker. Self-hosted instances use your own key by default. > **Info** > By default, only the platform admin can create new organization accounts. The > platform admin is the super-admin flag on a membership, not a role. > Any signed-in user can still join by invite or SSO. Opt out with > `taskstation self-host init --no-restrict-account-creation`, or re-enable the > admin-only default with `--restrict-account-creation`. Each API container has a 640 MiB memory limit by default. Keep this default on an 8 GiB host. Use a 1 GiB limit on a 16 GiB host when API traffic reaches the default limit: ```sh taskstation self-host env set TASKSTATION_API_MEMORY_LIMIT=1024m ``` Confirm the applied limit with `docker stats --no-stream`. ## Updates Every instance updates itself automatically. Pin an exact version instead: ```sh taskstation self-host update --tag 0.9.84 ``` Turn the updater off with `--auto-update off`. See [Self-hosting architecture](/docs/host/architecture) for the update schedule, the zero-downtime swap, and the channels. ## Backups TaskStation has no separate backup system. Each instance stores its data as two directories under `~/.config/taskstation/self-host//`: `volumes/db/data` (the Postgres database) and `volumes/storage` (file storage). The instance's `.env` file holds every secret and signing key it uses. Back up all three before you run a destructive command. > **Warn** > `taskstation self-host uninstall` stops the stack, deletes its containers and > volumes, and deletes the instance directory. This cannot be undone. ## Learn more - [Self-hosting architecture](/docs/host/architecture) — how the stack fits together. - [CLI reference](/docs/cli) — every `taskstation self-host` subcommand and flag. --- # TaskStation TaskStation is the AI command center for your company. Canonical page: https://taskstation.co/docs TaskStation is the Autonomous Company Operating System — a cloud computer where a workforce of AI agents runs your company, and everything is code you own. Each agent works inside your own git repo. A session runs an agent in an isolated sandbox, on its own branch. When the work is ready, the agent opens a change request. You review it and merge it to the default branch. ``` project (git repo + taskstation.yaml) └─ session ──> isolated sandbox on branch "" └─ agent commits + pushes └─ change request ──> merge ──> default branch ``` New projects use a v2 manifest and OpenCode REST. Existing v1 projects keep their compatibility behavior. - [Quickstart](/docs/quickstart): Create a project, start a session, and merge your first change request. - [Your project](/docs/project): Set up the manifest, agents, models, and secrets. - [Running work](/docs/work): Sessions, sandboxes, and change requests. - [Connect & automate](/docs/connect): Connectors, Slack, computers, and triggers. - [Feature flags](/docs/feature-flags): Per-project opt-ins, including Apps. - [CLI](/docs/cli): Develop and run TaskStation from your terminal. - [SDK](/docs/sdk): Build on TaskStation from TypeScript with the client SDK. - [Self-hosting](/docs/host): Install and operate TaskStation on your own infrastructure. --- # Agents An agent is a markdown file that OpenCode runs; the manifest only governs its access. Canonical page: https://taskstation.co/docs/project/agents An agent is a markdown file that defines how OpenCode acts in a [session](/docs/work/sessions). This page covers agent files, skills, and how a session picks an agent. ## OpenCode runs every session OpenCode, the open-source coding-agent runtime, runs every session. The `taskstation-agent` daemon starts it as `opencode serve`, with its config directory pointed at the project's `.taskstation/opencode/` folder. ## An agent is two files Each agent has two parts: - **Behavior** — a markdown file at `.taskstation/opencode/agents/.md`. Its frontmatter sets the model, mode, and tools. Its body is the system prompt. - **Governance** — an entry in `taskstation.yaml`'s `agents` map, keyed by the same name. It sets what the agent may access on TaskStation: connectors, secrets, skills, and CLI actions. The manifest never sets a prompt, mode, or tool. The `.md` file never sets platform access. See [the manifest reference](/docs/project/manifest) for every governance field. ## Governance is deny-by-default An agent with no `connectors`, `secrets`, `skills`, or `taskstation_cli` key gets none of that access. This grant is the second of the two bindings: the agent also holds roles as a `service_account` principal, and a session can only do what both allow — see [One vocabulary, two bindings](/docs/accounts#one-vocabulary-two-bindings). Set a field to `all` to grant full access, or list specific names. `default_agent` in `taskstation.yaml` names the agent a session starts when you request no agent. Legacy `taskstation.toml` (v1) projects list agents in an array, not a map, and grant full access by default. Set specific grants to restrict access. See [the manifest reference](/docs/project/manifest). ## Skills give an agent know-how A skill is a markdown file at `.taskstation/opencode/skills//SKILL.md`. OpenCode loads a skill on demand when the agent calls it — TaskStation does not inject skills into every prompt. An agent's `skills` grant in `taskstation.yaml` controls which skills it may load. ## How a session picks its agent Creating a session accepts an `agent_name`. TaskStation starts that agent for the session. If you omit `agent_name`, TaskStation uses the project's `default_agent`. Set the default agent from the dashboard, the API, or the SDK. Each writes `default_agent` to `taskstation.yaml` and commits the change to the project's default branch. An agent never exceeds the access of the person or token that started the session. New or changed agents and skills reach future sessions only after a [change request](/docs/work/change-requests) merges them into the default branch. Connect external tools before an agent can use them; see [Connectors](/docs/connect/connectors). --- # Your project A project is a git repository that holds your agent's config and state. Canonical page: https://taskstation.co/docs/project A project is one git repository. It holds a manifest, agent config, and the state your agent produces. No separate database exists to keep in sync. TaskStation backs a project two ways: - **TaskStation-managed repo** — TaskStation creates and hosts a private repo. Default. - **Imported GitHub repo** — connect an existing repo. TaskStation operates on it through the GitHub API. Either way, the project has a `default_branch`. Every [session](/docs/work/sessions) branches from it. Every [change request](/docs/work/change-requests) merges into it. Each branch of a GitHub repo can become its own project. `main` and `dev` can then run as separate projects, each with its own sessions and change requests. ## A project is not your codebase A project holds your agent's instructions, its [connectors](/docs/connect/connectors), its [triggers](/docs/connect/triggers), and its memory. Keep it small — TaskStation clones it into every session. Your code lives elsewhere. When a task needs a codebase, the agent clones that repository, does the work, and opens a change request back to it. Do not turn an existing codebase into a project by adding a manifest to it. Create a dedicated project instead, and point it at the repositories the agent works on. ## The manifest Every project has a manifest at its repo root, `taskstation.yaml` by default. The platform reads these top-level keys: | Key | Configures | |---|---| | `default_agent` | which agent runs by default | | `agents` | per-agent grants: connectors, secrets, skills, TaskStation CLI permissions | | `sandbox` | the sandbox image and hardware | | `triggers` | scheduled and webhook automation | | `connectors` | which external tools the project can reach | | `env` | env variable names the project expects (values come from secrets) | Unknown keys are ignored. Older projects may still run `taskstation.toml` (`taskstation_version: 1`). TaskStation reads both formats. See [legacy TOML](/docs/project/legacy-toml) for the migration path. ## In this section - [Manifest](/docs/project/manifest): The full taskstation.yaml reference. - [Agents](/docs/project/agents): Markdown personas with scoped tools. - [Models](/docs/project/models): Which model a session uses, and who pays. - [Secrets](/docs/project/secrets): Encrypted values injected into a sandbox. - [Legacy TOML](/docs/project/legacy-toml): Migrating from taskstation.toml. --- # Legacy taskstation.toml Support status and the migration path from v1 taskstation.toml to v2 taskstation.yaml. Canonical page: https://taskstation.co/docs/project/legacy-toml This page covers the v1 manifest (`taskstation.toml`, `taskstation_version: 1`). For the current manifest, see [Manifest reference](/docs/project/manifest). ## Support status TaskStation still supports v1 manifests. The platform resolves a project's manifest in this order: `taskstation.yaml`, then `taskstation.yml`, then `taskstation.toml`. The current starter creates one v2 `taskstation.yaml`. A v1 project keeps working with no forced upgrade. v1 accepts TOML or YAML syntax. Version 2 accepts YAML only. A v2 manifest written in TOML fails validation. ## Migrate to v2 Create `taskstation.yaml` at the repo root. Set `taskstation_version: 2`. Convert `[[agents]]` to an `agents:` map, keyed by agent name. Rename each agent's `env` field to `secrets`. List every secret the agent needs — v2 does not grant secrets by default. Add `default_agent`, naming the agent that runs by default. Delete `[[channels]]`. Reconnect each channel from the dashboard. If you use `[sandbox]` or `[[sandboxes]]`, move each image definition under `sandbox.templates`. Run `taskstation validate` against `taskstation.yaml`. Fix any error it reports. Delete `taskstation.toml`. Agent behavior — system prompt, `model`, `mode`, `temperature`, and more — never moves. It already lives in each agent's `.md` frontmatter under `.taskstation/opencode/agents/`, in both v1 and v2. See [Agents](/docs/project/agents). Check the result with `taskstation validate --file taskstation.yaml`. Fetch the full v2 schema with `taskstation schema --version 2`. See [CLI reference](/docs/cli) for both commands. ## Key differences | v1 (`taskstation_version: 1`) | v2 (`taskstation_version: 2`) | Change | |---|---|---| | `[[agents]]` (array) | `agents:` (map keyed by name) | Convert each array item to a map entry. | | `[[agents]].env` | `agents..secrets` | Renamed. Default flips from `all` to `none` — see the callout below. | | `[[agents]].model`, `[[agents]].file` | removed | Dead in v1 (parsed, never applied). Agent behavior always comes from the agent's `.md` frontmatter. | | no `default_agent` | `default_agent` (required) | Must name a declared, enabled agent. | | `[[channels]]` | removed | Channel routing is dashboard-managed. A connected channel appears as a `connectors:` entry with `provider: channel`. | | `[sandbox]` (singular, image keys directly on it) | `sandbox.templates` (list) | Move each image into a template entry. | | `[[sandboxes]]` | `sandbox.templates` | Renamed. | > **Secrets default to none in v2** > In v1, an agent with no `env` field gets every project secret by default. In v2, an agent with no `secrets` field gets none. When you migrate, list every secret each agent needs — an agent that silently loses a secret can fail mid-task. A few `taskstation_cli` actions are legacy-tolerated in v1 (`project.session.exec`, `project.gateway.routing.edit`, `project.schedule.read`, `project.schedule.write`, `project.webhook.read`, `project.webhook.write`, `channel.read`, `channel.connect`, `channel.send`, `channel.disconnect`) — the platform warns but allows them. The same actions are a hard error in v2. Remove them from any agent's `taskstation_cli` grant before you migrate. --- # Manifest reference Every taskstation.yaml v2 key, with defaults, limits, and validation rules. Canonical page: https://taskstation.co/docs/project/manifest The manifest is the file the platform treats as authoritative for a project. This page lists every `taskstation.yaml` (v2) key, its type, its default, and how the platform validates it. For the plain-language version, see [Your project](/docs/project). ## Two configuration surfaces A TaskStation project has two configuration surfaces with strict, non-overlapping ownership. - **TaskStation config** — `taskstation.yaml` at the repo root, plus `.taskstation/Dockerfile`. The platform reads this: the trigger sweep, the sandbox builder, session token minting, and the dashboard. - **OpenCode config** — everything under `.taskstation/opencode/` (`opencode.jsonc`, `agents/`, `skills/`, `commands/`, `tools/`, `plugins/`). The OpenCode runtime reads this, in the sandbox and locally. The join between the two halves is the agent name. A manifest key `agents.` must match the filename `.taskstation/opencode/agents/.md`. Every behavioral field lives only in that `.md` file's frontmatter and body, never in `taskstation.yaml`: - system prompt - `model`, `mode` - `temperature`, `top_p` - `steps`, `permission` The manifest's `agents:` block sets governance only: which agents may launch, and what each one may touch. The validator enforces this. It rejects a v2 agent block that contains any behavioral field, with an error that points at the agent's own `.md` file. ## Location and versions Any repo with a valid manifest at its root is a TaskStation project. `taskstation.yaml` (v2, YAML only) is the current format for new projects, created by the web "Create project" flow and by `taskstation init`. `taskstation.toml` (v1) still works for existing projects, but the platform accepts no new v2 features on it. See [Legacy TOML](/docs/project/legacy-toml) for the v1 schema and the v1-to-v2 migration steps. `taskstation_version: 2` requires YAML — a `.toml` file that declares `taskstation_version: 2` fails validation. A manifest that declares a version higher than `2` is rejected outright, so the platform never silently misreads a future field. Unknown top-level keys are ignored, so you can park your own metadata in the file. `validateManifest()` is the single gate behind `taskstation ship`, the change request merge check, and `taskstation validate`. The same rules apply everywhere, so anything that merges into `main` is structurally sound. The schema is public and generated from the same package: - [`taskstation.v2.schema.json`](/schema/taskstation.v2.schema.json) - [`taskstation.v1.schema.json`](/schema/taskstation.v1.schema.json) - [`taskstation.schema.json`](/schema/taskstation.schema.json) (dispatches on `taskstation_version`) Point a `taskstation.yaml` at the v2 schema for editor validation, or fetch it from the CLI: ```yaml # yaml-language-server: $schema=https://taskstation.co/schema/taskstation.v2.schema.json taskstation_version: 2 ``` ```bash taskstation schema --version 2 ``` ## Full example ```yaml taskstation_version: 2 default_agent: taskstation project: name: my-project description: What this project is. env: required: [DATABASE_URL] optional: [STRIPE_API_KEY, WEBHOOK_SLACK_SECRET] sandbox: templates: - slug: ml name: ML Development dockerfile: .taskstation/Dockerfile.ml cpu: 4 memory: 16 opencode: config_dir: .taskstation/opencode agents: taskstation: connectors: all secrets: all taskstation_cli: all skills: all release-bot: sandbox: ml connectors: [github] taskstation_cli: [project.gitops.push] secrets: [GITHUB_AGENT_TOKEN] triggers: - slug: daily-digest type: cron agent: taskstation cron: '0 0 9 * * 1-5' timezone: America/Los_Angeles prompt: | Summarize yesterday's commits. Open a CR against main. ``` ## Top-level keys | Key | Required | Notes | | --------------------------------------------------- | -------- | ------------------------------------------------------------------------------------------ | | `taskstation_version` | yes | must be `2` | | `default_agent` | yes | must name a declared, enabled agent | | `runtime` | no | only legal value is `"opencode"` (the default) | | `project.name` / `project.description` | no | display metadata; the platform does not read `project.name` for the project's display name | | `env.required` / `env.optional` | no | env var names; `required` is advisory only, never enforced at session start | | `sandbox` / `sandbox.templates` / `sandbox.default` | no | sandbox image(s) and hardware | | `opencode.config_dir` | no | default `.taskstation/opencode` | | `triggers` | no | list of cron, webhook, and monitor triggers | | `connectors` | no | list of connector definitions | | `policy.default_mode` | no | `allow_all` (default) or `risk`; project-wide connector approval mode | | `agents` | yes | name-to-block governance map; must not be empty | ## `project:` Optional, human-facing metadata: `name` and `description`. The platform does not read this table for anything — a project's display name comes from its own database record. Treat it as documentation for people reading the repo. ## `env:` Declares the env var names your sessions need. Values live in the dashboard's Environment variables page, never inline in the manifest. The platform decrypts and injects them as plain env vars at session start. ```yaml env: required: [DATABASE_URL] optional: [STRIPE_API_KEY] ``` | Field | Type | Notes | | ---------- | ---------- | -------------------------------------------------------------------------------------------------------- | | `required` | `string[]` | Advisory. The dashboard prompts the user for these, but session start does not block on a missing value. | | `optional` | `string[]` | Available to sessions if set. Absence is fine. | Env var names match `^[A-Z_][A-Z0-9_]*$`. The Secrets Manager caps a name at 64 characters — a longer name parses here but can never get a value. Names starting with `TASKSTATION_` can never get a value either. Keep names at 64 characters or fewer, and avoid the `TASKSTATION_` prefix. Full contract: [Secrets](/docs/project/secrets). ## `sandbox:` and `sandbox.templates` A list of named, bootable sandbox images. Optional — with no entries, every session boots the platform's default image. Each template needs exactly one of `dockerfile` (repo-relative) or `image` (a public Docker reference, tag- or digest-pinned). ```yaml sandbox: templates: - slug: ml name: ML Development dockerfile: .taskstation/Dockerfile.ml cpu: 4 memory: 16 disk: 50 ``` | Field | Type | Default | Notes | | ------------ | ------ | ------------------- | --------------------------------------------------------------------------------- | | `slug` | string | — | Required. Unique per project. `default` is reserved. | | `name` | string | slug | Display label in the dashboard picker. | | `dockerfile` | string | — | Repo-relative path. Mutually exclusive with `image`. | | `image` | string | — | Public Docker image, tag- or digest-pinned. Mutually exclusive with `dockerfile`. | | `entrypoint` | string | runtime layer's own | Overrides the container entrypoint. | | `cpu` | int | provider default | vCPU cores. Bound: 1–32. | | `memory` | int | provider default | RAM in GiB. Bound: 1–128. | | `disk` | int | provider default | Disk in GiB. Bound: 1–500. | Each value must be a positive integer. A value below the minimum fails validation with an error. A value above the maximum passes with a warning, then clamps at runtime. GPUs are not supported — declaring `gpu` produces a warning, not an error. See [Runtime](/docs/work/runtime) for what the runtime layer injects on top of your image. ### `sandbox.default` Set `default` on `sandbox` to make one template the project-wide default. Every session, trigger, and channel then boots it without naming a slug. ```yaml sandbox: default: dev templates: - slug: dev dockerfile: .taskstation/Dockerfile ``` `default` must name a template declared in this manifest, or the reserved value `"default"` for the platform image. An agent can select its environment with `agents..sandbox`. The value must name a manifest or dashboard template, or `"default"`. ```yaml agents: researcher: sandbox: ml ``` Session template resolution uses this order: 1. Explicit session `sandbox_slug`. 2. Selected agent `sandbox`. 3. Project `sandbox.default`. 4. Platform `"default"`. Cron triggers, webhook triggers, schedules, and channels use the selected agent's template. A spec change (`cpu`, `memory`, `disk`, or the image itself) rebuilds the project's snapshot. The new size applies on the next session, not a running one. ## `opencode:` Where the OpenCode config directory lives. Optional, with a default. ```yaml opencode: config_dir: .taskstation/opencode ``` | Field | Type | Default | Notes | | ------------ | ------ | ------------------ | ------------------------ | | `config_dir` | string | `.taskstation/opencode` | Repo-relative directory. | That directory holds agents, skills, commands, tools, plugins, and `opencode.jsonc`. `opencode.jsonc` stays the OpenCode-native registry for plugins, MCP servers, providers, and permissions — do not duplicate those settings in the manifest. See [Agents](/docs/project/agents). ## `triggers:` A list of cron, webhook, and monitor definitions. Each entry fires a session that runs `prompt` as its first message. ```yaml triggers: - slug: daily-digest type: cron cron: '0 0 9 * * 1-5' prompt: Summarize yesterday's commits. ``` | Field | Required | Default | Notes | | -------------- | --------------------- | --------------------- | ----------------------------------------------------------------------------------------------------- | | `slug` | yes | — | `[a-z0-9][a-z0-9_-]{0,127}`, unique among triggers. | | `type` | yes | — | `cron`, `webhook`, or `monitor` (experimental). | | `prompt` | yes | — | Non-empty. Supports templating. | | `name` | no | slug | Human label. | | `agent` | no | `default_agent` | Must name a key in `agents:`. Omit it to use `default_agent`; do not write the literal `default`. | | `enabled` | no | `true` | `false` skips the entry. | | `model` | no | resolves at fire time | Wire form `provider/model`. Pins the fired session to that model. See [Models](/docs/project/models). | | `session_mode` | no | `fresh` | `fresh`, `reuse`, `pinned`, or `keyed`. | | `session_id` | required for `pinned` | — | Exact session to re-prompt. | Cron triggers need exactly one of `cron` (a 6-field expression) or `run_at` (a one-off ISO-8601 timestamp), plus `timezone` as an IANA name. Webhook triggers need `secret_env`, the name of a secret that uses `broker` delivery with the `connector` consumer. Monitor triggers need `run` (a repo-relative command) and `mode` (`poll` or `stream`), and reject every cron and webhook field. Full field reference, credential setup, payload templating, endpoints, and session strategy: see [Triggers](/docs/connect/triggers). ## `connectors:` A list of external tools an agent can call. The definition lives in git; credentials live in the platform, never in the manifest. See [Connectors](/docs/connect/connectors) for the conceptual model. ```yaml connectors: - slug: gmail-read name: Gmail read only provider: pipedream app: gmail authorization_strategy: user policies: - match: search_email action: always_run ``` | Field | Required | Notes | | ------------------------ | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `slug` | yes | `[a-z0-9][a-z0-9_-]{0,127}`, unique among connectors. | | `provider` | yes | `pipedream`, `mcp`, `openapi`, `postman`, `graphql`, `http`, or `channel`. | | `name` | no | Display name. Defaults to slug. | | `authorization_strategy` | no | `project` or `user`. Defaults to `project` when omitted. A project strategy uses active project connections. A user strategy uses only the acting member's connection. | | `enabled` | no | Defaults to `true`. | | `credential` | no | Only `shared` is supported. The retired `per_user` mode is a hard error. | | `sensitive` | no | Defaults to `false`. `true` makes `require_approval` the unmatched-action default, including reads. Explicit project and connector rules still resolve first. | Provider-specific fields: | Provider | Required field | Notes | | ----------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | | `pipedream` | `app` | `account` optional, defaults to slug. | | `mcp` | `url` | `transport`: `http` (default) or `sse`. | | `openapi` | `spec` | A URL or repo-relative path. | | `postman` | `spec` | A Collection JSON URL or path, a `.postman/api` manifest, or a Postman workspace URL. | | `graphql` | `endpoint` | `spec` optional (SDL). | | `http` | `base_url` | `spec` optional. | | `channel` | `platform` | `slack`, `teams`, `email`, or `voice`. | | `computer` | — | API-managed. You cannot declare it by hand. A profile selects one or more paired machines. See [Computer Tunnel](/docs/connect/computers). | `channel` connectors rarely need a hand-written entry — connecting a channel from the dashboard creates one for you. See [Slack & channels](/docs/connect/slack). The slugs `taskstation_slack`, `taskstation_teams`, `taskstation_email`, and `computer` are platform-owned; using one with a different provider is a validation error. Computer Tunnel profiles are created through the connector API because tunnel ids are account control-plane identities and must not enter the manifest. `connectors.auth`: optional, for providers other than `pipedream`. `type` is `bearer`, `basic`, `custom`, `api_key`, `oauth1`, `hmac`, `aws_sigv4`, `mtls`, or `none` (default). `oauth1` is restricted to `openapi`, `postman`, and `http` providers. `in` is `header` (default), `query`, or `cookie`. `name` is required when `type` is `custom` or `api_key`. `connectors.policies`: a list of `{match, action}` pairs. `match` is a glob over tool names. `action` is `always_run`, `require_approval`, or `block`. Policies belong to the connector. Every connection under that connector uses the same policies. `policy.default_mode`: a top-level key, separate from `connectors:`, that sets the project-wide connector approval mode. It takes `allow_all` (default; every unmatched tool runs) or `risk` (require approval for write and destructive unmatched calls). Set it with `taskstation connectors policy set --default ` — see [CLI](/docs/cli). ## Channels v2 removes `channels:` from the schema. Channel-to-agent routing (Slack, Microsoft Teams, email, and voice today) is live operational state, not declarative config. You set it from the dashboard's Channels page or from chat commands, the same boundary that keeps credentials out of git. Connecting a channel still creates a `connectors` entry with `provider: channel` for the agent to call. See [Slack & channels](/docs/connect/slack). The v1 `[[channels]]` table is covered in [Legacy TOML](/docs/project/legacy-toml). ## `agents:` A name-to-block map, keyed by agent name. This map is governance only — it grants what an agent may touch, never what it says or does. This block is the **second of the two bindings**: a principal gets roles from TaskStation, and an agent additionally carries the TaskStation CLI scopes declared here. A session can only do what both allow. See [One vocabulary, two bindings](/docs/accounts#one-vocabulary-two-bindings). ```yaml default_agent: taskstation agents: taskstation: connectors: all secrets: all taskstation_cli: all skills: all release-bot: connectors: [github] connectors_required: [github] taskstation_cli: [project.gitops.push] secrets: [GITHUB_AGENT_TOKEN] ``` | Field | Default | Notes | | --------------------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | _(map key)_ | — | `[a-z0-9][a-z0-9_-]{0,127}`, unique per project. Matches the `.md` filename it governs. | | `enabled` | `true` | `false` treats the agent as undeclared — the platform will not launch it. | | `connectors` | `none` | Which connector slugs this agent may call. `all` means every connector. | | `connectors_required` | `none` | Connector slugs that must resolve an active, strategy-compatible authorization before sandbox startup. Must be a subset of `connectors`. A missing authorization returns `CONNECTOR_CONNECTION_REQUIRED`. | | `secrets` | `none` | Which project secrets this agent receives as sandbox env vars. Renamed from v1's `env`. | | `taskstation_cli` | `none` | Which TaskStation CLI and API permissions this agent may exercise. `all` means everything the launching user's roles allow — an agent can never exceed its launcher. | | `skills` | `none` | Which skills this agent may load. | | `workspace` | — | Git-boundary mode: `runtime`, `read`, or `branch`. Validated against the enum; any other value is rejected. | v2 is deny-by-default: an omitted `connectors`, `secrets`, `taskstation_cli`, or `skills` on a declared agent resolves to `none`. Grant every permission an agent needs explicitly, as the starter's `taskstation` agent does above. A project migrating from v1 must re-grant everything by hand — v1 defaults `env` (secrets) to `all` when omitted, the opposite of v2. `connectors_personal` remains a deprecated input alias for `connectors_required`. New manifests must use `connectors_required`. The connector's `authorization_strategy` decides whether the authorization is project-owned or member-owned. `taskstation_cli` grants come from a fixed set of project-scoped permissions (`project.read`, `project.gitops.push`, `project.gitops.merge`, `project.gitops.ref.any`, `project.trigger.*`, `project.secret.*`, `project.connector.*`, and more). Account-scoped permissions — `member.*`, `billing.*`, `project.create` — can never be granted to an agent. Run `taskstation validate --scopes` from inside a session to print the live, grantable list. ## When config changes take effect - A manifest or `.taskstation/opencode/` edit applies only after a change request merges to `main`. Sessions and the trigger sweep read the default branch, not session branches. - A trigger's cron or webhook change is picked up by the scheduler within seconds of the merge. - A sandbox image or hardware change rebuilds the snapshot. The new spec applies on the next session, not the current one. - A secret value change in the dashboard resolves at sandbox-create time. It applies on the next session, not a running one. ## Round-trip rules Dashboard edits are a read-modify-write on the same file. To keep diffs clean across UI and in-session edits: 1. Keep `taskstation_version` as the first key. 2. Inside a trigger entry, order fields `slug`, `name`, `type`, `agent`, `enabled`, then type-specific fields, then `prompt` last. 3. If you add a webhook trigger before its secret is set, list the secret name in `env.optional`. Leave the trigger `enabled: false` until the value is in. --- # Models How TaskStation picks a model, and how billing works for managed vs. your-own-key models. Canonical page: https://taskstation.co/docs/project/models TaskStation runs each [session](/docs/work/sessions) on a model. This page explains managed models vs. your own provider key (BYOK), how TaskStation picks a model automatically, and two billing gotchas to know. This page applies to projects with the LLM Gateway on. LLM Gateway is an experimental [feature flag](/docs/feature-flags), **on by default** where the platform offers it — check or toggle it in Settings → Experimental (operators can default a whole deployment off with `LLM_GATEWAY_DEFAULT_ENABLED=false`). Turning the flag off is a fully supported path. The project then runs **native OpenCode model management**: - Your provider API keys (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `OPENROUTER_API_KEY`, …) are injected into the sandbox as ordinary env vars — add them on the Model settings page or as [secrets](/docs/project/secrets). OpenCode connects each provider from its key automatically. - Model ids are OpenCode's native `provider/model` refs, like `anthropic/claude-opus-4-8`. Managed bare ids and `taskstation/…` refs do not exist off-gateway. - The model picker shows one list before and after the sandbox boots: the providers your keys connect, plus OpenCode Zen's free models (OpenCode connects those without a key). Thinking effort is the composer's thinking control (the model's own variants); the gateway's Generation defaults do not apply. - The gateway surfaces on this page — managed models, the model-defaults chain, budgets, logs — do not apply; OpenCode resolves the default model in the sandbox. ## Managed models and BYOK A model id has one of three shapes: - **Managed** — a bare id, like `grok-4.6` or `deepseek-v4-pro-0813`. TaskStation supplies the credentials. Cloud accounts pay with TaskStation credits. - **BYOK** — a `provider/model` id, like `anthropic/claude-opus-4-8`. You supply the key. Your provider account pays. - **ChatGPT** — a `codex/` id. You connect your ChatGPT plan once through OAuth, and it pays. Connect a BYOK key on the project's Model settings page, or set the provider's env var directly as a [secret](/docs/project/secrets). ## Thinking effort The composer's thinking control sets the session's model **variant** in both modes. The choices are the model's own published tiers (models.dev `reasoning_options`), never a fixed ladder; a model without a knob shows no control. `Auto` clears the variant. - Native (gateway off): OpenCode applies the variant as the provider's own request field. - Gateway on: the request carries `reasoning_effort`; the gateway maps it per upstream (OpenAI → `reasoning_effort`, Claude → adaptive thinking, OpenAI on Amazon Bedrock → Bedrock's `reasoning.effort` request field). For an upstream it cannot map yet (Nova, Grok on Bedrock today) the value is dropped and the model runs at its own default. - Amazon Bedrock refuses the bare in-region id of most current models ("on-demand throughput isn't supported"). The picker prefers the `global.` / regional inference-profile id when the catalog carries one, and the gateway retries a refused bare id once per profile prefix (`global.`, `us.`). - The sandbox learns the project's servable model set from the API at every boot (`GET /v1/llm/models?scope=picker`, the same composition as the web picker), so a model the picker offers always resolves in the runtime. - A project default per model lives in Customize → Gateway → Routing → **Generation defaults** (`model_generation_config`). It fills only a field the request left unset, so a session's variant always wins. ## How auto picks a model Set no model, and TaskStation resolves one through five layers, in order. (The id `auto` covers this same behavior, but it is not yet a selectable option in the model picker.) 1. An explicit pin — a session, channel, or trigger's own `model:` field. 2. The [agent's](/docs/project/agents) default for this project. 3. The project's default. 4. The account's default. 5. The platform default. TaskStation uses the first layer that has a value it can still serve. A saved default that stops working — a disconnected key, a retired model — is skipped automatically. A session never dies from a stale default. See the [manifest reference](/docs/project/manifest) for the trigger `model:` field. > **Billing surprises on BYOK** > Two costs are easy to miss on a paid cloud account: > > - **Platform fee.** TaskStation adds a 10% fee, billed as credits, on top of what > your own provider charges. Free-tier and self-hosted accounts are exempt. > - **Silent failover.** If your BYOK key hits a rate limit or billing error > mid-turn, TaskStation retries on a managed model and bills your credits instead > of failing the session. > > If you see credit charges on a BYOK-only project, check these two causes > before reporting a billing bug. ## Per-project model enablement The project controls which models its pickers offer. By default, the newest model of each family is offered automatically. TaskStation-managed models and any model your project's defaults or routing policy reference are always offered — a guard never prunes them. You can override the default for individual models on the **Manage models** page (Customize → Models). An exception is stored per project and takes effect immediately. The session model picker and the command palette hide anything you turn off; new models stay on by default as the catalog grows. Enablement governs what is offered, not what is served: a request that names a disabled model outright (for example through the raw API) still runs. The project's default model cannot be turned off — set a different default first. ## Shared, not private, keys A connected provider key applies to the whole project. There is no private, per-user key — setting a personal override for a provider key fails with a `llm_credentials_project_wide` error. Update the shared key on the [secrets](/docs/project/secrets) page instead. --- # Secrets How TaskStation stores project credentials, and how it controls whether agent code can read each value. Canonical page: https://taskstation.co/docs/project/secrets A secret is a per-project credential — an API key, token, or connection string — that a session needs but that must not live in the repository. TaskStation stores secrets on the project, never on the account, and encrypts every value with AES-256-GCM using a key derived per project. ## Identifier and name Each secret has two names. The `identifier` is the handle you use in the CLI and in an agent's `secrets` grant. The `name` is the uppercase environment-variable key in the sandbox, for example `STRIPE_API_KEY`. In most projects the identifier and the name match. They differ only when a project holds several candidate values for one name — for example, a primary and a backup Google Maps key can both resolve to `GOOGLE_MAPS_API_KEY`. Every access rule uses the identifier. None of them uses the name. ## Exposure and usage A secret carries two independent settings. Read them separately; they answer different questions. | Setting | Question it answers | Values | |---|---|---| | **Exposure** | Can agent code read the real value? | `environment`, `egress-enforced`, `none` | | **Usage** | Who spends it? | Agent code, LLM gateway, Connector, Git | Older TaskStation documentation presented one list that mixed the two. It is gone. There is no choice between a "network boundary" and an "HTTPS broker": one mechanism serves every egress-enforced secret on every sandbox provider. ### The three exposures | Exposure | What the sandbox holds | Use it for | |---|---|---| | **Environment** | The real value, as a plain environment variable | The default. Values the agent must compute with, and protocols that are not HTTPS | | **Egress-enforced** | A **handle** — a self-describing placeholder, worth nothing on its own | Experimental. HTTPS calls, once the `secrets_egress` flag is on | | **None** | Nothing at all | A credential only a TaskStation service spends, or a value kept on file and disabled | > **The working rule** > **Environment** is the default exposure. The real value loads into the sandbox, > where the agent can read, print, and forward it. **Egress-enforced** keeps the > value outside the sandbox, but it is experimental: enable the `secrets_egress` > feature flag (Settings → Feature flags) to use it. Until then, a secret loads > into the sandbox environment. ### Usages Most usages are assigned by TaskStation, not by you: - **Agent code** — implied whenever exposure is not `none`. - **LLM gateway** — assigned when the value is a recognized model-provider key. The gateway authenticates provider requests server-side. - **Connector:<slug>** — assigned by the connector binding flow. - **Git** — assigned by TaskStation for its own Git access. Read-only; you cannot set or clear it. A secret with exposure `none` and no usage renders as **Disabled**: stored, encrypted, and spent by nothing. ## Sent secrets and computed secrets Which exposure a secret can use is a property of the **upstream**, not of TaskStation. - **Sent secrets** — the value travels on the wire. API keys, bearer tokens, passwords. This is the vast majority. There is a moment where the value is bytes in an outbound request, so TaskStation can put it there itself, outside the sandbox. A sent secret can move to **egress-enforced** once the `secrets_egress` flag is on; by default it loads into the sandbox environment. - **Computed secrets** — the value is an ingredient in a calculation and never travels. AWS SigV4 signing keys, HMAC webhook-signing secrets, JWT client assertions, SSH private keys. Whoever computes must hold the value. No network boundary helps, because nothing on the wire contains the credential. Computed secrets must stay on **environment**. Environment is the default exposure, and the only one that can serve a credential the sandbox has to do math with. Non-HTTPS protocols — a Postgres connection string, SMTP credentials — are in the same position: they must stay on environment. When you save a value that looks like signing material, such as an `AKIA…` access-key pair or PEM/SSH material, TaskStation defaults it to **environment** and says why: this key signs requests locally, so egress enforcement cannot apply. ## Egress-enforced exposure > **Experimental — needs the secrets_egress flag** > Enforcement at the network is experimental. Enable the `secrets_egress` feature > flag (Settings → Feature flags) to use it. Until then, a secret loads into the > sandbox environment. With the flag off, creating an egress-enforced secret > returns `403` `feature_disabled`. The sandbox receives an environment variable whose value is a **handle**, not the credential. The agent uses that variable exactly as it would use the real key — in a header, a query string, or a body. `Authorization: Bearer $VAR`, `Cookie: …=$VAR`, an `X-Api-Key` header, a query parameter, and a JSON or form body field all work: TaskStation finds the handle wherever it appears (raw, URL-encoded, standalone base64, or JSON-escaped) and swaps in the real value. On the way out, TaskStation replaces the handle with the real value, but only for requests to hosts you approved. > **Send the handle as-is — do not base64 it yourself** > One thing does not work: hiding the handle inside a base64 blob you build > yourself, which is what **HTTP Basic auth** does (`curl -u $VAR:` → > `Authorization: Basic `). TaskStation cannot find a handle that is embedded, > unaligned, inside base64 you encoded, so the swap never happens and the upstream > answers `401`. Put the handle in a Bearer/token header, a query parameter, or a > body field instead. If the API only supports Basic auth, use **environment** > exposure for that secret. ```text agent's ordinary HTTP client └─▶ in-guest shim (terminates TLS for approved hosts only; holds no secret) └─▶ TaskStation, server-side host allow-list → resolve grant and session allowlist → decrypt → substitute handle → call upstream → redact echoes → audit ``` Facts that follow from that shape: - The real value is **never** in the sandbox: not an environment variable, not a file, not an alias. - A handle sent to a host you did not approve arrives as the literal handle string. The upstream rejects it. It is worth nothing. - A handle with a bad signature is never honored, and TaskStation records it as a forged handle. A valid handle for a secret this session may not spend is recorded as a stolen one. - The mechanism is identical on every sandbox provider — Daytona, E2B, and Platinum. There is no flag to turn on and no provider to pin. - Every relayed request writes a per-request audit record. - Hosts that are not on the list are tunnelled without being read. Pinned-TLS and mTLS clients to those hosts are unaffected. ### Hosts match exactly List every host you call, one exact hostname per line. | Rejected | Reason | |---|---| | A wildcard host, `*.example.com` | The agent must never choose the destination | | A URL scheme other than HTTPS | TaskStation terminates TLS to substitute | `api.example.com` does not cover `uploads.api.example.com`. Add the second host to the same secret. TaskStation rejects an unenforceable policy with `400` when you save it — it never stores a rule it cannot apply. Two egress-enforced secrets may share one host. Each handle maps to its own value, so both substitute correctly in the same request. ### Verify it with two probes Run both probes from inside the sandbox, against a host that is on the list. The example uses `postman-echo.com`, which serves one endpoint of each kind. Add it as an allowed host for the duration of the test. ```bash # 1. Reachability — an endpoint that does NOT echo request headers. curl -s -o /dev/null -w '%{http_code}\n' https://postman-echo.com/status/200 # expected: 200 # 2. Substitution — an endpoint that DOES echo request headers. curl -sS -H "authorization: Bearer $STRIPE_API_KEY" https://postman-echo.com/get # expected: 200, with "Bearer [REDACTED]" in the echoed headers ``` Read the pair together, probe 1 first: it is the only one that tells you the host is reachable, and probe 2 proves nothing until it passes. | Probe 2 result | Meaning | |---|---| | `200`, echoed header shows `[REDACTED]` | Working. The real value went upstream and the echo was redacted on the way back. | | `200`, echoed header shows the handle itself | The substitution did not run. Check the host list and the agent grant. | | `401` from a real API host | The substitution did not run. Same two checks. | | An empty reply or a connection error | A real failure. This is not a success symptom. | Confirm the real value is nowhere in the guest: ```bash env | grep '^STRIPE_API_KEY=' # expected: the handle, not the credential ``` The identifier also appears inside `TASKSTATION_SECRET_CAPABILITIES`, the value-free catalog that tells the agent which secrets exist, which variable holds each handle, and which hosts each one covers. It never carries a value. A host on the list presents a certificate TaskStation issued for this sandbox, which the sandbox already trusts, rather than the origin's own: ```bash curl -sv https://postman-echo.com/status/200 2>&1 | grep 'issuer:' # issuer: CN=TaskStation Egress CA (); O=TaskStation ``` ### What the relay changes An approved host is reached through TaskStation, so it behaves a little differently from a direct call: - Responses are not streamed. Server-sent events and websockets do not work through an approved host. - A request body is capped at 1 MiB, a response at 5 MiB, and the whole call at 30 seconds. TaskStation follows at most 3 redirects. - Only a fixed set of response headers comes back — content type and language, caching validators, the rate-limit family, `retry-after`, `x-request-id`. - Only clients that honour `https_proxy` are intercepted. TaskStation sets that variable, and the matching CA trust, in the agent's environment. A process started with a scrubbed environment calls the host directly, and its request leaves carrying the handle. For a request TaskStation cannot intercept, the agent has an explicit door to the same hosts under the same policy: ```bash taskstation secrets call STRIPE_API_KEY https://api.stripe.com/v1/customers ``` ## Environment exposure The sandbox receives the real value as a plain environment variable. Agent code can read it, print it, and forward it anywhere. TaskStation cannot redact it, cannot audit its use, and cannot stop it leaving. Environment is the default exposure. Every new secret uses it unless you move the secret to egress-enforced, which is experimental and needs the `secrets_egress` flag. A computed credential, and any non-HTTPS protocol, must stay on environment. ## Shared and personal scope A secret is either shared or a personal override: - **Shared** — the project-wide value. Every principal with the `project.secret.read` permission sees every shared secret. Who may read a secret is the permission; which agent receives it is the manifest grant below. - **Personal override** — your own value for one name, used instead of the shared row for sessions you start. Today TaskStation uses this only for one OAuth login credential. TaskStation never lets an LLM provider key, for example `ANTHROPIC_API_KEY`, become a personal override. The model gateway always reads the shared row. ## Add a secret ### Set the value from the CLI ```bash taskstation secrets set STRIPE_API_KEY=sk_live_... ``` TaskStation saves it as a shared secret for the project. To store more than one value under the same name, add `--identifier `. Names can't start with `TASKSTATION_` — TaskStation reserves that prefix for platform values. ### Or use the dashboard Open the project's Secrets page, enter the key and value, and save. TaskStation encrypts the value immediately and defaults every new secret to **environment** exposure — the real value loads into the sandbox. To move a secret to egress-enforced exposure, first enable the `secrets_egress` feature flag (Settings → Feature flags); it is experimental. With the flag on, the Secrets page shows the **"Can your code read this value?"** control: answering *no* moves the secret to egress enforcement and shows the host list; answering *yes* keeps environment exposure. With the flag off, a secret stays on environment. ### Grant it to an agent Pick an agent on the Secrets page, or add the identifier to that agent's `secrets` list in `taskstation.yaml` yourself. A session only receives the secrets its agent is granted. See [Grant a secret to an agent](#grant-a-secret-to-an-agent). ## Grant a secret to an agent A session receives a secret only when the agent it runs names the identifier in its `secrets` list. Matching uses the identifier, not the name, and ignores case. > **Egress-enforced and service usages need a named grant** > `secrets: all` grants environment exposure only. An egress-enforced secret, and > any secret a TaskStation service spends, needs its identifier written out in an > agent's list. A project with no `agents:` block in `taskstation.yaml` never receives > one. An ungranted secret is dropped silently: the session starts normally and > the first call to the host fails as though the credential were wrong. ### From the dashboard The Secrets page marks a secret no agent can receive: **No agent can receive this secret**. Choose an agent there and confirm. TaskStation edits `taskstation.yaml` and commits it as `chore(agents): grant to `. The grant works whether or not the manifest already declares that agent. An agent the manifest does not declare gets a new entry holding this one `secrets` list. An agent that is already declared keeps every other field — model, tools, connectors — and the identifier joins its existing list. An agent that already admits the identifier needs no commit, and TaskStation makes none. An agent on `secrets: all` is a special case. `all` cannot carry an egress-enforced secret, so TaskStation writes an explicit list: every identifier the project has today, plus this one. Nothing the agent receives today changes. A secret you add later needs its own grant. Two cases refuse the grant: - A project on `taskstation_version: 1` (`taskstation.toml`) has no agents map to edit. The request fails with `400` and `manifest_v1_unsupported`. Edit the manifest by hand, or move the project to `taskstation_version: 2`. - A secret that is **Disabled** has nothing to deliver. The request fails with `409` and `secret_not_grantable`. Give it an exposure first. If the project has no `agents:` block yet, read [The first `agents:` block changes the whole project](#the-first-agents-block-changes-the-whole-project) before you confirm. That one edit changes secret access for every other agent. ### By hand The same grant, written directly: ```yaml taskstation_version: 2 agents: my-agent: secrets: [STRIPE_API_KEY] ``` ### The first `agents:` block changes the whole project > **Declaring one agent denies the rest** > A project with no `agents:` block — or with no `taskstation.yaml` at all — is > ungoverned: every agent receives every environment-exposure secret, and no > agent receives an egress-enforced one. > > The moment the project declares its first agent, every agent that is **not** > listed receives no project secret at all — including environment secrets that > worked a minute earlier. Listing one agent revokes the rest. So list every agent that needs secrets, not only the one you are fixing. This is why the dashboard asks you to confirm the first time: after that commit, `agents:` is the project's allow-list, and an agent missing from it runs with no project secrets. ## Exposing capabilities to third-party users > **TaskStation secret policies are not a multi-tenant authorization system** > A secret policy protects **your project's own agent** from leaking your own > credential. It says nothing about which of your end users may spend it. If you are building something where **untrusted third-party users** reach a capability — a public chat surface, a shared app, an agent anyone on the internet can prompt — do not hand them secret policies at all. Every user of that surface shares one project agent, one grant, and one host list. TaskStation has no way to tell one of your customers from another, so an egress-enforced secret that any user's prompt can reach is a credential every user can spend, up to the full scope the upstream key carries. Build the boundary you actually need, on your side: 1. Stand up your own authorization and proxy service. It holds the upstream credential. 2. Point the agent at your service, not at the upstream. Give the agent only a credential for your service — that one can be egress-enforced to your own host. 3. Your service identifies the end user, applies your own authorization rules and per-user quotas, and only then makes the upstream call with the credential it holds. That service is where per-user rules belong: who may call what, how often, for which records. TaskStation secret exposure is one layer below it, and it does not substitute for it. ## List your secrets Run `taskstation secrets ls` to see which secrets a project declares and which ones have a value set. The list is configuration metadata. It never returns secret values. A scoped agent token sees only identifiers in its agent grant. A session-specific `secrets_allowlist` controls delivery into that session, but it does not hide configuration metadata that the agent grant permits. ## Rotate a secret ### Set a new value Run the same command with the new value, or set it again on the project's Secrets page: ```bash taskstation secrets set STRIPE_API_KEY=sk_live_new... ``` ### TaskStation pushes it to running sessions TaskStation pushes the change to every sandbox with an active session for the project, on a best-effort basis. For model or gateway credentials, TaskStation restarts the OpenCode compatibility process. Rotating an egress-enforced secret needs no push of the value at all: the sandbox holds a handle, and TaskStation reads the current value when the next request comes through. ## Remove a secret Run `taskstation secrets unset STRIPE_API_KEY` (or `unset `), or delete it from the project's Secrets page. > **Removal is immediate, propagation is not** > TaskStation deletes a shared secret right away. Push to already-running > sandboxes is best-effort, the same as rotation. ## Share a value without seeing it Run `taskstation secrets request STRIPE_API_KEY` to create a link. Anyone with the link can enter the value. You never see it. Links stay valid for 7 days by default; adjust with `--expires ` (max 30 days). An expired link shows a clear "expired" page — mint a fresh one with the same command. ## End-to-end example `taskstation.yaml`: ```yaml taskstation_version: 2 default_agent: my-agent agents: my-agent: secrets: [STRIPE_API_KEY] ``` Secret configuration. The first command stores the value with the default environment exposure. The second moves it to egress-enforced exposure, which is experimental and needs the `secrets_egress` flag (Settings → Feature flags); it returns `403` `feature_disabled` while the flag is off. ```bash taskstation secrets set STRIPE_API_KEY=sk_live_... taskstation secrets delivery STRIPE_API_KEY egress --allow-host api.stripe.com ``` The agent then calls Stripe with the variable it was given: ```bash curl -s -o /dev/null -w '%{http_code}\n' \ -H "authorization: Bearer $STRIPE_API_KEY" \ https://api.stripe.com/v1/customers # expected: 200 ``` `$STRIPE_API_KEY` holds a handle. Stripe receives the real key, because `api.stripe.com` is on the list. The same request to a host that is not on the list sends the handle, and the upstream rejects it. The same request from an agent whose `secrets` list omits `STRIPE_API_KEY` returns `401`. That session starts normally — an ungranted secret is not an error, it is simply never delivered. ## CLI commands | Command | What it does | |---|---| | `taskstation secrets ls` | List secrets declared and set for the project | | `taskstation secrets set KEY=VALUE [--identifier ]` | Create or update a secret. `KEY=-` reads the value from stdin | | `taskstation secrets unset IDENTIFIER` | Remove a secret | | `taskstation secrets delivery IDENTIFIER egress --allow-host ` | Egress-enforced exposure for the listed hosts. Experimental; needs the `secrets_egress` flag | | `taskstation secrets delivery IDENTIFIER runtime` | Environment exposure (the default) | | `taskstation secrets delivery IDENTIFIER denied` | Disabled | | `taskstation secrets call IDENTIFIER URL` | Send one policy-bound HTTPS request through TaskStation. Experimental; needs the `secrets_egress` flag | | `taskstation secrets sync` | Re-push project secrets to this session's sandbox | | `taskstation secrets request NAME [--scope runtime\|connector] [--expires ]` | Create a link so someone else can enter a value | | `taskstation env push --from ` | Upload a `.env` file as secrets | | `taskstation env pull [--out ] [--force]` | Export secret names, not values, to a `.env` file | The CLI keeps the stored vocabulary: `runtime` is environment exposure, `egress` is egress-enforced, `denied` is disabled. Run `taskstation secrets --help` for every flag. Setup links default to `connector`. Use `--scope runtime` only when the agent's shell must receive the value. A secret bound to a connector stays server-side. ## Names and permissions Format rules for the two names: | Name | Format | |---|---| | `identifier` | `^[A-Za-z0-9][A-Za-z0-9_.-]{0,127}$` | | `name` | `^[A-Z_][A-Z0-9_]{0,63}$` | Reading or writing a secret needs the `project.secret.read` or `project.secret.write` permission. A project manager holds both. A custom role only adds permissions — TaskStation has no deny rule — so no role can withhold either one from a manager. To restrict a manager, remove the manager role. Delivery to a session is the role verdict of the person who started it, intersected with the launched agent's manifest grant. Both must allow the secret. See [One vocabulary, two bindings](/docs/accounts#one-vocabulary-two-bindings). ## REST routes All routes sit under `/v1/projects/{projectId}`. | Method | Path | Description | |---|---|---| | GET | `/secrets` | List secrets. Scoped to the caller's grant if the caller is a scoped agent token. | | POST | `/secrets` | Create or update the shared value. Body: `{name, identifier?, value}`, plus an optional delivery policy. | | PUT | `/secrets/{identifier}/strategy` | Change the exposure and its host list. | | POST | `/secrets/{identifier}/grant` | Add the identifier to one agent's `secrets` list in `taskstation.yaml`. Body: `{agent}`. | | DELETE | `/secrets/{name}` | Delete the shared value. Personal overrides stay in place. | | PUT | `/secrets/{name}/personal` | Set or turn on the caller's personal override. | | DELETE | `/secrets/{name}/personal` | Remove the caller's personal override. | `POST /secrets` rejects names that start with `TASKSTATION_`. It returns `409` if the `identifier` already exists with a different `name`. It rejects the exact name `CODEX_AUTH_JSON` with `400` — TaskStation manages that secret through ChatGPT subscription onboarding. Both write routes return a `delivery_sync` object when the change had to reach running sandboxes. `ok: false` means the value is saved but at least one live session still uses the previous one; the listed sessions pick it up on restart. A secret in the list carries `delivery_blocked_reason`. The value `no_agent_grant` means no agent can receive this secret. `null` means it is granted, the exposure needs no grant, or TaskStation could not read the manifest. `POST /secrets/{identifier}/grant` clears that reason. It returns `already_granted: true` when the agent's list already admits the identifier, in which case TaskStation commits nothing. It returns `adopted_governance: true` when the edit added the project's first `agents:` block — the change [described above](#grant-a-secret-to-an-agent). It answers `400` `manifest_v1_unsupported` for a `taskstation.toml` project and `409` `secret_not_grantable` for a disabled secret. ## Rotation and propagation A secret write does not wait for a session restart. TaskStation pushes the change to every active sandbox in the project: 1. TaskStation builds a new environment snapshot, using the running agent's `secrets` grant. 2. The sandbox writes the snapshot to the live agent environment. New tool calls pick up the change right away. 3. If the changed secret is an LLM provider credential, TaskStation restarts OpenCode. This push is best-effort. The API call that changes the secret returns before the push finishes. A failed push is only logged, not retried. A sandbox with a failed push keeps the old value until the next successful push, or until the session restarts. ## Model credentials A project on TaskStation's managed model access needs no key of its own. To bring your own, set the provider variables your OpenCode provider config references. A recognized model key is assigned the **LLM gateway** usage, which spends it server-side; it needs no sandbox presence. Do not use a generic provider verification result as runtime proof. It cannot prove the selected model, region, entitlement, and API dialect. Send a real prompt through the exact model. --- # Quickstart Install the CLI, run a session, and merge your first change request. Canonical page: https://taskstation.co/docs/quickstart This page takes you from a fresh terminal to a merged change request (CR). Every command below is copy-pasteable. ### Install the CLI Run the install script. It downloads the `taskstation` binary for your OS. ```sh curl -fsSL https://taskstation.co/install | bash ``` TaskStation supports macOS and Linux. There is no Windows binary yet. ### Sign in ```sh taskstation login ``` This opens your browser for sign-in. After you sign in, `taskstation login` picks your account and a default [project](/docs/project) for you. ### Create or clone a project A new account has no projects yet, so `taskstation login` has none to pick. Run `taskstation init` with a name to create your first project. This scaffolds a project directory on your machine. ```sh taskstation init my-app cd my-app taskstation ship ``` `taskstation ship` creates the project in TaskStation Cloud on its first run, then pushes your code. Run `taskstation ship` again any time to sync local changes. To work on a project that already exists in TaskStation Cloud, clone it instead: ```sh taskstation projects clone ``` Find `` with `taskstation projects ls`. ### Run a session ```sh taskstation sessions new --prompt "Build the login page" ``` This starts a [session](/docs/work/sessions). The agent works in its own sandbox, on its own branch, so your project is not touched yet. Attach to the session to watch progress and reply: ```sh taskstation sessions chat ``` ### Review and merge the change request When the agent finishes, it opens a change request with a summary and the exact diff. ```sh taskstation cr ls ``` This lists the open CRs in your project. Review the one the agent opened, then merge it to land the work on your default branch: ```sh taskstation cr merge 1 ``` Replace `1` with the CR number from `taskstation cr ls`. Nothing reaches your project's default branch until you merge. ## Next steps - [How the pieces fit](/docs/work) Projects, sessions, and change requests, explained. - [CLI command reference](/docs/cli) Every command, flag, and subcommand. - [The CLI dev loop](/docs/cli) Link a repo, ship code, and manage sessions from your terminal. --- # Apps Create, deploy, and control TaskStation Apps from the SDK and from React. Canonical page: https://taskstation.co/docs/sdk/apps Apps are project-scoped serverless deployments. This page covers the SDK surface: the `apps` facade on a project handle, the artifact and deployment calls, the access calls, the exported types, and the React hooks. For what an App is, the source kinds, the CLI, the stable URL, and cold-wake behavior, read [Apps](/docs/feature-flags/apps). ```ts const apps = taskstation.project(projectId).apps; ``` > **Apps is a feature flag** > Every Apps route answers `403` with `{ error, code: "feature_disabled", feature: "apps" }` > until the project turns Apps on. Use `isFeatureDisabledError(error)` to branch > on it. See [Feature flags](/docs/feature-flags). ## The apps facade | Method | Wraps | What it does | |---|---|---| | `apps.list()` | `GET /projects/:pid/apps` | Lists the project's Apps | | `apps.create(input)` | `POST …/apps` | Creates an App and assigns its stable URL | | `apps.get(appId)` | `GET …/apps/:id` | Reads one App | | `apps.update(appId, input)` | `PATCH …/apps/:id` | Renames it or changes machine, idle timeout, or budget | | `apps.remove(appId)` | `DELETE …/apps/:id` | Deletes the App and its runtimes | | `apps.start(appId)` | `POST …/apps/:id/start` | Sets `desired_state` to `running` and warms the runtime | | `apps.stop(appId)` | `POST …/apps/:id/stop` | Suspends compute now; the next request resumes it | | `apps.rollback(appId, deploymentId)` | `POST …/apps/:id/rollback` | Moves traffic to a ready deployment | Artifacts are the immutable input to a deployment: | Method | Wraps | What it does | |---|---|---| | `apps.artifacts.register(input)` | `POST …/apps/artifacts` | Registers an `archive` or an `oci_image`; returns the upload URL for an archive | | `apps.artifacts.uploadArchive(bytes, options?)` | — | Registers, uploads, hashes, and finalizes one `.tar.gz` in a single call | | `apps.artifacts.finalize(artifactId, input)` | `POST …/apps/artifacts/:id/finalize` | Confirms `sha256` and `size_bytes` for a manual upload | Deployments are immutable and numbered: | Method | Wraps | What it does | |---|---|---| | `apps.deployments.create(appId, input)` | `POST …/apps/:id/deployments` | Starts a deployment from an artifact and a source | | `apps.deployments.list(appId)` | `GET …/apps/:id/deployments` | Lists the deployment history | | `apps.deployments.get(appId, deploymentId)` | `GET …/deployments/:did` | Reads one deployment plus its events | | `apps.deployments.logs(appId, deploymentId, options?)` | `GET …/deployments/:did/logs` | Reads runtime logs with a cursor | Access is the App's own authorization policy: | Method | Wraps | What it does | |---|---|---| | `apps.access.get(appId)` | `GET …/apps/:id/access` | Reads the policy. Needs `project.customize.write` | | `apps.access.update(appId, input)` | `PATCH …/apps/:id/access` | Replaces the policy and bumps its revision | | `apps.access.session(appId)` | `POST …/apps/:id/access-session` | Mints a five-minute URL that exchanges into a host-only cookie | ## Deploy a static site `uploadArchive` does the whole artifact handshake: it registers the artifact, checks it against `max_bytes`, `PUT`s the bytes, computes the SHA-256, and finalizes. ```ts const apps = taskstation.project(projectId).apps; const app = await apps.create({ slug: 'docs', name: 'Docs' }); const artifact = await apps.artifacts.uploadArchive(tarGzBytes, { onProgress: (uploaded, total) => console.log(`${uploaded}/${total}`), }); const deployment = await apps.deployments.create(app.app_id, { artifact_id: artifact.artifact_id, source: { kind: 'static', spa: true }, }); console.log(app.url, deployment.status); // https://…apps.taskstation.co queued ``` `create` accepts the machine and budget fields too: `cpu`, `memory_gb`, `disk_gb`, `idle_timeout_seconds`, and `monthly_budget_usd`. Omit them for the defaults. Wait for the deployment by polling its status: ```ts async function waitForReady(appId: string, deploymentId: string) { for (;;) { const { deployment } = await apps.deployments.get(appId, deploymentId); if (deployment.status === 'ready') return deployment; if (deployment.status === 'failed' || deployment.status === 'cancelled') { throw new Error(deployment.error ?? deployment.error_code ?? deployment.status); } await new Promise((resolve) => setTimeout(resolve, 2000)); } } ``` The status values are `queued`, `validating`, `building`, `provisioning`, `checking`, `ready`, `failed`, and `cancelled`. ## Deploy an OCI image Register the immutable image reference, then declare the process command and the public target port: ```ts const registered = await apps.artifacts.register({ kind: 'oci_image', image: 'ghcr.io/acme/service:2026-08-07', }); await apps.deployments.create(app.app_id, { artifact_id: registered.artifact.artifact_id, source: { kind: 'oci_image', image: 'ghcr.io/acme/service:2026-08-07', command: ['node', 'server.js'], port: 3000, readiness_path: '/health', }, }); ``` `register` returns `upload: null` for an `oci_image`. Only an `archive` gets an upload URL. `CreateAppDeploymentInput` also accepts `environment` (non-secret runtime values), `secrets` (runtime key to project secret name), and `provider` (`'daytona' | 'platinum' | 'e2b' | 'northflank'`). Omit `provider` to use the server policy. ## Read runtime logs ```ts let cursor = 0; for (;;) { const page = await apps.deployments.logs(app.app_id, deployment.deployment_id, { after: cursor, limit: 200, }); for (const entry of page.entries) console.log(entry.source, entry.line); cursor = page.next_cursor; if (page.entries.length === 0) break; } ``` Each entry carries `cursor`, `time`, `source` (`app`, `appd`, `caddy`), and `line`. ## Your App already knows who is looking An App hosted by TaskStation is opened by someone TaskStation **already signed in**. The Apps gate authenticates them before your first byte is served, so your App needs no login of its own — no second password, no consent screen, no redirect. In the browser: ```ts import { createTaskStation, taskstationAppViewerToken } from '@taskstation/sdk'; import { useTaskStationAppViewer } from '@taskstation/sdk/react'; const taskstation = createTaskStation({ backendUrl: 'https://api.taskstation.co/v1', getToken: taskstationAppViewerToken(), // the viewer's own App-scoped token }); function Header() { const { status, viewer } = useTaskStationAppViewer(); return {status === 'viewer' ? viewer.email : 'Signed out'}; } ``` On your App's server, the gate signs the identity into every request: ```ts import { readAppViewer, createAppViewerTaskStation } from '@taskstation/sdk/server'; const viewer = await readAppViewer(request); // { userId, email, groupIds, accountId, appId, accessMode, token } if (!viewer) return new Response('Not found', { status: 404 }); // and, for an `api`-scoped App, act as them: const taskstation = await createAppViewerTaskStation(request, { backendUrl }); await taskstation.projects.list(); // their projects, their role ``` `readAppViewer` verifies an HMAC over the header with `TASKSTATION_APP_VIEWER_SECRET`, which TaskStation injects into your App at deploy. A forged header never passes: the gate deletes any client-supplied copy before forwarding, and the signature is made with a secret derived per App. ### How much your App is told One setting on the App's access policy — **Settings → Access** in TaskStation, or `viewer_token_scope` on `PATCH /projects/:id/apps/:appId/access`: | Scope | Your App receives | |---|---| | `identity` (default) | The viewer's id, email and group ids, plus a `profile email` token. Enough to show each person their own data. | | `api` | The above, and a token that acts **as** that person on the TaskStation API — bounded by their own role. | | `off` | Nothing. | The token is never the user's TaskStation session: it lasts an hour, carries only those scopes, and every token an App minted dies when the App is deleted or its access policy changes. `public` and `password` Apps have no signed-in TaskStation viewer, so they receive none of this. An App served on its **own domain** (not `*.apps.taskstation.co`) has no gate in front of it — use [Sign in with TaskStation](/docs/sdk/sign-in) there instead. ## Manage access ```ts await apps.access.update(app.app_id, { mode: 'restricted', member_ids: [memberId], group_ids: [groupId], }); const preview = await apps.access.session(app.app_id); window.open(preview.url); // valid for five minutes ``` `AppAccessConfig` reports `password_configured`, never the password or its hash. Set a password with `{ mode: 'password', password }`. Each update increments `revision`, which revokes existing App cookies. ## Types Every type below is exported from `@taskstation/sdk`. | Type | What it holds | |---|---| | `App` | Identity, `url`, `access_mode`, `access_revision`, `desired_state`, `active_deployment_id`, `machine`, `idle_timeout_seconds`, `monthly_budget_usd`, `last_request_at`, `viewer_can_access` | | `AppDeployment` | `version`, `status`, `source_kind`, `hosting_provider`, `runtime_spec`, `build_spec`, `error_code`, `attempt_count`, `created_by`, `actor_type`, `source_session_id` | | `AppDeploymentDetail` | One `deployment` plus its `events` | | `AppAccessConfig` | `mode`, `revision`, `member_ids`, `group_ids`, `password_configured` | | `AppAccessMode` | `'private' \| 'project' \| 'restricted' \| 'public' \| 'password'` | | `AppSource` | `StaticAppSource \| BundleAppSource \| DockerfileAppSource \| OciImageAppSource` | | `AppArtifact` | `kind`, `status`, `sha256`, `size_bytes`, `image_reference` | | `AppLogEntry` · `AppLogsResponse` | One log line, and one page plus `next_cursor` | `viewer_can_access` answers whether the caller may OPEN the App, which is not the same as whether they can see it listed. A project manager sees every App in the project so a private one stays manageable when its creator leaves. Check this field before asking for an access session. Treat `undefined` as unknown, not as denied. `AppAccessMode` is a per-resource visibility setting on top of the role model, not a role. `restricted` names users and groups — the same principal types the role model uses. See [Accounts & access](/docs/accounts#per-feature-access-settings). `DockerfileAppSource` and `OciImageAppSource` require `command` and `port`. `StaticAppSource` and `BundleAppSource` do not. ## React hooks `@taskstation/sdk/react` exports three hooks for Apps. ### useProjectApps(projectId) The project's App inventory plus its lifecycle mutations. Every mutation invalidates the inventory on success. ```tsx import { useProjectApps } from '@taskstation/sdk/react'; function AppList({ projectId }: { projectId: string }) { const apps = useProjectApps(projectId); if (!apps.data) return null; return (

); } ``` It returns the query fields plus `create`, `update`, `start`, `stop`, and `remove`. ### useAppDeployments(projectId, appId) The immutable deployment history, refetched every 5 s so a running build advances on its own. ```tsx const deployments = useAppDeployments(projectId, appId); await deployments.deploy.mutateAsync({ artifact_id: artifact.artifact_id, source: { kind: 'static', spa: true }, }); await deployments.rollback.mutateAsync(previousDeploymentId); ``` Both mutations invalidate the deployment list and the App inventory. ### useAppAccess(projectId, appId, options?) The access policy and a short-lived access session. Both halves are separate queries, and each one is optional. ```tsx const access = useAppAccess(projectId, appId, { policy: canEditAccess, session: app.viewer_can_access, }); access.policy.data; // AppAccessConfig access.session.data; // { url, expires_at } await access.update.mutateAsync({ mode: 'project' }); ``` | Option | Default | Use `false` when | |---|---|---| | `policy` | `true` | The surface only previews the App. `GET …/access` is an administrative read and answers `403` for a caller without project-manager permissions. | | `session` | `true` | The caller may see the App but not open it. Pass `app.viewer_can_access`. | A grid of Apps that leaves both options at `true` fires one policy read and one session mint per App, and each is a `403` for a member who may not open that App. ## Errors ```ts import { featureDisabledKey, isFeatureDisabledError } from '@taskstation/sdk'; try { await taskstation.project(projectId).apps.list(); } catch (error) { if (isFeatureDisabledError(error)) { console.log(`${featureDisabledKey(error)} is off for this project`); } } ``` Other answers you should handle: `409` for a duplicate slug, `402` with `app_quota_exceeded` when the account is at its App limit, and `400` with `app_machine_out_of_range` or `app_budget_out_of_range` for a spec outside its bounds. --- # Authentication Authenticate the SDK with a personal access token or a service account. Canonical page: https://taskstation.co/docs/sdk/auth TaskStation accepts one bearer token per request. Pass it through `getToken` in `createTaskStation`. The SDK sends it as `Authorization: Bearer `. ```ts import { createTaskStation } from '@taskstation/sdk'; const taskstation = createTaskStation({ backendUrl: 'https://api.taskstation.co/v1', getToken: async () => process.env.TASKSTATION_API_KEY!, }); ``` `backendUrl` and `getToken` are required. The SDK caches nothing: it calls `getToken` on every request, so your app owns token storage and refresh. Set `clientSource` to `api`, `cli`, `mobile`, or `web` when your host needs a separate source in the centralized audit log. This value identifies the client surface. It does not change the authenticated actor or their permissions. ## Personal access tokens A personal access token (PAT) is the credential for the SDK, the CLI, and CI. Create one in your own settings, at **Settings → API keys** (`/settings/tokens`). The key starts with `taskstation_pat_` and shows only once, at creation. Store it as a secret. A PAT acts as the user who created it and holds exactly that user's role assignments. It adds no access of its own. Its scope only narrows the reach: | Scope | Reach | |---|---| | Account (default) | Every project in the account | | Project | One project only; every other project returns `403` | Choose the project scope for CI and other narrow-purpose credentials. `taskstation login` mints a PAT and stores it locally — it is the same credential type, not a separate token kind. ## Service accounts A service account is a separate credential family for non-human callers, prefixed `taskstation_sa_`. Create one at **Account → Tokens** (`/accounts/?tab=tokens`), the account-level surface for credentials that are not a person's. A service account is its own **principal** (`service_account`), not a person's credential. It has no membership, so it holds only the roles assigned to it directly. An agent's identity is a service account, which is how you assign a role to an agent. See [Accounts & access](/docs/accounts#one-access-model). > **Warn** > A new service account has no assignments and therefore no project access. If > you point `getToken` at one before you assign it a role, every call returns > `403 "You do not have access to this project"`. Assign it a project role > first, or use a personal access token for the SDK, the CLI, and demos > instead. ## OAuth access tokens (Sign in with TaskStation) A third-party app that signs users in through TaskStation receives a `taskstation_oat_` token per user. With the `taskstation` scope it acts as that user on the whole API, exactly like a personal access token, but it expires after an hour and rotates through a refresh token. `createTaskStationAuth` in `@taskstation/sdk/server` owns the whole lifecycle — see [Sign in with TaskStation](/docs/sdk/sign-in). ## Supabase JWT If your app uses TaskStation's own sign-in, return the live session token instead of a PAT: ```ts getToken: async () => (await supabase.auth.getSession()).data.session?.access_token ?? null, ``` The SDK calls `getToken` on every request, so a refreshed token takes effect automatically. ## Headless sign-in (email, password, magic link, social) Every ordinary sign-in flow is available through the TaskStation API, so a CLI, a native app, a script, or your own backend signs users up and in without a Supabase URL or key — on taskstation.co and on a self-host alike. ```ts import { createTaskStation } from '@taskstation/sdk'; const session = createTaskStation({ backendUrl, getToken: async () => null }).auth.session({ storage: { // optional: any get/set/remove get: () => localStorage.getItem('taskstation'), set: (v) => localStorage.setItem('taskstation', v), remove: () => localStorage.removeItem('taskstation'), }, }); const taskstation = createTaskStation({ backendUrl, getToken: session.getToken }); // refreshes itself const { session: s, user } = await taskstation.auth.signInWithPassword({ email, password }); await session.set(s, user); await taskstation.projects.list(); // as that user ``` | Call | Route | Notes | |---|---|---| | `auth.signUp({ email, password, redirect_to? })` | `POST /v1/auth/signup` | `requires_email_confirmation: true` → no session until the emailed link/code is used. | | `auth.signInWithPassword({ email, password })` | `POST /v1/auth/sign-in/password` | | | `auth.sendMagicLink({ email, redirect_to? })` → `auth.verifyOtp({ email, token, type: 'magiclink' })` | `/sign-in/magic-link`, `/verify-otp` | The email carries a link and a 6-digit code. | | `auth.signInWithProvider({ provider, redirect_to })` → `auth.exchangeCode({ code, code_verifier })` | `/sign-in/oauth`, `/oauth/exchange` | PKCE: keep `code_verifier` until the provider redirects back with `?code=`. `redirect_to` must be on the instance's redirect allow-list. | | `auth.refresh({ refresh_token })` | `POST /v1/auth/refresh` | `createTaskStationSession` calls it for you. | | `auth.resetPassword({ email, redirect_to? })` → `auth.verifyOtp({ type: 'recovery' })` → `auth.updatePassword({ password }, token)` | `/password/reset`, `/verify-otp`, `/password/update` | | | `auth.user(token)` / `auth.signOut(token)` | `GET /v1/auth/user`, `POST /v1/auth/sign-out` | Sign-out revokes at Supabase and in the TaskStation session gate. | Errors throw `HeadlessAuthError` with `code`, `message` and the upstream `status` (`invalid_credentials`, `over_request_rate_limit`, …). Each route is limited to 30 attempts per minute per IP. Multi-factor enrolment/challenge is not on the API yet — it stays on the TaskStation web app. ## Choose a credential | You are building | Use | |---|---| | A backend, script, or CI job | A personal access token, account-wide or project-scoped | | The TaskStation CLI | `taskstation login` (mints a personal access token) | | A web app, CLI, or native app signing users in itself | `taskstation.auth.*` + `createTaskStationSession` (headless sign-in, above) | | An automated caller with its own project assignment | A service account | | Your own app, signed in by its users with their TaskStation account | [Sign in with TaskStation](/docs/sdk/sign-in) — an OAuth access token the SDK manages for you | --- # Full example One file that lists projects, starts a session, and streams a reply. Canonical page: https://taskstation.co/docs/sdk/example This page shows the OpenCode REST SDK path in one file. It uses no framework and needs no build step beyond TypeScript. Use [`useSession`](/docs/sdk/react) for a React surface. ## The complete script ```ts import { ApiError, classifyTurn, createTaskStation, narrowChatEvent } from '@taskstation/sdk'; import type { MessageWithParts } from '@taskstation/sdk'; async function main() { // 1. One client, one auth seam. getToken returns your API key // (taskstation_pat_…) or a logged-in user's Supabase JWT — nothing else. const taskstation = createTaskStation({ backendUrl: 'https://api.taskstation.co/v1', getToken: async () => process.env.TASKSTATION_API_KEY!, }); // 2. Platform REST: list projects, pick one (or provision your first). const projects = await taskstation.projects.list(); const project = projects[0] ?? (await taskstation.projects.provision({ name: 'sdk-quickstart' })); console.log(`using project ${project.name} (${project.project_id})`); // 3. Create a session — a cheap platform call. No sandbox exists yet. const created = await taskstation.projects.createSession(project.project_id, { name: 'sdk full example', }); const session = taskstation.session(project.project_id, created.session_id); // 4. Ready the session. This provisions (or resumes) the real cloud // sandbox. ensureReady() polls /start (each call long-polls up to 30s) // until the runtime is ready or its deadline (~3 min) elapses, so a // cold boot just takes longer rather than throwing. The // retryUntilReady wrapper below is optional — keep it only if you want // a longer total budget than the default. const { opencodeSessionId } = await retryUntilReady(() => session.ensureReady()); // 5. Connect the event stream before you send, so no early events are // missed. narrowChatEvent() collapses the wire events into a small // typed union you can switch over. let resolveIdle!: () => void; const idle = new Promise((resolve) => (resolveIdle = resolve)); const stream = await session.stream({ onEvent: (event) => { const e = narrowChatEvent(event); if (!e) return; if (e.type === 'message.part.updated') process.stdout.write('.'); if (e.type === 'session.error') console.error('\nerror:', e.error); if (e.type === 'session.idle' && e.sessionID === opencodeSessionId) { resolveIdle(); // the turn is finished } }, }); // 6. Send. Per-send overrides pick the model and the agent for this // prompt only (ids come from projects.modelPicker() and // projects.detail().config.agents). await session.send('What files are in this repo?', { model: { providerID: 'taskstation', modelID: 'glm-5.3-flash' }, }); // 7. Wait for the turn to finish — the session.idle event, not a sleep. await idle; stream.close(); // 8. Render the transcript. classifyTurn() turns the wire part variants // into one union, so a renderer can switch on part.kind and // TypeScript proves no case is missed. const result = await session.runtime.session.messages({ sessionID: opencodeSessionId, }); for (const message of (result.data ?? []) as MessageWithParts[]) { for (const part of classifyTurn(message).parts) { if (part.kind === 'text') console.log(`\n[${message.info.role}] ${part.text}`); } } } /** Optional outer-budget wrapper — ensureReady() already polls internally. */ async function retryUntilReady(ensure: () => Promise): Promise { const deadline = Date.now() + 300_000; for (;;) { try { return await ensure(); } catch (error) { const provisioning = error instanceof ApiError && error.code === 'RUNTIME_UNAVAILABLE'; if (!provisioning || Date.now() > deadline) throw error; await new Promise((r) => setTimeout(r, 3_000)); } } } main().catch((error) => { console.error(error); process.exit(1); }); ``` Run it with Node 18 or later, Bun, or `tsx`: ```sh TASKSTATION_API_KEY=taskstation_pat_... npx tsx full-example.ts ``` The first `ensureReady()` call on a fresh session provisions a real cloud sandbox, so the ready step takes a while on the first run. Later runs resume the same sandbox and finish fast. ## What each step teaches | Step | Concept | Deep dive | | ---- | ---------------------------------------------------- | ----------------------------------- | | 1 | One client, one token, one auth seam | [Authentication](/docs/sdk/auth) | | 2–3 | The platform REST surface: projects and sessions | [Reference](/docs/sdk/reference) | | 4 | Session readiness, the bridge from platform to runtime | [Sessions](/docs/sdk/sessions) | | 5, 7 | Live SSE events, `narrowChatEvent`, `session.idle` | [Sessions](/docs/sdk/sessions) | | 6 | Per-send `{ model, agent }` overrides | [Sessions](/docs/sdk/sessions) | | 8 | `classifyTurn` and the exhaustive part union | [Reference](/docs/sdk/reference) | ## Going further from here The same client reaches the rest of the platform through one facade. ```ts const project = taskstation.project(projectId); // Workspace files inside the session's live sandbox const tree = await session.files.list('/workspace'); const readme = await session.files.read('/workspace/README.md'); // Runtime secrets are readable by selected sessions inside the sandbox. await project.secrets.upsert({ name: 'LOCAL_TOOL_TOKEN', value: 'secret-value', strategy: 'runtime', consumer: 'sandbox', }); // Managed provider credentials stay on the TaskStation LLM gateway. await project.secrets.upsert({ identifier: 'anthropic-primary', name: 'ANTHROPIC_API_KEY', value: 'sk-ant-…', strategy: 'broker', consumer: 'llm_gateway', }); // The project's agents and skills (config files in the repo) const { config } = await taskstation.projects.detail(projectId); console.log( config.agents.map((a) => a.name), config.skills.length, ); // LLM gateway observability — cost, latency, per-model breakdown const overview = await project.gateway.overview(7); const routing = await project.gateway.routing.get(); ``` The package ships runnable examples that cover each of these steps, in `packages/sdk/examples/`. They include a minimal client, streaming, a server wrapper, a TaskStation-as-a-Backend multi-tenant wrapper, transcript rendering, and files and secrets. --- # TypeScript SDK Install, authenticate, and send your first message with the typed SDK. Canonical page: https://taskstation.co/docs/sdk `@taskstation/sdk` is the typed client for the TaskStation platform. It wraps the TaskStation REST API and OpenCode REST runtime in one interface. The core client is fetch-based and runs in Node, Bun, and browsers. ## Install ```bash npm install @taskstation/sdk ``` `react` (18+) and `@tanstack/react-query` (5.75+) are optional peers, needed only for [React hooks](/docs/sdk/react). ## Create a client Call `createTaskStation` once, with your API base URL and a function that returns your token. ```ts import { createTaskStation } from '@taskstation/sdk'; export const taskstation = createTaskStation({ backendUrl: 'https://api.taskstation.co/v1', getToken: async () => process.env.TASKSTATION_API_KEY!, }); ``` `backendUrl` and `getToken` are the only required fields. The SDK calls `getToken` on every request and caches nothing — your host owns token storage and refresh. Create an API key in your own settings, at **Settings → API keys** (`/settings/tokens`). The key starts with `taskstation_pat_` and shows only once. Store it as a secret and return it from `getToken`. See [Auth](/docs/sdk/auth) for token types and scopes. ## Call a Connector A Connector defines callable tools. A Connection stores one authorization for that Connector. Credentials stay server-side. ```ts const connectors = taskstation.project(projectId).connectors; await connectors.catalog(); await connectors.search('send email'); await connectors.describe('gmail.send_email'); await connectors.call('gmail.send_email', { to, subject, body }); ``` An agent-minted session token already carries its project scope. Use `taskstation.connectors` when the agent does not have a separate `projectId` value. ## Start your first session 1. Create a session in your project. ```ts const created = await taskstation.project(projectId).sessions.create(); const session = taskstation.session(projectId, created.session_id); ``` 2. Wait for the sandbox to accept work. ```ts await session.ensureReady(); ``` `ensureReady()` starts or resumes the session sandbox. It polls the session's `/start` endpoint — each call long-polls up to 30 s — until the runtime is ready, hits a terminal stage, or its deadline elapses (default ~3 min, configurable via `{ readyTimeoutMs }`). On a cold boot it keeps polling while the sandbox reports `retriable: true`; it only throws an `ApiError` with `code: 'RUNTIME_UNAVAILABLE'` if the runtime is still not ready when the deadline expires. See [Sessions](/docs/sdk/sessions). 3. Send a message to the agent. ```ts await session.send('Add a README'); ``` `send()` calls `ensureReady()` for you, then sends the message. `createTaskStation` gives you an imperative client: call methods for every action, like projects, sessions, secrets, and triggers. `@taskstation/sdk/react` gives you hooks for live UI data — `useSession` runs a whole session in one hook. - [Full example](/docs/sdk/example) Zero to a streaming agent reply. - [Auth](/docs/sdk/auth) API keys and Supabase JWTs. - [Sessions](/docs/sdk/sessions) Lifecycle, streaming, and error handling. - [React hooks](/docs/sdk/react) `useSession` and other reactive hooks. - [Reference](/docs/sdk/reference) The full client, modules, turns, and distribution. --- # React hooks Run a TaskStation session in React with the useSession hook. Canonical page: https://taskstation.co/docs/sdk/react `@taskstation/sdk/react` adds React hooks on top of the SDK. This page covers `useSession`, the hook that runs a session end to end, and the other hooks confirmed stable for React apps. ## useSession(projectId, sessionId, options?) `useSession` starts the session, opens the server-selected event transport, and syncs messages, status, and pending prompts. Call it once per session view. ```tsx import { useSession } from '@taskstation/sdk/react'; function Chat({ projectId, sessionId }: { projectId: string; sessionId: string }) { const s = useSession(projectId, sessionId); if (s.phase !== 'ready') return ; return ( <> {s.messages.map(({ info, parts }) => ( ))} ); } ``` Readiness is server truth. The runtime is ready when `POST /start` returns `stage: 'ready'`. `useSession` does not run a separate client-side health check. ### Returns | Field | Type | What it holds | |---|---|---| | `phase` | `'starting' \| 'ready' \| 'error'` | Overall state. Render a boot screen until `ready`. | | `messages` | `{ info, parts }[]` | The message list. Parts stream in live. | | `status` | `SessionStatus` | The session status. | | `isBusy` | `boolean` | The agent is generating a reply. | | `questions`, `permissions` | array | Pending agent questions and tool-approval requests. A `permission` here is one runtime tool approval, not an IAM permission. | | `diffs`, `todos` | array | Live file diffs and todo items. | | `sendError` | `TaskStationSendError \| null` | The last `send` failure: `billing`, `runtime-not-ready`, or `runtime-error`. | | `rewindMessageId` | `string \| null` | The selected user message while a reversible rewind is staged. | | `rewindPending` | `boolean` | A rewind or restore request is in progress. | | `rewindError` | `TaskStationSendError \| null` | The last rewind or restore failure. | | `models`, `agents`, `defaultAgent`, `commands` | — | Selectable models, selectable agents, the default agent, and slash commands. Available before the runtime starts. | | `retry` | `() => void` | Force a re-check of `/start`. | ### Actions | Action | What it does | |---|---| | `send(text, override?)` | Send a prompt. `override` sets `{ model?, agent? }` for this message only. | | `sendParts(parts, override?)` | Send text and file prompt parts through the selected transport. | | `rewind(messageId)` | Rewind this canonical session to a user message. The selected message and later path become hidden and recoverable. | | `restoreRewind()` | Restore the removed path before another prompt commits its replacement. | | `cancel()` | Stop the current run and clear pending questions and permissions. | | `runCommand(command, args)` | Run a project slash command. | | `answerQuestion(id, answers)` | Answer a pending agent question. | | `rejectQuestion(id)` | Reject a pending agent question. | | `answerPermission(id, reply, message?)` | Answer a tool-approval request. `reply` is `'once'`, `'always'`, or `'reject'`. | `useSession` also returns `removeQuestion` and `removePermission`. Do not use them. They clear the prompt from local state but never notify the agent, so the run stays blocked. Use `answerQuestion`, `rejectQuestion`, or `answerPermission` instead. ### Options | Option | Default | What it does | |---|---|---| | `waitMs` | `15000` | The long-poll budget sent to `/start`. | | `replayStartStash` | `true` | Replay a prompt saved before the session existed, once the session is ready. | | `enabled` | `true` | Set `false` to delay the hook, for example until a billing check passes. | | `chatEngine` | `true` | Set `false` if your app mounts its own chat surface for this session, to avoid syncing messages twice. | Sending is optimistic. `send` shows your message right away, then stream events fill in the agent's reply. `rewind(messageId)` never creates a session. It uses the canonical session from `POST /start`. The runtime restores file state and keeps the removed transcript path recoverable. The next accepted prompt commits the replacement path. ## Other stable hooks `@taskstation/sdk/react` also exports React Query hooks for data that does not need a running session. Each mirrors a method on the [client](/docs/sdk/reference) and needs no provider. | Hook | Reads | |---|---| | `useProjectModels(projectId)` | Selectable models for the project. | | `useProjectModelPickerCatalog(projectId)` | The raw `/model-picker` record per wire model (`reasoning_options`, `temperature`, `limit`) for capability-gated controls. | | `useVisibleAgents({ projectId })` | The project's visible agents. | | `useProjectConfig(projectId)` | The project's runtime config: default agent, commands. | | `useProjectSecrets(projectId)` | Secrets: list, add, remove, and personal overrides. | | `useProjectTriggers(projectId)` | Triggers: list, create, update, remove, fire. | | `useChangeRequests(projectId, status?)` | Change requests: list, open, merge, close, request changes. | ## Next - [Sessions](/docs/sdk/sessions) — the session handle `useSession` wraps, and the `TaskStationSendError` kinds. - [Reference](/docs/sdk/reference) — the full REST surface these hooks read from. --- # SDK reference The full @taskstation/sdk API surface — client methods, modules, turns, and distribution. Canonical page: https://taskstation.co/docs/sdk/reference This page is the full `@taskstation/sdk` API surface: every client method, the framework-free modules, the turns helpers, and how the package ships. Use [SDK](/docs/sdk) to get started and [Sessions](/docs/sdk/sessions) for the session lifecycle in depth. ## The client `createTaskStation(config)` returns one client. Every method is a typed call to the platform API. The `project(id)` and `session(pid, sid)` handles bind ids so you never repeat them. ```ts const taskstation = createTaskStation({ backendUrl, getToken }); taskstation.accounts; // account / team operations taskstation.accountInvites; // invite accept/decline by token alone taskstation.projects; // top-level project operations taskstation.connectors; // Connector calls scoped by an agent-minted token taskstation.project(id); // id-bound project handle taskstation.session(pid, sid); // id-bound session handle → see Sessions taskstation.github; // GitHub App install + repo linking taskstation.billing; // credits, subscription, tier, transactions taskstation.sandboxShares; // public share links for a sandbox port taskstation.connectStatus; // easy-connect (Pipedream) status taskstation.marketplace; // public marketplace catalog taskstation.validateToken; // pasted-API-key check taskstation.config; // platform config in effect taskstation.runtime(); // OpenCode REST compatibility client ``` ### Accounts — `taskstation.accounts` | method | what | | --- | --- | | `list()` · `get(accountId)` | accounts you belong to · one account | | `create({ name })` · `updateName(accountId, name)` | create · rename an account | | `branding.get(accountId)` · `branding.update(accountId, { app_name })` · `branding.uploadAsset(accountId, kind, file)` · `branding.removeAsset(accountId, kind)` · `branding.reset(accountId)` | organization branding (Enterprise): own logo / icon / favicon and product name; `kind` is `logo` · `icon` · `favicon`, or `logo_dark` · `icon_dark` · `favicon_dark` for the dark-scheme variant | | `members(accountId)` · `invite(accountId, input)` | list members · invite one | | `updateMemberRole(accountId, userId, role)` · `removeMember(accountId, userId)` | assign the account role · remove a member | | `invites(accountId)` | pending invites | | `cancelInvite(accountId, inviteId)` · `resendInvite(accountId, inviteId)` | cancel · resend a pending invite | | `leave(accountId)` | leave the account | `accounts.tokens` mints account-scoped API keys (`taskstation_pat_...`). See [SDK auth](/docs/sdk/auth) for the full token model. | method | what | | --- | --- | | `tokens.list(accountId?, options?)` | list API keys — the whole account's, or only your own with `{ mine: true }` | | `tokens.create(input)` | mint one — `{ accountId?, name, expiresAt?, projectId? }` | | `tokens.revoke(tokenId, accountId?)` | revoke one | `accounts.audit` is the enterprise reconstruction log. It combines authenticated API requests with semantic session, connector, approval, and computer events. | method | what | | --- | --- | | `audit.log(accountId, filters?)` | list events by project, session, actor, source, outcome, request, correlation, resource, action, or time | | `audit.export(accountId, filters?)` | export the same filtered event stream as CSV or JSONL | | `audit.webhooks.list/create/update/remove(...)` | manage signed SIEM webhooks for the centralized stream | Each event includes `project_id`, `session_id`, `actor_type`, `source`, `outcome`, `request_id`, `trace_id`, and `correlation_id` when the action supplies them. The API does not store request bodies, prompts, secrets, credentials, or raw connector arguments in the centralized event. Connector events can include a bounded argument preview that redacts credential-shaped fields and opaque data. ### Access assignments TaskStation has one grant record: an **assignment**. It binds one principal (`user`, `group`, `service_account`, or `pending`) to one role, at one scope (`account`, or one `project`), optionally narrowed to one object (`agent`, `skill`, `secret`, `app`, or `trigger`) and optionally carrying an `expires_at`. Group access, per-resource access, and custom-role bindings are all assignments. See [Accounts & access](/docs/accounts#one-access-model) for the model. The canonical REST surface is: | Method + path | Does | | --- | --- | | `GET /v1/accounts/{accountId}/iam/assignments` | list assignments, filtered by principal, scope, object, or role | | `POST /v1/accounts/{accountId}/iam/assignments` | create one assignment | | `DELETE /v1/accounts/{accountId}/iam/assignments/{assignmentId}` | revoke one assignment | | `GET /v1/accounts/{accountId}/iam/permissions` | the permission catalog, as data | | `GET /v1/accounts/{accountId}/iam/roles` · `…/roles/{roleId}/permissions` | roles · one role's permissions | The SDK exposes them as `listAssignments`, `createAssignment`, `revokeAssignment`, and `listPermissions`. A catalog row carries `action`, `scope_type`, `resource_type`, `delegable`, `description`, `area`, `level`, and `implies` — read it instead of hardcoding action strings. Assigning a custom role needs the account's `rbac` entitlement; the route answers `402` with `code: "entitlement_required"` without it. ### Account invites — `taskstation.accountInvites` Reached by invite token alone — the invitee may not be a member yet. | method | what | | --- | --- | | `describe(inviteId)` · `accept(inviteId)` · `decline(inviteId)` | preview · accept · decline an invite | ### Projects — `taskstation.projects` | method | what | | --- | --- | | `list()` · `listForAccount(accountId)` | your projects · projects in an account | | `get(id)` · `detail(id)` | summary · full detail | | `create(input)` · `createRepo(input)` | from an existing `repo_url` · new empty GitHub repo | | `provision(input)` | new project on a new TaskStation-managed repo, seeded with a starter template — `{ name, account_id?, seed_starter?, starter_template?, marketplace_items?, source_item_id?, idempotency_key? }` | | `update(id, input)` · `archive(id)` | update settings · archive | | `llmCatalog(id)` · `modelPicker(id)` | full · compact model catalog for a selector | | `sandboxTemplates(id)` · `sandboxHealth(id)` | sandbox build templates · build health | | `sessions(id)` · `createSession(id, input?)` | list visible sessions · create a session | `provision` creates a new project; it does not start an existing project's sandbox. Start a session instead — see [Sessions](/docs/sdk/sessions). Send `idempotency_key` when a retry is possible — a reload, a second tab, a timeout you retried. `provision` mints a brand-new managed repo per call, so without a key those all create real duplicate projects. Reuse one key for every attempt at a single logical create and the repeats return the project the first attempt made (201, same `project_id`, `push_token: null`). The key identifies the attempt, not the payload — reusing one with a different `name` returns the first project and ignores the new value, so mint a fresh key per distinct create. Creating a second project with the same **name** and no key still works. A repeat that arrives while the first call is still provisioning gets `409` with `code: 'provision_in_flight'` rather than a `project_id` that call may still roll back. Retry with the same key. `project(id).sessions.list({ scope: 'project' })` is a lifecycle inventory for a caller with project-manager permissions. It adds accessible unavailable, warm, and soft-deleted sessions with ownership and runtime-state metadata. Both list scopes omit every session the caller cannot open. ### GitHub — `taskstation.github` Account-scoped GitHub App install and repo linking, not project-scoped. | method | what | | --- | --- | | `getInstallation(accountId)` · `listInstallations(accountId)` | this account's install · installs the user can reach | | `saveInstallation(input)` · `deleteInstallation(accountId, installationId?)` | record · unlink an install | | `listRepositories(accountId, installationId?)` · `listRepositoryBranches(...)` | repos the install can see · branches and the GitHub default | | `linkRepository(input)` | import a repo as a project | ### Billing — `taskstation.billing` Reads for credits, subscription, tier, and transaction history. Checkout, the customer portal, and credit purchases are Stripe flows, app-owned. | method | what | | --- | --- | | `accountState(accountId?)` · `accountStateMinimal(accountId?)` | full · minimal billing state | | `transactions(params?)` · `transactionsSummary(params?)` | history · summarized totals | | `creditBreakdown(accountId?)` · `usageHistory(params?)` | credit balance by source · usage over time | | `sessionCosts.list(options?)` · `sessionCosts.get(sessionId, options?)` | paginated session-cost records · one detailed session ledger | | `tierConfigurations()` | available plan tiers | | `checkout.createSession(input)` · `checkout.confirmSession(sessionId, accountId?)` | start · confirm a Stripe Checkout session | | `subscription.createPortalSession(...)` · `subscription.cancel(...)` · `subscription.reactivate(...)` | open the customer portal · cancel · reactivate | | `subscription.scheduleDowngrade(...)` · `cancelScheduledChange(...)` · `prorationPreview(...)` | schedule · cancel · preview a plan change | | `credits.purchase(input)` · `credits.autoTopupSettings(...)` · `credits.configureAutoTopup(...)` | one-off purchase · read · configure auto-topup | `sessionCosts.list()` accepts `accountId`, `projectId`, `limit`, and `offset`. Each row combines finalized LLM cost and billed sandbox compute cost. `sessionCosts.get()` adds model usage and the discriminated LLM/compute ledger. The list includes a reconciliation total for cost without a session. ### Sandbox shares — `taskstation.sandboxShares` Public share links for one exposed sandbox port. Sandbox-scoped, not project-scoped. | method | what | | --- | --- | | `list(sandboxId)` | active share links | | `create(input)` | create one — `{ sandboxId, port, ttl?, label? }` | | `revoke(sandboxId, token)` | revoke one | ### Marketplace catalog — `taskstation.marketplace` Public catalog browsing, read-only — distinct from `project(id).marketplace`, which installs an item onto a project. | method | what | | --- | --- | | `items(options?)` · `item(id)` · `itemFile(id, path)` | browse · one item · a file inside an item | | `marketplaces()` · `featured()` | all · featured marketplaces | | `sources.list()` · `sources.add(input)` · `sources.remove(id)` | list · add · remove a source | ### The project handle — `taskstation.project(id)` Binds the project id; every sub-resource hangs off it. ```ts const p = taskstation.project(projectId); await p.detail(); await p.update({ name }); await p.llmCatalog(); ``` | method | what | | --- | --- | | `get` · `detail` · `update` · `archive` | read · full detail · update · archive | | `llmCatalog` · `modelPicker` · `sandboxHealth` | model and sandbox-build reads | | `onboardingComplete` | mark project onboarding done | | `validateManifest(raw)` | validate a `taskstation.yaml` (or legacy `taskstation.toml`) manifest server-side | | `gitToken()` | mint a fresh scoped git push token (`409` for a bring-your-own repo) | | `setAgentScope(agentName, scope)` | set an agent's allowed secrets and connectors in the manifest — the second binding, not a role | #### `p.tokens` — project-scoped API keys Auto-minted at session create as `TASKSTATION_TOKEN`; can also be minted by hand. | method | what | | --- | --- | | `list()` | project API keys | | `create(input?)` | mint a new one | | `revoke(tokenId)` | revoke one | #### `p.setupLinks` — agent-minted setup links A link a person opens to enter a secret or connect an app, without full project access. | method | what | | --- | --- | | `requestSecret(input)` · `requestConnector(input)` | link to collect a secret · connect an app | #### `p.secrets` — project secrets | method | what | | --- | --- | | `list()` · `upsert(input)` | list metadata · create or update a write-only value and delivery policy | | `setStrategy(identifier, strategy, options?)` | change the exposure and its host list | | `broker(identifier, request)` | execute a session-authorized, policy-bound HTTPS request | | `remove(identifier)` | delete a secret | | `setPersonal(name, value)` · `removePersonal(name)` | set · remove a per-user override | | `setGitCredential(input)` | set a git auth credential | `runtime` with consumer `sandbox` is **environment** exposure — the default, and the only policy that puts a plaintext value in the session. `egress` with consumer `network` is **egress-enforced** exposure: the session holds a handle and TaskStation substitutes the real value outside the sandbox, for the exact HTTPS hosts the policy lists. Egress-enforced exposure is experimental; it needs the `secrets_egress` feature flag (Settings → Feature flags). With the flag off, `setStrategy(identifier, 'egress', …)` and `upsert(...)` with an egress policy return `403` `feature_disabled`. Every `broker` consumer has no session presence at all. See [Secrets](/docs/project/secrets). The `broker(...)` method requires a session-scoped token and an active session handle. #### `p.access` — project assignments, invites, requests Every method here reads or writes an assignment scoped to this project. Project roles are `manager` and `member`. | method | what | | --- | --- | | `list()` · `invite(email, role)` | principals with access · invite a user | | `update(userId, role)` · `revoke(userId)` | assign a project role · revoke the assignment | | `pendingInvites()` · `requests()` | outstanding invites · pending access requests | | `resendInvite(inviteId)` · `revokeInvite(inviteId)` · `approveRequest(id)` · `rejectRequest(id)` | resend/revoke an invite · approve/reject a request | | `groupGrants()` · `attachGroupGrant(...)` · `updateGroupGrant(...)` · `detachGroupGrant(...)` | the same assignments, with a `group` principal | `p.access.resourceGrants` is the **object assignment** view: it narrows a principal to one object in the project instead of the whole project. TaskStation enforces object assignments on agents and skills today. An agent is closed by default — a member reaches it only when an assignment names them or one of their groups. An object assignment carries no permissions of its own, and it restricts a project manager as much as a member. | method | what | | --- | --- | | `resourceGrants.list()` | object assignments in this project | | `resourceGrants.create(input)` | assign one object to a user or a group | | `resourceGrants.remove(grantId)` | revoke one object assignment | #### `p.connectors` — tool connectors | method | what | | --- | --- | | `catalog()` · `tools()` | callable Connector catalog · flattened `.` tools | | `search(query, options?)` · `describe(tool)` | find · inspect one callable tool | | `call(tool, args?)` | call one `.` tool through the server-side gateway | | `uploadAttachment(content, input)` | upload bytes and receive an opaque attachment handle for a later call | | `list()` · `config(connectorId)` | configured connectors · one connector's config | | `create(input)` | add a connector | | `auth.discover(input)` | preview auth from an OpenAPI spec, Postman collection, or endpoint | | `remove(connectorId)` · `sync()` | delete a connector · re-sync connectors | | `setName(connectorId, name)` · `setSensitive(connectorId, sensitive)` | rename · mark it sensitive (extra approval gating) | | `setAuthorizationStrategy(slug, strategy)` | select `project` or `user` connection ownership | | `setCredentialMode(connectorId, mode)` · `setCredential(connectorId, input)` | switch source · set the credential value | | `policies.get(connectorId)` · `policies.set(connectorId, policies)` | read · replace its tool policies | | `connections.list()` · `connections.reconcile(input)` | list · create/update connected accounts | | `connections.updateCredential(connectionId, input)` | rotate a connection credential | | `connections.revoke(connectionId)` · `connections.activate(connectionId)` | deny · restore a connection | `p.connectors.discover` browses the direct-connector catalog. It is **experimental** and off by default — enable it per project under [Settings → Experimental](/docs/feature-flags) → "Connectors API Discover". Easy Connect (Pipedream) remains the default connector marketplace. | method | what | | --- | --- | | `discover.list(query?, cursor?)` · `discover.detail(id)` | search OpenAPI/MCP/GraphQL/CLI entries · one entry's detail | `p.connectors.pipedream` is the optional managed-OAuth path; `listApps` returns OAuth apps only. Connect API-key apps directly instead. | method | what | | --- | --- | | `pipedream.listApps(params?)` · `pipedream.connect(input)` · `pipedream.finalize(input)` | browse the app catalog · start a connect flow · finalize it | A connector defines the tool, provider app, authorization strategy, and policies. `connections` stores its connected accounts. A session can select one with `connector_bindings: { alias: { connection_id } }`. Credentials stay encrypted and resolve per request. `connections` is the only active authorization facade. The retired `authorizations` and `profiles` names are not part of the current SDK surface. #### `p.policies` — project policies | method | what | | --- | --- | | `list()` · `set(policies)` | the project's policies · replace the set | #### `p.triggers` — cron and webhook automations A trigger starts an agent action on a schedule or an inbound webhook. See [Triggers](/docs/connect/triggers) for session strategy and payload templating. | method | what | | --- | --- | | `list()` | all triggers | | `create(input)` | create one | | `update(triggerId, input)` | edit a trigger | | `remove(triggerId)` | delete a trigger | | `fire(triggerId)` | run it now | | `setActivation(paused)` | pause or resume every trigger on the project | `create(input)` takes `{ name, type: 'cron' | 'webhook' | 'monitor', prompt_template, slug?, agent?, model?, enabled?, session_mode?, session_id?, cron?, run_at?, timezone?, secret_env?, session_access? }`. `name` and `prompt_template` are required. `cron`/`run_at` are mutually exclusive (`type: 'cron'`); `secret_env` (the webhook HMAC secret) applies to `type: 'webhook'`. `session_access` controls who can open sessions the trigger creates. It is `{ mode: 'private' | 'members' | 'project', memberIds: string[], groupIds: string[] }` and defaults to `private`. This policy is account-local runtime state. It does not enter the portable `taskstation.yaml` manifest. Updating only `session_access` creates no Git commit. A pinned session keeps its own sharing settings. A project manager can always open trigger-created sessions, including sessions that use `private` or selected-member access. `session_access` is a per-resource visibility setting on top of the role model, not a role. It decides who can open one trigger's sessions. It grants no permission the role verdict denies. See [Accounts & access](/docs/accounts#per-feature-access-settings). #### `p.marketplace` / `p.registry` — installed items Installs a catalog item's files onto the project's default branch. `registry.*` is an identical alias of `marketplace.*`. | method | what | | --- | --- | | `marketplace.list()` · `marketplace.install(id)` | installed items · install a catalog item | | `marketplace.updates()` · `marketplace.update(name)` · `marketplace.updateAll()` | available updates · update one · update all | | `marketplace.remove(name)` | uninstall an item | #### `p.files` — repo files (read) Read-only access to the project's git tree. To read and write files inside a running session, use the session's file operations — see [Files](#files) under Modules below. | method | what | | --- | --- | | `list(options?)` · `read(path, ref?)` | the repo tree · a file's contents at a git ref | | `search(query)` | search the repo | | `archive(options?)` · `history(path)` | download a tarball · a file's git history | #### `p.git` — history | method | what | | --- | --- | | `commits()` | the commit log | | `commit(sha)` · `commitDiff(sha)` | one commit · its diff | | `branches()` · `versionDiff(from, to)` | branches · diff between two refs | #### `p.changeRequests` — lifecycle and merge A change request (CR) is how a session's work merges into the default branch. | method | what | | --- | --- | | `list()` · `get(crId)` | open CRs · one CR | | `diff(crId)` · `mergePreview(crId)` | its diff · preview the merge result | | `open(input)` · `merge(crId, input?)` | open · merge a CR | | `close(crId, input?)` · `reopen(crId, input?)` | close without merging · reopen a closed one | | `requestChanges(crId, input)` | record feedback, optionally delivered back to the originating session | #### `p.sessions` — and the session handle | method | what | | --- | --- | | `list()` | the project's sessions | | `create(input?)` | create a session | | `session(sid)` | the session handle (same as `taskstation.session(id, sid)`) | The session handle is the heart of the runtime — see [Sessions](/docs/sdk/sessions). `create(input)` accepts `connector_bindings` keyed by connector-connection slug. Each value names a `connection_id`. It also accepts `secrets` for backend-origin secret narrowing and `require_connectors` for mandatory connectors. #### `p.approvals` — the connector approval inbox Pending connector-gated actions awaiting a decision — backs the permission-approval UX (`APPROVE` / `ASK` / `BLOCK`). | method | what | | --- | --- | | `list(options?)` · `sessionsNeedingInput(options?)` | pending approvals · sessions blocked on a decision | | `resolve(executionId, decision, scope?)` | approve or deny one — `decision: 'approve' \| 'deny'`, `scope: 'once' \| 'session' \| 'session_all'` | #### `p.gateway` — LLM observability Request logs, cost/latency rollups, budgets, and gateway API keys for this project's model traffic. | method | what | | --- | --- | | `logs(opts?)` · `log(logId)` | request log entries · one log entry | | `overview(days?)` · `series(days?)` · `breakdown(days?)` · `sessions(days?)` · `errors(days?)` | rollups, per-session cost, and errors over a window | | `budgets()` · `setBudget(input)` · `deleteBudget(budgetId)` | read · create/edit · remove a budget | | `keys()` · `createKey(name)` · `revokeKey(keyId)` | list · mint · revoke a gateway API key | | `playground(prompt, models)` | run one prompt against up to 6 models | | `routing.get()` · `routing.set(policy)` · `routing.reset()` | read · replace · inherit the routing policy | | `routing.preview(input)` | resolve a route without invoking a model | A routing policy holds a default model, a vision model, and an ordered fallback chain, each model attempted at most once; `fallbackOn` is `transient` or `any-error`. #### `p.channels` — Slack / email / voice Connector surfaces that let an agent act as a Slack app, an email address, or join a realtime voice call. | method | what | | --- | --- | | `slack.installation()` · `slack.mode()` · `slack.manifest()` | current install · mode · app manifest | | `slack.connect(input)` · `slack.disconnect()` | connect · disconnect | | `slack.getFile(url)` · `slack.uploadFile(input)` | download · upload a file via the server proxy | | `email.installation(connectorSlug?)` · `email.mode()` | current install · mode | | `email.connect(input)` · `email.disconnect(...)` · `email.updatePolicy(input)` | connect · disconnect · update the send/reply policy | | `voice.setBotName(name)` | rename the bot in a live call | #### `p.modelDefaults` — default model preferences Account, agent, and project-scoped model defaults, resolved by the gateway. | method | what | | --- | --- | | `get()` · `set(input)` · `clear(params)` | read · set a default · clear an override | #### `p.setDefaultAgent` — project default agent `p.setDefaultAgent(agentName)` checks that the agent is declared and enabled, then sets it as `default_agent` in the project's `taskstation.yaml`. New sessions prefer this agent unless a user picks another one. #### `p.updateFeatureFlag` — feature flags `p.updateFeatureFlag(feature, enabled)` turns one feature flag on or off for the project. Pass `enabled: null` to clear the override and fall back to the platform default. It calls the canonical `PATCH /v1/projects/:id/features`. `feature` is one of `FEATURE_FLAG_KEYS` (exported from `@taskstation/sdk`, typed as `FeatureFlagKey`). The caller needs the project's `project.customize.write` permission; the route answers `403` otherwise. Every flag-gated route rejects the same way while the flag is off: HTTP `403` with `{ error, code: "feature_disabled", feature }`. Use `isFeatureDisabledError(error)` to branch on it and `featureDisabledKey(error)` to read the flag key — never match on the message text. `p.updateExperimentalFeature(feature, enabled)` is the **deprecated** alias. It keeps calling the deprecated route alias `PATCH /v1/projects/:id/experimental` so consumers pinned to an older deployed API keep working. Use `p.updateFeatureFlag` in new code. #### `p.sandbox` — templates and snapshot builds Sandbox build config beyond `sandboxHealth`/`sandboxTemplates` on the project handle: Dockerfile/image/warm-pool templates and their snapshot builds. | method | what | | --- | --- | | `list()` · `snapshots()` | sandboxes for this project · built snapshots | | `rebuildSnapshot(slug?)` · `fixWithAgent()` | rebuild a snapshot · ask an agent to fix a broken build | | `createTemplate(input)` · `updateTemplate(...)` · `removeTemplate(...)` · `buildTemplate(...)` | add · edit · delete · build a template | | `setProvider(provider)` | request a provider switch. `null` (or the platform default / the already-active provider) applies immediately; switching to a different enabled provider starts a durable prepare→verify→activate transition — the current provider keeps serving while the target warm image is built and verified, then activated. The return is a tagged union: `kind:'project'` (immediate) or `kind:'preparation'` (poll `getProjectSandboxProviderTransition()` until `activated`/`failed`) | ### Escape hatch `taskstation.runtime()` returns the typed OpenCode REST client for the active sandbox. On a client created by `createScopedTaskStation` (`@taskstation/sdk/server`), it **throws** — the process-global "active" runtime is another request's sandbox in a multi-tenant server, a cross-tenant leak. Use the session-scoped `taskstation.session(pid, sid).runtime` (call `ensureReady()` first) instead, which resolves that session's own sandbox. ## Modules The framework-free modules behind [the client](#the-client) facade and the React hooks. Reach for them when you need one operation without the facade, a pure helper, or a Node-only isolation layer. Each module carries a stability tier so you know what to build on. | Tier | Meaning | | --- | --- | | Canonical | Import from the root `@taskstation/sdk`. Use this for all new code. | | Supported | A dedicated subpath (`@taskstation/sdk/react`, `@taskstation/sdk/server`). First-class, not deprecated. | | Deprecated alias | An old subpath that still works. It re-exports code the root already exports. Import from root instead. | | Internal | Outside semver. Do not import this in host code. | The root entry is canonical. Every framework-free name below is importable straight from `@taskstation/sdk`: ```ts import { files, getSessionHealth, getClient, authenticatedFetch, backendApi } from '@taskstation/sdk'; ``` ### Canonical modules | Module | What it does | | --- | --- | | Files | Workspace file operations: list, read, search, write | | Session runtime | Health probe and preview/proxy URL builders | | OpenCode client | The typed OpenCode REST client and its full type surface | | Auth | `authenticatedFetch` and token accessors | | Projects REST | The raw REST functions the facade wraps | | API client | `backendApi`, the low-level typed HTTP client | | Turns | Message-to-turn grouping, cost, and status math — see [Turns](#turns) | | Transcripts | `formatTranscript`, a client-side Markdown export | #### Files ```ts import { files } from '@taskstation/sdk'; const tree = await files.list('/workspace/src'); const { content } = await files.read('/workspace/README.md'); const hits = await files.findText('TODO'); await files.upload(file, '/workspace/uploads'); ``` `files` targets the globally active sandbox. If your host runs more than one session at a time, call `s.files` on the session handle instead. It always targets that session's own sandbox. See [Sessions](/docs/sdk/sessions). #### Session runtime helpers ```ts import { getSessionHealth, isRuntimeReady } from '@taskstation/sdk'; const result = await getSessionHealth(); if (result.ok && isRuntimeReady(result.health)) { // the sandbox daemon is ready } ``` `getSessionHealth` never throws on a non-2xx status. It returns `{ status, ok, health, body }` and lets you decide what a status means. The same module exports the URL helpers that rewrite an agent's `localhost` output into a reachable proxy URL: `detectLocalhostUrls`, `rewriteLocalhostUrl`, `proxyLocalhostUrl`, `parseLocalhostUrl`, and `buildWebProxyUrl`. #### OpenCode client ```ts import { getClient } from '@taskstation/sdk'; const client = getClient(); const { data } = await client.session.list({ limit: 100 }); ``` `getClient()` returns the typed OpenCode v2 compatibility client for the active sandbox, with auth already injected. Prefer `taskstation.session(pid, sid).runtime`, the same client scoped to one session, over the global `getClient()` when your host runs more than one session. #### Auth helpers ```ts import { authenticatedFetch, getAuthToken } from '@taskstation/sdk'; const res = await authenticatedFetch(`${runtimeUrl}/taskstation/health`); const token = await getAuthToken(); ``` The token comes from the `getToken` function you passed to `createTaskStation`. Most app code does not need this module — the file, session, and facade layers already authenticate for you. #### API client ```ts import { backendApi } from '@taskstation/sdk'; const data = await backendApi.get('/some/endpoint'); await backendApi.post('/some/endpoint', { name: 'x' }); ``` `backendApi` is the typed HTTP client every REST function builds on. Use it only for an endpoint that has no typed wrapper yet. ### Supported subpaths | Subpath | What it does | | --- | --- | | `@taskstation/sdk/react` | React hooks — see [React hooks](/docs/sdk/react) | | `@taskstation/sdk/server` | Request-scoped config for multi-tenant backends | #### Server-side isolation `createTaskStation` stores its config, including the token function, in one process-wide variable. That is fine for a browser tab, a CLI, or a single-tenant server. It is unsafe for a Node server that handles concurrent requests for different users, because the last `createTaskStation` call wins for every in-flight request. `@taskstation/sdk/server` fixes this with per-request isolation: ```ts import { createScopedTaskStation } from '@taskstation/sdk/server'; export async function handler(req: Request) { const taskstation = createScopedTaskStation({ backendUrl, getToken: () => tokenFor(req) }); return taskstation.projects.list(); } ``` `createScopedTaskStation` and `runWithTaskStation` isolate config per request with Node's `AsyncLocalStorage`. Never import `@taskstation/sdk/server` from a browser bundle — it statically pulls in `node:async_hooks`. A scoped client's top-level `runtime()` throws (it would resolve another tenant's sandbox). Reach a specific session's runtime via `taskstation.session(pid, sid).runtime` after `await s.ensureReady()`. ### Deprecated aliases About twenty old subpaths still work: `/files`, `/turns`, `/session`, `/auth`, `/projects-client`, `/api-client`, `/config`, `/event-stream`, `/opencode-client`, `/platform-client`, and more. Each one re-exports code the root `@taskstation/sdk` entry already exports. They stay working so no existing ```ts import { classifyPart, classifyTurn, toolInfo, toolViewModel } from '@taskstation/sdk'; ``` ```ts import { classifyPart, type ClassifiedPart } from '@taskstation/sdk'; for (const part of message.parts) { const classified: ClassifiedPart = classifyPart(part); switch (classified.kind) { case 'text': render(classified.text); break; case 'tool': render(classified.tool.title, classified.tool.status); break; } } ``` ```ts interface ToolView { name: string; title: string; status: 'pending' | 'running' | 'done' | 'error'; input?: Record; output?: string; error?: string; outputParsed?: unknown; // JSON.parse(output) when it parses, capped at 256KB outputText?: string; // the raw output text, always present } ``` ```ts toolInfo('bash'); // { label: 'Shell', category: 'shell' } getToolInfo('write', { filePath: '/workspace/main.go' }); // { icon: 'file-pen', title: 'Write', subtitle: 'main.go /workspace' } ``` ```ts const vm = toolViewModel(classifiedTool); if (vm.kind === 'shell') { render(vm.command, vm.stdout, vm.exitCode); } ``` ```ts import { groupMessagesIntoTurns, collectTurnParts, type TurnLike } from '@taskstation/sdk'; const turns: TurnLike[] = groupMessagesIntoTurns(messages); for (const turn of turns) { const parts = collectTurnParts(turn); } ``` ```ts import { isTextPart, isToolPart, getPartText } from '@taskstation/sdk'; if (isTextPart(part)) { // part.type narrowed to 'text' } const text = getPartText(part); // works for 'text' and 'reasoning' parts ``` ```ts import { getWorkingState, getTurnStatus, formatDuration } from '@taskstation/sdk'; const status = getTurnStatus(parts, childMessages); // "Running commands..." formatDuration(4300); // "4s" — durations under 1s return '' ``` ```ts import { getTurnError, getChildSessionError, unwrapError } from '@taskstation/sdk'; getTurnError(turn); // the first assistant error, unwrapped getChildSessionError(childMessages); // newest error in a sub-agent's messages unwrapError(rawError); // normalizes double-JSON and mixed error shapes ``` ```ts import { getTurnCost, getSessionCost, formatCost, formatTokens, COST_MARKUP } from '@taskstation/sdk'; const info = getTurnCost(partsWithMessage, modelPricingLookup); const sessionCost = getSessionCost(messages, modelPricingLookup); formatCost(0.0032); // "$0.003" formatTokens(12345); // "12k" ``` ```ts import { getChildSessionId, getChildSessionToolParts } from '@taskstation/sdk'; const childId = getChildSessionId(taskToolPart); const steps = getChildSessionToolParts(childMessages); ``` ```ts import { getPermissionForTool, getHiddenToolParts, isToolPartHidden } from '@taskstation/sdk'; const permission = getPermissionForTool(permissions, callID); const hidden = getHiddenToolParts(activePermission, activeQuestion); ``` ```ts import { getFilename, getDirectory, relativizePath, stripAnsi } from '@taskstation/sdk'; getFilename('/workspace/src/main.go'); // "main.go" getDirectory('/workspace/src/main.go'); // "/workspace/src" relativizePath('/workspace/src/main.go', '/workspace'); // "src/main.go" ``` ```ts import { sortSessions, childMapByParent, allDescendantIds } from '@taskstation/sdk'; sessions.sort(sortSessions(Date.now())); // pins sessions updated in the last 60s const childMap = childMapByParent(sessions); const descendants = allDescendantIds(childMap, sessionId); ``` ```sh npm install @taskstation/sdk ``` ```ts import { createTaskStation } from '@taskstation/sdk'; // framework-free core import { useSession } from '@taskstation/sdk/react'; // optional React layer import { createScopedTaskStation } from '@taskstation/sdk/server'; // Node and Bun servers ``` ```html ``` --- # Sessions Run a session, stream its events, and handle the errors it can throw. Canonical page: https://taskstation.co/docs/sdk/sessions A session is one agent run, in its own sandbox, on its own git branch. `taskstation.session(projectId, sessionId)` returns the handle for everything a session does: start it, send prompts, stream events, and read status. This page covers the handle, the readiness handshake, streaming, and the typed errors an SDK call can throw. ```ts const s = taskstation.session(projectId, sessionId); ``` `s` is the handle for everything a session does. The session ID, the sandbox ID, and the branch name are the same value. See [Sessions](/docs/work/sessions) for the concept. ## Session lifecycle | Method | Wraps | What it does | |---|---|---| | `s.get(opts?)` | `GET /projects/:pid/sessions/:sid` | Reads session details | | `s.update(input)` | `PATCH …/sessions/:sid` | Renames the session or updates metadata | | `s.start(waitMs?)` | `POST …/sessions/:sid/start` | Provisions and boots the runtime | | `s.restart()` | `POST …/sessions/:sid/restart` | Restarts the runtime; keeps the same sandbox | | `s.reloadConfig(input?)` | `POST …/sessions/:sid/reload` | Recompiles agent config and replaces the runtime after validation | | `s.reloadConfigStream(input, onEvent)` | `POST …/sessions/:sid/reload-stream` | Runs the same reload and emits server-confirmed progress phases | | `s.stop()` | `POST …/sessions/:sid/stop` | Stops the runtime; the session stays | | `s.delete()` | `DELETE …/sessions/:sid` | Deletes the session | | `s.setSharing(intent)` | `PUT …/sharing` | Sets sharing and visibility | | `s.cost()` | `GET /usage/session-costs/:sid` | Reads finalized LLM and compute cost without starting the runtime | | `s.scope()` | `GET …/sessions/:sid/scope` | Reads stored secret narrowing and materialized connection bindings | | `s.rescope(input)` | `PUT …/sessions/:sid/scope` | Replaces supplied scope fields for the next prompt or tool call | | `s.commit(input?)` | — | Commits the agent's work | > **Warn** > `s.delete()` deletes the session and its runtime. This cannot be undone. To pause a session > without losing it, call `s.stop()` instead. Use the streamed method when the caller displays reload progress: ```ts await s.reloadConfigStream({ refresh_repo: false }, (event) => { if (event.type === 'phase') console.log(event.phase); }); ``` The phases are `checking-session`, `refreshing-workspace`, `compiling-config`, `applying-config`, and `confirming-config`. The server omits `refreshing-workspace` when `refresh_repo` is `false`. The `applying-config` phase includes the daemon's validated runtime replacement. Three more read methods round out the handle: - `s.previews()` — candidate preview ports the runtime exposes. - `s.publicShares.list()` / `.create(input)` / `.revoke(shareId)` — public share links. - `s.audit(limit?)` — the session's audit trail of agent actions. - `s.transcript(options?)` — a compact server-side transcript (text and tool calls, no tool inputs or outputs). This works with a project-scoped session token. - `s.voiceTranscript(options?)` — this session's live voice-call transcript (spoken turns plus `ask_taskstation`/`run_command` worker tool calls). Returns an empty list when the session has no live call, not a 404. ## Readiness is a handshake Before you send a prompt, call `ensureReady()`. It provisions the sandbox if needed, waits for the runtime to boot, and returns the resolved runtime. ```ts const { opencodeSessionId, runtimeUrl, sandboxId } = await s.ensureReady(); ``` On a cold boot, `ensureReady()` can throw `RUNTIME_UNAVAILABLE`. See [Retry on a cold boot](#retry-on-a-cold-boot) for what that means and how to retry. `s.send()` and `s.abort()` call `ensureReady()` for you. ### Seed a server-authorized OpenCode pin A server-rendered React host can supply the OpenCode pin already persisted for the same TaskStation session: ```tsx const session = useSession(projectId, sessionId, { initialOpenCodeSessionId: persistedSession.opencode_session_id, }); ``` The seed only hydrates cached transcript content while `/start` runs. It does not override the runtime identity. The pin returned by `/start` is authoritative and replaces a stale seed. Do not accept this value from an untrusted tenant selector. Do not create an OpenCode session in the host. TaskStation creates and persists the root session. OpenCode query caches and transcript controllers are scoped to the sandbox runtime, so equal OpenCode ids from different sandboxes do not share cache entries. ## Send a prompt ```ts s.setModel({ providerID, modelID }); // sticky for later send() calls s.setAgent('build'); // sticky for later send() calls await s.send('Refactor the auth module'); await s.send('One-off task', { model, agent }); // overrides for this call only await s.abort(); // stop the current run ``` For OpenCode REST sessions, the first `send()` on a handle reads the model and agent persisted on the TaskStation session. This prevents a snapshot-inherited OpenCode session from reusing stale snapshot defaults. Prompt choice precedence is: 1. The `send()` call. 2. The handle's `setModel()` or `setAgent()` value. 3. The persisted TaskStation session default. `setModel` only chooses what the next local `send` asks for — it never leaves the handle. To **persist** a new model for a running session server-side, use `changeModel`: ```ts const { applied_live } = await s.changeModel('anthropic/claude-opus-4-8'); ``` Restarting the runtime is how the change takes effect, so an in-flight turn ends. `applied_live` is `true` when a running session took it now, `false` when it applies at the next start. Only the session owner, or a caller with project-manager permissions, may change the model; anyone else gets `403`. `send()` resolves the runtime, then prompts it. `abort()` stops the current run without deleting the session. ## Session scope and cost Read the stored secret narrowing and materialized connection bindings. `secrets_allowlist: null` means the agent's secret grant applies: ```ts const scope = await s.scope(); scope.connector_bindings_configured; // false = inherits the project defaults ``` `connector_bindings` is the RESOLVED map, so it looks the same for a session that overrode its connectors and one that inherits the project defaults. Read `connector_bindings_configured` to tell them apart before rendering the scope or sending it back. Replace one or both scope fields: ```ts await s.rescope({ secrets: ['DATABASE_URL'], connector_bindings: { github: { connection_id: connectionId }, }, }); ``` Each supplied field replaces its complete previous value. Omit a field to leave it unchanged. Connection changes apply to the next tool call. Secret removal stops future delivery but cannot remove an already disclosed value from model context or an existing process. Both axes have an explicit way back to the default. They are not the same as an empty value: ```ts await s.rescope({ secrets: null, // inherit the agent's secret grant connector_bindings: null, // drop the override; inherit the project defaults }); ``` `secrets: []` and `connector_bindings: {}` are the opposite instruction: an explicit "no project secrets" and "no connectors at all", project defaults included. A session that sends `{}` where it meant `null` fails closed on every alias it did not name. Read the unified cost record: ```ts const cost = await s.cost(); ``` The record combines finalized LLM cost, billed sandbox compute cost, model usage, token totals, compute duration, and ledger entries. `s.cost()` does not call `ensureReady()`. ## Runtime status and previews | Method | Returns | Use | | --------------------------- | ------------------------------ | -------------------------------------------- | | `s.health(init?)` | `{ status, ok, health, body }` | Check whether the runtime is alive | | `s.previewUrl(port, path?)` | `string` | Get a proxy URL for a port the agent exposed | | `s.proxyUrl(url?)` | `string \| undefined` | Rewrite a localhost URL the agent printed | ```ts const { ok, health } = await s.health(); const url = s.previewUrl(3000, '/docs'); ``` `s.health()` never throws. Call it any time, even before the session has a runtime. `s.previewUrl()` and `s.proxyUrl()` need a resolved runtime — call `s.ensureReady()` first, or they throw `SessionNotReadyError`. See [Session readiness errors](#session-readiness-errors). ## Streaming Use `s.stream()` to receive live events in a script or server. In a React app, use [`useSession`](/docs/sdk/react) instead — it manages the whole session lifecycle for you. `s.stream()` is the OpenCode REST compatibility event stream. The TaskStation API proxies it from the sandbox. There is no separate WebSocket endpoint. The transport is `fetch` with a streaming response body, read through `ReadableStream` and `TextDecoderStream`. The SDK handles reconnection, backoff, and a heartbeat check. Stream a session: 1. Call `ensureReady()` first. The runtime does not exist until the sandbox starts. 2. Open the stream before you send a message, so you do not miss early events. 3. Send the message. 4. Close the stream when you see `session.idle`. ```ts const session = taskstation.session(projectId, sessionId); const { opencodeSessionId } = await session.ensureReady(); const stream = await session.stream({ onEvent: (event) => { if (event.type === 'session.idle' && event.properties.sessionID === opencodeSessionId) { onTurnDone(); stream.close(); } }, }); await session.send('Refactor the auth module'); ``` Streaming needs `fetch` with a real `ReadableStream` body and `TextDecoderStream`. Browsers, Node 18 and later, Bun, and Cloudflare Workers all support it. React Native and Expo do not: their `fetch` has no `response.body`. On React Native, use `createHttpSessionSyncController` for bounded history and status synchronization. Use a platform-specific event transport for live events. The controller loads the newest 10 messages first. `loadOlder()` follows the server cursor. `loadHttpSessionHistory()` follows every cursor for explicit exports. ### Event types Each event has a `type` and a `properties` object that holds its data, for example `event.properties.sessionID`. | `type` | When it fires | | ----------------------------------------------- | --------------------------------------------------- | | `message.updated` / `message.removed` | A message changed or was deleted. | | `message.part.updated` / `message.part.removed` | A part (text, tool call, file) grew or was removed. | | `session.status` | The session's busy state changed. | | `session.idle` | The turn finished. | | `session.error` | The turn failed. The event carries the error. | | `question.asked` | The agent asked for input. | | `question.replied` / `question.rejected` | The answer to a question arrived. | Turn raw messages and parts into renderable output with `classifyTurn`. See [SDK reference](/docs/sdk/reference). ## Retry on a cold boot `ensureReady()` polls the session's `/start` endpoint — each call long-polls up to 30 s — until the runtime reaches a terminal `ready`/`failed`/`stopped` stage or its deadline (`readyTimeoutMs`, default ~180 s) elapses. On a warm session the first poll resolves `ready` immediately. On a cold boot it keeps polling while the sandbox reports `retriable: true`, so a slow start just takes longer rather than throwing. It only throws an `ApiError` with `code: 'RUNTIME_UNAVAILABLE'` if the runtime is still not `ready` when the deadline expires. `ensureReady()` is idempotent, so concurrent calls for the same session share one `/start` request instead of sending several. The `retryUntilReady` helper below is now optional — `ensureReady()` already retries internally — but stays useful if you want a longer total budget than the default `readyTimeoutMs`. ```ts async function retryUntilReady(ensure: () => Promise): Promise { const deadline = Date.now() + 300_000; for (;;) { try { return await ensure(); } catch (error) { const provisioning = error instanceof ApiError && error.code === 'RUNTIME_UNAVAILABLE'; if (!provisioning || Date.now() > deadline) throw error; await new Promise((r) => setTimeout(r, 3_000)); } } } ``` See [Error classes](#error-classes) for the full `ApiError` shape. In React, [`useSession`](/docs/sdk/react) retries `/start` for you, so you do not need this pattern. ## What `/start` tells you Every `/start` answer describes **that call**, not the row's accumulated history. Four fields carry it. | Field | Meaning | | --- | --- | | `observed_at` | One clock for the whole answer. | | `action` | What the server did: `inspected`, `checked_provider`, `resumed`, `provisioned`, `restored`, `reconciled`, `awaited_wake`, `cooling_down`. | | `observation` | What the server checked. `known: false` means **not checked on this call** — never "checked and found nothing". | | `boot` | `phase` (`provisioning` / `resuming` / `booting` / `ready` / `parked` / `failed`), `since`, and `actively_starting`. | `boot.actively_starting` answers "is a provider operation running for this session right now?". A `starting` payload with `actively_starting: false` means the server is waiting out a retry cooldown, not that a box is booting. ```jsonc { "stage": "starting", "retriable": true, "reason": "runtime_wake_cooldown", "observed_at": "2026-08-26T14:00:00.000Z", "action": "cooling_down", "boot": { "phase": "resuming", "since": "2026-08-26T13:58:00.000Z", "actively_starting": false }, "observation": { "provider": { "known": false, "status": null, "checked_at": null }, "runtime": { "known": false, "state": null, "boot_phase": null, "checked_at": null } }, "failure": { "category": "sandbox-provider", "message": "The runtime did not start (attempt 2). Retrying automatically.", "retryable": true, "evidence": { "check": "provider_not_running", "observed_at": "2026-08-26T13:58:00.000Z", "error": null, "attempts": 2, "next_retry_at": "2026-08-26T14:03:00.000Z" } } } ``` ### A failed start is retried for you A start that fails stamps a **cooldown**, not a permanent verdict. The next `/start` after the cooldown re-attempts the wake by itself. The cooldown grows with consecutive failures (2 min, 5 min, 10 min). After five consecutive failures `/start` answers `stage: "failed"` with the attempt count in `failure.message`; that verdict expires 30 minutes after the last failure, and `POST …/restart` clears it immediately. `retriable` is derived on every call. A state the server can still re-attempt never carries `retriable: false`. `failure.evidence` names the check that produced the negative, when it ran, and when the server retries. Every `/start` failure carries it. ## Files `s.files` reads and writes the session's sandbox: `list`, `read`, `readBlob`, `status`, `findFiles`, `findText`, `upload`, `create`, `copy`, `remove`, `mkdir`, `rename`. Every call resolves the runtime first, and always targets this session's own sandbox. See the [SDK reference](/docs/sdk/reference) for the full method list. ## The raw runtime `s.runtime` is the typed OpenCode REST client. Use it only for calls that `send`, `abort`, and `stream` do not cover. It requires a resolved OpenCode runtime — call `s.ensureReady()` first. ```ts const { opencodeSessionId } = await s.ensureReady(); await s.runtime.session.prompt({ sessionID: opencodeSessionId, parts: [{ type: 'text', text: 'Refactor the auth module' }], }); ``` The OpenCode `sessionID` here is not the session ID you pass to `taskstation.session(projectId, sessionId)`. The SDK resolves it during `ensureReady()` and caches it on the handle. ## Warm a project session Call `ensureWarm()` when a project landing page needs one runtime ready before the first prompt. ```ts const project = taskstation.project(projectId); const warm = await project.sessions.ensureWarm(); // An ORDINARY session. Prompt it like any other. await taskstation.session(projectId, warm.session.session_id).send("Build me a widget"); ``` `ensureWarm()` creates, or returns, one unused session for the current user. It is the same create `sessions.create()` runs, with the project's defaults: same billing gate, same concurrent-session cap, same connector requirements. The only difference is `metadata.warm`, which hides the session from `sessions.list()` until its first prompt lands. Treat it as speculative. A `409 WARM_SESSION_UNAVAILABLE` means the account has no concurrent-session headroom to spare or the project cannot be warmed right now — fall through to `sessions.create()`, which reports the real reason. The warm session carries the project's DEFAULT agent and sandbox. If the user picks a different one, abandon it and call `sessions.create()`: an unused warm session is hidden and reaped on its own. > **Warn** > `claimWarm()` is deprecated. A warm session is an ordinary session, so there is > nothing to claim — navigate to it and prompt it. The call still works for > consumers pinned to the older shape and is removed in the next major. ## Handling errors Every call through `createTaskStation` rejects with a typed `Error` subclass, never a plain object. Catch the error, check `instanceof`, and branch on `.status` or `.code`. ```ts import { ApiError, BillingError } from '@taskstation/sdk'; try { await taskstation.project(projectId).sessions.create(); } catch (err) { if (err instanceof BillingError) { // 402 — out of credits or over a plan limit } else if (err instanceof ApiError) { // any other failed request — err.status, err.code, err.detail } else { throw err; } } ``` ### Error classes | Class | Extends | When it throws | Key fields | | ---------------------- | ---------- | ------------------------------------------------------------------------------ | -------------------------------------------------------------------- | | `ApiError` | `Error` | Default for any failed request: bad status, network failure, timeout, or abort | `status`, `code`, `detail`, `response`, `url`, `endpoint`, `timeout` | | `AuthError` | `ApiError` | `getToken` returned `null`. TaskStation never sent the request | `code` is always `'NO_SESSION'` | | `BillingError` | `Error` | HTTP `402`. The only billing error class | `status` (`402`), `detail.message` | | `RequestTooLargeError` | `Error` | HTTP `431`. Usually too many files in one request | `detail.suggestion` | | `SessionNotReadyError` | `Error` | A runtime accessor ran before `ensureReady()` | `name` is `'SessionNotReadyError'` | `ApiError.name` is `'ApiError'` by default. Two cases override it: - `name: 'AbortError'`, `code: 'ABORTED'` — the request was cancelled, for example by navigation. This is not a failure. Ignore it. - `code: 'TIMEOUT'` — the request's own timeout elapsed. `url`, `endpoint`, and `timeout` show what timed out. For any other failure, `status` holds the HTTP status code. `code` comes from the backend's `error_code`, or falls back to the status as a string. `message` is an enumerable own property on `ApiError`, so it survives `JSON.stringify` and object spread. TaskStation retries some requests before your code sees an error. If a `GET` or `HEAD` request returns `502`, `503`, or `504`, TaskStation retries it up to 2 times, with a 250ms then 500ms delay. A transient transport failure on a `GET` or `HEAD` — a network error, not a status code — is retried the same way. A retry that succeeds never reaches `onError`. TaskStation never retries `POST`, `PUT`, `PATCH`, or `DELETE` requests, or a `500` response. TaskStation throws `AuthError` on the client, before it sends a request, when `getToken()` returns `null`. `AuthError` extends `ApiError`, so `err instanceof ApiError` still matches. Check `err instanceof AuthError`, or `err.code === 'NO_SESSION'`, to treat "not signed in" as a separate case from a backend failure. TaskStation throws `BillingError` for every HTTP `402` response: out of credits, over a plan limit, or another billing gate. `detail.message` holds the reason from the backend. TaskStation throws `RequestTooLargeError` for HTTP `431`. This usually means the request carried too many files. `detail.suggestion` holds a ready-to-show hint for the user. ### Session readiness errors Two errors mean the session's sandbox is not ready yet. Handle each one differently. `SessionNotReadyError` throws synchronously when you call a runtime accessor — `session.previewUrl()`, `session.proxyUrl()`, or `session.runtime` — before this session handle has resolved its sandbox. A session handle only resolves its own sandbox; it never falls back to another session's sandbox. ```ts import { SessionNotReadyError } from '@taskstation/sdk'; const s = taskstation.session(projectId, sessionId); try { const url = s.previewUrl(3000); // throws: not resolved yet } catch (err) { if (err instanceof SessionNotReadyError) { await s.ensureReady(); } } ``` Call `await session.ensureReady()` first, or call `send()`, which readies the session internally. `session.health()` is the one accessor that never throws this error, so you can poll it before the session boots. `RUNTIME_UNAVAILABLE` is the second error — it means `ensureReady()` itself timed out waiting for a cold boot. See [Retry on a cold boot](#retry-on-a-cold-boot) for the full pattern. In React, `useSession` retries this for you and exposes it through the `phase` value instead of throwing. ### Helpers | Helper | Signature | What it does | | -------------------------------- | ----------------------------------- | ------------------------------------------------------------------------------------------------------------- | | `parseBillingError(error)` | `(error) => Error` | Wraps a `402` response into a `BillingError`. Returns other errors unchanged | | `isBillingError(error)` | `(error) => boolean` | Returns `error instanceof BillingError` | | `formatBillingErrorForUI(error)` | `(error) => BillingErrorUI \| null` | Returns `null` for non-billing errors. Otherwise returns `{ alertTitle, alertSubtitle }` for an upgrade modal | ```ts import { formatBillingErrorForUI } from '@taskstation/sdk'; try { await taskstation.session(projectId, sessionId).start(); } catch (err) { const ui = formatBillingErrorForUI(err); if (ui) showUpgradeModal(ui.alertTitle, ui.alertSubtitle); } ``` ### In `@taskstation/sdk/react` `@taskstation/sdk/react` re-exports `BillingError`, `RequestTooLargeError`, `parseBillingError`, `isBillingError`, and `formatBillingErrorForUI`. It does not re-export `ApiError` or `AuthError` — import those from `@taskstation/sdk`. `useSession` classifies every `send`, `answerQuestion`, `answerPermission`, and `rejectQuestion` failure into one `sendError` object, so you do not need to write `instanceof` checks by hand: ```ts interface TaskStationSendError { kind: 'billing' | 'runtime-not-ready' | 'runtime-error'; message: string; billing?: BillingError; // set when kind is 'billing' cause: unknown; } ``` ```tsx const s = useSession(projectId, sessionId); if (s.sendError?.kind === 'billing') { const ui = formatBillingErrorForUI(s.sendError.billing); } ``` See [React hooks](/docs/sdk/react) for the rest of `useSession`. --- # Sign in with TaskStation Gate your own app behind TaskStation identity with one route, and act as the signed-in user through the SDK. Canonical page: https://taskstation.co/docs/sdk/sign-in "Sign in with TaskStation" makes TaskStation the identity provider for an app you run: a dashboard, an internal tool, a vertical product built on TaskStation. Your users sign in with their TaskStation account, your server knows who they are, and every TaskStation call your app makes runs as that user with that user's role assignments. The whole flow lives in `@taskstation/sdk`. Your app never stores a TaskStation token in the browser and never talks to Supabase. It is standard OAuth 2.1 (authorization code + PKCE) served by the TaskStation API, so it works the same against `api.taskstation.co` and against a self-hosted instance. > **Note** > Building an App **hosted by TaskStation** (`*.apps.taskstation.co`)? You need none of > this. The Apps gate already authenticated the viewer — read them with > `taskstationAppViewerToken()` / `readAppViewer()`. See > [Apps → Your App already knows who is looking](/docs/sdk/apps). ## 1. Register your app Go to **Account → Tokens → OAuth apps → Register app**, or call the SDK: ```ts const app = await taskstation.iam.oauthClients.create(accountId, { name: 'Dashboards', client_type: 'confidential', // 'public' for a browser/native app (PKCE only, no secret) redirect_uris: ['https://dashboards.example.com/api/taskstation/auth/callback'], scopes: ['profile', 'email', 'taskstation'], }); // app.client_id, app.client_secret (shown once) ``` Registration needs `token.create` on the account. Redirect URIs are compared byte for byte; `https` is required except on `localhost`. | Scope | Grants the app | |---|---| | `profile` | The user's id, email and account memberships (`GET /v1/accounts/me`). | | `email` | The email address (an alias for OIDC-shaped clients). | | `taskstation` | Acting as the user on the whole TaskStation API — projects, sessions, files, IAM probes. Without it the token is identity-only. | ## 2. Mount the handler ```ts // lib/taskstation-auth.ts import { createTaskStationAuth } from '@taskstation/sdk/server'; export const auth = createTaskStationAuth({ backendUrl: 'https://api.taskstation.co/v1', clientId: process.env.TASKSTATION_OAUTH_CLIENT_ID!, clientSecret: process.env.TASKSTATION_OAUTH_CLIENT_SECRET, // omit for a public client redirectUri: 'https://dashboards.example.com/api/taskstation/auth/callback', cookieSecret: process.env.TASKSTATION_AUTH_COOKIE_SECRET!, // ≥ 32 chars; encrypts the session cookie }); ``` ```ts // app/api/taskstation/auth/[...taskstation]/route.ts (Next.js App Router) import { auth } from '@/lib/taskstation-auth'; const handle = (request: Request) => auth.handler(request); export { handle as GET, handle as POST }; ``` The handler serves every route under `basePath` (derived from the redirect URI — `/api/taskstation/auth` above): | Path | Does | |---|---| | `/signin?return_to=/path` | Starts sign-in (PKCE S256 + state in a 10-minute cookie) and redirects to TaskStation. | | `/callback` | Exchanges the code, sets the encrypted `HttpOnly` session cookie, redirects to `return_to`. | | `/refresh?return_to=` | Rotates the token pair and redirects. Used by `requireViewer`. | | `/signout?return_to=` | Revokes the refresh token at TaskStation and clears the cookie. | | `/me` | The viewer as JSON, or `401`. Refreshes inline when the access token expired. | | `/proxy/*` | Forwards to the TaskStation API as the viewer. The browser SDK's `backendUrl`. | `return_to` is always confined to a same-origin path. ## 3. Gate pages and act as the user ```ts // middleware.ts — every page needs a viewer import { auth } from '@/lib/taskstation-auth'; export async function middleware(request: Request) { const gate = await auth.requireViewer(request); if (gate.response) return gate.response; // 302 → /refresh or /signin } export const config = { matcher: ['/((?!api/taskstation/auth|_next).*)'] }; ``` ```ts // a server component / route handler const viewer = await auth.viewer(request); // { userId, email, accounts, scopes, token, expiresAt } | null const taskstation = await auth.taskstation(request); // request-scoped client acting as the viewer const projects = await taskstation.projects.list(); const allowed = await taskstation.iam.can(accountId, viewer!.userId, { action: 'project.write', resourceType: 'project', resourceId }); ``` `viewer()` is read-only and never consumes the single-use refresh token; use `requireViewer()` in middleware so a page never renders signed-out for a user whose refresh token is still good. ## 4. The browser ```tsx import { createTaskStation } from '@taskstation/sdk'; import { SignInWithTaskStation, useTaskStationViewer } from '@taskstation/sdk/react'; const taskstation = createTaskStation(auth.clientConfig()); // backendUrl = '/api/taskstation/auth/proxy' function Header() { const { status, viewer } = useTaskStationViewer(); if (status === 'signed-in') return {viewer.email}; return ; } ``` The browser client sends a sentinel bearer; `/proxy` swaps it for the viewer's real token on the server. `useSession`, `taskstation.project(id).sessions.*` and every other SDK call work unchanged through it. ## What the user sees The first time, TaskStation shows a consent screen naming your app and the scopes. TaskStation remembers the decision per user and app, so later sign-ins redirect straight back. Revoking an app deletes every token it minted. ## Discovery `GET https://api.taskstation.co/.well-known/oauth-authorization-server` (also under `/v1/oauth/.well-known/…`) publishes the endpoints for a generic OAuth client. The SDK does not need it — it derives every endpoint from `backendUrl`. --- # Change requests How session work reaches the default branch through review. Canonical page: https://taskstation.co/docs/work/change-requests A change request (CR) merges one git branch into another. TaskStation creates a CR from a session's branch (`head_ref`) onto the project's default branch (`base_ref`, usually `main`). The CR row is metadata; the merge, diff, and conflict checks run as real git operations against the project's repository. A CR is the only way session work reaches the default branch. ## Why work goes through a CR A session runs in a sandbox on its own branch, named after the session ID. The sandbox does not last forever, but the branch does: git is the only durable record of a session's work. Every new session starts from the default branch. Until a CR merges, the work stays on its own branch, unreviewed and invisible to every other agent, trigger, and collaborator. This applies to every change: code, agents, skills, and the manifest (`taskstation.yaml`) — no exceptions. ## The agent mandate An agent must open a CR to land any change on the project's default branch. The agent does not merge its own CR — merging is the user's decision. Follow this contract: 1. Commit on the session branch (`$TASKSTATION_BRANCH_NAME`). Make small, working commits. Do not rewrite history or force-push. 2. Push the branch: `git push origin HEAD`. 3. Open the CR: `taskstation cr open --title "..." --description "..."`. Inside a session sandbox, `--head` and `--session` default to `$TASKSTATION_BRANCH_NAME` and `$TASKSTATION_SESSION_ID`. `--base` defaults to the project's default branch. 4. Tell the user the CR number, so they can review it. 5. Stop. Do not merge the CR yourself. ### Anti-patterns - **Force-pushing to the default branch.** This breaks the review contract, even where the backend allows it. - **"It's on my branch, pull it yourself."** The session branch is gone once the sandbox stops, unless a CR merged it first. - **Sending the change as a file, paste, or archive.** The CR system already solves this problem. ## Data model CRs live in the `change_requests` table. | Column | Type | Notes | | --- | --- | --- | | `cr_id` | uuid | Primary key. The REST API's identifier. | | `project_id` | uuid | The project the CR belongs to. | | `number` | integer | Per-project display number (`#1`, `#2`…). Unique per project. Never recycles. | | `title` | text | Required. | | `description` | text | Defaults to an empty string. | | `base_ref` | text | The branch merged into. Usually `main`. | | `head_ref` | text | The branch merged from. In a session, this is the session ID (a UUID). | | `status` | enum | `open`, `merged`, or `closed`. | | `head_commit_sha` | text | Refreshed against the live `head_ref` tip on every read, for open CRs. Captured at merge time for merged CRs. | | `base_commit_sha` | text | Same rule, for `base_ref`. | | `origin_session_id` | text | The session that opened the CR. Set to null if that session is deleted. | | `created_by` | uuid | The user who created the CR. | | `merged_at` / `merged_by` | timestamp / uuid | When and who merged the CR. | | `merge_commit_sha` | text | The merge commit. Equals `head_commit_sha` for a fast-forward. | | `closed_at` / `closed_by` | timestamp / uuid | When and who closed the CR without merging. | | `metadata` | jsonb | Holds `requested_changes`, a list of `{text, by, at}` entries added by `POST /:crId/request-changes`. CRs have no separate comment table. | | `created_at` / `updated_at` | timestamp | Set on creation. Updated on every status change or SHA refresh. | A unique index on `(project_id, number)` lets you reference a CR by its short number instead of its UUID. ## Lifecycle ``` open ──(merge)──▶ merged (terminal) open ──(close)──▶ closed ──(reopen)──▶ open ``` - `open` is the starting status. - `closed` is reversible. `POST /:crId/reopen` sets it back to `open`. - `merged` is terminal. You cannot reopen or close a merged CR. Open a new CR against the merged state instead. TaskStation refuses to create a CR whose branch has no commits ahead of the default branch. This usually means the agent committed locally but never pushed. Push the commits, then create the CR again. ## SHA refresh For an open CR, TaskStation refreshes `head_commit_sha` and `base_commit_sha` against the live branches on every `GET`. If the repository is unreachable, or a branch is missing, TaskStation skips the refresh and serves the CR's last known metadata. A merged CR keeps the SHAs captured at merge time — TaskStation never refreshes them again. ## Merge mechanics `POST /v1/projects/:projectId/change-requests/:crId/merge` runs these steps. 1. TaskStation reads the manifest (`taskstation.yaml`) from `head_ref` and validates it against the manifest schema. A branch with no manifest passes. An invalid manifest returns `422` with `code: "MANIFEST_INVALID"` and stops the merge. 2. TaskStation fast-forwards `base_ref` if `head_ref` is strictly ahead of it. 3. Otherwise, TaskStation creates a merge commit. The default message is `Merge CR #: `, and you can override it with `message` in the request body. The commit author is `TaskStation <noreply@taskstation.co>`. 4. A conflict returns `409` with the conflict list. Check the same list with `GET /:crId/merge-preview` before you merge. 5. On success, TaskStation sets `status` to `merged`, records `merged_at`, `merged_by`, and `merge_commit_sha`, and invalidates the project's git cache. Merging a CR that is not `open` returns `409`. ### Merge preview `GET /:crId/merge-preview` returns: | Field | Type | Meaning | | --- | --- | --- | | `base_sha` | string | Current tip of `base_ref`. | | `head_sha` | string | Current tip of `head_ref`. | | `merge_base` | string \| null | Common ancestor. Null if the histories are unrelated. | | `is_up_to_date` | boolean | `head_ref` is fully merged into `base_ref`. | | `can_merge` | boolean | No conflicts. | | `can_fast_forward` | boolean | `head_ref` is strictly ahead of `base_ref`. | | `conflicts` | string[] | File paths that would conflict. | ## REST API All routes sit under `/v1/projects/:projectId/change-requests`. | Method | Path | Notes | | --- | --- | --- | | GET | `/` | `?status=open\|merged\|closed\|all`. No filter returns every status. | | POST | `/` | Body: `{title, description?, head_ref, base_ref?, session_id?}`. Returns `201`. | | GET | `/:crId` | Returns the CR. Refreshes SHAs as a side effect. | | PATCH | `/:crId` | Edits `title` or `description`. `409` if not `open`. | | GET | `/:crId/diff` | Unified patch: file list, additions, deletions. | | GET | `/:crId/merge-preview` | See Merge preview above. | | POST | `/:crId/merge` | Body: `{message?}`. `422` on an invalid manifest. `409` on conflict or if not `open`. | | POST | `/:crId/close` | `409` if already `merged`. | | POST | `/:crId/reopen` | `409` if not `closed`. | | POST | `/:crId/request-changes` | Body: `{feedback}`. Appends to `metadata.requested_changes` and wakes the originating session's agent. `409` if not `open`. | `POST /` rejects a `head_ref` with no commits ahead of `base_ref`: `422 code: "CR_HEAD_NOT_AHEAD"`. This is the error an agent sees if it opens a CR before pushing its branch. ### Authorization Read routes need read access to the project. Write routes need write access. Each write action also needs a capability, shown below for a full token and for a session's scoped token. | Action | Capability | | --- | --- | | Open a CR | `project.gitops.push` | | Request changes | `project.review.act` | | Merge | `project.gitops.merge` | One capability per action, for a full token and a session's scoped token alike. Opening and merging are separate leaves, so a token can open change requests without the power to merge them — that is the mechanism behind the agent mandate above. `project.cr.open` and `project.cr.merge` were the pre-cutover names for `project.gitops.push` and `project.gitops.merge` — the same capability under a second name. A `taskstation.yaml` that still lists one keeps working (the grant is rewritten to the live spelling when it is resolved), but write the `gitops` name in anything new. A session can never merge a change request it opened itself, whatever it has been granted. That rule is structural, not a capability. --- <!-- /markdown/docs/work.md --> # Running work How work moves from prompt to merged change, and the three ways it starts. Canonical page: https://taskstation.co/docs/work TaskStation does work inside a [session](/docs/work/sessions): a branch and a sandbox for one unit of work. A session ends when its [change request](/docs/work/change-requests) (CR) merges back to the default branch. This page walks through the loop, then shows the three ways a session can start. - [Sessions](/docs/work/sessions): A branch and a sandbox for one unit of work. - [Change requests](/docs/work/change-requests): The reviewed merge back to the default branch. - [Runtime](/docs/work/runtime): Env vars, tokens, and the sandbox image a session runs in. ## What happens when a session starts 1. TaskStation creates the session row and cuts a branch from the default branch. The branch name is the session id. 2. TaskStation resolves a sandbox image: the default image, or your own `.taskstation/Dockerfile` if the manifest declares one. 3. The sandbox boots. Its daemon, `taskstation-agent`, clones the repo to `/workspace`. It starts OpenCode REST. Session status becomes `running`. 4. The agent works. It uses [secrets](/docs/project/secrets) through the environment variables TaskStation sets, then commits and pushes to the session branch. 5. The agent opens a [change request](/docs/work/change-requests). You review it and merge it — the only way work reaches the default branch. > **Git is the only durable record** > Stopping a session pauses the sandbox but keeps its files. Deleting a session > destroys the sandbox for good. Only work committed and pushed to the branch > survives, and only a merged change request makes it permanent. ## Three ways work runs A session starts one of three ways. | Mode | How it works | |---|---| | On-demand | You ask in chat and get the result now. | | Human-assisted | The agent works and checks in with you for the calls that matter. | | Automated | A [trigger](/docs/connect/triggers) — a schedule or webhook — starts the session end to end. | ## Related - [Projects](/docs/project): A git repo with a manifest. - [Agents](/docs/project/agents): A markdown persona with scoped tools. - [Models](/docs/project/models): Which model a session uses, and who pays. - [Triggers](/docs/connect/triggers): Schedules and webhooks that start sessions. - [Connectors](/docs/connect/connectors): Scoped reach into external apps. - [Slack & channels](/docs/connect/slack): Chat surfaces that start sessions. - [Computers](/docs/connect/computers): Machines distinct from session sandboxes. - [Accounts](/docs/accounts): Principals, roles, and assignments — who can do what. --- <!-- /markdown/docs/work/runtime.md --> # Runtime & sandbox Sandbox lifecycle, injected environment, token families, and the image a session boots from. Canonical page: https://taskstation.co/docs/work/runtime A session runs in an isolated sandbox on Daytona, Platinum, or E2B Cloud, built from a layered image. This page is the reference for the sandbox lifecycle, the environment TaskStation injects at boot, the token families, and the sandbox image itself. For the session concept, see [Sessions](/docs/work/sessions). ## Sandbox lifecycle A session row carries a `status`. The enum defines `queued`, `branching`, `provisioning`, `running`, `stopped`, `failed`, `completed`, but TaskStation only writes 4 of them. | Status | Set when | | --- | --- | | `provisioning` | At session create. TaskStation creates the session branch and requests the sandbox. | | `running` | Once the sandbox is live and reachable. | | `stopped` | On explicit stop, or by the idle sweep that hibernates inactive sandboxes. | | `failed` | If provisioning fails. | `queued`, `branching`, and `completed` exist in the enum but stay dead in the session flow. Do not treat them as live states. The sandbox itself carries a separate `status` in its own row, with its own enum. | Status | Set when | | --- | --- | | `provisioning` | The provider boots the sandbox. | | `active` | The provider confirms the sandbox is live. | | `stopped` | Explicit stop, idle auto-stop, or mid-restart. | | `error` | The provider reports a boot or runtime failure. | | `archived` | Terminal. You deleted the session and the provider destroyed the sandbox, not paused it. | TaskStation enforces a concurrent-session limit per account. Exceeding your tier's limit returns `429`. ### Active-turn protection While an OpenCode turn is `busy` or `retrying`, the sandbox daemon renews a short execution lease with the API every 60 seconds. The lease blocks the idle reaper. Each renewal also touches the provider, so the provider's own inactivity timer cannot hibernate the sandbox mid-run. `session.idle` and `session.error` release the lease. An open dashboard tab, preview, SSE connection, or health poll does not create a lease. A passive tab cannot keep an idle sandbox alive. ### Branch model - The session branch is named after the session id (a UUID). `TASKSTATION_SESSION_ID` and `TASKSTATION_BRANCH_NAME` carry the same value. - TaskStation cuts the branch from `base_ref`, which defaults to the project's default branch, at session-create time. - Triggers create their session branch the same way an interactive session does. - Nothing writes the default branch directly. Only a merged [change request](/docs/work/change-requests) does. ### Reconcile a session branch Use **Ask Agent: Sync Branch & Reload** from the session command palette when the base branch changed. The agent inspects the session branch, preserves local work, fetches the latest `base_ref`, resolves conflicts, runs the relevant tests, and commits the reconciliation. TaskStation does not choose one side of a conflict or reset the working tree. The agent finishes with `taskstation sessions reload "$TASKSTATION_SESSION_ID" --project "$TASKSTATION_PROJECT_ID" --no-repo --force --yes`. `--no-repo` is required because the agent already reconciled the branch. The reload replaces the OpenCode runtime after the replacement becomes healthy. It ends the current turn, so send `continue` after the runtime returns. The session header reports each server-confirmed reload phase in real time. A web reload checks the session, compiles the agent config, applies and validates the runtime replacement, then confirms the active config. It does not refresh the repository. The CLI can refresh the repository unless you pass `--no-repo`. ## Layout inside the sandbox ``` /workspace ← WORKDIR. The project repo is cloned here. /workspace/.taskstation/ ← Repo-internal TaskStation folder (Dockerfile + opencode config dir). /usr/local/bin/taskstation-agent ← The daemon (supervisor + reverse proxy). /usr/local/bin/taskstation-entrypoint ← The container ENTRYPOINT (PID 1). /opt/taskstation/home ← OpenCode's HOME — its object store lives here, off the repo. ``` OpenCode's `HOME` is `/opt/taskstation/home`, not `/workspace`. Its object store never lands among your repo files. ## Injected environment TaskStation injects these variables at boot. Only a project secret explicitly configured for `runtime` delivery enters the sandbox. Connector and model-provider credentials stay on the TaskStation server. | Variable | What | | --- | --- | | `TASKSTATION_PROJECT_ID` | UUID of this project. | | `TASKSTATION_SESSION_ID` | UUID of this session. Also the branch name. | | `TASKSTATION_BRANCH_NAME` | Same value as `TASKSTATION_SESSION_ID`. | | `TASKSTATION_REPO_URL` | Clone URL for the project repo. | | `TASKSTATION_DEFAULT_BRANCH` | The project's default branch. | | `TASKSTATION_BASE_REF` | The ref this session branched from. | | `TASKSTATION_SERVICE_PORT` | `8000` — the daemon's external port. | | `TASKSTATION_API_URL` | The platform API base (`.../v1`). | | `TASKSTATION_AGENT_NAME` | The agent the session was created with. | | `TASKSTATION_OPENCODE_MODEL` | The model to run, when set. | | `TASKSTATION_PROJECT_AUTO_CLONE` | `1` — tells the daemon to clone the repo on boot. | | `TASKSTATION_PROJECT_SECRET_NAMES` | Comma-separated names of the project's secrets. | | `TASKSTATION_PROJECT_SECRETS_REVISION` | Revision marker for the secret set. | | `TASKSTATION_BOOTSTRAP_OPENCODE_SESSION` | `1` — always set. Tells the daemon to create the OpenCode root on cold boot. | | `TASKSTATION_LLM_BASE_URL` | TaskStation LLM-gateway base URL. The gateway resolves provider credentials server-side. | | `TASKSTATION_TOKEN` | The session-bound TaskStation credential. See below. | TaskStation does not inject `TASKSTATION_WORKSPACE`. The image bakes in `/workspace` and no per-session step sets it. TaskStation does not inject a git token either — the daemon fetches a short-lived clone credential when it needs one; see [Pushing from a session](#pushing-from-a-session) below. TaskStation rejects a user secret named with the `TASKSTATION_*` prefix, because it reserves that prefix for platform variables. ## Session credential TaskStation has external tokens you create yourself and one credential inside each sandbox. ### External tokens A personal access token (PAT, prefixed `taskstation_pat_`) or a service account (prefixed `taskstation_sa_`) authenticates calls to the TaskStation API, the SDK, and the CLI from outside a sandbox. See [Authentication](/docs/sdk/auth) for PAT scope, service-account setup, and how to choose between them. ### In-sandbox token `TASKSTATION_TOKEN` is minted when TaskStation starts the session environment. It is bound to the launching user, project, session, and agent grant. The daemon, CLI, Git credential helper, LLM gateway, and connector gateway use this same credential. Each API route still applies its own capability check. The initial prompt and turn-ledger identifiers are not environment variables. The daemon claims them from the API with `TASKSTATION_TOKEN` after boot. ## Pushing from a session The daemon sends `TASKSTATION_TOKEN` only to the TaskStation Git proxy. The proxy resolves the upstream Git credential on the server. No upstream Git token enters the sandbox. `git push origin HEAD` sends commits to the session branch. Landing on the default branch requires a merged [change request](/docs/work/change-requests). ## The agent runtime The daemon launches OpenCode as `opencode serve --port 4096 --hostname 127.0.0.1`, with `OPENCODE_CONFIG_DIR` set to the project's config directory (default `.taskstation/opencode`) inside the cloned repo. See [Agents](/docs/project/agents) for how a session picks an agent and its config. ## Transcript attachments and memory OpenCode stores tool screenshots as base64 `data:` URLs inside its SQLite transcript. The daemon keeps that store small and the box alive: - **Attachment offload.** While no turn runs (every 5 minutes, at boot, and after a memory-guard event), attachment bytes older than the newest 12 per session — and every tool result OpenCode's compaction already cleared — move to `~/.local/share/taskstation/attachments/<id>`. The row keeps a 1×1 PNG placeholder plus a `taskstation.offloaded` marker; `/taskstation/part` serves the real bytes to the UI. Models never receive those old images anyway (the LLM proxy keeps the newest 12 per request). Set `TASKSTATION_ATTACHMENT_OFFLOAD=0` to disable. - **Resource telemetry.** `[resources]` in the daemon log every 60 s and on every OpenCode state change: box memory, cgroup limit, load, disk, daemon and OpenCode RSS. - **Memory guard.** Above 80 % memory the daemon samples every 10 s; at 92 % (`TASKSTATION_MEMORY_GUARD_PCT`) with a turn in flight it aborts the turn cleanly and reports `SandboxMemoryGuard` with the numbers as the turn's error, instead of letting the kernel kill OpenCode mid-turn. ## The daemon control surface The `taskstation-agent` binary runs as PID 1's child and fronts OpenCode on `TASKSTATION_SERVICE_PORT` (`8000`). Every route outside `/taskstation/*` requires the HMAC-signed `X-TaskStation-User-Context` header, validated against `TASKSTATION_TOKEN`. | Path | Purpose | | --- | --- | | `GET /taskstation/health` | Liveness check (no auth required). Reports daemon and OpenCode state, repo, branch, commit. | | `POST /taskstation/refresh` | Re-pull the session branch and restart OpenCode in place. | | `POST /taskstation/abort` | Abort the current run. | | `POST /taskstation/env` | Update the runtime environment. | | `/taskstation/pty` | Backs the in-dashboard terminal. | | `GET /taskstation/logs` | Tails the daemon's own log file (`/opt/taskstation/logs/daemon.log`, rotated at 32 MiB) or OpenCode's. `?source=daemon\|opencode\|all`, `?tail=N` (default 500, max 5000). Plain text. Same auth as `/taskstation/refresh`. | | `GET /taskstation/part/:sessionID/:messageID/:partID` | Attachment bytes on demand — a top-level file part or a tool result's `state.attachments[]` entry. Bytes the daemon offloaded to a sidecar file are served from there. | | `GET /taskstation/diag` | One JSON error report: OpenCode state/pid/port/port pair, boot timeline, a fresh resource snapshot (memory, cgroup limit, load, disk, RSS, duplicate opencode processes), the runtime-assets report, and the tail of both logs (`?tail=N`, default 200). Same auth as `/taskstation/logs`. | | `/proxy/{port}/*` | Reverse-proxy to another port inside the sandbox. The daemon's own port is blocked. | | `*` | Catch-all reverse-proxy to OpenCode on `127.0.0.1:4096`. Returns `503` while OpenCode boots. | Run `/taskstation/refresh` to apply an out-of-band change, such as a manifest edit committed from a parallel session, without re-provisioning the sandbox. ## The sandbox image Every sandbox boots from an image built in two layers. Your Dockerfile defines the base environment. The TaskStation runtime layer is added on top, so the dashboard can connect to the sandbox. ``` ┌─────────────────────────────────────────┐ │ TaskStation runtime layer (added on top) │ ← opencode + taskstation-agent + entrypoint ├─────────────────────────────────────────┤ │ Your Dockerfile │ ← .taskstation/Dockerfile └─────────────────────────────────────────┘ ``` If your project has no Dockerfile, TaskStation builds sessions from a bare `ubuntu:24.04` image plus the runtime layer below. ### Declare a template Reference your Dockerfile as a named template under `sandbox.templates` in `taskstation.yaml`. ```yaml sandbox: templates: - slug: dev name: Dev box dockerfile: .taskstation/Dockerfile default: dev ``` Set exactly one of `dockerfile` or `image` on each entry. Dockerfile paths must stay inside the repository. Set `default` to the template slug your sessions should use; omit it to use the platform default image. Older projects on `taskstation.toml` follow the same fields — see [legacy `taskstation.toml`](/docs/project/legacy-toml). Full field list: [manifest reference](/docs/project/manifest). ### What the runtime layer adds TaskStation appends this layer on top of your Dockerfile's final stage. - A system package floor with `git`, `curl`, `build-essential`, `ffmpeg`, and `tmux`. - pnpm-managed Node.js and npm, plus uv-managed Python 3. - A document-tools floor: LibreOffice, Pandoc, and OCR tools. - An exact uv-managed Python version exposed as `python` and `python3`. Use `uv run --with <package>` for third-party dependencies. - `opencode-ai`, the `bun` runtime, and `agent-browser` with a baked Chromium build. - The `taskstation-agent` daemon, the `taskstation` CLI, and the entrypoint script. - `ENV TASKSTATION_WORKSPACE=/workspace`, `WORKDIR /workspace`, `EXPOSE 8000`, and the entrypoint that starts the daemon. Everything you install in your own Dockerfile stays on `PATH`. TaskStation does not remove or relocate it. ### Constraints | Rule | Why | | --- | --- | | Don't set `ENTRYPOINT` or `CMD`. | TaskStation overrides both to start the daemon. | | Don't claim port `8000`. | Reserved for the daemon's reverse proxy. Run dev servers on other ports. | | `FROM` a Debian or Ubuntu base. | The runtime layer runs `apt-get`. Alpine, Fedora, and Arch fail the build. | | Don't run `apt-get clean` without `rm -rf /var/lib/apt/lists/*`. | The runtime layer re-runs `apt-get update`; a broken cache breaks the build. | | Don't bake credentials into the image. | Declare the name in `env:` and set the value as a project secret. TaskStation injects it at session start. | ### Hardware spec `cpu`, `memory`, and `disk` on a template entry set the sandbox size. All three are optional; an omitted field uses the platform default: 2 vCPU, 4 GiB memory, 20 GiB disk. ```yaml sandbox: templates: - slug: big image: ubuntu:24.04 cpu: 4 memory: 8 disk: 50 ``` `cpu` takes 1–32 cores, `memory` takes 1–128 GiB, and `disk` takes 1–500 GiB. TaskStation clamps any value above these limits down to the limit. TaskStation does not support GPUs; a `gpu` key on a template produces a warning, not an error. The spec is part of the template's snapshot, not a per-session setting. Changing it rebuilds the snapshot and applies to the next session. The current session keeps its already-booted spec. The default spec costs about $0.10 per hour on Daytona, the default sandbox provider. The same spec on Platinum or E2B costs about twice as much, because the Daytona rate includes a volume discount the other providers don't. TaskStation meters this cost only for Team accounts; free and self-hosted plans aren't billed for it. ### Ports and preview URLs The daemon listens on port `8000` and proxies any other port your app uses inside the sandbox. Only port `8000` itself is blocked from the proxy. The dashboard reaches a sandbox port two ways: - **Path-based**: `https://<api-host>/v1/p/<sandbox-id>/<port>/...`. This is the default form and the only one that supports WebSocket upgrades. - **Subdomain-based**: `https://p<port>-<sandbox-id>.<api-host>/...`. Use this form for apps that need root-relative paths or cookies, such as a Next.js or Vite dev server. WebSocket upgrades don't work on this form yet. The preview proxy strips `X-Frame-Options` and any `frame-ancestors` policy. A session can embed your app's preview in an iframe without any config on your side. A session's preview can also be shared through a public link, in view-only or interactive mode. Public share links block ports `22`, `4096`, `8000`, and the static file-share port. ### Snapshot rebuilds TaskStation content-addresses each snapshot: it hashes your Dockerfile's bytes, the hardware spec, and the platform's own runtime version. An unchanged hash reuses the existing snapshot. A changed hash triggers a rebuild, shown as "preparing image" on the first session that needs it; later sessions reuse that build. Editing the Dockerfile inside a session takes effect on the next session, not the current one. The edit reaches `main` only once its [change request](/docs/work/change-requests) merges. --- <!-- /markdown/docs/work/sessions.md --> # Sessions A session is a git branch and a sandbox where the agent works. Canonical page: https://taskstation.co/docs/work/sessions A session is one unit of agent work. TaskStation cuts a git branch and provisions a sandbox for it. The session id, the branch name, and the sandbox id are the same value. ## Status A session reaches one of 4 states in practice. | Status | Meaning | |---|---| | `provisioning` | TaskStation cuts the branch and requests the sandbox. | | `running` | The sandbox is live and reachable. | | `stopped` | The session is paused, by you or by idle auto-stop. | | `failed` | Provisioning failed. | The database defines 3 more values (`queued`, `branching`, `completed`). TaskStation does not write them for a session today. ## Stop, resume, and idle auto-stop You can stop a session yourself. Resume brings back the same sandbox with the same filesystem and runtime identity. Only the running processes and memory reset. The OpenCode conversation remains attached to the session. TaskStation also stops an idle session for you: - After 15 minutes idle, for a normal session. - After 5 minutes idle, for a session a trigger started. An open dashboard tab does not keep a session alive. A busy agent turn blocks the stop. The maintenance sweep runs every 5 minutes. A normal automatic Stop therefore occurs approximately 15 to 20 minutes after the terminal turn. Self-host operators can set `TASKSTATION_SANDBOX_AUTOSTOP_MINUTES` to change the normal idle grace. This setting does not change active-turn protection. ## What stop and delete keep > **Deletion is permanent** > Deleting a session destroys its sandbox for good. TaskStation keeps the session > record and the git branch, so you can still recover pushed work. Anything not > pushed is gone. Stop and resume keep the sandbox's identity and filesystem. Delete destroys the sandbox. Git is the only durable record: work the agent commits and pushes survives; everything else does not. ## Runtime Every session uses OpenCode REST. TaskStation stores the selected OpenCode agent and model when the session starts. Restart and resume keep the same session runtime. ## Session access A session is private to the person who created it. The owner opens **Session access** and picks one of 3 options. | Option | Who can open the session | |---|---| | Only you | The owner alone. This is the default. | | Specific people | The owner, plus the members and groups the owner picks. | | Whole project | Every member of the project. | Everyone with access reads the conversation and continues it. **Only the owner changes this.** A project manager who did not create the session can open it once it is shared with them, and can stop, restart, or delete it. They cannot rewrite who else can open it. Sharing a session with a manager is not handing them its access list. One kind of session has no human owner: the ones a trigger creates, which run under the trigger agent's identity. Project managers govern those. Set the policy for all of them on the trigger itself, under **Session access** on the trigger — saving there also updates the sessions that trigger already created. > **A session keeps its owner** > Removing someone from the account does not move their sessions to anyone else. > Their access policy freezes as it was. A project manager can still stop or > delete those sessions — deleting is the way to revoke a session nobody owns any > more. ## Sharing a preview You can share a session's live preview with a public link, in view-only or interactive mode. Minting that link is the session owner's call, for the same reason: the link is unauthenticated, so anyone holding the URL reads the session without signing in. A project manager can list and revoke a session's links without owning it — revoking only ever removes access. ## Providers TaskStation runs sessions on Daytona, Platinum, or E2B Cloud. A project follows the platform default, or requests a provider switch through the SDK — see [SDK reference](/docs/sdk/reference). A switch to a different provider is durable: the current provider keeps serving while the target warms, then activates. Every provider runs the same sandbox image. For the full status enum, injected environment variables, and daemon endpoints, see [Runtime](/docs/work/runtime). For how a session picks its agent, see [Agents](/docs/project/agents). For sessions a schedule or webhook starts, see [Triggers](/docs/connect/triggers). To land session work on the default branch, see [Change requests](/docs/work/change-requests). --- <!-- /markdown/use-cases/access-requests.md --> # How we handle access requests An access request in Slack triggers an agent to check policy, gather context, and prepare a least-privilege grant, with every grant requiring a human approval and logged. Canonical page: https://taskstation.co/use-cases/access-requests Access requests arrive constantly and informally. Someone needs a repo, a role in Okta, or a cloud IAM permission to finish a task, and they ask in Slack. Whoever holds the access has to check what the person's role should have, work out the narrowest grant that unblocks them, apply it, and remember to record it. Under time pressure the easy path is to grant broadly and move on, and the record of who has what drifts. We handle this by tying the grant to the request that starts it, and to the policy that governs it. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails — so you can set up the same flow. - **Team:** TaskStation - **Trigger:** An access request in Slack - **Connected systems:** Okta · GitHub · Cloud IAM - **Mode:** Policy-checked · Approval-gated · Logged ## The problem The common approaches each fall short. A ticket queue routes the request to a person who still does all the lookup and grant work by hand. A standing broad role avoids the back-and-forth but hands out more than the task needs. And a self-serve grant with no policy check trades safety for speed. None of them keep a clean record of what was granted and why. We wanted each request checked against policy, scoped to the least privilege that unblocks the work, and granted only after a person signs off, with a log left behind. ## What we built On TaskStation, an access request in Slack triggers an agent. Each request runs in its own isolated session — a cloud sandbox — with scoped access to Okta, GitHub, and cloud IAM. The agent reads the request, checks it against policy, gathers context on the person's role, team, and the least-privilege scope that fits, and prepares the grant. Every grant requires a human approval, and each one is logged. ## How it works ### Connect Slack as the trigger Slack is connected as a **channel**, so a request is the trigger. Post the access request and it spawns a fresh **session** in its own sandbox, seeded with who's asking and what they need. One request, one session, one disposable machine. ### Give the agent the access policy Which roles map to which grants, what least privilege means for each system, and which requests need extra scrutiny are stored as **skills** and **memory** that load into every session. The agent checks requests against that policy rather than improvising, and it updates as the policy changes. ### Connect what a grant can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read Okta** — the requester's role, team, and current group memberships. - **Prepare a GitHub grant** — repo or team access scoped to what the task needs. - **Prepare a cloud IAM grant** — the narrowest role or permission set that unblocks the work. ### Set the guardrails Granting access is the step that changes who can do what, so every grant stops at a **human approval gate** — no access is applied until a person signs off. Each approved grant is logged with the request, the policy check, and the scope. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each request come back scoped and logged With that in place, a request in Slack comes back as a policy-checked, least-privilege grant ready for a person to approve, with a record attached. "I need access to this repo" becomes a scoped GitHub grant; "I need this cloud role" becomes the narrowest IAM permission that fits, both held for sign-off. > **The pattern** > Connect Slack via a **channel** trigger, give the agent scoped **connectors** > into Okta, GitHub, and cloud IAM, encode the access policy as **skills** and > **memory**, and gate every grant behind a human with a log left behind. ## Guardrails Giving an agent a hand in access decisions is a trust question. The relevant controls on TaskStation: - **Isolation.** Every request runs in its own isolated sandbox on its own branch. The session is granted access only to the systems it's scoped to, and only what it's explicitly allowed to send is written back out. - **Scoped secrets.** Each credential is encrypted in the secrets manager, injected into the sandbox at runtime, and scoped to the agents you grant them to or the logs. - **Human approval gate.** No grant is applied until a person approves it, and each approved grant is logged. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Least privilege:** Every grant scoped to what the task needs - **Approval-gated:** No access applied without a person signing off - **Logged:** Every grant recorded with its request and scope Access requests that used to be granted broadly under time pressure now come back scoped, policy-checked, and ready for a person to approve, with a record left behind. Extending it to another system means connecting one more platform. --- <!-- /markdown/use-cases/ad-performance.md --> # How we monitor ad performance The ad-performance agent we run on TaskStation — connected to Google Ads and Meta Ads, with alerts in Slack. Every day it checks budget pacing, CPA and ROAS drift, and underperforming ads and keywords, then posts ranked findings and optimization recommendations — it never changes a budget, pauses a campaign, or edits an ad itself. Canonical page: https://taskstation.co/use-cases/ad-performance Paid ad spend drifts quietly. A campaign paces ahead of budget three days into the month. A cost-per-acquisition creeps up for a week before anyone notices the trend line. A keyword that used to convert keeps taking budget from ones that still do. None of it is a single dramatic failure — it's a slow leak across two ad platforms that nobody has time to check line by line every morning. We run an ad-performance agent on TaskStation that reads campaign spend and performance from Google Ads and Meta Ads every day and posts ranked findings and optimization recommendations to Slack. It never touches a budget, a campaign, or an ad — it only ever recommends. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Google Ads · Meta Ads · Slack - **Mode:** Read-only · recommend-only · one Slack post per day ## The problem Budget pacing, CPA/ROAS drift, and underperforming ads or keywords are each easy to catch on their own — if someone is looking. In practice they live split across two platforms with different dashboards, and checking both every day for every active campaign doesn't survive contact with a real workload. By the time a weekly review catches a campaign that's overspent its monthly budget by day ten, or a keyword whose CPA has doubled, the money is already spent. The usual fixes don't hold up. Each platform's native alerting fires on its own thresholds and only sees its own account. A weekly spreadsheet review is thorough but a week is a long time for a pacing problem or a CPA spike to compound. And automated bid/budget tools that act on their own remove the one thing a marketing team actually wants to keep: the decision. ## What we built On TaskStation, a daily cron spawns a fresh agent session. It reads campaign spend and performance from Google Ads and Meta Ads, checks budget pacing against each campaign's monthly target, computes CPA and ROAS drift against trailing performance, flags underperforming ads and keywords and any spend or click-through anomalies, and drafts prioritized optimization recommendations — pause a specific ad, shift budget toward a better performer, add a negative keyword — before posting the ranked list to Slack. It writes nothing back to either ad platform. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox. One day maps to one run on one disposable machine, so pacing and drift are recomputed from the current state of both platforms every time — nothing carries over. ### Give the agent the optimization playbook What counts as off-pace, how much CPA or ROAS drift is worth flagging, what makes an ad or a keyword "underperforming," and how a good recommendation is worded live as a **skill** that travels with the agent. When we tighten a threshold or add a new anomaly pattern, it's a change to the skill, not a one-off instruction. ### Connect the ad platforms read-only Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent reads: - **Campaign spend and performance from Google Ads** — daily spend, clicks, conversions, and cost per acquisition per campaign, ad, and keyword. - **Campaign spend and performance from Meta Ads** — the same, across Facebook and Instagram placements. - **Posts to Slack** — the ranked findings and recommendations, once a day. It has no write access to either platform. It cannot change a budget, pause a campaign, or edit an ad. ### Set the guardrails The agent's only actions are reading two ad platforms and posting to Slack. It **recommends** pausing an ad, shifting budget toward a better performer, or adding a negative keyword — it never pauses, edits, or reallocates spend itself. That decision, and the click that executes it, stays with the marketing team. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Post the ranked findings Each day's message lists what needs attention: campaigns off their budget pace, ads or keywords with CPA/ROAS drifting the wrong way, underperformers worth pausing, and anything that looks anomalous — each with the numbers behind it and a specific recommended action. The marketing team reads it and decides what to actually change. > **The pattern** > A daily **cron** spawns a fresh session with read-only **connectors** into > Google Ads and Meta Ads. The optimization rules live as a **skill**. The > agent reads spend and performance from both platforms and writes nothing but > the Slack post — every pause, budget shift, and keyword change is a > recommendation, never an action. ## Guardrails The agent reads live spend data across two ad platforms, so its access is scoped and one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to Google Ads and Meta Ads, and only the Slack post is written back out. - **Scoped secrets.** The Google Ads and Meta Ads credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Recommend-only.** The agent has no write access to either ad platform. It cannot change a budget, pause a campaign, or edit an ad — it can only report and suggest. - **Everything is code.** The agent's optimization rules, thresholds, and per-platform permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Spend and performance rechecked across both platforms - **Recommend-only:** Nothing is ever paused or reallocated automatically - **2 platforms:** Google Ads and Meta Ads watched in one daily post The pacing check, the CPA/ROAS trend, and the underperformer scan that used to require opening two dashboards now arrive as one ranked Slack post every morning, each finding carrying the numbers and a specific recommendation. The agent only reads and recommends; the marketing team decides what to pause, shift, or exclude. --- <!-- /markdown/use-cases/ap-invoice-processing.md --> # How we process vendor invoices as they land The accounts payable agent we run on TaskStation — connected to Gmail, Google Sheets, and Slack. Every 15 minutes it pulls new invoices from the inbox, extracts vendor, amount, and line items, matches them against POs and the ledger, and routes anything clean or flagged to Slack for approval — never scheduling or marking a payment itself. Canonical page: https://taskstation.co/use-cases/ap-invoice-processing Vendor invoices arrive by email all day, on no schedule anyone controls, and each one needs the same handling: open the attachment, read off the vendor and the amount and the line items, check it against the purchase order, check it isn't a duplicate of something already in the ledger, and check the price is what was agreed. Done by hand, that's five minutes of careful reading per invoice, and the five minutes is where duplicates and overcharges slip through — not because anyone is careless, but because checking against every prior invoice and every open PO doesn't scale to a person's afternoon. We run an accounts-payable agent on TaskStation that checks the inbox every 15 minutes, extracts every new invoice, matches it against POs and the AP ledger, and posts it to Slack for approval — flagged if it's a duplicate, an overcharge, or missing a PO. It records; it never pays. - **Team:** TaskStation - **Runs on:** Cron, every 15 minutes - **Connected systems:** Gmail · Google Sheets · Slack - **Mode:** Record + route for approval — never schedules or pays ## The problem An invoice by itself is just a number. Whether it's *right* depends on comparing it against things that live elsewhere: the PO that authorized the purchase, the price that was agreed, and every invoice from that vendor already in the ledger. A person doing this by hand checks the easy cases — is there a PO, is the total plausible — and the checks that actually catch problems, like a vendor resending last month's invoice or padding a line item by a few percent, get skipped when the inbox is long and the afternoon is short. The common shortcuts don't fix this. Paying on receipt catches nothing. A monthly batch review catches duplicates and overcharges after the fact, sometimes after the payment has already gone out. Neither one checks every invoice against every PO and every prior invoice at the moment it arrives. ## What we built On TaskStation, a cron fires every 15 minutes. Each firing spawns a fresh agent session — a cloud sandbox with no memory of the last run, because the AP ledger in Google Sheets *is* the record. The agent checks a Gmail label for new invoice emails, extracts the vendor, amount, and line items from each attachment, matches it against the known POs and prior invoices tracked in the sheet, flags anything that's a duplicate, an overcharge, or missing a PO, and records every invoice in the ledger. It then posts the batch to Slack for a person to approve. It never schedules a payment or marks anything paid. ## How it works ### Run every 15 minutes, fresh A **cron trigger** fires the agent every 15 minutes. Each firing spawns a fresh **session** in its own sandbox — nothing carries over in the agent's own memory. The AP ledger in Google Sheets is the continuity: every run reads it before doing anything, so duplicate and overcharge checks are judged against the same record every time. ### Give the agent the AP rules How we extract, match, and flag lives as a **skill** loaded into every session: what counts as a duplicate, the tolerance allowed before a line item counts as an overcharge, and what to do when no PO matches at all. The agent works to that standard instead of improvising it invoice by invoice. ### Connect what invoice processing needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read Gmail** — new invoice emails and their attachments in the watched label. - **Read and append to Google Sheets** — the POs and the AP ledger it matches against, and the new rows it records. - **Post to Slack** — the batch of processed invoices, flagged or clean, for approval. ### Set the guardrails The agent's only write is an appended row in the AP ledger and a Slack post. It has no access to a payment rail, and nothing in its instructions lets it schedule or mark an invoice as paid — that action belongs to a person, every time, with no exception for a clean match. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Route the batch for approval With that in place, every 15 minutes any new invoices are already extracted, matched, and recorded by the time a person looks at Slack — each one marked clean, duplicate, overcharge, or missing-PO, with the PO and prior-invoice references attached. A person reviews the batch and approves it for payment through the normal AP process. The agent's part ends at the Slack post. > **The pattern** > A **cron** every 15 minutes spawns a fresh session; the Google Sheet ledger is > the memory. The agent reads Gmail, matches against POs and prior invoices > through **connectors**, records every invoice in the sheet, and posts the batch > to Slack. It never schedules or marks a payment as paid — that's a human, always. ## Guardrails Accounts payable is where a mistake costs real money, so the agent's access is scoped to match: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to Gmail, the AP sheet, and Slack, and nothing else is written back out. - **Scoped secrets.** The Gmail, Sheets, and Slack credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **No payment action, ever.** The agent has no connector, tool, or instruction that schedules a payment or marks an invoice as paid. Every invoice, clean or flagged, is routed to a person for that decision. - **Flag, don't discard.** Duplicates, overcharges, and missing-PO invoices are recorded and flagged in the ledger, never silently dropped or auto-approved. - **Everything is code.** The agent's matching rules, skill, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every 15 min:** New invoices extracted, matched, and recorded - **0 payments:** Scheduled or marked paid by the agent - **3 checks:** Duplicate, overcharge, and missing-PO on every invoice The five careful minutes an invoice used to wait for now happen automatically, every 15 minutes, against the full history in the ledger rather than whatever a person remembers. What used to reach the payment step unchecked now arrives flagged, with the PO and the prior invoice it conflicts with already attached — and the decision to pay stays exactly where it was: with a person. --- <!-- /markdown/use-cases/ar-chaser.md --> # How we chase down overdue invoices An accounts-receivable agent we run on TaskStation — connected to Stripe, email, and the accounting ledger. It finds overdue invoices, sends the right reminder, logs every touch, and reconciles when payment lands. Canonical page: https://taskstation.co/use-cases/ar-chaser Chasing overdue invoices is steady, repetitive work that still needs judgment. Most reminders are routine: an invoice is a few days late, a polite note goes out, the customer pays. But the timing, the tone, and the escalation depend on how late the balance is and how large it is, and some accounts are sensitive enough that no automated email should go out without a person looking first. We run an accounts-receivable agent on TaskStation that does this chasing on a daily schedule. This is how we collect on our own invoices, including the connections and guardrails involved. - **Team:** TaskStation - **Runs on:** A daily cron - **Connected systems:** Stripe · Email · Accounting ledger - **Mode:** Read-mostly · reminders and status notes gate-able ## The problem Receivables slip because no one has time to work the list every day. An invoice goes a week late, then two, and the reminder that should have gone out on day one never does. The ones that need a firmer note get the same generic email as the ones that are barely late, and the sensitive accounts — a large balance, a disputed line, something heading to legal — get chased the same way as everything else. The common fixes are incomplete. Stripe's built-in reminders send on a fixed cadence regardless of balance size or account context. A spreadsheet and a calendar reminder depend on someone actually working it. A generic automation sends the same email to everyone and has no way to hold the risky ones back. ## What we built On TaskStation, a daily cron triggers an agent. Each morning it spawns an isolated session (a cloud sandbox) with scoped access to Stripe, our email, and the accounting ledger. It finds every invoice that's overdue or coming due soon, decides the right reminder for each based on how late and how large the balance is, sends it, logs the touch, and reconciles the invoice when payment lands. Anything sensitive stops at a human approval gate before it sends. ## How it works ### Trigger the run on a daily cron A **cron trigger** fires once a day and spawns a fresh **session** in its own sandbox. Each run works the full receivables list from scratch, so nothing carries over between days and a missed morning is just the next run picking up where it left off. One run maps to one session on one disposable machine. ### Give the agent the collections playbook Our collections policy lives as **skills** and **memory** that travel with the agent: the reminder cadence by days overdue, the escalating tone from a first notice to a final notice, the balance thresholds that change the approach, and which accounts are flagged as disputed or sensitive. When we adjust the policy, we write it down and the agent applies it on the next run. ### Connect Stripe, email, and the ledger Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read invoices from Stripe** — which are overdue, which are coming due, the balance and age of each. - **Send reminders by email** — the right escalating notice for each invoice, from the agent's own address. - **Read the accounting ledger** — to confirm balances and reconcile against what Stripe reports. - **Log every touch and status note** — each reminder and reconciliation recorded against the invoice. ### Set the guardrails The agent is **read-mostly**: it reads Stripe and the ledger, and the only writes it makes are reminder emails and status notes. Anything sensitive — a large balance, a disputed invoice, or an account heading to legal — stops at a **human approval gate** before it sends. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Let the list work itself each morning With that in place, the daily run finds the overdue and soon-due invoices, sends each account the reminder its timing and balance call for, and logs the touch. A routine day-three nudge goes out on its own. A large or disputed balance is held for a person to approve. When a payment lands, the agent matches it to the invoice and marks it settled. > **The pattern** > A daily **cron** spawns a session with scoped **connectors** into Stripe, email, > and the ledger. The collections policy is encoded as **skills** and **memory**. > The agent stays read-mostly, sensitive accounts wait for a human, and every > touch is logged. ## Guardrails The agent sends email on our behalf and touches invoice state, so the access is scoped and contained: - **Isolation.** Every run happens in its own per-task isolated sandbox. The session reads Stripe and the ledger and can send only the reminders and status notes it's scoped to; nothing else is written back out. - **Scoped secrets.** The Stripe, email, and ledger credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** A large balance, a disputed invoice, or an account heading to legal stops for a person to approve before any email goes out. - **Everything is code.** The agent's configuration, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every morning:** The full receivables list worked on schedule - **Read-mostly:** Only reminders and status notes are written - **3 systems:** Stripe, email, and the ledger in one agent Overdue invoices now get the right reminder on the day they should, with each touch logged and each payment reconciled when it lands. Routine chases run on their own, the sensitive accounts wait for a person, and the receivables list no longer depends on someone finding time to work it. --- <!-- /markdown/use-cases/brand-monitor.md --> # How we monitor brand mentions A daily agent that searches news, social platforms, and forums for new mentions of our brand, classifies sentiment, and posts a digest of the notable or negative ones to Slack — each with a suggested response drafted for a human to review. Canonical page: https://taskstation.co/use-cases/brand-monitor Brand mentions show up wherever people happen to be talking — a news writeup, a Reddit thread, a tweet, a forum post comparing us to a competitor. Most days nothing needs a response. Some days a single negative post with real reach needs one within the hour, and the team doesn't find out until it's already been seen by everyone else. We run a brand-monitor agent on TaskStation that checks the web every day for new mentions, classifies each one's sentiment, and posts a digest to our marketing Slack channel. It never replies or posts anywhere itself — for anything notable, it drafts a suggested response and leaves the sending to a person. - **Team:** TaskStation - **Runs on:** Daily cron, reusable session - **Connected systems:** The web · Slack - **Mode:** Search-only · draft, never post ## The problem Brand mentions aren't confined to one platform. They land in news coverage, social posts, and forum threads on no predictable schedule, and most of them don't matter — a passing mention, a neutral comparison, a repost. The ones that do matter are indistinguishable from the noise until someone actually reads them, and by the time a person gets around to checking, a negative post with real reach has often already had its worst hour. The usual workarounds don't fix this. A Google Alert fires on keyword matches with no sense of sentiment or reach, so every alert looks the same regardless of whether it's a five-word tweet or a widely shared negative review. Checking manually catches yesterday's mentions today, if someone remembers to look. Either way, nothing tracks what's already been seen, so the same viral post gets flagged again tomorrow. ## What we built A daily cron re-prompts a persistent agent session. Each run searches the web for new mentions of the brand across news, social platforms, and forums, checks each one against a ledger of mentions it has already reported, classifies the sentiment of anything new, and posts a digest to Slack — a one-line rollup for routine mentions, and a full callout with a suggested response draft for anything notable or negative. It never posts, replies, or comments anywhere; the only output is the Slack digest. ## How it works ### Run on a daily reusable session A **cron trigger** fires the project once a day, re-prompting the same **session** rather than starting from a blank slate. Because the session is reusable, it remembers what it already reported yesterday and only surfaces what's new today. ### Give the agent the brand to watch The brand name, aliases, product names, and common misspellings to search for live as **skills** and **memory**, along with the running ledger of every mention already seen. When the brand adds a product line or a new alias starts getting used, the watch list is updated the same way. ### Search and fetch the web Through the sandbox's **built-in web search and fetch** — no connector, no login, no credential, because these are public pages — plus a scoped **Slack connector** for the digest, the agent can: - **Search news, social platforms, and forums** — for the brand terms on the watch list, this run only. - **Fetch the full text of each candidate mention** — enough to classify sentiment and reach, not just a search snippet. - **Post to Slack** — one digest per run, in the marketing channel. ### Set the guardrails The agent only searches and reads public pages, and it only ever writes to one Slack channel. It never replies to a post, comments, or publishes anywhere on the open web — for anything flagged as notable, it drafts a suggested response and clearly marks it as a draft awaiting human approval. Any credential it needs is encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Post the digest Each run ends with one Slack message: how many new mentions were found, a sentiment breakdown, and — for anything negative or notable — the source, the excerpt, why it was flagged, and a suggested response draft. Routine mentions get a one-line rollup instead of a full callout. The team reads one digest and decides what, if anything, to say back. > **The pattern** > A daily **cron** re-prompts a reusable **session** that searches the web for > brand mentions, dedupes against a **memory** ledger of what's already been > reported, classifies sentiment, and posts a Slack digest. Anything notable > gets a suggested response — drafted, never sent. ## Guardrails The agent watches public conversation about the brand, so its access is scoped to searching and reading, with a single, narrow output: - **Isolation.** Every run happens in its own isolated sandbox, and only the Slack digest it's explicitly allowed to send is written back out. - **Scoped secrets.** Any credential the agent uses is encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Draft, never post.** The agent's only external action is one Slack digest per run. It never replies, comments, or posts on the platform where a mention appeared — a suggested response is a draft for a human to review and send themselves. - **Everything is code.** The brand terms, the watch scope, and the sentiment and notability rules are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** New mentions found without anyone remembering to check - **Dedupe by design:** Already-reported mentions never resurface - **Draft only:** A suggested response, never a reply the agent sends itself The team no longer finds out about a negative mention secondhand, and no one re-reads yesterday's viral post because it was flagged twice. The channel gets one digest a day — quiet on quiet days, specific and actionable when something needs a response — and every response that goes out is one a person chose to send. --- <!-- /markdown/use-cases/candidate-sourcing.md --> # How we source candidates for open roles The outbound-sourcing agent we run on TaskStation — daily, it finds candidates on LinkedIn for each open role, checks Greenhouse so no one already in the pipeline gets a duplicate reach-out, and drafts a personalized email for the recruiter to review and send. Canonical page: https://taskstation.co/use-cases/candidate-sourcing Filling an open role isn't just reading the applications that arrive — the best candidates for a lot of roles aren't applying anywhere, they're already employed and not looking. Finding them means searching LinkedIn by hand, checking the ATS to make sure whoever you found isn't already in the pipeline, and writing an intro that actually references their background instead of a generic template. Done well, it's hours of work per role. Done under time pressure, it gets skipped and the role stays open longer. We run a candidate-sourcing agent on TaskStation that does the search and the first draft every day. For an open role, it finds matching candidates on LinkedIn, checks Greenhouse so it never resurfaces someone already in the pipeline, and drafts a personalized outreach email referencing each candidate's actual background. It sources and drafts only — the recruiter reviews every draft, and nothing sends and no one is added to the ATS without them. This is how we source our own outbound pipeline. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Greenhouse · LinkedIn · Email - **Mode:** Sources and drafts only · recruiter reviews and sends ## The problem Outbound recruiting is a research problem before it's a writing problem, and neither half scales by hand. Searching LinkedIn for candidates who actually match a role's real bar takes real time, and doing it well across several open roles at once is more than one recruiter can sustain daily. Then there's the dedupe check: has this person already applied, already been contacted, already passed on an earlier round? Skipping that check means a candidate gets the same cold email twice, which reads as sloppy at best. The usual shortcuts make it worse. A generic sourcing tool exports a list of profiles with no reference to the pipeline, so recruiters re-discover people who are already in Greenhouse. A templated outreach message goes out to everyone in the list with a name swapped in, and candidates who get a lot of these can tell instantly. Neither problem gets fixed by working faster — they get fixed by connecting the systems that already have the answer. ## What we built On TaskStation, a daily cron triggers an agent. It spawns a fresh session (a cloud sandbox) scoped to an open role, searches LinkedIn for candidates matching that role's profile, checks the role's Greenhouse pipeline so it never proposes someone already there, and drafts a personalized outreach email for each new match — referencing that candidate's real background, not a template. Every draft lands in the recruiter's inbox for review; the agent never emails a candidate and never touches Greenhouse. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox scoped to one open role, so every day's batch is sourced and deduped against the current state of the pipeline — nothing carries over from the day before. ### Give the agent the role's sourcing profile What "a match" means for this role — the title, seniority, skills, and any other criteria worth searching for — lives as **skills** and **memory** that travel with the agent, alongside the outreach angles that make a message worth opening instead of ignoring. When the role's bar changes or a message pattern gets better replies, we write it down and the next day's batch reflects it. ### Connect Greenhouse, LinkedIn, and email Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Search LinkedIn** — for candidates matching the role's sourcing profile. - **Check Greenhouse** — the open role's pipeline, so anyone already applied, contacted, or passed on gets skipped rather than re-surfaced. - **Draft to email** — a personalized outreach message per new candidate, referencing their actual background, held for the recruiter to send. ### Set the guardrails The agent's job stops at the draft. It never sends an outreach email itself and never adds, moves, or tags anyone in Greenhouse — sourcing and drafting only. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Hand the batch to the recruiter With that in place, each morning brings a small batch of new candidates for the open role, each with a drafted email that references their real background and a note on why they matched. The recruiter reads each one, edits where they want, and sends it themselves — or adds the candidate to Greenhouse themselves, if they choose to. Nothing is written back out except the drafts. > **The pattern** > A daily **cron** spawns a fresh session scoped to one open role, with > **connectors** into LinkedIn to source and Greenhouse to dedupe. The > sourcing profile and outreach angles live as **skills** and **memory**. The > agent drafts to email and stops — the recruiter reviews, sends, and owns > the ATS. ## Guardrails The agent can search for and contact real people, so the boundary between "draft" and "send" is the control that matters: - **Isolation.** Every run happens in its own isolated sandbox, scoped to one open role. Only the drafted emails are written back out. - **Scoped secrets.** The Greenhouse and LinkedIn credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** The agent never sends outreach itself and never adds a candidate to Greenhouse. Every draft is held for the recruiter to review, edit, and send — and the ATS is theirs to update. - **Dedupe before draft.** The agent checks the role's Greenhouse pipeline before drafting anything, so a candidate already applied, contacted, or passed on never gets a duplicate reach-out. - **Everything is code.** The sourcing profile, outreach angles, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** New candidates sourced and deduped per open role - **Zero:** Duplicate reach-outs to candidates already in the pipeline - **3 systems:** Greenhouse, LinkedIn, and email in one agent Sourcing for an open role now starts with a batch already searched, deduped against the pipeline, and drafted with a real reference to each candidate's background. The recruiter's day starts with review instead of research, and every send and every ATS change still stays with them. --- <!-- /markdown/use-cases/churn-risk.md --> # How we flag at-risk accounts A daily cron scores accounts on leading churn signals across product usage, support, and billing, then posts a ranked at-risk list to Slack with a reason and a suggested next step for each. Canonical page: https://taskstation.co/use-cases/churn-risk Churn is usually visible before it happens. Usage tapers off, support threads get more frequent and more frustrated, a payment fails, a renewal approaches. The signals are there, but they sit in different systems, and no one is watching all of them at once. By the time an account cancels, the warning signs had been accumulating for weeks. We run a churn-risk agent on TaskStation that reads those signals every day and posts a ranked at-risk list to our customer-success Slack channel. It only reads customer data; the single output is the Slack post. This is how we watch our own accounts. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Postgres · Plain · Stripe · Slack - **Mode:** Read-only · one Slack post per day ## The problem The signals that predict churn live in separate systems: product usage in Postgres, support friction in Plain, payment health in Stripe, renewal dates in billing. Each one is a partial view. An account with declining usage might be fine; an account with declining usage, a rising support load, and a renewal next month is not. The common approaches don't combine them. A usage dashboard shows one signal and leaves the reader to correlate the rest. A health score baked into one tool only sees that tool's data. Manual account reviews are thorough but happen quarterly, long after the signals first appeared, and depend on someone remembering to look. ## What we built On TaskStation, a daily cron triggers an agent. It spawns an isolated session (a cloud sandbox) with read-only access to Postgres, Plain, and Stripe, scores every account on leading churn signals — declining usage, rising support friction, missed or failed payments, an upcoming renewal — and posts a ranked at-risk list to the customer-success Slack channel, with the reason for each account and a suggested next step. It writes nothing back to customer systems. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox. One day maps to one run on one disposable machine, so the score is recomputed from the current state every time and nothing carries over. ### Give the agent the scoring rules How we weigh the signals lives as **skills** and **memory** that travel with the agent: what counts as a usage decline, which support patterns matter, how a failed payment and an upcoming renewal combine, and what a good next step looks like for each kind of risk. When we learn which signals actually preceded a churn, we write it down and the scoring improves. ### Connect the signal sources read-only Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent reads: - **Product usage from Postgres** — activity trends per account, to catch a decline before it bottoms out. - **Support signals from Plain** — thread volume and tone, to catch rising friction. - **Billing from Stripe** — missed or failed payments and the upcoming renewal date. - **Posts to Slack** — the ranked at-risk list, with a reason and a suggested next step per account. ### Set the guardrails The agent is **read-only** across every customer system. It has no write access to Postgres, Plain, or Stripe, so it cannot change an account, a ticket, or a subscription. Its only output is the Slack post. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Post the ranked list With that in place, each morning brings one Slack post: accounts ranked by risk, each with the signals that put it there — the usage drop, the support thread, the failed payment, the renewal date — and a suggested next step. The customer-success team reads it and decides what to do. Nothing is written back automatically. > **The pattern** > A daily **cron** spawns a session with read-only **connectors** into Postgres, > Plain, and Stripe. The scoring lives as **skills** and **memory**. The agent > reads everything and writes nothing but the Slack post. ## Guardrails The agent reads across every customer system, so its access is scoped and one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to the systems it's scoped to, and only the Slack post is written back out. - **Scoped secrets.** The Postgres, Plain, and Stripe credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only.** The connectors into customer data are read-only. The agent cannot change an account, a ticket, or a subscription; it can only report. - **Everything is code.** The agent's scoring rules, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Accounts rescored on the current state of the data - **Read-only:** Nothing written back to any customer system - **4 signals:** Usage, support, billing, and renewal in one score Churn signals that used to sit in four separate systems now arrive as one ranked list in the channel where the team already works, each account carrying the reason it surfaced and a suggested next step. The agent only reads; the people decide what to do about each account. --- <!-- /markdown/use-cases/cloud-cost-anomaly.md --> # How we catch cloud cost spikes before the bill lands The cloud-cost anomaly agent we run on TaskStation — connected to AWS Cost Explorer and Slack. Every day it keeps a running spend baseline per service and account, flags whatever breaks out of that baseline, attributes the likely driver, and alerts with the delta. Read-only and alert-only; it never touches a resource or a budget. Canonical page: https://taskstation.co/use-cases/cloud-cost-anomaly Cloud bills surprise people because nobody is watching spend as it accrues. Cost Explorer shows the truth eventually, but by default a spike surfaces at the end of the month, in the invoice, long after the resource that caused it has been running for weeks. A new instance type left on overnight, a service that started scaling past its usual ceiling, a region nobody meant to deploy to — each one is a small decision that turns into a line item nobody recognizes. We run a cloud-cost anomaly agent on TaskStation that watches AWS spend every day, one persistent session that remembers what "normal" looks like per service and per account. It only reads billing data; the single output is a Slack alert with the delta and its best guess at what caused it. This is how we catch a cost problem while it's still one day old, not one invoice old. - **Team:** TaskStation - **Runs on:** Daily cron, reusable session - **Connected systems:** AWS Cost Explorer · Slack - **Mode:** Read-only · alert only, never modifies a resource or a budget ## The problem Cloud spend is easy to see and hard to watch. Cost Explorer will answer "how much did we spend" for any window you ask about, but it won't tell you, unprompted, that yesterday's spend on one service was 40% above its normal range. Nobody opens the console every morning to eyeball forty line items across a dozen accounts, so a spike sits there accruing until someone notices the invoice. The common approaches don't close the gap. A monthly budget alert fires only after the month's total crosses a threshold, by which point the overspend has already happened many times over. A flat per-service alert threshold treats a service that normally costs $50/day the same as one that normally costs $5,000/day, so it's either too noisy or too blind. And none of it tells you *why* — a dashboard shows the number moved, not what moved it. ## What we built On TaskStation, a daily cron re-prompts one persistent agent session with read-only access to AWS Cost Explorer. It pulls the prior day's spend broken out by service and by linked account, updates its running baseline for each, and flags anything that breaks out of its own normal range — not a flat threshold, but a deviation from what that specific service in that specific account usually costs. For each anomaly it works out the likely driver — a newly launched resource, a traffic surge, a region the spend wasn't previously in — and posts one alert to Slack with the delta and the suspected cause. It writes nothing back to AWS: no resource is touched, no budget is changed. ## How it works ### Run on a daily cron, one persistent session A **cron trigger** fires the agent once a day, but unlike a fresh-session scan, this is **one session re-prompted daily** (`session_mode: reuse`). It resumes its own memory of what each service and account normally spends instead of starting blind every morning, so the baseline gets sharper the longer the agent runs. ### Give the agent the anomaly rules What counts as a break from baseline, how much history to weigh a service's "normal" against, and how to read a spike's shape — sudden versus ramping, one account versus many — live as **skills** and **memory** that travel with the agent. When we learn a spike was actually a planned load test or a known seasonal pattern, we write it down and the agent stops flagging it. ### Connect AWS Cost Explorer read-only Through a scoped **connector**, brokered server-side so no raw credential reaches the model, the agent reads: - **Daily cost and usage by service and linked account** — the raw numbers the baseline and the anomaly check run against. - **Resource-level detail for the flagged period** — what actually launched, scaled, or shifted region, to support the driver attribution. - **Posts to Slack** — the delta, the suspected driver, and nothing else. ### Attribute the likely driver, not just the delta A number moving is not an explanation. For every anomaly, the agent checks what changed underneath it — a resource that came online in the window, usage metrics consistent with a traffic surge, spend appearing in a region that previously had none — and states its best-guess driver alongside the delta, so the alert is something a human can act on immediately. ### Set the guardrails The agent is **read-only** across AWS billing and Cost Explorer. It has no permission to modify or delete a resource, or to change a budget or spending control — it can only observe and alert. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Alert with the delta and the suspected cause With that in place, each day brings at most one Slack alert per anomaly: the service and account, the size of the deviation from baseline, and the suspected driver. No anomaly means no message. The engineering team reads it and decides whether to act — nothing changes in AWS on its own. > **The pattern** > A daily **cron** re-prompts one persistent **session** with a read-only > **connector** into AWS Cost Explorer. The baseline lives as **skills** and > **memory** that carry forward run to run. The agent reads everything and > writes nothing but the Slack alert. ## Guardrails The agent reads across every linked AWS account's billing data, so its access is scoped and strictly one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to AWS Cost Explorer, and only the Slack alert is written back out. - **Scoped secrets.** The AWS connector credential is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Read-only, no exceptions.** The AWS connector is read-only. The agent cannot launch, modify, or delete a resource, and it cannot change a budget or any spending control — it can only alert. - **Alert only.** The Slack post is the only output. No remediation, no auto-scaling change, no resource shutdown — a human decides what, if anything, to do about a spike. - **Everything is code.** The agent's baseline logic, thresholds, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Spend baselined per service and account against its own history - **Read-only:** Nothing modified, launched, or deleted in AWS, ever - **1 alert:** Delta + suspected driver, posted only when spend breaks the baseline A cost problem that used to surface as a surprising line item at the end of the month now surfaces the next morning, with the service, the account, the size of the spike, and a suspected cause already attached. The agent only reads and alerts; the engineering team decides what to do about each one. --- <!-- /markdown/use-cases/competitor-watch.md --> # How we monitor competitors A daily agent that checks competitor sites, changelogs, and pricing pages, diffs them against the last run, and posts a short summary of what changed to Slack. Canonical page: https://taskstation.co/use-cases/competitor-watch Competitors ship changes quietly. A pricing tier moves, a feature lands in the changelog, a landing page gets rewritten — and unless someone happens to check that week, the team finds out late. Manually visiting a dozen sites every morning is the kind of task that gets done for a while and then quietly stops. We watch our market with an agent that runs every day on TaskStation. It checks competitor sites, changelogs, and pricing pages, diffs them against the last run, and posts a short summary of what actually changed to a Slack channel. - **Team:** TaskStation - **Control surface:** Slack - **Connected systems:** The web · Slack - **Mode:** Daily cron · diff against last run ## The problem Keeping up with competitors is a standing task with no natural owner. The information is public, but it's spread across pages that change on no schedule, and most days nothing moves — so a person checking manually spends most of their time confirming that nothing happened. The workarounds thin out. A page-change alerting tool fires on every edit, including cosmetic ones, and buries the real signal in noise. A shared doc of "things to check" depends on someone remembering to check it. The task needs to run daily, compare against what it saw last time, and only surface what's worth reading. ## What we built A cron trigger runs an agent every day. Each run spawns an isolated session that fetches the competitor pages we track, compares them against the previous run's snapshot, and writes a short summary of what changed to Slack. Cosmetic edits are filtered out; pricing moves, new features, and messaging changes are called out. ## How it works ### Run it on a daily cron A **cron trigger** fires the project once a day. Each firing spawns a fresh **session** in its own isolated sandbox with web access. One run, one sandbox, torn down when it's done. ### Give the agent the watch list and what matters The list of competitors, the specific pages to check, and what counts as a meaningful change live as **skills** and **memory**. Each run also reads the previous run's snapshot from memory so it can diff against it rather than re-reporting the same state. The list is updated as the market shifts. ### Connect the web and Slack Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Fetch competitor pages** — sites, changelogs, and pricing pages on the watch list. - **Compare against the last run** — diff today's content against the stored snapshot to find what actually changed. - **Post to Slack** — a short summary in the growth channel, or nothing when nothing moved. ### Set the guardrails The agent only reads public pages and writes to one Slack channel, so its scope is narrow by design. It has no access to internal systems and takes no action beyond posting. Any credentials it needs are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each run report A run fetches the tracked pages, diffs them against yesterday, and posts what changed: a pricing tier that moved, a feature that shipped, a rewritten homepage. On a quiet day it says so briefly. The team reads one message instead of visiting a dozen sites. > **The pattern** > Run the check on a **cron trigger**, keep the watch list and last-run snapshot in > **skills** and **memory**, give the agent scoped **connectors** to the web and > Slack, and let each run diff and report. The team reads a summary instead of > browsing. ## Guardrails Even a read-only agent runs with the same controls as the rest of the platform: - **Isolation.** Each run executes in its own isolated sandbox, and only the Slack summary it's explicitly allowed to send is written back out. - **Scoped secrets.** Any credential the agent uses is encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Human approval gates.** The agent's only external action is posting a summary; anything beyond that scope would require a person to approve. - **Everything is code.** The watch list, the pages to check, and the definition of a meaningful change are files in the repo — versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Daily:** The market is checked without anyone remembering to - **Diff-only:** Cosmetic edits filtered out, real changes surfaced - **One message:** A summary instead of a dozen tabs The team learns about a competitor's pricing move or feature launch the morning after it happens rather than whenever someone next looks. Because each run diffs against the last, the channel stays quiet on quiet days and speaks up only when something actually changed. --- <!-- /markdown/use-cases/compliance-monitoring.md --> # How we monitor for compliance drift A daily sweep checks resources against policy for public buckets, untagged resources, and over-broad roles, files findings, and proposes remediation as a reviewed change rather than applying it. Canonical page: https://taskstation.co/use-cases/compliance-monitoring Infrastructure drifts out of compliance quietly. A bucket gets made public for a one-off, a resource ships without its tags, a role picks up a permission it no longer needs. Each change is small and reasonable in the moment, but they accumulate, and nobody notices until an audit or an incident surfaces them all at once. By then reconstructing when each drift happened is hard. We handle this by checking the infrastructure against policy every day, so drift surfaces the day it appears. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails — so you can set up the same sweep. - **Team:** TaskStation - **Trigger:** A daily scheduled sweep - **Connected systems:** AWS · Audit logs · Slack - **Mode:** Cron-driven · Review-gated ## The problem The common approaches each have limits. A quarterly manual audit finds drift late and in bulk, when the context is hardest to recover. A fixed rules engine catches the checks it was configured for and nothing beyond them. And auto-remediation that fixes drift on its own can break a service that depended on the very configuration it "corrected." We wanted the infrastructure checked against policy daily, with findings filed where the team already works and fixes proposed as reviewed changes rather than applied silently. ## What we built On TaskStation, a daily schedule triggers an agent. Each sweep runs in its own isolated session — a cloud sandbox — with scoped, read access to AWS and the audit logs. The agent checks resources against policy — public buckets, untagged resources, over-broad roles — and files what it finds to Slack. Remediation is proposed as a reviewed change, never applied automatically. ## How it works ### Connect the schedule as the trigger A **cron** trigger runs the sweep once a day. Each run spawns a fresh **session** in its own sandbox. One sweep, one session, one disposable machine. Nothing carries over between runs, so each day's check starts clean. ### Give the agent the compliance policy What counts as a violation — which buckets may be public, which tags are required, what a role should and shouldn't hold — is stored as **skills** and **memory** that load into every session. The agent checks against that policy rather than a fixed rule set, and it updates as the policy changes. ### Connect what the sweep can read Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read AWS resource state** — bucket policies, resource tags, and IAM roles. - **Read the audit logs** — when a configuration changed and what changed it. - **Post to Slack** — findings filed to the channel the team watches. ### Set the guardrails The sweep's access to AWS is **read-only** — it inspects state, it doesn't change it. Remediation is proposed as a reviewed **change request** and stops at a **human approval gate**; nothing is applied automatically. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each day surface the drift With that in place, the daily sweep checks the infrastructure against policy, files what drifted to Slack, and attaches a proposed fix for a person to review. A newly public bucket becomes a finding and a draft policy change. An untagged resource becomes a flag and a proposed tag set. An over-broad role becomes a narrower policy held for review. > **The pattern** > Run the sweep on a **cron** trigger, give the agent scoped read **connectors** > into AWS and the audit logs, encode the compliance policy as **skills** and > **memory**, and propose every fix as a reviewed **change request** rather than > applying it. ## Guardrails Giving an agent a standing view of the infrastructure is a trust question. The relevant controls on TaskStation: - **Isolation.** Every sweep runs in its own isolated sandbox on its own branch. The session is granted access only to what it's scoped to read, and only the findings it files are written back out. - **Scoped secrets.** Each credential is encrypted in the secrets manager, injected into the sandbox at runtime, and scoped to the agents you grant them to or the logs. - **Human approval gate.** Remediation is proposed, never auto-applied; a person reviews and applies each change. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Daily:** Infrastructure re-checked against policy every day - **Same-day:** Drift caught before it accumulates for an audit - **Review-gated:** Every fix proposed as a change, never auto-applied The drift that used to surface in bulk at audit time now arrives as small daily findings in Slack, each with a proposed fix a person reviews before it's applied. The team reviews a change instead of reconstructing months of drift, and the infrastructure stays close to policy. --- <!-- /markdown/use-cases/content-refresh.md --> # How we keep our content from going stale The content-refresh agent we run on TaskStation — a weekly agent that finds decaying marketing and blog pages via Search Console, refreshes the copy, stats, and internal links, and opens a PR for review. Canonical page: https://taskstation.co/use-cases/content-refresh Most published content doesn't die all at once. A comparison page that used to rank on page one slides to page two as a competitor updates their pricing. A blog post cites a stat from two years ago and a reader notices. A guide links to a feature we renamed six months back. None of it is broken enough to trigger an alert, so it just sits there, quietly losing impressions, until someone happens to look. We run a content-refresh agent on TaskStation that watches for that decay every week and does something about it. It reads Search Console for pages losing impressions and clicks, cross-references them against our content repo, refreshes the copy, the stats, and the internal links, and opens a PR. It never publishes anything itself — a human reviews and merges. - **Team:** TaskStation - **Runs on:** Weekly cron - **Connected systems:** Google Search Console · GitHub - **Mode:** Reusable session · PR only, never publishes ## The problem Nobody schedules time to re-read a two-year-old blog post. Content gets written once, published, and then left alone unless it's actively broken. Meanwhile the world underneath it keeps moving: prices change, product names change, competitors update their own pages, and search intent drifts. Search Console sees all of that as a slow decline in impressions and clicks per page, but nobody is watching that report on a schedule, and even when someone is, "this page is declining" doesn't tell you what to fix. Doing it manually doesn't scale either. A content team can audit a handful of pages a quarter if they're disciplined about it, but a site with hundreds of pages generates decay faster than any manual process can keep up with, and the same handful of high-traffic pages get all the attention while the long tail quietly rots. ## What we built On TaskStation, a weekly cron re-prompts a **persistent session** — the same agent, picking up where last week left off. It pulls the pages losing the most impressions and clicks from Search Console, cross-references them against our content repo, and rotates through them so the same three pages don't get all the attention while the rest decay untouched. For each page it picks that week, it refreshes the copy that's gone stale, updates any numbers or stats that have aged out, fixes internal links that point at renamed or superseded pages, and opens one PR per run. It never pushes to the live branch and never merges its own work. ## How it works ### Run on a weekly reusable session A **cron trigger** fires the agent once a week, and it re-prompts a **reusable session** rather than starting fresh. That matters here: refreshing every decaying page at once would be disruptive, so the agent works through a rotation, and the only way to rotate fairly is to remember what it already touched. ### Find what's decaying Through a scoped **connector**, the agent reads Google Search Console for the site: impressions and clicks per page over the trailing window compared to the window before it. Pages with the sharpest decline, weighted by how much traffic they still carry, become candidates. ### Cross-reference the content repo and rotate coverage The agent checks its **ledger** — which pages it refreshed and when — before picking this week's targets. A page that was refreshed three weeks ago drops in priority even if it's still declining; a page that's never been touched moves up. This is what keeps the rotation from fixating on the same few high-traffic pages every week. ### Refresh the copy, stats, and internal links For each page in this week's batch, the agent opens the source in **GitHub**, updates copy that reads as dated, refreshes numbers and stats that have gone stale, and repoints internal links that reference renamed or retired pages. It works on an isolated branch, never on the live one. ### Open a PR and stop The agent opens one PR per run summarizing what changed on which pages and why, using the **`gh` CLI** authenticated with a scoped `GH_TOKEN`. It never publishes and never merges. A human reviews the diff and decides what ships. > **The pattern** > A weekly cron re-prompts a **reusable session** that reads decay signals > from Search Console, rotates coverage using its own **ledger**, and edits > content on an isolated branch. The only output is a PR — never a push to the > live branch. ## Guardrails The agent edits marketing and blog copy, so what it can do without a human is tightly bounded: - **PR only, never publish.** The agent opens a pull request and stops. It never pushes to the live branch and never merges its own work. - **Isolation.** Every run happens in its own sandbox, on its own branch. The live content branch is never touched directly. - **Scoped secrets.** The `GH_TOKEN` used by the `gh` CLI is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only search data.** Search Console access is read-only; the agent can see decay signals but cannot change anything in Search Console itself. - **Everything is code.** The rotation ledger, the refresh rules, and the agent's permissions are files in the repo, changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every week:** A new batch of decaying pages picked up and refreshed - **PR only:** Nothing published without a human review - **0 repeats:** Ledger-driven rotation so coverage doesn't fixate on the same pages Pages that used to decay silently for months now get caught within weeks of losing traction, refreshed on a rotation that reaches the whole site instead of just the pages someone happened to remember. The agent finds the decay and does the editing; the team still decides what goes live. --- <!-- /markdown/use-cases/contract-review.md --> # How we do first-pass contract review A new contract in a Drive folder triggers an agent to summarize it, flag non-standard clauses against our playbook, and post the summary to the legal channel, drafting redlines for a lawyer to sign off. Canonical page: https://taskstation.co/use-cases/contract-review Contracts arrive faster than they can be read closely. A vendor agreement, an NDA, or an order form lands in a shared Drive folder, and someone on the legal team has to open it, read it against what we consider standard, flag the clauses that deviate, and summarize it for whoever's driving the deal. The first pass is mechanical but time-consuming, and it's the step that stands between a contract arriving and a lawyer being able to focus on what actually matters in it. We handle this by tying the first pass to the event that starts it: a new file in the folder. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails — so you can set up the same review. - **Team:** TaskStation - **Trigger:** A new contract in a Drive folder - **Connected systems:** Google Drive · Slack - **Mode:** Trigger-driven · Lawyer-gated ## The problem The common approaches each fall short. A shared inbox where contracts pile up means the first pass waits on whoever has time. A generic AI summarizer produces a readout that doesn't know our positions, so it flags nothing that matters to us. And skipping the first pass entirely puts a lawyer straight into a full read of every contract, including the routine ones. We wanted each contract summarized and checked against our own playbook the moment it lands, with the summary where the team works and the redlines left for a lawyer to sign off. ## What we built On TaskStation, a new contract in a Drive folder triggers an agent. Each contract runs in its own isolated session — a cloud sandbox — with scoped access to Drive and Slack. The agent reads the contract, summarizes it, flags clauses that deviate from our playbook, and posts the summary to the legal channel. It drafts redlines on the non-standard clauses, but a lawyer signs off before anything goes back to the counterparty. ## How it works ### Connect the Drive folder as the trigger A signed webhook watches the contracts folder in Drive. A new file fires it, and each firing spawns a fresh **session** in its own sandbox, seeded with the document. One contract, one session, one disposable machine. Nothing carries over between runs. ### Give the agent our contract playbook Our standard positions — which clauses are routine, which terms we push back on, what an acceptable liability cap or notice period looks like — are stored as **skills** and **memory** that load into every session. The agent checks each contract against that playbook rather than a generic notion of "standard," and it updates as our positions change. ### Connect what a review can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the contract in Drive** — the full document, not just an excerpt. - **Draft redlines** — proposed edits on the clauses that deviate from the playbook. - **Post to Slack** — the summary and flags filed to the legal channel. ### Set the guardrails The agent produces a first pass, not a decision. Redlines are drafted, never sent — every one stops at a **human approval gate** for a lawyer to sign off before it reaches the counterparty. It reads the contract and writes its summary and drafts; it doesn't act on the deal. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each contract come back reviewed With that in place, a contract landing in the folder comes back as a summary in the legal channel, the non-standard clauses flagged against our playbook, and draft redlines attached for a lawyer to review. A vendor agreement becomes a readout and a set of proposed edits; the lawyer starts from the flags instead of a blank read. > **The pattern** > Connect the Drive folder via a **trigger**, give the agent scoped > **connectors** into Drive and Slack, encode the contract playbook as **skills** > and **memory**, and gate every redline behind a lawyer. ## Guardrails Giving an agent a first pass on contracts is a trust question. The relevant controls on TaskStation: - **Isolation.** Every contract runs in its own isolated sandbox on its own branch. The session is granted access only to what it's scoped to, and only the summary and drafts it produces are written back out. - **Scoped secrets.** Each credential is encrypted in the secrets manager, injected into the sandbox at runtime, and scoped to the agents you grant them to or the logs. - **Human approval gate.** Redlines are drafted, never sent; a lawyer reviews and signs off before anything reaches the counterparty. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Every contract:** Summarized and checked against the playbook on arrival - **Lawyer-gated:** Redlines drafted, never sent without sign-off - **2 systems:** Drive and Slack — one agent The first pass that used to wait on whoever had time now arrives as a summary and a set of playbook-checked flags the moment a contract lands, with draft redlines attached. The lawyer starts from the flags instead of a cold read, and routine contracts stop consuming the time the tricky ones deserve. --- <!-- /markdown/use-cases/crm-hygiene.md --> # How we keep our CRM clean A nightly agent that dedupes contacts in HubSpot, fills missing fields from enrichment data, flags stale deals, and posts a data-quality summary — with bulk updates held for approval. Canonical page: https://taskstation.co/use-cases/crm-hygiene A CRM decays on its own. Duplicate contacts pile up from form fills and imports, records come in with missing fields, and deals sit untouched long after they've gone cold. Left alone, the pipeline reports drift away from reality and the sales team stops trusting the numbers. We keep our HubSpot clean with an agent that runs every night on TaskStation. It dedupes contacts, fills missing fields from enrichment data, flags stale deals, and posts a short data-quality summary. Bulk field updates over a threshold wait for a person to approve. - **Team:** TaskStation - **CRM:** HubSpot - **Connected systems:** HubSpot · Enrichment data · Slack - **Mode:** Nightly cron · human-gated ## The problem CRM hygiene is the work nobody schedules. Duplicate contacts split a company's history across two records, so the account owner sees half the story. Missing job titles, industries, and company sizes make segmentation unreliable. Deals that haven't moved in weeks still count toward the forecast. The usual fixes don't hold. A one-time cleanup helps until the next import undoes it. A paid dedupe add-on handles duplicates but not enrichment or stale deals. Asking reps to keep their own records tidy competes with selling and loses. The maintenance needs to run on its own, every night, without a person babysitting it. ## What we built A cron trigger runs an agent against HubSpot every night. Each run spawns its own isolated session with scoped access to HubSpot and our enrichment provider. It finds and merges duplicate contacts, fills missing fields from enrichment, flags deals that have gone stale, and posts a summary of what it changed. Any bulk field update above a set threshold stops for a person before it writes. ## How it works ### Run it on a nightly cron A **cron trigger** fires the project once a night. Each firing spawns a fresh **session** in its own isolated sandbox with the CRM and enrichment access a cleanup pass needs. One run, one sandbox, torn down when it finishes. ### Give the agent the hygiene rules Our data model lives as **skills** and **memory** loaded into every run: which fields are required, how we decide two contacts are the same record, what counts as a stale deal, and the merge rules to follow. The rules are updated as edge cases come up. ### Connect HubSpot and enrichment Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read and write HubSpot** — pull contacts and deals, merge duplicates, and update fields. - **Query the enrichment provider** — fill missing job titles, company sizes, and industries from external data. - **Post to Slack** — a short summary of every run in the RevOps channel. ### Set the guardrails The agent merges duplicates and fills individual gaps on its own, but a **bulk field update over a threshold** stops at a **human approval gate** before it writes. A sweeping change across hundreds of records is reviewed first. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each run report A run dedupes contacts, enriches what's missing, flags the deals that have gone quiet, and posts what it did: duplicates merged, fields filled, deals flagged. When a bulk change is pending, the summary says what's waiting and why. > **The pattern** > Run the cleanup on a **cron trigger**, give the agent scoped **connectors** into > HubSpot and enrichment, encode the hygiene rules as **skills** and **memory**, > and gate bulk changes behind a human. Each night's run then cleans and reports on > its own. ## Guardrails Giving an agent write access to the CRM is a data-integrity question. The controls on TaskStation: - **Isolation.** Each run executes in its own isolated sandbox, and only the HubSpot and enrichment changes it's explicitly allowed to make are written back out. - **Scoped secrets.** The HubSpot and enrichment credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gates.** Bulk field updates over the threshold require a person to approve before they write. - **Everything is code.** The dedupe rules, required fields, and stale-deal definition are files in the repo — versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Nightly:** Cleanup runs without anyone scheduling it - **3 passes:** Dedupe, enrichment, and stale-deal flags in one run - **Threshold-gated:** Large bulk changes reviewed before they write The pipeline reports stay closer to reality because the records behind them are maintained every night instead of in occasional cleanups. When a change is large enough to matter, a person sees it first, and the RevOps channel has a running record of what changed and when. --- <!-- /markdown/use-cases/customer-onboarding.md --> # How we drive new customers to activation The customer onboarding agent we run on TaskStation — a daily cron that checks every new account against its onboarding milestones in Postgres, drafts a nudge email when one stalls, and alerts the owning CSM in Slack with the context to act. Canonical page: https://taskstation.co/use-cases/customer-onboarding The first thirty days decide most renewals long before the renewal date shows up. An account that gets its workspace live, finishes setup, and hits its first real result in week one tends to stick. An account that signs and then goes quiet — no second login, no setup finished, no teammate invited — is already at risk, and usually nobody notices until the CSM happens to open HubSpot and wonder why an account they haven't heard from in three weeks is still on "customer since" last month. We run a customer-onboarding agent on TaskStation that checks every new account against its onboarding milestones every day, catches the ones that are stalling, drafts a nudge for the customer, and alerts the owning CSM with exactly what's stuck and why. It never changes an account and never sends anything itself — it hands the CSM a head start instead of a surprise churn conversation two quarters later. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Postgres · HubSpot · Slack · Email - **Mode:** Draft + alert only · CSM sends ## The problem A new account's onboarding progress lives in two places that don't talk to each other: the product events and activation milestones sit in Postgres, and the account record — who owns it, when it became a customer — sits in HubSpot. Neither system on its own tells a CSM whether an account is on track. A login count means nothing without knowing what day of onboarding the account is on; a HubSpot record means nothing without knowing whether the account has actually done anything since it signed. The usual fallback is a CSM manually checking their book of accounts, or a dashboard that shows raw usage numbers with no sense of what "on track" even means for a given account's age. By the time a stalled account gets noticed, it's often past the point where a simple nudge would have fixed it. ## What we built On TaskStation, a daily cron triggers an agent that reads HubSpot for every account inside its onboarding window and owned by a CSM, checks that account's activation milestones and product events in Postgres, and classifies each one as on track, stalled at a specific milestone, or overdue for full activation. For every stalled or overdue account, it drafts a nudge email addressed to the customer's own missed step and posts an alert to the owning CSM's Slack with the specific context — what's done, what's stuck, and what the nudge says. The draft waits in the CSM's inbox; the CSM decides whether to send it, edit it, or reach out a different way. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox with no memory of yesterday's run — the account list and every milestone are re-read from HubSpot and Postgres each time, and a marker property on the HubSpot record tracks which milestone was last nudged so the same stall doesn't get a duplicate email every day it persists. ### Give the agent the onboarding playbook The milestone framework — what counts as workspace-live, core setup, first activation event, team expansion, and full activation, and by what day each is expected — lives as a **skill** that travels with the agent, along with the stall/overdue thresholds and the tone for a nudge. When we learn which milestones actually predict a healthy account, we update the skill and every account benefits the next day. ### Connect Postgres and HubSpot Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent: - **Reads accounts from HubSpot** — which ones are inside the onboarding window, and who the owning CSM is. - **Reads activation milestones and product events from Postgres** — read-only, to see exactly what each account has and hasn't done since it signed. - **Writes back to HubSpot** only which milestone a nudge was last drafted for — never a plan, a seat count, or any other account setting. ### Post to Slack, draft to email Two channels, two purposes: a **Slack** post to the owning CSM with the account's onboarding status and the specific stall, and an **email** draft — the nudge to the customer — held for the CSM to open, edit, and send. ### Set the guardrails The agent tracks progress and drafts outreach; it never changes anything on the account itself and never contacts the customer directly. It cannot send the nudge email, and it cannot touch a plan, a seat, a billing setting, or any other account state. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Hand off to the CSM With that in place, a CSM opens Slack to a specific reason an account needs attention today — not a generic "check in on your accounts" reminder — and opens their inbox to a nudge already drafted around the exact step the customer hasn't taken. They read it, decide if it fits, and send it themselves. > **The pattern** > A daily **cron** spawns a session with read connectors into Postgres and > HubSpot. The milestone framework lives as a **skill**. The agent flags the > stall, drafts the nudge, and alerts the CSM; the CSM reviews, edits, and > sends. ## Guardrails The agent touches new, revenue-relevant accounts, so what it can do is narrow and explicit: - **Draft and alert only.** The agent never sends the nudge email itself and never contacts the customer directly. The email is a draft in the CSM's inbox until the CSM sends it. - **No account changes, ever.** The agent never touches a plan, a seat count, a billing setting, or any other account state — in HubSpot or anywhere else. It reports and drafts; it does not administer. - **Scoped writes.** The only thing the agent writes back to HubSpot is which milestone a nudge was last drafted for — never the deal stage, the owner, or any account property a CSM relies on. - **Isolation.** Every run happens in its own isolated sandbox. Credentials for Postgres and HubSpot are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. - **Everything is code.** The milestone framework, the stall thresholds, and the per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** New accounts rechecked against live milestone data - **0 sent:** Emails sent or account settings changed automatically - **5 milestones:** Tracked from first login to full activation Onboarding stalls that used to surface only when a CSM happened to look now arrive the same day they start, with the exact missed step and a nudge already written. The agent watches every new account and does the noticing; the CSM still decides what to send and when. --- <!-- /markdown/use-cases/customer-support.md --> # How we run customer support with an AI agent The support agent we run on TaskStation — connected to Plain, our codebase, and Stripe. It triages and resolves inbound threads, and stops for human approval on anything sensitive. Canonical page: https://taskstation.co/use-cases/customer-support We run our own customer support on TaskStation with an AI agent connected to Plain (our support tool), our codebase, and Stripe. Most support questions need product knowledge, engineering context, and billing data at the same time: a help-center chatbot can quote documentation but can't read the code behind a bug, confirm whether a subscription renewed, or take a scoped action on an account. This write-up covers how the setup works: the connections, the session model, and the guardrails. - **Team:** TaskStation — the company behind this platform - **Support stack:** Plain.com - **Connected systems:** Codebase · Stripe · Plain - **Mode:** Trigger-driven · human-gated ## The problem Support volume grows faster than the team. The questions that matter most are the ones a canned macro can't answer: "why was I charged twice?", "this button does nothing", "does your API support X?" Each one needs someone who understands the product, can read the code, and can check the customer's account — usually a senior engineer. A help-desk bot handles the easy questions but not these. Hiring scales linearly with tickets. And giving a generic AI assistant real access to production systems is hard to justify in a security review. ## What we built Our support inbox in **Plain** is connected to an agent running on TaskStation. Every inbound thread spawns its own isolated session with scoped access to the systems a support issue can touch: the product's knowledge, the codebase, and Stripe. It investigates, resolves what it can, and stops at a human for sensitive actions. ## How it works ### Connect Plain as the trigger A signed webhook from Plain points at the project. Every new thread or customer reply fires it, and each firing spawns a fresh **session** in its own isolated sandbox. One customer, one thread, one sandbox. Sessions don't share state, and a busy inbox means more sessions running in parallel. ### Give the agent the product's knowledge Product knowledge lives as **skills** and **memory** loaded into every session: how the product works, our support playbook, the reply tone, and resolutions that worked before. The knowledge is updated as threads are resolved. ### Connect the systems an issue can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Search the codebase** — for a reported bug, find the relevant code path, check recent changes, and see whether it's already known. - **Look up Stripe** — plan, invoices, and subscription state, to answer billing questions from account data. - **Read and reply in Plain** — full thread context in, a reply out. ### Set the guardrails The agent is **read-mostly by default**: it investigates freely, but writes are scoped. Refunds, plan changes, and anything touching a customer's money or account stop at a **human approval gate**. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Let each thread run An inbound thread triages, investigates across the three systems, and either resolves directly or hands off to a human with the work done and context attached. "Why was I charged twice?" becomes a Stripe lookup with an answer. "This button does nothing" becomes a codebase search that either explains the behavior or files a bug report for engineering. > **Summary** > Connect the support platform via a **trigger**, give the agent scoped > **connectors** into the systems an issue can touch, encode product knowledge as > **skills** and **memory**, and gate sensitive actions behind a human. Each > inbound thread then runs in its own session. ## Guardrails Giving an agent access to the codebase, Stripe, and the support inbox is a security question as much as a product one. The controls on TaskStation: - **Isolation.** Each thread runs in its own isolated sandbox on its own branch. A session can install, run, and experiment to reproduce a bug, and only what it's explicitly allowed to send is written back out. - **Scoped secrets.** The Plain, Stripe, and GitHub credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gates.** Irreversible actions (money, account state) require a person to approve. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **24/7:** Coverage without a night shift - **First-touch:** Many inbound threads resolved before a human opens them - **3 systems:** Product knowledge, Stripe, and the codebase in one agent The agent handles the investigation-heavy tickets that used to pull a senior engineer off their work, and when it escalates, it escalates with the diagnosis attached. Customers get an answer to the question at any hour. The setup relies on four pieces: sandbox isolation to contain each session, a secrets manager to broker tokens, human approval gates for irreversible actions, and memory that improves as threads are resolved. --- <!-- /markdown/use-cases/data-pipeline-monitor.md --> # How we monitor our data pipelines and warehouse health The pipeline-health agent we run on TaskStation — connected to the warehouse and GitHub. Every hour it checks table freshness, row-count anomalies, and schema drift, alerts Slack, and drafts a GitHub issue with the likely cause. It never touches the data. Canonical page: https://taskstation.co/use-cases/data-pipeline-monitor A pipeline can fail quietly. A load job silently stops appending rows, a source schema changes upstream and half the fields come through null, a table that should refresh every hour hasn't moved since yesterday. None of that throws an error anyone sees — the dashboards built on top of it just start being wrong, and usually the first person to notice is whoever's report looks off in a stakeholder meeting. We run a pipeline-monitor agent on TaskStation that checks the warehouse itself, every hour: is every table as fresh as its SLA requires, does today's row count look like the rest of the trend, has the schema changed under anyone, did a load job fail. It only reads the warehouse. When something is wrong, it posts to the data team's Slack channel and drafts a GitHub issue with its best guess at the cause. This is how we watch our own warehouse. - **Team:** TaskStation - **Runs on:** Hourly cron - **Connected systems:** Postgres warehouse · Slack · GitHub - **Mode:** Read-only · alert + draft issue, nothing else ## The problem Warehouse health doesn't announce itself. A load job can fail without an exception if it's built to skip a bad batch and move on. A table can go stale because an upstream cron got disabled, not because anything crashed. A schema can drift because a source API added a field, and the load still "succeeds" — it just drops or nulls what it doesn't recognize. Every one of these looks fine from the outside until a downstream query returns something wrong. The common approaches don't hold up. Dashboards built on the data can't tell you the data itself is broken — they just render whatever's there, wrong or not. A "did the job run" check confirms the process exited zero, not that the table it wrote actually looks right. Someone eyeballing row counts in a spreadsheet catches it eventually, usually after a downstream report has already gone out wrong. And this is a different failure mode from an application throwing an error mid-request — the pipeline can succeed and the warehouse can still be unhealthy. ## What we built On TaskStation, an hourly cron triggers an agent. It spawns a fresh session with read-only access to the warehouse and checks every monitored table against its freshness SLA, compares today's row count to the table's own trailing baseline, diffs the current schema against the last-seen shape, and looks for loads that failed or silently stalled. When it finds something out of bounds, it posts an alert to the data team's Slack channel with the affected table, the anomaly, and the most likely cause, and drafts a GitHub issue with the same diagnosis so the incident is tracked. It never writes to the warehouse — the alert and the draft issue are the only outputs. ## How it works ### Run on an hourly cron A **cron trigger** fires the agent once an hour. Each firing spawns a fresh **session** in its own sandbox. One hour maps to one run on one disposable machine, so every check is recomputed from the warehouse's current state and nothing carries over between runs. ### Give the agent the health rules What counts as stale, what a normal row-count trend looks like, and how a schema is expected to shape up live as **skills** and **memory** that travel with the agent: per-table freshness SLAs, the tables that are expected to grow in bursts versus steadily, and past incidents worth checking new anomalies against. When we tighten an SLA or learn a table has a legitimate weekly gap, we write it down and the checks adjust. ### Connect the warehouse read-only Through a scoped **connector**, brokered server-side so no raw credential reaches the model, the agent reads the warehouse directly: - **Table freshness** — the most recent load timestamp per monitored table, checked against its SLA. - **Row counts** — today's count against the table's own trailing baseline, to catch a load that ran short or doubled up. - **Schema state** — the current column set and types, diffed against the last-seen shape, to catch drift before it breaks a downstream query. - **Load history** — recent job runs, to catch a failed or silently stalled load. ### Alert and draft, never fix When a check comes back out of bounds, the agent posts to **Slack** — the table, the anomaly, and its best guess at the cause — and drafts a **GitHub** issue with the same diagnosis, using a scoped token so the incident lands as a trackable draft, not a merged change. It does not touch the pipeline, the schema, or the data. ### Set the guardrails The agent is **read-only** across the entire warehouse. It has no write access to any table, schema, or pipeline configuration — it can only observe and report. Its outputs are exactly two: the Slack alert and the drafted GitHub issue. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Let the warehouse watch itself With that in place, every hour brings a clean pass over the warehouse: freshness checked against SLA, row counts checked against trend, schema checked against the last-seen shape, loads checked for failures. Anything out of bounds shows up in Slack within the hour, with a GitHub issue already drafted and the likely cause attached. The data team decides what to do; nothing changes on its own. > **The pattern** > An hourly **cron** spawns a session with a read-only **connector** into the > Postgres warehouse. The SLAs and baselines live as **skills** and **memory**. > The agent reads everything and writes nothing but the Slack alert and a > drafted GitHub issue. ## Guardrails The agent reads across the entire warehouse, so its access is scoped and strictly one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to the warehouse it's scoped to, and only the Slack alert and the drafted issue are written back out. - **Scoped secrets.** The warehouse credential and the GitHub token are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only, no exceptions.** The connector into the warehouse is read-only. The agent cannot modify a row, alter a schema, or touch a pipeline — it can only alert and draft. - **Draft, not decide.** The GitHub issue is a draft with a diagnosis attached, never auto-closed or auto-assigned. A human owns the incident from there. - **Everything is code.** The agent's SLAs, baselines, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every hour:** Freshness, row counts, and schema rechecked against the live warehouse - **Read-only:** Nothing written to any table, schema, or pipeline - **2 outputs:** A Slack alert and a drafted GitHub issue, nothing else Warehouse problems that used to surface as a wrong number in someone's dashboard now show up as an hourly health check with a specific table, a specific anomaly, and a likely cause — already in Slack and already drafted as an issue. The agent only reads and reports; the data team decides what to fix. --- <!-- /markdown/use-cases/dependency-upgrades.md --> # How we keep dependencies up to date The upgrade agent we run on TaskStation — a weekly cron that opens dependency PRs, runs the full suite in a sandbox, and only opens the PR when it's green. Canonical page: https://taskstation.co/use-cases/dependency-upgrades Dependencies drift. Left alone, a project falls months behind, security patches pile up, and the eventual upgrade turns into a large, risky change nobody wants to own. The usual bots open a PR for every bump and leave a human to work out whether each one is safe, which mostly means the PRs sit unreviewed. We run an upgrade agent on TaskStation that does the checking before it asks for a review. A weekly cron proposes upgrades, applies them in an isolated sandbox, runs the full suite, and only opens a PR when the change is green. This write-up covers how the setup works: the trigger, the session model, and the guardrails. - **Team:** TaskStation - **Runs on:** A weekly cron - **Connected systems:** GitHub · CI - **Mode:** Cron-driven · PR opened only when green ## The problem Keeping dependencies current is work no one schedules. A version-bump bot opens a PR per package, but it can't tell whether the bump breaks anything — that check still falls to a person, so the PRs queue up and the project drifts anyway. The common fixes are incomplete. Ignoring upgrades until something forces the issue turns a routine bump into a migration. Merging bot PRs on green CI trusts whatever tests already exist, not that the upgrade is actually safe. Doing it by hand is reliable but slow, and it's the first thing dropped when the team is busy. ## What we built On TaskStation, a weekly cron triggers an upgrade agent. Each run spawns an isolated session (a cloud sandbox) with scoped access to the repository and CI. The agent checks which dependencies are behind, applies the upgrades on a branch, installs clean, and runs the full suite inside the sandbox. It opens a PR only when the change is green; a human merges. ## How it works ### Connect a weekly cron as the trigger A scheduled **trigger** fires the project once a week. Each firing spawns a fresh **session** in its own sandbox, seeded with a clean checkout of the default branch. One run, one disposable machine, so nothing carries over between weeks and independent upgrade sets can run in parallel. ### Give the agent the upgrade playbook How we handle upgrades lives as **skills** and **memory** that travel with the agent: which packages are pinned on purpose, how to run the suite, the order to apply major versions in, and migrations that have bitten us before. When an upgrade needs a manual step, we write it down and the agent applies it on the next run. ### Connect the systems the upgrade needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the manifests** — resolve which dependencies are behind and how far, separating patch and minor bumps from majors. - **Apply and install in the sandbox** — update the lockfile and install clean on a branch, with the resolution output captured in full. - **Run the full suite** — unit, integration, and e2e inside the sandbox, so a bump that breaks a path fails here rather than in review. - **Open a PR on GitHub** — the branch, the changelog for each bump, and the green result post as a pull request. ### Set the guardrails The agent opens a PR only when the suite passes; a failing upgrade is dropped or split, not pushed for a human to debug. It never merges — the merge is a **human approval gate**. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let the weekly run happen With that in place, each week the agent finds what's behind, applies the upgrades, runs the suite in the sandbox, and opens a PR that's already been proven green — with the bumps grouped and the changelog attached. A bump that breaks a test never becomes a PR; it comes back flagged with the failure instead. > **The pattern** > A weekly **trigger** spawns a session with scoped **connectors** into the repo > and CI. The upgrade playbook is encoded as **skills** and **memory**. The agent > proves the change green in its sandbox and a human owns the merge. ## Guardrails The agent changes dependencies and runs code, so the access is scoped and contained: - **Isolation.** Every run happens in its own isolated sandbox on its own branch. The session can install, resolve, and run the suite to prove an upgrade; only the branch and result are written back out. - **Scoped secrets.** The GitHub and CI credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **PR-gated.** The agent opens a pull request and stops. It never merges and never pushes to the default branch; a human owns the merge. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Weekly:** Upgrades proposed on a schedule, not when something breaks - **Green-only:** PRs opened only after the full suite passes - **Human merge:** The agent proves the change; the team decides Dependencies stay current without anyone scheduling the work, and the upgrade PRs that land in review have already been run against the full suite. Reviewers see a green change with the changelog attached instead of a bump they have to check out and test by hand. --- <!-- /markdown/use-cases/docs-maintainer.md --> # How we keep our docs in sync with the code The docs agent we run on TaskStation — connected to GitHub and our codebase. Once a day it checks the code that landed since its last run and updates the docs those changes affected, opening a PR for review. Canonical page: https://taskstation.co/use-cases/docs-maintainer Documentation tends to fall behind the code. The README, the setup guide, the API reference, and architecture notes drift a release or two back while the code keeps changing. The person who changes the code is usually not the person who owns the page it affects, so the two rarely get updated together. A renamed environment variable, a new setup step, or a removed endpoint is a small code change and a docs change that often goes unmade. We handle this by running a docs sweep close to when the code changes: once a day, over everything that merged since the last run. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails — so you can set up the same thing for your own repo. - **Team:** TaskStation - **Source of truth:** The codebase - **Connected systems:** GitHub · Codebase · Docs site - **Mode:** Daily sweep · PR-gated ## The problem The common fixes each have limits. "Docs are part of the PR" tends to get cut under deadline. A scheduled audit finds drift late and in bulk, when reconstructing what changed is hardest. And a generic AI writer pointed at the docs produces prose that doesn't match the code, because it never reads the code. We wanted the docs we already have to stay accurate to the code, updated on each merge rather than in periodic cleanups. ## What we built On TaskStation, a docs agent runs once a day in the same persistent session — a cloud sandbox — with scoped access to what a docs update needs: the commits that landed since its last run, the codebase for context, and the docs. It picks up from a checkpoint it kept from the previous run, determines what changed, rewrites the affected pages, and opens a single docs PR for review. Nothing publishes without a human merge. ## How it works ### Connect GitHub as the trigger A **cron trigger** fires once a day and resumes the same persistent **session** rather than spinning up a new one per merge. The session reads a checkpoint left by the previous run, then pulls every commit that landed on the default branch since that checkpoint — however many merges that turns out to be. Each affected doc page is handled as its own unit of work, so a problem with one page never blocks the rest of the sweep. One sweep, one PR, and the checkpoint advances at the end whether or not anything changed. ### Give the agent the codebase and the docs standard Our writing conventions are stored as **skills** and **memory** that load into every session: how the docs are structured, the terminology we use, which page covers which subject, and fixes that worked before. The agent writes to that standard rather than inventing one, and it updates as docs PRs get merged. ### Connect what a docs update can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the diff and the codebase** — it sees what changed, then reads the surrounding code to understand intent, not just the delta. - **Search the docs** — it finds every page, README, and reference section that mentions the changed behavior. - **Open a PR on GitHub** — the rewritten docs come back as a reviewable pull request, linked to the change that prompted it. ### Set the guardrails The agent never pushes to a branch anyone reads from: every change lands as a **pull request** gated on a human merge. It edits only files under the docs and README paths; the code itself is read-only to it. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Let each merge update the affected docs With that in place, each day's sweep triages what changed since the last checkpoint, finds the pages that drifted, rewrites them to match, and opens a docs PR with its reasoning attached. A renamed env var becomes an update to the setup guide. A new endpoint becomes a reference entry drafted from the actual handler. A removed feature becomes a PR that strips the stale section. > **The pattern** > Connect the repo via a **trigger** on merge, give the agent scoped > **connectors** into the diff, the codebase, and the docs, encode the writing > standard as **skills** and **memory**, and gate every change behind a reviewed > **PR**. ## Guardrails Giving an agent write access to documentation is a trust question. The relevant controls on TaskStation: - **Isolation.** The daily sweep runs in its own isolated sandbox, resuming the same session across runs via a durable checkpoint rather than persisting the raw repo state. The session can read the whole repo to understand a change, and only the docs PR it opens is written back out. - **Scoped secrets.** The GitHub credential is encrypted in the secrets manager, injected into the sandbox at runtime, and scoped to the agents you grant it to. - **PR-gated.** No change reaches a branch anyone reads without a person reviewing the diff and merging it. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Daily:** Docs re-checked against the code that landed since the last run - **Same-day:** Drift caught before it reaches a reader - **3 systems:** The diff, the codebase, and the docs — in one agent The backlog of "someone should update the README" changes now arrives as small, reviewable PRs within a day of the code landing, with the reasoning written down. The team reviews a diff instead of reconstructing months of drift, and readers stop hitting instructions that are no longer accurate. The setup relies on four pieces working together: sandbox isolation for the daily session, a secrets manager to broker the GitHub token, a PR gate on every change, and a durable checkpoint that carries the sweep forward run to run. --- <!-- /markdown/use-cases/employee-offboarding.md --> # How we offboard employees On a schedule, an agent checks for newly marked departures and runs the offboarding checklist across Okta, Google Workspace, Google Drive, and GitHub — revoking access, transferring ownership, and reclaiming licenses — holding the ownership transfer for approval and never deleting an account. Canonical page: https://taskstation.co/use-cases/employee-offboarding Offboarding is a checklist that has to run to completion. When someone leaves, their access has to be revoked everywhere, their documents and repositories have to change hands, and their licenses have to come back. Miss a step and a former employee keeps a login, a shared drive loses its owner, or a paid seat sits unused. The steps are simple; the risk is in the ones that get skipped. We run offboarding through an agent on TaskStation. On a schedule, the agent checks for newly marked departures and works the checklist across every connected tool, posting what it did and what is still pending. This is the security counterpart to how we onboard. - **Team:** TaskStation - **Runs on:** Scheduled Okta departure check - **Connected systems:** Okta/SSO · Google Workspace · Google Drive · GitHub - **Mode:** Schedule-driven · ownership transfer gated ## The problem When an employee leaves, their access has to be pulled from every system they touched, their work has to be handed off, and their licenses have to be reclaimed. That means Okta, Google Workspace, Slack, and GitHub, each with its own steps, done in the right order, on the same day. The common approaches leave gaps. A written runbook depends on someone working it by hand under time pressure, and a missed line is a live account no one notices. A provisioning tool covers SSO but not document ownership or repository access. IT tickets spread the work across people and days, and the parts that are easy to forget are exactly the ones that matter for security. ## What we built On TaskStation, an agent checks Okta on a schedule for newly marked departures. Each check spawns an isolated session (a cloud sandbox) with scoped access to Okta, Google Workspace, Google Drive, and GitHub, and works the offboarding checklist for every departure it finds: revoke SSO and app access, remove the person from the GitHub org, transfer document and drive ownership, and reclaim licenses. The ownership transfer waits at a human approval gate — account deletion is never attempted — and the agent posts a completed checklist showing what it did and what is still pending. ## How it works ### Check Okta on a schedule The agent checks the Okta group HR adds departing employees to on a schedule, so access comes down the same day a departure is marked. Each check spawns a fresh **session** in its own sandbox; if it finds more than one new departure, each is worked as its own independent case, and nothing carries over between runs. ### Give the agent the offboarding checklist Our offboarding policy lives as **skills** and **memory** that travel with the agent: the full list of systems, the order to work them in, which steps are reversible and which are not, and how ownership should be reassigned. When we add a tool or change a policy, we write it down and the agent picks it up on the next departure. ### Connect the systems access lives in Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Revoke SSO in Okta** — deactivate the account and pull the app assignments that hang off it. Access to Slack drops the moment SSO is revoked, where Slack is SSO-connected. - **Remove from GitHub** — remove the person from the GitHub org and its teams. - **Transfer ownership in Google Drive, then suspend in Google Workspace** — reassign document and shared-drive ownership so nothing is orphaned, then suspend the account. - **Reclaim licenses** — release paid seats across the connected tools so they return to the pool. ### Set the guardrails The reversible steps run on their own; transferring ownership away from a person stops at a **human approval gate** before it executes. Account deletion is out of scope for this agent entirely — it is never performed, gated or otherwise, and hands off to a human if one is ever genuinely needed. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Run the checklist to completion With that in place, each departure the scheduled check finds becomes one worked case: SSO revoked, GitHub access removed, ownership transferred, and licenses reclaimed, with the ownership transfer held for approval. The agent posts a completed checklist showing every step it took and anything still pending a human, so nothing is left half-done. > **The pattern** > A scheduled **trigger** spawns a session with scoped **connectors** into Okta, > Google Workspace, Google Drive, and GitHub. The offboarding policy is encoded > as **skills** and **memory**. The agent runs the reversible steps on its own > and holds the ownership transfer behind a human. ## Guardrails The agent revokes access and transfers ownership across every system, so the access is scoped and contained: - **Isolation.** Every departure case runs in its own isolated sandbox. The session is granted access only to the systems it's scoped to, and only the checklist result is written back out. - **Scoped secrets.** The Okta, Google Workspace, Google Drive, and GitHub credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** Transferring ownership away from a person requires a person to approve before it runs. - **Never deletes.** Account deletion is out of scope for this agent entirely — it is never performed, gated or otherwise. It hands off to a human if a deletion is ever genuinely needed. - **Everything is code.** The agent's checklist, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every departure:** The full checklist run across all systems - **Same day:** Access revoked once the scheduled check finds it - **4 systems:** Okta, Google Workspace, Google Drive, GitHub in one run Access that used to depend on someone working a runbook by hand now comes down the same day a departure is marked, with ownership transferred and licenses reclaimed in the same run. The ownership transfer waits for a person, and the completed checklist shows exactly what happened and what is still pending. --- <!-- /markdown/use-cases/employee-onboarding.md --> # How we onboard new hires A new-hire record triggers an agent to provision accounts, add the person to the right groups and channels, and file a first-week checklist, with account creation and group membership held for approval. Canonical page: https://taskstation.co/use-cases/employee-onboarding Onboarding a new hire spans several systems on day one. There's a Google Workspace account to create, groups and aliases to add them to, Slack channels to invite them into, and a first-week checklist someone has to remember to file. Each step is small, but they live in different tools, and the person doing the setup is rarely the person who owns every system. Steps get missed, and a new hire waits on access they should have had on their first morning. We handle this by tying the setup to the event that starts it: a new record in our HR system. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails — so you can set up the same flow for your own team. - **Team:** TaskStation - **Trigger:** A new-hire record - **Connected systems:** Google Workspace · Slack · Linear - **Mode:** Trigger-driven · Approval-gated ## The problem The common approaches each leave gaps. A written runbook depends on a person following every step by hand, and the steps drift as the tools change. A no-code automation handles the happy path it was built for but can't reason about which groups a given role needs. And handing each system to a different owner means the new hire's access lands piecemeal over their first few days. We wanted the accounts, memberships, and checklist a new hire needs to be prepared the moment their record exists, with the irreversible steps held for a person to confirm. ## What we built On TaskStation, a new-hire record triggers an agent. Each new hire runs in its own isolated session — a cloud sandbox — with scoped access to what onboarding needs: Google Workspace, Slack, and Linear. The agent reads the role, prepares the account, the group and channel memberships, and a first-week onboarding checklist, and files them. Account creation and group membership stop at an approval gate before anything is provisioned. ## How it works ### Connect the HR record as the trigger A signed webhook from our HR system points at the project. A new-hire record fires it, and each firing spawns a fresh **session** in its own sandbox, seeded with the role, team, and start date. One hire, one session, one disposable machine. Nothing carries over between runs. ### Give the agent the onboarding standard Which groups a role belongs to, which channels a team joins, and what a good first week looks like are stored as **skills** and **memory** that load into every session. The agent works to that standard rather than guessing, and it updates as we refine the process. ### Connect what onboarding can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Create the Google Workspace account** — mailbox, aliases, and the group memberships the role calls for. - **Add the hire to Slack** — the channels their team works in. - **File the checklist in Linear** — a first-week onboarding project with the tasks their role needs. ### Set the guardrails Account creation and group membership are the steps that grant access, so they stop at a **human approval gate** before anything is provisioned. A person confirms the account and the memberships in one place. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Let each new hire arrive set up With that in place, a new-hire record prepares the mailbox, the group and channel memberships, and the first-week checklist, and holds the access-granting steps for a person. A new engineer's record becomes a Workspace account, the right Slack channels, and a Linear onboarding project, ready on day one. > **The pattern** > Connect the HR record via a **trigger**, give the agent scoped **connectors** > into Google Workspace, Slack, and Linear, encode the onboarding standard as > **skills** and **memory**, and gate account creation and group membership > behind a human. ## Guardrails Giving an agent the ability to create accounts and grant access is a trust question. The relevant controls on TaskStation: - **Isolation.** Every new hire runs in its own isolated sandbox on its own branch. The session is granted access only to the systems it's scoped to, and only what it's explicitly allowed to send is written back out. - **Scoped secrets.** Each credential is encrypted in the secrets manager, injected into the sandbox at runtime, and scoped to the agents you grant them to or the logs. - **Human approval gate.** Account creation and group membership require a person to approve before anything is provisioned. - **Everything is code.** The agent's persona, skills, and per-system permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Day one:** Accounts and access prepared before the hire starts - **Approval-gated:** Every account and group membership confirmed by a person - **3 systems:** Google Workspace, Slack, and Linear — one agent The scramble to set up a new hire across several tools now arrives as one prepared, reviewable setup the moment their record exists, with the access-granting steps held for a person. Extending it to another system means connecting one more platform. --- <!-- /markdown/use-cases/error-triage.md --> # How we groom our production error backlog The error-triage agent we run on TaskStation — connected to Sentry and GitHub. Hourly, it groups new and spiking errors, dedupes against our existing GitHub issues, and drafts an issue with the stack trace and impact for the top offenders, then alerts Slack. It never resolves, ignores, or assigns an error itself. Canonical page: https://taskstation.co/use-cases/error-triage Sentry catches every error. Turning that stream into work someone actually picks up is a different problem. New error types pile up next to ones that have been spiking for days, half of them already have a GitHub issue somewhere and half don't, and by the time an engineer has the hour to sit down and sort it, the backlog is long enough that sorting it feels like its own project. Most weeks, no one has that hour, so the backlog just grows. We run an error-triage agent on TaskStation that grooms that backlog every hour. It reads the current state of Sentry, groups new and spiking errors, checks which ones already have a tracked GitHub issue, and drafts an issue with the stack trace and impact for the ones that don't — then posts a summary to Slack. This is grooming, not paging: it doesn't wake anyone up and it doesn't decide anything is fixed. - **Team:** TaskStation - **Runs on:** Hourly cron - **Connected systems:** Sentry · GitHub · Slack - **Mode:** Drafts issues + alerts only · never resolves or assigns ## The problem A Sentry project accumulates issues faster than anyone triages them. Some are brand new. Some have been sitting unresolved for weeks at a low, steady rate. Some just started spiking after today's deploy. Working out which of those deserve an engineer's attention — and whether that attention already exists as an open GitHub issue — means reading the whole backlog by hand, which is exactly the task that gets deferred every time something more urgent comes up. The common fixes don't groom the backlog, they just move it. Auto-creating a GitHub issue for every Sentry error floods the tracker with duplicates and one-offs no one will ever action. Leaving Sentry as the system of record means the backlog is only as current as the last time someone opened the dashboard. And on-call triage tools are built for the opposite end of the problem — one loud alert right now — not the quiet pile of errors that never paged anyone but still need a decision. ## What we built On TaskStation, an hourly cron triggers an agent. It spawns a fresh session with read-only access to Sentry, pulls the errors that are new or whose event frequency has spiked over their trailing baseline, and groups them by fingerprint. For each one, it checks our GitHub repo for an existing issue before doing anything else. For the top offenders that aren't already tracked, it drafts a GitHub issue with the stack trace and the impact — event count, affected users, first seen — and posts a summary of what's new, what's spiking, and what got drafted to Slack. It never touches Sentry's issue state and never assigns anything. ## How it works ### Run on an hourly cron A **cron trigger** fires the agent once an hour. Each firing spawns a fresh **session** in its own sandbox. One hour maps to one run on one disposable machine, so the backlog is re-read from Sentry's current state every time and nothing carries over between runs. ### Give the agent the grooming rules How we decide what's worth drafting lives as **skills** and **memory** that travel with the agent: what counts as a spike versus normal noise, how to rank by impact, the issue template we want the stack trace and impact written into, and labels/conventions specific to our tracker. When we tighten a threshold or change the template, we update the file and the next sweep follows it. ### Connect Sentry read-only and GitHub for drafts Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent reads Sentry. A separate, narrowly scoped GitHub token lets it draft issues via the `gh` CLI: - **Read the Sentry backlog** — new issues, and issues whose event frequency has spiked over their trailing baseline, with the full stack trace, first seen, and event/user counts. - **Check GitHub for an existing issue** — before drafting anything, search for an issue already tracking this Sentry error, by its Sentry issue ID. - **Draft a GitHub issue** — for the top offenders that aren't already tracked, with the stack trace and impact attached. - **Post to Slack** — a sweep summary: what's new, what's spiking, what got drafted, what was already tracked. ### Set the guardrails The agent never resolves, ignores, or mutes an error in Sentry, and it never assigns a GitHub issue to anyone — those are decisions for the team, not the agent. Its only writes are the drafted issue and the Slack post. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Let the backlog groom itself With that in place, every hour turns the raw Sentry stream into a short, ranked list of what actually needs a decision: new errors, real spikes, each one already checked against the tracker so nothing gets duplicated, and the ones that matter most already drafted as issues with the trace and impact attached. The team reads the Slack summary and decides what to prioritize; nothing is resolved, ignored, or assigned without them. > **The pattern** > An hourly **cron** spawns a session with a read-only **connector** into > Sentry and a scoped GitHub token for drafts. The grooming rules live as > **skills** and **memory**. The agent dedupes against existing issues, drafts > the top offenders, and alerts Slack — it never resolves, ignores, or assigns > an error. ## Guardrails The agent reads production error telemetry and writes issues others will act on, so its access is scoped and one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session can read Sentry and search GitHub; only the drafted issue and the Slack post are written back out. - **Scoped secrets.** The Sentry credential is brokered through the connector and the GitHub token is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Never resolves, ignores, or mutes.** The agent has no path to changing an error's state in Sentry. It reports; it doesn't manage the backlog's status. - **Never assigns.** A drafted GitHub issue has no assignee. Triage priority and ownership are decided by the team, not the agent. - **Everything is code.** The grooming rules, thresholds, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every hour:** The backlog re-read from Sentry's current state - **Deduped:** Checked against existing GitHub issues before drafting - **Draft-only:** Issues and a Slack alert — never a resolve, ignore, or assign The error backlog that used to need a dedicated hour to sort now arrives pre-groomed every hour: new errors and real spikes ranked by impact, checked against what's already tracked, and the top offenders already drafted as issues with the trace attached. The team spends its time deciding what to fix, not rediscovering what's broken. --- <!-- /markdown/use-cases/escalation-manager.md --> # How we route support escalations The escalation-manager agent we run on TaskStation — every 15 minutes it scans the Plain queue for SLA breaches, VIP accounts, and high-severity tickets, routes each to the right team, opens a linked Linear issue when engineering is needed, and posts to our escalation Slack channel. It never closes a ticket or promises a resolution. Canonical page: https://taskstation.co/use-cases/escalation-manager A ticket that needs to jump the line looks like any other ticket until someone reads it closely: the SLA clock that's about to run out, the account that happens to be one of our biggest, the report that's actually a full outage. In a busy queue, those tickets sit in first-in-first-out order next to everything else, and the person who should be routing them is also the person answering the queue. We run an escalation-manager agent on TaskStation that checks the Plain queue every 15 minutes, finds the tickets that meet our escalation criteria, routes each one to the right internal team, opens a linked Linear issue when it's an engineering problem, and posts the whole thing to our escalation Slack channel. It routes, links, and alerts. It never closes a ticket and never tells a customer when or how their issue will be fixed. - **Team:** TaskStation - **Runs on:** Every 15 minutes - **Connected systems:** Plain · Linear · Slack - **Mode:** Routes + alerts · never closes a ticket ## The problem Escalation-worthy tickets don't announce themselves. An SLA breach is a timestamp comparison someone has to actually run. A VIP account is a fact that lives in a spreadsheet or a CRM field, not in the ticket itself. High severity is a judgment call that depends on reading the ticket body, not just its tags. Any one of those, on its own, is easy to miss in a queue moving fast enough that most tickets get handled in the order they arrived. The common fallback is a person periodically scanning the queue for anything that looks urgent, on top of their regular ticket load. It works until the queue is busy, at which point the SLA-breaching ticket and the VIP account's ticket wait behind ten routine ones, and the bug report that should already be a Linear issue is still just a paragraph in a support thread nobody on engineering has seen. ## What we built On TaskStation, a cron fires every 15 minutes and spawns a fresh agent session. It reads the current state of the Plain queue, checks every open ticket against three criteria — SLA breach, VIP account, high severity — routes each qualifying ticket to the right internal team, opens a linked Linear issue when the ticket is an engineering problem, and posts one alert per escalation to the escalation Slack channel with the reason, the routed team, and the linked issue. It never closes a ticket, and it never tells the customer anything. ## How it works ### Run every 15 minutes, from scratch A **cron trigger** fires every 15 minutes. Each firing spawns a fresh **session** with no memory of the last one — the agent re-checks the entire open queue against Plain's current state rather than trusting what it concluded 15 minutes ago. Nothing carries over, and nothing is missed because a prior run's notes went stale. ### Give the agent the escalation rules What counts as an escalation lives as a **skill**: the SLA targets per severity tier, the VIP account list, what qualifies as high severity, and the routing table mapping ticket type to owning team. When we tighten an SLA target or add an account to the VIP list, we update the skill and the next sweep applies it. ### Connect the queue, the tracker, and the channel Through scoped **connectors** and **secrets**, brokered server-side so no raw credential reaches the model, the agent: - **Reads and routes tickets in Plain** — pulls the open queue, checks SLA timers and account metadata, and tags or assigns a qualifying ticket to the right internal team. It never resolves or closes a ticket. - **Opens a linked issue in Linear** — when a ticket needs engineering work, it files an issue in the engineering team's tracker with the ticket's context attached, and links the two records to each other. - **Posts to Slack** — one alert per escalation in the escalation channel: which ticket, why it escalated, which team it went to, and the linked issue if one was opened. ### Set the guardrails The agent's write access to Plain is scoped to **routing, not resolving** — it can tag and assign a ticket, but it cannot close one or mark it resolved. It never drafts or sends anything to the customer, and it never states or implies a resolution timeline. The only things that are written back out are the Plain routing update, the linked Linear issue, and the Slack alert. ### Route, link, and alert With that in place, every 15 minutes the queue gets rechecked against the current SLA clock, the current VIP list, and the current severity of every open ticket. Whatever qualifies gets routed to the right team, gets a linked Linear issue if engineering needs to see it, and shows up in the escalation channel with the reason attached. The team decides what happens from there. > **The pattern** > A 15-minute **cron** spawns a fresh session that reads the Plain queue > against **skill**-defined SLA targets, a VIP list, and a severity bar. It > routes the ticket, opens a **linked Linear issue** for engineering work, and > posts to Slack — it never closes a ticket or promises the customer anything. ## Guardrails The agent has write access to the support queue and the issue tracker, so its authority is scoped tightly to routing and alerting: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to Plain, Linear, and the escalation Slack channel — nothing else. - **Scoped secrets.** The Plain API key is encrypted in the Secrets Manager and injected into the sandbox at runtime; the Linear connector is brokered server-side. Every secret is scoped to the agents you grant it to. - **Route and alert, never resolve.** The agent can tag, assign, and open a linked issue. It cannot close a ticket, mark one resolved, or take any action that ends the customer's case. - **No promises to the customer.** The agent never drafts or sends a customer-facing message and never states a resolution timeline. Every output is internal: the routing tag, the linked issue, and the Slack alert. - **Everything is code.** The SLA targets, the VIP list, the severity bar, and the routing table are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every 15 min:** The open queue rechecked against the current SLA clock - **0 closures:** The agent routes and alerts; it never closes or resolves a ticket - **1 linked issue:** Per engineering escalation, traceable from ticket to issue and back Tickets that used to wait behind the rest of the queue for someone to notice now get caught within 15 minutes of qualifying, routed to the team that owns them, and — when it's an engineering problem — turned into a linked issue before anyone has to ask. The agent routes, links, and alerts; the team decides how the ticket actually gets resolved. --- <!-- /markdown/use-cases/expense-reconciliation.md --> # How we reconcile expenses every month The finance agent we run on TaskStation — connected to Stripe, the bank feed, and Google Sheets. Each month it matches transactions against invoices, flags mismatches, and posts a summary, escalating anything it can't match. Canonical page: https://taskstation.co/use-cases/expense-reconciliation Monthly reconciliation is the kind of task that's simple to describe and tedious to do: line up what came in and went out against what was supposed to, find the things that don't match, and explain them. On a small team it lands on one person for an afternoon, and it's the same afternoon every month. We handle this with a scheduled agent. Once a month it reconciles transactions against invoices, flags what doesn't line up, and posts a summary. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails. - **Team:** TaskStation - **Connected systems:** Stripe · Bank feed · Google Sheets - **Trigger:** Monthly cron - **Mode:** Scheduled · human-escalated ## The problem Reconciliation is mechanical until it isn't. Most lines match cleanly, and the value is entirely in the handful that don't: a payment that never landed, a fee that doesn't tie out, an invoice with no matching deposit. Finding those means going through everything, which is slow and easy to get wrong when it's the same repetitive comparison hundreds of times. A spreadsheet formula catches the exact matches but not the near ones, and the near ones are where the problems hide. So the check either runs shallow and misses things, or runs deep and eats a person's day. Either way it only happens as often as someone has time for it. ## What we built On TaskStation, a monthly **cron** triggers a reconciliation agent. Each run spawns its own isolated session — a cloud sandbox — with scoped access to what reconciliation needs: Stripe, the bank feed, and the Google Sheet we track in. It matches transactions against invoices, flags the mismatches, posts a summary to the sheet, and escalates anything it can't match to a person. ## How it works ### Connect a monthly cron as the trigger A **cron** trigger fires the project on a schedule — once a month, after the period closes. Each firing spawns a fresh **session** in its own sandbox. One run, one session, one disposable machine, with nothing carried over from the last month's run. ### Give the agent our reconciliation rules How we reconcile lives as **skills** and **memory** loaded into every session: what counts as a match, the tolerances we allow on fees and timing, the accounts we track, and mismatches we've explained before. The agent works to that standard rather than inventing one, and the memory updates as patterns recur. ### Connect what reconciliation can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read Stripe** — charges, payouts, and fees for the period. - **Read the bank feed** — deposits and withdrawals to match against Stripe and invoices. - **Read and write the Google Sheet** — the ledger it reconciles against and the summary it posts back. ### Set the guardrails The agent reads the financial systems and writes only to the tracking sheet — it never moves money or touches Stripe or the bank beyond reading. Anything it can't match within tolerance is **escalated to a person** rather than force-fit or written off. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each month run With that in place, the monthly run produces a reconciled sheet and a summary without anyone starting it. Clean matches tie out silently. A payout that's short by a fee gets explained. An invoice with no matching deposit gets flagged and sent to a person to chase, with the two records it couldn't reconcile attached. > **The pattern** > Fire the project on a monthly **cron**, give the agent read access to Stripe > and the bank feed and write access to the sheet through scoped **connectors**, > encode the reconciliation rules as **skills** and **memory**, and escalate > anything it can't match to a person. ## Guardrails Giving an agent access to financial systems is a security question first. The relevant controls on TaskStation: - **Isolation.** Each run happens in its own isolated sandbox. The session can read the accounts it needs, and only the summary it writes to the sheet is written back out. - **Scoped secrets.** The Stripe, bank, and Sheets credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** Anything the agent can't match within tolerance is escalated to a person rather than force-fit, and the agent never moves money. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Every month:** Reconciliation run on schedule, not when there's time - **Mismatches only:** A person looks at the exceptions, not every line - **3 systems:** Stripe, the bank feed, and the sheet in one agent The afternoon that used to go to comparing hundreds of lines now goes to resolving the few the agent couldn't match. The reconciliation runs every month whether or not anyone has time, and the summary is waiting when the finance person opens the sheet. The setup relies on four pieces: sandbox isolation per run, a secrets manager to broker the Stripe, bank, and Sheets tokens, an escalation path for anything unmatched, and memory that carries recurring explanations forward month to month. --- <!-- /markdown/use-cases/feature-flag-cleanup.md --> # How we remove stale feature flags The flag-cleanup agent we run on TaskStation — a weekly sweep that finds feature flags that are fully rolled out or long dead, deletes the flag and the dead code branch it guards, and opens a PR for a human to merge. Canonical page: https://taskstation.co/use-cases/feature-flag-cleanup Every feature flag we ship is supposed to be temporary. It guards a rollout, we watch it ramp to 100%, and then someone is supposed to go back and delete it. In practice that last step loses to whatever shipped after it. Flags pile up, each one leaving an `if` branch, a dead `else`, and a config entry that nobody reads anymore — until the codebase is carrying the weight of every rollout it ever did. We run a flag-cleanup agent on TaskStation that sweeps our repo every week, finds the flags that are safe to remove, deletes them and the branches they guard, and opens a PR. It never merges its own work, and it never touches a flag that's still in the middle of a rollout — those get flagged for a human instead. - **Team:** TaskStation - **Runs on:** Weekly cron, reusable session - **Connected systems:** GitHub - **Mode:** Opens PRs only · human merges ## The problem A feature flag has a natural lifecycle: ship behind it, ramp the rollout, hit 100%, delete it. The first three steps are urgent and have an owner. The fourth one isn't and doesn't — by the time a flag is fully rolled out, the team that shipped it has moved on, and removing a few lines of conditional logic loses every priority fight. The cost isn't obvious from any single flag. It's cumulative: every live flag is a branch some engineer has to reason about, a code path that might not even be tested anymore, and a source of bugs where the "off" branch silently rots. Left alone, the flag inventory only grows, and nobody wants to be the one who deletes a flag that turns out to still matter. ## What we built On TaskStation, a weekly cron re-prompts one **reusable session** that treats flag cleanup as an ongoing sweep rather than a one-off task. It clones our repo, inventories every feature flag, and classifies each one: fully rolled out and safe to remove, long dead with no live check left, still in partial rollout, or part of an active experiment. For the first two categories it deletes the flag and the dead branch it guards, proves the suite still passes, and opens a PR. For the rest, it does nothing but note them for a human to look at. ## How it works ### Run on a weekly reusable session A **cron trigger** fires once a week and re-prompts the same **session** rather than starting fresh each time. The agent keeps a ledger of every flag's age, rollout state, and whatever it already proposed, so it never re-litigates a flag it already handled or duplicates an open PR. ### Give the agent the classification rules How to tell a dead flag from a live one lives as a **skill** that travels with the agent: what "fully rolled out" looks like in our flag config, how long a flag has to go unchecked before it counts as long dead, and — just as important — what partial rollout and active-experiment signals look like, so the agent knows exactly when to stop. ### Work the codebase through GitHub The agent's only system is the codebase itself, reached through the **GitHub** connector via a scoped token: it clones the repo, greps for every flag definition and call site, removes the flag and the code branch it guards on an isolated branch, runs the verification suite in the sandbox, and pushes a PR through the `gh` CLI. ### Set the guardrails The agent never removes a flag still in partial rollout or an active experiment — those are reported, not touched. It never pushes to the default branch and never merges its own PR. The GitHub token is injected into the sandbox at runtime and scoped to the agents you grant them to. ### Open the cleanup PR Once the affected flags are classified and the dead code is removed, the agent runs the full suite. Only when it's green does it push the branch and open a PR listing exactly which flags it removed and why, plus a separate note for any flag it found still in partial rollout. A human reviews and merges; nothing lands on its own. > **The pattern** > A weekly **cron** re-prompts a reusable **session** that keeps a ledger of > every flag's age and rollout state. The classification rules live as a > **skill**. The agent removes only what's fully rolled out or long dead, and > hands anything still in partial rollout to a human. ## Guardrails The agent deletes code, so its judgment is bounded on both sides — what it's allowed to remove, and what it must leave alone: - **Partial rollout is untouchable.** Any flag still ramping, or part of an active experiment, is never removed. The agent logs it and moves on; a human decides its fate. - **No direct pushes, no self-merges.** Every change lands on an isolated branch and becomes a PR. The agent never pushes to the default branch and never merges its own work. - **Verification before the PR, not after.** The full suite runs inside the sandbox on the cleanup branch. A failing removal is dropped and logged, never pushed for a human to untangle. - **Scoped, ephemeral access.** The GitHub token is encrypted in the Secrets Manager and injected into the sandbox at runtime — scoped to the agents you grant them to, never written to disk or logs. - **Everything is code.** The classification rules, the ledger, and the agent's permissions are files in the repo, versioned and changed through a reviewed change request. ## The outcome - **Weekly:** Every flag re-checked against its current rollout state - **PR-only:** Every removal lands as a reviewable PR, never a direct push - **0:** Partial-rollout flags ever touched by the agent Flags that used to sit forgotten at 100% rollout now get deleted within a week of reaching it, with the dead branch cleaned up alongside them. The agent proposes; a human still owns the merge, and anything mid-rollout stays exactly where the team left it. --- <!-- /markdown/use-cases/flaky-test-triage.md --> # How we detect and quarantine flaky tests The flaky-test agent we run on TaskStation — connected to GitHub CI history and Slack. It scores every test's non-determinism, opens a quarantine PR for the worst offenders, and files a tracking issue for a human to review. Canonical page: https://taskstation.co/use-cases/flaky-test-triage Flaky tests erode trust in CI fast. A test fails for no reason anyone can pin down, gets rerun, goes green, and CI moves on — until the day a real regression hides behind the same "oh that one's just flaky" shrug. Nobody schedules time to fix the true flakes because nobody has a ranked list of which ones are actually costing the team reruns. We run a flaky-test-triage agent on TaskStation that reads the CI run history every day, scores every test on how often it flips outcome on unchanged code, and opens a quarantine PR for the worst offenders with a tracking issue attached. It only skips tests, never deletes one, and it never merges its own PR — a human reviews the quarantine and eventually retires it once the test is fixed. - **Team:** TaskStation - **Runs on:** Daily cron, one persistent session - **Connected systems:** GitHub · CI run history · Slack - **Mode:** Quarantine PR + tracking issue only — never merges, never deletes ## The problem Flakiness hides in plain sight. A test fails, the job reruns, it passes, and CI goes green — so the failure never becomes a signal anyone tracks. Spread across a few hundred tests and a few months, a handful of tests are quietly eating a rerun every week, and genuinely broken tests get the same "just rerun it" treatment as the flaky ones. The common responses don't fix this. Rerunning failed jobs until green hides the problem instead of measuring it. A channel where someone occasionally asks "is this one flaky again?" depends on a person noticing and remembering. Deleting a flaky test outright throws away whatever real coverage it had, and doing any of this by hand means first digging through weeks of CI logs to find which tests are actually the worst offenders. ## What we built On TaskStation, a daily cron re-prompts one persistent agent session. It resumes from a flakiness ledger, pulls the CI run history from GitHub since the last check, updates every test's non-determinism score, and once a test crosses the quarantine threshold, opens a PR that skips it with a reason and a link to the evidence — plus a running tracking issue listing every currently quarantined test. It never deletes a test and never merges its own PR. ## How it works ### Run on a daily cron, one persistent session A **cron trigger** fires once a day against the same **session**, not a fresh sandbox each time. Because the trigger runs in reusable-session mode, the per-test flakiness history survives from one run to the next instead of being recomputed from scratch, so a test's score reflects weeks of runs, not just today's. ### Give the agent the quarantine playbook How we score flakiness, which skip syntax each test framework uses, and what a quarantine PR and tracking issue should contain live as **skills** and **memory** that travel with the agent. When we tune the threshold or learn a flaky test's root cause, we write it down and the next run picks it up. ### Connect the systems the triage needs Through a scoped **connector**, brokered server-side so no raw token reaches the model, plus a **GH_TOKEN** secret for the `gh` CLI, the agent can: - **Read CI run history from GitHub** — every workflow run's per-test results over a rolling window, to see which tests flip outcome on unchanged code. - **Open a quarantine PR on GitHub** — skip markers on the worst offenders, each with a reason and a link to the run evidence. - **File a tracking issue on GitHub** — one running issue listing every currently-quarantined test, its flakiness score, and its status. - **Post to Slack** — a summary of what changed this run: newly quarantined tests, tests still flaky, and tests ready to be de-quarantined. ### Set the guardrails The agent can only **skip and mark**, never delete. It opens a PR and an issue and stops; it never merges its own PR and never pushes to the default branch. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Let the daily triage happen With that in place, each day the agent updates every test's flakiness score against the current run history, and when a test crosses the threshold, quarantines it in a PR with the evidence attached, rolls it into the tracking issue, and posts the day's summary to Slack. A human reviews the PR, merges it if the quarantine is warranted, and eventually removes the skip once the test is actually fixed. > **The pattern** > A daily **cron** re-prompts one persistent **session** so the flakiness > ledger survives across runs. The agent reads CI history through GitHub, > scores and ranks every test, and its only outputs are a quarantine PR, a > tracking issue, and a Slack summary — never a merge, never a deletion. ## Guardrails The agent changes what runs in CI, so its access is scoped and one-directional: - **Isolation.** Every run happens in the session's own isolated sandbox. Only the PR, the issue, and the Slack post leave it. - **Scoped secrets.** The GitHub connector and the `GH_TOKEN` used by the `gh` CLI are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. - **Skip, never delete.** The agent's only edit to a test file is a skip/quarantine marker with a reason; it never removes a test, its assertions, or its file. - **PR-gated.** The agent opens a PR and an issue and stops. It never merges and never pushes to the default branch; a human owns the merge and the eventual de-quarantine. - **Everything is code.** The agent's scoring rules, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Test flakiness re-scored against the latest CI history - **Ranked, not guessed:** Quarantine targets picked from evidence, not gut feel - **Human merge:** The agent proposes the quarantine; the team decides The tests that used to eat a silent rerun every week now show up ranked, with a PR and a tracking issue attached to the worst of them. CI gets quieter without losing coverage nobody meant to drop, and the team spends its fixing time on the flakes that are actually costing the most reruns. --- <!-- /markdown/use-cases/gdpr-dsar.md --> # How we handle GDPR data-subject requests on a deadline The GDPR-DSAR agent we run on TaskStation — it verifies each incoming access or deletion request, locates the subject's data across our product database, compiles the report inside the SLA, and flags it for legal to review before anything goes out. Canonical page: https://taskstation.co/use-cases/gdpr-dsar A GDPR data subject access request starts a clock. From the day we can verify who's asking, we have one calendar month to tell them what we hold on them, or to delete it — and that data is never in one place. It's spread across a users table, an orders table, a support history, an events table, a dozen joins a person has to know to write by hand. Miss the deadline or miss a table, and the exposure is the regulator's, not just the requester's. We run a GDPR-DSAR agent on TaskStation that reads each request out of our legal inbox, verifies it, locates the subject across the product database, and compiles the access-or-deletion report as a Google Doc — every day, well inside the SLA. It never deletes anything and never replies to the requester. It compiles and flags; legal decides and sends. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Gmail · Postgres · Google Docs - **Mode:** Compile + flag only · legal signs off before anything ships ## The problem A DSAR arrives as an email, but the data it's asking about lives in a relational schema built for the product, not for privacy law. "Everything you have on me" means rows in the users table, orders, invoices, support tickets, login history, marketing consent, and whatever else joins off a customer ID — tables that were never designed to be read together, by someone who has one month to do it and get it right every time. The common approaches don't hold up under repetition. A shared spreadsheet of "where subject data lives" goes stale the moment the schema changes. Assigning it to whoever's free that week means the verification step — confirming the requester actually is the subject, not someone phishing for their data — gets rushed or skipped. And a script that queries and deletes in the same run has no way to stop and ask a lawyer first. ## What we built On TaskStation, a daily cron triggers an agent. It spawns a fresh session that checks our legal inbox in Gmail for new DSAR emails, verifies each requester against the identity details we already hold, then queries Postgres read-only across every table keyed to that subject — account, orders, support history, consent records — and compiles the findings into a Google Doc formatted as an access report (or a deletion inventory, if that's what was asked for). The doc, and a recommended action, get flagged to legal for review. The agent never runs a delete and never emails the subject back. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox, so every DSAR is worked from a clean read of the inbox and the database — nothing from a prior day's run carries over, and a stale verification never gets reused. ### Verify before locating anything Verification lives as a **skill** that travels with the agent: what counts as proof the requester is the subject (matching the request email against the account on file, or the identity details supplied), what to do when verification fails or is ambiguous, and what the SLA clock actually is under GDPR — one month from a verified request, extendable once for complex cases. Nothing is located until a request passes this check. ### Connect the systems it needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent: - **Reads the DSAR out of Gmail** — the inbound request and any identity details attached, and labels the thread once it's been actioned. - **Queries Postgres read-only** — every table keyed to the subject: account, orders, invoices, support tickets, consent and login history — to build a complete picture of what we hold. - **Writes the report to Google Docs** — a structured access report or deletion inventory, one doc per request, ready for a lawyer to read start to finish. ### Set the guardrails The agent's only writes are the Google Doc it compiles and the label on the Gmail thread. Postgres access is read-only — there is no delete path the agent can reach, even for a request that explicitly asks for erasure. It never emails the requester. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Flag it for legal review With the report compiled, the agent posts it to the legal team's review channel with a recommended action — fulfill the access request, or proceed with deletion — and the SLA deadline it's working against. Legal reviews the report, decides, and is the one who replies to the subject or authorizes any deletion. The agent's job ends at the flag. > **The pattern** > A daily **cron** spawns a fresh session that verifies the request, reads Gmail for > the ask, queries Postgres read-only to locate the subject's data everywhere it > lives, and compiles a Google Doc report. The agent compiles and flags; legal > decides, deletes, and replies. ## Guardrails A DSAR touches the most sensitive data we hold, so the agent's access is scoped and its authority stops well short of the parts that matter: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to Gmail, Postgres, and Google Docs, and only the compiled report and the legal-channel flag are written back out. - **Read-only on the data itself.** Postgres access is read-only. The agent can locate a subject's data across every table; it cannot delete a row, update a record, or run anything but a `SELECT`. - **Never replies to the subject.** The agent's only outbound message is the flag to legal. It does not draft or send anything to the person who filed the request — that response is legal's to write and send. - **Deletion always waits for approval.** Even when a request explicitly asks for erasure, the agent only recommends it in the flagged report. No deletion happens until a lawyer approves it and someone else executes it. - **Everything is code.** The verification rules, the table map, and the agent's per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** DSAR inbox checked and any new request worked the same day - **0 deletions:** Every erasure held for a lawyer to approve and execute - **1 flag:** One compiled report and recommendation per request, in legal's queue A request that used to mean someone manually joining tables under deadline pressure now arrives in legal's queue as a compiled report, well inside the SLA, with the data located and the action recommended. The agent verifies and compiles; the people decide, delete, and reply. --- <!-- /markdown/use-cases/hr-policy-qa.md --> # How we answer HR policy questions automatically The HR-policy helpdesk agent we run on TaskStation — connected to Gmail, our Notion policy library, and Slack. Every 15 minutes it reads new questions from the HR inbox, answers strictly from documented policy, and drafts a reply for HR to send — escalating anything ambiguous, legal, personal, or compensation-related to Slack instead of answering it. Canonical page: https://taskstation.co/use-cases/hr-policy-qa The same handful of policy questions land in the HR inbox every week — how much PTO carries over, what the remote-work policy actually allows, how the expense limit works for a conference, when parental leave starts accruing. The answers are documented, sitting in a Notion policy library the whole company can already read. But finding the right page, quoting it accurately, and drafting a reply takes a person's attention away from the HR work that actually needs a person: the exception request, the accommodation, the sensitive conversation. We run an HR-policy helpdesk agent on TaskStation that reads the inbox every 15 minutes, answers what's documented, and hands off everything else. It never sends an email itself and it never decides a policy question the library doesn't already answer — it drafts, or it escalates, and a human takes it from there. - **Team:** TaskStation - **Runs on:** Every 15 minutes - **Connected systems:** Gmail · Notion · Slack - **Mode:** Fresh session · drafts only · escalates the rest ## The problem HR policy questions are high-volume and low-variance — most of them are already answered, word for word, somewhere in the policy library. But the inbox doesn't know that. Every question waits in the same queue whether it's "how many sick days do I have" or "I need to discuss a medical accommodation," and someone has to open each thread to find out which one it is before they can even start answering. A canned FAQ page only helps the employee who thinks to check it. A shared doc link in an auto-reply doesn't parse the actual question or point at the specific section that answers it. And routing everything to a person means the easy, already-documented questions compete for the same attention as the ones that genuinely need HR's judgment — the legal questions, the personal situations, the comp conversations — which is exactly backwards. ## What we built On TaskStation, a cron fires every 15 minutes and spawns a fresh session with read-only access to the HR Gmail inbox and the Notion policy library, plus a channel into Slack. It reads new questions, classifies each one, and for anything the library documents, searches for the exact policy section and drafts a grounded reply as a Gmail draft — never sent automatically. Anything ambiguous, legal, personal, or compensation-related gets posted to the HR Slack channel instead, with the question and the reason it needs a person, and the agent does not attempt an answer. ## How it works ### Run every 15 minutes, fresh each time A **cron trigger** fires every 15 minutes. Each firing spawns a fresh **session** in its own sandbox with no memory of the last run. The agent relies on Gmail's own labels — not its own memory — to know which threads it has already handled, so nothing is reprocessed and nothing is missed between runs. ### Give the agent the triage and answering rules How to tell a documented question from one that needs HR lives as a **skill** that travels with the agent: what counts as "answerable from policy," the exact bar for escalating (ambiguous, legal, personal, or compensation-related), and how to draft a reply that quotes or closely paraphrases the source text instead of paraphrasing loosely. ### Connect the inbox, the library, and the escalation channel Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent reads and writes: - **New questions from Gmail** — the HR inbox, read for unlabeled threads and written to only as a draft, never a send. - **Policy from Notion** — the documented policy library, read-only, searched for the specific page that answers each question. - **Escalations to Slack** — the HR channel, where anything outside documented policy gets posted with the question and the reason it needs a person. ### Set the guardrails The agent answers **only** what the policy library documents — it never infers, extends, or guesses at a policy that isn't written down, and it never makes an exception. It never sends an email; every answer is a Gmail draft for HR to review and send. Anything legal, personal, or compensation-related is escalated to Slack without an attempted answer, every time. ### Draft, escalate, and label With that in place, every 15 minutes the inbox gets worked: documented questions get a cited, drafted reply waiting in Gmail; everything else lands in the HR Slack channel with the question and why it was escalated. Each thread is labeled so the next run leaves it alone. HR reviews the drafts, sends what's right, and handles the escalations directly. > **The pattern** > A 15-minute **cron** spawns a fresh session with read-only access to Gmail and > Notion and a channel into Slack. The triage and answering rules live as a > **skill**. The agent drafts from documented policy or escalates — it never > sends and it never makes a policy call the library hasn't already made. ## Guardrails The agent touches an employee's inbox and the company's policy record, so its access is scoped and every output is held for a person: - **Answer from documented policy only.** The agent never invents, infers, or extends a policy beyond what's written in the Notion library. If the library doesn't clearly cover the question, it escalates instead of guessing. - **Drafts only, never sends.** Every reply the agent writes is a Gmail draft. HR reviews and sends it; the agent has no send authority. - **No exceptions, ever.** The agent cannot grant, waive, or bend a policy — that judgment call belongs to HR, not the agent. - **Hard escalation triggers.** Anything ambiguous, legal in nature, about an individual's personal situation, or related to compensation is routed to Slack for a human, with no attempted answer. - **Read-only into the policy library.** The agent never edits a Notion page. - **Scoped secrets.** Gmail and Notion access is brokered through connectors, scoped to the agents you grant them to. ## The outcome - **Every 15 min:** Inbox checked against the current policy library - **Drafts only:** Every reply is a Gmail draft; HR sends it - **0 exceptions:** Ambiguous, legal, personal, and comp questions always escalate The routine policy questions that used to sit in the queue behind everything else now get a cited, ready-to-send draft within 15 minutes, quoting the exact policy that answers them. The questions that actually need HR's judgment show up in Slack immediately, with nothing guessed at in between. The agent reads and drafts; HR decides and sends. --- <!-- /markdown/use-cases/inbox-triage.md --> # How we triage a shared inbox with an AI agent The inbox agent we run on TaskStation — connected to Gmail, a help doc, and Linear. It labels every inbound email, drafts a reply, or files a task, and stops for approval before anything goes to a customer. Canonical page: https://taskstation.co/use-cases/inbox-triage A shared inbox is where a small team's requests pile up: support questions, sales pings, bug reports, and the occasional invoice, all landing in one place with no owner. Sorting them by hand is the first thing that slips when the team is busy, and a message sitting unread for a day is a message the sender assumes was ignored. We handle this by putting an agent on the inbox. Every inbound email gets read, labelled, and either drafted a reply or turned into a task. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails. - **Team:** TaskStation - **Connected systems:** Gmail · Help doc · Linear - **Trigger:** New inbound email - **Mode:** Trigger-driven · human-gated ## The problem The messages that need a fast reply are mixed in with the ones that can wait, and telling them apart takes a person reading each thread. On a small team that person is also doing three other jobs, so triage happens in bursts — everything looks urgent when the inbox is opened once a day, and nothing looks urgent in between. Filters and rules help with the obvious cases but not the judgment calls: whether a message is a real support issue or a sales lead, whether it needs a reply or a ticket, and what a reasonable answer would be. Those are the parts that actually take time. ## What we built Our shared inbox in **Gmail** is connected to an agent running on TaskStation. Every inbound email spawns its own isolated session — a cloud sandbox — with scoped access to what triage needs: the message, our help doc, and Linear. It reads the email, applies a label, and then either drafts a reply in Gmail or files a task in Linear. Anything customer-facing waits for a person before it sends. ## How it works ### Connect Gmail as the trigger A signed webhook from Gmail points at the project. Every new inbound message fires it, and each firing spawns a fresh **session** in its own sandbox, seeded with the email. One message, one session, one disposable machine. Sessions don't share state, and a busy inbox means more sessions running in parallel. ### Give the agent our knowledge How we handle the inbox lives as **skills** and **memory** loaded into every session: our help doc, the categories we sort into, the reply tone, and answers that worked before. The agent triages to that standard rather than inventing one, and the memory updates as threads are resolved. ### Connect what triage can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the help doc** — it looks up the answer to a common question instead of guessing. - **Label and draft in Gmail** — it applies the right label and, where a reply fits, writes a draft on the thread. - **File a task in Linear** — a bug report or a request that needs follow-up becomes a ticket with the context attached. ### Set the guardrails The agent labels and drafts freely, but nothing sends on its own: every customer-facing reply stops at a **human approval gate** as a draft for a person to review and send. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each email run With that in place, an inbound email triages itself: it gets a label, and either a draft reply waiting for approval or a Linear ticket already filed. A common support question becomes a drafted answer pulled from the help doc. A bug report becomes a ticket. A sales ping becomes a label the right person can pick up. > **The pattern** > Connect the inbox via a **trigger** on new mail, give the agent scoped > **connectors** into the help doc, Gmail, and Linear, encode how we triage as > **skills** and **memory**, and gate every customer-facing reply behind a human. ## Guardrails Giving an agent access to the shared inbox is a trust question as much as a convenience one. The relevant controls on TaskStation: - **Isolation.** Each email runs in its own isolated sandbox. The session can read the thread and the help doc it needs, and only the draft or ticket it produces is written back out. - **Scoped secrets.** The Gmail and Linear credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** Nothing customer-facing sends without a person reviewing the draft first — replies land as drafts, not sent mail. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Every email:** Labelled and routed as it lands - **Drafted:** Common replies written and waiting for a send - **3 systems:** The help doc, Gmail, and Linear in one agent The inbox stops being a pile to sort and becomes a list already labelled, with replies drafted and tickets filed. The team reviews and sends instead of reading every thread cold, and no message sits unread waiting for someone to notice it. The setup relies on four pieces: sandbox isolation per email, a secrets manager to broker the Gmail and Linear tokens, a human approval gate before anything reaches a customer, and memory that improves as threads are resolved. --- <!-- /markdown/use-cases/incident-postmortem.md --> # How we draft incident postmortems When an incident resolves, an agent pulls the timeline from the incident channel, correlates deploys and log spikes, and drafts a structured postmortem as a doc PR for the team to review and edit. Canonical page: https://taskstation.co/use-cases/incident-postmortem A postmortem is worth writing and easy to skip. Once an incident is resolved and the pressure is off, someone has to reconstruct the timeline from a scrolling channel, line it up against the deploys and log spikes, work out the root cause, and write it all up. It's an hour of careful work at the moment everyone most wants to move on, so it often gets a thin summary or nothing. We run an agent on TaskStation that drafts the first version. When an incident resolves, it reads the incident channel, correlates the deploys and log spikes, and opens a structured postmortem as a doc PR. It drafts; humans review and finalize. This is how we write our own postmortems. - **Team:** TaskStation - **Runs on:** Incident resolved, the agent drafts - **Connected systems:** Incident channel · Logs · GitHub - **Mode:** Drafts a doc PR, humans finalize ## The problem The facts of an incident are scattered across places that don't line up on their own: the back-and-forth in the incident channel, the deploy history, and the log spikes. Turning them into a timeline means reading the channel top to bottom, matching each moment against what shipped and what the logs did, and inferring where the root cause sits. The common workarounds are thin. A blank template gets filled in from memory, so the timeline drifts and details get rounded off. A resolved incident with no writeup means the same failure mode can recur with nothing to point back to. The information to do it properly exists; assembling it by hand is exactly the work no one wants right after an incident. ## What we built On TaskStation, resolving an incident triggers an agent. It reads the incident channel for the timeline, pulls the deploy history and log spikes for the same window, correlates them, and drafts a structured postmortem, timeline, root cause, and action items, as a doc PR against the repo. The team reviews and edits the draft the way they'd review any change. The agent drafts; it never publishes a final postmortem on its own. ## How it works ### Trigger on the resolved incident The incident channel is connected as a **channel**, so marking an incident resolved is the trigger. That fires a fresh **session** in its own isolated sandbox, seeded with the incident. One incident maps to one session on one disposable machine, and the draft starts while the details are fresh. ### Give the agent the postmortem format What a good postmortem contains and how we structure it lives as **skills** and **memory** that travel with the agent: the section layout (timeline, root cause, impact, action items), the house style, and patterns from past incidents worth checking against. When we change how we write postmortems, we update the file and the next draft follows it. ### Connect the channel, the logs, and GitHub Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent reads: - **The incident channel** — the message timeline, who did what, and when the incident opened and resolved. - **The logs** — error-rate and latency spikes over the incident window, to line up against the timeline. - **GitHub** — the deploys and merges in the same window, to correlate what shipped with when things broke. The agent reads these to reconstruct the sequence; its one write is the doc PR. ### Set the guardrails The agent produces a **draft, not a decision**. Its output is a doc PR that a human reviews, edits, and merges: root cause and action items are proposed, never finalized by the agent. It doesn't publish a postmortem or assign owners on its own. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Open the draft for review With that in place, a resolved incident produces a first draft on its own: a timeline built from the channel and lined up against the deploys and log spikes, a proposed root cause with the evidence behind it, and a list of candidate action items. The team opens the PR, corrects what the agent inferred, adds the context only they have, and merges the version they stand behind. > **The pattern** > A resolved incident on the **channel** trigger spawns a session with read-only > **connectors** into the channel, the logs, and GitHub. The format lives as > **skills** and **memory**. The agent correlates the timeline and opens a doc > PR; humans finalize. ## Guardrails The agent reads an incident channel, logs, and deploy history and writes a document others will act on, so the access is scoped and contained: - **Isolation.** Every incident runs in its own isolated sandbox. The session reads the timeline, correlates the deploys and logs, and only the drafted doc PR is written back out. - **Scoped secrets.** The channel, logs, and GitHub credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **PR-gated.** The postmortem lands as a doc PR, not a published document. A human reviews the timeline, edits the root cause and action items, and owns the merge. - **Everything is code.** The postmortem format and the per-tool permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every incident:** A first draft the moment it resolves - **PR-gated:** The agent drafts, the team reviews and finalizes - **3 systems:** Incident channel, logs, and GitHub in one agent Resolved incidents now come with a drafted postmortem waiting as a doc PR, its timeline already correlated against the deploys and log spikes. The team spends its time judging the root cause and deciding the action items rather than reconstructing what happened, and fewer incidents close with no writeup at all. --- <!-- /markdown/use-cases/interview-scheduler.md --> # How we coordinate interview loops The interview-scheduling agent we run on TaskStation — daily, it finds candidates ready to schedule in Greenhouse, checks interviewer availability on Google Calendar, and drafts a slot proposal to the candidate and the calendar invites for a coordinator to confirm. Canonical page: https://taskstation.co/use-cases/interview-scheduler Once a candidate clears a screen, scheduling the next round becomes its own small project: read the interview plan to see who's on the panel, find a time when every one of those interviewers is actually free, propose it to the candidate, and get invites onto everyone's calendar before the slot goes stale. None of it is hard, but it's tedious across several interviewers and several candidates at once, and every day it sits undone is a day the loop stalls. We run an interview-scheduling agent on TaskStation that does the coordination every day. For every candidate marked ready to schedule in Greenhouse, it reads the interview plan, checks the panel's availability on Google Calendar, and drafts a slot proposal to the candidate plus the calendar invites for each interviewer. It never sends anything and never touches a hiring decision — a coordinator reviews every proposal and invite before it goes out. This is how we keep our own interview loops moving. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Greenhouse · Google Calendar · Email - **Mode:** Proposes and drafts only · coordinator confirms and sends ## The problem Scheduling an interview loop means combining two things that live in different systems: the interview plan — who's on the panel, how many rounds, how long each one runs — sits in Greenhouse, while whether those people are actually free sits in their calendars. Checking both by hand, for every candidate who clears a screen, is the kind of coordination work that's easy to fall behind on the moment more than one requisition is open at once. The usual workarounds don't hold up. A shared scheduling link only works if every interviewer keeps it current, and it says nothing about who the panel actually is for this candidate. A coordinator juggling several calendar tabs and a Greenhouse tab in parallel can do it, but it doesn't scale past a handful of loops, and a slipped invite reads as disorganized to a candidate who's interviewing elsewhere too. ## What we built On TaskStation, a daily cron triggers an agent. It spawns a fresh session (a cloud sandbox) that reads every candidate marked ready to schedule in Greenhouse, pulls that candidate's interview plan to see who's on the panel, checks each panelist's Google Calendar for open slots, and drafts a slot proposal for the candidate along with a calendar invite per interviewer. Every draft lands with the coordinator; the agent sends nothing and decides nothing about the candidate. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox, so the set of candidates ready to schedule and the panel's real availability are both recomputed from the current state — nothing carries over from the day before. ### Give the agent the scheduling conventions How we read an interview plan, how many slot options to propose, how far out to look for availability, and what a good candidate-facing proposal reads like all live as **skills** and **memory** that travel with the agent. When we change how loops are structured, we update the file and the next day's proposals follow it. ### Connect Greenhouse, Google Calendar, and email Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read Greenhouse** — which candidates are marked ready to schedule, and each one's interview plan: the panel, the round, and the duration. - **Read Google Calendar** — the availability of every interviewer on the plan, so a proposed time is one every panelist can actually make. - **Draft to email** — a slot proposal to the candidate and a calendar invite per interviewer, held for the coordinator. ### Set the guardrails The agent's job stops at the draft. It never emails a candidate directly, never sends a calendar invite, and never makes or implies a hiring decision — no rejecting a candidate, no advancing a stage, no offer. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Hand the batch to the coordinator With that in place, each morning brings a batch of ready-to-schedule candidates already matched against real interviewer availability: a proposed slot for the candidate and a drafted invite for each panelist. The coordinator reviews, confirms the time, and sends — or adjusts first. Nothing reaches a candidate or an interviewer's calendar without them. > **The pattern** > A daily **cron** spawns a fresh session with **connectors** into Greenhouse for > the interview plan and Google Calendar for panel availability. The scheduling > conventions live as **skills** and **memory**. The agent drafts the candidate > proposal and the invites and stops — the coordinator confirms and sends. ## Guardrails The agent touches a candidate's next step in the hiring process, so the boundary between "proposed" and "confirmed" is the control that matters: - **Isolation.** Every run happens in its own isolated sandbox. Only the drafted proposal and invites are written back out. - **Scoped secrets.** The Greenhouse and Google Calendar credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** The agent never sends the candidate proposal and never sends a calendar invite. Every draft is held for the coordinator to review, adjust, and send. - **No hiring decisions, ever.** The agent cannot reject a candidate, advance a stage, or send an offer. It schedules the next conversation; it never decides whether that conversation happens. - **Everything is code.** The scheduling conventions and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Ready-to-schedule candidates matched to real panel availability - **Zero:** Invites or candidate emails sent without coordinator review - **3 systems:** Greenhouse, Google Calendar, and email in one agent Scheduling an interview loop now starts with a proposal and a set of invites already drafted against the panel's real calendars, instead of a coordinator opening five tabs to find one time that works. The coordinator's day starts with a review instead of a hunt, and every send and every hiring call still stays with a person. --- <!-- /markdown/use-cases/investor-update.md --> # How we draft our monthly investor update A monthly agent we run on TaskStation — connected to Postgres, Stripe, and last month's update. It pulls the core metrics, compares them to last month and to plan, and drafts the update in our format for a founder to finalize. Canonical page: https://taskstation.co/use-cases/investor-update The monthly investor update is a small, recurring task that eats a founder's time. The numbers live in a few different places, they have to be pulled and compared to last month and to plan, and then the whole thing has to be written up in a consistent format. None of it is hard; it's just an hour or two of gathering and formatting that comes around every month. We run an agent on TaskStation that does the gathering and the first draft. This is how we draft our own investor update, including the connections and guardrails involved. - **Team:** TaskStation - **Runs on:** A monthly cron - **Connected systems:** Postgres · Stripe · Last month's update - **Mode:** Read-only · a founder finalizes and sends ## The problem Writing the update every month means pulling MRR and revenue from Stripe, active accounts and growth from the product database, and burn and runway from the finance numbers, then setting each against last month and against plan. The data is spread across systems, the comparisons are done by hand, and the write-up has to match the format investors are used to seeing. The common fixes are incomplete. A BI dashboard shows the current numbers but doesn't write the narrative or compare against plan. A saved template still needs someone to fill in every figure. Doing it by hand each month is reliable but it's the founder's time going into gathering and formatting rather than into the commentary that actually matters. ## What we built On TaskStation, a monthly cron triggers an agent. It spawns an isolated session (a cloud sandbox) with read-only access to the product database, Stripe, and last month's update. It pulls the core metrics — MRR, growth, burn, runway, active accounts — compares them to last month and to plan, and drafts the update in our usual format as a document. A founder edits it and sends it. ## How it works ### Trigger the draft on a monthly cron A **cron trigger** fires once a month and spawns a fresh **session** in its own sandbox. Each run pulls the current month's numbers and produces one draft. One run maps to one session on one disposable machine, and nothing carries over between months except what's read from the source systems. ### Give the agent our format and what matters The shape of our update lives as **skills** and **memory** that travel with the agent: which metrics we report, how we define each one, the format and section order investors expect, and the plan targets to compare against. Last month's update is the reference for tone and structure. When we change how we report, we write it down and the agent follows it next month. ### Connect Postgres, Stripe, and last month's update Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Query Postgres** — active accounts, growth, and the product metrics we track. - **Read Stripe** — MRR and revenue for the month. - **Read last month's update** — the prior numbers to compare against and the format to match. - **Draft into a document** — the numbers and the narrative assembled in our format for a founder to edit. ### Set the guardrails The agent is **read-only** across every connected system: it reads the database, Stripe, and last month's update, and it writes nothing back to any of them. Its one output is a draft document. A founder finalizes and sends; the agent never sends anything. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Hand a founder a finished first draft With that in place, the start of each month produces a draft with the metrics pulled, the month-over-month and against-plan comparisons filled in, and the narrative written in our format. The founder edits the commentary, checks the numbers, and sends it. The gathering and formatting are done; the judgment stays with a person. > **The pattern** > A monthly **cron** spawns a session with read-only **connectors** into Postgres, > Stripe, and last month's update. Our format and metric definitions are encoded as > **skills** and **memory**. The agent drafts; a founder finalizes and sends. ## Guardrails The agent reads revenue and product data to draft a document, so the access is scoped and contained: - **Isolation.** Every run happens in its own per-task isolated sandbox. The session reads only the systems it's scoped to, and only the draft document is written back out. - **Scoped secrets.** The Postgres and Stripe credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only.** The agent reads every source and writes back to none of them. Its only output is a draft, and it never sends. A founder owns the send. - **Everything is code.** The agent's configuration, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every month:** A finished first draft ready at the start of the month - **Read-only:** Nothing written back to any source system - **3 systems:** Postgres, Stripe, and last month's update in one agent The monthly update now arrives as a draft with the numbers pulled, the comparisons done, and the narrative written in our format. The founder spends their time on the commentary and the send rather than on gathering and formatting, and the data is only ever read. --- <!-- /markdown/use-cases/lead-follow-up.md --> # How we follow up with every inbound lead The sales agent we run on TaskStation — connected to HubSpot, email, and Google Calendar. It researches each new lead, drafts a personalized follow-up, and offers a call slot, with a human approving the send. Canonical page: https://taskstation.co/use-cases/lead-follow-up Inbound leads have a short half-life. A form filled out on Tuesday is worth far more if the follow-up goes out Tuesday than if it goes out Friday, but a good follow-up takes research — who the company is, what they do, why they might have signed up — and that research is exactly what gets skipped when the pipeline is full. We handle this by putting an agent on every new lead. It researches the company, drafts a personalized follow-up, and offers a time to talk. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails. - **Team:** TaskStation - **Connected systems:** HubSpot · Email · Google Calendar - **Trigger:** Scheduled HubSpot sweep (every 15 min) - **Mode:** Trigger-driven · human-gated ## The problem The follow-up that converts is the one that shows we read the form and understood the company. That takes a few minutes of research per lead, and a few minutes times every inbound is more time than a growing team has. So follow-ups either go out generic or go out late, and both lose deals. Templates make the sending fast but not the personalization, which is the part that matters. Booking the call adds another round of back-and-forth on top. The work that moves a lead forward is the work that keeps slipping. ## What we built Our lead pipeline in **HubSpot** is connected to an agent running on TaskStation. A scheduled sweep spawns a fresh, isolated session — a cloud sandbox — with scoped access to what a follow-up needs: the lead record, the web for research, our email, and the team calendar. It researches each new lead, drafts a personalized message, and proposes a call slot. A human approves before it sends. ## How it works ### Connect HubSpot as the trigger A scheduled **trigger** sweeps HubSpot every 15 minutes for new leads, and each firing spawns a fresh **session** in its own sandbox. A sweep that turns up several new leads at once still runs as a single session, but each lead is handled as an independent unit — a research or drafting failure on one lead never blocks or corrupts the others. Sessions don't share state between sweeps: the marker written back onto each HubSpot record is what carries forward. ### Give the agent our playbook How we follow up lives as **skills** and **memory** loaded into every session: our positioning, the tone we write in, what a good qualifying question looks like, and follow-ups that landed before. The agent writes to that standard rather than inventing one, and the memory updates as deals move. ### Connect what a follow-up can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Research the company** — it reads the lead's site and public information to understand what they do and why they signed up. - **Read and update HubSpot** — full lead context in, research notes back on the record. - **Draft the email** — a personalized follow-up written to our playbook, ready to send. - **Offer a slot on Google Calendar** — it checks real availability and proposes a specific time to talk. ### Set the guardrails The agent researches and drafts freely, but nothing goes out on its own: every follow-up stops at a **human approval gate** as a draft for a person to review and send. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each lead run With that in place, a new lead arrives already worked: the company researched, the record updated, a personalized email drafted, and a call slot proposed against real availability. The person on sales reviews the draft, adjusts if they want, and sends — instead of starting each follow-up from a blank page. > **The pattern** > Connect the CRM via a **trigger** on new leads, give the agent scoped > **connectors** into HubSpot, email, and the calendar, encode the sales playbook > as **skills** and **memory**, and gate every send behind a human. ## Guardrails Giving an agent the ability to reach out to leads is a brand question as much as a sales one. The relevant controls on TaskStation: - **Isolation.** Each sweep runs in its own isolated sandbox — a fresh session that can cover several new leads at once, each handled as an independent unit so a failure on one never blocks the rest. Only the drafts it produces are written back out. - **Scoped secrets.** The HubSpot, email, and Calendar credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** No message reaches a lead without a person reviewing the draft and sending it. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Every lead:** Researched and followed up, not just the easy ones - **Same-day:** Personalized follow-up ready while the lead is still warm - **3 systems:** HubSpot, email, and the calendar in one agent Every inbound lead now arrives with the research done, a personalized draft written, and a call slot offered — waiting for a person to approve rather than sitting untouched. The team spends its time deciding which leads to push instead of doing the same research over and over. The setup relies on four pieces: sandbox isolation per lead, a secrets manager to broker the HubSpot, email, and Calendar tokens, a human approval gate before anything reaches a lead, and memory that improves as deals move. --- <!-- /markdown/use-cases/lead-routing.md --> # How we route inbound leads to the right rep The lead-routing agent we run on TaskStation — every 15 minutes it scores each new HubSpot lead, assigns it to the right rep by territory, segment, or round-robin, and notifies the rep in Slack, flagging anything ambiguous for a human instead of guessing. Canonical page: https://taskstation.co/use-cases/lead-routing A new inbound lead is worth the most in the first few minutes after it lands. Every minute it sits unassigned is a minute a rep isn't calling it, and the minute it does get routed still depends on someone checking the queue, knowing which territory owns which region, and remembering whose turn it is in the rotation. That's a lot to ask of a manual process running every fifteen minutes, all day. We run a lead-routing agent on TaskStation that does exactly that check, every 15 minutes: read HubSpot for new leads, assign each one to the right rep, notify them in Slack, and flag anything it can't confidently route for a human. It never deletes or merges a lead — assignment and notification are the only actions it takes. - **Team:** TaskStation - **Runs on:** Cron, every 15 minutes - **Connected systems:** HubSpot · Slack - **Mode:** Fresh session · assigns + notifies, never deletes or merges ## The problem Routing rules live in more than one place: territory boundaries by region, segment specialists by company size or industry, and a round-robin pool for whatever's left over. A rep checking the lead queue has to hold all of that in their head, get it right every time, and still do it fast enough that the lead is still warm when the notification lands. The common fallback is a static routing workflow inside the CRM — a fixed if/then chain that breaks the moment a territory is split, a rep goes on leave, or a lead doesn't cleanly fit one bucket. Those leads get stuck in a queue, silently misrouted to whoever built the workflow's default, or followed up on days later once someone notices. None of that is fast, and none of it flags the leads that most need a human's judgment. ## What we built On TaskStation, a cron fires an agent every 15 minutes. It spawns a fresh session with scoped access to HubSpot, reads every new inbound lead that hasn't been routed yet, scores it for intent, matches it against the territory and segment rules, and falls back to whichever rep in the round-robin pool has taken the fewest leads so far this week. It assigns the HubSpot owner, posts a notification to the rep's Slack channel, and calls out anything high-intent so it gets worked first. A lead it can't confidently route — missing territory data, no segment match, a genuine tie — goes to a separate Slack channel for a human to assign by hand. It never deletes or merges a lead. ## How it works ### Run on a 15-minute cron A **cron trigger** fires the agent every 15 minutes. Each firing spawns a fresh **session** in its own sandbox — there's no persistent process and no local ledger. The HubSpot record itself is the memory: a routing-status property marks a lead as handled, so the next sweep only looks at what's genuinely new. ### Give the agent the routing rules The territory map, the segment-to-specialist mapping, and the round-robin pool live as a **skill** and **memory** that travels with the agent, not as a workflow buried in the CRM. When territories get redrawn or a rep joins the rotation, we update the skill and the very next sweep routes against the new rules — no CRM workflow to rebuild. ### Connect HubSpot and Slack Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent: - **Reads new leads from HubSpot** — lifecycle stage, region, company size, industry, title, and engagement signals. - **Writes the owner and a routing-status marker back to HubSpot** — the only writes it makes, and never a delete or a merge. - **Posts to Slack** — an assignment notification to the rep, a distinct flag for high-intent leads, and a separate escalation post for anything it can't confidently route. ### Set the guardrails The agent assigns and notifies strictly per the routing rules. It never deletes or merges a lead record under any circumstance, and it never guesses an owner for a lead that's ambiguous or doesn't match the rules — that lead goes to a human instead. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Assign, notify, and escalate With that in place, every 15 minutes brings one pass over the new-lead queue: each lead lands on the right rep's desk with the reason it was routed there, high-intent leads are called out so they get worked first, and anything genuinely ambiguous waits in a separate channel for a human to assign. No lead sits unrouted for more than one cycle without someone knowing why. > **The pattern** > A 15-minute **cron** spawns a fresh session with scoped **connector** access > to HubSpot and a **channel** into Slack. The territory, segment, and > round-robin rules live as **skills** and **memory**. The agent assigns and > notifies exactly per those rules, and hands anything it can't confidently > route to a human. ## Guardrails The agent writes to the lead record, so its write access is narrow and its escalation path is explicit: - **Never delete or merge.** The agent can set a lead's owner and a routing-status marker. It cannot delete a lead, merge two records, or touch any other field, no matter how confident the routing rules are. - **Ambiguous means human, not a guess.** A lead with no territory match, no segment match, and no clean tiebreak in the round-robin pool is flagged in a dedicated Slack channel for a person to assign. The agent never invents a fallback owner. - **Scoped secrets.** The HubSpot credential is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Everything is code.** The territory map, segment rules, round-robin pool, and intent-scoring criteria are files in the repo, versioned and changed through a reviewed **change request** rather than a CRM workflow builder. ## The outcome - **Every 15 min:** Sweep of the new-lead queue, all day - **0:** Leads deleted or merged by the agent — ever - **3-way routing:** Territory, segment, and round-robin in one pass Leads that used to sit in a queue waiting for someone to notice now land on the right rep's desk within 15 minutes, tagged with why they were routed there and whether they need to be worked first. The agent assigns and notifies; a human still owns every lead it can't confidently place. --- <!-- /markdown/use-cases/meeting-notes.md --> # How we turn meetings into notes and action items The meeting agent we run on TaskStation — connected to Google Calendar, the call transcript, and Linear. After each call it writes structured notes and files action items as tickets assigned to the right people. Canonical page: https://taskstation.co/use-cases/meeting-notes Most of what a meeting decides is lost within a day. The notes get taken by whoever remembers to, the action items live in someone's head, and the follow-up happens only if a person turns the conversation into tasks afterward — which is the step that gets skipped when the next meeting starts. We handle this by putting an agent on the calendar. After each call it writes structured notes and files the action items as tickets, assigned to the people who own them. This writes up how we run that on TaskStation — the connections, the steps, and the guardrails. - **Team:** TaskStation - **Connected systems:** Google Calendar · Transcript · Linear - **Trigger:** Meeting ends - **Mode:** Trigger-driven · assignee-routed ## The problem Turning a conversation into notes and tasks is real work, and it lands on whoever's least busy at the end of the call — which on a small team is nobody. Decisions get made, owners get named out loud, and then the meeting ends and none of it is written down. A week later the thing everyone agreed on hasn't started because it was never a task. Recording the call solves the record but not the follow-up: nobody rewatches an hour of video to find the three things they agreed to do. The gap is between the transcript and the tickets, and closing it by hand is exactly the chore that falls off. ## What we built Our calendar in **Google Calendar** is connected to an agent running on TaskStation. When a call ends, it spawns its own isolated session — a cloud sandbox — with scoped access to what note-taking needs: the event, the transcript, and Linear. It writes structured notes, pulls out the action items, and files each one as a ticket assigned to the person who owns it. ## How it works ### Connect the calendar as the trigger A signed webhook tied to Google Calendar points at the project. When a meeting ends, it fires, and each firing spawns a fresh **session** in its own sandbox, seeded with the event and its transcript. One meeting, one session, one disposable machine. Sessions don't share state, and back-to-back calls run as parallel sessions. ### Give the agent our conventions How we take notes lives as **skills** and **memory** loaded into every session: our notes structure, how we phrase an action item, who owns which area, and past meetings for context. The agent writes to that standard rather than inventing one, and the memory updates as projects move. ### Connect what note-taking can touch Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the event and transcript** — attendees, agenda, and the full record of what was said. - **Write structured notes** — decisions, discussion, and next steps in our format, attached where the team looks. - **File tickets in Linear** — each action item becomes a ticket assigned to its owner, linked back to the meeting. ### Set the guardrails The agent writes notes and files tickets, but it only creates — it doesn't close, reassign, or touch existing work. An action item it can't confidently assign gets flagged for a person rather than routed to a guess. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Let each meeting run With that in place, a finished call turns into a set of notes and a handful of tickets before the next meeting starts. "Sarah will handle the migration" becomes a ticket assigned to Sarah. "We decided to ship Friday" becomes a line in the notes. The follow-up exists as tasks instead of as a memory that fades. > **The pattern** > Connect the calendar via a **trigger** on meeting end, give the agent scoped > **connectors** into the transcript, the notes, and Linear, encode our notes > conventions as **skills** and **memory**, and route each action item to the > person who owns it. ## Guardrails Giving an agent the ability to file tickets and assign work is a trust question. The relevant controls on TaskStation: - **Isolation.** Each meeting runs in its own isolated sandbox. The session can read the transcript and context it needs, and only the notes and tickets it produces are written back out. - **Scoped secrets.** The Calendar and Linear credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** An action item the agent can't confidently assign is flagged for a person instead of routed to a guess. - **Everything is code.** The agent's persona, skills, and permissions are files in the repo — versioned and changed through a reviewed **change request**, not a dashboard setting. ## The outcome - **Every call:** Notes written and action items filed - **Assigned:** Tickets routed to the person who owns them - **3 systems:** The calendar, the transcript, and Linear in one agent The follow-up that used to depend on someone remembering now arrives as notes and tickets minutes after the call ends. The team reads a clean summary and finds their tasks already in Linear, instead of reconstructing the meeting from memory a week later. The setup relies on four pieces: sandbox isolation per meeting, a secrets manager to broker the Calendar and Linear tokens, a flag for anything the agent can't confidently assign, and memory that carries context forward across meetings. --- <!-- /markdown/use-cases/month-end-close.md --> # How we run our month-end close The finance agent we run on TaskStation — connected to Stripe, the close checklist and revenue ledger in Google Sheets, and Slack. Each month it walks the checklist, reconciles revenue against the ledger, flags anomalies and missing docs, and posts the open items — never touching a journal entry or closing the books itself. Canonical page: https://taskstation.co/use-cases/month-end-close Month-end close runs the same way every month and takes days anyway: walk a checklist of a few dozen items, reconcile revenue against the ledger, chase down whatever documentation is missing, and sanity-check the numbers against the trend before anyone signs off. None of it is hard. All of it is slow, and most of the time goes to confirming that the easy 90% is actually fine so the close team can focus on the 10% that isn't. We run a close agent on TaskStation that does that first pass every month. It reconciles our Stripe revenue against the ledger, works the rest of the checklist for missing docs, flags anything that looks off against prior months, and posts the open items to our finance Slack channel. It never writes a journal entry and it never marks a period closed — a human owns the close. - **Team:** TaskStation - **Runs on:** Monthly cron - **Connected systems:** Stripe · Google Sheets · Slack - **Mode:** Assemble and flag · human closes the books ## The problem The close checklist itself isn't the bottleneck — it's re-confirming the same categories of thing every month: does recognized revenue tie out to what Stripe actually processed, does every line have the supporting document it's supposed to, and does anything look different enough from the last few months to need an explanation before the numbers go out. Each check is quick in isolation. Doing all of them, consistently, on top of everything else close week involves, is where the days go. A reconciliation formula in the sheet catches the exact ties but not the near ones, and doesn't know what "looks anomalous" means without a person supplying that judgment fresh each time. So the checklist gets worked start to finish by hand, at the same cost, every single month, and it only happens as carefully as there's time for. ## What we built On TaskStation, a monthly **cron** triggers a close agent after the period closes. Each run spawns its own isolated session — a cloud sandbox — with read access to Stripe revenue and to the close checklist and ledger tabs in a Google Sheet. It reconciles Stripe revenue against what the ledger recorded, works the rest of the checklist for missing supporting docs, compares this month's figures against the trailing months for anomalies, and posts a close-status summary with every open item to Slack. It assembles and flags; it never enters a journal line and never marks the close done. ## How it works ### Run on a monthly cron, after the period closes A **cron** trigger fires the agent once a month, once the prior period has closed. Each firing spawns a fresh **session** in its own sandbox — the checklist and ledger sheet are the record, so nothing needs to carry over in memory between runs. ### Give the agent the close standard What counts as a reconciled line, what tolerance a timing lag gets, which checklist items need which supporting document, and how big a swing versus the trailing months counts as an anomaly worth flagging — all of it lives as **skills** and **memory** the agent loads before touching a single line, so every month is judged to the same standard. ### Connect what the close needs, read and flag only Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read Stripe** — revenue, refunds, and payouts for the closed period. - **Read the ledger tab** — what was recorded for the period, to reconcile against Stripe. - **Read and flag the checklist tab** — the status of every close item, and the notes and flags the agent adds; it never touches the ledger's recorded entries. - **Post to Slack** — the close-status summary with every open item. ### Set the guardrails The agent reconciles and flags; it does not close. It never writes a journal entry, never edits a recorded ledger amount, and never checks off the period as closed — that decision, and every action it implies, belongs to a person. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Post the open items With that in place, each month produces a checklist worked end to end and one Slack post: what's still unreconciled, which items are missing documentation, and which figures moved enough against the trend to need a look — each with enough detail that a person can act on it without re-pulling the numbers. The finance team reviews the list and closes the books themselves. > **The pattern** > A monthly **cron** spawns a fresh session with read access to Stripe and > read-and-flag access to the checklist and ledger in Google Sheets. The > reconciliation tolerances and anomaly rules live as **skills** and > **memory**. The agent assembles and flags every open item; a human closes > the books. ## Guardrails Touching a company's close process is a trust question first, so the access is scoped tightly and one direction only: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to the systems it's scoped to, and only the checklist flags and the Slack post are written back out. - **Scoped secrets.** The Stripe and Sheets credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **No journal entries, no closing the books.** The agent never writes a journal line, never edits a recorded ledger amount, and never marks a checklist item or period as closed. It flags; a human closes. - **Everything is code.** The agent's checklist rules, tolerances, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every month:** Checklist worked end to end on the current data - **Flag only:** No journal entries, no period ever marked closed by the agent - **One post:** Every open item — unreconciled, missing doc, anomaly — in one place The days that used to go to confirming the easy 90% of the checklist now go to the handful of items the agent actually flagged. The close still runs on the same schedule whether or not anyone has gotten to it yet, and the open-items list is waiting in Slack when the finance team sits down to close the books. --- <!-- /markdown/use-cases/nda-turnaround.md --> # How we turn NDAs around in minutes, not days The NDA agent we run on TaskStation — every 15 minutes it reviews a new inbound NDA against our standard playbook, redlines the non-standard terms in the document, and drafts a flag summary for counsel. It never signs or sends; counsel does. Canonical page: https://taskstation.co/use-cases/nda-turnaround An NDA is the most standardized document a deal touches, and also the one most likely to sit in an inbox for two days waiting for someone with a law degree to open it. The deal doesn't move until it's signed, but the review itself is usually mechanical: the same eight or nine clauses, checked against the same standard positions, over and over. We run an NDA-turnaround agent on TaskStation that watches for inbound NDAs and gives each one its first pass within fifteen minutes of arriving. It redlines what deviates from our standard positions and drafts a flag summary for counsel. It never signs anything and never sends anything to the counterparty — counsel does both. This is how we keep NDAs from being the thing a deal waits on. - **Team:** TaskStation - **Runs on:** Every 15 minutes - **Connected systems:** Gmail · Google Drive · Google Docs - **Mode:** Redline + flag only · counsel signs ## The problem NDAs are high volume and low variance — most deals need one, and most of them follow the same handful of templates with the same handful of predictable deviations: a one-way instead of mutual, a perpetual confidentiality term, a non-solicit that reaches further than it should. None of that requires a first read from a lawyer. It just requires someone to check. But "someone" means counsel, and counsel's queue is full of things that actually need judgment. An NDA that's 90% standard waits behind a contract renegotiation, and the deal it's gating waits with it. By the time it gets reviewed, the delay has nothing to do with the document's complexity and everything to do with queue position. ## What we built On TaskStation, a **cron** fires every 15 minutes and checks a watched Gmail label for a new inbound NDA. Each firing spawns a fresh session — a cloud sandbox — that saves the attached document into a Drive folder, opens it in Docs, checks it clause by clause against our standard NDA playbook, redlines every deviation as a suggested edit in the document, and drafts — never sends — a flag summary on the original email thread addressed to counsel. Counsel opens the thread to a redline and a summary instead of a blank document. ## How it works ### Run on a 15-minute cron A **cron trigger** checks the watched Gmail label every 15 minutes. Each firing spawns a fresh **session** in its own sandbox, seeded with whatever NDA is waiting. One NDA maps to one run on one disposable machine — nothing carries over between runs beyond what's already sitting in Gmail and Drive. ### Give the agent our standard positions Our standard NDA positions live as a **skill** loaded into every session: mutual vs. one-way, confidentiality duration, governing law, non-solicit scope, remedies, the residuals clause, return-or-destroy obligations, assignment, and indemnification. The agent checks against that standard instead of a generic sense of what an NDA should say. ### Connect what the review needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent: - **Reads the watched label in Gmail** — to find a new inbound NDA that hasn't already been processed. - **Saves the document to Google Drive** — into a dedicated NDA folder, so there's a durable copy to review and redline. - **Redlines in Google Docs** — every non-standard clause gets a suggested edit, never a direct change. ### Set the guardrails The agent's only outputs are a **Docs suggestion** and a **Gmail draft** — never a sent reply, never an applied edit, never a signature. Counsel reviews both and decides what goes back to the counterparty. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Counsel picks it up already redlined With that in place, an NDA that lands at 9:03am has a redline and a flag summary waiting by 9:15 — the routine terms noted, the non-standard ones marked with the section and the reason, a proposed replacement already drafted. Counsel reviews, adjusts if needed, and sends. Nothing reaches the counterparty without that review. > **The pattern** > A 15-minute **cron** spawns a fresh session with scoped **connectors** into > Gmail, Drive, and Docs. Our standard positions live as a **skill**. The agent > redlines and flags; it never signs, executes, or sends. ## Guardrails An agent that touches an unsigned legal document needs a hard line between drafting and executing: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to Gmail, Drive, and Docs, and only the redline and the draft leave it. - **Scoped secrets.** The Gmail, Drive, and Docs credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Suggesting mode only.** Redlines land as Docs suggested edits, never a direct write to the document, and the original attachment is never overwritten, deleted, or moved. - **Never sign, execute, or send.** The flag summary is a Gmail draft on the original thread, never sent. The agent has no path to countersign an NDA or message the counterparty — counsel owns both. - **Everything is code.** The agent's playbook, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **15 minutes:** Longest an inbound NDA waits before its first pass - **Redline only:** Every change is a suggested edit, never applied - **3 systems:** Gmail, Drive, and Docs in one agent NDAs stop being the paperwork a deal waits on. Counsel opens a thread to a redline already marked against our standard positions and a summary of exactly what's non-standard and why, instead of a cold document. The agent does the first pass; counsel still signs off on every word that goes back to the counterparty. --- <!-- /markdown/use-cases/nps-analysis.md --> # How we turn NPS responses into themes A weekly cron reads new NPS/CSAT survey responses from a Google Sheet, clusters them into themes with sentiment and detractor drivers, and posts the score trend and representative quotes to Slack. Canonical page: https://taskstation.co/use-cases/nps-analysis NPS and CSAT surveys generate a response every time someone answers, and most of it goes unread. The score gets tracked in a dashboard, but the comment next to it — the actual reason someone gave a 4 instead of a 9 — sits in a spreadsheet row and never gets aggregated with the hundred other comments saying the same thing. By the time a theme is obvious to a human skimming the sheet, it's been building for months. We run an NPS analysis agent on TaskStation that reads new survey responses every week and posts the themes, sentiment, and detractor drivers to our customer-success Slack channel, alongside how the score moved since last week. It only reads the survey sheet; the single output is the Slack post. This is how we keep our own NPS program from turning into an unread spreadsheet. - **Team:** TaskStation - **Runs on:** Weekly cron - **Connected systems:** Google Sheets · Slack - **Mode:** Read-only · one Slack post per week ## The problem Every NPS or CSAT response is really two signals: a number and a comment. The number is easy to track — most survey tools already chart it. The comment is where the "why" lives, and it's the part that gets lost. Free-text feedback piles up in a spreadsheet export, one row per response, and nobody reads all of it every week, so the same complaint from a dozen different detractors never gets counted as one thing. The common approaches don't fix this. A dashboard shows the score moving without saying why. Someone skimming the sheet manually catches the loudest complaint, not the most common one, and does it inconsistently from week to week. And because nothing tracks the theme over time, a driver that's been growing for a month looks identical to one that appeared once and went away. ## What we built On TaskStation, a weekly cron triggers an agent. It spawns an isolated session (a cloud sandbox) with read-only access to the Google Sheet the survey tool exports responses into, reads the full response history, clusters the free-text comments into themes, tags each response's sentiment and score band, isolates what's actually driving detractors this week, and computes how the score has moved week over week. It posts one summary to Slack with representative quotes. It writes nothing back to the sheet. ## How it works ### Run on a weekly cron A **cron trigger** fires the agent once a week. Each firing spawns a fresh **session** in its own sandbox. There's no memory carried between runs — the sheet itself is the record, so the agent recomputes themes, sentiment, and the score trend from the full response history every time. ### Give the agent the analysis rules How we cluster feedback and read the score lives as **skills** and **memory** that travel with the agent: what counts as a detractor driver versus a one-off complaint, how to pick a representative quote for a theme, where the promoter, passive, and detractor bands sit, and how to describe a week-over-week move. When we refine what a theme should look like, we write it down and the next run follows it. ### Connect the survey sheet and Slack Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent reads: - **Survey responses from Google Sheets** — the full export of NPS/CSAT scores, comments, and timestamps, read-only. - **Posts to Slack** — one weekly summary: the score and its trend, the top themes with quotes and counts, and the leading detractor drivers. ### Set the guardrails The agent is **read-only** on the survey sheet — it never edits a row, adds a column, or writes back a score. It never contacts a respondent. Its only output is the Slack post, and that post is a report, not an action. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Post the weekly summary With that in place, each Monday brings one Slack post: the current NPS/CSAT score and how it moved since last week, the themes behind the comments with a representative quote and count for each, and the drivers pulling detractors down this week. The customer-success team reads it and decides what, if anything, to act on. > **The pattern** > A weekly **cron** spawns a fresh session with a read-only **connector** into the > survey Google Sheet. The clustering and scoring rules live as **skills** and > **memory**. The agent reads the full history every time and writes nothing but > the Slack post. ## Guardrails The agent reads a live survey export, so its access is scoped and one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to the sheet it's scoped to, and only the Slack post is written back out. - **Scoped secrets.** The Google Sheets and Slack credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only, report-only.** The connector into the survey sheet is read-only — the agent cannot edit a response, add a row, or write back a score — and it never reaches out to a respondent. It reports; it doesn't act. - **Everything is code.** The agent's clustering rules, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every week:** Themes and score trend recomputed from the full response history - **Read-only:** Nothing written back to the survey sheet - **1 post:** Score trend, themes, and detractor drivers in one Slack message Survey comments that used to sit unread in a spreadsheet now arrive every week as a set of named themes with quotes and counts, next to the score's actual week-over-week move and what's driving it. The agent only reads and reports; the team still decides what to do about each theme. --- <!-- /markdown/use-cases/office-snacks.md --> # How we keep the office stocked People request snacks through a Slack shortcut, and an agent batches the requests weekly, prepares an order, and posts it for the office manager to approve before placing anything. Canonical page: https://taskstation.co/use-cases/office-snacks Keeping an office stocked is a small recurring chore that never quite has an owner. Requests come in over Slack, hallway conversations, and sticky notes; the office manager reconciles them into an order; and half the time a request is forgotten by the time the order goes in. It's low-stakes, but it's steady work. We handle it with an agent on TaskStation. People request snacks through a Slack shortcut, and once a week the agent batches every request, prepares an order, and posts it for the office manager to approve. Nothing is placed until a person signs off. - **Team:** TaskStation - **Control surface:** Slack - **Connected systems:** Slack · Ordering account - **Mode:** Weekly batch · human-gated ## The problem Snack requests arrive one at a time and out of band. Someone asks in a channel, someone else mentions it in passing, and the office manager is left assembling a list from memory. Requests get dropped, duplicates slip through, and the order goes in later than it should. The task is simple but constant. It needs one place to collect requests, a way to batch them on a schedule, and a person's sign-off before any money is spent — none of which a shared spreadsheet or a recurring calendar reminder actually provides. ## What we built Requests come in through a Slack shortcut, so there's one place to submit them. Once a week a cron trigger spawns an isolated session that collects the week's requests, prepares an order against the ordering account, and posts the draft in Slack for the office manager to approve. Only after approval does it place anything. ## How it works ### Collect requests through a Slack shortcut Slack is connected as a **channel**, and a shortcut is the way people submit requests. Each submission is collected against the week's batch. There's one place to ask, and nobody has to track requests in their head. ### Batch them on a weekly cron A **cron trigger** fires once a week and spawns a fresh **session** in its own isolated sandbox. The run gathers every request submitted since the last order. One run, one sandbox, torn down when it's done. ### Give the agent the ordering rules The preferred vendor, the budget, standing staples to always include, and how to consolidate duplicate requests live as **skills** and **memory**. The rules are updated as preferences and the budget change. ### Connect Slack and the ordering account Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read requests from Slack** — the week's submissions from the shortcut. - **Prepare an order in the ordering account** — build the cart against the vendor, but not check out. - **Post to Slack** — the draft order for the office manager to review. ### Approve before it places anything The prepared order stops at a **human approval gate** in Slack. The office manager sees the full cart and the total, and nothing is purchased until they approve. Credentials for the ordering account are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. > **The pattern** > Collect requests through a Slack **channel** shortcut, batch them on a **cron > trigger**, keep the vendor and budget rules in **skills** and **memory**, and hold > the order at a **human approval gate** before it spends. One place in, one approved > order out. ## Guardrails Because the agent can prepare an order that spends money, the controls matter even for a small chore: - **Isolation.** Each weekly run executes in its own isolated sandbox, and only the prepared order and Slack post it's explicitly allowed to send are written back out. - **Scoped secrets.** The ordering-account credential is encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Human approval gates.** No purchase is placed until the office manager approves the draft order in Slack. - **Everything is code.** The vendor, the budget, the staples, and the batching schedule are files in the repo — versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Weekly:** Requests batched into one order on a schedule - **One place:** Every request comes in through a Slack shortcut - **Approve first:** Nothing is purchased without a person signing off Requests stop getting lost, duplicates get consolidated, and the office manager reviews a finished order instead of assembling one. The spending step stays behind a person, so the automation saves the busywork without ever placing an order on its own. --- <!-- /markdown/use-cases/oncall-triage.md --> # How we triage on-call alerts before paging a human The triage agent we run on TaskStation — connected to Sentry, our logs, and GitHub. It works up a first-pass diagnosis on every alert and only pages a human when it can't resolve it. Canonical page: https://taskstation.co/use-cases/oncall-triage An alert fires at 3am. Before anyone can act on it, someone has to wake up, pull the stack trace, check what deployed recently, grep the logs, and work out whether this is a real incident or noise. Most of that is mechanical, and most of it happens before the human has any context — which is the worst time to be doing it. We run a triage agent on TaskStation that does the first pass. When an alert fires, the agent gathers the stack trace, the recent deploys, and the correlated logs, posts a first-pass diagnosis to the incident channel, and only pages a human when it can't resolve the alert or the severity is high. This write-up covers how the setup works: the trigger, the session model, and the guardrails. - **Team:** TaskStation - **Runs on:** Every alert, as it fires - **Connected systems:** Sentry · Logs · GitHub - **Mode:** Trigger-driven · pages a human only when needed ## The problem The first minutes of an incident are spent gathering context, not fixing anything. Which error is this, when did it start, what shipped just before, is it one user or all of them — the on-call engineer answers these by hand, half-awake, before they can even judge whether the page was warranted. The common fixes are incomplete. Paging on every alert burns the on-call rotation on noise and trains people to ignore the pager. Tuning thresholds cuts the noise but also hides real regressions. A runbook helps, but someone still has to be awake to follow it. None of it does the gathering for you. ## What we built On TaskStation, each alert triggers an agent. When Sentry fires, the alert spawns an isolated session (a cloud sandbox) with scoped, read-only access to Sentry, the logs, and GitHub. The agent pulls the stack trace, lists the deploys since the error first appeared, correlates the logs around the spike, and posts a first-pass diagnosis to the incident channel. It pages a human only when it can't resolve the alert or the severity is high. ## How it works ### Connect the alert as the trigger A signed webhook from Sentry points at the project. Every alert fires it, and each firing spawns a fresh **session** in its own sandbox, seeded with the alert payload. One alert maps to one session on one disposable machine, so nothing carries over between incidents and concurrent alerts triage in parallel. ### Give the agent the triage playbook How we triage lives as **skills** and **memory** that travel with the agent: which services are noisy, what a real regression looks like versus a known flake, the severity rules for when to page, and past incidents with their root causes. When an alert turns out to be benign, we write it down and the agent recognizes it next time. ### Connect the systems triage needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the Sentry issue** — the stack trace, the frequency, and the first-seen timestamp, to place the error in time. - **Correlate the logs** — pull the log lines around the spike and line them up against the trace. - **Check recent deploys on GitHub** — the commits and PRs that shipped just before the error appeared, to surface a likely cause. - **Post to the incident channel** — the diagnosis, the suspected deploy, and the evidence, as a single message. ### Set the guardrails The agent is **read-only**: it investigates across Sentry, the logs, and GitHub, but it does not deploy, roll back, or change anything. High-severity alerts and anything it can't resolve page a human straight away — the diagnosis is attached, not a substitute for the page. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each alert triage itself With that in place, an alert gathers its own context: the trace, the deploy window, the correlated logs, and a first-pass diagnosis land in the channel within the first minute. A known-benign spike is closed with the reasoning attached. A high-severity or unresolved alert pages the on-call engineer with the work already done, so they open the page to context instead of a blank terminal. > **The pattern** > A **trigger** on every alert spawns a session with scoped, read-only > **connectors** into Sentry, the logs, and GitHub. The triage playbook is encoded > as **skills** and **memory**. The agent diagnoses first and pages a human only > when it can't resolve the alert or the severity is high. ## Guardrails The agent reads production telemetry, so the access is scoped and contained: - **Isolation.** Every alert runs in its own isolated sandbox. The session can pull traces, logs, and deploy history to build the diagnosis; only the posted message and any page are written back out. - **Scoped secrets.** The Sentry, logging, and GitHub credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** The agent never deploys or rolls back. High-severity and unresolved alerts page a human, who owns any action taken. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **First-pass:** Diagnosis in the channel before the pager fires - **Less noise:** Benign alerts closed with reasoning, not paged - **3 systems:** Sentry, logs, and deploy history in one agent The mechanical first minutes of an incident happen before anyone is paged, and when the agent does page, it pages with the trace, the suspect deploy, and the correlated logs attached. The on-call engineer spends their time deciding and fixing rather than gathering context half-awake. --- <!-- /markdown/use-cases/outbound-outreach.md --> # How we run personalized outbound An agent connected to our CRM, enrichment, and email that researches each lead, drafts a genuinely personalized sequence, logs it to the CRM, and holds every send behind human approval with a daily cap. Canonical page: https://taskstation.co/use-cases/outbound-outreach Personalized outbound doesn't scale by hand, so most teams give up the personalization: a mail-merge template with a first name swapped in, sent to a list. It's fast and it's ignored. Real personalization means researching each account and writing to its actual context, which is exactly the part that doesn't scale. We run an agent on TaskStation that does the research and the writing per contact, then holds every send behind a person. It enriches each lead, drafts a first-touch and follow-up grounded in the account's real context, logs every draft to the CRM, and sends nothing without approval, capped daily. This is how we run our own outbound. - **Team:** TaskStation - **Runs on:** A new segment or lead list - **Connected systems:** CRM · Enrichment · Email - **Mode:** Every batch approved by a person · daily cap ## The problem Outbound has a tradeoff most teams resolve the wrong way. Personalized messages land, but researching and writing each one doesn't scale, so the list wins: a template, a merge field, a send button. The result is volume without relevance, and the reply rate shows it. The usual tools don't fix it. A sequencing tool blasts the same template on a schedule. An AI writer produces fluent copy with no real account context, which reads as personalized but isn't. And any tool that can mass-send is one wrong segment away from putting unreviewed email in front of thousands of people. ## What we built On TaskStation, a periodic sweep triggers the agent in a fresh session (a cloud sandbox) with scoped access to enrichment, the CRM, and email. The agent works through the batch of new contacts one by one, researches each account, drafts a first-touch and a follow-up sequence grounded in what it found, and logs every draft to the CRM. No message sends on its own: every batch waits for a person to approve, under a daily cap. ## How it works ### Connect the lead list as the trigger A new segment or lead list in the CRM is the trigger. On a schedule, the agent fires a fresh **session** in its own sandbox and works through the batch of contacts that list contains. Each contact is handled as an independent unit — a research or drafting failure on one contact is logged and skipped, and never blocks the rest of the batch — with the run capped so a sweep always stays a size a person can review. ### Give the agent the outreach playbook What genuine personalization looks like lives as **skills** and **memory** that travel with the agent: our angles, the proof points that resonate, what to avoid, and which signals in an account are worth writing to. When a message pattern works, we write it down and the agent reuses it on the next batch. ### Connect enrichment, the CRM, and email Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Enrich each contact** — role, company, recent signals, and the real context that makes a message specific. - **Draft a genuine sequence** — a first-touch and follow-ups grounded in the account's context, not a template with a name swapped in. - **Log every draft to the CRM** — each message written back against the contact, so the history is complete. - **Send on approval** — the email goes out only after a person clears the batch. ### Set the guardrails Nothing sends without a person approving it, and it never mass-sends unreviewed. Every batch stops at a **human approval gate**, and a **daily cap** limits how many messages can go out even once approved. A wrong segment can't turn into thousands of unreviewed emails, because the send is gated and capped regardless of how big the list is. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Run personalized outbound at a safe rate With that in place, loading a segment produces a batch of researched, account- specific drafts logged to the CRM and waiting for review. A person reads the batch, approves what's ready, and the daily cap meters the sends. The personalization is real because each message is written to a researched account, and nothing leaves without a person clearing it. > **The pattern** > A **trigger** on a new list fires a fresh session with scoped > **connectors** into enrichment, the CRM, and email. The outreach approach is > encoded as **skills** and **memory**. Every batch is approved by a person and > the daily cap meters the sends — nothing mass-sends unreviewed. ## Guardrails The agent can draft and send email, so the send is the tightly controlled step: - **Isolation.** Each sweep runs in its own fresh isolated sandbox, torn down when it finishes. The session reaches only enrichment, the CRM, and email, and only drafted messages leave it; nothing sends from inside the sandbox. A failure on one contact is logged and skipped — it never blocks or corrupts the rest of the batch. - **Scoped secrets.** The enrichment, CRM, and email credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** Nothing sends without a person approving it. Every batch is reviewed before send, and a **daily cap** limits volume even after approval, so a wrong list can never mass-send unreviewed. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every send:** Approved by a person and metered by a daily cap - **Per account:** Research and copy grounded in real context, not a merge field - **3 systems:** Enrichment, CRM, and email in one agent Outbound is now personalized per account and sent at a safe rate: each contact is researched, each message is written to its context and logged to the CRM, and every batch is cleared by a person under a daily cap. The personalization scales because the research does, and nothing mass-sends without review. --- <!-- /markdown/use-cases/payment-recovery.md --> # How we recover failed subscription payments An hourly agent tracks every failed Stripe invoice on a per-subscription ledger, escalating from a smart retry through a payment reminder, an update-your-card notice, and a final notice, alerting the revenue team on Slack and stopping the moment the invoice is paid. Canonical page: https://taskstation.co/use-cases/payment-recovery Most churn isn't a customer deciding to leave. It's a card that expired, a bank that flagged a charge, a payment that failed for a reason that has nothing to do with whether the customer still wants the product. Left alone, that failed invoice quietly turns into a canceled subscription. This is involuntary churn, and it's recoverable if someone catches it fast and follows up consistently. We run a payment-recovery agent on TaskStation that watches every failed Stripe invoice, works it through a fixed escalation ladder — smart retry, reminder, update-your-card notice, final notice — and stops the instant the invoice is paid. It never cancels a subscription, issues a credit, or refunds anything; those decisions stay with a human. - **Team:** TaskStation - **Runs on:** Hourly cron, reusable session - **Connected systems:** Stripe · Slack · Email - **Mode:** Smart-retry + dunning only · escalation, never cancellation ## The problem A failed payment isn't one event, it's the start of a countdown. Stripe will retry automatically for a while, but a generic retry schedule doesn't know that this customer just had a fraud hold, or that another one hasn't opened an email in a week. Handled with a single blunt reminder, some customers churn who would have paid if asked the right way at the right time; handled with silence, the subscription lapses and support finds out only when the customer complains that their access disappeared. The common approaches don't hold state well. Stripe's built-in retry logic is useful but generic — it doesn't loop in a human, escalate the tone as time passes, or tell the revenue team which accounts are about to fall off a cliff. A spreadsheet tracking who got which email decays within a week. And a blanket policy of "cancel after N failures" trades a recoverable customer for a clean queue. ## What we built On TaskStation, an hourly cron re-prompts one **persistent session** that keeps a per-subscription ledger of where every failed invoice sits on the recovery ladder: smart retry, payment reminder, update-your-card notice, or final notice. Each run it reads Stripe for failed invoices and subscription state, advances any subscription whose wait time has elapsed to the next rung, sends the matching email, and posts a summary of everything in flight to the revenue team's Slack channel. Any subscription that pays at any point is closed out on the ledger and never contacted again. ## How it works ### Run hourly on a reusable session A **cron trigger** fires every hour, but unlike a fresh-session use case, it re-prompts the **same session** each time. The ledger — which subscription is on which rung, and when it's next due to escalate — lives in that session's memory, so the agent always knows what it already sent and never repeats or skips a step. ### Read failed invoices from Stripe Through a scoped **connector**, brokered server-side so no raw key reaches the model, the agent reads failed invoices, subscription status, and payment-method state from Stripe. This is the only signal it needs to decide what happens next; the read covers both new failures and subscriptions already mid-ladder. ### Give the agent the escalation ladder The order and timing of the ladder live as a **skill**: attempt a smart retry first, then a friendly payment reminder, then a firmer "update your card" notice with a link to update the payment method, then a final notice — each rung gated on a minimum wait since the last one, so no customer gets two emails in the same day. ### Send the dunning email and alert the revenue team The agent sends the rung-appropriate **email** to the customer directly — this part doesn't wait for approval, because a smart retry or a dunning email is reversible and expected. It also posts to **Slack**: what advanced today, what's now on final notice, and what recovered since the last run. ### Stop at the guardrail, not at cancellation Reaching the final notice does not trigger a cancellation, a refund, or a credit. The ladder tops out there and the account sits, flagged in the Slack summary, waiting for a person on the revenue team to decide the next step. Credentials for Stripe and email are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Close the loop the moment it's paid Every run checks whether a ladder subscription has since paid. The moment it has, the agent marks it recovered on the ledger, stops emailing it, and reports it in the Slack summary as a win — no further action, no lingering reminder. > **The pattern** > An hourly cron re-prompts one **reusable session** holding a per-subscription > ledger. A read-mostly Stripe **connector** supplies the state; a **skill** > encodes the escalation ladder; **email** and **Slack** are the only outputs. > The ladder stops at a final notice — cancellation, credits, and refunds > always wait for a human. ## Guardrails The agent can retry a payment and send an email on its own; anything that moves money or ends a subscription is out of its hands: - **Smart-retry and dunning only.** The agent may retry a failed charge and send the four ladder emails. It never cancels a subscription, issues a credit, or processes a refund — those require a human, no exceptions. - **One rung at a time.** The skill enforces a minimum wait between rungs so no customer is double-messaged, and every subscription escalates on its own schedule rather than everyone moving in lockstep. - **Stops on payment.** The moment an invoice clears, the ledger closes it out and no further email goes out. - **Scoped secrets.** The Stripe credential is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Everything is code.** The ladder, the wait times, and the agent's grants are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Hourly:** Every failed invoice re-checked against the ladder - **4 rungs:** Smart retry → reminder → update-card → final notice - **0:** Cancellations, credits, or refunds issued by the agent Failed payments that used to sit unattended until a subscription quietly lapsed now work through a consistent, escalating recovery ladder the moment they happen, with the revenue team seeing exactly what's in flight and what needs their decision. The agent chases the payment; the people decide when to give up on it. --- <!-- /markdown/use-cases/phishing-triage.md --> # How we triage reported phishing emails The phishing-triage agent we run on TaskStation — every 15 minutes it pulls newly reported emails from our phishing-report inbox, inspects headers, links, and attachments for phishing indicators, and posts a risk verdict with a recommended action to our security channel. It only recommends; a human decides whether to block a sender or warn staff. Canonical page: https://taskstation.co/use-cases/phishing-triage Security awareness training tells every employee to forward anything suspicious to a phishing-report inbox, and they do — which means that inbox fills up faster than one person can open each message, view the raw headers, hover over every link, and decide whether it's real. The obvious spam never gets that far; the gateway already caught it. What lands in the report inbox is the ambiguous case, the one that needs an actual look, and during a real campaign it arrives in a burst. We run a phishing-triage agent on TaskStation that checks that inbox every 15 minutes, analyzes each reported email, and posts a verdict and a recommended action to our security Slack channel. It reads mail and it recommends; it never blocks a sender, deletes a message, or touches anything else in Gmail. - **Team:** TaskStation - **Runs on:** Every 15 minutes - **Connected systems:** Gmail · Slack - **Mode:** Read-only on Gmail · one Slack alert per report ## The problem A secure email gateway filters out the obvious junk, so what a human forwards to the phishing-report inbox is already the hard tier: a plausible-looking invoice, a lookalike login page, a reply-to that doesn't match the display name. Judging it well means checking SPF/DKIM/DMARC results, tracing where a shortened link actually resolves, and noticing an attachment that's a macro document dressed up as a PDF — several checks, every time, before anyone can say whether it's real. The common approaches don't hold up under volume. A person checks the inbox when they get to it, so a real campaign sits unflagged for hours. A static blocklist of known-bad domains misses this morning's freshly registered lookalike. A rule that auto-quarantines anything reported is fast but wrong often enough that people stop trusting it — and a false block on a legitimate sender is its own incident. ## What we built On TaskStation, a cron fires every 15 minutes and spawns a fresh agent session with read-only access to the phishing-report Gmail inbox. It pulls every message reported since the last check, examines the sender's authentication results, the reply-to and return-path against the display name, where every link actually resolves, and any attachments, then classifies the risk and drafts a verdict — block this sender, warn everyone who may have gotten the same email, or no action needed — and posts it to the security Slack channel. It changes nothing in Gmail itself. ## How it works ### Run on a 15-minute cron A **cron trigger** fires the agent every 15 minutes. Each firing spawns a fresh **session** in its own sandbox, so a burst of reports during an active campaign gets picked up on the next tick instead of waiting on a person, and nothing carries state between runs. ### Give the agent the indicator rulebook What counts as a red flag lives as a **skill**: how to read SPF/DKIM/DMARC results, what a reply-to/return-path mismatch means, how to trace a shortened or embedded link to its real destination, which attachment types are high-risk, and how those signals combine into a risk tier and a recommended action. ### Connect the report inbox read-only Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent reads: - **The phishing-report Gmail inbox** — every newly reported message, its full headers, its links, and its attachments. - **Posts to Slack** — one verdict per reported email, with the risk tier, the indicators that drove it, and the recommended action. ### Set the guardrails The agent **analyzes, recommends, and alerts — nothing more.** It never blocks a sender, deletes or moves a message, or takes any remediation action in Gmail or anywhere else. The recommended action is exactly that: a recommendation, for a person on the security team to act on. ### Post the verdict With that in place, every reported email produces one Slack post within 15 minutes: a risk tier, the specific indicators found — the failed authentication, the mismatched link, the suspicious attachment — and a recommended action. Security reads it and decides whether to block the sender, warn affected staff, or close it out as benign. > **The pattern** > A 15-minute **cron** spawns a fresh session with a read-only **connector** > into the phishing-report Gmail inbox. The indicator rulebook lives as a > **skill**. The agent classifies and recommends; it never blocks, deletes, or > remediates. ## Guardrails The agent reads live employee reports of suspected attacks, so its access is scoped and strictly one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to the phishing-report inbox and the security channel; nothing else is written back out. - **Scoped secrets.** The Gmail credential is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Recommend-only.** The agent analyzes, classifies, and recommends. It never blocks a sender, deletes or quarantines mail, or takes any other remediation action — every action it names is for a human to carry out. - **Everything is code.** The indicator rulebook, the risk tiers, and the per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **15 min:** Worst-case delay before a reported email gets a verdict - **Read-only:** Nothing blocked, deleted, or changed in Gmail automatically - **3 signals:** Headers, links, and attachments checked on every report Reports that used to wait for someone to have a free moment now get a verdict within 15 minutes, with the specific indicators laid out and a recommended action attached. The agent never touches a sender or a mailbox; the security team still makes every call, just with the analysis already done. --- <!-- /markdown/use-cases/pipeline-hygiene.md --> # How we keep deals moving The pipeline-hygiene agent we run on TaskStation — it scans HubSpot every day for stale deals, missing next steps, deals stalled in a stage, and overdue tasks, nudges the owning rep in Slack, and escalates the worst offenders to their manager. Canonical page: https://taskstation.co/use-cases/pipeline-hygiene A pipeline decays quietly. A deal goes a week without a call or an email and nobody notices. A rep forecasts a close date and never sets one. A deal sits in "Proposal Sent" for a month past when it should have moved. A follow-up task comes due and slides by unflagged. None of this shows up in a stage-by-stage pipeline report — the deal is still "open," still counted, still apparently fine. We run a pipeline-hygiene agent on TaskStation that reads HubSpot every day and flags exactly this: deals gone quiet, missing next steps and close dates, deals stuck in their current stage, and tasks past due. It nudges the owning rep in Slack and escalates the worst of the day to their manager. It is read-mostly — it flags and it nudges, and it never touches a deal's stage, owner, or amount. - **Team:** TaskStation - **CRM:** HubSpot - **Runs on:** Daily cron - **Mode:** Read-mostly · nudge + flag only ## The problem This is a different problem than dirty CRM data. Duplicate contacts and missing job titles are a data-quality issue; a deal that hasn't moved in three weeks is a process issue — the pipeline itself is not doing what it's supposed to do, which is progress toward a close. A stage-by-stage pipeline report can't tell the difference between a deal that's actively advancing and one that's been quietly parked, because both show up the same way: "open," sitting in a stage, counting toward the forecast. The usual fix is the weekly pipeline review, where a manager scrolls every open deal looking for the ones that have gone stale. By the time that meeting happens, a deal that went quiet on Monday has already lost a week, and the review only catches what someone remembers to ask about. Reps are heads-down selling, not auditing their own deal hygiene, and a next-step field or a close date left blank rarely gets attention until forecast time, when it's too late to fix cheaply. ## What we built A daily cron triggers an agent that spawns a fresh, isolated session with read-mostly access to HubSpot. It pulls every open deal, checks it against four hygiene rules — no logged activity in the configured window, a missing next step or close date, no stage movement past the configured window, and any overdue task tied to the deal — and flags what it finds. It nudges each owning rep in Slack with exactly what needs attention on their deals, and separately escalates the day's worst offenders to the sales manager's channel. It never changes a deal's stage, reassigns its owner, or edits its amount. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox — one day, one run, nothing carried over. The pipeline is re-scanned from its current state every morning. ### Give the agent the hygiene rules What counts as stale, stalled, or overdue lives as a **skill** that travels with the agent: the no-activity window, the stage-stall window, what a missing next step or close date looks like, and how the day's worst offenders are chosen for escalation. When the team's definition of "stuck" changes, the rule changes, not the agent's judgment call. ### Connect HubSpot read-mostly Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent reads deals, stage history, activities, and tasks, and writes only a hygiene flag back onto the deal — no other field: - **Deals and stage history** — current stage, amount, close date, next step, and how long the deal has sat in its current stage. - **Activities** — calls, emails, meetings, and notes logged against each deal, to determine staleness. - **Tasks** — open tasks tied to each deal and whether they're past due. - **Posts to Slack** — a nudge to the owning rep and an escalation to the manager's channel. ### Set the guardrails The agent is **read-mostly**. The only write it ever makes to HubSpot is an internal hygiene flag on the deal — never the stage, never the owner, never the amount. It cannot reassign a deal, advance it, or touch the number tied to it, no matter what the data suggests. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Nudge the rep, escalate the worst Each flagged deal gets a Slack nudge to its owning rep: what's wrong and what to do about it — log an activity, set a next step, move the deal, or clear the overdue task. At the end of the run, the deals in the worst shape — several rules tripped at once, largest amount at risk, longest gone quiet — get escalated to the sales manager's Slack channel. The rep and the manager decide what happens next; the agent only reports. > **The pattern** > A daily **cron** spawns a session with read-mostly **connector** access to > HubSpot. The hygiene rules live as a **skill**. The agent flags stale, > incomplete, stalled, and overdue deals, nudges the rep, escalates the worst > to the manager, and never changes a deal's stage, owner, or amount. ## Guardrails Reading and flagging every open deal every day is a trust question, so the agent's access is narrow and one-directional in the ways that matter: - **Isolation.** Every run happens in its own isolated sandbox. Only the hygiene flag write and the Slack messages are written back out. - **Scoped secrets.** The HubSpot credential is encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Read-mostly.** The agent's only write to HubSpot is an internal hygiene flag on the deal. It never changes a deal's stage, reassigns its owner, or edits its amount — those decisions stay with the rep and the manager. - **Nudge and flag only.** The agent's output is a Slack message or a flag; it never takes an action on a rep's behalf. - **Everything is code.** The staleness window, the stage-stall window, and the escalation criteria are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Open pipeline re-checked against the current HubSpot state - **4 rules:** Stale, missing fields, stalled stage, and overdue tasks in one pass - **Read-mostly:** Stage, owner, and amount are never touched by the agent The deals that used to sit quietly until the next pipeline review now surface the same day they go stale, with a nudge to the rep who owns them and an escalation for the ones a manager should already be watching. The agent only flags and nudges; the rep and the manager still decide what moves. --- <!-- /markdown/use-cases/pr-review-nudge.md --> # How we keep code review from stalling The pull-request review agent we run on TaskStation — it flags PRs stuck past a review SLA, gone stale, or sitting on unaddressed change requests, then nudges the author or reviewer in Slack with exactly what's blocking. Canonical page: https://taskstation.co/use-cases/pr-review-nudge A PR that sits unreviewed for three days doesn't announce itself. It just sits, under a growing pile of newer PRs, until the author pings someone directly or gives up and context-switches to something else. Multiply that across a team and a review SLA becomes a suggestion nobody enforces, and "requested changes" becomes a state a PR can live in indefinitely. We run a PR review nudge agent on TaskStation that reads every open pull request across our repos once a day and posts a Slack nudge for anything stuck: past the review SLA, gone stale, or sitting on unaddressed change requests. It only reads GitHub; the single output is the Slack nudge. It never merges, closes, or approves anything. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** GitHub · Slack - **Mode:** Read-only on GitHub · nudge + summarize only ## The problem A review SLA is easy to write down and hard to enforce. Nobody is watching the review queue full-time, so a PR opened Monday morning and still unreviewed Wednesday afternoon looks the same in the repo list as one opened five minutes ago — until someone happens to scroll far enough to notice. Stale PRs are worse: a branch with no commits, no comments, and no reviews in a week has usually just been forgotten, and it silently rots until a rebase conflict forces someone to deal with it. And "requested changes" is a state, not an alert — a reviewer asks for changes, the author sees it, life happens, and the PR sits there looking open while actually being blocked. The common approaches don't close this loop. GitHub's own notifications are easy to mute and don't distinguish a PR that's five minutes old from one that's five days old. A weekly standup catches the loudest blockers, not the quiet ones. Manually skimming the PR list is thorough the first time and skipped the second. ## What we built On TaskStation, a daily cron triggers an agent. It spawns a fresh session with read-only access to GitHub through the `gh` CLI, pulls every open PR across our repos, and classifies each one against three rules: awaiting review past the SLA, stale with no activity in N days, or carrying requested changes the author hasn't addressed. For each PR that trips a rule, it posts a Slack nudge to the person who owns the next move — the requested reviewer if it's overdue for review, the author if it's stale or has unaddressed feedback — with a one-line summary of what's actually blocking it. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox. One day maps to one run, so the review state is recomputed from GitHub's current state every time — nothing carries over from yesterday's nudges. ### Give the agent the SLA rules What counts as "past SLA," "stale," and "unaddressed" lives as a **skill** that travels with the agent: the review-SLA clock, the staleness window, how to tell a requested-changes review has actually been addressed versus just sitting there, and who to nudge for each case. ### Connect to GitHub read-only Through the `gh` CLI, authenticated with a scoped **GH_TOKEN** secret injected at runtime, the agent reads: - **Open pull requests** across the configured repos — age, author, requested reviewers, and current state. - **Reviews and comments** on each PR — whether a review was submitted, what it said, and whether the author has pushed or replied since. - **Posts to Slack** — the nudge, with the PR link and what's blocking it. ### Set the guardrails The agent has **no write access to GitHub**. It cannot merge a PR, close it, approve it, or dismiss a review. Its only output is the Slack nudge. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Nudge the right person With that in place, each morning brings a Slack nudge for every PR that's stuck: who it's nudging, why (overdue for review, stale, or unaddressed changes), how long it's been that way, and a link straight to the PR. The team decides what to do next — review it, ping the reviewer directly, or close it themselves. > **The pattern** > A daily **cron** spawns a fresh session that reads open PRs through the `gh` > CLI with a scoped **GH_TOKEN**. The SLA, staleness, and unaddressed-changes > rules live as a **skill**. The agent reads everything and writes nothing but > the Slack nudge. ## Guardrails The agent reads across every repo it's pointed at, so its access is scoped and strictly one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to GitHub (read-only) and Slack; nothing else is written back out. - **Scoped secret.** The `GH_TOKEN` is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Nudge and summarize only.** The agent never merges, closes, or approves a PR, and never dismisses or resolves a review. It reports what's blocking and leaves the decision to a human. - **Everything is code.** The SLA thresholds, the staleness window, and the agent's GitHub permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Every open PR rechecked against the current GitHub state - **3 checks:** Overdue for review, stale, and unaddressed changes, in one pass - **Zero:** Merges, closes, or approvals made by the agent A review queue that used to require someone remembering to scroll through it now surfaces its own blockers every morning, in Slack, addressed to whoever owns the next move. The agent only reads and nudges; the team still decides what happens to every PR. --- <!-- /markdown/use-cases/qa-agent.md --> # How we QA every pull request automatically The QA agent we run on TaskStation — connected to GitHub and our test environment. It checks out each PR, runs the suite, exercises the change, and posts the result. Canonical page: https://taskstation.co/use-cases/qa-agent Code review catches problems a person can find by reading a diff. It misses the ones you only find by running the change: a test that passes locally but flakes in CI, a path that works in the happy case and 500s on an edge case, a migration that reads fine but locks a table under load. These tend to surface in staging or production, after the PR is approved. We run a QA agent on TaskStation that does this checking when the PR opens. This is how we QA our own changes, including the connections and guardrails involved. - **Team:** TaskStation - **Runs on:** Every PR, on open and on push - **Connected systems:** GitHub · Test environment · Cloudflare - **Mode:** Trigger-driven · result posted on the PR ## The problem Some bugs don't show up in a diff. A reviewer reads the code and approves, but no one has checked out the branch, run the suite, hit the new endpoint, and watched what happens. CI runs the tests the author wrote; it doesn't cover the paths they missed. The common fixes are incomplete. Green CI shows the existing tests pass, not that the change is correct. A manual QA pass is thorough but slow and lands after the review. A generic AI reviewer only sees the diff; it can't run the branch, so it misses failures that only appear at runtime. ## What we built On TaskStation, each PR triggers an agent. On open and on every push, the PR spawns an isolated session (a cloud sandbox) with scoped access to the branch, the test suite, a deployable test environment, and the Cloudflare edge in front of it. The agent checks out the change, runs the suite, exercises the new behavior on a live deploy, and posts a pass/fail result on the PR. ## How it works ### Connect GitHub as the trigger A signed GitHub webhook points at the project. Every PR opened or updated fires it, and each firing spawns a fresh **session** in its own sandbox, seeded with the branch under test. One PR maps to one session on one disposable machine, so nothing carries over between runs and concurrent PRs run in parallel. ### Give the agent the branch and the test playbook The session checks out the branch and installs it clean. Our QA conventions live as **skills** and **memory** that travel with the agent: how to run the suite, which flows are critical, edge cases that have caused problems before, and what a result should contain. When a bug slips through, we write it down and the agent picks it up on the next run. ### Connect the systems QA needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Run the full suite** — unit, integration, and e2e inside the sandbox, with the failure output captured in full. - **Deploy to the test environment** — it stands the change up on an ephemeral deploy and exercises it end-to-end against the new paths. - **Reach the edge via Cloudflare** — it checks behavior through the edge (routing, headers, caching, redirects), not just localhost. - **Report on GitHub** — the pass/fail result and any reproduction post as a check and a comment on the PR. ### Set the guardrails The agent operates against the **test environment only**; production is out of scope. It does not merge or deploy to prod. Its output is a result, and a human owns the merge. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Let each PR arrive pre-QA'd With that in place, a new PR checks itself out, runs the suite, deploys to the test environment, exercises the change through Cloudflare, and posts a result: green when it passes, or a red check with the failing command, the logs, and steps to reproduce. A flaky test is flagged with the evidence. A broken edge route is caught before staging. > **Summary** > A **trigger** on every PR spawns a session with scoped **connectors** into the > branch, the test suite, the test environment, and Cloudflare. The QA playbook is > encoded as **skills** and **memory**. The agent stays on the test environment > and a human owns the merge. ## Guardrails The agent has access to the test environment and the edge, so the access is scoped and contained: - **Isolation.** Every PR runs in its own isolated sandbox on its own branch. The session can install, deploy, and exercise the change to reproduce a failure; only the reported result is written back out. - **Scoped secrets.** The GitHub, test-environment, and Cloudflare credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Test environment only.** The agent's access stops at the test plane: no production access, no prod deploy, no merge. It reports; the team decides. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every PR:** Run, deployed, and exercised before human review - **Pre-merge:** Runtime failures caught before they reach staging - **4 systems:** Branch, suite, test environment, and edge in one agent Runtime failures that a diff can't show now surface on the PR when it opens, with the failing command, the logs, and the repro attached. Reviewers spend their time on design rather than checking out branches by hand, and the changes that reach staging have already been run against the test environment. --- <!-- /markdown/use-cases/qbr-prep.md --> # How we draft QBR decks for key accounts The QBR-prep agent we run on TaskStation — connected to Postgres, HubSpot, Google Slides, and Google Docs. Weekly, for every account due a quarterly business review, it pulls usage trends, account health, support activity, and expansion signals into a draft deck and briefing doc for the CSM to review and present. Canonical page: https://taskstation.co/use-cases/qbr-prep A quarterly business review is only as good as its prep. The account's usage trend lives in the product database, its health score and CSM notes live in HubSpot, the support history is scattered across a quarter's worth of tickets, and the expansion angle is usually in someone's head. Pulling all of that into a deck by hand takes an afternoon per account, and with a book of a few dozen accounts, QBR week turns into a slide-building sprint instead of an account-strategy exercise. We run a QBR-prep agent on TaskStation that does that assembly every week. It reads across the source systems, drafts a deck and a companion briefing doc per account due a review, and stops there. It never writes back to a source system and never presents anything — the CSM opens the draft, edits the story, and walks into the room with it. - **Team:** TaskStation - **Runs on:** Weekly cron - **Connected systems:** Postgres · HubSpot · Google Slides · Google Docs - **Mode:** Read-only across sources · draft deck + briefing doc per account ## The problem The material for a QBR is real but scattered. Product usage trends sit in Postgres. Account health, CSM notes, and the deal history sit in HubSpot. Support friction is a quarter's worth of tickets no one has time to reread. Expansion opportunities are pattern-matched from all of the above, usually from memory, right before the meeting. Templates help with layout but not with assembly — someone still has to pull every number, read every note, and decide what's worth a slide. Doing that for one account is an afternoon. Doing it for every account due a review, every week, is the part that gets skipped, rushed, or handed to whoever has the least on their plate that week — which is exactly when a CSM walks in under-prepared for a renewal conversation. ## What we built On TaskStation, a weekly cron triggers an agent. It spawns an isolated session with read-only access to Postgres and HubSpot, finds every account whose next QBR falls due in the coming window, and for each one pulls usage trends, value delivered, account health, a support summary, and expansion opportunities. It assembles that into a draft deck in Google Slides and a companion briefing doc in Google Docs with the backup detail, then stops. Nothing is shared, presented, or sent — the CSM reviews both and takes it from there. ## How it works ### Run on a weekly cron A **cron trigger** fires the agent once a week. Each firing spawns a fresh **session** in its own sandbox — this run doesn't remember last week's, so it recomputes who's due and what their numbers look like from the current state of the data every time. ### Give the agent the QBR structure What goes in a QBR deck, and in what order, lives as a **skill** that travels with the agent: how we define a usage trend, what "value delivered" means for our product, how account health rolls up, how to summarize a quarter of support activity without burying the one issue that matters, and what counts as a credible expansion signal versus a guess. ### Connect the sources read-only, and the deck read-write Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent: - **Reads usage trends from Postgres** — activity and adoption over the quarter, per account. - **Reads account and CSM data from HubSpot** — health score, CSM notes, support ticket history, and the deal record, to find who's due and why. - **Writes the draft deck to Google Slides** — the QBR deck itself, one per due account. - **Writes the briefing doc to Google Docs** — the backup detail behind every slide, for the CSM to reference when a question goes deeper than the deck. ### Set the guardrails Postgres and HubSpot are **read-only** — the agent cannot change a usage record, a health score, a note, or a deal, even though it reads deeply from all of them. Its only writes are the new deck and the new doc it creates. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Land the draft, stop there Each due account ends the run with a draft deck and a briefing doc under a shared folder, named for the account and the quarter. The agent reports where each one landed and stops. The CSM reviews the story, edits the commentary, and is the one who opens the meeting and presents it. > **The pattern** > A weekly **cron** spawns a session with read-only **connectors** into > Postgres and HubSpot. The deck structure lives as a **skill**. The agent > reads the account's quarter and writes exactly two drafts — a deck and a > briefing doc — then stops for the CSM. ## Guardrails The agent reads deeply across account data, so its access is scoped and its output is bounded: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to the systems it's scoped to, and only the new deck and doc are written back out. - **Scoped secrets.** The Postgres and HubSpot credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only across the source systems.** Postgres and HubSpot are read-only connectors. The agent cannot change a usage record, a health score, a CSM note, or a deal — it can only draw from them. - **Draft only, nothing sent.** The deck and briefing doc are drafts. The agent never shares, presents, or emails either one — the CSM reviews and presents. - **Everything is code.** The deck structure, the skill, and the per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Weekly:** Every account due a review restaged before the CSM's meeting - **Read-only:** Nothing written back to Postgres or HubSpot - **2 drafts:** One deck and one briefing doc per account, every time The afternoon of pulling numbers and rereading tickets now happens automatically, every week, for every account that needs it. The CSM opens a draft deck that already has the trend, the health picture, the support summary, and a candidate expansion angle, edits the story into their own words, and presents it themselves. --- <!-- /markdown/use-cases/release-notes.md --> # How we generate release notes from merged PRs The release-notes agent we run on TaskStation — connected to GitHub. On each release it reads the merged PRs since the last one, groups them, writes the notes, and opens a changelog PR. Canonical page: https://taskstation.co/use-cases/release-notes Release notes are the first thing that gets skipped under a deadline. The information is all there in the merged PRs, but turning it into something a reader can follow — grouped by area, in plain language, with the noise dropped — is manual work that lands after the release is already out. We run a release-notes agent on TaskStation that does it from the PR history. On each release, the agent reads every PR merged since the last release, groups them by area, writes human-readable notes, and opens a PR to the changelog for review. This write-up covers how the setup works: the trigger, the session model, and the guardrails. - **Team:** TaskStation - **Runs on:** Every tag or release - **Connected systems:** GitHub - **Mode:** Trigger-driven · changelog PR opened for review ## The problem Writing release notes means reading back through everything that merged since the last release, deciding what a reader cares about, and phrasing it for someone who wasn't in the PRs. It's the kind of task that's easy to defer, so the changelog falls behind or gets a one-line summary that helps no one. The common fixes are incomplete. An auto-generated list of PR titles is accurate but unreadable — internal wording, no grouping, every dependency bump included. Writing them by hand is better but slow, and it competes with actually shipping. Either way the notes arrive after the release, if they arrive at all. ## What we built On TaskStation, each release triggers an agent. When a tag is pushed, the release spawns an isolated session (a cloud sandbox) with scoped access to the repository. The agent finds the previous release, reads every PR merged in between, groups them by area, writes notes in plain language, and opens a PR to the changelog. A human reviews and merges. ## How it works ### Connect the release as the trigger A signed GitHub webhook points at the project. A pushed tag or published release fires it, and each firing spawns a fresh **session** in its own sandbox, seeded with the new tag. One release maps to one session on one disposable machine, so nothing carries over between releases. ### Give the agent the notes playbook How we write release notes lives as **skills** and **memory** that travel with the agent: how to group changes by area, which labels mark internal-only work to drop, the voice the changelog is written in, and the format the file expects. When we adjust how a section reads, we write it down and the agent follows it on the next release. ### Connect the systems the notes need Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Find the range** — locate the previous release tag and the commit range up to the new one. - **Read the merged PRs** — titles, descriptions, labels, and authors for every PR merged in that range, to understand what actually changed. - **Group and write** — cluster the PRs by area and write plain-language notes, dropping internal churn and dependency noise. - **Open a changelog PR on GitHub** — the drafted notes land as a pull request against the changelog file. ### Set the guardrails The agent's output is a **PR against the changelog**, nothing more — it does not publish, tag, or announce. A human reviews the wording and merges. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Let each release draft its own notes With that in place, a pushed tag reads its own history: the agent finds the range, reads the merged PRs, groups them by area, writes the notes, and opens a changelog PR — grouped, readable, with internal churn dropped. The team reviews the wording instead of assembling the notes from scratch. > **The pattern** > A **trigger** on every release spawns a session with a scoped **connector** into > the repo. The notes playbook is encoded as **skills** and **memory**. The agent > drafts a changelog PR and a human owns the wording and the merge. ## Guardrails The agent reads the repository and writes a draft, so the access is scoped and contained: - **Isolation.** Every release runs in its own isolated sandbox. The session reads the PR history and drafts the notes; only the changelog branch is written back out. - **Scoped secrets.** The GitHub credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **PR-gated.** The agent opens a pull request against the changelog and stops. It never publishes or announces the release; a human reviews the wording and merges. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every release:** Notes drafted from the PR history automatically - **Grouped:** Changes clustered by area with internal churn dropped - **Human review:** The agent drafts; the team owns the wording The changelog stops falling behind, and the notes that reach review are already grouped and readable instead of a raw list of PR titles. Reviewers edit wording rather than reconstructing what shipped from the git history. --- <!-- /markdown/use-cases/renewal-manager.md --> # How we get ahead of contract renewals The renewal manager we run on TaskStation — a daily cron that surfaces every account 90, 60, and 30 days from renewal, preps a packet on usage and value delivered, drafts outreach, and flags at-risk accounts for the owner to review. Canonical page: https://taskstation.co/use-cases/renewal-manager Renewals used to surface the way most bad news does — late. An account owner would open HubSpot, notice a contract expiring in three weeks, and start from zero: pull the usage numbers, remember what shipped since the last check-in, guess at expansion room, and write the outreach the same afternoon. The accounts that renewed smoothly were the ones an owner happened to remember to check. The ones that churned quietly were usually the ones nobody looked at until it was too late to do anything but apologize. We run a renewal manager agent on TaskStation that watches every HubSpot deal's renewal date and gets ahead of it — at 90, 60, and 30 days out — with a packet already built and outreach already drafted. It never sends anything and never touches price. It hands the account owner a head start instead of a deadline. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** HubSpot · Google Calendar · Slack · Email - **Mode:** Prep + draft only · account owner sends ## The problem A renewal date sitting on a HubSpot deal is just a date. It doesn't say whether the account is healthy, whether the team has actually used what they bought, whether there's a natural expansion to raise, or whether the deal has gone quiet in a way that should worry someone. That context lives scattered across old notes, a memory of the last QBR, and whatever the owner can reconstruct under time pressure. The common fallback is a calendar reminder or a pipeline report sorted by close date — a flat list with no prep behind it. By the time an owner opens the account, the renewal is close, the packet doesn't exist yet, and the outreach gets written in a rush instead of shaped around what the account actually needs to hear. ## What we built On TaskStation, a daily cron triggers an agent that reads every HubSpot deal's renewal date and surfaces the ones crossing 90, 60, or 30 days out. For each one, it builds a renewal packet — usage and value delivered since the last renewal, and one or two concrete expansion ideas — checks Google Calendar to see whether a renewal conversation is already on the books, and drafts the outreach email. Anything that looks stalled, quiet, or shrinking gets flagged separately as at-risk. The packet and the draft wait for the account owner; nothing goes out on its own. ## How it works ### Run on a daily cron A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox with no memory of yesterday's run — the deal list and every signal are re-read from HubSpot itself, and a property on each deal record marks which renewal window it's already been surfaced for, so the same account doesn't get a duplicate packet every day it sits inside 90 days. ### Give the agent the renewal playbook What belongs in a renewal packet, how to read usage and engagement as a health signal, what counts as at-risk, and how the team phrases renewal outreach live as a **skill** that travels with the agent. When we learn what actually moves a renewal, we update the skill and every account benefits from it the next day. ### Connect HubSpot and Google Calendar Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent: - **Reads deals from HubSpot** — renewal dates, contract value, stage history, notes, and past activity, to find every account crossing a 90/60/30-day window and to spot the ones going quiet. - **Reads Google Calendar** — read-only, to check whether a renewal or QBR conversation is already scheduled near the renewal date, so the outreach can reference it instead of duplicating it. - **Writes back to HubSpot** only the window it last surfaced for each deal — never the stage, the amount, or anything a rep would use to track the actual negotiation. ### Post to Slack, draft to email Two channels, two purposes: a **Slack** post to the account owner with the day's renewal radar and any at-risk flags, and an **email** draft — the renewal packet and the outreach copy — held for that owner to open, edit, and send. ### Set the guardrails The agent preps and drafts; it never contacts a customer and never touches price. It cannot send the outreach email, and it cannot apply, offer, or even suggest a specific discount — pricing conversations stay with the account owner. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Hand off the packet With that in place, an account owner opens Slack to a ranked list of what's coming up and what looks at-risk, and opens their inbox to a drafted email already carrying the usage story, what's been delivered, and an expansion idea worth raising. They read it, edit it to fit the relationship, and send it themselves. > **The pattern** > A daily **cron** spawns a session with read connectors into HubSpot and > Google Calendar. The renewal playbook lives as a **skill**. The agent preps > the packet and drafts the outreach; the account owner reviews, edits, and > sends. ## Guardrails The agent touches revenue-relevant accounts, so what it can do is narrow and explicit: - **Prep and draft only.** The agent never sends the renewal outreach and never creates or moves a calendar event. The email is a draft in the owner's inbox until the owner sends it. - **No pricing decisions.** The agent never applies, offers, or proposes a specific discount or credit. Anything touching price is a conversation for the account owner, not a line in a draft. - **Scoped writes.** The only thing the agent writes back to HubSpot is which renewal window a deal has been surfaced for — never the deal stage, the amount, or the pipeline. - **Isolation.** Every run happens in its own isolated sandbox. Credentials for HubSpot and Google Calendar are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. - **Everything is code.** The renewal playbook, the at-risk criteria, and the per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **90/60/30:** Days out a renewal first gets a packet, not a deadline - **Every day:** Renewal list and at-risk flags recomputed from live HubSpot data - **0 sent:** Outreach and discounts the agent sends or applies on its own Renewals that used to start with a blank page now start with a packet already built and a draft already written, so the account owner's first move is a review, not a scramble. The agent watches every deal and does the prep work; the person still decides what to say and when to send it. --- <!-- /markdown/use-cases/resume-triage.md --> # How we screen inbound applicants An agent connected to our applications inbox, the role's written rubric, and Google Calendar that scores each resume against the rubric with evidence and proposes interview slots for strong matches. Canonical page: https://taskstation.co/use-cases/resume-triage Every open role brings in more applications than anyone can read carefully, so screening gets rushed: a few seconds per resume, inconsistent from one reviewer to the next, and easy to drift from the criteria the role was actually written against. Good candidates get skimmed past and the standard moves depending on who is reading. We run an agent on TaskStation that does the first read consistently. Each application triggers a session that scores the resume against the role's written rubric, writes a structured screen with evidence, and proposes interview slots for strong matches. A person makes every advance-or-reject call; the agent never rejects anyone. This is how we screen our own inbound applicants. - **Team:** TaskStation - **Runs on:** Every inbound application - **Connected systems:** Applications inbox / ATS · Role rubric · Google Calendar - **Mode:** Trigger-driven · a person makes every decision ## The problem The first read of a resume is where screening gets inconsistent. Volume forces it to be fast, and fast reads drift: the same resume gets a different verdict from a different reviewer, or against a criterion the role never listed. Strong candidates get missed and the bar moves with whoever is reading. The usual fixes trade one problem for another. Keyword filters reject on the wrong signal and quietly drop good people. A rushed human pass is inconsistent by the afternoon. An AI screener that scores against its own idea of "good" is a fairness problem: opaque, unaccountable, and impossible to check. ## What we built On TaskStation, each inbound application triggers an agent. The application spawns an isolated session (a cloud sandbox) with the role's written rubric and scoped access to Google Calendar. The agent reads the resume against the rubric, writes a structured screen — strengths, gaps, a score, and supporting quotes as evidence — and for strong matches proposes interview slots on the hiring manager's calendar. A person decides every case. ## How it works ### Connect the applications inbox as the trigger The applications inbox, or the ATS, is connected so a new application is the trigger. Each one fires a fresh **session** in its own sandbox, seeded with that resume. One application maps to one session on one disposable machine, so screens are independent and the pipeline processes in parallel. ### Give the agent the role's rubric The role's written rubric lives as **skills** and **memory** that travel with the agent: the required and preferred criteria, what strong evidence looks like for each, and how to score. The agent scores against this rubric and nothing else. When the rubric changes, we update the file and the agent screens against the new version. ### Connect the calendar Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the resume** — the full application, so the screen quotes what the candidate actually wrote. - **Score against the rubric** — strengths, gaps, and a score, each tied to supporting quotes as evidence. - **Propose interview slots** — for strong matches, open times on the hiring manager's Google Calendar, offered for a person to confirm. ### Set the guardrails The agent scores only against the written rubric and always surfaces the evidence behind every strength, gap, and score, so a decision can be checked. It never auto-rejects: a person makes every advance-or-reject decision. The agent produces the screen and the proposed slots; the hiring manager decides. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Let each application arrive pre-screened With that in place, an inbound application arrives already read against the rubric: a structured screen with a score and the quotes behind it, and for strong matches a set of proposed interview times. The hiring manager reviews the evidence, decides, and confirms a slot. The first read is consistent, and the decision stays with a person. > **The pattern** > A **trigger** on every application spawns a session with the role's rubric as > **skills** and **memory** and a scoped **connector** into Google Calendar. The > agent scores only against the written rubric and surfaces the evidence; it never > auto-rejects, and a person makes every decision. ## Guardrails The agent reads candidate applications and its output shapes hiring, so fairness and human judgment are built into the controls: - **Isolation.** Every application runs in its own isolated sandbox. The session reads only the resume it's seeded with and reaches only the calendar; only the screen and proposed slots are written back out. - **Scoped secrets.** The inbox, ATS, and calendar credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** The agent never auto-rejects. It scores only against the written rubric and always surfaces the evidence behind every score, and a person makes every advance-or-reject decision. - **Everything is code.** The rubric, the agent's configuration, and its permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every application:** Read against the same written rubric, with evidence - **A person decides:** No candidate is ever auto-rejected by the agent - **3 systems:** Applications inbox, rubric, and calendar in one agent The first read is now consistent and checkable: every application is scored against the same written rubric with the quotes that support each score, and strong matches arrive with interview times already proposed. The hiring manager spends their time deciding on evidence rather than skimming resumes, and every advance-or-reject call stays with a person. --- <!-- /markdown/use-cases/rfp-responder.md --> # How we draft RFP responses The RFP-responder agent we run on TaskStation — connected to the inbound RFP and our vetted answer library. It parses each question, drafts a response in a new Google Doc from the library and our product docs, and flags anything it can't answer confidently for a human to review before it goes out. Canonical page: https://taskstation.co/use-cases/rfp-responder An RFP or proposal questionnaire lands as a long document attached to an email — pricing, implementation timeline, support model, integrations, company background, references — and a submission deadline that doesn't move. Most of it we've answered before, in a past proposal or a product doc, but it's scattered across old submissions and no one remembers exactly where. A rep spends a day combing through past deals and pinging product and support for lines that already exist, before writing a single new sentence. We run an RFP-responder agent on TaskStation that drafts the response from our vetted answer library the moment the RFP arrives. This is how we answer our own RFPs and proposal questionnaires, including the connections and guardrails involved. - **Team:** TaskStation - **Runs on:** Every 15 minutes - **Connected systems:** Gmail · Google Drive · Google Docs - **Mode:** Drafts only · a human reviews and submits ## The problem Most of an RFP is questions we've answered dozens of times — pricing tiers, implementation timeline, integrations, support SLAs, company background, references — but each one arrives in its own document, a Word file, a PDF, a vendor's own spreadsheet, with its own structure and its own deadline. Someone has to read the whole thing, recall or dig up the answer we gave last time, and write it back into that document's layout, all before the submission window closes. The shortcuts don't hold up. Copying the last similar proposal wholesale risks carrying over stale pricing or a client-specific detail that doesn't apply here. Searching a shared drive of past submissions by hand is slow and depends on remembering which deal had the closest answer. A generic AI drafter with no grounding writes something that reads well and is often wrong — on a document that can become a contract exhibit, a confident wrong answer is worse than a blank one. ## What we built On TaskStation, a cron checks the inbox every 15 minutes for a new RFP or proposal questionnaire. Each check spawns a fresh session (a cloud sandbox) with access to our vetted past-answers library, our product docs, and the incoming RFP document. It parses the RFP into individual questions, matches each against the library and the docs, drafts the full response in a new Google Doc, and flags any question it can't answer confidently for a person. It never submits anything — the draft is the only output, and a human reviews and submits it through whatever channel the RFP requires. ## How it works ### Poll the inbox every 15 minutes A **cron trigger** checks Gmail every 15 minutes for a new RFP or proposal questionnaire. Each firing spawns a fresh **session** in its own sandbox — this is a fresh-session agent, so nothing carries over between checks; it finds new work from what's sitting in the inbox right now. ### Ground the agent in our vetted answers Our past-answers library lives as **skills** and **memory** that travel with the agent: the vetted response to each recurring question — pricing tiers, implementation timeline, integrations, support model, security basics, company background, references — cross-checked against the current product docs so an answer that's gone stale doesn't get reused. The agent drafts only from this grounded set. When we win a deal with a better answer, we write it down and the next RFP gets the improved version. ### Connect the RFP source and the drafting surface Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the inbound RFP from Gmail** — the email that carries the RFP or questionnaire and its deadline. - **Open the RFP document from Google Drive** — the attached or linked Word doc, PDF, or spreadsheet, whatever format the buyer sent. - **Build the draft in Google Docs** — a new document with the full response, mirroring the RFP's own question order and sections. - **Flag low-confidence questions** — anything without a confident match in the library, called out for a person to answer. ### Set the guardrails The agent **drafts only**: it writes the response doc and flags what it's unsure of, and it never submits or sends anything back to the buyer — no portal upload, no email, no attachment sent on our behalf. Anything without a confident match in the vetted library is left for a person rather than guessed. The finished draft stops at a **human approval gate** — a person reviews it and submits it. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Return a completed draft for review With that in place, a new RFP comes back within 15 minutes as a fully drafted Google Doc, answers pulled from our vetted library and product docs, and the uncertain questions called out by name. The rep reviews the draft, answers the flagged questions, and submits it through the channel the RFP calls for. The research and first draft are done; the review and submission stay with a person. > **The pattern** > A 15-minute **cron** checks Gmail and spawns a fresh session with scoped > **connectors** into Drive and Docs. The past-answers library lives as > **skills** and **memory**. The agent drafts every answer it can and flags what > it can't; a human reviews and submits. ## Guardrails The agent drafts a document that can end up in front of a customer, so its access is scoped and the output is contained: - **Isolation.** Every check runs in its own isolated sandbox. The session reads only the RFP source and the answer library it's scoped to, and only the draft Google Doc is the output. - **Scoped secrets.** The Gmail, Drive, and Docs credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Draft only, never submitted.** The agent has no path to submit an RFP portal, send an email, or upload a file to the buyer. A human reviews the draft and submits it themselves. - **Everything is code.** The agent's answer library, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every 15 min:** Inbox checked for a new RFP or questionnaire - **Grounded:** Every answer drawn from the vetted library and product docs - **Draft only:** A human reviews and submits every response RFPs that used to cost a rep a day of digging through old proposals now come back as a completed draft within 15 minutes, with the vetted answer already in place and the uncertain questions called out by name. The team spends its time on the questions that actually need judgment, and nothing goes to a buyer without a person reviewing and submitting it. --- <!-- /markdown/use-cases/saas-spend-audit.md --> # How we audit our SaaS spend for waste The SaaS spend audit agent we run on TaskStation — reconciles recurring card and bank charges against our subscription register every week and flags duplicate tools, unused seats, price hikes, and shadow IT, recommending only — it never touches a subscription. Canonical page: https://taskstation.co/use-cases/saas-spend-audit SaaS spend sprawls quietly. A team trials a tool and keeps paying for it after the trial ends. Two teams buy overlapping products without knowing it. A vendor raises its price and the invoice just changes. A card gets charged for something that was never logged anywhere. None of this shows up until someone sits down with a spreadsheet and a list of bank transactions, which happens rarely because it's tedious and the spreadsheet is usually stale by then. We run a SaaS spend audit agent on TaskStation that reconciles our card and bank charges against our subscription register every week and posts what it finds to Slack. It never touches a subscription — it only ever recommends. This is how we keep our SaaS spend honest. - **Team:** TaskStation - **Runs on:** Weekly cron - **Connected systems:** Plaid · Google Sheets · Slack - **Mode:** Read-only · recommend-only · never cancels ## The problem The two records that would catch SaaS waste — what's actually being charged and what's supposed to be — live in different places and drift apart. The bank feed shows every recurring charge; the subscription register shows what finance thinks it's paying for. A tool bought on a personal card during a trial never makes it into the register. A price increase shows up as a bigger number on the same line, easy to miss next to a hundred other charges. A quarterly spend review catches some of this, but it's a point-in-time snapshot done by hand, and three months is a long time for a duplicate tool or a stale renewal to keep charging. Anything faster than quarterly means someone doing the same manual comparison every week, which doesn't happen because nobody has a spare afternoon every week. ## What we built On TaskStation, a weekly **cron** re-prompts a persistent agent **session**. It reads recurring card and bank charges from Plaid, reads the subscription register we keep in Google Sheets, and reconciles the two: duplicate or overlapping tools, seats that look unused, price hikes since the last time it checked, renewals coming up soon, and shadow IT — charges with no matching row in the register at all. It posts what it finds to Slack, remembering what it already flagged so the same unresolved item doesn't repeat every week. ## How it works ### Run on a weekly, reusable session A **cron trigger** fires every week, but unlike a one-shot daily check this agent runs in a **reusable session** — the same session is re-prompted each week rather than starting fresh. That's what lets it remember which subscriptions it has already flagged and what it recommended last time, so a duplicate tool that's still unresolved gets one line in the digest, not a new alert every Monday. ### Give the agent the audit rules What counts as a duplicate, an unused seat, a meaningful price hike, or a renewal worth flagging lives as a **skill** that travels with the agent. It defines the matching logic between a bank charge and a register row, the threshold for a price change worth reporting, and the format of a recommendation. When we tune a threshold, it's a change to the skill, not a one-off instruction. ### Connect the systems read-only Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent reads: - **Recurring charges from Plaid** — the card and bank feed, grouped into recurring merchants over a trailing window. - **The subscription register from Google Sheets** — what finance believes we're paying for, and at what price. - **Posts to Slack** — the weekly digest of findings, new and still-open. It has no write access to Plaid or the register. It cannot add a line, remove one, or change a price in either system. ### Set the guardrails The agent's only actions are reading two systems and posting to Slack. It **recommends** cancellations, downgrades, and consolidations — it never cancels, pauses, downgrades, or otherwise modifies a subscription or a payment method itself. That decision, and the click that executes it, stays with a person. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Post the weekly digest Each week's message lists what's new — a fresh shadow-IT charge, a price that just moved, a renewal coming up — and what's still open from a prior week, without repeating anything unchanged. Every line carries the evidence (the charge, the register row, or the absence of one) and a suggested action. The finance team reads it and decides what to actually cancel or renegotiate. > **The pattern** > A weekly **cron** re-prompts a persistent **session** that reconciles Plaid > against a Google Sheet register through read-only **connectors**. The audit > rules live as a **skill**, a **ledger** tracks what's already been flagged, > and the agent only ever recommends — a person owns every cancellation. ## Guardrails Reconciling spend means reading two financial records, so the boundaries are explicit: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to Plaid and the register it's scoped to, and only the Slack digest is written back out. - **Scoped secrets.** The Plaid access token and Sheets credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Recommend-only.** The agent has no write access to Plaid, no billing API, and no way to cancel, downgrade, or pause a subscription. Its entire output is a Slack message with a suggested action; a human does anything that changes a subscription. - **No repeat noise.** A ledger persists across weeks, so an unresolved finding gets one entry, not a new alert every run — only new or changed waste is reported as new. - **Everything is code.** The agent's persona, its audit rules, and its per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every week:** Charges reconciled against the register on schedule - **Recommend-only:** Nothing is ever cancelled or changed automatically - **5 waste signals:** Duplicates, unused seats, price hikes, renewals, shadow IT The comparison that used to require someone opening a spreadsheet next to a bank statement now runs every week without anyone starting it, and only reports what's new or unresolved. The agent reads the charges and the register; the finance team decides which recommendation to act on. --- <!-- /markdown/use-cases/sales-call-followup.md --> # How we follow up after sales calls An agent connected to our call transcripts, HubSpot, and Linear that drafts the recap, updates the deal, and files the follow-up tasks, then holds the email for the rep to send. Canonical page: https://taskstation.co/use-cases/sales-call-followup After a sales call ends, someone has to write the recap email, answer the questions that came up, update the CRM, and file whatever the team committed to. It's an hour of work that usually happens late, half-remembered, or not at all. The deal stalls not because the call went badly but because the follow-up slipped. We run an agent on TaskStation that does the follow-up the moment a call ends. It reads the transcript, writes the recap, updates the HubSpot deal, and files the tasks in Linear, then holds the outbound email at an approval gate so the rep reviews and sends it. This is how we handle our own post-call follow-up. - **Team:** TaskStation - **Runs on:** Every sales call, when the recording ends - **Connected systems:** Call transcript · HubSpot · Linear - **Mode:** Trigger-driven · email held for the rep ## The problem The follow-up after a call is where deals leak. The rep just ran the meeting, so the context is in their head, but the writing-up is manual: recap the discussion, answer the open questions, restate the next steps, log it in the CRM, and file the tasks. Done well it takes an hour; done late it's vague; skipped, the deal goes quiet. The usual fixes don't close the gap. A CRM reminder tells you to follow up but doesn't do any of it. A note-taker that dumps a transcript into the deal leaves the recap, the CRM update, and the tasks for the rep to write by hand. A template email is fast but generic, and it still needs the account context filled in. ## What we built On TaskStation, the end of a call triggers an agent. When the recording finishes, the transcript spawns an isolated session (a cloud sandbox) with scoped access to HubSpot and Linear. The agent reads the transcript, drafts the recap email with the answers and next steps, updates the deal stage, notes, and contacts, and files the follow-up tasks. The recap email waits at an approval gate for the rep to review and send. ## How it works ### Connect the call as the trigger The meeting recorder is connected so the end of a call is the trigger. When a recording finishes, the transcript fires a fresh **session** in its own sandbox, seeded with the call. One call maps to one session on one disposable machine, so nothing carries between deals and concurrent calls process in parallel. ### Give the agent the recap playbook How we write a follow-up lives as **skills** and **memory** that travel with the agent: the shape of a good recap, how we phrase next steps, which objections need a considered answer, and the account's own history. When a rep improves a recap, we write it down and the agent applies it on the next call. ### Connect HubSpot and Linear Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the transcript** — the full call, including the questions raised and what was committed to. - **Update the HubSpot deal** — stage, notes, and contacts, written back from what the call actually covered. - **File follow-up tasks in Linear** — the concrete next steps, assigned and dated. - **Draft the recap email** — summary, answers to the open questions, and clear next steps, ready for the rep. ### Set the guardrails The agent reads the transcript and writes to the CRM and the task tracker on its own, but the outbound email never sends itself. Every recap stops at a **human approval gate**, so the rep reads it, edits if needed, and sends. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Let the follow-up be ready when the call ends With that in place, hanging up is the whole task. By the time the rep looks, the deal is updated, the tasks are filed, and a recap email is drafted and waiting: the summary, the answers, and the next steps. The rep reviews and sends. The follow-up happens while the call is still fresh instead of days later. > **The pattern** > A **trigger** on the end of each call spawns a session with scoped > **connectors** into HubSpot and Linear. How we write a recap is encoded as > **skills** and **memory**. The CRM and tasks update automatically; the outbound > email waits for the rep at an approval gate. ## Guardrails The agent reads customer conversations and writes to the CRM, so the access is scoped and the outbound step stays with a person: - **Isolation.** Every call runs in its own isolated sandbox. The session reads only the transcript it's seeded with and reaches only HubSpot and Linear; only the updates and the drafted email are written back out. - **Scoped secrets.** The transcript, HubSpot, and Linear credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** The recap email never sends on its own. Each one is held for the rep to review and send; the agent writes to the CRM and tasks, the rep owns the outbound. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every call:** Recap, CRM update, and tasks ready when it ends - **Rep sends:** The outbound email is always reviewed by a person - **3 systems:** Transcript, HubSpot, and Linear in one agent The follow-up now happens while the call is still fresh: the deal is updated, the tasks are filed, and a recap email is drafted and waiting. The rep spends a minute reviewing instead of an hour writing, and no email goes out without a person sending it. --- <!-- /markdown/use-cases/sales-forecast.md --> # How we roll the pipeline into a weekly forecast The sales-forecast agent we run on TaskStation — a weekly cron rolls up HubSpot's open pipeline into a weighted forecast vs quota by stage, rep, and segment, flags deals slipping the quarter, and posts it to Slack. Read-only — it never touches a deal's amount or close date. Canonical page: https://taskstation.co/use-cases/sales-forecast Every sales org has a number it's chasing every quarter, and every sales org reconciles that number the same way: someone pulls open deals out of HubSpot, weights them by gut feel or a spreadsheet formula nobody remembers building, and calls it the forecast. It's accurate for about a day — the day it was built — and it says nothing about the deal whose close date quietly slid past last week, or the six-figure deal that hasn't had a logged activity in a month. We run a sales-forecast agent on TaskStation that reads HubSpot every week and posts a weighted forecast to our sales Slack channel — broken down by stage, rep, and segment, measured against quota, with the deals putting that number at risk called out by name. It only reads the pipeline; the single output is the Slack post. This is how we forecast our own quarter. - **Team:** TaskStation - **CRM:** HubSpot - **Runs on:** Weekly cron - **Mode:** Read-only · one Slack post per week ## The problem A stage-by-stage pipeline view and a real forecast are different things. The pipeline view says a deal is "open" and sitting in "Proposal Sent"; it doesn't say whether that deal is on track to close this quarter, whether the rep's commit number matches what the stage actually implies, or whether a big deal has quietly gone stale while still counting toward the total. Turning open deals into a forecast means weighting every deal by how likely its stage really is to close, then rolling that up three different ways — by stage, by rep, by segment — and comparing it to quota. The usual fix is a spreadsheet rebuilt the week before the forecast call: someone exports deals, applies a weighting rule by hand, and reconciles it against every rep's stated commit. It's stale the moment it's built, it depends on someone remembering to run it, and it rarely catches a deal whose close date has already passed or one that's gone quiet in a late stage until the quarter is nearly over. ## What we built On TaskStation, a weekly cron triggers an agent. It spawns a fresh session with read-only access to HubSpot, pulls every open deal — stage, amount, close date, owner, and its segment — weights each one by the probability already configured on its HubSpot pipeline stage, and rolls the weighted amounts up by stage, by rep, and by segment. It compares the total to this quarter's quota, flags deals whose close date has already slipped or that don't have enough runway left in their current stage to realistically close this quarter, flags large deals that look at risk for other reasons — no recent activity, stalled in a late stage — and posts the whole forecast to Slack. It writes nothing back to HubSpot. ## How it works ### Run on a weekly cron A **cron trigger** fires the agent once a week. Each firing spawns a fresh **session** in its own sandbox — one week, one run, nothing carried over. The forecast is rebuilt from HubSpot's current state every time, so it never drifts from what's actually in the CRM. ### Give the agent the rollup rules How deals get weighted, which ones count toward this quarter, and what "slipping" and "at risk" mean live as a **skill** that travels with the agent: weight every deal by the probability HubSpot already has configured on its stage rather than a hardcoded guess, count a deal into the quarter by its close date, and call out a deal as slipping when its close date has already passed or its stage leaves it too little runway to close in time. ### Connect HubSpot read-only Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent reads: - **Open deals** — amount, stage, close date, owner, and the segment property used for the rollup. - **Pipeline stage configuration** — the win probability HubSpot already has set on each stage, used as the weighting instead of an assumption. - **Posts to Slack** — the weekly forecast, broken down by stage, rep, and segment, with slipping and at-risk deals called out. ### Set the guardrails The agent is **read-only** across HubSpot. It has no write access to a deal's amount, close date, stage, or owner — it cannot change the number tied to any deal, no matter what the rollup shows. Its only output is the Slack post. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Post the forecast Each week brings one Slack post: the total weighted forecast against quota, the breakdown by stage, by rep, and by segment, the deals that have slipped or are running out of runway this quarter, and the large deals showing other signs of risk. The sales team reads it and decides what to do. Nothing is written back to HubSpot automatically. > **The pattern** > A weekly **cron** spawns a session with read-only **connector** access to > HubSpot. The weighting and rollup rules live as a **skill**. The agent reads > the whole pipeline and writes nothing but the weekly Slack post. ## Guardrails Rolling up the entire pipeline every week is a trust question, so the agent's access is scoped and one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to HubSpot and Slack, and only the forecast post is written back out. - **Scoped secrets.** The HubSpot credential is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Read-only, no exceptions.** The connector into HubSpot is read-only. The agent never changes a deal's amount, close date, stage, or owner — it can only report on what it finds. - **Everything is code.** The weighting rule, the quarter window, and the slipping/at-risk criteria are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every week:** Forecast rebuilt from HubSpot's current pipeline state - **Read-only:** No deal amount, date, stage, or owner ever written - **3 cuts:** Stage, rep, and segment rolled up in one weekly post The forecast that used to get rebuilt by hand the week before the call now arrives every Monday in the channel where the team already works, weighted by HubSpot's own stage probabilities and carrying the specific deals putting the number at risk. The agent only reads and reports; the sales team still owns every deal. --- <!-- /markdown/use-cases/security-questionnaire.md --> # How we answer security questionnaires An agent we run on TaskStation — connected to the inbound questionnaire and our knowledge base of vetted answers and policies. It parses each question, drafts responses in the vendor's format, and flags anything it can't answer confidently. Canonical page: https://taskstation.co/use-cases/security-questionnaire Security questionnaires arrive in the middle of a sales cycle and hold the deal up until they're answered. The questions are mostly ones we've answered before — about our encryption, access controls, data handling, and policies — but they come in different formats each time, a SIG one deal, a CAIQ the next, a custom spreadsheet after that, and each has to be filled out in its own layout. We run an agent on TaskStation that drafts the answers from our vetted knowledge base. This is how we answer our own security questionnaires, including the connections and guardrails involved. - **Team:** TaskStation - **Runs on:** Each inbound questionnaire - **Connected systems:** Inbound questionnaire · Knowledge base - **Mode:** Drafts only · security reviews before it's sent ## The problem Most of a questionnaire is answers we already have. The same questions about encryption at rest, SSO, incident response, and data retention come up on nearly every one, and we've written vetted answers for them. But each questionnaire uses a different format, so someone has to read every question, find the matching approved answer, and paste it into the vendor's own layout — a SIG workbook, a CAIQ, or a custom spreadsheet. The common fixes are incomplete. A shared answer library still needs a person to match each question to it by hand. Copying last deal's responses risks pulling an answer that no longer fits. A generic AI drafter with no grounding will write confident answers that aren't the vetted ones, which is exactly what can't happen on a security document. ## What we built On TaskStation, each inbound questionnaire triggers an agent. It spawns an isolated session (a cloud sandbox) with access to our knowledge base of vetted answers and policy docs. It parses each question, matches it to our approved answers and policies, drafts responses in the vendor's own format — SIG, CAIQ, or a custom spreadsheet — and flags anything it can't answer confidently for a human. It returns a filled draft for security to review before it goes back. ## How it works ### Trigger on the inbound questionnaire A questionnaire arriving as an email, a spreadsheet, or a portal link is the **trigger**, and each one spawns a fresh **session** in its own sandbox. The agent parses the incoming document, whatever its format, into a list of questions to answer. One questionnaire maps to one session on one disposable machine. ### Ground the agent in our vetted answers Our approved answers and policy docs live as **skills** and **memory** that travel with the agent: the vetted response to each common question, the policies behind them, and the standards we map to. The agent answers only from this grounded set, so a drafted answer is one we've already approved rather than one the model invented. When a policy changes, we update it and the agent uses the new wording. ### Connect the questionnaire and the knowledge base Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read the inbound questionnaire** — from the email, spreadsheet, or portal it arrived in, parsed into individual questions. - **Search the knowledge base** — the vetted answers and policy docs to match each question against. - **Draft in the vendor's format** — writing each response back into the SIG, CAIQ, or custom layout it came in. - **Flag low-confidence questions** — anything without a confident match marked for a person to answer. ### Set the guardrails The agent **drafts only**: it fills the questionnaire and flags what it's unsure of, and it never sends. Anything it can't answer confidently from the vetted set is left for a person rather than guessed. The completed draft stops at a **human approval gate** — security reviews it before it goes back to the prospect. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Return a filled draft for review With that in place, an inbound questionnaire comes back as a filled draft in the vendor's own format, with each answer drawn from our vetted set and the low-confidence rows flagged. Security reviews the draft, answers the flagged questions, and sends it. The matching and formatting are done; the sign-off stays with a person. > **The pattern** > An inbound questionnaire is the **trigger** that spawns a session with scoped > **connectors** into the document and our knowledge base. The vetted answers and > policies are encoded as **skills** and **memory**. The agent drafts and flags; > security approves before anything is sent. ## Guardrails The agent drafts a security document from our own vetted answers, so the access is scoped and contained: - **Isolation.** Every questionnaire runs in its own per-task isolated sandbox. The session reads the inbound document and the knowledge base it's scoped to, and only the filled draft is written back out. - **Scoped secrets.** The credentials for the questionnaire source and the knowledge base are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Human approval gate.** The completed draft is reviewed by security before it goes back to the prospect, and any low-confidence question is left for a person rather than guessed. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every questionnaire:** Parsed, matched, and drafted in the vendor's format - **Grounded:** Answers drawn only from our vetted set - **Drafts only:** Security reviews before anything is sent Inbound questionnaires now come back as filled drafts in whatever format they arrived, with each answer drawn from our vetted knowledge base and the uncertain ones flagged. Security spends its time reviewing and answering the hard questions rather than matching and pasting, and nothing goes back to a prospect without a person's sign-off. --- <!-- /markdown/use-cases/slack-control-pane.md --> # How we run operations from Slack A single agent reachable from Slack, with scoped access to our database, Stripe, Linear, and GitHub. Ask in a thread and it runs the task across whatever platforms it needs. Canonical page: https://taskstation.co/use-cases/slack-control-pane Operations work usually spans several platforms: the database, Stripe, the sandbox providers, Linear, and GitHub. Onboarding a customer means provisioning, setting up billing, and filing a tracking issue. Shipping a fix means reviewing a PR, cutting a branch, and updating a ticket. No single tool covers the whole task, so the coordination falls to a person moving between several tabs. We run these tasks through a single agent reachable from Slack, with scoped access to the platforms we've connected. This is how we run our own operations on TaskStation, and the same setup works for any team. - **Team:** TaskStation - **Control surface:** Slack - **Connected systems:** Database · Stripe · Sandboxes · Linear · GitHub - **Setup per platform:** Connect once, then it can act ## The problem The cross-platform tasks are the ones no single tool owns. "Onboard this enterprise account" touches the database, Stripe, and Linear. "Get this fix out" touches GitHub, the sandbox that reproduces the bug, and the ticket that tracks it. Each step is simple; stitching them together means switching between tabs, copying IDs, and keeping the order straight. The common workarounds each fall short. Internal scripts each do one thing and break when an API changes. A no-code automation tool handles the flow it was built for and nothing beyond it. A chatbot that integrates with Slack can answer questions but can't run a multi-step task, because it has no safe way to hold credentials and take actions across systems. ## What we built Our Slack is wired to an agent with scoped access to every connected system. You @-mention it in a thread and describe the task in plain language. It spawns an isolated session, a cloud sandbox, with scoped access to the platforms we've connected, works out the steps, runs them, and replies in the thread as it goes. Adding a capability means connecting one more platform. ## How it works ### Make Slack the control surface Slack is connected as a **channel**, so a message is the trigger. Mention the agent in any thread and the request spawns a fresh **session** in its own isolated sandbox. The agent stays in the thread and replies as it works. One request, one session, one disposable machine. ### Connect a platform To give the agent a new capability, you **connect the platform to TaskStation once**. The database, Stripe, the sandbox providers, Linear, and GitHub are each a scoped **connector**, brokered server-side so no raw token reaches the model. Once a platform is connected, the agent can act on it. There's no integration to write or script to maintain. ### Run tasks across those platforms With the platforms connected, the agent handles cross-platform tasks: - **Invite a member to Linear**, create the project, and file the tracking issues. - **Review an open GitHub PR**, and open its own PRs to the codebase when a change is warranted. - **Query the database** to answer questions like how many accounts are on a given plan. - **Look up and adjust Stripe** state: plan, invoice, subscription. - **Spin up a sandbox** to reproduce a bug or run a one-off job. Whatever the task spans, it runs in one thread through one agent. ### Set the guardrails Because the agent has access to every connected platform, its scope is kept tight. Access is **read-mostly by default**. Actions that touch money, production data, or account state (a Stripe change, a merge, a destructive query) stop at a **human approval gate** in the thread. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Run operations from a thread With that in place, "onboard this account" is one message that provisions in the database, sets up billing in Stripe, and files the Linear project, with the irreversible steps held for approval. "Review this PR" is a code review plus, when needed, a follow-up PR. The work spans one thread instead of several tools. > **The setup** > Connect Slack as the **channel**, connect each platform to TaskStation as a scoped > **connector** (the only per-platform setup), and gate every sensitive action > behind a human. One agent then runs cross-platform tasks from a thread. ## Guardrails Giving one agent access to the database, Stripe, sandboxes, Linear, and GitHub is a security question. The controls that make it workable: - **Isolation.** Every request runs in its own isolated sandbox on its own branch. A session is granted access only to the platforms it's scoped to, and only what it's explicitly allowed to send is written back out. - **Scoped secrets.** Each platform credential is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Human approval gates.** Irreversible actions (money, production data, merges) require a person to approve in the thread. - **Everything is code.** The agent's persona, skills, and per-platform permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **One thread:** Any cross-platform task, asked in plain language - **Connect once:** A new platform is a new capability, no build - **5+ systems:** Database, Stripe, sandboxes, Linear, GitHub — one agent Cross-platform tasks such as onboarding, reviews, provisioning, and billing changes now run in the Slack thread where they're already discussed, with the risky steps held for approval. Extending the system means connecting the next platform. --- <!-- /markdown/use-cases/social-scheduler.md --> # How we draft social posts from our content calendar The social-scheduler agent we run on TaskStation — a daily agent that reads our Notion content calendar, drafts platform-appropriate posts for upcoming content, launches, and announcements, and holds them in Slack with the scheduled date for a human to approve and publish. Canonical page: https://taskstation.co/use-cases/social-scheduler Our content calendar lives in Notion: what's shipping, what's launching, what we're announcing, and when. Turning each of those rows into an actual LinkedIn post, a tweet, and an Instagram caption is a separate manual step that happens later, usually the day of, sometimes not at all. The calendar says what should go out; nobody had turned it into copy yet. We run a social-scheduler agent on TaskStation that closes that gap every day. It reads the content calendar in Notion, drafts platform-appropriate posts for anything launching or publishing soon, and holds every draft in our marketing Slack channel with its scheduled date. It never posts anywhere itself — a person reviews, edits, and publishes. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Notion · Slack - **Mode:** Fresh session · draft-only, human publishes ## The problem A content calendar tells you *what* and *when*, not the actual words that go on each platform. Writing the LinkedIn version, the X version, and the Instagram caption for every launch and announcement is a distinct task that someone has to remember to do, usually under time pressure right before the scheduled date. The common fixes don't hold up. A scheduling tool queues posts, but someone still has to write them first. A shared doc of "posts to write" depends on someone checking it daily and drafting ahead of the date, not the day of. And because drafting keeps sliding to "later," launches and announcements regularly go out with no social post at all, or one written in five minutes right before it's needed. ## What we built On TaskStation, a daily cron triggers an agent. It spawns a fresh session that reads our Notion content calendar for anything launching, publishing, or being announced in the days just ahead, drafts a platform-appropriate post for each one — LinkedIn, X, and Instagram get different lengths, tone, and structure — and posts the batch to our marketing Slack channel, each draft labeled with its scheduled date. Nothing is scheduled or published automatically; the drafts sit in Slack until someone approves them. ## How it works ### Run on a daily cron, fresh each time A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox — there's no memory to carry over, so the agent re-reads the calendar's current state every time rather than trusting what it drafted yesterday. ### Give the agent the platform playbook How each platform's copy should read — LinkedIn's longer, more explanatory tone; X's tight, single-idea framing; Instagram's caption-plus-hashtags structure — lives as a **skill** that travels with the agent. When we change our voice or add a platform, we edit the skill and every future draft follows it. ### Connect the calendar and the approval channel Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent reads the Notion content calendar for entries landing within the lookahead window. Through a **channel**, it posts the drafted batch to Slack. Neither connection touches any social platform. ### Set the guardrails The agent can read the calendar and write to one Slack channel. It has no connector to LinkedIn, X, Instagram, or any other social account, so there is no path from this agent to a live post — the guardrail isn't a rule the agent follows, it's a permission it was never given. ### Hold every draft for approval Each day's Slack post lists the upcoming calendar items with a drafted post per platform and the scheduled date attached. Someone on the team reviews each draft, edits it if needed, and schedules or publishes it through whatever tool they already use. The agent's job ends at the draft. > **The pattern** > A daily **cron** spawns a fresh session that reads the **Notion** content > calendar through a scoped connector, drafts platform-specific copy using a > **skill**, and posts the batch to **Slack** for approval. No connector to > any social account exists, so nothing can go out without a human. ## Guardrails The agent drafts copy that will eventually represent the company publicly, so what it can do on its own is narrow: - **No social connectors, period.** The agent has no credential and no connector to LinkedIn, X, Instagram, or any other platform. Publishing isn't blocked by a rule — the capability doesn't exist in this agent's scope. - **Draft only.** The only output that is written back out is a Slack message containing drafts. Nothing is scheduled, queued, or posted anywhere. - **Isolation.** Every run happens in its own isolated sandbox, torn down when the run ends. - **Scoped secrets.** The Notion connector is brokered server-side; no raw token enters the sandbox at all. - **Everything is code.** The platform playbook, the lookahead window, and the agent's permissions are files in the repo, changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** Upcoming calendar items turned into drafted copy ahead of time - **0 posts:** Published automatically — every draft waits for a human - **3 platforms:** LinkedIn, X, and Instagram drafted in one pass Launches and announcements that used to get a rushed post the day of, or none at all, now show up in Slack days ahead with a draft ready for each platform. The agent turns the calendar into copy; the team still decides what actually goes out. --- <!-- /markdown/use-cases/standup-summary.md --> # How we run async standups in Slack A daily agent that collects each person's Linear and GitHub activity, prompts anyone with nothing recorded, and posts a concise team standup to Slack. Read-only across every tool. Canonical page: https://taskstation.co/use-cases/standup-summary A standup is supposed to answer one question: what did everyone do, and what's next. In practice it costs a meeting, or a thread where half the team pastes a summary and the other half forgets. The information already exists in Linear and GitHub; someone just has to gather it, format it, and chase the people who didn't post. We run a daily agent on TaskStation that does the gathering. It reads each person's Linear and GitHub activity, nudges anyone with nothing recorded, and posts a concise standup to a Slack channel. This is how we run our own standups, and the same setup works for any team. - **Team:** TaskStation - **Runs on:** A daily cron, one post per weekday - **Connected systems:** Linear · GitHub · Slack - **Mode:** Read-only across the tools, no writes ## The problem A live standup takes a slot on everyone's calendar to relay information that's already recorded somewhere. An async thread saves the meeting but shifts the work onto people: each person has to remember to write, summarize their own day, and post it, and the ones who don't leave gaps. The common workarounds are partial. A recurring bot that asks "what did you do yesterday?" still depends on everyone answering. A dashboard shows activity but doesn't summarize it or tell you who's quiet. Neither reads the tools where the work actually happened, so the summary is only as complete as the people filling it in. ## What we built On TaskStation, a cron fires once each weekday morning and spawns an agent. It reads each team member's recent Linear issues and GitHub activity, builds a short per-person summary, checks who has nothing recorded, and posts the standup to a Slack channel. Anyone with a quiet day gets a light prompt so they can add context the tools can't see. The agent only reads; it never writes to Linear or GitHub. ## How it works ### Trigger it on a daily cron The standup runs on a **cron trigger**: once every weekday morning, the schedule spawns a fresh **session** in its own isolated sandbox. One run, one disposable machine, nothing carried over from the day before. No one has to start it and there's no bot sitting idle waiting for messages. ### Give the agent the team and the format Who's on the team, which repos and Linear projects to look at, and what a good standup entry looks like live as **skills** and **memory** that travel with the agent: the roster, the mapping from a person to their Linear and GitHub handles, and the house style for a summary. When the format changes, we update the file and the next run picks it up. ### Connect Linear and GitHub, read-only Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent reads: - **Linear** — each person's recently updated and closed issues, and what's in progress. - **GitHub** — commits, opened and merged PRs, and reviews since the last standup. Both connectors are scoped to read only. The agent can see the activity; it cannot change an issue or touch a branch. ### Set the guardrails The agent is **read-only** across Linear and GitHub by design: its only write is the Slack post. It doesn't move tickets, comment on PRs, or edit anyone's work. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Post the standup and prompt the gaps With that in place, the morning post writes itself: a per-person summary drawn from Linear and GitHub, grouped so the team can scan it in one read. Anyone with no recorded activity is flagged with a gentle prompt to add what the tools missed, meetings, planning, anything off-platform, so a quiet day shows context rather than a blank line. > **The pattern** > A **cron trigger** spawns a session each morning with read-only **connectors** > into Linear and GitHub. The roster and format live as **skills** and > **memory**. The agent summarizes, prompts the gaps, and posts to Slack, its > only write. ## Guardrails The agent reads everyone's activity across two tools and posts to a shared channel, so the access is scoped and contained: - **Isolation.** Every run happens in its own isolated sandbox. The session reads the day's activity, builds the summary, and only the Slack post is written back out. - **Scoped secrets.** The Linear, GitHub, and Slack credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only across the tools.** The Linear and GitHub connectors are scoped to read. The agent's only write is the standup post; it changes nothing in the source tools. - **Everything is code.** The roster, the format, and the per-tool permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **No meeting:** The standup is a post, not a slot on the calendar - **Read-only:** No writes to Linear or GitHub, only the Slack post - **3 systems:** Linear, GitHub, and Slack in one daily agent The standup now arrives each morning as a Slack post built from the work people already recorded, with quiet days flagged instead of skipped. No one spends a meeting relaying status, and no one has to write their own summary for the activity the tools already have. --- <!-- /markdown/use-cases/ticket-to-kb.md --> # How we turn resolved tickets into KB articles The ticket-to-KB agent we run on TaskStation — clusters resolved Plain support tickets for recurring questions with no matching help-center article, drafts the top gaps as real articles from the actual resolutions, and opens a PR for a human to review and publish. Canonical page: https://taskstation.co/use-cases/ticket-to-kb Every support team answers the same question more than once before anyone writes it down. An agent resolves a ticket, the customer is happy, and the answer disappears back into the support queue instead of becoming a help article. Weeks later the same question comes in again, gets solved the same way, and disappears again. The knowledge was generated a dozen times over; it was never captured once. We run a ticket-to-KB agent on TaskStation that reads our resolved Plain tickets every week, finds the questions that keep recurring with no article to answer them, and drafts the highest-value gaps as real articles pulled from how we actually solved them — then opens one PR. This is how new KB coverage gets written from what our support queue already knows. - **Team:** TaskStation - **Runs on:** Weekly cron · reusable session - **Connected systems:** Plain · GitHub (docs/KB repo) - **Mode:** Draft-only · one PR per run ## The problem The knowledge exists — it's sitting in hundreds of resolved Plain threads — but nobody has time to mine it. A support lead reviewing tickets for KB gaps manually has to remember which topics were already covered, which recurring questions are worth an article versus a one-off, and then actually write the draft from scratch. That review happens rarely, if at all, and by the time it does, the backlog of recurring-but-undocumented questions is large enough that picking where to start is its own project. A generic AI writer pointed at "write some help articles" doesn't fix this either — it has no memory of what's already covered and nothing grounding its drafts in how the team actually resolves the issue, so it either duplicates existing articles or writes generic copy that doesn't match a real resolution. ## What we built On TaskStation, a weekly cron re-prompts the same persistent session — a cloud sandbox that resumes its own ledger rather than starting cold. Each run it reads recently resolved threads from Plain, clusters them into recurring topics, checks each topic against both the KB repo's existing articles and its own ledger of topics already drafted, and picks the top gaps: recurring questions with real ticket volume and no article. It drafts each one from the actual resolutions used in those tickets, then opens a single PR against the docs repo. Nothing publishes without a human merge. ## How it works ### Run on a weekly reusable session A **cron trigger** fires once a week and resumes the same persistent **session** rather than starting fresh. The session reads its own ledger first — which topics already have an article, which are mid-review in an open PR, and which clusters were seen before but didn't yet have enough volume to justify a draft — so a topic already covered or already in flight never gets redrafted. ### Give the agent the clustering and drafting rules How we group tickets into a topic, how we decide a topic is a real KB gap rather than noise, and the article structure we want live as **skills** and **memory** that travel with the agent. When a draft gets merged as-is versus reworked heavily in review, we feed that back so the next batch of drafts gets closer to publishable. ### Connect the ticket source and the KB repo Through a scoped connection, the agent works with: - **Resolved tickets from Plain** — read-only, via a raw API key injected at runtime, scoped to the agents you grant them to. - **The docs/KB repo on GitHub** — read to see which topics already have an article, and write only as a pull request on an isolated branch, via a scoped GitHub token used by the `gh` CLI. ### Set the guardrails The agent never publishes an article and never merges its own PR. It can only open a pull request against an isolated branch; a human reviews the drafts, edits them, and decides what actually ships to the help center. Both credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Open one PR of drafted articles With that in place, each week brings one PR: a small batch of KB article drafts, each one built from the real resolutions of the tickets that raised the question, with the ticket volume and the topic's history noted for the reviewer. The support lead reads the PR, edits what needs editing, and merges what's ready. The ledger advances either way, so next week starts from what's still uncovered. > **The pattern** > A weekly **cron** resumes one persistent **session** that reads a durable > **ledger** of covered topics, clusters resolved Plain tickets read-only, and > opens a single GitHub **pull request** of drafted articles — never publishing > and never merging on its own. ## Guardrails Turning support history into published documentation is a one-way trust question — from tickets to drafts is safe, from drafts to a live article is not something the agent decides — so the guardrails are built around that line: - **Isolation.** Every run resumes in its own isolated sandbox. The session is granted access only to Plain and the KB repo, and only the pull request is written back out. - **Scoped secrets.** The Plain API key and the GitHub token are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Draft-only.** The agent opens a PR against an isolated branch and stops. It never pushes to the live docs branch and never merges its own work — a human reviews and publishes. - **Read-only tickets.** Plain access is read-only; the agent cannot edit, close, or reply to a ticket while mining it for KB content. - **Everything is code.** The agent's clustering rules, drafting standard, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every week:** Resolved tickets reclustered against the current KB gap set - **Draft-only:** Every article ships as a PR a human reviews and merges - **1 ledger:** Tracks which topics are covered so nothing is redrafted The backlog of "we keep answering this but never wrote it down" questions now arrives as a small, reviewable PR each week, with every draft traceable back to the real tickets that produced it. The agent only clusters and drafts; the support team decides what gets published. --- <!-- /markdown/use-cases/user-feedback.md --> # How we turn feedback into a roadmap On a schedule, an agent gathers feedback from support, public reviews, and a Slack channel, clusters it into themes with representative quotes and counts, and creates or updates a Linear issue per theme. Canonical page: https://taskstation.co/use-cases/user-feedback Product feedback arrives everywhere: support threads, public reviews, a Slack channel where people drop what they hear. The same request shows up in all three, worded differently each time, and never gets counted. A theme that a hundred people asked for looks the same as a one-off, because nothing pulls the mentions together. We run a feedback agent on TaskStation that gathers those sources on a schedule, clusters them into themes, and keeps a Linear issue per theme up to date. It reads the feedback and writes to Linear; people still own prioritization. This is how we keep our own roadmap grounded in what users actually ask for. - **Team:** TaskStation - **Runs on:** Scheduled cron - **Connected systems:** Plain · G2 / app stores · Slack · Linear - **Mode:** Reads feedback · writes Linear issues ## The problem Feedback is scattered across support threads in Plain, public reviews on G2 and the app stores, and an internal Slack feedback channel. The same underlying request appears in all of them, phrased differently, so it never gets counted. A theme that keeps recurring is indistinguishable from a single loud comment. The common approaches don't add up to a roadmap. Reading each source by hand is slow and inconsistent, and whoever reads it weighs it differently. A tag in the support tool captures support but not reviews or Slack. A spreadsheet of feature requests goes stale the moment someone stops maintaining it, and it still doesn't tell you how many people asked for the same thing. ## What we built On TaskStation, a scheduled cron triggers an agent. It spawns an isolated session (a cloud sandbox) with read access to Plain, the public review sources, and the Slack feedback channel, and write access to Linear. It gathers the feedback, clusters it into themes with representative quotes and counts, and creates or updates a Linear issue per theme, so the same request is deduplicated and quantified instead of scattered. ## How it works ### Run on a schedule A **cron trigger** fires the agent on a schedule. Each firing spawns a fresh **session** in its own sandbox. One run pulls the current feedback, reconciles it against the existing themes in Linear, and updates them. Nothing carries over between runs except what's written to Linear. ### Give the agent the clustering rules How we group feedback lives as **skills** and **memory** that travel with the agent: what makes two differently-worded requests the same theme, how to pick a representative quote, how to title an issue, and how to match new feedback to an existing theme instead of creating a duplicate. As our themes evolve, we write it down and the clustering stays consistent. ### Connect the sources and Linear Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Read support from Plain** — recent threads, for the requests and complaints users raise directly. - **Read public reviews** — G2 and the app stores, for what users say in the open. - **Read the Slack feedback channel** — the internal channel where the team drops what they hear. - **Write to Linear** — create a new issue for a new theme, or update the count and quotes on an existing one. ### Set the guardrails The agent is **read-only** on every source and its only write is to Linear, where it creates and updates issues. It does not set priority, assign owners, or close issues — those stay with people. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to. ### Keep one issue per theme With that in place, each run gathers the latest feedback, clusters it, and keeps one Linear issue per theme current: representative quotes, a running count, and the sources it came from. The same request stops being scattered across three systems and becomes one quantified issue the team can weigh against the rest. > **The pattern** > A scheduled **cron** spawns a session with read **connectors** into Plain, the > public reviews, and Slack, and a write connector into Linear. The clustering > lives as **skills** and **memory**. The agent quantifies the themes; people own > what to build. ## Guardrails The agent reads several feedback sources and writes to Linear, so its access is scoped to exactly that: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to the sources it's scoped to, and only the Linear writes are written back out. - **Scoped secrets.** The Plain, review-source, Slack, and Linear credentials are encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only sources, scoped writes.** The feedback sources are read-only; the only write is creating and updating Linear issues. The agent does not prioritize, assign, or close. - **Everything is code.** The agent's clustering rules, skills, and per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every run:** Feedback gathered, clustered, and reconciled to Linear - **One per theme:** The same request deduplicated instead of scattered - **4 sources:** Support, reviews, and Slack into one issue tracker Feedback that used to sit unread across support, reviews, and Slack now arrives as a set of quantified themes, each a single Linear issue with quotes and a count. The agent does the gathering and counting; the team still owns which themes become roadmap. --- <!-- /markdown/use-cases/vendor-onboarding.md --> # How we onboard new vendors The vendor-onboarding agent we run on TaskStation — connected to Gmail, Google Sheets, and Slack. Every day it checks new vendor requests for a completed W-9, banking form, and signed contract, records each vendor's status to our vendor register, and flags anything missing or invalid — never approving a vendor or touching payment setup itself. Canonical page: https://taskstation.co/use-cases/vendor-onboarding A new vendor shows up as an email thread: a name, a contact, and — if we're lucky — three attachments. A W-9, a banking form, and a signed contract. Each one needs to be there, complete, and consistent with the others before the vendor is ready to be set up. Done by hand, that check gets skipped under deadline pressure, and the gap that slips through is exactly the one that matters: a vendor entered into a payment system on an unsigned contract or a W-9 with no TIN. We run a vendor-onboarding agent on TaskStation that checks the inbox every day, validates the required documents for every new vendor request against a fixed checklist, records the result to our vendor register, and flags anything missing or invalid in Slack. It collects, validates, and records; it never approves a vendor and never touches payment or banking setup. - **Team:** TaskStation - **Runs on:** Daily cron - **Connected systems:** Gmail · Google Sheets · Slack - **Mode:** Collect + validate + record — never approves, never sets up payment ## The problem Vendor onboarding paperwork is simple in principle and inconsistent in practice. A W-9 arrives unsigned. A banking form is missing the routing number. A contract comes back with the wrong entity name because someone copy-pasted from a template. None of these are hard to catch individually, but catching all three, for every vendor, every time, is the kind of checklist work that a busy person does thoroughly on the first vendor of the week and loosely on the tenth. The cost of skipping it isn't visible until later — a vendor gets set up for payment on paperwork that turns out to be incomplete, and untangling it after the fact costs far more than the two minutes the check would have taken. What's missing isn't judgment about whether to onboard a vendor; it's someone reading every attachment against the same checklist, every time, without skipping ahead to the next request. ## What we built On TaskStation, a daily cron re-prompts a fresh session with no memory of the prior run — the vendor register in Google Sheets is the record it works from. Each run reads new vendor-request threads from a Gmail label, checks each attachment against a fixed checklist (a complete, signed W-9; a complete banking form; a signed contract with a matching entity name), records the result for every vendor whether clean or flagged, drafts a follow-up email for anything missing or invalid, and posts a summary to Slack for a person to act on. It never marks a vendor approved and never sets up a payment method or banking profile. ## How it works ### Run on a daily cron, fresh each time A **cron trigger** fires the agent once a day. Each firing spawns a fresh **session** in its own sandbox with no memory of the previous run — the vendor register it reads from Google Sheets is the only carry-over, so the same vendor is never processed twice and nothing depends on the agent remembering anything itself. ### Give the agent the intake checklist What counts as a complete W-9, a complete banking form, and a valid signed contract lives as a **skill** that travels with the agent — the exact fields each document needs, how the vendor name must match across all three, and what "invalid" looks like versus "missing." This is the fixed standard every vendor is checked against, every run. ### Connect the systems it needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent: - **Reads new vendor requests from Gmail** — the request thread and its attachments, filed under a dedicated label. - **Reads and writes the vendor register in Google Sheets** — the prior state of every vendor, and this run's recorded status for each one. - **Drafts follow-up emails in Gmail** — requesting a missing or corrected document, saved as a draft, never sent by the agent. - **Posts to Slack** — a summary of every vendor processed this run, flagged ones called out with exactly what's missing or invalid. ### Set the guardrails The agent's job stops at collect, validate, and record. It never marks a vendor approved, never initiates a payment method, and never touches a banking or ACH setup in any system — not even for a vendor whose paperwork is fully complete. Outbound email to a vendor is created as a draft only; a person reviews and sends it. Bank account and routing numbers are checked for presence and completeness but never copied into the register — only the document's status is recorded. ### Record every vendor and flag what's missing With that in place, each day the register gains one row per new vendor request — complete or flagged, with the specific missing or invalid item named — and Slack gets one summary post. Anything incomplete has a draft email already written and waiting. A person reviews the flag, sends the draft or requests something else, and makes the actual onboarding decision. > **The pattern** > A daily **cron** re-prompts a **fresh session** that reads new requests from > Gmail, checks the required documents against a **skill**-defined checklist, > and records the result in the Google Sheets vendor register — the register, > not the agent's memory, is the state. Anything missing gets a drafted email > and a Slack flag. The agent never approves a vendor or touches payment. ## Guardrails The agent handles vendor paperwork end to end except the decision that matters, so its access is scoped and its authority is capped below approval: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to Gmail, Google Sheets, and Slack, and only the register update and the Slack post are meant to persist past the run. - **Scoped, brokered credentials.** Gmail and Google Sheets access are injected into the sandbox at runtime, scoped to the agents you grant them to. - **Collect, validate, record — nothing further.** The agent checks documents and writes status to the register. It has no path to mark a vendor approved or to configure a payment method or banking profile, complete paperwork or not. - **Drafts, never sends.** A follow-up email to a vendor is created as a Gmail draft. A human reviews and sends it. - **No sensitive data at rest in the sheet.** Bank account and routing numbers are checked for completeness on the source document but never transcribed into the register — only a status. - **Everything is code.** The intake checklist, the register schema, and the agent's per-system permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every day:** New vendor requests checked against the same fixed checklist - **3 documents:** W-9, banking form, and signed contract validated per vendor - **0 approvals:** Vendor approval and payment setup always stay with a human Vendor paperwork that used to get a thorough read on the first request and a quick skim by the tenth now gets the same checklist every time, recorded in one register instead of scattered across a dozen email threads. The agent collects, validates, and records; a person still decides who gets approved and who gets paid. --- <!-- /markdown/use-cases/weekly-report.md --> # How we auto-generate the weekly metrics report The reporting agent we run on TaskStation — connected to our Postgres database and Slack. Every Monday it queries the metrics, writes the report with commentary on what moved, and posts it. Canonical page: https://taskstation.co/use-cases/weekly-report Every team has a weekly metrics report, and someone spends part of their Monday building it: running the same queries, dropping the numbers into a template, and writing a line or two about what changed. It's routine, it's on a schedule, and it's exactly the kind of thing that gets skipped the week it's most needed. We run a reporting agent on TaskStation that builds it. A Monday cron queries the metrics from our Postgres database, writes the weekly report with commentary on what moved and why, and posts it to Slack. Access to the database is read-only. This write-up covers how the setup works: the trigger, the session model, and the guardrails. - **Team:** TaskStation - **Runs on:** A Monday cron - **Connected systems:** Postgres · Slack - **Mode:** Cron-driven · read-only database access ## The problem The weekly report is low-skill, high-consistency work: the same queries, the same layout, every week. Done by hand it eats an hour of someone's Monday, and the week things are busiest is the week it's most likely to slip — which is usually the week the numbers most needed a look. The common fixes are incomplete. A static dashboard shows the numbers but doesn't say what moved or why, so someone still has to read it and write the summary. A scheduled SQL job can post the figures but not the commentary. The interpretation — what changed, whether it matters — is the part that takes a person, and it's the part that gets dropped. ## What we built On TaskStation, a Monday cron triggers a reporting agent. Each run spawns an isolated session (a cloud sandbox) with scoped, read-only access to the Postgres database and permission to post to one Slack channel. The agent runs the metric queries, compares them against the prior weeks, writes commentary on what moved, and posts the report to Slack. It cannot write to the database. ## How it works ### Connect a Monday cron as the trigger A scheduled **trigger** fires the project every Monday morning. Each firing spawns a fresh **session** in its own sandbox. One run, one disposable machine, so nothing carries over between weeks and the report is built from scratch each time. ### Give the agent the report playbook What the report contains lives as **skills** and **memory** that travel with the agent: the metric definitions, the queries, the layout, and what counts as a notable move worth calling out. As the metrics we care about change, we write it down and the agent picks it up on the next run. ### Connect the systems the report needs Through scoped **connectors**, brokered server-side so no raw token reaches the model, the agent can: - **Query Postgres read-only** — run the metric queries against a read-only role, pulling this week's numbers and the prior weeks for comparison. - **Compute the deltas** — compare against recent history to find what moved and by how much, inside the sandbox. - **Write the commentary** — turn the deltas into plain-language notes on what changed and whether it's worth attention. - **Post to Slack** — the report and its commentary land in the team channel as a single message. ### Set the guardrails The database connector is **read-only**: the agent queries but cannot insert, update, or delete, and its role is scoped to the metrics tables. It posts to one Slack channel and nothing else. Credentials are encrypted in the secrets manager and injected at runtime, scoped to the agents you grant them to. ### Let each Monday build its own report With that in place, every Monday the agent runs the queries, computes the deltas against recent weeks, writes commentary on what moved, and posts the report to Slack before the team logs on — the numbers and a plain-language read of them in one message. No one spends their Monday assembling it. > **The pattern** > A Monday **trigger** spawns a session with scoped **connectors**: a read-only > role into Postgres and a single Slack channel out. The report playbook is encoded > as **skills** and **memory**. The agent queries, interprets, and posts — with no > path to write to the database. ## Guardrails The agent reads production data, so the access is scoped and contained: - **Isolation.** Every run happens in its own isolated sandbox. The session queries the metrics and drafts the report; only the Slack message is written back out. - **Scoped secrets.** The Postgres and Slack credentials are encrypted in the secrets manager and injected into the sandbox at runtime, scoped to the agents you grant them to. - **Read-only database access.** The database role can select and nothing else — no insert, update, or delete — and it's scoped to the metrics tables. The report cannot change the data it reports on. - **Everything is code.** The agent's configuration, skills, and permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every Monday:** The report built and posted before the team logs on - **With commentary:** What moved and whether it matters, not just numbers - **Read-only:** Production metrics queried, never written The weekly report stops depending on someone having a free Monday, and it arrives with a read of what changed rather than a table to interpret. The team starts the week looking at the movement that matters instead of assembling the numbers by hand. --- <!-- /markdown/use-cases/win-loss-analysis.md --> # How we analyze why deals are won and lost The win-loss agent we run on TaskStation — a weekly cron reads every HubSpot deal closed-won or closed-lost since the last run and posts the patterns by segment, competitor, price, and stage of death to Slack. Canonical page: https://taskstation.co/use-cases/win-loss-analysis Every closed deal is a data point about why customers choose us or choose someone else, but the reason usually lives in one line of a HubSpot field, written by a rep in a hurry between calls. Read one at a time, those lines say nothing. Read across a hundred deals a quarter, they say exactly where we're strong, where a competitor is beating us, and where our own pipeline quietly falls apart. We run a win-loss analysis agent on TaskStation that reads closed deals from HubSpot every week and posts the patterns to our sales Slack channel. It only reads deal data; the single output is the Slack post. This is how we watch our own win rate. - **Team:** TaskStation - **Runs on:** Weekly cron - **Connected systems:** HubSpot · Slack - **Mode:** Read-only · one Slack post per week ## The problem Close reasons and competitor mentions get logged in HubSpot because the CRM asks for them, not because anyone reads them back. A closed-lost deal gets a one-line reason and maybe a competitor field, then disappears into the pipeline history. Nobody aggregates it, so the same loss pattern — a competitor consistently winning on price in one segment, a deal category that reliably dies at the same stage — can repeat for two quarters before a sales leader notices it by gut feel in a QBR. The common approaches don't fix this. A win-rate number on a dashboard tells you the score changed, not why. A quarterly deal-review deck is thorough but looks back three months too late to change this week's playbook. And nobody wants to be the person who reads two hundred close-reason fields by hand every Monday. ## What we built On TaskStation, a weekly cron triggers an agent. It spawns a fresh session (a cloud sandbox) with read-only access to HubSpot, pulls every deal closed-won or closed-lost since the last run, and breaks the outcomes down by segment, competitor, price band, and the stage where lost deals actually die. It synthesizes the patterns into themes and concrete recommendations, and posts the summary to Slack. It writes nothing back to HubSpot. ## How it works ### Run on a weekly cron A **cron trigger** fires the agent once a week. Each firing spawns a fresh **session** in its own sandbox. One week maps to one run on one disposable machine, so the analysis is recomputed from HubSpot's current state every time and nothing carries over between runs. ### Give the agent the analysis rules How we break down a win or a loss lives as a **skill** that travels with the agent: which fields hold the close reason and the competitor, how to band deal sizes, how to bucket segments, and what turns a pile of one-line reasons into a small number of real themes rather than a hundred unique snippets. ### Connect HubSpot read-only Through a scoped **connector**, brokered server-side so no raw token reaches the model, the agent reads: - **Closed-won and closed-lost deals** — every deal that closed since the last run, with amount, segment, and the stage it moved through. - **Close reasons and competitor fields** — the specific HubSpot properties a rep fills in at close, read verbatim. - **Stage history** — where a lost deal was sitting before it died, not just that it died. - **Posts to Slack** — the one weekly summary of themes and recommendations. ### Set the guardrails The agent is **read-only** on HubSpot. It has no write access to any deal, contact, or company record, so it cannot change a stage, a close reason, or an amount. Its only output is the Slack post. Credentials are encrypted in the Secrets Manager and injected at runtime, scoped to the agents you grant them to or written to logs. ### Post the themes and recommendations With that in place, each week brings one Slack post: win rate and loss count for the period, the breakdown by segment, competitor, and price, the stage where deals most often die, and a short list of themes with a recommendation attached to each. The sales team reads it and decides what to change. Nothing is written back to HubSpot automatically. > **The pattern** > A weekly **cron** spawns a fresh session with a read-only **connector** into > HubSpot. The breakdown rules live as a **skill**. The agent reads every > closed deal and writes nothing but the Slack post. ## Guardrails The agent reads every closed deal in HubSpot, so its access is scoped and one-directional: - **Isolation.** Every run happens in its own isolated sandbox. The session is granted access only to HubSpot, and only the Slack post is written back out. - **Scoped secrets.** The HubSpot credential is encrypted in the Secrets Manager and injected into the sandbox at runtime, scoped to the agents you grant them to or the logs. - **Read-only.** The connector into HubSpot is read-only. The agent cannot change a deal's stage, close reason, competitor field, or amount; it can only report on what's already there. - **Report only.** The agent never contacts a rep, a prospect, or a lost deal. It surfaces the pattern; a human decides what to do about it. - **Everything is code.** The agent's breakdown rules, skill, and HubSpot permissions are files in the repo, versioned and changed through a reviewed **change request** rather than a dashboard setting. ## The outcome - **Every week:** Closed deals rescored from HubSpot's current state - **Read-only:** Nothing written back to any HubSpot deal - **4 cuts:** Segment, competitor, price, and stage of death in one post Win-loss reasons that used to sit unread in a HubSpot field now arrive as one weekly summary in the channel where the sales team already works, with the themes named and a recommendation attached to each. The agent only reads; the people decide what to change in the playbook.