orgOS Agent
Why orgOS?
Because the goal is not a chattier bot. It is a cloud agent work system that can be organized, audited, and continuously improved.
中文版 / Chinese default
orgOS Agent is a cloud-native serverless agent runtime built on Cloudflare, and the current experimental vehicle for orgOSv8.
It is not a prompt demo with a chat box attached. It is a long-running, multi-tenant agent platform: real users sign in via OAuth and own their agent fleets; agents connect to real channels, call real tools, dispatch subagents, run on schedules of their own, accumulate their own memory — and every step leaves a verifiable evidence trail.
The platform develops itself: orgOS's day-to-day development, code review, price auditing, and acceptance testing are done by agents running on this very platform.
🟢 Highlights
- ⚡️ Edge-native runtime: runs on Cloudflare Workers, close to users and with minimal operations overhead.
- 🧑🤝🧑 Multi-tenant user product: Google OAuth sign-in (agentthursday.com gateway); per-user isolation of agents, credentials, archives, files, and memory. Owner-scoped fail-closed is a structural constraint, not routing-layer politeness.
- 🤖 Multi-agent orchestration: agents create subagents, fan out dispatches, and run declarative workflows (the plan lives in data, not in the model's head); lineage trees, phase flowcharts, and every subagent's own conversation render live in the UI.
- ⏰ Native scheduled tasks: owner-scoped schedules (interval / daily / weekly) driven by a platform alarm tick that survives deploys; double-fire-proof claiming, per-owner quotas, auto-disable on repeated failure; per-run execution history.
- 🧠 Six-layer memory: short-term context → deterministic anchor compaction → LLM consolidation (extraction / promotion / contradiction pruning / confidence gate) → semantic recall (bge-m3) → archive full-text search (FTS5, char-level CJK) → cross-agent insight promotion. Memory is earned, and inspectable.
- 🗄️ Conversation Archive + FTS5 search: history is never lost;
conversation_search runs on a bm25-ranked full-text index (Chinese natural-language queries work), falls back to LIKE on any failure, and leaves a retrieval audit every time.
- 💰 Cost transparency: per-turn in/out/cached tokens are persisted; the UI computes cost at display time from a maintained price table — which a Pricing Auditor agent on the platform re-verifies against official pages weekly.
- 🔌 BYO models and channels: bring your own provider keys at runtime (AES-256-GCM envelope encryption at rest) and your own Discord bot; live model discovery across providers (Anthropic / DeepSeek / xAI / Zhipu / Workers AI…).
- 🧩 Skillsets as data: capability sets are DB-stored manifests (prompt segments, safety policy, observability declarations); agents can author new skillsets themselves. Tool implementations stay in code; capability composition is data.
- 📡 ChannelHub: multi-channel messages enter a durable inbox and replies leave through an outbox; conversation↔agent binding routes work in real group chats.
- 📚 ContentHub: read external sources (GitHub etc.) and user-uploaded documents (random-nonce fencing against prompt injection) with revision / provenance / permission metadata throughout.
- 🔐 Truthfulness Guard: a claimed tool call with no matching dispatch in the trace gets flagged; multi-agent work is verified with hard evidence (workflow-runs + roster + content-audit), never narration.
- 📉 Model degradation awareness: unstable capability, missing tool calls, or unreliable output are surfaced explicitly instead of papered over.
- 🔎 Evidence / Inspect:
/api/inspect is the black-box replay surface — traces, tool events, retrieval audits, memory-layer state, schedule state, all queryable.
☁️ Cloudflare components
| Component |
Role |
| Workers |
console entry + the agentthursday user gateway (private worker-to-worker service binding). |
| Durable Objects |
per-agent state, registry, channel routing, content registry; alarms drive schedules and sweepers. |
| Durable Object SQL storage |
event log, archive (with its FTS5 index), memory, tasks/usage, schedules, encrypted credentials. |
| Workers AI |
model bindings (bge-m3 semantic recall, built-in chat models). |
| Workflows |
durable execution for agent-run / workflow-executor. |
| Browser Rendering |
headless browsing for agents (page verification, price audits — real browsing). |
| Containers / Sandbox binding |
heavier isolated execution. |
| R2 |
owner-keyed storage for uploaded documents. |
| Workers Assets |
both web UIs (console + user app). |
| Cron Triggers |
platform-level periodic jobs (e.g. the weekly price-table webhook). |
| Wrangler / secrets |
deployment, secrets, runtime configuration. |
🚀 Capabilities
🧑🤝🧑 1. Multi-tenant user product (agentthursday.com)
orgOS is not just an operator console — it is a real user product:
- Google OAuth sign-in with approval gating; the gateway reaches the console over a private service binding, so the admin surface is never on the public path
- Every user owns their agent fleet, model credentials, conversation archives, uploaded documents, and share links
- All reads and writes are owner-scoped fail-closed: an unresolvable owner yields empty/404, never a fallback to the admin view
- Chat-first UI: history cards, a streaming current card, subagent cards, token/cost chips, deep links, external share links
🤖 2. Multi-agent orchestration
- Agents can
agent_create subagents (inheriting the manager's model and owner by default), dispatch tasks, and collect pushed execution summaries
- Declarative workflows: the plan is data (phases / dependencies / budgets); the executor walks the graph and the run tree renders live
- Spawned-agent lifecycle is managed: task-scoped by default, idle auto-archival, lineage tree visualization — the roster does not grow without bound
- Subagent insights are promoted into the parent's memory at finalize, with
subagent:<id> provenance and fail-closed owner checks
⏰ 3. Native scheduled tasks
"Send me a digest every morning" is table stakes for an agent product; orgOS implements it platform-natively:
- An owner-scoped
scheduled_task table on the registry DO with a 60s alarm tick (idempotently re-armed; survives deploys)
- Claim-before-dispatch: overlapping ticks structurally cannot double-fire; intervals anchor on the due slot (no drift); missed slots are skipped, not replayed
- Built-in safety valves: per-owner cap, 900s minimum interval, auto-disable after 5 consecutive failures
- A scheduled run is a completely normal task — activity, traces, token/cost accounting all apply; per-run history (dispatched → ok/failed) is queryable
- One-click management from every agent's workspace title bar (local-time input, UTC storage)
🧠 4. Six-layer memory
Memory is not one big string stuffed into a prompt — it is layered, gated, and verifiable:
- Short-term context: the live window, deterministically anchor-compacted under pressure (no lossy summarization)
- Conversation Archive: archived before
new / reset; history is structurally unlosable
- LLM Consolidation: periodic extraction → confidence gate (≥0.8) → dedup → contradiction pruning (a supersede needs ≥2 non-trivial shared tokens to authorize a soft delete) → promotion into
agent_memories
- Semantic recall: bge-m3 vector search over agent memories via the
recall tool
- FTS5 archive search:
conversation_search on a bm25-ranked full-text index with char-level CJK segmentation and phrase queries; any failure falls back to LIKE
- Cross-agent promotion: subagent insights push up to the parent's store with provenance
Supporting cast: consolidation ledger, retrieval audits, per-layer inspect surfaces, and a dual-track (local vs. external shadow memory) comparison experiment.
💰 5. Cost transparency
- Per-turn in / out / cached-input tokens are persisted at finalize (API billing semantics: every step's input counts the full context)
- Cost is computed at display time — price changes never require data migration; unknown models show tokens only; cached and fresh input are priced separately
- The price table itself is re-verified weekly by a Pricing Auditor agent (real browsing, hard-evidence checks), with a webhook notification and a human gate on adoption
📡 6. Real channel collaboration
- Discord messages enter a durable inbox; the agent knows who is speaking and whether it is being addressed; replies return through the outbox to the original conversation
- Conversation↔agent binding routes (owner-scoped) let multiple agents share one channel with clear ownership
- Users can attach their own Discord bot by pasting a token at runtime — no redeploy
📚 7. ContentHub and documents
- External sources (GitHub etc.): list, read, search, multi-source fan-out — every read carries revision / provenance / permission / cache status
- User-uploaded documents: converted to markdown, stored owner-keyed in R2; every read is wrapped in a random-nonce fence so injected instructions are treated as data, not executed
- The agent's own workspace and external sources are strictly separated — no hallucinated reads
🛠️ 8. ToolHub and Skillsets
- Tool calls are executable, observable, reviewable: workspace I/O, code execution, sandbox, browser, memory, retrieval, dispatch… every significant action leaves an event
- Skillsets are data: manifests (skills / prompt segments / safety policy / evidence protocol) live in the DB; an agent can author a new skillset and attach it to another agent — the price-audit skillset was born exactly this way
- Safety policy is a first-class part of the manifest: path denylists and cross-repo write bans are declarative constraints
🔐 9. Truthfulness and degradation awareness
- Truthfulness Guard: claiming a tool call that has no dispatch → flagged; multi-agent claims are verified against workflow-runs + roster + content-audit hard evidence
- Degradation awareness: tool-calling regressions, unstable streaming, or failed structured output are recorded, aggregated, and surfaced to the user
- The goal is not for the agent to always look smart, but to tell you honestly when it is not reliable
🔎 10. Evidence / Inspect
/api/inspect is the black-box replay surface: trace events, actual tool calls, content-read chains, archive / retrieval / hygiene / consolidation audits, schedules and their run history, memory-layer state, FTS index watermarks — all queryable.
The point is not whether it sounds right, but whether it actually did the thing. orgOS Agent defaults to the verifiable side.
🧪 Try it
🛫 Deployment
Current deployment notes live here:
💻 Development
npm install
npm --prefix web install
npm --prefix gateway/web install
npm run typecheck # all tsconfigs
npm test # real node:sqlite unit tests (800+)
npm run build:web # console UI
npm run dev
npm run deploy # console; gateway deploys via wrangler separately
Secrets are managed through Wrangler and local development var files. Do not paste tokens into chat, logs, README examples, task reports, or commits.