Featured

Deploy OpenClaw in 60 seconds — 20% off logoDeploy OpenClaw in 60 seconds — 20% off

Launch OpenClaw on Hostinger in about 60 seconds and keep your agent live 24/7. Our referral link gives you 20% off, no coupon code needed.

Launch on Hostinger
Run your Hermes agent on Hostinger, fully managed logoRun your Hermes agent on Hostinger, fully managed

Launch Hermes on Hostinger in one click, fully managed, no VPS knowledge needed. Use code ZACAARON10 for 10% off.

Launch on Hostinger
Crawl and scrape any site into clean data, 10% off logoCrawl and scrape any site into clean data, 10% off

Firecrawl crawls and scrapes any site into clean markdown for your agent. Get 1,000 free credits, and new users get 10% off their first purchase.

Try Firecrawl free
6,000+ web scrapers for your AI agent, start free logo6,000+ web scrapers for your AI agent, start free

Apify gives your agent live web data: 6,000+ prebuilt scrapers and actors, MCP-ready. Sign up free with $5 in usage credits.

Try Apify free
One API to scrape, enrich, and extract the internet. logoOne API to scrape, enrich, and extract the internet.

Context.dev gives your agents a single API to scrape, enrich, and extract live web data — no proxies, no parsers, no maintenance.

Start building free
SetupClaw: done-for-you OpenClaw for founders & exec teams logoSetupClaw: done-for-you OpenClaw for founders & exec teams

White-glove OpenClaw for founders and exec teams (4–50+ employees): we install, harden, integrate your tools, and maintain it — secured from day one.

Get it set up for you
SEO data APIs for your agent, $1 free credit logoSEO data APIs for your agent, $1 free credit

DataForSEO gives your agent live access to SERP results, keyword data, backlinks, and on-page SEO data through one API. New accounts get a $1 credit, good for up to 20,000 keyword or backlink lookups.

Try DataForSEO free
Reach 48,000+ AI builders

A flat monthly placement in front of developers actively installing AI tools. No lock-in, cancel anytime.

Advertise here

Works with

Claude CodeClaude DesktopCursorVS CodeClineCodex CLIOpenClaw+ any MCP client

Install to Claude Code

This server doesn't publish a one-line install command. Follow the setup in the source repository.

Summary

An agent-native shell as an MCP server, designed for LLM agents (like Claude Code) to execute commands with structured output, lazy detail retrieval, and effect tracking, minimizing token usage.

README.md

veil-mcp

![CI](https://github.com/vkmtx/veil-mcp/actions/workflows/ci.yml) ![npm](https://www.npmjs.com/package/veil-mcp) ![License: MIT](LICENSE) ![Node](package.json) ![MCP](https://modelcontextprotocol.io)

A shell built for AI agents, not humans. veil is an MCP server that gives a coding agent (Claude Code, Cursor, …) a shell whose results come back as structured data — typed effects, one-call verification, addressable output, and a real undo — instead of a wall of scrollback text.

A normal terminal dumps everything and the agent re-greps fragile text, round-trips for state, and can't undo a mistake. veil turns each command into a quiet, structured result — and adds three things a plain shell simply can't.

quiet-by-default · effects-as-data · lazy detail · real safety net

Why it's good — in three numbers

You can approximate most of veil with Bash + truncation + careful prompting. The reason to actually adopt it is the three things a shell genuinely cannot do — each quantified, each reproducible with npm run metrics:

| | What you get | The number | |--|--|--| | ✅ Verify in one call | expect: { exit: 0, file_exists: "dist/index.js" } folds run → check → grep into a single call; effects come back typed, so "what changed?" needs no git status. | 55% fewer round-trips (11 → 5) — a scenario model over 5 hand-picked common tasks, not a live measurement | | ♻️ Checkpoint & roll back | sh_checkpoint / sh_restore wrap a risky refactor in an undo — a copy-on-write clone on APFS. | clone ~1.5× faster, ~0 MB vs a 60 MB rsync copy | | 🔒 Kernel sandbox | sandbox: true confines writes to cwd + temp (optionally no network) — and refuses to run rather than go unconfined. | 5 / 5 escape attempts blocked (in-cwd write still lands) |

And one honesty number — because quiet must never mean dishonest: a failure buried in the hidden middle of a long log is still surfaced, at 100% recall on a labeled corpus (SIGSEGV, CONFLICT, ! [rejected], timed out, …, none of which contain the word "error").

Everything else — quieter output, addressable detail, retry, blast-radius classification — is genuine convenience on top, not the moat.

Quickstart

No clone, no build — runs via npx:

claude mcp add veil -- npx -y veil-mcp
npx -y veil-mcp init     # adds the "prefer sh_run" nudge to this project's CLAUDE.md

<details><summary>Other agents (Cursor, Windsurf, Zed) · from source</summary>

// MCP server config for any MCP-speaking agent
{ "mcpServers": { "veil": { "command": "npx", "args": ["-y", "veil-mcp"] } } }
# from source
git clone https://github.com/vkmtx/veil-mcp && cd veil-mcp
npm install                          # builds dist/ via the prepare script
claude mcp add veil -- node "$(pwd)/dist/index.js"
# dev, no build step:  npm run dev    (tsx src/index.ts)

npx -y github:vkmtx/veil-mcp runs straight from GitHub. veil init is idempotent and touches only CLAUDE.md — see Adoption. </details>

The tools

| Tool | What it does | |------|--------------| | sh_run | Run a command → quiet structured result: exit, duration, files changed, token-aware stdout/stderr. background: true starts a long-running process (dev server, --watch) and returns { id, pid, status: "running" } immediately instead of blocking. The workhorse. | | sh_logs | Poll a background run's output — incremental, per-stream byte cursor (stdout_cursor/stderr_cursor), plus status/exit/signal. Never re-dumps what was already tailed. Omit id for the newest live run. | | sh_kill | Stop a background run. Signals the whole process group; SIGTERM escalates to SIGKILL after 2s. Killing an already-exited id is idempotent. Omit id for the newest live run. | | sh_detail | Pull the full stored output of a past run — no re-run. Disk-backed, so it survives a server restart. match=<regex> greps the stored stream for a value condensing hid. Omit id for the most recent run. | | sh_checkpoint / sh_restore | Snapshot a directory and roll back. Owner-only (0700) storage, published atomically. Restore refuses a target dir different from where the checkpoint was taken. Omit label to auto-number the checkpoint / restore the newest. | | sh_checkpoints | List checkpoint labels. |

Every id/label is optional and defaults to the run or checkpoint you almost certainly mean; a wrong one answers with the values that are addressable. sh_run also accepts cmd as an alias for command. These are not conveniences — a 30-day audit of real agent sessions found argument shape, not execution, behind most failed calls.

See it

// build AND verify the artifact exists — one call, no follow-up ls
sh_run { "command": "npm run build", "expect": { "exit": 0, "file_exists": "dist/index.js" } }

// confine a risky script to cwd, deny network, block reads of secret dirs
sh_run { "command": "./untrusted.sh", "sandbox": { "network": false, "protect_secrets": true } }

// dry-run in a CoW clone — see the cwd-relative diff, real cwd untouched
sh_run { "command": "rm -rf build && npm run generate", "preview": true }

// start a dev server detached, tail its output incrementally, stop it when done
sh_run  { "command": "npm run dev", "background": true }         // → { id: "cmd12", pid, status: "running" }
sh_logs { "id": "cmd12", "stdout_cursor": 0 }                     // poll again with the returned cursor for only NEW output
sh_kill { }                                                       // no id = the newest live run → { status: "terminating" }

// undo a refactor — label optional both ways (auto-N, then newest-first)
sh_checkpoint { "label": "pre-refactor" }
sh_restore   { "label": "pre-refactor" }

// find a value a condensed 50k-line log hid — no re-run, no full dump
sh_detail { "id": "cmd9", "selector": "stdout", "match": "ERROR|version=" }

<details><summary><b>All <code>sh_run</code> options</b> — expect, sandbox, retry, trace, …</summary>

| Option | Effect | |--------|--------| | command | The shell command (required). | | cwd | Working directory (defaults to the server's cwd). | | full | Return uncondensed stdout/stderr inline (escape hatch from condensing). | | timeout_ms | Per-command timeout (default 120s). On expiry the whole process group is killed (SIGTERM→SIGKILL), so a compound command's grandchildren (sleep 5; …) are reaped too. | | expect | Post-conditions verified in the same call: exit, stdout_contains, stdout_matches, stderr_empty, file_exists, file_absent, changed, max_ms. Failures surface in assert_ok + assertions_failed — no second ls/grep/git status. | | retries / retry_on_exit / backoff_ms | Declarative retry; attempts is reported when > 1. | | sandbox | Real OS sandbox. true confines file writes to cwd + temp; { network: false } also denies network; { writable: [...] } adds roots. { protect_secrets: true } or { deny_read: [...] } also blocks reads of configured secret dirs (~/.ssh, ~/.aws, …) — macOS deny file-read, Linux --tmpfs mask; sets secrets_protected: <n>. Scoped: it blocks the listed paths, not a proof against all exfiltration. Refuses to run if unavailable — never executes unconfined. Sets sandboxed: true. | | preview | Dry-run in a disposable CoW clone of cwd — the command runs inside the clone, you get the cwd-relative files_changed, and the real cwd is never touched (nothing is promoted). Honest scope: absolute-path / parent-dir / network effects are not captured and may happen for real — this is not a sandbox (combine with sandbox:true for containment). Refuses if the cwd can't be cloned. Sets preview: true + preview_warning; a diff too large to buffer reports preview_effects_incomplete rather than silently claiming nothing changed. | | trace | Structured FS/syscall trace (Linux strace). Surfaces trace_summary (paths read/written + syscall count); full trace via sh_detail selector=trace. Best-effort: no tracer → command still runs, trace_unavailable: true. Bounded by VEIL_MAX_STREAM_BYTES; an overflowing trace sets trace_truncated: true. | | scrub_env | Strip credential-shaped env vars (_TOKEN/_KEY/AWS_/…) from the child's environment before spawn. Auto-on whenever sandbox.protect_secrets/deny_read is set. Reports secrets_env_scrubbed — a count, values are never echoed. | | no_store | Keep this run memory-only: addressable via sh_detail for the session, but never written to disk. Sets stored: "memory-only". | | background | Start a long-running process (dev server, --watch) — returns immediately with { id, pid, status: "running" } instead of blocking until exit. Poll with sh_logs id=<id>, stop with sh_kill id=<id>. No stdin/TTY. Refused together with options that need completion (expect, preview, trace, retries, full, timeout_ms); still honors cwd/sandbox/scrub_env/no_store. Capped by VEIL_MAX_BG_PROCS (default 16); live children are reaped on server shutdown. |

</details>

<details><summary><b>Result fields</b> — emitted only when relevant (the quiet contract)</summary>

id, exit, ok, ms; then attempts, stdout_lines/stderr_lines (TRUE emitted counts), files_changed, timed_out, stdout_truncated/stderr_truncated, stdout_binary/stderr_binary, sandboxed, secrets_protected/secrets_unprotected, secrets_env_scrubbed, stored ("memory-only" under no_store), preview/ preview_method/preview_warning/preview_effects_incomplete, trace_summary/ trace_unavailable/trace_truncated, assert_ok/assertions_failed, advice, hint, and the condensed stdout/stderr. A background: true run instead returns id, pid, status: "running", and a hint pointing at sh_logs/sh_kill.

</details>

<details><summary><b>Configuration</b> — every tunable is an env var, no rebuild</summary>

| Env var | Default | Meaning | |---------|---------|---------| | VEIL_INLINE_MAX_LINES | 45 | stdout shorter than this (lines) is returned whole | | VEIL_HEAD_LINES | 20 | lines kept from the top when condensing | | VEIL_TAIL_LINES | 20 | lines kept from the bottom when condensing | | VEIL_MAX_LINE_CHARS | 1000 | max chars of any single inline line (longer → capped with a pointer) | | VEIL_STDERR_INLINE_ON_FAIL | 60 | on failure, show up to this many stderr lines inline | | VEIL_TIMEOUT_MS | 120000 | default per-command timeout (0 = none) | | VEIL_MAX_STREAM_BYTES | 5000000 | max bytes stored per stream (older dropped) | | VEIL_MAX_RECORDS | 500 | max addressable run records (oldest evicted) | | VEIL_MAX_STORE_BYTES | 268435456 | total disk-store byte budget (256MB), on top of VEIL_MAX_RECORDS — oldest evicted by mtime | | VEIL_STATE_DIR | auto | record store base ($XDG_STATE_HOME/veil~/.local/state/veil$TMPDIR/veil). none/off/memory/0 = memory-only | | VEIL_RECORD_TTL_MS | 86400000 | persisted records older than this are pruned on boot (0 = keep) | | VEIL_EFFECTS | true | compute the git effect-diff (set 0 to skip in huge repos) | | VEIL_MAX_BG_PROCS | 16 | max concurrent live background: true processes |

</details>

Output honesty

Condensing saves tokens, but it must never hide signal. So:

  • A failure buried mid-stream is surfaced — including crash idioms with no

error/fail keyword (Segmentation fault, SIGSEGV, CONFLICT, ! [rejected], undefined reference, timed out). More distinct signals than fit inline? The marker reports the true total with a +N more note, never a silent cap. Best-effort, but measured: 100% recall on a labeled corpus (see below).

  • A byte-capped stream is labeled and never shows its tail as the head.
  • stdout_lines/stderr_lines are the true emitted count; binary output is

base64-flagged, not mangled to mojibake.

  • advice never blocks — it nudges on the highest-signal issue (widen a sandbox

denial, checkpoint before an unconfined destructive command, use raw Bash for an interactive tool).

Safety

sh_run runs arbitrary shell commands with your privileges, and exposes the server's full environment (secrets included) to them. It's a shell — run it in trusted contexts. Two opt-in layers harden the risky cases:

  • Kernel sandbox (sandbox: true) — the real boundary. Confines writes to

cwd + temp via macOS sandbox-exec (Linux bubblewrap / Landlock, experimental), optionally denies network (Linux bwrap also masks /run//var/run, so a Docker/Podman socket isn't a bypass), blocks reads of secret dirs, and refuses to run rather than go unconfined. Honest scope: solid on macOS; Linux bwrap needs unprivileged user namespaces, which containers / Codespaces / Ubuntu 24.04+ often restrict — there veil falls back to a namespace-free Landlock backend (via landrun, kernel 5.13+) that write-confines where bwrap can't, and still reports unavailable (refusing) if neither works. The Landlock path is write-confine only: it refuses network-deny / secret-read-confine rather than fake them. The default non-sandboxed path works everywhere.

  • Guard hook (hooks/veil-guard.sh) — a **routing nudge, not a security

boundary.** It steers verbose/dangerous Bash toward sh_run, but it is fail-open and VEIL_BYPASS-able and never stops a command from running. Real containment is the sandbox above.

<details><summary>Enable the guard hook</summary>

A PreToolUse guard that hard-blocks only verbose (installs / builds / test runners — npm/pnpm/yarn/bun/deno/uv/pip/cargo/go/…, plus docker build/buildx/compose build) or dangerous (rm -rf, dd, mkfs, shred, find -delete, raw-device writes) Bash, steering it to sh_run. Commands sh_run can't help with are explicitly allowed through to raw Bash: long-running dev/watch/start servers (incl. bun run dev, docker compose up), backgrounded jobs (trailing &), process management (kill/pkill), and interactive/TTY tools (vim/less/top/tail -f). It is fail-open (any parse error → allow, so a bug can never block all Bash), with an escape hatch: prefix a command with VEIL_BYPASS=1 to force raw Bash.

It classifies what the shell will execute, not what the command string contains: heredoc bodies and quoted strings are stripped before matching (so a commit message mentioning "build", or grep -E '"(tsc|build)"' package.json, is not a build), and every tool name must sit at executable position — start of command, after an operator, or behind a runner like sudo/timeout/npx — so grep -rn "HttpApiGroup.make" src passes while npx vitest run still blocks. Deleting a regenerable build artifact (rm -rf .next|dist|build|out|coverage|.turbo|node_modules/.cache, relative or under an absolute project path) is not treated as dangerous, which keeps the dev-server restart idiom (pkill …; rm -rf .next; nohup next dev …) on the allow path; anything unresolvable — a glob, .., ~, $VAR, a root-level path, or one non-build target in the list — still blocks. Enable globally in ~/.claude/settings.json:

{ "hooks": { "PreToolUse": [
  { "matcher": "Bash",
    "hooks": [{ "type": "command",
      "command": "/bin/sh '/ABSOLUTE/PATH/veil-mcp/hooks/veil-guard.sh'" }] }
] } }

Takes effect on the next Claude Code restart. Remove the entry to disable. </details>

Adoption

veil is opt-in and complements Bash — its value lands only when the agent actually reaches for sh_run, and an agent left to itself often defaults to raw Bash. Two levers close that gap: the nudge (veil init writes a short CLAUDE.md block — soft, zero-friction) and the guard hook (stronger, per-machine). There's no native integration yet, so one must be configured; or skip both and call sh_run directly.

Reproduce every number

Don't take the numbers on trust — no account, all local:

git clone https://github.com/vkmtx/veil-mcp && cd veil-mcp && npm install
npm test          # 429+ smoke assertions over a live stdio server (prints its tally; some platform-gated)
npm run metrics   # the value numbers below
npm run backtest  # byte-savings regression (bulk-condense ratio + per-command overhead floor)
npm run bench     # detailed 5-dimension benchmark (economy, latency, per-feature, condense, session)

| Metric | Result | What it measures | |--------|--------|------------------| | Agent turns saved | 55% fewer round-trips (11 → 5) — a scenario model, not a live measurement | MCP calls collapsed by expect + effects + retry across 5 hand-picked common tasks (bench/metrics-data.ts) — counts calls, not bytes, so it holds as context windows grow | | Sandbox escapes blocked | 5 / 5 | adversarial outside-cwd / spawned-child / symlink / network writes denied by the kernel; a legitimate in-cwd write still lands (selective, not deny-all) | | Signal recall | 100% on 10 fixtures | buried failures surfaced from the elided middle, incl. non-keyword crash idioms | | Checkpoint cost | clone ~1.5× faster, ~0 MB vs rsync 60 MB | CoW clone latency + disk vs the rsync mirror (macOS / same-volume APFS) |

The deterministic rows (turns, recall) are asserted in the smoke suite from the same fixtures, so the published figures can't silently drift. Timing rows are machine-dependent. CI runs the whole suite on macOS and Linux (with bubblewrap + strace), so the Linux-only sandbox and trace paths are exercised too.

<details><summary>Roadmap &amp; architecture</summary>

| | Feature | Status | |---|---------|--------| | I / J / H | token-aware output · addressable detail (sh_detail, match) · effect diff | ✅ done | | G / M | inline assertions (expect) · declarative retry/timeout | ✅ done | | B / K-lite | blast-radius classification (read-only → destructive) gating every sh_run | ✅ done | | C / C+ | checkpoint / rollback · atomic CoW clone (same-volume APFS; cross-volume falls back to rsync, reported honestly) | ✅ done | | K | real sandbox (macOS sandbox-exec) | ✅ done | | J+ | disk-backed record store (survives restart, TTL-pruned) | ✅ done | | K-read / P | secret read-confine (sandbox.protect_secrets) · dry-run preview (CoW clone, real cwd untouched) | ✅ done | | — | tool surface pruned to what agents actually call (sh_plan, sh_history removed in 0.8.0 — 0 calls across a 30-day audit of 3.5k real sessions) | ✅ done | | K+ / A | Linux sandbox (bubblewrap) · structured trace (strace) | 🧪 experimental — validated on Linux CI | | K++ | namespace-free Linux sandbox (Landlock via landrun) — write-confine in containers/Codespaces where bwrap can't | 🧪 experimental — arg-builder unit-tested | | — | background / long-running processes (background: true, sh_logs, sh_kill) for dev servers / watchers | ✅ done | | — | streaming / PTY (interactive processes) | 🔭 planned |

See CHANGELOG.md for version history and ARCHITECTURE.md for the module/feature map. (Why an MCP server and not a shell fork? Most of the value is a presentation/orchestration layer that ships natively to how an LLM already consumes tools — in weeks, not a 200k-line C fork — and the kernel/FS bits, veil drives rather than reimplements.) </details>

Community

Early project, good time to shape it:

License

MIT — see LICENSE.

---

v0.7.1 — experimental, single-author. Adds background/long-running processes (sh_logs / sh_kill), env-secret scrubbing (scrub_env), memory-only runs (no_store), and two correctness/security audit passes (CHANGELOG): 429+ smoke assertions + backtest + value metrics, green on macOS and Linux CI. Judge it by the reproducible suite above, not its age.

See related servers & alternatives →

Related MCP servers

Browse all →

Related guides

Hand-picked reading to help you choose and use Developer Tools servers.