One Gateway.
All Harnesses.
Bridge your IDEs, code editors, and agent workflows to your local AI coding CLIs through a unified, OpenAI-compatible local API.
curl http://127.0.0.1:3500/v1/chat/completions \
-H "Authorization: Bearer afaq_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "REPLACE_WITH_ID_FROM_V1_MODELS",
"messages": [
{"role": "system", "content": "You are a code refactoring expert."},
{"role": "user", "content": "Refactor parse_stream() to handle reconnects cleanly."}
],
"stream": true
}'
Run Supported AI Harnesses on Your Machine
The gateway maps OpenAI-format requests directly to your local CLI subprocesses. Each harness is registered in the core registry.
claude
Runs claude -p with stream-json output and plan permission mode; exposes the adapter’s standard model choices when installed.
codex
Runs codex exec --json in a read-only sandbox; reads ~/.codex/models_cache.json with fallback choices.
opencode
Runs opencode run --format json; discovers models with opencode models.
cmd
Runs cmd --print with JSON output and plan permission mode; discovers models with cmd --list-models.
agy
Runs agy --print with JSON output and plan mode; discovers models with agy models and has documented fallback choices.
pi
Runs pi -p with JSON output; discovers provider/model pairs with pi --list-models.
minimaxRegistered generic text adapter for the minimax executable (install: npm install -g minimax-cli).
aiderRegistered generic text adapter for the aider executable.
clineRegistered generic text adapter for the cline executable (install: npm install -g @cline/cli).
cursor-agentRegistered generic text adapter for the cursor-agent executable.
grokRegistered generic text adapter for the grok executable (install: npm install -g @xai-official/grok).
kimiRegistered generic text adapter for the kimi executable (install: npm install -g @moonshot/kimi-code).
ompRegistered generic text adapter for the omp executable.
qodercliRegistered generic text adapter for the qodercli executable.
vibeRegistered generic text adapter for the vibe executable.
copilotRegistered generic text adapter for the copilot executable (install: npm install -g @github/copilot).
ozRegistered generic text adapter for the oz executable.
zcodeRegistered generic text adapter for the zcode executable.
From Application Call to CLI Execution
A deterministic, asynchronous request lifecycle that translates standard OpenAI payloads into isolated subprocess runs.
1. Standard API Ingestion
Client applications (Cursor, Cline, OpenAI SDKs, scripts) send standard POST /v1/chat/completions or GET /v1/models requests over HTTP.
2. Admission & Quota Gate
The gateway resolves authentication, verifies per-key model allow-lists, and atomically reserves usage quota in SQLite before touching the CLI.
3. Subprocess Execution
The specific HarnessAdapter spawns the local CLI as an isolated POSIX subprocess, draining stderr into a 64KB bounded ring buffer.
4. Streaming SSE Delivery
CLI stdout lines are parsed incrementally, mapped into standard delta chunks, streamed via SSE, and closed with atomic quota finalization.
Engineered for Reliability and Control
Grounded in verified architectural capabilities designed for developer security and production stability.
Streaming Chat Completions
Full Server-Sent Events (SSE) support with periodic SSE heartbeats, client cancellation propagation, and per-stream identity isolation.
Dynamic Model Discovery
Live CLI discovery runs at startup and on-demand via dashboard, serving verified model IDs from hot in-memory cache with Redis mirroring.
Atomic Quota Reservations
Two-phase reservation pattern (QuotaReservation + UsageRecord) guarantees concurrent requests cannot overrun daily or monthly limits.
Encrypted Credential Profiles
Provider secrets and tokens are encrypted at rest using Fernet symmetric encryption (CREDENTIALS_KEY). Plaintext secret material is never displayed again.
SQLite with Optional Redis
Runs on SQLite without an external database service for local single-node setups. An optional Redis instance can be configured for shared rate limits and cross-pod SSE replay.
Bilingual Web Dashboard
Intuitive browser dashboard in Arabic and English for conversations, harness installation, API key rotation, and usage metrics.
Administrator OS Terminal
Interactive browser PTY shell powered by WebSocket and os.posix_spawn. Enables direct CLI management and maintenance from the dashboard.
Subprocess Hardening
Deterministic subprocess timeouts, zombie process reaping, and bounded 64KB stderr drainers prevent OS pipe deadlocks.
Structured JSON & Tool Calling
Normalized function and tool call objects (response_format and tools) mapped transparently between client and harness CLI output.
Strict Trust Boundaries by Design
The gateway treats host security as a core architectural responsibility with defense-in-depth boundaries.
Authentication Boundary Separation
Hardened AuthAPI keys (afaq_...) are restricted exclusively to /v1/* inference routes. Dashboard, administration APIs, and terminal sessions require signed JWTs.
Loopback-First & Private Deployment
Network DefenseBinds to 127.0.0.1:3500 by default. Docker Compose strictly publishes to the host loopback interface, keeping internal Redis networks private.
Fail-Fast Production Secrets
Zero DefaultsWhen DEBUG=false, startup fails immediately if SECRET_KEY or CREDENTIALS_KEY are missing, shorter than 32 characters, or match example placeholders.
Trusted Proxy Resolution & Bcrypt Ceiling
IP & Crypto SafeClient IP resolution strictly uses socket peers unless explicitly whitelisted in TRUSTED_PROXIES. Bcrypt enforces a 72-byte ceiling to prevent silent password truncation.
Up and Running in Under Five Minutes
Choose your preferred workflow. Both paths guide you directly to the first-time administrator setup wizard.
-
1 Clone repository
Fetch the official repository and switch into the project directory.
git clone https://github.com/afaqhost/afaq-harness-gateway.git cd afaq-harness-gateway -
2 Run interactive setup wizard
The automated wizard verifies Python 3.10+, pip, venv, creates .env with strong secrets, and prepares dependencies.
make setup -
3 Start gateway development server
Launches Uvicorn on http://127.0.0.1:3500 with hot reload and automatic database initialization.
make dev
-
1 Run setup for Docker
Prepares configuration and container requirements using the setup script.
make setup ARGS="--path=docker" -
2 Build and start containers
Spins up the unprivileged Node user container with loopback publishing on port 3500.
docker compose up --build
Next steps after launch:
1. Open http://127.0.0.1:3500/setup to create the first administrator account.
2. In the dashboard, open Harnesses to review or install CLI tools.
3. Open API Keys, generate a key, and query GET /v1/models.
Beta Release Status (0.1.0-beta.1)
Core gateway and proxy workflows are tested (331 tests without warnings). APIs, configurations, and database schemas may still evolve prior to the 1.0 stable release. Always back up your SQLite database before updating.