Overview
AI Agent powers phone AI (text and voice), optional Live Chat assistance, and a public Runs API for backends. Create an AI Agent API token with permission ai:invoke so your server can only invoke AI—not email, chat minting, or billing.
| Item | Value |
|---|---|
| Service | AI Agent |
| Status | Public Runs API generally available · dashboard speak/voice live |
| Requires (public runs) | AI Agent entitlement |
| Requires (phone speak/voice) | Business Phone + AI Agent |
| Permissions | ai:invoke |
| Live token shape | umk_live_ai_<prefix>_<secret> |
| Test token shape | umk_test_ai_<prefix>_<secret> |
| Dashboard | Dashboard → AI |
| OpenAPI | /developers/openapi-ai-agent.json |
| Docs (AI) | /developers/ai-agent.md |
Get started
- Enable AI Agent in Dashboard → Billing. For phone speak/voice also enable Business Phone.
- Open Dashboard → AI to configure the agent name and automatic-agent behavior.
- Optional: enable Live Chat auto-replies under chat AI settings (requires AI Agent entitlement).
- Dashboard → Developers → create an API token. Product: AI Agent. Permission: ai:invoke. Environment: live or test. Optional IP allowlist.
- Copy the secret once. Store it only on your server. Fetch third-party data with that system's credentials, then call POST /api/v1/ai/runs with input + context.
API tokens
An API token lets an agent access AI Agent only—not Live Chat session minting, Business Email, or billing. Prefer ai:invoke alone for invoke jobs.
Authorization: Bearer umk_live_ai_<prefix>_<secret>
# or
X-API-Key: umk_live_ai_<prefix>_<secret>
API base: https://umneyconnect.com/api| Permission | Allows |
|---|---|
| ai:invoke | Start and read AI runs on the public AI API (when live) |
Full token guide: /developers/api-tokens
Concepts
| Resource | Description |
|---|---|
| Speak | Text in → AI text response (tenant knowledge + agent config) |
| Speak stream | Same speak result delivered as SSE (text/event-stream) |
| Voice (REST) | audioBase64 or text in → text + audioBase64 out (STT → generate → TTS) |
| Voice (WebSocket) | Duplex audio/text turns over /api/ai/voice/ws with a dashboard JWT |
| Chat AI settings | Auto-reply mode for Live Chat (needs AI Agent entitlement to enable) |
| Inbox suggestion | Agent-requested reply suggestion on a Live Chat channel |
| Agent run (public) | umk_* unit of work under POST/GET /v1/ai/runs |
| Third-party context | Caller-supplied knowledge (CRM/Nest/…). Connect does not crawl external systems |
Third-party integration
Connect does not crawl Nest, CRM, or other SaaS. Your backend authenticates to those systems, fetches records, and injects facts into AI via context. The AI Agent token only authorizes Connect AI — never the third party.
- Issue an AI Agent token (ai:invoke) and store it on your server only.
- In your Nest/CRM backend, authenticate to the third-party system with that system's own credentials.
- Fetch the records the user needs (order, booking, profile, etc.) and serialize a short factual summary.
- Call POST /api/v1/ai/runs with input (user question) and context (those facts). Optional metadata for correlation IDs only.
- Return outputText to your client. Never embed the Connect token in the mobile/browser app.
| Field | Rule |
|---|---|
| input | Required · user question · max 4000 chars |
| context | Optional · third-party facts · max 24000 · treated as untrusted data |
| metadata | Optional · string values only · correlation · not auto-fetched |
| mode | text only today |
| Execution | Synchronous — create response includes final status + outputText |
Nest / server recipe
# Your server
CONNECT_API_BASE=https://umneyconnect.com/api
AI_API_KEY=umk_live_ai_…
# 1) Fetch from third party with THEIR credentials
# order = await nestCrm.getOrder(orderId)
# 2) Invoke Connect AI with injected context
curl -X POST "$CONNECT_API_BASE/v1/ai/runs" \
-H "Authorization: Bearer $AI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "text",
"input": "Where is my order?",
"locale": "en",
"context": "Order ORD-123 status=shipped carrier=DHL tracking=JD014…",
"metadata": { "orderId": "ORD-123", "source": "nest-safety" }
}'
# 3) Response (succeeded)
# {
# "id": "…",
# "status": "succeeded",
# "outputText": "…",
# "input": "Where is my order?",
# "context": "Order ORD-123 …",
# "metadata": { "orderId": "ORD-123", "source": "nest-safety" }
# }Dashboard AI API (available now)
Uses the signed-in workspace session (JWT). Phone speak/voice require both Business Phone and AI Agent entitlements. Models run on Workers AI (text, Whisper STT, Aura TTS).
| Method | Path | Purpose |
|---|---|---|
| POST | /api/tenants/{tenantId}/ai/speak | Text AI response |
| POST | /api/tenants/{tenantId}/ai/speak/stream | SSE text response |
| POST | /api/tenants/{tenantId}/ai/voice | Voice or text → text + audioBase64 |
POST …/ai/speak
{
"text": "Where is my order?",
"locale": "en",
"automaticAgent": true,
"context": "Optional third-party facts your backend already fetched"
}Response: { text, done: true }. Empty text defaults to a greeting prompt. Inference budget ~30s; misconfigured AI binding returns 503. sessionId may appear on the wire but is not used for persistence today.
POST …/ai/speak/stream
Content-Type: text/event-stream
# One SSE event with the full result (not token-by-token streaming)
data: {"text":"…","done":true}POST …/ai/voice
{
"text": "optional if audioBase64 set",
"audioBase64": "PCM s16le mono base64",
"sampleRate": 24000,
"locale": "en"
}
// Response: { text, audioBase64, done: true }
// Pipeline: STT (if audio) → generate → TTS
// Neither text nor audio → { text: "", audioBase64: "", done: true }Voice WebSocket (dashboard JWT)
Upgrade to WebSocket at /api/ai/voice/ws?token=<dashboard JWT>. Requires AI configured on the platform. Pipeline asserts Business Phone + AI Agent (same as REST voice). Duplex turns: client sends audio or text; server returns text, audio, then done.
Client → server
{ "type": "audio" | "text" | "interrupt" | "config", "data"?: "…" }Server → client
{ "type": "audio" | "text" | "done" | "error", "data"?: "…", "done"?: true }Live Chat AI
Enabling Live Chat auto-reply and inbox suggestions requires an active AI Agent entitlement only (Business Phone is not asserted on these chat paths). Session minting and chat REST still use Live Chat tokens—see /developers/live-chat.
| Method | Path | Auth / notes |
|---|---|---|
| GET | /api/tenants/{tenantId}/chat/ai-settings | Workspace admin JWT; returns settings + entitlementActive + optional usage rate |
| PUT | /api/tenants/{tenantId}/chat/ai-settings | Workspace admin JWT; enabling asserts AI Agent |
| POST | /api/tenants/{tenantId}/chat/channels/{channelId}/ai-suggestion | Agent JWT; AI Agent required; rate limit 20/min |
| POST | /api/tenants/{tenantId}/chat/channels/{channelId}/ai-suggestions/{id}/accept | Agent JWT; marks suggestion accepted |
PUT …/chat/ai-settings (shape)
{
"enabled": true,
"mode": "when_no_agent_online",
"businessProfile": "We sell …",
"handoffKeywords": ["human", "agent", "person"],
"maxRepliesPerConversation": 20
}Public AI API (token — live)
Requires an AI Agent umk_* token with ai:invoke. Runs execute synchronously today (status succeeded or failed on the create response). Connect merges workspace FAQ/terms with optional caller context — it does not crawl third-party systems.
| Method | Path | Permission | Status | Notes |
|---|---|---|---|---|
| POST | /v1/ai/runs | ai:invoke | Live | Create + execute text run |
| GET | /v1/ai/runs/{id} | ai:invoke | Live | Read run status/output |
| POST | /v1/ai/runs/{id}/cancel | ai:invoke | Planned | Cancel async run |
POST /v1/ai/runs
curl -X POST "https://umneyconnect.com/api/v1/ai/runs" \
-H "Authorization: Bearer $AI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "text",
"input": "Where is my order?",
"locale": "en",
"context": "Order ORD-123 status=shipped tracking=JD014",
"metadata": { "orderId": "ORD-123" }
}'
# Response includes id, status, outputTextErrors and limits
- 403 — missing Business Phone or AI Agent entitlement (phone speak/voice/WS), wrong product token, or missing ai:invoke
- 403 — AI Agent missing when enabling chat auto-reply or requesting inbox suggestions (Business Phone not required on those paths)
- 502 — inbox suggestion returned empty AI output
- 503 — AI binding not configured, or inference timeout after retry
- Empty voice body (REST) — returns { text: '', audioBase64: '', done: true }
- Inbox ai-suggestion: 20 requests/minute
- Public /api/v1/* rate limit: 600 requests/minute per key when those handlers exist
- Text inference budget ~30s; STT ~20s; TTS ~30s (platform timeouts)
Agent recipe
# 1. Enable AI Agent in Billing (Business Phone also required for phone speak/voice)
# 2. Developers → create token → product "AI Agent" → ai:invoke
CONNECT_API_BASE=https://umneyconnect.com/api
AI_API_KEY=umk_live_ai_…
# 3. Fetch third-party data with that system's credentials (Nest/CRM/…)
# 4. POST /v1/ai/runs with input + context
# 5. Or JWT dashboard: POST /api/tenants/{id}/ai/speak with context
# 6. Never embed the token in clients
# 7. Connect does not crawl third parties — you inject context