AI¶
AI¶
{
"id": "<string>",
"customer_id": "<string>",
"name": "<string>",
"detail": "<string>",
"type": "<string>",
"is_insight_active": <boolean>,
"engine_model": "<string>",
"parameter": "<object>",
"engine_key": "<string>",
"rag_id": "<string>",
"init_prompt": "<string>",
"current_prompt_history_id": "<string>",
"tts_type": "<string>",
"tts_voice_id": "<string>",
"stt_type": "<string>",
"stt_language": "<string>",
"vad_config": {
"confidence": <number>,
"start_secs": <number>,
"stop_secs": <number>,
"min_volume": <number>
},
"smart_turn_enabled": <boolean>,
"auto_aicall_audit_enabled": <boolean>,
"tool_names": ["<string>"],
"mcp_server_ids": ["<string>"],
"direct_hash": "<string>",
"tm_create": "<string>",
"tm_update": "<string>",
"tm_delete": "<string>"
}
id(UUID): The AI configuration’s unique identifier. Returned when creating an AI viaPOST /aisor when listing AIs viaGET /ais.customer_id(UUID): The customer that owns this AI configuration. Obtained from theidfield ofGET /customers.name(String, Required): A human-readable name for the AI configuration (e.g.,"Sales Assistant").detail(String, Optional): A description of the AI’s purpose or additional notes.type(enum string, Optional): The AI’s operating mode.normal(default) is a general-purpose AI usable in calls, tasks, and conversations.insightrestricts the AI to the Insight tool set (AllInsightToolNames) and uses a dedicated system prompt tailored for agent-facing Q&A over a contact-manager Case. See Type.is_insight_active(Boolean): Whether this is the customer’s active Insight AI (the one the Case Insight Assistant panel auto-attaches to a Case). Only meaningful whentypeisinsight; alwaysfalsefortype=normalAIs. A customer may keep any number of Insight AIs (e.g. to prepare a new prompt or model before switching to it), but at most one may be active at a time. Newly created Insight AIs are always inactive. Activate one withPOST /ais/{id}/activate_insight. When a customer has no active Insight AI, the most recently created one is used. See Insight AI activation.engine_model(String, Required): The LLM provider and model. Format:<provider>.<model>(e.g.,openai.gpt-4o,anthropic.claude-3-5-sonnet). See Engine Models.parameter(Object, Optional): Custom key-value parameter data for the AI configuration. Supports flow variable substitution at runtime. Typically left as{}.engine_key(String, Required): The API key for the LLM provider. Must be a valid key from the provider’s dashboard.rag_id(UUID, Optional): The knowledge base ID for thesearch_knowledgetool. Obtained from theidfield ofGET https://api.voipbin.net/v1.0/rags. When set, the AI assistant can search this knowledge base during voice calls. Set to00000000-0000-0000-0000-000000000000or omit to disable.init_prompt(String, Required): The system prompt that defines the AI’s behavior, persona, and instructions. No enforced length limit.current_prompt_history_id(string/UUID): UUID of the most-recentai_ai_prompt_historiesentry for this AI. Included in webhook events so callers can correlate each AI event with the exact prompt version that was active. Zero UUID (00000000-0000-0000-0000-000000000000) means no versioned prompt history has been recorded yet.tts_type(enum string, Required): Text-to-Speech provider. See TTS Types.tts_voice_id(String, Optional): Voice ID for the selected TTS provider. If omitted, the default voice for the chosen TTS type is used. See default voices in TTS Types.stt_type(enum string, Required): Speech-to-Text provider. See STT Types.stt_language(String, Optional): STT language in BCP-47 format (e.g.,ko-KR,en-US). Controls which language the Speech-to-Text engine listens for. When set, the STT provider is configured to recognize this specific language, improving accuracy for non-English calls. Empty string or omitted means auto-detect (provider default).vad_config(Object, Optional): Voice Activity Detection configuration. All fields are optional — omitted fields use Pipecat defaults. See VAD Config.smart_turn_enabled(Boolean, Optional): Enable smart turn detection using Pipecat’s LocalSmartTurnAnalyzerV3 for more natural turn-taking. Whentrue, the VADstop_secsparameter is automatically forced to0.2regardless ofvad_configsettings. Defaults tofalse. See Smart Turn.auto_aicall_audit_enabled(Boolean, Optional): Whentrue, any AICall that finishes while using this AI configuration automatically triggers an AICall audit. Defaults tofalse(opt-in).tool_names(Array of String, Optional): List of enabled tool functions. Use["all"]to enable all tools,[]to disable all tools, or list specific tool names. Fortype=insightAIs, only Insight tool names are permitted (currentlyget_contact_interactions,get_conversation_content,get_related_cases,get_case_notes,get_contact_profile,get_call_transcript,emit_info_card,notify_agent);["all"]is not valid for Insight AIs. Fortype=normalAIs, any Normal tool name or["all"]is permitted; Insight-only tool names are rejected. Mismatched combinations return400. See Tool Functions.mcp_server_ids(Array of UUID, Optional): List of customer-owned MCP (Model Context Protocol) server IDs whose tools are made available to this AI, in addition totool_names. Adding an ID requires an MCP Server that exists, is owned by the same customer, and has not been deleted; a cross-customer, nonexistent, or deleted ID is rejected with400. Tool use scope: a whitelisted server’s tools are presented to this AI’s Normal-type, single-AI sessions –type=insightAIs and team-typed AI calls never receive MCP tools regardless of this list, and realtime voice call sessions do not either. Sessions that do discover and advertise include chat,ai_taskflow actions and API-created single-AI sessions, so expecttools/list(and, on a tool call,tools/call) requests authenticated per the server’sauth_typein its logs. Of what is discovered, the tool names are stored on the AI session record and outlive the request, while the input schemas and descriptions are not stored; see MCP Server for retention. Tools discovered from a usable MCP server are namespaced asmcp_<first 8 hex chars of the server id>_<tool name>when presented to the LLM, up to 256 tools per session across this AI’s whitelisted servers. At call time a server is skipped silently, rather than failing the call, when itsstatusis notactive, when it has been deleted, or when it is no longer owned by this customer; a result the remote server itself reports as an error surfaces to the model as a failure, never a success. Deleting a server does not rewrite the AIs referencing it; see the note below. Defaults to[].direct_hash(String): Hash for direct AI access. Empty string when direct access is disabled. When enabled, this hash forms the direct SIP URI:sip:direct.<hash>@sip.voipbin.net. Regenerate viaPOST /ais/{id}/direct-hash-regenerate.tm_create(String, ISO 8601): Timestamp when the AI configuration was created.tm_update(String, ISO 8601): Timestamp when the AI configuration was last updated.tm_delete(String or null, ISO 8601): Timestamp when the AI configuration was deleted.nullwhile the AI configuration has not been deleted.
Note
AI Implementation Hint
The engine_key field contains the LLM provider’s API key. This key is write-only: it is accepted on POST /ais and PUT /ais but is never returned in GET responses for security. If you need to change the key, send a full PUT update with the new key.
Note
AI Implementation Hint
AI configurations do not use the 9999-01-01 00:00:00.000000 sentinel that some older VoIPBin resources use for “not yet occurred.” tm_update and tm_delete are null until the configuration is first updated or deleted, so test for null rather than comparing against a sentinel date.
Note
Deleting an MCP server an AI still references
Deleting an MCP server does not rewrite the AIs that reference it. The ID stays in mcp_server_ids so that saving an unrelated field on the AI keeps working, but the server contributes no tools from the moment it is deleted, and the ID is no longer offered in the admin UI’s server picker. To remove it, submit mcp_server_ids without that ID.
Example¶
{
"id": "a092c5d9-632c-48d7-b70b-499f2ca084b1",
"customer_id": "5e4a0680-804e-11ec-8477-2fea5968d85b",
"name": "Sales Assistant AI",
"detail": "AI assistant for handling sales inquiries",
"type": "normal",
"is_insight_active": false,
"engine_model": "openai.gpt-4o",
"parameter": {},
"engine_key": "sk-...",
"rag_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"init_prompt": "You are a friendly sales assistant. Help customers find the right products.",
"tts_type": "elevenlabs",
"tts_voice_id": "EXAVITQu4vr4xnSDxMaL",
"stt_type": "deepgram",
"stt_language": "en-US",
"vad_config": {
"stop_secs": 0.5
},
"smart_turn_enabled": true,
"auto_aicall_audit_enabled": false,
"tool_names": ["connect_call", "send_email", "stop_service"],
"direct_hash": "",
"tm_create": "2024-02-09 07:01:35.666687",
"tm_update": "9999-01-01 00:00:00.000000",
"tm_delete": "9999-01-01 00:00:00.000000"
}
Type¶
The type field controls the AI’s operating mode.
Type |
Description |
|---|---|
normal |
General-purpose AI. Usable in calls, tasks, and conversations, with the tool set configured via |
insight |
Restricted to the Insight tool set ( |
Note
AI Implementation Hint
Updating an AI via PUT /ais/{id} without sending type leaves the existing type unchanged — it does not reset the AI to normal. To change an AI’s type, explicitly send the desired type value in the PUT body.
Insight AI activation¶
A customer may keep multiple type=insight AI configurations (for example, to prepare a second Insight assistant with a different prompt, model, or tool set before switching to it), but at most one may be active at a time. The is_insight_active field marks which one.
The active Insight AI is the one the Case Insight Assistant panel auto-attaches when an agent opens a Case. Testing an Insight AI directly (by passing its assistance_id explicitly) works against any Insight AI regardless of its active status, so a draft configuration can be tried out without disturbing the one in production use.
Activating one
$ curl --location --request POST 'https://api.voipbin.net/v1.0/ais/a092c5d9-632c-48d7-b70b-499f2ca084b1/activate_insight?token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...'
{
"id": "a092c5d9-632c-48d7-b70b-499f2ca084b1",
"customer_id": "5e4a0680-804e-11ec-8477-2fea5968d85b",
"name": "Case Insight Assistant v2",
"type": "insight",
"is_insight_active": true,
...
}
Activating an AI automatically deactivates whichever Insight AI was active before. There is no separate deactivate call. Activating the already-active AI succeeds and changes nothing.
Behavior rules
Event |
Effect on |
|---|---|
AI created ( |
Always |
AI activated |
Set to |
AI deleted ( |
Cleared to |
AI type changed away from |
Cleared to |
No Insight AI is active |
The Case panel falls back to the most recently created Insight AI. |
Errors
Status / reason |
Cause |
|---|---|
|
The target AI is not |
|
The target AI does not exist or has been deleted. |
|
Another activation for the same customer was in flight. Retry the request. |
Engine Model¶
The engine_model field specifies which LLM provider and model to use. Format: <provider>.<model>.
Supported Providers
Provider |
Format |
Examples |
|---|---|---|
OpenAI |
|
openai.gpt-4o, openai.gpt-4o-mini |
Anthropic |
|
anthropic.claude-3-5-sonnet |
AWS Bedrock |
|
aws.claude-3-sonnet |
Azure OpenAI |
|
azure.gpt-4 |
Cerebras |
|
cerebras.llama3.1-8b |
DeepSeek |
|
deepseek.deepseek-chat |
Fireworks |
|
fireworks.llama-v3-70b |
Google Gemini |
|
gemini.gemini-1.5-pro |
Grok |
|
grok.grok-1 |
Groq |
|
groq.llama3-70b-8192 |
Mistral |
|
mistral.mistral-large |
NVIDIA NIM |
|
nvidia.llama3-70b |
Ollama |
|
ollama.llama3 |
OpenRouter |
|
openrouter.meta-llama/llama-3-70b |
Perplexity |
|
perplexity.llama-3-sonar-large |
Qwen |
|
qwen.qwen-max |
SambaNova |
|
sambanova.llama3-70b |
Together AI |
|
together.meta-llama/Llama-3-70b |
Dialogflow |
|
dialogflow.cx, dialogflow.es |
Common OpenAI Models
Model |
Description |
|---|---|
gpt-4o |
Latest GPT-4 Omni model (recommended) |
gpt-4o-mini |
Smaller, faster GPT-4 Omni variant |
gpt-4-turbo |
GPT-4 Turbo with vision capabilities |
gpt-4 |
Original GPT-4 model |
gpt-3.5-turbo |
Fast and cost-effective model |
o1 |
OpenAI o1 reasoning model |
o1-mini |
Smaller o1 reasoning model |
o3-mini |
Latest o3 mini reasoning model |
TTS Type¶
Text-to-Speech provider for converting AI responses to audio.
Type |
Description |
|---|---|
elevenlabs |
ElevenLabs high-quality voice synthesis (recommended) |
deepgram |
Deepgram Aura voices |
openai |
OpenAI TTS (alloy, echo, fable, etc.) |
aws |
AWS Polly voices |
azure |
Azure Cognitive Services TTS |
Google Cloud Text-to-Speech |
|
cartesia |
Cartesia TTS |
hume |
Hume AI emotional TTS |
playht |
PlayHT voice synthesis |
Default Voice IDs by TTS Type
TTS Type |
Default Voice ID |
|---|---|
elevenlabs |
EXAVITQu4vr4xnSDxMaL (Rachel) |
deepgram |
aura-2-thalia-en (Thalia) |
openai |
alloy |
aws |
Joanna |
azure |
en-US-JennyNeural |
en-US-Wavenet-D |
|
cartesia |
71a7ad14-091c-4e8e-a314-022ece01c121 |
STT Type¶
Speech-to-Text provider for converting incoming audio to text.
Type |
Description |
|---|---|
deepgram |
Deepgram speech recognition (recommended) |
cartesia |
Cartesia speech recognition |
elevenlabs |
ElevenLabs speech recognition |
STT Language¶
The stt_language field specifies which language the Speech-to-Text engine should recognize. The value must be in BCP-47 format (e.g., en-US, ko-KR, ja-JP).
When set, the STT provider is explicitly configured for the specified language, which improves recognition accuracy — especially for non-English conversations. When omitted or set to an empty string, the STT provider uses its default auto-detection behavior.
Common BCP-47 Language Codes
Language Code |
Language |
|---|---|
en-US |
English (United States) |
en-GB |
English (United Kingdom) |
ko-KR |
Korean |
ja-JP |
Japanese |
zh-CN |
Chinese (Simplified) |
de-DE |
German |
fr-FR |
French |
es-ES |
Spanish (Spain) |
pt-BR |
Portuguese (Brazil) |
it-IT |
Italian |
nl-NL |
Dutch |
ru-RU |
Russian |
ar-SA |
Arabic |
hi-IN |
Hindi |
pl-PL |
Polish |
Note
AI Implementation Hint
The stt_language is configured on the AI resource itself, not per-call. If you need different STT languages for different call scenarios, create separate AI configurations — one per language — and reference the appropriate ai_id in each flow action.
VAD Config¶
Voice Activity Detection configuration for tuning speech detection sensitivity and timing.
All fields are optional. Omitted fields use Pipecat’s native defaults.
Field |
Default |
Min |
Max |
Description |
|---|---|---|---|---|
confidence |
0.7 |
0.0 |
1.0 |
Minimum confidence threshold to detect voice. |
start_secs |
0.2 |
0.0 |
30.0 |
Duration in seconds of continuous speech needed to confirm speaking started. |
stop_secs |
0.2 |
0.0 |
30.0 |
Duration in seconds of silence needed to confirm speaking stopped. |
min_volume |
0.6 |
0.0 |
1.0 |
Minimum audio volume for voice detection. |
Note
AI Implementation Hint
When vad_config is null or omitted, Pipecat’s native defaults apply (confidence=0.7, start_secs=0.2, stop_secs=0.2, min_volume=0.6). To keep the AI responsive but avoid cutting off speech mid-sentence, increase stop_secs (e.g., 0.5). To make the AI more patient before responding, increase both stop_secs and start_secs.
Smart Turn¶
When smart_turn_enabled is true, the Pipecat pipeline uses LocalSmartTurnAnalyzerV3 — a local ONNX model that analyzes speech and transcription context to detect when the user has truly finished their turn, rather than pausing mid-sentence. This results in more natural conversations with fewer premature interruptions.
Note
AI Implementation Hint
Smart Turn detection requires VAD stop_secs=0.2. When smart_turn_enabled is true, any stop_secs value in vad_config is silently overridden to 0.2. This value matches the model’s training data and allows Smart Turn to dynamically adjust timing.
Tool Names¶
The tool_names field controls which tool functions the AI can invoke during conversations.
Configuration Options
Value |
Description |
|---|---|
|
Enable all available tool functions |
|
Disable all tool functions (AI can only converse) |
|
Enable only the specified tools |
Available Tools
See Tool Functions for the complete list of tools and their descriptions.
Example configurations:
// Enable all tools
"tool_names": ["all"]
// Enable only call transfer and email
"tool_names": ["connect_call", "send_email"]
// Enable conversation control tools only
"tool_names": ["stop_service", "stop_flow", "set_variables"]
// Disable all tools (conversation-only AI)
"tool_names": []