Skip to content

Known Limitations

ChatGPT Limitations

  • Session expiry: ChatGPT sessions expire periodically and require admin reconnection. Requests on an expired account fail fast with 503 provider_session_expired / no_healthy_provider — they never silently degrade.
  • Request latency is upstream-bound: typical chat completions take ~5–15 s; streaming time-to-first-token is ~5–10 s on a healthy account. Long prompts and project retrieval can extend this. There is no low-latency path — plan UX accordingly.
  • Parameter subset: temperature, top_p, max_tokens are accepted for SDK compatibility but ignored (ChatGPT web does not expose them). tools, tool_choice, n > 1, logprobs are rejected with 400 UNSUPPORTED_PARAMETER.
  • Image generation is implicit and quota-dependent: it is triggered by the prompt, not a parameter. A connected account can still fail at runtime when the upstream plan/quota is exhausted — GhostMind then returns 503 IMAGE_GENERATION_UNAVAILABLE (not retryable) instead of a silent text fallback. Image failures are isolated from the account circuit breaker.
  • Project context from the first message only: a conversation must be created inside a project to retrieve its files — you cannot attach a project to an existing general conversation.
  • No self-service capture: browser capture for ChatGPT is not yet available — admin session connection is the current method.
  • Rate limits: subject to ChatGPT’s own rate limiting and throttling.

Google Savio Limitations

  • TTS only: Savio supports speech generation only — no chat or transcription.
  • Browser-based execution: requires a browser worker; a single clip typically takes ~30–120 s. Always use the async job path (/v1/audio/generations), not the bounded sync endpoint.
  • WAV output: response_format is accepted but the provider may coerce to WAV. MP3/OGG transcoding is not currently offered.
  • Result expiry: generated audio is retained ~24 h by default (operator-configurable via GHOSTMIND_AUDIO_JOB_TTL_HOURS) — download and persist what you need.
  • Voice availability: 200+ voices across Arabic dialects; the catalog may change based on provider updates.

General Limitations

  • Usage tracking, not billing: /v1/usage reports request/token aggregates per API key. GhostMind does not compute provider costs or pricing.
  • No multi-turn audio: speech generation is single-request (no conversation context).
  • Transcription privacy: transcripts are not stored — if you need persistence, save them client-side.
  • UAC Actions are beta-gated: /v1/actions* requires an enabled UAC service, an authorized beta workspace, and actions:read/actions:run key scopes. It is not part of the general chat surface.

Next Steps