Generalize the Codex Responses-to-Chat tool media mechanism from 6c9d444c
to the remaining protocol bridges, so image-bearing tool results are never
tokenized as base64 text on any conversion path:
- Claude-to-Chat: tool_result images become native image_url parts,
batched into one synthetic user message after the tool message batch.
- Claude-to-Responses: JSON-string, MCP, and nested variants are restored
as native input_image parts inside function_call_output.
- Codex/GrokBuild-to-Anthropic: non-standard tool images are restored as
Anthropic image blocks (case-insensitive data:/http prefix parsing).
- Claude-to-Gemini: Gemini 3 uses multimodal functionResponse.parts;
older models get inlineData parts in the same user turn. The new
InlineImagesOnly scope keeps remote URLs and malformed data URLs in the
legacy text form instead of emitting fileData the API would reject.
Shared changes in tool_media:
- Centralize the plan/queue/flush helpers used by both Chat bridges.
- strip_and_clamp_media_from_tool_value clamps residual base64 inside
parsed JSON strings at any nesting depth before re-serialization.
- Emitted Chat parts no longer carry cache_control or
prompt_cache_breakpoint, preserving the strip-all-cache_control
contract that strict upstreams (GLM/Qwen) depend on.
- Media detection now requires full convertibility, keeping detection
and extraction symmetric by construction.
The media sanitizer detects and strips the new shapes symmetrically
(Chat string tool results, Anthropic string tool results, Gemini
inlineData/fileData and functionResponse.parts); the legacy typed-block
replacement runs first so replacement stays a superset of detection and
cache_control survives on replaced Anthropic blocks.
No-media tool outputs keep byte-identical legacy representations on all
bridges to protect prompt-cache prefixes.
The Responses->Chat conversion serialized image-bearing *_output items
into role:"tool" text via canonical_json_string, so view_image results
were tokenized as base64 text (~9000x inflation). Codex replays full
history every turn, so sessions hit context-limit 400s and wedged
(#4465, #5663).
- add proxy/tool_media: shared detection/strip/clamp walker for tool
output media (typed input_image / image_url / input_file /
input_audio, Anthropic source and MCP data+mimeType image shapes,
untyped data: image_url, whole-string bare data URLs)
- transform_codex_chat: replace media blocks in place with marker text
so tool content stays a plain string, and flush the extracted media
as one synthetic role:"user" message after each consecutive tool
batch; media-free traffic stays byte-identical to keep prompt-cache
prefixes stable
- media_sanitizer: detect and strip tool-output media symmetrically
(including JSON-string outputs) so reactive image stripping can heal
upstream modality rejections
- forwarder: regression tests pinning the reactive trigger and the
context-limit-400 non-trigger
E2E against Kimi K3 through the proxy: the replayed turn stays ~12k
input tokens with 99% cache hit, versus ~85k+ of base64 text per
replay before.
The cost backfill hardcoded codex|gemini as cache-inclusive apps, so
grokbuild TOTAL-semantics rows were priced on full input tokens with
cache reads double-counted. Converge the writer (proxy logger and
calculator) and the backfill recompute onto a single
sql_helpers::is_cache_inclusive_app predicate backed by the existing
CACHE_INCLUSIVE_APP_TYPES constant, and add a regression test for the
grokbuild backfill path.
Add a managed "xAI (Grok) OAuth" Codex provider that routes through the
local proxy to api.x.ai via the shared Grok CLI OAuth identity, plus the
native-Responses compatibility layer that makes Codex 0.142+ traffic work
against xAI's strict upstream serde parser.
Provider:
- codex.rs: recognize the xai_oauth placeholder in extract_auth, hard-pin
the base URL to api.x.ai and the tool profile to native Responses
- forwarder.rs: treat xAI OAuth auth failures as non-retryable
- presets + ProviderForm/CodexFormFields: managed OAuth preset that hides
the api key/endpoint fields and derives the provider type across apps
Native Responses compatibility (gated on is_xai_oauth, so no other
provider is affected):
- transform_codex_responses_namespace: flatten Codex's private
namespace/plugin tool declarations into top-level function tools on the
request; restore the flat function_call names back to {name, namespace}
on the response (streaming and non-streaming) so the client matches its
own namespaced tool registry
- transform_codex_responses_xai_sanitize: strip the OpenAI-backend-private
fields xAI rejects (external_web_access, prompt_cache_retention,
safety_identifier, the additional_tools carrier, tool_search, ...) with
deterministic removals that keep the prompt-cache prefix stable
- wire both into the native passthrough after the request transform;
response restore runs in a dedicated handler so the generic passthrough
hot path is untouched
Ports the proven approach of sub2api's Grok Responses gateway. Verified
with a 4-round codex -> xAI OAuth workload: all tasks green, zero upstream
errors.
Derive request ids from upstream envelope ids (Codex/OpenAI top-level
id, Gemini responseId, Claude message id with non-empty filtering)
scoped as session:{app_type}:{provider_id}:{id} for non-Claude sources;
Claude keeps bare session:{id} to preserve session-log convergence.
The logger now queries and conditionally writes under a single
connection guard: identical semantic replays return without writing or
notifying, session_log rows may be upgraded by proxy, and same-id
different-semantic responses land on a deterministic SHA-256 collision
fallback key instead of being overwritten. Fixes the random-UUID
INSERT OR REPLACE duplication behind #5496.
Add an xAI OAuth manager using the OAuth 2.0 Device Authorization Grant
with endpoints resolved from xAI's OIDC discovery document. All HTTP goes
through the app-managed proxy client.
- Managed provider kind xai_oauth: forced openai_responses wire format,
pinned api.x.ai base URL, bearer injection gated to the xAI origin,
tokens registered for log redaction, single-auth-key takeover policy.
- Token cache cannot bypass account state: cache hits re-validate account
usability, refresh commits run under the mutation lock with a
refresh-token CAS check, and pending logins are re-checked before an
account is persisted.
- Refresh classification: 401/403 with any body and 400 with a non-JSON
body mark the account for re-auth; 429/5xx stay transient.
- Shared auth_* commands dispatch to xAI with guard types mirroring the
Copilot/Codex branches.
Backend half of the logging overhaul.
Retention:
- Keep the last 4 rotated files at 20MB each and stop deleting logs on
startup, so a crash's prior-run logs survive a restart.
- Apply the persisted log level right after plugin registration instead
of at the end of setup.
- Bound crash.log with size-based rotation (5MB x 2).
Secret redaction on every log path:
- Proxy: strip userinfo/query from upstream URLs (keep path for
diagnostics), exact-match redact known auth.api_key/access_token,
classify request/response bodies instead of logging them, header
allowlist.
- Redact the Gemini `?key=` in cache-trace endpoints.
- Omit MCP custom-field values (headers may carry tokens).
- Redact deeplink and model-fetch URLs before logging.
Docs:
- FAQ points users to the persistent crash.log for support workflows.
Normalize function tool parameters on the Codex Responses -> Chat Completions bridge: default null/missing parameters (direct and nested function forms) to {"type":"object","properties":{}}, coerce explicit type:null, and add a root type for top-level oneOf schemas so strict OpenAI-compatible upstreams (DeepSeek etc.) no longer reject built-in Codex tools like codex_app__automation_update. Also hardens the Codex -> Anthropic tool path with the same object-type guarantee and adds regression coverage for all reported shapes.
Keep non-empty call IDs across continuation deltas and release parallel tool calls in Chat index order when identity fields arrive late. Preserve valid sparse and later calls during finalization.
* fix: normalize function parameters type to "object" for strict OpenAI-compatible providers
Some Responses tools carry parameters with `type: null` (e.g.
codex_app__automation_update), causing HTTP 400 from strict
OpenAI-compatible providers like DeepSeek that require
`{"type": "object", "properties": {...}}`.
This adds normalize_function_parameters() to ensure the type
field is always "object" in both branches of
responses_function_tool_to_chat_tool.
Closes#4705
* style: fix cargo fmt issues
* fix: handle null function parameters
The forced 1-hour cache_control TTL (schema v14) was a mistake. Injected
breakpoints return to Anthropic's standard 5-minute TTL, caller-owned
markers are preserved verbatim instead of having their TTLs rewritten,
and the 5m/1h cache-write buckets are removed from usage parsing and
pricing (back to the single aggregate cache-creation rate). The cache
TTL selector is removed from the rectifier settings panel along with its
i18n keys in all four locales.
SCHEMA_VERSION returns to 13: the unreleased cache_creation_1h_tokens
column and the v13->v14 migration are removed. The feature never shipped
in a release and the introducing commit was never pushed, so no
databases were stamped v14 outside this machine (local DB verified at
user_version 13).
Mapped GPT models were rejected by Codex clients with "model does not
support image inputs". Two root causes:
- Catalog entries for native-Responses/Anthropic providers cloned a
template whose input_modalities defaulted to ["text"], so every mapped
model was advertised text-only. model_catalog_json replaces Codex's
built-in model table wholesale, and both the TUI and the IDE extension
block images pre-send when the current model is found without "image".
- Editing the current Codex provider during proxy takeover only refreshed
the DB backup, so removing the mapping left a stale model_catalog_json
pointer (and its text-only catalog file) active in live config.
Changes:
- New shared model_capabilities module: explicit row declaration first,
then a confirmed text-only registry (exact tail match only — prefix
matching removed, variants enumerated so future -vl/-vision models fail
open), everything else unknown.
- Catalog generation writes input_modalities from that inference for all
tool profiles: unknown models fail open to ["text","image"]; only
confirmed text-only models are advertised as ["text"], giving users a
clear client-side prompt instead of silent image stripping.
- Live catalog reverse-import collapses modalities that match current
inference, so registry corrections are not frozen into hidden row
overrides and the rectifier's heuristic opt-out keeps working.
- Saving the current Codex provider while takeover owns live now
re-projects the live config (mirrors the hot-switch path), so mapping
edits and removals take effect immediately.
- Media rectifier delegates to the shared module; its preflight toggle is
documented (4 locales) as proxy-request-only, never affecting catalog
capability declarations.
ChatGPT's Codex backend routes model availability by the originator and version header pair. Requests identifying as cc-switch without a version were assigned to a cohort where gpt-5.6-luna resolved to an unavailable internal engine, resulting in a misleading 404 Model not found response.
Identify takeover requests as codex_cli_rs and send version 0.144.1, satisfying luna's minimal_client_version requirement of 0.144.0. A direct HTTP A/B test confirmed the existing headers returned 404 while the aligned identity completed successfully, so no WebSocket transport workaround is required.
Allow the built-in Codex official provider to participate in takeover mode while preserving Codex's native OAuth or API-key credentials instead of persisting them into provider records.
Project official routing into a dedicated TOML provider, normalize inline tables, clean stale managed placeholders, and fail closed when the live configuration cannot be transformed safely.
Validate forwarded authorization, make official 401/403 responses non-retryable, avoid circuit-breaker pollution, and share the first-party ChatGPT endpoint across the Codex and Claude adapters.
Parse and retain Anthropic's ephemeral 5-minute and 1-hour cache-creation token buckets while preserving the existing aggregate cache-write metric for compatibility.
Price 1-hour writes at the documented premium relative to the configured 5-minute write rate, clamp inconsistent provider details safely, and include TTL buckets in usage diagnostics.
Persist 1-hour cache-write tokens with schema version 14 so zero-cost backfills and later pricing updates retain the original TTL semantics. Keep session import paths compatible through zero-valued detail fields.
Honor both the optimizer master switch and the cache-injection sub-switch before mutating native optimizer requests.
Use the available four-breakpoint budget across tools, system content, the latest cacheable message, and an older user anchor for long tool-heavy conversations.
Preserve caller-owned breakpoint limits, avoid thinking blocks as cache targets, normalize configured TTLs, and add regression coverage for disabled optimization and long histories.
Fail closed on HTTP 2xx failure envelopes and pre-output SSE failures so semantic upstream errors can trigger failover instead of becoming empty successful replies.
Finalize incomplete and truncated streams explicitly, handle clean EOF and whole JSON responses, and keep tool-call stop reasons and terminal event ordering consistent.
Preserve structured tool results, URL images, documents, system roles, and signed thinking across both conversion directions. Drop incomplete historical tool calls safely and classify malformed completed arguments as non-retryable client requests.
Keep Codex-to-Anthropic prompt caching enabled by default while honoring the dedicated cache-injection switch.
Codex always sends prompt_cache_key in its Responses requests, but the
Responses -> Chat Completions conversion dropped it, breaking session
cache affinity on upstreams that route by key (e.g. Kimi Coding).
- Re-inject prompt_cache_key after conversion in the forwarder: an
explicit client key wins, otherwise a client-provided session ID;
generated per-request UUIDs are never sent upstream.
- Provider-aware gating: "auto" enables only known-compatible upstreams
(api.openai.com, api.kimi.com/coding) because strict gateways reject
unknown fields with HTTP 400 (e.g. Fireworks); an advanced
Auto/Enabled/Disabled override is available on the Codex form in all
four locales.
- Kimi For Coding preset opts in explicitly.
Warn when caller-provided cache breakpoints already exceed the supported total of four while preserving the original markers and upgrading their TTLs.
Clarify that automatic injection is governed by the remaining breakpoint budget, and add regression coverage proving excess caller markers are never deleted or reordered.
Round-trip encrypted Responses reasoning items through bridge-owned Anthropic thinking blocks, discard orphaned reasoning-only history, and consume the official reasoning text event vocabulary.
Track concurrent streaming items by stable IDs and output indexes, reuse a dedicated fallback block for keyless legacy reasoning, recover tool arguments from done events, and close blocks in protocol order.
Normalize empty or incomplete non-streaming tool arguments, reject malformed completed calls, and persist upstream usage before returning terminal conversion errors.
Parse cache_write_tokens from OpenAI usage details and preserve cache creation data across Chat, Responses, and Anthropic conversion paths.
Add explicit input-token semantics to request logs and rollups so legacy rows subtract cache reads only while new total-inclusive rows subtract both cache reads and writes. Migrate v12 databases, normalize rollups to fresh input, and cover historical backfill behavior with regression tests.
* feat(codex): support native Anthropic Messages protocol as upstream
Allow gateways that only expose the native Anthropic Messages protocol
(/v1/messages) to be used by Codex: the local proxy performs bidirectional
request/response/streaming conversion between Responses and Anthropic.
Backend:
- Add two conversion modules: transform_codex_anthropic / streaming_codex_anthropic
- codex.rs: add routing detection and auth: ANTHROPIC_AUTH_TOKEN→Bearer (default),
ANTHROPIC_API_KEY→x-api-key, mutually exclusive
- handlers.rs: add handle_codex_anthropic_to_responses_transform
- forwarder.rs: support optional Claude Code client fingerprint impersonation
(User-Agent / anthropic-beta / x-app / system prompt first-line injection)
and /responses→/v1/messages rewriting
- codex_config: the anthropic format reuses the NativeResponses profile to strip
custom tools
- ProviderMeta: add impersonateClaudeCode
Frontend:
- CodexApiFormat: add "anthropic"; the form adds auth field selection and an
impersonation toggle
- Add en/ja/zh/zh-TW copy
Robustness:
- Downgrade when tool history / forced tool_choice conflicts with extended
thinking, avoiding upstream 400s
- Emit cache_creation_input_tokens in usage and use saturating_add to guard
against overflow
- Append a unique suffix to non-streaming/streaming output-item ids to avoid
multi-segment text/thinking overwriting each other
* fix(codex): harden Anthropic bridge against empty text blocks and truncated streams
- Drop empty/whitespace-only text content blocks when rebuilding Anthropic
messages from Responses history; Anthropic 400s on empty text blocks (e.g. an
empty assistant text emitted alongside a tool_use), which broke follow-up and
tool-result requests. Also drop messages left without content.
- Do not report a truncated Anthropic stream as completed: when the SSE
connection ends before message_stop with no stop_reason, emit an incomplete
response if partial output exists, or a failed (stream_truncated) response
otherwise, mirroring the chat converter's EOF handling.
* fix(codex): resolve Anthropic-bridge review findings
Blocking:
1. Defer stripping the [1m] long-context marker until after catalog matching and model write-back, re-stripping it on the final Codex→Anthropic body and setting a flag to emit the context-1m-2025-08-07 beta header, so the marker is no longer lost or overridden by the default model.
2. Gate Anthropic thinking on the trailing turn only (via trailing_turn_allows_thinking) instead of scanning full history, so a Codex session that resends history each round no longer permanently loses thinking after the first tool call.
3. Inject 5m ephemeral cache_control on the Codex→Anthropic body by reusing cache_injector (handling the system string→array conversion), so system/tools/history are cached instead of re-sent at full price every round.
4. Add a shared base_url_is_full_endpoint helper (normalizing whitespace/query/fragment/trailing slash) used by both the Anthropic and Chat paths, so a base URL already ending in /v1/messages is treated as a full endpoint instead of double-appending to /v1/messages/v1/messages.
5. Align the catalog tool-profile predicate with the routing predicate so resolve_codex_catalog_tool_profile returns the Anthropic profile whenever the request converts to Anthropic, preventing freeform tools like apply_patch from being silently filtered by a ProxyChat catalog.
6. Explicitly disable native web_search for the Anthropic profile (including the no-catalog branch) via set_codex_native_web_search_field, so Codex no longer treats it as available while the transform silently drops it.
7. Only forward tool_choice when tools survive filtering, dropping it otherwise, to avoid a non-retryable 400 from upstream when tool_choice is sent with no tools.
8. Lower the fallback default to max_tokens=8192 (only when max_output_tokens is omitted) and clamp the thinking budget to max_tokens/2, disabling thinking below the 1024 floor, to avoid hard 400s on low-output-ceiling models/gateways.
Minor:
9. Centralize the Codex/OpenAI fingerprint-header denylist in is_codex_client_fingerprint_header so impersonating Claude Code uniformly drops originator/session_id/conversation_id/chatgpt-account-id/x-client-request-id/openai-* and the x-stainless-*/x-codex-* prefixes.
10. Retain content_block_start.input as start_input and fall back to it at block close when no input_json_delta arrived, so a gateway that carries the full tool input on the start event no longer yields empty tool arguments.
11. Extract a shared codex_responses_sse module as the single Responses SSE envelope builder that both the chat and anthropic streaming emitters delegate to, with byte-for-byte-unchanged wire output, so future event-format fixes touch one place instead of two.
* fix(codex): add per-provider max_output_tokens override for Codex→Anthropic path
Codex does not forward model_max_output_tokens in the request body,
causing the proxy to fall back to a conservative 8192 default. This
truncates long or thinking-heavy responses (stop_reason=max_tokens).
- Add maxOutputTokens field to ProviderMeta (Rust + TypeScript)
- Inject the value into the Anthropic request body before transform,
taking precedence over request-supplied and default values
- Add numeric input in Codex form fields (Anthropic format only)
- Add i18n entries for label, placeholder, and hint (en/zh/zh-TW/ja)
- Include roundtrip and omission unit tests for the new field
* fix(codex): harden Anthropic protocol bridge
* fix(codex): address Anthropic bridge review
* fix(codex): preserve flattened Anthropic inputs
* fix(codex): harden Anthropic recovery paths
* fix(ci): satisfy Rust clippy
---------
Co-authored-by: Jason <farion1231@gmail.com>
Fixes#5025. The rectifier's media fallback missed Volcano Coding Plan's
GLM 5.2 on both paths:
- Preventive: known_text_only_model had glm-5.1 but not glm-5.2, so image
blocks were forwarded verbatim. Added glm-5.2 as an exact tail match
(not a prefix, to avoid stripping images from a future glm-5.2v
multimodal variant following Zhipu's 4v/5v naming).
- Reactive: the upstream error "Model only support text input" never
mentions image/media, so the mentions_image gate rejected it before
hints ran; and the existing "only supports text" hint misses the
gateway's missing third-person "s". Added a self-evident phrase list
("only support text" / "only supports text") that asserts a modality
rejection on its own and bypasses the image-mention gate.
Includes regression tests using the verbatim #5025 error body and the
glm-5.2[1M] mapped-model form, plus a glm-5.2v negative assertion.
* feat: add Claude subagent takeover config
* feat: add Claude subagent model field
* i18n: add Claude subagent model labels
* fix(proxy): preserve configured subagent model mapping
* fix(providers): exclude subagent model from Claude common config
* style: format rust code
* Update Longcat presets to LongCat-2.0
* fix(proxy): classify LongCat-2.0 as text-only for media sanitizer
The Longcat presets now use LongCat-2.0, but the known_text_only_model
allowlist still only matched the retired longcat-flash-chat tail. Without
this, images pasted into a text-only LongCat-2.0 session are forwarded
upstream instead of being replaced with the unsupported-image marker,
causing a hard rejection. Add longcat-2.0 (keeping the retired name for
saved configs) and a regression test.
---------
Co-authored-by: chengzifeng <chengzifeng@meituan.com>
Co-authored-by: Jason <farion1231@gmail.com>
* fix(proxy): decompress Codex request body before forward, support zstd
Codex Desktop sends zstd-compressed request bodies when authenticated
against the Codex backend, which broke local proxy routing because the
handlers parsed the raw bytes with serde_json directly.
Reworked on top of current main so it preserves the response_processor
behavior that landed after this PR was first opened:
- Extract content-encoding helpers into a shared proxy::content_encoding
module. decompress_body keeps returning Option<Vec<u8>> so unknown
encodings stay pass-through with their content-encoding header intact,
and keeps the deflate zlib-then-raw fallback (RFC 9110).
- Add zstd/zst support (zstd 0.13) and disable reqwest's auto zstd
decompression via .no_zstd() for parity with gzip/br/deflate.
- Decompress the request body before JSON parsing in the three Codex
handlers (chat_completions / responses / responses_compact) and strip
the stale content-encoding / content-length / transfer-encoding headers
so the forwarder regenerates them.
- Support stacked codings (e.g. "gzip, zstd") by decoding in reverse
order and merge repeated Content-Encoding headers via get_all.
Fixes#3764Fixes#3696
Co-authored-by: chenx-dust <16610294+chenx-dust@users.noreply.github.com>
* fix(proxy): decompress upstream error bodies before reading them
The forwarder error branch consumes non-2xx responses via String::from_utf8
directly, bypassing read_decoded_body. reqwest has no auto-decompression
feature enabled, so a compressed error body (gzip/br/deflate/zstd) arrives
as raw bytes, fails from_utf8, and gets dropped, hiding upstream rate-limit
and auth details from the client.
Decode the error body with the shared proxy::content_encoding helper,
mirroring the success path. Falls back to the raw bytes when the encoding is
unsupported or decoding fails.
Co-authored-by: chenx-dust <16610294+chenx-dust@users.noreply.github.com>
---------
Co-authored-by: Jason <farion1231@gmail.com>
Co-authored-by: chenx-dust <16610294+chenx-dust@users.noreply.github.com>
CI cargo fmt --check failed on transform_codex_chat.rs (the wrapping
introduced upstream did not match rustfmt 1.95.0). Apply cargo fmt so the
chat_legacy_function_call_to_response_item call uses the expected layout.
* Chat API: skip tool calls with missing function names
Some providers send empty or absent function names in streaming
tool call deltas. Previously these produced invalid output items.
- Don't overwrite accumulated state.name with empty deltas
- Skip tool calls that never received a valid name (instead of
falling back to 'unknown_tool')
- Apply the same defensive guard in finalize_tools and the
non-streaming path
* Address review: defer empty-name skip to finalization
Require both call_id and name before triggering should_add,
instead of skipping eagerly when name is absent in the first
delta. This handles providers that send id before name,
as suggested in the Codex review.
* Guard legacy function_call against empty name
Return Option<Value> from chat_legacy_function_call_to_response_item,
returning None when function_call.name is missing or empty. This covers
the legacy message.function_call path that the original guard missed.
* Remove unreachable unknown_tool fallback
---------
Co-authored-by: Jason <farion1231@gmail.com>
* fix(codex): restore cached tool call fields
* refactor(codex): merge duplicate enrich loops in chat history
enrich_call_item_from_cache copied the fill-if-empty loop for
reasoning_content/reasoning. The two loops are identical and key
order is irrelevant, so fold both key sets into a single loop.
Pure refactor, no behavior change; codex_chat_history tests pass.
---------
Co-authored-by: Jason <farion1231@gmail.com>
DeepSeek's Anthropic-compatible endpoint rejects requests where
thinking.type=disabled coexists with effort parameters, returning
HTTP 400. This breaks Claude Code 2.1.166+ sub-agents (Workflow/Dynamic
Workflow), which hardcode thinking:disabled.
Rather than overriding thinking:disabled, remove the conflicting effort
parameters (output_config.effort / reasoning_effort) to respect Claude
Code's intent — sub-agents don't need to display reasoning.
Fixes: https://github.com/deepseek-ai/DeepSeek-V3/issues/1397
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The provider panel health check sent a real streaming model request, which many third-party providers block (401/403/WAF), causing false negatives while only stable official endpoints passed. Replace it with a lightweight reachability probe: GET the provider base_url and treat any HTTP response (200/4xx/5xx) as reachable; only DNS/connect/TLS/timeout count as failure. Latency is the probe's TTFB.
Backend (services/stream_check.rs): rewrite ~2200 -> ~350 lines, dropping real-request building, format conversion, auth and API-path resolution while keeping per-app base_url extraction. Defaults: 8s timeout, 1 retry, 1500ms degraded threshold.
Failover invariant: the reachability check must never reset the circuit breaker (reachable != usable; a 403 host is reachable but broken for real traffic). Remove the resetCircuitBreaker call from useStreamCheck; failover failure detection stays driven solely by real proxy traffic (forwarder/circuit_breaker untouched). useResetCircuitBreaker is kept dormant for a future manual-recovery entry.
Open the check to all providers: drop the official/copilot/codex-oauth/third-party gating and the 'sends a real request' confirm dialog. For official providers whose base_url is intentionally empty, fall back to the endpoint the client actually uses (Claude -> api.anthropic.com, Codex -> chatgpt.com/backend-api/codex, Gemini -> generativelanguage). Non-official providers with a missing base_url still error to avoid a false green light. Claude Desktop Official is native 1P mode (talks to claude.ai, cc-switch not in the request path, no reliable endpoint) so its button stays hidden.
Slim StreamCheckConfig and per-provider testConfig to timeout/threshold/retries (drop test model + prompt); sync zh/en/ja/zh-TW. Retain the now-unused anthropic_to_openai/anthropic_to_gemini transform utilities and their test suites.
- Wire claude-fable-5 as a fourth tier on both proxy paths, with a
fable -> opus -> default fallback mirroring the official downgrade.
- Whitelist the fable- prefix for the Desktop 1.12603.1+ validator.
- Clarify fallbackModelHint (zh/en/ja/zh-TW): a blank tier on
third-party endpoints forwards the literal model name and 404s.
Refs #3980, #4026, #4049.
The Claude/Codex format-transform non-stream branch returned an opaque 422
"Failed to parse upstream response" whenever a 2xx upstream body was not
valid JSON. The common case: MaaS gateways force-stream a stream:false
request and return an SSE body with a non-SSE Content-Type, defeating the
header-only is_sse() check.
On serde failure, sniff for SSE and aggregate the chunks into a single
JSON, then run the existing converter so clients still receive a valid
non-stream response.
- chat_sse_to_response_value: aggregate chat.completion.chunk SSE
(content / reasoning / refusal / tool_calls / legacy function_call),
tool_calls index-keyed via BTreeMap to avoid unbounded densification,
first-wins finish_reason, message-snapshot override, completeness and
error-event guards; synthesize an id when the upstream omits one
- responses_sse_to_response_value: process the residual trailing block,
tolerating truncation and skipping it once a completed event was seen
- enrich remaining parse failures with content-type / content-encoding /
body-snippet diagnostics
- deflate: try zlib (RFC 9110) before raw; keep the content-encoding
header for unsupported encodings
- gate zero-usage rows on the Claude transform path
Extract a shared `parse_custom_user_agent` helper in provider.rs returning
`Result<Option<HeaderValue>>`, and reuse it in the forwarder, stream check,
and model fetch paths so detection, forwarding, and model listing all apply
the same provider-level User-Agent. Previously only the forwarder honored it,
so stream check could fail (or model listing 403) on UA-gated upstreams that
the proxy itself handled fine.
- stream_check injects the provider's custom UA on the claude/codex paths and
still skips the GitHub Copilot fingerprint UA.
- model_fetch service + command and the model-fetch.ts wrapper thread an
optional UA through to GET /v1/models.
- runtime callers silently ignore invalid values via `.ok().flatten()`
(no save-time block, so deeplink imports stay lenient).
The model mapped for takeover (env mapping, Claude Desktop routes,
Copilot normalization, Codex chat override) was discarded inside the
forwarder, so usage attribution depended entirely on the upstream
echoing it back. When the upstream omitted the model or mirrored the
client alias, kimi/glm tokens were recorded and priced as claude-*
(roughly 5-25x overstatement).
- capture the final outbound model in forward(), return it via
ForwardResult, and store it on the request context
- attribution fallback order is now: upstream echo (empty string
treated as missing) -> outbound model -> client-requested model
- 'request' pricing mode anchors to the outbound model instead of the
pre-mapping client alias; unchanged when no mapping applies
- persist the resolved pricing_model on every usage row
- Claude Desktop rows now log app_type "claude-desktop" on streaming
and transform paths too (was hardcoded "claude", silently dropping
desktop provider pricing overrides and splitting the cost basis by
the stream flag); its global pricing defaults inherit the claude
config since proxy_config only allows claude/codex/gemini rows
Codex /responses requests routed to text-only OpenAI-chat upstreams
(e.g. DeepSeek deepseek-v4-flash) failed with HTTP 400 "unknown variant
image_url" when images were sent: the responses->chat conversion turns
input_image items into image_url blocks the model rejects. The media
rectifier previously covered only the Claude adapter, so neither the
proactive strip nor the reactive retry fired for Codex.
- media_retry_should_trigger: accept "Codex" adapter, not just "Claude"
- contains_image_blocks / replace_images: also scan responses `input`
(input_image) in addition to chat `messages`
- is_image_block_type: match image | image_url | input_image
- is_unsupported_image_error: add "unknown variant" hint for the
deserialize error
- forward(): proactively run apply_media_prevention for Codex after the
responses->chat conversion
Proactively strips images for known text-only models (heuristic on by
default) and reactively retries with images replaced on upstream
image-unsupported errors. Adds tests for chat image_url, codex
input_image, the reactive trigger, and the deserialize error match.
Builds on #2774 (which fixed cache_read for the streaming openai_chat path).
Two gaps remained, both double-counting cache tokens when a Claude client
meters as app_type="claude" (input_includes_cache_read=false):
1. cache_read was still added to input on the non-streaming openai_chat path
(transform.rs openai_to_anthropic) and the whole openai_responses family
(transform_responses.rs build_anthropic_usage_from_responses, covering the
non-streaming call site and both streaming_responses call sites).
2. cache_creation was never subtracted on any converted path, including the
streaming openai_chat path #2774 had already touched. Claude billing treats
cache_creation as a separate bucket, so an inclusive upstream carrying a
direct cache_creation_input_tokens field billed it twice.
All four metering points now compute:
input = prompt_tokens - cache_read - cache_creation
restoring the invariant input + cache_read + cache_creation == prompt_tokens.
Pure OpenAI upstreams are unaffected (no cache_creation concept/field).
Tests: update direct-cache assertions (40->20), add a streaming conservation
regression test, and pin prompt<cache underflow (saturating clamp to 0) for all
three metering functions. cargo test 1573 pass, clippy clean.
Note: fix is forward-only; historical rows are not recomputed (cost is frozen at
log time and app_type="claude" mixes native + converted rows).
Audited all proxy format-conversion paths (Chat<->Message, Chat<->Response,
Gemini<->Message) for usage/cache metering. Five issues found and fixed.
The dedup mechanism (request_id PK, proxy/session source isolation) is
untouched, so no double-counting is introduced.
- A (Claude + openai_chat, streaming): inject stream_options.include_usage
so OpenAI-compatible upstreams emit usage in the SSE tail. Without it the
converted Anthropic message_delta was all-zero and the whole request's
input/output/cache was dropped. Same root cause as the already-fixed
Codex Chat path; the injection is extracted into a shared helper
(transform::inject_openai_stream_include_usage) reused by both paths.
- C (Claude + gemini_native): subtract cachedContentTokenCount from
input_tokens in build_anthropic_usage so input becomes fresh input
(Anthropic semantics). Previously the cache-hit tokens were billed twice
because this path meters as app_type="claude" (input_includes_cache_read
= false) while Gemini's promptTokenCount includes the cache.
- D (Codex + openai_chat, streaming): gate log_usage on
has_billable_tokens() to skip the synthetic all-zero usage the converter
emits when a non-compliant upstream omits usage, preventing empty-row
request-count inflation.
- P2 (from_claude_stream_events): use has_billable_tokens() for the return
gate instead of input>0||output>0, so a fully-cached streamed request
(cache_read>0, input==output==0) is still recorded. Affects all
Claude-streaming paths, not just Gemini.
- P3 (Codex Chat->Responses, non-streaming): apply the same
has_billable_tokens() filter the streaming branch got, since the
synthesized all-zero usage makes from_codex_response return Some and
bypass the `if let Some` guard.
Add TokenUsage::has_billable_tokens() as the unified predicate. New tests
cover include_usage injection, gemini input subtraction, the gate itself,
cache-only stream recording, and synthetic all-zero codex usage.
Full lib suite: 1569 passed.
* fix(proxy): strip cache_control from OpenAI format conversion (#3805)
- Remove cache_control passthrough from system messages, text blocks,
and tools to prevent 400 errors on strict OpenAI-compatible endpoints
- Always simplify single text block content to plain string format
- Fixes two format conversion bugs reported in issue #3805
* fix(proxy): apply cargo fmt to fix CI formatting check
Convert Responses input_file (requiring file_id or file_data, never file_url which Chat file parts do not support) and input_audio parts into their Chat Completions equivalents, and handle top-level input_* items that previously fell through and were dropped, clearing stale pending reasoning for non-assistant messages.
Replace the unconditional finalize at chat-to-responses stream end with a three-way guard: complete normally when finish_reason or [DONE] arrived, emit an incomplete response when substantive output exists without a finish_reason, and emit a failed (stream_truncated) event for empty truncation instead of masking it as completed. Also propagate late-arriving reasoning_content onto still-active tool-call items.
Generalize the cross-turn reasoning cache in codex chat history from function_call only to the full tool-call triad (function_call, custom_tool_call, tool_search_call) and their *_output counterparts, so apply_patch and tool-search calls keep their reasoning_content when restored via previous_response_id.