mirror of
https://github.com/farion1231/cc-switch.git
synced 2026-07-26 14:35:22 +08:00
5376ea042b
* i18n: update cache terminology across all languages
- Change 'Cache Read' to 'Cache Hit' in all languages
- Change 'Cache Write' to 'Cache Creation' in all languages
- Update zh: 缓存读取 → 缓存命中, 缓存写入 → 缓存创建
- Update en: Cache Read → Cache Hit, Cache Write → Cache Creation
- Update ja: キャッシュ読取 → キャッシュヒット, キャッシュ書込 → キャッシュ作成
Affected keys: cacheReadTokens, cacheCreationTokens, cacheReadCost,
cacheWriteCost, cacheRead, cacheWrite
* feat(usage): add cache metrics to trend chart
- Add cache creation tokens visualization (orange line)
- Add cache hit tokens visualization (purple line)
- Add gradient definitions for new cache metrics
- Include cache data in hourly aggregation
- Display cache metrics alongside input/output tokens
This provides better visibility into cache usage patterns over time.
* fix(usage): fix timezone handling in datetime picker
- Add timestampToLocalDatetime() to convert Unix timestamp to local datetime
- Add localDatetimeToTimestamp() with validation for incomplete input
- Fix issue where typing hours/minutes would jump to previous day
- Validate datetime format completeness before conversion
- Use local timezone instead of UTC for datetime-local input
This resolves the issue where users couldn't fine-tune time selection
and the input would jump unexpectedly when editing hours or minutes.
* feat(usage): add auto-refresh for usage statistics
- Add 30-second auto-refresh interval for all usage queries
- Disable background refresh to save resources
- Apply to: summary, trends, provider stats, model stats, request logs
- Queries automatically update when tab is active
- Pause refresh when user switches to another tab
This keeps usage data fresh without manual refresh.
* fix(proxy): improve usage logging and cache token parsing
- Log requests even when usage parsing fails (with default values)
- Add detailed debug logging for usage metrics
- Support cache_read_input_tokens field in Codex responses
- Fallback to input_tokens_details.cached_tokens if needed
- Add test case for cached_tokens in input_tokens_details
- Ensure all requests are tracked in database for analytics
This fixes missing request logs when API responses lack usage data
and improves cache token detection across different response formats.
* style(rust): use inline format args in format! macros
- Replace format!("...", var) with format!("...{var}")
- Update universal provider ID formatting
- Update error message formatting
- Update config.toml generation in Codex provider
Fixes clippy::uninlined_format_args warnings.
* feat(proxy): enhance provider router logging
- Add debug logs for failover queue provider count
- Log circuit breaker state for each provider check
- Add logs for missing current provider scenarios
- Log when no current provider is configured
- Use inline format args for better readability
This improves debugging of provider selection and failover behavior.
* feat(database): update model pricing data
- Update Claude models to full version format (e.g. claude-opus-4-5-20251101)
- Add GPT-5.2 series model pricing (10 models)
- Add GPT-5.1 series model pricing (10 models)
- Add GPT-5 series model pricing (12 models)
- Add Gemini 3 series model pricing (2 models)
- Update Gemini 2.5 series model ID format (use dot separator)
- Unify display names by removing thinking level suffixes
* fix(usage): correct Gemini output token calculation
Fix Gemini API output token parsing to use totalTokenCount - promptTokenCount
instead of candidatesTokenCount alone. This ensures thoughtsTokenCount is
included in output statistics.
- Update from_gemini_response to calculate output from total - input
- Update from_gemini_stream_chunks with same logic for consistency
- Fix from_codex_stream_events to use adjusted token calculation
- Add test case for responses with thoughtsTokenCount
- Update existing tests to match new calculation logic
* fix(usage): correct cache token billing and add Codex format auto-detection
- Avoid double-billing cache tokens by subtracting from input before calculation
- Add smart Codex parser that auto-detects OpenAI vs Codex API format
- Extract model name from Codex responses for accurate tracking
* fix(proxy): improve takeover detection with live config check
- Add live config takeover detection for hot-switch decision
- Rebuild takeover when backup is missing or placeholder remains
- Make detect_takeover_in_live_config_for_app public
- Fix is_takeover_active to use actual takeover status
* refactor(usage): simplify model pricing lookup by removing suffix fallback
Replace complex suffix-stripping fallback with direct prefix/suffix cleanup.
Model IDs are now cleaned by removing vendor prefix (before /) and colon
suffix (after :), then matched exactly against pricing table.
* feat(database): add Chinese AI model pricing data
Add pricing for domestic AI models (CNY/1M tokens):
- Doubao-Seed-Code (ByteDance)
- DeepSeek V3/V3.1/V3.2
- Kimi K2/K2-Thinking/K2-Turbo (Moonshot)
- MiniMax M2/M2.1/M2.1-Lightning
- GLM-4.6/4.7 (Zhipu)
- Mimo V2 Flash (Xiaomi)
Also fix test case to use correct model ID and remove invalid currency column.
* refactor(proxy): improve header forwarding with blacklist approach
Change from whitelist to blacklist mode for request header forwarding.
Only skip headers that will be overridden (auth, host, content-length).
This preserves client's original headers and improves compatibility.
* fix(proxy): bypass timeout and retry configs when failover is disabled
When auto_failover_enabled is false, timeout and retry configurations
should not affect normal request flow. This change ensures:
- create_forwarder: passes 0 for all timeout/retry params when failover
is disabled, effectively bypassing these checks
- streaming_timeout_config: returns 0 for both first_byte_timeout and
idle_timeout when failover is disabled
This prevents unnecessary timeout errors and retry attempts when users
have explicitly disabled the failover feature.
* fix(proxy): handle zero value input in failover config fields
* refactor(proxy): remove retry logic and add enabled check for failover
* refactor(proxy): distinguish circuit-open from no-provider errors
* Align usage stats to sliding windows
* feat(proxy): add body and header filtering for upstream requests
* feat(proxy): enable transparent passthrough for headers
- Passthrough anthropic-beta header as-is from client
- Passthrough anthropic-version header from client
- Passthrough client IP headers (x-forwarded-for, x-real-ip) by default
- Filter private params (underscore-prefixed fields) from request body
- No database changes required
* feat(proxy): extract session ID from client requests for logging
- Add SessionIdExtractor to parse session ID from Claude/Codex requests
- Support extraction from metadata.user_id, headers, previous_response_id
- Pass session_id through RequestContext to usage logger
- Enable request correlation by session in proxy_request_logs
196 lines
6.5 KiB
Rust
196 lines
6.5 KiB
Rust
use axum::{
|
|
http::StatusCode,
|
|
response::{IntoResponse, Response},
|
|
Json,
|
|
};
|
|
use serde_json::json;
|
|
use thiserror::Error;
|
|
|
|
#[derive(Debug, Error)]
|
|
pub enum ProxyError {
|
|
#[error("服务器已在运行")]
|
|
AlreadyRunning,
|
|
|
|
#[error("服务器未运行")]
|
|
NotRunning,
|
|
|
|
#[error("地址绑定失败: {0}")]
|
|
BindFailed(String),
|
|
|
|
#[error("请求转发失败: {0}")]
|
|
ForwardFailed(String),
|
|
|
|
#[error("无可用的Provider")]
|
|
NoAvailableProvider,
|
|
|
|
#[error("所有供应商已熔断,无可用渠道")]
|
|
AllProvidersCircuitOpen,
|
|
|
|
#[error("未配置供应商")]
|
|
NoProvidersConfigured,
|
|
|
|
#[allow(dead_code)]
|
|
#[error("Provider不健康: {0}")]
|
|
ProviderUnhealthy(String),
|
|
|
|
#[error("上游错误 (状态码 {status}): {body:?}")]
|
|
UpstreamError { status: u16, body: Option<String> },
|
|
|
|
#[error("超过最大重试次数")]
|
|
MaxRetriesExceeded,
|
|
|
|
#[error("数据库错误: {0}")]
|
|
DatabaseError(String),
|
|
|
|
#[error("配置错误: {0}")]
|
|
ConfigError(String),
|
|
|
|
#[allow(dead_code)]
|
|
#[error("格式转换错误: {0}")]
|
|
TransformError(String),
|
|
|
|
#[allow(dead_code)]
|
|
#[error("无效的请求: {0}")]
|
|
InvalidRequest(String),
|
|
|
|
#[error("超时: {0}")]
|
|
Timeout(String),
|
|
|
|
/// 流式响应空闲超时
|
|
#[allow(dead_code)]
|
|
#[error("流式响应空闲超时: {0}秒无数据")]
|
|
StreamIdleTimeout(u64),
|
|
|
|
/// 认证错误
|
|
#[allow(dead_code)]
|
|
#[error("认证失败: {0}")]
|
|
AuthError(String),
|
|
|
|
#[allow(dead_code)]
|
|
#[error("内部错误: {0}")]
|
|
Internal(String),
|
|
}
|
|
|
|
impl IntoResponse for ProxyError {
|
|
fn into_response(self) -> Response {
|
|
let (status, body) = match &self {
|
|
ProxyError::UpstreamError {
|
|
status: upstream_status,
|
|
body: upstream_body,
|
|
} => {
|
|
let http_status =
|
|
StatusCode::from_u16(*upstream_status).unwrap_or(StatusCode::BAD_GATEWAY);
|
|
|
|
// 尝试解析上游响应体为 JSON,如果失败则包装为字符串
|
|
let error_body = if let Some(body_str) = upstream_body {
|
|
if let Ok(json_body) = serde_json::from_str::<serde_json::Value>(body_str) {
|
|
// 上游返回的是 JSON,直接透传
|
|
json_body
|
|
} else {
|
|
// 上游返回的不是 JSON,包装为错误消息
|
|
json!({
|
|
"error": {
|
|
"message": body_str,
|
|
"type": "upstream_error",
|
|
}
|
|
})
|
|
}
|
|
} else {
|
|
json!({
|
|
"error": {
|
|
"message": format!("Upstream error (status {})", upstream_status),
|
|
"type": "upstream_error",
|
|
}
|
|
})
|
|
};
|
|
|
|
(http_status, error_body)
|
|
}
|
|
_ => {
|
|
let (http_status, message) = match &self {
|
|
ProxyError::AlreadyRunning => (StatusCode::CONFLICT, self.to_string()),
|
|
ProxyError::NotRunning => (StatusCode::SERVICE_UNAVAILABLE, self.to_string()),
|
|
ProxyError::BindFailed(_) => {
|
|
(StatusCode::INTERNAL_SERVER_ERROR, self.to_string())
|
|
}
|
|
ProxyError::ForwardFailed(_) => (StatusCode::BAD_GATEWAY, self.to_string()),
|
|
ProxyError::NoAvailableProvider => {
|
|
(StatusCode::SERVICE_UNAVAILABLE, self.to_string())
|
|
}
|
|
ProxyError::AllProvidersCircuitOpen => {
|
|
(StatusCode::SERVICE_UNAVAILABLE, self.to_string())
|
|
}
|
|
ProxyError::NoProvidersConfigured => {
|
|
(StatusCode::SERVICE_UNAVAILABLE, self.to_string())
|
|
}
|
|
ProxyError::ProviderUnhealthy(_) => {
|
|
(StatusCode::SERVICE_UNAVAILABLE, self.to_string())
|
|
}
|
|
ProxyError::MaxRetriesExceeded => {
|
|
(StatusCode::SERVICE_UNAVAILABLE, self.to_string())
|
|
}
|
|
ProxyError::DatabaseError(_) => {
|
|
(StatusCode::INTERNAL_SERVER_ERROR, self.to_string())
|
|
}
|
|
ProxyError::ConfigError(_) => (StatusCode::BAD_REQUEST, self.to_string()),
|
|
ProxyError::TransformError(_) => {
|
|
(StatusCode::UNPROCESSABLE_ENTITY, self.to_string())
|
|
}
|
|
ProxyError::InvalidRequest(_) => (StatusCode::BAD_REQUEST, self.to_string()),
|
|
ProxyError::Timeout(_) => (StatusCode::GATEWAY_TIMEOUT, self.to_string()),
|
|
ProxyError::StreamIdleTimeout(_) => {
|
|
(StatusCode::GATEWAY_TIMEOUT, self.to_string())
|
|
}
|
|
ProxyError::AuthError(_) => (StatusCode::UNAUTHORIZED, self.to_string()),
|
|
ProxyError::Internal(_) => {
|
|
(StatusCode::INTERNAL_SERVER_ERROR, self.to_string())
|
|
}
|
|
ProxyError::UpstreamError { .. } => unreachable!(),
|
|
};
|
|
|
|
let error_body = json!({
|
|
"error": {
|
|
"message": message,
|
|
"type": "proxy_error",
|
|
}
|
|
});
|
|
|
|
(http_status, error_body)
|
|
}
|
|
};
|
|
|
|
(status, Json(body)).into_response()
|
|
}
|
|
}
|
|
|
|
/// 错误分类
|
|
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
|
pub enum ErrorCategory {
|
|
/// 可重试错误(网络问题、5xx)
|
|
Retryable, // 网络超时、5xx 错误
|
|
/// 不可重试错误(4xx、认证失败)
|
|
NonRetryable, // 认证失败、参数错误、4xx 错误
|
|
#[allow(dead_code)]
|
|
ClientAbort, // 客户端主动中断
|
|
}
|
|
|
|
/// 判断错误是否可重试
|
|
#[allow(dead_code)]
|
|
pub fn categorize_error(error: &reqwest::Error) -> ErrorCategory {
|
|
if error.is_timeout() || error.is_connect() {
|
|
return ErrorCategory::Retryable;
|
|
}
|
|
|
|
if let Some(status) = error.status() {
|
|
if status.is_server_error() {
|
|
ErrorCategory::Retryable
|
|
} else if status.is_client_error() {
|
|
ErrorCategory::NonRetryable
|
|
} else {
|
|
ErrorCategory::Retryable
|
|
}
|
|
} else {
|
|
ErrorCategory::Retryable
|
|
}
|
|
}
|