feat: tiered LLM (GLM free / Claude paid) + rate limits + quota enforcement
All checks were successful
Deploy to Production / deploy (push) Successful in 53s
All checks were successful
Deploy to Production / deploy (push) Successful in 53s
The free tier was hemorrhaging Anthropic cost with no abuse cap (no rate limit on /preview, Opus default in the build worker, 5-min cache TTL that made cache-miss the common case). This switches free users to GLM, paid users to Claude tiers, and tightens every leak found in the audit. Backend: - @bmm/llm: GLM provider via Zhipu's OpenAI-compatible endpoint, pickPreviewModel + pickBuildModel helpers, plan-aware ModelChoice - preview-cache TTL 5min -> 24h (kills the cache-miss path) - /v1/servers/preview: picks model from caller's plan, returns model name to UI - /v1/servers POST: enforces SERVER_LIMITS per plan (402), rate-limits builds - daily rate-limit on preview (5/40/150/1000) and build (3/20/100/500) - /v1/auth/me returns plan so the wizard can show the right model name - generator worker: GLM default, Anthropic Sonnet fallback if GLM errors Frontend: - Wizard fetches plan, shows "<model> is drafting the tool spec" pre-emptively, upgrade hint for hobby users, friendly errors for 402 / 429 - Pricing page: AI-model line per tier (Open-tier / Haiku / Sonnet / Opus), Team €149 -> €199, Enterprise €499 -> €999, daily-preview limit per tier - Privacy + Security: explicit subprocessor disclosure for Anthropic (US) / Zhipu (CN) and which tier uses which Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1,12 +1,40 @@
|
||||
import { generateSpec as sharedGenerate, type GenerationResult } from '@bmm/llm';
|
||||
import { type GenerationResult, generateSpec as sharedGenerate } from '@bmm/llm';
|
||||
import { config } from '../config.js';
|
||||
|
||||
export type { GenerationResult };
|
||||
|
||||
/**
|
||||
* Build-worker spec generation (cache-miss path). Runs async in a BullMQ
|
||||
* worker — no proxy timeout. Defaults to GLM to keep this rare path cheap;
|
||||
* falls back to Anthropic Sonnet on GLM failure so a temporary outage at one
|
||||
* provider doesn't break builds.
|
||||
*/
|
||||
export async function generateSpec(prompt: string): Promise<GenerationResult> {
|
||||
if (config.GLM_API_KEY) {
|
||||
try {
|
||||
return await sharedGenerate(prompt, {
|
||||
provider: 'glm',
|
||||
glmApiKey: config.GLM_API_KEY,
|
||||
model: config.MODEL_GENERATE,
|
||||
maxTokens: 8192,
|
||||
timeoutMs: 180_000,
|
||||
});
|
||||
} catch (err) {
|
||||
console.warn(
|
||||
'[generator] GLM failed, falling back to Anthropic Sonnet:',
|
||||
(err as Error).message,
|
||||
);
|
||||
}
|
||||
}
|
||||
if (!config.ANTHROPIC_API_KEY) {
|
||||
// No keys at all → @bmm/llm returns mockSpec, which keeps builds working
|
||||
// in dev without any provider configured.
|
||||
return sharedGenerate(prompt, { provider: 'anthropic' });
|
||||
}
|
||||
return sharedGenerate(prompt, {
|
||||
provider: 'anthropic',
|
||||
apiKey: config.ANTHROPIC_API_KEY,
|
||||
model: config.MODEL_GENERATE,
|
||||
model: 'claude-sonnet-4-6',
|
||||
maxTokens: 8192,
|
||||
});
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user