feat: tiered LLM (GLM free / Claude paid) + rate limits + quota enforcement
All checks were successful
Deploy to Production / deploy (push) Successful in 53s

The free tier was hemorrhaging Anthropic cost with no abuse cap (no rate
limit on /preview, Opus default in the build worker, 5-min cache TTL that
made cache-miss the common case). This switches free users to GLM, paid
users to Claude tiers, and tightens every leak found in the audit.

Backend:
- @bmm/llm: GLM provider via Zhipu's OpenAI-compatible endpoint, pickPreviewModel
  + pickBuildModel helpers, plan-aware ModelChoice
- preview-cache TTL 5min -> 24h (kills the cache-miss path)
- /v1/servers/preview: picks model from caller's plan, returns model name to UI
- /v1/servers POST: enforces SERVER_LIMITS per plan (402), rate-limits builds
- daily rate-limit on preview (5/40/150/1000) and build (3/20/100/500)
- /v1/auth/me returns plan so the wizard can show the right model name
- generator worker: GLM default, Anthropic Sonnet fallback if GLM errors

Frontend:
- Wizard fetches plan, shows "<model> is drafting the tool spec" pre-emptively,
  upgrade hint for hobby users, friendly errors for 402 / 429
- Pricing page: AI-model line per tier (Open-tier / Haiku / Sonnet / Opus),
  Team €149 -> €199, Enterprise €499 -> €999, daily-preview limit per tier
- Privacy + Security: explicit subprocessor disclosure for Anthropic (US) /
  Zhipu (CN) and which tier uses which

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Marco Sadjadi
2026-05-23 23:50:00 +02:00
parent 66128c73d8
commit bc174c1302
14 changed files with 537 additions and 58 deletions

View File

@@ -1,12 +1,40 @@
import { generateSpec as sharedGenerate, type GenerationResult } from '@bmm/llm';
import { type GenerationResult, generateSpec as sharedGenerate } from '@bmm/llm';
import { config } from '../config.js';
export type { GenerationResult };
/**
* Build-worker spec generation (cache-miss path). Runs async in a BullMQ
* worker — no proxy timeout. Defaults to GLM to keep this rare path cheap;
* falls back to Anthropic Sonnet on GLM failure so a temporary outage at one
* provider doesn't break builds.
*/
export async function generateSpec(prompt: string): Promise<GenerationResult> {
if (config.GLM_API_KEY) {
try {
return await sharedGenerate(prompt, {
provider: 'glm',
glmApiKey: config.GLM_API_KEY,
model: config.MODEL_GENERATE,
maxTokens: 8192,
timeoutMs: 180_000,
});
} catch (err) {
console.warn(
'[generator] GLM failed, falling back to Anthropic Sonnet:',
(err as Error).message,
);
}
}
if (!config.ANTHROPIC_API_KEY) {
// No keys at all → @bmm/llm returns mockSpec, which keeps builds working
// in dev without any provider configured.
return sharedGenerate(prompt, { provider: 'anthropic' });
}
return sharedGenerate(prompt, {
provider: 'anthropic',
apiKey: config.ANTHROPIC_API_KEY,
model: config.MODEL_GENERATE,
model: 'claude-sonnet-4-6',
maxTokens: 8192,
});
}