Vision-LLM pipeline · determinism + QC
AutoGrade - AI Paper Evaluation
A fire-and-forget edge chain - extract → grade → qc-check - sidesteps the 150s serverless timeout by treating Postgres status as the orchestrator. Vision-LLM reads handwritten + Gujarati scripts; self-consistency voting at temp=0 makes grading deterministic; an arithmetic QC gate catches AI hallucinations before they reach the teacher.
- Role
- Solo full-stack + AI engineer
- Status
- Shipped
- Year
- 2025
Metrics
- Grading time
- ~90% saved
- Per paper
- ~2 min
- Re-runs
- Deterministic
- Edge functions
- 4
Stack
Problem
A teacher with 60 students × 20 questions burns 6–10 hours per exam, with inconsistent strictness across the batch and zero traceability of why marks were deducted. No off-the-shelf product handles handwritten, multi-page, mixed-script papers with auditable evidence.
Solution
React/Vite SPA → Supabase (Postgres + RLS + Auth + Storage + Edge Functions) → Lovable AI Gateway (Gemini Pro, vision). Fire-and-forget edge chain with DB-status orchestration. react-pdf canvas rendering enables bounding-box evidence overlays.
Architecture
Key Engineering Decisions
DB-as-orchestrator instead of Inngest/SQS
Work is self-bounded (one student paper) with no fan-out. Postgres status is the source of truth - inspectable with `SELECT status FROM evaluations`. The right level of complexity for the problem.
Fire-and-forget edge chain
AI pipeline is 60–120s; edge timeout is 150s. Chaining functions via `fetch keepalive:true` decouples the long job from the HTTP request lifecycle.
Determinism contract across every AI call
temp=0, top_p=0.1, self-consistency voting. Same paper graded twice produces the same marks - non-negotiable for teacher trust.
Canvas (react-pdf), not iframe
Bypasses browser PDF sandboxing AND enables bounding-box overlays for evidence - required for the product UX.
Build Notes
- 4 Deno edge functions: parse-question-paper, extract, grade, qc-check - JWT-verified at the boundary, service-role for internal chaining
- Per-question parallel grading at `temp=0`, `top_p=0.1` with self-consistency voting for reproducible marks
- qc-check re-derives totals, verifies arithmetic, evaluates evidence quality, flags low-confidence questions; emits PASS / FLAG / FAIL
- react-pdf canvas (not iframe) renders pages so deductions can highlight the exact bounding box on click
- UI polls `evaluations.status` every 3s; progress bar maps pending→processing→grading→qc_check→completed
- Multi-tenant security: RLS on every table, private buckets with short-lived signed URLs, HIBP leaked-password protection
Results
- ~90% reduction in grading time per batch
- Deterministic re-runs: identical paper → identical marks
- Zero arithmetic errors reach the teacher (QC gate)
- Every deduction defensible: click → bounding box on the rendered PDF
Why this matters
- Vision LLM in production on a messy real-world dataset
- Self-consistency + arithmetic QC as guardrails - not just prompting
- Timeout-driven architecture: design the workflow around the runtime constraint
- Evidence UX: every AI decision is traceable to a region on the page