Back to projects

Vision-LLM pipeline · determinism + QC

AutoGrade - AI Paper Evaluation

A fire-and-forget edge chain - extract → grade → qc-check - sidesteps the 150s serverless timeout by treating Postgres status as the orchestrator. Vision-LLM reads handwritten + Gujarati scripts; self-consistency voting at temp=0 makes grading deterministic; an arithmetic QC gate catches AI hallucinations before they reach the teacher.

Role
Solo full-stack + AI engineer
Status
Shipped
Year
2025
autograde---ai-paper-evaluation.app
60 × 20 QUESTIONSQ18/10Q26/10Q39/10Q47/10QC · PASSPAPER + SCRIPT → PARALLEL GRADEEVIDENCEEXTRACTGRADEQCRESULT

Metrics

Grading time
~90% saved
Per paper
~2 min
Re-runs
Deterministic
Edge functions
4

Stack

React 18ViteTypeScriptTailwindshadcn/uiSupabaseEdge Functions (Deno)Lovable AI GatewayGemini 2.5 / 3.1 Pro Preview (vision)react-pdf

Problem

A teacher with 60 students × 20 questions burns 6–10 hours per exam, with inconsistent strictness across the batch and zero traceability of why marks were deducted. No off-the-shelf product handles handwritten, multi-page, mixed-script papers with auditable evidence.

Solution

React/Vite SPA → Supabase (Postgres + RLS + Auth + Storage + Edge Functions) → Lovable AI Gateway (Gemini Pro, vision). Fire-and-forget edge chain with DB-status orchestration. react-pdf canvas rendering enables bounding-box evidence overlays.

Architecture

QP Upload
→
Parse
→
Student PDF
→
Extract
→
Grade (parallel)
→
QC Check
→
Results

Key Engineering Decisions

  • DB-as-orchestrator instead of Inngest/SQS

    Work is self-bounded (one student paper) with no fan-out. Postgres status is the source of truth - inspectable with `SELECT status FROM evaluations`. The right level of complexity for the problem.

  • Fire-and-forget edge chain

    AI pipeline is 60–120s; edge timeout is 150s. Chaining functions via `fetch keepalive:true` decouples the long job from the HTTP request lifecycle.

  • Determinism contract across every AI call

    temp=0, top_p=0.1, self-consistency voting. Same paper graded twice produces the same marks - non-negotiable for teacher trust.

  • Canvas (react-pdf), not iframe

    Bypasses browser PDF sandboxing AND enables bounding-box overlays for evidence - required for the product UX.

Build Notes

  • 4 Deno edge functions: parse-question-paper, extract, grade, qc-check - JWT-verified at the boundary, service-role for internal chaining
  • Per-question parallel grading at `temp=0`, `top_p=0.1` with self-consistency voting for reproducible marks
  • qc-check re-derives totals, verifies arithmetic, evaluates evidence quality, flags low-confidence questions; emits PASS / FLAG / FAIL
  • react-pdf canvas (not iframe) renders pages so deductions can highlight the exact bounding box on click
  • UI polls `evaluations.status` every 3s; progress bar maps pending→processing→grading→qc_check→completed
  • Multi-tenant security: RLS on every table, private buckets with short-lived signed URLs, HIBP leaked-password protection

Results

  • ~90% reduction in grading time per batch
  • Deterministic re-runs: identical paper → identical marks
  • Zero arithmetic errors reach the teacher (QC gate)
  • Every deduction defensible: click → bounding box on the rendered PDF

Why this matters

  • Vision LLM in production on a messy real-world dataset
  • Self-consistency + arithmetic QC as guardrails - not just prompting
  • Timeout-driven architecture: design the workflow around the runtime constraint
  • Evidence UX: every AI decision is traceable to a region on the page