CourseCompass AI
A degree-planning assistant for UNC students covering 200+ majors and minors plus general education. It answers from the official catalog with clickable citations and checks uploaded schedules against requirements.
- Eval pass rate 11% → 100%
- 200+ programs and tracks
- 80+ offline tests in CI

What it’s for
I built CourseCompass AI to help UNC students understand what they still need to graduate. They can ask about a major in plain language, or upload a photo of their schedule, and get an answer grounded in the official course catalog.
How I built it
A Python scraper turns catalog pages for 202 majors, minors, and tracks plus the IDEAs in Action general education pages into structured documents, indexed with Titan embeddings in Amazon S3 Vectors through a Bedrock Knowledge Base. Most questions are routed to the right program in code; only unclear ones use a Claude classifier call. The routed program’s requirements are always included, sources are numbered, and every citation is verified on the server before it reaches the student.
What it does
Answers lead with the direct answer, cite the exact catalog page, and ask one focused question when the program or degree is ambiguous. Students can add courses by photo, PDF, or pasted text; minimum-grade rules are checked in code, so a C- is never counted toward “C or better.” Rate limits and an AWS Budget action that blocks the app at $10 a month cap spending.
The interesting part
Measuring changed how I built it. An 18-case live eval graded in code against catalog facts scored the first working version at 2 of 18 cases; the rebuilt pipeline passes all 18. The biggest fixes came from moving judgment the model kept getting wrong, like comparing letter grades, into deterministic code, and from sending less but more relevant context: about 58% fewer input tokens per answer.
Guardrails
Answers are grounded only in retrieved catalog text, and the bot says so when the catalog doesn’t cover a question. Retrieved text and uploads are treated as data, never instructions. Citation numbers that don’t match a real source are removed, off-topic requests are declined, and grade comparisons are computed rather than generated.
Testing and CI/CD
Offline suites with mocked Bedrock calls cover routing, retrieval assembly, citations, grade checks, uploads, and spending limits. GitHub Actions runs lint, type checks, the tests, and a production build on every push and pull request, and Vercel deploys from GitHub. The live eval runs only on demand because it calls the model.
VibeSafe
HAVE A QUESTION OR AN IDEA?