
Essay Grading with
Artificial Intelligence
Product | Machine Learning | UX | Accessibility
The Product
BARBRI Bar Review helps aspiring attorneys prepare for the Bar Exam through expert-led legal education, including unlimited essay submissions professionally evaluated by licensed attorneys.
The Problem
With courses lasting just 6–8 weeks and more than 80,000 essays submitted each season, graders faced an enormous workload while maintaining the consistency and quality students relied on. BARBRI needed a way to accelerate evaluation without compromising confidence in every score.
The Project
(1) Design an intuitive grading experience that enables attorneys to efficiently review, prioritize, and evaluate thousands of essays.
(2) Integrate AI-assisted rubric analysis to improve scoring consistency while preserving instructor judgment.
(3) Provide meaningful, actionable feedback that helps students strengthen their writing and prepare with confidence for the Bar Exam.
Tools | Figma, Adobe Suite, Miro, JIRA
Duration | May 2024 - August 2024
Team | Society Product Squad (1 Product Owner, 1 PM, 1 UX, 7-8 Devs) + ML/AI Data Team
My Role | Lead Product Designer
Background
Gathering Requirements
Creating Cohesion
Classifying Feedback
Working with the ML Team
Testing with Graders
Creating Confidence
Making Adjustments
Outcomes
The Process | TOC


Notable Impacts & Outcomes
-
Reduced average grading time from 20–30 minutes during initial testing to under 15 minutes per essay (averaging approximately 12 minutes).
-
Centralized the management of thousands of essay submissions through a searchable grading workspace designed for high-volume evaluation.
-
Introduced BARBRI's first AI-assisted grading experience, establishing the foundation for future machine learning capabilities across the LMS.
-
Created a scalable human-in-the-loop evaluation workflow that balanced AI-assisted scoring with professional grader oversight.
-
Designed to support more than 125,000 essay submissions each season, enabling BARBRI to scale its unlimited essay feedback program.
Background
As BARBRI modernized its learning platform, essay grading became a critical opportunity to improve both operational efficiency and the student experience. While AI offered the potential to accelerate evaluation, introducing it into a high-stakes educational environment required balancing speed with consistency, transparency, and instructor trust. Every design decision had to support BARBRI's promise of delivering meaningful, high-quality feedback to every student preparing for the Bar Exam.
I led the 0-to-1 design of BARBRI's first AI-assisted grading experience—balancing efficiency with the trust, transparency, and human oversight required to support future attorneys and legal professionals.
Gathering Requirements
Before exploring solutions, my Product Owner and I partnered to define the foundational requirements for BARBRI's modernized essay management platform. Rather than designing AI as a standalone feature, we mapped the end-to-end ecosystem across students, professional graders, and administrators to understand where technology could meaningfully improve the experience.
Working collaboratively with product, engineering, and machine learning teams, we identified opportunities where AI could reduce repetitive evaluation tasks while preserving the expertise and judgment of professional graders. Using Miro to map workflows, Figma to explore concepts, and Jira to coordinate cross-functional efforts, we established a shared vision that balanced operational efficiency with BARBRI's commitment to delivering consistent, high-quality feedback at scale.


Creating Cohesion
Although students and graders interacted with different interfaces, they were participating in the same learning experience. I designed both applications as a unified system, establishing shared interaction patterns, navigation, and visual language that created continuity across the platform while supporting each user's unique responsibilities.
Maintaining consistency across experiences reduced the learning curve, accelerated development through reusable components, and reinforced BARBRI's broader LMS ecosystem—allowing students and graders to focus on meaningful feedback rather than learning new interfaces.
Student-Essay UI
Essay AI-Grader UI


Classifying Feedback
Designing an AI-assisted grading experience began with understanding how professional graders evaluate essays. Together with legal experts and the machine learning team, we identified two distinct types of feedback—one that determined a student's score and another that supported their learning.
-
Quantitative (Graded): Rubric-based evaluations that contribute to a student's score and overall performance.
-
Qualitative (Instructional): Personalized feedback that guides, encourages, and helps students improve without affecting their final score.
Separating these feedback models shaped both the user experience and the AI itself. Clear rubric structures improved usability for graders while providing the machine learning team with the consistent training data needed to generate accurate, explainable scoring recommendations.
Working with the Machine Learning Team
Designing the grader experience required close collaboration with BARBRI's Machine Learning team. Together, we translated the grading expertise of legal professionals into structured evaluation models the AI could learn from—ensuring recommendations reflected established rubric criteria while remaining grounded in human judgment.
Because every essay assignment followed its own jurisdiction-specific rubric, the evaluation process was highly nuanced. Individual essays could contain more than 80 rubric criteria, each carrying unique point values across 6-, 10-, or 100-point scoring scales. Rather than replacing graders, the AI analyzed rubric responses to generate scoring recommendations that accelerated evaluation while preserving instructor oversight.

To train the model, we leveraged historical essay submissions with validated scoring outcomes. These examples were categorized by response quality to build a reliable training dataset, enabling the AI to recognize grading patterns before being validated by professional graders. This iterative partnership between machine learning and subject matter experts laid the foundation for a more consistent, trustworthy evaluation process.
Testing with Graders
Introducing this new platform required a shift in grader workflows. Previously, graders manually scored each rubric, tallied the total, and left it to students to interpret their jurisdictional scale. With the new system, graders instead reviewed AI-generated scores for accuracy—actively training the AI with every correction—while continuing to provide instructional feedback.
We conducted user testing with graders to evaluate whether the full scoring and feedback process could be completed in under 15 minutes per essay. Initial sessions averaged 20–30 minutes due to the learning curve of the new interface and the cognitive load of parsing numerous rubric items.
Building Confidence
One of the most impactful design decisions was introducing confidence scores at the rubric level. Rather than treating every AI recommendation equally, graders could immediately identify low-confidence evaluations that required closer review while quickly validating high-confidence responses.
This human-in-the-loop approach reduced unnecessary review time, prioritized expert attention where it mattered most, and reinforced trust in AI-assisted evaluation.


Making Adjustments
Following usability testing, I introduced keyboard shortcuts and workflow optimizations to reduce repetitive interactions and accelerate score entry. These refinements helped graders process high-confidence essays more efficiently while preserving time for thoughtful review of more complex submissions.
Score Report (Self-Grade Only)
Score Report (Graded)



Check out the full experience below!
Reflection / What I learned...
Looking back, this project marked the beginning of my journey designing AI-powered experiences. While the goal was to help graders process thousands of essays within a demanding six-to-eight-week grading window, the real challenge wasn't automation—it was trust.
BARBRI's promise to students was built on fairness, consistency, and helping aspiring attorneys succeed. Introducing AI into that process meant every design decision had to reinforce confidence rather than replace human expertise. Features like confidence scoring, human review workflows, and grader-trained models weren't simply technical capabilities—they were essential guardrails that ensured instructors remained accountable for every evaluation.
Working across engineering, data science, educational experts, and business stakeholders also taught me that successful AI products are as much organizational challenges as they are design challenges. Building alignment around how AI should support people required as much collaboration as designing the experience itself.
More than any individual feature, this project shaped the philosophy I continue to carry into every AI product I design today: the most effective AI doesn't replace human judgment—it strengthens it. When AI is transparent, accountable, and thoughtfully integrated into existing expertise, it becomes more than an automation tool; it becomes a trusted partner.
Responsible AI by Design
Educational AI isn't simply about automation—it directly influences how students learn and instructors teach. Throughout the project, we intentionally designed experiences that kept humans accountable while allowing AI to accelerate repetitive evaluation.

Human Oversight
AI generated recommendations.
Faculty retained final grading authority.

Creating Confidence
Students and instructors understood how scores were generated.

Consistency without Rigidity
Improved consistency while preserving instructor judgment.

Continuous Learning
Continuously improved through human review.
Applied to AI Grading
From secure data handling and transparent recommendations to human oversight and role-based access, Champion AI was designed to empower expert consultants - never replace them.