About the Client
The client is a globally recognized university system specializing in postgraduate medical and nursing education. It operates one of the largest clinical simulation centers in North America, conducting thousands of high-stakes Objective Structured Clinical Examinations (OSCEs) and Simulated Patient Encounters (SPEs) each year.
These simulations are an essential part of clinical education, allowing students to demonstrate medical knowledge, procedural skills, communication, empathy, and decision-making in controlled, realistic environments.
Background
During clinical simulations, students interact with trained actors portraying patients or medical mannequins designed for hands-on procedures. Faculty evaluators assess performance across a broad range of technical and interpersonal competencies—from clinical reasoning and diagnostic accuracy to communication, confidence, empathy, and patient engagement.
As enrollment increased and simulation scenarios became more sophisticated, the evaluation process became increasingly difficult to scale.
Faculty members had to manually review lengthy video recordings, evaluate multiple criteria, compare performance against detailed rubrics, and prepare individualized feedback for students. With thousands of sessions taking place each year, this placed substantial demands on instructors while making consistency across evaluators increasingly challenging.
The university needed a scalable evaluation approach that could reduce manual effort while maintaining the rigor, fairness, and educational quality required for high-stakes clinical training.
The Challenge
The university identified three primary challenges affecting its simulation assessment process.
Time-Intensive Manual Evaluation
Reviewing, scoring, and documenting a single simulation could require 45–60 minutes of instructor time.
As simulation volume grew, maintaining this level of manual review became increasingly difficult and consumed time faculty could otherwise dedicate to teaching, mentoring, and personalized student guidance.
Inconsistent Evaluation Across Simulation Formats
Clinical assessments ranged from straightforward one-on-one encounters to multi-student, multi-patient, and multi-room simulations.
Applying consistent evaluation standards across these different formats was challenging, creating potential variation in scoring and feedback.
Subjectivity in Soft-Skill Assessment
Competencies such as empathy, confidence, communication, and patient engagement are critical to clinical practice but inherently more difficult to assess consistently.
Even with standardized rubrics, differences in evaluator interpretation could introduce variability into scoring.
The university’s objective was to reduce instructor workload, improve evaluation consistency, and provide students with faster, more actionable feedback—without removing faculty judgment from the assessment process.
The Solution
The university partnered with Supply Medium to design and implement a secure, cloud-native Multimodal Generative AI Evaluation System built on AWS.
The solution analyzes video, audio, transcripts, and structured assessment criteria to create comprehensive, rubric-aligned evaluations of clinical simulations.
A Retrieval-Augmented Generation (RAG) architecture provides the AI with relevant institutional rubrics, guidelines, historical context, and evaluator feedback, helping maintain alignment with the university’s assessment standards as those standards evolve.
The application was deployed on Amazon EKS to support scalable, containerized workloads, while Amazon RDS supported backend application data and the knowledge required by the RAG workflow.
Implementation was structured across three key phases.
Phase 1: Multimodal Data Integration & Processing
The first phase established the data foundation required to analyze complex clinical simulations.
Synchronized Audio & Video Ingestion
Video and audio streams from simulation rooms were time-synchronized and securely ingested into the cloud.
Synchronization allowed the system to correlate speech, physical actions, patient responses, and other events across the entire encounter.
Speech & Sentiment Analysis
Speech recognition transformed conversations into structured transcripts for downstream analysis.
Audio-based analysis evaluated factors such as tone and clarity, providing additional signals for assessing communication and interaction patterns within the encounter.
Multimodal Video Analysis
Video streams were processed using Gemini Pro to identify relevant observable behaviors, including:
- Body posture and positioning
- Hand gestures
- Eye contact
- Interaction patterns
- Adherence to defined procedural steps
Combining visual observations with audio and transcript data provided a richer representation of student performance than any individual modality could provide independently.
RAG Knowledge Foundation
Extracted transcripts, structured observations, and relevant contextual information were stored through the platform’s data layer using Amazon RDS.
The RAG workflow enabled retrieval of relevant institutional rubrics, domain guidelines, historical context, and prior evaluator feedback when generating assessments.
This ensured evaluations were grounded in the university’s defined standards rather than relying solely on general model knowledge.
Phase 2: Intelligent Evaluation & Scoring
Once multimodal information had been extracted, the system synthesized the signals into a structured evaluation.
AI-Assisted Performance Analysis
Outputs from video, audio, transcript, and sentiment analysis were brought together and evaluated using Claude 3 Sonnet.
The model analyzed performance across multiple dimensions and applied the university’s detailed evaluation criteria to produce a coherent assessment of each simulation.
Rubric-Aligned Evaluation
Institutional grading rubrics were incorporated into the evaluation workflow, helping ensure that AI-generated assessments followed the same competency framework used by faculty.
Through RAG, the system could retrieve current rubric information, domain guidance, and relevant evaluator feedback stored within the knowledge layer.
As institutional standards evolved, updated information could be incorporated into future evaluations without redesigning the entire assessment workflow.
Structured Evaluation Reports
For each simulation, the system generated a structured report containing quantitative scoring and qualitative feedback.
For example, rubric categories could include:
Clinical Reasoning — 4/5
Empathy — 5/5
Alongside scores, the platform generated contextual feedback highlighting strengths and specific opportunities for improvement.
A report might note that a student demonstrated strong diagnostic reasoning while identifying a missed opportunity to summarize the patient’s concerns before proceeding with the examination.
This gave faculty a structured starting point for review while providing students with feedback tied to specific aspects of their performance.
Phase 3: Adaptability Across Simulation Environments
The platform was designed to support the wide variety of simulation formats used across the university.
One-on-One Simulations
For traditional student-patient encounters, the platform evaluated the complete interaction across clinical, procedural, and communication criteria.
Multi-Student & Multi-Patient Scenarios
During group simulations, performance signals could be associated with individual students, allowing each participant to receive a separate evaluation while preserving the context of collaborative activity.
This supported more consistent assessment even when multiple participants were involved in the same scenario.
Multi-Room Simulations
For distributed simulations, synchronized metadata connected activity captured across multiple cameras and microphones into a common timeline.
This enabled the system to evaluate events across different locations as components of a single simulation rather than isolated recordings.
Scalable Cloud Architecture
Deployment on Amazon EKS allowed the platform to scale containerized workloads based on simulation volume and processing demand.
This architecture supported parallel evaluation workloads and helped accommodate peak assessment periods without requiring equivalent increases in manual infrastructure management.
Continuous Refinement
Evaluator corrections and institutional updates could be incorporated into the RAG knowledge workflow, helping future assessments remain aligned with evolving rubrics and educational expectations.
AI therefore functioned as an adaptive evaluation assistant while faculty retained authority over assessment decisions.
The Outcome
The implementation transformed a highly manual assessment workflow into a standardized, AI-supported evaluation process designed for greater speed, consistency, and scale.
| Metric | Before: Manual | After: GenAI | Improvement |
|---|---|---|---|
| Instructor Evaluation Time | 45–60 min | 10–15 min | 67% faster |
| Total Instructor Hours Annually | ~9,500 hrs | ~2,375 hrs | 7,000+ hrs saved |
| Evaluation Consistency (ICC) | 0.65 | 0.89 | 37% higher reliability |
| Feedback Delivery Time | 3–5 days | <1 hour | 95% faster feedback |
| Cost per Session | ~$45 | ~$5.50 | 87% cost reduction |
| Pass/Fail Variance | ~12% | <1% | Near-perfect alignment |
Reduced Instructor Workload
Evaluation time decreased from 45–60 minutes to approximately 10–15 minutes per session, significantly reducing the amount of repetitive review required from faculty.
Across annual simulation volumes, this translated into more than 7,000 instructor hours saved, allowing educators to dedicate more time to mentoring, teaching, and individualized student development.
Greater Evaluation Consistency
Evaluation consistency improved from an ICC of 0.65 to 0.89, representing a 37% increase in reliability.
Standardized rubric application and multimodal analysis helped reduce variation across evaluators and simulation formats.
Faster Student Feedback
Feedback that previously required 3–5 days could be delivered in less than one hour, enabling students to review their performance while the simulation experience was still fresh.
This created a faster feedback loop between assessment, reflection, and improvement.
Lower Evaluation Costs
The estimated cost per simulation decreased from approximately $45 to $5.50, representing an 87% reduction while supporting greater evaluation scale.
Stronger Alignment
Pass/fail variance decreased from approximately 12% to less than 1%, demonstrating significantly stronger alignment in assessment outcomes.
Lasting Impact
The Multimodal GenAI Evaluation System established a scalable framework for supporting simulation-based assessment across the university.
By combining multimodal analysis, RAG, institutional rubrics, cloud-native infrastructure, and faculty oversight, the university can evaluate growing simulation volumes while maintaining consistent assessment standards.
Faculty spend less time on repetitive review and more time providing meaningful educational guidance. Students receive faster, detailed, and actionable feedback. Administrators gain an evaluation model capable of scaling alongside enrollment and increasingly complex simulation environments.
Through its partnership with Supply Medium, the university demonstrated how Generative AI can strengthen clinical education when implemented as a collaborative tool—augmenting human expertise rather than replacing it.
The result is a faster, more consistent, and scalable approach to clinical competency assessment while keeping educational judgment and accountability firmly in the hands of faculty.