Business Overview

Our client operates a leading medical exam preparation platform in the United States, providing high-fidelity mock exams and learning tools for medical students. Previously, their internal editorial team relied on a highly manual, fragmented workflow. Editors used public web interfaces like Anthropic’s Claude chat to generate board-style questions—and it required uploading lengthy prompts and raw source documents. This single-shot approach involved tedious manual prompt handling, caused massive data-entry bottlenecks, and risked exposing proprietary question banks to public LLM training models.
To safely scale the content pipeline, the client required a secure, purpose-built, internal question-generation solution.

They envisioned a platform that would serve as a centralized workspace and transform raw medical guidelines, PDFs, and source documents into rigorous, board-style exam questions—complete with complex multi-choice options, detailed rationale explanations, and clinical takeaways.
The client approached NIX to build such a solution from the ground up. Their primary request was to engineer an automated, centralized workspace that integrates directly with their proprietary database, maximizes human-in-the-loop editing efficiency, and ensures complete data isolation to eliminate data leakage.

Project Scope

img 02@2x

Our team engineered and delivered the AI-based learning platform completely from scratch, moving from initial concept to a production-grade internal application. The core scope of work included:

  • End-to-end system architecture
  • Multi-agent engine development
  • Primary product feature integration
  • Flexible containerized deployment
  • Agile collaboration and refinement

Challenge

The core engineering hurdle of this project was achieving an ultra-fast generation time without sacrificing the strict educational standards required for medical board exams. To deliver the project successfully, our team had to overcome the following specific challenges:

  • 01

    Strict Latency Constraints

    The platform was bound to a rigid performance requirement to generate complex clinical vignettes, multiple-choice options, and detailed rationales in under one minute.

  • 02

    Time-consuming Verifications

    To ensure high-quality, accurate, and reliable questions, the system needed to run multiple validation steps. Typically, these multi-turn AI workflows add substantial processing overhead and delays.

  • 03

    Bottlenecks of Linear Processing

    Standard, line-by-line text generation proved too slow to meet the client’s needs. We overcame this bottleneck by completely shifting away from linear generation to an advanced, concurrent, parallel map-reduce architecture.

Solution

Intelligent Ingestion and Dynamic Context Enrichment

Our team custom-engineered a sophisticated data ingestion pipeline, designed to transform raw, unstructured medical data into highly structured, context-aware prompts while entirely eliminating manual copy-pasting for the editorial team.

  • Dynamic guideline and prompt generation: Rather than hardcoding rules for a single exam, NIX built a flexible PDF-processing pipeline that handles varying institutional standards. It extracts formatting, style, and questioning constraints directly from uploaded exam blueprints or guideline documents. The system automatically maps these into a master prompt template that editors can quickly review, fine-tune, and save.
  • Live clinical web verification: To prevent inaccuracies and ensure all generated content aligns with the most current medical data, we engineered and integrated a real-time clinical web search engine into the generation pipeline. It actively queries the web for the latest peer-reviewed research and current clinical practice guidelines, injecting fresh, verified data directly into the prompt context before generation begins.
img 03@2x
img 04@2x

Orchestration: High-speed Parallel Multi-agent Engine

To balance the strict latency requirement with the rigor of medical practice, NIX bypassed standard linear text generation in favor of a stateful, concurrent architecture that transforms raw AI capabilities into a scalable, enterprise-grade production tool.

  • Stateful cyclic orchestration: Powered by LangGraph and LangChain, the intelligence layer replaces unpredictable AI guesses with a highly disciplined, multi-turn workflow. The system tracks context across multiple stages so the AI strictly adheres to complex guidelines, ensuring consistent, standardized content quality.
  • Parallel map-reduce framework: Acting like an automated medical board, a central orchestrator drafts the core clinical scenario (the map phase) and concurrently fans out sub-tasks to specialized threads operating in parallel. This operational shift eliminates production bottlenecks, allowing the client to scale their content library at a fraction of the traditional time and cost.
  • Specialized concurrent execution threads: The system fires simultaneous LLM API calls to independently handle rationales, verify distractors, and extract takeaways. It seamlessly merges these pieces into a board-style question while embedding automated quality control directly into the code. Because these threads cross-verify components behind the scenes, editors receive highly accurate drafts requiring minimal rewriting—drastically reducing human editing time, slashing operational overhead, and accelerating time-to-market.

Human-in-the-loop and Interactive Revision Engine

Recognizing that medical education requires absolute professional oversight, NIX designed the system to ensure that human experts always retain total control over the AI’s output.

  • Streamlit review dashboard: An intuitive Streamlit frontend, backed by FastAPI, allows the internal editorial team to evaluate and edit generated content without context switching.
  • Targeted revision graph: If an editor requests a localized change (e.g., altering a patient’s age or a specific choice), the system rewrites only the necessary components rather than regenerating the entire question from scratch.
  • Advanced prompt management: Users maintain complete control over the AI’s behavior by editing system prompts, saving custom rules, and tuning formatting preferences to match their unique pedagogical voice.
img 05@2x
img 06@2x

Production-grade Deployment and Secure Ingestion

Data privacy and platform scalability were foundational to the application’s physical architecture.

  • Isolated database integration: Approved questions are automatically packaged with their metadata and tables, and saved directly into the organization’s database and secure storage (Question Bank), ready for immediate export or deployment to medical students.
  • Containerized infrastructure: To maximize availability and code portability, the application is packaged via Docker and deployed using AWS EC2 and AWS Lightsail, ensuring rapid rollout and low infrastructure overhead.

Outcome

As a result of this collaboration with NIX, the client yielded a highly successful, production-grade automation workspace delivered completely from scratch.

 

The standalone platform has been fully built, rigorously tested, and successfully deployed for internal use. As the project team wraps up the final features, the project is smoothly transitioning into the platform integration phase. The next milestone focuses on embedding the core AI loops directly into the client’s primary commercial platform, enabling their team to instantly push newly vetted, board-style questions straight into production for their medical students.

 

By transitioning from fragmented public AI tools to the tailored intelligent question generation solution, the client achieved substantial business value across security, velocity, and quality control

  • Eliminated data leakage risks: The private, containerized cloud environment ensures complete data isolation, preventing proprietary question banks from leaking or being used for public model training.
  • Accelerated content pipeline: High-speed parallel processing slashes question generation time to under 60 seconds, completely removing data-entry bottlenecks and driving faster time-to-market.
  • Guaranteed medical rigor: Built-in automated quality gates and verification loops ensure clinical accuracy, maintaining strict educational standards while drastically reducing manual rewriting time.
  • Seamless professional oversight: The custom interactive dashboard and revision graph give editors total control to tweak specific clinical details on the fly without regenerating entire questions.
img 07@2x
Team:

Team:

3 experts ( AI engineer, Tech Lead, Project Manager )
Tech stack:

Tech stack:

FastAPI, LangChain, LangGraph, Streamlit, AWS App Runner, AWS Lightsail, Docker, Generative AI

REQUEST A CONSULTATION

Contact us   

Relevant Case Studies

View all case studies

Ending the Search Bottleneck: Custom RAG Chatbot Development for an EdTech Leader

Education

Success Story Ending the Search Bottleneck: Custom RAG Chatbot Development for an EdTech Leader image

65% Less Manual Work: Healthcare Document Automation With AI

Healthcare

Success Story 65% Less Manual Work: Healthcare Document Automation With AI image

AI-powered Search Solution for a Healthcare Company

Healthcare

Success Story AI-powered Search Solution for a Healthcare Company image

AI Chatbot for Personalized Healthcare

Healthcare

Success Story AI Chatbot for Personalized Healthcare image

vSentry—AI Web App for Vehicle Monitoring

Cybersecurity

Electronics

Success Story vSentry—AI Web App for Vehicle Monitoring image

Voice AI in Healthcare: Faster Specialty Medication Access

Healthcare

Success Story Voice AI in Healthcare: Faster Specialty Medication Access image
01

Contact Us

Accessibility Adjustments
Adjust Background Colors
Adjust Text Colors