EDULYTICS

Designing Async OMR Processing

What happens when hundreds of answer sheets need processing without slowing down the main application?

The Problem

OMR answer-sheet processing can take significantly longer than a normal application request. If every upload is processed synchronously, users are forced to wait while the server performs CPU-heavy work, making the overall product feel slower and less responsive.

The Constraint

Bursty Uploads

Many answer sheets may be uploaded within a short period of time.

Heavy Processing

Each job can require significantly more work than a standard application request.

Failures

Individual processing jobs can fail without affecting the entire application.

Retries

Failed jobs need a safe way to be processed again.

Many Jobs

The system needs to handle multiple jobs without blocking normal product traffic.

  • Bursty UploadsMany answer sheets may be uploaded within a short period of time.
  • Heavy ProcessingEach job can require significantly more work than a standard application request.
  • FailuresIndividual processing jobs can fail without affecting the entire application.
  • RetriesFailed jobs need a safe way to be processed again.
  • Many JobsThe system needs to handle multiple jobs without blocking normal product traffic.

The Options

Process Inside the Request

Not Chosen

Background Processing in the App

Not Chosen

Queue Processing Jobs

Chosen

Why We Didn't Choose It

The API would have to keep the request open until the full OMR processing job finished.

Benefit

Simple flow with fewer moving parts.

Drawback

Long waits, request timeouts, and heavy processing tied directly to user traffic.

Decision Rationale

Why We Chose The Approach

Asynchronous OMR processingThe API acknowledges uploaded answer sheets quickly and queues the processing jobs. Independent workers consume the queued jobs. Job lifecycle tracking records completion, failure and retry states. A failed job can be retried through the queue. Select a stage to highlight its decision rationale.RETRYAnswer sheetsAPIacknowledgementJob queueIndependentworkersJob lifecycleCOMPLETIONFAILURERETRY

The Result

OMR processing ran independently, keeping the application responsive during heavy workloads.

In Retrospect

If I were refining this architecture again, I’d focus on these areas:

  1. How do we make jobs safe to retry without duplication?
  2. How should failures and dead-letter handling work?
  3. How should queue backlog and worker scaling be monitored?