Skip to content
All work
Case Study — Core AI Systems

AI Interview Platform

An asynchronous AI interviewer that holds a real technical conversation — streaming the candidate's speech, forming contextual follow-up questions from what they actually said, and producing a structured evaluation afterwards.

Role
Core AI Dev
Mode
Async Voice
Pipeline
STT → LLM → TTS
Output
Scored Report

The Friction

First-round technical screening is expensive, scheduling-bound, and inconsistent — two candidates rarely get the same interview. Existing automated tools ask fixed questions from a list, which candidates quickly learn to game and which cannot probe an interesting answer.

The Adaptive Screen

A context-aware interviewer that evaluates candidate speech as it arrives, decides whether an answer warrants a follow-up, and generates that follow-up from the specific claim the candidate made. Candidates authenticate by UUID link — no account creation — and the session records audio and video alongside the transcript.

UUID AccessAdaptive Follow-upsSession Recording

System architecture

AI Interview Platform architecture: candidate audio and video feeds a real-time speech-to-text stream, into an LLM interviewer engine handling dialogue management and question formulation, out through text-to-speech response generation back to the candidate, with a parallel branch into an automated evaluation module producing scoring and reports.
AI Interview Platform — pipeline

Hover or focus a stage to see why it’s built that way.

Turn-taking as a first-class problem

The hardest part was not generation — it was knowing when the candidate had finished speaking. Interrupting someone mid-thought destroys the experience, so silence detection is tuned deliberately conservative.

UUID links over accounts

Requiring a signup before an interview costs candidates. A signed, single-use UUID link removes the barrier while keeping each session isolated.

Evaluation decoupled from the interview

Scoring runs after the session rather than inline. It keeps interview latency low and means the rubric can be revised and re-run over past transcripts.

Stack snapshot

  • React
  • Node.js
  • Express
  • WebSockets
  • Speech-to-Text API
  • Text-to-Speech API
  • LLM APIs
  • MongoDB

Challenges & learnings

Engineering challenges
  • Latency budget across four hops

    Speech in, transcription, generation, speech out — each hop adds delay and the sum is what the candidate feels. Streaming at every stage instead of waiting for complete results is what made the conversation feel natural.

  • Keeping the interviewer on track

    Given full freedom, the model drifted into tangents or accepted vague answers. Constraining it to a competency plan while leaving follow-up phrasing free kept sessions comparable between candidates.

Key learnings
  • Conversation design is systems design

    The perceived intelligence of a voice agent comes from timing and turn-taking far more than from the quality of any single generated sentence.

  • Deterministic scaffolding around a non-deterministic core

    The reliable pattern: let the model handle language, and let ordinary code handle state, sequencing and the guarantees you actually need to hold.