AI Interview Platform
An asynchronous AI interviewer that holds a real technical conversation — streaming the candidate's speech, forming contextual follow-up questions from what they actually said, and producing a structured evaluation afterwards.
- Role
- Core AI Dev
- Mode
- Async Voice
- Pipeline
- STT → LLM → TTS
- Output
- Scored Report
The Friction
First-round technical screening is expensive, scheduling-bound, and inconsistent — two candidates rarely get the same interview. Existing automated tools ask fixed questions from a list, which candidates quickly learn to game and which cannot probe an interesting answer.
The Adaptive Screen
A context-aware interviewer that evaluates candidate speech as it arrives, decides whether an answer warrants a follow-up, and generates that follow-up from the specific claim the candidate made. Candidates authenticate by UUID link — no account creation — and the session records audio and video alongside the transcript.
System architecture

Hover or focus a stage to see why it’s built that way.
Turn-taking as a first-class problem
The hardest part was not generation — it was knowing when the candidate had finished speaking. Interrupting someone mid-thought destroys the experience, so silence detection is tuned deliberately conservative.
UUID links over accounts
Requiring a signup before an interview costs candidates. A signed, single-use UUID link removes the barrier while keeping each session isolated.
Evaluation decoupled from the interview
Scoring runs after the session rather than inline. It keeps interview latency low and means the rubric can be revised and re-run over past transcripts.
Stack snapshot
- React
- Node.js
- Express
- WebSockets
- Speech-to-Text API
- Text-to-Speech API
- LLM APIs
- MongoDB
Challenges & learnings
Latency budget across four hops
Speech in, transcription, generation, speech out — each hop adds delay and the sum is what the candidate feels. Streaming at every stage instead of waiting for complete results is what made the conversation feel natural.
Keeping the interviewer on track
Given full freedom, the model drifted into tangents or accepted vague answers. Constraining it to a competency plan while leaving follow-up phrasing free kept sessions comparable between candidates.
Conversation design is systems design
The perceived intelligence of a voice agent comes from timing and turn-taking far more than from the quality of any single generated sentence.
Deterministic scaffolding around a non-deterministic core
The reliable pattern: let the model handle language, and let ordinary code handle state, sequencing and the guarantees you actually need to hold.