Ticker

6/recent/ticker-posts

Gen AI interview questions and answers

 .  An enterprise chat assistant you built for a Fortune 500 client gets less accurate the longer a conversation runs. How would you find the cause and fix it?

  A N S W E R G UI D E   

CORE ANSWER

Accuracy that decays with conversation length is rarely the model 'getting tired'. Almost always it is a context-management problem: the prompt keeps growing, important facts get diluted or silently truncated, early instructions lose influence, and contradictory turns pile up. The strong answer is to prove this with data first, then redesign how context is assembled instead of just buying a bigger window.

STEP-BY-STEP APPROACH

1.      Measure the decay: Build an evaluation set of long, realistic conversations and plot accuracy against turn number and prompt token count. A curve that bends after a certain token count points straight at context handling.

2.      Inspect the real prompt: Log the fully assembled prompt for failing turns. Check whether the system prompt, policies or key user facts are being truncated, and whether retrieved documents are crowding out conversation history.

3.      Test position effects: Plant a fact early, ask about it later, and vary where it sits in the prompt. Models attend less reliably to the middle of long contexts ('lost in the middle'), so placement matters.

4.      Redesign memory in layers: Keep the last few turns verbatim, compress older turns into a rolling summary, and extract durable facts (names, IDs, decisions, constraints) into a structured store that is re-injected on every turn.

5.      Pin what must never be lost: Keep the system prompt and critical rules at the top, restate key constraints close to the current question, and allocate a token budget per prompt section.

6.      Retrieve fresh context per turn: Re-run retrieval on the current question rather than carrying stale chunks from ten turns ago.

TRADE-OFFS TO MENTION

•     Summaries are cheaper than raw history but can drop specifics; protect numbers and entities explicitly in the summarization prompt.

•     A larger context window is a quick patch, yet it raises cost and latency and does not cure dilution.

•     Fact extraction adds an extra LLM call; run it asynchronously after the response so it does not hurt latency.

METRICS TO TRACK

Accuracy bucketed by turn number  •  Contradiction rate against earlier turns  •  Prompt tokens per request  •  p95 latency  •  Re-ask and thumbs-down rate in long sessions

ILLUSTRATIVE EXAMPLE

An HR policy assistant dropped from roughly 90% accuracy at turn 3 to about 70% by turn 20. Logs showed the history-truncation logic was cutting the oldest messages first, including the policy instructions. Pinning the system prompt, keeping a six-turn window plus a rolling summary and a small fact store flattened the accuracy curve and cut prompt tokens by around 40%.

ANSWERING THE FOLLOW-UPS

▸ Is a bigger context window the answer, or does the memory design need rework?

Redesign memory first. A bigger window postpones truncation but does not fix dilution, and it costs more on every call. Use a larger window as a complement for genuinely long documents, not as the memory strategy.

▸ How would you prove the fix actually worked?

Run the same long-conversation eval set before and after and confirm accuracy stays flat across turn buckets. Then A/B test online and watch thumbs-down, re-asks and escalations specifically for long sessions.

▸ Which trade-offs would you weigh?

Cost and latency versus recall, summary fidelity versus compression, the privacy implications of storing extracted facts, and the extra engineering complexity of a multi-layer memory system.


Post a Comment

0 Comments