Interview preparation
12 generative AI interview questions: RAG, agents and evaluation
Prepare with 12 practical generative AI interview questions and answer guides covering RAG, agents, evaluation, security and system design.
By TruTalent · Updated · 5 min read
For: AI engineers, analysts and product candidates

The short answer
What generative AI interview questions should I prepare for?
Prepare to explain a system’s behaviour, not just define AI terms. Useful practice topics include retrieval versus fine-tuning, evaluation, tool permissions, latency, cost and failure recovery. For each answer, connect the concept to a project you can inspect and discuss honestly. The questions below are original practice prompts, not a claim about any employer’s interview bank.
What to take away
- Explain the tradeoff and a concrete example in each answer.
- Distinguish a working demo from a tested production system.
- Use your own results; never invent experience or measurements.
1. What is retrieval-augmented generation?
RAG retrieves relevant information and supplies it as context for generation. It is useful when an answer needs a specific collection of documents or changing external knowledge. Describe the stages—ingestion, retrieval, context assembly and answer generation—and where errors can occur. Google Cloud provides a technical overview.
Follow-up to practise: how would you tell whether the wrong answer came from missing evidence or from the model misusing good evidence?
2. When would you choose RAG instead of fine-tuning?
Use retrieval when the main need is access to external knowledge with source traceability. Fine-tuning changes model behaviour through training and may help with patterns or task behaviour. They can be combined. First define the problem and compare against a baseline; neither approach automatically fixes inaccurate answers.
3. How would you choose a document-chunking strategy?
Start with document structure and the kinds of questions users ask. Preserve headings and source metadata; avoid splitting information that needs to be read together. Compare candidate approaches on representative questions rather than declaring one chunk size universally best.
A good answer explains the tradeoff between enough context and irrelevant material, and how retrieval and final-answer quality would be checked separately.
4. How do you evaluate a RAG application?
Define representative questions, expected evidence and a scoring rubric. Test retrieval separately from the final answer. Include unsupported, ambiguous and permission-restricted questions. Google Cloud’s evaluation guidance is useful background.
State sample size, selection method and failure categories. Hold back some cases during development, and compare changes with a baseline. If a model grades answers, validate its judgments against human review.
5. How would you reduce unsupported answers?
Improve evidence quality, make source boundaries clear, evaluate retrieval and teach the application to abstain when evidence is missing. Check that citations actually support the answer. Avoid promising that a prompt or RAG eliminates hallucinations.
Example: when two policy versions conflict, a good system should use explicit version rules or ask for review, rather than silently combining them.
6. When does a task need an agent?
Consider an agent when the next step depends on intermediate results and a model’s dynamic choice adds value. Prefer a fixed workflow when the sequence is predictable and easier to verify. Anthropic’s workflow/agent distinction is useful background. Explain the extra cost and failure paths introduced by dynamic tool use.
7. How should tool permissions work?
Enforce permissions in the application and underlying service. Validate arguments, restrict available actions and require approval for consequential changes where appropriate. A model’s instruction to behave safely is not an authorisation boundary.
Be ready to explain what happens if retrieved material tries to redirect the agent or if a user requests another person’s information. OWASP’s risk guidance provides relevant security context.
Sources: OWASP — Top 10 for LLM applications
8. How do you handle a failing tool or repeated action?
Set timeouts and bounded retries. Record task state, distinguish retryable errors from permanent failures and design operations to avoid duplicate effects. Stop and escalate when progress is uncertain. In a portfolio, demonstrate the failure path rather than merely listing these controls.
9. How do agent evaluations differ from answer checks?
An agent may produce an attractive explanation while leaving the wrong final state. Evaluate the completed task, intermediate actions, permission use and resource consumption. Anthropic’s agent-evaluation article discusses why traces and grading design matter.
For a support agent, checking the final draft is only one part: also inspect whether it touched the correct record, used permitted tools and respected the approval step.
10. How would you reduce latency and cost?
Measure where time and spending go before changing the architecture. Consider reducing unnecessary context, caching where freshness and permissions allow it, choosing simpler paths for easy tasks and avoiding redundant calls. Re-run quality checks after each change.
Explain which tradeoff is acceptable for the user. A cheaper answer that requires repeated corrections may increase total workflow cost.
11. Design an internal knowledge assistant. Where do you start?
Clarify users, data sources, access rules, update frequency and acceptable failure modes. Sketch ingestion, retrieval, generation, evidence display and monitoring. Define how deleted or changed documents propagate and how the system responds when it cannot answer.
Discuss a baseline such as search plus a curated FAQ. Present evaluation and rollout as part of the design rather than adding them at the end.
12. Tell me about an AI project that failed.
Describe the goal, your role, the observed failure, the evidence and the change you made. Explain what remained unresolved. If the project was simulated or personal, say so. A useful answer demonstrates learning and ownership without inventing production scale.
Prepare a short walkthrough with one architecture sketch, one result table and one failure example. For a mock interview, have a peer challenge assumptions and ask you to reproduce a result. Use technical vocabulary only when you can explain what it means in your system.
Frequently asked questions
- Are these actual questions asked by a particular company?
- No. These are original practice questions based on common engineering concepts and the cited technical guidance. Interview content varies by employer, team and seniority.
- How should a fresher answer questions about production systems?
- Be explicit about what you have built and tested. Explain how a prototype would need to change for production, but do not imply that you have operated systems you have not worked on.
- Should I memorise these answers?
- Use them to structure understanding. Practise explaining a concrete example, tradeoff and failure case in your own words, then connect the answer to inspectable work.
Sources & further reading
Sources accessed 11 October 2026. The learning plans and practice projects are TruTalent’s editorial examples. Source dates and scopes are noted below.
- Google Cloud — Retrieval-augmented generation
Technical explanation of retrieval grounding for language models.
- Google Cloud — Evaluating retrieval in RAG systems
Engineering guidance on measuring retrieval and answer quality.
- Anthropic — Building effective agents
Engineering guidance on workflows, agents and practical design tradeoffs.
- Anthropic — Demystifying evals for AI agents
January 2026. Practical evaluation concepts; not a hiring survey.
- OWASP — Top 10 for LLM applications
Security risk taxonomy for generative AI applications.