Nettms · Building TomorrowAcademy of AI &
Engineering
Jobs
Free Learnings
For companiesTeach
Sign inStart free →
Interview Questions

Applied AI & Gen AI Interview Questions & Answers

Common Applied AI and Generative AI interview questions — LLMs, RAG, agents, embeddings and evaluation — with clear answers. Practise them in a free AI mock interview.

Practise these in a free AI mock interview
  1. 1. What is RAG (Retrieval-Augmented Generation)?

    RAG retrieves relevant documents from a knowledge base (via embeddings + a vector database) and feeds them into the LLM prompt so it answers from your data. It reduces hallucination and uses fresh or private information without retraining.

  2. 2. RAG vs fine-tuning — when to use which?

    Use RAG when the model needs current or proprietary facts and easy updates. Fine-tune to change style, format or task behaviour that examples can teach. Most systems try RAG first (cheaper, updatable) and fine-tune only when needed.

  3. 3. What is an embedding?

    A numeric vector that captures the meaning of text (or images) so similar items sit close together in vector space. Embeddings power semantic search and RAG retrieval.

  4. 4. How do you reduce hallucinations in an LLM app?

    Ground answers with RAG, tell the model to say "I don't know" when unsure, lower temperature for factual tasks, add citations, and validate outputs with schemas/checks or an evaluator. Test on a fixed eval set before shipping.

  5. 5. What is the difference between a token and a word?

    LLMs process text as tokens — sub-word chunks, so a common word may be one token and a rare word several. Cost and context limits are measured in tokens, not words.

  6. 6. What is prompt engineering?

    Designing the instructions and context you give an LLM to get reliable output — being specific, giving examples (few-shot), setting a role, requesting a format like JSON, and breaking complex tasks into steps.

  7. 7. What is an AI agent?

    An LLM that decides and takes actions using tools (search, code, APIs) in a loop to reach a goal, rather than just returning text. Agents plan, call tools, observe results and iterate.

  8. 8. How do you evaluate an LLM feature before shipping?

    Build a representative test set, define metrics (accuracy, relevance, format-correctness, safety), run automated evals including LLM-as-judge, review edge cases manually, and monitor in production. Never ship on a single good demo.

  9. 9. How do you control LLM costs in production?

    Use smaller/cheaper models where they suffice, cache repeated calls, trim prompt and context length, cap max tokens, batch requests, and add per-user rate limits.

  10. 10. What is a vector database and why use one?

    A database optimised to store embeddings and find the nearest (most similar) vectors quickly. It is the retrieval backbone of RAG and semantic search.

  11. 11. What is temperature in an LLM?

    A setting that controls randomness. A low temperature near 0 gives focused, repeatable answers — good for factual tasks. A higher temperature gives more varied, creative output.

  12. 12. What is a context window?

    The maximum amount of text, measured in tokens, a model can consider at once — prompt plus response. Go over it and older content is dropped, so long inputs need summarising or retrieval.

  13. 13. What is the difference between zero-shot and few-shot prompting?

    Zero-shot gives only instructions; few-shot adds a few worked examples so the model infers the pattern. Examples improve reliability on formatting and tricky edge cases.

  14. 14. What is chain-of-thought prompting?

    Asking the model to reason step by step before answering. It improves accuracy on multi-step problems, though in production you often keep only the final answer and hide the reasoning.

  15. 15. What is tool or function calling?

    Letting the model output a structured request to call a function or API with arguments, which your code runs and returns. It is how agents take real actions like searching or querying a database.

  16. 16. What is prompt injection and how do you defend against it?

    When untrusted input tricks the model into ignoring its instructions, for example "ignore previous instructions". Defend by separating system instructions from user data, validating inputs, limiting tool permissions, and checking outputs.

  17. 17. How do you keep answers grounded and citable?

    Use RAG to supply source passages, instruct the model to answer only from them and cite the source, and have it say it does not know when the passages do not contain the answer.

  18. 18. What is LLM-as-a-judge?

    Using an LLM to score or compare other outputs against criteria. It scales evaluation cheaply but can be biased or inconsistent, so you validate it against human labels.

  19. 19. What is fine-tuning and when is it worth it?

    Training a base model further on your own examples to change its style or behaviour. It is worth it when prompting and RAG cannot reach the consistency you need, you have quality data, and the task is stable.

  20. 20. What is quantization?

    Reducing the numeric precision of a model's weights, for example from 16-bit to 4-bit, to shrink memory and speed up inference with a small accuracy trade-off. It lets larger models run on smaller hardware.

  21. 21. How do you reduce latency in an LLM app?

    Stream the response, use a smaller or faster model where it suffices, cache, run retrieval in parallel, keep prompts short, and show partial output so the app feels responsive.

  22. 22. What is overfitting?

    When a model learns the training data too closely, including its noise, and then performs poorly on new data. Prevent it with more data, regularisation, simpler models and proper validation.

  23. 23. What is the difference between classification and regression?

    Classification predicts a category (spam or not spam); regression predicts a continuous number (a price). The choice shapes the model, the loss function and the metrics.

  24. 24. How would you build a chatbot over a company's documents?

    Ingest and chunk the documents, embed them into a vector store, retrieve the most relevant chunks per question (RAG), prompt the LLM to answer from them with citations, add guardrails, and evaluate on real questions.

  25. 25. What are system, user and assistant messages?

    The system message sets the rules and role, user messages are the human input, and assistant messages are the model's replies. Keeping instructions in the system message and data in user messages improves control and safety.

Knowing the answers isn’t enough — say them out loud

Practise these Applied AI questions in a free AI mock interview: answer by voice, get instant feedback on your strengths and the gaps to fix.

Start a free mock interviewBuild a free resume

Admissions open · free to apply

Request your admission

Attended a masterclass or have a friend's referral code? You get ₹15,000 off. Fill this and our team takes it from here — pay by cash or online.

Nettms · Building TomorrowAcademy of AI &
Engineering

Free masterclasses, live cohorts, and pan-India placement support. Building Tomorrow.

ISO CertifiedStartup IndiaPractitioner-taught

Programs

  • Applied AI Engineer
  • Data Analysis with Gen AI
  • BIM
  • All programs

Free Tools

  • Code Compiler
  • Python Playground
  • Pandas Playground
  • SQL Playground
  • AI Glossary
  • Success Stories

Free Learning

  • Free Masterclass
  • Free Learnings
  • Free Admission
  • Blog
  • Events
  • WhatsApp Channel
  • Newsletter

Workshops

  • AI Workshop
  • BIM Workshop
  • Refer & Earn

Company

  • About
  • Hire with us
  • Careers at Nettms
  • Become a trainer
  • Contact
  • Privacy
  • Terms
  • Refund policy

Get the weekly drop.

One email a week — career insights, free masterclass invites, and what India's top employers are hiring for.

© 2026 Nettms Urban Habitat Pvt. Ltd. · 🌱 Building Tomorrow

Hyderabad, India

  • Home
  • Programs
  • Practice
  • Free
  • Account