Interactive AI Learning System

AI Master Map

From the foundations of Artificial Intelligence to Machine Learning, Deep Learning, Transformers, LLMs, RAG, Agents, Multimodal AI and production AI engineering.

AlgorithmsLibraries & PackagesModels TechniquesPipelinesInterview Prep TechHero Case Study

01 · The Complete AI Hierarchy

Think of AI as the large umbrella. The branches below are related disciplines, not always strict parent-child categories.

AI

Artificial Intelligence

Machines performing tasks that normally require human intelligence: reasoning, perception, planning, learning and decision-making.

Example: navigation, recommendation, chess, conversational assistants.

ML

Machine Learning

A branch of AI where systems learn patterns from data instead of being given every rule explicitly.

Example: spam classification.

DL

Deep Learning

Machine learning using multi-layer neural networks that learn representations from data.

Example: image recognition.

GenAI

Generative AI

Models that generate new content such as text, images, audio, video or code.

Example: an LLM writing Python code.

Foundation Models

Foundation Models

Large pretrained models adaptable to many downstream tasks.

Example: language, vision-language and multimodal models.

Agentic AI

AI Agents

Systems that combine models with tools, state, planning and feedback to accomplish goals.

Example: an agent checking weather, querying a database and sending a message.

Also learn: symbolic AI, search/planning, expert systems, fuzzy logic, evolutionary algorithms, probabilistic models, knowledge graphs, robotics, computer vision, speech, recommender systems and causal AI.

02 · AI History → Modern AI

1940s–1950s

Foundations

Formal logic, early neural models, computation and the question of machine intelligence.

1956

Dartmouth Workshop

The term Artificial Intelligence became established as a research field.

1960s–1980s

Symbolic AI & Expert Systems

Rules, logic, search, knowledge representation and hand-built expert systems dominated.

AI Winters

Expectation vs Reality

Funding and interest dropped when systems failed to meet ambitious promises and computing/data were limited.

1990s–2000s

Statistical Machine Learning

Probabilistic methods, decision trees, SVMs, ensembles and data-driven prediction became central.

2012+

Deep Learning Era

GPUs, large datasets and neural networks produced major advances in vision, speech and language.

2017

Transformer Architecture

Self-attention became the key architecture behind modern language and multimodal foundation models.

2020s → Today

Foundation Models, RAG & Agents

Large models, retrieval, tool use, multimodality and agentic workflows moved AI into application engineering.

03 · Mathematics You Need

Linear Algebra

Vectors & Matrices

Core representation for features, embeddings, neural weights and transformations.

Learn: vectors, matrices, dot product, norms, matrix multiplication, eigen concepts.

Probability

Uncertainty

Models likelihood and uncertainty.

Learn: conditional probability, Bayes theorem, distributions, expectation, variance.

Statistics

Data Understanding

Learn sampling, correlation, hypothesis testing, confidence intervals and distributions.

Calculus

Optimization

Derivatives and gradients tell neural networks how to change parameters to reduce loss.

Optimization

Minimize the Loss

Gradient descent, SGD, momentum, Adam, learning rate schedules and regularization.

04 · Python AI Ecosystem

Data

NumPy arrays and numerical computing.

Pandas / Polars tabular data.

Matplotlib / Plotly visualization.

Classical ML

scikit-learn preprocessing, algorithms, pipelines and metrics.

XGBoost, LightGBM, CatBoost for gradient boosting.

Deep Learning

PyTorch, TensorFlow, Keras, JAX.

Language AI

transformers, datasets, tokenizers, sentence-transformers.

Vector Search

Qdrant, FAISS, Milvus, Weaviate, pgvector.

Web AI

Django, FastAPI, Flask, Pydantic.

05 · Machine Learning

The practical pipeline is more important than memorizing algorithms.

Problem→Collect Data→Clean→Features→Train→Validate→Test→Deploy
Supervised

Learn from Labels

Input + known target.

Algorithms: linear/logistic regression, trees, random forest, SVM, KNN, boosting.

Examples: spam detection, price prediction.

Unsupervised

Find Structure

No target labels.

Algorithms: K-means, hierarchical clustering, PCA, anomaly detection.

Self-Supervised

Generate Training Signals

The data creates its own labels.

Example: predict a masked or next token.

Reinforcement Learning

Learn by Rewards

An agent acts in an environment and learns from rewards.

Learn: MDP, policy, value, Q-learning, DQN, policy gradient, actor-critic, PPO.

Ensembles

Combine Models

Bagging reduces variance; boosting builds models sequentially; stacking combines different learners.

Tuning

Hyperparameters

Grid search, random search and Bayesian optimization select settings such as depth, learning rate and regularization.

Important ML Algorithms

AlgorithmUseCore idea
Linear RegressionRegressionFit a linear relationship between features and target.
Logistic RegressionClassificationPredict class probability.
Decision TreeClassification/RegressionSplit data using feature rules.
Random ForestClassification/RegressionEnsemble of decision trees.
Gradient BoostingClassification/RegressionSequentially correct previous errors.
SVMClassificationFind a separating boundary with maximum margin.
KNNClassification/RegressionUse nearby examples.
K-MeansClusteringGroup points around learned centroids.
PCADimensionality reductionFind lower-dimensional directions preserving variance.
Naive BayesClassificationBayesian classification with conditional-independence assumption.
Isolation ForestAnomaly detectionOutliers are easier to isolate using random splits.

06 · Deep Learning

Neural Network

Neuron → Layer → Network

Weights transform inputs. Activations introduce non-linearity.

Learn: perceptron, MLP, forward pass, loss, backpropagation, gradient descent.

Activations

ReLU, Sigmoid, Tanh, GELU, Softmax

Control non-linear transformations or probabilities.

CNN

Convolutional Networks

Learn spatial patterns using filters.

Use: image classification, detection, segmentation.

RNN/LSTM/GRU

Sequence Models

Process sequential information. Important historically for language before Transformers.

Autoencoder/VAE

Representation & Generation

Encode data into latent representations and reconstruct or generate samples.

GAN

Generative Adversarial Network

Generator creates samples while discriminator tries to distinguish real from generated.

Diffusion

Modern Generative Vision

Learn to reverse a gradual noising process to generate samples.

Learn: noise schedule, denoising, U-Net, conditioning, latent diffusion.

ViT

Vision Transformer

Images are divided into patches and processed with Transformer-style attention.

Deep-learning concepts: epochs, batches, learning rate, initialization, normalization, dropout, L1/L2 regularization, batch norm, layer norm, residual/skip connections, gradient clipping, early stopping and transfer learning.

07 · NLP — Natural Language Processing

Text→Clean→Tokenize→Represent→Model→Task

Text preprocessing

Normalization, punctuation handling, tokenization, stemming, lemmatization and stop-word decisions.

TF-IDF

Weights terms by how important they are to a document relative to a collection.

Use: classic search and text classification.

Word Embeddings

Represent words as vectors.

Models: Word2Vec, GloVe and contextual embeddings.

NER

Named Entity Recognition identifies entities such as people, organizations, locations and dates.

Classic NLP tasks

Sentiment, classification, summarization, translation, question answering, information extraction and search.

Modern NLP

Transformers replaced many task-specific sequence architectures with pretrained representations and fine-tuning/instruction methods.

08 · Transformers

Core idea: instead of processing tokens only sequentially, self-attention lets each token dynamically use information from other relevant tokens.
Tokens→Embeddings→Position→Self-Attention→FFN→Output

Q, K, V

Query asks what this token needs. Key describes what another token offers. Value is the information aggregated.

Multi-Head Attention

Multiple attention heads can learn different relationships simultaneously.

Encoder

Builds contextual representations using bidirectional attention. Example family: BERT-style models.

Decoder

Generates autoregressively using causal masking. Example family: GPT-style models.

Encoder–Decoder

Useful for sequence-to-sequence tasks such as translation and summarization.

Modern internals

Residual connections, LayerNorm/RMSNorm, positional methods such as RoPE, causal masks, KV cache and sometimes Mixture-of-Experts.

09 · LLMs & Generative AI

Tokenizer

Text → Token IDs

Subword tokenization reduces vocabulary problems.

Learn: BPE, WordPiece, SentencePiece and token limits.

Pretraining

Next-token Prediction

The model learns statistical structure from huge datasets by predicting missing/next tokens.

Instruction Tuning

Follow Tasks

SFT teaches a pretrained model to follow examples and instructions.

Alignment

Preference Optimization

RLHF and methods such as DPO train models toward preferred behavior.

Fine-tuning

Adapt the Model

Full fine-tuning changes many weights; PEFT/LoRA updates a smaller set of parameters.

Quantization

Smaller/Faster Inference

Reduce numerical precision such as FP16 → INT8/INT4. Common formats/tools include GGUF, GPTQ and AWQ.

Inference controls

Temperature, top-k, top-p, max tokens, stop sequences and structured output affect generation.

Performance

KV caching, batching, speculative decoding, efficient serving and model distillation improve cost/latency.

MoE

Mixture-of-Experts routes different tokens through selected expert networks instead of activating every parameter.

Prompt structure: role/instructions → task → context → constraints → examples → required output format. Learn zero-shot, few-shot, decomposition, self-consistency, structured output and tool/function calling.

10 · RAG — Retrieval-Augmented Generation

RAG connects an LLM to external knowledge at inference time instead of requiring the model to memorize every answer.

Documents→Parse→Chunk→Embed→Qdrant→Retrieve Top-K→Rerank→LLM→Answer + Citations

1. Ingestion

Read PDF, DOCX, HTML, TXT or database records. Tools can include pypdf, python-docx or specialized parsers.

2. Chunking

Split documents into useful retrieval units.

Strategies: fixed-size, recursive, semantic, structure-aware, parent-child.

Use overlap carefully; too little loses context, too much adds redundancy.

3. Embeddings

Convert text into vectors representing semantic meaning.

Your stack: sentence-transformers → 384-dimensional vectors.

4. Vector DB

Store vectors + metadata for fast similarity search.

Your project: Qdrant collection telecom_knowledge.

5. Retrieval

Query embedding → similarity search → top-k chunks.

Similarity: cosine similarity is a common measure.

6. Reranking

A stronger model scores retrieved candidates for query relevance. Cross-encoders are a common approach.

7. Hybrid Search

Combine dense semantic retrieval with sparse keyword retrieval such as BM25.

Advanced Retrieval

Query rewriting, multi-query, HyDE, contextual compression, metadata filters and graph RAG.

RAG Evaluation

MetricWhat it tells you
Recall@K / Hit RateDid retrieval find the relevant chunk?
Precision@KHow much of the retrieved set is relevant?
MRR / nDCGHow highly were useful results ranked?
Context Precision/RecallIs the retrieved context useful and sufficiently complete?
Faithfulness / GroundednessDoes the answer actually follow the retrieved evidence?
Answer RelevanceDoes the answer address the user's question?
RAG vs Fine-tuning: RAG is usually better when knowledge changes frequently or needs citations. Fine-tuning is mainly for behavior, style, task specialization or learned patterns. They can be combined.

11 · AI Agents

Goal→Reason/Plan→Choose Tool→Tool Result→Observe→Next Action→Final

Workflow vs Agent

A workflow follows mostly predefined steps. An agent dynamically decides which tools/actions to use based on state and goals.

Tool Calling

The model selects a structured function such as weather lookup, database query, calculator or email sender.

ReAct

A common conceptual pattern: reason about the next action, act with a tool, observe the result, then continue.

State & Memory

Short-term conversation state plus optional long-term semantic, episodic or procedural memory.

Agent Patterns

Planner-executor, supervisor-worker, sequential agents, peer-to-peer and graph/state-machine workflows.

MCP

Model Context Protocol is a standard approach for connecting AI systems with tools and resources through defined interfaces.

Guardrails

Permissions, schemas, validation, human approval, tool restrictions, timeouts and sandboxing prevent dangerous actions.

Multi-Agent

Different specialized agents can collaborate, for example Weather Agent → Decision Agent → Communication Agent.

12 · Multimodal AI & Other AI Branches

Computer Vision

Classification, object detection, segmentation, OCR, image embeddings, tracking and visual question answering.

Tools: OpenCV, PyTorch, Transformers.

Speech AI

Speech-to-text, text-to-speech, speaker identification and audio classification.

Vision-Language

Models combine image and language representations for captioning, visual QA and image understanding.

Multimodal RAG

Retrieve text, images, tables or other modalities before generating an answer.

Generative Images

Diffusion-based text-to-image, image-to-image, inpainting and conditioning techniques.

Knowledge Graphs

Represent entities and relationships as nodes and edges. Useful for reasoning, entity linking and graph RAG.

Recommender Systems

Content-based filtering, collaborative filtering, matrix factorization and neural recommenders.

Causal AI

Studies cause-and-effect rather than only correlation. Learn interventions, confounding, DAGs and counterfactual reasoning.

13 · AI Engineering & Production

Application Architecture

UI → API → authentication → business logic → AI orchestration → model/RAG/agent → database/tools.

Serving

Learn REST, WebSocket, async processing, queues, streaming responses, batching, retries and rate limiting.

Local Inference

Learn CPU/GPU inference, Apple Silicon/MPS basics and local serving tools such as Ollama, llama.cpp or vLLM.

Observability

Trace prompts, retrieval, tool calls, latency, token usage, failures and user feedback.

Evaluation-Driven Development

Create golden datasets, regression tests and automated evaluation before changing prompts, retrievers or models.

Cost & Latency

Use caching, smaller models, routing, batching, quantization and retrieval limits to control cost.

Data Engineering

ETL/ELT, data validation, labeling, data versioning, feature pipelines and data quality monitoring.

Deployment

Docker, CI/CD, environment management and optional cloud/container orchestration are production skills.

14 · Responsible AI & Security

Hallucination

Generated content can be unsupported or incorrect. Mitigate with retrieval, grounding, validation and explicit uncertainty.

Prompt Injection

Untrusted text attempts to manipulate model instructions. Treat retrieved/user content as data, not authority.

Tool Security

Use least privilege, allowlists, schemas, sandboxing, approval gates and safe defaults.

Data Privacy

Minimize sensitive data, control access, redact where needed and understand retention policies.

Bias & Fairness

Measure model behavior across relevant groups and understand limitations of training data.

Supply Chain

Validate models, packages, datasets and dependencies before using them in production.

15 · Your TechHero Project — Where Everything Connects

Use this as your portfolio case study. It demonstrates normal web development plus real AI engineering.

Django→MySQL Ticket→Knowledge Docs→Chunking→Sentence Transformers→Qdrant→Retriever→LLM→Technician Answer

Web Engineering

Django + MySQL handles authentication, technicians, tickets and business workflows.

RAG

Troubleshooting documents are chunked, embedded and stored in Qdrant for semantic retrieval.

Embeddings

Your current sentence-transformers setup produces 384-dimensional embeddings.

LLM Generation

The retrieved context is supplied to an LLM with instructions to answer from the evidence.

Feedback Loop

Technician feedback can become reviewed knowledge, creating an evaluation and knowledge-improvement loop.

Next Upgrades

Hybrid search → reranking → query rewriting → citations → RAG evaluation → tool-calling agent → multimodal support.

20 · AI Frameworks & Orchestration

Frameworks do not replace the underlying AI concepts. They help you build, connect, observe and control AI pipelines.

LangChain

LLM Application Framework

Provides building blocks for prompts, models, retrievers, tools, document loaders, structured output and agent workflows.

Learn: models → prompts → parsers → retrievers → tools → agents → memory/state.

LangGraph

Stateful Agent Workflows

Build graph-based, stateful AI workflows where nodes perform work and edges control what happens next.

Useful for: multi-step agents, human approval, retries, loops, branching and durable state.

LlamaIndex

Data + RAG Framework

Focuses strongly on connecting LLMs to private data, indexes, retrievers and knowledge sources.

Haystack

Search & RAG Pipelines

Framework for retrieval, document processing, pipelines and question-answering systems.

DSPy

Programmatic LLM Optimization

Treat prompts and LLM calls as programmable components and optimize them against evaluation data.

Plain Python

Do not assume a framework is mandatory. For your TechHero project, implementing RAG directly with Python, Qdrant and an LLM is excellent for understanding the fundamentals.

Interview point: Explain the architecture first and the framework second. “LangChain” is not RAG itself; “LangGraph” is not an agent by itself. They are engineering tools used to implement these patterns.

21 · Embeddings, Vector Search & Vector Databases

Text→Embedding Model→Vector→Index→Similarity Search

Embedding

An embedding converts an item such as text, image or code into a numerical vector representing useful relationships.

Important: embedding dimensions depend on the model.

Dense Embeddings

Most vector dimensions contain learned continuous values. Sentence Transformers is a common Python ecosystem for semantic text embeddings.

Sparse Embeddings

Representations emphasize specific terms/features. Traditional TF-IDF and BM25 are sparse retrieval approaches.

Cosine Similarity

Measures the angle between vectors. Frequently used when comparing semantic embeddings.

Dot Product

Measures vector alignment through multiplication and summation. Some embedding systems are optimized for dot-product retrieval.

Euclidean Distance

Measures straight-line distance between vectors. Whether it is appropriate depends on how the embedding model was trained.

HNSW

Hierarchical Navigable Small World graphs provide fast approximate nearest-neighbor search.

IVF

Inverted File indexes partition vectors into clusters and search selected regions rather than the entire collection.

PQ

Product Quantization compresses vectors into smaller representations to reduce memory and speed search.

Metadata Filtering

Combine vector similarity with fields such as product, location, language, date, ticket type or permissions.

Vector Database Comparison

TechnologyTypical strengthLearn
QdrantVector search + payload filtering + modern RAG applicationsCollections, points, payloads, filters, HNSW
FAISSLocal/library-level similarity searchIndexes, ANN, clustering, vector distance
pgvectorVector search inside PostgreSQLSQL + vectors + relational metadata
MilvusLarge-scale vector infrastructureCollections, indexes, distributed retrieval
WeaviateVector database with application-oriented featuresObjects, vectors, filters, retrieval

22 · Chunking — A Critical RAG Skill

Fixed-size

Split text after a target number of characters or tokens.

Good: simple baseline. Risk: may split meaning.

Recursive

Try separators such as paragraphs, lines and sentences before forcing smaller pieces.

A strong general-purpose baseline.

Semantic

Split where the meaning changes rather than using only length.

Can improve retrieval but costs more processing.

Structure-aware

Respect headings, sections, tables, code blocks, FAQ question/answer boundaries and document structure.

Parent-child

Retrieve a small child chunk but return a larger parent section to the LLM for additional context.

Chunk overlap

Repeat a small amount between neighboring chunks so information at boundaries is not lost.

Metadata

Store document ID, title, section, page, source, timestamp, category and access rules alongside each chunk.

Chunk quality test

Ask real questions and inspect whether the correct answer is contained in one or more retrieved chunks. Evaluate retrieval rather than guessing a universal chunk size.

23 · Advanced Retrieval Techniques

Dense Retrieval

Query and documents become embeddings and are compared semantically.

BM25

A strong lexical retrieval algorithm that rewards useful term matches and handles document length.

Hybrid Search

Combine sparse keyword retrieval and dense semantic retrieval for better coverage.

Reranking

Retrieve a broader candidate set, then use a stronger relevance model to reorder candidates.

Cross-Encoder

Reads the query and candidate document together to estimate relevance. More accurate but usually slower than bi-encoder retrieval.

Query Rewriting

Transform an unclear user question into a retrieval-friendly query.

Multi-Query Retrieval

Generate several search formulations to improve recall.

HyDE

Generate a hypothetical answer/document and use its embedding to search for real supporting documents.

Contextual Compression

Retrieve documents, then reduce them to the portions relevant to the current question.

Graph RAG

Use entities and relationships in a graph to retrieve connected knowledge, especially for multi-hop questions.

24 · Agent Engineering in Detail

Tool Schema

Define a tool name, description and typed parameters so the model can request it reliably.

Tool Router

Decide which available tool should receive a request.

Planner-Executor

A planner creates a strategy while an executor performs individual steps.

Supervisor Pattern

A supervisor routes tasks to specialized workers such as search, database, coding or communication agents.

Human-in-the-loop

Pause before risky actions and request approval from a person.

Memory Types

Short-term: current state. Semantic: facts. Episodic: past events. Procedural: learned procedures/workflows.

State Machine / Graph

Represent agent states and transitions explicitly. This makes complex workflows easier to debug and control.

Retries & Recovery

Handle tool failures, timeouts, malformed outputs and partial completion without blindly repeating dangerous actions.

Agent Evaluation

Measure tool selection, task completion, correctness, number of steps, latency, cost and unsafe-action rate.

MCP Architecture

Learn the concepts of MCP hosts, clients, servers, tools and resources, and why standardized tool/resource interfaces can simplify integrations.

25 · Model Families You Should Recognize

Family / ArchitectureWhat to understandTypical use
Linear / LogisticSimple statistical baselinesPrediction / classification
Tree / Forest / BoostingNonlinear tabular learningBusiness datasets
CNNConvolution and spatial featuresComputer vision
RNN / LSTM / GRURecurrent sequence modelingHistorical NLP/time series
Autoencoder / VAELatent representation and generationCompression / generation
GANGenerator vs discriminatorSynthetic media
DiffusionDenoising generative processImages/audio/video research
BERT-style EncoderBidirectional contextual representationClassification / extraction / embeddings
GPT-style DecoderCausal next-token generationLLMs / generation
Encoder-Decoder TransformerInput sequence → output sequenceTranslation / transformation
Vision TransformerImage patches + attentionVision
Multimodal Foundation ModelMultiple modalities in one systemText + image/audio/video
Mixture-of-ExpertsSparse expert activationEfficient large models

26 · Data Engineering for AI

Collect→Validate→Clean→Label→Transform→Version→Train/Retrieve→Monitor

Data Cleaning

Missing values, duplicates, invalid records, inconsistent formats and outliers.

Feature Engineering

Transform raw variables into useful model inputs.

Data Leakage

Information from validation/test/future data accidentally enters training, producing unrealistically good results.

Class Imbalance

One class is much more frequent. Learn class weights, thresholding and methods such as SMOTE with proper validation.

Data Labeling

Human annotation creates supervised targets. Learn quality checks, inter-annotator agreement and active learning.

ETL vs ELT

ETL transforms before loading; ELT loads first and transforms in the target data system.

27 · MLOps / LLMOps

Experiment Tracking

Record datasets, parameters, metrics, model versions and results.

Model Registry

Track approved model versions and lifecycle status.

Prompt Registry

Version prompts just like application code.

Evaluation Sets

Maintain stable test questions and expected behaviors to catch regressions.

Monitoring

Track quality, latency, cost, failures, drift, retrieval quality and user feedback.

Tracing

Follow an AI request through prompt → retrieval → reranking → tool calls → model → final answer.

CI/CD for AI

Automate unit tests, integration tests, evaluation tests and safe deployment.

Governance

Access control, audit logs, data retention, model approvals and reproducibility.

28 · AI Terms You Must Be Able to Explain

Parameter

A value learned by a model during training, such as a neural-network weight.

Hyperparameter

A configuration selected outside normal parameter learning, such as learning rate or tree depth.

Epoch

One complete pass through the training dataset.

Batch

A subset of training examples processed together.

Inference

Using a trained model to produce a prediction or generation.

Context Window

The amount of input/output token context a model can handle for a request.

Grounding

Constraining an answer to trusted external evidence.

Hallucination

Unsupported or incorrect generated content presented as if it were true.

Fine-tuning

Further training a pretrained model on task/domain data.

Transfer Learning

Reuse learned representations from one task/domain for another.

Zero-shot

Perform a task without task-specific examples in the prompt.

Few-shot

Provide examples in the prompt to demonstrate the desired behavior.

Structured Output

Require machine-readable output such as JSON matching a schema.

Function Calling

Allow the model to request a structured application function/tool.

Guardrail

A rule, validator or control that limits unsafe or invalid behavior.

Ground-truth

The reference answer/label used to evaluate a system.

16 · Recommended Learning Roadmap

Phase 1

Python Foundation

Syntax, functions, OOP, exceptions, modules, typing, virtual environments, files, JSON, HTTP, async basics.

Phase 2

Data + Math

NumPy, Pandas, visualization, linear algebra, probability, statistics and optimization basics.

Phase 3

Classical ML

Preprocessing, regression, classification, clustering, evaluation, feature engineering and tuning.

Phase 4

Deep Learning

PyTorch, tensors, training loops, CNNs, sequence models, normalization and regularization.

Phase 5

NLP + Transformers

Tokenization, embeddings, attention, BERT/GPT concepts, Hugging Face and inference.

Phase 6

LLM Engineering

Prompting, structured output, function calling, fine-tuning, PEFT, quantization and evaluation.

Phase 7

RAG

Chunking, embeddings, Qdrant, retrieval, hybrid search, reranking, evaluation and citations.

Phase 8

Agents

Tools, state, memory, workflows, ReAct, MCP, multi-agent patterns and guardrails.

Phase 9

Multimodal

Vision, OCR, speech, vision-language models and multimodal RAG.

Phase 10

Production AI

FastAPI/Django APIs, queues, observability, security, evaluation, deployment and cost optimization.

17 · AI Interview Cheat Sheet

What is an embedding?

A numerical vector representation of data such as text that lets us compare semantic relationships using vector similarity.

What is RAG?

A pattern where relevant external knowledge is retrieved and supplied to an LLM before generation.

RAG vs Fine-tuning?

RAG changes the knowledge available at inference time; fine-tuning changes model behavior/parameters.

Vector DB vs MySQL?

MySQL excels at relational records and transactions. Vector databases optimize similarity search over embeddings.

Why cosine similarity?

It compares the direction of vectors and is commonly used to measure semantic similarity between embeddings.

What is Top-K?

The number of highest-ranked retrieval results returned for a query.

Why reranking?

Initial retrieval is fast and broad; a reranker can more precisely order the candidate documents.

What is temperature?

A generation control that changes the distribution of next-token probabilities; higher values generally increase randomness.

Transformer vs RNN?

Transformers use attention and can process token relationships efficiently in parallel during training; RNNs process recurrent state sequentially.

Precision vs Recall?

Precision asks “of predicted positives, how many were correct?” Recall asks “of actual positives, how many did we find?”

Overfitting?

The model learns training-specific patterns too closely and performs poorly on unseen data.

Agent vs RAG?

RAG retrieves knowledge. An agent can decide actions and use tools iteratively to accomplish a goal. An agent can also use RAG.

18 · Your Practical Project Ladder

01

Python data analysis dashboard

02

Scikit-learn classification API

03

PyTorch neural-network project

04

NLP sentiment / NER project

05

Semantic search using embeddings

06

Production RAG API

07

TechHero: Django + MySQL + Qdrant + LLM

08

Tool-calling weather/communication agent

09

Multi-agent workflow

10

Multimodal RAG + production evaluation

19 · Learning Progress

Tick topics as you study. Progress is saved in your browser.

0%0 / 20 completed