01 · The Complete AI Hierarchy
Think of AI as the large umbrella. The branches below are related disciplines, not always strict parent-child categories.
Artificial Intelligence
Machines performing tasks that normally require human intelligence: reasoning, perception, planning, learning and decision-making.
Example: navigation, recommendation, chess, conversational assistants.
Machine Learning
A branch of AI where systems learn patterns from data instead of being given every rule explicitly.
Example: spam classification.
Deep Learning
Machine learning using multi-layer neural networks that learn representations from data.
Example: image recognition.
Generative AI
Models that generate new content such as text, images, audio, video or code.
Example: an LLM writing Python code.
Foundation Models
Large pretrained models adaptable to many downstream tasks.
Example: language, vision-language and multimodal models.
AI Agents
Systems that combine models with tools, state, planning and feedback to accomplish goals.
Example: an agent checking weather, querying a database and sending a message.
02 · AI History → Modern AI
Foundations
Formal logic, early neural models, computation and the question of machine intelligence.
Dartmouth Workshop
The term Artificial Intelligence became established as a research field.
Symbolic AI & Expert Systems
Rules, logic, search, knowledge representation and hand-built expert systems dominated.
Expectation vs Reality
Funding and interest dropped when systems failed to meet ambitious promises and computing/data were limited.
Statistical Machine Learning
Probabilistic methods, decision trees, SVMs, ensembles and data-driven prediction became central.
Deep Learning Era
GPUs, large datasets and neural networks produced major advances in vision, speech and language.
Transformer Architecture
Self-attention became the key architecture behind modern language and multimodal foundation models.
Foundation Models, RAG & Agents
Large models, retrieval, tool use, multimodality and agentic workflows moved AI into application engineering.
03 · Mathematics You Need
Vectors & Matrices
Core representation for features, embeddings, neural weights and transformations.
Learn: vectors, matrices, dot product, norms, matrix multiplication, eigen concepts.
Uncertainty
Models likelihood and uncertainty.
Learn: conditional probability, Bayes theorem, distributions, expectation, variance.
Data Understanding
Learn sampling, correlation, hypothesis testing, confidence intervals and distributions.
Optimization
Derivatives and gradients tell neural networks how to change parameters to reduce loss.
Minimize the Loss
Gradient descent, SGD, momentum, Adam, learning rate schedules and regularization.
04 · Python AI Ecosystem
Data
NumPy arrays and numerical computing.
Pandas / Polars tabular data.
Matplotlib / Plotly visualization.
Classical ML
scikit-learn preprocessing, algorithms, pipelines and metrics.
XGBoost, LightGBM, CatBoost for gradient boosting.
Deep Learning
PyTorch, TensorFlow, Keras, JAX.
Language AI
transformers, datasets, tokenizers, sentence-transformers.
Vector Search
Qdrant, FAISS, Milvus, Weaviate, pgvector.
Web AI
Django, FastAPI, Flask, Pydantic.
05 · Machine Learning
The practical pipeline is more important than memorizing algorithms.
Learn from Labels
Input + known target.
Algorithms: linear/logistic regression, trees, random forest, SVM, KNN, boosting.
Examples: spam detection, price prediction.
Find Structure
No target labels.
Algorithms: K-means, hierarchical clustering, PCA, anomaly detection.
Generate Training Signals
The data creates its own labels.
Example: predict a masked or next token.
Learn by Rewards
An agent acts in an environment and learns from rewards.
Learn: MDP, policy, value, Q-learning, DQN, policy gradient, actor-critic, PPO.
Combine Models
Bagging reduces variance; boosting builds models sequentially; stacking combines different learners.
Hyperparameters
Grid search, random search and Bayesian optimization select settings such as depth, learning rate and regularization.
Important ML Algorithms
| Algorithm | Use | Core idea |
|---|---|---|
| Linear Regression | Regression | Fit a linear relationship between features and target. |
| Logistic Regression | Classification | Predict class probability. |
| Decision Tree | Classification/Regression | Split data using feature rules. |
| Random Forest | Classification/Regression | Ensemble of decision trees. |
| Gradient Boosting | Classification/Regression | Sequentially correct previous errors. |
| SVM | Classification | Find a separating boundary with maximum margin. |
| KNN | Classification/Regression | Use nearby examples. |
| K-Means | Clustering | Group points around learned centroids. |
| PCA | Dimensionality reduction | Find lower-dimensional directions preserving variance. |
| Naive Bayes | Classification | Bayesian classification with conditional-independence assumption. |
| Isolation Forest | Anomaly detection | Outliers are easier to isolate using random splits. |
06 · Deep Learning
Neuron → Layer → Network
Weights transform inputs. Activations introduce non-linearity.
Learn: perceptron, MLP, forward pass, loss, backpropagation, gradient descent.
ReLU, Sigmoid, Tanh, GELU, Softmax
Control non-linear transformations or probabilities.
Convolutional Networks
Learn spatial patterns using filters.
Use: image classification, detection, segmentation.
Sequence Models
Process sequential information. Important historically for language before Transformers.
Representation & Generation
Encode data into latent representations and reconstruct or generate samples.
Generative Adversarial Network
Generator creates samples while discriminator tries to distinguish real from generated.
Modern Generative Vision
Learn to reverse a gradual noising process to generate samples.
Learn: noise schedule, denoising, U-Net, conditioning, latent diffusion.
Vision Transformer
Images are divided into patches and processed with Transformer-style attention.
07 · NLP — Natural Language Processing
Text preprocessing
Normalization, punctuation handling, tokenization, stemming, lemmatization and stop-word decisions.
TF-IDF
Weights terms by how important they are to a document relative to a collection.
Use: classic search and text classification.
Word Embeddings
Represent words as vectors.
Models: Word2Vec, GloVe and contextual embeddings.
NER
Named Entity Recognition identifies entities such as people, organizations, locations and dates.
Classic NLP tasks
Sentiment, classification, summarization, translation, question answering, information extraction and search.
Modern NLP
Transformers replaced many task-specific sequence architectures with pretrained representations and fine-tuning/instruction methods.
08 · Transformers
Q, K, V
Query asks what this token needs. Key describes what another token offers. Value is the information aggregated.
Multi-Head Attention
Multiple attention heads can learn different relationships simultaneously.
Encoder
Builds contextual representations using bidirectional attention. Example family: BERT-style models.
Decoder
Generates autoregressively using causal masking. Example family: GPT-style models.
Encoder–Decoder
Useful for sequence-to-sequence tasks such as translation and summarization.
Modern internals
Residual connections, LayerNorm/RMSNorm, positional methods such as RoPE, causal masks, KV cache and sometimes Mixture-of-Experts.
09 · LLMs & Generative AI
Text → Token IDs
Subword tokenization reduces vocabulary problems.
Learn: BPE, WordPiece, SentencePiece and token limits.
Next-token Prediction
The model learns statistical structure from huge datasets by predicting missing/next tokens.
Follow Tasks
SFT teaches a pretrained model to follow examples and instructions.
Preference Optimization
RLHF and methods such as DPO train models toward preferred behavior.
Adapt the Model
Full fine-tuning changes many weights; PEFT/LoRA updates a smaller set of parameters.
Smaller/Faster Inference
Reduce numerical precision such as FP16 → INT8/INT4. Common formats/tools include GGUF, GPTQ and AWQ.
Inference controls
Temperature, top-k, top-p, max tokens, stop sequences and structured output affect generation.
Performance
KV caching, batching, speculative decoding, efficient serving and model distillation improve cost/latency.
MoE
Mixture-of-Experts routes different tokens through selected expert networks instead of activating every parameter.
10 · RAG — Retrieval-Augmented Generation
RAG connects an LLM to external knowledge at inference time instead of requiring the model to memorize every answer.
1. Ingestion
Read PDF, DOCX, HTML, TXT or database records. Tools can include pypdf, python-docx or specialized parsers.
2. Chunking
Split documents into useful retrieval units.
Strategies: fixed-size, recursive, semantic, structure-aware, parent-child.
Use overlap carefully; too little loses context, too much adds redundancy.
3. Embeddings
Convert text into vectors representing semantic meaning.
Your stack: sentence-transformers → 384-dimensional vectors.
4. Vector DB
Store vectors + metadata for fast similarity search.
Your project: Qdrant collection telecom_knowledge.
5. Retrieval
Query embedding → similarity search → top-k chunks.
Similarity: cosine similarity is a common measure.
6. Reranking
A stronger model scores retrieved candidates for query relevance. Cross-encoders are a common approach.
7. Hybrid Search
Combine dense semantic retrieval with sparse keyword retrieval such as BM25.
Advanced Retrieval
Query rewriting, multi-query, HyDE, contextual compression, metadata filters and graph RAG.
RAG Evaluation
| Metric | What it tells you |
|---|---|
| Recall@K / Hit Rate | Did retrieval find the relevant chunk? |
| Precision@K | How much of the retrieved set is relevant? |
| MRR / nDCG | How highly were useful results ranked? |
| Context Precision/Recall | Is the retrieved context useful and sufficiently complete? |
| Faithfulness / Groundedness | Does the answer actually follow the retrieved evidence? |
| Answer Relevance | Does the answer address the user's question? |
11 · AI Agents
Workflow vs Agent
A workflow follows mostly predefined steps. An agent dynamically decides which tools/actions to use based on state and goals.
Tool Calling
The model selects a structured function such as weather lookup, database query, calculator or email sender.
ReAct
A common conceptual pattern: reason about the next action, act with a tool, observe the result, then continue.
State & Memory
Short-term conversation state plus optional long-term semantic, episodic or procedural memory.
Agent Patterns
Planner-executor, supervisor-worker, sequential agents, peer-to-peer and graph/state-machine workflows.
MCP
Model Context Protocol is a standard approach for connecting AI systems with tools and resources through defined interfaces.
Guardrails
Permissions, schemas, validation, human approval, tool restrictions, timeouts and sandboxing prevent dangerous actions.
Multi-Agent
Different specialized agents can collaborate, for example Weather Agent → Decision Agent → Communication Agent.
12 · Multimodal AI & Other AI Branches
Computer Vision
Classification, object detection, segmentation, OCR, image embeddings, tracking and visual question answering.
Tools: OpenCV, PyTorch, Transformers.
Speech AI
Speech-to-text, text-to-speech, speaker identification and audio classification.
Vision-Language
Models combine image and language representations for captioning, visual QA and image understanding.
Multimodal RAG
Retrieve text, images, tables or other modalities before generating an answer.
Generative Images
Diffusion-based text-to-image, image-to-image, inpainting and conditioning techniques.
Knowledge Graphs
Represent entities and relationships as nodes and edges. Useful for reasoning, entity linking and graph RAG.
Recommender Systems
Content-based filtering, collaborative filtering, matrix factorization and neural recommenders.
Causal AI
Studies cause-and-effect rather than only correlation. Learn interventions, confounding, DAGs and counterfactual reasoning.
13 · AI Engineering & Production
Application Architecture
UI → API → authentication → business logic → AI orchestration → model/RAG/agent → database/tools.
Serving
Learn REST, WebSocket, async processing, queues, streaming responses, batching, retries and rate limiting.
Local Inference
Learn CPU/GPU inference, Apple Silicon/MPS basics and local serving tools such as Ollama, llama.cpp or vLLM.
Observability
Trace prompts, retrieval, tool calls, latency, token usage, failures and user feedback.
Evaluation-Driven Development
Create golden datasets, regression tests and automated evaluation before changing prompts, retrievers or models.
Cost & Latency
Use caching, smaller models, routing, batching, quantization and retrieval limits to control cost.
Data Engineering
ETL/ELT, data validation, labeling, data versioning, feature pipelines and data quality monitoring.
Deployment
Docker, CI/CD, environment management and optional cloud/container orchestration are production skills.
14 · Responsible AI & Security
Hallucination
Generated content can be unsupported or incorrect. Mitigate with retrieval, grounding, validation and explicit uncertainty.
Prompt Injection
Untrusted text attempts to manipulate model instructions. Treat retrieved/user content as data, not authority.
Tool Security
Use least privilege, allowlists, schemas, sandboxing, approval gates and safe defaults.
Data Privacy
Minimize sensitive data, control access, redact where needed and understand retention policies.
Bias & Fairness
Measure model behavior across relevant groups and understand limitations of training data.
Supply Chain
Validate models, packages, datasets and dependencies before using them in production.
15 · Your TechHero Project — Where Everything Connects
Use this as your portfolio case study. It demonstrates normal web development plus real AI engineering.
Web Engineering
Django + MySQL handles authentication, technicians, tickets and business workflows.
RAG
Troubleshooting documents are chunked, embedded and stored in Qdrant for semantic retrieval.
Embeddings
Your current sentence-transformers setup produces 384-dimensional embeddings.
LLM Generation
The retrieved context is supplied to an LLM with instructions to answer from the evidence.
Feedback Loop
Technician feedback can become reviewed knowledge, creating an evaluation and knowledge-improvement loop.
Next Upgrades
Hybrid search → reranking → query rewriting → citations → RAG evaluation → tool-calling agent → multimodal support.
20 · AI Frameworks & Orchestration
Frameworks do not replace the underlying AI concepts. They help you build, connect, observe and control AI pipelines.
LLM Application Framework
Provides building blocks for prompts, models, retrievers, tools, document loaders, structured output and agent workflows.
Learn: models → prompts → parsers → retrievers → tools → agents → memory/state.
Stateful Agent Workflows
Build graph-based, stateful AI workflows where nodes perform work and edges control what happens next.
Useful for: multi-step agents, human approval, retries, loops, branching and durable state.
Data + RAG Framework
Focuses strongly on connecting LLMs to private data, indexes, retrievers and knowledge sources.
Search & RAG Pipelines
Framework for retrieval, document processing, pipelines and question-answering systems.
Programmatic LLM Optimization
Treat prompts and LLM calls as programmable components and optimize them against evaluation data.
Plain Python
Do not assume a framework is mandatory. For your TechHero project, implementing RAG directly with Python, Qdrant and an LLM is excellent for understanding the fundamentals.
21 · Embeddings, Vector Search & Vector Databases
Embedding
An embedding converts an item such as text, image or code into a numerical vector representing useful relationships.
Important: embedding dimensions depend on the model.
Dense Embeddings
Most vector dimensions contain learned continuous values. Sentence Transformers is a common Python ecosystem for semantic text embeddings.
Sparse Embeddings
Representations emphasize specific terms/features. Traditional TF-IDF and BM25 are sparse retrieval approaches.
Cosine Similarity
Measures the angle between vectors. Frequently used when comparing semantic embeddings.
Dot Product
Measures vector alignment through multiplication and summation. Some embedding systems are optimized for dot-product retrieval.
Euclidean Distance
Measures straight-line distance between vectors. Whether it is appropriate depends on how the embedding model was trained.
HNSW
Hierarchical Navigable Small World graphs provide fast approximate nearest-neighbor search.
IVF
Inverted File indexes partition vectors into clusters and search selected regions rather than the entire collection.
PQ
Product Quantization compresses vectors into smaller representations to reduce memory and speed search.
Metadata Filtering
Combine vector similarity with fields such as product, location, language, date, ticket type or permissions.
Vector Database Comparison
| Technology | Typical strength | Learn |
|---|---|---|
| Qdrant | Vector search + payload filtering + modern RAG applications | Collections, points, payloads, filters, HNSW |
| FAISS | Local/library-level similarity search | Indexes, ANN, clustering, vector distance |
| pgvector | Vector search inside PostgreSQL | SQL + vectors + relational metadata |
| Milvus | Large-scale vector infrastructure | Collections, indexes, distributed retrieval |
| Weaviate | Vector database with application-oriented features | Objects, vectors, filters, retrieval |
22 · Chunking — A Critical RAG Skill
Fixed-size
Split text after a target number of characters or tokens.
Good: simple baseline. Risk: may split meaning.
Recursive
Try separators such as paragraphs, lines and sentences before forcing smaller pieces.
A strong general-purpose baseline.
Semantic
Split where the meaning changes rather than using only length.
Can improve retrieval but costs more processing.
Structure-aware
Respect headings, sections, tables, code blocks, FAQ question/answer boundaries and document structure.
Parent-child
Retrieve a small child chunk but return a larger parent section to the LLM for additional context.
Chunk overlap
Repeat a small amount between neighboring chunks so information at boundaries is not lost.
Metadata
Store document ID, title, section, page, source, timestamp, category and access rules alongside each chunk.
Chunk quality test
Ask real questions and inspect whether the correct answer is contained in one or more retrieved chunks. Evaluate retrieval rather than guessing a universal chunk size.
23 · Advanced Retrieval Techniques
Dense Retrieval
Query and documents become embeddings and are compared semantically.
BM25
A strong lexical retrieval algorithm that rewards useful term matches and handles document length.
Hybrid Search
Combine sparse keyword retrieval and dense semantic retrieval for better coverage.
Reranking
Retrieve a broader candidate set, then use a stronger relevance model to reorder candidates.
Cross-Encoder
Reads the query and candidate document together to estimate relevance. More accurate but usually slower than bi-encoder retrieval.
Query Rewriting
Transform an unclear user question into a retrieval-friendly query.
Multi-Query Retrieval
Generate several search formulations to improve recall.
HyDE
Generate a hypothetical answer/document and use its embedding to search for real supporting documents.
Contextual Compression
Retrieve documents, then reduce them to the portions relevant to the current question.
Graph RAG
Use entities and relationships in a graph to retrieve connected knowledge, especially for multi-hop questions.
24 · Agent Engineering in Detail
Tool Schema
Define a tool name, description and typed parameters so the model can request it reliably.
Tool Router
Decide which available tool should receive a request.
Planner-Executor
A planner creates a strategy while an executor performs individual steps.
Supervisor Pattern
A supervisor routes tasks to specialized workers such as search, database, coding or communication agents.
Human-in-the-loop
Pause before risky actions and request approval from a person.
Memory Types
Short-term: current state. Semantic: facts. Episodic: past events. Procedural: learned procedures/workflows.
State Machine / Graph
Represent agent states and transitions explicitly. This makes complex workflows easier to debug and control.
Retries & Recovery
Handle tool failures, timeouts, malformed outputs and partial completion without blindly repeating dangerous actions.
Agent Evaluation
Measure tool selection, task completion, correctness, number of steps, latency, cost and unsafe-action rate.
MCP Architecture
Learn the concepts of MCP hosts, clients, servers, tools and resources, and why standardized tool/resource interfaces can simplify integrations.
25 · Model Families You Should Recognize
| Family / Architecture | What to understand | Typical use |
|---|---|---|
| Linear / Logistic | Simple statistical baselines | Prediction / classification |
| Tree / Forest / Boosting | Nonlinear tabular learning | Business datasets |
| CNN | Convolution and spatial features | Computer vision |
| RNN / LSTM / GRU | Recurrent sequence modeling | Historical NLP/time series |
| Autoencoder / VAE | Latent representation and generation | Compression / generation |
| GAN | Generator vs discriminator | Synthetic media |
| Diffusion | Denoising generative process | Images/audio/video research |
| BERT-style Encoder | Bidirectional contextual representation | Classification / extraction / embeddings |
| GPT-style Decoder | Causal next-token generation | LLMs / generation |
| Encoder-Decoder Transformer | Input sequence → output sequence | Translation / transformation |
| Vision Transformer | Image patches + attention | Vision |
| Multimodal Foundation Model | Multiple modalities in one system | Text + image/audio/video |
| Mixture-of-Experts | Sparse expert activation | Efficient large models |
26 · Data Engineering for AI
Data Cleaning
Missing values, duplicates, invalid records, inconsistent formats and outliers.
Feature Engineering
Transform raw variables into useful model inputs.
Data Leakage
Information from validation/test/future data accidentally enters training, producing unrealistically good results.
Class Imbalance
One class is much more frequent. Learn class weights, thresholding and methods such as SMOTE with proper validation.
Data Labeling
Human annotation creates supervised targets. Learn quality checks, inter-annotator agreement and active learning.
ETL vs ELT
ETL transforms before loading; ELT loads first and transforms in the target data system.
27 · MLOps / LLMOps
Experiment Tracking
Record datasets, parameters, metrics, model versions and results.
Model Registry
Track approved model versions and lifecycle status.
Prompt Registry
Version prompts just like application code.
Evaluation Sets
Maintain stable test questions and expected behaviors to catch regressions.
Monitoring
Track quality, latency, cost, failures, drift, retrieval quality and user feedback.
Tracing
Follow an AI request through prompt → retrieval → reranking → tool calls → model → final answer.
CI/CD for AI
Automate unit tests, integration tests, evaluation tests and safe deployment.
Governance
Access control, audit logs, data retention, model approvals and reproducibility.
28 · AI Terms You Must Be Able to Explain
Parameter
A value learned by a model during training, such as a neural-network weight.
Hyperparameter
A configuration selected outside normal parameter learning, such as learning rate or tree depth.
Epoch
One complete pass through the training dataset.
Batch
A subset of training examples processed together.
Inference
Using a trained model to produce a prediction or generation.
Context Window
The amount of input/output token context a model can handle for a request.
Grounding
Constraining an answer to trusted external evidence.
Hallucination
Unsupported or incorrect generated content presented as if it were true.
Fine-tuning
Further training a pretrained model on task/domain data.
Transfer Learning
Reuse learned representations from one task/domain for another.
Zero-shot
Perform a task without task-specific examples in the prompt.
Few-shot
Provide examples in the prompt to demonstrate the desired behavior.
Structured Output
Require machine-readable output such as JSON matching a schema.
Function Calling
Allow the model to request a structured application function/tool.
Guardrail
A rule, validator or control that limits unsafe or invalid behavior.
Ground-truth
The reference answer/label used to evaluate a system.
16 · Recommended Learning Roadmap
Python Foundation
Syntax, functions, OOP, exceptions, modules, typing, virtual environments, files, JSON, HTTP, async basics.
Data + Math
NumPy, Pandas, visualization, linear algebra, probability, statistics and optimization basics.
Classical ML
Preprocessing, regression, classification, clustering, evaluation, feature engineering and tuning.
Deep Learning
PyTorch, tensors, training loops, CNNs, sequence models, normalization and regularization.
NLP + Transformers
Tokenization, embeddings, attention, BERT/GPT concepts, Hugging Face and inference.
LLM Engineering
Prompting, structured output, function calling, fine-tuning, PEFT, quantization and evaluation.
RAG
Chunking, embeddings, Qdrant, retrieval, hybrid search, reranking, evaluation and citations.
Agents
Tools, state, memory, workflows, ReAct, MCP, multi-agent patterns and guardrails.
Multimodal
Vision, OCR, speech, vision-language models and multimodal RAG.
Production AI
FastAPI/Django APIs, queues, observability, security, evaluation, deployment and cost optimization.
17 · AI Interview Cheat Sheet
What is an embedding?
A numerical vector representation of data such as text that lets us compare semantic relationships using vector similarity.
What is RAG?
A pattern where relevant external knowledge is retrieved and supplied to an LLM before generation.
RAG vs Fine-tuning?
RAG changes the knowledge available at inference time; fine-tuning changes model behavior/parameters.
Vector DB vs MySQL?
MySQL excels at relational records and transactions. Vector databases optimize similarity search over embeddings.
Why cosine similarity?
It compares the direction of vectors and is commonly used to measure semantic similarity between embeddings.
What is Top-K?
The number of highest-ranked retrieval results returned for a query.
Why reranking?
Initial retrieval is fast and broad; a reranker can more precisely order the candidate documents.
What is temperature?
A generation control that changes the distribution of next-token probabilities; higher values generally increase randomness.
Transformer vs RNN?
Transformers use attention and can process token relationships efficiently in parallel during training; RNNs process recurrent state sequentially.
Precision vs Recall?
Precision asks “of predicted positives, how many were correct?” Recall asks “of actual positives, how many did we find?”
Overfitting?
The model learns training-specific patterns too closely and performs poorly on unseen data.
Agent vs RAG?
RAG retrieves knowledge. An agent can decide actions and use tools iteratively to accomplish a goal. An agent can also use RAG.
18 · Your Practical Project Ladder
01
Python data analysis dashboard
02
Scikit-learn classification API
03
PyTorch neural-network project
04
NLP sentiment / NER project
05
Semantic search using embeddings
06
Production RAG API
07
TechHero: Django + MySQL + Qdrant + LLM
08
Tool-calling weather/communication agent
09
Multi-agent workflow
10
Multimodal RAG + production evaluation
19 · Learning Progress
Tick topics as you study. Progress is saved in your browser.