Dockerize Everything: ML → LLM → Agents
This revision guide covers the full Docker lesson. Three real projects, every command decoded word by word, and one tiffin-delivery analogy to remember. 🍱 Read it, run it, then test yourself at the bottom.
The Dabbawala Analogy + Mac Setup
How to use this guide
This is a revision document, not a video transcript. It works best if you go through it three times, with your terminal open beside it.
- Pass 1 — Read Skim top to bottom without typing anything. Goal: recognise the vocabulary. Image, container, volume, compose, service name.
- Pass 2 — Build Rebuild all three projects from scratch, copying commands from here. Use the self-check lists as you go.
- Pass 3 — Recall Close this page. Review the questions and answers and interview questions at the bottom from memory. Whatever you miss, that's your revision list.
Every command block has a Copy button. Every command block is followed by a decoder that explains each flag — don't skip those, they're where the details are.
Why Docker exists at all
What you'll get:
a clear mental model of Docker before you write a single
command,
plus a working setup on your Mac. By the end of Port 0 you should have run
hello-world
and be
able
to explain what happened.
You've said it, or you will: "But it works on my machine!" The code crashes on your teammate's laptop, crashes on staging, and let's not talk about production.
Mumbai's dabbawalas deliver 200,000 lunch boxes every single day with Six Sigma accuracy — without an app. How? They standardised the packaging and the delivery system . It doesn't matter where a tiffin comes from or where it's going; the system is identical. Docker is the dabbawala system for software.
In this guide you'll dockerize three real projects — a classic ML model, an LLM-powered RAG app, and a multi-agent system. Same skill, three difficulty levels.
The Master Analogy — The Tiffin System 🍱
Come back to this table every time a new term confuses you. If you memorise one thing from the whole session, memorise this:
| Docker Concept | Tiffin World | One-line meaning |
|---|---|---|
| Dockerfile | Recipe card 📝 | Step-by-step written instructions for how to prepare the meal. |
| Image | Master tiffin (sealed, ready) 🍱 | The packed box produced from the recipe — frozen in time, ready to ship. |
| Container | One delivered tiffin 🚚 | A running copy of the image. One image can produce a hundred tiffins. |
| Docker Hub | Central kitchen / warehouse 🏭 | Where ready-made boxes for every recipe live — Python's, Ubuntu's, Redis's. |
| Port mapping | The building's gate number 🚪 |
The box lives in flat 8000 inside the building, but deliveries come through gate
8000 —
-p 8000:8000
.
|
| Volume | The steel box that comes back ♻️ | Delete the container, the data survives — like the reusable steel dabba. |
| docker compose | Ordering a full thali 🍽️ | One order gets you dal, rice, sabzi, roti — all containers together. |
⭐ The most important rule in this entire guide
"An image is a photograph, not a mirror." When you build an image, Docker takes a photo of your code at that moment. If you edit your code afterwards, the photo does NOT update by itself — you must take a new photo (rebuild). Forgetting this causes 90% of beginner confusion, so it's repeated throughout.
🖼️ VM vs Docker — the one diagram to hold in your head
Picture two kinds of housing:
- Virtual Machine = a standalone bungalow. Every app gets an entire house: its own kitchen, bathroom, and security guard (a full operating system). Heavy, slow to start, expensive.
- Docker = apartment flats. One building (the host operating system's core) is shared, but every flat (container) has its own lock, its own belongings, its own privacy. Lightweight, starts in seconds.
The punchline: a VM takes minutes to boot; a container takes milliseconds. That's why at Swiggy/Zomato scale you don't run VMs — you run containers. Traffic spike? Open 50 new flats in seconds.
💻 Mac setup — do this first
Open
Terminal
(press
Cmd + Space
, type "Terminal", press Enter). If the terminal is new
to you: it's just a way to talk to your computer with text instead of clicks. You type a
command, press
Enter,
the computer replies.
brew install --cask docker
docker --version
docker compose version
docker run hello-world
Command decoder — what each piece means
- brew
- Homebrew — the "app store for your terminal" on Mac. It downloads and installs software for you. No Homebrew? Download the .dmg from docker.com/products/docker-desktop instead.
- --cask
- Tells brew this is a full desktop application (with an icon), not just a command-line tool.
- docker --version
- Asks Docker "which version are you?" If it answers, Docker is installed correctly.
- docker compose version
- Checks the second tool you'll use — Compose — which manages multiple containers at once. It comes bundled with Docker Desktop.
- docker run hello-world
- "Run a container from the image called hello-world." Docker looks for it on your Mac, doesn't find it, downloads it from Docker Hub, and runs it. Your first tiffin, delivered from the central kitchen.
⚠️ The step most people miss: after installing, open the Docker Desktop app once (Cmd+Space → "Docker" → Enter). A whale 🐳 icon appears in the menu bar at the top of the screen. Docker commands only work while that whale is there — the app runs the Docker "engine" in the background. No whale = every command fails with "Cannot connect to the Docker daemon" . ("Daemon" is just an old Unix word for a background program.)
When
hello-world
prints
"Hello from Docker!"
— you have just pulled an image from
Docker
Hub and turned it into a running container. First tiffin delivered. 🎉
Two Mac traps to know before they hit you
Trap 1: Apple Silicon chips speak a different language
M-series Macs (M1/M2/M3/M4) use the ARM64 chip architecture; most cloud servers use AMD64 (x86) . Think of it as two languages: an image "written in AMD64" may not run on an ARM Mac. Sometimes an image downloads fine but refuses to start, or prints a platform warning.
The fix when you hit a platform errordocker run --platform linux/amd64 <image-name>
Command decoder
- --platform linux/amd64
- "Pretend to be an AMD64 machine." Your Mac translates on the fly (via Rosetta). Like watching a dubbed movie — slightly slower, but it works.
Trap 2: zsh and the stuck
quote>
prompt
The Mac terminal uses a shell called
zsh
. If you paste a command that has a
#
comment
at
the end AND that comment contains an apostrophe (like
you'll
), zsh gets confused and shows
quote>
, waiting forever.
Fix: press Ctrl+C and re-type the command without the
comment.
Remember this one — it saves ten minutes the first time it happens.
✅ Self-check — Port 0
- Docker Desktop is running (whale 🐳 steady in my menu bar)
-
docker run hello-worldsucceeded on my machine - I can recite 3 mappings from the tiffin analogy without looking
- I can explain "an image is a photograph, not a mirror" to someone else
- I can say the difference between a VM and a container in one sentence
Dockerize an ML Project — "QuickBite ETA" 🛵
The situation you're solving
What you'll build: a sklearn model (a food-delivery ETA predictor) served via FastAPI, packed into a container. This is where the core fundamentals live: Dockerfile anatomy, layers, caching, port mapping, .dockerignore.
Imagine you're an ML engineer at a Zomato-style startup. You've built a model that predicts how many minutes until an order arrives — distance, restaurant prep time, rider availability, rain as inputs; ETA as output.
The model runs beautifully on your laptop. Then DevOps says: "Ship it to the server." The server has Python 3.9; you have 3.12. Different sklearn version. NumPy may differ too. This is dependency hell. Docker is the air conditioning because it gives the app a controlled environment.
📁 Project structure
Build the skeleton first:
Terminalmkdir quickbite-eta && cd quickbite-eta
touch train.py app.py requirements.txt Dockerfile .dockerignore
Command decoder
- mkdir
- "Make directory" — creates a new folder.
- &&
- "Then" — run the next command only if the first one succeeded.
- cd
- "Change directory" — step inside that folder.
- touch
- Creates empty files with these names. You'll fill them in next.
scikit-learn==1.5.2
pandas==2.2.3
fastapi==0.115.6
uvicorn==0.34.0
joblib==1.4.2
==
is like writing "1 cup of rice" in a recipe
instead of "some rice" — you get the same dish every time.
🤖 train.py — a 60-second model
import pandas as pd, numpy as np, joblib
from sklearn.ensemble import RandomForestRegressor
# ① make the fake training data reproducible
np.random.seed(42)
n = 5000
# ② generate sample orders with the same features the API will receive
df = pd.DataFrame({
"distance_km": np.random.uniform(0.5, 12, n),
"prep_time_min": np.random.uniform(5, 30, n),
"rider_available": np.random.randint(0, 2, n),
"is_raining": np.random.randint(0, 2, n),
})
# ③ compute a noisy ETA target so the model has a pattern to learn
# ETA = base + distance*3 + prep + rain penalty + rider penalty + noise
df["eta_min"] = (8 + df.distance_km*3 + df.prep_time_min*0.7
+ df.is_raining*9 + (1-df.rider_available)*6
+ np.random.normal(0, 2, n))
# ④ train the model to predict eta_min from the order features
X, y = df.drop(columns=["eta_min"]), df["eta_min"]
model = RandomForestRegressor(n_estimators=60, random_state=42).fit(X, y)
# ⑤ save the trained model so the API can load it without retraining
joblib.dump(model, "eta_model.pkl")
print("Model saved: eta_model.pkl ✅")
joblib.dump
saves the trained brain into a single file,
eta_model.pkl
, so the API
can load it later without retraining. The
seed(42)
line means "use the same randomness every
time" so your results match everyone else's.
🚀 app.py — FastAPI serving
from fastapi import FastAPI
from pydantic import BaseModel
import joblib, pandas as pd
# ① create the API and load the trained model once at startup
app = FastAPI(title="QuickBite ETA")
model = joblib.load("eta_model.pkl")
# ② describe the order fields every prediction request must send
class Order(BaseModel):
distance_km: float
prep_time_min: float
rider_available: int
is_raining: int
@app.get("/")
def health():
return {"status": "QuickBite ETA is live 🛵"}
@app.post("/predict")
def predict(order: Order):
# ① convert validated JSON into the dataframe shape the model expects
X = pd.DataFrame([order.model_dump()])
# ② run the model and round the ETA for display
eta = round(float(model.predict(X)[0]), 1)
# ③ send the prediction back as JSON for the caller
return {"eta_minutes": eta, "message": f"Your food arrives in {eta} min 🍔"}
/predict
, run this Python
function." The
Order
class is the order form: it declares exactly which fields a request must
contain and their types — send text where a number belongs and FastAPI politely rejects it
for free
.
At
startup you load the saved model brain from
eta_model.pkl
once; then every request is: read
form
→ ask model → return the answer as JSON (the universal "key: value" text format APIs speak). The
/
route is a health check — a doorbell to confirm the shop is open.
Dockerfile anatomy — learning to read the recipe
This is the most important block in the guide. A Dockerfile is a plain text file (no extension!) with instructions Docker executes top to bottom to produce an image. Map every line to the tiffin analogy:
Dockerfile# ① Base image = rent a ready-made kitchen (from Docker Hub)
FROM python:3.12-slim
# ② Set up your counter inside that kitchen
WORKDIR /app
# ③ Copy ONLY the shopping list first (caching trick — explained below)
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# ④ NOW copy the rest of your code
COPY . .
# ⑤ Train the model INSIDE the image (baked in at build time)
RUN python train.py
# ⑥ Declare which window the food comes out of (documentation)
EXPOSE 8000
# ⑦ What runs the moment the tiffin is opened
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
Line-by-line decoder
- FROM python:3.12-slim
- Don't start from an empty computer — start from a ready-made image that already has Linux + Python 3.12 installed. "slim" = the lightweight version. Docker downloads it from Docker Hub automatically.
- WORKDIR /app
- "From now on, work inside the folder /app" (inside the container). Creates it if missing. Like choosing which counter you'll cook on.
- COPY requirements.txt .
-
Copy one file from your Mac into the image. The
.means "into the current folder" (which is /app because of WORKDIR). - RUN pip install ...
-
RUN executes a command
while building
the image. pip is Python's package installer;
-r requirements.txtmeans "install everything on this list".--no-cache-dirtells pip not to keep downloaded files around — smaller image. - COPY . .
- "Copy everything in the current folder on my Mac → into /app in the image." (Everything except what .dockerignore excludes.)
- RUN python train.py
- Trains the model during the build, so the finished image already contains eta_model.pkl. The container never needs to train anything.
- EXPOSE 8000
-
A label saying "this app listens on port 8000". It's documentation — it doesn't open
anything by
itself
(that's
-p's job later). - CMD [...]
-
The one command that runs when a container starts. Here: start the uvicorn web
server, serving the
appobject fromapp.py.--host 0.0.0.0means "accept connections from outside the container, not just from inside" — without it, port mapping silently fails. A classic gotcha.
🔑 The caching trick — why requirements.txt goes first
Picture a stack of layers (like a stack of parathas 🫓). Docker turns each instruction into a layer and caches it. When you rebuild, Docker checks each layer top-down: "did anything this layer depends on change?" If not, it reuses the cached layer instantly. But the moment one layer changes, every layer below it must rebuild too .
Now the logic is clear: if you did
COPY . .
first and installed dependencies after, then
every
tiny code edit
would invalidate the copy layer — and force the slow 2-minute pip install to re-run
below
it. By copying only
requirements.txt
first, the pip layer only rebuilds when the shopping list
itself changes.
The rule to remember: if the shopping list hasn't changed, why go back to the store? Code changes daily; dependencies change monthly. Rarely-changing things at the top, frequently-changing things at the bottom. This one trick makes builds 10x faster.
🙈 .dockerignore
__pycache__/
*.pyc
venv/
.venv/
.git/
.env
*.ipynb
eta_model.pkl
data/raw/
COPY . .
, it grabs
everything
in
the folder — unless it's listed here. Same idea as .gitignore. You exclude: Python's junk cache
files,
virtual
environments (can be 500MB!), git history, secrets (.env), and notebooks. You also exclude
eta_model.pkl
— if you ever trained locally, you do NOT want that stale local file copied in;
the
image trains its own fresh copy in step 5. "Only the food goes into the tiffin — not your diary
and house
keys."
Milestone #1 — Build, run, predict
docker build -t quickbite-eta:v1 .
docker images
docker run -d -p 8000:8000 --name eta-service quickbite-eta:v1
curl -X POST http://localhost:8000/predict \
-H "Content-Type: application/json" \
-d '{"distance_km": 4.5, "prep_time_min": 15, "rider_available": 1, "is_raining": 1}'
Command decoder — slow down and read every line here
- docker build
- "Follow the Dockerfile recipe and produce an image."
- -t quickbite-eta:v1
- "Tag" — give the image a name and a version label (name:version). Without it you get an unmemorable ID like a7f3c9.
- .
- The lonely dot means "the recipe and files are in THIS folder." Forgetting the dot is the #1 first-day error.
- docker images
- List all images (master tiffins) stored on your machine.
- docker run
- "Create and start a container from this image."
- -d
- "Detached" — run in the background and give my terminal back. Without -d, the logs take over your terminal until you press Ctrl+C (which also stops the container!).
- -p 8000:8000
- Port mapping, format host:container — "connect gate 8000 of my Mac to flat 8000 inside the container." Traffic to localhost:8000 gets forwarded inside.
- --name eta-service
- A friendly name so you can say "eta-service" in later commands instead of a random ID.
- curl
- A terminal tool for sending web requests — a browser without the window.
- -X POST
- The request type. GET = "give me something", POST = "here's data, process it."
- -H "Content-Type..."
- A header telling the server "the data I'm sending is JSON."
- -d '{...}'
-
The data itself — one order, as JSON. The backslash
\at line ends just means "command continues on the next line."
The response comes back as
{"eta_minutes": 41.2, ...}
. Now open
http://localhost:8000/docs
in your browser: FastAPI auto-generates a Swagger page. It lists the API and lets you send the same request in the browser. Set is_raining to 1; the ETA becomes higher. 🌧️
Notice what you did NOT do on your Mac: no Python environment, no pip install, no version checking. Everything lives inside the box. And this exact box will run on AWS, on a Windows laptop, anywhere — identically .
🛠️ Peek inside the container + your everyday commands
docker ps
docker ps -a
docker logs -f eta-service
docker exec -it eta-service bash
docker stop eta-service
docker rm eta-service
Command decoder
- docker ps
- List running containers only. Comes from "process status".
- docker ps -a
-
List ALL containers, including dead/exited ones. ⚠️ Memorise this: if a container crashed,
plain
pshides it and you'll think it never existed.-a= "all". - docker logs -f eta-service
-
Show everything the container has printed.
-f= "follow" — keep streaming new lines live (Ctrl+C to stop watching; the container keeps running). - docker exec -it ... bash
-
"Execute a command inside a running container." The command here is
bash— a shell — so you get a terminal INSIDE the box.-it= interactive + terminal, i.e. "let me type." Trylsandcat app.pyinside, thenexitto come back out. - docker stop / rm
- stop = pause the delivery (container still exists, restartable). rm = remove the stopped container entirely. The image is untouched — you can always run a fresh one.
docker exec
makes the idea concrete: you are standing inside a tiny, separate Linux world
living
inside your Mac.
The photograph rule in action — try this deliberately
Edit the message string in
app.py
, then restart the container.
Nothing changes.
Why?
The
container runs the image — the photograph — and the photo was taken before your edit. The fix
is always:
docker build -t quickbite-eta:v1 .
docker rm -f eta-service
docker run -d -p 8000:8000 --name eta-service quickbite-eta:v1
Rebuild → remove old container (
rm -f
= force-remove even if running) → run fresh. Losing ten minutes to a stale image is a common beginner mistake. Do it once here on purpose so you remember the fix.
✅ Self-check — Port 1
-
Image built: quickbite-eta shows up in
docker images -
/predictworks for me via both curl and Swagger - I can explain layer caching in one sentence
-
I stepped inside with
docker execand came back out - I can answer: "I edited my code, why doesn't the container see it?"
-
I can explain why
--host 0.0.0.0is required in the CMD
Dockerize an LLM Project — "ScalerGPT" RAG bot 📚
What changes at level 2
What you'll build: a RAG chatbot (FastAPI + OpenAI API + ChromaDB). New concepts: secrets/env vars, docker compose, volumes, multi-container networking, startup readiness . This code is tested — it includes fixes for two real bugs that show up every time.
This time the model isn't yours — it's OpenAI's. LLM projects bring three problems classic ML didn't have:
1️⃣
Secrets
— bake your API key into the image and you've written your PIN on your ATM card.
2️⃣
Multiple services
— an app plus a vector database. Two boxes, one order.
3️⃣
State
— the database's data must survive even after its container dies.
The three solutions you'll learn here: .env files, docker compose, and volumes.
RAG in 60 seconds
RAG = Retrieval Augmented Generation — an open-book exam for the LLM. Instead of answering from memory (where it hallucinates), the model is handed the relevant pages from your notes and told "answer using only this."
- Embedding: turning text into a list of numbers such that similar meanings get similar numbers. It's how the machine "feels" that "container" and "Docker box" are related even though they share no words.
- Vector database (Chroma): a library that stores those numbers and can instantly find "the 3 most similar passages to this question."
- The pipeline: question → find relevant chunks (retrieve) → paste them into the prompt (augment) → let the LLM write the answer (generate).
📁 Project structure
mkdir scalergpt && cd scalergpt
mkdir docs
touch app.py ingest.py requirements.txt Dockerfile docker-compose.yml .env.example .dockerignore .gitignore
requirements.txt
fastapi==0.115.6
uvicorn==0.34.0
openai==1.59.7
chromadb-client==0.6.3
python-dotenv==1.0.1
chromadb-client
— the
thin client
—
not
the full
chromadb
package. The full package IS the database (heavy); the client just
talks
to a database running elsewhere. Since Chroma will live in its own container, your app only
needs the phone,
not the whole telephone exchange. This keeps the app image small — the #1 problem with LLM
images is size.
Drop some
.txt
or
.md
notes into
docs/
— they become the bot's
knowledge.
Secrets 101 — the ATM PIN rule
The golden rule: "Image = ATM card (safe to share). Env var = PIN (inject at runtime, never write it on the card)."
An environment variable is a named value the operating system hands to a program when it starts — like a sticky note passed to the chef as they walk in, rather than printed in the recipe book everyone can read. A .env file is simply a text file of such notes, one per line.
.env.example → copy to .env and add your real key# Copy this file: cp .env.example .env — then paste your real key.
OPENAI_API_KEY=sk-paste-your-real-key-here
-
No quotes, no spaces around the
=. Get a key at platform.openai.com/api-keys. The whole demo costs less than one US cent. -
Add
.envto BOTH.dockerignoreAND.gitignore. You ship a safe.env.exampletemplate instead; each person copies it to.envlocally. -
If you ever write
ENV OPENAI_API_KEY=sk-...in a Dockerfile, anyone can read it back withdocker history. This is a favourite interview question.
🧠 app.py — RAG with a retry loop ✅ fixed version
This is the corrected code. The naive version crashes at startup — the next card explains exactly why, because that bug is the best lesson in the whole guide.
app.pyimport os, sys, time
import chromadb
from chromadb.utils import embedding_functions
from fastapi import FastAPI, HTTPException
from openai import OpenAI
from pydantic import BaseModel
app = FastAPI(title="ScalerGPT")
# ① fail loudly if the key is missing instead of showing a cryptic traceback
# Fail LOUDLY and clearly if the key is missing - not with a cryptic traceback
API_KEY = os.getenv("OPENAI_API_KEY", "").strip()
if not API_KEY or API_KEY.startswith("sk-paste"):
sys.exit("[ScalerGPT] OPENAI_API_KEY missing. Put a real key in .env")
llm = OpenAI(api_key=API_KEY)
# ② create the embedder Chroma will use for similarity search
# The thin client has no built-in embedder - we must supply one explicitly
openai_ef = embedding_functions.OpenAIEmbeddingFunction(
api_key=API_KEY, model_name="text-embedding-3-small")
# ③ read the Chroma network address from environment variables
CHROMA_HOST = os.getenv("CHROMA_HOST", "localhost")
CHROMA_PORT = int(os.getenv("CHROMA_PORT", "8000"))
def connect_to_chroma(retries=30, delay=2):
# ① retry because Chroma may be started before it is ready
# Chroma takes a few seconds to boot. depends_on only waits for its
# container to START, not to be READY - so we knock politely and retry.
for attempt in range(1, retries + 1):
try:
# ② create a client and prove the service answers a heartbeat
client = chromadb.HttpClient(host=CHROMA_HOST, port=CHROMA_PORT)
client.heartbeat()
print(f"[ScalerGPT] Connected to chroma at {CHROMA_HOST}:{CHROMA_PORT}", flush=True)
# ③ return the ready client so the app can continue startup
return client
except Exception as e:
# ④ wait before trying again so startup races can self-heal
print(f"[ScalerGPT] Waiting for chroma ({attempt}/{retries}): {type(e).__name__}", flush=True)
time.sleep(delay)
# ⑤ stop with a clear message if Chroma never becomes reachable
sys.exit(f"[ScalerGPT] Could not reach chroma at {CHROMA_HOST}:{CHROMA_PORT}")
# ④ connect to Chroma and open the notes collection
chroma = connect_to_chroma()
collection = chroma.get_or_create_collection(name="notes", embedding_function=openai_ef)
# ⑤ define the request body shape for /ask
class Question(BaseModel):
query: str
@app.get("/")
def health():
return {"status": "ScalerGPT is live 📚", "docs_indexed": collection.count(),
"chroma_host": CHROMA_HOST, "chroma_port": CHROMA_PORT}
@app.post("/ask")
def ask(q: Question):
# ① stop early if no documents were indexed
if collection.count() == 0:
raise HTTPException(status_code=400,
detail="No documents indexed. Run: docker compose exec app python ingest.py")
# ② RETRIEVE - find the 3 most relevant chunks
hits = collection.query(query_texts=[q.query], n_results=3)
documents = hits.get("documents") or [[]]
context = "\n\n---\n\n".join(documents[0])
# ③ AUGMENT - paste those chunks into the prompt
system_prompt = ("You are ScalerGPT, a helpful teaching assistant. "
"Answer using ONLY the context below. If it does not contain the answer, "
f"say you don't know.\n\nCONTEXT:\n{context}")
# ④ GENERATE - let the LLM write the final answer
resp = llm.chat.completions.create(model="gpt-4o-mini",
messages=[{"role": "system", "content": system_prompt},
{"role": "user", "content": q.query}])
# ⑤ return the model answer plus how many chunks supported it
return {"question": q.query, "answer": resp.choices[0].message.content,
"sources_used": len(documents[0])}
/ask
is the three-step RAG
dance: find the 3 most relevant passages, paste them into the instructions, let GPT write the
answer. The
if count() == 0
guard gives a helpful "you forgot to ingest" message instead of a confusing
empty
answer.
You'll also need
ingest.py
— it reads every file in
docs/
, splits them into
paragraph chunks, and loads them into Chroma, using the same retry pattern.
The bug that WILL bite you: started ≠ ready
Here's the exact sequence, because you will hit it:
-
You run
docker compose up -d. Thendocker compose psshows… only chroma . The app has vanished. -
Lesson 1:
pshides dead containers.docker compose ps -areveals the app: Exited (1) . -
Lesson 2:
docker compose logs appshows the reason: "Connection refused… Could not connect to a Chroma server." -
Lesson 3 (the real one): the compose file says
depends_on: chroma— so why did it fail? Becausedepends_ononly waits for chroma's container to START, not for the database inside it to be READY. Chroma needs a few seconds to boot. The app knocked immediately, got no answer, and gave up.
Analogy:
the restaurant unlocked its door (container started) but the chef hasn't tied his apron yet
(service not ready). If you shout your order at the locked kitchen and storm out, that's a
crash. The retry
loop in app.py is the polite customer who waits and knocks again. (The alternative fix — a
compose
healthcheck
+
depends_on: condition: service_healthy
— is in the practice
challenges.)
docker-compose.yml — the thali order system ✅ fixed version
So far you ordered tiffins one at a time —
docker run
this,
docker run
that.
Compose says:
write one menu, then just say "serve the thali."
The file format is YAML — a way of
writing structured settings where
indentation shows what belongs to what
, like a neatly indented
shopping list. (Careful: YAML is picky — use spaces, never tabs.)
# ① define the containers that make up the RAG stack
services:
app:
# ② build and run the FastAPI app from this folder
build: .
container_name: scalergpt-app
ports:
- "8000:8000"
# ③ pass secrets and Chroma's internal address to the app
env_file: .env
environment:
- CHROMA_HOST=chroma
- CHROMA_PORT=8000 # INTERNAL port - see the port trap below!
# ④ start Chroma first and restart the app if it crashes
depends_on:
- chroma
restart: unless-stopped
chroma:
# ⑤ run the vector database and expose it to the host on port 8001
image: chromadb/chroma:0.6.3
container_name: scalergpt-chroma
ports:
- "8001:8000"
# ⑥ persist Chroma data outside the container filesystem
volumes:
- chroma_data:/chroma/chroma
restart: unless-stopped
# ⑦ create the named volume used by the Chroma service
volumes:
chroma_data:
Line-by-line decoder
- services:
- The menu. Each entry below is one container you want.
- build: .
- "Build this service's image from the Dockerfile in this folder" (the app's Dockerfile is the same 6-liner as Port 1, minus the training step).
- image: chromadb/chroma:0.6.3
- Don't build — download this ready-made image from Docker Hub. You never write a Dockerfile for Chroma; someone already packed that tiffin.
- env_file: .env
- "At startup, hand this container all the sticky notes from .env" — the PIN gets injected at runtime, never baked into the image.
- environment:
- More sticky notes, written directly here (fine for non-secrets like hostnames).
- CHROMA_HOST=chroma
- Not an IP address — the service name ! Compose creates a private network where every service is reachable by its name, like flats on an intercom. The app dials "chroma" and Docker connects the call.
- depends_on:
- "Start chroma before me." ⚠️ Start — not ready. That's why app.py retries.
- restart: unless-stopped
- "If I crash, bring me back automatically" — unless a human explicitly stopped me. A free safety net.
- volumes: (on chroma)
- "Mount the storage box named chroma_data at the path /chroma/chroma inside the container" — that's where Chroma keeps its data, so the data now lives OUTSIDE the disposable container.
- volumes: (bottom)
- Declares the storage box itself so Docker creates and tracks it.
The port trap: published vs internal
Look at chroma's line:
"8001:8000"
. Two different numbers — this is where everyone gets
burned:
- Chroma listens on port 8000 inside its own container (the flat number).
- You publish it as 8001 on the Mac (the street gate) — only so YOU can poke it from outside for debugging, and because the app already took the Mac's 8000.
-
The app container is
already inside the building
— it's a neighbour, not a street visitor.
Neighbours use flat numbers. So the app must use
CHROMA_PORT=8000.
Setting
CHROMA_PORT=8001
is the single most common bug in this setup
— it produces
"connection refused" and an exited app container. Say it twice:
containers talking to containers use
INTERNAL ports. Only your Mac uses the published port.
Milestone #2 — Serve the thali
cp .env.example .env # then put your REAL key inside .env
docker compose up -d --build
docker compose ps -a
docker compose logs app
docker compose exec app python ingest.py
curl -X POST http://localhost:8000/ask \
-H "Content-Type: application/json" \
-d '{"query": "What is Docker?"}'
Command decoder
- docker compose up
- "Read docker-compose.yml and make reality match it" — create the network, the volume, and all containers, in dependency order.
- -d
- Detached — background, same as before.
- --build
-
"Rebuild my images first if code changed." ⚠️ Make this a habit: after ANY file edit,
up -d --build. Without it, Compose happily reuses the old photograph and your edit never arrives. This single mistake costs most people 15 minutes. - docker compose ps -a
- List this project's containers including dead ones . Expect both "Up". If app says "Exited", read its logs.
- docker compose logs app
- The app's diary. You want to see: [ScalerGPT] Connected to chroma at chroma:8000 — you may first see a few "Waiting for chroma (1/30)" lines. That shows the retry loop is working.
- docker compose exec app python ingest.py
-
"Inside the already-running
app
container, execute
python ingest.py." This is how you run one-off jobs (migrations, imports) in production — you don't start a new container, you step into the live one.
Then open http://localhost:8000/docs and ask questions from the Swagger UI. The test worth doing: ask "What is the capital of France?" — ScalerGPT says it doesn't know, because the answer isn't in your docs. That's proof the answers are grounded in YOUR documents, not the model's memory.
Milestone #2.5 — The volume persistence test
docker compose down
docker compose up -d
curl http://localhost:8000/
Command decoder
- docker compose down
- The opposite of up: stop and DELETE all containers and the network. The thali is cleared. But volumes survive by default.
- docker compose up -d
- Brand-new containers from scratch.
- curl http://localhost:8000/
-
docs_indexedis STILL > 0 — you never re-ingested, yet the data is there. It lives in the volume, not the container. ♻️ - docker compose down -v
-
Know it, don't run it casually:
-valso deletes volumes. THIS is how you actually lose the data. The steel dabba goes to the scrapyard.
The takeaway: containers die all the time — deployments, crashes, scaling. Volumes are the steel box that comes back after every delivery. Container = disposable. Volume = permanent.
✅ Self-check — Port 2
-
Both services show "Up" in
docker compose ps -a -
I ingested docs and got a grounded answer from
/ask -
I proved the volume works:
down→up→ data still there - I can explain why my API key must never go in the Dockerfile
- I can explain the difference between the published port and the internal port
-
I can explain why
depends_onalone didn't save me
Dockerize an Agentic Project — "DeskBuddy" 🤖
What an agent actually is
What you'll build: a 3-container agentic system — an agent service (LLM + tool-calling loop), a tools service (separate microservice), and Redis (conversation memory). New concepts: microservice separation, private networking (no published port!), and one compose file running it all.
An agent is an LLM that doesn't just answer — it does things . It thinks, picks up a tool, checks the result, and keeps going.
Think of the architecture as an office: the agent = the manager (decides what needs doing), tools = the departments (they do the actual work), and Redis = the office register (remembers who said what).
Why three separate boxes? Because in production, the tools team is different from the agent team. You update the tools without touching the agent. That's microservices — and without Docker, microservices are just a slide in a deck.
What is Redis, in 30 seconds?
Redis is a very fast "sticky-note board" database: you store values under names (
key → value
) and
read them back in microseconds. Here it remembers each conversation: key = the session ID,
value = the
message
history. Why not a Python variable? Because containers die and restart — a variable dies
with them. Redis in
its own container (with a volume) means the agent can crash, restart, and still remember
you. Someone
already
packed the Redis tiffin: you just write
image: redis:7-alpine
.
📁 Structure + the tools service
mkdir deskbuddy && cd deskbuddy
mkdir agent tools
touch docker-compose.yml .env.example
touch agent/app.py agent/requirements.txt agent/Dockerfile
touch tools/app.py tools/requirements.txt tools/Dockerfile
from fastapi import FastAPI
from pydantic import BaseModel
import datetime
# ① create a tiny API that exposes utility tools
app = FastAPI(title="DeskBuddy Tools")
# ② describe the calculator request body
class Calc(BaseModel):
expression: str
@app.post("/calculator")
def calculator(c: Calc):
try:
# ① evaluate the expression in a stripped-down namespace for the demo
# demo only - never use eval in production!
return {"result": eval(c.expression, {"__builtins__": {}})}
except Exception as e:
# ② report tool errors as data so the agent can continue
return {"error": str(e)}
@app.get("/datetime")
def now():
# ① return the current timestamp for the agent's clock tool
return {"now": datetime.datetime.now().isoformat()}
eval
is fine for a classroom demo and dangerous in production — it can execute arbitrary code. In a
real service
you'd use a safe expression parser instead.
🧠 The agent — the think → act → observe loop
# ① connect to shared Redis memory and the private tools service
r = redis.Redis(host=os.getenv("REDIS_HOST", "redis"), port=6379, decode_responses=True)
TOOLS_URL = os.getenv("TOOLS_URL", "http://tools:7000")
@app.post("/chat")
def chat(req: Chat):
# ① load this session's memory and append the new user message
key = f"history:{req.session_id}"
history = [json.loads(m) for m in r.lrange(key, 0, -1)] # load memory
history.append({"role": "user", "content": req.message})
# ② let the model decide whether it needs a tool, up to five rounds
for _ in range(5): # the agent loop
# ③ ask the model for the next assistant message or tool call
resp = llm.chat.completions.create(
model="gpt-4o-mini", messages=history, tools=TOOL_DEFS)
msg = resp.choices[0].message
if not msg.tool_calls: # no tool needed?
break # then we're done
# ④ save the assistant tool request before running the tool
history.append(msg.model_dump(exclude_none=True))
for tc in msg.tool_calls: # run each requested tool
# ⑤ execute each requested tool and feed its result back to the model
result = call_tool(tc.function.name, json.loads(tc.function.arguments))
history.append({"role": "tool", "tool_call_id": tc.id,
"content": json.dumps(result)})
# ⑥ save memory back to Redis, then return the final answer
# save memory back to redis, return final answer
TOOL_DEFS
describes
each tool's name and inputs — the menu card); (2) the LLM either answers in words — done, break
— or replies
"please run calculator with 23*47 for me"; (3) you actually call the tools service over HTTP,
paste the
result
back into the conversation, and go around again so the LLM can see what happened. Max 5 laps so
a confused
model can't loop forever (a "safety fuse"). Memory: before the loop you load this session's
history from
Redis; after, you save it back — that's how a follow-up like "now double it" works.
Look at the two addresses:
http://tools:7000
and
host="redis"
. Again —
service names, not IPs. By project 3 this should feel natural.
The full thali — 3-service compose
# ① define the agent, tool, and memory containers
services:
agent:
# ② build the public chat API from the agent folder
build: ./agent
ports:
- "9000:9000"
# ③ give the agent secrets and private service names
env_file: .env
environment:
- TOOLS_URL=http://tools:7000
- REDIS_HOST=redis
# ④ start the helper services before the agent
depends_on: [tools, redis]
restart: unless-stopped
tools:
# ⑤ build the internal tools API without exposing host ports
build: ./tools
restart: unless-stopped
# NOTE: no ports! Explained below - this is the security gem
redis:
# ⑥ run Redis and keep its data in a named volume
image: redis:7-alpine
volumes:
- agent_memory:/data
restart: unless-stopped
# ⑦ create the named volume Redis uses for memory
volumes:
agent_memory:
What's new here
- build: ./agent
- Each service builds from its own subfolder's Dockerfile. One compose file, two custom images, one ready-made.
- tools: (no ports!)
-
The security gem. No
portssection = no street gate = the outside world cannot reach it at all . Only fellow residents of the private network (the agent) can call it attools:7000. Trycurl localhost:7000from your Mac — connection refused, by design . "Never give internal departments a public entrance." One tiny omission; gold in interviews and in production. - redis:7-alpine
- "alpine" = built on a tiny 5MB Linux. Whole Redis image ≈ 40MB.
- agent_memory volume
- Same trick as Chroma: conversation memory survives container death.
Note what you did
not
need here: the retry-loop lesson doesn't bite, because the Redis client only
connects when first used (lazily), and Redis boots in under a second.
restart: unless-stopped
is
the seatbelt anyway.
Milestone #3 — The agent in action + memory proof
Split your terminal in two. Left:
docker compose logs -f
. Right: the curls. The logs show
requests ripple across three containers live.
cp .env.example .env # real key inside, same as before
docker compose up -d --build
docker compose ps -a
curl -X POST http://localhost:9000/chat \
-H "Content-Type: application/json" \
-d '{"session_id": "demo", "message": "What is 23*47, and what time is it right now?"}'
curl -X POST http://localhost:9000/chat \
-H "Content-Type: application/json" \
-d '{"session_id": "demo", "message": "Now double that multiplication result"}'
docker compose logs -f
What to watch for
- First curl
- A two-tool task on purpose: the agent must call BOTH calculator and clock, then compose one answer. Watch the loop go around twice in the logs.
- Second curl
- Same session_id = same Redis key = the agent remembers 1081 from a minute ago and answers 2162. Change session_id to "other" and ask again — no memory. That's isolation per user, for free.
- docker compose logs -f
- All three services' diaries interleaved, colour-coded by name. This is how you debug multi-service systems.
The point: three services, one command, real memory, proper isolation. And this exact compose file works unchanged on an EC2 instance — same commands, same result . That bridge from laptop to production is Docker's main value.
✅ Self-check — Port 3
-
3 containers running
(
docker compose ps -a) - The agent answered using a tool call
- The memory demo worked — my follow-up used prior context
-
I can explain why
curl localhost:7000fails on purpose - I can draw the 3-service architecture from memory on paper
Cheatsheet, Golden Rules & Debugging
📋 The cheatsheet — full command list
This is the section to keep open on a second screen while you work. Use the full cheatsheet, follow the debug tree when something breaks, and re-read the golden rules before any interview.
| Command | What it does | Tiffin translation |
|---|---|---|
docker build -t name:tag .
|
Build an image from the Dockerfile here | Pack the master box from the recipe |
docker run -d -p 8000:8000 img
|
Start a container, background, map ports | Deliver the tiffin, set the gate |
docker ps
/
docker ps -a
|
Running containers / ALL incl. dead ones | Today's deliveries / the full register |
docker logs -f name
|
Stream a container's output live | The box's diary |
docker exec -it name bash
|
Open a shell inside a running container | Step inside the box |
docker images
|
List all images on your machine | The shelf of master tiffins |
docker rm -f name
|
Force-remove a container, even if running | Recall the delivery immediately |
docker compose up -d --build
|
Rebuild if needed + start all services | Serve the thali (fresh) |
docker compose ps -a
|
This project's containers, incl. exited | Which dishes made it, which didn't |
docker compose exec app CMD
|
Run a one-off command in a live service | Ask the chef mid-service |
docker compose down
|
Stop + delete containers (volumes safe) | Clear the thali, keep the steel boxes |
docker compose down -v
|
…and delete volumes too ⚠️ data gone | Scrap the steel boxes |
docker compose logs -f
|
All services' logs together | CCTV over the whole kitchen |
docker compose up -d --force-recreate
|
Recreate containers (e.g. after .env edits) | Fresh boxes, same recipe |
docker volume ls
|
List volumes on your machine | Count the steel dabbas in stock |
docker system df
|
How much disk Docker is eating | Check the pantry weight |
docker system prune -a
|
Delete unused images/containers (careful!) | Diwali deep-clean 🧹 |
docker run --platform linux/amd64
|
The Mac ARM fix | Dubbing for another audience |
docker history image
|
Show every layer + the command that made it | Read the recipe backwards (and find leaked secrets) |
lsof -i :8000
|
Find which process is holding a port | Who's blocking the gate? |
The golden rules — read these before any interview
-
The image is a photograph, not a mirror.
Edited a file? Rebuild:
docker compose up -d --build. -
pshides the dead. Service missing?ps -a, thenlogs <service>. -
Started ≠ ready.
depends_onwaits for the container, not the service inside. Retry in code or add a healthcheck. - Internal ports for neighbours, published ports for visitors. Container→container uses the internal port.
- Secrets are PINs. .env + env_file at runtime; never ENV in a Dockerfile, never commit .env.
-
.env changed? Recreate.
Env vars load at container start:
--force-recreate. -
No
ports:= no public entrance. Internal services should stay internal. - Rarely-changing layers on top. That's the whole caching game.
-
See
quote>in zsh? You pasted a # comment. Ctrl+C, re-run without it.
🧭 The debug decision tree — follow it top to bottom
When something breaks, don't guess. Walk this:
Something is broken.
├─ Does docker ps show my container?
│ ├─ NO → run docker ps -a (or compose ps -a)
│ │ ├─ Status "Exited" → docker logs <name> ← the answer is in here
│ │ │ ├─ "Connection refused" → started ≠ ready, or wrong port/host
│ │ │ ├─ "KeyError / missing key" → .env not loaded → --force-recreate
│ │ │ └─ Python traceback → your code, not Docker. Fix, then --build
│ │ └─ Not listed at all → the build failed. Scroll up in the build output.
│ └─ YES → keep going ↓
├─ Can I reach it from my Mac (curl localhost:PORT)?
│ ├─ NO → check three things, in order:
│ │ 1. Is there a -p / ports: line at all? (no ports = private, by design)
│ │ 2. Is the LEFT number the one I'm curling? (host:container)
│ │ 3. Does the app bind to 0.0.0.0, not 127.0.0.1?
│ └─ YES → keep going ↓
├─ Can container A reach container B?
│ ├─ NO → am I using the SERVICE NAME as host (not localhost)?
│ │ am I using the INTERNAL port (not the published one)?
│ │ are both services in the same compose file?
│ └─ YES → keep going ↓
└─ Is my code change showing up?
├─ NO → you rebuilt? docker compose up -d --build
│ edited .env? docker compose up -d --force-recreate
└─ YES → it's a logic bug now. Now debug the app logic.
Top Mac errors and their fixes
- "Cannot connect to the Docker daemon" → Docker Desktop isn't open. Launch it, wait for the whale.
-
"port is already allocated"
→ find the culprit with
lsof -i :8000, or map another port:-p 8080:8000. -
App container missing from
ps→ it crashed:ps -athenlogs app. -
Platform warning (arm64/amd64)
→ add
--platform linux/amd64. -
Build very slow / disk full
→
docker system df, thendocker system prune. -
Edits not showing up
→ rebuild:
up -d --build. -
Stuck at
quote>→ Ctrl+C, re-run without the trailing comment.
🏋️ Practice challenges — do at least two
-
Warm-up
— Push the QuickBite ETA image to Docker Hub (
docker tag+docker push). -
Medium
— Replace ScalerGPT's retry loop with a compose
healthcheckon chroma +depends_on: condition: service_healthy. Write two lines on which approach you'd pick and why. - Hard — Add a weather tool to DeskBuddy by updating ONLY the tools service — no agent rebuild. This proves you understood microservices.
-
Advanced
— Multi-stage builds on all three projects; cut image sizes by 40%. Record
before/after from
docker images. - Bonus — Break something on purpose (wrong port, missing .env, edit without rebuild), then fix it using only the debug tree above. This is a fast way to remember the steps.
Self-test — 30 questions
30 questions and answers
-
Question: What is the difference between an image and a container?
Answer: An image is a sealed, read-only package built from a Dockerfile — the master tiffin. A container is a running instance of that image — one delivered tiffin. One image can produce many containers.
-
Question: You edited app.py and restarted the container. Nothing changed.
Why?
Answer: The container runs the image, and the image is a photograph taken at build time. Your edit isn't in it. Rebuild:
docker compose up -d --build(ordocker build+rm -f+run). -
Question: In
-p 8080:8000, which number is the container's?Answer: The right one. Format is host:container . Your Mac listens on 8080 and forwards to 8000 inside the container.
-
Question: Why must the CMD use
--host 0.0.0.0?Answer: 127.0.0.1 inside a container means "only accept traffic that originated inside this container." Port mapping then silently delivers nothing. 0.0.0.0 means "accept from any interface," which lets the mapped port work.
-
Question: Why does requirements.txt get copied before the rest of the
code?
Answer: Layer caching. Docker rebuilds a layer and everything below it when its inputs change. Dependencies change rarely, code changes constantly — so put the pip install above
COPY . .and you skip the slow install on most rebuilds. -
Question: What does EXPOSE actually do?
Answer: Almost nothing at runtime — it's documentation/metadata saying "this app listens here." Actually publishing the port is
-por the composeports:key. -
Question: Difference between RUN and CMD?
Answer: RUN executes at build time and its result is baked into a layer. CMD executes at container start and is the default process. A Dockerfile can have many RUNs and one effective CMD.
-
Question: Your container isn't in
docker ps. What are the next two commands?Answer:
docker ps -ato see it as Exited, thendocker logs <name>to read why it died.pshides the dead. -
Question: What is a volume, and what does it survive?
Answer: Docker-managed storage that lives outside the container's writable layer. It survives
stop,rm, andcompose down. It does NOT survivecompose down -vor an explicitdocker volume rm. -
Question: Where do you put an API key, and where must you never put it?
Answer: Put it in a
.envfile loaded at runtime viaenv_file:. NeverENV KEY=...in a Dockerfile —docker historyexposes it — and never commit .env to git. -
Question: The app says "connection refused" to chroma even though
depends_onis set. Why?Answer:
depends_ononly waits for the container to start , not for the service inside to be ready . Fix with a retry loop in code, or a composehealthcheck+condition: service_healthy. -
Question: Chroma is published as
"8001:8000". Which port does the app container use?Answer: 8000 — the internal one. Containers on the same compose network talk flat-to-flat. 8001 is the street gate, useful only from your Mac.
-
Question: How does one container find another by name?
Answer: Compose creates a private network and registers each service name in its DNS.
http://tools:7000resolves to the tools container. No IPs, ever. -
Question: A service in compose has no
ports:key. Is that a bug?Answer: No — it's a security feature. Without published ports the service is unreachable from outside the Docker network. Only sibling containers can call it. Internal services should look exactly like this.
-
Question: What's the difference between
docker stopanddocker rm?Answer:
stophalts a running container but keeps it (restartable, logs intact).rmdeletes the container object. Neither touches the image. -
Question: What does
docker compose downdelete, and what does it keep?Answer: Deletes containers and the network. Keeps images and named volumes. Add
-vand the volumes go too. -
Question: You edited .env but the container still uses the old value.
Fix?
Answer: Environment variables are read at container start.
docker compose up -d --force-recreateto build fresh containers with the new values. -
Question: What is .dockerignore for, and name three things that belong in
it?
Answer: It excludes files from the build context that
COPY . .would otherwise pull in. Typical entries:venv/,.git/,__pycache__/,.env, notebooks, large raw data. -
Question: VM vs container in one sentence?
Answer: A VM virtualises the whole machine and ships an entire OS (bungalow); a container shares the host kernel and isolates only the process and its filesystem (flat). Minutes vs milliseconds to start.
-
Question: Why
chromadb-clientinstead ofchromadbin requirements?Answer: The full package includes the database engine itself — heavy. The thin client only talks to a Chroma running elsewhere (its own container). Smaller app image, cleaner separation.
-
Question: What are the three steps of RAG?
Answer: Retrieve the most relevant chunks from a vector store, augment the prompt by pasting them in, generate the answer with the LLM constrained to that context.
-
Question: Why does the agent need Redis rather than a Python dict?
Answer: A dict lives in the container's memory and dies with it. Containers restart constantly. Redis + a volume keeps conversation history across crashes, restarts, and redeploys — and lets you scale to multiple agent replicas.
-
Question: Describe the agent loop in one breath.
Answer: Send conversation + tool menu to the LLM; if it returns no tool call, you're done; otherwise execute the requested tool, append the result to the conversation, and loop again — capped at N iterations so it can't spin forever.
-
Question: Why cap the agent loop at 5 iterations?
Answer: A confused model can call tools forever, burning tokens and money. The cap is a safety fuse — the same reasoning as a timeout or a circuit breaker.
-
Question: What does
restart: unless-stoppedbuy you?Answer: Docker restarts the container automatically after a crash or a host reboot, but respects a deliberate
docker stop. Cheap resilience for free. -
Question: Two containers, same compose file. One says
localhost:8000to reach the other. What happens?Answer: It fails. Inside a container,
localhostis that container itself. You must use the other service's name as the hostname. -
Question: What does
docker exec -it name bashgive you, and when do you use it?Answer: A shell inside a running container. Use it to inspect files, check env vars, or confirm what actually got copied in — the quickest way to answer "is my file even in there?"
-
Question: Your Mac says "port is already allocated". Two ways out?
Answer: Find and stop whatever holds it (
lsof -i :8000), or publish on a different host port:-p 8080:8000. The container side doesn't need to change. -
Question: Why does an M-series Mac sometimes need
--platform linux/amd64?Answer: Apple Silicon is ARM64; many published images are built for AMD64 only. The flag runs the image under emulation — slower, but it works when no ARM variant exists.
-
Question: One sentence: why does Docker matter for ML and AI work
specifically?
Answer: Because ML/AI stacks are dependency nightmares (CUDA, sklearn, torch, vector DBs) and multi-service by nature — Docker makes the whole stack reproducible and identical from laptop to EC2, which is the difference between a notebook and a product.
Interview questions you can now answer
Read each question. Say your answer out loud before reading the notes. Recall builds memory better than re-reading. Anything you get wrong, go back to that Port.
These come up in real ML/AI engineering interviews. If you can answer eight of these cleanly, this session did its job.
-
Explain the Docker build cache and how you'd optimise a slow Dockerfile.
Talk about layers,
ordering,
COPY requirements.txtfirst,--no-cache-dir, and multi-stage builds. -
How do you handle secrets in a containerised app?
Runtime injection via env vars / secret
managers,
never
ENVin the Dockerfile,.envin both ignore files, anddocker historyas the attack you're preventing. -
Your service depends on a database that takes 20s to boot. How do you handle startup
ordering?
depends_on isn't enough; use healthchecks with
condition: service_healthy, or application-level retry with backoff. Mention that retry-in-code is more portable across orchestrators. - How do containers discover each other? Compose/Kubernetes DNS by service name, internal ports, and why hardcoding IPs is wrong.
- Container vs VM — when would you still choose a VM? Different kernel required, strong hardware-level isolation for untrusted workloads, or legacy OS dependencies.
- How do you persist state in a stateless system? Volumes and external stores; containers are cattle, not pets.
- How would you make this compose stack production-ready? Healthchecks, resource limits, non-root user, pinned image digests, logging driver, secrets manager, and moving to an orchestrator.
- What's in your image that shouldn't be? Build tools, .git, credentials, test data, dev dependencies — and how multi-stage builds and .dockerignore fix it.
-
How do you debug a container that exits immediately?
ps -a→logs→ check the CMD → run it interactively with an overridden entrypoint. - Why doesn't your internal service publish a port? Attack surface. Only the edge service is reachable; everything else is on the private network.
📚 Glossary — the vocabulary you're now expected to use
- Image
- A read-only, layered package of your app plus everything it needs to run.
- Container
- A running instance of an image, isolated from the host and from other containers.
- Layer
- The filesystem diff produced by one Dockerfile instruction. Cached and reused across builds.
- Build context
-
The folder you pass to
docker build(the lonely.). Everything in it, minus .dockerignore, is sent to the Docker engine. - Registry / Docker Hub
- The remote store where images are pushed and pulled from.
- Tag
-
The human-readable
name:versionlabel on an image. - Daemon
- The background engine that actually does the work. Docker Desktop starts it; the whale icon means it's alive.
- Port publishing
-
Mapping a host port to a container port so the outside world can reach in
(
-p host:container). - Volume
- Docker-managed storage that outlives containers. For anything you can't afford to lose.
- Bind mount
- Mapping a host folder straight into a container. Great for live-reload in development, avoided in production.
- Compose
- A tool that reads one YAML file and runs a whole multi-container application.
- Service
- One entry in a compose file — and also the DNS name other containers use to reach it.
- Healthcheck
- A command Docker runs periodically to decide whether a container is actually ready, not just started.
- Multi-stage build
- Using one image to build and a second, smaller one to run — the standard way to shrink production images.
- Orchestrator
- The system that runs containers across many machines (Kubernetes, ECS). Compose is the single-machine version of the same idea.
🎬 The one line to remember
"Writing code is half the job. Making it run anywhere in the world — that's engineering. Docker is the box that carries your work to the world."
Where to go next: Kubernetes — for when you have 10,000 boxes to manage instead of three. Everything you learned here (images, ports, volumes, service names, readiness) maps directly onto it. You already understand the core ideas. 🐳