By the end, you will have shipped two AI apps.#
You do not need a PhD, advanced math or a machine learning background. In one sitting, you will learn how AI actually works. You will also build two real apps you can share: an AI Website Summarizer and a mini LLM Arena.
The promise. You start with curiosity. You finish with two working AI apps that you can share. That puts you ahead of 99% of people who only talk about AI.
Scaler Academy β each topic includes worked examples with real inputs and outputs.
Why AI Is Everywhere Now#
70 Years of AI in Short#
AI is not new. It is 70 years old. So why did it show up in your feed only recently? Two things changed. They changed what you can build.
- 1950sβ2000s β AI is a research lab thing. Decades of slow progress. You needed a PhD, a university, and years to do anything useful.
- 2017 β The βTransformerβ is invented. A new model design learns language very well. This is the engine inside GPT, Claude, Gemini.
- Nov 2022 β ChatGPT launches. The fastest-adopted product in history. Suddenly everyone can talk to AI.
- Now β you β anyone can build on top of it. Giant labs do the hardest part. You call their AI with a few lines of code. That is what you will learn here.
Why this matters now. Something that used to need a team + 6 months + a data warehouse can now be built by one person in an afternoon with an internet connection. This cost drop is why βAI Engineerβ is one of the fastest-growing jobs on the planet.
The Two Big Shifts#
Old AI did one task, such as spam-or-not. Today's AI does thousands β write, summarise, translate, code, answer β all from the same model. One brain, endless uses.
You do not build the AI. You use it through a simple API. It is like ordering food through an app instead of running a kitchen. If you can call a function, you can build with AI.
Teaching note. Make this personal: ask the room "what took you forever last week that an AI could've drafted in 30 seconds?" Everyone has an answer. That gap is the opportunity we are learning to fill.
What an LLM Is#
Super-autocomplete#
LLM = Large Language Model. In simple terms, it does one thing very well: it predicts the next chunk of text. Chat, code and summaries all build on that one trick.
Your phone suggests the next word as you type. An LLM is that idea scaled up a billion times and trained on much of the internet. Its guesses are good enough to write essays, code and answers.
First: Words Become Tokens#
A model does not read letters or full words. It reads tokens, which are pieces roughly ΒΎ of a word. It thinks in tokens and you pay per token. Here is a worked example.
AI Engineering is surprisingly approachable!
AI Engineering is surprisingly approachable!
Rare-word exampleunbelievableness becomes unbelievableness.
But why tokens β why not just whole words, or single letters?
This is an engineering trade-off. The model needs a fixed list of "things it knows" (its vocabulary). Whole words and single letters both cause problems. Tokens sit in the useful middle:
- β One token per wordThere are millions of words across all languages, plus names, typos, slang and new words every day. The vocabulary would be too large, and it would fail on a word it had never seen.
- β One token per letterThere are only ~100 symbols, so the vocabulary is small. But every sentence becomes hundreds of tokens. That is slow and expensive, and the model has to learn spelling from scratch.
- β Tokens (sub-words)A vocabulary of ~100k common chunks. Frequent words stay whole. Rare words are built from pieces. This is compact, and it can spell any new word.
Here is a concrete example. A rare word still works because it is assembled from familiar pieces:
How a rare word gets built:
"unbelievableness"βunbelievablenessFour known tokens are enough. Meanwhile
"the","cat","is"are each a single token. That is why tokens were chosen. One fixed vocabulary can handle every word in every language, even words invented tomorrow.
Why this matters when you build. English is "cheap" (βΒΎ word per token). Code, emojis, and many non-English scripts use more tokens per character β so the same sentence in Hindi or Tamil can cost noticeably more tokens than in English. Keep this in mind when you price an app.
Then It Predicts the Next Token Again and Again#
Given the tokens so far, the model gives every possible next token a probability. It picks one, adds it, and repeats. That loop is βgenerating textβ.
Low means the model usually picks the most likely word. The output is focused and repeatable. High gives surprising words more chance. The output is more creative and varied. The table shows the probabilities.
In this example, the model gives six possible next tokens these raw scores (logits): Python=3.2, JavaScript=2.4, Rust=1.6, public=0.9, spreadsheets=0.3, bananas=-1.2.
| Temperature | What the label means | Probability of each token |
|---|---|---|
0.20 | βοΈ focused | Python 98.2% JavaScript 1.8% Rust 0.0% public 0.0% spreadsheets 0.0% bananas 0.0% |
0.70 | βοΈ balanced | Python 67.8% JavaScript 21.6% Rust 6.9% public 2.5% spreadsheets 1.1% bananas 0.1% |
1.50 | π₯ creative | Python 42.7% JavaScript 25.0% Rust 14.7% public 9.2% spreadsheets 6.2% bananas 2.3% |
Sampling chooses one token using these probabilities. At low temperature, Python almost always wins. At high temperature, lower-ranked words get a real chance.
This is a real engineering choice. A support bot wants temperature β 0.2 (consistent). A brainstorming tool wants β 1.0 (varied). You set this early in any real project, including this one.
One More Fact: The Model Has No Memory#
Between API calls, an LLM forgets everything. Anything it should "know" in a conversation must be re-sent every time, inside a limited context window (measured in tokens). This fact explains a large part of AI engineering.
- π€ Tokens, not wordsIt reads, reasons and bills in tokens.
- π― Predicts next tokenAll abilities emerge from this loop.
- π‘οΈ TemperatureLow = focused, high = creative.
- πͺ No memoryYou resend context each time.
The Fill-in-the-Blank Trick#
Self-Supervision: Fill-in-the-Blank#
Here is the core idea. How do you teach a machine language without humans labelling billions of examples? You let it play fill-in-the-blank with the entire internet.
Humans hand-label data: "this email = spam", "this review = positive". Accurate, but you need millions of labelled examples. Very slow.
Take any sentence from the web, hide a word, and ask the model to guess it. The answer is already in the text, so the internet becomes its own teacher. No humans are needed. The model gets trillions of practice questions.
Here is the hidden-word example:
| Guess | Result | Feedback from the original example |
|---|---|---|
coffee | Correct | Exactly. To know this, the model learned context, grammar and facts about the world. That is self-supervision. |
tea | Close | Reasonable. βTeaβ also works, so the model learns that several answers can be good, with different probabilities. |
bicycle | Wrong | Not quite. Wrong guesses are part of how the model learns. |
sadness | Wrong | Not quite. After billions of these practice questions, the model gets very good at choosing likely words. |
Why this changes everything. The model does this billions of times. To guess βcoffeeβ, it has to learn grammar, context, what baristas do, what is hot and what fits in a cup. Understanding emerges as a side effect of getting very good at fill-in-the-blank. That is the main idea behind ChatGPT.
Keep it honest. Because it learned by predicting plausible text, an LLM can sometimes produce confident-sounding wrong answers β called hallucinations. A big part of our job as AI engineers is designing around that (e.g. feeding it real documents β we will see that later). Trust, but verify.
What You Can Build With AI#
Multimodal: Models With More Senses#
First, update your mental model. Modern AI is not just text. It can also see, hear and speak. We call these foundation models: one giant general-purpose model that you adapt to many jobs.
flowchart LR
T[π Text] --> M((π§ Foundation Model))
I[πΌοΈ Images] --> M
A[π€ Audio] --> M
V[π¬ Video] --> M
M --> O1[π¬ An answer]
M --> O2[ποΈ An image]
M --> O3[π A voice]
M --> O4[π A report]
The 8 Things People Build With AI#
Almost every AI product you have seen fits one of these groups. Find your idea here:
- π» CodingWrite, explain & fix code. The number 1 use case today.
Copilot Β· Cursor - βοΈ WritingEmails, blogs, marketing copy, rewriting.
Jasper Β· Notion AI - π¨ Image & VideoGenerate art, edit photos, make clips.
Midjourney - π EducationPersonal tutors that explain anything at your pace.
Khanmigo - π¬ ChatbotsSupport, sales & assistants that converse.
Intercom Fin - π₯ Info AggregationSummarise and search across large amounts of text.
β this project - ποΈ Data OrganizationTag, sort & structure messy information.
classification - βοΈ Workflow / AgentsAI that takes multi-step actions for you.
the frontier
Industry spotlight Β· same skill, different product. A coding tool, a legal-doc reader, a Swiggy-order chatbot β under the hood they follow the same pattern: connect to a model, give it the right context, get a useful answer, wrap it in a UI. Learn that once and each of these 8 becomes buildable.
Here we build a π₯ Info-Aggregation tool β an AI Website Summarizer. It is simple, useful and easy to share.
Where You Fit In#
What AI Engineering Means#
One line: AI Engineering is building real products on top of pre-trained models (GPT, Claude, Gemini) β without training those models yourself.
The car analogy. A Machine Learning engineer builds the engine. An AI Engineer builds the car around an engine someone else already built β wiring it to your data, your users, your problem, so it actually ships. You do not forge the engine. You drive.
Three Roles, Separated Clearly#
Trains models from raw data. Cares about datasets, GPUs, accuracy. "How do I create a model?"
Takes a powerful existing model and makes it a product. Cares about prompts, APIs, cost, reliability. "How do I use GPT to summarise 10,000 articles a day?"
The app, the buttons, the database. The AI Engineer is increasingly a software engineer who also speaks fluent "model".
Why this is accessible. Giant labs have already done the hardest and most expensive part: training the model. You can focus on the product: turning that model into something people use. Python + an API key is enough to begin.
Your First Six Lines That Talk to an AI#
Step 0 β Set Up Once (~10 min)#
Before any code runs, three quick bits of plumbing. Do not worry β it is a one-time setup.
β Install a code editor β Cursor or VS Code
You need one editor to write and run code. Both are free and work identically for this course β pick whichever you like:
VS Code is the world's most popular editor (by Microsoft) β rock-solid and widely used. Cursor is built on top of VS Code with an AI assistant baked in. If you want AI help while coding, pick Cursor; if you want the classic standard, pick VS Code. Either is totally fine β you only need one.
- Download & install. Cursor: go to
cursor.comβ Download. VS Code: go tocode.visualstudio.comβ Download. Install it like any normal app (Windows / Mac / Linux). - Add the Python & Jupyter extensions. Open your editor β click
the Extensions icon on the left β search
Python(by Microsoft) andJupyterβ Install both. (Same steps in Cursor and VS Code.) - Open a notebook & pick the kernel. Open a
.ipynbfile β click Select Kernel (top-right) β choose your Python environment. You run a cell with Shift + Enter.
β‘ Get your OpenAI API key
An API key is a secret password that lets your code use OpenAI's models (you pay only for what you use β cents for everything here). Here are the setup steps:
- Sign in to the OpenAI Platform.
- Add a small amount of credit under Billing.
- Create a new secret key.
- Copy the key immediately. OpenAI shows a secret key only once.
Sample key shape: sk-proj-A7mQ9rT2vX5nB8cD1eF4gH6jK0pL3sZy
The original generator used the prefix sk-proj- plus 32 random letters or numbers from A-Z, a-z and 0-9.
β’ Put the key in a .env file (never in your code)
Create a file literally named .env in your project folder, and paste your
key inside. Your code reads it from there β so the secret never appears in the code you
share.
# .env β keep this file private! Add it to .gitignore
OPENAI_API_KEY=sk-proj-xxxxxxxxxxxxxxxxxxxxxxxx
load_key.py β read it in your code
# pip install python-dotenv
from dotenv import load_dotenv
# β load secret values from .env into your environment
load_dotenv() # loads everything from .env
# now OpenAI() finds the key automatically β no key in your code π
π The habit that prevents expensive mistakes. Add .env to your
.gitignore so it is never uploaded to GitHub. Bots can find leaked keys in minutes and create real bills. Secrets live in .env, never
in code.
The Six Lines That Talk to the AI#
After setup, calling an LLM is an API call. It is like fetching the weather, but the reply is generated by a model. Here is the full example:
first_call.py
# pip install python-dotenv
from dotenv import load_dotenv
# pip install openai
from openai import OpenAI
# β load the API key from .env so the client can find it
load_dotenv() # loads everything from .env
# β‘ create the OpenAI client that will send chat requests
client = OpenAI() # reads your API key from the environment
# β’ ask the model with a system role and a user question
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are a witty travel guide."},
{"role": "user", "content": "Suggest one thing to do in Bangalore."},
],
)
# β£ print the assistant message from the first choice
print(response.choices[0].message.content)
# now OpenAI() finds the key automatically β no key in your code π
The Three Roles in Chat Models#
systemβ sets the personality and rules β "You are a careful tutor. Never give the full answer." Written once, applies throughout.userβ what the human asks. The actual request.assistantβ the model's reply. To continue a chat, append it back and resend the whole list (remember: no memory!).
- Your app sends the request. The request body is
{ "Bangalore?" }. - The model receives it. The model changes from idle to thinking.
- The service returns a status. The status is
200 OK. - Your app reads the assistant message. Sample response: βSip filter coffee on an MTR rooftop and watch the city wake up β Bangalore's best 10-minute holiday. ββ
π Security rule. Never paste your API key into shared or public code. It is a password to your wallet. Keep it in a .env file or environment
variable, on the server side only.
Industry spotlight Β· the model is swappable. Change one string: "gpt-4o-mini" β a Claude or
open-source model β and your whole product runs on a different engine. There's even
a free, runs-on-your-laptop option (Ollama) that uses this exact
same code, just pointed at a local address. You will use it soon.
Bonus: Use the Same Code on Groq + Llama#
If you do not want to add billing yet, Groq gives you a free API key and runs open models like Meta's Llama very fast. Because Groq speaks the same "language" as OpenAI, you change just two things β the key and one line β and everything else stays identical:
groq_call.py
# pip install openai (yes β the same library!)
import os
from openai import OpenAI
from dotenv import load_dotenv
# β load environment variables for the Groq key
load_dotenv()
# β‘ point the SAME client at Groq instead of OpenAI π
client = OpenAI(
api_key=os.getenv("GROQ_API_KEY"),
base_url="https://api.groq.com/openai/v1",
)
# β’ ask the Llama model the same travel-guide question
response = client.chat.completions.create(
model="llama-3.3-70b-versatile", # a free Llama model on Groq
messages=[
{"role": "system", "content": "You are a witty travel guide."},
{"role": "user", "content": "Suggest one thing to do in Bangalore."},
],
)
# β£ print the assistant message from the first choice
print(response.choices[0].message.content)
Where to get the key. Sign up free at
console.groq.comβ API Keys β Create key, then add it to your.envasGROQ_API_KEY=gsk_.... Notice that little changed β just the key, thebase_url, and the model name. This is a key idea in AI Engineering: the model is swappable.
Mini-Project: AI Website Summarizer#
How It Works β the Whole App in One Picture#
Now build the app. Our app: give it any web page URL β it reads the page and hands you a clean summary. Think of it as a short digest for the internet. It is simple, genuinely useful, and a good first project to share.
flowchart LR
U[π URL<br>user pastes it] --> S[π·οΈ Scrape<br>grab page text]
S --> P[π§© Prompt<br>text + instructions]
P --> L[π§ LLM<br>summarise]:::hl
L --> O[π Summary<br>shown to the user]:::good
Step 1 β Grab the Page Text (the "Scraper")#
π¦ Treat this one as a black box. You do not need to understand a single line below. This is plain web-scraping (not AI), and there's a library for it. All it does: take a URL β return the page's readable text as a string. That is it. Copy it, trust it, move on β the AI part is in Step 2.
# pip install requests beautifulsoup4
import requests
from bs4 import BeautifulSoup
HEADERS = {
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/120.0 Safari/537.36",
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.9",
}
def fetch_website_contents(url):
# β add scheme if the user forgot it
if not url.startswith(("http://", "https://")):
url = "https://" + url
# β‘ download the page and return a friendly error if it fails
try:
response = requests.get(url, headers=HEADERS, timeout=15)
response.raise_for_status()
except requests.exceptions.RequestException as e:
return f"Could not fetch the website. Error: {e}"
# β’ parse the HTML and keep the page title for context
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.string if soup.title else "No title found"
# β£ remove page chrome so the summary focuses on useful text
for tag in soup(["script", "style", "nav", "footer", "header", "img", "input"]):
tag.decompose()
# β€ turn the cleaned page into text for the model
text = soup.get_text(separator="\n", strip=True)
return f"Title: {title}\n\nPage contents:\n{text}"
Step 2 β The Brain (Prompt + LLM Call)#
from openai import OpenAI
from dotenv import load_dotenv
from scraper import fetch_website_contents
# β load the .env file so OpenAI can read the API key
load_dotenv() # <-- this reads your .env file
# β‘ create the model client once so summarize can reuse it
client = OpenAI()
# β’ define the website-summary instructions for the model
system_prompt = """You analyze the contents of a website and
give a short, friendly summary. Ignore navigation menus.
Respond in markdown."""
def summarize(url):
# β scrape the page so the model receives text instead of a URL
website = fetch_website_contents(url)
# β‘ send the scraped text to the chat model with the system prompt
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role":"system", "content": system_prompt},
{"role":"user", "content": f"Summarize this website:\n\n{website}"},
],
)
# β’ return only the assistant's written summary
return response.choices[0].message.content
πͺ The "wow" lever β change the personality. Edit one line of the
system_prompt β "give a snarky, humorous summary" or "explain
it to a 10-year-old" or "respond in Hindi" β and the whole app behaves differently.
That is prompt engineering, and you just learned it. The example below shows those personalities.
Steps 3 and 4 β Run It on Your Machine#
Step 3 β The entry point
A tiny main.py asks for a URL and prints the summary:
from summarizer import summarize
# β ask the learner which website to summarize
url = input("Website URL: ")
# β‘ summarize that URL and print the model response
print(summarize(url))
Step 4 β Run it on your machine
Save the three files (scraper.py, summarizer.py,
main.py) and your .env in one folder. Then open the terminal
in Cursor (Terminal β New Terminal) and run:
# 1. install everything you need (one time)
$ pip install openai requests beautifulsoup4 python-dotenv
# 2. run the app
$ python main.py
# 3. paste a URL when asked
Website URL: https://example.com
- Paste a URL and press Enter. Try a blog or news site. In a second or two, the summary prints in your terminal. You just used your own AI app.
- Stop it anytime. Press
Ctrl + Cin the terminal to shut the app down; runpython main.pyagain for another page.
If you hit an error, 90% of the time it is a missing pip install or a
key not loaded from .env β check those two first.
Example Run of the Working Version#
Below is a sample run of that app. It shows the sample sites, the personalities and the outputs.
| Preset | Full sample text |
|---|---|
| π Startup site | NimbusPay β Payments that just work. NimbusPay is a payments platform built for small Indian businesses who are tired of clunky tools and hidden fees. Accept UPI, cards, and wallets with a single integration that takes ten minutes to set up. Our flat 1% fee means no surprises at the end of the month. NimbusPay also gives you a real-time dashboard so you can see every transaction as it happens. Over 12,000 shops already use NimbusPay to get paid faster. We just launched instant settlements, so your money reaches your bank account the same day instead of waiting three days. Sign up today and your first month is completely free. |
| π° News article | Governments race to regulate AI as adoption surges. Lawmakers around the world are scrambling to write rules for artificial intelligence as the technology spreads into hospitals, banks, and classrooms. Supporters say clear regulation will build public trust and prevent harm. Critics worry that heavy rules could slow innovation and hand an advantage to larger companies who can afford compliance. A new draft framework focuses on transparency, requiring companies to disclose when content is AI-generated. It also demands that high-risk systems, such as those used in hiring or medicine, be tested for bias before launch. Industry groups have asked for more time to adapt. The debate is expected to continue for years as the technology keeps evolving. |
| βοΈ Personal blog | My journey from teacher to AI engineer. Three years ago I had never written a line of Python. I was a high-school teacher who felt stuck and curious about the AI everyone kept talking about. I started small: one tiny project every weekend, even when they barely worked. The first thing I built was a tool that summarized news articles for my students. It was ugly, but it worked, and that little win changed everything. I kept shipping projects and sharing them online, and slowly people started noticing. Last month I started my first job as an AI engineer at a startup. The lesson I keep repeating to anyone who will listen: you do not need permission or a perfect plan, you just need to build small things often and share them. |
| Tone | Intro line | TL;DR label |
|---|---|---|
| π Friendly | Here's the gist, in plain English: | In short: |
| π Snarky | Fine, I read it so you do not have to: | The honest TL;DR: |
| π§ Explain like I'm 5 | Okay, imagine I'm explaining this to a 10-year-old: | So basically: |
| πΌ Professional | Executive summary: | Bottom line: |
| Preset | Top sentences | TL;DR sentence |
|---|---|---|
| π Startup site |
| NimbusPay is a payments platform built for small Indian businesses who are tired of clunky tools and hidden fees. |
| π° News article |
| Critics worry that heavy rules could slow innovation and hand an advantage to larger companies who can afford compliance. |
| βοΈ Personal blog |
| I started small: one tiny project every weekend, even when they barely worked. |
The sample outputs come from a simple extractive summarizer, not a model. It normalizes whitespace, keeps sentences with at least four words, ignores common stop words, scores each sentence by word frequency divided by the square root of sentence length, returns up to three top sentences in original order, and uses the highest-scoring sentence as the TL;DR. If no text is supplied, it says βPaste some text or pick a sample first.β If the text is too short, it says βHmm, that text was too short to summarize.β
βοΈ This demo summarizes in your browser (a simple method) so
it runs with no API key. Your real summarizer.py sends the text to GPT for
a smarter summary β but the shape, the flow, and the personality switch are exactly
what you see here.
Industry spotlight Β· summarization is a real product category. Summarizing news, earnings calls, support threads, legal contracts, research papers, meeting transcripts β it is one of the most-used AI features in companies today. Swap "website" for "PDF", "email thread", or "YouTube transcript" and you've got a dozen more apps from the same code.
π₯ Bonus Build: an LLM Arena (One Prompt, Two Models, You Judge)#
Here is a second mini-project. Inspired by arena.ai, where people send one prompt to two AIs and vote on the better answer β that is how many people rank AI models. You can build a tiny version with what you already know: send the same prompt to two models, show both answers side by side.
arena.py β the whole idea
def battle(prompt):
# β wrap the learner prompt in the chat format both models expect
msgs = [{"role": "user", "content": prompt}]
# β‘ ask Model A β OpenAI's GPT
a = openai_client.chat.completions.create(model="gpt-4o-mini", messages=msgs)
# β’ ask Model B β Llama on Groq (same code, different brain!)
b = groq_client.chat.completions.create(model="llama-3.3-70b-versatile", messages=msgs)
# β£ return both answers so the app can compare them
return a.choices[0].message.content, b.choices[0].message.content
# β¦now we turn this into a tiny terminal app with π/π voting π
The full app with A/B voting
arena_app.py
# pip install openai python-dotenv
import os
from openai import OpenAI
from dotenv import load_dotenv
# β load API keys for both providers
load_dotenv()
# β‘ create one client for OpenAI and one for Groq
openai_client = OpenAI() # uses OPENAI_API_KEY
groq_client = OpenAI(api_key=os.getenv("GROQ_API_KEY"),
base_url="https://api.groq.com/openai/v1")
def ask(client, model, prompt):
# β send the prompt to whichever model client was passed in
r = client.chat.completions.create(
model=model, messages=[{"role": "user", "content": prompt}])
# β‘ return just the assistant text from the first choice
return r.choices[0].message.content
def battle(prompt):
# β ask OpenAI and Groq the same prompt
a = ask(openai_client, "gpt-4o-mini", prompt)
b = ask(groq_client, "llama-3.3-70b-versatile", prompt)
# β‘ return both answers so the learner can compare them
return a, b
def vote(label):
return f"π³οΈ Thanks! You voted: {label}" # in real apps, save this to a file/DB
# β’ read one prompt to send to both models
prompt = input("Ask both models the same thing: ")
# β£ collect both model answers
a, b = battle(prompt)
# β€ show both answers before asking for a vote
print("\nπ€ Model A:\n" + a)
print("\nπ€ Model B:\n" + b)
# β₯ collect the winner and show a thank-you message
choice = input("\nWhich answer was better? (A/B): ").strip().upper()
print(vote(f"π Model {choice}"))
Run it the same way
bash β your project folder
# install once, then launch
$ pip install openai python-dotenv
$ python arena_app.py
Ask both models the same thing: Explain recursion in 2 sentences.
- Type a prompt. Run the script, type any question, and press Enter β both models answer.
- Read both answers, then vote π / π. Type
AorBfor the answer you liked better β just like on arena.ai. - Record or share it. A short clip of the battle is perfect
for sharing. Stop anytime with
Ctrl + C.
Why this is a strong first project. It uses both keys you set up (OpenAI + Groq), proves the "brain is swappable" idea in practice, and the side-by-side A/B format is easy to understand. The example below shows the same idea.
The demo compares GPT-4o-mini from OpenAI with Llama-3.3-70b from Groq. The arena randomly puts GPT on the left or right and hides the names until you vote. It then adds +1 for a good vote or -1 for a bad vote. Both scores start at 0.
| Prompt | GPT-4o-mini answer | Llama-3.3-70b answer |
|---|---|---|
| Explain recursion to a total beginner, in 2 sentences. | Recursion is when a function solves a problem by calling itself on a smaller piece of the same problem, until it hits a simple βbase caseβ it can answer directly. Like a dictionary entry that points to another word β you follow the chain until you reach one you already know. | Sure! π Picture standing between two mirrors β you see a smaller copy of yourself, then a smaller one inside that, on and on. Recursion is just a function that keeps calling itself on a tinier version of the task until it is small enough to answer in one go. |
| Write a catchy tagline for a coffee shop. | βBrewed for the moments that matter.β Clean, premium, and it reads well on a cup or a storefront sign. | Ooh fun one! Try: βLife's short. The coffee shouldn't be.β β Punchy, a little cheeky, and very shareable. |
| Is a hotdog a sandwich? Give a definitive answer. | Definitive answer: No. A sandwich needs a filling between two separate pieces of bread; a hot dog uses a single, hinged bun β so it forms its own category. | Haha, the eternal debate! My verdict: NOT a sandwich. π The bun is one connected piece, not two slices β so the hot dog stands proudly in a class of its own. π |
| Give me a fun fact about space. | A day on Venus is longer than its year: it rotates so slowly that one Venusian day takes ~243 Earth days, while a full orbit of the Sun takes only ~225. | Here's a wild one π β Venus spins so slowly that a single day there lasts LONGER than its whole year! Your birthday could come round before the sun sets. Space is wonderfully weird. β¨ |
For other prompts, the sample answers describe each style: GPT-4o-mini is βtight and well-structuredβ; Llama-3.3 on Groq is βwarmer, chattier, and lightning fast.β After a vote, the arena shows: βIn a real arena, votes like yours pile up across thousands of people to build a public leaderboard.β
βοΈ The sample answers on this page are written examples.
In your real arena_app.py, the two answers come from the actual models β
and the voting is exactly how arena.ai builds its public leaderboard.
Industry spotlight Β· voting helps rank models. Side-by-side "blind taste tests" like this β millions of human votes on anonymous model pairs β are how the AI world decides which model is actually best, beyond marketing claims. Companies use the same technique internally to choose which model to ship. This small arena is a real evaluation method in miniature.
Make the Project Your Own#
Same 3-step recipe (input β prompt β output), different idea. Pick whichever feels fun β all are beginner-simple:
- βοΈ Email Subject LinerPaste an email β get 5 catchy subject lines.
- π CV β Cover LetterPaste your CV + a job ad β a tailored draft.
- π’ Company BrochureGive a company site β a fun marketing brochure.
- πΊ YouTube SummarizerPaste a transcript β the key takeaways.
- π³ Recipe FormatterMessy recipe text β clean steps & a shopping list.
- βοΈ Travel Planner"3 days in Goa" β a day-by-day itinerary.
π‘ These are drawn from real beginner projects in your course's community folder β proof that "simple + shipped" beats "complex + someday".
A Peek at Where This Is Going#
π LangChain β the Toolkit for Bigger Apps#
You've built an app that answers. The rest of this journey is about apps that do. Here's a taste of three things coming up β no need to master them today, just get excited.
Right now you call the API by hand. The moment you want reusable prompts, multi-step pipelines, memory, or to plug in your own documents, a framework like LangChain saves you re-inventing the wheel. Same idea as today β just with handy connectors:
langchain_taste.py
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
# β describe how the website text should be summarized
prompt = ChatPromptTemplate.from_template("Summarize this website: {content}")
# β‘ choose the chat model that will do the writing
model = ChatOpenAI(model="gpt-4o-mini")
# β’ pipe the prompt into the model to make a reusable chain
chain = prompt | model # "|" = send the prompt INTO the model
# β£ run the chain with one page's text
chain.invoke({"content": website_text}) # reuse it for any page!
"Chat with your own PDFs." You retrieve the relevant snippets from your documents and paste them into the prompt β so the AI answers from your data, not just its memory. It is the most common real-world pattern, and it tames hallucinations.
π€ Agents β When the AI Can Take Action#
An agent is an LLM given tools (a calculator, web search, your database) and a loop: it thinks, acts, looks at the result, and repeats until done. Step through one:
get_weather(city="Bangalore"){ rain_chance: 78%, temp: "24Β°C" }The leap. Nobody hard-coded "check the weather" β the agent decided to, because we gave it the tool and the goal. That autonomy is the jump from "chatbot" to "agent".
π₯ Multi-Agent β a Team of AIs#
For bigger jobs, companies split work across specialists that hand off to each other. This example shows a content team writing a short blog post:
| Agent | Role | Message from the original run |
|---|---|---|
| π§ Manager | Plans & delegates | Breaking goal into research β write β review. |
| π¬ Researcher | Finds the facts | Found: agents = LLM + tools + loop; used by Cursor, support bots. |
| βοΈ Writer | Drafts the post | Drafted a hook + insight + CTA. |
| π Editor | Polishes & checks | Tightened wording, approved. β |
π€ AI βagentsβ aren't sci-fi β they're just an LLM given tools and a loop.
That is how Cursor fixes code and support bots resolve tickets end-to-end.
The shift from chatbots β agents is the biggest change in software this year.
What will you let an agent do for you? π
π Where this path ends. By the end of the journey, you will build a solution where several agents collaborate to solve a real business problem. Companies hire for this work. It rests on the basics here: connect to a model, give it context, get something useful, ship it.
- π§ LLMsPredict next token; tokens, temperature, context.
- πͺ Self-supervisionLearned the internet via fill-in-the-blank.
- π API callsystem / user / assistant.
- ποΈ You shippedTwo real AI apps.