Enroll in our TKTIS Academy.

TerenceNot sure where to begin? Terence knows the full catalogue and helps you choose.
Find your path
The track6 courses · 12 days

Microsoft Agents.

The Microsoft stack end to end: stand up Foundry, ground it permission-aware with Foundry IQ and Azure AI Search, design with the patterns, then evaluate, ship secure MCP tools and a production-ready agent frontend.

View track →
The track6 courses · 12 days

Claude Agents.

The Claude stack end to end: choose your runtime, from Messages API and Agent SDK to Managed Agents, ground it with pgvector and hybrid search, design with the patterns, then evaluate, ship secure MCP tools and a production-ready agent frontend.

View track →
The track4 courses · 7 days

Copilot Studio Agents.

Building agents low-code on Microsoft 365: from a first agent through knowledge, actions, MCP and flows to LLM-judged evaluation, analytics, ALM and governance.

View track →
The track4 courses · 6 days

Ship Faster with Claude.

Start at the user interface, tune Claude Code into an instrument, put reviews on every diff, and finish with a spec-driven harness with validation and feedback gates, the way we build ourselves.

View track →
The track5 courses · 7 days

Ship Faster with GitHub Copilot.

Start at the user interface, tune Copilot into an instrument across IDE, CLI and the app, delegate issues to agents from mission control, put reviews on every pull request, and finish with a spec-driven harness on Spec Kit, Actions and Agentic Workflows.

View track →
The track3 courses · 3 days

Leading with AI.

Choose, control, redesign: where AI creates value, under which conditions it may act, and how work, roles and accountability change. You bring one process of your own; it travels through all three days.

View track →
The track4 courses · 4 days

AI Governance & Enablement.

A light but workable governance function: make AI use visible, controllable and responsibly scalable through register and triage, controls and agent authority, vendors and data flows, and operating model and literacy, without creating a paper tiger.

View track →
The track2 courses · 2 days

Microsoft Copilot at Work.

Use Microsoft Copilot effectively, critically and safely, then delegate bounded work to Copilot Cowork without losing control of data, actions and quality. Two spread days; between them you collect one recurring task from your own role.

View track →
The track2 courses · 2 days

Claude at Work.

Use Claude effectively, critically and safely for knowledge work, then delegate bounded work to Claude Cowork with clear boundaries around files, connectors, actions and quality. Two spread days; between them you collect one recurring task from your own role.

View track →
The track5 courses · 10 days

AI Systems Architecture.

Translate a business problem into a defensible end-to-end AI architecture: master the building blocks first, from rules and optimization to RAG and agents, then decide, combine and design to production. One running case defended before a review board, vendor-neutral throughout.

View track →
The track4 courses · 8 days

Microsoft Fabric Data Engineering.

Design, build and operate production data products in Microsoft Fabric: OneLake, Lakehouse, Warehouse, Data Factory, Delta Lake and Real-Time Intelligence, from architecture and ingestion to controlled production operation. Loading one source into bronze is an early lab, not the end result.

View track →
The track4 courses · 8 days

Databricks Lakehouse Engineering.

Design, build and operate governed lakehouse data products with Databricks, Delta Lake, Unity Catalog and the Lakeflow platform: managed ingestion, CDC, batch and streaming pipelines, orchestration, governance and cost monitoring. Taught with current product names and patterns, not the curriculum of several years ago.

View track →
The track7 courses · 14 days

Applied Data Science.

Develop reliable predictive and forecasting models, from raw data to validated, explainable and governed model assets. Baselines first, leakage discipline throughout, and current tools without tying the competence to one library.

View track →
The track1 courses · 2 days

Practical AI Platform Engineering.

Choose, deploy, secure and operate a practical AI platform: one managed provider and one private model behind a single governed gateway, with keys, budgets, fallbacks, observability and a runbook included. Private hosting is one option, not the default answer.

View track →
The track4 courses · 8 days

Mathematics for AI Research.

The mathematical language of modern AI in four courses: linear algebra, calculus and automatic differentiation, probability and statistics, and information theory, each taught through the models they explain, from attention and backpropagation to sampling and loss design.

View track →
The track3 courses · 6 days

Deep Learning Foundations.

Neural networks built, trained and diagnosed from first principles: backpropagation by hand before frameworks take over, optimization and training dynamics as one coherent system, and architectures compared on the structure they model. One experiment discipline runs through the track: baseline, controlled change, measurement, ablation, conclusion.

View track →
The track7 courses · 14 days

Language Models & Frontier Model Research.

Language models from tokenization and the transformer block to frontier scale: efficient attention and state-space memory, sparse Mixture-of-Experts, pretraining data and optimization, post-training and reasoning, closing with a research lab that reconstructs and partly reproduces a current frontier architecture. The current capstone edition is The Road to K3.

View track →
Practitioner2 days

Microsoft Foundry Agent Platform: Fundamentals.

Stand up Microsoft Foundry properly and deploy a bounded agent: models, Agent Service, toolboxes, identity, tracing and evaluations, in the portal or from the SDK.

Explain the current Foundry resource and project model, and select and deploy the right model for the job
Pro or no code · Microsoft AgentsView course →
Practitioner2 days

Foundry IQ & Azure AI Search: RAG Fundamentals.

Grounded agents on the Microsoft stack, built twice: a direct Azure AI Search index tool against Foundry IQ knowledge bases, permission-aware, cited and measured.

Distinguish search, RAG and agentic retrieval, and reject RAG when something simpler wins
Pro code · Microsoft AgentsView course →
Practitioner2 days

Claude Agent Platform: Fundamentals.

The Claude stack is not one agent builder: Messages API with your own loop, the Agent SDK, or Managed Agents. Build a bounded agent and make the runtime choice deliberately.

Position the Claude platform components and choose a model on quality, latency and cost
Pro or low code · Claude AgentsView course →
Practitioner2 days

Supabase pgvector & Hybrid Search: RAG Fundamentals.

Permission-aware RAG on the open stack: Postgres full-text plus pgvector, fused with RRF and a reranker, Row Level Security end to end, and the judgment to know when RAG is the wrong tool.

Design an ingest and embedding pipeline: parsing, chunking, metadata, deduplication, updates and re-embedding
Pro code · Claude AgentsView course →
Foundations1 day

Copilot Studio: The Fundamentals.

Build, test and publish your first Copilot Studio agent in a day: an Agent Brief, instructions, topics, knowledge, all live in Teams and Microsoft 365 Copilot before you leave.

Position Copilot Studio against Microsoft 365 Copilot, Copilot Cowork and Microsoft Foundry
Low code · Copilot Studio AgentsView course →
Practitioner2 days

Copilot Studio: Knowledge, Actions & Flows.

Where Copilot Studio agents earn their keep: permission-aware knowledge, actions through connectors and MCP, deterministic agent flows, and event triggers kept safely in draft-and-approve.

Select, describe and secure knowledge sources, with permission-aware access and citations checked
Low code · Copilot Studio AgentsView course →
Practitioner2 days

Copilot Studio: Evaluation, Analytics & Governance.

From demo to managed production: LLM-judged test sets with Agent Evaluation, the Copilot Agent Kit, transcripts and Application Insights, solutions and ALM, DLP, and a release decision you can defend.

Formulate measurable success criteria and failure modes, and build single-response and conversation test sets
Low code · Copilot Studio AgentsView course →
Practitioner2 days

Agentic Design Patterns.

Every reliable agent is built from the same twenty-one patterns. This course walks Antonio Gulli's Agentic Design Patterns end to end and applies them to a workflow of your own.

Decide when a task calls for a single prompt, a fixed workflow, or a genuine agent, and justify the choice
No code · Microsoft AgentsView course →
Mastery2 days

Evaluation and Observability for LLM and Agent Systems.

The flagship systems, CI, and online-eval course: instrument with tracing, curate datasets, score with calibrated judges, wire a CI eval gate, and monitor live traffic. Course-local eval elsewhere stays narrowly scoped; this owns the production loop.

Design a layered evaluation strategy (deterministic assertions, LLM-as-judge, production sampling) and choose the right layer per failure mode
Hands-on · Microsoft AgentsView course →
Advanced2 days

Model Context Protocol: Building Secure Remote Tool Servers.

Build and deploy a secure stateless remote MCP server on the 2026 spec, hardened, least-privilege, and shipped as a Cloudflare Worker, so any AI client can call your tools.

Explain what MCP is, the problem it solves, and decide when a standards-based server is the right choice versus a direct SDK tool
Hands-on · Microsoft AgentsView course →
Advanced2 days

Building Agent Experiences with AG-UI and CopilotKit.

Build a production agent experience on AG-UI with CopilotKit: streaming, generative UI, shared state, and human-in-the-loop on high-impact actions.

Explain the AG-UI protocol and its event model, and why an open agent-to-UI standard matters across different backends
Hands-on · Microsoft AgentsView course →
Practitioner1 day

Claude Design to Claude Code: The Frontend Handover.

Start at the user interface: design the flow in Claude Design, make every state explicit in a handover contract, then sync it natively into Claude Code for controlled implementation.

Use Claude Design as a design space with your own design system and tokens as constraints
Pro code · Ship Faster with ClaudeView course →
Advanced2 days

Claude Code: The Toolbox.

Claude Code as a tuned instrument instead of a chat box: CLAUDE.md and rules, skills and plugins, hooks, subagents, MCP, permissions and sandboxing, worktrees, /goal and /loop, headless and routines. Every object, put to work.

Structure repository context with CLAUDE.md, scoped rules and auto memory, and keep it hygienic
Pro code · Ship Faster with ClaudeView course →
Advanced1 day

Code Reviews with Claude.

Reviews that scale with the code you now produce: /code-review on every diff, managed review on every pull request, /security-review on demand, standards that get enforced, and a sharp line around what stays human.

Set up local review with /code-review and managed review on pull requests, matched to your plan
Pro code · Ship Faster with ClaudeView course →
Mastery2 days

The Spec-Driven Coding Harness.

The capstone: a spec-driven harness where a feature moves from contract and plan to implementation, independent validation and human feedback gates, the way we build software ourselves.

Design a spec as a lasting contract, with requirements traceable to tests and code
Pro code · Ship Faster with ClaudeView course →
Practitioner1 day

Claude Design to GitHub Copilot: The Frontend Handover.

Start at the user interface: design the flow in Claude Design, capture it in a versionable handover contract in the repository, and let GitHub Copilot implement it. Controlled, verifiable, reviewable.

Explore a frontend flow fast in Claude Design, with your design system and existing components as constraints
Pro code · Ship Faster with GitHub CopilotView course →
Advanced2 days

GitHub Copilot: The Toolbox.

GitHub Copilot as a tuned instrument across IDE, CLI, GitHub.com and the app: instructions, AGENTS.md, prompt files, skills, custom agents, hooks, MCP, Spaces and Memory. Every object, put to work.

Choose the right Copilot surface per task and read the feature support matrix
Pro code · Ship Faster with GitHub CopilotView course →
Advanced1 day

GitHub Agents: Mission Control.

Stop typing along, start delegating: parallel agent sessions started from issues, prompts or the CLI, steered from the Agents page and the Copilot app, with automations and governance from prompt to pull request.

Start sessions from prompts, issues, PRs, the IDE and /delegate, and pick local or cloud execution deliberately
Pro code · Ship Faster with GitHub CopilotView course →
Advanced1 day

Code Reviews with GitHub Copilot.

Reviews that scale with the code you now produce: Copilot on every pull request, standards in instructions that get read, Fix with Copilot on findings, and the deterministic gates that actually block a merge: CodeQL, Code Quality, rulesets.

Set up local /review and manual or automatic PR review, with review effort tuned
Pro code · Ship Faster with GitHub CopilotView course →
Mastery2 days

The Spec-Driven Coding Harness on GitHub.

The capstone: a spec-driven delivery harness on GitHub, with Spec Kit from constitution to tasks, coding agents implementing, deterministic Actions and Agentic Workflows correctly separated, and rulesets and human feedback gates on the way to merge.

Grow a constitution and feature specs with the full Spec Kit flow, traceable to tasks and issues
Pro code · Ship Faster with GitHub CopilotView course →
Leadership1 day

AI Strategy & Portfolio Decisions.

Most AI initiatives never reach measurable value. Pick the ones that will: the right form of AI per problem, a ranked opportunity portfolio, and a Pilot Charter with stop-or-scale criteria fixed up front.

Explain why most AI initiatives stall before measurable value, and which management decisions separate the ones that do not
No code · Leading with AIView course →
Leadership1 day

AI Governance, Risk & the EU AI Act.

The AI Act is only part of the risk picture. Build governance as a business process: know what AI you run, your role as provider or deployer, and the controls, owners and evidence that make your approach defensible.

Place your organization in its AI Act role, provider or deployer, and know which obligations actually follow
No code · Leading with AIView course →
Leadership1 day

Designing and Leading Human-Agent Work.

Not how much work agents can take over, but what must stay human, and what can be delegated under which conditions. Author, Editor, Director and Orchestrator as a diagnostic, applied to one process of your own.

Decide per task what stays human and what can be delegated, and under which conditions
No code · Leading with AIView course →
Foundations1 day

Microsoft Copilot Essentials & Responsible Use.

Not a tour of the buttons: choose the right Copilot work mode per task, turn prompts into proper work briefings, work within existing access rights, and check what comes back, critically and safely.

Choose the right Copilot work mode: Chat, Search, the Office apps, Pages, Notebooks, Researcher, Analyst, agents or Cowork
No code · Microsoft Copilot at WorkView course →
Practitioner1 day

Microsoft Copilot Cowork & Workflow Automation.

Chat for an answer, Cowork for work: know when to delegate, brief with milestones and a final human review, automate low-risk tasks draft-and-approve, and control what runs in your name.

Judge when Cowork is the right tool, and when one answer, a deterministic workflow or no delegation at all fits better
No code · Microsoft Copilot at WorkView course →
Foundations1 day

Claude Essentials & Responsible Use.

Choose the right Claude work mode per task, whether Chat, research, Projects, Artifacts, connectors or Cowork, brief it properly, build context that lasts, and check output critically before it leaves your hands.

Choose between Chat, research, Projects, Artifacts, connectors and Cowork per task
No code · Claude at WorkView course →
Practitioner1 day

Claude Cowork & Workflow Automation.

Know when to delegate to Cowork, brief with boundaries around files, connectors and actions, schedule recurring work safely, and turn results into work products with clear sharing rules.

Judge when Cowork is the right tool, and when one answer, a deterministic workflow or no delegation at all fits better
No code · Claude at WorkView course →
Practitioner1 day

AI Inventory, Risk Triage & Lifecycle Controls.

Find all the AI in your organization, formal and informal, and move every application through one flow: discover, register, triage, assess, decide, deploy, monitor, change, respond, retire.

Discover formal and informal AI use, and register systems and use cases with owner, provider, model, data and integrations
No code · AI Governance & EnablementView course →
Practitioner1 day

AI Data, IP & Vendor Governance.

May this AI service be used for this data? Draw the full data flow, judge providers and deployment models, and decide with conditions instead of a flat yes or no.

Draw the full data flow of an AI application, from user to provider to downstream systems
No code · AI Governance & EnablementView course →
Practitioner1 day

AI Operating Model, Literacy & Adoption.

Governance as a service: an operating model with roles and decision rights, role-based literacy, and adoption that gets measured, because blocking alone breeds shadow AI.

Choose an operating model, central, federated or hub-and-spoke, and staff it with roles and decision rights
No code · AI Governance & EnablementView course →
Practitioner2 days

AI Technology Foundations.

No buzzword catalogue: for every technology you learn the input and output, the data it needs, how quality is measured, its typical failures, and when something simpler is better.

Explain per technology the input and output, required data, quality measurement, typical failures and needed infrastructure
Hands-on · AI Systems ArchitectureView course →
Practitioner2 days

Foundation Models, Retrieval & RAG Architectures.

RAG is not a button: from ingestion and chunking to reranking, context assembly and citations. Build a minimal vendor-neutral pipeline, measure it, and defend whether RAG is even needed.

Read foundation models at system level: context windows, inference, quality, latency, cost, routing, fallback, structured outputs
Hands-on · AI Systems ArchitectureView course →
Advanced2 days

AI Solution Discovery & Architecture Decisions.

Not every problem is an LLM problem. Decompose the business problem, compare solution categories on their merits, and record the decisions as ADRs, before a product choice dictates the problem definition.

Decompose a business problem: stakeholders, workflow, decisions, exceptions, outcome and baseline
No code · AI Systems ArchitectureView course →
Mastery2 days

Production AI Systems & Architecture Capstone.

Bring every choice together in a production-ready architecture, covering security, reliability, observability, scale, cost and governance, and defend it before a review board.

Design end to end: diagrams, data and control flows, trust boundaries, APIs, queues and integrations
Hands-on · AI Systems ArchitectureView course →
Practitioner2 days

AI Platform Engineering: Managed APIs, Gateways & Private Inference.

Most small teams don't need a GPU cluster. They need one controlled access layer. Put a managed provider and a private model behind a single gateway, with keys, budgets, fallbacks and a runbook.

Compare direct APIs, managed platforms, an internal gateway, private inference and hybrid setups, and decide when not to self-host
Hands-on · Practical AI Platform EngineeringView course →
Practitioner2 days

Fabric Data Architecture, OneLake & Delta Engineering.

One platform does not mean zero architecture. Decide between Lakehouse, Warehouse and Eventhouse, structure OneLake and your workspaces, and engineer Delta tables that hold up under real change.

Choose between Lakehouse, Warehouse, Eventhouse and hybrid designs based on workload, data shape and consumers
SQL and light Python · Microsoft Fabric Data EngineeringView course →
Practitioner2 days

Fabric Data Movement, CDC & Orchestration.

Getting one file into bronze is a demo. This course builds the ingestion framework around it: the right movement pattern per source, incremental loads and CDC, and pipelines that survive reruns, backfills and bad records.

Choose the right movement pattern per source: Shortcuts, Mirroring, Copy Job, Copy Activity, Dataflow Gen2, Eventstreams or notebooks
Hands-on · Microsoft Fabric Data EngineeringView course →
Practitioner2 days

Fabric Transformation & Data Product Engineering.

Silver is not just cleaned bronze. Build incremental, tested transformations with merge, SCD Type 2 and materialized lake views, and publish gold data products consumers can actually rely on.

Choose per transformation between code-first notebooks, SQL and declarative materialized lake views
Pro code · Microsoft Fabric Data EngineeringView course →
Advanced2 days

Fabric DataOps, Security & Real-Time Operations.

A data product that only runs on the happy path is a liability. Ship it through Git and deployment pipelines, lock down access, wire up monitoring, add a real-time flow, then break it on purpose and recover.

Deploy Fabric items through Git and deployment pipelines across dev, test and production, with variable libraries and service principals
Hands-on · Microsoft Fabric Data EngineeringView course →
Practitioner2 days

Databricks Lakehouse Architecture, Delta & Unity Catalog.

The Databricks you learned a few years ago is not the platform you will run today. Build the foundation on current patterns: Unity Catalog for governance, Delta engineered with liquid clustering and predictive optimisation, and a deliberate position on Iceberg.

Design a workspace, catalog, schema and environment structure and decide between managed and external data and storage
SQL and light Python · Databricks Lakehouse EngineeringView course →
Practitioner2 days

Lakeflow Connect, Auto Loader & CDC.

Ingestion is where lakehouse projects quietly fail: schema drift, duplicates, late events and half-loaded tables. Build it on the current services instead: Lakeflow Connect for managed sources, Auto Loader for cloud files, and AUTO CDC in place of legacy APPLY CHANGES.

Decide per source between Lakeflow Connect managed connectors, Auto Loader and a custom path, cloud and region availability included
Hands-on · Databricks Lakehouse EngineeringView course →
Practitioner2 days

Lakeflow Pipelines & Data Product Engineering.

Delta Live Tables is not the name anymore, and APPLY CHANGES is not the API. Build declarative batch and streaming data products with Lakeflow pipelines: streaming tables, materialized views and expectations, with orchestration handled by the platform.

Model a pipeline declaratively: flows, streaming tables, materialized views, views and sinks in SQL and Python
Pro code · Databricks Lakehouse EngineeringView course →
Advanced2 days

Lakeflow Jobs, DataOps & Production Governance.

A pipeline that runs in a notebook is not a production service. Ship it through Declarative Automation Bundles, orchestrate it with Lakeflow Jobs, watch cost and health in system tables, and prove you can recover when it breaks.

Orchestrate production workflows with Lakeflow Jobs: task graphs, triggers, continuous jobs, retries and alerts
Hands-on · Databricks Lakehouse EngineeringView course →
Practitioner2 days

Data Preparation, Statistical Foundations & Pipelines.

Data science is not a sequence of calls to fit() and predict(). Before any model, you need a sound question, honest splits and a baseline worth beating: that discipline starts here.

Translate a business question into a measurable prediction problem with a target, unit of analysis, prediction window and acceptance criteria
Python · Applied Data ScienceView course →
Practitioner2 days

Regression, Classification & Probabilistic Models.

Linear and logistic regression are not warm-up exercises: they are the transparent benchmark every ensemble must beat. Learn what these models optimise and what cross-entropy actually means.

Fit and diagnose linear regression: residuals, assumptions, interactions and multicollinearity
Pro code · Applied Data ScienceView course →
Practitioner2 days

Trees, Ensembles & Gradient Boosting.

On tabular data, boosted trees are still the models to beat. Train forests, XGBoost and LightGBM properly, then judge them on calibration, latency and complexity, not one score.

Grow, prune and regularise decision trees and explain their instability
Pro code · Applied Data ScienceView course →
Mastery2 days

Model Selection, Cross-Validation & Decision Metrics.

A single accuracy number is how optimistic models reach production. Learn which split is valid when, which metric answers which question, and how calibration and thresholds turn probabilities into decisions.

Select K-fold, stratified, grouped, nested or temporal validation and explain when shuffling is invalid
Hands-on · Applied Data ScienceView course →
Advanced2 days

Feature Engineering, Unsupervised Learning & Anomaly Detection.

Unlabeled results are the easiest to oversell. Engineer features with leakage discipline, cluster and reduce dimensions with restraint, and detect anomalies without calling every rare case wrong.

Engineer domain-informed features: interactions, ratios, aggregates, frequency and recency, with online and offline consistency in mind
Pro code · Applied Data ScienceView course →
Practitioner2 days

Time Series Forecasting.

A forecast validated with a shuffled split is fiction. Model trend, seasonality and stationarity, backtest with rolling origins, and make foundation models earn their place against a seasonal naive baseline.

Frame a forecasting problem: horizon, frequency, known future variables, and the cost of over- and under-forecasting
Pro code · Applied Data ScienceView course →
Mastery2 days

Explainability, Responsible ML & Model Lifecycle.

A model nobody can explain, trace or monitor is not finished. Turn a validated model into a governed asset: SHAP explanations, subgroup assessment, a registered pipeline and a defensible go or no-go.

Create global and local explanations and compare permutation importance, PDP, ICE and SHAP
Hands-on · Applied Data ScienceView course →
Practitioner2 days

Linear Algebra for AI Research.

Attention is a matrix operation, embeddings are geometry and low-rank adaptation is a statement about rank. This course builds the linear algebra that lets you derive and check those concepts instead of following them as recipes.

Describe neural operations with vectors, matrices and tensors, and read broadcasting, batched operations and Einstein notation fluently
Mathematics and Python · Mathematics for AI ResearchView course →
Practitioner2 days

Calculus & Automatic Differentiation.

Backpropagation is not a framework function, it is the chain rule organised over a graph. Build your own autodiff engine and you will never read loss.backward() the same way again.

Derive gradients by hand with the product, quotient and chain rule, and check them numerically
Mathematics and Python · Mathematics for AI ResearchView course →
Practitioner2 days

Probability & Statistics for AI Research.

A benchmark difference means nothing until you know the variance across seeds. This course builds the probabilistic and statistical machinery to state, estimate and doubt results properly.

Formulate uncertainty with probability rules, Bayes' rule, random variables and the standard distributions
Mathematics and Python · Mathematics for AI ResearchView course →
Advanced2 days

Information Theory for Machine Learning.

Cross-entropy is not an arbitrary loss and perplexity is not a magic number: both fall out of a handful of information-theoretic definitions. This course derives them and shows where the intuition breaks.

Compute and interpret entropy, joint entropy and conditional entropy in bits and nats
Mathematics and Python · Mathematics for AI ResearchView course →
Practitioner2 days

Neural Networks from First Principles.

PyTorch will happily train a network you do not understand. Build the perceptron, the multilayer network and backpropagation yourself, verify every gradient, and only then let autograd take over.

Implement a perceptron and a multilayer network in NumPy without a high-level model API
Mathematics and Pro code · Deep Learning FoundationsView course →
Advanced2 days

Optimization, Training Dynamics & Generalization.

An optimizer choice is a claim, not a preference. Compare SGD, momentum, AdamW and Muon under fair baselines, break and repair gradient flow, and report results over multiple seeds.

Compare SGD, momentum, AdamW and Muon technically, including parameter groups and optimizer state
Mathematics and Hands-on · Deep Learning FoundationsView course →
Advanced2 days

Deep Learning Architectures & Representation Learning.

Every architecture is an inductive bias made concrete. Build CNNs, residual blocks, recurrent networks and autoencoders, probe what they learn, and see precisely why recurrence gave way to attention.

Build a CNN and explain convolution, receptive fields, parameter sharing and translation equivariance
Pro code · Deep Learning FoundationsView course →
Advanced2 days

Attention & Transformers from First Principles.

Every frontier model still runs on the operations of one 2017 paper. Build a transformer from first principles: derive attention, keep every tensor shape explicit, train a small language model and measure exactly what it costs.

Implement scaled dot-product attention, causal and padding masks, and multi-head attention with explicit tensor shapes
Mathematics and Pro code · Language Models & Frontier Model ResearchView course →
Mastery2 days

Transformer Depth, Residual Streams & Normalization.

A transformer only trains at depth because its residual and normalization design allows it. Follow information and gradients through deep networks, and put standard residuals, Pre-LN, Post-LN and Attention Residuals to a controlled test.

Explain the degradation problem, identity paths and the additive residual stream, from ResNet to the Transformer
Mathematics and Pro code · Language Models & Frontier Model ResearchView course →
Mastery2 days

Efficient Attention, State Space Models & Long Context.

Full attention pays a quadratic price for exact recall, and long context makes the bill visible. Compare linear attention, state-space models, MLA and the KDA hybrid, and measure which efficiency claims survive your own benchmarks.

Explain the quadratic limits of full attention: compute, memory, KV cache growth and decode bottlenecks
Mathematics and Pro code · Language Models & Frontier Model ResearchView course →
Mastery2 days

Sparse Scaling & Mixture-of-Experts.

A trillion-parameter model that activates a fraction of itself per token is a routing problem as much as a modeling one. Build a sparse Mixture-of-Experts, break its load balancing on purpose, and trace the sparse scaling path from Switch Transformers to DeepSeek and Kimi K3.

Explain the split between total and activated parameters and what sparse conditional computation buys
Mathematics and Pro code · Language Models & Frontier Model ResearchView course →
Mastery2 days

Language Model Pretraining, Data, Optimization & Scaling.

A base model is decided before the first gradient step: tokenizer, data mixture, compute budget and optimizer set the ceiling. Design the full pretraining chain and train a compact base model you can account for, token by token and FLOP by FLOP.

Train a tokenizer and analyse vocabulary size, special tokens and multilingual token fertility
Pro code · Language Models & Frontier Model ResearchView course →
Mastery2 days

Post-Training, Reasoning, Multimodality & Agentic Models.

A base model predicts tokens; instruction following, reasoning, vision and tool use are built in post-training. Follow the path from SFT through preference learning to reinforcement learning, with the Kimi K-series as the running research case.

Run a supervised fine-tuning experiment with chat templates, response masking and before-and-after evaluation
Mathematics and Pro code · Language Models & Frontier Model ResearchView course →
Mastery2 days

AI Research Methods & Frontier Model Architecture Lab.

Frontier reports make claims; research method decides which ones hold. Reconstruct the paper dependency graph behind Kimi K3, reproduce one component at reduced scale and defend your conclusion before a research review panel.

Dissect a frontier technical report from abstract and claims to appendices and limitations
Research capstone · Language Models & Frontier Model ResearchView course →