The AI Engineering Skills Map In Detail: Building and Deploying AI Applications
Andrew Ng's deep dive into the first AI engineering skill — what it really takes to build and deploy AI applications that work in production, from LLM foundations to evaluation-driven development.
Andrew Ng recently unpacked the first skill from his AI Engineering Skills Map — Building and deploying AI applications — in more detail. This is a continuation of the overview we mapped against the AI-First Talent framework. Where that post looked at how all four skills line up with how organizations hire, this one zooms in on the single skill Ng calls the foundation of working with AI: making unpredictable systems reliable.
Why AI applications are a different kind of software
The key difference between AI applications and non-AI software is that the former's output is less predictable. You don't know in advance what an LLM will output, or what predictions a supervised learning algorithm will make. Because of this uncertainty, building AI systems is a much more iterative process than building traditional software — it is harder to plan the process in advance. Skilled AI engineers repeatedly build a piece of software, examine it, and decide what to try next, taking a sequence of steps that are highly influenced by the intermediate results. Being able to skillfully decide what to do next allows you to create reliable software systems based on unreliable AI components.
Ng breaks this skill into six sub-skills. Each one matters on its own, but together they describe what it actually takes to ship AI that holds up in production.
LLM foundations
Understanding how large language models tokenize input and generate output allows you to understand when to count on them and when they may fail. It also lets you reason about when to use a multimodal model, how to make tradeoffs on what to include in the context window, cache hits, knowledge cutoff, reasoning effort level, sampling parameters, and when to use special features such as tool calling. These foundations help you choose the right model or mix of models and apply specialized techniques when needed, such as fine-tuning or self-hosting models.
Grounding models with data
LLMs require good input context to produce useful outputs. RAG using vector search was an early attempt to give LLMs relevant context, but the set of techniques for grounding models with data has grown significantly. You have to decide what to include in a prompt vs. what to let an LLM retrieve on demand using tools, and which representation fits the data and search queries: a vector index, a knowledge graph, or a semantic layer over structured data such as customer records. You also turn documents (text, PDFs, HTML, images) into LLM-ready inputs and engineer pipelines to keep data clean and fresh. When you understand the menu of techniques available to get data, you are better able to give your LLM relevant context.
Building agentic systems
Agentic systems range from workflows that execute a predefined sequence of LLM calls to ones based on an agent harness that lets an LLM repeatedly decide its own next step. You choose the architecture — what steps to chain, what to parallelize, when to use code and when to use an LLM — and engineer the workflow or harness, with fallbacks. When designing the agent loop, you decide what tools the model can call (including MCP, CLI and sandbox execution environments), what memory architecture to use, how to manage context over long sessions, and when a task needs multi-agent orchestration instead of a single-agent architecture. You also turn promising prototypes into reliable, safe and secure agents for production, which requires understanding guardrails, adversarial inputs, and identifying and working around key risks such as data exfiltration, and governance.
Agentic workflows are evolving rapidly, and you also benefit from understanding any cutting-edge techniques relevant to your application area, such as voice agents, computer-use agents, or generative UI.
Evaluation-driven development
In Ng's experience, the most important trait that distinguishes someone great at building AI systems is whether they can drive a disciplined evals/error analysis loop to drive development. This allows you to repeatedly focus your effort on directions that are more likely to be fruitful. It's a tricky skill to master, because the right approach varies significantly by project and even according to the stage of the project.
Building good evals is a deep technical skill. You look at a system's traces and outputs, carry out exploratory data analysis, and combine that with product and business insight to decide what to measure. You understand the menu of options for evals — when to use deterministic (code-based) evaluations, when to use an LLM-as-a-judge, when to have a human in the loop — and how to evaluate your evals so as to keep evolving them. These evaluations then feed into an iterative process that drives further development, and makes progress systematic rather than random.
Operating in production
Operating AI software is different from traditional software because of its unpredictability, cost, and latency. You build observability mechanisms to understand the system's performance on real usage: track performance, detect drift, and respond quickly to model failures and security incidents such as adversarial prompt injections. Regression testing and CI/CD require more statistical evaluations than traditional software, and the testing effort should be calibrated relative to the risk of a mistake. You also select the right mix of techniques — model choice optimization, distillation and fine-tuning, agentic workflow simplifications — to optimize for cost and latency, especially if your application reaches many users.
Machine learning foundations
Modern LLMs are built using machine learning techniques including supervised learning and reinforcement learning. Every engineer Ng knows that is good at building with LLMs also understands machine learning and deep learning at some depth. Many applications still require knowing how to use machine learning — either a model someone else trained or one you train yourself. This requires knowing the popular machine learning and deep learning models and tradeoffs in accuracy, training speed, inference speed, and understanding how to engineer the data needed to train and evaluate these models. The machine learning concepts of bias/variance, error analysis, and engineering your data — all core mental frameworks for navigating how to work with systems with uncertain output — remain key to making a wide range of decisions in AI system development.
What this means for hiring
There is a lot to learn to become good at building and deploying AI systems, and this is a field with significant technical depth. But every bit you learn helps you become better at AI engineering and build more exciting applications. A strong complement to these skills is software engineering, which Ng will cover next.
For hiring teams, this detail matters because it shows why "AI fluency" can't be screened as a single checkbox. The six sub-skills above are what real AI engineering capability looks like — and most of them (evals, production ops, ML foundations, grounding) are craft skills built on top of solid engineering judgment, not tools you pick up in a weekend. This is exactly why the AI-First Talent framework screens for depth in the craft and business context first, then checks whether AI-native workflows have been layered on top. The tools amplify whatever judgment is already there; they don't substitute for it.
Read the original letter from Andrew Ng in full at The Batch.
Want this installed in your hiring process?
We run a 30-minute intro call to map your roles against the AI-first framework.
Book a 30-min intro callRelated reading
The AI Engineering Skills Map: Andrew Ng's Four Skills and What They Mean for Hiring
A complete guide to Andrew Ng's AI Engineering Skills Map — all four skills explained, how they map to the AI-First Talent framework, and how to screen for them.
AI Engineering: How to Land Your Dream Job (Using the Skills Map)
A practical playbook for landing an AI engineering job — built on Andrew Ng's four-skill map, with the portfolio, interview answers and signals hiring teams actually reward.
Software Engineering Fundamentals in the AI Engineering Skills Map
Andrew Ng's skills map puts fundamentals second for a reason. Here are the four interview prompts, a scoring rubric, and the answers that separate judgment from vibe-coding.