Skip to main content

Command Palette

Search for a command to run...

Agentic AI

Updated
β€’7 min readβ€’View as Markdown
Y

Developer | Adept in software development | Building expertise in machine learning and deep learning

Key benefits of Agentic Workflows

Agentic workflows achieve better performance because they give the LLM Tools, Memory, and the ability to Plan and Act, turning it into an iterative problem-solver instead of a one-shot guesser.

πŸ€– Agentic Workflow vs. Direct LLM Call (Simple Note)


Direct LLM Call (Simple)

  • What it is: The LLM acts like a calculator. You give it a single prompt, and it gives you a single, final answer.

  • Limit: It relies only on its internal training data and cannot fix its own mistakes or break a big problem into steps.

  • Best for: Simple, short-answer questions (e.g., "What is the capital of France?").


Agentic Workflow (Smart)

  • What it is: The LLM acts like an intelligent assistant. It uses a cycle of thinking, planning, acting, and checking.

  • Advantage: It can use external tools (like search or code) to get real-time data and break complex goals into smaller, manageable steps. It can also self-correct errors.

  • Best for: Complex, multi-step tasks that require real-world interaction, external data, or logical reasoning (e.g., "Plan a travel itinerary including real flight prices").

parallelization faster than human, which have to do sequentially.

Modular: can add or update tools, swap out models

What tasks suitable for Agentic AI

The following is easier for the Agent.

clear and step by step process

standard procedure to follow

text assets only

The following is hard for Agent

Steps not known ahead of time

Plan solve as you go

Multimodal sound and vision

Agentic Design Patterns

Without Reflection

Direct generation:

zero-shot prompting, no example given in the prompt.

One or two shot prompting, few shot

one(two) example given in the prompt. Few shot prompting, multiple examples given in the prompt.

Reflection

The agent examines its own output and figures out how to improve it. iteratively improve the output. consistently outperforms direct generation on a variety of tasks. shown that reflection improves the performances of direction generation on a variety of tasks.

Agentic design pattern perspective

Reflection loop usually goes:

  1. Generate answer

  2. Evaluate answer

  3. Revise answer based on evaluation

  4. Repeat if needed

Evaluation is step 2. Reflection is the whole dance.

Code Generation: using GPT 4o

Reflection: then using a reasoning model to do: you are an expert data analyst who provides constructive feedback on visualizations.

Step 1: critique the attached chart for readability, clarify and completeness.

Step 2: write new code to implement your improvements.

Evaluating the impact of reflection

  1. objective evals: code-based evals are easier, build a dataset of ground truth examples(human).

  2. subjective(another LLM) evals : use LLM as a judge, and rubric-based grading is better

using one model to generate, and using another model to Judge

Grading with a rubric gives more consistence results.

Sum up the scores

Using External Feedback

Patterns: No reflection, with reflection, and reflection with external feedback; increasing performance better and better. So when no reflection reaching performance bottleneck, then considering to introduce reflections, and then if it reaches the bottleneck again, then should consider to introduce using external feedback.

For instance: prompting: Write code for task x β†’ LLM β†’ code v1 β†’ execute code β†’ code output errors β†’ LLM (reflection on errors) β†’ code v2;

Tool Use

LLMs can be given tools, meaning that functions that they can call in order to get work done. For example, if you ask an LLM, what's the best coffee maker according to reviewers, and you give it a web search tool, then it can actually search the internet to find much better answers.

Simple tool execution:

Let LLM call functions, or more precisely, let an LLM request to call functions, that is what we mean by tool use, and the tools are just functions provided to LLM to call.

We leave LLM to decide to when appropriate to use tools. (does it mean we leave LLM a tool box, and it will choose among tools to solve a task or many tasks)

Planning

Rather than developer hard coding the sequence of steps in advance, this actually lets the LLM decide what are the steps to take.

Multi-agent collaboration

Just as a human manager might hire a number of others to work together on a complex project, in some cases it might make sense for you to hire a set of multiple agents, maybe each of which specializes in a different role, and have them work together to accomplish a complex task.

Multi-agent workflows are difficult to control since you don't always know ahead of time what the agents will do, but research has shown that they can result in better outcomes for many complex tasks, including things like writing biographies or deciding on chess moves to make in the game.

Practical Tips for Building Agentic AI

This lecture provides practical tips for building effective agentic AI workflows, emphasizing the need for an iterative, data-driven approach, primarily through the use of evaluations (evals).

Evaluation (evals)

The Role of Evals in Driving Improvement

  • Create Focused Evals: Once a common error mode is identified, create a small evaluation set (eval), often 10-20 examples, to specifically measure and track progress for that problematic aspect (e.g., date extraction accuracy).

  • Evals Track Progress: The eval metric provides an objective way to see if changes to prompts or system components are leading to improvement.

  • Iterate on Evals: Evals should also be iterated on over time. Start simple ("quick and dirty") and then collect more examples or change the evaluation method if the current eval fails to reflect human judgment about system improvement.

Examples of Error Modes and Evals

  • | Example Workflow | Identified Error Mode | Evaluation Strategy | Ground Truth | Evaluation Type | | --- | --- | --- | --- | --- | | Invoice Processing | Confusing invoice date and due date. | Objective Code Eval (RegEx matching a specific date format). | Per-Example Ground Truth: Manually annotate the correct due date for each invoice. | Objective, Per-Example GT | | Marketing Copy Assistant | Not adhering to the 10-word length limit. | Objective Code Eval (Word count function). | No Per-Example Ground Truth: The target (10 words) is the same for every example. | Objective, No Per-Example GT | | Research Agent | Missing important, high-profile discussion points. | LLM-as-a-Judge (Prompting an LLM to score how many pre-defined "gold standard" points are present). | Per-Example Ground Truth: Gold standard talking points differ for each research topic. | Subjective, Per-Example GT |

Workflow for Building Agentic AI Systems

Agentic AI systems are built through an iterative loop: start by creating a simple end-to-end version of the workflow, then inspect its outputs and build small evals to measure performance. Use these evals and error analysis to identify which components are failing, refine or replace those parts, and then re-test the entire system. By repeatedly cycling through building, analyzing, and improving, the system gradually reaches the desired performance metrics. This disciplined iterationβ€”rather than trying to perfect components upfrontβ€”is the key to developing effective agentic AI.

Andrew describes the process of building agentic AI systems as a continuous cycle between building and analyzing.

  1. Start with a quick end-to-end prototype
    Build a simple, possibly rough version of the whole workflow. This gives you immediate visibility into outputs and traces.

  2. Analyze early outputs
    Look at system traces and results to get an intuition for what’s working and what isn’t. This helps identify which components need improvement.

  3. Add small evaluation sets
    As the system becomes more stable, create small datasets (even 10–20 examples) to compute basic metrics and guide improvement.

  4. Perform structured error analysis
    When the system matures further, analyze which components frequently cause failures. This lets you focus on the highest-impact fixes.

  5. Introduce component-level evals
    For a mature system, build dedicated evaluations for individual components to improve them efficiently.

  6. Iterate β€” not linear
    The workflow loops: tune end-to-end β†’ error analysis β†’ fix components β†’ update component evals β†’ tune again.

  7. Use tools, but expect to build custom evals
    Although monitoring tools exist, most real workflows are custom, so tailored evaluations are usually necessary.

Andrew emphasizes that less experienced teams over-build and under-analyze. Skilled teams use disciplined analysis to focus their engineering time where it matters.


Agentic AI Workflow Diagram

             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ 1. Quick End-to-End    β”‚
             β”‚      Prototype          β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ 2. Inspect Outputs &   β”‚
             β”‚      Read Traces       β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ 3. Add Small Evals     β”‚
             β”‚  (10–20 examples)      β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ 4. Structured Error    β”‚
             β”‚       Analysis         β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ 5. Improve Components  β”‚
             β”‚   (Targeted Fixes)     β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ 6. Build Component-    β”‚
             β”‚     level Evals        β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚   Iterate & Loop Back  β”‚
             β”‚  (Not a Linear Process)β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

This diagram represents Andrew Ng's iterative cycle of building and analyzing agentic AI systems. Let me know if you'd like a colored version, more detailed flow, or a downloadable image.

Patterns for Highly Autonomous Agents