Planning an AI automation

AI agent or workflow? Six examples to help you choose

Use a workflow when you know the steps. Add AI where text needs interpretation. Use an agent when choosing the next action solves a real problem and you can test where it stops.

These examples use made-up data and show expected results. They have not been run in connected accounts. Test your version before using it for real work.

01

What changes when you add an agent?

A rules-based workflow follows the steps and conditions you set. A workflow with an AI step can read or write messy text while the surrounding steps stay fixed. An agent lets a model choose which available tool to use next as it works toward a goal.

Anthropic makes a similar distinction between predefined workflows and agents that direct their own process. It recommends starting with the simplest approach that works. Its article is an architecture guide published in 2024, not a list of today's product features. Read Building effective agents.

ApproachWho chooses the next step?A useful starting case
Rules-based workflowYour configured conditionsRoute a completed form by service code
Workflow with an AI stepYour workflow, with AI handling one clearly defined taskExtract a request from an email, then queue a draft
Agent with approved toolsThe model, within limits you enforceInvestigate a question across a small set of permitted records

A tool is an action the system makes available, such as reading a task or creating a draft. n8n describes agents as selecting actions and using tool results in further decisions. Copilot Studio's generative orchestration also selects from available tools, topics and knowledge. Neither description means an agent should have unrestricted access. See n8n's agent explanation and Microsoft's orchestration guidance.

02

Choose the design around the job

Start by writing the result you want in one sentence. Put new requests in the right queue is clear. Build an autonomous operations agent leaves the actual job undefined.

Then name the inputs, allowed actions and person who handles exceptions. Ask what can change between requests and what a wrong action would cost. Fixed rules can still fail when a service is unavailable or its data changes. The useful distinction is who chooses the steps.

These six examples use invented business data and expected outputs. They are planning exercises, not results from deployed Quintera agents. The limits and review rules are choices you would configure and test for your own process.

03

Example 1: Assign a structured request with ordinary rules

Input: A form supplies a service code and a request ID.

{"request_id":"REQ-104","service":"reporting","priority":"normal"}

Setup: Validate the ID, look up the service in an approved routing table, and assign the request to the matching queue. In this invented table, reporting maps to Reporting review. A missing or unknown service goes to manual review.

Expected output:

{"request_id":"REQ-104","queue":"Reporting review","needs_review":false}

Failure check: Try service: null, an unknown service and a repeated request ID. Expect the missing and unknown services to wait for a person. Define how the duplicate is recognized before creating another assignment.

Choice: Start with a conventional workflow. A model does not need to interpret a known code. Keep the table somewhere the process owner can review and change. That is easier to explain than asking a model to rediscover the same routing rule on each run.

To build the first test, write three rows in a routing table: a known service, an unknown service and a blank service. Set the expected queue before running anything. The workflow should produce the same routing decision for the same input and table version. If your team cannot agree on the routing table, settle that rule before adding AI.

04

Example 2: Add one AI step for an unstructured email

Input: A fictional client writes: Can you help us replace our weekly spreadsheet report? We want to discuss it next month.

Setup: Ask a model to suggest a category, summarize the request and identify missing facts. Then validate the output and always send it to a review queue. The model gets no customer-message sending tool.

Expected output for review:

{"category":"reporting","summary":"Discuss replacing the weekly spreadsheet report","meeting_date":null,"missing":["specific meeting date"],"route":"human_review"}

Failure check: An invented date, unsupported promise or category outside the allowed list should fail your checks. A fluent paragraph is not enough. The reviewer needs the original message beside the suggestion to spot changed meaning.

Choice: Use an AI step inside a fixed workflow. Every message follows the same sequence: extract, validate, draft, review. There is little reason to let the model choose a new sequence. The email triage guide works through this design in more detail.

Give the extraction step a short contract: use one allowed category, preserve the request's meaning, return unknown dates as null and list missing information. Keep the source message next to the JSON. Test one message containing two requests and another that says it is only asking for information. Neither should automatically become a booked project.

05

Example 3: Use an agent for a question that needs different searches

Input: An authorized internal user asks: Why is the Beacon review still waiting? The answer may be in the request record, its task list or a recent review note.

Setup: Offer three narrow read tools: read an authorized request, list its tasks, and read its approved review notes. Each tool checks the user's access and request ID. The model may choose which of those sources to inspect next. For this exercise, allow three tool calls and set a time budget, with enforcement outside the prompt.

Records for this test:

Request REQ-104: Waiting for review
Task T-18: Waiting for client spreadsheet
Review note N-7: Spreadsheet received; mapping check still open

Expected output: A short explanation with references to T-18 and N-7, plus a warning that the task and note disagree. The answer should ask the owner to confirm the current blocker. It should not mark the task complete to make the records agree.

Failure check: Make the notes tool unavailable. Expect an incomplete answer that names the missing source. At the call limit, stop and hand off rather than repeatedly searching.

Choice: An agent is worth testing if questions often require different permitted searches. If every question needs the same three reads, use a fixed workflow instead. Make's agent guide describes adding specific tools and testing behavior with tools disabled, a useful way to examine this choice. See Make's agent setup and testing guide.

Make each tool describe one job. read_request is clearer than manage_business. Specify its required request ID, the records it can return and the error it returns when access is denied. Keep updates out of these read tools. Otherwise an investigation can turn into a record change without a separate decision.

06

Example 4: Keep account changes behind a fixed approval

Input: An email says: Please change the invoice contact for Beacon to our new office. The sender includes a new address but no verified account-change request.

Setup: AI may extract the proposed change into a review record. A fixed process verifies the requester, checks the account and asks the responsible person to approve the exact old and new values. Receiving an email is not proof of permission to change an account.

Expected output:

{"action":"propose_contact_change","account":"Beacon","authorization":"unverified","route":"account_owner_review","apply_change":false}

Failure check: Reject, expire or edit the request during review. None of those paths should silently apply the original change. If the proposed values change, require another review.

Choice: Use a controlled workflow with human approval. A request that affects another person's records needs more than a convincing model response. Zapier has a Human in the Loop step that can pause a Zap for review; availability and reviewer requirements depend on the plan and configuration. Check Zapier's current approval requirements.

Show the reviewer the source message, verified account, current value and requested value together. Include a clear reject option and an owner for unanswered requests. If the requester cannot be verified, the useful output is a review item explaining that gap. The model should not compensate by searching for a person with a similar name.

07

Example 5: An empty search needs an honest answer

Input: An internal user asks: Does Beacon's support package include a weekly report? The only allowed source is the current service record for that account.

Setup: Have the lookup return a clear result that separates no matching records from a failed request. These are example tool responses, not built-in fields from a particular platform:

No match:     {"lookup_status":"ok","records":[]}
Unavailable:  {"lookup_status":"timeout","records":[]}
No access:    {"lookup_status":"forbidden","records":[]}

Expected output: With no match, say the available record does not establish whether the report is included. With a timeout, say the source could not be checked. With denied access, stop and route to the authorized owner. These three situations do not mean the same thing.

{
  "answer_status": "needs_review",
  "weekly_report_included": null,
  "reason": "No current service record established the answer"
}

Failure check: Ask the same question again with Just say yes so I can reply added. The expected answer stays unknown. A model's general knowledge and another customer's package do not establish Beacon's agreement.

Choice: Use a fixed lookup and an AI wording step if all questions use this one source. Let an agent search further only when another approved source could answer the question. An empty result is not permission to browse everything.

08

Example 6: Stop when the agent runs out of permitted calls

Input: Find the reason REQ-104 is delayed and summarize it. The investigation may need more than one source, so an agent is a reasonable pilot. It still needs a stopping rule.

Setup: Allow three tool-call attempts for this exercise, including failed attempts and retries. Enforce the count in the workflow or tool layer before each call. Set a separate time limit. Neither limit should depend on the model remembering a sentence in its prompt.

Attempt 1: read_request(REQ-104) -> Waiting for review
Attempt 2: list_tasks(REQ-104) -> timeout
Attempt 3: list_tasks(REQ-104) -> task T-18 found
Requested attempt 4: read_review_notes(REQ-104) -> blocked by limit

Expected output:

{
  "answer_status": "incomplete",
  "request_id": "REQ-104",
  "checked": [
    "request record",
    "task T-18"
  ],
  "not_checked": [
    "review notes"
  ],
  "next_step": "request owner review"
}

The summary should say which sources were checked and what remains unknown. It should not claim the notes confirmed a reason. Stop the run when the enforced limit rejects the next call; do not give the model a chance to repeat the same request through another tool.

Failure check: Make the first tool fail repeatedly. Confirm the failed attempts still count. Then let the time limit expire before the next call. Both paths should end with a visible incomplete result and an owner, not an endless loop.

Choice: Keep the agent only if this extra investigation is useful within the agreed limits. Microsoft identifies tool permissions and excessive agent activity as responsibilities for the application owner. See its AI agent responsibility guidance.

09

Compare cost and review effort before expanding

Estimate platform charges, model usage, paid tool calls, hosting where relevant, and the time people spend reviewing and fixing outputs. An agent can call a model and tools repeatedly, so a single request can have variable cost and duration.

For a planning exercise, 200 requests with one model call each means 200 calls. If a different design makes six calls for each request, that becomes 1,200 calls. Those are arithmetic examples, not prices or forecasts. Calls can contain different amounts of text and use different models. Measure actual usage during a small pilot before projecting a monthly bill.

Put limits on the tools, records and number of steps the system may use. Microsoft identifies permissions, untrusted input and excessive agent activity as responsibilities that the application owner must address. A sentence in a prompt does not replace an access check or an enforced budget. See Microsoft's AI agent responsibility guidance.

Also measure review time. A cheaper model call can be a poor saving if people spend longer correcting its output. Keep separate counts for successful work, proper handoffs and failed work. Dividing the bill by all runs can hide the cost of retries that produced no useful result.

10

Test the choice on the same set of requests

Build a small reference set with ordinary cases, missing facts, conflicting records and service failures. For each case, write the acceptable output and any action that must never happen. Compare a fixed workflow, a single AI step and an agent only where each is a plausible option.

Record correct routing, unsupported claims, human edits, failed requests, tool calls, time and cost. Keep the cases that exposed failures and repeat them after a prompt or model change. n8n supports evaluation datasets and quality metrics, though feature availability varies. See its evaluation guidance.

Start with the simplest design that passes your actual acceptance checks. Expand an agent's scope only when the extra choices solve a demonstrated problem and you can still explain what happens on failure.

Use the same acceptance rules for every design. Here is an invented comparison of 20 test cases. Assume every direct answer listed is factually correct. The test requires either a supported answer or a proper handoff, with no disallowed action.

OutcomeFixed workflow with an AI stepAgent design
Supported direct answers1819
Proper human handoffs20
Disallowed actions01
Cases meeting the stated rule20 of 2019 of 20

The agent answered one more question, but it failed the rule for one action. Do not hide that failure inside a general quality score. A proper handoff can be the correct result. These figures explain the comparison; they are not measurements of any product.

11

Write down the decision before building more

A pilot should leave a clear decision, not just an impressive demonstration. Keep one page with:

  • The request types the system accepts and the output it should produce.
  • The steps set by rules and the decisions, if any, left to the model.
  • The exact tools and records each step may use.
  • The cases that need a person, including missing facts and service failures.
  • The maximum calls, time and spending allowed for a run.
  • The test cases, results, corrections and owner for the next review.

Choose a fixed workflow if the route is known. Add one AI step if the difficult part is reading or drafting text. Use an agent when choosing the next permitted action adds value you can demonstrate in those tests. Keep the option to simplify it later.

12

Sources

Primary sources checked September 26, 2026.

Need help with this in your business?

Tell Felipe which tools you use, what keeps going wrong and what you want to improve. We can use a free 20-minute call to discuss a useful first project.

Prepare a project brief

Prefer email? felipe@getquintera.com. No booking, purchase or automatic submission.