How long does it take to build an AI agent?

By Chad Upton, Section CTO
Michael is on PTO this week, so our CTO Chad Upton is guest posting on why some agents take days to build and others take weeks. If you want help building your own agents, book a time with our team.
–
Teams often talk about “building an agent” as though it describes one kind of project. From the engineering side, I’d love for it to be that simple.
At Section, we can get a useful question-answering agent running in a day or two. An agent that organizes files for our advisory team has taken weeks of interviews, process decisions, integration work, and testing.
Organizing files doesn’t sound especially ambitious. But the effort has less to do with how impressive the task sounds than how much work needs to be defined before an agent can do it.
Does the information already exist? Are the steps clear? Can one person tell you whether the result is right, or do several people need to agree? Those questions aren’t as exciting as choosing tools and starting to build, but they have a lot to do with how long the project will take.
Here are the three levels of effort we’re seeing at Section, what drives the timeline for each, and what someone needs to own afterward.
1. Agents that answer questions
Effort level: 1–2 days, assuming the underlying information and access are in place
At this level, you’re making existing knowledge easier to access. The agent answers questions that would otherwise interrupt someone’s work, or sit unanswered until the person comes back from PTO.
At Section, we have Slack channels where people ask questions about our product, services, compliance, and Section’s point of view. Claude draws on documents, transcripts, and other relevant context to help answer.

Most of the setup effort goes into deciding what the agent should know, which sources it should trust, and what a good answer looks like.
That sounds simple until you give it both Product Strategy 2026 and Product Strategy 2026 H1_revised. Your colleagues may know which document is current. The agent needs a way to make that distinction, too.
You also need to test questions the source material can’t answer. I’d rather the agent tell someone it doesn’t know than give them a confident answer we have to correct later.
This is why the 1-2 day estimate comes with a condition: the information needs to be ready. If someone first has to reconcile conflicting documents, find missing material, or sort out access, add that work to the estimate.
What ownership looks like:
- Owner: A subject-matter expert who can judge answer quality (not necessarily the person who configured the agent)
- Ongoing work: Keep sources current, review incorrect answers, and clarify when a question needs a human.
- Handoff: Document the sources and give others access to maintain them.
2. Agents that execute individual tasks
Effort level: A week or more to build, test, and refine
These agents complete a defined task, such as turning a call transcript into a follow-up email or producing a presentation using your company’s template. The extra effort comes from specifying how the work should happen, not just what the agent should know.
For a presentation-building agent, “use our template” is only the beginning. You need to define how to structure the content, what the finished output should look like, and how to generate the file. There may be a reusable skill to configure, tools to connect, or code to write.
A lot of this work is making your own expectations explicit. You or I can look at a presentation and immediately know it’s wrong. But explaining what the agent should do differently (and turning that feedback into instructions that work next time) takes more effort. Frankly a lot of people skip this part because it takes too long, and just fix the problems themselves - which doesn’t scale.
Then you need to test against real work. Budget for several rounds of use, feedback, and adjustments, rather than treating the first good output as the finish line.
What keeps the project relatively contained is a narrow scope and a direct feedback loop. One person usually understands the task and can check the result. You’re also not trying to get several teams to agree on how the work should happen. More integrations or exceptions can still extend the timeline, but you have fewer decisions to coordinate.
What ownership looks like:
- Owner: Often the person using the workflow. As others adopt it, distinguish between maintaining the shared version and reviewing each output.
- Ongoing work: Refine instructions, fix recurring failures, and maintain the tools and connections.
- Handoff: Give reusable work a shared home and someone who can maintain it. At Section, we’ve built an internal plugin marketplace and an authoring plugin that helps employees publish through GitHub. That makes the work easier to find and reuse, but a team dependency still needs a backup owner.
3. Agents that run part of a shared process
Effort level: Weeks of process design, integration, testing, and rollout
At this level, you’re no longer just helping someone complete their own task. You’re changing how work happens across a team. That adds decisions, dependencies, and people to the build.
Riley, an agent we’re developing for our advisory team, helps organize client files, create consistent folder structures, and put signed agreements where they belong.

The technical work is only part of that assignment. Before Riley can organize anything, we need to agree on what “organized” means. What should the folder structure look like? What happens when a file doesn’t fit? Which changes can Riley make independently, and which need approval?
Answering those questions has required interviews and process decisions alongside the integration work, which of course means coordinating people’s schedules and going back and forth.
Testing also takes longer because the builder isn’t the only person who can judge success. We’ve needed one or two people who encounter this work every day to check Riley’s results. Did it create the right structure? Did the agreement end up in the right place? Did a proposed change make sense?
This is where I’d be careful about a timeline that only accounts for engineering. Your engineer might be available, but is the person who knows where the agreements belong? Don’t plan around users testing whenever they have a spare minute. They need time set aside, because their feedback determines what gets built and fixed next.
What ownership looks like:
- Owner: A product owner who understands the work, defines success, and prioritizes fixes, with dedicated users to test alongside them. This does not have to be the builder and probably shouldn’t be unless they’re a PM-type.
- Day-to-day management: Give the team closest to the work control over routine decisions. For Riley, advisory can add employees or turn monitoring on and off without routing every change through engineering.
- Maintenance and handoff: Separate technical maintenance from daily operation. Riley’s workflows and integrations live in n8n, its code in GitHub, and its operating rules and configuration in Notion. Engineering can maintain the underlying system while advisory manages how it operates.
A simple task isn’t always a simple build
A task that takes someone five minutes can still take weeks to automate reliably, for two reasons:
One, so much of human knowledge is implicit. You already know when to gutcheck with a colleague or ask for a V2 - and you probably don’t think of those as decisions anymore. They’re just how you do the job.
And our human brains are great at nuance. We naturally know the difference between two similar documents, why one client’s annoyance is another’s approval, etc.
When you build an agent, you have to make that knowledge usable: turn it into instructions, provide reliable context, or establish when the agent should ask a person.
So before committing to a timeline, I’d ask the people doing the work to walk through a few real examples, especially the ones that didn’t go according to plan. Where did they make a judgment call? What did they know that wasn’t written down? Who did they need to ask?
Those answers tell you whether you’re automating a well-defined task or using the build to finally define it. Both can be worthwhile, but the second is a much bigger project - and more human too.


.webp)