The development process in 5 phases
Building a custom AI agent follows a structured five-phase process: Discovery, Design, Build, Test and Deploy. It is not a matter of writing code and shipping: every agent is a software product that has to integrate with existing business systems, respect regulatory constraints and produce results that are reliable 100% of the time.
Yellow Tech has refined this process across more than 300 AI agents in production for 500+ Italian organizations. The approach is iterative: you start from a working MVP and refine it step by step on real data. That lowers project risk and speeds up time-to-value.
The overall duration ranges from 4-8 weeks for a single use case to 6-12 months for enterprise multi-agent programs. Here is each phase in detail.
Phase 1 - Discovery: understanding the problem
Discovery is the most important phase. The business process you want to automate is analyzed at a level of granularity that goes well beyond a simple description: all the inputs (documents, emails, events, data), all the expected outputs, the possible exceptions, the business rules and the systems involved get mapped.
Our team runs workshops with the process owners at the client company. Two deliverables come out of it: a detailed process map of the current flow (as-is) and of the target flow (to-be), and a requirements document with measurable success KPIs (processing time, error rate, volume handled).
Discovery takes 1-2 weeks. It is the moment when you decide whether an AI agent is the right solution for that specific problem. Not every process benefits from AI automation: some are too simple (a rule-based automation is enough), others too complex or too low in volume to justify the investment. Good AI consulting also knows how to say no.
Phases 2 and 3 - Design and Build
In the Design phase the technical architecture of the agent takes shape. The key decisions include: which LLM to use (the choice depends on cost, latency, reasoning ability and data residency requirements), which tools and APIs to integrate, how to structure the workflow, how to handle errors and exceptions, and which guardrails to put in place for security.
The phase produces a design document that the client approves before work continues. That document covers the system architecture, the flow diagrams, the integration specifications, the security plan and the operating cost model (API costs, infrastructure, monitoring).
The Build phase is development itself. The team works with weekly releases to a staging environment. Every release is testable by the client, who can give feedback right away. The typical tech stack includes Python or TypeScript for the agent logic, dedicated AI frameworks for reasoning and tool use, and cloud infrastructure (AWS, GCP or Azure depending on the client ecosystem).
The Build takes 2-4 weeks for a single use case. This phase also produces the technical documentation and the operational runbook for the client team.
Technologies and tech stack
We take a model-agnostic approach: the choice of technology depends on the use case, not on commercial partnerships. The 30+ specialists on the team have cross-cutting expertise across all the major providers and frameworks.
- Foundation Model - OpenAI (GPT), Anthropic (Claude), Google (Gemini), Meta (Llama), Mistral. The choice depends on performance, cost per token, latency and data residency requirements.
- Orchestration - Custom frameworks in Python and TypeScript (LangChain, CrewAI, Vercel AI SDK) for building agents with multi-step reasoning and advanced tool use. Platforms like n8n are also used for automation workflows.
- Voice AI - ElevenLabs for voice agents with latency under 500ms and speech quality indistinguishable from a human. Used for phone customer service and internal voice assistants.
- Sales Intelligence - Clay for prospect data enrichment, scoring and outreach automation. Integrated with CRMs (HubSpot, Salesforce, Pipedrive) for end-to-end pipeline management.
- Infrastructure - AWS, GCP or Azure depending on the client ecosystem. Containerization with Docker, orchestration with Kubernetes for enterprise deployments. Monitoring with Datadog or Grafana.
The team you need
Building an AI agent calls for different skills from those of a traditional software project. The typical team for a single use case includes: an AI Solution Architect who designs the architecture and picks the technologies, one or two AI Engineers who build the agent, and a Project Manager who coordinates the work and manages the client relationship.
Enterprise projects add: an AI Governance Specialist for regulatory compliance, a Data Engineer for complex data integrations, and an AI Trainer who trains the client team on using and maintaining the agent. AI training for the internal team is essential for long-term success.
The client team has to make available at least one process owner (the person who knows the process) and one IT contact (for the technical integrations). How available these people are has a direct effect on how fast the project moves.
Testing and iteration
Testing an AI agent is more complex than testing traditional software, because the output is not deterministic: the same input can produce slightly different answers. That is why the approach works on three levels.
The first level is automated testing: unit tests on individual functions, integration tests on the API connections, and regression tests to check that changes do not break existing behavior. The second level is testing with real data: hundreds of actual cases from the client company are processed and the results verified. The third is UAT (User Acceptance Testing): end users test the agent in a controlled environment and give feedback.
After go-live the iteration cycle keeps going. Agent performance (accuracy, latency, escalation rate) is monitored, and adjustments come through prompt tuning, knowledge base updates or workflow optimization. The first month after deployment is included in the project; after that a maintenance contract kicks in.
Post-go-live maintenance
An AI agent in production needs ongoing maintenance. LLMs get updated (new versions of GPT, Claude, Gemini), the APIs of business systems change, business processes evolve. Without maintenance, agent performance degrades over time.
We offer maintenance contracts at three levels. The basic tier covers performance monitoring, alerting on anomalies and security updates. The standard tier adds monthly prompt tuning, knowledge base updates and technical support with a guaranteed SLA. The premium tier includes a dedicated team, continuous evolution of the agent and priority access to new technologies.
Maintenance costs run from 10% to 20% of the initial development cost on an annual basis. It is an investment that protects the value of the agent over time. For more on AI consulting costs, see the dedicated guide.
Frequently asked questions
For a single use case, 4 to 8 weeks from kickoff to go-live. For enterprise multi-agent programs, 6 to 12 months. Yellow Tech follows an iterative approach with weekly releases, so the client sees concrete progress from the second week of the project.
No, Yellow Tech runs the entire development end-to-end with a team of 30+ dedicated specialists. What is needed is the availability of a process owner (whoever knows the process) and an IT contact (for the integrations). For the long term, we recommend pairing development with an AI training program for the internal team.
The approach is model-agnostic: the best technologies for each use case get selected. The main LLMs used are OpenAI GPT, Anthropic Claude, Google Gemini, Meta Llama and Mistral. Development happens in Python and TypeScript with dedicated AI frameworks. Yellow Tech has cross-cutting expertise across all the major providers.
With a three-level testing approach: automated tests (unit, integration, regression), testing with the client real data, and UAT with end users. After go-live, Yellow Tech monitors accuracy, latency and escalation rate, stepping in with continuous tuning. 98% of Yellow Tech clients rate the service positively (CSAT).
Annual maintenance costs run from 10% to 20% of the initial development cost, depending on the service level chosen (basic, standard, premium). The first month after go-live is always included in the project. Yellow Tech offers guaranteed SLAs and 24/7 monitoring for mission-critical agents.
Related guides
- AI Agents for Business: What They Are, How They Work, What They Cost
- AI for Customer Service: Automating Customer Support
- AI for Sales: Sales Automation and Lead Generation
- Document Automation with AI: Invoices, Contracts and Compliance
- AI Consulting in Italy: The Complete Guide for Businesses
- Artificial Intelligence Courses for Companies: The 2026 Guide
- How Much AI Consulting Costs: What Drives the Price
Want to see how AI can help your company?
