The New Backbone of Project Management in the Age of AI
Value in AI projects comes not only from model selection, but from problem framing, trusted data, RAG architecture, agile learning loops, governance, and measurable delivery discipline. This column connects LLM and RAG solutions with project management and organizational execution capacity.
Aycan Altınöz · 2026-08-26
The New Backbone of Project Management in the Age of Artificial Intelligence
On the AI agenda we've long been talking about models, tools, and licenses. Every week we encounter a new model that is faster, carries more context, and produces more impressive outputs. Demos are prepared in minutes, efficiency figures are shared in presentations, and organizations enter into the expectation that they will undergo a major transformation quickly. Then the reality on the ground reveals itself: data is scattered, processes are not clearly defined, responsibilities are ambiguous, success criteria are lacking, and the produced output is not integrated into daily workflows.
For me, the most critical point of AI projects begins exactly here. Technology makes the potential visible. Value, however, is created through choosing the right problem, reliable data, controlled integration, measurable objectives, and disciplined delivery.
For more than 13 years I've worked within product, project, agile transformation, and delivery processes across banking, fintech, insurance, automotive, transportation, and technology. In this journey—during which I had the opportunity to work with organizations of different scales and cultures such as Allianz, Yemeksepeti, Ford Otosan, Asis, and Burgan Bank—I observed a recurring pattern: strong technology investments fail to deliver the expected impact when the implementation system isn't built with equal strength. AI makes this truth more visible because an institution's decision-making and implementation capacity reflects on the outcome just as directly as the model's performance.
Today LLM, RAG, agent architectures, and Agile methodologies are treated as separate topics. Yet enterprise value emerges when these areas meet within the same delivery system. LLMs process language and context. RAG brings corporate knowledge into the model. Agents plan specific tasks and interact with tools. The Agile approach manages uncertainty through small learning loops. Project management aligns strategy, risk, resources, stakeholders, and measurement dimensions around a single goal.
The real work that begins after the impressive demo
Preparing an AI demo is now quite easy. You upload a document, write a few instructions, and within seconds get an answer that looks proper. When moving to enterprise use, much more is expected of the same system. It must find the correct information among thousands of documents, respect user permissions, choose the up-to-date source, express uncertainty when appropriate, limit erroneous guidance, maintain audit logs, and operate at an acceptable cost.
This difference defines the gap between a prototype and a product.
A prototype answers the question "Will this idea work?" A product, on the other hand, is tested by the question "Can this solution reliably generate value every day for different users and under changing conditions?" From a project management perspective, the second question requires addressing scope, architecture, data, security, operations, change management, and measurement together.
The first mistake I frequently see in organizations is putting tool selection ahead of problem definition. The team chooses a model or platform first, then looks for a need to apply that tool to. With such a start, the project scope is shaped by technological capabilities. The user problem, business outcome, and process change are pushed into the background.
For a healthy start, the following questions need clear answers:
Which decision do we want to make faster or more accurately?
In which workflow do we want to reduce the costs of waiting, rework, or errors?
How does the user complete this task today?
What level of error is acceptable?
In which cases will human approval be required?
Which dataset and which metrics will we use to measure the output's accuracy?
How will the safe fallback mechanism operate if the solution fails?
An AI backlog prepared without clarity on these questions can quickly turn into a list of disconnected experiments. The team works hard and produces many outputs, yet the business outcome remains unchanged.
Positioning the LLM correctly
Large language models are radically changing how we access information and produce work through natural language. Many tasks—text summarization, classification, content generation, coding, document comparison, routing customer requests, and providing decision support—can be carried out through the same interface. This flexibility presents a major opportunity. The same flexibility also expands the risk surface in projects whose boundaries are not well defined.
An LLM is a probabilistic language system trained on large amounts of data. It generates its response by composing the most likely sequence of words given the context and instructions. Therefore, a fluent answer is not automatic proof of correctness. The model's confident tone can reduce a user's skepticism toward misinformation. In enterprise use, this effect can lead to financial loss, regulatory risk, customer dissatisfaction, or reputational damage.
At this point, the success criteria for an LLM project should be defined in more detail than the phrase "gives good answers." Accuracy, alignment with sources, completeness, consistency, latency, operational cost, security, and user behavior should all be evaluated together. In some use cases, 90% accuracy is a strong start. In high-risk areas such as credit decisions, legal interpretation, health advice, or access authorization, that same rate can produce unacceptable outcomes.
Each use case carries its own risk profile. Therefore, while a single corporate AI policy may provide an adequate framework, product-level decisions should be made on a use-case basis. An assistant that drafts content cannot be governed by the same control mechanisms as an agent that initiates transactions on a customer's account.
RAG: Operationalizing corporate memory
One of the expectations I encounter most in LLM projects is that the model should be familiar with all the information within the organization. Procedures, contracts, product documents, past project decisions, support records, technical manuals, and regulatory documents are kept in different systems. The information exists, access is fragmented. Employees sift through numerous folders, screens, and message histories to reach the correct document.
Retrieval-Augmented Generation, i.e., RAG, offers a powerful architecture to solve this problem. When a user's question arrives, relevant content is retrieved from corporate sources, selected passages are provided to the model as context, and the answer is generated based on that content. This allows a general-purpose model to operate with the organization's up-to-date and authorized information.
RAG's corporate value is evident in three areas. First, it shortens the time to access information. Second, it increases auditability by making the source of the answer visible. Third, it enables the use of information that is not present in the model's training data or that changes over time.
Treating RAG as merely a document-upload feature creates serious design gaps. A high-quality RAG system requires many decisions, from the ingestion of the document into the system to the presentation of the answer to the user:
Which sources will be considered reliable?
How often will content be updated?
Into what size chunks, and with what degree of semantic coherence, will documents be split?
How will semantic similarity, keyword, and hybrid methods be used together in search?
Will the retrieved results be re-ranked?
How will the user's role and access permissions be enforced during the search phase?
How will the model respond when the source material is insufficient?
In what format will citations and references be presented to the user?
With which evaluation set will the quality of the answers be tested regularly?
If even one of these questions is left weak, the system may appear to work technically, but user trust will quickly erode. An outdated procedure surfacing at the top, incorrect matching of similar concepts, or using a document in an answer to which the user lacks access can overshadow the entire value of the project.
Even after RAG, the possibility of error persists. Finding the relevant document does not mean the model will interpret that document flawlessly. Therefore, retrieval quality and generation quality should be measured separately. "Was the correct source found?", "Does the source answer the question?", "Is the generated answer consistent with the source?" and "Did the answer help the user complete their task?" should be tied to different metrics.
I see RAG as much the technical layer of corporate memory as a mirror of the discipline of knowledge management. If document ownership is unclear, content freshness isn’t tracked, and there are conflicting records about the same topic, RAG makes those problems visible. The project team can turn that visibility into an opportunity for cleanup and governance.
Why is Agile more valuable in artificial intelligence projects?
Uncertainty is high in artificial intelligence projects. The same model can produce different results on different datasets. An answer a user finds satisfactory may not align with the one the technical team scores highly. You must constantly balance cost, latency, and accuracy. A change of model or provider can affect a workflow that previously worked. This environment makes large delivery packages developed behind closed doors risky.
The Agile approach offers a disciplined learning-and-decision system here. You pick a small scope, test it with real users, measure the results, and make the next decision based on evidence. The value of a sprint is measured less by the number of completed tasks and more by the uncertainty reduced.
Agile work in AI projects requires extending the classic software backlog. Alongside User Story and technical tasks, data preparation, building evaluation sets, prompt and context design, model comparison, security testing, observability, cost analysis, and user feedback should also be part of the backlog.
The definition of "Done" should be revisited. A feature functioning on the screen is insufficient for completion. The produced answer's alignment with sources, critical error rate, response time, cost limit, access control, logging, and rollback scenario should be included in the acceptance criteria.
Basing sprint goals on outputs in these projects can mislead teams. "Completing the chatbot screen" is a deliverable. "Validating a flow that will reduce the support team's time to reach the correct procedure by 30%" is a sprint goal tied to a business outcome. The second statement brings design and technology decisions together under the same objective.
Agile ceremonies should be managed with the same perspective. When the Daily becomes a status report, the Sprint Review a product demo, and the Retrospective a general satisfaction chat, learning capacity declines. In AI projects, the focus of these meetings should be sharper:
Which assumption did we test?
Which outcome did we observe in which user group?
What were the highest-risk error examples?
Which issues were we able to attribute to the model, the data, or the process?
Which uncertainty will we reduce in the next sprint?
These questions turn the Agile approach from a ritual-level practice into a management system.
The project manager's new role
In the age of artificial intelligence, the project manager's role is becoming more technical. The required technical depth includes technology literacy at a level that enables understanding how architectural choices affect business outcomes, risks, and the delivery plan. Responsibility for model development, however, remains with the relevant specialist teams.
A project manager should now be able to connect the following headings:
Business objective and use case
Use case and required data
Data, the model, and the RAG architecture
Architecture and security, cost, and performance
Model output and user decision
User decision and measurable business outcome
Pilot findings and the decision to scale
When these connections are not established, teams produce successful outputs within their own areas of expertise, but the product as a whole cannot achieve the same success. The data team may reach a good retrieval rate, the software team a working integration, and the design team a smooth interface. If the user still prefers the old method, the project has not passed the threshold of value creation.
One of the project manager's key responsibilities is to create a shared language of success. The technical team tracks "precision", "recall", "groundedness" and latency metrics. The business unit looks at processing time, cost of errors, sales conversion, or customer satisfaction. Risk teams assess data privacy, explainability, and regulatory compliance. Senior management wants to see the return on investment. Healthy governance combines these metrics into a single decision framework.
I find the supervisory engineering approach valuable here. As AI systems take on more tasks, the human role is evolving from executing every step manually to defining boundaries, monitoring performance, managing exceptions, and intervening when necessary. This approach prevents automation from creating a responsibility gap. Which decisions can be made by the system, at what threshold human approval is required, and who will act in case of an error are defined at the start of the design.
In this new arrangement the project manager acts like a delivery architect. They translate strategy into the backlog, risk into control mechanisms, user needs into acceptance criteria, and pilot data into the decision to scale.
How should a RAG use case be handled?
Suppose a RAG-based assistant is being developed in a customer support organization to speed access to procedure and product information. The traditional approach might proceed by preparing technical requirements and opening the solution to users several months later. The uncertainties in an AI project make such a long feedback interval expensive.
I would approach this scenario with a phased validation model.
In the first stage, the current state is measured. How many minutes does a support agent take to reach the correct information? Which topics cause the greatest time loss? How many tickets are reopened due to incorrect or outdated information? Which sources are actually used? The success target is defined based on these baseline values.
In the second stage a narrow subject area is chosen. For example, a dataset limited to the three highest-volume products and approved procedure documents is prepared. An evaluation set composed of real questions is created. Expected answers, acceptable alternatives, and examples of critical errors are marked by experts.
In the third stage, retrieval performance is tested. Can the system find the correct document and the correct section? Are access controls applied before search? Does the current version take precedence over the old version? At this stage, the quality of information access is verified before the model's answer.
In the fourth stage, answer generation and user experience are addressed. Is the answer concise and actionable? Is the source link visible? If the system is unsure, does it state that and direct the user to the right channel? Is it easy for the user to report an incorrect answer and initiate the correction process?
In the fifth stage, a controlled pilot is conducted. A small group of users uses the system in real workflows. Average processing time, rate of reopened records, user acceptance, number of critical errors, response cost, and latency are monitored together. At the end of the pilot, a decision is made to scale, improve, or narrow the scope.
The strength of this flow is that it breaks a large investment into small, measurable decisions. Each stage produces evidence for the next spending and scope decision.
Recurring patterns that lead to failure
Some problems in AI projects recur regardless of industry or organizational scale.
The first pattern is treating demo success as production success. A system that performs well on controlled examples may behave differently with real user language, incomplete queries, conflicting documents, and high traffic. The production environment is where exceptions occur.
The second pattern is treating data preparation as a minor technical task in the project plan. In RAG systems, document quality, metadata, version management, permission structure, and content ownership are central to the solution. When data preparation is delayed, teams often end up repeatedly tweaking prompts. Resource problems are then attempted to be solved at the interface or model layer.
The third pattern is limiting success metrics to user count or the number of answers produced. High usage does not always mean high value. If a user asks the same question several times, verifies the answer in another system, or heavily edits the AI output, additional operational costs accumulate behind the apparent usage.
The fourth pattern is leaving human approval undefined. "Human in the loop" is a phrase frequently used in presentations. In practice, when it isn't defined which role will approve, at what threshold, within what time frame, and based on what information, this control becomes a bottleneck. Human approval should also be designed as a measurable process.
The fifth pattern is reducing change management to a training session. It's not enough for users to know how to access the new tool. Role definitions, performance measurement, decision authority, escalation channels, and feedback mechanisms must be updated for the new way of working. People want to understand how the system affects their jobs, responsibilities, and risks.
The sixth pattern is managing the AI component separately from the entire process. The model generates an answer, and then the user copies that answer and moves it to another system. In that case, a few minutes' gain can be lost to new steps and control overhead. Value should be measured across the end-to-end flow.
Governance is the infrastructure of speed
AI governance is perceived in most organizations as a control performed at the final stage by security and legal teams. This approach slows the project because critical constraints become visible after investments have been made. When governance is brought into the beginning of the design, the team knows where it can experiment freely and at which thresholds additional controls are required.
An effective governance model separates use cases by risk level. An assistant that prepares internal communications can be positioned in a low-risk area. Systems that handle customer data, provide financial recommendations, or take automated actions require stricter evaluation. Data classification, access control, logging, model selection, evaluation frequency, and human approval are determined according to that risk level.
Authorization is particularly critical in RAG projects. A document that a user cannot access must not be included in a retrieval result. Trying to hide information at the response layer is a late and fragile control. Authorization filters should be applied during the search phase, and logs should be able to record which sources were used.
Prompt injection, data leakage, inappropriate content generation, and source manipulation should also be included in the test plan. Security testing should not remain a one-off check; as the model, data, or prompts change, critical scenarios should be rerun. The equivalent of regression testing in software projects expands in AI systems into evaluation sets and security scenarios.
Which metrics describe success?
A single metric does not explain the whole picture in AI projects. A balanced measurement system can be built across four layers.
The first layer is technical quality. Retrieval accuracy, consistency with sources, critical error rate, response time, usability, and cost per operation are monitored here.
The second layer is user behavior. The acceptance rate of the system's recommendations, user corrections, repeated questions, types of feedback, and the rate of reverting to old ways of working are tracked.
The third layer is process performance. Cycle time, wait time, rework, escalations, errors, and operational costs are measured.
The fourth layer is business outcome. Targets such as revenue impact, customer satisfaction, risk reduction, service level, employee capacity, or time-to-market are selected according to the use case.
A causal relationship must be established between these layers. If model accuracy increases but processing time does not change, there may be another bottleneck in the process. If user acceptance is high but the cost of errors is increasing, the trust mechanism may be weak. If business outcomes have improved but the model cost per transaction is rising rapidly, the scaling decision should be reassessed.
Senior management needs visibility that produces decisions beyond hundreds of technical metrics. Which use case is generating value? Which one is approaching the risk threshold? Where is data investment needed? Which pilot should be scaled? Which work should be stopped? Project management should design reporting to answer these questions.
The human and organizational dimension
Task distribution is changing with artificial intelligence. Some portions of repetitive information retrieval, initial drafting, classification, and control activities can be carried out by the system. People can devote more time to areas that require contextualization, exception handling, ethical assessment, stakeholder communication, and decision responsibility.
This transition does not happen spontaneously. Employees need new competencies, managers need new performance indicators, and the organization needs clear responsibility definitions. AI literacy, beyond good prompt writing, includes skills such as questioning outputs, verifying sources, recognizing types of errors, protecting sensitive data, and making correct escalation decisions.
One of the most important questions for leaders is how the capacity gain will be used. When a team saves 100 hours a week, will that capacity be directed toward customer experience, quality, innovation, or higher business volume? If this decision is not made, productivity gains may dissipate before converting into financial or strategic outcomes.
Agile leadership regains importance here. The leader clarifies the objective, boundaries, and decision principles and gives the team room to act within the solution space. A space is created where teams can safely experiment. Learnings from failed attempts are recorded. Successful use cases are scaled with shared components.
An actionable roadmap for organizations
When building an AI portfolio, choosing among hundreds of ideas can become difficult. I find it helpful to evaluate use cases across five dimensions: business value, data readiness, technical feasibility, risk level, and capacity for change.
Scenarios with high business value and strong data readiness where risk is manageable are good candidates for early pilots. In areas with high value but poor data quality, priority should be given to knowledge and data infrastructure. Low-value but flashy ideas clutter the portfolio and create loss of focus.
An actionable roadmap could proceed in the following order:
Identify strategic objectives and process bottlenecks.
Score use cases in terms of value, data, risk, and feasibility.
Define narrow-scope pilots and clear success criteria.
Design the RAG, model, integration, and security components to be reusable.
Establish short feedback loops with real users.
Monitor technical, user, process, and business metrics together.
Decide at the end of the pilot whether to scale, improve, or stop.
Turn the learnings into corporate standards and new backlog items.
This approach keeps innovation under control and can progress without expanding bureaucracy. Standard evaluation sets, approved data connections, security templates, shared logging infrastructure, and clear decision thresholds enable teams to move faster. The organization's shared AI competency reduces the burden of solving the same basic problems from scratch in every project.
Final word: Competitive advantage will arise from implementation capacity
AI tools are becoming increasingly accessible. Many organizations can access the same models, similar platforms, and powerful infrastructure. The area that will create a lasting difference is implementation capacity: selecting the right problem, making information reliable, learning in small cycles, incorporating risk into design, and the ability to deploy a working solution end-to-end into the process.
LLM forms the language layer of this transformation, RAG the enterprise knowledge layer, Agile the learning rhythm, and project management the holistic delivery discipline. When these four areas work together, AI moves beyond a presentation headline and turns into operational capability.
From my perspective, the project manager of the future is a delivery leader who establishes a decision system between business, technology, data, people, and governance. They maintain calendar and budget management; monitor what the model can do, understand what the organization needs, and bring these two areas together into measurable outcomes.
In the coming period, institutions' success will be determined as much by their answer to the following question as by the name of the model they use: Can we transform AI capacity into a reliable, repeatable, and scalable operational system?
Organizations that give a strong answer to this question will be fueled by rapid learning, clear decision-making, and disciplined execution. The number of tools will remain a secondary indicator behind these capabilities.
