Post

Managing Non-Deterministic Systems: A PgMP's Guide to Delivering Enterprise AI Programs

Why traditional software delivery playbooks break down when managing Generative AI programs, and how technical program managers can govern probabilistic systems without stalling innovation.

Managing Non-Deterministic Systems: A PgMP's Guide to Delivering Enterprise AI Programs

For most of my career across software engineering and program management, successful delivery followed a reassuringly predictable formula:

  1. Define clear requirements and acceptance criteria.
  2. Design deterministic system architecture.
  3. Build, write unit and integration tests with binary outcomes (assert expected == actual).
  4. Commit to milestones, track critical paths, and release.

If a bug appeared in production, you could reproduce it with the same inputs, trace the stack trace, patch the logic, and verify that the test suite passed with 100% confidence.

Then came enterprise Generative AI and large language models (LLMs).

Suddenly, software is no longer deterministic. Feed the exact same prompt with the exact same payload into a model at 9:00 AM and 2:00 PM, and you might get two subtly different responses. A minor tweak to a system prompt meant to fix an edge case in customer support can inadvertently degrade classification accuracy across billing inquiries.

When business stakeholders ask: “Can you guarantee this AI agent will never give wrong advice to a customer?”—traditional project management instincts want to say “Once QA is done.”

As a Program Manager with both a computer science background and PgMP governance experience, here is the honest reality: Traditional delivery playbooks break when applied blindly to probabilistic systems.

Here is how we need to reshape our program management strategy to successfully deliver enterprise AI initiatives.


1. Shift from Binary QA to Continuous Evals

In standard software programs, quality is binary: tests pass or fail, builds are green or red.

In AI programs, quality is a statistical distribution. There is no such thing as “zero errors” in an open-ended LLM application. If you wait for 100% accuracy before launching, your program will never leave the pilot phase.

flowchart TD
    subgraph Traditional["Traditional Deterministic Delivery"]
    A[Requirements & Specs] --> B[Code Implementation]
    B --> C[Unit / Integration Tests]
    C -->|Pass 100%| D[Predictable Production Release]
    end

    subgraph AI["Probabilistic AI Delivery"]
    E[Prompt & Model Strategy] --> F[Golden Eval Benchmark]
    F --> G{Confidence Interval?}
    G -->|>= 94% Accuracy| H[Controlled Deployment]
    G -->|< 94% / Edge Failure| I[Tune Prompt / Middleware / Few-Shot]
    I --> F
    end

What to do instead:

  • Establish an Evaluation Harness (Evals) early: Treat your evaluation dataset as first-class program infrastructure. Work with your engineering leads to curate a golden benchmark dataset (200–1,000 real-world edge cases and representative user queries).
  • Define acceptable confidence intervals: Align with product and legal leadership on metric thresholds (e.g., Accuracy ≥ 94%, Hallucination Rate ≤ 1.5%, Latency p95 ≤ 2.2s).
  • Automate regression runs on prompt & model changes: Every time a prompt changes, a temperature parameter is tuned, or an underlying model version updates, run the automated eval suite to measure the net delta across the entire benchmark.

[!NOTE] Program velocity in AI is directly bounded by your eval feedback loop. If running an evaluation takes three days of manual review, your iteration speed is three days. If it’s automated via LLM-as-a-judge and ground-truth comparisons, your team can test hypotheses daily.


2. Re-architect Risk Management for “Graceful Failure”

In PgMP terms, risk management isn’t just about identifying what might go wrong—it is about designing the operational tolerance of the enterprise.

When building AI programs, we must assume the model will occasionally fail. The goal of technical program governance is ensuring that when it fails, it fails safely and gracefully.

flowchart LR
    User([User / Client Request]) --> Guardrail[Deterministic Input Guardrail<br/>Regex / Schema Validation]
    Guardrail --> Router{Model Router}
    
    Router -->|Triage & Summary| LightModel[Lightweight Model]
    Router -->|Complex Synthesis| FrontierModel[Frontier Reasoning Model]
    
    LightModel --> EvalCheck{Confidence Check}
    FrontierModel --> EvalCheck
    
    EvalCheck -->|High Confidence| Output([System Action / Response])
    EvalCheck -->|Low Confidence / Anomaly| HITL[Human-in-the-Loop<br/>Specialist Review]

Architectural Guardrails Every TPM Should Track:

  1. Deterministic Guardrails & Fallbacks: Never let an unconstrained model talk directly to sensitive backend systems. Put deterministic validation layers, regex filters, and schema parsers (like Pydantic / Zod) between the LLM and your APIs.
  2. Human-in-the-Loop (HITL) Escalation Paths: Design workflows where confidence scores below a specific threshold automatically route the task to a human specialist rather than hallucinating an answer.
  3. Strict Blast Radius Containment: Limit agent tool permissions to read-only where possible. Any state-changing action (deleting data, initiating a wire transfer, sending an external email) must require explicit confirmation.

3. Dynamic Cost & Token Governance (FinOps for AI)

In traditional web applications, compute and infrastructure costs scale relatively linearly with user sessions and database queries.

In LLM-based programs, infrastructure costs can spike unpredictably based on:

  • Context window bloat (sending thousands of uncompressed tokens per request).
  • Sub-agent loops (an autonomous agent getting stuck in a reasoning retry loop).
  • Uncached repetitive prompts across distributed microservices.
1
2
Traditional Cost: Users -> Fixed CPU/Memory -> Predictable Monthly Run-rate
AI Program Cost: Users -> Dynamic Token Depth * Model Tier * Agent Tool Iterations -> Volatile Variance

Strategic Governance Actions:

  • Implement Prompt Caching & Middleware: Ensure architecture teams leverage prompt caching (e.g., Anthropic Prompt Caching) to cut input token costs by up to 80–90% on static context payloads.
  • Tiered Model Routing: Don’t route simple data extraction tasks to the largest flagship frontier models. Use lightweight models for triage, classification, and summarization, and reserve expensive reasoning models for complex synthesis.
  • Set Hard Budget Quotas & Anomaly Alerts: Build monitoring dashboards tracking daily token consumption per tenant or feature track.

4. Stakeholder Alignment: Managing the “Magic vs Reality” Gap

Perhaps the biggest challenge for an AI Program Manager isn’t technical—it’s managing executive expectations.

Stakeholders often view AI either through extreme optimism (“This will automate our entire operations by Q3”) or paralyzing fear (“One mistake and our brand reputation is destroyed”).

How to Bridge the Gap:

1
2
3
Step 1: Shift the goalpost from "Automation" to "Augmentation"
Step 2: Tie milestones to measurable Benefits Realization, not model hype
Step 3: Define clear Stage-Gate criteria from Prototype -> Pilot -> Production
  • Measure Net Efficiency, Not Just Direct Accuracy: If an internal AI copilot drafts 80% of an analyst’s report and reduces turnaround time from 4 hours to 45 minutes, that represents massive program value—even if the analyst still spends 15 minutes reviewing and editing the draft.
  • Adopt Phased Stage-Gates:
    • Stage 1 (Internal Dogfooding): Deployed only to a friendly internal cohort with continuous qualitative feedback.
    • Stage 2 (Shadow Mode / Co-pilot): The AI generates recommendations in parallel with human operators to benchmark live performance against human ground truth.
    • Stage 3 (Supervised Autonomous): The system handles routine low-risk tiers directly while escalating edge cases.

Key Takeaway for Technical Leaders

Delivering enterprise AI is not just data science, and it is not just standard software development. It is the art of building deterministic scaffolding around probabilistic engines.

As technical program leaders, our role is not to eliminate uncertainty—which is impossible with generative models—but to measure it, bound it, and structure the organization around it so our teams can innovate with confidence.


What challenges has your team encountered when moving generative AI projects from proof-of-concept into production governance? I’d love to hear your thoughts.

This post is licensed under CC BY 4.0 by the author.