AI/ML Engineering
Introduction
While AI Platform Engineering establishes the hardened, reusable infrastructure of the "Hub," the actual translation of business requirements into production software happens at the application tier. This is the domain of AI/ML Engineering.
For technology executives, an immediate operational challenge is defining the boundaries of this role. In many early-stage initiatives, enterprises make the mistake of assigning classical data scientists to build enterprise software or expecting traditional full-stack developers to master the nuances of non-deterministic runtimes.
To scale an AI enterprise, leaders must treat AI/ML Engineering as a specialized, decoupled software engineering discipline. This section outlines the functional demarcation of the role, the foundational shifts in engineering mindset required for probabilistic software development, and the core technical disciplines necessary to build resilient, production-grade AI applications.
1. Demarcation of Roles: The Enterprise Reality
To eliminate operational friction and avoid talent misalignment, the enterprise target operating model must clearly differentiate the roles of the Data Scientist, the Traditional Software Engineer, and the AI/ML Engineer.

The Data Scientist
The Data Scientist focuses on the mathematics underneath the system. They spend their time cleaning data, analyzing distributions, selecting architectures, fine-tuning open-weight models, and tracking loss curves. They work within research sandboxes and training clusters. Their core deliverable is an optimized model artifact, a custom embedding model, or a highly specialized classification weight.
The Traditional Software Engineer
The Traditional Software Engineer operates entirely within a deterministic framework. They build user interfaces, write business microservices, maintain API schemas, manage transactional state, and orchestrate standard databases. They expect an input to consistently map to an immutable, predictable output.
The AI/ML Engineer
The AI/ML Engineer bridges the gap between these two worlds. They treat the foundation model as a raw compute engine. They do not typically train models from scratch. Instead, they build the execution context, design stateful agent orchestrations, implement semantic parsing pipelines, and manage the runtime integration between deterministic enterprise systems and probabilistic model endpoints. Their primary goal is building reliable software on top of fundamentally unpredictable components.
2. The Probabilistic Software Engineering Mindset
The introduction of large language models breaks classical software patterns. Traditional testing and deployment strategies rely on the assumption that code behaves consistently over time. When an application's core logic shifts to an LLM, it inherits a unique set of challenges:
- Model Variability: An upstream API update or an unexpected token sequence can cause a model's output formatting to change instantly, bypassing standard exception handlers.
- Context Sensitivity: Small adjustments to a system prompt, or changes in the sorting order of injected RAG documents, can lead to wild shifts in output quality, relevance, or factual accuracy, including hallucinations.
AI/ML Engineering replaces standard imperative programming with defensive probabilistic engineering. Every function call to an intelligence endpoint must be treated as a network operation that could fail, return invalid structures, time out, or produce misaligned logic. This requires engineers to build validation, fallback, and structural checking directly into the core application layer.
3. Core Technical Disciplines
To transition applications successfully through TOGAF Phase G (Implementation Governance), AI/ML Engineers must master four core software practices.
3.1 Advanced Context Window Management
Context windows are constrained, expensive computing real estate. Naively stuffing thousands of raw database tokens into an LLM request causes cost inflation and performance degradation.

- Mitigating "Lost in the Middle" Degradation: Models pay the highest attention to information at the absolute beginning and the absolute end of an injected context window. If critical reference information sits in the middle of a massive token block, comprehension drops significantly. AI/ML Engineers use advanced ranking systems, such as Cross-Encoders, to reposition the most relevant information blocks to the outer edges of the prompt payload before dispatching it to the model gateway.
- Token Optimization Strategies: Engineers implement strict structural text pruners, removing boilerplate HTML tags, Markdown elements, and repetitive phrases from source text chunks. Keeping context windows clean directly controls latency and optimizes the enterprise Cost per Successful Task (CPST).
3.2 Structured Output Enforcement
An enterprise application cannot parse a conversational response like "Sure, here is the account information you requested...". The system requires strict, type-safe data schemas, such as valid JSON objects containing expected keys and correctly cast data types, to trigger downstream automated actions.
AI/ML Engineers use verification tools, such as Pydantic in Python or type guards in TypeScript, alongside model-level grammar constraints. Rather than simply asking the model to return JSON in a text prompt, they pass hard schema limits directly to the platform runtime gateway, forcing the model's token selection engine to emit outputs that align with the required syntax structure.
3.3 Agentic Design Patterns & State Machine Engineering
Giving autonomous agents unconstrained freedom to choose their own execution loop is a major risk factor for production enterprise environments. An open-ended agent loop can trigger infinite recursive calls, quickly burning through token budgets without resolving the user's task.

To prevent runaway behavior, AI/ML Engineers structure agent workflows using Directed Acyclic Graphs (DAGs) or strict state charts. The agent's autonomous choices are limited to branching paths explicitly defined by the application developer. If an agent needs to access an enterprise tool, such as searching a database or modifying an account record, it must transition through defined state nodes that include automated pre-execution validation checks and explicit human approval gates.
3.4 Defensive Error Handling
When interacting with non-deterministic components, classical try/catch blocks are insufficient. AI/ML Engineering introduces multi-tiered recovery patterns to preserve system uptime:
- Structural Parsing Retries: If a model returns an invalid JSON string that violates schema limits, the application catches the validation error. It passes the malformed string and the parsing exception back to a lightweight model tier, instructing it to fix the formatting error instantly.
- Graceful Degradation: If a premium frontier model endpoint experiences an outage or hits a rate limit, the application downshifts the request to a local, open-weight Small Language Model (SLM) cluster managed by the internal platform. While the response may lack advanced reasoning depth, the application continues to function safely, ensuring overall business continuity.
4. Engineering Blueprints: Structured Output Implementation
The following Python blueprint demonstrates how an AI/ML Engineer implements Defensive Structured Output Enforcement. It highlights context-window sanitization, strict Pydantic validation, and automated recovery loops to handle malformed model responses cleanly.
import json
import logging
from typing import List, Optional, Dict, Any
from pydantic import BaseModel, Field, ValidationError
# Configuration setup for enterprise application trace insights
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger("SpokeAIMLEngineerRuntime")
# =====================================================================
# 1. Declarative Data Layer Definition (Target Enterprise Schema)
# =====================================================================
class CustomerRiskEntity(BaseModel):
"""Target output schema enforced by the application layer."""
organization_name: str = Field(description="Fully qualified formal entity name.")
tax_identifier: str = Field(description="Validated country-specific corporate tax identification number.")
sanction_check_passed: bool = Field(description="True if the entity has zero flag matching on international lists.")
identified_risk_factors: List[str] = Field(default=[], description="List of discrete risk anomalies spotted.")
confidence_score: float = Field(description="Mathematical certainty assessment bounded between 0.0 and 1.0.")
# =====================================================================
# 2. Mock Infrastructure Dependency Layers
# =====================================================================
class MockPlatformGatewayClient:
"""Simulates an internal platform gateway model call runtime interface."""
def __init__(self):
self.trigger_malformed_response = True
def dispatch_inference(self, prompt: str, schema: Dict[str, Any]) -> str:
# Scenario A: Simulate a typical non-deterministic parsing failure (Emitting toxic prefix prose)
if self.trigger_malformed_response:
self.trigger_malformed_response = False # Set flag to resolve on the subsequent recovery call
return "Here is the raw data analysis array you requested: ```json\n{\n \"organization_name\": \"Global Logistics Inc\",\n \"tax_identifier\": \"TX-998811\",\n \"sanction_check_passed\": \"TRU_VALUE_INVALID\",\n \"identified_risk_factors\": [\"Regional shipping delays\"],\n \"confidence_score\": 0.94\n}\n```"
# Scenario B: Secure Clean Recovery Path Output
return json.dumps({
"organization_name": "Global Logistics Inc",
"tax_identifier": "TX-998811",
"sanction_check_passed": True,
"identified_risk_factors": ["Regional shipping anomalies managed safely"],
"confidence_score": 0.94
})
# =====================================================================
# 3. Defensive AI/ML Application Implementation Execution Layer
# =====================================================================
class RiskEvaluationOrchestrator:
def __init__(self, platform_client: MockPlatformGatewayClient):
self.client = platform_client
def _sanitize_context_window(self, raw_input_payload: str) -> str:
"""Optimizes the token footprint by stripping out boilerplate text junk."""
# Strips out formatting noise to minimize token consumption rates
clean_text = raw_input_payload.strip().replace(" ", " ")
return clean_text
def process_risk_profile(self, target_input_data: str) -> CustomerRiskEntity:
"""
Processes unstructured corporate insights defensively to extract a type-safe Pydantic entity record.
"""
# Context window management optimization step
sanitized_payload = self._sanitize_context_window(target_input_data)
base_prompt = f"Analyze customer notes and output data schema records. Context: {sanitized_payload}"
target_schema_dict = CustomerRiskEntity.model_json_schema()
# Step 1: Initial Inference Attempt
logger.info("Dispatching initial application transaction to the platform gateway.")
raw_output = self.client.dispatch_inference(base_prompt, target_schema_dict)
try:
# Step 2: Attempt standard parsing extraction
cleaned_json = self._extract_json_block_if_present(raw_output)
validated_record = CustomerRiskEntity.model_validate_json(cleaned_json)
return validated_record
except (ValidationError, json.JSONDecodeError) as parsing_exception:
logger.warning(f"Initial validation check failed: {str(parsing_exception)}. Initiating recovery loops.")
# Step 3: Trigger the repair loop fallback path
repair_prompt = (
f"You returned an invalid schema output that failed Pydantic parsing guidelines. "
f"Error details: {str(parsing_exception)}. Fix the json structure to match the "
f"required parameters. Raw data to fix: {raw_output}"
)
# Request clean fix from platform gateway endpoint
repaired_output = self.client.dispatch_inference(repair_prompt, target_schema_dict)
try:
final_clean_json = self._extract_json_block_if_present(repaired_output)
final_validated_record = CustomerRiskEntity.model_validate_json(final_clean_json)
logger.info("Successfully recovered structural data schema integrity via secondary path.")
return final_validated_record
except Exception as system_breakdown:
logger.critical("Critical Error: The structural formatting error could not be recovered automatically.")
raise RuntimeError("System degradation: Failed to extract structured outputs safely.") from system_breakdown
def _extract_json_block_if_present(self, response_text: str) -> str:
"""Helper to isolate raw JSON code blocks from conversational text wraps."""
if "```json" in response_text:
parts = response_text.split("```json")
actual_code_block = parts[1].split("```")[0].strip()
return actual_code_block
return response_text.strip()
# =====================================================================
# 4. Production Execution Test Driver Run
# =====================================================================
if __name__ == "__main__":
mock_gateway = MockPlatformGatewayClient()
orchestrator = RiskEvaluationOrchestrator(mock_gateway)
unstructured_notes = """
CUSTOMER DOSSIER DATA INTAKE REPORT:
Target Entity name is Global Logistics Inc. Operating under registration token identifier TX-998811.
The compliance screening desk performed local checks and no active lists items matched.
Some regional shipping delays were noted in Q3 but do not threaten structural operational viability.
"""
try:
result_entity = orchestrator.process_risk_profile(unstructured_notes)
print("\n--- Final Validated Enterprise Object Output ---")
print(f"Organization: {result_entity.organization_name}")
print(f"Tax Identifier: {result_entity.tax_identifier}")
print(f"Sanction Screening Passed: {result_entity.sanction_check_passed}")
print(f"Confidence Level: {result_entity.confidence_score}")
except Exception as err:
print(f"Process terminated abruptly due to structural exceptions: {err}")
5. State Machine and DAG Orchestration for Agent Workflows
To bridge the gap between abstract architectural design and concrete implementation, engineering teams must leverage orchestration frameworks specifically designed to enforce stateful predictability. Rather than writing raw, unmanaged loop code, AI/ML engineers should implement these Directed Acyclic Graphs (DAGs) and state charts using modern, industry-standard toolsets.
Orchestrators like LangGraph and LlamaIndex Workflows excel at modeling agent interactions as formal state machines with clear nodes and conditional edges. For high-throughput, enterprise-wide orchestration where infrastructure-level state persistence is mandatory, decoupled cloud primitives like AWS Step Functions or Azure Durable Functions should be used to govern the execution path.
Grounding the application layer in these structured frameworks transforms non-deterministic agent loops into auditable, deterministic state charts that can be easily monitored and controlled.
Unconstrained autonomous agents exhibit unpredictable execution paths, non-deterministic token consumption patterns, and a high risk of getting stuck in infinite recursive loops.
Framing an agentic application as a formal state machine dictates legal transitions through state-bound execution nodes, deterministic transition edges, and bounded autonomy, where the agent only reasons within local context windows and explicit business logic rules.

To deliver the predictability required to clear TOGAF Phase G (Implementation Governance), AI/ML Engineers replace unconstrained execution loops with structured architectures based on State Machines or Directed Acyclic Graphs (DAGs).
Structuring Workflows with State Machines and DAGs
By framing an agentic application as a formal state machine, the developer explicitly dictates the legal transitions between distinct operational phases:
- State-Bound Execution Nodes: Each step in the business process is isolated within a specific state node, such as Data Ingestion, Schema Validation, Risk Scoring, or Report Generation. The agent's reasoning capabilities are confined entirely within that node's local context window and restricted tool subset.
- Deterministic Transition Edges: The paths between nodes are controlled by explicit logical conditions. For example, the system cannot transition from the
Risk Scoringstate to theReport Generationstate unless a deterministic validation function returns apassstatus. - Bounded Autonomy: The agent is permitted to make autonomous decisions only within the boundaries of its current state node, such as determining the best linguistic formatting for a specific extract. It is completely barred from autonomously deciding what core business step to take next.
When NOT to Give an Agent Autonomous Freedom
Enterprise architects must enforce absolute boundaries on agent autonomy. As a strict rule, complete autonomous freedom must be denied in the following operational scenarios:
-
State-Changing Operations on Systems of Record: Agents must never be given unvetted write access to core corporate databases, financial ledgers, or customer profile records. All state-modifying actions must be routed through a deterministic queue that includes an explicit Human-in-the-Loop (HITL) approval gate.
-
External-Facing Communication Channels: Systems that automatically dispatch emails, generate legal notifications, or publish customer-facing advice must have their final payloads validated by structural schemas and toxicity guardrails before transmission, rather than allowing an agent to broadcast directly.
-
Cross-Domain Security and Privilege Boundaries: An agent must never be allowed to dynamically evaluate its own authorization limits or decide whether it has permission to view a sensitive file. Access control lists (ACLs) must be checked and enforced outside the model context by the application runtime.
-
Multi-Step Compounding Compliance Workflows: If an international regulation mandates that Step A must be fully completed and audited before Step B can begin, the orchestrator framework must enforce this sequence through a hard-coded DAG workflow. The agent cannot be allowed to bypass or merge steps to optimize speed.
By enforcing a state machine architecture, the AI CoE ensures that application logs are cleanly auditable, operational token expenses remain predictable, and the overall system behaves safely under peak production enterprise workloads.
6. Mapping AI/ML Engineering Duties to TOGAF Phase G (Implementation Governance)
To achieve true enterprise scalability, the code produced by AI/ML Engineers cannot be deployed using ad-hoc DevOps scripts. It must be governed by a formal lifecycle framework. Within the TOGAF 10 Framework, Phase G (Implementation Governance) provides the structural oversight necessary to verify that the implemented systems conform strictly to the defined AI Architectural North Star.
The primary objective of the AI Center of Excellence (CoE) during Phase G is to shift from open-ended innovation to rigorous, evidence-based validation. The AI/ML Engineer's implementation code must pass through four distinct governance verification gates managed by the central platform before receiving production deployment sign-off.

1. Context Window Integrity and Cost Governance
- Engineering Duty: Optimizing token placement strategies and context window packing frameworks to avoid model degradation.
- Phase G Alignment: The governance framework runs automated checks on the engineer's ingestion code. It validates that retrieval functions use appropriate chunk sizes and that sorting logic places high-priority context at the outer edges of the payload. Any pipeline that naively passes unpruned data blocks into expensive frontier endpoints is rejected to protect the enterprise token budget from cost spikes.
2. Deterministic Schema Enforcement Audits
- Engineering Duty: Building strict type-safe structural schemas, such as Pydantic validation decorators and JSON parsing rules, into application runtimes.
- Phase G Alignment: Phase G compliance mandates that no probabilistic application may output unparsed conversational text directly to an enterprise system of record. Automated code scans inspect the engineer's application repositories to verify that all interface parameters are tightly constrained by grammar rules or structural models at the platform gateway level, guaranteeing syntactic predictability.
3. State Chart and DAG Structural Boundary Verification
- Engineering Duty: Restricting agent autonomy by coding workflows inside Directed Acyclic Graphs (DAGs) and explicit state charts.
- Phase G Alignment: The CoE reviews the application's state transition matrices. The governance team verifies that all state-modifying actions, such as updates, payments, or file modifications, are bounded by deterministic edges that include an explicit Human-in-the-Loop (HITL) approval step. Applications utilizing open-ended, unconstrained autonomous loops are quarantined and blocked from production deployment.
4. Non-Deterministic Regression Benchmarking
- Engineering Duty: Curating domain-specific test sets and evaluating system accuracy using automated grading setups.
- Phase G Alignment: As part of the formal Phase G sign-off, the application is subjected to adversarial evaluation within the staging framework. The system must process historical test batches, and its output is graded by an independent LLM-as-a-Judge scoring engine. The codebase is only cleared for production if it achieves a faithfulness score ( \geq 0.95 ) and successfully stops simulated prompt injection attacks via its semantic firewall layer.
By aligning these development tasks with TOGAF Phase G, the enterprise ensures that the deployment of probabilistic code remains as structured, measurable, and safe as classical IT infrastructure.
Architectural Disclaimer
This architectural guide and its referenced governance structures are intended exclusively for educational and strategic planning purposes. Generative AI runtimes introduce non-deterministic behaviors that vary based on environmental data context, model versioning shifts, and downstream dependencies. Implementing comprehensive deployment pipelines requires extensive, independent security verification and compliance auditing tailored to your specific organizational requirements.