The OBD-AI project proposes an agentic AI diagnostic platform that combines vehicle telemetry, diagnostic trouble codes (DTCs), service information, historical repair knowledge, machine learning and large language models.
The central research question is:
How can an AI system combine authoritative automotive knowledge with real-time vehicle data while remaining grounded, explainable, auditable and safe?
This paper proposes that the answer should not be a conventional chatbot.
OBD-AI should instead be designed as a RAG-LLM + MCP agentic diagnostic system.
The architecture separates two fundamentally different capabilities:
- RAG-LLM answers: What does the technical knowledge say?
- MCP tools answer: What is the vehicle reporting right now, and what authorized system operation can be performed?
- Agent orchestration answers: What evidence should be collected, what knowledge should be retrieved, and what diagnostic hypothesis best explains the evidence?
IAS RESEARCH | KEENSOFTWARE
Research & Development White Paper
Model Context Protocol (MCP) and Retrieval-Augmented Generation (RAG-LLM) for OBD-AI
An Agentic, Grounded and Safety-Governed Architecture for Intelligent Vehicle Diagnostics
Version: 1.0 — Research & Development Draft
Date: October 2026
Location: Winnipeg, Manitoba, Canada
Organizations: IAS-Research.com | KeenSoftware / KeenComputer.com | KeenDirect.com
Application Domain: ICE | Hybrid | PHEV | BEV | Fleet Diagnostics | Predictive Maintenance
Research status: This paper is an architectural and R&D proposal. Vehicle-protocol capabilities, OEM diagnostic access, service-manual licensing, MCP implementations, and production safety requirements must be validated against the exact hardware, software and standards versions used in deployment.
Executive Summary
The OBD-AI project proposes an agentic AI diagnostic platform that combines vehicle telemetry, diagnostic trouble codes (DTCs), service information, historical repair knowledge, machine learning and large language models.
The central research question is:
How can an AI system combine authoritative automotive knowledge with real-time vehicle data while remaining grounded, explainable, auditable and safe?
This paper proposes that the answer should not be a conventional chatbot.
OBD-AI should instead be designed as a RAG-LLM + MCP agentic diagnostic system.
The architecture separates two fundamentally different capabilities:
- RAG-LLM answers: What does the technical knowledge say?
- MCP tools answer: What is the vehicle reporting right now, and what authorized system operation can be performed?
- Agent orchestration answers: What evidence should be collected, what knowledge should be retrieved, and what diagnostic hypothesis best explains the evidence?
The design therefore follows a fundamental principle:
RAG retrieves knowledge; MCP retrieves and operates on structured reality; the agent connects the two under explicit policy and safety controls.
For example, when a vehicle reports P0420, the AI should not simply ask an LLM to explain the code.
Instead:
- MCP reads the DTC.
- MCP retrieves VIN and vehicle configuration.
- MCP retrieves freeze-frame and relevant live PIDs.
- The diagnostic agent determines the diagnostic context.
- RAG retrieves the correct service-manual procedure.
- Hybrid/vector/graph retrieval identifies related components and failure modes.
- The LLM produces a diagnostic hypothesis.
- Evidence is attached to each recommendation.
- The system identifies the next verification test.
- Any vehicle-changing operation requires explicit authorization.
This transforms OBD-AI from an AI chatbot into an AI diagnostic engineering platform.
The research architecture also incorporates:
- hybrid RAG;
- metadata-filtered retrieval;
- GraphRAG / Neo4j;
- structured automotive knowledge graphs;
- MCP tool servers;
- agentic planning;
- BDD/Cucumber regression testing;
- human-in-the-loop verification;
- provenance and citations;
- cybersecurity controls;
- telemetry privacy;
- safety gates;
- predictive maintenance;
- fleet management;
- digital-twin and simulation opportunities.
MCP is particularly relevant because the July 28, 2026 specification introduced a stateless protocol core, authorization hardening, caching of list results and other capabilities aimed at scalable agentic infrastructure. (Model Context Protocol Blog)
1. Research Objectives
The OBD-AI R&D program has six primary objectives.
Objective 1 — Ground AI diagnostic reasoning
Reduce hallucination by grounding responses in:
- OEM service manuals;
- technical service bulletins;
- recall information;
- diagnostic standards;
- DTC definitions;
- vehicle-specific procedures;
- validated repair cases;
- engineering knowledge.
RAG originated as an approach for combining parametric language-model knowledge with retrieved external knowledge, particularly for knowledge-intensive tasks. (UCL NLP)
Objective 2 — Connect AI to live vehicle information
The LLM cannot independently know:
- current DTCs;
- freeze-frame values;
- vehicle VIN;
- coolant temperature;
- engine RPM;
- fuel trims;
- oxygen-sensor readings;
- battery information;
- trip history.
MCP provides a standardized mechanism for exposing external tools and resources to AI applications. The current MCP architecture includes tools, resources and prompts as core server primitives. (Model Context Protocol)
Objective 3 — Build an agentic diagnostic workflow
The system should not simply answer questions.
It should:
Observe → Retrieve → Correlate → Hypothesize → Verify → Recommend → Record
This represents a transition from:
Question answering
to:
AI-assisted diagnostic reasoning.
Objective 4 — Establish safety boundaries
The architecture must distinguish between:
Read-only operations
and
vehicle-changing operations.
For example:
|
Operation |
Risk |
|---|---|
|
Read VIN |
Low |
|
Read DTCs |
Low |
|
Read PIDs |
Low |
|
Read freeze frame |
Low |
|
Retrieve manual |
Low |
|
Create work order |
Medium |
|
Clear DTCs |
High |
|
Actuator test |
High |
|
ECU programming |
Very high |
|
HV-system control |
Critical |
This distinction must be implemented in software architecture—not merely described in an LLM prompt.
Objective 5 — Create measurable engineering quality
OBD-AI should be testable like conventional safety-sensitive software.
The project should therefore combine:
- BDD;
- Cucumber/Gherkin;
- unit testing;
- integration testing;
- RAG evaluation;
- MCP tool validation;
- adversarial testing;
- technician review;
- regression testing.
Objective 6 — Establish a scalable commercial architecture
The platform should eventually support:
- individual vehicle owners;
- independent repair shops;
- automotive technicians;
- dealerships;
- fleet operators;
- used-vehicle inspection;
- warranty organizations;
- automotive engineering organizations;
- EV service organizations.
2. The OBD-AI Research Hypothesis
The central hypothesis of this research is:
A vehicle diagnostic AI system will provide more reliable and actionable assistance when real-time structured vehicle evidence is obtained through controlled tools and domain knowledge is retrieved through an auditable RAG architecture, rather than relying on an LLM's parametric knowledge alone.
The architecture can therefore be represented as:
Diagnostic Intelligence=Vehicle Evidence+Domain Knowledge+Agentic Reasoning+VerificationDiagnostic\ Intelligence = Vehicle\ Evidence + Domain\ Knowledge + Agentic\ Reasoning + Verification
where:
Vehicle Evidence=MCP(OBD/CAN/ECU/Fleet)Vehicle\ Evidence = MCP(OBD/CAN/ECU/Fleet)
and:
Domain Knowledge=RAG(Manuals/TSBs/Repair Knowledge)Domain\ Knowledge = RAG(Manuals/TSBs/Repair\ Knowledge)
and:
Diagnostic Decision=LLM(Evidence+Retrieved Knowledge+Policy)Diagnostic\ Decision = LLM(Evidence + Retrieved\ Knowledge + Policy)
This is the foundation of OBD-AI.
3. Why Conventional OBD Diagnostics Are Insufficient
Traditional diagnostic workflows generally resemble:
Vehicle → Scan Tool → DTC → Technician
The scan tool is extremely useful, but a DTC is not necessarily a diagnosis.
For example:
P0420 — Catalyst System Efficiency Below Threshold
does not automatically mean:
Replace catalytic converter.
Possible causes can include:
- exhaust leaks;
- oxygen-sensor problems;
- wiring;
- air/fuel imbalance;
- misfire history;
- exhaust contamination;
- catalyst degradation;
- incorrect operating conditions;
- other upstream faults.
Therefore:
DTC≠Root CauseDTC \neq Root\ Cause
A more realistic diagnostic relationship is:
DTC+VehicleContext+LiveEvidence+ServiceKnowledge→DiagnosticHypothesisDTC + VehicleContext + LiveEvidence + ServiceKnowledge \rightarrow DiagnosticHypothesis
This is precisely where RAG and MCP become important.
4. RAG-LLM: The Knowledge Layer
4.1 Basic architecture
The RAG architecture consists of:
Documents → Ingestion → Chunking → Embedding → Index → Retrieval → Reranking → LLM
For OBD-AI, this should become:
Manual → Document Intelligence → Automotive Metadata → Hybrid Retrieval → Evidence Set → Diagnostic Agent
5. OBD-AI Knowledge Corpus
The knowledge corpus should be divided into authority levels.
Tier 1 — Manufacturer information
- service manuals;
- workshop manuals;
- diagnostic procedures;
- wiring documentation;
- technical service bulletins;
- recalls;
- manufacturer diagnostic information.
Tier 2 — Standards
Examples include:
- SAE J1979;
- SAE J2012;
- ISO 15765;
- ISO 14229;
- ISO 15031;
- related CAN/diagnostic standards.
SAE J1979 defines diagnostic test modes and their communication between OBD systems and test equipment; the current SAE listing identifies J1979_202505 as reaffirmed in May 2025. (SAE Mobilus)
Tier 3 — Engineering knowledge
- textbooks;
- training manuals;
- engineering publications;
- diagnostic theory;
- component behavior.
Tier 4 — Curated field knowledge
- validated repair cases;
- technician observations;
- internal troubleshooting notes;
- fleet maintenance records.
Tier 5 — Community/public information
Potentially useful, but lower authority.
Such information should never automatically override an OEM procedure.
6. Automotive RAG Metadata
Metadata becomes one of the most important components of the system.
Each knowledge object should include:
|
Metadata |
Example |
|---|---|
|
Manufacturer |
Toyota |
|
Model |
Prius |
|
Model year |
2018 |
|
Generation |
Gen 4 |
|
Engine |
2ZR-FXE |
|
Powertrain |
HEV |
|
System |
Emissions |
|
ECU |
Engine ECU |
|
DTC |
P0420 |
|
Component |
Catalytic converter |
|
Procedure |
Catalyst efficiency test |
|
Document |
Service Manual |
|
Section |
Engine Control |
|
Page |
543 |
|
Source |
OEM |
|
Version |
2025 |
|
Region |
North America |
|
License |
Licensed |
|
Confidence |
Verified |
The metadata filter should be applied before semantic ranking whenever possible.
This prevents a highly similar procedure for the wrong vehicle from winning the retrieval competition.
7. Hybrid RAG Architecture
OBD-AI should not rely on vector search alone.
The recommended architecture is:
Layer 1 — Lexical retrieval
Useful for:
- P0420;
- P0171;
- sensor numbers;
- connector IDs;
- part numbers;
- ECU names.
Layer 2 — Vector retrieval
Useful for:
- natural-language questions;
- symptom descriptions;
- similar diagnostic procedures;
- semantic relationships.
Layer 3 — Metadata filtering
Used for:
- make;
- model;
- year;
- engine;
- powertrain;
- system;
- region.
Layer 4 — Reranking
A cross-encoder or equivalent model evaluates candidate passages.
Layer 5 — Graph retrieval
Neo4j/GraphRAG can connect:
Vehicle → ECU → DTC → Sensor → Component → Failure Mode → Procedure → Manual
GraphRAG research demonstrates the value of graph-based retrieval for connecting entities and answering questions across large private corpora. (arXiv)
8. Proposed OBD-AI Knowledge Graph
A representative ontology is:
Vehicle ├── hasVIN ├── hasEngine ├── hasPowertrain ├── containsECU │ └── DiagnosticSession ├── reportsDTC │ └── DTC │ ├── affectsSystem │ ├── relatesToComponent │ └── hasFailureMode │ ├── containsPID ├── containsFreezeFrame └── producesEvidence Component ├── hasSensor ├── hasFailureMode ├── hasInspectionProcedure └── referencedByManual
This allows questions such as:
Which components are associated with P0420 on this engine?
or:
Which previous cases showed this combination of P0171 + high fuel trim + MAF deviation?
Graph retrieval can complement conventional vector RAG rather than replacing it.
9. Model Context Protocol for OBD-AI
MCP provides the tool and context integration layer.
The current MCP documentation defines three important primitives:
- Tools — executable functions;
- Resources — contextual information;
- Prompts — reusable interaction templates. (Model Context Protocol)
For OBD-AI, the most important primitive is the tool.
10. Proposed MCP Server Architecture
Instead of one enormous MCP server:
OBD-AI MCP └── everything
the recommended architecture is:
OBD-AI Agent | +----------+----------+ | | MCP Gateway RAG Gateway | | +------+------+ +-----+------+ | | | | | Telemetry Diagnostics Fleet Knowledge Graph MCP MCP MCP MCP DB | Vehicle / Puck
This gives each bounded context a clear security boundary.
11. Proposed MCP Servers
11.1 telemetry-mcp
Tools:
get_live_pids() get_freeze_frame() get_trip_history() get_sensor_history()
Purpose:
Provide real-time and historical vehicle measurements.
11.2 diagnostics-mcp
Tools:
read_dtcs() decode_dtc() get_readiness_monitors() get_diagnostic_status()
This server should primarily be read-only.
11.3 vehicle-info-mcp
Tools:
decode_vin() get_vehicle_profile() get_engine_configuration() get_powertrain_configuration() get_recall_status()
11.4 knowledge-mcp
Tools:
search_manuals() get_procedure() search_tsbs() get_source() get_citation()
This server becomes the controlled interface to RAG.
11.5 predictive-mcp
Tools:
get_component_health() forecast_failure() get_anomaly_score() get_maintenance_prediction()
11.6 fleet-mcp
Tools:
list_vehicles() get_service_history() get_vehicle_status() create_work_order() update_work_order()
Business-system writes should require their own authorization.
11.7 control-mcp
This is the most restricted server.
Potential tools:
clear_dtcs() run_actuator_test() execute_diagnostic_routine()
Future functionality might include ECU programming, but that should be treated as a substantially different safety and cybersecurity category.
12. RAG and MCP: Division of Responsibility
The most important architectural decision is:
|
Question |
RAG |
MCP |
|---|---|---|
|
What does P0420 mean? |
✓ |
✓ structured lookup |
|
What is the manufacturer's procedure? |
✓ |
MCP exposes RAG |
|
What vehicle is connected? |
✓ |
|
|
What DTCs are currently stored? |
✓ |
|
|
What is coolant temperature now? |
✓ |
|
|
What happened during last trip? |
✓ |
|
|
What does the manual say? |
✓ |
✓ gateway |
|
What component is associated with a failure mode? |
✓ |
✓ graph/tool |
|
Create work order |
✓ |
|
|
Clear DTC |
✓ gated |
|
|
Run actuator test |
✓ gated |
The design principle is:
Use RAG for knowledge. Use MCP for authoritative structured state and controlled actions.
13. Agentic Diagnostic Loop
The OBD-AI agent should implement an iterative diagnostic loop.
OBSERVE ↓ IDENTIFY VEHICLE ↓ READ DTC / TELEMETRY ↓ BUILD DIAGNOSTIC CONTEXT ↓ RETRIEVE KNOWLEDGE ↓ CORRELATE EVIDENCE ↓ GENERATE HYPOTHESES ↓ SELECT NEXT TEST ↓ VERIFY ↓ UPDATE HYPOTHESIS ↓ RECOMMEND REPAIR ↓ RE-TEST
This is substantially more powerful than:
Question → LLM → Answer
14. Example: P0420 Diagnostic Workflow
A technician asks:
"Why is my check-engine light on?"
Step 1 — MCP
read_dtcs()
Result:
P0420
Step 2 — Vehicle context
decode_vin() get_vehicle_profile()
Result:
2018 Toyota Prius 2ZR-FXE HEV
Step 3 — Freeze frame
get_freeze_frame()
Example:
RPM: 1,850 Coolant: 88°C Load: 42% Vehicle speed: 72 km/h
Step 4 — Structured DTC interpretation
decode_dtc("P0420")
Step 5 — RAG
Query:
Catalyst efficiency below threshold Toyota Prius 2018 2ZR-FXE P0420
Metadata:
make=Toyota model=Prius year=2018 engine=2ZR-FXE powertrain=HEV system=emissions DTC=P0420
Step 6 — Graph reasoning
Potential relationships:
P0420 | +-- catalyst +-- oxygen sensors +-- exhaust leakage +-- fuel mixture +-- misfire +-- emissions ECU
Step 7 — Agent reasoning
The agent should not immediately say:
Replace the catalytic converter.
Instead:
P0420 is present. Based on the vehicle-specific diagnostic procedure and current evidence, several causes remain possible. The next recommended verification is X because the manufacturer procedure identifies it as a prerequisite before catalyst replacement.
This is an important distinction between AI-generated speculation and evidence-based diagnostic assistance.
15. Agentic RAG
Traditional RAG:
Question ↓ Retrieve ↓ Generate
OBD-AI should use:
Question ↓ Vehicle context ↓ Tool selection ↓ Structured evidence ↓ Query decomposition ↓ Multi-source retrieval ↓ Reranking ↓ Evidence validation ↓ Reasoning ↓ Next diagnostic action ↓ Verification
This can be described as:
Agentic Diagnostic RAG
rather than simple RAG.
16. Self-Reflective and Corrective RAG
Research such as Self-RAG demonstrates the value of adaptive retrieval and reflection for improving factuality and citation accuracy. (ICLR Proceedings)
OBD-AI can adapt the concept without necessarily requiring the exact Self-RAG training architecture.
For example:
Retrieval confidence
High ↓ Generate Medium ↓ Retrieve additional sources Low ↓ Abstain
Corrective RAG research similarly proposes evaluating retrieval quality before allowing the generation stage to rely on potentially poor evidence. (arXiv)
This is particularly valuable for automotive diagnostics.
17. Evidence Fusion
OBD-AI should combine four evidence categories.
E1 — Vehicle evidence
DTC PID Freeze frame VIN Trip history
E2 — Manufacturer evidence
Service manual TSB Recall Diagnostic procedure
E3 — Historical evidence
Previous repairs Fleet cases Technician observations
E4 — AI inference
Hypothesis Probability Next test
The UI should distinguish them.
For example:
|
Evidence |
Source |
|---|---|
|
P0420 active |
Vehicle |
|
1,850 RPM |
Vehicle |
|
Catalyst test procedure |
OEM manual |
|
Similar previous case |
Fleet history |
|
Probable cause |
AI inference |
This prevents the dangerous impression that every statement came from the manufacturer.
18. MCP Security Architecture
MCP creates a new attack surface because an AI agent can interact with external tools.
The security architecture should therefore include:
User ↓ Mobile Authentication ↓ AI Gateway ↓ Policy Engine ↓ MCP Authorization ↓ MCP Server ↓ Vehicle / Enterprise System
Controls should include:
- OAuth/OIDC where appropriate;
- scoped credentials;
- per-tool authorization;
- tenant isolation;
- rate limiting;
- audit logging;
- server allowlists;
- input validation;
- schema validation;
- replay protection;
- network segmentation.
The July 2026 MCP specification specifically expanded authorization and infrastructure-oriented capabilities, making protocol-version pinning and security review important implementation requirements. (Model Context Protocol Blog)
19. Prompt Injection
A particularly important RAG threat is indirect prompt injection.
Suppose an attacker inserts malicious text into:
- a service document;
- uploaded diagnostic notes;
- fleet history;
- an external webpage;
- a knowledge-base record.
The LLM might interpret that text as instructions.
OWASP identifies prompt injection as a major LLM application risk and distinguishes direct and indirect forms of manipulation. (OWASP Gen AI Security Project)
Therefore:
Retrieved documents must be treated as data, not instructions.
A retrieved passage should never be allowed to:
- change MCP permissions;
- authorize an action;
- override system policy;
- execute arbitrary code;
- bypass confirmation.
20. Control-MCP Safety Gate
Vehicle-changing tools require a fundamentally different workflow.
Unsafe architecture
LLM ↓ clear_dtcs()
Proposed architecture
LLM ↓ Prepare action ↓ Policy Engine ↓ Safety Preconditions ↓ Mobile Confirmation ↓ Snapshot ↓ control-mcp ↓ Vehicle ↓ Verify result ↓ Audit record
The LLM should request an operation.
The policy engine should authorize it.
The vehicle interface should execute it.
21. Pre-Action Snapshot
Before clearing DTCs:
Vehicle ID VIN Timestamp DTC list Pending DTCs Permanent DTCs Freeze-frame Readiness status Relevant PIDs Diagnostic session Technician identity
should be captured.
This creates an evidence chain:
Before repair ↓ Diagnosis ↓ Repair ↓ Clear ↓ Drive cycle ↓ After repair
This is extremely valuable for warranty, fleet and professional workshop environments.
22. High-Voltage EV Safety
The architecture must treat:
- BEV;
- HEV;
- PHEV;
- high-voltage battery;
- inverter;
- DC/DC converter;
- electric motor;
- HV interlock;
as a separate safety domain.
OBD-AI should never infer that a generic diagnostic instruction is sufficient for HV work.
Instead:
EV detected ↓ HV system involved ↓ Safety procedure required ↓ Qualified technician requirement ↓ Manufacturer procedure ↓ Human verification
The AI should provide decision support rather than authorize hazardous physical work.
23. Privacy and Data Governance
Vehicle telemetry may become personal information when associated with:
- VIN;
- driver identity;
- GPS;
- trip history;
- driving behavior;
- fleet employee;
- service history.
OBD-AI should therefore implement:
- data minimization;
- retention policies;
- tenant separation;
- encryption;
- access control;
- consent;
- deletion policies;
- audit trails.
NIST AI RMF provides a useful governance framework for trustworthy AI development, while NIST CSF 2.0 provides a broader cybersecurity risk-management framework. (NIST)
24. Bounded Context Architecture
The OBD-AI domain can be structured into five major bounded contexts.
|
Context |
Responsibility |
|---|---|
|
Vehicle Telemetry |
PIDs, sensor data, trips |
|
Diagnostics |
DTCs, readiness, diagnostic state |
|
Vehicle Information |
VIN, model, configuration |
|
Knowledge & Advisory |
Manuals, TSBs, procedures |
|
Fleet Management |
Vehicles, maintenance, work orders |
A sixth restricted context is recommended:
|
Context |
Responsibility |
|---|---|
|
Vehicle Control |
Controlled diagnostic actions |
This follows Domain-Driven Design principles and prevents the AI agent from becoming a monolithic software component.
25. Recommended OBD-AI Technology Stack
A practical R&D architecture can use:
Vehicle / Edge
- OBD-II;
- CAN;
- ISO 15765-4;
- appropriate diagnostic interfaces;
- ARM/STM32-class edge hardware;
- OBD-AI Puck.
SAE J1979's current specification explicitly covers OBD diagnostic communication and references DoCAN/ISO 15765-4 among the underlying communication technologies. (SAE Mobilus)
Mobile
- Android/iOS application;
- secure device authentication;
- technician UI;
- human confirmation;
- telemetry visualization.
Backend
- Python;
- FastAPI;
- Docker;
- MQTT where appropriate;
- PostgreSQL;
- Redis.
RAG
- RAGFlow;
- document parsing;
- vector database;
- hybrid search;
- reranking;
- metadata filtering.
LLM
Potential R&D environments:
- Ollama;
- Hugging Face;
- domain-specific or general instruction models.
Graph
- Neo4j;
- Cypher;
- graph-based retrieval;
- vehicle/component/DTC ontology.
Agent
- MCP;
- agent orchestration;
- tool policies;
- structured outputs;
- state management.
Engineering
- Git;
- Docker;
- CI/CD;
- Cucumber;
- BDD;
- automated evaluation.
The prior OBD-AI architecture work provides a strong foundation for combining RAGFlow/Ollama/Hugging Face with Neo4j and vehicle diagnostic data.
26. Eight Core OBD-AI Use Cases
UC1 — Guided DTC Diagnosis
Input
P0420
MCP
read_dtcs decode_vin get_freeze_frame decode_dtc
RAG
Vehicle-specific diagnostic procedure.
Output
- probable causes;
- evidence;
- recommended test;
- manual citation.
UC2 — Powertrain-Aware Diagnosis
Same DTC, different:
- gasoline;
- diesel;
- HEV;
- PHEV;
- BEV.
The agent retrieves different procedures based on vehicle metadata.
UC3 — Technician Manual Assistant
Example:
"What is the torque sequence for this cylinder head?"
MCP establishes the vehicle.
RAG retrieves:
- correct engine;
- correct procedure;
- torque specification;
- sequence.
UC4 — Pre-Purchase Vehicle Health
OBD-AI evaluates:
- DTCs;
- pending codes;
- readiness;
- live PIDs;
- recalls;
- historical data.
Output:
Vehicle Health Report
UC5 — Predictive Maintenance
MCP supplies:
trip history component health service history
RAG supplies:
failure modes maintenance intervals TSBs inspection procedures
Agent generates:
Predicted maintenance recommendation + supporting evidence.
UC6 — Fleet Triage
Hundreds of vehicles can be prioritized:
STOP SERVICE NOW SERVICE SOON MONITOR NORMAL
The classification should be policy-based rather than purely LLM-generated.
UC7 — Intermittent Fault Analysis
The system correlates:
time RPM temperature load speed DTC sensor values
with known diagnostic patterns.
This creates an AI-assisted engineering workflow for faults that technicians cannot reproduce easily.
UC8 — Post-Repair Verification
The system:
- records pre-repair evidence;
- validates repair;
- optionally clears DTCs after confirmation;
- retrieves manufacturer drive-cycle requirements;
- monitors readiness;
- produces a repair-verification report.
27. Predictive Maintenance Research
The platform can eventually move from:
Reactive diagnostics
to:
Predictive diagnostics.
For example:
HealthScore=f(DTCFrequency,Temperature,OperatingHours,SensorDrift,Mileage,ServiceHistory)HealthScore = f( DTCFrequency, Temperature, OperatingHours, SensorDrift, Mileage, ServiceHistory )
The model can estimate:
P(Failure∣ObservedEvidence)P(Failure|ObservedEvidence)
But the prediction must remain distinct from the manufacturer's confirmed diagnostic procedure.
Thus:
Prediction is not diagnosis.
The system should explicitly label:
- measured;
- retrieved;
- inferred;
- predicted.
28. Digital Twin Extension
A future research direction is the OBD-AI digital twin.
Architecture:
Physical Vehicle ↓ OBD/CAN ↓ Digital Vehicle State ↓ Knowledge Graph ↓ Simulation Model ↓ AI Agent
The project could eventually integrate:
- SystemC/TLM;
- virtual ECUs;
- CAN simulation;
- diagnostic replay;
- synthetic vehicle data.
This would allow the team to test diagnostic agents without always requiring a physical vehicle.
29. BDD and Cucumber as the AI Safety Net
The existing BDD direction is especially valuable.
A diagnostic scenario can become an executable AI test.
Example:
Feature: P0420 diagnostic reasoning Scenario: P0420 on 2018 hybrid vehicle Given a vehicle identified as a 2018 hybrid vehicle And the vehicle reports P0420 And freeze-frame data is available When the technician requests a diagnosis Then the agent retrieves the vehicle-specific procedure And the retrieved procedure matches the engine configuration And the answer cites the source And the answer does not recommend catalyst replacement without the required verification steps
This is more powerful than manually checking chatbot responses.
30. MCP Tool-Calling Tests
BDD should also test the agent's tools.
Example:
Scenario: Clearing DTCs requires confirmation Given P0420 is stored When the technician asks to clear the code Then the agent prepares a clear request And a pre-action snapshot is created And confirmation is requested And control-mcp is not called before confirmation
This turns security policy into executable software requirements.
31. RAG Evaluation Framework
The RAG system should be evaluated independently of the LLM.
Retrieval metrics
- Recall@k;
- Precision@k;
- MRR;
- nDCG;
- wrong-vehicle retrieval rate.
Grounding metrics
- citation coverage;
- citation correctness;
- unsupported-claim rate;
- abstention accuracy.
Agent metrics
- correct tool selection;
- tool ordering;
- parameter validity;
- unnecessary tool calls;
- unsafe tool calls.
Operational metrics
- latency;
- token usage;
- infrastructure cost;
- MCP response time;
- retrieval latency.
32. Proposed Quality Score
A research-level composite score can be defined as:
OBD-AI Quality=wRR+wGG+wTT+wSS+wHHOBD\text{-}AI\ Quality = w_RR + w_GG + w_TT + w_SS + w_HH
where:
- RR = retrieval quality;
- GG = grounding;
- TT = tool correctness;
- SS = safety;
- HH = human-rated usefulness.
However, safety should not simply be averaged into general quality.
A single unsafe action should be treated as a blocking failure.
Therefore:
SafetyFailure⇒ReleaseFailureSafetyFailure \Rightarrow ReleaseFailure
for critical control scenarios.
33. RAG Failure Modes
|
Failure |
Mitigation |
|---|---|
|
Wrong vehicle manual |
Metadata filtering |
|
Wrong model year |
VIN filtering |
|
Missing procedure |
Abstention |
|
Bad PDF extraction |
Document QA |
|
Diagram lost |
Figure-aware ingestion |
|
Wrong ranking |
Reranking |
|
Outdated manual |
Version metadata |
|
Conflicting documents |
Authority ranking |
|
Hallucination |
Citation + verification |
|
Prompt injection |
Treat retrieval as data |
34. MCP Failure Modes
|
Failure |
Control |
|---|---|
|
Wrong tool |
Tool policy |
|
Wrong parameter |
Schema validation |
|
Unauthorized action |
Authorization |
|
Replay |
Request identifiers |
|
Tool compromise |
Server isolation |
|
Excessive calls |
Rate limits |
|
Dangerous command |
Safety policy |
|
Stale tool metadata |
Versioning |
|
Protocol change |
Version pinning |
|
Audit gap |
Immutable logging |
35. Agent Failure Modes
The agent may:
- retrieve too much;
- retrieve too little;
- select the wrong tool;
- infer unsupported causes;
- confuse correlation with causation;
- recommend replacement too early;
- ignore safety constraints;
- over-trust historical cases;
- fail to abstain.
Therefore:
The agent must be treated as an uncertain reasoning component operating inside a deterministic engineering envelope.
36. Recommended Agent Policy
The agent should follow a hierarchy:
1. Safety policy 2. Authorization policy 3. Vehicle facts 4. Manufacturer evidence 5. Engineering standards 6. Curated knowledge 7. Historical cases 8. Statistical prediction 9. LLM reasoning
This hierarchy prevents the LLM from overriding stronger sources.
37. Explainability Model
Every diagnostic conclusion should ideally have an evidence chain:
Conclusion ↓ Evidence ↓ Vehicle data ↓ Diagnostic procedure ↓ Knowledge source ↓ Citation
For example:
Hypothesis: Possible catalyst-efficiency problem.
Evidence
- P0420 active;
- freeze-frame conditions;
- vehicle configuration;
- oxygen-sensor observations.
Knowledge
- manufacturer diagnostic procedure;
- relevant TSB.
Next verification
- manufacturer-defined test.
This is much more useful to a professional technician than a paragraph of unsupported AI reasoning.
38. Human-in-the-Loop Architecture
OBD-AI should implement three operating levels.
Level 1 — Informational
AI answers
No vehicle action.
Level 2 — Assisted
AI recommends Human verifies
Examples:
- diagnostic test;
- service procedure;
- work order.
Level 3 — Controlled
AI proposes Policy validates Human confirms System executes System verifies
Examples:
- clear DTC;
- actuator test.
This hierarchy should be central to product architecture.
39. R&D Roadmap
Phase 1 — Foundation
Build:
- Puck data interface;
- diagnostics-mcp;
- vehicle-info-mcp;
- knowledge-mcp;
- small licensed document corpus;
- metadata model;
- vector retrieval;
- citations.
Target: UC1 + UC3.
Phase 2 — Contextual Intelligence
Add:
- telemetry-mcp;
- hybrid retrieval;
- reranking;
- Neo4j;
- GraphRAG;
- BDD evaluation;
- automated regression.
Target: UC2 + UC4.
Phase 3 — Agentic Diagnostics
Add:
- diagnostic planning;
- evidence fusion;
- predictive-mcp;
- historical cases;
- self/corrective retrieval;
- confidence scoring.
Target: UC5 + UC7.
Phase 4 — Fleet Intelligence
Add:
- fleet-mcp;
- service history;
- work orders;
- fleet dashboards;
- predictive maintenance.
Target: UC6.
Phase 5 — Controlled Vehicle Actions
Add:
- control-mcp;
- confirmation workflow;
- snapshots;
- safety policy engine;
- actuator testing;
- comprehensive audit trail.
Target: UC8.
40. Research Work Packages
The project can be organized into eight R&D work packages.
WP1 — Automotive Knowledge Engineering
Develop:
- ontology;
- DTC knowledge;
- service-manual ingestion;
- metadata.
WP2 — RAG Engineering
Research:
- hybrid retrieval;
- reranking;
- GraphRAG;
- corrective retrieval;
- citation systems.
WP3 — MCP Engineering
Develop:
- MCP servers;
- schemas;
- authorization;
- versioning;
- gateway.
WP4 — Agentic Reasoning
Research:
- diagnostic planning;
- tool selection;
- evidence fusion;
- confidence;
- abstention.
WP5 — Embedded Vehicle Interface
Develop:
- Puck;
- CAN/OBD;
- telemetry acquisition;
- secure communications.
WP6 — Safety and Security
Research:
- prompt injection;
- tool abuse;
- authorization;
- HV safety;
- auditability.
WP7 — AI Testing
Develop:
- BDD;
- Cucumber;
- synthetic cases;
- replay datasets;
- adversarial tests.
WP8 — Commercialization
Develop:
- technician application;
- fleet platform;
- API;
- SaaS architecture;
- enterprise deployment.
41. Strategic Role of IAS-Research, KeenSoftware and KeenDirect
The OBD-AI program naturally supports a three-part R&D-to-commercialization model.
IAS-Research.com
Research and intellectual property
Responsibilities:
- AI/RAG research;
- MCP architecture;
- diagnostic ontology;
- GraphRAG research;
- evaluation methodology;
- embedded/AI research;
- white papers;
- patents/IP exploration;
- academic and industrial collaboration.
KeenComputer.com / KeenSoftware
Engineering and deployment
Responsibilities:
- software engineering;
- mobile applications;
- cloud/server infrastructure;
- cybersecurity;
- DevOps;
- AI integration;
- MCP implementation;
- fleet applications;
- customer deployment.
KeenDirect.com
Hardware and supply-chain layer
Responsibilities:
- OBD interfaces;
- Puck hardware;
- CAN components;
- ARM/embedded hardware;
- sensors;
- accessories;
- prototype and production supply.
This produces a complete:
Research → Engineering → Hardware → Product → Commercialization
pipeline.
42. Potential Commercial Products
The research architecture can lead to several products.
OBD-AI Consumer
- vehicle health check;
- DTC explanation;
- pre-trip check;
- used-car inspection.
OBD-AI Technician
- AI diagnostic assistant;
- service-manual search;
- evidence-driven troubleshooting;
- repair verification.
OBD-AI Fleet
- fleet monitoring;
- predictive maintenance;
- vehicle triage;
- work-order integration.
OBD-AI Enterprise
- APIs;
- dealer integration;
- warranty;
- fleet analytics;
- enterprise knowledge bases.
43. Intellectual Property Opportunities
Potential IP areas include:
- Vehicle-context-aware agentic RAG
- MCP-based automotive diagnostic orchestration
- DTC-to-evidence graph reasoning
- Safety-gated vehicle control agents
- Diagnostic evidence provenance architecture
- BDD-based autonomous diagnostic-agent evaluation
- Predictive-maintenance + service-manual reasoning
- Vehicle digital-twin + RAG diagnostic systems
The research program should document novel mechanisms carefully before public disclosure if patent protection is contemplated.
44. Research Questions for Future Publications
The OBD-AI program can generate multiple academic/industrial research papers.
RQ1
Does vehicle-specific metadata filtering significantly improve diagnostic RAG accuracy?
RQ2
Does GraphRAG improve multi-component fault reasoning compared with vector RAG?
RQ3
Does MCP improve tool interoperability and auditability in automotive AI agents?
RQ4
Can BDD scenarios provide an effective regression framework for agentic RAG systems?
RQ5
Does evidence fusion reduce premature component replacement?
RQ6
Can predictive maintenance models be improved by combining telemetry with repair knowledge?
RQ7
How should AI agents be safety-gated when they can invoke vehicle-control operations?
RQ8
Can digital-twin environments provide sufficient synthetic data for automotive diagnostic-agent testing?
45. Proposed Reference Architecture
┌──────────────────────┐ │ Technician │ │ Mobile / Web │ └──────────┬───────────┘ │ ▼ ┌──────────────────────┐ │ OBD-AI Agent │ │ Planner + Reasoner │ └──────────┬───────────┘ │ ┌──────────────┴──────────────┐ │ │ ▼ ▼ ┌───────────────┐ ┌───────────────┐ │ MCP Gateway │ │ RAG Gateway │ └───────┬───────┘ └───────┬───────┘ │ │ ┌────────────┼────────────┐ ┌────────┼────────┐ ▼ ▼ ▼ ▼ ▼ ▼ Telemetry Diagnostics Vehicle Vector Graph Knowledge MCP MCP MCP RAG RAG MCP │ │ │ │ │ └────────────┴─────┬──────┘ └────────┴──────┐ │ │ ▼ ▼ ┌───────────┐ ┌─────────────┐ │ OBD-AI │ │ Manuals / │ │ Puck │ │ TSB / Docs │ └─────┬─────┘ └─────────────┘ │ ▼ ┌─────────────┐ │ Vehicle ECU │ │ CAN / OBD │ └─────────────┘ Restricted path: Agent │ ▼ Policy Engine │ Human Confirmation │ ▼ control-mcp │ ▼ Vehicle Control
46. Core Architectural Principle
The architecture can ultimately be summarized by the following equation:
OBD-AI=MCPVehicle+RAGKnowledge+GraphRelationships+LLMReasoning+BDDVerification+HumanSafetyOBD\text{-}AI = MCP_{Vehicle} + RAG_{Knowledge} + Graph_{Relationships} + LLM_{Reasoning} + BDD_{Verification} + Human_{Safety}
or conceptually:
Observe with MCP → Understand with RAG → Connect with Graph → Reason with LLM → Verify with BDD → Approve with Human Governance.
That is the central research contribution of the proposed OBD-AI architecture.
47. Conclusion
OBD-AI should not be designed as a conventional generative-AI chatbot connected to a vehicle.
It should be engineered as an agentic diagnostic system in which:
- MCP provides controlled access to live vehicle and enterprise systems;
- RAG provides authoritative technical knowledge;
- metadata prevents wrong-vehicle retrieval;
- hybrid search combines lexical and semantic retrieval;
- GraphRAG connects DTCs, components, ECUs, symptoms and procedures;
- the LLM performs contextual reasoning;
- BDD converts diagnostic requirements into executable tests;
- human approval governs safety-sensitive operations;
- audit logs establish evidence and accountability.
The resulting architecture addresses a fundamental weakness of conventional LLM applications:
An LLM can generate a plausible answer without actually knowing what the vehicle is doing.
OBD-AI addresses this by connecting the model to evidence.
Likewise, conventional RAG can retrieve a technically correct paragraph while lacking the current vehicle state.
OBD-AI addresses this by connecting RAG to MCP-derived vehicle context.
The resulting system is therefore not merely:
RAG + MCP
but:
A vehicle-context-aware, evidence-grounded, agentic diagnostic architecture.
The immediate R&D priorities should be:
- finalize the automotive ontology and metadata model;
- implement diagnostics-mcp;
- implement vehicle-info-mcp;
- build the first licensed service-manual RAG corpus;
- integrate hybrid vector + GraphRAG retrieval;
- define typed MCP schemas;
- convert existing diagnostic scenarios into Cucumber/BDD tests;
- establish citation and provenance requirements;
- implement agent safety policies;
- defer control-mcp until the read-only architecture and security/evaluation framework are proven.
This creates a technically defensible path from OBD-II data acquisition to AI-assisted diagnosis, and ultimately toward predictive maintenance, fleet intelligence and automotive digital twins.
References and Research Sources
- Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 / arXiv:2005.11401. (UCL NLP)
- Gao, Y., Xiong, Y., Gao, X., et al. (2023/2024). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997. (DOI)
- Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2024). Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. ICLR 2024. (ICLR Proceedings)
- Yan, S-Q., Gu, J-C., Zhu, Y., & Ling, Z-H. (2024). Corrective Retrieval Augmented Generation. arXiv:2401.15884. (arXiv)
- Edge, D., Trinh, H., Cheng, N., et al. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130. (arXiv)
- Model Context Protocol Specification, Model Context Protocol project. The July 28, 2026 specification introduced the current stateless protocol-core direction, authorization improvements and related infrastructure capabilities. (Model Context Protocol Blog)
- Model Context Protocol — Server Specification, describing MCP tools, resources and prompts. (Model Context Protocol)
- Model Context Protocol TypeScript SDK v2, implementation documentation for the 2026-07-28 specification. (MCP TypeScript SDK)
- SAE International. SAE J1979_202505 — E/E Diagnostic Test Modes, reaffirmed May 23, 2025. (SAE Mobilus)
- SAE International. SAE J1979 / ISO 15031-5 — E/E Diagnostic Test Modes. Diagnostic services and OBD communication framework. (SAE Mobilus)
- ISO. ISO 15765-4 — Road vehicles — Diagnostic communication over Controller Area Network (DoCAN).
- ISO. ISO 14229 — Road vehicles — Unified Diagnostic Services (UDS).
- Smart, J. F. (2014). BDD in Action: Behavior-Driven Development for the Whole Software Lifecycle. Manning.
- Evans, E. (2003). Domain-Driven Design: Tackling Complexity in the Heart of Software. Addison-Wesley.
- OWASP. OWASP Top 10 for Large Language Model Applications 2025. (OWASP Gen AI Security Project)
- NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. (NIST)
- NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. (NIST)
- Pascoe, C., Quinn, S., & Scarfone, K. (2024). The NIST Cybersecurity Framework (CSF) 2.0. NIST CSWP 29. (NIST)
Recommended R&D Deliverables
The next logical engineering artifacts derived from this paper are:
- OBD-AI MCP Server Specification — complete tool names, JSON schemas, permissions and error codes.
- OBD-AI RAG Knowledge Model — metadata schema, chunking strategy, vector/graph indexes and provenance.
- OBD-AI Neo4j Ontology — Vehicle–ECU–DTC–Sensor–Component–Failure–Procedure graph.
- OBD-AI Agent Specification — planner, tool-selection policy, retrieval policy and abstention logic.
- OBD-AI BDD/Cucumber Test Specification — 50–100 executable diagnostic scenarios.
- OBD-AI Security Threat Model — MCP, RAG, prompt injection, vehicle-control and telemetry threats.
- OBD-AI Puck-to-MCP Reference Implementation — OBD/CAN → Puck → mobile → MCP → RAG-LLM.
- OBD-AI Research Dataset — DTC + PID + freeze-frame + vehicle context + manual evidence + expected diagnostic outcome.
These deliverables would turn this white paper from an architecture proposal into a research prototype and engineering roadmap suitable for an IAS-Research/KeenSoftware proof-of-concept and subsequent commercialization program.