Generative AI Output Validation With Source Mapping Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generative AI models are susceptible to hallucinations and errors in their outputs, and existing validation methods are manually intensive and require excessive human intervention.

Innovation Solution

A system tracks data sources used for training generative AI models and provides transparent exposure of these sources to users, using metadata to map and validate responses, optionally with additional AI models for verification and confidence scoring.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If generative AI models are used to produce outputs, then productivity and automation are improved, but reliability and accuracy deteriorate due to hallucinations and errors

Engineering Contradiction:
Improveautomation capabilityVSAvoidoutput accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces an audit log as an intermediary component that records the provenance of training data and sources used by the generative AI model. This audit log serves as a mediator between the AI model's outputs and the user, providing transparent exposure of data sources without interfering with the model's generative capabilities. The audit log enables verification of output accuracy while maintaining automation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual validation methods are used to verify AI outputs, then reliability is improved, but productivity and efficiency deteriorate due to excessive human intervention

Engineering Contradiction:
Improveoutput validationVSAvoidvalidation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements a self-service validation mechanism where the system automatically generates and exposes audit logs containing data source information and provenance metadata. This allows the system to validate its own outputs without requiring manual human intervention. The audit log provides sufficient information for users to independently verify AI outputs, transforming validation from a manual process to an automated self-verification process.

Inventive Principle:
Principle #25Self-service

3Reliability

If transparent exposure of data sources is implemented, then trust and validation capability are improved, but device complexity increases due to additional tracking and metadata generation

Engineering Contradiction:
Improvetrust in outputVSAvoidsystem structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by generating and storing audit log metadata during the training and data processing phases, before the AI model produces outputs. The system pre-records the provenance of training data, data sources, and processing steps in structured format. This preliminary documentation eliminates the need for complex real-time tracking during inference, as the audit information is already prepared and stored for immediate exposure when needed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12608545B2Validating generative artificial intelligence output
Publication Date: 2026.04.21 SALESFORCE INC
  • US12608545B2 patent drawing
  • US12608545B2 patent drawing
  • US12608545B2 patent drawing

AI summary

Methods, apparatuses, and computer-program products are disclosed. The method may include training a generative artificial intelligence (AI) model on a plurality of data sources and generating, based on the training, training log metadata indicating individual data sources of the plurality of data sources. The method may include receiving, from a user device, a generative AI query and generating, using the trained generative AI model and based on one or more data sources of the plurality of data sources, a response to the generative AI query. The method may include mapping one or more portions of the response to the one or more data sources of the plurality of data sources based on the training log metadata and transmitting, to the user device, the response and one or more indications of the one or more data sources based on the mapping and the training log metadata.