Generative AI Model Safeguarding Through Relevance Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generative AI systems often produce hallucinations, leading to mistrust, safety concerns, and ethical issues, particularly in critical fields like healthcare and autonomous driving, where inaccuracies can cause harm and perpetuate misinformation.

Innovation Solution

Implement a model safeguarding process that determines the relevance of user-generated and model-generated content during interactive sessions with generative AI systems to ensure alignment with analysis protocols, preventing inappropriate or irrelevant content by requesting or filtering out non-relevant inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If generative AI systems are used to produce content, then productivity and creativity are improved, but reliability deteriorates due to hallucinations and inaccuracies

Engineering Contradiction:
Improvecontent generation capabilityVSAvoidaccuracy of generated content
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the analysis protocol continuously monitors and evaluates generated content against predefined criteria. The protocol receives generated content, analyzes it for relevance and accuracy, and provides feedback to either allow or block the content from being displayed, thereby improving reliability while maintaining productivity

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The analysis protocol performs preliminary evaluation of generated content before it is displayed to users. By analyzing content in advance against the established protocol criteria, the system prevents inaccurate or irrelevant content from reaching users, thus improving reliability without compromising the speed of content generation

Inventive Principle:
Principle #10Preliminary action

2Reliability

If content is filtered and monitored to improve reliability, then safety and accuracy are improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of generated contentVSAvoidcomplexity of safeguarding system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The analysis protocol acts as an intermediary component between the generative AI system and the user interface. It sits in the middle of the content flow, receiving generated content, analyzing it against protocol criteria, and making decisions to allow or block content. This intermediary approach improves reliability while adding minimal complexity compared to integrating safeguards directly into the generation process

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The safeguarding system is segmented into a separate analysis protocol module that operates independently from the generative AI system. This segmentation allows the filtering and monitoring functions to be isolated in a dedicated component, improving reliability without requiring complex integration with the generation process

Inventive Principle:
Principle #1Segmentation

3Reliability

If analysis protocol is applied to all content, then reliability is improved, but processing time increases

Engineering Contradiction:
Improveaccuracy of generated contentVSAvoidprocessing time for content analysis
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The analysis protocol applies partial analysis by evaluating only the critical aspects of generated content against the protocol criteria. Rather than performing exhaustive analysis, the protocol focuses on key relevance and accuracy checks, improving reliability through targeted evaluation while minimizing processing time and computational overhead

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250315521A1System and Method for Safeguarding Models
Publication Date: 2025.10.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250315521A1 patent drawing
  • US20250315521A1 patent drawing
  • US20250315521A1 patent drawing

AI summary

A computer-implemented method, computer program product and computing system for: enabling a generative AI system to effectuate an analysis protocol; enabling a user to utilize the generative AI system during an interactive session concerning the analysis protocol; receiving content during the interactive session, thus defining received content; and determining whether the received content is relevant with respect to the analysis protocol.