Generative AI Data Leak Detection via Watermarking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in detecting unauthorized use of confidential or proprietary data by generative artificial intelligence platforms, making it difficult to determine if such systems have used sensitive information in generating output.

Innovation Solution

A computing platform with a processor, memory, and communication interface generates questions for generative AI platforms to identify unauthorized dissemination of confidential data files, using unique identifying features such as watermarks or invisible text changes, and transmits alerts to administrative devices upon detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If generative artificial intelligence platforms process confidential or proprietary data, then the utility and functionality of the AI system is improved, but the risk of unauthorized data dissemination and data leaks increases

Engineering Contradiction:
Improvefunctionality of AI systemVSAvoidrisk of data leak
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary actions by embedding unique identifying features (watermarks) into confidential data files before they are processed by the generative AI platform. This proactive measure enables subsequent detection of unauthorized dissemination without interfering with the AI's functional processing of the data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary detection system that acts as a mediator between the confidential data and the generative AI platform. This system monitors AI outputs for embedded watermarks and generates alerts, allowing the AI to maintain its functionality while providing a protective layer against data leaks.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If traditional data security measures are implemented to prevent data leaks, then the protection level is improved, but the ability to detect unauthorized use of confidential data by AI systems deteriorates

Engineering Contradiction:
Improvedata protection levelVSAvoiddetection capability of unauthorized AI data use
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

Instead of applying uniform security measures across all data, the system applies local quality by embedding unique identifying features (watermarks) specifically within confidential data files. This targeted approach maintains data usability while enabling precise detection of unauthorized AI-generated outputs containing sensitive information.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent employs watermarking technology that embeds unique identifying features into confidential data, analogous to color changes. These embedded markers are imperceptible in normal data processing but become detectable when the data is improperly disseminated or appears in AI-generated outputs, enabling reliable detection without affecting data functionality.

Inventive Principle:
Principle #32Color changes

3Measurement precision

If comprehensive monitoring of AI outputs is performed to detect data leaks, then the detection accuracy is improved, but the computational resources and system complexity increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts only the essential detection function from comprehensive monitoring by focusing specifically on identifying embedded watermarks in AI outputs. This extraction approach achieves high detection accuracy for confidential data leaks while avoiding the complexity of analyzing all aspects of AI-generated content, thereby reducing overall system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250165755A1Data Leak Detection in Generative Artificial Intelligence Model Output
Publication Date: 2025.05.22 BANK OF AMERICA CORP
  • US20250165755A1 patent drawing
  • US20250165755A1 patent drawing
  • US20250165755A1 patent drawing

AI summary

Aspects of the disclosure relate to detection of confidential information used in generative artificial intelligence platforms. An artificial intelligence computing platform having at least one processor, a memory, and a communication interface may generate questions and ask other generative artificial intelligence platforms the generated questions to determine if the other generative artificial intelligent platforms answers indicate that restricted access, confidential, or proprietary information data has been leaked and disseminated. In an embodiment, prompt injection may be used to determine if external generative artificial intelligence platforms are utilizing an enterprises confidential or proprietary information. A computing platform may transmit, via the communication interface, to an administrative computing device, information regarding the unauthorized dissemination which, when processed by the administrative computing device causes a notification to be displayed on the administrative computing device.