Automated Compliance Analysis for Data Breach Events

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying and managing compliance-related information in data breach events are inefficient and error-prone, particularly due to the reliance on manual review of unstructured data, which can lead to missed notifications and non-compliance with stringent regulations like GDPR and CCPA, especially when dealing with complex and varied data types such as image files.

Innovation Solution

A method involving a combination of machine learning and human review to analyze structured, unstructured, and semi-structured data files, including image files, to identify protected information elements, and generate compliance-related databases for timely notifications, using a machine learning framework to automate the process and integrate human validation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual review methods are used to identify protected information in data breach events, then human judgment can validate results, but the process is time-consuming and error-prone

Engineering Contradiction:
Improveaccuracy of protected information identificationVSAvoidtime to complete data breach analysis
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical review processes with an automated machine learning system that uses natural language processing and image recognition algorithms to identify protected information elements in structured, unstructured, and semi-structured data files, dramatically reducing analysis time while maintaining high accuracy through automated validation mechanisms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces a hybrid system that acts as an intermediary between fully automated machine learning analysis and human review, using the ML system to pre-process and flag potential protected information, then selectively engaging human reviewers only for ambiguous cases, thus optimizing both speed and accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated machine learning methods are used to analyze data files, then processing speed increases, but accuracy may decrease without human validation

Engineering Contradiction:
Improvespeed of data breach responseVSAvoidaccuracy of protected information detection
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback loops where the machine learning system continuously learns from human reviewer corrections and validations, refining its algorithms to improve detection accuracy over time while maintaining high processing speeds through automated iteration and model retraining on validated datasets

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies partial automation where the machine learning system performs comprehensive initial analysis of all data files to identify potential protected information, then applies selective human review only to cases with lower confidence scores or ambiguous classifications, achieving both speed and accuracy

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If comprehensive review of all data files is conducted, then compliance accuracy improves, but resource requirements increase

Engineering Contradiction:
Improvecompliance notification accuracyVSAvoidcomputational resources required
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by directing comprehensive analysis resources specifically to data files and sections that contain or are likely to contain protected information elements, while using lighter processing for clearly non-sensitive files, thus achieving high compliance accuracy without proportionally increasing resource consumption across all files

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the data breach analysis process into distinct phases: initial automated scanning to identify potential protected information, detailed analysis of flagged sections, and validation steps, allowing resources to be concentrated on critical analysis tasks rather than uniformly processing all data files

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If human reviewers manually analyze unstructured data including image files, then complex data types can be validated, but the process becomes highly complex and slower

Engineering Contradiction:
Improvevalidation accuracy of unstructured dataVSAvoidcomplexity of review process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual human review of unstructured data and image files with specialized machine learning components including optical character recognition (OCR) for images, natural language processing for unstructured text, and automated classification algorithms that can validate complex data types with high precision while simplifying the overall process through automation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP4049161B1Systems and methods for identifying compliance-related information associated with data breach events
Publication Date: 2025.10.08 CANOPY SOFTWARE INC
  • EP4049161B1 patent drawingFigure 1A
  • EP4049161B1 patent drawingFigure 1B
  • EP4049161B1 patent drawingFigure 2

AI summary

Various examples are provided related to identification and management of compliance-related information associated with data breach events. In one example, a method includes receiving a first data file collection associated with a first data breach event; generating information associated with presence or absence of protected information elements of all or part of the first data file collection and incorporating data files including the protected information elements in a second data file collection; analyzing data files selected from the second data file collection; and incorporating the information associated with the analysis into machine learning information that may be used for subsequent analysis of data file collections.