Automated Data Exposure Detection Using ML Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting data leaks are reactive, resource-intensive, and often costly, making it difficult for companies to promptly identify and mitigate data exposure across vast online data sources.

Innovation Solution

A process for automated detection of data leaks involves determining the degree of data exposure by analyzing access paths and content associated with documents, using machine-learning algorithms to generate scores based on predefined characteristics, and aggregating data from multiple sources to identify potential security threats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods are used to search for data leaks by looking for keywords in billions of results on the Internet, then data leak detection capability is improved, but calculating power requirements and data center infrastructure increase significantly

Engineering Contradiction:
Improvedata leak detection capabilityVSAvoiddata center infrastructure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the data leak detection process into multiple independent components: (1) collecting data from multiple sources (search engines, social networks, dark web), (2) processing data through different analysis methods (keyword search, machine learning, natural language processing), and (3) aggregating results to determine exposure degree. This segmentation allows each component to be optimized independently and reduces the need for a single massive data center infrastructure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that aggregates data from multiple sources and applies various analysis methods before final evaluation. This intermediary layer (the automated detection system itself) mediates between raw data sources and the final detection output, enabling efficient processing without requiring direct access to billions of results simultaneously.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If conventional keyword search methods are used to detect data leaks across billions of Internet results, then detection coverage is improved, but electric consumption becomes prohibitive

Engineering Contradiction:
Improvedetection coverageVSAvoidelectric consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by selectively processing only the most relevant data sources and using different analysis intensities for different data types. Not all data requires full machine learning analysis - some can be processed with simpler keyword matching or heuristics, reducing overall energy consumption while maintaining detection coverage.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameters of the search and analysis process dynamically, adjusting the depth of analysis, data source priorities, and processing intensity based on the specific detection context. This allows the system to consume less energy for routine monitoring while allocating more resources when suspicious patterns are detected.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If companies monitor all data accessible online to detect data leaks, then detection completeness is improved, but search complexity increases

Engineering Contradiction:
Improvedetection completenessVSAvoidsearch complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the online data landscape into distinct categories (search engine results, social media platforms, dark web markets, pastebin sites) and applies tailored processing methods to each segment. This segmentation reduces search complexity by treating different data sources with appropriate specialized techniques rather than applying a single complex search strategy to all data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal automated detection system that can handle multiple data sources and analysis methods through a single integrated platform. This multi-functional system reduces overall complexity by providing a unified interface and coordination layer, even though it processes diverse data types through various specialized components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If data leaks are identified after significant time has passed (average 205 days in France), then detection thoroughness may be improved, but response time worsens significantly

Engineering Contradiction:
Improvedetection thoroughnessVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by continuously monitoring data sources and maintaining readiness to detect leaks before they cause significant damage. The system performs ongoing surveillance and preliminary analysis of potentially exposed data, enabling early detection and rapid response rather than waiting for traditional lengthy investigation processes.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent incorporates feedback mechanisms where detection results are continuously analyzed and used to refine future search strategies. The system learns from detected patterns and adjusts its monitoring priorities, enabling faster detection of similar future leaks while maintaining thoroughness through adaptive optimization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12306968B2Process for determining a degree of data exposure
Publication Date: 2025.05.20 CYBELANGEL
  • US12306968B2 patent drawing
  • US12306968B2 patent drawing
  • US12306968B2 patent drawing

AI summary

A process for determining a degree of data exposure based on receiving entries associated with documents, where each of these entries includes an access path and information about a server hosting the documents. The process may include generating subsets from the set of entries, and, for at least one of the subsets, determining at least one of a first and a second score. The first score may be determined by generating a value as a function of the access paths, and using a machine-learning algorithm to determine the first score based on this value. The second score may be determined by receiving content associated with each entry, generating a value as a function of the associated content, and determining, with a machine-learning algorithm, the second score based on the value. A degree of exposure of the data present on the associated server may be determined from one or both scores.