Targeted Data Discovery Through Metadata and Graph Field Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As data volumes increase across multiple data sources, identifying and managing targeted data becomes significantly difficult due to varying handling processes and mechanisms, making it challenging to locate and retrieve specific data portions efficiently.
Innovation Solution
A computational framework that analyzes data sources to identify targeted data types, records metadata for location and identification methods, and uses machine learning to match and query data fields, enabling efficient retrieval and management of targeted data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored across multiple data sources with varying handling processes, then data coverage and completeness are improved, but data identification and retrieval complexity increases
Solution Approach 1:
The patent segments the complex data identification task into distinct components: metadata generation for each data source, graph data structure construction for representing data relationships, and targeted querying mechanisms. This segmentation allows each component to be optimized independently while maintaining overall system effectiveness across multiple data sources.
Solution Approach 2:
The patent introduces metadata as an intermediary layer between the querying system and the actual data sources. This metadata layer abstracts the heterogeneity of different data sources, providing a unified interface for data identification and retrieval without requiring direct complexity management of each underlying data source.
2Adaptability or versatility
If the number of data sources increases to handle more data types, then data versatility is improved, but the difficulty of locating specific data portions increases
Solution Approach 1:
The patent performs preliminary actions by pre-generating metadata for each data source and pre-constructing graph data structures that map data relationships. This preliminary organization enables efficient data location during querying operations, reducing the difficulty of finding specific data portions across diverse data sources.
Solution Approach 2:
The patent transforms the data identification problem by changing parameters from direct data searching to metadata-based querying. By utilizing graph data structures with defined relationships and properties, the system changes the search space from raw data to structured metadata, making data location more efficient across diverse sources.
3Device complexity
If manual data identification methods are used across multiple systems, then system simplicity is maintained, but productivity and efficiency decrease
Solution Approach 1:
The patent implements self-service mechanisms where the system automatically generates metadata, constructs graph data structures, and performs data identification without manual intervention. This automation maintains relative system simplicity while dramatically improving productivity, as the system serves itself in organizing and locating data across multiple sources.
Data Source
AI summary
Various embodiments provide methods, apparatus, systems, computing devices, computing entities, and/or the like for identifying targeted data for a data subject across a plurality of data objects in a data source. In accordance with one embodiment, a method is provided comprising: receiving a request to identify targeted data for a data subject; identifying a first data object using metadata for a data source that identifies the first data object as associated with a first targeted data type for a data portion from the request; identifying a first data field from a graph data structure of the first data object that identifies the first data field as used for storing data having the first targeted data type; and querying the first data object based on the first data field and the data for the first targeted data type to identify a first targeted data portion for the data subject.


