Unstructured Data Redaction via Identity Graph Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computing tools for managing sensitive data face performance degradation due to processing extraneous unstructured data in requests, leading to resource wastage and inaccurate responses, especially when requests are received as unstructured communications like emails or text messages.
Innovation Solution
A method and system that analyze unstructured data within requests to categorize and map relevant data to personal data using an identity graph, redacting irrelevant data and processing the request with only relevant information, employing techniques like natural language processing and machine learning for efficient data matching and retrieval across multiple data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If computing tools process all unstructured data in requests, then complete information is available for analysis, but system performance degrades due to resource wastage
Solution Approach 1:
The patent extracts and removes extraneous unstructured data from requests before processing. The system identifies and separates relevant personal data from irrelevant unstructured content, processing only the extracted relevant portions. This extraction principle resolves the contradiction by eliminating harmful data elements that waste resources while preserving the necessary information for accurate request handling.
Solution Approach 2:
The patent applies different processing quality levels to different portions of the request data. Relevant personal data receives thorough, accurate processing while extraneous unstructured data is either discarded or processed with minimal attention. This local quality differentiation maintains processing accuracy for critical elements while reducing overall resource consumption.
2Loss of information
If computing tools process extraneous unstructured data, then no information is lost, but resource usage increases leading to performance degradation
Solution Approach 1:
The system extracts only the necessary personal data from unstructured requests, discarding extraneous information. This extraction approach prevents information loss of relevant data while eliminating the processing of energy-wasting extraneous content, thus resolving the contradiction between information completeness and resource conservation.
Solution Approach 2:
The patent discards extraneous unstructured data that does not contribute to request processing while recovering and retaining only the essential personal data elements. This selective discarding and recovering process eliminates unnecessary resource consumption without compromising the completeness of relevant information processing.
3Measurement precision
If computing tools process all data in unstructured requests, then comprehensive analysis is possible, but response accuracy decreases due to extraneous information
Solution Approach 1:
The patent extracts relevant personal data from unstructured requests with high precision, separating it from extraneous information. This extraction process improves measurement precision by focusing processing on identified relevant elements while maintaining completeness of necessary information through targeted recovery of personal data elements.
Solution Approach 2:
The system applies high-quality, precise processing only to relevant personal data portions while using minimal or no processing on extraneous unstructured data. This local quality approach ensures accurate identification and handling of critical information without being degraded by irrelevant content, resolving the contradiction between precision and information completeness.
Data Source
AI summary
System and methods are disclosed for redacting analyzing unstructured data in a request for data associated with a data subject to determine whether the unstructured data is relevant to the request. The relevancy of pieces of the unstructured data may be determined by determining a categorization for each such piece of unstructured data and comparing them to known personal data associated with the data subject having the same categorization. Pieces of the unstructured data that do not match known personal data having the same categorization are redacted from the request before the request is processed.


