User Data Processing for Privacy Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Natural language processing systems face challenges in identifying and handling user-identifiable data within spoken language user inputs, which can contain sensitive information, and existing methods lack effective mechanisms to detect and manage such data appropriately.
Innovation Solution
A system is configured to detect user-identifiable data in natural language user inputs using multiple components that employ semantic similarity, entity recognition, and machine learning models to identify and flag sensitive information, allowing for appropriate data handling and deletion post-processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If natural language processing systems process spoken language inputs, then user interaction capability is improved, but user privacy and data security deteriorate due to potential exposure of sensitive information
Solution Approach 1:
The system performs preliminary detection of user-identifiable data in spoken language inputs before the data is processed or stored. By detecting sensitive information upfront using multiple components (semantic similarity, entity recognition, machine learning models), the system can take preventive actions such as redacting or deleting the data, thereby protecting user privacy while maintaining natural language processing functionality
Solution Approach 2:
The system extracts and isolates user-identifiable data from the spoken language input using specialized detection components. By separating the sensitive information from the rest of the input data, the system can handle the non-sensitive portions for processing while securely managing or removing the extracted sensitive data, thus resolving the contradiction between processing capability and privacy protection
2Measurement precision
If multiple detection components are used to identify user-identifiable data, then detection accuracy is improved, but system complexity increases
Solution Approach 1:
The detection system is segmented into multiple specialized components, each responsible for a specific aspect of user-identifiable data detection: semantic similarity analysis, entity recognition, and machine learning-based detection. This segmentation allows each component to focus on its specialized function, improving overall detection accuracy while organizing complexity into manageable, modular units that can be independently optimized and maintained
Data Source
AI summary
Techniques for determining attributable data in a natural language user input that can be used to identify a specific user are described. A system may use various data signals determined using different components. The system may process the various signals to make a final determination on whether the input includes attributable data. The system may use a first component to detect user-identifiable data in the input. The system may use a second component to determine whether the input is potentially attributable to a particular user.


