NLP Consent Mechanism for Privacy-Aware Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition systems lack effective mechanisms for protecting user privacy by requesting consent before using sensitive natural language inputs for training data generation, potentially compromising personal data.

Innovation Solution

A system that includes a feedback skill to request consent from users before generating training data, allowing users to opt-in for improving speech processing accuracy while ensuring privacy protection by distinguishing between public and private data interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems use natural language inputs for training data generation, then speech processing accuracy is improved, but user privacy is compromised

Engineering Contradiction:
Improvespeech processing accuracyVSAvoiduser privacy compromise
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary classification of natural language inputs into public and private data categories before training data generation. This advance categorization ensures that only appropriate data is used for training, preventing privacy compromise while maintaining training quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary classification mechanism that acts as a mediator between user inputs and training data generation. This intermediary layer evaluates and categorizes inputs, ensuring privacy protection is implemented systematically in the data pipeline.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If consent mechanisms are implemented for training data usage, then user privacy is protected, but system complexity increases

Engineering Contradiction:
Improveuser privacy protectionVSAvoidsystem complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically classifying natural language inputs into public and private categories without requiring external consent mechanisms. This automation maintains privacy protection while avoiding the complexity of manual consent management.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The classification system provides feedback about data categorization decisions, enabling the system to adaptively manage training data selection. This feedback mechanism simplifies privacy protection by creating an automated decision-loop rather than requiring complex external consent protocols.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If all natural language inputs are used for training, then training data quantity increases, but private data protection is weakened

Engineering Contradiction:
Improvetraining data quantityVSAvoidprivate data protection
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by treating different types of natural language inputs differently based on their sensitivity. Public data is freely used for training while private data is protected, creating localized quality variations in data usage that maximize training quantity without compromising privacy.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system segments the natural language input stream into distinct public and private data categories. This segmentation enables selective usage where public data contributes to training quantity while private data remains protected, resolving the contradiction between data quantity and privacy protection.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11195534B1Permissioning for natural language processing systems
Publication Date: 2021.12.07 AMAZON TECH INC
  • US11195534B1 patent drawing
  • US11195534B1 patent drawing
  • US11195534B1 patent drawing

AI summary

Techniques for improving privacy protection by requesting consent to review and analyze a natural language interaction. A natural language processing (NLP) system may process a user query, provide a system response, and then request feedback indicating whether the system response was responsive to the user query. To improve an accuracy of the NLP system, the NLP system may request consent to review interaction data and generate training data in order to train a model. If consent is granted, the NLP system may store confirmation of the consent and share the interaction data for training. The NLP system may process the user query using a first skill but request the feedback using a separate feedback skill. Based on the feedback, the feedback skill may pass the interaction back to the first skill differently and/or cause the first skill to perform one or more different actions.