De-identifying Personal Information in Chatbot Conversations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Chatbots often require personal user information during customer service interactions, which can lead to privacy invasions due to the presence of personal information in text data, and existing technologies lack effective methods to de-identify such information for secure data utilization.
Innovation Solution
An apparatus and method that includes a sentence detection unit to identify personal information in conversations between user devices and chatbots, a personal information identification model to detect de-identification target sentences, and a search unit to de-identify specific tokens from conversational data, generating training data by removing or replacing personal information to protect user privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If personal information is collected during chatbot customer service interactions, then service quality and user assistance are improved, but user privacy is compromised and data security risks increase
Solution Approach 1:
The patent extracts personal information from conversational text data using a personal information identification model. The model detects and separates personal information tokens (such as names, phone numbers, addresses) from the original conversation data, allowing the useful conversational content to be retained while removing the harmful personal information elements.
Solution Approach 2:
The patent introduces a de-identification processing system as an intermediary between data collection and data storage/usage. This intermediary process includes a personal information identification model that acts as a mediator to filter and redact personal information before the data is stored or used for training, thus protecting privacy while maintaining service quality.
2Measurement precision
If personal information is stored for training data purposes, then model accuracy and service performance are improved, but data security and privacy protection deteriorate
Solution Approach 1:
The patent applies preliminary de-identification action to training data before it is stored or used for model training. The personal information identification model pre-processes the conversational data to remove personal information tokens, ensuring that the training data is cleaned in advance before entering the storage or training pipeline, thus preventing privacy leaks while maintaining model accuracy.
Solution Approach 2:
The patent extracts personal information from training data using automated detection and redaction processes. The personal information identification model identifies and removes personal information tokens from conversational text, allowing the training data to retain its useful linguistic patterns and contextual information while eliminating security risks associated with personal data storage.
3Adaptability or versatility
If text data containing personal information is used for training, then conversational service functionality is improved, but privacy protection and compliance with data protection regulations worsen
Solution Approach 1:
The patent introduces a de-identification processing system as an intermediary layer between raw conversational data and training data. This intermediary process uses a personal information identification model to detect and redact personal information tokens, allowing the system to comply with privacy regulations while preserving the functional and linguistic qualities needed for conversational service training.
Solution Approach 2:
The patent extracts personal information from conversational text data through automated detection and removal processes. The personal information identification model identifies personal information tokens (such as names, contact information, addresses) and removes them from the training data, enabling the system to maintain service functionality while avoiding privacy violations and regulatory compliance issues.
Data Source
AI summary
An apparatus for generating de-identified training data for conversational service includes a sentence detection unit configured to detect at least one sentence including personal information in a conversation between a user device and a chatbot; a de-identification target sentence detection unit configured to input conversational data including the at least one sentence into a personal information identification model and detect a de-identification target sentence through the personal information identification model; a search unit configured to search a predefined de-identification target token from the conversational data when a de-identification target sentence is detected from the conversational data; and a training data generation unit configured to generate training data on the conversational data by de-identifying text corresponding to the searched de-identification target token.


