AI Speech De-identification via Pitch Modulation and Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for speech recognition technology in IoT devices raises concerns about personal information protection, as large amounts of speech data collected for improved accuracy often contain sensitive personal information, necessitating de-identification to ensure privacy.
Innovation Solution
An artificial intelligence device that de-identifies speech signals by modulating pitch in the frequency region, dividing signals into voiceless and voiced components, and adjusting signal length in time units, allowing for speech recognition without revealing personal information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech data is collected to accumulate a large amount of data for accurate speech recognition, then speech recognition accuracy is improved, but personal information protection problems occur
Solution Approach 1:
The patent extracts and removes personal information from speech data through de-identification processing. The system separates identifiable personal characteristics from the speech signal while retaining the linguistic content, allowing accurate speech recognition without compromising personal information protection.
Solution Approach 2:
The patent introduces de-identification processing as an intermediary step between speech data collection and speech recognition. This intermediary process transforms the speech data into an anonymized form that preserves recognition accuracy while eliminating personal information risks.
2Measurement precision
If speech data classified as personal information is collected, then speech recognition accuracy is improved, but prior consent procedures are required causing troublesomeness
Solution Approach 1:
The patent performs de-identification processing preliminarily on speech data before it is used for training or recognition. By removing personal information in advance, the system eliminates the need for complex consent procedures while maintaining data utility for speech recognition.
Solution Approach 2:
The patent converts the potentially harmful personal information in speech data into a beneficial anonymized form. The de-identification process transforms data that would require consent procedures into safe, usable training data that improves speech recognition without legal or procedural burdens.
3Object-affected harmful factors
If de-identification is performed on speech signals, then personal information protection is improved, but speech recognition accuracy may deteriorate
Solution Approach 1:
The patent applies de-identification selectively to specific features of the speech signal that carry personal information, while preserving the linguistic and semantic content. By targeting only the identifiable characteristics for modification, the system maintains speech recognition accuracy while achieving personal information protection.
Solution Approach 2:
The patent modifies specific parameters of the speech signal (such as pitch, timbre, or spectral characteristics) that are responsible for personal identification, while keeping the phonetic and linguistic parameters intact. This selective parameter transformation protects personal information without degrading recognition performance.
Data Source
AI summary
An artificial intelligence device for learning a de-identified speech signal includes a memory configured to store a speech recognition model, a microphone configured to acquire an original speech signal, and a processor configured to perform de-identification with respect to the acquired original speech signal and perform speech recognition with respect to the de-identified speech signal through the speech recognition model.


