Anonymization of audio files by automatically replacing sensitive data with neutral audio content
An iterative audio processing method using neural networks and neutral audio replacement addresses the inefficiencies of current audio file data removal, ensuring secure and cost-effective removal of sensitive data across languages and accents.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-16
- Publication Date
- 2026-03-20
AI Technical Summary
Current methods for removing sensitive data from audio files, such as those used in call centers, are costly, environmentally inefficient, and not feasible for all languages due to limitations in automatic speech recognition technology, particularly for languages without transcription mechanisms and varying speech patterns.
An iterative process that detects audio signals representative of sensitive data and replaces them with neutral audio content, using a convolutional neural network to analyze MFCCs and replace segments with duration-specific neutral audio, optionally including voice distortion.
Effectively removes sensitive data from audio files across languages and accents, reducing computational and financial costs while maintaining data protection, allowing for secure storage and analysis without generating risks.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Anonymization of audio files by automatic replacement of sensitive data with neutral audio content FIELD OF INVENTION
[0001] The present invention relates to the removal of sensitive data within an audio file. It is particularly applicable to the field of call centers and to conversations recorded by these call centers.
[0002] Current digital technologies make it easy to produce audio files containing conversations. These audio files may need to be stored permanently for various reasons.
[0003] However, these audio files may contain sensitive data. Typically, when a conversation concerns the commercial domain, for example the relations between a company and its customers, these conversations contain personal data: names, addresses, telephone numbers, etc.
[0004] This data is considered sensitive because it can be used for malicious behavior (cyberbullying, fraud...).
[0005] Various regulations govern the uses of this sensitive information, and in particular the conditions for its management and storage. One example is the General Data Protection Regulation (GDPR) issued by the European Parliament (Regulation (EU) 2016 / 679).
[0006] It can be considered that there are millions of audio files containing such sensitive data. These numerous files can constitute just as many potential targets for malicious actors, and this large number makes it practically difficult to ensure technical protection for each of these files, distributed across various technology platforms throughout the world.
[0007] However, the detection and protection of this sensitive data in audio files remains a complex technical challenge.
[0008] One possibility is to transcribe the conversations contained in the audio files into a text document. Text analysis techniques can be used to detect sensitive data and remove it. This modified text file can be stored in place of the original audio file, or used to generate a new audio file by speech synthesis.
[0009] This mechanism does indeed make it possible to delete sensitive data, but it relies on an audio-to-text transcription technology (or ASR for "Automatic Speech Recognition") which is not always technically feasible. Indeed, there are several thousand languages worldwide, while the most popular transcription mechanism currently available, Google's, only allows transcriptions from approximately 200 languages. Therefore, there are languages for which no automatic transcription mechanism (ASR) is available.
[0010] Furthermore, the performance of these transcription mechanisms is highly sensitive to the speech patterns of the participants in the conversation, and in particular to accents. Even within the same language, there can be very strong dialectal or accentual differences. In some cases, the text transcriptions may be unusable.
[0011] Another major drawback is the cost of such a process. Indeed, the use of an automatic transcription platform generates a significant financial cost as well as a prohibitive environmental cost due to the highly computational nature of this processing.
[0012] There is therefore a need to improve current state-of-the-art proposals with a mechanism for removing sensitive data from audio files that is less expensive than transcription methods (both financially and ecologically) and that can have high performance for all languages. Summary of the invention
[0013] The invention aims to avoid or improve the situation compared to the proposals of the prior art, and in particular, to avoid the use of automatic transcription of audio files into texts.
[0014] To this end, according to a first aspect, the present invention can be implemented by a method for processing an audio file to remove sensitive data, by a processing device comprising iterative steps on the entire audio file of: - detection within said audio file of an audio signal representative of a type of sensitive data, - replacement of a segment of said audio file corresponding to said audio signal with neutral audio content.
[0015] Thus, the audio file is modified by the removal of sensitive data. It can be stored, transmitted, etc., without generating any particular risks, especially with regard to the protection of personal data. The necessary processing is far less extensive than in prior art proposals and is notably compatible with the requirements of mechanisms that respect environmental constraints.
[0016] According to preferred embodiments, the invention comprises one or more of the following features which can be used separately or in partial combination with each other or in total combination with each other: - The segment of said audio file also corresponds to a duration associated with that type of sensitive data. This allows us to optimize the audio portion that is deleted and replaced to an estimated duration for sensitive data, based on its type. - Neutral audio content also depends on the type of sensitive data, which makes it possible to find the different types of sensitive data that were present in an audio file (without having access to their content). - The process includes a subsequent step of analyzing the audio file to detect the absence of a particular type of sensitive data, based on the neutral audio content. It is thus possible, for example, to detect a non-conformity in the conversation corresponding to that audio file. - Neutral audio content corresponds to a signal of fixed frequency. - the process also includes a voice distortion step, which allows for further improvement of the anonymization of the audio file if necessary.
[0017] Another aspect of the invention relates to a gateway, a call center processing device adapted to implement the steps of the process as previously defined.
[0018] Another aspect of the invention relates to a call center comprising a processing device as previously described, and a control device adapted to store a conversation between at least one user and an operator in a database in the form of said audio file.
[0019] Another aspect of the invention relates to a computer program suitable for implementation on a processing device, the program comprising code instructions which, when executed by a processor, carries out the steps of the process as previously defined.
[0020] Another aspect of the invention relates to a data carrier on which at least one series of program code instructions for the execution of a process as previously defined has been stored.
[0021] Other features and advantages of the invention will become apparent from the following description of a preferred embodiment of the invention, given by way of example and with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE FIGURES
[0022] Fig. 1 represents an example of implementation of the invention in a call center.
[0023] Figure [Fig.2] illustrates an illustrative flowchart of a process according to an embodiment of the invention.
[0024] DETAILED DESCRIPTION OF EMBODIMENT METHODS OF THE INVENTION
[0025] The proposed method can be applied to any context where audio files are generated or stored and are likely to contain sensitive data.
[0026] It can be applied to any system for managing communications between a company and its customers, to call centers, to tele- or videoconferencing management systems, etc.
[0027] A call center (also called a contact center, or "call center" in English) can be considered as a set of technical means enabling the management of audio or audio-video communications between external users and internal operators.
[0028] This type of system can be deployed internally within a company to manage customers, or it can be managed by a specialized company.
[0029] In the latter case, these call centers are shared by several companies, each company entering into a contract with the call center manager for a defined mission.
[0030] The term "call center" therefore has a certain degree of ambiguity, as it can refer both to a set of technical resources and to a company managing such a set of technical resources. In the following, we will consider a call center to be a software system, deployed on a hardware platform (computer or set of computers), designed to manage audio or audio / video communications between parties and to store audio files containing the conversations held during these communications.
[0031] Fig. 1 illustrates an implementation for such a call center, but the explanations and principles set forth are transposable to other contexts.
[0032] A call center 30 enables communication and the management of this communication between users 20 and teleoperators IA, IB, IC from a set of teleoperators 10.
[0033] The communication setup includes the management of an incoming call from a user 20 in order to assign an available teleoperator and establish the communication between the latter and the calling user.
[0034] According to one embodiment, a control device 31 can be adapted to store at least some of the conversations between users (customers...) 20 and operators 10 in a database 32 in the form of an audio file 33.
[0035] In the case of an audio / video session, only the audio track can be stored. In the case where an audio / video capture is performed, only the audio track 33 of the audio / video recording is subsequently of interest.
[0036] The call center 30 (or any other related system) also includes a processing device 34 which is adapted to process the audio files 33 contained in the database 32.
[0037] This treatment is illustrated by the flowchart shown in [Fig. 2]. This flowchart represents one proposed embodiment of this treatment, or process. It is indeed possible to subdivide this process according to different levels of detail and to make choices regarding the definition of each step of the process.
[0038] According to this embodiment, this process is iterative, in the sense that it consists of sequentially processing the data of the audio file 33 until the entire file has been processed. Typically, the processing begins at the beginning of the audio file and ends at the end of the audio file.
[0039] In a step S1, the processing device 34 searches the audio file for an audio signal representative of a sensitive data type. If such an audio signal is detected, the processing device 34 implements a step S2 of replacing a segment of the audio file corresponding to this audio signal and a duration associated with the sensitive data type, with neutral audio content.
[0040] Sensitive data can indeed be categorized into different types (or classes). For example, personal data may include a first type corresponding to people's names, a second type corresponding to telephone numbers, a third type corresponding to postal addresses, a fourth type corresponding to email addresses, etc.
[0041] It is understood that other types of personal data, or more generally sensitive data (passwords, bank details, etc.) that may be communicated in a conversation between a customer and a company (in particular) can thus be defined.
[0042] In a conversation, this personal data is very generally announced by a specific audio signal. This audio signal may correspond to an utterance, in the speaker's language, of the type of sensitive data that will follow. For example, these audio signals may correspond to expressions such as "my number is," "my address is," "I live at," etc. In other words, when these expressions are detected in a conversation, it can be determined that the following sounds correspond to sensitive data.
[0043] The detection of such an audio signal can therefore be based on the detection of keywords (or words), “keyword spotting” in English.
[0044] Keyword detection is a problem that was historically first defined in the context of speech processing. In speech processing, keyword detection consists of identifying keywords in utterances.
[0045] The problem of keyword detection appears to have been first presented in the article by Rohlicek, J.; Russell, W.; Roukos, S.; Gish, H. (1989). "Continuous hidden Markov modeling for speaker-independent word spotting". Proceedings of the 14th IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). 1:627-630.
[0046] Much work has been carried out in this field, resulting in publications, products, patents or patent applications, among which examples can be cited: US11361763B1, US11514901B2...
[0047] Among the algorithms used for this task, we can mention: - Sliding window methods and waste model - K-best hypotheses - Iterative Viterbi decoding - Convolutional neural networks, as described in Sainath, Tara N; Parada, Carolina (2015). "Convolutional neural networks for small-footprint keyword spotting". Sixteenth Annual Conference of the International Speech Communication Association. arXiv: 1711.00333 - Detection of keywords with a reduced footprint based on transformers, etc.
[0048] According to a particular embodiment, the data from the audio file 33 are converted into a spectrogram or into MFCC (for "Mel-Frequency Cepstral Coefficients"). This conversion is generally carried out on the basis of segmenting the audio file into frames of, for example, 20 to 40 ms.
[0049] MFCCs are a representation of audio features that capture the most relevant aspects of the speech signal for tasks such as speech recognition. They are derived from the frequency spectrum of an audio signal and are based on human frequency perception, particularly suited to speech sounds.
[0050] These MFCCs are obtained by frequency analysis on the data sequences of the audio file (corresponding to the segmentation): - Apply the FFT to each sequence to obtain a power spectrum. - Filter the frequencies using the Mel scale. - Apply the logarithm to the results. - Perform a direct cosine transform to obtain the MFCC coefficients
[0051] A convolutional neural network (CNN) can then be used on this spectrogram or MFCC, which can be considered as a two-dimensional image, thus taking advantage of the good suitability of this type of technology for the extraction of features from image data.
[0052] Thus, this keyword detection is based on characteristic elements of the audio signals to be detected and of time windows of the audio file. These elements are characteristic of the words to be detected.
[0053] This mechanism makes it easy to integrate new languages, new accents, new pronunciations by capturing the corresponding vocalizations of these keywords and to perform new learning to update the models that allow their detection in the audio files.
[0054] Step S2 consists of replacing a segment of the audio file corresponding to the audio signal that was detected in step SI and to a duration associated with the type of sensitive data, with neutral audio content.
[0055] This segment may begin immediately, or substantially immediately, after the audio signal representing a type of sensitive data. This is explained by the considerations mentioned above, according to which sensitive data generally follows a voice announcement (i.e., this audio signal), for semantic reasons.
[0056] This segment may have a duration that depends on the type of sensitive data.
[0057] Thus, it can be predicted that sensitive data such as a "postal address" will take longer than sensitive data such as a "telephone number". Therefore, an estimated duration can be associated with each type of sensitive data. This duration can be estimated beforehand, based on a sample of representative audio files, and may depend on the language in question.
[0058] These durations may be average durations or maximum durations, or average durations with an additional margin, etc., depending on specific embodiments. It should be noted that replacing a substantial portion of sensitive data may be sufficient to render it unusable by a malicious third party. Therefore, the durations may be reduced to avoid deleting non-sensitive content from the audio file.
[0059] The processing device may have access to a lookup table associating types of sensitive data with durations.
[0060] According to one embodiment, the neutral audio content also depends on the type of sensitive data. The lookup table can then associate each type of data with a duration and neutral audio content.
[0061] By way of example, this neutral audio content may be a fixed frequency signal. This fixed frequency may, according to one embodiment, be dependent on the type of sensitive data.
[0062] According to other embodiments, frequency and / or volume modulations may be provided as neutral audio content, capable of forming short, recognizable melodies. Also, vocalization of the sensitive data type can be carried out by means of a prior recording by a person or by speech synthesis.
[0063] The following table illustrates one embodiment of such a lookup table.
[0064] [Tables 1] Type of sensitive data Frequency of neutral audio content Duration Address 400 Hz 20 seconds Date of birth 500 Hz 8 seconds Phone number 600 Hz 5 seconds
[0065] Preferably, these two steps SI, S2 are iterated until the entire audio file has been scanned to ensure that all sensitive data has been detected and replaced with neutral audio content. A step S3 verifies this stopping criterion to either initiate a new SI detection or consider that the entire file has been processed and stop the process.
[0066] Replacing the segment of the audio file containing sensitive data with neutral audio content naturally removes this sensitive data. The modified audio file 33 is thus purged of all sensitive data. It can be stored without generating risks related to the protection of personal data and can be used for malicious purposes by a cyber attacker.
[0067] Making the neutral audio content dependent on the type of sensitive data allows some of the semantic information initially present in the audio file to be retained. It remains possible to determine which types of sensitive data are present and where in the audio file. This information can be useful during the analysis of the audio file, or a set of audio files, for statistical purposes.
[0068] This analysis can be based on neutral audio content: insofar as this depends on the type of sensitive data, it is sufficient to determine the segments containing fixed frequencies (or any other arrangement provided) and, based on this frequency, determine the type of sensitive data.
[0069] In particular, the method may include an S4 step of analyzing the audio file in order to detect the absence of a particular type of sensitive data.
[0070] In the case of a company's relationship with its customers (or prospects), a process may be imposed on operators for conducting conversations. This process may include, in particular, requesting certain information from the Client: for example, name and / or address and / or contact information (telephone number, email address, etc.). It is thus possible to scan the audio file and, by detecting the various neutral audio contents, determine which types of sensitive data are present and compare them to a list of required sensitive data. This allows us to detect the absence of a type of sensitive data that should have been present.
[0071] This absence can be reported as an anomaly, and be subject to statistics or corrective actions.
[0072] In addition, a set of audio files can be analyzed to perform various statistics on the basis of neutral audio content (occurrences of different types of sensitive data, location at the beginning / middle / end of the conversation, etc.).
[0073] These processing operations are based on anonymized audio files (expurgated of sensitive data), and can therefore be carried out by various processing platforms, including open ones, without requiring additional security mechanisms.
[0074] According to a particular embodiment, the process may further comprise a voice distortion step. The audio file 33 may undergo digital processing to render the speakers' voices unrecognizable in the conversation. Various techniques may be used, such as vocoders, voice modulators, the application of different filters, etc.
[0075] Vocoder technology is described in particular in the article by Panariello, M., Todisco, M., & Evans, N., “Vocoder drift in x-vector-based speaker anonymization”, 2023, arXiv preprint arXiv:2306.02892.
[0076] This embodiment allows for additional anonymization of audio files 33.
[0077] Of course, the present invention is not limited to the examples and embodiment described and illustrated. In particular, it is susceptible to numerous variations accessible to those skilled in the art, some of which have been described previously or simply mentioned.
Claims
Demands
1. A method for processing an audio file (33) for the removal of sensitive data, by a processing device (34), comprising iterative steps over the whole of said audio file of: - detection (S1) within said audio file of an audio signal representative of a type of sensitive data, - replacement (S2) of a segment of said audio file corresponding to said audio signal by neutral audio content.
2. A method according to the preceding claim, wherein said segment of said audio file also corresponds to a duration associated with said type of sensitive data.
3. A method according to any one of claims 1 or 2, wherein said neutral audio content also depends on said type of sensitive data.
4. A method according to the preceding claim, comprising a subsequent analysis step (S4) of said audio file in order to detect the absence of a particular type of sensitive data, from said neutral audio content.
5. A method according to any one of the preceding claims, wherein said neutral audio content corresponds to a fixed frequency signal.
6. A method according to any one of the preceding claims, further comprising a voice distortion step.
7. A computer program capable of being implemented on a processing device (34), the program comprising code instructions which, when executed by a processor, carries out the steps of the process defined in claims 1 to 6.
8. Data carrier on which at least one set of program code instructions for the execution of a method according to any one of claims 1 to 6 has been stored.
9. A processing device (34) for a call center (30) adapted to carry out the steps of the process according to any one of claims 1 QA
10. 1 d 0. Call center (30) comprising a processing device (34) according to the preceding claim, and a control device (31) adapted to store a conversation between at least one user (20) and an operator (IA, IB, IC) in a database (32) in the form of said audio file (33).
Citation Information
Patent Citations
Detecting system-directed speech
US11361763B1
Anchored speech detection and speech recognition
US11514901B2
Method of retaining a media stream without its private audio content
EP2169669A1
Filtering sensitive information
US10529336B1
Cybersecurity for sensitive-information utterances in interactive voice sessions
US20220199093A1