MASKING THE SPEAKER'S LANGUAGE

DE602023015553T2Active Publication Date: 2026-04-22MUSICIENS ARTISTES INTERPRÈTES ASSOCIÉS M A I A
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
MUSICIENS ARTISTES INTERPRÈTES ASSOCIÉS M A I A
Filing Date
2023-05-31
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing voice masking techniques are easily reversible and can compromise speaker privacy, particularly in applications like investigative journalism and voice input services, due to advancements in speech recognition technology.

Method used

A method that alters the pitch and timbre of a speaker's voice by applying opposing frequency shifts, making it irreversible and resistant to reverse engineering, using digital signal processing techniques to ensure the masked voice remains intelligible and unique.

Benefits of technology

The method effectively protects speaker identity and privacy by creating a masked voice that is unintelligible to speaker recognition systems and resistant to reversal, while maintaining audio quality.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

Domaine technique

[0001] The present invention relates to masking the voice of a speaker, in particular to protect the identity of the speaker by restricting the possibility of identifying him by analyzing an original recording of his voice.

[0002] It finds applications, in particular, in audio-phonic or audio-visual editing and / or mixing systems in which it can be implemented by audio-phonic processing software. Arrière-plan technologique

[0003] In certain areas of the audiovisual industry, for example, it is useful to be able to broadcast programs (audio, video, and / or multimedia content) while concealing the speaker's identity in order to protect them from any potentially harmful consequences of such broadcasting. For instance, in investigative journalism, it is common practice to anonymize the recording of a witness interview that could be used against their interests, either by the perpetrators of the crimes they are reporting, or by individuals or a competent judicial body if the witness admits to having violated any regulations.

[0004] Techniques for anonymizing a speaker's voice in an audio signal—that is, making it difficult to identify the speaker through audio analysis—have been known for a long time. The oldest and most widespread technique simply transforms the speaker's voice by shifting the harmonics. This shift can be made either towards higher frequencies, i.e., towards the treble, or towards lower frequencies, i.e., towards the bass. Referring to characters from well-known fictional audiovisual programs, the transformed voice is sometimes said to resemble the voice of "Mickey Mouse"™ or the voice of "Darth Vader"™, which are generated using such techniques from the voice of a real person. However, the resulting voice transformation is easily reversible with technical means now accessible to many, if not everyone.

[0005] Furthermore, speech recognition software, or more specifically speaker recognition software (" speaker recognition Voice signatures (in English) are sometimes used by the police to identify individuals making threatening phone calls or anonymous calls. However, some software of this type now exists that can identify a speaker with such a high degree of reliability that it can lead to a conviction. While such use may seem commendable from a public perspective, malicious use of such software can have far less desirable consequences for individuals, potentially causing irreparable harm, such as violations of privacy. Therefore, audio-phonic processing techniques can be used to protect speakers whose voice recordings are likely to be disseminated or intercepted on communication networks.

[0006] Finally, another case where speaker protection is desirable is that of voice input applications using speech recognition (“ speech recognition (in English) to allow users to access services. Speech recognition is used to recognize what is said. Therefore, it allows speech to be transformed into text, which is why it is also known as speech-to-text conversion.

[0007] The article "Voice Mask: Anonymize and Sanitize Voice Input on Mobile Devices," published in the scientific journal COMPUTER SCIENCE, CRYPTOGRAPHY AND SECURITY, Cornell University, US, November 30, 2017, pages 1-10, by Jianwei Qian et al., reveals that with hands-free communication, voice input has largely replaced the use of traditional keyboards (e.g., the virtual keyboards Google®, Microsoft®, Sougou™, and iFlytek™). These techniques are used daily by many users for voice search (with, for example, applications like Microsoft Bing®, Google Search®) and AI-based personal assistants (e.g., Apple's Siri®, and Amazon Echo®), across a wide range of mobile devices.In these applications, due to the limited resources on mobile devices, the speech recognition operation is generally offloaded to a cloud computing server (“ . cloud computing " for greater accuracy and increased efficiency. It follows that privacy (" privacy User privacy can be compromised. Indeed, even though in these applications only the speech content needs to be recognized by speech recognition, it has become easy to perform speaker recognition to identify regular mobile users by their voice, using machine learning techniques that leverage recurring usage. This allows for the analysis of sensitive content in their input via speech recognition, and then the creation of user profiles based on this content. The aim is to provide biased responses to users' queries and / or targeted marketing offers. The authors of the article propose a voice neutralization application that ensures good protection of user identity and private speech content, at the cost of minimal degradation in speech recognition quality.It adopts a voice conversion mechanism that resists multiple attacks.

[0008] The article "Speaker Anonymization for Personal Information Protection Using Voice Conversion Techniques," published in Proceedings of Access 2020, Digital Object Identifier, Vol. 8, 2020, pages 198637-198645, IEEE, US, by In-Chul Yoo et al., discloses voice conversion techniques for anonymizing a speaker. The goal is to preserve the linguistic content of the speech while removing biometric data from the original speaker's voice. The proposed method modifies conventional speaker identity vectors into anonymized speaker identity vectors using various techniques.

[0009] The article entitled “Speaker Anonymization Using X-vector and Neural Waveform Models,” published in the Proceedings of the 10th ISCA Speech Synthesis Workshop, September 20–22, 2019, Vienna, Austria, pages 155–160, by Fuming Fan et al., proposes a speaker anonymization approach to conceal the speaker's identity while maintaining high-quality anonymous speech. This approach is based on extracting the speaker's linguistic and identity characteristics from an utterance and then using them with acoustic neural and waveform models to synthesize anonymous speech. The speaker's original identity, in the form of timbre, is removed and replaced with that of an anonymous pseudo-identity. The approach leverages the speaker's peak representations as X-vectors. These representations are used to derive anonymous speech.These are used to derive pseudo-identities of anonymous speakers by combining several vectors X of random speakers.

[0010] The authors of the article entitled "Exploring the Importance of FO Trajectories for Speaker Anonymization using x-vectors and Neural Waveform Models," International Audio Laboratories, Erlangen, 2021, Workshop on Machine Learning in Speech and Language Processing (MLSLP), September 6, 2021, ISCA, Germany, pages 1-6, EU Gaznepoglu et al., considering the presence of personal information in the different components of the fundamental frequency F0 of a speaker's voice and the availability of various approaches to modify the F0 component, propose an exploration of their potential in the context of voice anonymization. They suggest that decomposing the F0 component, modifying the speaker-related characteristics, possibly introducing noise during the process, and then resynthesizing could increase anonymization performance and / or improve intelligibility.It is mentioned that the approaches proposed so far, such as shifting and scaling, all depend on the identity of the person to be protected.

[0011] The article "Speaker anonymization using the McAdams coefficients" in the journal COMPUTER SCIENCE, AUDIO AND SPEECH PROCESSING, Cornell University, US, September 2021, pages 1-5, by Patino J et al., discusses the reversibility of anonymization. The authors present their work exploring the potential of well-known signal processing techniques as a solution to the anonymization problem, as opposed to more complex and demanding solutions that require training data. They suggest optimizing an original solution based on McAdams coefficients to modify the spectral envelope (i.e., the timbre) of speech signals. They sought to confirm that different values ​​of the McAdams coefficient α (alpha), which modify voice timbre, can produce different pseudo-voices for the same speaker.This results in a stochastic approach to anonymization in which the McAdams coefficient is sampled from a uniformly distributed range, i.e. α ∈ . U (αmin,αmax). However, in the proposed applications, the article merely teaches that the coefficient α can be changed randomly from one speaker to another, indicating that a malicious third party would then need to know the exact McAdams coefficient used to anonymize the speech of any particular speaker, in order to reverse the transformation.

[0012] US document 10 141 008 B1 proposes a method for intentionally modifying a voice by altering a recording of the voice, 2 parameters corresponding to pitch and timbre, in an ascending manner on one parameter and descending on the other parameter.

[0013] Therefore, there remains a need for a technique to mask a speaker's voice that cannot be easily circumvented. Résumé de l'invention

[0014] A first aspect of the proposed invention relates to a method according to claim 1 of masking the voice of a speaker to protect his / her identity and / or privacy by intentionally altering the pitch and timbre of the voice.

[0015] Thanks to this process, the voice is masked, meeting the speaker's privacy requirements, as it can easily be implemented in the very first piece of equipment in the audio acquisition and processing chain. At the same time, the process allows for a final output that remains intelligible; that is, it is neither a "Mickey Mouse" voice nor a "Darth Vader" voice, due to the two alterations applied to each audio segment. These alterations produce opposing changes in the frequency content. Specifically, one alteration applies a rising effect (towards higher pitches) and the other a falling effect (towards lower pitches), so that these two effects combine in terms of the frequency content of the audio segment in question.The resulting masked audio segment has a frequency content that remains overall closer in spectral dynamics to that of the original audio segment, despite the masking of the voice that is achieved.

[0016] Advantageously, the frequency alterations are limited, with one always being upwards while the other is always downwards. Therefore, the software or device used to implement the solution cannot itself perform the reverse operation.

[0017] According to the method, two components of the audio signal spectrum are simultaneously altered compared to the original recording of the speaker's voice. This is a key factor contributing to the irreversibility of the process, as a malicious third party wishing to revert to the original voice would have to manipulate these two voice characteristics in combination, making their task more difficult than masking by pitch shifting alone.

[0018] Another advantage is that the variability of the alterations is not stationary. It varies over time. Thus, there can be several variations within a single second of processing.

[0019] Ultimately, the proposed implementation methods provide a voice masking that is irreversible in audio, i.e. by inverse audio processing.

[0020] Furthermore, the voice masked by the proposed method is not analyzable by known speaker recognition techniques, and does not expose the speaker to commercial practices that infringe on their privacy by using voice recognition techniques, given that the masked voice of the same speaker is never masked twice in the same way.

[0021] In advantageous implementation modes, the temporal segmentation of the audio signal into a series of successive audio segments of determined duration can be achieved by time windowing independent of the content of the audio signal.

[0022] In advantageous implementation modes, the temporal segmentation of the audio signal can be configured so that the duration of an audio segment is equal to a fraction of a second, so that successive variations from one pair of segments to another of the parameters varying the first and second alterations occur several times per second.

[0023] In advantageous implementation modes, the alteration of the pitch of the audio signal corresponds to a variation of the fundamental frequency of any of the following values: ± 6.25%, ± 12.5%, ± 25%, ± 50% and ± 100%.

[0024] In advantageous implementations, the first and second alterations are interdependent, fluctuating jointly to satisfy a specific criterion regarding their respective effects on the frequency content of the audio segment's timbre and pitch. For example, this criterion might consist of maintaining a minimal difference between the respective effects of the two alterations, thus avoiding a temporary reversion to the original voice.

[0025] A second aspect of the invention relates to a computer program comprising instructions which, when the computer program is loaded into the memory of a computer and executed by a processor of that computer, are adapted to implement all the steps of the process according to the first aspect of the invention above.

[0026] The computer program for implementing the process can be recorded non-transiently on a tangible recording medium that is readable by a computer.

[0027] The computer program for implementing the process can be advantageously sold as a plugin, suitable for integration into a "host" software application, such as audio-visual or audio-visual production and / or processing software like Pro Tools™, Media Composer™, Premiere Pro™, or Audition™, among others. This approach is particularly well-suited to the audiovisual industry. Indeed, it eliminates the need to transfer the original (unmasked, and therefore unencrypted) audio signal to a remote server or another computer. The user's computer alone holds the source file of the original voice, that is, the audio before the masking process is executed. This significantly reduces the risk of malicious interception of the original audio signal.However, the process can be implemented using audio-phonic processing software that can easily run on standalone hardware with standard processing capabilities, such as a general-purpose computer, because its processing is performed in real time. It does not require the implementation of any artificial intelligence, voice database, or machine learning process, unlike many prior art solutions, including some of those presented in the introduction.

[0028] Alternatively, the computer program for implementing the process can advantageously be integrated, either ab initio This can be done either through a software update to the internal software embedded in equipment dedicated to the production and / or processing of audio or audiovisual content (referred to as "media" in industry jargon), such as an audio and / or video mixing and / or editing console. Such equipment is primarily intended for producers, mixers, and other media post-production professionals.

[0029] A third aspect of the invention relates to an audio-phonic or audio-visual processing device, comprising means for implementing the process. This device can be implemented, for example, in the form of a general-purpose computer capable of executing the computer program according to the second aspect above.

[0030] Finally, a fourth and final aspect of the invention relates to an audio-phonic or audio-visual processing device such as an editing and / or mixing console enabling the production of media (namely audio, audiovisual, or multimedia content) corresponding to or incorporating a speech signal from a speaker, in particular from a speaker to be protected, the device comprising a device according to the third aspect. Brève description des figures

[0031] The following description, with reference to the accompanying drawings, given by way of non-limiting examples, will clearly explain the nature of the invention and how it can be implemented. The drawings depict: Figure 1 : an organizational chart illustrating the main steps of the process according to implementation methods; Figure 2 : a very simplified schematic representation of an audio-phonic system in which the process can be implemented; Figure 3A And figure 3B : diagrams illustrating implementation methods of down-shifting and up-shifting, respectively, which can be applied to the timbre and pitch of an audio signal segment according to implementation methods; Figure 4A et figure 4B : frequency diagrams of a recorded audio sequence, showing the distribution of energy as a function of frequency before and after, respectively, the implementation of a voice masking process according to the prior art; Figure 5A et figure 5B : frequency diagrams of the same audio sequence as in the figure 4A showing the distribution of energy as a function of frequency before and after, respectively, the implementation of a voice masking process according to the proposed process; Figure 6 : a detailed flowchart illustrating the steps of the process according to implementation methods. Description de mode(s) de réalisation

[0032] In the figures, and unless otherwise specified, identical elements shall bear the same reference symbols.

[0033] The human voice is the collection of sounds produced by the friction of air from the lungs against the folds of the larynx. The pitch and resonance of the sounds produced depend on the shape and size not only of the vocal cords but also on the rest of the person's body. The size of the vocal cords is one source of the difference between men's and women's voices, but it is not the only one. The trachea, mouth, and pharynx, for example, define a cavity in which the sound waves emitted by the vocal cords resonate. Furthermore, genetic factors contribute to differences in vocal cord size even among individuals of the same sex.

[0034] Given all these characteristics that are unique to each person, the voice of each human being is singular.

[0035] The process allows a speaker's voice to be masked in order to protect their identity and / or privacy.

[0036] In what follows, the original speech signal refers to the audio signal corresponding to an acquired, undistorted sequence of the speaker's voice. The masked audio signal refers to the result of processing the original speech signal obtained by implementing the method.

[0037] According to the proposed implementation methods, the protection of the speaker's identity and / or privacy is achieved through the intentional alteration of not only the pitch but also the timbre of the speaker's voice. This alteration is carried out using digital signal processing techniques, based on computer-implemented processing algorithms.

[0038] A complex sound of fixed pitch can be analyzed into a series of elementary vibrations, called natural harmonics, whose frequency is a multiple of that of the reference frequency, or fundamental frequency. For example, if we consider a fundamental frequency with a value of f, the waves having the frequencies 2f, 3f, 4f, ..., j × f , ...and so on are considered harmonic waves. The fundamental frequency (from which the frequencies are derived) j × f (harmonics) characterize the perceived pitch of a note, for example, an "A". The distribution of the intensities of the different harmonics according to their order j Characterized by their envelope, this defines the timbre. The same applies to a speech signal as to musical notes, speech being merely a succession of sounds produced by the vocal apparatus of a human being.

[0039] It should be noted that the timbre of a musical instrument or voice refers to the set of sonic characteristics that allow an observer to identify the sound produced by ear, independently of its pitch and intensity. Timbre, for example, allows us to distinguish the sound of a saxophone from that of a trumpet playing the same note with the same intensity. These two instruments have their own resonances, which differentiate the sounds to the ear: the sound of a saxophone contains more energy in the relatively lower frequency harmonics, resulting in a relatively "muted" timbre, whereas the timbre of a trumpet has more energy in the relatively higher frequency harmonics, resulting in a "brighter" sound, even though they have the same fundamental frequency.For the voice, the vocal register refers to the set of frequencies emitted with identical resonance, that is, the part of the vocal range in which a singer, for example, emits sounds of respective pitches with an almost identical timbre.

[0040] The organizational chart of the figure 1 schematically illustrates the main steps in the process of masking a speaker's voice. The process can be implemented in an audio-phonic system 20 as very schematically represented in the figure 2 This system may include hardware 201 and software 202 enabling this implementation.

[0041] It should be noted that even though the invention relates to masking a speech signal, which by nature is an audio signal, this signal may belong to an audiovisual program (combining sound and images), such as a video of an interview with a witness who wishes and / or must remain anonymous, filmed, for example, with a hidden camera or with the image of the witness to be protected blurred. In other words, the speech signal may correspond to all or part of the soundtrack of a video, and more generally to any audio, radio, audiovisual, or multimedia program.

[0042] The audio-phonic system 20 is, for example, an audiovisual mixing device, used to edit video sequences in order to produce an audiovisual program from various video sequences and their respective "soundtracks".

[0043] The hardware 201 of the audio-phonic system 20 includes at least one computer, such as a microprocessor associated with random access memory (or RAM, from the English " Random Access memory "), and means for reading and recording digital data on digital recording media (mass storage such as an internal hard drive), and data interfaces for exchanging data with external devices. At the figure 2 A symbolically represented audio signal acquisition device 31, such as a microphone (or mic), and a data storage device 22, such as a USB flash drive, are also shown. Alternatively, the system 20 can communicate in read and / or write mode with other external data storage media in order to read the data of an audio signal to be processed and / or to record the audio signal data after processing. Furthermore, the system 20 may include communication means such as a modem or a network card for Ethernet, 4G, 5G, etc., or a Wi-Fi or Bluetooth® communication interface.

[0044] The software means 201 of the audio-phonic system 20 include a computer program which, when loaded into RAM and executed by the processor of the audio-phonic system 20, is adapted to perform the steps of the process of masking a speaker's signal.

[0045] With reference to the organizational chart of the figure 1 , in step 11 the sound of the speaker's voice is captured via microphone 31 of system 20, either for immediate processing in system 20, or for delayed processing.

[0046] Immediate processing refers to processing performed during the acquisition of the audio signal, without an intermediate step of fixing this audio signal onto any permanent recording medium. The data of the original audio signal then simply passes through the system's RAM (non-permanent memory).

[0047] Conversely, delayed processing refers to processing performed on a recording, made within or under the control of the audio-phonic system 20, of the speaker's speech signal acquired via the microphone 31. This recording is fixed onto a mass data storage medium, for example, a hard drive internal to the system 20. It could also be a peripheral, i.e., external, hard drive connected to this system. It could also be another peripheral data storage device with permanent memory capable of permanently storing the audio data of the speech signal, such as a USB flash drive, a memory card (Flash or other type), or an optical or magnetic recording medium (audio CD, CD-ROM, DVD, Blu-ray disc, etc.).

[0048] The mass data storage medium can also be a data server with which the audio-phonic system 20 can communicate to download (" upload » in English) the audio signal data so that it can be stored there, and later downloaded (“ download (in English) for subsequent processing. This server can be local, that is, part of a local area network (LAN) (from the English " Local Area Network " to which also belongs the audio-phonic system 20. The data server can also be a remote server, such as a data server in the Cloud which is accessible via the open Internet network.

[0049] Alternatively, the speech signal corresponding to the speaker's speech sequence may have been acquired via other equipment, separate from the audio-phonic system 20 that implements the speaker's voice masking process. In this case, an audio data file encoding the speaker's voice may have been recorded on removable data storage, which can then, in step 11, be connected to the audio-phonic system 20 for playback of the audio data. This audio data file may also have been uploaded to a data server in the cloud, which the audio-phonic system 20 can also access in order to download the audio data of the audio signal to be processed. In all these situations, step 11 of the process then consists solely, for the audio-phonic system 20, of accessing the audio data of the speaker's speech signal.

[0050] In all cases, step 11 of the process includes a (temporal) segmentation of the original speech signal into a series of successive audio segments of a fixed duration, which is constant from one segment to the next in the resulting series of segments. Preferably, the segmentation of the audio signal into a series of successive audio segments of the same fixed duration is performed by a temporal windowing that is independent of the content of the audio signal and can be done "on the fly".

[0051] The expression "independent of the audio signal's content" means that the windowing is independent of both the frequency content—that is, the distribution of energy across the audio signal's frequency spectrum—and the informational or linguistic content—that is, the semantics and / or grammatical structure of the speech contained in that audio signal, in the language spoken by the speaker. The process is therefore very simple to implement, since no physical or linguistic analysis of the signal is necessary to generate the signal segments for processing.

[0052] In signal processing, a time windowing operation allows the processing of a signal of deliberately limited length to a duration τ knowing that any calculation can only be performed on a finite number of values. To observe or process a signal over a finite duration, it is multiplied by an observation window function, also called a weighting window and denoted h(t). The simplest, but not necessarily the most used or preferred, is the rectangular window (or door) of size m defined as follows: h t = 1 , si t ∈ 0 m 0 , sinon

[0053] By multiplying (by numerical calculation) the digitized audio signal S(t) by the door function h(t) above, then shifting, we obtain a finite series made up of a determined number N of audio signal segments s k ( τ ), each of the same fixed duration D, and indicated by the letter k noted: s k τ k = 1 , 2 , 3 , … N Or τ denotes the relative index of time in the segment.

[0054] Advantageously, the duration D of an audio segment s k ( τ ) is equal to a fraction of a second, for example between 10 milliseconds (ms) and 100 ms (in other words, D ∈ [10 ms, 100 ms ]) . An audio segment then has a duration shorter than that of a word in the language spoken by the speaker, regardless of the language in which they are speaking. This duration is a fortiori shorter than the length of a sentence or even a portion of a sentence in that language. The length of an audio segment s k ( τ ) is then, at most, on the order of the duration of a phoneme, that is to say, the duration of the smallest speech unit (vowel or consonant). An audio segment s k ( τ Therefore, it does not, in itself, carry any informational content with regard to spoken language, as its duration is far too short for that. This gives the masking process the advantage of simplicity, and also good robustness against the risk of reversion.

[0055] It should be noted that such a decomposition of the audio signal S(t) in a series { s k (τ)} k =1,2,3,... N of segments also called elementary frames and indexed by the letter k The following, obtained by windowing and shifting, is classic in signal processing, as it allows the signal to be processed in successive time slices.

[0056] Step 11 also involves the formation of a series of audio segment pairs, each comprising a primate and a duplicate of an audio segment from the series of audio segments above. As will be seen in more detail later, with reference to the step diagram of the figure 6 These pairs can be defined more specifically in the frequency domain, after applying a Fourier transform (FT) to the segments s k ( τ ) of the audio signal in the time domain. In each pair formed by a primate and a duplicate of a segment of the speaker's original speech signal, these two elements are identical to each other and originate from the same segment of the speaker's original speech signal. In what follows and in the figures of the accompanying drawings, the series of primates and the series of duplicates of the audio segments of the speech signal thus produced undergo processing for each primate and each duplicate of the audio segment in a pair, to extract, on the one hand, the envelope of the harmonics characterizing the timbre of the audio segment, and on the other hand, the signal characterizing the pitch of the audio segment. In the figure 1 (the series of stamps and the series of heights are designated interchangeably by the letters A and B, or vice versa).

[0057] For each pair of segments, the signals characterizing the pitch and timbre extracted from the original and duplicate undergo parallel processing, essentially independent of each other. This processing is illustrated by steps 12a and 13a of the left branch and by steps 12b and 13b of the right branch, respectively, of the algorithm schematically illustrated by the flowchart of the figure 1 .

[0058] Step 12a is a first upward distortion (denoted MOD A in what follows and in the diagrams), applied to each element of the audio segments in series A. This upward distortion is not identical from one element of series A to another. Rather, it evolves according to at least one first masking parameter. However, regardless of how the first masking parameter evolves, this first upward distortion always has the effect of raising a specific portion of the frequency content of the primacy of the audio segment to which it is applied. This means that all or part of the primacy frequencies of the segment in question are shifted towards higher frequencies, relative to the corresponding audio segment of the original speech signal. The application of the first distortion generates an altered timbre (here, upward) of the audio segment.

[0059] Step 12b is a second, downward alteration (denoted MOD B in what follows and in the diagrams) applied to each element of series B of the audio segments. Just like the upward alteration MOD A applied to the elements of series A, this downward alteration MOD B is not identical from one element of series B to another. This means that it changes according to at least one second masking parameter. However, regardless of how this second masking parameter changes, this downward alteration always lowers a specific portion of the frequency content of the audio segment element to which it is applied. This means that all or part of the frequencies of the audio segment in question are shifted towards lower frequencies, relative to the corresponding audio segment of the original speech signal. Applying the second alteration results in a pitch shift (here, downward) of the audio segment.

[0060] It should be noted that it is advantageous for each of the MOD A and MOD B alterations to be limited in terms of the evolution of the frequency content of the elements of the audio segment to which it is applied. This means that these alterations of the frequency spectrum are each either exclusively upward or exclusively downward, without any change in the direction of the shift of the relevant frequencies within the spectrum. Indeed, this prevents the audio-phonic system 20 from being used by malicious individuals to whom it has been supplied or made available, or who might gain access to it by any other means, in order to reverse the alteration of the audio signal.Such a reversal could indeed consist of applying alterations to the masked audio signal (which the malicious third party had copied or intercepted in any way) using judiciously chosen masking parameters to revert to the original speech signal, that is, the audio signal corresponding to the speaker's natural voice. However, thanks to the implementation methods described above, such a maneuver is not possible with the audio-phonic system 20 according to the invention itself. In fact, no change to the masking parameter values ​​of the upward alteration MOD A and the downward alteration MOD B that the malicious third party might attempt can reverse the unidirectional shifts of the pitch and timbre, respectively, of the original speech signal. In other words, the audio system 20 does not offer the possibility of reversing the alteration it produces.This does not prevent a malicious third party from attempting this fraud with other means, but at least the system used to mask the audio signal containing a speaker's natural voice cannot be diverted from its function, in fact "reversed", in order to undermine the protection of the speaker that it provides.

[0061] The process then includes a step 15 combining the timbre of the audio segment, altered by the MOD A alteration and obtained in step 12a, with the pitch of the audio segment, altered by the MOD B alteration and obtained in step 12b, to form a single resulting altered audio segment. By combination, we mean an operation that, from a physical point of view, recombines the respective altered spectra, that is, merges the respective frequency contents of the altered timbre of the audio segment and the altered pitch of said audio segment, possibly with averaging and / or smoothing. In signal processing, this can be achieved by multiplication (symbol "×") or by convolution (symbol "*"), either in the time domain or in the frequency domain after transforming the audio signal(s) from the time domain to the frequency domain by a Fourier transform.

[0062] The process further includes, from one pair of audio segments to another in the series of pairs of audio segments: In step 13a for the elements of series A, a variation of at least one parameter of the alteration MOD A, for example a variation of this alteration within a configurable width interval, this variation being symbolically denoted VAR A in what follows and in the figures, and; in step 13b for the elements of series B, a variation of at least one parameter of the alteration MOD B, for example a variation of this alteration within a configurable width interval, this variation being symbolically denoted VAR B in what follows and in the figures, said variations of alterations being variable from one pair of segments to another in the series of pairs of audio segments.

[0063] Those skilled in the art will appreciate that, in practice, steps 12a and 12b on the one hand, and steps 13a and 13b on the other hand, can be carried out in the reverse order of that presented in the figure 2 In other words, they are interchangeable: steps 13a and 13b can be executed after (as shown) or before steps 12a and 12b.

[0064] Preferably, steps 13a and 13b cause a local perturbation, around time τ, of the (spectral) characteristics of the timbre and pitch, said perturbation varying from one segment to another in the series { s k (τ)} k =1,2,3,... N (therefore as a function of k) randomly, non-stationarily (for example, by random walks) and independently on each of the two spectral components, namely pitch and timbre.

[0065] In one example of implementation, which is not exhaustive, the alteration of the audio signal's pitch can thus correspond to a "directional" variation, namely a rise or fall, of the fundamental frequency of the audio signal, which can take any of the following predetermined values: ± 6.25%, ± 12.5%, ± 25%, ± 50%, and ± 100%. These example values ​​correspond approximately to variations of a semitone, a tone, a third, a fifth, or an octave, respectively, of the pitch (i.e., the fundamental frequency, or "pitch") of the original speech signal.

[0066] Repeating step 14 sequentially for successive pairs of primates and duplicates of the audio segments generated in step 11 generates a series of altered audio segments.

[0067] The process finally includes, in step 15, the recomposition of the masked audio signal from the series of altered audio segments obtained by repeating the previous steps, 12a-12b, 13a-13b and 14. This recomposition is carried out by superposition-addition, in the time domain, of the successive elements of the series of altered audio segments produced in step 14, as they are transformed.

[0068] It should be noted that, in the resulting altered audio segment, the frequency content is doubly altered compared to the spectrum of the original speech signal segment. This results from the cumulative effects of the respective MOD A and MOD B functions.

[0069] In some implementation modes, successive changes in the first masking parameter and the second masking parameter that occur at each occurrence of steps 13a and 13b, respectively, induce random variations of said first and second parameters, from one pair to another in the series of pairs of audio segments generated in step 11.

[0070] Since MOD A and MOD B alterations affect different components of the spectrum of the segment of the original speech signal under consideration, since they also use distinct masking parameters, and since their respective masking parameters evolve independently of each other in a random manner, the masking effect that is obtained is very difficult, if not impossible, to reverse.

[0071] Thus, the variations of the first and second masking parameters are themselves randomly fluctuating from one pair of segments to another in the series of audio segment pairs. In other words, the variations denoted VAR A and VAR B in steps 13a and 13b of the parameters of the modifications denoted MOD A and MOD B introduced in steps 12a and 12b fluctuate over time. In particular, this fluctuation occurs from one segment to another of the original speech signal. Therefore, on the figure 1 , this fluctuation is symbolized by an operation denoted VAR A+B in step 14.

[0072] There figure 3A and the figure 3B illustrate a method of implementing down-segment and up-segment alteration, respectively, which can be applied to the timbre and pitch of an audio signal segment, in steps 12a and 12b, respectively, of the process illustrated by the flowchart of the figure 1 .

[0073] In this example, the ascending alteration MOD A is applied to the pitch of the voice, symbolized by the figure 3A by a tuning fork. A tuning fork is known as an object whose acoustic resonance produces a sound with a pure frequency, such as, in principle, the fundamental frequency (or pitch) of a human voice. Furthermore, the descending alteration MOD B is applied to the timbre of the voice, symbolized by the figure 4A by the frequency spectrum envelope of an audio signal. Of course, the example represented by the figures 3A And 3B is not limiting. The upward alteration MOD A can conversely be applied to the fundamental frequency (pitch) while the downward alteration MOD B would be applied to all or part of the harmonic envelope (timbre).

[0074] In all cases, the two alterations MOD A and MOD B each produce shifts of certain frequencies (namely, in the example considered here, the pitch for one, and the harmonic envelope for the other) along opposite directions in the frequency spectrum (namely, an upward shift towards the higher frequencies for one, and a downward shift towards the lower frequencies for the other). In the resulting protected audio signal, these effects operating in two different directions provide good protection while preserving a certain intelligibility of the audio signal. Indeed, the "masculinizing" effect of a frequency shift towards the lower frequencies resulting from the upward alteration MOD A is partially counterbalanced by the "feminizing" effect of a frequency shift towards the higher frequencies resulting from the downward alteration MOD A.This avoids generating a masked signal close to the voice of "Darth Vader" ™< or close to the voice of "Mickey Mouse" ™< .

[0075] The audio file obtained after implementing the process of figure 1 can be transmitted by email, posted on social media or a website, broadcast over the airwaves, or distributed on any recording medium. The process allows the voice to be masked. posteriori, on a recording of the speaker's voice, as can easily be done with audio editing software. As offered as a computer program as a plugin to be integrated into audio-sound or audio-visual processing software, the method does not allow for making audio or video calls with a masked voice.

[0076] Once the process is implemented on the audio-phonic or audio-visual platform with which the speaker's voice is acquired, the original voice does not circulate on any computer network, thus avoiding the risk of interception by a malicious third party of the data corresponding to the unmasked voice.

[0077] The computer program that implements the masking process, by performing the corresponding digital processing calculations, can be included in a host software, for example the operational software of an audio-phonic processing environment, such as an audio mixing console or audio-visual editing console.

[0078] The result obtained by implementing the process, namely the masked audio signal, can be fixed, that is to say recorded: either on a separate track, added "as an insert" in the program being composed on the audio-phonic or audio-visual processing system; or directly on the original audio file which has been processed, for example by replacing the data of the original speech signal in order to remove the original recording of the speaker's voice and thus guarantee its perpetual protection.

[0079] This result is irreversible in audio and cannot be analyzed by speech recognition. It is immediately readable, meaning that the audio data file or the corresponding audio track can be played to listen to the masked audio signal, particularly to verify by ear or by any other available technical means that the speaker's original voice is no longer recognizable.

[0080] There figure 4A is a frequency diagram of a recorded audio sequence, showing the distribution of energy as a function of time (on the x-axis) and frequency (on the y-axis). figure 4B is a frequency diagram of the audio sequence of the figure 4A After implementing a voice masking technique according to prior art, by simply shifting the pitch. The term "pitched" is sometimes used to refer to a signal that has undergone such a shift. Comparing these two frequency diagrams clearly shows a very strong similarity in the harmonics of the signal between the original signal and the pitched signal.

[0081] There figure 5A and the figure 5B allow comparison of the frequency diagrams of the same audio sequence as at the figure 4A by showing the energy distribution as a function of frequency before and after, respectively, the implementation of a voice masking process according to the proposed method. This comparison shows that the harmonics of the signal have undergone significant transformations. A clear distinction is made at the figure 5B than harmonics of the figure 5A have undergone significant modifications, thus masking the harmonics of the original signal. This masking makes it extremely difficult, if not impossible, to compare the spectrograms of the original speech signal and the masked speech signal.

[0082] Implementation methods for the process presented above schematically and in its main stages only will now be described in more detail with reference to the flowchart of the figure 6 .

[0083] The implementation of the process consists of applying a numerical treatment, here for example in the time-frequency domain which is better suited to this type of processing by calculations, following { s k (τ)} k= 1, 2, 3, ... segments s k (τ) of the digitized speech signal S(t). Such a segment is noted s k ( τ ) at the top of the figure 6 Those skilled in the art will appreciate that, in practice, the treatment illustrated by the step diagram in this figure is obviously applied successively to each segment. s k ( τ ) indexed by the letter k.

[0084] The segment s k ( τ ) is subjected in step 61 to a Fourier Transform (FT), for example a short-term Fourier transform known by the acronym STFT (or TFCT, from the English " Short Term Fourier Transform ») in order to move into the time-frequency domain. Each segment s k ( τ ) of duration τ in the temporal domain, is thus converted to give a segment denoted S k ( t, f ) which takes complex values ​​in the time-frequency domain.

[0085] In step 62, a decomposition of the segment is performed S k ( t, f ) in a module term noted X k ( t, f ) and a phase term denoted Q k (t , f These terms are such that: S k t f = X k t f × Q k t f Or : X k t f = S k t f ; And, Q k t f = exp i × Arg S k t f , où Arg denotes the argument of a complex number.

[0086] The term X k ( t, f ) corresponds to the Power Spectral Density (PSD) of the audio signal in the vicinity of the instant t. From this term X k ( t, f ), it is then possible, on the one hand, to determine the fundamental frequency of speech (or "pitch") of speech, that is to say the height, and on the other hand, to estimate the envelope of the Power Spectral Density, that is to say the timbre.

[0087] More specifically, in step 63, a pair of segments is formed that are initially equal to each other and equal in terms of modulus. X k ( t, f ) of the segment S k ( t, f ), and which, for the purposes of this discussion, is called the primate and duplicate of the segment S k ( t, f ) . We will also sometimes speak of a series of pairs, each formed (that is, for each value of the index k) by this primacy and this duplicate of the segment S k ( t, f Differentiated treatments applied to the primate and duplicate, respectively, of the segment thus allow us to separate the module term X k ( t, f ) in two components A k ( t, f ) And B k ( t, f ) distinct such that, in the time-frequency domain, we have: X k t f = A k t f × B k t f , Or : A k ( t, f ) corresponds, for the segment of the index signal kconsidered, to the signal characterizing the timbre of the audio signal; and, B k ( t, f ) corresponds, for this segment, to the signal characterizing the height (or pitch) of the audio signal.

[0088] For example, the timbre component A k ( t, f ) can be obtained by the cepstri method. For this purpose, an Inverse Fourier Transform (IFFT, denoted as " Inverse Fast Fourier Transform (in English), and we then obtain the cepstrum, which is a dual-time form of the logarithmic spectrum (the spectrum in the frequency domain becomes the cepstrum in the time domain). After this transformation, the fundamental frequency can be calculated from the cepstral signal by determining the index of the main peak of the cepstrum, and by windowing the cepstrum, we obtain the envelope of the spectrum that corresponds to the timbre component. A k ( t, f ).

[0089] The height (or pitch) component B k ( t, f ) ,As for it, it can then be obtained by dividing the signal point by point X k ( t, f ) by the value of the stamp component A k ( t, f In other words, to obtain the pitch component B k ( t, f ), we can "subtract" (which is achieved by a division calculation in the time-frequency space) from the modulus term X k ( t, f ) of the segment S k ( t, f ) the contribution A k ( t, f ) of the spectrum envelope to obtain "what remains" which is treated as the (spectrum of the) signal characterizing the height (or pitch) or more generally what is called the fine structure of the Power Spectral Density (PSD).

[0090] In steps 64a and 65a, on the one hand, and in steps 64b and 65b, on the other hand, ascending or descending alterations are then applied to the envelope A k ( t, f ) of the spectrum corresponding to timbre and fine structure B k ( t, f ) of the spectrum corresponding to the pitch, according to a preferably monotonic transformation along the frequency axis, these alterations being distinct from one another in their implementation methods, and also being each randomly variable from one audio signal segment to another. These alterations allow the timbre and pitch of the signal to be modified independently and variably over time (non-stationary), more specifically from one audio signal segment to another, that is, as a function of the index k. For each of the timbre and pitch, this result is obtained globally by multiplying in the time-frequency domain the component A k ( t, f ) Or B k ( t, f ), respectively, of the power spectral density X k ( t, f ) : on the one hand, through a function of altering the frequency scale Γ A ( f ) Or Γ B ( f ) , at step 65a for the stamp component A k ( t, f ) and at step 65b for the height component B k ( t, f ), respectively, which are preferably monotonic, one of which is ascending while the other is descending relative to its effects on the frequency content of the original audio segment S k ( t, f ) ; and, on the other hand, by a time variation function γ A ( t ) ou γ B ( t ) applied globally to the frequency scale, at step 64a for the timbre component A k ( t, f) and in step 64b for the height component, B k ( t, f ) respectively.

[0091] The order in which these operations are performed in the time-frequency domain is irrelevant. In the implementation shown in the figure 6 First, the components are multiplied. A k ( t, f ) And B k ( t, f ) by the time variation functions γ A ( t ) And c B ( t ) , respectively, and then we perform the multiplications of the respective results of these first multiplications by the frequency alteration functions C A ( f ) And C B ( f ) , à Step 65a and step 65b, respectively. But these two sets of multiplications could just as easily be performed in reverse order. In other words, steps 65a and 65b could be performed before steps 64a and 64b, respectively.

[0092] As you will have understood, and as shown on the left of the blocks illustrating steps 64a, 64b, 65a and 65b in the figure 6 , in these implementation modes the frequency alteration functions C A ( f ) And C B ( f ) correspond to MOD A and MOD B alterations, respectively, which were presented above with reference to the figure 1 Similarly, the time variation functions c A ( t ) And c B ( t ) , correspond to the VAR A and VAR B variations, respectively, which were presented above with reference to the figure 1 To avoid any ambiguity, it should be noted that, from the point of view of the frequency spectrum of the original audio segment S k ( t, f ), which varies over time due to the effect of time variation functions c A ( t ) And c B ( t ) ,This is the overall effect on this spectrum, and more specifically on timbre and pitch respectively, of the combination of frequency alteration functions. C A ( f ) And C B ( f ) and time variation functions c A ( t ) And c B ( t ) , respectively, that is, the summation (or addition) of their respective effects. These respective effects of the combination of frequency alteration functions C A ( f ) And C B ( f ) and time variation functions c A ( t ) And c B ( t ) on the frequency spectrum of the original audio segment is more specifically related to frequency alteration functions C A ( f ) And C B ( f ) , respectively, the time variation functions c A ( t ) And c B ( t) having only the effect of varying them, preferably randomly, in order to strengthen the robustness of the masking against attempts to reverse due to malicious intent.

[0093] In the example shown in the figure 6 Step 64a includes the application to the signal A k ( t, f ) which corresponds to the timbre component, on the frequency scale f, of the time variation function c A ( t ) , to generate an intermediate signal, noted A' k (t, f ) , of the stamp component A k ( t, f This operation can be written as a multiplication in the time-frequency domain, as follows: A ′ k t f = A k t , f × γ A t

[0094] The function c A ( t) is a linear function. Preferably, and as already mentioned above, it fluctuates randomly over time, varying from one segment of the original audio signal to another in the series of segments S k ( t, f ) which are processed sequentially. In other words, it changes according to the value of the index k, according to a random process whose refresh is governed by a parameter i , so that the alteration of the timbre is not stationary.

[0095] Similarly, step 64b includes the application to the signal B k ( t, f ), which corresponds to the pitch component (or pitch=, on the frequency scale f of the time variation function c B ( t ) , to generate an intermediate signal, noted B'k ( t, f This operation can be written as a multiplication in the time-frequency domain, as follows: B ′ k t f = B k t , f × γ B t

[0096] The function c B ( t ) is a linear function. Preferably, and as already mentioned above, it fluctuates randomly over time, varying from one segment of the original audio signal to another in the series of segments S k ( t, f ) which are processed sequentially. In other words, it changes according to the value of the index k according to a random process whose refresh is governed by a parameter i so that the alteration of height is not stationary.

[0097] The fluctuations, as a function of time, of the time variation function c A ( t ) applied to the timbre component and / or the time variation function c B ( t) applied to the pitch component (height), and all the more so when one and / or the other of these fluctuations are random, helps to reinforce the irreversibility of the voice masking process.

[0098] For example, the time variation function c A ( t ) can vary according to a random walk within an amplitude range [ δ A min< , δ A max< ] determined and with a temporal refresh rate corresponding to the parameter i as mentioned above, where δ A min< , δ A max< And i are the first masking parameters, associated with the time variation function c A ( t ) .

[0099] Similarly, the time variation function c B ( t ) can, for example, vary according to a random walk within an amplitude range [ δ B min< , δ B max< and with a temporal refresh rate corresponding to the parameter i as mentioned above, where δ B min< , δ B max< And i are second parameters, associated with the time variation function c B ( t ) . The fluctuations of the two time variation functions c A ( t ) And c B ( t ) are preferably independent of each other, in order to reinforce the irreversibility of the alterations. In other words, the temporal variation functions c A ( t ) And c B ( t ) are uncorrelated.

[0100] We will appreciate that the parameter i is the parameter of the fluctuation denoted VAR A+B at the figure 1 This parameter defines, for example, the number of random variations per second in the spectral alterations of an audio segment. For example, if i If VAR was equal to zero, the variations VAR A and VAR B are stationary, so the results of the alterations MOD A and MOD B would be fixed, which is not the case in practice. In an example i has a value between 1 and 10. Since this value has the dimensions of a frequency, we can say that i is between 1 and 10 Hz. This value is lower than the frequency of the temporal segmentation of the original speech signal into audio segments (by windowing), which is more in the order of 100 Hz.

[0101] Next, in steps 65a and 65b, frequency alteration functions are applied C A ( f ) And C B ( f ) , respectively, to the stamp component A k ( t, f) and the height component B k ( t, f ), respectively, to generate a timbre component of the masked audio segment, denoted A"k ( t, f ), and a pitch component of the masked audio segment, denoted B"k ( t, f ), respectively. These frequency alteration functions C A ( f ) And C B ( f) correspond to the alterations noted MOD A and MOD B on the figure 1 .

[0102] Each of these operations can be written as a multiplication in the time-frequency domain, in the following way: A " k t f = A ′ k t , Γ A f B " k t f = B ′ k t , Γ B f

[0103] The function C A ( f ) and the function C B ( f ) can be linear or nonlinear deformation functions of the frequency axis. In the case where one and / or the other are linear, we have: Γ A f = f × Γ A and / or, respectively Γ B f = f × Γ B

[0104] Preferably, the alteration functions C A ( f ) And C B ( f ) are monotonic, meaning that the distortion they introduce on the frequency axis is either upward, with the effect of raising a specific part of the frequency content of the audio segment sk ( t ) ,either downward, with the effect of lowering a specific part of the frequency content of the audio segment sk ( t ) . Furthermore, they are constrained in opposite directions, in that if one is ascending monotonically, the other is descending monotonically, and vice versa. This prevents the software implementing the masking process from being used itself to attempt a reversal of the speaker's voice masking process, as already explained above with reference to steps 12a and 12b of the figure 1 .

[0105] Furthermore, the fact that among the alteration functions C A ( f ) And C B ( f) one is an upward alteration function, while the other is a downward alteration function, which helps to preserve the intelligibility of the voice after masking, since the frequency shift(s) towards the high end, on the one hand, and the frequency shift(s) towards the low end that they produce, on the other hand, partially compensate for each other by avoiding too much distortion of the voice, which would otherwise be predominant in the masked audio signal.

[0106] One of the advantages of the process stems from implementations in which these MOD A and MOD B modifications are varied for the indices ksuccessive, random, uncorrelated sequences (one for timbre and the other for pitch) are used to continuously modify these two voice characteristics independently, unpredictably, and non-stationarily. Unlike methods where the modification is constant, this makes it impossible to reverse the process once the frequency variations have been made. The protection is all the stronger when the random variations VAR A and VAR B are significant.

[0107] The next two steps allow us to preserve the temporality of the original by resynthesizing the audio signal masked by the index k.

[0108] Thus, step 67 includes the reconstruction of each modified audio segment, noted X"k ( t, f ), in the time-frequency domain, by recombination of the new envelope A"k ( t, f ) and the new fine structure of the frequency spectrum B"k ( t, f) of the audio segment under consideration. The term "new" used here in reference to the envelope and fine structure means that it refers to the envelope and fine structure after masking, that is, after the application of frequency alteration functions. C A ( f ) And C B ( f ) correspond to MOD A and MOD B alterations, respectively, and time variation functions c A ( t ) And c B ( t ) , respectively. This reconstruction can be obtained by multiplying the new timbre component in the time-frequency domain. A"k ( t, f ) by the new pitch component B"k ( t, f ) of the frequency spectral density (FSD) of the masked audio segment, as follows: X " k t f = A " k t f × B " k t f

[0109] Step 68 involves the reconstruction of each masked audio segment noted S"k ( t, f ) ,in the time-frequency domain. This recomposition can be obtained by multiplying the modulus component in the time-frequency domain. X"k ( t, f ) by the corrected phase component Q" k ( t, f ) of the hidden audio segment S"k ( t, f ), in the following manner: S " t f = X " t f × Q " t f

[0110] The corrected phase component Q" k ( t, f ) of the hidden audio segment S"k ( t, f ) is obtained, in the example shown in the figure 6 , at step 66 from the phase term Q k ( t, f ) of the audio segment considered S k ( t, f ), which phase term was generated in step 62. Step 66 is intended to correct the phase term Q k ( t, f ) of the audio segment S k ( t, f ) depending on random variations c B ( t ) and the alteration function C B ( f) which were applied to the pitch term B ( t, f ) . This ensures the temporal continuity of the phase F"k ( t, f ) , of the hidden audio segment S"k ( t, f ), that is to say, the continuity of the phase F"k ( t, f ) of this segment with F"k ( t - 1, f ) , Or F"k ( t, f ) corresponds to Arg S"k ( t, f ) .

[0111] It should be noted that such phase correction is known in itself and is generally implemented in any signal transformation process whenever the power spectral density of a signal is modified. In the implementation methods proposed here, it is generated at step 66 solely based on modifications made to the pitch component. B"k ( t, f ) of the power spectral density of the masked audio segment S"k ( t, f) compared to the pitch component B k ( t, f ) of the power spectral density of the original audio segment S k ( t, f Indeed, essentially, it is the changes made to the pitch that necessitate phase alignment of the frequency components of the spectrum. However, those skilled in the art will appreciate that the phase alignment in step 66 could also take into account the changes made to the timbre component. A"k ( t, f ) of the frequency spectral density of the masked audio segment S"k ( t, f ) compared to the timbre component A k ( t, f ) of the frequency spectral density of the original audio segment S k ( t, f This is not shown on the organizational chart of the figure 6in order not to overload it, which would impair its readability, but the person skilled in the art understands, based on their usual knowledge and in view of the indications provided here, how this can be implemented in practice.

[0112] Once the audio segment is hidden S"k ( t, f The signal was obtained by calculations in the time-frequency domain as described above; it only remains to bring it back into the time domain, which is done in step 69. This step consists of generating the masked signal. s"k (τ) in the time domain, from the signal S"k ( t, f ) in the time-frequency domain. For example, this can be achieved by an OLA method (from the English " Overlap-and-Add " on the successive inverse Fourier transforms of s"k(τ). The OLA method, also called the superposition and addition method, is based on the linearity property of linear convolution. The principle of this method consists of decomposing the linear convolution product into a sum of linear convolution products. Of course, other methods can be considered by a person skilled in the art to perform this inverse Fourier transform, in order to generate s"k (τ) in the time domain from S"k ( t, f ) in the time-frequency domain.

[0113] The process that has been presented in the preceding description can be implemented by a computer program, for example as a plugin that can be integrated into audio-phonic or audio-visual processing software.

[0114] To the figure 6 Reference 60 collectively designates the parameters for masking a speaker's voice, namely δ A min< , δ A max< , δ B min< , δ B max< , θ , C A And C Bwhich can be adjusted by a user, via a suitable human-machine interface of the device on which the speaker's voice masking software is run.

Claims

1. Method for masking the voice of a speaker in order to protect their identity and / or their privacy by intentionally altering the pitch and the timbre of their voice, comprising: - temporally dividing an audio signal corresponding to an original recording of the voice of the speaker into a series of successive audio segments of a determined constant duration, and forming a series of pairs of audio segments each comprising a primary version and a duplicate of an audio segment of said series of audio segments; and, for each pair of audio segments: - processing the primary version of the audio segment and processing the duplicate of the audio segment in order to extract therefrom a signal characterizing the pitch of the audio segment and comprising the fundamental frequency of said audio segment, on the one hand, and a signal characterizing the timbre of the audio segment and comprising the envelope of the harmonics of said audio segment, on the other hand; - a first alteration (12a; 65a) that evolves as a function of at least a first parameter, called masking parameter, applied to the signal characterizing the timbre of the audio segment, and having the effect of altering all or part of the envelope of the harmonics, so as to generate a first altered signal; - a second alteration (12b; 65b) that evolves as a function of at least a second parameter, called masking parameter, applied to the signal characterizing the frequency pitch of the audio segment, and having the effect of altering the value of the fundamental frequency, so as to generate a second altered signal; one of the alterations out of the first alteration (12a; 65a) and the second alteration (12b; 65b) being a rising alteration, having the effect of raising a determined portion of the frequency content of the audio segment to which it is applied, while the other alteration is a falling alteration, having the effect of lowering a determined portion of the frequency content of the audio segment to which it is applied, and - combining (15; 67) the two altered signals, so as to form a resulting altered audio segment, the method furthermore comprising, from one pair of audio segments to another in the series of pairs of audio segments: - varying (13a; 64a) the first masking parameter of the first alteration; and - varying (13b; 64b) the second masking parameter of the second alteration, said variations of said first and second alterations fluctuating randomly from one pair of segments to another in the series of pairs of audio segments, and the method furthermore comprising: - recomposing a masked audio signal from the series of altered audio segments.

2. Method according to Claim 1, wherein the temporal division of the audio signal into a series of successive audio segments of a determined duration is carried out by time windowing independent of the content of the audio signal.

3. Method according to Claim 1 or 2, wherein the temporal division of the audio signal is configured such that the duration of an audio segment is equal to a fraction of a second, such that successive variations from one pair of segments to another, of the first masking parameter and of the second masking parameter occur multiple times per second.

4. Method according to any one of the preceding claims, wherein the second alteration corresponds to a variation of the fundamental frequency by any one of the following values: ±6.25%, ±12.5%, ±25%, ±50% and ±100%.

5. Computer program comprising instructions that, when the computer program is loaded into the memory of a computer and is executed by a processor of said computer, cause the computer to implement all of the steps of the method according to any one of Claims 1 to 4.

6. Audio or audiovisual processing device comprising means for implementing all of the steps of the method according to any one of Claims 1 to 4.

7. Audio or audiovisual processing apparatus such as an editing and / or mixing console for producing audio, audiovisual or multimedia content corresponding to or incorporating a speech signal of a speaker, in particular of a speaker to be protected, the apparatus comprising a device according to Claim 6.