Personalization of a neural network for speech enhancement
A compact AI model with dual-stage architecture and voice print integration addresses neural network limitations in noisy environments, enhancing target voices efficiently on portable devices.
Patent Information
- Application Number
- FR2024003534
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-05
- Publication Date
- 2025-10-10
AI Technical Summary
Neural networks for speech enhancement perform poorly in environments with overlapping voices and require significant computing power, leading to degraded call clarity and privacy issues.
A personalized speech enhancement system using a compact artificial intelligence model with a dual-stage architecture and voice print integration, which includes an encoder block and decoder, to adapt to individual vocal characteristics, requiring limited computing power and energy consumption.
The system effectively enhances target voices in noisy environments, improving call clarity and privacy by personalizing speech enhancement on portable devices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Personalization of a neural network for speech enhancement
[0001] The invention relates to the field of speech enhancement methods which use an artificial intelligence model.
[0002] BACKGROUND OF THE INVENTION
[0003] Speech enhancement (or SE) is a feature that involves isolating a voice in a noisy sound environment, to reproduce this voice while improving its intelligibility.
[0004] This is a feature that is very popular in audio technologies today, particularly due to the increase in video conferencing and teleworking.
[0005] The implementation of this functionality in portable audio devices is particularly relevant.
[0006] For example, in the case of a headset, it is very interesting to be able to isolate the voice of the headset user during a telephone call, and to transmit it enhanced and denoised to the user's interlocutor.
[0007] Neural networks are currently the algorithms that achieve the best performance for speech enhancement.
[0008] However, neural networks exhibit degraded performance and generally perform poorly in environments where interfering voices overlap with the user's voice. In addition to degrading call clarity, this also poses a privacy issue as background conversations (e.g., private conversations) may be transmitted during the call.
[0009] We therefore envisage implementing personalized speech enhancement (PSE), which aims to extract from a noisy environment including interfering voices, a predefined “target voice” belonging to an identified speaker.
[0010] In the application example just given, the target voice is that of the headset user.
[0011] In the case of the headset and, more generally, of any portable audio device, the personalized speech enhancement process requires “on-board” processing to enable the restitution of the target voice in real time.
[0012] The known traditional methods were quite ineffective. However, recently, the Deep Noise Suppression (DNS) challenge organized a test corresponding to the task of personalizing speech enhancement. This challenge has has given rise to a lot of work on this subject, which has led to the design of relatively efficient architectures. However, even if these architectures are considered "real-time" because they are causal, they are often very heavy and cannot be implemented in an embedded product.
[0013] SUBJECT OF THE INVENTION
[0014] The subject of the invention is a personalized speech enhancement system which is effective and whose implementation requires limited computing power and energy consumption. Summary of the invention
[0015] In order to achieve this goal, a personalized speech enhancement system is proposed, comprising: • at least one memory, in which a previously recorded voice print of a speaker and a previously trained speech enhancement artificial intelligence model are stored; • a processing unit designed to, in real time: • acquire at least one input audio signal comprising a target speaker voice and ambient noise; • produce at least one characterizing vector from the at least one input audio signal; • perform an inference of the artificial intelligence model, by applying the at least one characterizing vector as input to said artificial intelligence model, and by introducing the voice print into said artificial intelligence model to personalize it; • acquire a reproduction of the target voice at the output of said artificial intelligence model.
[0016] The introduction of the speaker's voiceprint into the speech enhancement artificial intelligence model, when performing the inference thereof, makes it possible to personalize the speech enhancement. The model is thus adapted to the speaker's vocal characteristics. The effectiveness of the target voice enhancement is therefore significantly increased, particularly when the ambient noise includes interfering voices. By using a compact model, a personalized speech enhancement system is therefore obtained which is very effective, and which requires limited computing power and energy consumption. The personalized speech enhancement system can thus be implemented on-board, particularly in a portable audio device (headset, earphones, hearing aid, etc.).
[0017] We further propose a system as previously described, in which the artificial intelligence model comprises:
[0018] - an encoder block, at the input of which is applied at least one character vector terizing, and generating at least one output vector;
[0019] - at least one decoder, at the input of which is applied at least one vector of exit ;
[0020] the encoder block generating at least one intermediate vector and comprising at least one operator arranged to merge the at least one intermediate vector and the voice print to produce the at least one output vector.
[0021] A system as previously described is further provided, wherein the artificial intelligence model comprises a first stage performing a low-precision analysis producing a low-precision reproduction of the target voice, and a second stage using results from the first stage and performing a more precise analysis to obtain the reproduction of the target voice.
[0022] A system as previously described is further proposed, in which the processing unit produces a first characterizing vector and a second characterizing vector from the input audio signal, and in which:
[0023] - the encoder block comprises a first branch belonging to the first stage and at the input of which the first characterizing vector is applied, and a second branch belonging to the second stage and at the input of which the second characterizing vector is applied;
[0024] - the model comprises a first decoder belonging to the first stage and a second decoder belonging to the second floor.
[0025] We further propose a system as previously described, in which the operator is arranged to merge a first intermediate vector produced by the first branch of the encoder block, a second intermediate vector produced by the second branch of the encoder block, and the voice print.
[0026] A system as previously described is further proposed, in which the first branch of the encoder block comprises a first operator arranged to merge a first intermediate vector produced by the first branch of the encoder block and the voice print, and / or the second branch of the encoder block comprises a second operator arranged to merge a second intermediate vector produced by the second branch of the encoder block and the voice print, the first branch producing a first output vector applied to the input of the first decoder and the second branch producing a second output vector applied to the input of the second decoder.
[0027] A system as previously described is further provided, in which the processing unit produces, from the at least one input audio signal, an audio signal current which is a signal of the frequency domain, the first characterizing vector being obtained by calculating an ERB of the current audio signal, and the second characterizing vector being obtained by calculating a complex spectrogram of the current audio signal.
[0028] We further propose a system as previously described, in which the at least one operator comprises at least one concatenation operator arranged to concatenate along a frequency axis the at least one intermediate vector and the voice print.
[0029] A system as previously described is further proposed, in which a structure of the model is based on a structure of a U-Net network.
[0030] We further propose a system as previously described, in which the structure of the model is based on a structure of a DeepFilterNet2 network.
[0031] We further propose a system as previously described, the voice print being obtained using a voice encoder of the ECAPA-TDNN type.
[0032] A system as previously described is further proposed, in which the voice print is generated by the processing unit of the system.
[0033] A system as previously described is further proposed, in which the processing unit is arranged to generate, from the at least one incoming audio signal, a current voice print, to calculate a similarity value representative of a similarity between the current voice print and the voice print stored in the memory, and to introduce into the model a new voice print formed from the voice print stored in the memory, the current voice print and the similarity value.
[0034] We further propose a system as previously described, in which the processing unit is arranged to implement a feedback loop arranged to generate, from the reproduction of the target voice at the output of the model, a new voice print, and to execute another inference of the model this time using the new voice print.
[0035] We further propose a system as previously described, the speaker being a user of the system, who communicates with a remote interlocutor.
[0036] We further propose a system as previously described, the speaker being an individual who communicates with the user of the system.
[0037] An audio device incorporating a system as previously described is further provided.
[0038] We further propose a method for personalized speech enhancement, implemented in the processing unit of the system as previously described, and comprising the steps, implemented in real time, of: • acquire at least one input audio signal comprising a target voice of the speaker, and ambient noise; • produce at least one characterizing vector from the input audio signal; • perform an inference of the artificial intelligence model, by applying the at least one vector characterizing as input said artificial intelligence model, and introducing the voice print into said artificial intelligence model to personalize it; • acquire a reproduction of the target voice at the output of said artificial intelligence model.
[0039] A computer program is further provided comprising instructions which cause the processing unit of the system as previously described to execute the steps of the personalized speech enhancement method as previously described.
[0040] A computer-readable recording medium is further provided, on which the computer program as previously described is recorded.
[0041] The invention will be better understood in light of the following description of particular non-limiting embodiments of the invention. Brief description of the drawings
[0042] Reference will be made to the attached drawings, among which:
[0043] [Fig-1] [Fig.l] represents a headset integrating the enhancement system personalized speech;
[0044] [Fig.2] [Fig.2] schematically represents the personalized speech enhancement system;
[0045] [Fig.3] [Fig.3] illustrates the generation of the voice print;
[0046] [Fig.4] [Fig.4] represents the transmission of the user's audio extract, and the reception of the voice print by the headset;
[0047] [Fig.5] [Fig.5] represents the artificial intelligence model used in the personalized speech enhancement system;
[0048] [Fig.6] [Fig.6] represents the model of [Fig.5] according to a first embodiment;
[0049] [Fig.7] [Fig.7] represents a concatenation operation;
[0050] [Fig.8] [Fig.8] represents the model of [Fig.5] according to a second mode of realization lization;
[0051] [Fig.9] [Fig.9] illustrates the generation of a voiceprint on the fly;
[0052] [Fig. 10] [Fig. 10] represents the implementation of a feedback loop to improve a voice print;
[0053] [Fig. 11] [Fig. 11] illustrates the constitution of the training data;
[0054] [Fig. 12] [Fig. 12] illustrates the general principle of the invention. DETAILED DESCRIPTION OF THE INVENTION
[0055] With reference to [Fig.l], a headset 1 conventionally comprises one or more microphones 2 which make it possible to capture the voice of the user 3 of the headset 1 during a telephone call. This or these microphones 2 are for example positioned on a microphone boom 4 or on the earphones 5 of the headset 1. The voice of the user (the speaker) is then transmitted to his remote interlocutor.
[0056] The headset 1 includes a personalized speech enhancement system 6. The system is here integrated into one of the earphones 5 of the headset 1.
[0057] With reference to [Fig.2], the system 6 comprises a first processing unit 7.
[0058] The first processing unit 7 is an electronic and software unit. The first processing unit 7 comprises one or more first processing components 8, and for example any processor or microprocessor, general or specialized (for example a DSP, for Digital Signal Processor, or a GPU, for Graphics Processing Unit, or even an NPU, for Neural Processing Unit), a microcontroller, or even a programmable logic circuit such as an FPGA (for Field Programmable Gate Arrays) or an ASIC (for Application Specific Integrated Circuit)•
[0059] The system 6 also comprises one or more first memories 9 (and in particular one or more non-volatile memories), connected to or integrated into the first processing component(s) 8. At least one of these first memories 9 forms a computer-readable recording medium, on which is recorded at least one computer program comprising instructions which cause the first processing unit 7 to execute the steps of the personalized speech enhancement method which will be described.
[0060] The microphones 2, here three in number, therefore capture the sound signals So present in the environment of the headset 1, and produce input audio signals Sae. The sound signals So, and therefore the input audio signals Sae, comprise a useful signal, in this case the user's voice, but also ambient noise Ba. By "ambient noise", we mean here any non-useful sound signal captured by the microphones 2, coming from any type of "parasitic" source. The ambient noise Ba therefore possibly comprises interfering voices, which are not those of the user.
[0061] The purpose of the personalized speech enhancement system 6 is to improve, for the user's interlocutor, the intelligibility of the user's voice during the telephone call. The system 6 is called "personalized" because it is adjusted and optimized, as will be seen, to reproduce the voice of this particular user. The user's voice is called the "target voice" Vc.
[0062] The personalized speech enhancement method comprises a preliminary phase and a phase carried out in real time. The preliminary phase is carried out in upstream of the real-time phase.
[0063] The preliminary phase consists first of all in acquiring an audio extract produced by microphones from a capture of the target voice of the user, then in generating a voice print from this audio extract.
[0064] The voice print makes it possible to characterize the target voice of the user and forms a voice representation of the user.
[0065] It is this voice print which will allow, during the real-time phase, to personalize the enhancement of speech.
[0066] With reference to Figures 3 and 4, the preliminary phase may use the headset 1 (but not necessarily), as well as another device (but not necessarily). Here, the preliminary phase uses the microphones 2 of the headset 1, and the user's smartphone 10, which is connected to the headset 1.
[0067] The smartphone 10, running a dedicated application, communicates with the user to implement the preliminary phase.
[0068] The smartphone 10 asks the user to speak. The user's voice is then captured by the microphones 2 of the headset 1, which therefore produce an audio extract Ea. This audio extract is transmitted by the headset 1 to the smartphone 10 on the application associated with the headset 1.
[0069] The smartphone 10 comprises a second processing unit 11 which, like the first processing unit 7 of the headset 1, comprises at least one second processing component. The smartphone 10 also comprises one or more second memories 12. At least one of these second memories 12 forms a computer-readable recording medium, on which at least one computer program (and in particular the application mentioned) is recorded.
[0070] The second processing unit 11 of the smartphone implements a voice encoder 14.
[0071] The voice encoder 14 that is used here is of the ECAPA-TDNN type (for Emphasized Channel Attention, Propagation, and Aggregation Time Delay Neural Network). This voice encoder 14 uses an artificial intelligence model. This model comprises layers implementing a time delay neural network, followed by an attention layer (attention pooling rent). The model has been previously trained. Here, the weights are fixed during its use in the enhancement process. The second processing unit 11 of the smartphone 10 executes an inference of this model.
[0072] The audio extract Ea is therefore applied as input to the voice encoder 14. The audio extract Ea forms an observation Nibs.
[0073] The observation is then transformed by the voice encoder 14 into a voice print Ev of the user.
[0074] This is also called a voice print, and we have:
[0075] xe = Encocleur(xobs),
[0076] where Encoder is the encoding function implemented by the voice encoder.
[0077] The voiceprint is such that:
[0078]
[0079] where D is the dimension of the latent space of the encoder model 14 and therefore of the voice print Ev.
[0080] The smartphone 10 checks the quality of the voice print Ev. If it is not satisfactory, the smartphone asks the user to re-record his voice. If the quality of the voice print Ev is satisfactory, the smartphone 10 transmits the voice print Ev to the headset 1, which is recorded in one of the first memories 9 of the system 6.
[0081] The preliminary phase is completed. We note here: - that the capture of the user's voice, to produce the audio extract, could be done by equipment other than the headset 1 (by the smartphone 10 for example, or by a computer 15); - that the voice encoder 14 could be implemented in equipment other than the smartphone 10, in the computer 15 for example or in the headset 1 itself. The entire preliminary phase could therefore be carried out by the headset 1. The first processing unit 7 could therefore implement the voice encoder and produce the voice print.
[0082] The personalized speech enhancement functionality can then be used in real time.
[0083] We are now interested in the real-time phase.
[0084] The user is making a telephone call. The microphones 2 therefore capture in real time sound signals So which include the target voice Vc of the user and the ambient noise Ba which prevails in the user's environment (and which may include interfering voices).
[0085] The first processing unit 7 of the headset 1 comprises, for each microphone 2, a buffer 17 and a calculation block 18.
[0086] The first processing unit 7 therefore acquires, for each microphone 2, the input audio signal Sae produced by said microphone from the sound signal So. The input audio signals Sae therefore comprise the target voice Vc of the user and the ambient noise B a.
[0087] Each calculation block 18 then calculates the Short Term Fourier Transform (TFCT, or STFT for Short Time Fourier Transform) of the input audio signal Sae.
[0088] The output of each calculation block 18 therefore represents the spectrum of an audio signal entrance.
[0089] The first processing unit 7 then produces, from the input audio signals Sae, a current audio signal Sac which is a frequency domain signal and which is a spectral representation of the input audio signals.
[0090] The first processing unit 7 of the headset comprises at least one pre-processing block, at the input of which the current audio signal is applied, to produce at least one characterizing vector from the current audio signal and therefore from the input audio signals.
[0091] Here, the first processing unit 7 comprises a first pre-processing block 20a and a second pre-processing block 20b.
[0092] The first pre-processing block 20a generates a first vector characterizing Vcl. The second pre-processing block 20b generates a second vector characterizing Vc2.
[0093] The first vector characterizing Vcl is obtained here by estimating the ERB (Equivalent Rectangular Bandwidth, which can be translated as “equivalent rectangular bandwidth”) of the current audio signal.
[0094] The second vector characterizing Vc2 is here obtained by estimating the complex spectrogram of the current audio signal Sac. The second vector characterizing Vc2 is obtained by converting the TFCT to a logarithmic scale.
[0095] We therefore obtain:
[0096] ERB, Spec = Preprocessing^x^,
[0097] where xa is the current audio signal, ERB is the first characterizing vector, Spec is the second characterizing vector, and Preprocessing is the preprocessing function implemented by the preprocessing blocks 20a, 20b.
[0098] Then, the first processing unit 7 executes an inference of a previously trained artificial intelligence model 21 for speech enhancement, by applying the first vector characterizing Vcl and the second vector characterizing Vc2 as input to the model 21.
[0099] The inference model (and therefore, in particular, its weights) is stored in one of the first memories 9 of the system 6 of the headset 1.
[0100] The model 21 used here is based on the U-Net network and, more specifically, on the DeepFilterNet2 model which has a so-called "cascade" or "dual-stage" architecture. This type of "cascade" architecture was developed for standard speech enhancement and has greatly improved enhancement performance. It is therefore a very efficient architecture.
[0101] We first describe a model according to a first embodiment, with reference to [Fig.5]. This model 21 is a two-stage model (dual-stage architecture): 21a and 21b
[0102] Model 21 includes: • an encoder block 23, at the input of which is applied at least one characterizing vector (and therefore here the first characterizing vector and the second characterizing vector), and generating at least one output vector; • at least one decoder 24, at the input of which is applied the at least one output vector, and generating a reproduction of the target voice.
[0103] As we will see, the voice print Ev is also applied as input to the encoder block 23.
[0104] Model 21 also includes a separator.
[0105] The encoder block 23 makes it possible to analyze the characterizing vectors by relating the frequency information. The encoder 23 comprises convolution layers making it possible in particular to reduce the size of the characterizing vectors by compressing the information and obtaining new characteristics in a latent space. These characteristics are then used by the separator, whose role is to separate the voice from the rest in the latent space. The separator is made up of recurrent networks in order to relate the temporal information. Finally, the decoder makes it possible to transform the information from the latent space to the space of the input characteristics (and therefore from the first characterizing vector and the second characterizing vector).
[0106] More precisely, with reference to [Fig.6], the encoder block 23 comprises a first branch 23a at the input of which the first vector characterizing Vcl (ERB) is applied, and a second branch 23b at the input of which the second vector characterizing Vc2 (complex spectrogram) is applied.
[0107] The first branch 23a belongs to the first stage 21a of the model 21. The second branch 23b belongs to the second stage 21b of the model 21.
[0108] The first branch 23a of the encoder block 23 successively comprises a first convolution layer 25, a second convolution layer 26, a third convolution layer 27 and a fourth convolution layer 28.
[0109] The second branch 23b of the encoder block 23 successively comprises a first convolution layer 29, a second convolution layer 30 and a grouped linear layer 31 (GLineaf).
[0110] The encoder block 23 generates at least one intermediate vector and comprises at least one operator 32 arranged to merge the at least one intermediate vector and the voice print Ev to produce the at least one output vector.
[0111] Here, the fourth convolution layer 28 of the first branch 23a produces a first intermediate vector Vil. The grouped linear layer 31 of the second branch 23b produces a second intermediate vector Vi2.
[0112] The first intermediate vector Vil is:
[0113] XERB = EncERB(ERS),
[0114] where EncERB is the function implemented by the first branch 23a of the encoder block 23.
[0115] The second intermediate vector Vi2 is:
[0116] xSpec = Em:Spec (Spec),
[0117] where EflCSpec is the function implemented by the second branch 23b of the encoder block 23.
[0118] The at least one operator 32 comprises a concatenation operator arranged to concatenate along a frequency axis at least one intermediate vector and the voice print.
[0119] The concatenation operator 32 therefore merges by concatenation the first intermediate vector Vil, the second intermediate vector Vi2 and the voice print Ev.
[0120] [Fig.7] illustrates the concatenation.
[0121] The concatenation is carried out in the following manner. The voice print Ev is duplicated along the time axis in order to obtain as many occurrences as there are frames in the intermediate vectors (after passing through the first layers of the encoder), then each voice print Ev is concatenated with each frame of the two intermediate vectors.
[0122] In [Fig.7], all frames of the intermediate vectors are different, while all frames of the voiceprint are equal. The unified architecture of the model in [Fig.5] and [Fig.6] therefore corresponds to the addition of the voiceprint Ev in the encoder block, thus making it possible to combine together the ERBs, the complex spectrograms and the information on the target voice of the speaker.
[0123] The concatenation operator 32 therefore generates the concatenated vector Vc:
[0124] Conca^XER^ XSpe^ xj
[0125] The encoder block 23 further comprises a grouped linear layer 33 at the input of which the output of the concatenation operator 32 is applied.
[0126] The separator, which is here integrated into the encoder block 23, comprises a closed recurrent unit 34 (GRU layer, for Gated Recurrent Unit).
[0127] The closed recurrent unit 34 therefore produces an output vector Vs from the concatenated vector Vc, and therefore from the first intermediate vector Vil, the second intermediate vector Vi2 and the voice print Ev.
[0128] The output vector Vs is:
[0129] xEnc= GRUiConca^X^
[0130] where GRU is the closed recurrent unit function 34.
[0131] The at least one decoder 24 here comprises a first decoder 24a and a second decoder 24b.
[0132] The first decoder 24a belongs to the first stage 21a of the model 21. The second decoder 24b belongs to the second stage 21b of the model 21. The convolution layers of the first branch 23a of the encoder block 23 are connected to the first decoder 24a.
[0133] The output vector Vs is applied to the input of the first decoder 24a and to the input of the second decoder 24b.
[0134] The first decoder 24a generates from the output vector Gerb gains called “ERB gains”:
[0135] Gekb= DecE^XEJ
[0136] These gains are applied to the first vector characterizing Vcl.
[0137] We therefore obtain:
[0138] XG = ERB x
[0139] The second decoder 14b, for its part, generates filtering coefficients Cdf of the deep filter 36 (Deep Filtering):
[0140] CDF= Dee^X^)
[0141] The first vector characterizing Vcl, to which the ERB gains have been applied, is then applied to the input of the deep filter 36:
[0142] Xdf = DF(Xg | CDF).
[0143] The first stage 21a of the model 21 works in amplitude and therefore makes it possible to estimate an envelope of the target voice. It performs a coarse, imprecise analysis, producing an imprecise reproduction of the target voice. The first stage 21a produces real-value gains. The second stage 21b performs a more precise analysis. The second stage 21b operates in the complex domain and implements deep filtering, using the results of the first stage (the gains), which makes it possible to obtain denoised content forming the reproduction Rv of the target voice, which is a more precise reproduction than that of the first stage.
[0144] This two-stage structure allows the model to be compacted while maintaining its efficiency.
[0145] The second stage 21b operates on the lower part of the spectrogram (here up to the frequency f5 kHz).
[0146] Note that a mask is generated and applied to the output of the deep filter 36 in order to mask the noise.
[0147] An inverse Fourier transform 37 is then calculated to obtain the reproduction of the target voice Rv.
[0148] The first processing unit 7 of the headset 1 then acquires the reproduction Rv of the target voice and transmits it to the user's interlocutor, who thus receives the enhanced voice of the user.
[0149] We speak of a “unified” encoder block to designate the encoder block 23 of the first embodiment. Indeed, the two branches 23a, 23b of the encoder 23 merge within it via the concatenation operator 32.
[0150] We are now interested, with reference to [Fig.8], in a model 40 according to a second embodiment and, more precisely, in a first alternative of this second embodiment.
[0151] The model 40 again comprises an encoder block 41 and at least one decoder 42, in this case a first decoder 42a and a second decoder 42b.
[0152] The encoder block 41 comprises a first branch 41a and a second branch 41b.
[0153] The first branch 41a of the encoder block 41 successively comprises a first convolution layer 43, a second convolution layer 44, a third convolution layer 45, a fourth convolution layer 46, then a first concatenation operator 47, a grouped linear layer 48 (GLinear) and a first closed recurrent unit 49.
[0154] The second branch 41b of the encoder block 41 successively comprises a first convolution layer 50, a second convolution layer 51, a first grouped linear layer 52 (GLineaf), then a second concatenation operator 53, a second grouped linear layer 54 and a second closed recurrent unit 55.
[0155] The first concatenation operator 47 concatenates along the frequency axis a first intermediate vector Vcl (output of the fourth convolution layer 46) produced by the first branch of the encoder block, and the voice print Ev. The second concatenation operator 53 concatenates along the frequency axis a second intermediate vector Vc2 (output of the third convolution layer 52) produced by the second branch 41b of the encoder block 41, and the voice print Ev.
[0156] The first closed recurrent unit 49 generates a first output vector Vsl which is applied as input to the first decoder 42a. The second closed recurrent unit 55 generates a second output vector Vs2 which is applied as input to the second decoder 42b.
[0157] In this first alternative of the second embodiment of the model 40, the voice print Ev is therefore integrated into each of the branches of the encoder block.
[0158] We speak of a “double” or “dual” encoder block to designate the encoder block of the second embodiment. Indeed, the two branches of the encoder do not merge within it but are independent and each connected to a separate decoder.
[0159] More generally, the first branch 41a of the encoder block comprises a first operator 47 arranged to merge a first intermediate vector produced by the first branch of the encoder block and the voice print, and / or the second branch 41b of the encoder block comprises a second operator 53 arranged to merge a second intermediate vector produced by the second branch of the encoder block and the voice print.
[0160] It would therefore also be possible, according to a second alternative of the second embodiment of the model, to integrate the voice print only into the first branch 41a of the encoder block 41 (which alone then includes a concatenation operator). It would also be possible, according to a third alternative of the second embodiment, to integrate the voice print only into the second branch 41b of the encoder block 41 (which alone then includes a concatenation operator).
[0161] We are now again interested in the generation of the voice print Ev.
[0162] It is possible to improve the user's voice print in real time. This improvement process is implemented in the headset.
[0163] For this, with reference to [Fig.9], the first processing unit 7 comprises an on-the-fly fingerprint generation block 60, comprising a fingerprint generation block 61 and a similarity calculation block 62.
[0164] The first processing unit 7 acquires the at least one input audio signal Sae of the user which is applied as input to the fingerprint generation block 61. The fingerprint generation block 61 generates from the at least one input audio signal a current voiceprint Eve. Then, the calculation block 62 calculates a similarity value representative of a similarity between the current voiceprint Eve and the voiceprint Ev in memory, which was produced during the preliminary phase.
[0165] This similarity value is here a distance between two vectors: the first vector is the memory fingerprint and the second vector is the current voiceprint. The distance has a value which is for example between "0" and "1", where "0" means "completely dissimilar" and "1" means "identical".
[0166] If the two voiceprints are sufficiently similar (i.e. if the similarity value is greater than a predefined threshold), the first processing unit 7 generates a new voiceprint En which is formed from the voiceprint Ev stored in the memory, the current voiceprint Eve and the similarity value. This new voiceprint En is introduced into the model 21 (or 40).
[0167] This method has several advantages. First of all, it makes it possible to obtain an increasingly robust voice print and to overcome the problem of domain adaptation. Indeed, by changing the sound environment, the nature of the voice print may change slightly. Here, this voice print is updated in real time, which eliminates this problem. This method also makes it possible to overcome the difficulty relating to temporary changes in voice, for example if the speaker has a broken, hoarse or other voice. This method therefore makes it possible to adapt to this type of situation.
[0168] The generation of the voice print on the fly also makes it possible, as this print is regularly updated, to produce a less precise voice print during the preliminary phase, thus using a less cumbersome voice encoder. The reduction in precision is compensated by the real-time improvement of the print. This can make it possible to carry out the voice encoding of the preliminary phase on board, and therefore in the first processing unit 7 of the headset 1. Interaction with the smartphone (or any other equipment) is then no longer necessary.
[0169] Alternatively, with reference to [Fig. 10], the first processing unit 7 may implement a feedback loop 64 to generate the voice print on the fly.
[0170] The first processing unit 7 executes a first inference of the model 21, 40 using the voice print Ev stored in memory. The model produces a reproduction of the target voice (enhanced voice).
[0171] The first processing unit 7 again comprises an on-the-fly fingerprint generation block 65, comprising a fingerprint generation block 66 and a similarity calculation block 67. The block 66 generates a new voice print En from the reproduction Rv of the target voice, then the calculation block 67 calculates the similarity between the voice print Ev in memory and the new voice print En.
[0172] If the similarity value is greater than a predefined threshold, then the block 65 takes into account this new voice print Ev, which is then integrated into the model 21, 40. The first processing unit 7 executes another inference of the model, this time using the new voice print En. The reproduction Rv of the target voice is then again applied as input to the on-the-fly print generation block 65.
[0173] We are now interested, with reference to [Fig. 11], in the constitution of the training data.
[0174] As seen, the model used is a speech enhancement model that has been transformed to be personalized according to the user (more precisely, according to the user's voice).
[0175] Three datasets constitute the training data used to train this model.
[0176] A first data set 71 comprises voice prints obtained from audio extracts of identified speakers. The audio extracts here come from a first data corpus 72, for example the LibriSpeech corpus. For each identified speaker 73, an audio extract 75 was taken, then this audio extract 75 was encoded by the voice encoder 14 previously described. The first data set 71 is therefore obtained, which therefore comprises voice prints of identified speakers.
[0177] As seen, the model must be able to effectively isolate a target voice in an environment including noise and interfering voices. We therefore used a second corpus of data 76 containing audio extracts of unidentified speakers, to produce interfering voices 77. The second corpus of data 76 is for example from LibriSpeech and Mozilla Common Voice.
[0178] A third corpus of data 78 containing noise extracts was also used to obtain audio noise 79. The third corpus of data 78 here includes data from the DNS challenge (themselves coming from the Freesound and AudioSet databases).
[0179] An audio mixer 80 was used to mix the interfering voices 77, the noise 79, and audio clips of identified speakers to form the second data set 81. The audio clips of identified speakers, although relating to the same speakers, are different from those used to form the first data set 71.
[0180] Finally, a third data set 82 was generated from voice extracts of the identified speakers. Again, these audio extracts, although concerning the same speakers, are different from those used to form the first data set 71.
[0181] The model is therefore trained using as inputs the data from the second data set 81. For each data item from the second data set 81 corresponding to the same speaker, the voice print of said speaker from the first data set 71 is introduced into the model. The model output, i.e. the reproduction of the target voice, is compared with the audio extract from the third data set 82 to check whether the reproduction of the target voice is correct, then the weights of the model are updated (by gradient backpropagation for example).
[0182] Approximately 2400 different speakers were used for the target voices and approximately 7000 unwanted voices, resulting in over 800 hours of data.
[0183] A combination of cost functions (or loss functions) was used for training: compressed spectral loss, multi-resolution loss, and over-suppression loss.
[0184] The compressed spectral loss corresponds to the sum of two terms. The first term is the mean square error of the spectrogram amplitude compressed by a factor c, and the second is the same term but taking into account the phase.
[0185] The compressed spectral loss can be written: 101861 L,Il |Yf-14 II 2 + II |Yf> y -14^'' He 2
[0187] Multi-resolution loss supports voice reconstruction. This involves calculating a cost function for multiple TFCTs generated with different window sizes (e.g., 5ms, 10ms, 20ms, and 40ms). For each TFCT, the mean square error of the compressed amplitude is multiplied with the mean square error of the compressed complex TFCT.
[0188] The multi-resolution loss can be written:
[0189] = Il M' -MVx II H> y - II
[0190] The loss of over-deletion is specific to customization and is intended to avoid suppress the target voice too much. It simply penalizes the network when it misses pieces of the target voice in the reconstructed spectrogram.
[0191]
[0192]
[0193] The over-suppression loss can be written: 0, 1^. / )01 2
[0194] The overall cost function is therefore written:
[0195] L = AspecLspec + ^MR^MR + ^OS^OS,
[0196] where Lspec is the compressed spectral loss, LMR is the multi-resolution loss, Los is the over-suppression loss, and where Àspec, ^mr, are the loss weights defined during training.
[0197] The results obtained for the four models presented are presented. Here, we call “first model” the model according to the first embodiment, “second model” the model according to the first alternative of the second embodiment, “third model” the model according to the second alternative of the second embodiment. embodiment, and “fourth model” the model according to the third alternative of the second embodiment.
[0198] Reference is made to the two tables in the Appendix: Table 1 and Table 2.
[0199] The enhancement results obtained are grouped in Table 1. The Results were obtained on a "synthetic" test base, which was made by the inventors, and whose characteristics are grouped on the "raw test base" line. For all the metrics used, a higher result means better performance. The metrics used here are metrics known to those skilled in the art: PESQ, STOI, CSIG, CB AK and CO VL.
[0200] It is observed that the different customized models (first, second, third and fourth models) improve the performances well and clearly exceed those of the original DeepFilterNet2 model. The first model, with unified encoder block, is the one which presents the best results.
[0201] Furthermore, as we have seen, the size of the model used is critical because the inference is carried out on-board. We therefore adapted a compact model (DeepFilterNet2') to the customization while impacting its complexity as little as possible.
[0202] Table 2 allows us to compare the first model (the most efficient) with some prior art networks on a computational level. This table shows the number of parameters and the number of MAC operations (for Multiply-Accumulât™ e Operation, which can be translated as multiplication-accumulation operation) associated with the different models. The lower the value obtained, the more favorable it is.
[0203] Table 2 also specifies whether or not each model is a cascade model (dual-stage).
[0204] Two main lessons can be drawn from this table. First, the customization method used did not increase the computational load of the model. The model was customized without increasing its complexity. Second, we note that the model used is significantly smaller, both in number of parameters and in MAC. This very compact model can therefore be implemented on an embedded product, which is not the case for other state-of-the-art models. We also note that the closest network, E3Net, does not correspond to a cascaded architecture (dual-stage).
[0205] With reference to [Fig. 12], the method described therefore consists, during a preliminary phase, in applying an audio extract Ea as input to a voice encoder 14 to obtain a voice print Ev of the target voice. The voice print Ev is stored in the memory of the embedded product.
[0206] Then, in real time, the first processing unit acquires at least one input audio signal Sae, which is a mixture between the target voice Vc and an ambient noise Ba comprising noises B and one or more interfering voices Vint. The model 21, 40 then generates a reproduction Rv of the target voice.
[0207] This method, more complex than simple noise reduction, makes it possible to greatly improve the clarity of a call in the presence of several interlocutors.
[0208] It has been described here that the target voice is that of the headset user.
[0209] However, the system and method may also be used to enhance a target voice that is that of the user's interlocutor. There may be multiple target voices, from multiple interlocutors.
[0210] The method is similar to that used when one wishes to isolate the user's voice, but in the opposite direction. To do this, an audio extract of the voice of frequently called people is required.
[0211] The method comprises the following steps.
[0212] Frequent callers record their voice.
[0213] From these recordings, a voice print is generated for each speaker and is stored in the user's headset.
[0214] During a call, if the interlocutor is identified, the incoming audio stream is retrieved and transmitted to the neural network, previously trained on the dataset presented above. Again, the stream is divided into overlapping windows and each window is applied as input to the neural network. The voiceprint of the interlocutor is introduced into the template to personalize it. The reproduction of the target voice is then transmitted to the headset user.
[0215] Alternatively, for each window, the audio signal can be compared to known voiceprints to identify the target voice. The corresponding voiceprint is then used to personalize the model.
[0216] Similarly, within the same household for example, the voice prints of several interlocutors can be generated and stored. When the voice of one of these interlocutors is recognized, the personalized enhancement process is carried out using the voice print of said interlocutor.
[0217] Generating and storing multiple voiceprints of multiple speakers allows for personalization for multiple speakers at the same time. For example, when a conference speakerphone is used, everyone in the room records their voice so that the system can perform personalized enhancement for each voice.
[0218] The system and method can also be used in a face-to-face conversation application, and for example in a hearing aid audio device.
[0219] The target voice may be that of the interlocutor of the user of the audio device.
[0220] The target voice may also be that of the user of the audio device. The method allows the reproduction of the target voice to be obtained, which is then removed from the captured signal.
[0221] OWS (Open Wearable Stereo) or open-ear type headphones have the disadvantage that, in the case of directional enhancement, the user's voice will also be enhanced with an echo due to processing latency, which will be counterproductive.
[0222] It may therefore be interesting to enhance the useful signal (voice of the speakers opposite) by removing interfering noise (background noise) and removing the user's own voice from the device in order to avoid an echo phenomenon.
[0223] Of course, the invention is not limited to the embodiments described but encompasses any variant falling within the scope of the invention as defined by the claims.
[0224] The voice encoder used to generate the voice print may be different from that described here. The voice encoding could for example be a representation in the form of “MFCC” coefficients (for Mel-F requency Cepstrum Coefficients).
[0225] A first characterizing vector and a second characterizing vector are used, but there could be only one characterizing vector, or a different number of characterizing vectors. The at least one characterizing vector can be obtained from different representations of the at least one input audio signal.
[0226] The input audio signal(s), or the current audio signal, can also be introduced as is, without processing, into the model. The at least one vector ca characterizing can then be the current audio signal itself, or the at least one input audio signal itself (or a combination of the input audio signals). Passing through an intermediate frequency representation is not obligatory.
[0227] It has been described that the speech enhancement artificial intelligence model has a structure that is based on the structure of a U-Net network and, more specifically, that the structure of the model is based on the structure of a DeepFilterNet2 network. This is not mandatory, however. Any speech enhancement model can be used and customized by introducing the voiceprint into said model. The model is not necessarily an encoder-decoder model and, in an encoder-decoder model, the introduction of the voiceprint into the model is not necessarily performed in the encoder at the output thereof.
[0228] The introduction of the voiceprint can be done by using an operator arranged to merge one or more intermediate vectors and the voiceprint. This fusion is not necessarily a concatenation, it could be an addition, a multiplication, a cross-attention mechanism, etc.
[0229] The personalized speech enhancement system can be integrated into a portable audio device other than a headset: earphones, hearing aid, tablet, laptop, smartphone, etc. The personalized speech enhancement system can also be integrated into a non-portable device: videoconferencing system, hands-free car kit, virtual personal assistant, any device equipped with a voice recognition module, etc. The personalized speech enhancement system can also be integrated with existing software, in particular videoconferencing. APPENDIX Model PESQ STOI CSIG CBAK COVL Raw test base 1.81 0.75 2.99 2.45 2.37 DeepFilterNet2 2.10 0.75 3.11 2.66 2.58 First model 2.30 0.78 3.59 2.86 2.94 Second model 2.24 0.78 3.50 2.79 2.86 Third model 2.23 0.76 3.50 2.76 2.85 Fourth model 2.18 0.77 3.43 2.76 2.79
[0231] [Tables 1] Dual-stage model Params (M) MAC (G) DeepFilterNet2 y es 2.31 0.36 First model yes 2.31 0.36 TEA-PSE 3.0 yes 22.24 19.66 NPU ELEVOC yes 12.49 8.50 pBSRNN yes 5.91 5.54 pPercepNet no 8.5 - E3Net no 4.5 -
[0232] Table 2
Claims
Claims
1. System (6) for personalized speech enhancement, comprising: • at least one memory (9), in which are stored a voice print (Ev) of a speaker, previously recorded, and a previously trained artificial intelligence model (21; 40) for speech enhancement; • a processing unit (7) arranged to, in real time: • acquire at least one input audio signal (Sae) comprising a target voice (Vc) of the speaker and ambient noise; • produce at least one characterizing vector (Vcl, Vc2) from the at least one input audio signal; • execute an inference of the artificial intelligence model, by applying the at least one characterizing vector as input to said artificial intelligence model, and by introducing the voice print into said artificial intelligence model to personalize it;• acquire a reproduction (Rv) of the target voice at the output of said artificial intelligence model.;
2. System according to claim 1, in which the artificial intelligence model comprises: - an encoder block (23; 41), at the input of which is applied the at least one characterizing vector, and generating at least one output vector; - at least one decoder (24; 42), at the input of which is applied the at least one output vector; the encoder block generating at least one intermediate vector and comprising at least one operator (32; 47, 53) arranged to merge the at least one intermediate vector and the voice print to produce the at least one output vector.
3. A system according to one of the preceding claims, wherein the artificial intelligence model comprises a first stage performing a low-precision analysis producing a low-precision reproduction of the target voice, and a second stage using results of the first stage and performing a more precise analysis to obtain the reproduction (Rv) of the target voice.
4. System according to claims 2 and 3, wherein the processing unit (7) produces a first characterizing vector (Vcl) and a second characterizing vector (Vc2) from the input audio signal, and wherein: - the encoder block comprises a first branch (23a; 41a) belonging to the first stage and at the input of which the first characterizing vector is applied, and a second branch (23b; 41b) belonging to the second stage and at the input of which the second characterizing vector is applied; - the model comprises a first decoder (24a; 42a) belonging to the first stage and a second decoder (24b; 42b) belonging to the second stage.
5. System according to claim 4, wherein the operator (32) is arranged to merge a first intermediate vector (Vil) produced by the first branch (23a) of the encoder block, a second intermediate vector (Vi2) produced by the second branch (23b) of the encoder block, and the voice print (Ev).
6. System according to claim 4, in which the first branch (41a) of the encoder block comprises a first operator (47) arranged to merge a first intermediate vector produced by the first branch of the encoder block and the voice print, and / or the second branch (41b) of the encoder block comprises a second operator (53) arranged to merge a second intermediate vector produced by the second branch of the encoder block and the voice print, the first branch producing a first output vector (Vsl) applied as input to the first decoder (42a) and the second branch producing a second output vector (Vs2) applied as input to the second decoder (42b).
7. System according to one of claims 4 to 6, in which the processing unit produces, from the at least one input audio signal (Sae), a current audio signal (Sac) which is a frequency domain signal, the first characterizing vector (Vcl) being obtained by calculating an ERB of the current audio signal, and the second characterizing vector (Vc2) being obtained by calculating a complex spectrogram of the current audio signal.
8. System according to one of claims 2 to 7, in which the at least one operator (32; 47, 53) comprises at least one conca- operator tenation arranged to concatenate along a frequency axis the at least one intermediate vector and the voice print.
9. System according to one of claims 2 to 8, wherein a structure of the model is based on a structure of a U-Net network.
10. The system of claim 9, wherein the structure of the model is based on a structure of a DeepFilterNet2 network.
11. System according to one of the preceding claims, the voice print (Ev) being obtained using a voice encoder of the ECAPA-TDNN type.
12. System according to one of the preceding claims, in which the voice print (Ev) is generated by the processing unit (7) of the system (6).
13. System according to one of the preceding claims, in which the processing unit (7) is arranged to generate, from the at least one incoming audio signal, a current voiceprint (Eve), to calculate a similarity value representative of a similarity between the current voiceprint (Eve) and the voiceprint (Ev) stored in the memory, and to introduce into the model a new voiceprint formed from the voiceprint stored in the memory, the current voiceprint and the similarity value.
14. System according to one of claims 1 to 12, in which the processing unit (7) is arranged to implement a feedback loop (64) arranged to generate, from the reproduction (Rv) of the target voice at the output of the model, a new voice print (En), and to execute another inference of the model this time using the new voice print.
15. System according to one of the preceding claims, the speaker being a user of the system, who communicates with a remote interlocutor.
16. System according to one of claims 1 to 14, the speaker being an individual who communicates with the user of the system.
17. Audio apparatus (1) incorporating a system according to one of the preceding claims.
18. Method for personalized speech enhancement, implemented in the processing unit (7) of the system (6) according to one of claims 1 to 16, and comprising the steps, implemented in real time, of: • acquiring at least one input audio signal (Sae) comprising a target voice (Vc) of the speaker, and ambient noise; • produce at least one characterizing vector (Vcl, Vc2) from the input audio signal; • execute an inference of the artificial intelligence model, by applying the at least one characterizing vector as input to said artificial intelligence model, and by introducing the voice print into said artificial intelligence model to personalize it; • acquire a reproduction (Rv) of the target voice as output from said artificial intelligence model.
19. Computer program comprising instructions which cause the processing unit (7) of the system (6) according to one of claims 1 to 16 to execute the steps of the personalized speech enhancement method according to claim 18.
20. A computer-readable recording medium on which the computer program according to claim 19 is recorded.
Citation Information
Patent Citations
Voice quality enhancement method and related device
US20240096343A1