Personalisation of a neural network for speech enhancement

A personalized speech enhancement system using a compact AI model with dual-stage U-Net architecture addresses the challenge of overlapping voices in noisy environments, enhancing target voices efficiently in portable devices.

WO2025210186A1PCT designated stage Publication Date: 2025-10-09OROSOUND +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/059185
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-05
Filing Date
2025-04-03
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Neural networks used for speech enhancement perform poorly in environments with overlapping voices and require significant computing power, making them unsuitable for real-time implementation in portable devices.

Method used

A personalized speech enhancement system that includes a reference voiceprint and a compact AI model, utilizing a U-Net network architecture with dual stages for efficient voice isolation and denoising, adapted for real-time operation with limited computing resources.

Benefits of technology

The system effectively enhances target voices in noisy environments with overlapping voices, requiring minimal computing power and energy, suitable for embedded systems like headsets and earphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025059185_09102025_PF_FP_ABST
    Figure EP2025059185_09102025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a personalised speech enhancement system arranged, in real time, to: o acquire at least one input audio signal (Sae) comprising a target voice (Vc) of the speaker and the noise; o produce at least one characterising vector (Vc1, Vc2) on the basis of the at least one input audio signal; o perform inference of a speech enhancement artificial intelligence model by applying the at least one characterising vector as input to the model, and by introducing a voiceprint of the speaker into the model in order to personalise it; o obtain a reproduction (Rv) of the target voice as the output of the artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Personalization of a Neural Network for Speech Enhancement DESCRIPTION

[0002] The invention relates to the field of speech enhancement methods that use an artificial intelligence model.

[0003] BACKGROUND OF THE INVENTION

[0004] Speech enhancement (or SE) is a feature that isolates a voice in a noisy sound environment, to reproduce this voice while improving its intelligibility.

[0005] This is a feature that is very popular in audio technologies today, particularly due to the increase in video conferencing and teleworking.

[0006] The implementation of this functionality in portable audio devices is particularly relevant.

[0007] For example, in the case of a headset, it is very interesting to be able to isolate the voice of the headset user during a telephone call, and to transmit it enhanced and denoised to the user's interlocutor.

[0008] Neural networks are currently the algorithms that achieve the best performance for speech enhancement.

[0009] However, neural networks exhibit degraded performance and generally perform poorly in environments where interfering voices overlap with the user's voice. In addition to degrading call clarity, this also poses a privacy issue as background conversations (e.g., private conversations) may be transmitted during the call. Therefore, Personalized Speech Enhancement (PSE) is being considered, which aims to extract a predefined "target voice" belonging to an identified speaker from a noisy environment including interfering voices.

[0010] In the example application just given, the target voice is that of the headset user.

[0011] In the case of the headset and, more generally, any portable audio device, the personalized speech enhancement process requires “on-board” processing to enable the reproduction of the target voice in real time.

[0012] Traditionally known methods were relatively ineffective. However, recently, the Deep Noise Suppression (DNS) challenge organized a challenge corresponding to the task of personalizing speech enhancement. This challenge has given rise to a lot of work on this subject, which has led to the design of relatively efficient architectures. However, even if these architectures are considered "real-time" because they are causal, they are often very heavy and cannot be implemented in an embedded product.

[0013] SUBJECT OF THE INVENTION

[0014] The invention relates to a personalized speech enhancement system which is effective and whose implementation requires limited computing power and energy consumption.

[0015] SUMMARY OF THE INVENTION

[0016] To achieve this goal, a personalized speech enhancement system is proposed, comprising: - at least one memory, in which are stored a reference voice print of a speaker, previously recorded, and a previously trained artificial intelligence speech enhancement model;

[0017] - a processing unit arranged to, in real time: o acquire at least one input audio signal comprising a target voice of the speaker and ambient noise; o generate, from the at least one input audio signal, a current voiceprint, calculate a similarity value representative of a similarity between the current voiceprint and the reference voiceprint, and adapt the reference voiceprint as a function of the current voiceprint and the similarity value; o produce at least one characterizing vector from the at least one input audio signal; o execute an inference of the artificial intelligence model, by applying the at least one characterizing vector as input to said artificial intelligence model, and by introducing the reference voiceprint into said artificial intelligence model to personalize it; o acquire a reproduction of the target voice as output from said artificial intelligence model.

[0018] The introduction of the speaker's reference voiceprint into the speech enhancement artificial intelligence model, when performing its inference, allows for personalized speech enhancement. This adapts the model to the speaker's vocal characteristics. This significantly increases the effectiveness of target voice enhancement, particularly when the ambient noise includes interfering voices. By using a compact model, we therefore obtain a personalized speech enhancement system that is very effective, and which requires limited computing power and energy consumption. The personalized speech enhancement system can thus be implemented in embedded systems, particularly in portable audio devices (headsets, earphones, hearing aids, etc.).

[0019] We further propose a system as previously described, in which the artificial intelligence model comprises:

[0020] - an encoder block, at the input of which at least one characterizing vector is applied, and generating at least one output vector;

[0021] - at least one decoder, at the input of which is applied the at least one output vector; the encoder block generating at least one intermediate vector and comprising at least one operator arranged to merge the at least one intermediate vector and the reference voice print to produce the at least one output vector.

[0022] Further provided is a system as previously described, wherein the artificial intelligence model comprises a first stage performing a low-precision analysis producing a low-precision reproduction of the target voice, and a second stage using results from the first stage and performing a more precise analysis to obtain the reproduction of the target voice.

[0023] We further propose a system as previously described, in which the processing unit produces a first characterizing vector and a second characterizing vector from the input audio signal, and in which: the encoder block comprises a first branch belonging to the first stage and at the input of which the first characterizing vector is applied, and a second branch belonging to the second stage and at the input of which the second characterizing vector is applied;

[0024] - the model includes a first decoder belonging to the first stage and a second decoder belonging to the second stage.

[0025] We further propose a system as previously described, in which the operator is arranged to merge a first intermediate vector produced by the first branch of the encoder block, a second intermediate vector produced by the second branch of the encoder block, and the reference voiceprint.

[0026] We further propose a system as previously described, in which the first branch of the encoder block comprises a first operator arranged to merge a first intermediate vector produced by the first branch of the encoder block and the reference voiceprint, and / or the second branch of the encoder block comprises a second operator arranged to merge a second intermediate vector produced by the second branch of the encoder block and the reference voiceprint, the first branch producing a first output vector applied as input to the first decoder and the second branch producing a second output vector applied as input to the second decoder.

[0027] We further propose a system as previously described, in which the processing unit produces, from the at least one input audio signal, a current audio signal which is a frequency domain signal, the first characterizing vector being obtained by calculating an ERB of the current audio signal, and the second characterizing vector being obtained by calculating a complex spectrogram of the current audio signal.

[0028] We further propose a system as previously described, in which the at least one operator comprises at least one concatenation operator arranged to concatenate along a frequency axis the at least one intermediate vector and the reference voiceprint.

[0029] We further propose a system as previously described, in which a structure of the model is based on a structure of a U-Net network.

[0030] We further propose a system as previously described, in which the structure of the model is based on a structure of a DeepFilterNet2 network.

[0031] We further propose a system as previously described, an initial reference voice print being obtained using an ECAPA-TDNN type voice encoder.

[0032] Further provided is a system as previously described, wherein the initial reference voiceprint is generated by the processing unit of the system.

[0033] We further propose a system as previously described, in which the processing unit is arranged to implement a feedback loop arranged to generate, from the reproduction of the target voice at the output of the model, a new reference voiceprint, and to execute another inference of the model this time using the new reference voiceprint.

[0034] We further propose a system as previously described, the speaker being a user of the system, who communicates with a remote interlocutor.

[0035] We further propose a system as previously described, the speaker being an individual who communicates with the user of the system.

[0036] We also propose an audio device integrating a system as previously described.

[0037] A method for personalized speech enhancement is further proposed, implemented in the processing unit of the system as previously described, and comprising the steps, implemented in real time, of: o acquiring at least one input audio signal comprising a target voice of the speaker, and ambient noise; o generating, from the at least one input audio signal, a current voiceprint, calculating a similarity value representative of a similarity between the current voiceprint and the reference voiceprint, and adapting the reference voiceprint as a function of the current voiceprint and the similarity value; o producing at least one characterizing vector from the input audio signal;o performing an inference of the artificial intelligence model, by applying the at least one characterizing vector as input to said artificial intelligence model, and by introducing the reference voiceprint into said artificial intelligence model to personalize it; o acquiring a reproduction of the target voice as output from said artificial intelligence model.;

[0038] Further provided is a computer program comprising instructions that cause the processing unit of the system as previously described to execute the steps of the personalized speech enhancement method as previously described.

[0039] A computer-readable recording medium is further provided, on which the computer program as previously described is recorded. The invention will be better understood in light of the following description of particular non-limiting embodiments of the invention.

[0040] BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Reference will be made to the attached drawings, including:

[0042] [Fig. 1] Figure 1 shows a headset incorporating the personalized speech enhancement system;

[0043] [Fig. 2] Figure 2 schematically represents the personalized speech enhancement system;

[0044] [Fig. 3] Figure 3 illustrates the generation of the initial reference voiceprint;

[0045] [Fig. 4] Figure 4 represents the transmission of the user's audio extract, and the reception of the initial reference voiceprint by the headset;

[0046] [Fig. 5] Figure 5 represents the artificial intelligence model used in the personalized speech enhancement system;

[0047] [Fig. 6] Figure 6 represents the model of Figure 5 according to a first embodiment;

[0048] [Fig. 7] Figure 7 represents a concatenation operation;

[0049] [Fig. 8] Figure 8 represents the model of Figure 5 according to a second embodiment;

[0050] [Fig. 9] Figure 9 illustrates the generation of a voiceprint on the fly;

[0051] [Fig. 10] Figure 10 represents the implementation of a feedback loop to improve a voiceprint;

[0052] [Fig. 11] Figure 11 illustrates the constitution of the training data;

[0053] [Fig. 12] Figure 12 illustrates the general principle of the invention. DETAILED DESCRIPTION OF THE INVENTION

[0054] With reference to Figure 1, a headset 1 conventionally comprises one or more microphones 2 which make it possible to capture the voice of the user 3 of the headset 1 during a telephone call. This or these microphones 2 are for example positioned on a microphone boom 4 or on the earphones 5 of the headset 1. The voice of the user (the speaker) is then transmitted to his remote interlocutor.

[0055] The headset 1 includes a personalized speech enhancement system 6. The system is here integrated into one of the earphones 5 of the headset 1.

[0056] Referring to Figure 2, the system 6 comprises a first processing unit 7.

[0057] The first processing unit 7 is an electronic and software unit. The first processing unit 7 comprises one or more first processing components 8, and for example any processor or microprocessor, general or specialized (for example a DSP, for Digital Signal Processor, or a GPU, for Graphics Processing Unit, or even an NPU, for Neural Processing Unit), a microcontroller, or even a programmable logic circuit such as an FPGA (for Field Programmable Gate Arrays) or an ASIC (for Application Specific Integrated Circuit).

[0058] The system 6 also comprises one or more first memories 9 (and in particular one or more non-volatile memories), connected to or integrated into the first processing component(s) 8. At least one of these first memories 9 forms a computer-readable recording medium, on which is recorded at least one computer program comprising instructions which cause the first processing unit 7 to execute the steps of the personalized speech enhancement method which will be described.

[0059] The microphones 2, here three in number, therefore capture the sound signals So present in the environment of the headset 1, and produce input audio signals Sae. The sound signals So, and therefore the input audio signals Sae, include a useful signal, in this case the user's voice, but also ambient noise Ba. By "ambient noise", we mean here any non-useful sound signal captured by the microphones 2, coming from any type of "parasitic" source. The ambient noise Ba therefore possibly includes interfering voices, which are not those of the user.

[0060] The purpose of the personalized speech enhancement system 6 is to improve, for the user's interlocutor, the intelligibility of the user's voice during the telephone call. The system 6 is called "personalized" because it is adjusted and optimized, as will be seen, to reproduce the voice of this particular user. The user's voice is called the "target voice" Vc.

[0061] The personalized speech enhancement process involves a preliminary phase and a real-time phase. The preliminary phase is performed before the real-time phase.

[0062] The preliminary phase consists of first acquiring an audio sample produced by microphones from a capture of the user's target voice, and then generating an initial reference voiceprint from this audio sample. The initial reference voiceprint characterizes the user's target voice and forms a voice representation of the user.

[0063] It is this voice print that will allow, during the real-time phase, to personalize the speech enhancement.

[0064] As discussed below, the initial reference voiceprint is generated during the preliminary phase and then optimized in real time during the real-time phase to produce a reference voiceprint. The initial reference voiceprint is therefore the reference voiceprint that has not yet been updated.

[0065] Referring to Figures 3 and 4, the preliminary phase may use headset 1 (but not necessarily), as well as another device (but not necessarily). Here, the preliminary phase uses microphones 2 of headset 1, and the user's smartphone 10, which is connected to headset 1.

[0066] The smartphone 10, running a dedicated application, communicates with the user to implement the preliminary phase.

[0067] The smartphone 10 asks the user to speak. The user's voice is then captured by the microphones 2 of the headset 1, which therefore produce an audio extract Ea. This audio extract is transmitted by the headset 1 to the smartphone 10 on the application associated with the headset 1.

[0068] The smartphone 10 comprises a second processing unit 11 which, like the first processing unit 7 of the headset 1, comprises at least one second processing component. The smartphone 10 also comprises one or more second memories 12. At least one of these second memories 12 forms a computer-readable recording medium, on which at least one computer program (and in particular the application mentioned) is recorded.

[0069] The second processing unit 11 of the smartphone implements a voice encoder 14.

[0070] The voice encoder 14 used here is of the ECAPA-TDNN type (for Emphasized Channel Attention, Propagation, and Aggregation Time Delay Neural Network). This voice encoder 14 uses an artificial intelligence model. This model comprises layers implementing a time delay neural network, followed by an attention pooling layer. The model has been previously trained. Here, the weights are fixed during its use in the enhancement process. The second processing unit 11 of the smartphone 10 performs an inference of this model.

[0071] The audio extract Ea is therefore applied as input to the voice encoder 14. The audio extract Ea forms an observation %obs•

[0072] The observation is then transformed by the voice encoder 14 into an initial reference voice print Evi of the user.

[0073] We also call x e this voiceprint, and we have: x e = Encoder(xobs ), where Encoder is the encoding function implemented by the speech encoder.

[0074] Voiceprint x e is such that: x e G $R D , where D is the dimension of the latent space of the model of the encoder 14 and therefore of the initial reference voiceprint Evi. The smartphone 10 checks the quality of the initial reference voiceprint Evi. If it is not satisfactory, the smartphone asks the user to re-record his voice. If the quality of the initial reference voiceprint Evi is satisfactory, the smartphone 10 transmits to the headset 1 the initial reference voiceprint Evi, which is recorded in one of the first memories 9 of the system 6.

[0075] The preliminary phase is over. We note here:

[0076] - that the capture of the user's voice, to produce the audio extract, could be done by equipment other than the headset 1 (by the smartphone 10 for example, or by a computer 15);

[0077] - that the voice encoder 14 could be implemented in equipment other than the smartphone 10, in the computer 15 for example or in the headset 1 itself. The entire preliminary phase could therefore be carried out by the headset 1. The first processing unit 7 could therefore implement the voice encoder and produce the voice print.

[0078] The personalized speech enhancement feature can then be used in real time.

[0079] We are now interested in the real-time phase.

[0080] Before the first update of the reference voiceprint Ev, it corresponds to the initial reference voiceprint Evi.

[0081] The user is making a phone call. The microphones 2 therefore capture in real time sound signals So which include the target voice Vc of the user and the ambient noise Ba which prevails in the user's environment (and which may include interfering voices).

[0082] The first processing unit 7 of the headset 1 comprises, for each microphone 2, a buffer 17 and a calculation block 18.

[0083] The first processing unit 7 therefore acquires, for each microphone 2, the input audio signal Sae produced by said microphone from the sound signal So. The input audio signals Sae therefore comprise the target voice Vc of the user and the ambient noise Ba.

[0084] Each calculation block 18 then calculates the Short Time Fourier Transform (STFT) of the input audio signal Sae.

[0085] The output of each calculation block 18 therefore represents the spectrum of an input audio signal.

[0086] The first processing unit 7 then produces, from the input audio signals Sae, a current audio signal Sac which is a frequency domain signal and which is a spectral representation of the input audio signals.

[0087] The first processing unit 7 of the headset comprises at least one pre-processing block, at the input of which the current audio signal is applied, to produce at least one characterizing vector from the current audio signal and therefore from the input audio signals.

[0088] Here, the first processing unit 7 comprises a first pre-processing block 20a and a second pre-processing block 20b.

[0089] The first pre-processing block 20a generates a first vector characterizing Vcl. The second pre-processing block 20b generates a second vector characterizing Vc2.

[0090] The first vector characterizing Vcl is obtained here by estimating the ERB (Equivalent Rectangular Bandwidth) of the current audio signal.

[0091] The second vector characterizing Vc2 is obtained here by estimating the complex spectrogram of the current audio signal Sac. The second vector characterizing Vc2 is obtained by converting the TFCT to a logarithmic scale.

[0092] We therefore obtain:

[0093] ERB.Spec = PreprocessingÇxa), where x a is the current audio signal, ERB is the first characterizing vector, Spec is the second characterizing vector, and Preprocessing is the preprocessing function implemented by the preprocessing blocks 20a, 20b.

[0094] Then, the first processing unit 7 executes an inference of a previously trained artificial intelligence model 21 for speech enhancement, by applying the first vector characterizing Vcl and the second vector characterizing Vc2 as input to the model 21.

[0095] The inference model (and therefore, in particular, its weights) is stored in one of the first memories 9 of the system 6 of the headset 1.

[0096] The model 21 used here is based on the U-Net network and, more specifically, on the DeepFilterNet2 model which has a so-called "cascade" or "dual-stage" architecture. This type of "cascade" architecture was developed for standard speech enhancement and has greatly improved enhancement performance. It is therefore a very efficient architecture.

[0097] We first describe a model according to a first embodiment, with reference to figure 5. This model 21 is a two-stage model (dual-stage architecture): 21a and 21b.

[0098] Model 21 includes:

[0099] - an encoder block 23, at the input of which is applied at least one characterizing vector (and therefore here the first characterizing vector and the second characterizing vector), and generating at least one output vector;

[0100] - at least one decoder 24, at the input of which is applied the at least one output vector, and generating a reproduction of the target voice.

[0101] As we will see, the reference voice print Ev is also applied as input to the encoder block 23.

[0102] Model 21 also has a separator.

[0103] The encoder block 23 makes it possible to analyze the characterizing vectors by relating the frequency information. The encoder 23 includes convolution layers making it possible in particular to reduce the size of the characterizing vectors by compressing the information and obtaining new characteristics in a latent space. These characteristics are then used by the separator, whose role is to separate the voice from the rest in the latent space. The separator is made up of recurrent networks in order to relate the temporal information. Finally, the decoder makes it possible to transform the information from the latent space to the space of the input characteristics (and therefore from the first characterizing vector and the second characterizing vector).

[0104] More precisely, with reference to FIG. 6, the encoder block 23 comprises a first branch 23a at the input of which the first vector characterizing Vcl (ERB) is applied, and a second branch 23b at the input of which the second vector characterizing Vc2 is applied.

[0105] (complex spectrogram).

[0106] The first branch 23a belongs to the first stage 21a of the model 21. The second branch 23b belongs to the second stage 21b of the model 21.

[0107] The first branch 23a of the encoder block 23 successively comprises a first convolution layer 25, a second convolution layer 26, a third convolution layer 27 and a fourth convolution layer 28.

[0108] The second branch 23b of the encoder block 23 successively comprises a first convolution layer 29, a second convolution layer 30 and a grouped linear layer 31 (GLinear).

[0109] The encoder block 23 generates at least one intermediate vector and comprises at least one operator 32 arranged to merge the at least one intermediate vector and the reference voiceprint Ev to produce the at least one output vector.

[0110] Here, the fourth convolution layer 28 of the first branch 23a produces a first intermediate vector Vil. The grouped linear layer 31 of the second branch 23b produces a second intermediate vector Vi2.

[0111] The first intermediate vector Vil is:

[0112] %ERB = Enc ERB (ERB), where Enc ERB is the function implemented by the first branch 23a of the encoder block 23.

[0113] The second intermediate vector Vi2 is:

[0114] ^spec Enc Spec (Spec), where Enc Specis the function implemented by the second branch 23b of the encoder block 23. The at least one operator 32 comprises a concatenation operator arranged to concatenate along a frequency axis at least one intermediate vector and the reference voice print Ev.

[0115] The concatenation operator 32 therefore merges by concatenation the first intermediate vector Vil, the second intermediate vector Vi2 and the reference voiceprint Ev.

[0116] Figure 7 illustrates concatenation.

[0117] The concatenation is carried out as follows. The reference voiceprint Ev is duplicated along the time axis in order to obtain as many occurrences as frames in the intermediate vectors (after passing through the first layers of the encoder), then each reference voiceprint Ev is concatenated with each frame of the two intermediate vectors.

[0118] In Figure 7, all frames of the intermediate vectors are different, while all frames of the voiceprint are equal. The unified architecture of the model in Figure 5 and Figure 6 therefore corresponds to the addition of the reference voiceprint Ev in the encoder block, thus allowing to combine together the ERBs, the complex spectrograms and the information on the target voice of the speaker.

[0119] The concatenation operator 32 therefore generates the concatenated vector Vc:

[0120] Concat[X ERB ;X Spec ;x e ]

[0121] The encoder block 23 further comprises a grouped linear layer 33 at the input of which the output of the concatenation operator 32 is applied. The separator, which is here integrated into the encoder block 23, comprises a closed recurrent unit 34 (GRU layer, for Gated Recurrent Unit).

[0122] The closed recurrent unit 34 therefore produces an output vector Vs from the concatenated vector Vc, and therefore from the first intermediate vector Vil, the second intermediate vector Vi2 and the reference voiceprint Ev.

[0123] The output vector Vs is:

[0124] X Enc = GRU(Concat[X ERB ;X Spec ;x e ]) where GRU is the closed recurrent unit function 34.

[0125] The at least one decoder 24 here comprises a first decoder 24a and a second decoder 24b.

[0126] The first decoder 24a belongs to the first stage 21a of the model 21. The second decoder 24b belongs to the second stage 21b of the model 21. The convolution layers of the first branch 23a of the encoder block 23 are connected to the first decoder 24a.

[0127] The output vector Vs is applied to the input of the first decoder 24a and to the input of the second decoder 24b.

[0128] The first decoder 24a generates from the output vector Gerb gains called “ERB gains”:

[0129] GERB = Dec E R B (X Enc )

[0130] These gains are applied to the first vector characterizing Vcl.

[0131] We therefore obtain:

[0132] X G = ERB xG ERB

[0133] The second decoder 14b, for its part, generates filtering coefficients Cdf of the deep filter 36 (Deep Filtering):

[0134] GDF = DeCDF(XEne) The first vector characterizing Vcl, to which the ERB gains have been applied, is then applied to the input of the deep filter 36:

[0135] The first stage 21a of the model 21 works in amplitude and therefore allows to estimate an envelope of the target voice. It performs a coarse, imprecise analysis, producing an imprecise reproduction of the target voice. The first stage 21a produces real-valued gains. The second stage 21b performs a more precise analysis. The second stage 21b operates in the complex domain and implements a deep filtering, using the results of the first stage (the gains), which allows to obtain a denoised content forming the reproduction Rv of the target voice, which is a more precise reproduction than that of the first stage.

[0136] This two-story structure allows the model to be compacted while maintaining its efficiency.

[0137] The second stage 21b operates on the lower part of the spectrogram (here up to the frequency f d f = 5kHz).

[0138] Note that a mask is generated and applied to the output of the deep filter 36 in order to mask the noise.

[0139] An inverse Fourier transform 37 is then calculated to obtain the reproduction of the target voice Rv.

[0140] The first processing unit 7 of the headset 1 then acquires the reproduction Rv of the target voice and transmits it to the user's interlocutor, who thus receives the enhanced voice of the user.

[0141] We speak of a “unified” encoder block to designate the encoder block 23 of the first embodiment. Indeed, the two branches 23a, 23b of the encoder 23 merge within it via the concatenation operator 32. We are now interested, with reference to FIG. 8, in a model 40 according to a second embodiment and, more precisely, in a first alternative of this second embodiment.

[0142] The model 40 again comprises an encoder block 41 and at least one decoder 42, in this case a first decoder 42a and a second decoder 42b.

[0143] The encoder block 41 comprises a first branch 41a and a second branch 41b.

[0144] The first branch 41a of the encoder block 41 successively comprises a first convolution layer 43, a second convolution layer 44, a third convolution layer 45, a fourth convolution layer 46, then a first concatenation operator 47, a grouped linear layer 48 (GLinear) and a first closed recurrent unit 49.

[0145] The second branch 41b of the encoder block 41 successively comprises a first convolution layer 50, a second convolution layer 51, a first grouped linear layer 52 (GLinear), then a second concatenation operator 53, a second grouped linear layer 54 and a second closed recurrent unit 55.

[0146] The first concatenation operator 47 concatenates along the frequency axis a first intermediate vector Vcl (output of the fourth convolution layer 46) produced by the first branch of the encoder block, and the reference voiceprint Ev. The second concatenation operator 53 concatenates along the frequency axis a second intermediate vector Vc2 (output of the third convolution layer 52) produced by the second branch 41b of the encoder block 41, and the reference voiceprint

[0147] Ev. The first closed recurrent unit 49 generates a first output vector Vsl which is applied as input to the first decoder 42a. The second closed recurrent unit 55 generates a second output vector Vs2 which is applied as input to the second decoder 42b.

[0148] In this first alternative of the second embodiment of model 40, the reference voice print Ev is therefore integrated into each of the branches of the encoder block.

[0149] The encoder block of the second embodiment is referred to as a "double" or "dual" encoder block. In fact, the two branches of the encoder do not merge within it but are independent and each connected to a separate decoder.

[0150] More generally, the first branch 41a of the encoder block comprises a first operator 47 arranged to merge a first intermediate vector produced by the first branch of the encoder block and the reference voiceprint, and / or the second branch 41b of the encoder block comprises a second operator 53 arranged to merge a second intermediate vector produced by the second branch of the encoder block and the reference voiceprint.

[0151] We could therefore also, according to a second alternative of the second embodiment of the model, integrate the reference voiceprint only in the first branch 41a of the encoder block 41 (which alone then includes a concatenation operator). We could also, according to a third alternative of the second embodiment, integrate the reference voiceprint only in the second branch 41b of the encoder block 41 (which alone then includes a concatenation operator). We are now interested in the real-time updating of the reference voiceprint Ev.

[0152] The first processing unit 7 generates, from the at least one input audio signal Sae, a current voiceprint Eve, calculates a similarity value representative of a similarity between the current voiceprint Eve and the reference voiceprint Ev, and adapts the reference voiceprint Ev as a function of the current voiceprint Eve and the similarity value.

[0153] In a first embodiment, with reference to FIG. 9, the first processing unit 7 comprises an on-the-fly fingerprint generation block 60, comprising a fingerprint generation block 61 and a similarity calculation block 62.

[0154] The first processing unit 7 acquires the at least one input audio signal Sae of the user which is applied as input to the fingerprint generation block 61. The fingerprint generation block 61 generates from the at least one input audio signal a current voice print Eve. Then, the calculation block 62 calculates a similarity value representative of a similarity between the current voice print Eve and the reference voice print Ev in memory.

[0155] If this update of the reference voiceprint is the first update, the reference voiceprint Ev in memory is the initial reference voiceprint Evi, which was produced during the preliminary phase. If this update is not the first, the reference voiceprint Ev in memory is the reference voiceprint Ev that was adapted during the previous print update. The similarity value here is a distance between two vectors: the first vector is the reference voiceprint Ev in memory and the second vector is the current voiceprint Eve. The distance has a value that is for example between "0" and "1", where "0" means "completely dissimilar" and "1" means "identical".

[0156] If the two voiceprints are sufficiently similar (i.e., if the similarity value is greater than a predefined threshold), the first processing unit 7 generates a new reference voiceprint which is calculated from the reference voiceprint Ev stored in the memory, the current voiceprint Eve and the similarity value. This new reference voiceprint becomes the reference voiceprint Ev, is stored in the memory and is introduced into the template 21 (or 40).

[0157] This method has several advantages. First, it allows for an increasingly robust voiceprint and overcomes the problem of domain adaptation. Indeed, when changing the sound environment, the nature of the voiceprint may change slightly. Here, this reference voiceprint is updated in real time, which eliminates this problem. This method also overcomes the difficulty of temporary voice changes, for example if the speaker has a broken, hoarse, or other voice. This method therefore makes it possible to adapt to this type of situation.

[0158] On-the-fly voiceprint generation also allows, as this print is regularly updated, to produce a less precise initial reference voiceprint during the preliminary phase, thus using a less cumbersome voice encoder. The reduction in precision is compensated by the real-time improvement of the print. This can allow the voice encoding of the preliminary phase to be carried out on-board, and therefore in the first processing unit 7 of the headset 1. Interaction with the smartphone (or any other equipment) is then no longer necessary.

[0159] Alternatively, with reference to FIG. 10, the first processing unit 7 can implement a feedback loop 64 to generate the current voice print Eve (“on the fly”) from the reproduction Rv of the target voice output from the model.

[0160] The first processing unit 7 performs a first inference of the model 21, 40 using the reference voice print Ev stored in memory. The model produces a reproduction Rv of the target voice (enhanced voice).

[0161] The first processing unit 7 again comprises an on-the-fly fingerprint generation block 65, comprising a fingerprint generation block 66 and a similarity calculation block 67. The block 66 generates the current voiceprint Eve from the reproduction Rv of the target voice, then the calculation block 67 calculates the similarity between the reference voiceprint Ev in memory and the current voiceprint Eve.

[0162] If the similarity value is greater than a predefined threshold, then the first processing unit 7 generates a new reference voiceprint which is calculated from the reference voiceprint Ev stored in the memory, the current voiceprint Eve and the similarity value. This new reference voiceprint becomes the reference voiceprint Ev and is stored in the memory.

[0163] The first processing unit 7 performs another inference of the model, this time using the new reference voiceprint Ev. The reproduction Rv of the target voice is then again applied as input to the on-the-fly fingerprint generation block 65.

[0164] We are now interested, with reference to Figure 11, in the constitution of the training data.

[0165] As seen, the model used is a speech enhancement model that has been transformed to be personalized to the user (more precisely, to the user's voice).

[0166] Three datasets constitute the training data used to train this model.

[0167] A first data set 71 comprises voice prints obtained from audio extracts of identified speakers. The audio extracts here come from a first data corpus 72, for example the LibriSpeech corpus. For each identified speaker 73, an audio extract 75 was taken, then this audio extract 75 was encoded by the voice encoder 14 previously described. The first data set 71 is therefore obtained, which therefore comprises voice prints of identified speakers.

[0168] As we have seen, the model must be able to effectively isolate a target voice in an environment including noise and interfering voices. We therefore used a second dataset 76 containing audio extracts of unidentified speakers, to produce interfering voices 77. The second dataset 76 is for example from LibriSpeech and Mozilla Common Voice. We also used a third dataset 78 containing noise extracts, to obtain audio noise 79. The third dataset 78 here includes data from the DNS challenge (themselves from the Freesound and AudioSet databases).

[0169] An audio mixer 80 was used to mix the interfering voices 77, the noise 79, and audio clips of identified speakers to form the second dataset 81. The audio clips of identified speakers, although of the same speakers, are different from those used to form the first dataset 71.

[0170] Finally, a third dataset 82 was generated from voice excerpts of the identified speakers. Again, these audio excerpts, although concerning the same speakers, are different from those used to form the first dataset 71.

[0171] The model is therefore trained using as inputs the data from the second data set 81. For each data item from the second data set 81 corresponding to the same speaker, the voice print of said speaker from the first data set 71 is introduced into the model. The model output, i.e. the reproduction of the target voice, is compared with the audio extract from the third data set 82 to check whether the reproduction of the target voice is correct, then the weights of the model are updated (by gradient backpropagation for example).

[0172] Approximately 2,400 different speakers were used for the target voices and approximately 7,000 unwanted voices, resulting in over 800 hours of data.

[0173] A combination of cost functions (or loss functions) was used for training: compressed spectral loss, multi-resolution loss, and over-suppression loss.

[0174] The compressed spectral loss is the sum of two terms. The first term is the mean square error of the spectrogram amplitude compressed by a factor c, and the second is the same term but taking into account the phase.

[0175] The compressed spectral loss can be written:

[0176] Multi-resolution loss supports voice reconstruction. This involves calculating a cost function for multiple TFCTs generated with different window sizes (e.g., 5ms, 10ms, 20ms, and 40ms). For each TFCT, the mean square error of the compressed amplitude is multiplied by the mean square error of the compressed complex TFCT.

[0177] The multi-resolution loss can be written:

[0178] The over-suppression loss is specific to personalization and is intended to avoid over-suppressing the target voice. It simply penalizes the network when pieces of the target voice are missing from the reconstructed spectrogram.

[0179] The over-suppression loss can be written: The overall cost function is therefore written: where L spec is the compressed spectral loss,L MR is the multi-resolution loss, L osis the over-suppression loss, and where the are the loss weights defined in during training.

[0180] The results obtained for the four models presented are presented. Here, the model according to the first embodiment is called "first model", the model according to the first alternative of the second embodiment is called "second model", the model according to the second alternative of the second embodiment is called "third model", and the model according to the third alternative of the second embodiment is called "fourth model".

[0181] Reference is made to the two tables in the Appendix: Table 1 and Table 2.

[0182] The enhancement results obtained are grouped in Table 1. The results were obtained on a "synthetic" test base, which was made by the inventors, and whose characteristics are grouped on the "raw test base" line. For all the metrics used, a higher result means better performance. The metrics used here are metrics known to those skilled in the art: PESQ, STOI, CSIG, CBAK and COVL.

[0183] We can see that the different customized models (first, second, third and fourth models) improve the performances and clearly exceed those of the original DeepFilterNet2 model. The first model, with unified encoder block, is the one that presents the best results. Furthermore, as we have seen, the size of the model used is critical because the inference is carried out on-board. We have therefore adapted a compact model (DeepFilterNet2) to the customization while impacting its complexity as little as possible.

[0184] Table 2 allows us to compare the first (most efficient) model with some prior art networks on a computational level. This table shows the number of parameters and the number of MAC operations (for Multiply-Accumulative Operation) associated with the different models. The lower the value obtained, the more favorable it is.

[0185] Table 2 also specifies whether or not each model is a cascade (dual-stage) model.

[0186] Two main lessons can be drawn from this table. First, the customization method used did not increase the computational load of the model. The model was customized without increasing its complexity. Second, we note that the model used is significantly smaller, both in number of parameters and in MAC. This very compact model can therefore be implemented on an embedded product, which is not the case for other state-of-the-art models. We also note that the closest network, E3Net, does not correspond to a cascaded architecture (dual-stage).

[0187] With reference to Figure 12, the method described therefore consists, during a preliminary phase, in applying an audio extract Ea as input to a voice encoder 14 to obtain an initial reference voice print Evi of the target voice. The initial reference voice print Evi is stored in the memory of the embedded product. Then, in real time, the first processing unit acquires at least one input audio signal Sae, which is a mixture between the target voice Vc and an ambient noise Ba comprising noises B and one or more interfering voices Vint. The model 21, 40 then generates a reproduction Rv of the target voice. The reference voice print Ev is updated and optimized in real time.

[0188] This method, more complex than simple noise reduction, can greatly improve the clarity of a call when there are several people involved.

[0189] Here it has been described that the target voice is that of the headset user.

[0190] However, the system and method can also be used to enhance a target voice that is that of the user's interlocutor. There may be multiple target voices, from multiple interlocutors.

[0191] The process is similar to that used when you want to isolate the user's voice, but in reverse. To do this, you need an audio sample of the voices of frequently called people.

[0192] The process includes the following steps.

[0193] Frequent callers record their voice.

[0194] From these recordings, an initial reference voiceprint is generated for each speaker and is stored in the user's headset.

[0195] During a call, if the speaker is identified, the incoming audio stream is retrieved and transmitted to the neural network, previously trained on the dataset presented above. Again, the stream is divided into overlapping windows and each window is applied as input to the neural network. The voiceprint of the speaker is introduced into the model to personalize it. The reproduction of the target voice is then transmitted to the headset user.

[0196] Alternatively, for each window, the audio signal can be compared to known voiceprints to identify the target voice. The matching voiceprint is then used to personalize the model.

[0197] Similarly, within the same household, for example, the voiceprints of several speakers can be generated and stored. When the voice of one of these speakers is recognized, the personalized enhancement process is carried out using the voiceprint of said speaker.

[0198] Generating and storing multiple voiceprints for multiple speakers allows for personalization for multiple speakers at the same time. For example, when using a conference speakerphone, everyone in the room records their voices so the system can perform personalized enhancement for each voice.

[0199] The system and method can also be used in a face-to-face conversation application, and for example in a hearing aid audio device.

[0200] The target voice can be that of the user's interlocutor on the audio device.

[0201] The target voice can also be that of the audio device user. The process allows for the reproduction of the target voice, which is then removed from the captured signal.

[0202] OWS (Open Wearable Stereo) or open-ear headphones have the disadvantage that, in the case of directional enhancement, the user's voice will also be enhanced with an echo due to processing latency, which will be counterproductive.

[0203] It may therefore be interesting to enhance the useful signal (voice of the speakers opposite) by removing interfering noise (background noise) and removing the user's own voice from the device in order to avoid an echo phenomenon.

[0204] Of course, the invention is not limited to the embodiments described but encompasses any variant falling within the scope of the invention as defined by the claims.

[0205] The speech encoder used to generate the initial reference voiceprint may be different from the one described here. For example, the speech encoding could be a representation in the form of "MFCC" coefficients (for Mel-Frequency Cepstrum Coefficients).

[0206] A first characterizing vector and a second characterizing vector are used, but there could be only one characterizing vector, or a different number of characterizing vectors. The at least one characterizing vector can be obtained from different representations of the at least one input audio signal.

[0207] The input audio signal(s), or the current audio signal, can also be introduced as is, without processing, into the model. The at least one characterizing vector can then be the current audio signal itself, or the at least one input audio signal itself (or a combination of the input audio signals). Passing through an intermediate frequency representation is not mandatory.

[0208] It has been described that the speech enhancement artificial intelligence model has a structure that is based on the structure of a U-Net network and, more specifically, that the structure of the model is based on the structure of a DeepFilterNet2 network. This is not mandatory, however. Any speech enhancement model can be used and customized by introducing the voiceprint into said model. The model is not necessarily an encoder-decoder model and, in an encoder-decoder model, the introduction of the voiceprint into the model is not necessarily performed in the encoder at the output of the model.

[0209] The introduction of the voiceprint can be done by using an operator arranged to merge one or more intermediate vectors and the voiceprint. This fusion is not necessarily a concatenation, it could be an addition, a multiplication, a cross-attention mechanism, etc.

[0210] The personalized speech enhancement system can be integrated into a portable audio device other than a headset: earphones, hearing aid, tablet, laptop, smartphone, etc. The personalized speech enhancement system can also be integrated into a non-portable device: videoconferencing system, hands-free car kit, virtual personal assistant, any device equipped with a voice recognition module, etc. The personalized speech enhancement system can also be integrated with existing software, in particular videoconferencing. APPENDIX Table 1

[0211] Table 2

Claims

CLAIMS 1. System (6) for personalized speech enhancement, comprising: - at least one memory (9), in which are stored a reference voice print (Ev) of a speaker, previously recorded, and an artificial intelligence model (21; 40) for speech enhancement, previously trained; - a processing unit (7) arranged to, in real time: o acquire at least one input audio signal (Sae) comprising a target voice (Vc) of the speaker and ambient noise; o generate, from the at least one input audio signal, a current voiceprint (Eve), calculate a similarity value representative of a similarity between the current voiceprint (Eve) and the reference voiceprint (Ev), and adapt the reference voiceprint according to the current voiceprint and the similarity value; o produce at least one characterizing vector (Vcl, Vc2) from the at least one input audio signal; o execute an inference of the artificial intelligence model, by applying the at least one characterizing vector as input to said artificial intelligence model, and by introducing the reference voiceprint into said artificial intelligence model to personalize it;o acquire a reproduction (Rv) of the target voice at the output of said artificial intelligence model.; 2. System according to claim 1, in which the artificial intelligence model comprises: - an encoder block (23; 41), at the input of which is applied at least one characterizing vector, and generating at least one output vector; - at least one decoder (24; 42), at the input of which is applied the at least one output vector; the encoder block generating at least one intermediate vector and comprising at least one operator (32; 47, 53) arranged to merge the at least one intermediate vector and the reference voice print to produce the at least one output vector.

3. System according to one of the preceding claims, in which the artificial intelligence model comprises a first stage carrying out a less precise analysis producing a less precise reproduction of the target voice, and a second stage using results of the first stage and carrying out a more precise analysis to obtain the reproduction (Rv) of the target voice.

4. System according to claims 2 and 3, in which the processing unit (7) produces a first characterizing vector (Vcl) and a second characterizing vector (Vc2) from the input audio signal, and in which: - the encoder block comprises a first branch (23a; 41a) belonging to the first stage and at the input of which the first characterizing vector is applied, and a second branch (23b; 41b) belonging to the second stage and at the input of which the second characterizing vector is applied; the model comprises a first decoder (24a; 42a) belonging to the first floor and a second decoder (24b; 42b) belonging to the second floor.

5. System according to claim 4, in which the operator (32) is arranged to merge a first intermediate vector (Vil) produced by the first branch (23a) of the encoder block, a second intermediate vector (Vi2) produced by the second branch (23b) of the encoder block, and the reference voiceprint (Ev).

6. System according to claim 4, wherein the first branch (41a) of the encoder block comprises a first operator (47) arranged to merge a first intermediate vector produced by the first branch of the encoder block and the reference voiceprint, and / or the second branch (41b) of the encoder block comprises a second operator (53) arranged to merge a second intermediate vector produced by the second branch of the encoder block and the reference voiceprint, the first branch producing a first output vector (Vsl) applied as input to the first decoder (42a) and the second branch producing a second output vector (Vs2) applied as input to the second decoder (42b).

7. System according to one of claims 4 to 6, in which the processing unit produces, from the at least one input audio signal (Sae), a current audio signal (Sac) which is a frequency domain signal, the first characterizing vector (Vcl) being obtained by calculating an ERB of the current audio signal, and the second characterizing vector (Vc2) being obtained by calculating a complex spectrogram of the current audio signal.

8. System according to one of claims 2 to 7, in which the at least one operator (32; 47, 53) comprises at least one concatenation operator arranged to concatenate along a frequency axis the at least one intermediate vector and the reference voiceprint.

9. System according to one of claims 2 to 8, wherein a structure of the model is based on a structure of a U-Net network.

10. The system of claim 9, wherein the structure of the model is based on a structure of a DeepFilterNet2 network.

11. System according to one of the preceding claims, an initial reference voice print (Evi) being obtained using an ECAPA-TDNN type voice encoder.

12. System according to claim 11, wherein the initial reference voiceprint is generated by the processing unit (7) of the system (6).

13. System according to one of the preceding claims, in which the processing unit (7) is arranged to implement a feedback loop (64) arranged to generate the current voice print from the reproduction (Rv) of the target voice at the output of the model.

14. System according to one of the preceding claims, the speaker being a user of the system, who communicates with a remote interlocutor.

15. System according to one of claims 1 to 13, the speaker being an individual who communicates with the user of the system.

16. Audio device (1) integrating a system according to one of the preceding claims.

17. A method for personalized speech enhancement, implemented in the processing unit (7) of the system (6) according to one of claims 1 to 15, and comprising the steps, implemented in real time, of: o acquiring at least one input audio signal (Sae) comprising a target voice (Vc) of the speaker, and ambient noise; o generating, from the at least one input audio signal, a current voiceprint (Eve), calculating a similarity value representative of a similarity between the current voiceprint (Eve) and the reference voiceprint (Ev), and adapting the reference voiceprint as a function of the current voiceprint and the similarity value; o producing at least one characterizing vector (Vcl, Vc2) from the input audio signal;o performing an inference of the artificial intelligence model, by applying the at least one characterizing vector as input to said artificial intelligence model, and by introducing the reference voiceprint into said artificial intelligence model to personalize it; o acquiring a reproduction (Rv) of the target voice as output from said artificial intelligence model.; 18. Computer program comprising instructions which cause the processing unit (7) of the system (6) according to one of claims 1 to 15 to execute the steps of the personalized speech enhancement method according to claim 18.

19. Computer-readable recording medium on which the computer program according to claim 18 is recorded.

Citation Information

Patent Citations

  • Voice quality enhancement method and related device

    US20240096343A1