Voice synthesizer, voice synthesis method, and computer program therefor

WO2025186271A8PCT designated stage Publication Date: 2025-10-02THALES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/055876
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2025-03-04
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Conventional low-bitrate speech synthesizers suffer from significant degradation in speech quality due to reduced input parameters, while AI-based solutions require extensive computational resources and are not suitable for real-time implementation.

Method used

A neural source-filter model using monodirectional recurrent neural networks for voice synthesis, which projects input parameters into a latent space and generates harmonics, allowing for real-time operation with improved quality and reduced computational footprint.

Benefits of technology

The proposed solution achieves high-quality voice synthesis with minimal input parameters, enabling real-time operation and reduced memory and computational costs, while mitigating the limitations of existing AI and low-bitrate synthesizers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025055876_02102025_PF_FP_ABST
    Figure EP2025055876_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a voice synthesizer (10) comprising: - a projection module (12) configured to project, via a first unidirectional recurrent neural network, input parameters characterizing each frame of an input sound signal comprising a human voice, in a latent space having a parameterizable dimension; - a harmonic signal generator (14) configured to generate, via a second neural network, a signal representative of the harmonics of said input sound signal; - a voice synthesizer (16) configured to synthesize, via a third unidirectional recurrent neural network, an encoded voice associated with said human voice from: - the projection of the projected input parameters into said latent space having a parameterizable dimension, and - the signal representative of the harmonics of the input sound signal comprising the human voice and provided by the generator.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TITLE: Speech synthesizer, speech synthesis method and associated computer program

[0002] The present invention relates to a speech synthesizer having an architecture according to a neural source-filter model.

[0003] The present invention also relates to a voice synthesis method implemented by said voice synthesizer.

[0004] The invention also relates to a computer program comprising software instructions which, when implemented by said speech synthesizer, implement such a method.

[0005] The invention relates to the field of vocoding, voice synthesis and voice transmission.

[0006] Subsequently, according to the present invention "vocoding" and by extension the corresponding device "voice synthesizer" is associated with the contraction of the English words Voice Coding respectively Voice coder, in other words voice coding by an electronic device for processing the sound signal, corresponding substantially to a method according to which, from a decomposition into important properties (eg main spectral components) of an input voice or another sound, synthesizes an associated synthetic voice at the output.

[0007] In recent years, vocoding solutions including AI capabilities have been proposed. However, such solutions are generally very demanding in terms of input parameters, requiring, for example, an input number of around eighty Mel frequencies to provide voice synthesis. In addition, such AI solutions are generally not designed to allow real-time implementation, favoring operation adapted to uses such as text-to-speech transformation or voice assistants. Furthermore, the number of parameters generally associated with the implementation of such AI solutions, in the order of tens of millions of parameters, reduces the possibilities of implementation in embedded hardware with strong memory and computing capacity constraints.

[0008] In conventional low-bitrate digital speech synthesizers, the input parameters are often limited in number and compressed so as to use the least bit rate possible, however such a reduction in the number of input parameters used by the speech synthesizer is generally accompanied by a degradation in the quality of the speech synthesis. Thus, conventional low-bitrate solutions generally suffer from a fairly significant deterioration of the synthesized voice, these conventional low-bitrate solutions cannot compensate for the lack of information with a priori information, as is the case for the aforementioned AI solutions.

[0009] For example, for these classic low-bitrate solutions, reducing the number of input frequencies leads to a deterioration of the high-frequency components, and without a compensation mechanism, classic low-bitrate speech synthesizers often provide a very robotic voice.

[0010] To summarize, currently, the decrease in the number of input parameters is therefore systematically accompanied by a decrease in the quality of speech synthesis, and conventional low-bitrate speech synthesizers cannot compensate for the loss of information linked to the decrease in the number of parameters, while AI speech synthesizers compensate for the loss of information but are often characterized by heavy numerical calculations, as well as by a large number of implementation parameters, these properties being inherent to the fact that AI speech synthesizers learn a priori distribution to compensate for areas without information.

[0011] The aim of the invention is therefore to propose a voice synthesizer capable of synthesizing a human voice of sufficient quality for voice communication while exploiting a very small number of input parameters.

[0012] To this end, the invention relates to a voice synthesizer having an architecture according to a neural source-filter model comprising at least:

[0013] - a projection module configured to project, via a first monodirectional recurrent neural network, input parameters, characterizing each frame of an input sound signal comprising a human voice, into a latent space of configurable dimension;

[0014] - a harmonic signal generator configured to generate, via a second neural network, a signal representative of the harmonics of said input sound signal comprising said human voice;

[0015] - a voice synthesizer configured to synthesize, via a third monodirectional recurrent neural network, an encoded voice associated with said human voice from:

[0016] - the projection of said projected input parameters into said latent space of parameterizable dimension provided by said projection module, and

[0017] - said signal representative of the harmonics of the input sound signal comprising said human voice and provided by said generator. Such a voice synthesizer is advantageous because unlike the neural source-filter architectures in the literature, the use of monodirectional recurrent neural networks is compatible with real-time use, i.e. with a latency of less than 50ms.

[0018] Furthermore, the speech synthesizer according to the present invention is configured to provide a frugal speech synthesis solution, because it is based on a low-data-intensive architecture. Indeed, the source-filter model is certainly a classic model for decomposing the voice into source and filter components, however the use of this model, boosted with “modern” monodirectional recurrent neural networks (i.e. capable of restoring causality) makes it possible to offer a framework conducive to voice modeling. Such “artificial intelligence boosting” leads to better voice reconstruction with lower memory and computational costs.

[0019] According to other advantageous aspects of the invention, the voice synthesizer comprises one or more of the following characteristics, taken individually or in all technically possible combinations:

[0020] - said signal representative of the harmonics of said input sound signal comprising said human voice is a weighted sum of the signal corresponding to the pitch of said human voice and its harmonics, said weighting being learned beforehand by said second neural network;

[0021] - white noise is added to said weighted sum;

[0022] - said input parameters characterizing said input sound signal comprising said human voice are frequency parameters corresponding to one of the elements belonging to the group comprising at least:

[0023] - a Fourier spectrum of said input sound signal comprising said human voice;

[0024] - a cepstrum defined by Mel frequency cepstral coefficients calculated by a discrete cosine transform applied to the power spectrum of said input sound signal comprising said human voice;

[0025] - the constant Q transform of said input sound signal comprising said human voice;

[0026] - the synthesizer further comprises an adversarial training module comprising two discriminators configured to be trained simultaneously with said voice synthesizer during its supervised learning phase, according to a zero-sum game, said discriminators being trained to differentiate real signals from those suitable for being synthesized by said voice synthesizer, said voice synthesizer being trained to fool said two discriminators; - one of said two discriminators is sensitive to low-frequency information and the other of said two discriminators is sensitive to high-frequency information;

[0027] - the synthesizer further includes:

[0028] - a reconstruction module configured to reconstruct missing frames of said synthesized encoded voice from frames before or after said missing frames, and / or

[0029] - a band extension module configured to reconstruct a signal comprising said synthesized encoded voice with a sampling frequency higher than that of said input sound signal.

[0030] - at least one of said neural networks of said voice synthesizer is also configured to extract the vocal biomarkers of the speaker having generated said input sound signal.

[0031] The invention also relates to a voice synthesis method implemented by such a voice synthesizer having an architecture according to a neural source-filter model, said method comprising the following steps:

[0032] - projection, via a first monodirectional recurrent neural network of said voice synthesizer, of the input parameters, characterizing each frame of an input sound signal comprising a human voice, into a latent space of configurable dimension

[0033] - generation, via a second neural network of said voice synthesizer, of a signal representative of the harmonics of said input sound signal comprising said human voice;

[0034] - voice synthesis, via a third monodirectional recurrent neural network, of an encoded voice associated with said human voice from:

[0035] - the projection of said projected input parameters into said latent space of parameterizable dimension provided by said projection module, and

[0036] - said signal representative of the harmonics of the input sound signal comprising said human voice and supplied by said generator.

[0037] The invention also relates to a computer program comprising software instructions which, when executed by a speech synthesizer, implement a speech synthesis method, as defined above.

[0038] The invention will appear more clearly on reading the description which follows, given solely by way of non-limiting example, and made with reference to the drawings in which: [Fig. 1] Figure 1 is a general schematic representation of a voice synthesizer according to the present invention;

[0039] [Fig. 2] Figure 2 illustrates the use of the speech synthesizer of Figure 1;

[0040] [Fig. 3] Figure 3 illustrates an example of the architecture of the speech synthesizer of Figure 1;

[0041] [Fig. 4] Figure 4 illustrates the implementation of an optional module of the voice synthesizer of Figure 1 during its training phase; points of interest of said geographical area

[0042] Fig. 5] Figure 5 is a flowchart of a speech synthesis method implemented by the speech synthesizer of Figure 1 according to one embodiment.

[0043] In the remainder of the description, the expression "substantially equal to" is understood as a relationship of equality to plus or minus 10%, that is to say with a variation of at most 10%, more preferably as a relationship of equality to plus or minus 5%, that is to say with a variation of at most 5%.

[0044] Furthermore, subsequently, we consider that a neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.

[0045] More precisely, each layer consists of neurons taking their inputs from the outputs of the neurons in the previous layer, or from the input variables for the first layer.

[0046] Alternatively, more complex neural network structures can be considered with a layer that can be connected to a layer further away than the immediately preceding layer.

[0047] Each neuron is also associated with an operation, that is, a type of processing, to be carried out by said neuron within the corresponding processing layer.

[0048] Each layer is connected to other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a connection between two neurons. It is often a real number, which takes both positive and negative values. In some cases, the synaptic weight is a complex number.

[0049] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the previous layer, each value then being multiplied by the respective synaptic weight of each synapse, or link, between said neuron and the neurons of the previous layer, then applying an activation function, typically a non-linear function, to said weighted sum, and delivering as output of said neuron, in particular to the neurons of the following layer connected to it, the value resulting from the application of the activation function. The activation function makes it possible to introduce non-linearity into the processing carried out by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function are examples of activation functions.

[0050] As an optional addition, each neuron is also able to apply, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.

[0051] A fully connected layer of neurons is one in which the neurons in that layer are each connected to all the neurons in the previous layer.

[0052] Such a type of layer is more often referred to as "fully connected" and sometimes referred to as a "dense layer".

[0053] The electronic device 10 for processing the sound signal, also called subsequently a voice synthesizer, is illustrated in FIG. 1. As can be seen in this FIG. 1, the voice synthesizer 10 comprises a projection module 12 configured to project, via a first monodirectional recurrent neural network RNi, input parameters, characterizing each frame of an input sound signal comprising a human voice, into a latent space (from the English embedding) of parameterizable dimension.

[0054] Such a projection module 12 is also called an encoder, and is therefore capable of “immersing” the input parameters, characterizing each frame of an input sound signal comprising a human voice in a latent space (i.e. a vector space) of parameterizable dimension, for example of size equal to 128 (i.e. capable of comprising 128 components). The RNi neural network, previously trained as will be seen later, supports the non-linear projection of said input parameters into the latent space potentially containing more information than the input parameters since temporal relationships can be incorporated therein thanks to the recurrent and unidirectional nature of the RNi neural network.

[0055] The input parameters suitable for being supported by said projection module 12 are parameters which characterize said input sound signal comprising said human voice and correspond in particular to frequency parameters corresponding to one of the elements belonging to the group comprising at least:

[0056] - a Fourier spectrum of said input sound signal comprising said human voice - a cepstrum defined by Mel frequency cepstral coefficients calculated by a discrete cosine transform applied to the power spectrum of said input sound signal comprising said human voice;

[0057] - the constant Q transform CQT (from the English constant-Q transform) of said input sound signal comprising said human voice

[0058] - etc.

[0059] As illustrated by FIG. 1, the voice synthesizer 10 further comprises a harmonic signal generator 14 configured to generate, via a second neural network RN2, a signal representative of the harmonics of said input sound signal comprising said human voice.

[0060] As an optional addition, such a harmonic signal generator 14 is configured to generate said signal representative of the harmonics of said input sound signal comprising said human voice which is a weighted sum of the signal corresponding to the pitch of said human voice and its harmonics, said weighting being learned beforehand by said second neural network.

[0061] According to a variant of this optional complement, white noise is added to said weighted sum.

[0062] In other words, said harmonic signal generator 14 is configured to produce a signal composed of a sinusoidal signal at the pitch of the voice and its harmonics. The output signal of this generator 14 is a weighted sum of the pitch of the voice and its harmonics. The weighting is learned by the second neural network RN2 during a prior training phase. White noise is added to this harmonic signal to serve as a basis for the stochastic components of the voice. The unvoiced frames or frames not containing voices are composed only of white noise without a harmonic component.

[0063] According to a speech synthesizer architecture according to the known “source-filter” model, the output generated by the harmonic signal generator 14 plays the role of the “source” part of this “source-filter” architecture.

[0064] As illustrated by FIG. 1, the voice synthesizer 10 further comprises a voice synthesizer 16 configured to synthesize, via a third monodirectional recurrent neural network RN3, an encoded voice associated with said human voice from:

[0065] - the projection of said projected input parameters into said latent space of parameterizable dimension provided by said projection module 12, and

[0066] - said signal representative of the harmonics of the input sound signal comprising said human voice and provided by said generator 14. Such a voice synthesizer 16 is also called a decoder, and like the preceding elements 12 and 14 of the voice synthesizer 10 is also based on a recurrent neural network architecture. Such a voice synthesizer 16 therefore has as inputs the output of the projection model 12 and the output of the harmonic signal generator 14. The projection model 12 and the harmonic signal generator 14 are capable of operating in parallel or successively in any order, provided that the output of the projection model 12 and the output of the harmonic signal generator 14 associated with the same frame of the input sound signal are processed simultaneously by said voice synthesizer 16.

[0067] According to a speech synthesizer architecture according to the known “source-filter” model, the speech synthesizer 16 plays the role of the “filter” part of this “source-filter” architecture.

[0068] Thus, unlike the neural source-filter architectures according to the current state of the art, the use of unidirectional recurrent neural networks proposed according to the present invention makes it possible to operate in real time.

[0069] For example, the voice synthesizer 10 according to the present invention is capable of operating with an input sound signal previously broken down into 40ms frames with a look ahead buffer zone of 40ms, essentially due to the extraction of the spectral parameters at the input of the projection module 12 (i.e. the encoder).

[0070] As an optional addition, as shown in dotted lines, said voice synthesizer 10 further comprises an adversarial training module 18 comprising two discriminators Di and D2 configured to be trained simultaneously with said voice synthesizer 10 during its supervised learning phase, according to a zero-sum game, said discriminators Di and D2 being trained to differentiate the real signals from those suitable for being synthesized by said voice synthesizer 10, said voice synthesizer 10 being trained to fool said two discriminators D1 and D2.

[0071] It should be noted that once the learning phase has been carried out, such an adversarial training module 18 is no longer used during the inference phase using the voice synthesizer 10 thus previously trained in an adversarial manner.

[0072] According to an advantageous variant of this optional complement, one of said two discriminators is sensitive to low frequency information and the other of said two discriminators is sensitive to high frequency information.

[0073] In other words, the two discriminators D1 and D2 are trained at the same time as the speech synthesizer 10 to differentiate the real signals from the signals synthesized by the speech synthesizer 10 by creating a zero-sum game where the discriminators D1 and D2 try to differentiate the signals created by the speech synthesizer 10 from the original signals and where the speech synthesizer 10 is trained to fool the two discriminators Di and D2.

[0074] Such a zero-sum game allows the speech synthesizer resulting from said adversarial training to be balanced, and the use of two discriminators, sensitive to low and high frequency information, allows the speech synthesizer 10 to correct its defects over extended frequency ranges, and is particularly effective in removing metallic artifacts from the voice synthesized via the speech synthesizer 10.

[0075] For example, the two discriminators correspond to the discriminators D1 and D2 described in the article by J. Kong et al entitled “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis” NeurIPS 2020, but applied in another context, without combined implementation of a speech synthesizer architecture according to a source-filter model, and the zero-sum game according to the present invention where said discriminators D1 and D2 are trained to differentiate real signals from those suitable for being synthesized by said speech synthesizer 10, said speech synthesizer 10 being trained to fool said discriminators D1 and D2.

[0076] The adversary training implemented via said optional adversary training module 18 is further described below in relation to FIG. 4.

[0077] Such an optional adversary training module 18 therefore makes it possible to improve the quality of synthesis.

[0078] As an optional addition, said voice synthesizer 10 further comprises a reconstruction module 20 configured to reconstruct (inpainting) missing frames of said synthesized encoded voice from frames before or after said missing frames, and / or a band extension module 22 configured to reconstruct a signal comprising said synthesized encoded voice with a sampling frequency higher than that of said input sound signal.

[0079] The optional reconstruction module 20 makes it possible to improve the overall transmission of the synthesized signal via said voice synthesizer 10 by compensating for micro-cuts which would occur without its implementation due to missing frame(s), while the optional band extension module 22 makes it possible to increase the output sampling frequency.

[0080] As an optional addition, at least one of said neural networks of said voice synthesizer 10, for example the first monodirectional recurrent neural network RN1 of the projection module 12 or the second neural network RN2, is also configured to extract the vocal biomarkers of the speaker having generated said input sound signal, after having been previously trained for this purpose in a supervised manner. In a variant not shown, the voice synthesizer 10 is also capable of optionally comprising another neural network dedicated to the extraction of the vocal biomarkers of the speaker having generated said input sound signal after having been previously trained for this purpose in a supervised manner.

[0081] In the example of figure 1, the voice synthesizer 10 comprises an information processing unit 26 formed for example of a memory 28 and a processor 30 associated with the memory 28.

[0082] In the example of Figure 1, the projection module 12, the generator 14 and the voice synthesizer 16, as well as optionally the adversary training module 18, the reconstitution module 20 and the band extension module 22, are each implemented in the form of software, or a software brick, executable by the processor 30. The memory 28 of the voice synthesizer 10 is then capable of storing projection software, harmonic signal generation software, and voice synthesis software, as well as optionally additionally adversary training software, reconstitution software and band extension software. The processor is then capable of executing each of the software among the projection software, the harmonic signal generation software, and the voice synthesis software, as well as optionally additionally the adversary training software, the reconstitution software, and the band extension software.

[0083] In a variant not shown, the projection module 12, the generator 14 and the voice synthesizer 16, as well as, as an optional addition, the adversary training module 18, the reconstitution module 20 and the band extension module 22, are each produced in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array), or an integrated circuit, such as an ASIC (Application Specific Integrated Circuit).

[0084] When the voice synthesizer 10 is produced in the form of one or more software programs, that is to say in the form of a computer program, also called a computer program product, it is furthermore capable of being recorded on a medium, not shown, readable by a computer. The computer-readable medium is for example a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example FLASH or NVRAM) or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.

[0085] Figure 2 illustrates the use 32 of the voice synthesizer of Figure 1. According to this use, a speaker 34 produces a sound signal 36 comprising his own voice. Such a sound signal 36 is first decomposed by means of a known decomposition tool into input parameters 38 characterizing it.

[0086] These input parameters 38 are in particular frequency parameters corresponding to one of the elements belonging to the group comprising at least:

[0087] - a Fourier spectrum of said input sound signal comprising said human voice;

[0088] - a cepstrum defined by Mel frequency cepstral coefficients calculated by a discrete cosine transform applied to the power spectrum of said input sound signal comprising said human voice;

[0089] - the constant Q transform of said input sound signal comprising said human voice.

[0090] These input parameters 38 are provided as input to the projection module 40 (corresponding to the projection module 12 of FIG. 2) which comprises the first monodirectional recurrent neural network RNi configured for said input parameters in a latent space of parameterizable dimension.

[0091] In parallel, said input sound signal 36 comprising said human voice is provided as input to the harmonic signal generator 42 (corresponding to the generator 14 of FIG. 1), said generator 42 providing as output a signal 44 representative of the harmonics of said input sound signal comprising said human voice.

[0092] Said signal 44 representative of the harmonics of said input sound signal comprising said human voice as well as the result provided by the projection module 40 are associated with the same frame of the input sound signal 36 and provided simultaneously as input to the voice synthesizer 46 configured to synthesize, via a third monodirectional recurrent neural network RN3, an encoded voice 48 associated with said human voice of the input sound signal 36.

[0093] Figure 3 illustrates an example 50 of architecture of the voice synthesizer of Figure 1. According to this example 50, the voice synthesizer receives as input the sound signal 52 comprising a human voice.

[0094] As illustrated by the architectural example 50, the projection module 12 optionally firstly comprises a module 54 capable of extracting the parameters which characterize said input sound signal 52 comprising said human voice and correspond in particular to frequency parameters.

[0095] According to example 50 of figure 3 these parameters correspond to the cepstrum defined by Mel-Frequency Cepstral Coefficients (MFCC) calculated by a discrete cosine transform applied to the power spectrum of said input sound signal 52 comprising said human voice.

[0096] The projection module 12 further comprises the first monodirectional recurrent neural network RN1 configured to project these MFCC coefficients into a latent space of configurable dimension, for example of dimension 128 (i.e. capable of comprising 128 components).

[0097] As illustrated by example 50 of figure 3, the first monodirectional recurrent neural network RNi successively comprises a normalization layer LN (from the English Layer-norm), followed by three dense layers, a gate recurrent unit GRU (from the English Gate Recurrent Unit) and a dense layer, so as to provide as output the projection Z of said MFCC parameters in the latent space of parameterizable dimension.

[0098] As an alternative to the GRU unit, a layer with long short-term memory such as LSTM (Long Short Term Memory) or even more classically a recurrent RNN (Recurrent Neural Network) layer is used.

[0099] Note that according to another implementation option, not shown, the projection module 12 does not include the module 54 suitable for extracting the parameters which characterize said input sound signal 52 but directly receives these parameters, determined outside the voice synthesizer, as input to the monodirectional recurrent neural network RNi.

[0100] Furthermore, as illustrated by the example architecture 50, the harmonic signal generator 14 optionally firstly comprises a module 56 capable of extracting the pitch from the input sound signal 52.

[0101] The harmonic signal generator 14 further comprises a second neural network RN2 comprising in particular a Sin Gen layer, configured to generate a sinusoidal signal at the pitch of the voice, followed by a dense layer such that the signal produced by the generator 14 is a signal composed of a sinusoidal signal at the pitch of the voice and its harmonics, this signal corresponding to a weighted sum E of the pitch of the voice and its harmonics. The weighting is learned by the second neural network RN2 during a prior training phase. The output generated E by the harmonic signal generator 14 plays the role of the “source” part of this “source-filter” architecture.

[0102] The projection model 12 and the harmonic signal generator 14 are capable of operating in parallel or successively in any order, provided that the output Z of the projection model 12 and the output E of the harmonic signal generator 14 associated with the same frame of the input sound signal are provided as input and processed simultaneously by said voice synthesizer 16.

[0103] From said inputs Z and E, the voice synthesizer 16, via its third monodirectional recurrent neural network RN3, is configured to synthesize an encoded voice associated with said human voice of said input signal 52. More precisely, the voice synthesizer 16 plays the role of the “filter” part of this “source-filter” architecture of the voice synthesizer according to the present invention. According to example 50 of FIG. 3, the third monodirectional recurrent neural network RN3 of the voice synthesizer 16 successively comprises a first dense layer corresponding in particular to a conventional perceptron and having only the input E as input.

[0104] This first dense layer of the third monodirectional recurrent neural network RN3 is followed by two successive filter units F1 and F2 (from the English Filter Unit). The first filter unit F1 receives two inputs, namely both the Z output of the projection model 12 and the output of the first dense layer, while the second filter unit F2, placed after the first filter unit F1, also receives two inputs, namely both the Z output of the projection model 12 and the output of the first filter unit F1.

[0105] These two filter units F1 and F2 have an identical structure as illustrated at zoom 57 inside said filter unit structure.

[0106] More precisely, each filtering unit F1 or F2 firstly comprises a concatenation layer capable of concatenating the two received inputs represented on the one hand by the letter “y” (corresponding to the output S of the first dense layer of the third monodirectional recurrent neural network RN3 for the first filtering unit F1 or to the output Su of the first filtering unit F1 for the second filtering unit F2), and on the other hand by the letter “Z” corresponding to the output Z of the projection model 12.

[0107] At the output of said concatenation layer, each filtering unit F1 or F2 comprises a GRU gate recurrent unit (from the English Gate Recurrent Unit).

[0108] As an alternative to the GRU unit, a layer with long short-term memory such as LSTM (Long Short Term Memory) or even more classically a recurrent RNN (Recurrent Neural Network) layer is used.

[0109] The GRU unit of each filtering unit F1 or F2 provides two outputs noted in zoom 57 Si and S2 as input each of a dense layer, the input represented by the letter “y”, as specified previously, being added as illustrated within zoom 57 of figure 3 to the output Si of one of these two dense layers.

[0110] For the first filtering unit F1, the outputs Si and S2 indicated in the zoom 57 correspond respectively to the outputs Su and S12, while for the second filtering unit F2, the outputs Si and S2 indicated in the zoom 57 correspond respectively to the outputs S21 and S22. Finally, as illustrated by example 50 of figure 3, the output S21 of the second filtering unit F2 is provided as input to a last dense layer of said third monodirectional recurrent neural network RN3.

[0111] Optionally, as illustrated by Figure 3, the last dense layer of said third monodirectional recurrent neural network RN3 also receives a second input equal to the sum of: the output S12 of the first filtering unit F1 to which is added the output S22 of the second filtering unit F2, to which is further added a random white noise represented by the letter a in particular to serve as a basis for the stochastic components of the voice.

[0112] The output 58 of the last dense layer of said third monodirectional recurrent neural network RN3 corresponds to the encoded channel (i.e. the synthetic voice) provided by the voice synthesizer 16.

[0113] Figure 4 illustrates the principle of adverse training optionally implemented according to the present invention.

[0114] More specifically, in Figure 4 the adversarial training module 60 comprises two discriminators (not shown in Figure 4) configured to be trained simultaneously with said voice synthesizer 10 during its supervised learning phase, according to a zero-sum game, said discriminators being trained to differentiate the real signals 36 generated by a human speaker 34 from those 48 suitable for being synthesized by said voice synthesizer 10, said voice synthesizer 10 being trained to fool said two discriminators.

[0115] For example, the two discriminators correspond to the discriminators D1 and D2 described in the article by J. Kong et al entitled “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis” NeurIPS 2020, but applied in another context without combined implementation of a speech synthesizer architecture according to a source-filter model, and the zero-sum game according to the present invention where said discriminators D1 and D2 are trained to differentiate real signals from those suitable for being synthesized by said speech synthesizer 10, said speech synthesizer 10 being trained to fool said discriminators D1 and D2.

[0116] As indicated previously, advantageously one of said two discriminators is sensitive to low-frequency information and the other of said two discriminators is sensitive to high-frequency information, the use of two discriminators, sensitive to low- and high-frequency information, allows the voice synthesizer 10 to correct its defects over extended frequency ranges, and is particularly effective for removing metallic artifacts from the voice synthesized via the voice synthesizer 10. More precisely, each discriminator of the adversarial training module 60 is configured to provide a PR prediction when the incoming signal is real 36 or a Ps prediction when the incoming signal has been synthesized by the voice synthesizer 10 and the prediction result is then used in a loss function of the voice synthesizer during training. The discriminators and voice synthesizer are trained jointly (i.e.joint learning).

[0117] The loss function of the speech synthesizer is the average, over the different frames of the sample, of the (1- Ps) 2 (Ps being one if a discriminator “says” that the signal is real and zero otherwise).

[0118] The loss function of each discriminator is the sum of the averages of: (1- PR) 2 and Ps 2 (Ps and PR being one if the discriminator predicts that the signal is real and zero otherwise).

[0119] The loss function taking this prediction into account during the supervised learning phase of the adversarial training module, comprising the two discriminators, and jointly of the voice synthesizer 10 is advantageously adaptive. In particular, during learning, a time average of the spectrograms is performed, followed by a sliding average on the frequencies (i.e. second axis of the spectrogram) in order to smooth the frequency distribution of the signal used for each iteration of the supervised learning.

[0120] Then, this smoothed spectrum x is transformed into v using the following activation function: v = 1,3 - sigmoid (Iog10(amplitude(x))) in order to calculate, weights of weighting by frequency, in the loss function using the spectrogram, this loss function using the spectrogram corresponding to the sum for each time / frequency value of the amplitude of the spectrogram of the L1 norm (absolute value) of the difference between the predicted spectrogram and the reference spectrogram.

[0121] Such processing allows an adaptation, at each stage of learning, to the signal actually used and therefore allows learning to focus on the relevant frequencies of the frame in question (i.e. the current packet of sound data (from the English batch) in order to have learning resulting in a better quality model.

[0122] A voice synthesis method 70 implemented via said voice synthesizer 10 will now be explained with reference to FIG. 5 showing a flowchart of the steps of this method 70.

[0123] Generally, said method 70 comprises during its inference phase 72, implemented by a trained voice synthesizer V_E having an architecture according to a neural source-filter model, a projection step 74 P, via a first monodirectional recurrent neural network of said voice synthesizer, input parameters, characterizing each frame of an input sound signal S_E comprising a human voice, in a latent space of configurable dimension.

[0124] Said method 70 also comprises a step 76 of generation G, via a second neural network of said trained voice synthesizer V_E, of a signal representative of the harmonics of said input sound signal S_E comprising said human voice.

[0125] The steps of projection 74 and generation 76 of the harmonic signal are suitable for being implemented in parallel as illustrated according to FIG. 5, or in a manner not shown, successively in any order, provided that the output of the projection step 74 and the output of the step 76 of generation of the harmonic signal associated with the same frame of the input sound signal S_E are processed simultaneously during a subsequent step 78 of voice synthesis SYNTH as such via a third monodirectional recurrent neural network, of an encoded voice associated with said human voice of the input sound signal S_E.

[0126] Optionally, as shown in dotted lines, said inference phase 72 further comprises a step 80 of reconstitution R_T_M of the missing frames of said synthesized encoded voice from frames before or after said missing frames, and / or a step 82 of band extension E_B to reconstitute a signal comprising said synthesized encoded voice with a sampling frequency higher than that of said input sound signal S_E, and / or a step 84 of extraction BIO of the vocal biomarkers of the speaker having generated said input sound signal via at least one of said neural networks of said voice synthesizer.

[0127] As an optional addition, when the voice synthesizer used has not been previously trained, the method 70 also comprises a prior learning phase 88 (i.e. a training phase for each element of the voice synthesizer).

[0128] According to an optional aspect of this optional supplement, said learning phase 88 comprises a step 90 of adversary training E_A via said adversary training module previously described.

[0129] Those skilled in the art will understand that the invention is not limited to the embodiments described, nor to the particular examples of the description, the embodiments and variants mentioned above being suitable for being combined with each other to generate new embodiments of the invention.

[0130] The present invention thus makes it possible, as experimentally verified, by virtue of the specific structure of the voice synthesizer according to the present invention, to provide an improvement over current vocoding with equivalent input parameters. In particular, the source-filter type neural architecture of the voice synthesizer according to the present invention makes it possible to have a "light" (i.e. frugal) solution in computational footprint and memory compared to current AI solutions, and opens up the possibility of implementing other "functionalities" such as, for example, reconstruction by Inpainting, bandwidth extension (this is an addition of information by prediction) or even "voice conversion" (i.e. passage from one voice to another voice (a priori associated with existing speakers) or even "voice transformation) (i.e.switching from one voice to another voice (which may more generally be that of a virtual speaker)) By delegating voice synthesis to AI approaches with recurrent neural networks at each stage of the source-filter model rather than to solutions designed by "experts" (i.e. solutions resulting from modeling), the solution according to the present invention makes it possible in particular to compensate for a potential loss of information linked to the reduction in the number of input parameters, particularly in the context of a low-speed application.

[0131] The choice of a "biased" source-filter architecture (i.e. a neural network architecture that has been "biased" or rather specialized towards voice synthesis by adding speech modeling by a sum of harmonics) towards voice synthesis makes it possible to reduce the digital footprint, in terms of memory and calculation, of the implemented neural networks.

[0132] Finally, optionally, the practical difficulty of training neural networks in such an open task with few input parameters for a large number of possible temporal signals is advantageously mitigated by the use of adversarial training where two discriminators are trained to recognize the signals generated by the speech synthesizer according to the present invention or jointly said speech synthesizer is trained to fool the discriminators.

[0133] Thus, the voice synthesizer and the corresponding voice synthesis method are suitable for integration into digital radiocommunication applications, voice over IP (Internet Protocol) Vol P, audio compression, etc.

Claims

CLAIMS 1. Voice synthesizer (10) having an architecture according to a neural source-filter model characterized in that it comprises at least: - a projection module (12) configured to project, via a first monodirectional recurrent neural network, input parameters, characterizing each frame of an input sound signal comprising a human voice, into a latent space of configurable dimension; - a harmonic signal generator (14) configured to generate, via a second neural network, a signal representative of the harmonics of said input sound signal comprising said human voice; - a voice synthesizer (16) configured to synthesize, via a third monodirectional recurrent neural network, an encoded voice associated with said human voice from: - the projection of said projected input parameters into said latent space of parameterizable dimension provided by said projection module, and - said signal representative of the harmonics of the input sound signal comprising said human voice and provided by said generator; said signal representative of the harmonics of said input sound signal comprising said human voice being a weighted sum of the signal corresponding to the pitch of said human voice and its harmonics, said weighting being learned beforehand by said second neural network.

2. A speech synthesizer (10) according to claim 1, wherein white noise is added to said weighted sum.

3. Voice synthesizer (10) according to any one of the preceding claims, wherein said input parameters characterizing said input sound signal comprising said human voice are frequency parameters corresponding to one of the elements belonging to the group comprising at least: - a Fourier spectrum of said input sound signal comprising said human voice; - a cepstrum defined by Mel frequency cepstral coefficients calculated by a discrete cosine transform applied to the power spectrum of said input sound signal comprising said human voice; - the constant Q transform of said input sound signal comprising said human voice.

4. Voice synthesizer (10) according to any one of the preceding claims, further comprising an adversarial training module (18) comprising two discriminators configured to be trained simultaneously with said voice synthesizer during its supervised learning phase, according to a zero-sum game, said discriminators being trained to differentiate real signals from those suitable for being synthesized by said voice synthesizer, said voice synthesizer being trained to fool said two discriminators.

5. A speech synthesizer (10) according to claim 4, wherein one of said two discriminators is sensitive to low frequency information and the other of said two discriminators is sensitive to high frequency information.

6. A speech synthesizer (10) according to any preceding claim, further comprising: - a reconstruction module (20) configured to reconstruct missing frames of said synthesized encoded voice from frames before or after said missing frames, and / or - a band extension module (22) configured to reconstruct a signal comprising said synthesized encoded voice with a sampling frequency higher than that of said input sound signal.

7. A speech synthesizer (10) according to any preceding claim, wherein at least one of said neural networks of said speech synthesizer is also configured to extract speech biomarkers of the speaker who generated said input sound signal.

8. Method (70) of voice synthesis implemented by a voice synthesizer having an architecture according to a neural source-filter model according to any one of the preceding claims, said method comprising the following steps: - projection (74), via a first monodirectional recurrent neural network of said voice synthesizer, of the input parameters, characterizing each frame of an input sound signal comprising a human voice, into a latent space of configurable dimension; - generation (76), via a second neural network of said voice synthesizer, of a signal representative of the harmonics of said input sound signal comprising said human voice; - voice synthesis (78), via a third monodirectional recurrent neural network, of an encoded voice associated with said human voice from: - the projection of said projected input parameters into said latent space of parameterizable dimension provided by said projection module, and - said signal representative of the harmonics of the input sound signal comprising said human voice and provided by said generator; said signal representative of the harmonics of said input sound signal comprising said human voice being a weighted sum of the signal corresponding to the pitch of said human voice and its harmonics, said weighting being learned beforehand by said second neural network.

9. Program comprising software instructions which, when executed by said voice synthesizer, implement a voice synthesis method according to claim 8.