Speech synthesis method and associated synthesis device
Neural networks, particularly MDN networks, are used to enhance voice synthesis after MELP coding, addressing the challenge of reconstructing natural voice signals by modeling residual linear prediction signals, achieving improved signal fidelity and naturalness.
Patent Information
- Application Number
- PCT/EP2025/058945
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2025-04-02
- Publication Date
- 2025-10-09
AI Technical Summary
Existing voice synthesis methods using MELP coding fail to reconstruct voice signals with all the characteristics of the original signal after compression, leading to synthetic signals that are not sufficiently natural.
A voice synthesis method employing neural networks, specifically MDN networks, to model the probability density of residual linear prediction signals, generating an excitation signal that closely resembles the original voice signal, using techniques like Gaussian mixture models and preprocessing to enhance the synthesis process.
The method effectively reconstructs a voice signal closer to the original, incorporating natural temporal variability and reducing the robotic aspect, while maintaining compatibility with MELP decoding processes.
Smart Images

Figure EP2025058945_09102025_PF_FP_ABST
Abstract
Description
[0001] Voice synthesis method and associated synthesis device
[0002] This patent application claims the benefit of document FR 24 03377 filed on 02 / 04 / 2024 which is incorporated by reference.
[0003] The present invention relates to a voice synthesis method. The present invention relates to a voice synthesis device. The present invention also relates to a computer program product and a corresponding information medium.
[0004] In the field of radio communications, voice signal exchanges often have to be carried out.
[0005] To do this, the voice signal is often compressed to limit the amount of information to be exchanged.
[0006] In particular, it is known to use mixed-excitation linear prediction coding as a compression technique.
[0007] Mixed excitation linear prediction coding is more often referred to as MELP coding, which refers to the corresponding English term for “Mixed Excitation Linear Prediction”.
[0008] This coding is notably described in the document “The 1200 and 2400 bit / s NATO Interoperable Narrow Band Voice Coder”, STANAG N°4591, NATO Standardization Agency.
[0009] However, the reconstruction of the signal thus encoded by MELP decoding leads to a synthetic signal which does not present all the characteristics of the original voice signal.
[0010] There is therefore a need for a voice synthesis method that can reconstruct a voice signal closer to the original voice signal within the framework of compression of the original voice signal by MELP coding.
[0011] For this purpose, the description describes a method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of:
[0012] - obtaining a coded flow according to a linear prediction coding with mixed excitation,
[0013] - extraction of parameters from the coded stream,
[0014] - application of at least one neural network on the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream,
[0015] - generation of an estimate of the residual linear prediction signal from the obtained probability density, to obtain an excitation signal, and
[0016] - voice synthesis from the excitation signal. According to particular embodiments, the synthesis method has one or more of the following characteristics, taken in isolation or in all technically possible combinations:
[0017] - during the application step, at least one neural network is applied to all the extracted parameters.
[0018] - the probability density of the residual linear prediction signal is modeled by a Gaussian mixture model, the at least one neural network predicting, for each Gaussian of the Gaussian mixture model, at least one distribution parameter and the weight of the Gaussian.
[0019] - one neural network predicts the distribution parameters and another neural network predicts the weights of each Gaussian.
[0020] - each neural network predicts a respective parameter of each Gaussian.
[0021] - the method comprises a step of determining the fundamental frequency, the neural network used during the application step depending on a frequency interval in which the determined fundamental frequency is found.
[0022] - the neural network used during the application step depends on the type of voicing.
[0023] - the at least one neural network comprises first layers forming an initial network, the initial network being a convolutional neural network, a long short-term memory neural network or a gated recurrent neural network.
[0024] The description also describes a device for synthesizing a voice signal, the synthesis device being suitable for:
[0025] - obtain a coded flow according to a mixed excitation linear prediction coding,
[0026] - extract parameters from the coded stream,
[0027] - apply at least one neural network to the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream,
[0028] - generate an estimate of the residual linear prediction signal from the obtained probability density, to obtain an excitation signal, and
[0029] - synthesize the voice from the excitation signal.
[0030] The description also provides a computer program product comprising instructions which, when the program is executed by a computer.
[0031] The description also describes a computer-readable medium comprising instructions which, when executed by a computer. In this description, the expression "suitable for" means indifferently "adapted for", "adapted to" or "configured for".
[0032] Characteristics and advantages of the invention will appear on reading the description which follows, given solely by way of non-limiting example, and made with reference to the appended drawings, in which:
[0033] - Figure 1 is a representation of a voice synthesis device,
[0034] - Figure 2 is an example of a neural network used by the speech synthesis device,
[0035] - Figure 3 is another example of a neural network used by the speech synthesis device, and
[0036] - Figure 4 is yet another example of a neural network used by the speech synthesis device.
[0037] A speech synthesis device is shown schematically in Figure 1.
[0038] The voice synthesis device is capable of synthesizing a voice signal from a coded incident stream.
[0039] In the example described, the voice synthesis device comprises a receiver 12, a calculator 14 and a transmitter 16.
[0040] The receiver 12 is suitable for receiving external signals.
[0041] For example, receiver 12 is an antenna.
[0042] The computer 14 is capable of implementing a method of synthesizing a voice signal which will be described later.
[0043] The calculator 14 is an electronic circuit designed to manipulate and / or transform data represented by electronic or physical quantities in registers of the calculator and / or memories into other similar data corresponding to physical data in the memories of registers or other types of display devices, transmission devices or storage devices.
[0044] As specific examples, the computer 14 is implemented in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array), or an integrated circuit, such as an ASIC (Application Specific Integrated Circuit). It could also be envisaged to use a digital processor or DSP (Digital Signal Processor), a GPP processor (General Purpose Processor) or a neural processor.
[0045] Alternatively, when the method is carried out in the form of one or more software programs, i.e. in the form of a computer program, also called a computer program product, it is also capable of being recorded on a medium, not shown, that is readable by a computer. The computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example FLASH or NVRAM) or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.
[0046] The transmitter 16 is suitable for transmitting the signal synthesized by the computer 14.
[0047] For example, transmitter 16 is a loudspeaker.
[0048] An example of a method for synthesizing a voice signal is now described.
[0049] The synthesis method is a method aimed at obtaining a synthesized voice signal from a coded incident stream.
[0050] For this, according to the example described, the synthesis method comprises an obtaining step, an extraction step, an application step, a generation step and a synthesis step.
[0051] During the obtaining step, the calculator 14 obtains a stream coded according to a MELP coding.
[0052] The coded stream here corresponds to the result of applying MELP coding to a voice signal that we wish to transmit.
[0053] The transmitted voice signal is the signal that the synthesis process aims to synthesize.
[0054] In one embodiment, the receiver 12 receives the coded stream and transmits it to the computer 14. The computer 14 thus obtains the coded stream by interaction with the receiver 12.
[0055] The coded stream is a set of frames, each frame being encoded on a predefined number of bits according to the type of MELP coding applied.
[0056] For example, for 2400 bps MELP encoding, one frame is quantized to 54 bits; for 1200 bps MELP encoding, three frames are quantized to 81 bits; and for 600 bps MELP encoding, four frames are quantized to 54 bits.
[0057] During the extraction step, the calculator 14 extracts parameters from the coded stream. The parameters of a frame that can be extracted are chosen from the following parameters:
[0058] • the LSF coefficients representing the spectral envelope of the coded stream. The abbreviation LSF refers to the corresponding English term for “Une Spectral Frequencies” which means spectral line frequencies. The LSF coefficients are therefore the spectral line frequency coefficients. • the amplitude Fourier coefficients,
[0059] • two gains, each gain corresponding to a respective half-frame,
[0060] • two jointly quantified parameters, namely the pitch parameters (literally height parameter, the pitch parameter corresponding to an estimate of the fundamental frequency associated with the vibration frequency of the vocal cords) and overall voicing parameters (literally global voicing parameter),
[0061] • a jitter parameter (jitter parameter in French), a jitter being added to avoid an overly robotic aspect of the synthesized signal, and
[0062] • the voicing indicator parameters by frequency band, more often referred to as bandpass voicing parameters.
[0063] These parameters and their expressions are defined by the standard mentioned above.
[0064] Depending on the embodiments, the calculator 14 extracts one or more of these parameters.
[0065] The calculator 14 generally extracts at least the coefficients, the pitch or voicing and the gains for each frame.
[0066] During the application step, the calculator 14 applies at least one neural network to the extracted parameters.
[0067] As described later with reference to Figures 2 to 4, several topologies are considered with one or more neural networks and different output parameters.
[0068] As a note, in the presence of several neural networks, it could be considered that each neural network is a neural subnetwork of a neural network consisting of all the neural subnetworks. In the following, this term is not used to simplify the reader's reading.
[0069] The at least one neural network makes it possible, in the example described, to obtain the probability density of the residual linear prediction signal associated with the coded stream.
[0070] As will become clearer later, this result is obtained by using one or more MDN-type neural networks and the output of this or these networks provides a parametric description of the probability density of the target data (more precisely the LPC residual signal).
[0071] The size of the residual excitation signal is 2xN0, NO being the size of a period associated with the value of the pitch fO.
[0072] Linear prediction modeling allows to obtain a set of linear prediction coefficients by minimizing the prediction error. The residual signal corresponds to the difference between the speech signal and the predicted signal, and is denoted LPC residual signal in the following.
[0073] Linear prediction coefficients are more commonly referred to as LPC coefficients. The abbreviation LPC refers to the corresponding English term "Linear Prediction Coefficients" which literally means "linear prediction coefficients".
[0074] The model used for probability is a Gaussian mixture model.
[0075] From a notational point of view, Ng will denote the number of Gaussians (more precisely the number of weights, namely 1 per Gaussian), size L of a vector will be 2xN0 (NO being the size of a period) and the size P of the output vector of the neural network will be Ng + 2 x Ng x L.
[0076] It can be noted here that Gaussians are multidimensional (size L), and are represented in the figures as one-dimensional Gaussians for ease of understanding.
[0077] Such a model is often referred to by the acronym GMM, which refers to the corresponding English term “Gaussian Mixture Model”.
[0078] A Gaussian mixture model is a statistical model expressed as a mixture density. It is usually used to parametrically estimate the distribution of random variables by modeling them as the sum of several Gaussians, called kernels. The aim is to determine the variance, mean, and amplitude of each Gaussian. These parameters are optimized according to the maximum likelihood criterion in order to approximate the desired distribution as closely as possible.
[0079] Due to such an output, the at least one neural network considered here is a probabilistic neural network.
[0080] From a very schematic point of view, a classic neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.
[0081] More precisely, each layer consists of neurons taking their inputs from the outputs of the neurons in the previous layer, or from the input variables for the first layer.
[0082] Alternatively, more complex neural network structures can be considered with a layer that can be connected to a layer further away than the immediately preceding layer.
[0083] Each neuron is also associated with an operation, that is, a type of processing, to be performed by the neuron within the corresponding processing layer. Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a connection between two neurons. It is often a real number, which takes both positive and negative values. In some cases, the synaptic weight is a complex number.
[0084] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the previous layer, each value then being multiplied by the respective synaptic weight of each synapse, or link, between said neuron and the neurons of the previous layer, then applying an activation function, typically a non-linear function, to said weighted sum, and delivering as output of said neuron, in particular to the neurons of the following layer connected to it, the value resulting from the application of the activation function. The activation function makes it possible to introduce non-linearity into the processing carried out by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function are examples of activation functions.
[0085] As an optional addition, each neuron is also able to apply, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.
[0086] According to the example described, each neural network here is a mixed-density neural network.
[0087] Such a neural network is often referred to by the acronym MDN, which stands for Mixture Density Network.
[0088] Such networks are notably described in the document entitled "Mixture Density Networks", CM Bishop, Neural Computing Research Group Report, Feb. 1994.
[0089] As shown above, unlike a traditional neural network, the MDN network does not produce as output an estimate of the target vector of size L, but the most likely generative statistical model (here a GMM).
[0090] Three topologies are possible: a first topology consists of using a single network to predict all the parameters of the GMM model, a second topology consists of using one network per type of parameters (weights, means, variances), and a third topology consists of using one network for the weights and one network for the parameters of the Gaussians (means and variances).
[0091] A more complete description is now referred to Figures 2 to 4.
[0092] According to a first topology illustrated by Figure 2, a single neural network is applied to all the extracted parameters. According to the example shown, the single neural network determines at output, for each Gaussian (three in the case shown), the variance, the mean and the weight. These quantities are noted o, p and W with an index i in Figure 2.
[0093] The probability density is modeled by a weighted sum of Gaussians (GMM model). Each Gaussian (characterized by its mean and variance) is associated with a weight.
[0094] From these parameters, the probability density of the residual LPC signal associated with the coded stream is generated.
[0095] The generative model thus obtained at output makes it possible to synthesize a residual synchronous LPC pitch excitation.
[0096] This last stage of preprocessing being the same for all the topologies presented, it is not repeated for the second and third topologies.
[0097] According to a second topology corresponding to Figure 3, two distinct neural networks are used.
[0098] The first neural network predicts the weights of each Gaussian and the second neural network predicts the distribution parameters of each Gaussian.
[0099] More precisely, in the example illustrated with three Gaussians, the first neural network applied to the set of extracted parameters gives as output the weight of the first Gaussian, the weight of the second Gaussian and the weight of the third Gaussian.
[0100] The second neural network outputs the variance and mean of the first Gaussian, the variance and mean of the second Gaussian, and the variance and mean of the third Gaussian.
[0101] According to a second topology corresponding to Figure 4, three distinct neural networks are used, each neural network predicting a respective parameter of each Gaussian.
[0102] More precisely, in the example illustrated with three Gaussians, the first neural network applied to the set of extracted parameters gives as output the weight of the first Gaussian, the weight of the second Gaussian and the weight of the third Gaussian.
[0103] The second neural network outputs the average of the first Gaussian, the average of the second Gaussian and the average of the third Gaussian.
[0104] It can be specified here that each Gaussian is multidimensional in the sense that the variance and the mean are vectors. The weights are nevertheless scalars. Finally, the third neural network gives as output the variance of the first Gaussian, the variance of the second Gaussian and the variance of the third Gaussian.
[0105] In each of the topologies, the applied neural network can be obtained by any known learning technique.
[0106] Learning is done from an aligned corpus of MELP parameters and periods of the LPC residual signal calculated from the original signal.
[0107] More precisely, the training corpus is built using MELP modeling.
[0108] The residual LPC signal is cut into 2xN0 vectors, using the pitch provided by the MELP coding and after synchronization by maximizing the inter-correlation between the residual LPC signal and the excitation signal synthesized by the MELP coding.
[0109] Such an approach makes it possible to associate with each set of MELP parameters a portion of the residual signal of duration 2 x NO, and thus to constitute the learning corpus.
[0110] Alternatively, any other approach suitable for achieving pitch-synchronous segmentation of the residual signal can be used here.
[0111] According to one embodiment, to improve prediction performance, a stage for preprocessing the inputs of the neural network may also be provided.
[0112] The preprocessing stage is therefore located upstream of the input layer of the neural network.
[0113] As the name suggests, the preprocessing stage implements a preprocessing operation of some of the extracted parameters.
[0114] In addition, the preprocessing stage here performs a conversion which may vary from one embodiment to another.
[0115] In one example, the preprocessing operation involves a conversion of LSF coefficients into cepstral coefficients. In practice, such a conversion is usually done in two steps with an intermediate passage to an LPC representation.
[0116] Cepstral coefficients are the coefficients of an LPCC representation of the envelope.
[0117] Cepstral coefficients are therefore more often referred to as LPCC coefficients in reference to the English term “linear prediction cepstral coefficients”.
[0118] According to another example or in addition, the preprocessing operation involves a conversion of incident coefficients into the frequency domain. By incident coefficients, it is understood that the set of at least one layer takes as input incident coefficients of a representation.
[0119] The incident coefficients here are the cepstral coefficients.
[0120] Also, in the example described, the coefficients obtained after passage into the frequency domain are log-spectral envelope coefficients, more simply called here DCT coefficients. These coefficients are obtained by DCT transformation of the LPCC coefficients.
[0121] The abbreviation DCT refers to the technique that enabled the conversion to the frequency domain and refers to the corresponding English name of “Discrete Cosine Transform” which means discrete cosine transform.
[0122] Thus, with the proposed examples, instead of the LSF coefficients, we can use the LPCC coefficients or the DCT coefficients as input to the neural network.
[0123] These changes in representation for the envelope allow for a more easily representable representation than the representation using LSF coefficients.
[0124] In each case, the application of one or each neural network thus makes it possible to obtain the probability density of the residual LPC signal.
[0125] During the generation step, the computer 24 generates the excitation signal (estimated signal of the residual prediction signal) as a function of the probability density obtained.
[0126] More precisely, the calculator 24 generates an estimate of the residual LPC signal over the horizon of two pitch periods based on the probability density obtained.
[0127] This signal is to be compared to the case of the synthesis carried out during MELP decoding. Indeed, the excitation signal of the LPC model is obtained from a simplifying hypothesis, namely the presence of a localized pulse and the random nature of the phase for the unvoiced frequency bands.
[0128] In such decoding, a phase dispersion filter is used to simulate the shape of the glottal pulse, this filter being a whitened triangular filter.
[0129] The neural network proposed in this method can simultaneously predict all the information, so that the dispersion filter is no longer necessary.
[0130] MELP decoding is thus based on the assumption of an invariant triangular shape, whereas the use of the network makes it possible to predict a shape closer to a natural signal.
[0131] During the synthesis step, the computer 14 synthesizes the voice from the residual signal obtained during the generation step. Such synthesis is done by applying MELP decoding to the excitation signal obtained at the end of the generation step.
[0132] More precisely, the modeling over two consecutive periods of the residual linear prediction signal is exploited.
[0133] The target vectors (and thus the vectors generated from the GMM models predicted by the MDN neural network) are weighted by an analysis window during training, which allows for an addition-overlap type synthesis.
[0134] Such a synthesis is often referred to by the acronym OLA, which refers to the corresponding English term "OverLap-and-Add".
[0135] An example of such a window is a square root of a Hanning window.
[0136] This window is again applied to the predicted vector before the OLA synthesis, which ultimately amounts to applying a Hanning window to the entire analysis-synthesis cycle, the analysis here being associated with the learning phase.
[0137] The process thus makes it possible to obtain the voice signal to be synthesized from the original MELP stream (coded by the MELP technique).
[0138] In the described method, the use of an MDN network makes it possible to generate a generative model of the excitation signal and not the excitation signal directly.
[0139] This feature allows temporal variability to be taken into account during synthesis, which contributes to the natural rendering of the synthesized voice.
[0140] Indeed, the use of the generative model makes it possible to generate a controlled temporal variability of the excitation signal from the same set of MELP parameters and thus allows the synthesis of an excitation signal closer to a natural signal by proposing a better modeling of its different components.
[0141] The implementation of the synthesis method is also relatively easy since the decoding remains identical to that of a MELP decoder with the exception of the generation of the excitation signal.
[0142] Other embodiments benefiting from the same advantages can be envisaged.
[0143] In the topologies described above, the neural network has a multi-layer perceptron-like structure.
[0144] Such a structure is often referred to by the acronym MLP, which refers to the corresponding English term “Multilayer Perceptron”.
[0145] To better take into account the temporal trajectories of the excitation periods, a recurrent neural network structure could also be used, particularly for certain layers. Advantageously, the layers concerned are the first layers, so that the neural network used during the application step comprises a neural network having an initial neural network, this initial network forming a recurrent neural network.
[0146] A recurrent neural network is an artificial neural network that exhibits recurrent connections.
[0147] A recurrent neural network is thus made up of interconnected neurons interacting non-linearly and for which there is at least one cycle in the structure. The units are connected by synapses which have a weight. The output of a neuron is a non-linear combination of its inputs.
[0148] Such a neural network is often referred to by the acronym RNN, which refers to the corresponding English term for “Recurrent Neural Network”.
[0149] In this case, this means that the neural network takes into account past and future frames to make a prediction at a given time in addition to the current frame (corresponding to the prediction time).
[0150] According to another embodiment, the or each initial neural network is a long short-term memory network.
[0151] Such a network is more often referred to by the acronym LSTM, which corresponds to the corresponding English term for “Long Short-Term Memory”.
[0152] According to yet another embodiment, the or each initial neural network is a gated recurrent neural network.
[0153] Such a network is more often referred to by the acronym GRU, which corresponds to the corresponding English term for “Gated Recurrent Unit”.
[0154] Convolutional neural network-type structures are also possible to form part of the network.
[0155] Such a network is more often referred to by the acronym CNN, which corresponds to the corresponding English term for “Convolutional Neural Networks”.
[0156] Thus, according to the embodiments, the at least one initial neural network comprises a CNN network, an LSTM network or a GRU network.
[0157] It is also possible to take into account other elements at the level of neural networks.
[0158] For example, according to one embodiment, the method comprises a step of determining the fundamental frequency, the neural network used during the application step depending on a frequency interval in which the determined fundamental frequency is located. This amounts to using several MDN networks according to predefined pitch ranges (which condition the size of the excitation periods.
[0159] According to one embodiment, the neural networks differ for all of the coefficients.
[0160] According to another embodiment, the neural network differs, for example, for the output layer which can then be optimized for a data subspace (conditioned by the pitch and / or voicing information).
[0161] For example, in this context, learning includes two learnings.
[0162] The first learning is a global learning carried out on the entire corpus using an output layer sized for the longest pitch period (when the target vector is shorter, the target vector is completed on both sides with zeros).
[0163] The second learning is performed by subspace by freezing all or part of the first layers of the network. The last layer is then resized at the output using the largest size of the target vector for the interval considered.
[0164] The splitting according to the pitch value can be done as follows: the boundaries of the intervals are defined so as to have a balanced distribution of the training data in each interval.
[0165] In some cases it may be advantageous for the defined intervals to have an overlap.
[0166] In practice, a number of 4 intervals can be chosen. Such a division makes it possible in particular to avoid the presence in the same interval of a pitch frequency and its double or half frequency.
[0167] In another example, the neural network used in the application step depends on the type of voicing from the MELP.
[0168] More precisely, an optimization of the last layer based on voicing information can be achieved.
[0169] For each neural network obtained by one of the previous techniques (global learning or pitch interval learning), the neural network is re-optimized by freezing all or part of the first layers (at least the last layer is optimized) for the data corresponding to a predefined voicing criterion.
[0170] A possible voicing criterion consists of defining, for example, three voicing levels based on the MELP voicing vector (5 sub-bands), namely:
[0171] • unvoiced or weakly voiced: (0,0,0,0,0) and (1,0,0,0,0)
[0172] • voiced (1,1,0,0,0) and (1,1,1,0,0), and • strongly voiced (1,1,1,1,0) and (1,1,1,1,1).
[0173] In this specific example, neural networks can be considered for each different pair of pitch intervals and voicing types. Since there are 4 pitch intervals and 3 voicing types, this leads to 12 neural networks that differ only in the output layer. This helps to limit the total number of coefficients.
[0174] Alternatively, it may be possible to consider using separate neural networks but at the cost of much greater memory space involved.
[0175] In each case, the voice synthesis process which allows to reconstruct a voice signal closer to the original voice signal within the framework of a compression of the original voice signal by MELP coding.
Claims
CLAIMS 1. Method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of: - obtaining a coded flow according to a linear prediction coding with mixed excitation, - extraction of parameters from the coded stream, - application of at least one neural network on the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream, - generation of an estimate of the residual linear prediction signal from the obtained probability density, to obtain an excitation signal, and - voice synthesis from the excitation signal.
2. Synthesis method according to claim 1, in which, during the application step, the at least one neural network is applied to all of the extracted parameters.
3. Synthesis method according to claim 1 or 2, in which the probability density of the residual linear prediction signal is modeled by a Gaussian mixture model, the at least one neural network predicting, for each Gaussian of the Gaussian mixture model, at least one distribution parameter and the weight of the Gaussian.
4. Synthesis method according to claim 3, in which a neural network predicts the distribution parameters and another neural network predicts the weights of each Gaussian.
5. Synthesis method according to claim 3, in which each neural network predicts a respective parameter of each Gaussian.
6. Synthesis method according to any one of claims 1 to 5, in which the method comprises a step of determining the fundamental frequency, the neural network used during the application step depending on a frequency interval in which the determined fundamental frequency is found.
7. Synthesis method according to any one of claims 1 to 6, in which the neural network used during the application step depends on the type of voicing.
8. Synthesis method according to any one of claims 1 to 7, in which the at least one neural network comprises first layers forming an initial network, the initial network being a convolutional neural network, a long short-term memory neural network or a gated recurrent neural network.
9. Device for synthesizing a voice signal, the synthesis device being suitable for: - obtain a coded flow according to a mixed excitation linear prediction coding, - extract parameters from the coded stream, - apply at least one neural network to the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream, - generate an estimate of the residual linear prediction signal from the obtained probability density, to obtain an excitation signal, and - synthesize the voice from the excitation signal.
10. Computer program product comprising instructions which, when the program is executed by a computer, cause the latter to implement the steps of a method according to any one of claims 1 to 8.
11. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the steps of a method according to any one of claims 1 to 8.
Citation Information
Patent Citations
METHOD AND DEVICE FOR THE PRESSURE GASIFICATION OF PULVERULENT FUELS
FR2403377A1
Voicing analysis in a linear predictive speech coder
EP1037197A2
Artificial intelligence based audio coding
US11437050B2
Process and device for creating comfort noise in a digital speech transmission system
US5812965A