Voice synthesis method and associated synthesis device

The voice synthesis method uses neural networks to enhance MELP decoding by predicting residual linear prediction signals, achieving a more natural-sounding voice synthesis by generating an excitation signal with improved temporal variability.

FR3160801A1Pending Publication Date: 2025-10-03THALES SA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
FR2024003377
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing voice synthesis methods using MELP coding fail to reconstruct voice signals with all the characteristics of the original signal after compression, leading to synthetic signals that are not sufficiently close to the original.

Method used

A voice synthesis method utilizing neural networks to extract parameters from a coded stream, apply a Gaussian mixture model to predict the residual linear prediction signal, and generate an excitation signal for improved voice synthesis, incorporating preprocessing steps like conversion to cepstral coefficients and frequency domain representation.

Benefits of technology

The method effectively reconstructs a voice signal closer to the original by generating a natural-sounding excitation signal, accounting for temporal variability and improving the quality of synthesized voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for synthesizing voice and associated synthesis device The present invention relates to a method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of: - obtaining a stream coded according to a mixed excitation linear prediction coding, - extracting the parameters of the coded stream, - applying at least one neural network to the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream, - generating an estimate of the residual linear prediction signal estimated from the probability density obtained, to obtain an excitation signal, and - synthesizing the voice from the excitation signal. Figure: figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Voice synthesis method and associated synthesis device

[0001] The present invention relates to a voice synthesis method. The present invention relates to a voice synthesis device. The present invention also relates to a computer program product and a corresponding information medium.

[0002] In the field of radio communications, voice signal exchanges often have to be carried out.

[0003] For this, a compression of the voice signal is often carried out to limit the quantity of information to be exchanged.

[0004] In particular, it is known to use mixed excitation linear prediction coding as a compression technique.

[0005] Mixed excitation linear prediction coding is more often referred to as MELP coding, which refers to the corresponding English term “Mixed Excitation Linear Prediction”.

[0006] This coding is notably described in the document “The 1200 and 2400 bit / s NATO Interoperable Narrow Band Voice Coder”, STANAG No. 4591, NATO Standardization Agency.

[0007] However, the reconstruction of the signal thus coded by MELP decoding leads to a synthetic signal which does not have all the characteristics of the original voice signal.

[0008] There is therefore a need for a voice synthesis method which makes it possible to reconstruct a voice signal closer to the original voice signal within the framework of compression of the original voice signal by MELP coding.

[0009] For this purpose, the description describes a method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of:

[0010] - obtaining a coded stream according to a mixed excitation linear prediction coding,

[0011] - extraction of parameters from the coded stream,

[0012] - application of at least one neural network on the extracted parameters for obtain the probability density of the residual linear prediction signal associated with the coded stream,

[0013] - generation of an estimate of the residual linear prediction signal from the probability density obtained, to obtain an excitation signal, and

[0014] - voice synthesis from the excitation signal.

[0015] According to particular embodiments, the synthesis method has one or more of the following characteristics, taken in isolation or in all technically possible combinations:

[0016] - during the application step, the at least one neural network is applied to all extracted parameters.

[0017] - the probability density of the residual linear prediction signal is modeled by a Gaussian mixture model, the at least one neural network predicting, for each Gaussian of the Gaussian mixture model, at least one distribution parameter and the weight of the Gaussian.

[0018] - a neural network predicts the distribution parameters and another network neurons predict the weights of each Gaussian.

[0019] - each neural network predicts a respective parameter of each Gaussian.

[0020] - the method comprises a step of determining the fundamental frequency, the neural network used in the application step depending on a frequency interval in which the determined fundamental frequency is located.

[0021] - the neural network used during the application step depends on the type of voicing.

[0022] - the at least one neural network comprises first layers forming a initial network, the initial network being a convolutional neural network, a long short-term memory neural network, or a gated recurrent neural network.

[0023] The description also describes a device for synthesizing a voice signal, the synthesis device being capable of:

[0024] - obtain a stream coded according to a mixed excitation linear prediction coding,

[0025] - extract parameters from the coded stream,

[0026] - apply at least one neural network to the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream,

[0027] - generate an estimate of the residual linear prediction signal from the density of probability obtained, to obtain an excitation signal, and

[0028] - synthesize the voice from the excitation signal.

[0029] The description also provides a computer program product comprising instructions which, when the program is executed by a computer.

[0030] The description also describes a computer-readable medium comprising instructions which, when executed by a computer.

[0031] In the present description, the expression “suitable for” means indifferently “adapted for”, “adapted to” or “configured for”.

[0032] Characteristics and advantages of the invention will appear on reading the description which follows, given solely by way of non-limiting example, and made with reference to the appended drawings, in which:

[0033] - [Fig.l] [Fig.l] is a representation of a voice synthesis device,

[0034] - [Fig.2] [Fig.2] is an example of a neural network used by the device of voice synthesis,

[0035] - [Fig.3] [Fig.3] is another example of a neural network used by the device voice synthesis, and

[0036] - [Fig.4] [Fig.4] is yet another example of a neural network used by the voice synthesis device.

[0037] A voice synthesis device is shown schematically in [Fig.l].

[0038] The voice synthesis device is capable of synthesizing a voice signal from a coded incident flow.

[0039] In the example described, the voice synthesis device comprises a receiver 12, a calculator 14 and a transmitter 16.

[0040] The receiver 12 is capable of receiving external signals.

[0041] For example, receiver 12 is an antenna.

[0042] The computer 14 is capable of implementing a method for synthesizing a voice signal which will be described later.

[0043] The computer 14 is an electronic circuit designed to manipulate and / or transform data represented by electronic or physical quantities in registers of the computer and / or memories into other similar data corresponding to physical data in the memories of registers or other types of display devices, transmission devices or storage devices.

[0044] As specific examples, the calculator 14 is produced in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array), or an integrated circuit, such as an ASIC (Application Specific Integrated Circuit). It could also be envisaged to use a digital processor or DSP (Digital Signal Processor), a GPP (General Purpose Processor) processor or a neural processor.

[0045] Alternatively, when the method is carried out in the form of one or more software programs, i.e. in the form of a computer program, also called a computer program product, it is furthermore capable of being recorded on a medium, not shown, readable by a computer. The computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. By way of example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example FLASH or NVRAM) or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.

[0046] The transmitter 16 is capable of transmitting the signal synthesized by the computer 14.

[0047] For example, transmitter 16 is a loudspeaker.

[0048] An example of a method for synthesizing a voice signal is now described.

[0049] The synthesis method is a method aimed at obtaining a synthesized voice signal from of a coded incident stream.

[0050] For this, according to the example described, the synthesis method comprises an obtaining step, an extraction step, an application step, a generation step and a synthesis step.

[0051] During the obtaining step, the computer 14 obtains a stream coded according to a MELP coding.

[0052] The coded stream here corresponds to the result of applying MELP coding to a voice signal that one wishes to transmit.

[0053] The transmitted voice signal is the signal that the synthesis method aims to synthesize.

[0054] In one embodiment, the receiver 12 receives the coded stream and transmits it to the computer 14. The computer 14 thus obtains the coded stream by interaction with the receiver 12.

[0055] The coded stream is a set of frames, each frame being coded on a number of bits predefined according to the type of MELP coding applied.

[0056] For example, for 2400 bits / s MELP coding, one frame is quantized to 54 bits; for 1200 bits / s MELP coding, three frames are quantized to 81 bits, and for 600 bits / s MELP coding, four frames are quantized to 54 bits.

[0057] During the extraction step, the computer 14 extracts parameters from the coded stream. The parameters of a frame that can be extracted are chosen from the following parameters: • the LSF coefficients representing the spectral envelope of the coded stream. The abbreviation LSF refers to the corresponding English term for “Line Spectral Frequencies” which means spectral line frequencies. The LSF coefficients are therefore the spectral line frequency coefficients. • the Fourier coefficients in amplitude, • two gains, each gain corresponding to a respective half-frame, • two jointly quantified parameters, namely the pitch parameters (literally height parameter, the pitch parameter corresponding to an estimate of the fundamental frequency associated with the vibration frequency of the vocal cords) and overall voicing parameters (literally global voicing parameter), • a jitter parameter (jitter parameter in French), a jitter being added to avoid an overly robotic aspect of the synthesized signal, and • the voicing indicator parameters by frequency band, more often referred to as bandpass voicing parameters.

[0058] These parameters and their expressions are defined by the standard mentioned above.

[0059] According to the embodiments, the computer 14 extracts one or more of these parameters.

[0060] The calculator 14 generally extracts at least the coefficients, the pitch or voicing and the gains for each frame.

[0061] During the application step, the calculator 14 applies at least one neural network to the extracted parameters.

[0062] As described later with reference to Figures 2 to 4, several topologies are envisaged with one or more neural networks and different output parameters.

[0063] As a remark, in the presence of several neural networks, it could be considered that each neural network is a sub-neural network of a neural network constituted by all the sub-neural networks. In the following, this term is not used to simplify the reading for the reader.

[0064] The at least one neural network makes it possible, in the example described, to obtain the probability density of the residual linear prediction signal associated with the coded stream.

[0065] As will appear later with more precision, this result is obtained by the use of one or more MDN type neural networks and that the output of this or these networks makes it possible to obtain a parametric description of the probability density of the target data (more precisely the residual LPC signal).

[0066] The size of the residual excitation signal is 2xN0, N0 being the size of a period associated with the value of the pitch fO.

[0067] Linear prediction modeling makes it possible to obtain a set of linear prediction coefficients by minimizing the prediction error. The residual signal corresponds to the difference between the speech signal and the predicted signal, and is denoted LPC residual signal in the following.

[0068] Linear prediction coefficients are more often referred to as LPC coefficients. The abbreviation LPC refers to the corresponding English term "Linear Prediction Coefficients" which literally means "linear prediction coefficients".

[0069] The model used for probability is a Gaussian mixture model.

[0070] From the notational point of view, Ng will designate the number of Gaussians (more precisely the number of weights, namely 1 per Gaussian), size L of a vector will be 2xN0 (NO being the size of a period) and the size P of the output vector of the neural network will be Ng + 2 x Ng x L.

[0071] It may be noted here that the Gaussians are multidimensional (size L), and are represented in the figures as one-dimensional Gaussians to facilitate understanding.

[0072] Such a model is often designated by the acronym GMM which refers to the corresponding English name of “Gaussian Mixture Model”.

[0073] A Gaussian mixture model is a statistical model expressed as a mixture density. It is usually used to parametrically estimate the distribution of random variables by modeling them as the sum of several Gaussians, called kernels. This involves determining the variance, mean and amplitude of each Gaussian. These parameters are optimized according to the maximum likelihood criterion in order to approach the desired distribution as closely as possible.

[0074] Due to such an output, the at least one neural network considered here is a probabilistic neural network.

[0075] From a very schematic point of view, a classic neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.

[0076] More precisely, each layer comprises neurons taking their inputs from the outputs of the neurons of the previous layer, or from the input variables for the first layer.

[0077] Alternatively, more complex neural network structures can be envisaged with a layer that can be connected to a layer further away than the immediately preceding layer.

[0078] Each neuron is also associated with an operation, i.e. a type of processing, to be carried out by said neuron within the corresponding processing layer.

[0079] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a connection between two neurons. It is often a real number, which takes both positive and negative values. In some cases, the synaptic weight is a complex number.

[0080] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the previous layer, each value then being multiplied by the respective synaptic weight of each synapse, or link, between said neuron and the neurons of the previous layer, then of applying an activation function, typically a non-linear function, to said weighted sum, and of delivering at the output of said neuron, in particular to the neurons of the following layer which are connected to it, the value resulting from the application of the activation function. The activation function allows non-linearity to be introduced into the processing performed by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function are examples of activation functions.

[0081] As an optional addition, each neuron is also capable of applying, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.

[0082] According to the example described, each neural network here is a mixed-density neural network.

[0083] Such a neural network is often referred to by the acronym MDN, which corresponds to the corresponding English term for “Mixture Density Network”.

[0084] Such networks are notably described in the document entitled “Mixture Density Networks”, CM Bishop, Neural Computing Research Group Report, Feb. 1994.

[0085] As appears above, unlike a traditional neural network, the MDN network does not produce as output an estimate of the target vector of size L, but the most likely generative statistical model (here a GMM).

[0086] Three topologies are possible: a first topology consists of using a single network to predict all the parameters of the GMM model, a second topology consists of using a network per type of parameters (weights, means, variances), and a third topology consists of using a network for the weights and a network for the parameters of the Gaussians (means and variances).

[0087] A more complete description is now referred to Figures 2 to 4.

[0088] According to a first topology illustrated by [Fig.2], a single neural network is applied to all the extracted parameters.

[0089] According to the example shown, the single neural network determines at output, for each Gaussian (three in the case shown), the variance, the mean and the weight. These quantities are noted o, q and W with an index i on [Fig.2].

[0090] The probability density is modeled by a weighted sum of Gaussians (GMM model). For each Gaussian (characterized by its mean and its variance) a weight is associated.

[0091] From these parameters, the probability density of the residual LPC signal associated with the coded stream is generated.

[0092] The generative model thus obtained at output makes it possible to synthesize a residual synchronous pitch LPC excitation.

[0093] This last pre-processing stage being the same for all the topologies presented, it is not repeated for the second and third topologies.

[0094] According to a second topology corresponding to [Fig.3], two distinct neural networks are used.

[0095] The first neural network predicts the weights of each Gaussian and the second neural network predicts the distribution parameters of each Gaussian.

[0096] More precisely, in the example illustrated with three Gaussians, the first neural network applied to the set of extracted parameters gives as output the weight of the first Gaussian, the weight of the second Gaussian and the weight of the third Gaussian.

[0097] The second neural network gives as output the variance and the mean of the first Gaussian, the variance and the mean of the second Gaussian and the variance and the mean of the third Gaussian.

[0098] According to a second topology corresponding to [Fig.4], three distinct neural networks are used, each neural network predicting a respective parameter of each Gaussian.

[0099] More precisely, in the example illustrated with three Gaussians, the first neural network applied to the set of extracted parameters gives as output the weight of the first Gaussian, the weight of the second Gaussian and the weight of the third Gaussian.

[0100] The second neural network gives as output the average of the first Gaussian, the average of the second Gaussian and the average of the third Gaussian.

[0101] It can be specified here that each Gaussian is multidimensional in the sense that the variance and the mean are vectors. The weights are nevertheless scalars.

[0102] Finally, the third neural network gives as output the variance of the first Gaussian, the variance of the second Gaussian and the variance of the third Gaussian.

[0103] In each of the topologies, the applied neural network can be obtained by any known learning technique.

[0104] Learning is done from an aligned corpus of MELP parameters and periods of the LPC residual signal calculated from the original signal.

[0105] More precisely, the constitution of the learning corpus is done from MELP modeling.

[0106] The residual LPC signal is cut into 2xN0 vectors, using the pitch provided by the MELP coding and after synchronization by maximizing the inter-correlation between the residual LPC signal and the excitation signal synthesized by the MELP coding.

[0107] Such an approach makes it possible to associate with each set of MELP parameters a portion of the residual signal of duration 2 x N0, and thus to constitute the learning corpus.

[0108] Alternatively, any other approach suitable for performing pitch-synchronous segmentation of the residual signal may be used here.

[0109] According to one embodiment, to improve the prediction performance, a stage for preprocessing the inputs of the neural network can also be provided.

[0110] The preprocessing stage is therefore located upstream of the input layer of the neural network.

[0111] As its name suggests, the preprocessing stage implements a preprocessing operation of some of the extracted parameters.

[0112] In addition, the preprocessing stage here performs a conversion which may vary from one embodiment to another.

[0113] According to one example, the preprocessing operation comprises a conversion of the LSF coefficients into cepstral coefficients. In practice, such a conversion is generally done in two stages by an intermediate passage to an LPC representation.

[0114] Cepstral coefficients are the coefficients of an LPCC representation of the envelope.

[0115] Cepstral coefficients are thus more often referred to as LPCC coefficients in reference to the English term “linear prediction cepstral coefficients”.

[0116] According to another example or in addition, the preprocessing operation comprises a conversion of incident coefficients into the frequency domain.

[0117] By incident coefficients, it is understood that the set of at least one layer takes as input incident coefficients of a representation.

[0118] The incident coefficients here are the cepstral coefficients.

[0119] Also, in the example described, the coefficients obtained after passage into the frequency domain are log-spectral envelope coefficients, more simply called here DCT coefficients. These coefficients are obtained by DCT transformation of the LPCC coefficients.

[0120] The abbreviation DCT designates the technique which enabled the conversion to the frequency domain and refers to the corresponding English name of “Discrete Cosine Transform” which means discrete cosine transform.

[0121] Thus, with the proposed examples, instead of the LSF coefficients, we can therefore use the LPCC coefficients or the DCT coefficients as input to the neural network.

[0122] These changes in representation for the envelope make it possible to obtain a representation that is more easily representable than the representation using the LSF coefficients.

[0123] In each case, the application of one or each neural network thus makes it possible to obtain the probability density of the residual LPC signal.

[0124] During the generation step, the computer 24 generates the excitation signal (estimated signal of the residual prediction signal) as a function of the probability density obtained.

[0125] More precisely, the calculator 24 generates an estimate of the residual LPC signal over the horizon of two pitch periods as a function of the probability density obtained.

[0126] This signal is to be compared to the case of the synthesis carried out during MELP decoding. Indeed, the excitation signal of the LPC model is obtained from a simplifying hypothesis, namely the presence of a localized pulse and the random nature of the phase for the unvoiced frequency bands.

[0127] In such decoding, a phase dispersion filter is used to simulate the shape of the glottal pulse, this filter being a whitened triangular filter.

[0128] The neural network proposed in the present method makes it possible to simultaneously predict all the information, so that the dispersion filter is no longer necessary.

[0129] MELP decoding is thus based on the assumption of an invariant triangular shape, whereas the use of the network makes it possible to predict a shape closer to a natural signal.

[0130] During the synthesis step, the computer 14 synthesizes the voice from the residual signal obtained during the generation step.

[0131] Such a synthesis is done by applying MELP decoding to the excitation signal obtained at the end of the generation step.

[0132] More precisely, the modeling over two consecutive periods of the residual linear prediction signal is used.

[0133] The target vectors (and therefore the vectors generated from the GMM models predicted by the MDN neural network) are weighted by an analysis window during learning, which allows an addition-overlap type synthesis.

[0134] Such a synthesis is often referred to by the acronym OLA, which refers to the corresponding English term “OverLap-and-Add”.

[0135] An example of such a window is a square root of a Hanning window.

[0136] This window is again applied to the predicted vector before the OLA synthesis, which ultimately amounts to applying a Hanning window to the entire analysis-synthesis cycle, the analysis here being associated with the learning phase.

[0137] The method thus makes it possible to obtain the voice signal to be synthesized from the original MELP stream (coded by the MELP technique).

[0138] In the method described, the use of an MDN network makes it possible to generate a generative model of the excitation signal and not the excitation signal directly.

[0139] This feature makes it possible to take into account temporal variability at the time of synthesis, which contributes to the natural rendering of the synthesized voice.

[0140] Indeed, the use of the generative model makes it possible to generate a controlled temporal variability of the excitation signal from the same set of parameters. MELP and thus allows the synthesis of an excitation signal closer to a natural signal by offering better modeling of its different components.

[0141] The implementation of the synthesis method is also relatively easy insofar as the decoding remains identical to that of a MELP decoder with the exception of the generation of the excitation signal.

[0142] Other embodiments benefiting from the same advantages can be envisaged.

[0143] In the topologies described above, the neural network has a multi-layer perceptron type structure.

[0144] Such a structure is often referred to by the acronym MLP which refers to the corresponding English name of “Multilayer Perceptron”.

[0145] To better take into account the temporal trajectories of the excitation periods, a recurrent neural network structure could also be used, particularly for certain layers.

[0146] Advantageously, the layers concerned are the first layers, so that the neural network used during the application step comprises a neural network having an initial neural network, this initial network forming a recurrent neural network.

[0147] A recurrent neural network is an artificial neural network exhibiting recurrent connections.

[0148] A recurrent neural network is thus made up of interconnected neurons interacting non-linearly and for which there is at least one cycle in the structure. The units are connected by synapses which have a weight. The output of a neuron is a non-linear combination of its inputs.

[0149] Such a neural network is often designated by the acronym RNN which refers to the corresponding English name of “Recurrent Neural Network”.

[0150] In this case, this means that the neural network takes into account past and future frames to make a prediction at a given time in addition to the current frame (corresponding to the prediction time).

[0151] According to another embodiment, the or each initial neural network is a long short-term memory network.

[0152] Such a network is more often designated by the acronym LSTM which corresponds to the corresponding English name of “Long Short-Term Memory”.

[0153] According to yet another embodiment, the or each initial neural network is a gated recurrent neural network.

[0154] Such a network is more often designated by the acronym GRU which corresponds to the corresponding English name of “Gated Recurrent Unit”.

[0155] Convolutional neural network type structures are also conceivable to form part of the network.

[0156] Such a network is more often designated by the acronym CNN which corresponds to the corresponding English name of “Convolutional N eural N etworks”.

[0157] Thus, according to the embodiments, the at least one initial neural network comprises a CNN network, an LSTM network or a GRU network.

[0158] It is also possible to take into account other elements at the level of neural networks.

[0159] For example, according to one embodiment, the method comprises a step of determining the fundamental frequency, the neural network used during the application step depending on a frequency interval in which the determined fundamental frequency is found.

[0160] This amounts to using several MDN networks based on predefined pitch ranges (which condition the size of the excitation periods.

[0161] According to one embodiment, the neural networks differ for all of the coefficients.

[0162] According to another embodiment, the neural network differs, for example, for the output layer which can then be optimized for a data subspace (conditioned by the pitch and / or voicing information).

[0163] According to one example, in this context, learning comprises two learnings.

[0164] The first learning is a global learning carried out on the entire corpus using an output layer sized for the longest pitch period (when the target vector is shorter, the target vector is completed on both sides with zeros).

[0165] The second learning is carried out by subspace by freezing all or part of the first layers of the network. The last layer is then resized at the output using the largest size of the target vector for the interval considered.

[0166] The division according to the pitch value can be carried out as follows: the limits of the intervals are defined so as to have a balanced distribution of the learning data in each interval.

[0167] It may be advantageous in certain cases for the defined intervals to have an overlap.

[0168] In practice, a number of 4 intervals may be chosen. Such a division makes it possible in particular to avoid the presence in the same interval of a pitch frequency and its double or half frequency.

[0169] According to another example, the neural network used during the application step depends on the type of voicing from the MELP.

[0170] More precisely, an optimization of the last layer based on the voicing information can be carried out.

[0171] For each neural network obtained by one of the preceding techniques (global learning or pitch interval learning), the neural network is re-optimized by freezing all or part of the first layers (at least the last layer is optimized) for the data corresponding to a predefined voicing criterion.

[0172] A possible voicing criterion consists of defining, for example, three voicing levels as a function of the MELP voicing vector (5 sub-bands), namely: • unvoiced or weakly voiced: (0,0,0,0,0) and (1,0,0,0,0) • voiced (1,1,0,0,0) and (1,1,1,0,0), and • strongly voiced (1,1,1,1,0) and (1,1,1,1,1).

[0173] In this specific example, neural networks can be considered for each different pair of pitch intervals and voicing types. Since there are 4 pitch intervals and 3 voicing types, this leads to 12 neural networks that differ only in the output layer. This allows limiting the total number of coefficients).

[0174] Alternatively, it may be considered to use separate neural networks but at the cost of much greater memory space involved.

[0175] In each case, the voice synthesis method which makes it possible to reconstruct a voice signal closer to the original voice signal within the framework of a compression of the original voice signal by MELP coding.

Claims

Claims

1. Method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of: - obtaining a stream coded according to a mixed excitation linear prediction coding, - extracting the parameters of the coded stream, - applying at least one neural network to the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream, - generating an estimate of the residual linear prediction signal from the probability density obtained, to obtain an excitation signal, and - synthesizing the voice from the excitation signal.

2. Synthesis method according to claim 1, wherein, during the application step, the at least one neural network is applied to all of the extracted parameters.

3. A synthesis method according to claim 1 or 2, wherein the probability density of the linear prediction residual signal is modeled by a Gaussian mixture model, the at least one neural network predicting, for each Gaussian of the Gaussian mixture model, at least one distribution parameter and the weight of the Gaussian.

4. A synthesis method according to claim 3, wherein one neural network predicts the distribution parameters and another neural network predicts the weights of each Gaussian.

5. A synthesis method according to claim 3, wherein each neural network predicts a respective parameter of each Gaussian.

6. A synthesis method according to any one of claims 1 to 5, wherein the method comprises a step of determining the fundamental frequency, the neural network used during the application step depending on a frequency interval in which the determined fundamental frequency is located.

7. A synthesis method according to any one of claims 1 to 6, wherein the neural network used in the applying step depends on the voicing type.

8. A synthesis method according to any one of claims 1 to 7, wherein the at least one neural network comprises first layers forming an initial network, the initial network being a convolutional neural network, a long short-term memory neural network or a gated recurrent neural network.

9. Device for synthesizing a voice signal, the synthesis device being capable of: - obtaining a stream coded according to a mixed excitation linear prediction coding, - extracting parameters from the coded stream, - applying at least one neural network to the extracted parameters to obtain the probability density of the residual linear prediction signal associated with the coded stream, - generating an estimate of the residual linear prediction signal from the probability density obtained, to obtain an excitation signal, and - synthesizing the voice from the excitation signal.

10. A computer program product comprising instructions which, when the program is executed by a computer, cause the latter to implement the steps of a method according to any one of claims 1 to 8.

11. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the steps of a method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voicing analysis in a linear predictive speech coder

    EP1037197A2

  • Artificial intelligence based audio coding

    US11437050B2

  • Process and device for creating comfort noise in a digital speech transmission system

    US5812965A