Voice synthesis method and associated devices

The voice synthesis method addresses the issue of MELP's inadequate signal reconstruction by using neural networks to correct and extend frequency bands, enhancing the quality of synthesized voice signals.

FR3155617B1Active Publication Date: 2025-10-10THALES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023012801
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-10-10
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

Existing voice synthesis methods using mixed excitation linear prediction coding (MELP) fail to reconstruct voice signals with the necessary characteristics of the original signal after compression, leading to a synthetic signal that is not sufficiently close to the original.

Method used

A voice synthesis method employing a neural network to extract and correct spectral envelope parameters, applying mixed excitation linear prediction techniques, and merging signals across frequency bands to extend the frequency band from 4 kHz to 8 kHz, improving the quality of the synthesized voice.

Benefits of technology

The method enhances the quality of synthesized voice signals by making them closer to the original, utilizing neural networks to predict and correct spectral envelopes, and merging signals effectively, resulting in improved voice synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000020_0000
    Figure 00000020_0000
  • Figure 00000021_0000
    Figure 00000021_0000
  • Figure 00000022_0000
    Figure 00000022_0000
Patent Text Reader

Abstract

Method for synthesizing voice and associated device The present invention relates to a method for synthesizing a voice signal comprising the steps of: - obtaining a stream coded according to a mixed excitation linear prediction coding, - extracting the parameters of the coded stream, - applying a neural network to extracted parameters to obtain a set of output parameters comprising a first set of gains and first spectral envelope parameters, the gain and the spectral envelope being the gain and the spectral envelope of the voice signal to be synthesized in the frequency band 4 kHz - 8 kHz, - generating a first signal from parameters derived from the output parameters, - applying a mixed excitation linear prediction synthesis technique to the extracted parameters to obtain a second signal, and - merging the signals to obtain the voice signal to be synthesized. Figure: figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Voice synthesis method and associated devices

[0001] The present invention relates to a voice synthesis method. The present invention relates to a voice synthesis device. The present invention also relates to a computer program product and a corresponding information medium.

[0002] In the field of radio communications, voice signal exchanges often have to be carried out.

[0003] For this, a compression of the voice signal is often carried out to limit the quantity of information to be exchanged.

[0004] In particular, it is known to use mixed excitation linear prediction coding as a compression technique.

[0005] Mixed excitation linear prediction coding is more often referred to as MELP coding, which refers to the corresponding English term “Mixed Excitation Linear Prediction”.

[0006] This coding is notably described in the document “The 1200 and 2400 bit / s NATO Interoperable Narrow Band Voice Coder”, STANAG No. 4591, NATO Standardization Agency.

[0007] However, the reconstruction of the signal thus coded by MELP decoding leads to a synthetic signal which does not have all the characteristics of the original voice signal.

[0008] There is therefore a need for a voice synthesis method which makes it possible to reconstruct a voice signal closer to the original voice signal within the framework of compression of the original voice signal by MELP coding.

[0009] For this purpose, the description describes a method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of:

[0010] - obtaining a coded stream according to a mixed excitation linear prediction coding,

[0011] - extraction of parameters from the coded stream,

[0012] - application of a neural network on parameters from the extracted parameters to obtain a set of output parameters, the set of output parameters comprising a first set of gains and first spectral envelope parameters, the gain and the spectral envelope being respectively the gain and the spectral envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz, the set of output parameters comprising the contrast of the envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz,

[0013] - correction of the first spectral envelope parameters, using contrast to obtain first corrected spectral envelope parameters,

[0014] - generation of a first signal from generation parameters, the parameters generation comprising parameters from the first set of gains and the first corrected spectral envelope parameters,

[0015] - application of a mixed excitation linear prediction synthesis technique on the parameters extracted to obtain a second signal corresponding to the voice signal to be synthesized for frequencies below 4 kHz, and

[0016] - merging the first signal and the second signal to obtain the voice signal to syn to thetize.

[0017] The method described exploits a principle of artificial band extension applied to the MELP decoder.

[0018] The method thus extends the frequency band of the MELP decoder from 4 kHz (Narrow-Band) to 8 kHz (Wide-Band).

[0019] The quality of the synthesized voice is then improved, and thus, closer to the original voice.

[0020] This improvement in voice synthesis and easy integration can also be achieved with a different implementation of the synthesis method.

[0021] According to particular embodiments, the synthesis method has one or more of the following characteristics, taken in isolation or in all technically possible combinations:

[0022] - the merging step includes an oversampling operation followed by a operation of applying a quadrature mirror filter, the two operations are applied to the first signal and the second signal, to obtain respectively a first intermediate signal and a second intermediate signal, the merging step comprising the addition of the first intermediate signal and the second intermediate signal to obtain the voice signal to be synthesized.

[0023] - the extracted parameters include the frequency coefficients of spectral lines representing the spectral envelope of the codestream, the parameters from the extracted parameters including the spectral line frequency coefficients representing the spectral envelope of the codestream. This means that the LSF parameters are from the MELP stream and can be used as input to the network.

[0024] - the extracted parameters include the frequency coefficients of spectral lines representing the spectral envelope of the coded stream and in which, during the step of applying the neural network, an operation of converting the frequency coefficients of spectral lines representing the spectral envelope of the coded stream into cepstral coefficients representing the spectral envelope of the coded stream is implemented, the parameters from the extracted parameters comprising the cepstral coefficients re presenting the spectral envelope of the coded stream. This means that the cepstral coefficients can be used at the input of the network.

[0025] - the extracted parameters include the frequency coefficients of spectral lines representing the spectral envelope of the coded stream and in which, during the step of applying the neural network, operations of conversion of the frequency coefficients of spectral lines representing the spectral envelope of the coded stream are implemented to obtain coefficients in the frequency domain representing the spectral envelope of the coded stream, the parameters resulting from the extracted parameters comprising the coefficients in the frequency domain representing the spectral envelope of the coded stream.

[0026] - during the application step, the neural network also predicts parameters voicing, the generation parameters including the predicted voicing parameters.

[0027] - the neural network is a recurrent neural network, preferably a network with long short-term memory or a gated recurrent neural network.

[0028] - the coded stream is organized into frames, the neural network being applied to four frames, two previous frames, one current frame and one future frame.

[0029] The description also describes a device for synthesizing a voice signal, the synthesis device being capable of:

[0030] - obtain a stream coded according to a mixed excitation linear prediction coding,

[0031] - extract parameters from the coded stream,

[0032] - apply a neural network to parameters from the extracted parameters to obtain a set of output parameters, the set of output parameters comprising a first set of gains and first spectral envelope parameters, the gain and the spectral envelope being respectively the gain and the spectral envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz, the set of output parameters comprising the contrast of the envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz,

[0033] - correct the first spectral envelope parameters, using contrast to obtain first corrected spectral envelope parameters,

[0034] - generate a first signal from generation parameters, the generation parameters generation comprising parameters from the first set of gains and the first corrected spectral envelope parameters,

[0035] - apply a mixed excitation linear prediction synthesis technique on the parameters extracted to obtain a second signal corresponding to the voice signal to be synthesized for frequencies below 4 kHz, and

[0036] - merge the first signal and the second signal to obtain the voice signal to be syn- to thetize.

[0037] For this purpose, the description describes a method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of:

[0038] - obtaining a coded stream according to a linear prediction coding with mixed excitation,

[0039] - extraction of parameters from the coded stream,

[0040] - application of a neural network on parameters from the extracted parameters to obtain a set of output parameters, the set of output parameters comprising a first set of gains and first spectral envelope parameters, the gain and the spectral envelope being respectively the gain and the spectral envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz, the set of output parameters comprising the contrast of the envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz,

[0041] - correction of the first spectral envelope parameters, using contrast to obtain first corrected spectral envelope parameters,

[0042] - generation of a first signal from generation parameters, the parameters generation comprising parameters from the first set of gains and the first corrected spectral envelope parameters,

[0043] - application of a mixed excitation linear prediction synthesis technique on the parameters extracted to obtain a second signal corresponding to the voice signal to be synthesized for frequencies below 4 kHz, and

[0044] - merging the first signal and the second signal to obtain the voice signal to syn to thetize.

[0045] According to particular embodiments, the synthesis method has one or more of the following characteristics, taken in isolation or in all technically possible combinations:

[0046] - the fusion step includes an operation of calculating the impulse response associated with the output parameters and an impulse response calculation operation associated with the extracted parameters.

[0047] - the merging step includes an oversampling operation applied to the impulse responses followed by a recombination operation.

[0048] - the extracted parameters include the frequency coefficients of spectral lines representing the spectral envelope of the coded stream, the parameters from the extracted parameters including the spectral line frequency coefficients representing the spectral envelope of the coded stream.

[0049] - the extracted parameters include the frequency coefficients of spectral lines representing the spectral envelope of the coded stream and in which, during the application step, an operation of converting the line frequency coefficients spectral coefficients representing the spectral envelope of the coded stream into cepstral coefficients representing the spectral envelope of the coded stream is implemented, the parameters from the extracted parameters including the spectral coefficients representing the spectral envelope of the coded stream.

[0050] - the extracted parameters include the frequency coefficients of spectral lines representing the spectral envelope of the coded stream and in which, during the application step, operations for converting the frequency coefficients of spectral lines representing the spectral envelope of the coded stream are implemented to obtain coefficients in the frequency domain representing the spectral envelope of the coded stream, the parameters from the extracted parameters comprising the coefficients in the frequency domain representing the spectral envelope of the coded stream.

[0051] - the neural network is a recurrent neural network, preferably a network with long short-term memory or a gated recurrent neural network.

[0052] - the coded stream is organized into frames, the neural network being applied to four frames, two previous frames, one current frame and one future frame.

[0053] The description also describes a voice signal synthesis device, the device being suitable for:

[0054] - obtain a stream coded according to a mixed excitation linear prediction coding,

[0055] - extract parameters from the coded stream,

[0056] - apply a neural network to parameters from the extracted parameters to obtain a set of output parameters, the set of output parameters comprising a first set of gains and first spectral envelope parameters, the gain and the spectral envelope being respectively the gain and the spectral envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz, the set of output parameters comprising the contrast of the envelope of the speech signal to be synthesized in the frequency band between 4 kHz and 8 kHz,

[0057] - correct first spectral envelope parameters, using contrast to obtain first corrected spectral envelope parameters,

[0058] - merge parameters from the output parameter set and pass them parameters extracted from the coded stream, the parameters from the output parameters including the first corrected spectral envelope parameters, and

[0059] - apply a mixed excitation linear prediction synthesis technique on the parameters obtained after the fusion step to obtain the voice signal to be synthesized.

[0060] The description also provides a computer program product comprising instructions which, when the program is executed by a computer, cause the latter to implement the steps of a method as previously described.

[0061] The description also describes a computer-readable medium comprising ins instructions which, when executed by a computer, cause it to implement the steps of a method as previously described.

[0062] In the present description, the expression “suitable for” means indifferently “adapted for”, “adapted to” or “configured for”.

[0063] Characteristics and advantages of the invention will appear on reading the description which follows, given solely by way of non-limiting example, and made with reference to the appended drawings, in which:

[0064] - [Fig.l] [Fig.l] is a representation of a voice synthesis device,

[0065] - [Fig.2] [Fig.2] is a flowchart illustrating an example of a synthesis process of a voice signal,

[0066] - [Fig.3] [Fig.3] is a schematic representation of an example of implementation implementation of a step of the synthesis process, and

[0067] - [Fig.4] [Fig.4] is a flowchart illustrating another example of a method of synthesis of a voice signal.

[0068] A voice synthesis device is shown schematically in [Fig.l].

[0069] The voice synthesis device is capable of synthesizing a voice signal from a coded incident flow.

[0070] In the example described, the voice synthesis device comprises a receiver 12, a calculator 14 and a transmitter 16.

[0071] The receiver 12 is capable of receiving external signals.

[0072] For example, receiver 12 is an antenna.

[0073] The computer 14 is capable of implementing a method for synthesizing a voice signal which will be described later.

[0074] The computer 14 is an electronic circuit designed to manipulate and / or transform data represented by electronic or physical quantities in registers of the computer and / or memories into other similar data corresponding to physical data in the memories of registers or other types of display devices, transmission devices or storage devices.

[0075] As specific examples, the calculator 14 is produced in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array), or an integrated circuit, such as an ASIC (Application Specific Integrated Circuit). It could also be envisaged to use a digital processor or DSP (Digital Signal Processor), a GPP (General Purpose Processor) or a neural processor.

[0076] Alternatively, when the method is carried out in the form of one or more software programs, i.e. in the form of a computer program, also called a product computer program, it is further capable of being recorded on a medium, not shown, readable by a computer. The computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, the readable medium is an optical disc, a magneto-optical disc, a ROM memory, a RAM memory, any type of non-volatile memory (for example FLASH or NVRAM) or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.

[0077] The transmitter 16 is capable of transmitting the signal synthesized by the computer 14.

[0078] For example, transmitter 16 is a loudspeaker.

[0079] It is now described with reference to [Fig.2] which is a flowchart illustrating an example of a method for synthesizing a voice signal.

[0080] The synthesis method is a method aimed at obtaining a synthesized voice signal from a coded incident stream.

[0081] For this, according to the example described, the synthesis method comprises an obtaining step E20, an extraction step E22, a first application step E24, a correction step E26, a second application step E28, a conversion step E29, a generation step E30 and a fusion step E40.

[0082] During the obtaining step E20, the computer 14 obtains a stream coded according to a MELP coding.

[0083] The coded stream here corresponds to the result of applying MELP coding to a voice signal that we wish to transmit.

[0084] The transmitted voice signal is the signal that the synthesis method aims to synthesize.

[0085] In one embodiment, the receiver 12 receives the coded stream and transmits it to the computer 14. The computer 14 thus obtains the coded stream by interaction with the receiver 12.

[0086] The coded stream is a set of frames, each frame being coded on a number of bits predefined according to the type of MELP coding applied.

[0087] For example, for 2400 bits / s MELP coding, one frame is quantized to 54 bits; for 1200 bits / s MELP coding, three frames are quantized to 81 bits, and for 600 bits / s MELP coding, four frames are quantized to 54 bits.

[0088] Furthermore, each frame corresponds to a 22.5 ms speech frame.

[0089] During the extraction step E22, the computer 14 extracts parameters from the flow coded. The parameters of a frame that can be extracted are chosen from the following parameters: • the LSF coefficients representing the spectral envelope of the coded stream. The abbreviation LSF refers to the corresponding English term for “Line Spectral Frequencies” which means spectral line frequencies. The LSF coefficients are therefore the coefficients of spectral line frequencies. • the Fourier coefficients in amplitude, • two gains, each gain corresponding to a respective half-frame, the two gains thus correspond to a gain vector for the frame, • two joint quantification parameters, namely the jointly quantified pitch parameters (literally pitch parameter, the pitch parameter corresponding to an estimate of the fundamental frequency associated with the vibration frequency of the vocal cords) and overall voicing (literally global voicing parameter), • a jitter parameter (jitter parameter in French), a jitter being added to avoid an overly robotic aspect of the synthesized signal, and • the voicing indicator parameters by frequency band, more often referred to as bandpass voicing parameters.

[0090] These parameters and their expressions are defined by the standard mentioned above.

[0091] According to the embodiments, the calculator 14 extracts one or more of these parameters.

[0092] The calculator 14 generally extracts at least the coefficients, the pitch or voicing and the gains for each frame.

[0093] During the first application step E26, the computer 14 applies a neural network to parameters derived from the extracted parameters.

[0094] Depending on the case, the parameters resulting from the extracted parameters are the extracted parameters themselves or parameters obtained by preprocessing the extracted parameters.

[0095] The neural network makes it possible, in the example described, to obtain a set of output parameters.

[0096] According to the described example, the output parameters comprise a first set of gains, first spectral envelope parameters and the envelope contrast.

[0097] The gain is the gain of the voice signal to be synthesized in the frequency band between 4 kHz and 8 kHz.

[0098] More specifically, in the example described, the first set of gains corresponds to a first vector of gains for each frame bringing together a gain for a first part of the frame and another gain for a second part of the frame. The parts of frames are delimited similarly to the delimitations of the MELP coder.

[0099] These gains forming the first gain vector are part of the LPC modeling of a voice signal and represent the energy of the excitation signal (or residual signal) of the LPC model.

[0100] A speech signal in an LPC model is represented by a set of linear prediction coefficients.

[0101] Linear prediction coefficients are more often referred to as LPC coefficients. The abbreviation LPC refers to the corresponding English term for "Linear Prediction Coefficients" which literally means "linear prediction coefficients".

[0102] In the example described, a 10th order LPC model is used. This is also the LCP modeling order used for MELP coding / decoding.

[0103] The frequency band between 4 kHz and 8 kHz corresponds, in this context, to a high frequency (HF) band and can therefore be referred to as the HF band in the following.

[0104] The first spectral envelope parameters designate parameters making it possible to represent the envelope of the voice signal to be synthesized in the HF band in a model.

[0105] The contrast of the envelope is defined here as the value of SFM, the acronym SFM referring to the English term “Spectral Flatness Measure” which literally means measurement of the flatness of the spectrum.

[0106] For example, contrast is here defined as the ratio between the arithmetic and geometric sums of the coefficients representing the envelope.

[0107] From a very schematic point of view, a classic neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.

[0108] More precisely, each layer comprises neurons taking their inputs from the outputs of the neurons of the previous layer, or from the input variables for the first layer.

[0109] Alternatively, more complex neural network structures can be envisaged with a layer that can be connected to a layer further away than the immediately preceding layer.

[0110] Each neuron is also associated with an operation, i.e. a type of processing, to be carried out by said neuron within the corresponding processing layer.

[0111] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a connection between two neurons. It is often a real number, which takes both positive and negative values. In some cases, the synaptic weight is a complex number.

[0112] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the previous layer, each value then being multiplied by the respective synaptic weight of each synapse, or link, between said neuron and the neurons of the previous layer, then applying an activation function, typically a non-linear function, to said weighted sum, and delivering as output of said neuron, in particular to the neurons of the next layer connected to it, the value resulting from the application of the activation function. The activation function makes it possible to introduce non-linearity into the processing carried out by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function are examples of activation functions.

[0113] As an optional addition, each neuron is also capable of applying, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.

[0114] In the example described, the neural network is a recurrent neural network.

[0115] A recurrent neural network is an artificial neural network exhibiting recurring connections.

[0116] A recurrent neural network is thus made up of interconnected neurons interacting non-linearly and for which there is at least one cycle in the structure. The units are connected by synapses which have a weight. The output of a neuron is a non-linear combination of its inputs.

[0117] Such a neural network is often designated by the acronym RNN which refers to the corresponding English name of “Recurrent Neural Network”.

[0118] In this case, this means that the neural network takes into account past and future frames to make a prediction at a given time in addition to the current frame (corresponding to the prediction time)

[0119] As a specific example, the neural network takes as input the parameters of two previous frames, the current frame and a subsequent frame, thus using a time sequence of four frames.

[0120] According to one embodiment, the neural network is a long short-term memory network.

[0121] Such a network is more often designated by the acronym LSTM which corresponds to the corresponding English name of “Long Short-Term Memory”.

[0122] According to yet another embodiment, the neural network is a gated recurrent neural network.

[0123] Such a network is more often designated by the acronym GRU which corresponds to the corresponding English name of “Gated Recurrent Unit”.

[0124] The applied neural network can be obtained by any known learning technique.

[0125] It may be noted here that the formation of the training set may be accomplished by using the steps of the present method to calculate the output that the neural network should have for a given MELP stream.

[0126] According to a particular embodiment, to improve the performance of prediction, a stage of preprocessing of the neural network inputs can also be provided.

[0127] The preprocessing stage is therefore located upstream of the input layer of the neural network.

[0128] As its name suggests, the preprocessing stage implements a preprocessing operation of some of the extracted parameters.

[0129] In addition, the preprocessing stage here performs a conversion which may vary from one embodiment to another.

[0130] According to one example, the preprocessing operation comprises a conversion of the LSF coefficients into cepstral coefficients. In practice, such a conversion is generally done in two stages by an intermediate passage to an LPC representation.

[0131] Cepstral coefficients are the coefficients of an LPCC representation of the envelope.

[0132] Cepstral coefficients are thus more often referred to as LPCC coefficients in reference to the English term “linear prediction cepstral coefficients”.

[0133] According to another example or in addition, the preprocessing operation comprises a conversion of incident coefficients into the frequency domain.

[0134] By incident coefficients, it is understood that the set of at least one layer takes as input incident coefficients of a representation.

[0135] The incident coefficients here are the cepstral coefficients.

[0136] Also, in the example described, the coefficients obtained after passing into the frequency domain are log-spectral envelope coefficients, more simply called DCT coefficients.

[0137] The abbreviation DCT designates the technique which enabled the conversion to the frequency domain and refers to the corresponding English name of “Discrete Cosine Transform” which means discrete cosine transform.

[0138] Thus, with the proposed examples, instead of the LSF coefficients, we can therefore use the LPCC coefficients or the DCT coefficients as input to the neural network.

[0139] These changes in representation for the envelope make it possible to obtain a representation that is more easily representable than the representation using the LSF coefficients.

[0140] This also makes it possible to make the neural network compatible with cases where coefficients other than LSF coefficients are used.

[0141] In the example described, from a functional point of view, the neural network therefore takes as input parameters from the extracted parameters and then predicts as output the gain, the spectral envelope and the contrast in the HF band from the converted parameters and the unconverted parameters.

[0142] During the correction step E26, the computer 14 corrects the first spectral envelope using the contrast to obtain a first corrected spectral envelope.

[0143] Such a step is schematically represented in [Fig.3], the left graph corresponding to the situation before implementation of the correction step and the right graph to the situation after implementation of the correction step.

[0144] Each graph represents the signal amplitude for each DCT coefficient.

[0145] In this diagram, curve C1 corresponds to the first envelope while curve C2 corresponds to the contrast.

[0146] The first curve Cl extends on average around an affine line (curve C3 in [Fig.3]). The affine line is obtained by linear regression.

[0147] The calculator 14 removes the affine line from the first envelope (affine normalization operation) then adjusts the contrast of the first normalized envelope so that it corresponds to the predicted contrast.

[0148] The adjustment is made by applying a multiplicative factor to each DCT coefficient so that the contrast obtained is equal or close (in the sense of minimizing an error criterion) to that of the C2 curve.

[0149] This allows us to obtain the new curves Cl, C2 and C3 visible on the graphs on the right.

[0150] During the second application step E28, the computer 14 applies a mixed excitation linear prediction synthesis technique (MELP synthesis forming part of the MELP decoding) to the coded stream to obtain a second signal from the extracted parameters.

[0151] Among these extracted parameters, we find in particular a second gain vector and parameters relating to a second spectral envelope.

[0152] Each second gain of the second gain vector is the gain of the voice signal to be synthesized for frequencies below 4 kHz.

[0153] In the example described, as for the case of the first gain, the second gain corresponds to a second vector of gains bringing together a gain for a first part of the frame and another gain for a second part of the frame.

[0154] Frame portions are delimited similarly to MELP encoder delimitations.

[0155] Frequencies below 4 kHz correspond in this context to a low frequency (LF) band and can therefore be referred to as the LF band in the following.

[0156] The second spectral envelope is the spectral envelope of the voice signal to be synthesized in the BF band.

[0157] During the conversion step E29, the computer 14 performs a conversion of the corrected envelope parameters to obtain a set of prediction coefficients linear LPC model representing the speech signal to be synthesized.

[0158] During the first generation step E30, the computer 14 generates a first signal from generation parameters.

[0159] The generation parameters comprising parameters from the first set of gains and the first spectral envelope parameters.

[0160] In this case, here, the generation parameters are the first set of gains and the first corrected spectral envelope parameters.

[0161] For this, the computer 14 uses an LPC type synthesis under the assumption of an unvoiced signal. The excitation signal of the LPC model is then white noise.

[0162] According to the example described, the computer 14 applies to each signal an oversampling operation (042 in the diagram of [Fig.2]) and an operation of applying a quadrature mirror filter (044 in the diagram of [Fig.2]).

[0163] More precisely, the computer 14 applies an oversampling operation on the first signal (by inserting zeros) to obtain a first intermediate signal on which the computer 14 then applies the quadrature mirror filter, thus obtaining a first signal to be recombined.

[0164] A quadrature mirror filter is often referred to as a QMF filter, the abbreviation QMF referring to the corresponding English name of “Quadrature Mirror Filter”.

[0165] Similarly, the computer 14 applies an oversampling operation to the second signal to obtain a second intermediate signal to which the computer 14 then applies the quadrature mirror filter, thus obtaining a second signal to be recombined.

[0166] The computer 14 then recombines the two signals to be recombined.

[0167] For this, according to the example described, the computer 14 implements an addition of the first signal to be recombined and the second signal to be recombined (operation 046 in [Fig.2]).

[0168] The result of the addition, that is to say the sum of the first signal to be recombined and the second signal to be recombined, is the synthesized voice signal.

[0169] Due to the operations used when implementing the merge step, the merge step here corresponds to a QMF synthesis.

[0170] The method thus makes it possible to obtain the voice signal to be synthesized from the original MELP stream (coded by the MELP technique).

[0171] The method described exploits a principle of artificial band extension applied to the MELP decoder.

[0172] More specifically, the method uses the prediction of the parameters of the HF band voice signal. This makes it possible to obtain a synthesis scheme over a wide band (the frequency band between 0 kHz and 8 kHz) of the voice signal.

[0173] This scheme is easily integrated into a conventional MELP decoder, i.e. a MELP decoder in which the signal to be decoded is only analyzed, transmitted and synthesized up to 4 kHz, which limits the quality of the signals generated.

[0174] The method thus extends the frequency band of the MELP decoder from the BF band to the meeting of the BF and HF bands.

[0175] The quality of the synthesized voice is then improved, and thus closer to the original voice.

[0176] This improvement in voice synthesis and easy integration can also be achieved with a different implementation of the synthesis method.

[0177] To this end, another synthesis method is now described with reference to [Fig.3] which illustrates the implementation of such another method.

[0178] The synthesis method comprises an obtaining step E120, an extraction step E122, a first application step E124, a correction step E126, a conversion step E129, a fusion step E140 and a third application step E150.

[0179] The same remarks as for the synthesis process described with reference to [Fig.2] apply to this process and are not repeated. Only the differences are underlined.

[0180] More precisely, the obtaining step E120, the extraction step E122, the first application step E124, the correction step E126 and the conversion step E129 are similar to the steps E20 to E28 described previously.

[0181] At the end of the implementation of these steps E120 to E129, the computer 14 has the first gain vector and the first corrected spectral envelope parameters (description of the voice signal to be synthesized on the 4 kHz - 8 kHz band) and the second gain vector and the second spectral envelope parameters (description of the voice signal to be synthesized on the 0 kHz - 4 kHz band).

[0182] During the merging step E140, the computer 14 merges parameters from the set of output parameters and the parameters extracted from the coded stream.

[0183] For this, the computer 14 calculates the impulse response on each BF and HF band from the aforementioned parameters (operations 0142 in [Fig.4]), which is the impulse response associated with the coefficients of the linear prediction filter.

[0184] More specifically, the computer 14 calculates the impulse response on the HF band from the first set of gains and the first corrected spectral envelope parameters.

[0185] The impulse response on the HF band is calculated using the parameters extracted from the coded stream.

[0186] On each of the impulse responses, the computer 14 applies an oversampling operation to obtain signals to be recombined (operations 0144 on the [Fig.l]).

[0187] The computer 14 then applies a recombination operation to the signals to be recombined.

[0188] As represented schematically in [Fig.4] by operations 0146 and 0148, the recombination operation consists, according to the example described, in adding the first recombined signal (HF recombined signal) translated in frequency with the second recombined signal (LF recombined signal).

[0189] The calculator 14 then performs a calculation operation 0149 aimed at obtaining the LPC coefficients in wide band (combination of the BF and HF bands).

[0190] For this, the calculator 14 implements, for example, an autocorrelation calculated on the impulse response obtained previously followed by an LPC analysis.

[0191] LPC analysis is, for example, implemented by applying a Levinson-Durbin technique.

[0192] A set of LPC coefficients is thus obtained representing the voice signal to be synthesized in an LPC model.

[0193] Any other technique for obtaining the LPC coefficients can be used during this calculation operation 0149.

[0194] In contrast to the technique corresponding to the embodiment of [Fig.2], the present fusion step corresponds to a broadband fusion synthesis in the parametric domain.

[0195] During the third application step E150, the computer 14 applies a mixed excitation linear prediction synthesis technique to the parameters obtained after the fusion step to obtain the voice signal to be synthesized.

[0196] As visible in [Fig.4] with arrows 160 and 170, the synthesis technique can also be applied to additional parameters such as output parameters and / or other extracted parameters.

[0197] The method according to the embodiment of [Fig.4] has the same advantages as the method according to the embodiment of [Fig.2].

[0198] Furthermore, thanks to the fusion technique used, the synthesized signal does not include any dissociation of the two bands and a significant energy dip around 4 kHz inherent in the QMF synthesis process.

[0199] Furthermore, it is easy to apply voicing to the HF signal.

[0200] For example, the method according to the embodiment of [Fig.2] can be used. with an unvoiced hypothesis.

[0201] According to another example, an extension of the voicing cutoff frequency beyond 4 kHz may be carried out if the MELP voicing indicator in the frequency band of 3 to 4 kHz is active, using an arbitrarily defined frequency (e.g., voicing extension up to 6 kHz).

[0202] According to yet another example, the neural network is capable of predicting, in addition to the output parameters already mentioned such as gain or spectral envelope parameters, voicing indicators.

[0203] Several voicing indicators may be used, each corresponding to a frequency band. As an illustration, for 4 indicators, one indicator could be used per 1 kHz band between 4 kHz and 8 kHz.

Claims

Claims

1. A method for synthesizing a voice signal, the method being implemented by computer and comprising the steps of: - obtaining a stream coded according to a mixed excitation linear prediction coding, - extracting the parameters of the coded stream, - applying a neural network to parameters derived from the extracted parameters to obtain a set of output parameters, the set of output parameters comprising a first set of gains and first spectral envelope parameters, the gain and the spectral envelope being respectively the gain and the spectral envelope of the voice signal to be synthesized in the frequency band between 4 kHz and 8 kHz, the set of output parameters comprising the contrast of the envelope of the voice signal to be synthesized in the frequency band between 4 kHz and 8 kHz, - correcting the first spectral envelope parameters,using the contrast to obtain first corrected spectral envelope parameters, - generating a first signal from generation parameters, the generation parameters comprising parameters from the first set of gains and the first corrected spectral envelope parameters, - applying a mixed excitation linear prediction synthesis technique to the extracted parameters to obtain a second signal corresponding to the speech signal to be synthesized for frequencies below 4 kHz, and - merging the first signal and the second signal to obtain the speech signal to be synthesized.,

2. A synthesis method according to claim 1, wherein the merging step comprises an oversampling operation followed by an operation of applying a quadrature mirror filter, the two operations are applied to the first signal and the second signal, to obtain respectively a first intermediate signal and a second intermediate signal, the merging step comprising the addition of the first intermediate signal and the second intermediate signal to obtain the voice signal to be synthesized.

3. A synthesis method according to claim 1 or 2, wherein the pa- extracted parameters include the spectral line frequency coefficients representing the spectral envelope of the coded stream, the parameters from the extracted parameters including the spectral line frequency coefficients representing the spectral envelope of the coded stream.

4. A synthesis method according to any one of claims 1 to 3, wherein the extracted parameters comprise the spectral line frequency coefficients representing the spectral envelope of the coded stream and wherein, during the step of applying the neural network, an operation of converting the spectral line frequency coefficients representing the spectral envelope of the coded stream into cepstral coefficients representing the spectral envelope of the coded stream is implemented, the parameters derived from the extracted parameters comprising the cepstral coefficients representing the spectral envelope of the coded stream.

5. A synthesis method according to any one of claims 1 to 3, wherein the extracted parameters comprise the frequency coefficients of spectral lines representing the spectral envelope of the coded stream and wherein, during the step of applying the neural network, operations of converting the frequency coefficients of spectral lines representing the spectral envelope of the coded stream are implemented to obtain coefficients in the frequency domain representing the spectral envelope of the coded stream, the parameters resulting from the extracted parameters comprising the coefficients in the frequency domain representing the spectral envelope of the coded stream.

6. A synthesis method according to any one of claims 1 to 5, wherein, in the applying step, the neural network also predicts voicing parameters, the generation parameters comprising the predicted voicing parameters.

7. A synthesis method according to any one of claims 1 to 6, wherein the neural network is a recurrent neural network, preferably a long short-term memory network or a gated recurrent neural network.

8. A synthesis method according to claim 7, wherein the coded stream is organized into frames, the neural network being applied to four frames, two previous frames, one current frame and one future frame.

9. Device (10) for synthesizing a voice signal, the synthesis device (10) being capable of: - obtain a coded flow according to a mixed excitation linear prediction coding, - extract parameters from the coded stream, - applying a neural network to parameters from the extracted parameters to obtain a set of output parameters, the set of output parameters comprising a first set of gains and first spectral envelope parameters, the gain and the spectral envelope being respectively the gain and the spectral envelope of the voice signal to be synthesized in the frequency band between 4 kHz and 8 kHz, the set of output parameters comprising the contrast of the envelope of the voice signal to be synthesized in the frequency band between 4 kHz and 8 kHz, - correcting the first spectral envelope parameters, using the contrast to obtain first corrected spectral envelope parameters, - generating a first signal from generation parameters, the generation parameters comprising parameters from the first set of gains and the first corrected spectral envelope parameters, - apply a mixed excitation linear prediction synthesis technique on the extracted parameters to obtain a second signal corresponding to the voice signal to be synthesized for frequencies below 4 kHz, and - merge the first signal and the second signal to obtain the voice signal to be synthesized.

10. A computer program product comprising instructions which, when the program is executed by a computer, cause the latter to implement the steps of a method according to any one of claims 1 to 8.

11. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the steps of a method according to any one of claims 1 to 8.