Audio encoding and decoding method and device
The audio encoding and decoding method employs a machine learning-trained model with a recurrent neural network to enhance audio quality and reduce resource consumption, addressing the challenge of high compression ratios in limited-resource platforms.
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- THALES SA
- Filing Date
- 2025-10-13
- Publication Date
- 2026-04-15
AI Technical Summary
Existing audio encoding and decoding systems face challenges in achieving high compression ratios with good reconstruction quality while minimizing computing and energy resource consumption, particularly in platforms with limited resources.
An audio encoding and decoding method using a machine learning-trained compression and quantization model combined with a bitrate reduction model, specifically a recurrent neural network, to predict and add refinement layers to the compressed signal, enhancing quality while reducing computational and energy consumption.
The method achieves improved audio reconstruction quality and reduced resource consumption, enabling efficient encoding and decoding in low-power, low-computing environments.
Smart Images

Figure IMGAF001_ABST
Abstract
Description
[0001] The present invention relates to a method and device for audio encoding and decoding, as well as an associated computer program.
[0002] The invention lies in the field of audio coding and decoding, for the realization of audio coding and decoding devices, also called audio codecs.
[0003] More specifically, the invention lies in the field of low or very high bitrate audio codecs, which can be embedded on platforms with limited computing and energy resources.
[0004] Many practical applications require audio communication, for example commands or information spoken by an operator to a remote operator, in real time and of sufficient quality for comprehension.
[0005] It is well known that audio signals, especially speech signals, need to be compressed to enable their transmission over a data rate-constrained communication network. Audio encoding and decoding devices have been developed for this purpose. Such devices are used in many communication systems, for example, radio communication systems or videoconferencing systems.
[0006] One problem to solve is to obtain the most faithful possible reproduction of the audio signal before compression, with the lowest possible transmission rate, or in other words, the best possible compression ratio.
[0007] To this end, numerous audio encoding and decoding systems have been developed. Recently, systems based on artificial intelligence have been proposed.
[0008] In addition to the constraints of throughput and quality of reproduction, it is also important to consider the limitation of the consumption of computing and energy resources of an audio encoding and decoding device, in order to allow it to be embedded on platforms with reduced computing and / or energy resources.
[0009] The aim of the invention is therefore to remedy this problem, by proposing a method and device for audio encoding and decoding that combines a high compression ratio, good reconstruction quality and low resource consumption.
[0010] To this end, the invention relates to an audio encoding and decoding method implemented in an audio encoding and decoding device comprising an audio encoder implementing a machine learning-trained compression and quantization model to provide, as output, for an input audio signal, a compressed signal consisting of symbols represented on N successive quantization layers, the first N1 quantization layers of said N quantization layers corresponding to a first compression rate, and an audio decoder configured to reconstruct an audio signal from a compressed signal consisting of symbols. This method comprises the following steps implemented by a computing processor: application of a bitrate reduction model, previously trained by machine learning on a database of compressed audio signals, to predict from the first N1 quantization layers, a number N2 of subsequent quantization layers, called refinement layers, the bitrate reduction model being a recurrent neural network, said bitrate reduction model being implemented on a received compressed signal, formation of an enriched compressed signal by adding said refinement layers, the enriched compressed signal being provided to an audio decoding module to obtain the reconstructed audio signal.
[0011] Advantageously, using a bitrate reduction model allows for the prediction, during decoding, of the refinement layers required to obtain an enriched compressed signal and, consequently, a higher-quality reconstructed audio signal. In other words, starting from a low-bitrate compressed signal, the bitrate reduction model yields an enriched compressed signal equivalent to high-bitrate compression. Furthermore, the use of a recurrent neural network allows for implementation with low computational and energy consumption, which is particularly useful for implementing such a process in embedded communication systems.
[0012] The audio encoding and decoding process according to the invention may also have one or more of the following characteristics, taken independently or in all technically feasible combinations.
[0013] The application of a rate reduction model involves applying a transformation of each symbol of the received compressed signal into a vector of predetermined dimension, said vectors forming a time sequence of vectors which are provided as input to said recurrent neural network.
[0014] The recurrent neural network is formed from a succession of recurrence layers, each recurrence layer containing a plurality of recurrence blocks.
[0015] The number of recurrence layers is less than or equal to 10.
[0016] An internal dimension and an output dimension of said recurrent neural network are reduced by applying a division parameter.
[0017] The division parameter has a chosen value, preferably between 5 and 30.
[0018] According to another aspect, the invention relates to an audio encoding and decoding device comprising an audio encoder implementing a machine learning-trained compression and quantization model to provide, as output, for an input audio signal, a compressed signal consisting of symbols represented on N successive quantization layers, the first N1 quantization layers of said N quantization layers corresponding to a first compression rate, and an audio decoder configured to reconstruct an audio signal from a compressed signal consisting of symbols, comprising a processing unit configured to implement: a bitrate reduction module, implementing a bitrate reduction model previously trained by machine learning on a database of compressed audio signals, to predict from the first N1 quantization layers, a number N2 of subsequent quantization layers, called refinement layers, the bitrate reduction model being a recurrent neural network, said bitrate reduction model being implemented on received compressed signal, a training module for an enriched compressed signal by adding said refinement layers, the enriched compressed signal being provided to a decoding module to obtain the reconstructed audio signal.
[0019] According to another aspect, the invention relates to an information recording medium, on which software instructions are stored for the execution of an audio encoding and decoding process as briefly described above, when these instructions are executed by a programmable electronic device.
[0020] According to one particular aspect, the rate reduction module is configured to apply a transformation of each symbol of the received compressed signal into a vector of predetermined dimension, said vectors forming a time sequence of vectors which are provided as input to said recurrent neural network.
[0021] According to one particular aspect, the recurrent neural network is formed of a succession of recurrence layers, each recurrence layer comprising a plurality of recurrence blocks, the number of recurrence layers being less than or equal to 10.
[0022] According to another aspect, the invention relates to a computer program comprising software instructions which, when implemented by a programmable electronic device, implement an audio encoding and decoding process as briefly described above.
[0023] The invention will become clearer upon reading the following description, given solely by way of non-limiting example, and made with reference to the drawings in which: there figure 1 is a schematic representation of an audio encoding / decoding system according to one embodiment; the figure 2 is a synoptic diagram of the main functional blocks of a flow reduction module according to one embodiment; the figure 3 is a flowchart of an audio encoding / decoding process according to one embodiment; the figure 4 is a functional diagram of an audio encoding / decoding device according to one embodiment.
[0024] The invention applies in the context of encoding / decoding audio signals with low and very low bitrate compression, and finds particular applications in audio communication devices embedded on any type of platform, for example portable communication devices.
[0025] Low bitrate audio compression refers to compression between 1 and 5 kbps (kilobits per second). Very low bitrate audio compression corresponds to a bitrate below 1 kbps.
[0026] The audio signals processed are typically speech signals, spoken by an operator.
[0027] With reference to the figure 1 , an audio encoding and decoding system 2 comprises an encoder and transmitter 4, referred to hereafter as encoder 4, and a receiver and decoder 6, referred to hereafter as decoder 6.
[0028] An audio signal SE is provided as input to a compression block 8 of the encoder 4, which implements a machine learning-trained compression model on a database of audio signals. For example, the compression model is a neural network.
[0029] As is known, a neural network consists of an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.
[0030] More specifically, each layer comprises neurons taking their inputs from the outputs of the neurons in the previous layer, or from the input variables for the first layer.
[0031] Alternatively, more complex neural network structures can be considered with a layer that can be linked to a layer further away than the immediately preceding layer.
[0032] Each neuron is also associated with an operation, that is, a type of processing, to be carried out by said neuron within the corresponding processing layer.
[0033] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a link between two neurons. This is often a real number, which takes on both positive and negative values. In some cases, the synaptic weight is a complex number.
[0034] Each neuron performs a weighted summation of the value(s) received from the neurons in the preceding layer. Each value is then multiplied by the respective synaptic weight of each synapse, or connection, between that neuron and the neurons in the preceding layer. Next, an activation function, typically a non-linear function, is applied to this weighted summation. The resulting value is then delivered to the neuron's output, particularly to the neurons in the next layer connected to it. The activation function introduces non-linearity into the processing performed by each neuron. The sigmoid function, the hyperbolic tangent function, and the Heaviside function are examples of activation functions.
[0035] As an optional complement, each neuron is also capable of applying, in addition, a multiplicative factor, and an additive bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the value of the multiplicative factor and the value from the activation function, plus the bias.
[0036] A convolutional neural network is also sometimes called a convolutional neural network or by the acronym CNN, which refers to the English term " Convolutional Neural Networks ».
[0037] In a convolutional neural network, each neuron in the same layer has exactly the same connection pattern as its neighboring neurons, but at different input positions. The connection pattern is called the convolution kernel or, more often, " kernel » in reference to the corresponding English name.
[0038] A fully connected layer of neurons is a layer in which the neurons of said layer are each connected to all the neurons of the preceding layer.
[0039] This type of layer is more often referred to by the English term " fully connected » , and sometimes referred to as the "dense layer".
[0040] The values of the weights, multipliers and biases where applicable are learned during a machine learning phase to perform the classification or prediction task.
[0041] At the output of the compression block 8, a quantization block 10 is applied.
[0042] For example, quantization block 10 is implemented by learning a dictionary representing the space of compressed audio signals. The quantization block then associates an input vector with the index of the dictionary word closest to that vector. This quantization method is called "vector quantization".
[0043] At the output of the quantization block 10, for the input audio signal SE, a corresponding compressed and quantized signal Sc is obtained, comprising a series of symbols (or "tokens" in English) s-c1,...,s-ct, each symbol having a quantized value on a given number N of quantization levels or depth levels, represented on a depth axis.
[0044] As a non-limiting example, in the figure 1 , N=8.
[0045] Denoting M as the number of possible values per level of quantization, the number of possible values of the symbols is then equal to M N< .
[0046] Depending on the depth dimension, the symbols forming the compressed signal are represented in N successive quantization layers, each layer corresponding to a level of quantization, starting from a first level of quantization (or depth) 1, corresponding to the coarsest quantization, up to the level of quantization (or depth) N corresponding to the finest quantization.
[0047] Thus, the first N1 quantization layers, for example N1=3, correspond to a first compression rate, which is a low rate, while the N quantization layers correspond to a second compression rate which is higher than the first compression rate.
[0048] In other words, the first compression rate corresponds to a first compression ratio and the second compression rate corresponds to a second compression ratio, the first compression ratio being higher than the second compression ratio.
[0049] In order to facilitate transmission over a communication network 5, preferably the first compression ratio is applied.
[0050] A portion SC1 of the compressed signal, corresponding to the first N1 quantization layers, for example N1 between N / 4 and 3*N / 4, is transmitted via the communication network 5 to the receiver and decoder 6. The portion SC1 is a compressed signal corresponding to a first rate compression, low or very low rate, of the audio signal SE.
[0051] The decoder 6 includes a rate reduction module 14, which implements a rate reduction model 15, previously trained on a database of compressed audio signals, the compressed audio signals in this database being able to be different from the compressed audio signals corresponding to the audio signals in the database used for training the encoder, the rate reduction model being trained by machine learning to predict from the first N1 quantization layers, the next N2 quantization layers, called refinement layers, respecting the condition N1+N2 less than or equal to N (i.e. N2 less than or equal to N-N1).
[0052] For example, N2=N-N1.
[0053] In the example illustrated in the figure 1 , N2=5.
[0054] Thus, advantageously, the flow reduction module 14 allows us to predict the refinement layers 12 which have not been transmitted for all transmitted symbols, the flow corresponding to these refinement layers then being saved.
[0055] The decoder 6 further includes a recomposition module 16, which recomposes an enriched compressed signal, with N1+N2 quantization layers, from the received compressed signal and the predicted refinement layers.
[0056] The enriched compressed signal is provided as input to a decoding module 18 which provides a reconstructed SR audio signal.
[0057] Advantageously, the reconstructed SR audio signal is of better quality than a reconstructed audio signal by providing at input to the decoder 6 only the S C1 portion of the compressed signal corresponding to the first N1 quantization layers only.
[0058] Advantageously, the rate reduction module 14 implements a recurrent neural network, for example an LSTM neural network (for "Long Short Term Memory").
[0059] There figure 2 illustrates an implementation of a module 14 flow reduction based on a recurrent neural network.
[0060] The rate reduction module 14 receives as input a portion of compressed signal corresponding to the first N1 quantization layers, and comprising T successive symbols, each symbol being coded on N1 quantization layers.
[0061] Module 14 is configured to implement a transformation 30 of each s-cj symbol into a vector Xj of predetermined dimension Dim. Such a transformation is known as "embedding" in English.
[0062] The predetermined dimension Dim is typically equal to 64, 128 or 256, preferably equal to 64 in order to limit the size of the vectors and consequently, to limit the number of calculations.
[0063] A time sequence of vectors X 1 ,.. XT represented in a discrete space of dimension Dim, also called latent space, is obtained.
[0064] The rate reduction module 14 is configured to apply a recurrent LSTM model at step 32. The vectors X1...XT are then provided, in ascending time order, to the recurrent neural network.
[0065] Step 32 of applying an LSTM recurrence pattern implements 34 recurrence blocks 1 to 34 L on L levels or layers of recurrence. Each recurrence block applies predefined operations.
[0066] For a given recurrence layer, each 34j recurrence block takes as input the results of corresponding calculations applied to temporally preceding elements. For a subsequent recurrence layer, the outputs of the previous recurrence layer are used as input.
[0067] In the example, for a recurrence layer of index I, a recurrence block takes elements as input c t − 1 l , h t − 1 l which are results of the previous temporal recurrence block (time index t-1) of the same recurrence layer and an element h t l − 1 output of the previous recurrence layer of index I-1.
[0068] For the first recurrence layer, the elements X 1 ... XT are provided as input.
[0069] Advantageously, the recurrent neural network model used is sized to have a limited number of parameters and a limited number of operations to perform, so as to consume less memory and computing resources.
[0070] Preferably, in order to limit the number of computing resources used, the number L of recurrence layers is limited, for example L is less than or equal to 10.
[0071] Preferably, the internal dimension (dimension of intermediate results, vectors C and H) and the output dimension of the LSTM recurrent neural network is reduced to M / div, where "div" is a division parameter to be adjusted to meet computational complexity constraints.
[0072] For example, the division parameter "div" has a value chosen so as to limit the number of parameters implemented in the recurrent neural network to a predetermined maximum number.
[0073] For example, the division parameter "div" has a value between 5 and 30
[0074] The throughput reduction module 14 is configured to then apply a projection step 36 (commonly called a dense layer) to transform the latent output space of the recurrent neural network into non-normalized probability predictions of the corresponding refinement layers. The number of predictions is equal to (N² x M) x T.
[0075] At the output, the flow reduction module provides the predicted N2 refinement layers (reference 12), from the probabilities obtained in step 36.
[0076] Advantageously, the described flow reduction module 14 performs a prediction of all refinement layers simultaneously, which improves the time complexity.
[0077] The main steps of an audio encoding and decoding process in an embodiment are described with reference to the figure 3 .
[0078] On the encoder side, the process involves, on an input SE audio signal, the implementation 40 of a compression and quantization model 42.
[0079] Preferably the compression and quantization model is a machine learning model trained on a database of audio signals, preferably of good quality audio signals (intelligible, without noise...).
[0080] For example, in one embodiment the compression and quantization model is formed of a neural network and a vector quantization layer as described above.
[0081] The process then includes a transmission step 44 of a compressed signal formed from the first N1 quantization layers, a crude representation of the input audio signal corresponding to low or very low bitrate coding.
[0082] On the decoder side, the compressed signal is received at reception stage 46.
[0083] Next, the previously trained rate reduction model is implemented in step 48 on the received compressed signal to predict N2 non-transmitted quantization layers. For example, a rate reduction module such as the one described above with reference to the figure 2 is being implemented.
[0084] The process then includes a step 50 of formation of an enriched compressed signal, with N1+N2 quantization layers.
[0085] This enriched compressed signal is provided as input to a decoder 52, which allows the reconstruction of an audio signal called the reconstructed audio signal, SR, from a signal compressed by the compression and quantization model 42.
[0086] In one embodiment, the audio encoding and decoding process is implemented by an audio encoding and decoding device, also called an audio codec.
[0087] There figure 4 schematically illustrates an audio codec 60 configured to implement an audio encoding and decoding process, according to one embodiment.
[0088] The 60 audio codec is a programmable electronic device.
[0089] The device 60 includes at least one computing processor 62, at least one electronic memory unit 64, a communication interface 66 and optionally, a human-machine interface 68, adapted to communicate via a communication bus 65.
[0090] The electronic memory 64 is configured to store the compression and quantization model 42, previously trained for this task, and the rate reduction model 15, previously trained for this task.
[0091] The computing processor 62 is configured to run an encoder module 70, which applies the compression and quantization model to an input audio signal, to provide a compressed signal, applied when encoding an audio signal.
[0092] The compute processor 62 is also configured to execute the following modules when decoding an audio signal: a rate reduction module 72, configured to apply the rate reduction model previously trained on a database of compressed audio signals, to predict from the first N1 quantization layers, an N2 number of subsequent quantization layers, a module 74 for training an enriched compressed signal, a decoding module 76, which implements a suitable decoding to reconstruct an audio signal from a signal compressed by the audio encoder 70.
[0093] In one embodiment, modules 70, 72, 74, 76 are implemented in the form of software instructions forming a computer program, which, when executed by a programmable electronic device, implements an automatic unlearning analysis method for at least one class by means of a data classification model according to the invention.
[0094] In an alternative not shown, modules 70, 72, 74, and 76 are each implemented as programmable logic components, such as FPGAs (from the English Field Programmable Gate Array ), microprocessors, GPGPU components (from English General-purpose processing on graphies processing ), or even dedicated integrated circuits, such as ASICs (from the English Application Specific Integrated Circuit ).
[0095] The computer program, containing software instructions, is also capable of being stored on a non-transient, computer-readable information storage medium. This computer-readable medium is, for example, a medium capable of storing electronic instructions and being connected to a bus of a computer system. Examples of such media include optical discs, magneto-optical discs, ROMs, RAM, any type of non-volatile memory (e.g., EPROM, EEPROM, FLASH, NVRAM), magnetic cards, or optical cards.
[0096] Advantageously, the described process allows for a gain in throughput by going from 2400 bits / second to 600 bits / second while maintaining sufficient audio quality and intelligibility.
[0097] Advantageously, the described audio encoding and decoding process can be implemented in a low power consumption codec, typically less than 50mW, through the implementation of neural network models with a limited number of parameters, typically less than 2 million.
Claims
1. An audio encoding and decoding method, implemented in an audio encoding and decoding device comprising an audio encoder (4) implementing a machine learning-trained compression and quantization model to provide, as output, for an input audio signal, a compressed signal formed of symbols represented on N successive quantization layers, the first N1 quantization layers of said N quantization layers corresponding to a first compression rate, and an audio decoder (6) configured to reconstruct an audio signal from a compressed signal formed of symbols, the method being characterized in thatIt comprises the following steps implemented by a computing processor: - application (48) of a bitrate reduction model (15), previously trained by machine learning on a database of compressed audio signals, to predict, from the first N1 quantization layers, a number N2 of subsequent quantization layers, called refinement layers, the bitrate reduction model (15) being a recurrent neural network, said bitrate reduction model (15) being implemented on a received compressed signal (S C1 ), - formation (50) of a compressed signal enriched by adding said refinement layers, the compressed enriched signal being provided to an audio decoding module to obtain the reconstructed audio signal.
2. A method according to claim 1, wherein the application of a rate reduction model (15) comprises an application of a transformation (30) of each symbol of the received compressed signal into a vector of predetermined dimension, said vectors forming a time sequence of vectors which are provided as input to said recurrent neural network.
3. Method according to claim 1 or 2, wherein said recurrent neural network is formed of a succession of recurrence layers, each recurrence layer comprising a plurality of recurrence blocks.
4. A method according to claim 3, wherein the number of recurrence layers is less than or equal to 10.
5. A method according to any one of claims 3 or 4, wherein an internal dimension and an output dimension of said recurrent neural network are reduced by applying a division parameter.
6. Method according to claim 4, wherein said division parameter has a chosen value, preferably between 5 and 30.
7. Computer program comprising software instructions which, when executed by an audio encoding and decoding device, implement an audio encoding and decoding method in accordance with claims 1 to 6.
8. Audio encoding and decoding device, comprising an audio encoder (4) implementing a machine learning-trained compression and quantization model to provide as output, for an input audio signal, a compressed signal formed of symbols represented on a number N of successive quantization layers, the first N1 quantization layers of said N quantization layers corresponding to a first compression rate, and an audio decoder (6) configured to reconstruct an audio signal reconstructed from a compressed signal formed of symbols, characterized in that It includes a computing processor (62) configured to implement: - a rate reduction module (72), implementing a rate reduction model (15) previously trained by machine learning on a database of compressed audio signals, to predict from the first N1 quantization layers, a number N2 of subsequent quantization layers, called refinement layers, the rate reduction model being a recurrent neural network, said rate reduction model (15) being implemented on the received compressed signal, - a module (74) for forming an enriched compressed signal by adding said refinement layers, the enriched compressed signal being provided to a decoding module (76) to obtain the reconstructed audio signal.
9. Device according to claim 8, wherein said rate reduction module (72) is configured to apply a transformation (30) of each symbol of the received compressed signal into a vector of predetermined dimension, said vectors forming a time sequence of vectors which are supplied as input to said recurrent neural network.
10. Device according to claim 8 or 9, wherein said recurrent neural network is formed of a succession of recurrence layers, each recurrence layer comprising a plurality of recurrence blocks, the number of recurrence layers being less than or equal to 10.
Citation Information
Patent Citations
Methods and apparatus for rate quality scalable coding with generative models
US20220044694A1