Modular nesc

The modular neural speech codec addresses computational inefficiencies by using sub-band decomposition and multiple neural encoders/decoders, achieving scalable complexity, bitrate, and audio bandwidth adaptation for diverse devices.

WO2025237540A1PCT designated stage Publication Date: 2025-11-20FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/063769
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

State-of-the-art neural speech coders are computationally expensive and lack scalability to cater to devices with varying computational limitations and requirements.

Method used

A modular neural speech codec (NESC) with sub-band decomposition and multiple neural encoders/decoders, allowing for scalability in complexity, bitrate, and quality by encoding and decoding individual sub-bands, and incorporating guided or blind bandwidth extension.

Benefits of technology

The modular approach enhances efficiency by enabling scalable complexity, bitrate, and audio bandwidth adaptation to different devices, improving computational efficiency and perceptual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024063769_20112025_PF_FP_ABST
    Figure EP2024063769_20112025_PF_FP_ABST
Patent Text Reader

Abstract

Encoder (500) for encoding an audio signal (AS), comprising: a sub-band decomposition entity (510) configured to obtain individual groups of sub-bands (SBa, SBb, SBc) of the audio signal (AS) based on the audio signal (AS); and several neural encoders (514a, 514b, 514c), each involving learnable layers and configured to code individual groups of sub- bands (SBa, SBb, SBc) or a single full band signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Modular NESC

[0002] Description

[0003] Embodiments of the present invention refer to an encoder and a decoder. Further embodiments refer to the corresponding methods for encoding and decoding and to the computer programs. In general, embodiments are in the field of neural end-to-end speech codecs.

[0004] State-of-the-art (SOTA) neural speech coders [1 ,2,3] typically encode the time-domain waveform with a neural encoder, quantize the resulting representation, and decode the quantized representation to reconstruct the original signal.

[0005] Fig. 1 shows an exemplarily block diagram of a neural end-to-end speech codec (NESC). It comprises a DNN encoder 110 and a DNN decoder 120. In between quantization entity 115 may be arranged. The DNN encoder 110 receives an input signal IS, e.g. a 16 kHz input signal IS. The DNN decoder 120 is configured to output a plurality of decoded signals to a PQMF synthesis entity, here comprising four bands. This entity is marked by the reference numeral 125 and outputs the output (audio) signal AS’.

[0006] In recent times, such neural codecs have seen unprecedented growth and adaptation. They can produce good quality WB(Wideband) / SWB(Super-Wideband) speech signals [3, 4] at very low bitrates. These codecs are extremely powerful but are also computationally expensive. With a plethora of available devices and their varying computational limitations, there is a need for a neural codec that can cater to different requirements. In this solution, we envision modularity in neural codec which is scalable in terms of complexity, bitrate and quality.

[0007] It is and objective of the present invention to improve the efficiency of the codec, i.e. on the encoder side and the decoder side.

[0008] This objective is solved by the subject-matter of the independent claims.

[0009] An embodiment provides an encoder for encoding an audio signal. The encoder comprises sub-band decomposition entity and several neural encoders. The sub-band decomposition entity is configured to obtain individual groups of sub-bands of the audio signal. Each of the several neural encoders involve learnable layers and are configured to code individual groups of sub-bands, or a single full band signal.

[0010] According to another embodiment, a decoder for decoding an audio signal is provided. The decoder comprises several neural decoders and a filter-bank synthesis. Each of the several neural decoders involve learnable layers and are configured for decoding individual groups of sub-bands of the audio signal. The filter-bank synthesis is configured to generate the audio signal based on the individual group of sub-bands.

[0011] Another embodiment provides a system comprising the encoder and the decoder as discussed above. Here, one of the several neural encoders forms together with one of the several neural decoders a module. Note, according to further embodiments each module may be treated separately or together. This means that the encoder and the decoder may be trained together.

[0012] Embodiments of the present invention are based on the principle that it has been found that modular NESC codecs comprising of multiple any S-like sub-modules that can encode and / or decode multiple sub-bands enable to increase the efficiency. According to embodiments, a respective neural encoder of the several neural encoders and a respective neural decoder of the several neural decoders form a module. According to embodiments, each module may comprise of

[0013] • An encoder: Takes as input either directly the signal at different sampling rates (e.g. at 16, 24 or 32 kHz) or the sub-bands corresponding to the signal obtained from PQMF analysis filter or from any tine-frequency decomposition. It outputs a learned representation.

[0014] • A quantizer: Takes as input the learned representation from the encoder and outputs a quantized version of it.

[0015] • A decoder: Takes as input the quantized representation from the quantizer and uses it to generate the sub-band signal

[0016] According to further embodiments, each module is configured to encode and decode multiple sub-bands of the PQMF synthesis filter of any time-frequency decomposition. This means that the output side of the decoder a PQMF synthesis filterbank or any inverse transformation may be used. According to embodiments, the output of the different decoders may additionally be combined via a learnable convolutional layer (to compensate for overlapping and aliasing) and finally a PQMF synthesis filter produce the output speech.

[0017] According to embodiments, especially on the encoder side the individual groups of subbands have partially overlapping frequency range.

[0018] According to embodiments, especially on the encoder side the sub-band decomposition entity comprises an analysis filter, especially a PQMF analysis filter or a block transform or a filterbank.

[0019] According to embodiments, especially on the encoder side the several neural encoders are part of several modules. According to embodiments, especially on the encoder side the several neural encoders are part of distinct modules.

[0020] According to embodiments, especially on the encoder side the encoder comprises several quantizers for the several neural encoders.

[0021] According to embodiments, especially on the encoder side the several neural encoders are configured for varying bitrate and / or differ from each other with respect to its complexity. According to further embodiments, especially on the encoder side the several neural encoders can be activated or deactivated for achieving varying bitrate and / or varying complexity.

[0022] According to embodiments, especially on the encoder side each (distinct) module is trained separately or together with baseband module.

[0023] According to embodiments, especially on the encoder side at least one module being always active is configured to code a baseband audio signal.

[0024] According to embodiments, especially on the encoder side at least one module being configured to be adaptively activated is configured to extend the baseband audio signal encoder.

[0025] According to embodiments, especially on the encoder side at least one module is configured for guided bandwidth extension. According to embodiments, especially on the decoder side the individual groups of subbands have partially overlapping frequency range.

[0026] According to embodiments, especially on the decoder side the filter-bank synthesis is configured to generate the audio signal at different possible sampling rates and / or bandwidths.

[0027] According to embodiments, especially on the decoder side the decoder further comprises a base band decoder.

[0028] According to embodiments, especially on the decoder side the decoder comprises a learned conventional layer configured to combine at least two sub-bands.

[0029] According to embodiments, especially on the decoder side the filter-bank synthesis comprises a PQMF synthesis filter or an inverse block transform.

[0030] According to embodiments, especially on the decoder side the several neural decoders form parts of several or distinct modules.

[0031] According to embodiments, especially on the decoder side at least one module is configured for extended the bandwidth of the decoded signal obtained by the base-band decoder. According to embodiments, especially on the decoder side at least one module is configured for guided bandwidth extension or blind bandwidth extension.

[0032] Further embodiments provide a method for encoding an audio signal comprising the following steps:

[0033] • obtaining individual groups of sub-bands of the audio signal based on the audio signal using sub-band decomposition; and

[0034] • coding individual groups of sub-bands, or a single full band signal using several neural encoders, each involving learnable layers.

[0035] Another embodiment provides a method for decoding an audio signal, comprising the following steps: decoding individual groups of sub-bands of the audio signal using several neural decoders, each involving learnable layers;

[0036] • generating the audio signal based on the individual groups of sub-bands using filterbank synthesis.

[0037] Of course, embodiments of the present invention may be computer-implemented. Thus, another embodiment provides a computer program code for performing, when running on a computer, the steps of the method for decoding and / or the method for encoding.

[0038] Embodiments of a side aspect provide an encoder for encoding an audio signal comprising a sub-band decomposition entity, neural encoders and a BWE encoder. The sub-band decomposition entity may be configured to obtain individual groups of sub-bands of the audio signal based on the audio signal. The neural encoders, each involving a learnable layer, is configured to code individual groups of sub-bands of the audio signals, wherein the different groups of the sub-bands have partially overlapping frequency ranges. The BWE encoder is configured to perform guided bandwidth extension.

[0039] For this side aspect, another embodiment provides a decoder for decoding an audio signal comprising neural decoders, a filter-bank synthesis and a BWE decoder. Each of the neural decoders involve learnable layers and are configured to decoder individual groups of subbands of the audio signal. The filter-bank synthesis is configured for generating the audio signal based on the individual groups of sub-bands. The BWE decoder is configured to perform blind or guided bandwidth extension.

[0040] Another embodiment refers to a system comprising the encoder and the decoder of the side aspect.

[0041] Furthermore, embodiments may provide a method for encoding and a method for decoding using the slightly different approach. The method for decoding may comprise the steps:

[0042] - decoding individual groups of sub-bands of the audio signal using neural decoders, each involving learnable layers;

[0043] - generating the audio signal based on the individual groups of sub-bands; - performing blind or guided bandwidth extension.

[0044] The method for encoding may comprise the steps of obtaining individual groups of subbands of the audio signal based on the audio signal using sub-band decomposition;

[0045] - encoding individual groups of sub-bands of the audio signal using neural encoders each involving learnable layers, wherein different groups of sub-bands have partially overlapping frequency ranges;

[0046] - performing guided bandwidth extension.

[0047] It should be noted that according to embodiments, the encoder and the decoder of this side aspect may be combined with the features discussed of the encoder and the decoder of the main aspect.

[0048] Below, embodiments of the present will subsequently discussed referring to the enclosed figures, wherein:

[0049] Fig. 1 shows a schematic block diagram for illustrating neural end-to-end speech codec (NESC);

[0050] Fig. 2 shows a schematic block diagram for illustrating a modular NESC according to an embodiment;

[0051] Fig. 3 shows a schematic block diagram for illustrating a guided bandwidth extension according to an embodiment;

[0052] Fig. 4 shows a schematic block diagram for illustrating blind bandwidth extension according to an embodiment;

[0053] Fig. 5 shows a schematic block diagram of an encoder according to a basic aspect; and

[0054] Fig. 6 shows a schematic block diagram of a decoder according to a basic aspect.

[0055] Below, embodiments of the present invention will subsequently be discussed referring to the enclosed figures, wherein identical reference numerals are provided with objects having identical or similar functions, so that the description thereof is mutually applicable and interchangeable.

[0056] Fig. 5 shows an encoder 500 for encoding an audio signal AS. The encoder 500 comprises a sub-band decomposition entity 510 and several neural encoders 514a, 514b and 514c. Note, here, exemplarily three neural encoders are shown, wherein according to further embodiments also two or more than three may be used.

[0057] The sub-band decomposition entity 510 is configured to obtain individual groups of subbands of the audio signal based on the audio signal AS. The individual groups of sub-bands are marked by the reference numeral SBa, SBb and SBc. Dependent on the number of neural encoders 514a, 514b and 514c, the number of sub-bands SBA, SBB and SBC output by the entity 510 may vary. Each of the neural encoders 514a, 514b and 514c involves a learnable layer and is configured to code the individual groups of sub-bands SBa, SBb and SBn. Alternatively, a single full band may be coded by the several neural encoders 514a, 514b and 514c.

[0058] Regarding the sub-band decomposition entity, it should be noted that same may be implemented using an analysis filter, especially a PQMF analysis filter or a block transform or another filterbank. It should be noted that according to embodiments, the individual groups sub-bands SBA; SBB and SBC may have overlapping frequency ranges.

[0059] The encoder 500 outputs from the several neural encoders 514a, 514b and 514c several or a combined datastream including the content of the audio signal AS. The datastream is marked by DS and sent to the decoder or sent to the several neural decoders of the decoder.

[0060] Fig. 6 shows a decoder 600 receiving the datastream DS or the plurality of datastreams DS. The decoder 600 comprises a plurality of neural decoders 614a, 614b and 614c (here three, two or more than three are possible as well) in combination with a filter-bank synthesis entity 610.

[0061] The several neural decoders, where each involves a learnable layer are configured for decoding individual groups of sub-bands SBa, SBb and SBc of the audio signal AS. These are output to the filter-bank synthesis 610 which is configured to generate the audio signal AS’ based on the individual groups of sub-bands. Note, the generated audio signal AS’ substantially complies to the input audio signal AS but is synthesized and, thus, carries an own reference numeral. According to embodiments, the filter-bank synthesis may be configured to generate the audio signal AS’ at different possible sampling rates and / or audio bandwidths.

[0062] With respect to Fig. 2, a Modular NESC will be discussed using the encoder architecture 500 and the decoder architecture 600. According to embodiments, a neural encoder 514a, 514b or 514c may be combined with a neural decoder 614a, 614b and 614c, respectively, so as to form a module. The respective modules are marked by the reference numeral 599a, 599b, 599c and 599d.

[0063] The outputs of the four modules 599a to 599d are forwarded to the PQMF synthesis 612 as marked by the reference numeral SB-1 to SB-n. According to embodiments, so-called convolutional layers 616a and 616b may be arranged between the decoders 614a, 614b, 614c and the synthesis entity 612. Here, a convolutional layer 616a receives from two decoders 614a and 614b a respective signal and outputs the signal SB-2 to SB-3. The same combination is made by the convolutional layer 616b receiving the signal from 614b and 614c and outputting the signals SB-4 and SB-5 to 612.

[0064] The Modular NESC comprises multiple NESC-like sub-modules 599a to 599d that can encode and decode the multiple sub-bands SB-1 to SB-n of the PQMF synthesis filter 612. Each module 699a to 699d comprises according to embodiments:

[0065] • An encoder 514a-514d: Takes as input either directly the signal at different sampling rates (e.g. at 16, 24 or 32 kHz) or the sub-bands corresponding to the signal obtained from PQMF analysis filter. It outputs a learned representation.

[0066] • A quantizer 115: Takes as input the learned representation from the encoder 514a- 514d and outputs a quantized version of it.

[0067] • A decoder 614a-614d: Takes as input the quantized representation from the quantizer 115 and uses it to generate the sub-band signal.

[0068] In addition, the output of different decoders is combined via a learned convolutional layer 616a and 616b (to compensate for overlapping and aliasing), and finally a PQMF synthesis filter 612 produce the output speech IS’. Given that each sub-band has an audio bandwidth of 2kHz, each module can generate signal of 4kHz, this leads to scalability of generated audio bandwidth as each additional module can generate higher band signal. The choice of the number of sub-modules to use dictates:

[0069] • Bandwidth: since the addition of each module will produce higher bands.

[0070] • Complexity: since each module has a computational cost.

[0071] • Bit rate: since each quantized representation needs to be sent through the transmission channel.

[0072] By this configuration, it is possible that the Modular NESC is characterized by scalability in terms of complexity, bitrate and audio bandwidth of the handle signal of varying sampling rates.

[0073] As discussed above, the encoder 514a to 514d, the quantizer 115 and the decoder 614a to 614d of each sub-module 599a to 599d in the Modular NESC follow the same structure. According to embodiments, there may be some optional additions that improve the interplay between the modules 599a to 599d. For example, an optional connection between the modules 599a to 599d may be available as illustrated by the hatched line. This forms a so- called latent skip connection. Here, the NESC decoder 614a to 614d takes the quantized learned representation as an input. For Modular NESC, when generating with the n-th module, we can additionally use the learned representation from all the previous modules (n-1)-th, ... , 1st (dotted arrows in the plot).

[0074] According to further embodiments, the sub-band encoder 514a to 514d may be enhanced. The NESC encoder and its DPCRNN frontend take a time-domain signal as an input. For Modular NESC we can either take the input signal directly or the sub-bands components of it.

[0075] According to embodiments, adaptions for the final convolutional layer 616n may be possible. For example, additional convolutional layers are added to combine the outputs of different modules.

[0076] The PQMF synthesis entity 612 may be enhanced as follows: synthesis filters with increasing number of sub-bands are used to generate the final speech waveform. According to embodiments, all modules in Modular NESC are trained jointly end-2-end, but iterative training is also possible.

[0077] More widely used techniques of BWE. Encoder encodes the side-information that are representative of higher bands and transmit it to the decoder. Decoder utilizes the additional side-information and performs BWE. The bitrate used for side-information are very low compared to core-band bitrate. It has the potential to achieve better quality than Blind-BWE.

[0078] In recent times, DNN based solutions for BWE have been on the rise and are well exploited to produce state-of-the-art results for BWE. Such methods are mostly stand-alone solution [6][7] that can postprocess speech signal to increase their bandwidth. Our proposed system provides a well-integrated solution for BWE for neural speech codec and can also be integrated with conventional speech codecs. The solution is extension of Modular NESC and has been designed for both categories of BWE.

[0079] According to further embodiments, a so-called bandwidth extension is possible.

[0080] For this, parallel to the neural decoder and the neural encoder, the bandwidth extension encoder and / or the bandwidth extension decoder is available. Bandwidth Extension (BWE) has been used extensively in conventional speech codec. It is a method of extending the bandwidth of the speech signal beyond its original bandwidth. By the virtue of such extension, it can enhance the perceptual quality of decoded speech. BWE can be performed either with no information of higher bands or with small side information from the encoder, which is transmitted at very low bitrate. Depending upon the usage of side-information BWE can be broadly into two categories:

[0081] 1. Blind bandwidth extension

[0082] 2. Guided bandwidth extension

[0083] This means that such a bandwidth extension may be used in two different versions according to embodiments. A first bandwidth extension using guided bandwidth extension is illustrated by Fig. 3, wherein a second bandwidth extension using blind bandwidth extension is illustrated by Fig. 4. Fig. 3 shows a schematic block diagram of a proposed system for Guided-BWE. Fig. 3 shows a module 599a having the encoder 514a and the decoder 614a, the convolutional layer 616a and the PQMF synthesis 612. Parallel to the module 599a, the BWE module 399 is arranged. It comprises a DNN encoder 314, a DNN decoder 324 and a quantizer 115 in between. At the input of the encoder 314, a PQMF analysis 312 is arranged which outputs sub-band portions, e.g. for six bands. The synthesis of the DNN decoder 324 and the DNN decoder 614a is comparable to the embodiment as discussed in context of Fig. 2, namely using the convolutional layers 616a and 612. Although just one module 599a is illustrated, a plurality of modules may be used according to embodiments.

[0084] In other words, this means that the solution is similar to Modular NESC, but it used exclusively for BWE of NESC. The module 399 is iteratively trained such that a neural codec (NESC) is trained to generate a wide-band signal followed by training of the BWE module 399 to extend the bandwidth. The module 399 comprises of the following:

[0085] • An encoder 314, that takes interleaved higher sub-bands as input and output a learned representation.

[0086] • A quantizer 115, that quantizes learned representation at very low bitrate.

[0087] • A decoder 324, that generates the sub-band for BWE.

[0088] • A learnable conv layer 616a to ensure that the overlapping sub-bands are compensated, and no aliasing artefacts are injected in the filter-bank synthesis.

[0089] Finally, a synthesis filter bank 612 can combine the output to retrieve higher band signal. The solution can also be extended to any conventional wideband speech codec. For conventional codec, the lower sub-bands can be obtained from decoded speech by passing through a PQMF analysis filter 312 which is then combined with higher sub-bands generated by BWE module 399 in similar fashion.

[0090] Blind-BWE is done without any side-information and / or especially without side-information about the higher bands. No extra bit consumption. In conventional codec, it was mostly done through spectrum extrapolation and addition of sinusoids and pulses. The extended signal produced by such method are generally or at least sometimes limited in perceptual quality and can sometimes sound very unpleasant to the user. Fig. 4 illustrates the blind bandwidth extension. The BWE module 399 is just replaced by a DNN decoder 324’, wherein the module 599a, the convolutional layer 616a and the PQMF synthesis 612 is comparable to the embodiment discussed in context of Fig. 3. The Blind- BWE is similar to the guided bandwidth extension, but only uses the decoder network 324’ to generate sub-band signal. The decoder 324’ is conditioned with the quantized latent information from lower band codec or previous modules. It can also be conditioned with additional features obtained from the decoded signal of pretrained codec.

[0091] It should be noted that in above embodiments, features have been discussed in context of a module comprising the neural encoder and the neural decoder. However, aspects discussed in context of modules are of course applicable to the encoder side or the decoder side as well.

[0092] Further embodiments provide a (neural) encoder comprising:

[0093] • A sub-band decomposition,

[0094] • One or several neural encoders, involving learnable layers, for coding individual group of sub-bands,

[0095] • Where the different groups of sub-bands overlap in frequency.

[0096] Another embodiment provides a (neural) decoder (scalable in bitrate / complexity) comprising:

[0097] • One or several neural decoders, involving learnable layers, for decoding individual group of sub-bands (able to generate multi-rate signal)

[0098] • A base-band decoder

[0099] A filter-bank synthesis generating the output signal at different possible sampling rates / bandwidth, having as input the different decoded signals, processing version of them of the one or several neural decoders.

[0100] According to embodiments, the neural decoder comprises learnable (e.g. Conv.) layers to merge and compensate aliasing of the sub-band components of a synthesis filter-bank. Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some one or more of the most important method steps may be executed by such an apparatus.

[0101] The inventive encoded audio signal can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0102] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.

[0103] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0104] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.

[0105] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0106] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non- transitionary.

[0107] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.

[0108] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0109] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0110] A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver .

[0111] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0112] The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0113] References

[0114] [1] Pia, Nicola et al., “NESC: Robust Neural End-2-End Speech Coding with GANs”, https: / / arxiv.org / abs / 2207.03282

[0115] [2] Zeghidour, Neil et al. “Soundstream: An End-to-End Neural Audio Codec”, https: / / arxiv.orq / abs / 2107.03312

[0116] [3] Defossez Alexandre et al., “High Fidelity Neural Audio Compression”, https: / / arxiv.orq / abs / 2210.13438

[0117] [4] Ritesh Kumar et al., “High-Fidelity Audio Compression with Improved RVQGAN”

[0118] [5] F. Nagel and S. Disch, "A harmonic bandwidth extension method for audio codecs,"

[0119] A harmonic bandwidth extension method for audio codecs | IE iterance Publication |

[0120] [6] Seungu Han et al. “ Nil-Wave 2: A General Neural Audio Upsampling Model for Various

[0121] Sampling Rates"

[0122] Rates (arxiv.org’)

[0123] [7] J. Su, Y. Wang, A. Finkelstein and Z. Jin, "Bandwidth Extension is All You Need,

Claims

Claims1. Encoder (500) for encoding an audio signal (AS), comprising: a sub-band decomposition entity (510) configured to obtain individual groups of subbands (SBa, SBb, SBc) of the audio signal (AS) based on the audio signal (AS); and several neural encoders (514a, 514b, 514c), each involving learnable layers and configured to code individual groups of sub-bands (SBa, SBb, SBc) or a single full band signal.

2. Encoder (500) according to claim 1 , wherein the individual groups of sub-bands (SBa, SBb, SBc) have partially overlapping frequency range.

3. Encoder (500) according to one of the previous claims, wherein the sub-band decomposition entity (510) comprises an analysis filter, especially a PQMF analysis filter or a block transform or any learned or deterministic filterbank.

4. Encoder (500) according to one of the previous claims, wherein the several neural encoders (514a, 514b, 514c) are part of several modules (599a, 599b, 599c, 599d).5 Encoder (500) according to one of the previous claims, wherein the several neural encoders (514a, 514b, 514c) are part of distinct modules (599a, 599b, 599c, 599d).

6. Encoder (500) according to one of the previous claims, wherein the encoder comprises several quantizers for the several neural encoders (514a, 514b, 514c).

7. Encoder (500) according to one of the previous claims, wherein the several neural encoders (514a, 514b, 514c) are configured for varying bitrate and / or differ from each other with respect to its complexity.

8. Encoder (500) according to one of the previous claims, wherein the several neural encoders (514a, 514b, 514c) can be activated or disactivated for achieving varying bitrate and / or varying complexity. final9. Encoder (500) according to one of the previous claims, wherein each (distinct) module (599a, 599b, 599c, 599d) is trained separately or together with baseband module (599a, 599b, 599c, 599d).

10. Encoder (500) according to one of the previous claims, wherein at least one module (599a, 599b, 599c, 599d) being always active is configured to code a baseband audio signal (AS).

11. Encoder (500) according to one of the previous claims, wherein at least one module (599a, 599b, 599c, 599d) being configured to be adaptively activated is configured to extend the baseband audio signal (AS) encoder (500).

12. Encoder (500) according to one of the previous claims, wherein at least one module (599a, 599b, 599c, 599d) is configured for guided bandwidth extension.

13. A decoder (600) for decoding an audio signal (AS), comprising: several neural decoders (614a, 614b, 614c), each involving learnable layers and configured for decoding individual groups of sub-bands (SBa, SBb, SBc) of the audio signal (AS); a filter-bank synthesis (612) configured for generating the audio signal (AS) based on the individual groups of sub-bands (SBa, SBb, SBc).

14. Decoder (600) according to claim 13, wherein the individual groups of sub-bands (SBa, SBb, SBc) have partially overlapping frequency range.

15. Decoder (600) according to claim 13 or 14, wherein the filter-bank synthesis (612) is configured to generate the audio signal (AS) at different possible sampling rates and / or audio bandwidths.

16. Decoder (600) according to one of claims 13, 14 or 15, wherein the decoder (600) further comprises a base band decoder (600). final17. Decoder (600) according to one of claims 13 to 16, wherein the decoder (600) comprises a learned conventional layer (616a, 616b) configured to combine at least two sub-bands (SBa, SBb, SBcs).

18. Decoder (600) according to one of claims 13 to 17, wherein the filter-bank synthesis (612) comprises a PQMF synthesis filter or an inverse block transform or any learned or deterministic filterbank.

19. Decoder (600) according to one of claims 13 to 18, wherein the several neural decoders (614a, 614b, 614c) form parts of several or distinct modules (599a, 599b, 599c, 599d).

20. Decoder (600) according to one of claims 13 to 19, wherein at least one module (599a, 599b, 599c, 599d) is configured for extended the bandwidth of the decoded signal obtained by the base-band decoder (600).

21. Decoder (600) according to one of claims 13 to 20, wherein at least one module (599a, 599b, 599c, 599d) is configured for guided bandwidth extension or blind bandwidth extension.

22. System comprising an encoder (500) according to one of claims 1 to 12 and a decoder (600) according to one of claims 13 to 21 , wherein one of the several neural encoders (514a, 514b, 514c) forms together with one of the several neural decoders (614a, 614b, 614c) a module (599a, 599b, 599c, 599d).

23. System according to claim 22, wherein each module (599a, 599b, 599c, 599d) is trained separately or together.

24. An encoder (500) for encoding an audio signal (AS), comprising a sub-band decomposition entity (510) configured for obtaining individual groups of sub-bands (SBa, SBb, SBc) of the audio signal (AS) based on the audio signal (AS) using sub-band decomposition; finalneural encoders (514a, 514b, 514c), each involving learnable layers and configured to code individual groups of sub-bands (SBa, SBb, SBc) of the audio signal (AS), where the different groups of sub-band have partially overlapping frequency ranges; and a BWE encoder (500) configured to perform guided bandwidth extension.

25. A decoder (600) for decoding an audio signal (AS), comprising: neural decoders (600), each involving learnable layers and configured to decode individual groups of sub-bands (SBa, SBb, SBc) of the audio signals (AS); a filter-bank synthesis (612) configured for generating the audio signal (AS) based on the individual groups of sub-bands (SBa, SBb, SBc); and a BWE decoder (600) configured to perform blind or guided bandwidth extension.

26. Method for encoding an audio signal (AS), comprising: obtaining individual groups of sub-bands (SBa, SBb, SBc) of the audio signal (AS) based on the audio signal (AS) using sub-band decomposition; and coding individual groups of sub-bands (SBa, SBb, SBc), or a single full band signal using several neural encoders (514a, 514b, 514c), each involving learnable layers.

27. Method for decoding an audio signal (AS), comprising: decoding individual groups of sub-bands (SBa, SBb, SBc) of the audio signal (AS) using several neural decoders (614a, 614b, 614c), each involving learnable layers; generating the audio signal (AS) based on the individual groups of sub-bands (SBa, SBb, SBc) using filter-bank synthesis (612).

28. Digital storage medium having stored there on a computer program code for performing, when running on a computer, the steps of claim 26 or 27. final

Citation Information

Patent Citations

  • Device and method for encoding / decoding audio signal using filter bank

    US20210166701A1