Scalar quantization for audio coding

By employing scalar quantization and a convolutional encoder-decoder approach to learn a low-dimensional representation of audio signals, this method addresses the high computational complexity and training instability issues of existing neural audio coding methods, achieving efficient audio signal reconstruction and scalable data rates.

CN121605474APending Publication Date: 2026-03-03FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380100903.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing neural audio coding methods suffer from problems such as a large number of parameters, high computational complexity, unintuitive calculation of large-dimensional vector distances, difficulty in balancing data rate and quality, unstable training, high computational cost, and complex codebook selection when learning useful discrete representations of latent signals.

Method used

The method employs scalar quantization, which uses a convolutional encoder and decoder to learn the low-dimensional representation of the audio signal. It then uses a scalar quantization module and learnable segments for quantization and dequantization, and combines entropy coding technology to further compress the index, thereby achieving end-to-end training and scalability.

Benefits of technology

It achieves high-quality audio signal reconstruction, reduces computational complexity and training costs, improves training stability and data rate scalability, and simplifies the codebook selection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

Techniques for encoding and decoding audio signals are described. A decoder (10) configured to generate an audio signal (16) from an encoded signal (3) representing the audio signal may comprise: an encoded signal reader (560) configured to read the encoded signal (3), thereby providing a plurality of indices (556); a scalar dequantization module (500) comprising: a plurality of quantization index converters (355), each quantization index converter (555) configured to convert an index (556) of the plurality of indexes to a corresponding potential scalar value (551) such that the plurality of potential scalar values (551) form a first potential audio signal representation (550) of the audio signal; and a first learnable section (540) for providing a second potential representation (530) from the first potential audio signal representation (550); a second learnable section (520) comprising at least one learnable layer and configured to generate an audio signal (16) from a second potential audio signal representation (530).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Encoders and decoders are disclosed herein. For example, vocoders and related methods are disclosed.

[0002] Learning an intermediate discrete representation of an audio signal for efficient signal transmission in communication applications is central to any neural audio encoder (NAC). This application proposes an efficient model for such discrete representations, also known as a scalar quantizer (SQ), and an associated training technique that allows a tradeoff between quality and the desired data transmission rate. The quantization method maps features from the NAC encoder output to a set of representation values, resulting in a discrete representation of the input signal. The model may include (e.g., convolutional) encoder-decoder pairs that learn a low-dimensional representation of the NAC output, which is quantized channel-by-channel and may be transformed for subsequent decoding by the NAC decoder. The SQ can be trained end-to-end with the NAC by approximating the non-differentiable quantizer with an identity and associated MSE loss or by simulating the quantization process by adding uniformly distributed noise. By adjusting the encoding hierarchy and latent dimensions using transfer learning from a previously trained NAC, rate scalability is allowed without (expensively) retraining the NAC and storing the resulting weights for each target data rate.

[0003] The proposed method demonstrates advantages in interpretability, computational efficiency, and scalability of data rate for discrete representations, without exhibiting serious drawbacks compared to competing traditional methods. Background Technology

[0004] Due to the excellent quality of the reconstructed audio signal achievable at low required data rates, NACs have attracted significant research interest from industry [1,2] (Soundstream, Encodec) and academia [4]. One component of these NACs is a learned, data-efficient, and discrete representation of the input signal, which forms the basis of the transmitted signal. Most NACs consist of a convolutional encoder, which provides a tight signal representation, also known as the latent signal, to the quantizer module. From this signal representation, the transmitted signal is formed, which is then reconstructed by the NAC decoder at the receiver.

[0005] Most traditional NACs utilize a variant of the vector quantizer (VQ) [5-7] to learn the aforementioned discrete intermediate representation. Here, a set of template vectors is appropriately learned or selected such that it represents the latent as accurately as possible. For each incoming signal frame, the most accurate codebook vector is selected as the replacement for the frame-by-frame latent vector, and the corresponding codebook vector index is transmitted. At the receiver, the NAC decoder reconstructs the input signal based on the representation provided by the selected codebook vector.

[0006] A drawback of training such VQs via backpropagation is their non-differentiability. Therefore, many methods have been proposed to address this issue, including:

[0007] ● Skip the quantizer in the reverse path and enforce the codebook structure in the latent by additional loss (VQ-VAE) [6].

[0008] ● The quantizer is approximated by a smooth substitution (Softmax)[7].

[0009] ●Sampling is performed from the continuous relaxation of the discrete distribution (Gumbel Softmax)[8].

[0010] ● Soft to hard scheduling, which causes smooth substitutions to oriented towards hard quantization during training[7].

[0011] Although scalar quantization methods are only poorly understood in neural audio coding (e.g., [3]), they are quite popular in neuroimaging and audio coding [8]. In this application, we propose a method for learning discrete audio representations based on scalar quantization.

[0012] Deficiencies of traditional methods

[0013] There are drawbacks associated with the traditional methods mentioned:

[0014] 1. In order to learn a useful discrete representation of the latent signal, the codebook vector of VQ must be chosen to be of high dimension. This essentially increases the number of parameters required and the computational complexity of the resulting NAC.

[0015] 2. Discrete representations learned by VQ are often difficult to interpret and exhibit unintuitive behavior due to the computation of distances between large-dimensional vectors.

[0016] 3. It is difficult to balance data rate and quality without retraining NAC and storing different weights for each trained model.

[0017] 4. Skipping the quantizer in backpropagation (VQ-VAE) often leads to unintuitive results. Training the codebook requires additional mechanisms (e.g., training via recursive averaging), which necessitates additional parameterization, and if chosen incorrectly, this can lead to unstable training.

[0018] 5. Several VQ modules must be trained together, i.e., residual VQ, to obtain convincing results.

[0019] 6. The quality of NAC obtained using VQ does not scale well with the data rate used.

[0020] 7. Training the VQ codebook is not mandatory, but it is crucial for achieving an acceptable convergence rate for NAC training.

[0021] 8. Choosing the smoothness level for the softmax approximation of a quantizer is difficult because high smoothness provides well-performing gradients but provides poor approximations of hard quantization, and vice versa. The scheduling of smoothness selection complicates this further.

[0022] 9. Decisions about the best-fit codebook vector require calculating the distance between the high-dimensional latent vector and all codebook vectors, which can be computationally expensive.

[0023] References

[0024] [1] Zeghidour, Neil, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. „SoundStream: An End-to-End Neural Audio Codec“. arXiv,7. Juli 2021. http: / / arxiv.org / abs / 2107.03312.

[0025] [2] Défossez, Alexandre, Jade Copet, Gabriel Synnaeve, und Yossi Adi. „High Fidelity Neural Audio Compression“. arXiv, 24. Oktober 2022. http: / / arxiv.org / abs / 2210.13438.

[0026] [3] Zhen, Kai, Jongmo Sung, Mi Suk Lee, Seungkwon Beack, und MinjeKim. „Scalable and Efficient Neural Speech Coding: A Hybrid Design“. IEEE / ACMTransactions on Audio, Speech, and Language Processing 30 (2022): 12-25.

[0027] [4] Jiang, Xue, Xiulian Peng, Huaying Xue, Yuan Zhang, und Yan Lu. „Cross-Scale Vector Quantization for Scalable Neural Speech Coding“. arXiv, 6.July 2022. http: / / arxiv.org / abs / 2207.03067.

[0028] [5] Pia, Nicola, Kishan Gupta, Srikanth Korse, Markus Multrus, undGuillaume Fuchs. „NESC: Robust Neural End-2-End Speech Coding with GANs“.arXiv, 7. July 2022. http: / / arxiv.org / abs / 2207.03282.

[0029] [6] Oord, Aaron van den, Oriol Vinyals, and Koray Kavukcuoglu. „Neural Discrete Representation Learning“. arXiv, 30. May 2018. http: / / arxiv.org / abs / 1711.00937.

[0030] [7] Agustsson, Eirikur, Fabian Mentzer, Michael Tschannen, LukasCavigelli, Radu Timofte, Luca Benini, und Luc Van Gool. „Soft-to-Hard VectorQuantization for End-to-End Learning Compressible Representations“. arXiv, 8.June 2017. http: / / arxiv.org / abs / 1704.00648.

[0031] [8] Jang, Eric, Shixiang Gu, und Ben Poole. „CategoricalReparameterization with Gumbel-Softmax“. arXiv, 5. August 2017. http: / / arxiv.org / abs / 1611.01144.

[0032] [9] Balle, Johannes, Philip A. Chou, David Minnen, Saurabh Singh,Nick Johnston, Eirikur Agustsson, Sung Jin Hwang, und George Toderici. „Nonlinear Transform Coding“. IEEE Journal of Selected Topics in SignalProcessing 15, Nr. 2 (Februar 2021): 339-53.

[0033]

[10] Agustsson, Eirikur, Fabian Mentzer, Michael Tschannen, LukasCavigelli, Radu Timofte, Luca Benini, und Luc Van Gool. „Soft-to-Hard VectorQuantization for End-to-End Learning Compressible Representations“. arXiv, 8.Juni 2017. http: / / arxiv.org / abs / 1704.00648.

[0034]

[12] Oord, Aaron van den, Oriol Vinyals, und Koray Kavukcuoglu. „Neural Discrete Representation Learning“. arXiv, 30. Mai 2018. http: / / arxiv.org / abs / 1711.00937.

[0035]

[13] Zhen, Kai, Jongmo Sung, Mi Suk Lee, Seungkwon Beack, und MinjeKim. „Scalable and Efficient Neural Speech Coding: A Hybrid Design“. IEEE / ACMTransactions on Audio, Speech, and Language Processing 30 (2022): 12-25.https: / / doi.org / 10.1109 / TASLP.2021.3129353.

[0036]

[14] Zeghidour, Neil, Alejandro Luebs, Ahmed Omran, Jan Skoglund, undMarco Tagliasacchi. „SoundStream: An End-to-End Neural Audio Codec“. arXiv,7. Juli 2021. http: / / arxiv.org / abs / 2107.03312.

[0037]

[15] M. H. Vali and T. Bäckström, "NSVQ: Noise Substitution in VectorQuantization for Machine Learning," in IEEE Access, vol. 10, pp. 13598-13610,2022, doi: 10.1109 / ACCESS.2022.3147670.

[0038]

[16] J. Ballé, V. Laparra, and E. P. Simoncelli, “End-to-endoptimized image compression,” in Proc. 5th Int. Conf. Learn. Represent.,2017, pp. 1-27

[0039]

[17] Défossez, Alexandre, Jade Copet, Gabriel Synnaeve, und YossiAdi. „High Fidelity Neural Audio Compression“. arXiv, 24. Oktober 2022. http: / / arxiv.org / abs / 2210.13438. Summary of the Invention

[0040] According to one aspect, a decoder is provided, configured to generate an audio signal from an encoded signal representing an audio signal, the decoder comprising:

[0041] The encoded signal reader is configured to read encoded signals, thereby providing multiple indices;

[0042] The scalar solution quantization module includes:

[0043] Multiple quantization index converters, each configured to convert an index among multiple indices into a corresponding latent scalar value, such that the multiple latent scalar values ​​form a first latent audio signal representation of the audio signal; and

[0044] A first learnable segment is used to provide a second potential representation from a first potential audio signal representation;

[0045] The second learnable segment includes at least one learnable layer and is configured to generate an audio signal from a second potential audio signal representation.

[0046] According to one aspect, an encoder is provided for generating an encoded signal, wherein an input audio signal is encoded in the encoded signal, and the encoder includes:

[0047] The first learnable segment includes at least one learnable layer to provide a first potential representation of the input audio signal.

[0048] The scalar quantization module, used to quantize the first latent representation, includes:

[0049] A second learnable segment is used to provide multiple latent scalar values ​​to be quantized from a first latent representation; and

[0050] Multiple quantizers are used to provide multiple indices, each quantizer being configured to quantize a single latent scalar value to be quantized and providing an index from multiple indices from a single latent scalar value; and

[0051] The encoded signal writer is configured to write multiple indices into the encoded signal.

[0052] According to one aspect, a decoding method is provided for generating an audio signal from an encoded signal representing an audio signal, the method comprising:

[0053] Read the encoded signal to obtain multiple indices;

[0054] Performing scalar solution quantization includes:

[0055] The conversion is performed through multiple quantization index converters, each converting an index among multiple indices into a corresponding latent scalar value, such that the multiple latent scalar values ​​form a first latent audio signal representation of the audio signal; and

[0056] A second potential audio signal representation is provided from a first potential audio signal representation via a first learnable segment; and

[0057] An audio signal is generated from a second potential audio signal representation by means of a second learnable segment comprising at least one learnable layer.

[0058] According to one aspect, a method for generating an encoded signal is provided, wherein an input audio signal is encoded in the encoded signal, the method comprising:

[0059] A first potential representation of the input audio signal is provided through a first learnable segment comprising at least one learnable layer.

[0060] The first latent representation is quantized using the scalar quantization module through the following operations:

[0061] Multiple latent scalar values ​​to be quantized are provided from the first latent representation through the second learnable segment; and

[0062] Multiple indices are obtained through multiple quantizers, each of which quantizes a single latent scalar value and provides an index among multiple indices from a single latent scalar value; and

[0063] Multiple indices are written into the encoded signal.

[0064] According to one aspect, a non-transitory storage unit is provided, which stores instructions that, when executed by a computer, cause the computer to execute and / or control the execution of the above methods. Attached Figure Description

[0065] Figure 1a , Figure 1b and Figure 1c This shows an example of an encoder based on the example.

[0066] Figure 2a , Figure 2b , Figure 2c , Figure 2d and Figure 2eThis section demonstrates examples of encoder operations and mode selection.

[0067] Figure 3a , Figure 3b and Figure 3c This example demonstrates the decoder based on the example.

[0068] Figure 4 This demonstrates an example of the operations and mode selection at the decoder.

[0069] Figure 5a , Figure 5b and Figure 5c This demonstrates the technology used to control mode selection at the encoder.

[0070] Figure 6a and Figure 6b This demonstrates the technology used to control mode selection at the decoder.

[0071] Figure 7a Show the residual quantizer at the encoder.

[0072] Figure 7b Show the residual quantizer at the decoder.

[0073] Figure 8 shows an example based on existing technology.

[0074] Figure 9 Examples of the technology according to the present invention are shown.

[0075] Figure 10 and Figure 11 Examples of encoders and decoders (e.g., Figures 2a to 3c (A more detailed version of the optional features of the encoder and decoder).

[0076] Figures 12 to 15 exhibit Figures 2a to 3c A detailed example of optional features for the decoder. Detailed Implementation

[0077] Figure 1a Encoder 2 is shown (some of its instances 2b and 2c are shown in...) Figure 1b , Figure 1cExample (in the text). Encoder 2 (2b, 2c) can generate an encoded signal 3 (e.g., a bitstream or a portion thereof) to encode the input audio signal 1 in the encoded signal 3. The input audio signal 1 can be in the frequency domain or the time domain (a time / frequency converter can be provided at the input of encoder 2 or within encoder 2). The input audio signal 1 can be subdivided, for example, into a series of frames, which may be different from or overlap each other. The encoded signal 3 generated by encoder 2 can be a bitstream (or a portion thereof). The input audio signal 1 can be a mono signal. In the case of encoding spatial multichannel signals (e.g., stereo signals), in some examples, multiple input audio signals 1 can be encoded in parallel and independently using the processing described herein, thereby generating multiple encoded signals 3, for example, written into the same bitstream. Alternatively, multiple spatial channels can be linearly combined before being fed to multiple instances of encoder 2, such as the middle spatial channel and side spatial channels in the case of stereo. The encoded signal 3 may be transmitted, for example, via a transmission device such as a wired or wireless communication device (e.g., in a communication network) to a decoder (e.g., via a client / server connection or a peer-to-peer connection) and / or stored in a storage unit, for example, for subsequent reading by a decoder (e.g., decoder 10, see below).

[0078] Encoder 2 may include a first learnable segment 20. The first learnable segment 20 may include at least one learnable layer (e.g., a neural network having, for example, convolutional layers and / or recurrent units and / or fully connected layers) to provide a first latent representation 330 (also indicated as 469 in some examples) of the input audio signal 1. In some examples, the first latent representation 330 (469) may be represented as a matrix (e.g., an M×N matrix), where M>1 and N≥1 (in the case of a vector, i.e., an M×1 matrix, M can be understood as the number of latent channels). In some examples, the first latent representation 330 (469) may be represented as a vector (having M entries, each a latent scalar value or a latent channel, e.g., where M>1). However, the number of rows M in the matrix may be less than the original number of samples per frame. However, the reduction in the number of rows relative to the number of samples can be compensated by a number of columns N greater than 1. For example, M may be the number of latent channels, and N may be the length of the frame. It should be noted that each frame can be subdivided into multiple vectors, and each vector can be subdivided into multiple latent channels (in some examples, the latent channel dimension may correspond to the column dimension), each latent channel having a single latent scalar value to be encoded. In some examples, the first learnable segment 20 may produce a first latent representation 330 that is independent of the bit rate (e.g., for any bit rate of the encoded signal 3, the first latent representation may remain the same, where the resolution is the same and the number of latent channels M is the same for each frame), and / or independent of the resolution to be given to the encoded signal 3, and / or independent of other executable options, and / or independent of the input audio signal itself.

[0079] The encoder 2 may include a scalar quantization (SQ) module 300 that receives a first latent representation 330 (469). The scalar quantization module 300 may have the task of quantizing the first latent representation 330, for example, per latent channel (per latent scalar value).

[0080] Scalar quantization module 300 may include a second learnable segment 340. The second learnable segment 340 may provide a second latent representation 350 from a first latent representation 330. The second latent representation 350 may include a plurality of latent scalar values ​​351 (e.g., each scalar value for each latent channel, i.e., each scalar value for latent representation 330). The second learnable segment 340 may include at least one learnable layer. The output of the second learnable segment 340 may be a plurality of latent scalar values ​​351. In some examples, the number of latent channels (latest scalar values) per frame may vary (e.g., decrease) or more generally vary (e.g., selectively decrease), for example, depending on a selection applied by controller 358 (see below).

[0081] (It should be noted that in some examples, the first learnable segment 20 can be set for all bit rates, while the second learnable segment 340 can be closely associated with a given set / codebook of the scalar quantizer. In other words, the first latent representation 330 can be a general representation of the input audio signal 1, while the second latent representation 350 can be a specific latent representation for a given set of scalar quantizers 355, for example, for a specific bit rate.)

[0082] The scalar quantization module 300 includes multiple quantizers 355. Each quantizer 355 can provide a single index 356 for a corresponding potential scalar value 351 (e.g., for a corresponding channel). (In a multi-level example, e.g., in...) Figure 7a The diagram will demonstrate that multiple indices 356R1 and 356R2 can exist from a single scalar value 351. Therefore, multiple quantizers 355 can complexly provide multiple indices for each frame of the input audio signal 1 (and for potential representations 330 and 350). Each index 356 can have a length of several bits, with a bit count that is, for example, between 3 and 8, or more specifically, between 3 and 6, or even more specifically, between 2 and 5.

[0083] Each quantizer 355 implements a mapping from (e.g., approximate) real-valued representation 351 to discrete-valued representation 356 derived from a set of finite values. The mapping is applied channel-by-channel and can differ per channel: for example, different numbers of codebooks / quantization levels may exist for different quantizers 355. The number and position of the "quantization levels" are parameters of each scalar quantizer 355, and the quantizer and quantization levels are distinct objects (the latter being the building block of the first).

[0084] There are multiple quantizers 355 because there are multiple potential channels (potential scalar values) 351 (e.g., for each frame, there are multiple channels, i.e., multiple scalar values 351 representing 330 or 350), and each quantizer 355 converts a single specific scalar value (in a specific potential channel) into a single specific index.

[0085] There may be a fixed relationship between a specific quantizer 355 to be used and a specific position in the vector of the latent representation of the audio signal 1. Thus, each specific quantizer 355 can be applied to a specific potential channel (potential scalar value) according to a specific position in the latent representation (330 or 350) representing the audio signal 1.

[0086] Since the index 356 is obtained from the potential scalar values 351 of the second latent representation 350, the set of indices formed by the potential scalar values 351 of each channel represents the quantized latent version of the audio signal 1 (e.g., for one frame).

[0087] See Figure 2a and Figure 4 , the selections 353 and 553 can be performed at the encoder and decoder respectively so as to select one coding mode in at least one first coding mode 341 and one second coding mode 342 and select one decoding mode in at least one first decoding mode 541 and one second decoding mode 542 respectively.

[0088] In some examples, the second learnable section 340 can vary the number of potential channels (potential scalar values) between its input 330 and its output 350: the second learnable section 340 can change (e.g., reduce) the number of potential channels (potential scalar values) for each frame such that the second latent representation 350 can have a different number of potential channels (potential scalar values) relative to the number of potential channels (potential scalar values) of the first latent representation 330. By changing the number of channels (potential scalar values) in the latent representation (from 330 to 350), the number of quantizers 355 also changes and the number of indices 356 also changes accordingly. In some examples, this can be a selection of a mode (see also below). Generally, the second learnable section 340 can convert a first number N1 of potential channels (potential scalar values) of the first latent representation 330 (469) into a second number N2 of potential channels (potential scalar values) 351 of the second latent representation 330 (where typically N2 < N1) (e.g., for each frame). Examples can be N1 = 16 and N2 = 8 (or N1 = 64 and N2 = 16 or N2 = 32; other values are possible).

[0089] In some examples, the number of potential channels (potential scalar values) per frame can vary (e.g., optionally), such as based on selection (e.g., user selection or selection controlled by automatic components), such as adaptively (e.g., to adapt to a particular audio signal 1, specifically, to adapt to a particular frame or frame sequence of audio signal 1).

[0090] Encoder 2 may include an encoded signal writer 360 that writes (e.g., via encapsulation) multiple indices 356 to the encoded signal 3. Even though the encoded signal writer 360 is shown as part of the scalar quantization module 300 in the figures, it may also be external to the scalar quantization module 300. However, for simplicity, the encoded signal writer 360 is shown as being inside the scalar quantization module 300 in the figures, but in any of the following examples, it may be external.

[0091] The encoded signal writer 360 may also include additional encoding tools, such as an entropy encoder, which is designed to further compress the quantization index without loss by using variable-length codes that depend on the estimated and / or pre-calculated occurrence probabilities of different quantization indices. For example, entropy coding may use at least one of Huffman codes, arithmetic range codes, and Golomb-Rice codes.

[0092] See Figure 3a This section presents a decoder (audio generator) 10 (e.g., capable of decoding the encoded signal 3 generated by encoder 2, such as 2b, 2c). Decoder 10 (which can, for example, be transmitted via...) Figure 3b Decoder 10b or Figure 3c A decoder 10c (instantiated) can generate an output audio signal 16, which is intended to be a potentially reliable copy or high-fidelity approximation of the input audio signal 1. For example, the output audio signal 16 can be rendered, for example, by a speaker downstream of or included in the decoder 10. Alternatively, the decoder 10 (e.g., 10b, 10c) can encode the generated audio signal 16 into another encoded signal representation and can therefore operate as a codec.

[0093] Decoder 10 (e.g., 10b, 10c) may include an encoded signal reader 560 capable of reading the encoded signal 3. The encoded signal reader 560 may output multiple indices 556, which may be the same as the indices 356 output by the quantizer 355 of encoder 2.

[0094] The encoded signal reader 560 may also include additional inverse encoding tools, such as an entropy decoder, designed to decode the entropy-encoded quantization index. For example, the entropy decoder may support at least one of Huffman codes, arithmetic range codes, and / or Columbus-Rice codes.

[0095] Decoder 10 (e.g., 10b, 10c) may include scalar dequantization module 500, which provides a second potential representation 530 (also referred to as code 112) of the audio signal 16 to be generated. Even though the encoded signal reader 560 is shown as part of the scalar dequantization module 500 in the figures, the encoded signal reader 560 may also be external to the scalar dequantization module 500. Here, reference numeral 513 is provided to indicate a scalar dequantization module 500 without the encoded signal reader 560.

[0096] Dequantization module 500 (513) may include multiple quantization index converters (dequantizers) 555, each configured to convert a single index 556 (or, in residual techniques, such as in...) Figure 7b In this process, more than one index (556R1, 556R2) is converted into a single latent scalar value 551. Therefore, in some examples, each frame of the input audio signal 1 (and the output audio signal 16 to be generated) can be mapped by multiple latent scalar values ​​551. The latent scalar values ​​551 can form a first latent representation 550 of the audio signal 16 to be generated. The latent scalar value 551 can be considered as corresponding to the latent scalar value 351 at encoder 2, and the first latent representation 550 of the audio signal 16 to be generated can be considered as corresponding to a second latent representation 350 of the input audio signal 1 at encoder 2. In some examples, the number and / or configuration of the quantization index converters (dequantizers, dequantization levels) 555 depend on the number N2 of indices 356 written by encoder 2 (e.g., 2b, 2c): for example, if the encoder has encoded a frame with 16 indices (and therefore 16 quantizers 355), decoder 10 (e.g., 10b, 10c) will therefore use 16 quantization index converters 555, and if the encoder has encoded a frame with 8 indices (and therefore 8 quantizers 355), decoder 10 (e.g., 10b, 10c) will therefore use 8 quantization index converters 555, and so on. Therefore, in some examples, the number of converters 555 for each frame can be varied, for example, depending on the number of indices 556 written in the encoded signal 3 for each frame (e.g., optionally, by selection, and specifically, adaptively, for example, based on a particular audio signal encoded in the encoded signal 3, for example, based on signaling written in the encoded signal 3 as side information).

[0097] There are multiple quantization index converters 555 because there are multiple potential scalar values 551 to be generated (e.g., for each frame, there are multiple scalar values 551), and each quantization index converter 555 converts each index 556 into a specific scalar value of a matrix (e.g., a vector). More generally, there may be a fixed relationship between a particular quantization index converter 555 to be used and a particular position in the vector of the potential representation of the audio signal 16 (and the input audio signal 1) to be generated. Information about the relationship may be signaled in the encoded signal 3 or may otherwise be obtained from the encoded signal 3 (e.g., it may be assumed from a particular position of the index in the encoded signal 3).

[0098] The dequantization module 500 (513) may include a first learnable section (scalar dequantization learnable section) 540 that receives the potential scalar values 551 of the first potential representation 550 and generates a second potential representation 530 (e.g., in the form of code 112, e.g., two-dimensional). The first learnable section (scalar dequantization learnable section) 540 may be regarded as corresponding to the second learnable section 340 of the encoder 2, and the second potential representation 530 may be regarded as corresponding to the first potential representation 330 (469) at the encoder 2. However, it should be noted that even if the second learnable section 340 on the encoder side has (e.g., optionally, e.g., adaptively) changed (e.g., reduced) the number of scalar values from the first number N1 of potential channels (potential scalar values) in the first potential representation 330 on the encoder side to the second number N2 of potential channels (potential scalar values) 351 in the second potential representation 350 on the decoder side, the number of potential channels 551 in the first potential representation 550 on the decoder side may also remain N2, but the number N3 of potential channels in the second potential representation 530 on the decoder side can generally be independent of N1 (N3 > N1, N3 < N1, or N3 = N1 may make no difference) (N3 > N2 may be advantageous to give the output audio signal 16 good resolution while the encoded signal 3 preferably has a small number of indexes). Thus, the second potential representation 530 on the decoder side can convert the number of potential channels (potential scalar values) from the number N2 of indexes 556 obtained from the encoded signal 3 to the number N3 that is required by the second learnable section 520 and is generally independent of the encoded signal 3 and / or the bit rate and / or other selections. Thus, the second learnable section 340 can convert the number of potential channels (potential scalar values) from the number N1 of the first potential representation 330 (which is on the application side and can generally be independent of the bit rate and / or selections) to the number N2 (usually N2 ≤ N1) and can adapt to conditions such as the target bit rate, selections, etc.

[0099] Decoder 10 (e.g., 10b, 10c) may include a second learnable segment (e.g., a Neural Audio Coding (NAC) decoder) 520. The second learnable segment 520 may output an audio signal 16. The second learnable segment 520 can be considered as corresponding to the first learnable segment 20 of encoder 2. However, it should be noted that the decoder-side second learnable segment 520 does not necessarily need to be a mirror image of the encoder-side first learnable segment 20. Essentially, it is not strictly required that the decoder mirror the operation of the encoder.

[0100] Generally, the operation of the second learnable segment 520 of the decoder 10 (e.g., 10a, 10b) can be considered as independent of at least one of the bit rate of the encoded signal 3 and / or independent of the characteristics and / or selection of the encoded signal 3.

[0101] The quantizer 355 in encoder 2 and the quantization-index converter (dequantizer) 555 in decoder 10 can utilize technologies not shown in [the original text]. Figure 1a and Figure 3a At least one codebook in the encoder 2 and / or decoder 10. At least one codebook may be learnable or deterministic. Some examples are described here.

[0102] At least one (or each) codebook can perform the association between scalar values ​​and indices, for example by mapping each scalar value 351 to a specific index 356 and vice versa at the encoder, and by mapping index 556 to a specific scalar value 551 at the decoder. In some cases, at least one codebook may have a fixed length (i.e., the codebook has a variable-length bitstream representation), meaning that all indices 356 and 556 have the same length (e.g., all indices have 4 bits). In other cases, the codebook may have a variable bit length (i.e., the codebook has a fixed-length bitstream representation) such that different indices 356 and 556 can have different lengths (e.g., more frequent scalar values ​​can be mapped to more compact indices, such as those with smaller extended lengths, while less frequent indices or scalar values ​​can be mapped to less compact indices, such as those with larger extended lengths; this can be effective, for example, for at least two indices, or for multiple indices, or for most indices, or for all indices). At least one (or each) codebook may have variable precision, meaning that some indices approximate scalar values ​​better (e.g., with less uncertainty) than others: for example, a range with more frequent scalar values ​​may be mapped to several indices, the number of which is greater than the number of indices to which a range with less frequent scalar values ​​is mapped, for example, with regard to range extension: thus, for scalar values ​​in highly frequent ranges, the approximate uncertainty decreases, thereby improving precision, while less frequent ranges of scalar values ​​have fewer indices, each with greater uncertainty. In summary, the codebook or quantization may be non-uniform, where the range of values ​​to be quantized is divided into unequal intervals such that more frequent intervals are smaller than less frequent intervals. This subdivision may differ between different codebooks and may be constrained during training.

[0103] At the quantizer 355 of encoder 2, at least one codebook may permit the conversion of a single latent scalar value 351 into a single index 356. At each of the quantization-index converters 555 of decoder 10, at least one codebook may permit the conversion of a single index 556 into a single latent scalar value 551.

[0104] In some cases, at least one quantizer 355 (e.g., all quantizers) or at least one quantization index converter 555 (e.g., all quantization index converters) may have multiple levels, such as Figure 7a and Figure 7b As shown in the example. In some examples, each level may be associated with a specific codebook, or all levels may share the same codebook.

[0105] Figure 1b An example of encoder 2b (which may be an instance of encoder 2) is shown, in which a single codebook 357 is shared by multiple quantizers 355 in encoder 2b. Therefore, in Figure 2bIn the example, the equivalent latent scalar value 351 quantized by different quantizers 355 will be mapped by the same index 356 using different quantizers with the same codebook 357. For example... Figure 3b As shown in decoder 10b (which may be an instance of decoder 10), a single codebook 557 can be shared by multiple quantization index converters 555: equivalent indices 556, once converted by different quantization index converters 555, will be mapped to the same underlying scalar value 551. Using a single codebook (e.g., 357, 557) for all quantizers 355 (each a quantization index converter 555) has several advantages, as storing a single codebook requires less storage space and less computational effort during training.

[0106] Figure 1c An example of encoder 2c (which may be an instance of encoder 2) is shown, wherein at least one quantizer 355a (e.g., all quantizers) uses a quantizer-specific codebook 357a. Although Figure 1c A shared codebook 357b is also shown, shared by multiple quantizers 355 in encoder 2c (e.g., proper subsets of multiple quantizers) (in the example of the graph, instantiated by quantizers 355b and 355c), but this is not required: each of the quantizers may have a quantizer-specific codebook. Since each latent scalar value (quantized by the corresponding quantizer to provide the corresponding index) may have a specific relationship to a specific location in the latent representation (e.g., a matrix, such as a vector), each quantizer-specific codebook may also have a specific relationship to a specific location in the latent representation. For example, different codebooks may exist for different locations. Figure 3cAn example of a decoder 10c (which may be an instance of decoder 10) is shown, which can apply at least one quantizer-specific codebook: at least one quantization index converter 555a (e.g., all quantization index converters) can use one quantization index converter-specific codebook 557a, while the remaining quantization index converters 555b, 555c can use at least one shared codebook 557b. Since each latent scalar value (to be generated from a corresponding index by the corresponding quantization index converter) can have a specific relationship with a specific position in the latent representation (e.g., a matrix, such as a vector), each quantization index converter-specific codebook can also have a specific relationship with a specific position in the latent representation. For example, different codebooks can exist for different positions. Multiple quantizer-specific codebooks (e.g., 357a) can be used. (Each quantizer-specific codebook, for each quantizer-index converter, has several advantages, including improved accuracy: probabilistically, some intervals of scalar values ​​may be more frequent in the first position of the latent representation, while other intervals of scalar values ​​may be more frequent in the second position of the latent representation. For this reason, each quantizer-specific codebook (each quantizer-index converter-specific codebook) can define different associations with different positions in the latent representation. For example, at each position of the latent representation, the first interval of highly frequent scalar values ​​will be mapped by a first large number of indices (thus having low approximation error), while at the same first position, the second interval of infrequent scalar values ​​will be mapped by a smaller number of indices (thus having high approximation error). Therefore, during training, each position of the latent representation is assigned an index distribution representing the probability of each interval of scalar values. In other words, each index approximates a small segment of scalar values ​​in the highly frequent intervals and a long segment of scalar values ​​in the infrequent intervals. In general, the global range of values ​​to be quantized can be divided into unequal intervals.

[0107] like Figure 2a and Figure 4 As shown, encoder 2 (2b, 2c) and / or decoder 10 (10b, 10c) can operate in different modes, but at least in two modes:

[0108] - First mode (first encoding mode 341 at the encoder; first decoding mode 541 at the decoder); and

[0109] - Second mode (second encoding mode 342 at the encoder; second decoding mode 542 at the decoder);

[0110] -Optionally, other modes may be defined (e.g., at least one other encoding mode and / or at least one other decoding mode).

[0111] Typically, different modes offer different qualities. For example, first modes 341 and 541 may offer reduced quality (e.g., reduced resolution) compared to second modes 342 and 542, but may also mean a lower bit rate and / or a reduction in computational power compared to second modes 342 and 352. A selection (353 at the encoder; 553 at the decoder) can be performed to choose between modes.

[0112] However, in some examples, the different modes may only behave differently. For example, different modes may be used for different situations, such as different classification results for audio signal 1, and are therefore called classification modes. For example, if the frame is classified as a voiced frame, the voiced guided classification mode can be selected, while if the frame is classified as an unvoiced frame, the unvoiced guided classification mode can be selected.

[0113] In some examples, the pattern is uniquely located within the quantization module 300 (for the encoder) and the dequantization module 500 (513) of the decoder, and is completely ignored by the first learnable segment 20 of the encoder and / or by the second learnable segment 520 of the decoder.

[0114] Different modes can be obtained using different training sessions (which can be independent of, for example, training sessions for the first learnable segment 20 and the second learnable segment 520). Different modes may mean different instances of quantization module 300 or dequantization module 500 (513).

[0115] Figures 5a to 5c Example of providing a choice between the first encoding mode 341 and the second encoding mode 342:

[0116] -like Figure 5a As shown, selection 553 (e.g., via selection command 353a from quantization controller 358) can be based at least in part on signal 359a indicating manual selection or application of selection;

[0117] -like Figure 5b As shown, selection 353 (e.g., notified via selection command 353b, for example, from quantization controller 358) may be based at least in part on signal 359b indicating the state 359a of a communication link (e.g., a communication network), for example, as measured by channel state meter 359, in order to adapt to the state of the communication link; and

[0118] -like Figure 5cAs shown, selection 353 (e.g., notified by selection command 353c from quantization controller 358) can be based at least in part on classification result 360c indicating classification of input audio signal 1 (e.g., for a specific frame) (e.g., performed by signal classifier 360) in order to adapt to the specific input audio signal 1 (classification result 360c can, for example, distinguish voiced frames from unvoiced frames).

[0119] It is worth noting that Figures 5a to 5c Examples can be combined with each other: Option 535 can be based on any one of the following criteria (or any combination thereof): classification result 360c, state 359a, and option 359a.

[0120] Figures 6a to 6b An example is shown at decoder 10 (10b, 10c). In general, in at least some examples, selection 553 can be performed as above (e.g., via dequantization controller 558). Figure 6a (and Figure 5a (Relative) Examples of selection 553 based on user selection 559a and / or application selection 559a are shown. Not shown but implementable with... Figure 5b and Figure 5c Examples of relative examples. Figure 6b The example demonstrates how selection 553 is performed based on the side information in encoded signal 3 (e.g., notified via command signal 553d) (e.g., after selection 353 is performed by the encoder and signaled as side information in encoded signal 3).

[0121] Figure 2b An example of selection as a quantizer number-oriented selection 353 is shown. Here, selection 353 is performed between a first encoding mode 341 (indicated here by 341N) and a second encoding mode 342 (indicated here by 342N). Here, selection 353 can be performed between at least the following:

[0122] - A first low quantizer number encoding pattern 341, wherein the number N2' (e.g., N2'=8) of the potential channels (potential scalar values) 351 of the first potential representation 350 is low (the number of quantizers 355 is also low); and

[0123] - Second high quantizer number encoding pattern 342, wherein the number N2" (e.g., N2"=16) of the potential channels (potential scalar values) 351 of the first potential representation 350 (and the number of quantizers 355) is higher than the number N2' in the first encoding pattern.

[0124] Similarly, even if not explicitly shown, at the decoder, a choice 553 may exist between the first decoding mode 541 and the second decoding mode 542 (e.g., the side information required by the encoded signal 3, such as regarding...). Figure 6b Here, option 553 can be selected from at least the following:

[0125] - A first low-quantization index converter number decoding mode 541, wherein the number N2' (e.g., N2'=8) of the index 556 and the number N2' of the potential channels (potential scalar values) 551 of the first potential representation 550 is low (and the number N2' of the quantizers 555 is also low); and

[0126] - The second high-quantization index converter number decoding mode 542, wherein the number N2" (e.g., N2"=16) of the index 556 and the number of potential channels (potential scalar values) 556 of the first potential representation 350 (and the number of quantizers 355) is higher than the number N2' in the first encoding mode.

[0127] It is noteworthy that, in this case, the selection between the first mode and the second mode (at the encoder and / or at the decoder) can be performed independently of the first learnable segment 20 on the encoder side and / or the second learnable segment 520 on the decoder side. Generally, the first low quantizer number encoding mode 341 and the first low quantizer number decoding mode 541 provide lower quality (e.g., worse resolution) than the second high quantizer number encoding mode 342 and the second low quantizer number decoding mode 542. However, the first low quantizer number encoding mode 341 and the first low quantizer number decoding mode 541 typically require a lower bit rate than the second high quantizer number encoding mode 342 and the second low quantizer number decoding mode 542, and are therefore more suitable, for example, in the case of a busy communication link. Advantageously, the selection 353 at the encoder can be based, for example, on a measurement 359b of the state 359a of the communication link (e.g., a communication network), as in Figure 5b For example, in the case of low network performance (e.g., high busy conditions and / or high error rates), a first mode can be selected (at both the encoder and decoder) to provide a low bit rate version of the encoded signal 3, while in the case of good network performance (e.g., low busy conditions and / or low error rates), a second mode can be selected (at both the encoder and decoder) to provide satisfactory audio quality.

[0128] It is also possible to have more than two modes, each of which is associated with a corresponding bit rate and / or a corresponding resolution, in order to perform quantizer-guided selection to provide a good trade-off between the requested quality and the available bit rate.

[0129] Therefore, in some examples, in quantizer-oriented selection, it can be summarized that the higher the available bit rate of the transmitted encoded signal 3, the higher the number of selectable quantizers (and coherently, the higher the dimension of the second potential representation 350 on the encoder side and the higher the dimension of the first potential representation 550 on the decoder side).

[0130] It should be noted that, Figure 2b In the case of quantizer-oriented selection, the first low quantizer number mode 341 and the first low quantizer index converter number mode 541 can also be regarded as examples of the first low index number mode, because the low number of indexes N2' is accompanied by the low number of quantizers 355 and 555. Similarly, the second high quantizer number mode 341 and the second high quantizer index converter number mode 541 can also be regarded as examples of the second low index number mode, because the higher number of indexes N2" is accompanied by the higher number of quantizers 355 and 555 N2".

[0131] Figure 2c An example of selection 353 as codebook selection 353C is shown. Here, selection 353C is performed between the first encoding mode 341 (indicated here by 341C) and the second encoding mode 342 (indicated here by 342C). Here, selection 353 can be performed between at least the following:

[0132] - First encoding mode 341 (341C), in which the first codebook 357C1 is used; and

[0133] - Second encoding mode 342 (342C), in which the second codebook 357C2 is used.

[0134] Similarly, at the decoder, codebook selection 553 can be performed between the first decoding mode 341 and the second decoding mode 342. Here, selection 553 can be performed between at least the following:

[0135] - First decoding mode, in which the first codebook is used; and

[0136] - Second decoding mode, in which a second codebook is used.

[0137] For example, the second codebook used for the second encoding / decoding mode may have more indices (or at least most of them) and / or higher bit lengths than the first codebook used for the first encoding / decoding mode. Therefore, the second encoding / decoding mode can theoretically achieve better resolution (but with a higher bit length). In this example, a higher bit rate results in higher resolution, and the first mode is preferred.

[0138] even though Figure 2c The display shows the option to select a shared codebook, but (for example, in...) Figure 1c(In the middle) The quantizer-specific codebook can be selected (that is, the selection is between a first set of quantizer-specific codebooks and quantization index converter-specific codebooks and a second set of quantizer-specific codebooks and quantization index converter-specific codebooks).

[0139] It should also be noted that in some examples, the first mode can mean both quantizer number-oriented selection and codebook selection. For example:

[0140] - The first low-resolution encoding / decoding mode may use a first low-bit-length codebook (or a set of first low-bit-length codebooks) and a small number of quantizers (quantization index converters, respectively), for example, to achieve low resolution and high bit rate by maintaining a low bit length (e.g., in the event of a poor communication link); and

[0141] - The second high-resolution encoding / decoding mode may use a second high-bit-length codebook (or a set of second high-bit-length codebooks) and a large number of quantizers (quantization index converters, respectively) to achieve high resolution at a low bit rate (e.g., in the case of the execution state of a communication link), with a higher bit length than in the first low-resolution encoding / decoding mode.

[0142] However, it should be noted that the selection of 353 and 553 is not necessarily between high-resolution and low-resolution modes. In some examples, different selectable codebooks 357C1 and 357C2 may be used for different applications and / or different audio signals. For example, as determined by the result 360c of classification 360 ( Figure 5c This allows selecting a first encoding / decoding mode for voiced frames and a second encoding / decoding mode for unvoiced frames. In this case, there may not be a difference in bitrates between different resolutions / quality levels, but only differences in codebooks, where one might be more suitable than the other.

[0143] Figure 2eAnother mode of operation for encoder 2 (2b, 2c) is shown. Here, a first latent representation 330 (469) is provided to quantization module 300 (indicated here by 300e). In this case, a first encoding mode 341' and a second encoding mode 342' can be executed in parallel. The first encoding mode 341' outputs a first encoded signal 3' and the second encoding mode 342' outputs a second encoded signal 3". Subsequently, the first encoded signal 3' can be re-decoded at block 341e' to obtain a first re-decoded version 1' of the input audio signal 1, and the second encoded signal 3" can be re-decoded at block 341e" to obtain a second re-decoded version 1" of the input audio signal 1. The selection at block 353e can be based on comparison, for example, by comparing the original version 1 of the input audio signal with the first re-decoded version 1' and the second re-decoded version 1" to determine a first distortion measure indicating the distortion of the first re-decoded version 1' compared to the original audio signal 1 and a second distortion measure indicating the distortion of the second re-decoded version 1" compared to the original audio signal 1, and by further comparing the first distortion measure with the second distortion measure. The mode that minimizes distortion is selected, as shown by switch 353e': the encoded signal 3 to be provided to the bitstream will thus be the encoded signal 3' and 3" with the minimized distortion compared to the original audio signal 1. Instead of encoded version 3, the same operation can be performed on the version formed by the index upstream of the encoded signal writer 360 (indicated by 556' and 556"). Advantageously, it is not necessary to perform a technique at the decoder that executes the two modes 341' and 342" in parallel (or, in other examples, more than two modes in parallel).

[0144] Instead of comparing the first re-decoded version 1' and the second re-decoded version 1" of the input audio signal 1 with the input audio signal 1, Figure 2e An example can be configured to execute both a first encoding mode (341') ​​for providing a first encoded signal version (3') and a second encoding mode (341") for providing a second encoded signal version (3") in parallel, but by selecting the encoded signal version (3', 3") that maximizes processing efficiency (e.g., minimizes computational cost) between the first and second encoding modes in which the encoded signal version (3', 3") is written into the encoded signal 3.

[0145] exist Figure 2dThe diagram illustrates an example of an encoder (which in some examples may be one of 2, 2b, or 2c) that can select between a first encoding mode 341 and a second encoding mode 342. Thus, the first encoding mode 341 represents a combination of at least one scalar quantizer (355) that maps a single per-latent channel scalar value (351S) to a per-latent channel index (356) and at least one vector quantizer (355S) that maps a subset of all per-latent channel scalar values ​​(351S) to a single index (356S). Even if not shown in the diagram, a decoder (in some examples, decoder 10) that can select between a first decoding mode (e.g., corresponding to the first encoding mode 341) and a second decoding mode (e.g., corresponding to the second encoding mode 342) may also exist. The different modes may be controlled, for example, by bit rate, to adapt to the connection state (e.g., a busy connection state implies a low bit rate, which requires the use of the first encoding mode, while a less busy connection state may imply a higher bit rate, which requires the use of the second encoding mode).

[0146] exist Figure 2d In the example, in the first low quantizer number encoding mode 341, there are fewer quantizers (e.g., for each frame or for each latent representation) than in the second encoding mode 342. For example, in the first encoding mode 341, there is at least one vector quantizer 355S that quantizes a vector formed by at least two scalar values ​​(e.g., a first scalar value 351S' and a second scalar value 351S) to a single index 356S, while in the second encoding mode 342, each of the first scalar value 351S' and the second scalar value 351S" is quantized independently of the scalar quantizers 355' and 355" respectively to provide two independent indices 355' and 355" respectively. Conversely, at the decoder, in the first low quantizer number decoding mode, there are fewer quantization index converters (e.g., for each frame or for each latent representation) than in the second high quantization number decoding mode. For example, in the first decoding mode, there is at least one quantization index converter that converts at least one single index 556 into at least two scalar values, while in the second decoding mode, there are two indices that map to at least two scalar values.

[0147] In several examples, it can be assumed that the length (e.g., time length) of indices 356, 356S is fixed (e.g., all indices 356, 356S require the same number of bits or other symbols written in the encoded signal 3), or that the length of the indices is variable (e.g., some indices have a shorter length in terms of their bitstream representation compared to others). For example, index 356S (generated in the first encoding mode 341) can have a shorter representation in the bitstream (encoded signal) 3 compared to the two indices 356', 356" that would be needed in the second encoding mode 342. For this reason, in the first encoding mode 341, the length of the encoded signal 3 is reduced, which is advantageous, for example, when a low bit rate is required (e.g., when the transmission channel is noisy), while the second encoding mode 342 allows for better resolution (because the codebook provides more indices), but the payload and computational power consumption increase. Therefore, in Figure 2d In the example, both the encoder and decoder are better adapted to the characteristics of the communication link.

[0148] exist Figure 2d In the example, selection can be understood as a choice between a first vector (or at least a partial vector) mode (with at least one vector quantizer 355S) and a second scalar mode for the same frame, with all scalar quantizers 355S.

[0149] It should be noted that, in parallel, the encoder and / or decoder may select between at least a second high-index-number encoding mode and a first low-index-number encoding mode. Compared to the first low-index-number encoding mode, at least one codebook with a higher number of indices, higher resolution, and / or higher bit length (at least on average, e.g., at least for most indices or at least for the most frequent indices) may be used in the second high-index-number encoding mode. Examples can be provided by... Figure 2dProvided: The second high index number encoding pattern is represented by the second pattern 342, while the first low index number encoding pattern is represented by the first pattern 341 (and in fact, the second pattern 342 has fewer encoded indices than the first pattern 341, because in the first pattern 341, only one index 365S is generated from multiple scalar values ​​351S' and 351S", while in the second pattern 342, there are a larger number of quantizers 355, each of which provides exactly one index, thus providing more indices than in the first pattern 341). However, in another example, even without any vector quantizers, the first low index number encoding pattern can also be achieved by enabling the second learnable region... The segment is represented by producing fewer scalar values ​​351 in the first mode than in the second mode. Conversely, at decoder 10, even without any vector quantizers processing multiple scalar values, the first low-index-number decoding mode can also be represented by providing fewer scalar values ​​551 to the first learnable segment 510 in the first mode than in the second mode. It should be noted that, in general, reducing the number of indices also reduces resolution. However, the first low-index-number encoding (or decoding) mode allows for a reduction in the length of the encoded signal 3, thus adapting to poorly performing communication links, while the second high-index-number encoding (or decoding) mode allows for improved quality, for example, when the communication link is satisfactory.

[0150] Figure 7a and Figure 7b Examples of multi-level (residual) quantization and multi-level (residual) dequantization are shown respectively.

[0151] Figure 7a Examples of multi-level (residual) quantization applicable to encoder 2 (e.g., 2b, 2c) are shown. Specifically, a residual quantizer 355R is shown (which may be an instance of any of the quantizers 355 discussed above), but in one example, other residual quantizers (e.g., N2 in number) are the same as residual quantizer 355R. Specifically, residual quantizer 355R may include a family with at least two sub-quantizers, such as a basic sub-quantizer (first sub-quantizer) 355R1 and a residual sub-quantizer (second sub-quantizer) 355R2, but in some examples, there are more than two sub-quantizers. Here, the potential scalar value (potential channel) 351 output by the second learnable segment 350 can be input into the basic sub-quantizer (first sub-quantizer) 355R1. Thus, a first index (basic index) 356R1 is obtained, and can therefore be encapsulated into the encoded signal 3 (or in Figure 2eIn the example, it is encapsulated in version 3' or 3"). Next, the first index (basic index) 356R1 can be input to the dequantizer (quantization index converter) 555R1 (which is essentially the quantization index converter at the analog decoder). Then, the dequantizer (quantization index converter) 555R1 generates a dequantized version 551R1 of the first index 356R1. The dequantized version 551R1 of the first index 356R1 thus represents how the decoder will dequantize the analog of the potential scalar value 351. Next, at comparison block 353R1, the potential scalar value 351 is compared with the dequantized version 551R1 of the first index 356R1, thereby providing a first residual potential scalar value 351R1, which represents the quantization error that impairs the dequantized version 551R1 of the first index 356R1. The first residual potential scalar value 351R1 can then be input to the second subquantizer 355R2 to obtain the second index (residual index). 356R2, the second index can also be encapsulated into the encoded signal 3 (or... Figure 2e In the example, encapsulated in version 3' or 3"). If only the first index 355R1 and the second index 355R2 are required, the inverse subquantizer 555R2 and the second comparison block 353R2 can be avoided; otherwise, the inverse subquantizer 555R2 and the second comparison block 353R2 (similar to the inverse subquantizer 555R1 and the first comparison block 353R1) can be used to obtain the third index from the second residual value 351R2.

[0152] Figure 7b This example demonstrates multi-stage (residual) dequantization (inverse quantization) at decoder 10 (e.g., 10b, 10c), specifically showcasing two quantization index converters 555R (which can be... Figures 3a to 3c Example of a quantization index converter 555). Here, the encoded signal reader 560 provides a first index (basic index) 556R1 (corresponding to...). Figure 7a The first index (356R1) and the second index (residual index) (556R2, corresponding to...) Figure 7a The second residual index 356R2). Next, the indices 556R1 and 556R2 are dequantized at the two corresponding dequantizer instances 555R1' and 555R2' to obtain the first component (basic component) 551'R1 and the second component (residual component) 551'R2 of the dequantized potential scalar value 551 to be obtained. Then, in the additional block 553R, the first component (basic component) 551'R1 and the second component (residual component) 551'R2 are added to each other to obtain the dequantized potential scalar value 551, which is a part of the dequantized first potential representation 550 to be input into the first learnable segment 540.

[0153] even though Figure 7a and Figure 7b Not shown, each of the subquantizers 355R1, 355R2 and the inverse subquantizers 555R1, 555R2, 555R1' and 555R2' is also intended to utilize a codebook. The codebook may be the same for each quantization step (or for each inverse quantization step separately), or different codebooks may be used (e.g., a basic codebook may be used for the first subquantizer 355R1, and different residual codebooks may be used for the second residual subquantizer 355R2). This also applies to the inverse subquantizer.

[0154] It is also possible to select (353, 553) based on at least the first mode and the second mode (in the example, other selectable modes are possible). Figure 7a and Figure 7b The operation of the encoder and decoder. For example:

[0155] - In the first non-residual mode, the encoder generates a single index 356 for each latent scalar value 351, and the decoder uses only a single index 356 (with its decoder-side version 556) to generate a single dequantized latent scalar value 556.

[0156] - In the second residual mode, the encoder (e.g., in...) Figure 7a For each potential scalar value 351, a first basic index 356R1 and at least one second residual index 356R2 are generated, and the decoder (e.g., in...) generates a first basic index 356R1 and at least one second residual index 356R2 ... Figure 7b The first basic index 356R1 (in decoder-side version 556R1) is used to generate the first component 551R1' and at least one second residual index 356R2 (in decoder-side version 556R2) is used to generate the second component 551R2'.

[0157] Generally, the first non-residual mode achieves lower quality (e.g., lower resolution) but has a low bit rate, while the second residual mode achieves higher quality but requires a high bit rate. This can be achieved at the encoder. Figures 5a to 5c One technique illustrated involves selecting between a first non-residual mode and a second residual mode 353, while at the decoder, the selection can be performed based on signaling in the encoded signal 3, for example... Figure 7a In the example, selecting (353, 553) can be performed between at least a first residual mode with a lower level of residual quantization and at least a second residual mode with a higher level of residual quantization.

[0158] In summary, the selection can be made between at least two (but in some examples, more than two) encoding / decoding modes. At least one (e.g., some, all) of the following statements may apply:

[0159] 1) The first mode is a low quantizer number mode (341, 342) and the second encoding mode is a high quantizer number mode (341, 342), for example, because the number of potential scalar values ​​(potential channels) (N2', N2") varies between the first and second modes (e.g.) Figure 2b (in the middle), or because the first pattern is at least a partial vector pattern (e.g. Figure 2d (in Chinese) and the second mode is a scalar mode (e.g.) Figure 2d middle);

[0160] 2) The first mode uses at least one first codebook, and the second mode uses at least one second codebook, with different resolutions, bit lengths, and / or index numbers (see...). Figure 2c );

[0161] 3) The first mode is a low index number mode (341, 342) and the second mode is a high index number mode, for example... Figure 2c and Figure 2d middle;

[0162] 4) The first mode is the first potential mode that reduces the potential, while the second mode is the second potential mode that increases the potential.

[0163] 5) The first mode is the basic mode (single-level mode) and the second mode is the residual mode (multi-level mode, such as...). Figure 7a and Figure 7b (in the middle), or the first mode is multi-level but has a first number of levels, the first number being less than the second number of levels in the second mode;

[0164] 6) At the encoder only, the first encoding mode and the second encoding mode can be executed in parallel, so as to select between the first encoding mode and the second encoding mode by selecting to write the encoded signal version with minimal distortion into the encoded signal 3;

[0165] 7) The first mode offers a higher resolution than the second mode;

[0166] 8) One mode can be a classification mode for a specific category, such as a classification mode for voiced consonants, while the other mode can be a classification mode for unvoiced consonants.

[0167] The above statements can be combined with each other in different combinations.

[0168] The pattern can be accessed through selectable options (e.g.) Figure 5a (middle), channel status measurement (e.g.) Figure 5b (in Chinese) and signal classification (e.g.) Figure 5c The selection is based on at least one of the criteria in ( ).

[0169] As explained above, latent representations (e.g., 330, 350, 530, 550, etc.) can be represented by matrices (e.g., M×N matrices), where M can be the number of latent channels and N can be the length of a frame. Generally, when explaining the selection of different encoding / decoding modes, examples are frame-based: for example, low-index-number modes and high-index-number modes have different numbers of indices per frame; a first low-quantizer-number mode and a second high-quantizer-number mode have different numbers of quantizers per frame, and so on.

[0170] Discussion

[0171] Model: The model of this invention may include (not necessarily symmetrical) a pair of convolutional encoders (e.g., 2, such as 2b or 2c) and decoders (e.g., 10, such as 10b or 10c) and quantization modules (e.g., 300, 313). The encoders may transform a first latent representation 330 into a (typically) lower-dimensional representation 350 input to a plurality of quantizers 355. Each quantizer 355 independently approximates each element (latent scalar value, latent channel) 351 of the latent vector, for example, by the closest match from a set of candidate values ​​(e.g., stored in a codebook). This set of candidate values ​​may be learned or may be fixed. The index 356 of the candidate values ​​for each latent dimension is stored or transmitted to a receiver (e.g., decoder 10, such as 10b or 10c), from which the decoder reconstructs the corresponding quantizer input vector. (e.g., convolutional) The decoder may reconstruct the latent (e.g., its versions 550 and 530), which is then used to reconstruct the input signal 16 by a NAC decoder 520.

[0172] Training: The SQ 300 (313) can be trained by approximating each quantizer 355 with an identity in the reverse path and enforcing the quantizer structure with an additional loss, or by simulating the effect of the quantizer by adding uniform noise to each potential dimension (scaled to match the target quantization resolution). In both cases, the intermediate (second) representation 350 of the SQ module can be quantized element-wise during inference.

[0173] Scalability: Data rate can be traded off against signal quality, for example by reducing the quantizer resolution of the pre-trained NAC during inference or by training a new SQ for the pre-trained NAC. Here, the corresponding pre-trained SQ module 300 (313) can be approximated by a weaker SQ, which provides an NAC that works at a lower data rate simply by minimizing the loss of measuring the difference between the student and teacher outputs using a student-teacher approach.

[0174] Benefits derived from the proposed technology (inference)

[0175] 1. The SQ module 300 (313) allows for a latent representation 350, which is orders of magnitude smaller than the latent representation of a conventional VQ providing comparable quality. Thus, the overall design of the NAC encoder (first learnable segment) 20 can be accomplished with fewer parameters and higher computational efficiency.

[0176] 2. The discrete representation learned by SQ 300 (313) is easily interpreted as a latent approximation because the distance between the latent representation and the SQ approximation is computed between scalars.

[0177] 3. Quantization can be achieved extremely efficiently through integer casting or other efficient methods.

[0178] 4. SQ module 300 (313) also works with a fixed codebook.

[0179] 5. During inference, different SQ modules can be used (e.g., instantiating different encoding / decoding modes) with varying numbers and distributions of encoding layers and / or encoder / decoder pairs, allowing for a different number of dimensions in the potential. This provides an easy and efficient way to balance data rate and quality. Furthermore, it eliminates the need to store several complete NACs corresponding to different data rates, requiring only a single NAC and several (tiny) SQ models.

[0180] 6. A single SQ module 300 (313) is sufficient, while most VQ methods must rely on residual quantization.

[0181] Benefits from the proposed technology (training)

[0182] 1. Strategies for achieving scalability have been proposed for SQ (see above). These techniques avoid the costly retraining of the entire NAC (approximately several weeks on a large GPU) and only require retraining 300 and / or 500 SQ modules (approximately several hours on a small GPU) or only require adjusting the quantizer levels.

[0183] 2. No additional mechanisms are needed for training the codebook.

[0184] 3. Robust training methods, namely, pass-through training and noise-based training.

[0185] 4. The proposed quantization technique converges faster than the competing VQ method.

[0186] Applications and benefits of the present invention

[0187] ● Computing efficient speech coding

[0188] ●Scalable speech coding

[0189] ● Storing efficient speech codes

[0190] ● It may have better quality at higher data rates (to be verified).

[0191] For aspects that are extremely important for deduction (for readability, the "patent template" is not used),

[0192] 1. Efficient quantization through integer casting

[0193] 2. Regarding the scalability of quantizer resolution

[0194] 3. By switching the scalability of the retrained SQ with respect to dimension

[0195] For aspects that are extremely important in training:

[0196] 1. Train the SQ function by approximating the quantizer with appropriately scaled, uniformly distributed white noise.

[0197] 2. Train SQ using a direct approximation, that is, by using an identity approximation quantizer with additional loss.

[0198] 3. Train SQ using a smooth substitution (Softmax) approximation quantizer.

[0199] 4. Training the codebook using moving averages

[0200] 5. Training the codebook via backpropagation

[0201] 6. Training the residual codebook

[0202] 7. Dictionary of training codebooks

[0203] 8. Retrain the SQ class pre-trained with NAC

[0204] Select the data rate by switching between scalar quantizer (SQ) modules.

[0205] Training: For all of the following data rate adjustment options, the NAC encoder / decoder (20, 520) and SQ encoder / decoder (300, 500) are trained together with SQ using a distribution of the codebook hierarchy (CL) (user-defined and fixed or learnable).

[0206] Inference: For all of the following options, the trained SQ module 300 (313) or a portion thereof of the NAC encoder / decoder (20, 520) and SQ encoder / decoder (300, 500) is replaced by a user-defined or retrained alternative that allows the data transmission rates of the NAC 20 and / or 520 to be different.

[0207] 1) Adjust the user-defined CL (parametric quantizer 355) of the trained neural audio encoder (NAC) during application.

[0208] a) A certain number of CLs of the parameterized quantizer 355 of the SQ module 300 (313) are selected by the user in a certain distribution (e.g., uniform or with higher resolution for smaller values).

[0209] b) Train a certain number of CLs of the parameterized quantizer 355 of the SQ module 300 (313), while keeping the NAC encoder / decoder (20, 520) and the SQ encoder / decoder (300, 500) fixed. During application, the application / user can switch between these trained codebooks.

[0210] c) Options a) and b) can be applied globally (using a single codebook for all potential channels) or per potential channel (using a different codebook for each potential channel). The resolution can be the same or different for all potential channels (providing better resolution to more important channels, and vice versa).

[0211] 2) Switching between the retrained SQ module, including the SQ encoder / decoder and SQ.

[0212] a) Train a new SQ encoder / decoder pair (340, 540), where the bottleneck dimension may differ from the dimension in the NAC training with the SQ (deterministic or trained) including a user-defined number of CLs. During application, combine the NAC encoder / decoder (20, 520) with different retrained SQ modules 300 (313).

[0213] b) Perform 2a) and then apply methods 1a) to 1c), that is, keep the NAC encoder / decoder (20, 520) and SQ encoder / decoder (300, 500) from 2a) fixed and only readjust the CL of SQ.

[0214] The following are non-restrictive examples of specific parts of the above examples.

[0215] Examples of some features

[0216] Figure 10An example of a vocoder (or more generally, a system for processing audio signals) system is shown. The vocoder system may include, for example, an encoder 2 (e.g., 2b, 2c) and / or a decoder 10 (e.g., 10b, 10c). As explained above, the encoder 2 may include a first encoder-side learnable layer (NAC encoder) 20, also referred to as an audio signal representation generator, to generate a first latent representation (audio signal representation) 330 (469) of the input audio signal 1. The input audio signal 1 may be processed by the first encoder-side learnable layer 20. The first latent representation 330 of the input audio signal 1 may be stored (and, for example, for the purpose of processing audio signals), or may be quantized (e.g., by a quantizer 300) to obtain a bitstream 3. The decoder 10 (audio generator) may read the bitstream 3 and generate an output audio signal 16.

[0217] Each of the first encoding side learnable layer 20, encoder 2, and / or decoder 10 may be a learnable system and may include at least one learnable layer and / or learnable block.

[0218] The input audio signal 1 (which may be obtained, for example, from a microphone or from other sources such as a storage unit and / or a synthesizer) may belong to a type having a sequence of audio signal frames. For example, different input audio signal frames may represent sound in a fixed time length (e.g., 10 ms or milliseconds, but in other examples, different lengths may be defined, such as 5 ms and / or 20 ms). Each input audio signal frame may include a sequence of samples (e.g., at 16 kHz or kilohertz, and 160 samples will be present in each frame). In this case, the input audio signal is in the time domain, but in other cases, it may be in the frequency domain. The input audio signal 1 may be provided to a learnable block 200, which may be a portion of a first learnable segment. The learnable block 200 may belong to a type having two paths (e.g., dealing with at least one residual). The learnable block 200 may provide a processed version 269 of the input audio signal 1 to a second learnable block 290 (in some cases, this block may be avoided). Subsequently, learnable block 200 or learnable block 290 can provide a processed version of its output input audio signal 1 to quantization module 300. Quantization module 300 can provide an encoded signal (bitstream) 3. It will be seen that quantization module 300 can be a learnable quantization module.

[0219] The learnable block 200 can process the input audio signal 1 (or one of its processed versions) after it has been converted into a multidimensional representation. A format definer 210 can therefore be used. The format definer 210 can be a deterministic block (e.g., a non-learnable block). Downstream of the format definer 210, the processed version 220 (also referred to as the first audio signal representation of the input audio signal 1) output by the format definer 210 can be processed by at least one learnable layer (e.g., 230, 240, 250, 290). At least the learnable layers within the learnable block 200 (e.g., layers 230, 240, 250) are learnable layers for processing the first audio signal representation 220 of the multidimensional version (e.g., a two-dimensional version) of the input audio signal 1. As will be shown, this can be achieved, for example, by a scrolling window that moves along a single dimension (temporal domain) of the input audio signal 1 and produces the multidimensional version 220 of the input audio signal 1. As can be seen, the first audio signal representation 220 of the input audio signal 1 may have a first dimension (inter-frame dimension) such that multiple consecutive frames (e.g., relative to each other immediately following a frame) are ordered according to (along) the first dimension. It should also be noted that a second dimension (intra-frame dimension) such that samples of each frame are ordered according to (along) the second dimension. Figure 10 or Figure 11 As can be seen, in some examples, frames t can then be organized along the second direction (inter-frame direction) using two samples 0' and 0'. As can be seen, this sequence of frames t, t+1, t+2, t+3, etc., can be followed along the first dimension, and in the second dimension, a sample sequence is also followed for each frame. Format definer 210 can be configured to insert input audio signal samples for each given frame along the second dimension (e.g., intra-frame dimension) of the first multi-dimensional audio signal representation of the input audio signal. Alternatively, format definer 210 can be configured to insert additional input audio signal samples (e.g., in a predefined number, such as application-specific, such as user- or application-defined) of one or more additional frames immediately following a given frame along the second dimension (e.g., intra-frame dimension) of the first multi-dimensional audio signal representation 220 of the input audio signal 1. Format definer 210 is configured to insert additional input audio signal samples (e.g., in a predefined number, such as application-specific, such as user- or application-defined) of one or more additional frames immediately preceding a given frame along the second dimension of the first multi-dimensional audio signal representation 220 of the input audio signal 1. However, in some examples, this is not necessary and can be used to avoid inserting samples from other frames.

[0220] Downstream of format definer 210, at least one learnable layer (230, 240, 250) can be input to an audio signal representation 220 with input audio signal 1. Notably, in this case, at least one learnable layer 230, 240, and 250 can follow a residual technique. For example, at point 248, a residual value can be generated from the audio signal representation 220. Specifically, the audio signal representation 220 can be subdivided into a main portion 259a' of the audio signal representation 220 of the input audio signal and a residual portion 259a. Therefore, the main portion 259a' of the audio signal representation 220 can be left unprocessed until point 265c, where the main portion 259a' of the audio signal representation 220 is added (summed) to a processed residual version 265b' output by, for example, at least one learnable layer 230, 240, and 250 cascaded together. Thus, a processed version 269 of the input audio signal 1 can be obtained.

[0221] At least one residual learnable layer 230, 240, 250 may include at least one of the following:

[0222] - An optional first learnable layer (230), such as a first convolutional learnable layer, is a convolutional learnable layer configured to generate a second multidimensional audio signal representation of the input audio signal (1) by sliding along a second direction [e.g., an intra-frame direction] of a first multidimensional audio signal representation (220) of the input audio signal (1);

[0223] - A second learnable layer (240) may be a recursive learnable layer (e.g., a gated recursive learnable layer) configured to generate a third multidimensional audio signal representation of the input audio signal (1) by operating along a first direction (e.g., an inter-frame direction) of a second multidimensional audio signal representation (220) of the input audio signal (1) [e.g., using a 1×1 kernel, such as a 1×1 learnable kernel, or another kernel, such as another learnable kernel];

[0224] - A third learnable layer (250) [which may be, for example, a second convolutional learnable layer], which is a convolutional learnable layer configured to generate a fourth multidimensional audio signal representation (265b') of the input audio signal by sliding along a second direction [e.g., an intra-frame direction] of a first multidimensional audio signal representation of the input audio signal [e.g., using a 1×1 kernel, for example, a 1×1 learnable kernel].

[0225] Notably, the first learnable layer 230 may be a first convolutional learnable layer. It may have a 1×1 kernel. The 1×1 kernel can be applied by sliding the kernel along the second dimension (i.e., for each frame). A recursive learnable layer 240 (e.g., a gated recursive unit GRU) may be input with the output from the first convolutional learnable layer 230. The recursive learnable layer (e.g., a GRU) may be applied along the first dimension (i.e., by sliding from frame t to frame t+1, to frame t+2, and so on). As will be explained later, in the recursive learnable layer 240, each value of the output for each frame may also be based on previous frames (e.g., the immediately preceding frame, or another number of n frames immediately preceding a particular frame; for example, in the case of n=2, the output of the recursive learnable layer 240 for frame t+3 will take into account the values ​​of samples from frames t+1 and t+2, but will not take into account the values ​​of samples from frame t). A processed version of the input audio signal 1 output by the recursive learnable layer 240 can be provided to the second convolutional learnable layer (third learnable layer) 250. The second convolutional learnable layer 250 may have a kernel (e.g., a 1×1 kernel) that slides along a second dimension (along the second intra-frame dimension) of each frame. The output 265b' of the second convolutional learnable layer 250 may then be added, for example at point 265c, to the principal portion 259a' of the audio signal representation 220 of the input audio signal 1, which has bypassed learnable layers 230, 240, and 250.

[0226] Next, the processed version 269 of the input audio signal 1 can be provided (as latent signal 269) to at least one learnable block 290. At least one convolutional learnable block 290 can provide, for example, a version with 256 samples (but different numbers can be used, such as 128, 516, etc.).

[0227] like Figure 11 (which can be regarded as) Figure 11 As shown in the example), at least one convolutionally learnable block 290 may include a convolutionally learnable layer 429 to perform convolution (e.g., using a 1×1 kernel) on a signal (potential) 269 (e.g., output by learnable block 200). The convolutionally learnable layer 429 may be a non-residual learnable layer. The convolutionally learnable layer 429 may output a convolutional version 420 of the signal 269, and may also be a processed version of the input audio signal 1.

[0228] At least one convolutional learnable block 290 may include at least one residual learnable layer. At least one convolutional learnable block 290 may include at least one learnable layer (e.g., 440, 460). Learnable layers 440, 460 (or at least one or more of them) may follow residual techniques. For example, at point 448, residual values ​​may be generated from the audio signal representation or latent representation 269 (or its convolutional version 420). Specifically, the audio signal representation 420 may be subdivided into a principal portion 459a' of the audio signal representation 420 of the input audio signal 1 and a residual portion 459a. Therefore, the principal portion 459a' of the audio signal representation 420 of the input audio signal 1 may not undergo any processing until point 465, where the principal portion 459a' of the audio signal representation 420 of the input audio signal 1 is added (summed) to the processed residual version 465b' output by at least one cascaded learnable layer 440 and 460. Therefore, a potential representation 469 (330) of the input audio signal 1 can be obtained and the output of the first learnable segment 20 (audio representation generator) can be represented.

[0229] At least one residual learnable layer in at least one convolutional learnable block 290 may include at least one of the following:

[0230] - First layer (430), which is configured to generate a residual multidimensional audio signal representation of the input audio signal (1) from the audio signal representation 420 (the first layer 430 may be an activation function, such as Leaky ReLu, see below);

[0231] - A second learnable layer (440), which is a convolutional learnable layer configured to generate a residual multidimensional audio signal representation of the input audio signal 1 by means of the audio signal representation output from the first learnable layer (430) through convolution [e.g., kernel 3 may be used];

[0232] - The third layer (450) is used to generate a residual multidimensional audio signal representation of the input audio signal 1 from the audio signal representation output from the second learnable layer (440) (the learnable layer 450 may be an activation function, such as Leaky ReLu, see below);

[0233] - A fourth learnable layer (460) is a convolutional learnable layer configured to generate a residual multidimensional audio signal representation 456b' of the input audio signal 1 from the output of the third learnable layer (450) via convolution [e.g., kernel 1×1].

[0234] The output 465b' of the second convolutional learnable layer 460 (fourth learnable layer) can then be added (summed) at point 465 to the principal portion 459a' of the audio signal representation 420 (or 269) of the input audio signal 1, which has bypassed layers 430, 440, 450, and 460.

[0235] It should be noted that output 469 (330) can be considered as being generated by the first coding-side learnable layer 20 (e.g., in...). Figures 1a to 1c The first latent representation of the output (in Chinese).

[0236] Subsequently, a quantization module 300 may be provided if it is necessary to write the encoded signal 3. The quantization module 300 may be a learnable quantization module [e.g., a quantization module using at least one learnable codebook], which is discussed in detail above. The quantization module (e.g., a learnable quantization module) 300 may associate the index of at least one codebook with each frame of the latent representation (e.g., 220 or 469) of the input audio signal (1) or a processed version of the first multidimensional audio signal representation in order to produce the encoded signal 3 [at least one codebook may be, for example, a learnable codebook].

[0237] It is noteworthy that the cascades formed by learnable layers 230, 240, 250 and / or by layers 430, 440, 450, 460 may include more or fewer layers, and different choices may be made. However, it is noteworthy that these layers are residual learnable layers, and they are bypassed by the main portion 259' of the audio signal representation 220.

[0238] Figure 12 exhibit Figures 3a to 3cThe decoder (audio generator) 10 (e.g., 10b, 10c) may be any example (but different examples may be used), and is therefore indicated by 10d. The encoded signal 3 may include frames (e.g., encoded as an index, for example encoded by encoder 2, for example after quantization by quantization module 300). An output audio signal 16 may be obtained. The decoder 10 (10d) may include a first data provider 702. The first data provider 702 may be input with an input signal (input data) 14 (e.g., from an internal source, such as a noise generator or storage unit, or from an external source, such as an external noise generator or external memory unit, or even data obtained from the encoded signal 3). The input signal 14 may be noise, such as white noise, or a deterministic value (e.g., a constant). The input signal 14 may have multiple channels (e.g., 128 channels, but other numbers of channels are possible, such as a number greater than 64). The first data provider 702 may output first data 15. The first data 15 may be noise or obtained from noise. First data 15 may be input into at least one first processing block 50 (40). First data 15 may (e.g., when acquired from noise, it thus corresponds to input signal 14) be independent of the output audio signal 16, but in some cases, it may be acquired from encoded signal 3, such as LPC parameters, or other parameters acquired from encoded signal 3; noteworthyly, an advantage of this example is that first data 15 need not be a specific acoustic feature, and first data 15 can more easily be noise. At least one first processing block 50 (40) may adjust first data 15 to obtain first output data 69, for example, using adjustments obtained by processing encoded signal 3. First output data 69 may be provided to a second processing block 45. From the second processing block, audio signal 16 may be obtained (e.g., synthesized via PQMF). First output data 69 may be in multiple channels. The first output data 69 can be provided to a second processing block 45, which can combine multiple channels of the first output data 69 to provide an output audio signal 16 in a single signal channel (e.g., after PQMF synthesis, such as in...). Figure 14 and Figure 10 The Chinese use 110 as an indicator, but... Figure 12 (Not shown in the text).

[0239] As explained above, the output audio signal 16 (and the original audio signal 1 and its encoded version, encoded signal 3 or its representation 20, or any other processed version thereof, such as 269, or residual versions 259a and 265b', or major version 259a', and any intermediate version output by layers 230, 240, 250, or any intermediate version output by any of layers 429, 430, 440, 450, 460) is generally understood to be subdivided according to the frame sequence (in some examples, the frames do not overlap with each other, while in some other examples they may overlap). Each frame may include a sequence of samples. For example, each frame may be subdivided into 16 samples (but other resolutions are possible). It should also be noted that multiple frames may be grouped into a single packet of encoded signal 3, for example, for transmission or storage. Although the duration of a frame is generally considered fixed, the number of samples per frame may vary, and upsampling operations may be performed.

[0240] Decoder 10 (10d) is available for:

[0241] - A first branch (e.g., a frame-by-frame branch) 10a', which can be updated for each frame, for example using a frame obtained from the encoded signal 3 (e.g., the frame can be in the form of an index quantized by quantization module 300 and / or in the form of a code (such as a scalar, vector) 112 (530) converted, for example, from dequantization module 500 (513), which is also an inverse quantization module or a dequantization module); and / or

[0242] - Second branch (e.g., sample-by-sample branch) 10b'.

[0243] The second branch 10b' may contain at least one of blocks 702, 77 and 69.

[0244] like Figure 12 As shown, an index 556 can be obtained from the dequantization module 500 (513) to obtain a first (decoder-side) latent representation 550. The first latent representation 550 may be multidimensional (e.g., two-dimensional, three-dimensional, etc.). The dequantization module 500 (513) may include (e.g., a learnable codebook).

[0245] The per-sample branch 10b' can be updated for each sample, for example, at the output sampling rate and / or for each sample at a sampling rate lower than the final output sampling rate, for example, using noise 14 or another input obtained from an external or internal source.

[0246] The first processing block 40 can operate like a conditional neural network, providing data (e.g., codes 112, 530) from the encoded signal 3 to generate conditions that modify the input data 14 (input signal). The input data (input signal) 14 (in any of its evolutions) will undergo several processing steps to obtain an output audio signal 16, which is intended to be a version of the original input audio signal 1. The conditions, the input data (input signal) 14, and its subsequent processed versions can all be represented, for example, as activation maps subjected to learnable layers via convolution. Notably, during its evolution toward speech 16, signal 1 can undergo upsampling (e.g., in...). Figure 14 In this case, from one sample 49 to multiple samples, such as thousands of samples), but the number of its channels 47 can be reduced (e.g., in...). Figure 14 In the meantime, the number of channels ranges from 64 or 128 to a single channel.

[0247] The first data 15 can be obtained, for example, from an input (such as noise or a signal from an external signal) or from other internal or external sources (e.g., per-sample branch 10b'). The first data 15 can be considered as input to the first processing block 40 and can be an evolution of the input signal 14 (or can be the input signal 14). Basically, the first data 15 is modified according to conditions set by the first processing block 40 to obtain the first output data 69. The first data 15 can be in multiple channels, for example, in a single sample. Also, the first data 15 provided to the first processing block 40 can have a single sample resolution, but in multiple channels. The multiple channels can form a set of parameters that can be associated with the encoded parameters encoded in the encoded signal 3. However, generally, during processing, in the first processing block 40, the number of samples per frame increases from a first number to a second higher number (i.e., the sampling rate, also referred to herein as the bit rate, increases from the first sampling rate to the second higher sampling rate). On the other hand, the number of channels can decrease from the first number of channels to the second lower number of channels. The conditions for the first processing block (discussed in detail below) can be indicated by 74 and 75 and generated by target data 12, which in turn is generated from target data 12 obtained from encoded signal 3 (e.g., via dequantization modules 500, 513). It will be shown that the conditions (adjustment feature parameters) 74 and 75 and / or target data 12 can also be upsampled to conform to (e.g., adapt) the dimensions of the version of target data 12. The unit providing first data 15 (from an internal source, an external source, encoded signal 3, etc.) is referred to herein as first data provider 702.

[0248] As from Figure 12As can be seen, the first processing block 40 may include a pre-tunable learnable layer 710, which may be or include a recursive learnable layer, such as a recursive learnable neural network, such as a GRU, but this is not required. The pre-tunable learnable layer 710 may generate target data 12 for each frame. The target data 12 may be at least 2-dimensional (e.g., multi-dimensional): each frame may have multiple samples in the second dimension and each frame may have multiple channels in the first dimension. The target data 12 may be in the form of a spectrogram, which may be a Mel spectrogram (but this is not strictly required), for example, in cases where the frequency scale is non-uniform and / or driven by cognitive principles. When the sampling rate corresponding to the tunable learnable layer to be fed is different from the frame rate, the target data 12 may be the same for all samples of the same frame, for example at the layer sampling rate. Another upsampling strategy may also be applied. The target data 12 may be provided to at least one tunable learnable layer, at least one tunable learnable layer is indicated herein as having layers 71, 72, 73 (see also...) Figure 15 (And below). Adjustable learnable layers 71, 72, and 73 can generate conditions (some of which may be indicated as beta β and gamma γ, or numbered 74 and 75), which are also referred to as adjustment feature parameters to be applied to the first data 12 and any oversampled data derived from the first data. Adjustable learnable layers 71, 72, and 73 can be in the form of a matrix with multiple channels and multiple samples per frame. The first processing block 40 may include an inverse normalization (or stylization component) block 77. For example, stylization component 77 can apply adjustment feature parameters 74 and 75 to the first data 15. Examples may be element-wise multiplication of the value of the first data with condition β (which may act as a bias operation) and addition with condition γ (which may act as a multiplier operation). Stylization component 77 can generate first output data 69 sample-by-sample.

[0249] Decoder 10 (10d) may include a second processing block 45. The second processing block 45 may combine multiple channels of the first output data 69 to obtain an output audio signal 16 (or its precursor audio signal 44', such as...). Figure 14 (As shown in the image).

[0250] For now, please refer to Figure 13 The encoded signal 3 is subdivided into multiple frames; however, these frames are encoded in the form of indices 356, 556 (e.g., obtained from the quantization module 300 of encoder 2). The quantization module 500 (513) obtains a first latent representation 550 from the indices 356, 556 of the encoded signal 3 to obtain a scalar value 551 to be grouped into the code. The first and second dimensions are shown in... Figure 13In code 112 (530) (other dimensions may exist). Each frame is subdivided into multiple samples in the horizontal axis direction (first inter-frame dimension). The first latent representation 550 may be used by a pre-tuned learnable layer 710 (e.g., a recursive learnable layer) to generate target data 12, which may also be at least two-dimensional (e.g., multi-dimensional), such as in the form of a spectrogram (e.g., a Mel spectrogram, but this is not strictly necessary). Each target data 12 may represent a single frame and the frame sequence may evolve over time in the horizontal axis direction (from left to right) along the first inter-frame dimension. For each frame, several channels may be in the vertical axis direction (second intra-frame dimension). For example, different coefficients will appear in different entries in the columns associated with the coefficients, which are associated with frequency bands. Tunable learnable layers 71, 72, 73 generate feature parameters 74, 75 (β and γ). The horizontal axis (second intra-frame dimension) of β and γ is associated with different samples in the same frame, while the vertical axis (first inter-frame dimension) is associated with different channels. In parallel, the first data provider 702 can provide first data 15. The first data 15 can be generated for each sample and can have multiple channels. At the stylization component 77 (and more generally, at the first adjustment block 40), adjustment feature parameters β and γ (74, 75) can be applied to the first data 15. For example, element-wise multiplication can be performed between a series of stylization conditions 74, 75 (adjustment feature parameters) and the first data 15 or its evolutions. It will be shown that this process can be repeated many times.

[0251] As clearly seen above, the first output data 69 generated by the first processing block 40 can be obtained as a 2D matrix, where the horizontal axis (first inter-frame dimension) represents samples and the vertical axis (second intra-frame dimension) represents channels. Through the second processing block 45, an audio signal 16 with a single channel and multiple samples can be generated (e.g., in a shape similar to the input audio signal 1), particularly in the time domain. More generally, at the second processing block 45, the number of samples per frame (bit rate, also known as the sampling rate) of the first output data 69 can evolve from a second number of samples per frame (second bit rate or second sampling rate) to a third number of samples per frame (third bit rate or third sampling rate) that is higher than the second number of samples per frame (second bit rate or second sampling rate). On the other hand, the number of channels of the first output data 69 can evolve from a second number of channels to a third number of channels that is less than the second number of channels. In other words, the bit rate or sampling rate (third bit rate or third sampling rate) of the output audio signal 16 can be higher than the bit rate (or sampling rate) (first bit rate or first sampling rate) of the first data 15 and the bit rate or sampling rate (second bit rate or second sampling rate) of the first output data 69, while the number of channels of the output audio signal 16 can be lower than the number of channels of the first data 15 (first channel number) and the number of channels of the first output data 69 (second channel number).

[0252] Examples of convolutions are discussed below and will be understood to be used at any of the preconditional learnable layer 710 (e.g., a recursively learnable layer), at least one conditional learnable layer 71, 72, 73, and more generally, at the first processing block 40 (50). Generally, the resulting set of conditional parameters (e.g., for a frame) may be stored in a queue (not shown) for subsequent processing by the first or second processing block when the previous frame is processed, respectively.

[0253] A discussion is now provided regarding the operations performed primarily in blocks downstream of the pre-tuned learnable layer 710 (e.g., a recursive learnable layer). We consider target data 12 obtained from the pre-tuned learnable layer 710 and applied to tunable learnable layers 71 to 73 (which in turn are applied to style component 77). Blocks 71 to 73 and 77 may be represented by a generator network layer 770. The generator network layer 770 may include multiple learnable layers (e.g., multiple blocks 50a to 50h, see below).

[0254] Figure 12 (and its in) Figure 14 The embodiments shown in the text illustrate an audio decoder (generator) 10 (10d), such as examples 10b and 10c, which can decode (e.g., generate, synthesize) an audio signal (output signal) 16 from an encoded signal 3, for example, according to the technology of the present invention (also known as StyleMelGAN). The output audio signal 16 can be generated based on the input signal 14, which may be noise, such as white noise (“first option”), or may be obtained from another source. As explained above, the target data 12 may include (e.g., a spectrogram) (e.g., a Mel spectrogram), which provides, for example, a time sample sequence to a Mel scale (e.g., obtained from a pre-tuned learnable layer 710). The target data 12 and / or the first data 15 are typically processed to obtain speech that can be recognized as natural by a human listener. In the decoder 10d, the first data 15 obtained from the input is styled (e.g., at block 77) to have vectors with acoustic features tuned by the target data 12. Finally, the output audio signal 16 will be recognized as speech by human listeners. For example, in Figure 14 In this context, input vector 14 and / or first data 15 (e.g., noise, obtained from an internal or external source) can be a 128×1 vector (a single sample, such as a time-domain sample or a frequency-domain sample, and 128 channels). Figure 14The input signal 14 to be provided to channel mapping 30 is shown; the first data provider 702 is not shown or is considered to be the same as channel mapping 30. In other examples, input vectors 14 of different lengths may be used. The input vector 14 may be processed in the first processing block 40 (e.g., under the conditioning of the target data 12 obtained from the encoded signal 3 via the pre-conditioning layer 710). The first processing block 40 may include at least one, such as multiple processing blocks 50 (e.g., 50a…50h). Figure 14 In this example, eight blocks 50a…50h are shown (each of which is also labeled “TADE Residual Blocks”), but in other examples, a different number may be chosen. In many examples, processing blocks 50a, 50b, etc., provide progressive upsampling of the signal evolving from the input signal 14 to the final audio signal 16 (e.g., at least some processing blocks, such as 50a, 50b, 50c, 50d, 50e, increase the sampling rate such that each of them increases the sampling rate (also referred to as bit rate) in the output relative to the sampling rate in its input), while some other processing blocks (e.g., 50f to 50h) (e.g., downstream of the blocks that increase the sampling rate (e.g., 50a, 50b, 50c, 50d, 50e)) do not increase the sampling rate (or bit rate). Blocks 50a to 50h can be understood as forming a single block 40 (e.g., Figure 12 (as shown in the block). In the first processing block 40, a set of adjustable layers (e.g., 71, 72, 73, but different numbers are possible) can be used to process the target data 12 and the input signal 14 (e.g., the first data 15). Thus, adjustable feature parameters 74, 75 (also referred to as gamma γ and beta β) can be obtained during training, for example, through convolution. Thus, the learnable layers 71 to 73 can be part of the weight layers of the learning network. As explained above, the first processing blocks 40, 50 may include at least one stylization component 77 (normalization block 77). At least one stylization component 77 can output first output data 69 (when multiple processing blocks 50 exist, multiple stylization components 77 can produce multiple components, which can be added to each other to obtain the final version of the first output data 69). At least one stylization component 77 can apply adjustable feature parameters 74, 75 to the input signal 14 (potential) or the first data 15 obtained from the input signal 14.

[0255] The first output data 69 may have multiple channels. The generated audio signal 16 may have a single channel.

[0256] Decoder 10 (10d) may include a second processing block 45 (in Figure 14The diagram shows blocks 42, 44, and 46. The second processing block 45 can be configured to combine multiple channels of the first output data 69 (input as second input data or second data). Figure 14 (Indicated by 47), so that in a single channel but in the sample sequence (in Figure 14 In the middle, the output audio signal 16 is obtained by using 49 as the indicator sample.

[0257] The term "channel" should not be understood in the context of stereo, but rather in the context of neural networks (e.g., convolutional neural networks) or more generally, in the context of learnable units. For example, the input signal (e.g., latent noise) 14 may be in 128 channels (in its representation in the time domain), because a channel sequence is provided. For example, when the signal has 40 samples and 64 channels, it can be understood as a matrix of 40 columns and 64 rows, and when the signal has 20 samples and 64 channels, it can be understood as a matrix of 20 columns and 64 rows (other representations are possible). Thus, the resulting audio signal 16 can be understood as a mono signal. In the case of generating a stereo signal, the disclosed technique is simply repeated for each stereo channel to obtain multiple audio signals 16 that are subsequently mixed.

[0258] At least the original input audio signal 1 and / or the generated speech 16 can be a time-domain value sequence. Conversely, the outputs of each (or at least one) of blocks 30 and 50a to 50h, 42, 44 can typically have different dimensions. In at least some of blocks 30 and 50a to 50e, 42, 44, upsampling can be performed on the signals (14, 15, 59, 69) that evolve from input 14 (e.g., noise or LPC parameters, or other parameters obtained from the encoded signal) toward speech 16. For example, double upsampling can be performed at the first block 50a in blocks 50a to 50h. Examples of upsampling can include sequences such as: 1) repetition of the same value; 2) insertion of zeros; 3) another repetition or insertion of zeros + linear filtering; and so on.

[0259] The resulting audio signal 16 is typically a single-channel signal. In cases where multiple audio channels are required (e.g., for stereo playback), the proposed procedure can, in principle, be repeated multiple times.

[0260] Similarly, the target data 12 may also have multiple channels generated by the pre-tuned learnable layer 710 (e.g., in a spectrogram such as a Mel spectrogram). In some examples, the target data 12 may be upsampled (e.g., according to a factor of 2, a power of 2, a multiple of 2, or a value greater than 2, such as according to different factors, such as 2.5 or multiples thereof) to adapt the dimensions of the signal (59a, 15, 69) that evolves along subsequent layers (50a to 50h, 42), for example to obtain tunable feature parameters 74, 75 that adapt the dimensions to the signal dimensions.

[0261] If the first processing block 40 is instantiated as multiple blocks (e.g., 50a to 50h), the number of channels may be maintained, for example, in at least some of the multiple blocks (e.g., from 50e to 50h, and in block 42, the number of channels remains unchanged). The first data 15 may have a first dimension or at least one dimension lower than the dimension of the audio signal 16. The first data 15 may have a total number of samples in all dimensions lower than the audio signal 16. The first data 15 may have one dimension lower than the audio signal 16 but have a greater number of channels than the audio signal 16.

[0262] As explained by the phrase "tuning set of learnable layers," the audio decoder 10 (10d) can be obtained according to the paradigm of conditional neural networks, such as based on conditional information. For example, the conditional information may consist of target data (or its upsampled version) 12, from which the tuning set of layers 71 to 73 (weight layers) is trained and tuning feature parameters 74, 75 are obtained. Therefore, the stylized component 77 is tuned by learnable layers 71 to 73. This can also be applied to the preconditioning layer 710.

[0263] Examples at encoder 2 (or the first learnable layer 20 on the coding side) and / or decoder 10 (10d) can be based on a convolutional neural network. For example, a small matrix (e.g., a filter or kernel) of 3×3 (or 4×4, or 1×1, or less than 10×10, etc.) can be convolved (rotated) along a larger matrix (e.g., channel × sample latent or input signal and / or spectrogram and / or upsampled spectrogram, or more generally, target data 12), meaning, for example, a combination (e.g., multiplication and sum of products; dot product, etc.) between elements of the filter (kernel) and elements of the larger matrix (activation map or activation signal). During training, elements of the filter (kernel) are obtained (e.g., learned) that minimize the loss. During inference, elements of the filter (kernel) obtained during training are used. Examples of convolution can be used at at least one of blocks 71 to 73, 61b, 62b (see below), 230, 250, 290, 429, 440, and 460. Notably, matrices can be used alternatively. In the case of conditional convolution, the convolution need not be applied to the signal evolving from the input signal 14 toward the audio signal 16 through intermediate signals 59a(15), 69, etc., but can be applied to the target signal 14 (e.g., for generating adjustment feature parameters 74 and 75, which are subsequently applied to the first data 15 or the latent or previous signal, or the signal evolving from the input signal toward the speech 16). In other cases (e.g., at blocks 61b and 62b, see below), the convolution can be unconditional and can be applied, for example, directly to the signals 59a(15), 69, etc., evolving from the input signal 14 toward the audio signal 16. Both conditional and unconditional convolution can be performed.

[0264] In some examples (at the decoder or encoder), the activation function downstream of the convolution can be different depending on the desired effect (ReLU, TanH, softmax, etc.). ReLU maps to the maximum value between 0 and the value obtained during convolution (in effect, it maintains the same value if it is positive and outputs 0 if it is negative). Leaky ReLU outputs x if x > 0, and 0.1*x if x ≤ 0, where x is the value obtained through convolution (instead of 0.1, in some examples, another value can be used, such as a predetermined value within 0.1 ± 0.05). TanH (which can be implemented, for example, at blocks 63a and / or 63b) provides the hyperbolic tangent of the value obtained during convolution, e.g., TanH(x) = (e^(-1 / 2)). x -e -x ) / (e x +e -x), where x is the value obtained during convolution (e.g., at block 61b, see below). Softmax (e.g., applied at, for example, block 64b) applies an exponent to each element of the convolution result and normalizes it by dividing by the sum of the exponents. Softmax provides a probability distribution of entries in the matrix produced by the convolution (e.g., at 62b). After applying the activation function, a pooling step (not shown in the figure) may be performed in some examples, but in others, this step may be avoided. It is also possible to have a softmax-gated TanH function, for example, by multiplying the result of the TanH function (e.g., obtained at 63b, see below) with the result of the softmax function (e.g., obtained at 64b) (e.g., at 65b, see below). In some examples, multiple convolutional layers (e.g., a tunable set of learnable layers or at least one tunable learnable layer) may be one downstream of another and / or parallel to each other to improve efficiency. If an activation function and / or pooling is provided, it can also be repeated in different layers (or, for example, different activation functions can be applied to different layers) (this can also be applied to the encoder).

[0265] At decoder 10 (10d), the input signal 14 is processed at different steps to become the generated audio signal 16 (e.g., under conditions set by the set of adjustable layers or learnable layers 71 to 73, and based on parameters 74, 75 learned by the set of adjustable layers or learnable layers 71 to 73). Therefore, the input signal 14 (or its evolved form, i.e., the first data 15) can be understood as being in the processing direction (in... Figure 4 The direction from 14 to 16 in Figure 7 changes to the resulting audio signal 16 (e.g., speech). The conditions will be substantially based on the target signal 12 and / or based on the preconditions in the encoded signal 3 and generated based on training (in order to obtain the optimal set of parameters 74, 75).

[0266] It should also be noted that multiple channels of input signal 14 (or any of its evolutions) can be viewed as a set of learnable layers and associated stylized components 77. For example, each row of matrices 74 and 75 can be associated with a specific channel of the input signal (or any of its evolutions), the specific channel being obtained, for example, from a specific learnable layer associated with that specific channel. Similarly, stylized components 77 can be viewed as being formed by multiple stylized components (each stylized component for each row of input signals x, c, 12, 76, 76', 59, 59a, 59b, etc.).

[0267] Figure 14 An example of an audio decoder 10 (10d) is shown. Figure 14 Pre-tuned learnable layer 710 is now shown (shown in...) Figure 12The target data 12 is obtained from the encoded signal 3 through a pre-conditioning layer 710 (see above). The target data 12 may be a Mel spectrogram obtained from the pre-conditioning learnable layer 710; the input signal 14 may be a signal obtained from an internal or external source, and the output 16 may be speech. The input signal 14 may have only one sample and multiple channels (indicated by "x", because it is variable, for example, the number of channels may be 80 or others). The input vector 14 may be obtained as a vector with 128 channels (but other numbers are possible). In the case that the input signal 14 is noise ("first option"), it may have a zero-mean normal distribution and follow the formula It can be 128-dimensional random noise with a mean of 0 and a correlation matrix (squared 128 × 128) equal to the identity matrix I (different choices can be made). Therefore, in the example where the noise is used as input signal 14, it can be completely decorrelated between channels and has a variance of 1 (energy). This can be implemented at every 22,528 generated samples (or other numbers for different examples); therefore, the dimension can be 1 on the time axis and 128 on the channel axis. In the example, the input signal 14 can be a constant value.

[0268] The input vector 14 can be processed step by step (e.g., at blocks 702, 50a to 50h, 42, 44, 46, etc.) to evolve into speech 16 (the evolved signal will be indicated, for example, by different signals 15, 59a, x, c, 76', 79, 79a, 59b, 79b, 69, etc.).

[0269] At block 30, channel mapping can be performed. This can consist of simple convolutional layers or include simple convolutional layers to change the number of channels, for example, from 128 to 64 in this case. Therefore, block 30 can be learnable (in some examples, it can be deterministic). As can be seen, at least some of the processing blocks 50a, 50b, 50c, 50d, 50e, 50f, 50g, and 50h (together embodying the first processing block 50 of FIG. 6) can increase the number of samples, for example, by performing upsampling (e.g., up to 2x upsampling) on ​​each frame. The number of channels can remain the same (e.g., 64) along blocks 50a, 50b, 50c, 50d, 50e, 50f, 50g, and 50h. The number of samples can be, for example, the number of samples per second (or other time units): we can obtain sound at 16 kHz or higher (e.g., 22 kHz) at the output of block 50h. As explained above, a sequence of multiple samples can constitute a frame. Each of blocks 50a to 50h (50) can also be a TADE residual block (a residual block in the context of temporally adaptive denormalized TADE). Notably, each block 50a to 50h (50) can be modulated by target data (e.g., code) 12 and / or by encoded signal 3. At the second processing block 45 (Figures 1 and 6), multiple samples can be obtained in a single channel and in a single dimension (see also...). Figure 13 As can be seen, another TADE residual block 42 (further to blocks 50a to 50h) can be used (which reduces the dimension to four single channels). Next, a convolutional layer 44 and an activation function (which can be, for example, TanH 46) can be executed. A group 110 of (pseudo-orthogonal mirror filters) can also be applied to obtain the final signal 16 (which can be stored, rendered, etc.).

[0270] At least one of blocks 50a to 50h (or each of them in a particular example) and 42, as well as encoder layers 230, 240, and 250 (and 430, 440, 450, and 460), may be, for example, residual blocks. The residual learnable blocks (layers) can predict the residual components of the signal that evolves from the input signal 14 (e.g., noise) to the output audio signal 16. The residual signal is only a portion (residual component) of the main signal that evolves from the input signal 14 toward the output signal 16. For example, multiple residual signals can be added to each other to obtain the final output audio signal 16. However, other architectures may be used.

[0271] Figure 15An example of one of blocks 50a through 50h (50) is shown. Blocks 50a through 50h (50) can be copied from each other, but this may be the case when they are trained. As can be seen, each block 50 (50a through 50h) is input with first data 59a, which is first data 15 (or its upsampled version, such as the version output by upsampled block 30) or the output from a previous block. For example, block 50b may be input with the output of block 50a; block 50c may be input with the output of block 50b, and so on. In the example, different blocks can be operated on in parallel with each other, and the results are added together. Figure 15 As can be seen, the first data 59a provided to block 50 (50a to 50h) or 42 is processed and its output is output data 69 (which will be provided as input to subsequent blocks). As indicated by line 59a', the major component of the first data 59a actually bypasses most of the processing of the first processing blocks 50a to 50h (50). For example, the major component 59a' bypasses blocks 60a, 900, 60b and 902 and 65b. The residual component 59a of the first data 59 (15) can be processed to obtain the adder 65c (which is in Figure 15 The residual portion 65b' is added to the main component 59a' at (indicated in the text, but not shown). The bypass of the main component 59a' and the addition at adder 65c can be understood as instantiating the fact that each block 50 (50a to 50h) processes the operation on the residual signal, which is then added to the main portion of the signal. Therefore, each of the blocks 50a to 50h can be considered a residual block. The addition at adder 65c does not necessarily need to be performed within residual blocks 50 (50a to 50h). Multiple residual signals 65b' (each through the output of each of the residual blocks 50a to 50h) can be added individually (e.g., at a single adder block in, for example, a second processing block 45). Therefore, different residual blocks 50a to 50h can operate in parallel with each other. Figure 15In the example, each block 50 (50a to 50h) may repeat its convolutional layer twice. A first denormalization block 60a and a second denormalization block 60b may be used in a cascaded manner. The first denormalization block 60a may include an instance of a style component 77 to apply adjusted feature parameters 74 and 75 to the first data 59 (15) (or its residual version 59a). The first denormalization block 60a may include a normalization block 76. The normalization block 76 may perform normalization along the channels of the first data 59 (15) (e.g., its residual version 59a). Thus, a normalized version c (76') of the first data 59 (15) (or its residual version 59a) is obtained. Therefore, the style component 77 may be applied to the normalized version c (76') to obtain a denormalized (adjusted) version of the first data 59 (15) (or its residual version 59a). The denormalization at component 77 can be obtained, for example, by element-wise multiplication of the elements of matrix γ (which embodies condition 74) with signal 76' (or another signal version between the input signal and speech) and / or by element-wise addition of the elements of matrix β (which embodies condition 75) with signal 76' (or another signal version between the input signal and speech). Thus, a denormalized version 59b of the first data 59 (15) (or its residual version 59a) can be obtained (adjusted by adjusting feature parameters 74 and 75).

[0272] Next, a gated activation 900 can be performed on the denormalized version 59b of the first data 59 (e.g., its residual version 59a). Specifically, two convolutions 61b and 62b can be performed (e.g., each with a 3×3 kernel and an expansion factor of 1). Different activation functions 63b and 64b can be applied to the results of convolutions 61b and 62b, respectively. Activation 63b can be TanH. Activation 64b can be softmax. The outputs of the two activations 63b and 64b can be multiplied together to obtain a gated version 59c of the denormalized version 59b of the first data 59 (or its residual version 59a). Subsequently, a second denormalization 60b can be performed on the gated version 59c of the denormalized version 59b of the first data 59 (or its residual version 59a). The second denormalization 60b can be similar to the first denormalization and is therefore not described here. Subsequently, a second activation 902 can be performed. Here, the kernel can be 3×3, but the expansion factor can be 2. In any case, the expansion factor of the second gate activation 902 can be greater than the expansion factor of the first gate activation 900. The set of conditioning of learnable layers 71 to 73 (e.g., obtained from pre-conditioned learnable layers) and the stylization component 77 can be applied to signal 59a (e.g., applied twice for each block 50a, 50b...). Upsampling of target data 12 can be performed at upsampling block 70 to obtain upsampled version 12' of target data 12. Upsampling can be obtained by nonlinear interpolation and can use, for example, a factor of 2, a power of 2, a multiple of 2, or another value greater than 2. Thus, in some examples, the spectrogram (e.g., a Mel spectrogram) 12' can be made to have the same (e.g., conformal) dimensions as the signal (76, 76', c, 59, 59a, 59b, etc.) to be conditioned by the spectrogram. In the example, the first and second convolutions at 61b and 62b, respectively downstream of TADE blocks 60a or 60b, can be performed at the same number of elements in the kernel (e.g., 9, e.g., 3×3). However, the second convolution in block 902 can have a dilation factor of 2. In the example, the maximum dilation factor used for the convolution can be 2 (ii).

[0273] As explained above, the target data 12 may be upsampled, for example, to conform to the input signal (or signals derived from it, such as 59, 59a, 76', also referred to as the latent signal or activation signal). Here, convolutions 71, 72, 73 (intermediate values ​​of the target data 12 are indicated by 71') may be performed to obtain parameters γ (gamma 74) and β (beta 75). The convolution at any of 71, 72, 73 may also require a Rectified Linear Unit (ReLU) or a Leaky ReLU. Parameters γ and β may have the same dimensions as the activation signal (the signal processed to evolve from the input signal 14 into the resulting audio signal 16, which, when normalized, is represented here as x, 59, 59a, or 76'). Therefore, when the activation signal (x, 59, 59a, 76') has two dimensions, γ and β (74 and 75) also have two dimensions, and each of them can be superimposed on the activation signal (the length and width of γ and β can be the same as the length and width of the activation signal). At style component 77, the adjustment feature parameters 74 and 75 are applied to the activation signal (which can be the first data 59a or 59b output by multiplier 65a). However, it should be noted that the activation signal 76' can be a normalized version of the first data 59, 59a, 59b (15) (at instance normalization block 76), with normalization performed on the channel dimension. It should also be noted that the formula (γ*c+β, in style component 77) shown in style component 77 Figure 15 Also used in China The indications () can be element-wise multiplications, and in some examples, not convolutional multiplications or dot products. Convolutions 72 and 73 may not have activation functions downstream of them. Parameter γ (74) can be understood as having a variance value and β (75) can be understood as having a bias value. It should be noted that for each block 50a to 50h, 42, learnable layers 71 to 73 (e.g., along with the similarly formatted component 77) can be understood as weighted layers. Also, Figure 14 Block 42 can be instantiated as Figure 15 Block 50. Next, for example, convolutional layer 44 reduces the number of channels to 1, and then TanH 46 is performed to obtain speech 16. The output 44' of blocks 44 and 46 may have a reduced number of channels (e.g., 4 channels instead of 64), and / or may have the same number of channels as the previous blocks 50 or 42 (e.g., 40).

[0274] Pseudo-orthogonal mirror filter (PQMF) synthesis (see also below) 110 can be performed on signal 44' to obtain audio signal 16, for example, in one channel (other techniques may be used).

[0275] In the example, the encoded signal 3 may be transmitted (e.g., via a communication medium, such as a wired connection and / or a wireless connection) and / or stored (e.g., in a storage unit). Therefore, the encoder 3 and / or the first encoding-side learnable layer 20 may include and / or be connected to and / or configured to control transmission units (e.g., modems, transceivers, etc.) and / or storage units (e.g., mass memory, etc.). To permit storage and / or transmission, other means for processing the encoded signal for storage and / or transmission and for reading and / or receiving purposes may exist between the quantization module 300 and the dequantization module 500 (513).

[0276] Bias towards scalar quantization in the context of traditional neural audio coding:

[0277] ●Reference

[10] , a groundbreaking paper on discrete representation learning, argues that VQ achieves stronger compression than SQ, but this is not the case.

[0278] ●VQ-VAE

[12] : Advocates that soft-to-hard technology is unrealizable.

[0279] ● Minje Kim’s paper

[13] also uses soft to hard, but with scalar quantization; however, in our experiments, soft to hard training never worked well.

[0280] The training technique requires a trainable codebook, but this is not mandatory for our scalar quantization technique.

[0281] ○ The target bit rate is much higher than our target: 12, 20, 32 kbps.

[0282] ● Traditional competitor Soundstream

[14] : The authors argue that VQ is the most commonly used technology for neural audio codecs, which is also our understanding.

[0283] ● Similarly, a recent journal article explicitly considers quantization techniques for neural audio coding

[15] , mentioning only VQ-related methods in its review while omitting SQ. Therefore, this paper implicitly argues that VQ is more suitable for neural audio coding.

[0284] ● Scalar quantization has been successfully used for neural image coding at high bit rates (e.g.,

[16] ), but has not yet been successfully used for neural audio coding at low bit rates (below 4 kbps).

[0285] ● Competitor codecs

[17] : VQ is compared to a simpler version of SQ and VQ is claimed to be superior to SQ in preliminary tests. The authors did not follow up on SQ and did not even provide their results.

[0286] Other examples

[0287] Typically, an example may be implemented as a computer program product having program instructions that, when run on a computer, are operationally used to perform one of the methods described herein. The program instructions may, for example, be stored on a machine-readable medium. Other examples include a computer program stored on a machine-readable medium for performing one of the methods described herein. In other words, an example of a method is therefore a computer program having program instructions for performing one of the methods described herein when the computer program is run on a computer. Another example of a method is therefore a data carrier medium (or digital storage medium, or computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein. The data carrier medium, digital storage medium, or recording medium is tangible and / or non-transitory, rather than intangible and transient signals. Thus, another example of a method is a stream of data or a sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be transmitted, for example, via a data communication connection, such as via the Internet. Another example includes a processing component, such as a computer or programmable logic device, that performs one of the methods described herein. Another example includes a computer having a computer program mounted thereon for performing one of the methods described herein. Another example includes an apparatus or system for transmitting (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, mobile device, storage device, or the like. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver. In some examples, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some examples, a field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods can be performed by any suitable hardware device. The above examples are merely illustrative of the principles discussed above. It should be understood that modifications and variations to the configurations and details described herein will be readily apparent. Therefore, it is intended to be limited by the scope of the claims, and not by the specific details presented through the description and interpretation of the examples herein. Even if equivalent or equivalent components or components having equivalent or equivalent functions appear in different figures, they are indicated in the following description by equivalent or equivalent reference numerals.

Claims

1. A decoder (10) configured to generate an audio signal from an encoded signal (3) representing an audio signal (16), the decoder (10) comprising: The encoded signal reader (560) is configured to read the encoded signal (3), thereby providing multiple indices (556); Scalar dequantization module (500), including: Multiple quantization index converters (355), each configured to convert an index (556) of the multiple indices into a corresponding latent scalar value (551), such that the multiple latent scalar values ​​(551) form a first latent audio signal representation (550) of the audio signal; and A first learnable segment (540) is used to provide a second potential representation (530) from a first potential audio signal representation (550); as well as The second learnable segment (520) includes at least one learnable layer and is configured to generate an audio signal (16) from a second potential audio signal representation (530).

2. The decoder of claim 1, wherein each quantization index converter is configured to provide a single potential scalar value using at least one codebook, the at least one codebook being a quantization index converter-specific codebook.

3. The decoder of claim 1, wherein all quantization index converters or at least a subset thereof are configured to provide corresponding multiple scalar values ​​using at least one codebook, wherein the at least one codebook is a common codebook.

4. The decoder as described in any one of claims 2-3, wherein at least one quantization index converter is a residual or multi-level quantization index converter.

5. The decoder as described in any one of claims 2-4, wherein at least one codebook is learnable.

6. The decoder as described in any one of claims 2-5, wherein at least one codebook is deterministic.

7. The decoder as described in any one of claims 2-6, wherein at least one codebook has a fixed-length representation in the bitstream.

8. The decoder of any one of claims 2-7, wherein at least one codebook has a variable-length representation in the bitstream.

9. The decoder of claim 8, configured such that for at least two or all of the potential scalar values, the more frequent potential scalar values ​​are converted from an index having a more compact representation in the encoded signal than the index mapped to the less frequent scalar values.

10. The decoder as claimed in any of the preceding claims, wherein the encoded signal reader (560) is configured to entropy decode the index in the encoded signal (3).

11. The decoder of any one of claims 2-10, wherein at least one codebook or quantization is non-uniform, wherein the range of values ​​to be quantized is divided into unequal intervals such that more frequent intervals are smaller than less frequent intervals.

12. The decoder as claimed in any of the preceding claims is configured to select between at least one first decoding mode and a second decoding mode, wherein the first decoding mode is a first low quantization index converter number decoding mode and the second decoding mode is a second high quantization index converter number decoding mode, wherein in the first decoding mode, the decoder is configured to receive fewer potential scalar values ​​in the first decoding mode than in the second decoding mode, thereby the decoder uses fewer quantization index converters in the first decoding mode than in the second decoding mode.

13. The decoder of claim 12, configured to use at least one single quantization index converter in a first low quantization index converter number decoding mode to quantize a plurality of scalar values ​​of a first potential representation into a single index.

14. The decoder of claim 13, wherein at least one single quantization index converter is configured in a first low quantization index converter number decoding mode to convert a codebook into a plurality of scalar values, the codebook having a reduced resolution, bit length, and / or index number relative to the resolution used by the respective quantization index converter in a second high quantization index converter number decoding mode.

15. The decoder of any one of claims 12-14, configured to select between at least one first decoding mode and a second decoding mode, wherein the first decoding mode is a first low index number decoding mode and the second decoding mode is a second high index number decoding mode, and configured to use at least one codebook in the second high index number decoding mode, the at least one codebook having a higher index number, higher resolution and / or higher bit length than in the first low index number decoding mode.

16. The decoder as claimed in any of the preceding claims is configured to select between at least one first decoding mode and a second decoding mode, such that: In the second decoding mode, multiple quantization index converters are used to provide multiple scalar values, each quantization index converter being configured to provide a single scalar value or a component thereof from a corresponding index among the multiple indices; and In the first decoding mode, a vector quantization index converter is used to provide multiple scalar values ​​from a single index among multiple indices.

17. The decoder as claimed in any of the preceding claims is configured to select between at least one first decoding mode and a second decoding mode, such that: In the first decoding mode, multiple parallel quantization index converters are used to provide different scalar values ​​from multiple indices, each quantization index converter being configured to provide a single scalar value from a corresponding single index among the multiple indices according to the basic codebook; and In the second decoding mode, at least one quantization index converter is used to provide at least one scalar value from at least two indices: a first index that approximates at least one scalar value using a basic codebook; and at least one second index that approximates at least one residual scalar value using at least one residual codebook.

18. The decoder as claimed in any of the preceding claims, configured to select between at least one first decoding mode and a second decoding mode, wherein the second decoding mode is a multi-level mode having a second number of levels and the first decoding mode is a single-level mode or a multi-level mode having a first number of levels, the first number being less than the second number, such that: In the first decoding mode, multiple quantization index converters are used to provide multiple indices, each quantization index converter being configured to convert a single index into a single scalar value or to convert a first number of multiple indices into a scalar value; and In the second decoding mode, at least one quantization index converter is used to convert a second number of indices to provide at least one scalar value.

19. The decoder as claimed in any of the preceding claims is configured to select between at least one first classification decoding mode and a second classification decoding mode based on the classification (360) of the input audio signal (1), wherein the first classification decoding mode is trained for a first category of the classification and the second classification decoding mode is trained for a second category of the classification.

20. The decoder as claimed in any one of claims 12-19, configured to select between at least one first decoding mode and a second decoding mode based on manual selection.

21. The decoder as claimed in any one of claims 12-20, configured to select between at least one first decoding mode and a second decoding mode based on a request from an application.

22. The decoder as claimed in any one of claims 12-21 is configured to select between at least one first decoding mode and a second decoding mode based on the signaling written in the encoded signal (3).

23. The decoder as claimed in any of the preceding claims, wherein the second learnable segment (540) is configured to change the dimension of the latent representation from the first latent representation (530) to the second latent representation (550).

24. The decoder as described in any of the preceding claims, comprising: The first data provider (702) is configured to provide first data (15) derived from the input signal (14); The first processing block (40, 50, 50a-50h) is configured to receive first data (15) and output first output data (69) in a given frame. The decoder further includes: At least one adjustable learnable layer (71, 72, 73) is configured to process target data (12) from a second latent representation to output adjustable feature parameters (74, 75); and The stylization component (77) is configured to apply the adjustment feature parameters (74, 75) to the first data (15, 59a) or the normalized first data (59, 76').

25. The decoder of claim 24 is configured to obtain the input signal from noise (14).

26. The decoder of any one of claims 23-25, further comprising at least one pre-adjustable learnable layer (710), the at least one pre-adjustable learnable layer (710) being configured to receive a second latent representation and output target data (12) representing an audio signal.

27. The decoder of claim 26, wherein at least one pre-tuned learnable layer (710) is configured to provide the target data (12) as a spectrogram or a decoded spectrogram.

28. The decoder of any one of claims 23-27, wherein the first convolutional layer (71-73) is configured to convolve the target data (12) or the upsampled target data using a first activation function to obtain the first convolutional data (71').

29. The decoder of any one of claims 24-28, further comprising a normalization component (76) configured to normalize the first data (59a, 15).

30. The decoder of any one of claims 24-29, wherein the target data (12) comprises a spectrogram.

31. The decoder as described in any of the preceding claims is configured to render the generated audio signal (16).

32. The decoder as claimed in any of the preceding claims is configured to encode the generated audio signal (16) into a second encoded representation.

33. The decoder as claimed in any of the preceding claims, wherein the second learnable segment is pre-trained relative to the first learnable segment.

34. An encoder (2) for generating an encoded signal (3), wherein an input audio signal (1) is encoded in the encoded signal (3), the encoder (2) comprising: The first learnable segment (20) includes at least one learnable layer to provide a first latent representation (330) of the input audio signal (1). A scalar quantization module (300), used for quantizing the first latent representation (330), includes: A second learnable segment (340) is used to provide multiple latent scalar values ​​(351) to be quantized from a first latent representation (330); and Multiple quantizers (355) are configured to provide multiple indices (356), each quantizer (355) being configured to quantize a single latent scalar value (351) to be quantized and to provide an index (356) among multiple indices from a single latent scalar value (351); and The encoded signal writer (360) is configured to write multiple indices (356) into the encoded signal (3).

35. The encoder of claim 34, wherein each quantizer (355a) or at least one quantizer is configured to quantize the corresponding potential scalar value (351) using at least one codebook (357a), wherein the at least one codebook is a quantizer-specific codebook.

36. The encoder of claim 34 or 35, wherein all of the plurality of quantizers (355, 355b, 355c) or at least a subset thereof are configured to quantize the respective latent scalar values ​​using at least one codebook (357), wherein the at least one codebook is a common codebook.

37. The encoder of any one of claims 34-36, wherein at least one quantizer is a residual or multi-level quantizer.

38. The encoder of any one of claims 34-37, wherein at least one codebook (357, 357a, 357b) is learnable.

39. The encoder of any one of claims 34-38, wherein at least one codebook (357, 357a, 357b) is deterministic.

40. The encoder of any one of claims 34-36, wherein at least one codebook (357, 357a, 357b) has a fixed-length bitstream representation.

41. The encoder of any one of claims 34-40, wherein at least one codebook (357, 357a, 357b) has a variable-length bitstream representation.

42. The encoder of claim 41, configured such that for at least two of the potential scalar values, or for a plurality of potential scalar values, or for all potential scalar values, the more frequent potential scalar values ​​are mapped to an index in the encoded signal representation that is more closely packed than the index mapped by the less frequent scalar values.

43. The encoder of any one of claims 34-42, wherein the encoded signal writer (360) is configured to entropy encode at least one index provided by the codebook (357, 357a, 357b).

44. The encoder of any one of claims 34-43, wherein at least one codebook or quantization is non-uniform, wherein the range of values ​​to be quantized is divided into unequal intervals such that more frequent intervals are smaller than less frequent intervals.

45. The encoder of any one of claims 34-44, configured to select between at least a first encoding mode and a second encoding mode, wherein the first encoding mode is a first low quantizer number encoding mode (341) and the second encoding mode is a second high quantizer number encoding mode (342), the encoder being configured to provide a plurality of indices in the first low quantizer number encoding mode (341) having fewer indices than in the second high quantizer number encoding mode (342), the encoder (2) thereby using fewer quantizers (355, 355s) in the first low quantizer number encoding mode (341) than in the second high quantizer number encoding mode (342).

46. ​​The encoder of claim 45, wherein the scalar quantization module (300) is configured to use at least one vector quantizer (355S, VQ) in a first low quantizer number encoding mode (341) to quantize a vector formed by a plurality of scalar values ​​(351S', 351S") of the second latent representation (350) to a single index (356S), and in a second high quantizer number encoding mode (342) to quantize the plurality of scalar values ​​(351S', 351S") of the second latent representation (350) to two different indices (356S', 356S") by two quantizers (355) of a plurality of quantizers.

47. The encoder of claim 46, wherein at least one vector quantizer (355S, VQ) quantizing a plurality of scalar values ​​(351S', 351S") is configured in a first low quantizer number encoding mode (341) to quantize the plurality of scalar values ​​(351S) using a codebook, the codebook having a reduced resolution, bit length and / or index number relative to the resolution used by the respective quantizer in a second high quantizer number encoding mode (342).

48. The encoder of any one of claims 34-47, configured to select between at least one first encoding mode and a second encoding mode, wherein the second encoding mode is a second high index number encoding mode (342) and the first encoding mode is a first low index number encoding mode (341), wherein the encoder is configured to use at least one codebook in the second high index number encoding mode (341), the at least one codebook having a higher index number, higher resolution and / or higher code length and / or more quantization levels and / or higher index bit length than in the first low index number encoding mode (341).

49. The encoder of any one of claims 34-48, configured to select between at least one first coding mode and a second coding mode, wherein the second coding mode is a second extended latent coding mode and the first coding mode is a first reduced latent coding mode, wherein the second learnable segment (340) is configured to provide more latent scalar values ​​in the second coding mode than in the first coding mode (341).

50. The encoder as claimed in any one of claims 34-49, configured to select between at least a first encoding mode (341) and a second encoding mode (342), such that: In the second encoding mode (342), multiple quantizers (355) are used to provide multiple indices (356), each quantizer (355) being configured to quantize a single latent scalar value (351) to provide one of the multiple indices (356); and In the first encoding mode (351), at least one quantizer (355S) is used to quantize multiple potential scalar values ​​(351S) onto a single index (356S).

51. The encoder of any one of claims 34-50, configured to select between at least one first encoding mode and a second encoding mode, wherein the second encoding mode is a multi-level encoding mode having a second number of levels and the first encoding mode is a single-level encoding mode or a multi-level encoding mode having a first number of levels, the first number being less than the second number, such that: In the first encoding mode, multiple quantizers are used to provide multiple indices, each quantizer (355) being configured to quantize a single scalar value according to the basic codebook (357) to provide one of the multiple indices (356), which is a single index; and In the second encoding mode, at least one quantizer is used to quantize at least one scalar value to provide at least two indices: a first index indicating at least one potential scalar value using a basic codebook; and at least one second index indicating at least one residual potential scalar value using at least one residual codebook.

52. The encoder as claimed in any one of claims 34-51, configured to perform in parallel a first encoding mode (341') ​​for providing a first encoded signal version (3') and a second encoding mode (341") for providing a second encoded signal version (3"), and to select between the first encoding mode and the second encoding mode by selecting an encoded signal version (3', 3") that minimizes distortion with respect to the input audio signal (1) in the first encoded signal version and the second encoded signal version (3', 3") written in the encoded signal (3).

53. The encoder of claim 52 is configured to determine a first distortion metric of a first re-decoded version (1') of an input audio signal (1) obtained by a first encoding mode (341') ​​and a second distortion metric of a second re-decoded version (1") of an input audio signal (1) obtained by a second encoding mode (341"), so as to perform a selection (353e) by selecting the encoded signal version (3', 3") with the lowest distortion metric between the first distortion metric and the second distortion metric in the encoded signal (3).

54. The encoder as claimed in any one of claims 34-53 is configured to execute in parallel a first encoding mode (341') ​​for providing a first encoded signal version (3') and a second encoding mode (341') ​​for providing a second encoded signal version (3"), and to select between the first encoding mode and the second encoding mode by selecting the encoded signal version (3', 3") that maximizes processing efficiency among the first encoded signal version and the second encoded signal version (3', 3") written into the encoded signal (3).

55. The encoder as claimed in any one of claims 34-54, configured to select between a first encoding mode and a second encoding mode based on manual selection.

56. The encoder as claimed in any one of claims 34-55, configured to select between a first encoding mode and a second encoding mode based on a request from an application.

57. The encoder as claimed in any one of claims 34-56, configured to select between a first encoding mode and a second encoding mode based on the state (359a, 359b) of the communication link through which the encoded signal (3) is transmitted, so as to: When the communication link performance is relatively high, a coding mode that provides higher resolution but higher bit length is selected; and When the performance of the communication link is relatively poor, an encoding mode that provides lower resolution but lower bit length is selected.

58. The encoder of any one of claims 34-57, configured to select between at least one first classification coding mode and a second classification coding mode based on a classification (360) of an input audio signal (1) or a processed version thereof, wherein the first classification coding mode is trained for a first category of the classification and the second classification coding mode is trained for a second category of the classification.

59. The encoder of claim 58, wherein the first category is an unvoiced category and the second category is a voiced category, wherein the first classification coding mode is an unvoiced guided mode and the second classification coding mode is a voiced guided mode.

60. The encoder of any one of claims 34-59, wherein the first learnable segment (340) is configured to reduce the dimension from the first latent representation (330) to the second latent representation (350).

61. The encoder of any one of claims 34-60, wherein the first learnable segment includes a format definer (210), the format definer (210) being configured to define a multidimensional audio signal representation (220) of the input audio signal, the multidimensional audio signal representation of the input audio signal including at least: The first dimension allows multiple consecutive frames to be ordered according to the first dimension; as well as The second dimension allows multiple samples from at least one frame to be ordered according to the second dimension to define multiple channels. The multidimensional audio signal represents at least one learnable layer input to the first learnable segment.

62. The encoder of any one of claims 34-61, wherein the first learnable segment is pre-trained relative to the second learnable segment.

63. A decoding method for generating an audio signal from an encoded signal (3) representing an audio signal (16), the method (10) comprising: Read the encoded signal to obtain multiple indices (556); Perform scalar dequantization (500), including: The conversion is performed by a plurality of quantization index converters (555), each quantization index converter (555) converting an index (556) of the plurality of indices into a corresponding latent scalar value (551), such that the plurality of latent scalar values ​​(551) form a first latent audio signal representation (550) of the audio signal; and A second potential audio signal representation (530) is provided from a first potential audio signal representation (550) via a first learnable segment (540); as well as An audio signal (16) is generated from a second potential audio signal representation (530) through a second learnable segment (520) including at least one learnable layer.

64. A method for generating an encoded signal (3), wherein an input audio signal (1) is encoded in the encoded signal (3), the method comprising: A first latent representation (330) of the input audio signal (1) is provided through a first learnable segment (20) including at least one learnable layer. The first latent representation (330) is quantized using the scalar quantization module (300) through the following operations: Through the second learnable segment (340), multiple latent scalar values ​​(351) to be quantized are obtained from the first latent representation (330); and Multiple indices (356) are obtained through multiple quantizers (355), each of the multiple quantizers (355) quantizes a single latent scalar value (351) and provides an index (356) among the multiple indices from a single latent scalar value (351); and Multiple indices (356) are written into the encoded signal (3).

65. A non-transitory storage unit storing instructions that, when executed by a computer, cause the computer to perform and / or control the method as described in claim 63 or 64.