Error resilience tools for audio encoding / decoding

The audio signal representation decoder and encoder system addresses the lack of effective error resilience in neural network-based voice codecs by performing packet loss concealment and forward error correction in the quantization domain, enhancing audio quality and reducing latency in VoIP communications.

JP2026502158APending Publication Date: 2026-01-21FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025536665
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-23
Filing Date
2023-12-14
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Existing neural network-based voice codecs lack effective error resilience mechanisms for packet loss and forward error correction, particularly at low bitrates, leading to complexity, latency, and suboptimal performance in real-time VoIP communications.

Method used

An audio signal representation decoder and encoder system that performs packet loss concealment and forward error correction in the quantization domain using learnable predictor layers and codebooks, allowing for efficient reconstruction of audio signals even with packet loss.

Benefits of technology

The system provides improved error resilience and reduced complexity by integrating packet loss concealment and forward error correction within the neural coding scheme, maintaining audio quality and reducing latency in VoIP communications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502158000011
    Figure 2026502158000011
  • Figure 2026502158000012
    Figure 2026502158000012
  • Figure 2026502158000013
    Figure 2026502158000013
Patent Text Reader

Abstract

In particular, examples of an audio signal representation encoder, an audio encoder, an audio signal representation decoder, and an audio decoder are provided, eg, using error resilience tools for trainable applications (eg, using neural networks). In one example, an audio signal representation decoder (1810, 1810a, 1810b) is provided that is configured to decode an audio signal representation (1820a, 1820b) from a bitstream (1830, 1630) that has been divided into a sequence of packets. a bitstream reader (1802a, 1892b) configured to sequentially read the sequence of packets (1830, 1630); a packet loss controller (1806a, 1806b) configured to check whether a current packet (1830, 1630) has been successfully received or should be considered lost; and a quantization index converter (1818a, 1818b) configured to convert at least one index (1804a, 1804b) extracted from the current packet (1830, 1630) into at least one current code (1820a, 1820b) from at least one codebook, when the packet loss controller (1806a, 1806b) determines that the current packet (1830, 1630) has been successfully received, thereby forming at least a portion of the audio signal representation (1820a, 1820b). The audio signal representation decoder (1810, 1810a, 1810b) is configured to generate at least one current code by prediction (1810a, 1810b) from at least one previous code or index via at least one learnable predictor layer when the packet loss controller (1806a, 1806b) determines that the current packet should be considered lost, thereby forming at least a part of the audio signal representation (1820a, 1820b).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] In particular, examples of audio signal representation encoders, audio encoders, audio signal representation decoders, and audio decoders are provided that use error resilience tools for trainable applications (e.g., using neural networks). In particular, error resilience tools for neural end-to-end speech codecs, such as forward error correction (FEC) and packet loss concealment (PLC), are described. [Background technology]

[0002] Error resilience tools such as packet loss concealment (PLC) and forward error correction (FEC) have been implemented for traditional voice codec systems. In applications such as VoIP, where frequent packet loss and delay are unavoidable, such tools play a critical role in maintaining quality of service to end users. In recent years, deep neural network (DNN)-based voice codecs have significantly evolved due to their ability to transmit voice signals at very low bitrates. The recently proposed neural end-to-end voice codec (NESC) efficiently encodes voice signals at low bitrates below 3.2 kbps and is robust to noise and reverberant voice signals (NESC is described in particular in Figures 9-13 and related discussions). Extending NESC's robustness to packet loss, we propose an autoregressive neural network to perform packet loss concealment along with low-bitrate forward error correction at additional bitrates, which can be as low as 0.8 kbps. Our method operates on the latent representation of NESC and is trained independently of the codec.

[0003] Real-time VoIP communications are highly susceptible to network conditions and congestion, resulting in packet loss or long delays in packet arrival. Decoders must be able to handle such losses and conceal lost packets to maintain good quality of service. Basic packet loss concealment (PLC) techniques have included methods such as silencing lost frames, repeating pitch lags, or some form of extrapolation. State-of-the-art communication codecs, such as Enhanced Voice Services (EVS), support two types of error resilience tools: packet loss concealment, which extrapolates coded parameters from previous frames, such as line spectral frequency (LSF), and pitch information for future frames transmitted for lost frames with additional transmitted information; and forward error correction (FEC), in which features from distant frames are coarsely quantized and piggybacked onto future frames (see [1], [2]). Transmitting redundant information in anticipation of losses must be done carefully, as it can place additional strain on network connections and introduce additional latency.

[0004] In recent years, neural network-based systems have shown unprecedented progress, outperforming traditional systems in various fields such as speech enhancement, speech coding, and speech synthesis. Similarly, DNN-based PLC models such as WaveNetEQ ([3]), PLAAE ([4]), LPCNet-based PLC ([5]), and ([6]) have shown superiority over traditional concealment methods for large bursts and higher error rates. While most of these methods perform concealment directly on the speech signal using post-processing techniques, a recently proposed LPCNet-based PLC model predicts the features of future frames and generates a concealment signal using an autoregressive LPCNet ([8]).

[0005] Limitations of post-processing (DNN-based) PLC: · The coding-agnostic PLC model requires tweaking and tuning to optimally support different codecs. · Good received frames may be affected by post-processing steps. · Complexity and latency overhead. · Joint training between the PLC and the encoding module is not possible, as the recovery capability is limited by the quality of the codec.

[0006] [References] [1]Anssi Ramo and Antti Kurittu and Henri Toukomaa, “EVS Channel Aware Mode Robustness to Frame ErasureEVS Robustness to Frame Erasures”.2553-2557.10.21437 / Interspeech.2016-917. [2] C. Rao and S. Zhao, “Multiple additional bit-rate channel-aware modes in EVS codec for packet loss recovery,” 2019 IEEE International Conference on Signal, Information and Data Processing (ICSIDP), 2019, pp. 1-5, doi:10.1109 / ICSIDP47821.2019.9173341. [3] F. Stimberg et al., “WaveNetEQ-Packet Loss Concealment with WaveRNN,” 2020 54th Asilomar Conference on Signals, Systems, and Computers, 2020, pp. 672-676, doi:10.1109 / IEEECONF51394.2020.9443419. [4]Pascual,Santiago,Joan Serra,and Jordi Pons.“Adversarial auto-encoding for packet loss concealment.”2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics(WASPAA).IEEE,2021. [5]Valin,J.M.,Mustafa,A.,Montgomery,C.,Terriberry,T.B.,Klingbeil,M.,Smaragdis,P.,&Krishnaswamy,A.(2022).Real-Time Packet Loss Concealment With Mixed Generative and Predictive Model.arXiv preprint arXiv:2205.05785. [6]Xue,Huaying,Xiulian Peng,Xue Jiang,and Yan Lu.“Towards Error-Resilient Neural Speech Coding.” arXiv preprint arXiv:2207.00993(2022). [7]Pia,N.,Gupta,K.,Korse,S.,Multrus,M.,&Fuchs,G.(2022).NESC:Robust Neural End-2-End Speech Coding with GANs.arXiv preprint arXiv:2207.03282. [8]J.Valin and J.Skoglund,A Real-Time Wideband Neural Vocoder at 1.6 kb / s Using LPCNet.arXiv:1903.12087. [9] J. Wang, Y. Guan, C. Zheng, R. Peng, and X. Li, “A temporal spectral generative adversarial network based end-to-end packet loss concealment for wideband speech transmission,” The Journal of the Acoustical Society of America, vol. 150, no. 4, pp. 2577-2588, 2021. Summary of the Invention [Problem to be solved by the invention]

[0007] We propose a less complex and cumbersome solution that is more integrated within the neural coding scheme than [6] by performing the concealment in the quantization domain within the inverse quantization scheme. On the other hand, for neural coders, no or very few specific FEC solutions for neural coding have been proposed so far. [Means for solving the problem]

[0008] According to the present invention there is provided an audio signal representation decoder adapted to decode an audio signal representation from a bitstream, the bitstream being divided into a sequence of packets, the audio signal representation decoder comprising: a bitstream reader configured to sequentially read the sequence of packets; a packet loss controller configured to check whether a current packet has been successfully received or should be considered lost; a quantization index converter configured to convert at least one index extracted from the current packet to at least one current code from at least one codebook, thereby forming at least a portion of the audio signal representation, if the packet loss controller determines that the current packet has been successfully received; The audio signal representation decoder is configured to generate at least one current code from at least one previous code via at least one learnable predictor layer, or to form at least a part of the audio signal representation, if the packet loss controller determines that the current packet should be considered lost.

[0009] According to one aspect, at least one codebook associates an index with a code or a portion of a code, such that the quantization index converter converts at least one index extracted from the current packet into at least one transformed code or at least a portion of a transformed code.

[0010] According to one aspect, the at least one codebook comprises: a base codebook that associates indices with key parts of the code; at least one low-rank codebook that associates an index with a residual portion of the code; the at least one index extracted from the current packet includes at least one high-rank index and at least one low-rank index; The quantization index converter is configured to convert at least one high-rank index into a major part of the current code and convert at least one low-rank index into at least one residual part of the current code; The quantization index transformer is further configured to reconstruct the current code by adding the dominant portion to the at least one residual portion.

[0011] According to one aspect, the at least one codebook comprises: a base codebook that associates indices with key parts of the code; at least one low-rank codebook; the at least one index extracted from the current packet includes at least one high-rank index and at least one low-rank index; The quantization index converter is configured to convert at least one high-rank index into a main part of the current code or a high-rank sub-code, and convert at least one low-rank index into at least one residual part of the current code or a high-rank sub-code; The quantization index converter is further configured to reconstruct the current code by adding the dominant part to at least one residual part, or by combining at least one high-rank subcode with at least one low-rank subcode, or by combining at least one high-rank subcode with at least one low-rank subcode, or by combining a high-rank subcode with at least one low-rank subcode.

[0012] According to one aspect, the audio signal representation decoder may be configured to predict at least one current code from at least one higher-ranked index of at least one preceding or following packet, but not from the lowest-ranked index of at least one preceding or following packet. According to one aspect, the audio signal representation decoder may be configured to predict the current code from at least a high-rank index and at least one medium-rank index of at least one previous packet, but not from a lowest-rank index of at least one previous packet.

[0013] According to one aspect, an audio signal representation decoder configured to store redundancy information written to packets of a bitstream but referencing different packets is configured to store the redundancy information in a temporary storage unit; The audio signal representation decoder searches the temporary storage unit if at least one current packet is to be considered lost, and if redundant information referencing the at least one current packet is retrieved, Retrieving at least one index from the redundant information by referring to the current packet; causing a quantization index transformer to transform at least one index retrieved from the at least one codebook into a permutation code; The processing block may be configured to generate at least a portion of the audio signal by transforming at least one substitution code into at least a portion of the audio signal. According to one aspect, the redundant information provides at least a high-rank index of the at least one preceding or subsequent packet, but not at least one of the low-rank indexes of the at least one preceding or subsequent packet. According to one aspect, the at least one trainable predictor may be configured to perform prediction, the at least one trainable predictor having at least one trainable predictor layer.

[0014] According to one aspect, at least one trainable predictor is trained by sequentially predicting a predicted current code or a current index, respectively, from a preceding and / or following packet, and by comparing the predicted current code or the current code obtained from the predicted index with a transformed code converted from a successfully received packet, to learn trainable parameters of at least one trainable predictor layer that minimize an error of the predicted current code relative to a transformed code converted from a packet having the correct format. According to one aspect, the at least one trainable predictor layer includes at least one recurrent trainable layer. According to one aspect, at least one learnable predictor layer includes at least one gated recurrent unit.

[0015] According to one aspect, at least one trainable predictor layer has at least one state; at least one learnable predictor layer, to predict a current code, such that the current learnable predictor layer instantiation receives state from at least one previous learnable predictor layer instantiation that predicted at least one previous code for at least one previous packet; It is instantiated iteratively along successive instantiations of multiple learnable predictor layers.

[0016] According to one aspect, to predict a current code, a current learnable predictor layer instantiation receives, at input: at least one preceding translation code if at least one preceding packet is deemed to have been successfully received; and at least one preceding predictive code when at least one preceding packet is deemed lost. According to one aspect, to predict the current code, the current learnable predictor layer instantiation receives state from at least one previous iteration both when at least one previous packet is considered to have been successfully received and when at least one previous packet is considered to have been lost.

[0017] According to one aspect, at least one learnable predictor layer is configured to predict a current code and / or receive state from at least one previous learnable predictor layer instantiation and provide a predicted code and / or output state to at least one subsequent learnable predictor layer instantiation both when at least one previous packet is deemed to be received successfully and when at least one previous packet is deemed to be lost. According to one aspect, the current instantiation of the learnable predictor layer includes at least one learnable convolutional unit. According to one aspect, the current instantiation of the learnable predictor layer includes at least one learnable recurrent unit.

[0018] According to one aspect, at least one recurrent unit of the current learnable layer receives state from at least one corresponding recurrent unit from at least one preceding learnable predictor layer instantiation and outputs state to at least one corresponding recurrent unit of at least one subsequent learnable predictor layer instantiation. According to one aspect, the current learnable predictor layer instantiation has a sequence of learnable layers.

[0019] According to one aspect, for a current instantiation of a learnable predictor layer, the sequence of learnable layers includes at least one reduced-dimensionality learnable layer and at least one increased-dimensionality learnable layer following the at least one reduced-dimensionality learnable layer. According to one aspect, the at least one reduced-dimensionality learnable layer includes at least one learnable layer having states. According to one aspect, the at least one dimensionality expansion learnable layer includes at least one stateless learnable layer. According to one aspect, the sequence of learnable layers is gated. According to one aspect, the sequence of learnable layers is gated through a softmax activation function.

[0020] According to the present invention there is provided an audio signal representation decoder adapted to decode an audio signal representation from a bitstream, the bitstream being divided into a sequence of packets, the audio signal representation decoder comprising: a bitstream reader configured to sequentially read a sequence of packets and, from at least one current packet, at least one index of at least one current packet; a bitstream reader configured to extract redundant information about at least one preceding or following packet, the redundant information enabling at least one index in the at least one preceding or following packet to be reconstructed; a packet loss controller, PLC, configured to check whether at least one current packet has been successfully received or should be considered lost; a quantization index converter configured to convert at least one index of at least one current packet to at least one current transformed code from at least one codebook, thereby forming part of the audio signal representation; and a redundancy information storage unit configured to store redundancy information when the PLC determines that the at least one current packet should be considered lost, and to provide the stored redundancy information on the at least one current packet to form part of the audio signal representation via the redundancy information.

[0021] According to one aspect, the redundant information storage unit is configured to store at least one index from a preceding or subsequent packet as redundant information, so as to provide the stored at least one index to the quantization index converter when the controller determines that the at least one current packet should be considered lost. According to one aspect, the redundancy information storage unit is configured to store at least one code previously extracted from a preceding or subsequent packet as redundancy information, in order to bypass the quantization index converter using the stored code when the controller determines that the at least one current packet should be considered lost. According to one aspect, at least one codebook associates an index with a code or a portion of a code, such that the quantization index converter converts at least one index extracted from the current packet into at least one transformed code or at least a portion of a transformed code.

[0022] According to one aspect, the at least one codebook comprises: a base codebook that associates indices with key parts of the code; at least one low-rank codebook that associates an index with a residual portion of the code; the at least one index extracted from the current packet includes at least one high-rank index and at least one low-rank index; The quantization index converter is configured to convert at least one high-rank index into a major part of the current code and convert at least one low-rank index into at least one residual part of the current code; The quantization index transformer is further configured to reconstruct the current code by adding the dominant portion to the at least one residual portion.

[0023] According to one aspect, the audio signal representation decoder may be configured to generate or retrieve at least one current code from at least one high-rank index of at least one preceding or subsequent packet, but not to generate or retrieve at least one current code from a lowest-rank index of at least one preceding or subsequent packet. According to one aspect, the audio signal representation decoder may be configured to generate or retrieve a current code from at least a high-rank index of at least one preceding or subsequent packet and from at least one medium-rank index, but not generate or retrieve a current code from a lowest-rank index of at least one preceding or subsequent packet.

[0024] According to one aspect, there is provided an audio generator for generating an audio signal from a bitstream, the audio generator comprising an audio signal representation decoder; It is further configured to generate the audio signal by converting the audio signal representation into an audio signal. According to one aspect, the audio signal may be further configured to render the generated audio signal.

[0025] According to one aspect, the first data provider may be configured to provide first data derived from an input signal for a given frame (e.g., a portion of an audio signal to be generated). For the given frame, there may be a first processing block configured to receive the first data and output first output data in the given frame; The first processing block is at least one training learnable layer configured to process, for a given frame, target data from the decoded audio signal representation to obtain training feature parameters for the given frame; a styling element configured to apply the adjustment feature parameters to the first data or the normalized first data. According to one embodiment, the audio generator may be configured such that the bit rate of the audio signal is greater than the bit rate of both the target data and / or the first data and / or the second data. According to one aspect, the second processing block may be configured to increase the bit rate of the second data to obtain the audio signal.

[0026] According to one aspect, the first processing block is configured to upsample the first data from a number of samples of a given frame to a second number of samples of the given frame that is greater than the first number of samples. According to one aspect, the second processing block is configured to upsample the second data obtained from the first processing block from a second number of samples for a given frame to a third number of samples for the given frame that is greater than the second number. According to one aspect, the audio generator may be configured to reduce the number of channels of the first data from a first number of channels to a second number of channels of the first output data that is less than the first number of channels.

[0027] According to one aspect, the second processing block may be configured to reduce the number of channels of the first output data obtained from the first processing block from the second number of channels of the audio signal to a third number of channels, the third number of channels being less than the second number of channels. According to one aspect, the audio signal is a mono audio signal. According to one aspect, the audio generator may be configured to derive the input signal from an audio signal representation. According to one aspect, the audio generator may be configured to derive the input signal from noise.

[0028] According to one aspect, the training set of learnable layers comprises one or at least two convolutional layers. According to one aspect, at least one pre-trained learnable layer is configured to receive an audio signal representation or a processed version thereof and, for a given frame, output target data representative of the audio signal in the given frame. According to one aspect, at least one pre-conditioned learnable layer is configured to provide target data as a spectrogram or a decoded spectrogram. According to one aspect, the first convolutional layer is configured to convolve the target data or the upsampled target data using a first activation function to obtain first convolved data. According to one aspect, the training set and styling elements of a learnable layer are part of a weight layer within a residual block of a neural network that includes one or more residual blocks.

[0029] According to one aspect, the audio generator further comprises a normalization element configured to normalize the first data. According to one aspect, the audio generator further comprises a normalization element configured to normalize the first data in a channel dimension. According to one aspect, the audio signal is a vocalized audio signal. According to one embodiment, the target data is upsampled by a factor that is a power of 2, or another factor such as 2.5 or a multiple of 2.5. According to one aspect, the target data is upsampled by non-linear interpolation.

[0030] According to one aspect, the first processing block comprises: a further set of learnable layers configured to process data derived from the first data using a second activation function; The second activation function is a gated activation function. According to one aspect, the further set of learnable layers comprises one or more convolutional layers. According to one embodiment, the second activation function is a softmax-gated hyperbolic tangent (TanH) function. According to one aspect, the first activation function is a leaky rectified linear unit (leaky ReLu) function. According to one embodiment, the convolution operation is performed with a maximum expansion factor of two. According to one embodiment, the audio generator comprises eight first processing blocks and one second processing block. According to one embodiment, the first data has one dimension less than the audio signal. According to one aspect, the target data is a spectrogram.

[0031] According to one aspect, an encoder is provided, the encoder comprising: an audio signal representation generator configured to generate, via at least one learnable layer, an audio signal representation as a representation of the audio signal, the audio signal representation comprising a sequence of tensors; a quantizer configured to convert each current tensor of the sequence of tensors into at least one index, each index being obtained from at least one codebook that associates a plurality of tensors with a plurality of indices; a bitstream writer configured to write packets to the bitstream such that the current packet includes at least one index of a current tensor of the sequence of tensors; The encoder is configured to write redundant information of a current tensor in at least one preceding or subsequent packet of the bitstream that is different from the current packet, and / or to write redundant information of a tensor that is different from the current packet in the current packet. According to one aspect, at least one codebook associates portions of a tensor with indices such that a quantizer transforms a current tensor into multiple indices.

[0032] According to one aspect, the at least one codebook comprises: a base codebook that associates key parts of a tensor with an index; at least one low-rank codebook that associates residual portions of the tensor with indices; At least one current tensor has at least one principal part and at least one residual part; The quantizer is configured to transform a principal portion of the at least one current tensor into at least one high-rank index and to transform at least one residual portion of the at least one tensor into at least one low-rank index; The bitstream writer thereby writes both the high-rank index and at least one low-rank index into the bitstream.

[0033] According to one aspect, the encoder may be configured to provide redundant information that has a high-rank index of at least one preceding or following packet but does not have a lowest-rank index of at least the same at least one preceding or following packet. According to one aspect, the encoder may be configured to transmit the bitstream to a receiver over a communication channel. According to one aspect, the encoder may be configured to monitor payload conditions of the communication channel so as to increase the amount of redundant information if the payload conditions of the communication channel exceed a predetermined threshold.

[0034] According to one aspect, the encoder comprises: When the payload in the communication channel is below a predetermined threshold, for each current packet, transmit only the high-rank index of at least one preceding or subsequent packet as redundant information; The method may be configured to transmit, for each current packet, both a high-rank index of at least one preceding or subsequent packet and at least some low-rank indexes of at least one preceding or subsequent packet as redundant information if the payload of the communication channel exceeds a predetermined threshold. According to one aspect, the encoder may be configured to calculate a packet offset between a current packet and at least one preceding or succeeding packet having redundant information depending at least on the payload of the communication channel.

[0035] According to one aspect, the encoder may be configured to calculate a packet offset between the current packet and at least one preceding or succeeding packet having redundant information depending at least on the envisaged application. According to one aspect, the encoder may be configured to calculate a packet offset between a current packet and at least one preceding or succeeding packet having redundant information depending on at least an input provided by an end user. According to one aspect, the at least one codebook includes a redundant codebook that associates multiple tensors with multiple indices, and the encoder is configured to write redundant information of a current tensor in at least one preceding or subsequent packet of the bitstream that is different from the current packet as an index received from the at least one quantization codebook.

[0036] According to one aspect, there is provided a method for decoding an audio signal representation from a bitstream, the method comprising: Reads the sequence of packets contained in the bitstream, starting with the current packet, At least one index of the current packet; extracting redundant information about at least one preceding or following packet, the redundant information enabling at least one index in the at least one preceding or following packet to be reconstructed; checking whether the current packet has been successfully received or should be considered lost; Transforming at least one index of the current packet into at least one current transformed code from at least one codebook, thereby forming part of the audio signal representation; storing the redundant information to form part of the audio signal representation via the redundant information, and providing the stored redundant information in at least one current packet if the check determines that the at least one current packet should be considered lost.

[0037] According to one aspect, there is provided a method for decoding an audio signal representation from a bitstream, the bitstream being divided into a sequence of packets, the audio signal representation decoder comprising: sequentially reading a sequence of packets; checking whether the current packet was successfully received or should be considered lost; if the check determines that the current packet has been successfully received, converting at least one index extracted from the current packet into at least one current code from at least one codebook, thereby forming at least a portion of the audio signal representation; and generating at least one current code by prediction from at least one previous code or index via at least one learnable predictor layer if the packet loss controller determines that the current packet should be considered lost.

[0038] According to one aspect, the method provided herein comprises: generating an audio signal representation as a representation of the audio signal via at least one learnable layer, the audio signal representation comprising a sequence of tensors; converting each current tensor of the sequence of tensors into at least one index, each index being obtained from at least one codebook that associates a plurality of tensors with a plurality of indices; writing a packet in the bitstream such that the current packet includes at least one index of a current tensor in the sequence of tensors; writing redundant information of the current tensor to at least one preceding or subsequent packet of a bitstream different from the current packet, and / or writing redundant information of at least one tensor to be written to at least one preceding or subsequent packet of a bitstream different from the current packet to the current packet.

[0039] According to one aspect, a non-transitory storage unit is provided, which, when executed by a computer, causes the computer to: From the current packet, At least one index of the current packet; extracting redundant information about at least one preceding or following packet, the redundant information enabling at least one index in the at least one preceding or following packet to be reconstructed; Checks whether the current packet has been received well or should be considered lost, Transforming at least one index of the current packet into at least one current transformed code from at least one codebook, thereby forming part of the audio signal representation; Stores instructions for controlling the storage of redundant information and retrieving the stored redundant information on the at least one current packet when the check determines that the at least one current packet should be considered lost to form part of the audio signal representation via the redundant information.

[0040] According to one aspect, a non-transitory storage unit is provided, which, when executed by a computer, causes the computer to: Have the sequence of packets read sequentially, causes the current packet to be checked to see if it was received well or if it should be considered lost, if the check determines that the current packet was successfully received, converting at least one index extracted from the current packet into at least one current code from at least one codebook, thereby forming at least a portion of the audio signal representation; If the check determines that the current packet should be considered lost, instructions are stored to generate at least one current code by prediction from at least one previous code or index via at least one learnable predictor layer.

[0041] According to one aspect, a non-transitory storage unit is provided, which, when executed by a computer, causes the computer to: generating an audio signal representation as a representation of the audio signal via at least one learnable layer, the audio signal representation comprising a sequence of tensors; transforming each current tensor of the sequence of tensors into at least one index, each index being taken from at least one codebook that associates multiple tensors with multiple indices; causing a packet to be written to the bitstream such that the current packet includes at least one index of the current tensor in the sequence of tensors; Stores instructions to cause redundant information of a current tensor to be written to at least one preceding or subsequent packet of a bitstream different from the current packet and / or to cause the current packet to write redundant information of at least one tensor to be written to at least one preceding or subsequent packet of a bitstream different from the current packet.

[0042] In the above aspects, references are often made to parts of the code, for example, they may refer to components (e.g., addends) or subcodes (e.g., high-rank and low-rank subcodes). [Brief explanation of the drawings]

[0043] [Figure 1a] 1 illustrates an example of a PLC according to the present disclosure. [Figure 1b] 1 illustrates an example of a PLC according to the present disclosure. [Figure 2] 1 illustrates a technique in an audio signal representation decoder. [Figure 3a] Demonstrates bitstream buffering techniques. [Figure 3b] Demonstrates bitstream buffering techniques. [Figure 4] The evaluation results of this example are shown below. [Figure 5] The evaluation results of this example are shown below. [Figure 6a] 1 shows an example of an audio encoder and an audio signal representation encoder. [Figure 6b] 1 shows an example of an audio encoder and an audio signal representation encoder. [Figure 7] 1 shows an example of an audio decoder and an audio signal representation decoder. [Figure 8a] 1 shows an example of an audio decoder and an audio signal representation decoder. [Figure 8b] 1 shows an example of an audio decoder and an audio signal representation decoder. [Figure 9] 1 shows an example of an audio decoder and techniques for an audio decoder and an audio signal representation decoder. [Figure 10] 1 shows an example of an audio decoder and techniques for an audio decoder and an audio signal representation decoder. [Figure 11] 1 shows an example of an audio decoder and techniques for an audio decoder and an audio signal representation decoder. [Figure 12] 1 shows an example of an audio decoder and techniques for an audio decoder and an audio signal representation decoder. [Figure 13] 1 shows an example of an audio decoder and techniques for an audio decoder and an audio signal representation decoder. DETAILED DESCRIPTION OF THE INVENTION

[0044] Note that in the following, we often refer to learnable layers, which may be implemented, for example, within a neural network.

[0045] 6a and 6b show two examples of encoders 1600, specifically encoder 1600a of FIG. 6a and encoder 1600b of FIG. 6b. Referring to FIG. 6a, the encoder 1600a of FIG. 6a encodes an input audio signal 1602 onto a bitstream 1630. The input audio signal 1602 may be, for example, an uncompressed analog or digital representation of an audio signal recorded from a microphone and / or stored in a storage unit and / or received remotely. The encoder 1600a may operate sequentially, for example, by sequentially generating packets (or portions of packets, or multiple packets) of a bitstream from a portion of the input audio signal 1602. The encoder 1600a may include an audio signal representation generator 1604. The audio signal representation generator 1604 may include at least one learnable layer and may therefore be considered a learnable audio signal representation generator 1604. The audio signal representation generator 1604 may generate (e.g., via at least one learnable layer) an audio signal representation 1606, which may be a sequence of tensors (codes). Each tensor may be a vector or a matrix, or a generalized matrix (e.g., one with more than two dimensions, e.g., an n×m×p tensor with at least one of n, m, and p greater than 1). If a tensor is a vector, it shall have at least two dimensions (e.g., an n×1 matrix with n greater than 1).

[0046] The encoder 1600a may include a quantizer 1608. The quantizer 1608 may convert each current tensor 1606 in the sequence of tensors to at least one index 1626. Thus, a sequence of indices may be output by the quantizer 1608. Each index may be received from at least one codebook, collectively indicated in FIG. 6a by reference numeral 1620. In general terms, the quantizer 1608 may search the at least one codebook 1620 for an index that represents a particular code (or portion thereof) in the bitstream 1630.

[0047] In some examples, there may be several codebooks. Figure 6a shows a high-rank codebook 1622. The high-rank codebook may output at least one high-rank index 1623 to the quantizer 1608. Figure 6a also shows a low-rank codebook 1624 (which may be optional), which may output a low-rank index 1625 to the quantizer 1608. This is because, in some examples, multiple indices may be associated with a tensor to increase resolution, with higher-rank indexes 1623 being given to the most important parts of the tensor 1606, lower-rank indexes 1625 being given to less important parts of the tensor 1606, and so on, down to the lowest-rank index being given to the least important parts of the tensor 1606. Thus, there may be more than one low-rank codebook (in which case there may be a ranking between different codebooks so that each codebook has a different ranking from the other codebooks, and there may be a base codebook that is the highest-ranked codebook, and a low-rank codebook). In some examples, there are three codebooks (e.g., a base codebook that is the highest-ranked codebook, a medium-ranked codebook, and a lowest-ranked codebook). In other examples, there may be four codebooks (e.g., a base codebook that is the highest-ranked codebook, a first highest-ranked codebook, a second highest-ranked codebook, and a lowest-ranked codebook). The index output by the base codebook is the highest-ranked index, the index output by the lowest-ranked codebook is the lowest-ranked index, and so on. However, in some examples, there may be a single codebook 1620 (and thus no low-ranked codebook 1624). In either case, each codebook 1620 (whether there is a single codebook or multiple codebooks 1622, 1624, etc.) provides an index 1626 for each tensor or portion of a tensor. Thus (e.g., if there is only one codebook and no low-rank codebook 1624), each tensor 1606 is mapped to a single index 1626. Alternatively, each tensor 1606 may be mapped to multiple indices 1626 (e.g., 1623, 1625), e.g., in which case there are multiple codebooks (e.g., 1622, 1624, etc.). For each tensor input to the quantizer 1608, the output indices 1626 may be identified, for example, by their position.

[0048] The quantizer 1608, when using several codebooks, can include a technique known as split vector quantization and a multi-stage vector quantization technique, also known as residual vector quantization. In split vector quantization, the tensor to be quantized is divided into multiple subvectors (or, more generally, subtensors), which are then quantized independently. This allows for finer control over the quantization process, since different subvectors (or, more generally, subtensors) can be quantized using different bit widths or precision levels. Split vector quantization design can be performed manually by selecting the optimal bit width for each subvector (or, more generally, subtensor), or automatically using machine learning techniques. Multi-stage vector quantization, on the other hand, involves quantizing a tensor from a low-precision representation to a high-precision representation in iterative stages, with each stage further reducing quantization distortion. This is achieved as described above by first encoding the tensor using the highest-ranked codebook and then further encoding the resulting quantization error using the second-highest-ranked codebook. The process is repeated until the final stage, which uses the lowest-ranked codebook. Again, the quantization design can be done manually by selecting the optimal bit-width for each stage, or automatically using machine learning techniques.

[0049] The encoder 1600a may include a bitstream writer 1628. The bitstream writer 1628 may write packets to the bitstream 1630. For example, an index 1626 (e.g., 1623, 1625) may be encapsulated in the current packet and / or in a predetermined location according to a predetermined syntax. As shown in FIG. 3a (see below), a packet may include one main frame (which may be a main frame and carries an index 1626 written by the quantizer 1608 based on at least one codebook) and one auxiliary frame (which may carry redundant information 1612 or 1612b, as described below). Thus, the encoder 1600a may write redundant information 1612 for the current sensor 1606 to the bitstream 1630. However, the redundant information 1612 may be written to at least one preceding or following packet of the bitstream 1630 that is different from the current packet. Similarly, there may be redundant information for further packets in the preceding or following packets of the bitstream.

[0050] Furthermore, the current packet may be associated with redundant information of a different packet (e.g., in the same frame). The bitstream 1630 may also include further information for each packet, such as a packet identifier and syntactic redundancy check information (e.g., cyclic redundancy check (CRC) information or other syntactic redundancy check information such as parity / disparity bits), which helps the receiver distinguish between packets that should be considered correctly received (and therefore used to render the audio signal) and packets that should be considered lost (and therefore used to render the audio signal). However, advantageously, even if a packet should be considered lost by the receiver, redundant information exists in at least one other packet that allows for at least partial reconstruction of a portion of the audio signal.

[0051] According to one example, the redundant information 1612 may be output by the redundant information storage 1610, for example, to be provided to the bitstream writer 1628. The redundant information storage 1610 may store an index 1626 (e.g., 1623, 1625) associated with the current tensor 1606 and provide the index to the bitstream writer 1628 in a packet different from the current packet. Note that in FIG. 6a, both the high-rank index 1623 (received from the high-rank codebook 1622) and the low-rank index 1625 (received from the low-rank codebook 1624) are shown to be provided to the redundant information storage 1610. Nevertheless, in some examples, only the high-rank codebook 1623 is provided to the redundant information storage 1610, while the low-rank codebook 1625 is refrained from being provided to the redundant information storage 1610, in which case the low-rank codebook 1625 consequently does not contribute to the redundant information 1612 written to the bitstream 1630. In examples where there are more than two codebooks (e.g., a high-rank, base codebook, a medium-rank codebook, a low-rank codebook with a lower rank than the base codebook, etc.), only indices from the high-rank codebook may be provided to the redundant information store 1612, and at least one low-rank codebook may provide indices that are not provided to the redundant information store. In some examples, at least a portion (e.g., all) of at least one low-rank index is also provided to the redundant information store 1610, but it is dynamically determined whether to also provide the low-rank indexes to the bitstream writer 1628 based on the network payload conditions (e.g., if the network is busy beyond a certain threshold, at least a portion of the low-rank indexes are not written to the bitstream 1630, while if the network is busy beyond a certain threshold, the low-rank indexes are also written to the bitstream 1630; see also below). Thus, each codebook 1620 (1622, 1624, etc.) may associate portions of a tensor with indices such that the quantizer 1608 transforms the current tensor 1606 into multiple indices.

[0052] As described above, each codebook 1620 (e.g., 1622, 1624) may include (in some examples) a base codebook (high-rank codebook) 1622 that associates a prime portion of a tensor with an index, and at least one low-rank codebook 1624 that associates an index with a residual portion of the tensor. This is because each tensor may have at least one prime portion and at least one residual portion (the residual portion may be one or more and may be ranked exactly the same as the codebook). Thus, the quantizer 1608 may convert the prime portion of at least one current tensor into at least one high-rank index 1623 and convert at least one residual portion of at least one tensor into at least one low-rank index 1625. Thus, the bitstream writer 1628 may write both the high-rank index and the at least one low-rank index 1625 to the bitstream 1620. As described above, in some examples, only at least one high-rank index 1623 (obtained from a high-rank codebook 1622) of at least one preceding or following packet is written to the bitstream, and at least the lowest-rank index 1625 (or, in some examples, other low-rank indexes having intermediate rankings between the highest-rank codebook and the lowest-rank codebook) is not written to the bitstream 1630.

[0053] In the encoder 1600b of Figure 6b, the trainable audio signal representation generator 1604 may be the same as that in the encoder 1600a of Figure 6a. Also, the bitstream writer 1628 may, in principle, be the same as the analogous bitstream writer in the encoder 1600a of Figure 6a. There may be at least one codebook 1620a (e.g., 1622, 1624), which may, in principle, be identical to at least one codebook 1620 (e.g., 1622, 1624) in the encoder 1600a of Figure 6a. However, in the encoder 1600b, redundant information (here denoted 1612b) may be obtained from an index 1623b derived from at least one codebook 1620b ("redundant codebook") that is different from the codebook 1620a. The indices 1623b provided by the codebook 1620b may be a codebook with limited resolution and reduced length (thus reducing the payload and speeding up transmission). Essentially, the codebook 1620b is generally different from the main codebook 1620a (1622, 1624), which outputs indices 1626 (e.g., 1623, 1625, etc.) to the quantizer 1608. The indices 1623b provided by the redundant codebook 1620b may advantageously have worse resolution than the indices 1626 (e.g., 1623, 1625, etc.) provided by at least one codebook 1620a (e.g., 1622, 1624, etc.). Thus, at least one codebook 1620a may be considered as at least one main codebook, while at least one redundant codebook 1620b (which may have worse resolution) may provide approximate information about the indices 1626 (e.g., 1623, 1625, etc.) provided to the bitstream 1630. Thus, the redundant codebook 1620b may provide indices 1623b that occupy fewer bit lengths than the indices output by the main codebook 1620a.Furthermore, the redundant codebook can be designed and trained for a particular size and redundant information storage needs, resulting in better redundant information 1612b than partially retaining the indices derived from the quantizer 1608. Alternatively, the at least one codebook 1620b (and indexes 1623b) may have the same design as the at least one primary codebook 1620a. For example, there may be one high-rank redundant codebook and at least one low-rank redundant codebook, and there may be different approximations. Thus, the fact that the arrow 1623b is a single arrow does not necessarily mean that there is only a single index 1623b output by the redundant codebook 1620b. Thus, the redundant codebook 1620b can be described in the same manner as the ranking codebook 1620a in any aspect. Therefore, any features described in principle for the at least one codebook 1620 (1622, 1624) or 1624a may also be used to describe any example of the redundant codebook 1610, and that description will not be repeated for the sake of brevity.

[0054] 3a shows an example of encapsulating both the index 1626 converted by the quantizer 1608 and the redundant information 1612 or 1612b for either of the encoders 1600a and 1600b. The redundant information 1612 or 1612b may be generated by the bitstream writer 1628, for example, in a jitter buffer 1628j. Thus, the bitstream writer 1628 (e.g., the jitter buffer 1628j) may generate packets for the primary frame (which may include the index 1626 output by the quantizer 1608) and the redundant frame (which may include the redundant information 1612 or 1612b). The figure shows the nth packet with primary frame n and redundant frame n-5, the (n+1)th packet with the (n+1)th primary frame and the (n-4)th redundant frame (which is also the redundant information 1612 or 1612b), and up to the (n+5)th primary frame and the nth redundant frame. Essentially, redundant information is retrieved from the nth packet (e.g., a reduced version) so that the n+5th packet only has high-rank indices 1623 and no low-rank indices 1625. Thus, when the nth packet is encoded by bitstream writer 1628, information about primary frame n is provided to redundant information store 1610 (in the case of encoder 1600a of FIG. 6a) and is then used as redundant information 1612 for encoding the (n+5)th packet. Another case is the example of FIG. 6b, where when n+5 is encoded, the redundant information for the nth packet (redundant information 1612b) is retrieved from a different codebook relative to redundant codebook 1620b, which is different from primary codebook 1620. Note also that primary frame 1626 is shown to include the i-th frame indices C0(i), C1(i), C2(i), C3(i) (C0 is the index 1623 obtained from the high-rank codebook 1622, C1 is the first low-rank index 1625 obtained from the low-rank codebook 1624, and so on). For the example of Figure 6b, redundant frame C0(i-5) (1612) may be replaced with C'(i-5).

[0055] As mentioned above, in other cases, the bitstream 1630 may be transmitted to a receiver. FIG. 3b shows an example of a technique that may be implemented in the encoder 1600a or 1600b. In this case, the bitstream 1630 is transmitted over a network 1640, or more generally, over a communication channel. Here, the amount of redundant information written into the current packet may vary based on the state of the network 1640. FIG. 3b shows a detector 1642 that may detect the state of the network 1640 (the detector may, for example, measure the latency of a transmission, as provided, for example, in an acknowledgement from the receiver, and / or measure the amount of transmissions sent simultaneously by other devices, and / or have knowledge of corrupted frames acknowledged by the receiver, or may detect or sense any other quantity that allows determining the state of the communication channel). The controller 1644 may read the state 1643 of the network 1640. As can be seen, controller 1644 is shown controlling switch 1645, which can select whether to allow or prevent encoding of at least one low-rank index 1625, but which does not affect the encoding of high-rank index 1625. Switch 1624 may be open, for example, when the payload condition of network 1640 is below some predetermined threshold, so that there is no encoding of low-rank index 1625 when the network is busy, while at least one high-rank index 1623 is nevertheless written to bitstream 1630. When the payload condition of network (communication channel) 1640 is below said predetermined threshold (meaning the network is relatively clear), switch 1645 may be closed, allowing low-rank index 1625 to be written to bitstream 1630.

[0056] Additionally or alternatively, the controller 1644 may exercise a control 1645′ to control the offset between the current packet and the packet for which the redundant information 1612 or 1612b was received based on the payload status 1643 of the communication channel 1640. Thus, the offset between the currently written packet and the packet for which the redundant information is provided may change dynamically depending on the payload. Referring to FIG. 3a, there may be a situation in which the (n−4)th packet is encoded as redundant information instead of the (n−5)th packet, thereby dynamically changing the offset in the payload. In another situation, the control 1645′ may be based on a selection or preselection from a user rather than (or entirely based on) the payload 1643 of the network 1640. As an example, if the network is congested, the packet loss rate may be higher, meaning that bursts of packet loss are larger, and the offset may need to be increased. Alternatively, another layer of redundant information may be added, such that in the signal, the nth packet adds two pieces of redundant information with two different offsets. For example, the nth packet can be given (n-5)th redundant information and (n-7)th redundant information. By adding additional bit rates, more of the bitstream can be protected.

[0057] Thus, the encoder may calculate a packet offset between a current packet and at least one preceding or following packet having redundant information depending at least on the payload of the communication channel, e.g., the higher the payload in the communication channel or the higher the error rate in the communication channel, the larger the packet offset. The packet offset may be signaled in the bitstream.

[0058] In an example, the packet offset between the current packet and at least one preceding or subsequent packet with redundant information may be defined by the encoder depending at least on the intended application, in an example, the packet offset between the current packet and at least one preceding or subsequent packet with redundant information depending at least on input provided by an end user.

[0059] By virtue of the above, it can be seen that redundant information 1612 or 1612b may be provided in the bitstream 1630, which, in the event of a packet being lost, helps to reconstruct the packet from the redundant information 1612 or 1612b written in different packets.

[0060] 7 shows an example of an audio generator 1700. The audio generator 1700 may convert a bitstream 1630 (which may, in some examples, be the same bitstream generated by the encoders 600a or 600b of FIG. 6a or 6b). The audio generator 1700 may generate an audio signal 1724 from the bitstream 1630. Note that the audio signal 1724 is generally meant to be a reliable representation of the input audio signal 1602, e.g., as provided to the encoders 1600a and 1600b; e.g., here, forward error correction (FEC) may be implemented.

[0061] This example can be implemented in an audio signal representation decoder 1710 (which may or may not be part of the audio generator 1700). The audio signal representation decoder 1710 may decode an audio representation 1720 that represents the audio signal 1602 (which is then transformed in the audio signal 1724). Accordingly, we will now describe how the audio signal representation decoder 1710 may be configured according to some examples. The audio signal representation decoder 1710 may first decode the audio representation 1720 from a bitstream 1630. The bitstream 1630 is divided into a sequence of packets, for example, as described above. The audio signal representation decoder 1710 (or more specifically the audio generator 1700) may comprise a bitstream reader 1702 (e.g., an index extractor). The bitstream reader 1702 may sequentially read the sequence of packets (forming the bitstream 1630). The bitstream reader 1702 may extract at least one index 1704 (e.g., multiple indexes) for the at least one current packet from the at least one current packet. From the at least one current packet, redundancy information 1714 providing information of at least one preceding or subsequent packet may be provided to a redundancy information storage unit 17100 (see below). The redundancy information 1714 may subsequently be provided as redundancy information 1712 for subsequent packets (if that packet is deemed lost), see below. The index 1704 extracted by the bitstream reader 1702 may be an index 1626 (1623, 1625) or 1623b, or a representation thereof, inserted into the bitstream 1630 by the encoder 1600a or 1600b. The redundant information 1714 may be redundant information 1612 and / or 1612b inserted by the redundant information storage 1610 or 1610b of the encoder 1600a or 1600b, respectively.

[0062] The indices 1704 extracted by the index reader 1702 may then be transformed by a quantization index transformer 1718 .

[0063] The audio signal representation decoder 1710 (or, specifically, the audio generator 1700) may include a packet loss controller (PLC) 1706 (which may operate as an FEC controller). The PLC 1706 may check whether at least one current packet has been received successfully or should be considered lost. For example, the PLC may perform a syntax check on a redundancy code inserted in the bitstream 1630 in relation to the current packet (or any other check, such as a check of syntactic redundancy check information). The PLC 1706 may thus distinguish between a current packet that should be considered correct and a current packet that should be considered lost. The output 1708 of the PLC 1706 may therefore be referred to as correctness information. If the correctness information 1708 indicates that the current packet should be considered correct, a code (tensor) is decoded from the index of the correct packet. Otherwise, if the correctness information 1708 indicates that the current packet should be considered lost, the index of the current packet in the bitstream is not decoded, or at least is not used at all. 7 via a switch 1716 connecting a quantization index converter 1718 (which outputs a transformed code of the audio signal representation 1720) to retrieve at least one index 1704 of the current packet as a lead from redundancy information 1712 provided by the bitstream reader 1702 or by the redundancy information storage unit (index predictor) 17100. Thus, if the correctness information 1708 indicates that the current packet should be considered valid (i.e., correct), the switch 1716 connects the output of the bitstream reader 1702 to the input of the quantization index converter 1718. If the correctness information 1708 indicates that the current packet should be considered lost, the switch 1716 switches to connect the output 1712 of the redundancy information storage unit (index provider) 17100 to the input of the quantization index converter 1718.The output 1712 of the redundancy information storage unit 17100 is the redundancy information (eg, 1612, 1612b) of the current packet that was previously obtained from another packet.

[0064] The quantization index converter 1718 may convert at least one index 1704 (or redundant information 1712) into a code or portion of a code 1720. The converted code may be a tensor (e.g., a vector, but if a vector, it shall be at least two-dimensional). The converter code 1720 may, in some examples, be meant to be a copy of the audio signal representation 1606 of Figures 6a and 6b, if possible. The audio signal representation 1720 (a sequence of codes, such as a tensor) may be the output of the audio signal representation decoder 1710. The audio signal representation 1720 may be input to a processing / rendering block 1722, which may generate an audio signal 1724.

[0065] If the first packet in the sequence is deemed correct by PLC 1706, then redundancy information 1714 may include one index (e.g., 1626, 1623, 1625, 1623b) of a second, different packet in bitstream 1630. If the current packet is deemed lost, then redundancy information 1714 is not provided to redundancy information storage unit 17100.

[0066] It may have the following sequence: 1. The first packet in the bitstream 1630 is received and deemed correct by the PLC 1706. The correctness information 1708 indicates that the current packet is correct. 2. A switch 1716 then connects the output of the bitstream reader 1702 (which is at least one index 1704 representing at least one index 1623, 1625, 1626, 1623b in Figures 6a, 6b) to a quantization index converter 1718. 3. Further, at least one piece of redundant information 1714 for the second packet of the bitstream 1630 is stored (as 1714) in the redundant information storage unit 17100 (the redundant information 1714 may be redundant information 1612 or 1612b provided to the bitstream reader 1628 by the redundant information storage device 1610 or 1610b). 4. For subsequent frames, if the PLC 1706 detects that the format is incorrect and the packet should be considered lost, the validity information 1708 indicates that the packet should be considered lost. 5. Next, switch 1716 is operated to connect the output 1712 of the redundancy information storage unit (index predictor) 17100 to the input of the quantization index converter 1718. In this case, input 1714 is not provided to the redundancy information storage unit 17100.

[0067] Storing redundant information 1714 (eg, indices or high-rank indices of key parts of tensors of audio signal representation 1606) allows for the audio signal representation 1606, or at least key parts thereof, to be reconstructed.

[0068] FIG. 7 also shows at least one codebook 1620, which may be a copy of at least one codebook 1620 or 1620a of either of FIGS. 6a and 6b. The at least one codebook 1620 may also include the main codebook 1620a and at least one redundant codebook 1620b of FIG. 6b. Thus, the information obtained from the codebooks may be provided in the same manner (it should be noted, however, that the codebooks in the encoders of FIGS. 6a and 6b provide indices based on codes, while the codebooks of FIGS. 7-8b provide codes based on indices). Even in this case, it is possible to have at least one high-rank codebook 1622 and / or at least one low-rank codebook 1624. In general, the techniques suggested for the redundant information storage unit are the same as those described in FIGS. 6a and 6b and therefore will not be repeated here. It should be noted that the codebooks of the audio signal representation decoder 1700 are generally identical to those of the encoder to enable correct decoding of the audio signal representation 1720. As the technology is the same, the same features will not be repeated.

[0069] A processing and / or rendering block 1722 may be used, for example, to process and / or render the audio signal 1724 represented by the transformed code 1720. It should also be noted that the redundant information 1712 used when a packet is considered lost may be information obtained from a packet that has an offset relative to the current packet, for example, as defined by control 1645' in Figure 3b. The offset used may be signaled, for example, in the bitstream 1630. The audio signal representation decoder 1710 may read a signal indicating a packet offset between the current packet and at least one preceding or following packet having redundant information, depending at least on the payload of the communication channel, to reconstruct the packet to which the redundant information refers and to store the redundant information associated with the packet to which the redundant information refers.

[0070] 7, it is assumed that the redundancy information storage unit 17100 is an index provider that stores an index. However, a variant is also possible in which the redundancy codes have already been obtained in the quantization index converter 1718 and are therefore already converted and stored in the redundancy information storage unit 17100. In this case, when the redundancy information 1714 is needed (because the current packet is considered lost), the quantization index converter 1718 may be bypassed and the redundancy information storage unit 17100 directly provides part of the audio signal representation 1720.

[0071] 8a illustrates an example of an audio generator 1800, referred to herein as audio generator 1800a. The audio generator 1800a may comprise an audio signal representation decoder 1810 (referred to herein as 1810a). The audio signal representation decoder 1810a may be separate from the audio signal generator 1800a. The audio generator 1800a may generate at least one audio signal 1824a from a bitstream 1830. The audio signal representation decoder 1810a may generate an audio signal representation 1820a from the bitstream 1830. A processing and / or rendering block 1822a of the audio generator 1800a may be input with the audio signal representation 1820a. Because the audio signal representation decoder 1810a may be separate from the processing and / or rendering block 1822a, the audio signal representation decoder 1800a will be described herein separate from the processing and / or rendering block 1822a.

[0072] The bitstream 1830 may, in some examples, be the same bitstream 1630 described above (e.g., as generated by the encoder 1600a and / or the encoder 1600b and / or as input to the audio signal representation decoder 1710). However, in some examples, the bitstream 1830 may differ from the bitstream 1630, and it is not strictly necessary to have the redundant information 1612 written into the bitstream 1830.

[0073] The audio signal representation 1810a may include a bitstream reader (or index extractor) 1802a. This bitstream reader 1802a may be of the same type as the bitstream reader 1702 of FIG. 7 in some examples. However, in this case, the redundant information is not necessarily read (sometimes it is, sometimes it is not, but in some examples it is not necessary to have it). The bitstream reader 1802a may output an extracted index 1804a. The audio signal representation decoder 1810a may include a packet loss controller 1806a, which may be of the same type as the PLC 1706. As mentioned above, the packet loss controller 1806a may perform a check (e.g., based on cyclic redundancy coding (CRC) or similar technique) to verify whether the format of a received packet of the bitstream 1830 should be considered correct or lost. The output of the PLC 1806a may therefore be validity information 1808a (which may be of the same type as the validity information 1708 described above). The audio signal representation decoder 1810a may include a quantization index converter 1818a. The quantization index converter 1818a may output transformed codes 1820a, for example, from the indices 1804a. The transformed codes 1820a may be of the same type as the transformed codes 1720 described above (e.g., they may be tensors, and in some cases in certain vector cases may be n×1 vectors, for example, where n>1), and / or in some examples may be a representation of the audio signal representation 1606 of FIGS. 6a and 6b above. The transformed codes 1820a may be the output of the audio generator 1800a, at least for codes transformed from indices extracted from bitstream packets having the correct format (as identified by the PLC 1806a).If the PLC 1806a establishes that the current packet has an incorrect format (and is therefore considered lost), the PLC 1806a may cause (e.g., via correctness information 1808a) to omit conversion of the index 1804a extracted from the current packet. As seen in FIG. 8a, this is indicated by a switch 1816a (controlled by correctness information 1808a) that can selectively prevent the quantization index converter 1818a from receiving an index from the bitstream reader 1802a. If the current packet is considered lost, a trainable code predictor 1810aa may be used (FIG. 8a shows the output of the trainable code predictor as provided based on the correctness information 1808a). The output of the trainable code predictor 1810aa may be a predicted code 1811a. Thus, when the current packet is considered lost, instead of the conversion code 1820, the prediction code 1811a may be provided to the processing and / or rendering block 1822a, or more generally, represent the output for a particular packet of the audio signal representation decoder 1810a.

[0074] Variations on audio generator 1800a and audio signal representation decoder 1810b are represented in FIG. 8b as audio signal representation 1800b (collectively referred to as audio generator 1800) and audio signal representation decoder 1810b (collectively referred to as 1810), where the same elements as in FIG. 8a are numbered the same but with the suffix "b" instead of "a." As can be seen, PLC 1806b (which may therefore be similar to PLC 1806a) may output validity information 1808b regarding the format of the current packet of bitstream 1830. Bitstream reader (index extractor) 1802b may be the same type of bitstream reader (index extractor) 1802a as in FIG. 8a and may output extracted index 1804b (which may be similar to extractor index 1804a of FIG. 8). Thus, the quantization index converter 1818b may be input with the extracted index 1804b if the PLC 1806b confirms that the current packet is correct, and may be input with the predicted index 1811b if the PLC 1806b determines that the current packet should be considered lost. This is the main difference between the audio signal representation decoder 1810b and the audio signal representation decoder 1810a. Here, the index 1816b' is predicted by a learnable index predictor 1810bb (there is no learnable code predictor 1810aa), but there is a learnable index predictor 1810bb to which the extracted index 1804b may be input, and the extracted index may then be used to perform prediction (1811b) of the subsequent and / or preceding index when the packet is considered lost (e.g., via correctness information 1808b). This is represented by switch 1816b, which shifts the output 1804b (extracted index) from bitstream reader 1802 towards quantization index 1818b and the output 1811b of learnable index predictor 1810bb and the input of quantization index converter 1818b.

[0075] The quantization index converter 1818b may be similar to the quantization index converter 1818a of Figure 18. As can be seen by comparing Figure 8b with Figure 8a, the quantization index converter 1818b receives a code 1804b if the packet is valid, and a predicted index 1811b if the packet is believed to be lost.

[0076] The examples of both Figures 8a and 8b may utilize a codebook 1820. The codebook 1820 may, in some examples, be one of or a copy of codebooks 1620, 1620b, 1620a, 1622, 1624, etc. The codebook 1820 may provide codes 1826 (e.g., 1626, 1623, 1625, 1623b) to the learnable code predictor 1810aa, the quantization index converter 1818a or 1818b, and / or the learnable index predictor 1810bb. Any of the above examples may be used to implement the codebook 1820 of Figure 8a or 8b.

[0077] For example, there may be a high-rank codebook and a low-rank codebook (also shown here as 1622 and 1624 for simplicity). Thus, in the example of FIG. 8a, the learnable code predictor 1810aa may predict the code 1811a from an index (obtained from the codebook 1820). When there is a high-rank codebook and a low-rank codebook, the codebook 1820 may, in some examples, only have a high-rank codebook (e.g., 1622), thereby providing only high-rank indices 1623 to the learnable code predictor 1810. Thus, predictions in the learnable code predictor 1810aa may be limited to only high-rank codebooks (e.g., only the highest-ranked codebook). In some examples, the learnable code predictor 1810aa may learn predictions from the currently transformed code 1820a, such that it performs predictions based on previous transformations of the correct packet. This is the meaning of the arrow 1820a' from the transformer code 1820a output to the trainable code predictor 1810aa by quantization in the quantization index transformer 1818a.

[0078] A similar strategy may be performed in the audio signal representation decoder 1810b, where at least one code 1820 (which may be the same as in FIG. 8a and may be any of 1620, 1620a, 1620b, and 1624, etc.) may be input to the learnable index predictor 1810bb. Here, the codebook 1820 may provide an index to the quantization index converter 1818b (to correct the index 1804b extracted from the correct packet) and / or the learnable index predictor 1810bb (e.g., to predict the index 1811b if the packet is not correctly retained). As shown in FIG. 8b, an arrow 1816b' connects the extracted index 1816b (if correct) to the input of the learnable index predictor 1810bb, so that the learnable index predictor 1810bb can learn the correct index. Of course, when predicting an index by the learnable index predictor 1810bb, it is not necessary to provide the input 1816b' to the learnable index predictor 1810bb.

[0079] As described above and below, high-rank vs. low-rank codebooks may be used in the case of split quantization or residual quantization. For example, a base codebook (high rank) may be used to decode the prime part of a prime part code (or prime sub-code), and a low-rank codebook may be used to decode the residual part of the code (or low-rank sub-code). Then, the prime part of the code and the residual part of the code can be added (e.g., by addition) to combine different sub-codes into a transformed code. In some cases, there are no different rankings for different subcodes, but still different codebooks.

[0080] 2 shows an example of a learnable code predictor 1200, which may be, for example, the learnable code predictor 1810aa of the audio signal representation decoder 1810a of FIG. 8a. FIG. 2 shows a sequence of previously transformed codes 1202, which may be the previously transformed codes 1820a' portion of the audio signal representation 1820a transformed by the quantization index converter 1818a. The previously transformed codes 1202 (which may be the codes 1820a' shown in FIG. 8a) are shown here in sequence. If the current code is n in the sequence (and therefore considered to be at time t=n), the previously predicted codes may be:

[0081] · The 0th previously converted code 1820a'0 (packet at time t=0). First previously translated code 1820a'1 (translated from the packet at time t=1). · Second converted code 1828'2 (converted from the packet at time t=2). ·… · Current (last) converted code 1820a'n (converted from the packet at time t=n-1).

[0082] To predict the current nth code, a prediction is obtained as estimated code 1811a for the current code (t=n). Previous predicted codes (1811a3, 1811a2, and 1811a1) are also obtained for previous time instances (t=3, t=2, t=1). Note that in some examples, the sequence may be limited to a predetermined number of previous time points (e.g., the last 5, or 10, or 20 packets). The output of the trainable code predictor 1200 (1810aa) may be a sequence of predicted codes 1204 (e.g., may be the predicted code 1811a predicted by the trainable code predictor 1810aa of FIG. 8a). As can be seen, the trainable code predictor 1200 (1810aa) may comprise at least one trainable predictor layer. The at least one trainable predictor layer may include at least one recurrent trainable layer (e.g., a recurrent neural network). The at least one trainable predictor layer may include at least one gated recurrent unit. More generally, the audio signal representation decoder 1800 may be autoregressive.

[0083] At least one trainable predictor layer may be instantiated repeatedly along the sequential instantiations of multiple predictor layers and along the sequence of packets whose codes are predicted sequentially. Examples of trainable predictor layer instantiations (collectively referenced 1210) include: · Instantiate 12101 a learnable predictor layer to predict the first code 1811a1 (t=1). · Instantiating 12102 a learnable predictor layer to predict a second code 1811a2 (t=2). · Instantiate 12103 a learnable predictor layer to predict the third code 1811a3 (t=3). ·… · The current (last) learnable predictor layer instantiation 1210n that predicts the current nth code 1811an (t=n).

[0084] In the example, the instantiations (12101, 12102, 12103, ..., 1210n) of the learnable predictor layer are meant to be executed sequentially and / or iteratively for the sequence of codes 1811a1, 1811a2, 1811a3, ..., 1811an that must be predicted. Thus, after the current instantiation 1210n for predicting code 1811an, there is a new instantiation 1210(n+1) for predicting the subsequent code 1811a(n+1).

[0085] As shown in FIG. 8a, each instantiation 1210 may have inputs 1211 that may be selectably either: at least one (e.g., immediately preceding) preceding predicted code (e.g., 1811a1 is provided as input 1220'0 from the first instantiation 12101 to the second instantiation 12102, 1811a2 is provided as input 1220'1 from the second instantiation 12102 to the third instantiation 12103, ..., 1811a(n-1) is provided as input 1220'(n-1) from the (n-1)th instantiation (not shown) to the current nth instantiation 1210n); At least one (e.g., immediately preceding) preceding transformation code (e.g., 1820a'0 is provided to the first instantiation 12101, 1820a'1 is provided to the second instantiation 12102, ..., 1820a'(n-1) is provided to the current n-th instantiation 1210n). As can be seen in Figure 2, a learnable predictor layer instantiation 1210 may include at least one (e.g., two) learnable layers (e.g., 1212, 1214) having states, collectively referred to as 1222, including state 1 of the first layer 1212, referred to here as 12221, and state 2 of the second layer 1214, referred to as 12222.

[0086] State may be provided from a preceding instantiation (e.g., the immediately succeeding instantiation) to a succeeding instantiation (e.g., up to the current instantiation 1210n). For example, state 1222 of instantiation 12101 is provided to instantiation 12102 (in this case, state 12221 of the first layer 1212 of instantiation 12101 is provided to the first layer 1212 of the immediately succeeding instantiation 12102, and state 12222 of the second layer 1214 of instantiation 12101 is provided to the second layer 1214 of the immediately succeeding instantiation 12102). Similarly, the state of the predictor 1222 of instantiation 12102 (specifically, layers 1212 and 1214) is provided to instantiation 12103 (specifically, layers 1212 and 1214). Similarly, current instantiation 1210n receives state 1222 from a preceding instantiation (this is not shown in FIG. 2). Thus, when a code is predicted, it is predicted through an instantiation of a learnable predictor layer whose state takes into account the state of a preceding instantiation (e.g., the immediately preceding instantiation). For example, the current n-th predicted code 1811an is obtained through layers 1212 and 1214 of the current n-th instantiation 1210n, taking into account the state 1222 of the immediately preceding iteration 1210(n-1) (specifically, layers 1212 and 1214 of the immediately preceding iteration 1210(n-1)).

[0087] To predict a current code (e.g., 1811an), a current learnable predictable layer instantiation 1210n receives at input 1211 a selection from among the following: At least one preceding conversion code 1820'a(n-1) if at least one preceding packet is deemed to have been successfully received (thereby activating connection 1820a' of FIG. 8a). or at least one preceding prediction code 1220'(n-1) if at least one preceding packet is considered lost.

[0088] However, to predict the current code 1811an, the last learnable predictor layer instantiation 1210n receives states 1222 (12221, 12222) from at least one previous (e.g., immediately previous) iteration both when at least one previous packet is considered to have been received successfully and when at least one previous packet is considered to have been lost.

[0089] Thus, as can be seen, each instantiation 1210 has at its input 1211 either a previously transformed code 1202 (1820a', such as 1820a'0, 1820a'1, 1820a'2, 1820a'(n-1) etc.) or a previously predicted code (e.g., 1811a1 is provided to input 1211 of instantiation 12102 as 1220'0, 1811a2 is provided to input 1211 of instantiation 12103 as 1220', and 1220'(n-1) is provided as input 1211 of current instantiation 1210n). Thus, for each input 1211 of each instantiation 1210, either a previously corrected code 1820a' or a previously predicted code 1204 (1811a) is provided as an input to each iteration. In this example, it is primarily assumed that each duration receives the code and state from the previous iteration, even though some generalization is possible for previous iterations that are not the previous instantiation (iteration). Thus, when the current code (e.g., 1811an) is predicted, the most recently converted code (obtained from a correct packet) is taken into account, and if some previously received packets are not correctly retained, the previously predicted code is taken into account. In either case, state 1222 may be provided from each instantiation to the subsequent instantiation (e.g., the immediately following instantiation), so that something is inherited independently from other previous packets, regardless of whether the previous packets are corrected or not.

[0090] Consider a situation where the previous code n-1 has already been transformed by the quantization index transformer 1818a (because the previous packet was received correctly) to predict the current n-th code 1811an. In this case, to predict the current n-th code, the trainable predictor layer instantiation 1210n receives (as input to the latent value 1211) the previous transformed code 1820a'(n-1) (output by the transformer 1818a) rather than the previous predicted code 1120'(n-1) output by the previous iteration. However, the instantiation 1210(n-1) to predict the (n-1)th code is still executed. Because the (n-1)th code is obtained from the correct packet, it may be considered unnecessary to provide state 1222 from the (n-1)th instantiation 1210(n-1) to the current nth instantiation 1210n, since the immediately preceding (n-1)th code 1820a'(n-1) is converted from the correct packet, and therefore it may be considered unnecessary to inherit state 1222 from the previous iteration (e.g., 1210(n-1)). However, it is understood that by passing (bypassing) state 1222 from the previous iteration 1210(n-1) to the current iteration 1210n (even if the previous code was obtained from the correct packet), something from an earlier iteration (e.g., 1210(n-2), 1210(n-3), etc.) can be passed to the current iteration 1210n (more importantly, something from the previous (n-2), (n-3), etc. codes is inherited by the nth code). It is understood that to generate at least a part of the audio signal representation 1820a, the prediction may advantageously take into account not only the immediately preceding code (transformed or predicted), but also some further preceding codes that precede the immediately preceding code, thus increasing reliability since states are derived from preceding codes that are not the immediately preceding code.

[0091] For example, assume the following: The 0th and 1st previously converted codes (1820a'0 and 1820a'1) are obtained from correctly received packets, and therefore instantiations 12101 and 12102 provide the correct state 1222 to their immediately following instantiations 12102 and 12103, respectively. The second, third, ..., (n-1)th codes are taken from the corrupted packet, and therefore the second, third, ..., (n-1)th instantiations 12103, 12104, ..., 1210(n-1) (because they are taken from the incorrect packet, the transformed codes 1810'2, 1820'3, 1820'(n-2) cannot be input to these instantiations, but the previously predicted codes 1220'1, 1220'1, 1220'(n-2), respectively) provide the state 1222 for the immediately following instantiations 12103, 1204, ..., 1210n, respectively.

[0092] At first glance, these states provided to instantiations 12103, 1204, ... 1210n may be considered invalid because they are based on predictions and not on corrected data. However, it is understood that in this way, instantiations 12103, 12104, ... 1210(n-1) may provide (to subsequent instantiations) a state that has a good amount of accuracy despite being associated with a corrupted packet, because this state nevertheless inherits some amount of the previous correct state.

[0093] Each learnable predictor layer instantiation 1210 n may include at least a learnable convolutional unit 1216. That is, as may be obtained, at least one recurrent unit 1212, 1214 of a current learnable layer 1210 n may receive as input a state from at least one corresponding recurrent unit 1212, 1214 from at least one preceding learnable predictor layer instantiation and may output a state to at least one corresponding recurrent unit 1212, 1214 of at least one subsequent learnable predictor layer instantiation.

[0094] In some examples, each current learnable predictor layer instantiation has a series of learnable layers (e.g., each learnable layer in the series, except for the last, outputs a processed code to the immediately succeeding layer in the series, and the last learnable layer in the series outputs a code to the immediately succeeding learnable predictor layer instantiation), (e.g., for each learnable predictor layer instantiation, except for the last learnable predictor layer instantiation, each learnable layer in the series outputs its state to the corresponding learnable layer of the immediate learnable predictor layer instantiation).

[0095] In some examples, in each learnable predictor layer instantiation, the sequence of learnable layers includes at least one dimensionality-reduced learnable layer (1214) (e.g., GRU2) and at least one dimensionality-expanding learnable layer 1216 (e.g., FC) following the at least one dimensionality-reduced learnable layer (e.g., such that the output of the learnable predictor layer instantiation has the same dimensionality as the input of the learnable predictor layer instantiation).

[0096] In some examples (e.g., FIG. 2), at least one reduced-dimensionality learnable layer 1214 (e.g., GRU2) includes at least one learnable layer with state (e.g., such that each learnable predictor layer instantiation, except for the last learnable predictor layer instantiation, provides the state of at least one reduced-dimensionality learnable layer to at least one reduced-dimensionality learnable layer of the immediately succeeding learnable predictor layer instantiation).

[0097] In some examples (e.g., FIG. 2), the at least one dimension-expanded learnable layer (1216) (e.g., FC) includes at least one learnable layer without state (e.g., such that an instantiation of a predictor layer does not provide the state of the at least one dimension-expanded learnable layer to the at least one dimension-expanded learnable layer of an immediately subsequent instantiation of the learnable predictor layer). In some examples (e.g., Figure 2), the sequence of learnable layers is gated. In some examples (e.g., Figure 2), a series of learnable layers are gated through a softmax activation function.

[0098] Here, below we show a possible sequence of instantiations (e.g., 1210n) of a series of trainable predictor layers: -The input (input potential) 1211 may be either a previously converted code 1202, 1820a' (e.g., 1820a'(n-1)) or a previously predicted code 1204, 1811a (e.g., 1220'(n-1)). A first recurrent unit (e.g., iterative recurrent unit 1212) may transform input potential values ​​1211 from a first invention (e.g., 1,1,256) to a second dimension (the same dimension) to obtain an output 1215. The output 1215 is reduced from the dimension 1,1,256 to the second dimension 1,1,128.

[0099] In some examples, a gated unit is defined that has the following (e.g., state 1, 12222 is entered from the previous iteration): A convolutional layer 1216 (e.g., a layer with a state) that can have input values ​​1215 and dimensionally expanded output values ​​1217 (1, 1, 256) An activation function 1218 (e.g., softmax) to derive an estimated latent value 1220 that is used as the predicted code for the current packet (e.g., 1811an) and provided to the instantiation of the learnable predictor layer immediately following the code to be predicted. Of course, state may be provided from recurrent layers 1212 and 1214 to the corresponding recurring layer of the immediately subsequent learnable predictable layer instantiation.

[0100] 1a illustrates an example of a decoder 1100, which may be one of the encoders 1600 (such as 1600a and 1600b, as described above), which may include, for example, a trainable audio signal representation generator (denoted herein as an “NESC encoder”) 1104 (which may be instantiated, for example, by a trainable audio signal representation generator 1604). A residual quantization step (an example of which may be, for example, the quantizer 1608, described above or in FIGS. 6a and 6b) may utilize at least one codebook (e.g., a base codebook 1122, which may be the base codebook 1622 as described above; a first-ranking codebook 1124, which may be the low-rank codebook 1624 described above; a further low-rank codebook 1224a and a lowest-rank codebook 1124b). A bitstream is shown as bitstream 1630, but may also be bitstream 1830 in some examples. The decoder is shown at 1300 and may be one of the decoders 1600 and 1700 described above. The result may be a rendered audio signal 1724.

[0101] Figure 1b shows a more conceptual example of Figure 2, showing a codebook 1820, a transformed code 1810' (which may be provided to a trainable code predictor 1820aa via 1820a'), trainable layers 1212, 1214, which are gated recurrent units, and a convolutional layer 1216.

[0102] The examples of audio signal representation decoders 1710, 1810a, 1810b in Figures 7-8b are not necessarily separate from one another. For example, an audio signal representation decoder may simultaneously embody both audio signal representation decoders 1710 and 1810a. This is because the audio signal representation decoder may include both a redundancy storage unit 17100 (e.g., for storing redundant information 1714 from the bitstream 1630), as in Figure 8a, and a trainable code predictor 1818aa (e.g., for predicting code 1811a from the bitstream 1630, as in Figures 8a and 2, if redundant information has not yet been stored).

[0103] It should be noted that the examples of Figures 7 to 8b may be mixed. For example, a phonetic representation decoder may have both the trainable code predictor 1810aa of Figure 8a and the redundant information storage unit 17100 of Figure 7. In this case, the trainable code predictor 1810aa may be activated only when no redundant information is retrieved in the redundant information storage unit 17100. For example, if at least one current packet is to be considered lost, the redundant information storage unit 17100 may be retrieved. If the redundant information referencing the at least one current packet is retrieved, at least one index is retrieved from the redundant information referencing the current packet, and the quantization index converter converts the at least one retrieved index from the at least one codebook into a substitution code. Thus, the processing block may generate at least a portion of the audio signal by converting the at least one substitution code into at least a portion of the audio signal. Alternatively, if the redundant information is not retrieved, prediction is performed by the trainable code predictor 1810aa, and the predicted code is used as the code for the audio signal representation 1820a.

[0104] The same can be achieved, for example, by implementing an audio representation decoder having both the learnable index predictor 1810bb of FIG. 8b and the redundant information storage unit 17100 of FIG. 7. In this case, index prediction is performed only if no redundant information is retrieved in the redundant information storage unit 17100. In the above examples, reference is often made to the fact that redundant information is written to the bitstream. In some examples (e.g., in some examples implementing split quantization or residual quantization), the encoder 1600a or 1600b may write only the highest-ranked index to the bitstream as redundant information and skip low-ranked indexes. In this way, the audio signal representation decoder (e.g., 1710, 1810a, 1810b) may store and use the redundant information. The payload is reduced, and an acceptable reliability is nevertheless achieved.

[0105] Additionally or alternatively, for example in the example of FIG. 2, higher rank indices may be used (e.g., as previously transformed codes 1202 such as 1820a'0...1820a'n), thus reducing the computational complexity of the trainable predictor 1200 (1810a) since lower rank indices are not predicted. It should be noted that the learnable predictors (1200, 1810a, 1810b) may be trained to learn learnable parameters of at least one learnable predictor layer that minimizes the error of the predicted current code relative to the transformed code converted from a packet having the correct format by sequentially predicting the predicted current code or the respective current index from preceding and / or subsequent packets and comparing the predicted current code or the current code obtained from the predicted current code or the predicted index with the transformed code converted from a packet that has been successfully received.

[0106] 9 shows an example of an audio generator 10 (which may be one of 1700, 1800a, 1800b), in which the following may be recognized: An audio generator 10 that may be an example of audio generator 1700 and / or 1800a Audio signal representation decoder 1710 or 1810a Codebook 1820 (this may be one of codebooks 1620, 1622, 1624, 1620b, 1620a, 1122, 1124, 1124a, 1124b, etc.) Trainable code predictor 1810aa of FIG. 8a (but may also be implemented in the examples of FIG. 8b or FIG. 7) The redundancy information storage unit 17100 of FIG. 7 (but may also be implemented in the example of FIG. 8a or FIG. 8b) Bitstreams 1630, 1830 (also shown as 3 for all examples) Quantization index converters 1718, 1818a, 1818b (also called 313 for all examples) Decoded audio signal representations (codes, tensors, vectors) 1820a, 1820b, 1720 (also called 112 for all examples) Processing and / or rendering blocks 1722, 1822a, 1822b Output audio signals 1724, 1824a, 1824b (also referred to as 16 in all examples)

[0107] It is important to note that the examples of Figures 9-13 do not necessarily have to be implemented and that there are other techniques that may be implemented. The bitstream 3 (e.g., 1630 or 1830) (obtained at the input) may include frames (e.g., encoded as indices, e.g., encoded by the encoder 1600a or 1600b). An output audio signal 16 (e.g., one of 1724, 1824a, 1824b) may be obtained. The audio generator 10 (1700, 1800a, 1800b) may include a first data provider 702. The first data provider 702 may receive an input signal (input data) 14 (e.g., from an internal source, e.g., a noise generator or a storage unit, or from an external source, e.g., an external noise generator or an external storage unit, or even from data obtained from the bitstream 3). The input signal 14 may be noise, e.g., white noise, or a deterministic value (e.g., a constant). The input signal 14 may have multiple channels (e.g., 128 channels, but other numbers, e.g., greater than 64 channels, are also possible). The first data provider 702 may output first data 15. The first data 15 may be noise or may be derived from noise. The first data 15 may be input to at least one first processing block 50 (40). The first data 15 may be unrelated to the output audio signal 16 (e.g., 1724, 1824a, 1824b) (e.g., if it is derived from noise and therefore corresponds to the input signal 14). The at least one first processing block 50 (40) may adjust the first data 15, for example, using adjustments obtained by processing the bitstream 3 (e.g., 1630 or 1830), to obtain first output data 69. The first output data 69 may be provided to the second processing block 45. From the second processing block 45, the audio signal 16 (e.g., 1724, 1824a, 1824b) may be obtained (e.g., by PQMF synthesis). The first output data 69 may be in multiple channels.The first output data 69 may be provided to a second processing block 45 which may combine multiple channels of the first output data 69 to provide an output audio signal 16 (e.g., 1724, 1824a, 1824b) into one signal channel (e.g., after PQMF synthesis, e.g., indicated at 110 in Figures 4 and 10, but not shown in Figure 9).

[0108] As shown in Figure 9, in frame-by-frame branch 10a', an index may be provided to quantization index converter 313 (which may be one of 1718, 1818a, 1818b) to obtain a code (e.g., a vector, or more generally, a tensor) 112 (e.g., one of 1704, 1804a, 1804b). Code 112 (1704, 1804a, 1804b) may be multidimensional (e.g., two-dimensional, three-dimensional, etc.) and may here be understood to be in the same format (or a similar or similar format) as the format of the audio signal representation output by audio signal representation generator 1604 of Figure 6a or 6b. Thus, quantization index converter 313 (1718, 1818a, 1818b) may be understood as performing the inverse operation of quantizer 1608 of Figures 6a and 6b. As described above, the quantization index converter 313 (1718, 1818a, 1818b) may be connected to a trainable codebook (e.g., 1820, 1620, 1620a, 1620b, 1624, etc.), for example, as illustrated in Figures 7-8b. The quantization index converter 313 (1718, 1818a, 1818b) may be trained together with the quantizer, or more generally, together with other elements of the encoder 1600a, 1600b and / or audio generator 10 (1700, 1800a, 1800b). The quantization index converter 313 (1718, 1818a, 1818b) may operate on a frame-by-frame basis, for example, by considering a new index for each new frame it generates. Thus, each code (e.g., vector or more generally, tensor...) 112 (1704, 1804a, 1804b) has the same structure as each of the quantized latent representations and does not necessarily share the exact same values, but rather an approximation thereof.

[0109] The sample-by-sample branch 10b' may be updated for each sample, for example, at the output sampling rate and / or at a sampling rate lower than the final output sampling rate, using, for example, noise 14 or another input obtained from an external or internal source. It should also be noted that bitstream 3 (e.g., 1630 or 1830) is considered here to encode a mono signal, and that output audio signal 16 (e.g., 1724, 1824a, 1824b) and original audio signal 1602 are also considered to be mono signals. In the case of a stereo or multi-channel signal, such as a speaker signal or an Ambisonics signal, all of the techniques here are repeated for each audio channel (in the stereo case, there are two input audio channels 1, two output audio channels 16, etc.).

[0110] "Channel" is understood here in the context of convolutional neural networks, according to which a signal is viewed as an activation map having at least two dimensions: a number of samples (e.g., in the abscissa dimension, or e.g., the time axis), and a number of channels (e.g., in the ordinate direction, or e.g., the frequency axis). The first processing block 40 may operate like a conditional network (e.g., a conditional neural network) that is provided with data (e.g., code vectors or, more generally, tensors 112) from the bitstream 3 (e.g., 1630 or 1830) to generate conditions that modify the input data 14 (input signal). The input data (input signal) 14 (or any of its expansions) undergoes some processing to derive an output audio signal 16 (e.g., 1724, 1824a, 1824b) that is intended to be a version of the original input audio signal 1. The conditions of both the input data (input signal) 14 and their subsequent processed versions may be represented as activation maps that are applied to a learnable layer, for example by convolution. In particular, during its development towards speech (e.g., 1724, 1824a, 1824b), or more generally towards a generated audio signal 16, the signal may undergo upsampling (e.g., in Figure 10, from one sample 49 to multiple samples, e.g., several thousand samples), while the number of its channels 47 may be reduced (e.g., from 64 or 128 channels to a single channel).

[0111] The first data 15 may be obtained, for example, from an input (such as noise or a signal from an external signal) or from other internal or external sources (e.g., sample-by-sample branch 10b'). The first data 15 may be considered as an input to the first processing block 40 and may be an expansion of the input signal 14 (or may be the input signal 14). The first data 15 may be considered as a latent signal or a prior signal in the context of a conditional neural network (or more generally, a conditional trainable block or layer). Essentially, the first data 15 is modified according to conditions set by the first processing block 40 to obtain the first output data 69. The first data 15 may be in multiple channels, e.g., a single sample. Alternatively, the first data 15 provided to the first processing block 40 may have a single sample resolution, albeit with multiple channels. The multiple channels may form a set of parameters that may be associated with coding parameters encoded in the bitstream 3 (e.g., 1630 or 1830). Generally speaking, however, during processing in the first processing block 40, the number of samples per frame increases from a first number to a second higher number (i.e., the bit rate increases from a first bit rate to a second higher bit rate). On the other hand, the number of channels may be reduced from a first number of channels to a second lower number of channels. The conditions (described in greater detail below) used in the first processing block may be indicated by 74 and 75 and are generated by target data 12, which is then generated from target data 12 obtained from bitstream 3 (e.g., 1630 or 1830). It is noted that the conditions (adjustment feature parameters) 74 and 75 and / or the target data 12 may also be upsampled to fit (e.g., adapt) the dimensions of the version of target data 12. The unit providing first data 15 (either from an internal source, an external source, bitstream 3 (e.g., 1630 or 1830), etc.) is referred to herein as a first data provider 702.

[0112] As can be seen in FIG. 9 , the first processing block 40 may include a pre-trained trainable layer 710, which may be or include a recurrent trainable layer, e.g., a recurrent trainable neural network, e.g., a GRU. The pre-trained trainable layer 710 may generate target data 12 for each frame. The target data 12 may be at least two-dimensional (e.g., multi-dimensional), with multiple samples in each frame in the second dimension and multiple channels in each frame in the first dimension. In some examples, the target data 12 may be in the form of a spectrogram, e.g., a mel-spectrogram if the frequency scale is non-uniform and / or motivated by perceptual principles. If the sampling rate corresponding to the trained trainable layer to be fed is different from the frame rate, the target data 12 may be the same for all samples of the same frame, e.g., at the layer sampling rate. Other upsampling strategies may also be applied. The target data 12 may be provided to at least one conditioning learnable layer, shown here as having layers 71, 72, and 73 (see also FIG. 12 and below). The conditioning learnable layers 71, 72, and 73 may generate conditions, also referred to as conditioning feature parameters, to be applied to the first data 12 (some of which may be shown as β (beta) and γ (gamma) or numbers 74 and 75), and any upsampled data derived from the first data. The conditioning learnable layers 71, 72, and 73 may be in the form of a matrix having multiple channels and multiple samples for each frame. The first processing block 40 may include a denormalization (or styling element) block 77. For example, the styling element 77 may apply the conditioning feature parameters 74 and 75 to the first data 15. One example may be element-wise multiplication of the values ​​of the first data by condition β (which may act as a bias) and addition with condition γ (which may act as a multiplier). The styling element 77 may generate first output data 69 for each sample.

[0113] The decoder (audio generator) 10 (1700, 1800a, 1800b) may include a second processing block 45. The second processing block 45 may combine multiple channels of the first output data 69 to obtain an output audio signal 16 (e.g., 1724, 1824a, 1824b) (or its predecessor audio signal 44').

[0114] Reference is now made primarily to FIG. 11. Bitstream 3 (e.g., 1630 or 1830), which may be subdivided into multiple frames, is coded in the form of indices (e.g., as obtained from a quantizer). From the indices of bitstream 3 (e.g., 1630 or 1830), a code (e.g., a scalar, vector, or more generally, a tensor) 112 is obtained via a quantization index converter 313 (1810a, 1810b, 1718). The first and second dimensions are shown in code 112 in FIG. 11 (other dimensions may also exist). Each frame is subdivided into multiple samples in the abscissa direction (first inter-frame dimension). Different terms may refer to the "frame index" and "feature map depth" in the abscissa direction (first direction), as well as the "latent dimension or coding parameter dimension." In the ordinate direction (second intra-frame dimension), multiple channels are provided. The codes 112 (1820a, 1820b, 1720) may be used by a pre-conditioning learnable layer 710 (e.g., a recurrent learnable layer) to generate target data 12, which may be at least two-dimensional (e.g., multi-dimensional), such as in the form of a spectrogram (e.g., a mel-spectrogram). Each target data 12 may represent a single frame, and a sequence of frames may evolve over time in the abscissa direction (left to right) along the first inter-frame dimension. Several channels may be in the ordinate direction (second intra-frame dimension) of each frame. For example, different coefficients occur in different entries of each column, relative to coefficients associated with frequency bands. The conditioning learnable layers 71, 72, 73 generate feature parameters 74, 75 (β and γ). The abscissas (second intra-frame dimension) of β and γ are associated with different samples of the same frame, and the ordinates (first inter-frame dimension) are associated with different channels. In parallel, the first data provider 702 may provide the first data 15. The first data 15 may be generated sample by sample and may have many channels.In the styling element 77 (or more generally in the first adjustment block 40), the adjustment feature parameters β and γ (74, 75) may be applied to the first data 15. For example, an element-wise multiplication may be performed between the sequence of styling conditions 74, 75 (adjustment feature parameters) and the first data 15 or an expansion thereof. It is noted that this process may be repeated multiple times.

[0115] As is clear from the above, the first output data 69 generated by the first processing block 40 may be obtained as a two-dimensional matrix (or even as a tensor with more than two dimensions) having samples on the abscissa (first inter-frame dimension) and channels on the ordinate (second intra-frame dimension). Through the second processing block 45, an audio signal 16 having a single channel and multiple samples (e.g., similar in shape to the input audio signal), particularly in the time domain, may be generated. More generally, in the second processing block 45, the number of samples per frame (bit rate) of the first output data 69 may evolve from a second number of samples per frame (second bit rate) to a third number of samples per frame (third bit rate) that is greater than the second number of samples per frame (second bit rate). On the other hand, the number of channels of the first output data 69 may evolve from the second number of channels to a third number of channels that is less than the second number of channels. In other words, the bit rate (third bit rate) of the output audio signal 16 (e.g., 1724, 1824a, 1824b) may be higher than the bit rate (first bit rate) of the first data 15 and the bit rate (second bit rate) of the first output data 69, and the number of channels of the output audio signal 16 (e.g., 1724, 1824a, 1824b) may be lower than the number of channels (first channel number) of the first data 15 and the number of channels (second channel number) of the first output data 69. Models that process coded parameters frame-by-frame by juxtaposing the current frame to the previous frame already in that state are also called streaming or stream-wise models, and may be used as convolution maps for real-time and stream-wise applications such as speech coding.

[0116] Examples of convolutions are described herein below, and it can be understood that they may be used in the first processing block 40 (50) in any of the pre-conditioning trainable layers 710 (e.g., recurrent trainable layers), and generally in at least one of the conditioning trainable layers 71, 72, 73 and more. Generally speaking, a passed set of conditional parameters (e.g., for one frame) may be stored in a queue (not shown) to be subsequently processed by the first or second processing block while the first or second processing block, respectively, processes the previous frame.

[0117] Next, we will describe operations primarily performed in blocks downstream of the pre-conditioning trainable layer 710 (e.g., recurrent trainable layers). We consider target data 12 already obtained from the pre-conditioning trainable layer 710 and applied to the conditioning trainable layers 71-73 (which are then applied to styling elements 77). Blocks 71-73 and 77 may be embodied by a generator network layer 770. The generator network layer 770 may include multiple trainable layers (e.g., multiple blocks 50a-50h, see below).

[0118] FIG. 9 (and its embodiment in FIG. 10 ) illustrates an example of an audio decoder (generator) 10 that can decode (e.g., generate, synthesize) an audio signal (output signal) 16 from a bitstream 3 (e.g., 1630 or 1830) according to, for example, the present technology (also referred to as StyleMelGAN). The output audio signal 16 (e.g., 1724, 1824a, 1824b) may be generated based on an input signal 14 (also referred to as a latent signal, which may be noise, e.g., white noise (“first option”), or may be obtained from another source). The target data 12, as described above, may include (e.g., be) a spectrogram (e.g., a mel spectrogram), which provides, for example, a mapping of a sequence of time samples to a mel scale (e.g., obtained from a pre-trained trainable layer 710). The target data 12 and / or the first data 15 are generally processed to obtain speech that is recognizable as natural to a human listener. In the decoders 1700, 1800a, 1800b, the first data 15 obtained from the input is styled (e.g., in block 77) to have a vector (or more generally, a tensor) with acoustic features conditioned by the target data 12. Finally, the output audio signal 16 (e.g., 1724, 1824a, 1824b), if it is speech, is recognized as speech by a human listener. The input vector 14 and / or the first data 15 (e.g., noise obtained from internal or external sources) may be a 128×1 vector (single sample, e.g., time-domain sample or frequency-domain sample, and 128 channels), as in FIG. 10 (FIG. 10 shows the input signal 14 provided to the channel mapping 30; the first data provider 702 is not shown or considered to be the same as the channel mapping 30). In other examples, input vectors 14 of different lengths may be used.The input vector 14 may be processed in a first processing block 40 (e.g., under conditioning of target data 12 obtained from bitstream 3 (e.g., 1630 or 1830) via preconditioning layer 710). The first processing block 40 may include at least one, e.g., a plurality of processing blocks 50 (e.g., 50a...50h). In FIG. 10, eight blocks 50a...50h (each of which is also identified as a "TADEResBlock") are shown, although a different number may be selected in other examples. In many examples, processing blocks 50a, 50b, etc. provide gradual upsampling of the signal evolving from input signal 14 to final audio signal 16 (e.g., 1724, 1824a, 1824b) (e.g., at least some processing blocks, e.g., 50a, 50b, 50c, 50d, 50e, increase the sampling rate such that each of them increases the sampling rate (bit rate) at its output relative to the sampling rate at its input), while some other processing blocks (e.g., 50f-50h) (e.g., downstream with respect to those that increase the sampling rate, e.g., 50a, 50b, 50c, 50d, 50e) do not increase the sampling rate (bit rate). Blocks 50a-50h may be understood as forming a single block 40 (e.g., as shown in FIG. 9). In the first processing block 40, a training set of learnable layers (e.g., 71, 72, 73, although a different number is possible) may be used to process the target data 12 and input signal 14 (e.g., first data 15). Accordingly, training feature parameters 74, 75 (also referred to as gamma (γ) and beta (β)) may be obtained during training, for example, by convolution. Thus, the learnable layers 71-73 may be part of the weight layer of the training network. As mentioned above, the first processing blocks 40, 50 may include at least one styling element 77 (normalization block 77).At least one styling element 77 may output first output data 69 (if there are multiple processing blocks 50, multiple styling elements 77 may generate multiple components that may be added together to obtain a final version of the first output data 69). At least one styling element 77 may apply adjustment feature parameters 74, 75 to the input signal 14 (latent values) or first data 15 derived from the input signal 14.

[0119] The first output data 69 may have multiple channels. The generated audio signal 16 (e.g., 1724, 1824a, 1824b) may have a single channel. The audio generator (e.g., decoder) 10 may include a second processing block 45 (shown in FIG. 10 as including blocks 42, 44, 46, and 110). The second processing block 45 may be configured to combine multiple channels (shown in FIG. 10 as 47) of the first output data 69 (input as second input data or second data) to obtain an output audio signal 16 (e.g., 1724, 1824a, 1824b) in a single channel but in a sequence of samples (shown in FIG. 10 as 49).

[0120] The term "channel" should not be understood in the context of stereo audio, but rather in the context of a general neural network of trainable units (e.g., a convolutional neural network) or higher. For example, a sequence of channels may be provided, so that the input signal (e.g., latent noise) 14 may have 128 channels (in a time-domain representation). For example, if the signal has 40 samples and 64 channels, it may be understood as a matrix with 40 columns and 64 rows, and if the signal has 20 samples and 64 channels, it may be understood as a matrix with 20 columns and 64 rows (other representations are possible). Thus, the generated audio signal 16 (e.g., 1724, 1824a, 1824b) may be understood as a mono signal. If a stereo signal is generated, the disclosed technique is simply repeated for each stereo channel to obtain multiple audio signals 16 that are subsequently mixed.

[0121] At least the original input audio signal and / or the generated speech 16 may be a sequence of time-domain values. Conversely, the output of each (or at least one) of blocks 30, 50a-50h, 42, 44 may generally have different dimensions (e.g., two-dimensional or other multidimensional tensors). At least some of blocks 30, 50a-50e, 42, 44 may upsample signals (14, 15, 59, 69) evolving from input 14 (e.g., noise) to speech 16. For example, in the first block 50a of blocks 50a-50h, two upsampling operations may be performed. Examples of upsampling operations may include, for example, the following sequences: (1) repeating the same value, (2) inserting zeros, and (3) another repeat or inserting zeros plus linear filtering.

[0122] The generated audio signal 16 (e.g., 1724, 1824a, 1824b) may generally be a single channel signal. If multiple audio channels are required (e.g., for stereo sound reproduction), the procedure may in principle be repeated multiple times. Similarly, the target data 12 may also have multiple channels (e.g., in a spectrogram such as a mel-spectrogram) as generated by the pre-conditioning learnable layer 710. In some examples, the target data 12 may be upsampled (e.g., by a factor of 2, a power of 2, a multiple of 2, or a different factor greater than 2, e.g., 2.5 or a multiple thereof) to match the dimensionality of the signals (59a, 15, 69) evolving along subsequent layers (50a-50h, 42), e.g., to obtain dimensionally adjusted feature parameters 74, 75 that match the dimensionality of the signals.

[0123] When the first processing block 40 is instantiated in multiple blocks (e.g., 50a-50h), the number of channels may remain the same for at least some of the multiple blocks (e.g., from 50e to 50h, the number of channels remains unchanged in block 42). The first data 15 may have a first dimension or at least one dimension lower than the audio signal 16 (e.g., 1724, 1824a, 1824b). The first data 15 may have a total number of samples across all dimensions lower than the audio signal 16 (e.g., 1724, 1824a, 1824b). The first data 15 may have one dimension lower than the audio signal 16 (e.g., 1724, 1824a, 1824b), but may have a larger number of channels than the audio signal 16 (e.g., 1724, 1824a, 1824b).

[0124] The example may be implemented according to the paradigm of a generative adversarial network (GAN). The GAN includes a GAN generator 11 (FIG. 10) and a GAN discriminator 100 (FIG. 10). The GAN generator 11 attempts to generate an audio signal 16 (e.g., 1724, 1824a, 1824b) that is as close to a real audio signal as possible. The GAN discriminator 100 recognizes whether the generated audio signal 16 (e.g., 1724, 1824a, 1824b) is real or fake. Both the GAN generator 11 and the GAN discriminator 100 may be implemented as neural networks (or by other learnable techniques). The GAN generator 11 minimizes loss (e.g., by a gradient method or other method) and updates the tuning feature parameters 74, 75 (and / or a codebook) by taking into account the results of the GAN discriminator 100. The GAN discriminator 100 reduces its discrimination loss (e.g., by gradient methods or other methods) and updates its internal parameters. Thus, the GAN generator 11 is trained to generate a better audio signal 16, while the GAN discriminator 100 is trained to distinguish between a fake audio signal generated by the GAN generator 11 and a real signal 16. The GAN generator 11 may include the functions of the decoders 1700, 1800a, and 1800b, but not the functions of the GAN discriminator 100. Thus, in most of the above, the GAN generator 11 and the audio decoders 1700, 1800a, and 1800b may have more or less the same characteristics except for the characteristics of the discriminator 100. The audio decoders 1700, 1800a, and 1800b may include the discriminator 100 as an internal component. Thus, the GAN generator 11 and the GAN discriminator 100 may collaborate in constructing the audio decoders 1700, 1800a, 1800b. In examples where the GAN discriminator 100 is not present, the audio decoders 1700, 1800a, 1800b can be constructed by the GAN generator 11 alone.

[0125] As explained by the expression "training set of trainable layers," the audio decoders 1700, 1800a, 1800b may be obtained, for example, based on conditional information, according to the paradigm of a conditional neural network (e.g., a conditional GAN). For example, the conditional information may consist of target data (or an upsampled version thereof) 12 on which the training set of layers 71-73 (weight layers) is trained and from which training feature parameters 74, 75 are obtained. Thus, the styling element 77 is trained by the trainable layers 71-73. The same may be true for the pre-training layer 710.

[0126] An example of the encoder 1600a, 1600b (or audio signal representation generator 1610a, 1610b) and / or encoded audio signal representation decoder 1710, 1810a, 1810b (or more generally, audio generator) 10 may be based on a convolutional neural network. A small matrix (e.g., a filter or kernel), which may be, for example, a 3x3 matrix (or a 4x4 matrix, or a 1x1 matrix, or even less than 10x10, etc.), is convolved (convolved) with a larger matrix (e.g., channel x samples latent values ​​or input signals and / or spectrograms and / or spectrograms or upsampled spectrograms or more general target data 12), which means, for example, combinations (e.g., multiplications and sums of products, dot products, etc.) between the elements of the filter (kernel) and the elements of the larger matrix (activation map, or activation signal). During training, the element of the filter (kernel) that minimizes the loss is obtained (learned). During inference, elements of the filter (kernel) obtained during training are used. Examples of convolutions may be used in at least one of blocks 71-73, 61b, 62b (see below), 230, 250, 290, 429, 440, and 460. Note that instead of matrices, three-dimensional tensors (or tensors with more than three dimensions) may also be used. If the convolution is conditional, the convolution may not necessarily be applied to a signal evolving from the input signal 14 through intermediate signals 59a (15), 69, etc. toward the audio signal 16 (e.g., 1724, 1824a, 1824b), but may be applied to the target signal 14 (e.g., to generate adjustment feature parameters 74 and 75, successively applied to the first data 15, or to latent values, or to prior values, or to the signal evolving from the input signal toward the audio signal 16). In other cases (e.g., in blocks 61b, 62b, see below), the convolution may be unconditional, e.g., applied directly to signals 59a(15), 69, etc. evolving from input signal 14 to audio signal 16 (e.g., 1724, 1824a, 1824b).Both conditional and unconditional convolutions may be performed.

[0127] In some examples (in the decoder or encoder), it is possible to have an activation function (ReLu, TanH, softmax, etc.) downstream of the convolution, which may vary depending on the intended effect. ReLu may map the maximum value between 0 and the value obtained by the convolution (in practice, it keeps the same value if it is positive and outputs 0 if it is negative). Leaky ReLu may output x if x>0 and 0.1*x if x≦0, where x is the value obtained by the convolution (in some examples, instead of 0.1, another value may be used, such as a predetermined value within 0.1±0.05). TanH (which may be implemented, for example, in blocks 63a and / or 63b) may provide the hyperbolic tangent of the value obtained by the convolution, e.g., TanH(x)=(e x -e -x ) / (e x +e -x), where x is the value obtained in the convolution (e.g., in block 61b, see below). A softmax (e.g., applied in block 64b) may apply an exponential function to each element of the convolution result and normalize it by dividing by the sum of the exponential functions. A softmax may provide a probability distribution of the entries in the matrix resulting from the convolution (e.g., provided in 62b). After application of the activation function, a pooling step may be performed in some examples (not shown), while this may be avoided in other examples. For example, it is also possible to have a softmax-gated TanH function by multiplying (e.g., in 65b, see below) the result of the TanH function (e.g., obtained in 63b, see below) by the result of the softmax function (e.g., obtained in 64b). Multiple convolutional layers (e.g., training sets of learnable layers) may, in some examples, be downstream of one another and / or parallel to one another for efficiency. If application of activation functions and / or pooling is provided, they may be repeated at different layers (or, for example, different activation functions may be applied at different layers) (this may also be applied at the encoder).

[0128] In the audio signal representation decoder 1710, 1810a, 1810b (or audio generator 1700, 1800a, 1800b), the input signal 14 is processed in different steps to become a generated audio signal 16 (e.g., 1724, 1824a, 1824b) (e.g., under conditions set by the training set of the trainable layer or layers 71-73, and with respect to parameters 74, 75 learned by the training set of the trainable layer or layers 71-73). Thus, the input signal 14 (or its expanded version, i.e., first data 15) can be understood to evolve in a direction of processing (14 to 16) to become a generated audio signal 16 (e.g., 1724, 1824a, 1824b) (e.g., speech). The conditions are generated substantially based on preconditions (eg, 1630 or 1830) in the target signal 12 and / or bitstream 3 and training (to arrive at the most favorable set of parameters 74, 75).

[0129] It should also be noted that multiple channels of input signal 14 (or any of its developments) may be considered to have a set of learnable layers and associated styling elements 77. For example, each row of matrices 74 and 75 may be associated with a particular channel of the input signal (or one of its developments), e.g., obtained from a particular learnable layer associated with the particular channel. Similarly, styling elements 77 may be considered to be formed by a number of styling elements (each row of input signals x, c, 12, 76, 76', 59, 59a, 59b, etc.).

[0130] FIG. 10 illustrates an example of an audio decoder 10 (or, more generally, an audio generator) (which may embody audio decoders 1700, 1800a, 1800b), which may also comprise (e.g., be) a GAN generator 11 (see below). FIG. 10 illustrates a pre-conditioning trainable layer 710 (shown in FIG. 9) even if target data 12 is obtained from bitstream 3 (e.g., 1630 or 1830) via a pre-conditioning layer 710 (see above). The target data 12 may be a mel-spectrogram (or other tensor) obtained from the pre-conditioning trainable layer 710 (but may also be other types of tensor), the input signal 14 may be latent (prior) noise or a signal obtained from an internal or external source, and the output 16 may be speech. The input signal 14 may have only one sample and multiple channels (e.g., the number of channels may vary, such as 80, and thus is denoted as “x”). The input vector 14 may be obtained as a vector of 128 channels (although other numbers are possible). If the input signal 14 is noise ("first choice"), it has a zero-mean normal distribution and is given by the formula TIFF2026502158000001.tif1032 which may be random noise of dimension 128 with mean 0 and an autocorrelation matrix (square matrix 128×128) equal to the identity matrix I (although different choices may be made). Thus, in the example where noise is used as input signal 14, the noise can be completely decorrelated between channels and have variance 1 (energy). TIFF2026502158000002.tif721 may be realized for every 22528 generated samples (or other numbers may be chosen for different examples), and therefore the dimensions may be 1 on the time axis and 128 on the channel axis. In the example, the input signal 14 may be a constant value.

[0131] The input vector 14 may be processed in stages (e.g., in blocks 702, 50a-50h, 42, 44, 46, etc.) to evolve into speech 16 (the evolved signals are shown, for example, as different signals 15, 59a, x, c, 76', 79, 79a, 59b, 79b, 69, etc.).

[0132] Channel mapping may be performed in block 30. This may consist of or include a simple convolutional layer to change the number of channels, for example, from 128 to 64 in this case. Thus, block 30 may be trainable (in some instances, it may be deterministic). As can be seen, at least some of processing blocks 50a, 50b, 50c, 50d, 50e, 50f, 50g, and 50h (which collectively embody the first processing block 50 of FIG. 6) may increase the number of samples, for example, by performing upsampling (e.g., up to 2 upsampling) for each frame. The number of channels may remain the same (e.g., 64) along blocks 50a, 50b, 50c, 50d, 50e, 50f, 50g, and 50h. Samples may be, for example, samples per second (or other time unit), and at the output of block 50h, a sound above 16 kHz (e.g., 22 kHz) may be obtained. As mentioned above, a sequence of multiple samples may constitute one frame. Each of the blocks 50a-50h (50) may also be a TADEResBlock (a residual block in the context of TADE (Temporal Adaptive DEnormalization)). In particular, each block 50a-50h (50) may be conditioned by target data (e.g., code, which may be a tensor, such as a multidimensional tensor, having two, three, or more dimensions) 12 and / or a bitstream 3 (e.g., 1630 or 1830). In the second processing block 45, only a single channel may be acquired, and multiple samples are acquired in a single dimension (see also FIG. 11). As can be seen, another TADEResBlock 42 (further blocks 50a-50h) may be used (reducing the dimension to four single channels). Then, a convolutional layer 44 and an activation function (which may be, for example, TanH 46) may be executed. A (quasi-quadrature mirror filter) bank 110 may also be applied (possibly stored, rendered, etc.) to obtain a final output 16 .

[0133] At least one of the blocks 50a-50h (or each of them, specific examples) and 42 and the encoder layers 230, 240 and 250 (and 430, 440, 450, 460) may be, for example, a residual block. The residual trainable block (layer) may perform prediction on the residual component of a signal evolving from the input signal 14 (e.g., noise) to the output audio signal 16 (e.g., 1724, 1824a, 1824b). The residual signal is only a portion (residual component) of the main signal evolving from the input signal 14 to the output signal 16. For example, multiple residual signals may be added together to obtain the final output audio signal 16 (e.g., 1724, 1824a, 1824b). Other architectures may also be used.

[0134] FIG. 12 shows an example of one of the blocks 50a-50h (50). The blocks 50a-50h (50) may be replicas of each other, but as can be seen, once trained, each block 50 (50a-50h) may receive input of first data 59a, which may be either the first data 15 (or an upsampled version thereof, e.g., output by the upsampling block 30) or the output from a preceding block. For example, block 50b may receive input of the output of block 50a, block 50c may receive input of the output of block 50b, and so on. In the example, different blocks may operate in parallel with each other, and the results are added together. From FIG. 12, it can be seen that the first data 59a provided to block 50 (50a-50h) or 42 is processed, and its output is output data 69 (provided as input to the subsequent block). As indicated by line 59a', the principal component of the first data 59a actually bypasses most of the processing of the first processing blocks 50a-50h (50). For example, blocks 60a, 900, 60b, 902, and 65b are bypassed by principal component 59a'. The residual component 59a of the first data 59(15) may be processed to obtain a residual portion 65b' that is added to the principal component 59a' at adder 65c (described in FIG. 12 but not shown). The bypassed principal component 59a' and the addition at adder 65c may be understood as instantiating the fact that each block 50(50a-50h) processes an operation on a residual signal, which is then added to the principal portion of the signal. Thus, each of blocks 50a-50h can be considered a residual block. The addition at adder 65c does not necessarily have to be performed within the residual block 50(50a-50h). A single addition of multiple residual signals 65b' (each output by a respective residual block 50a-50h) may be performed (e.g., in a single adder block within the second processing block 45). Thus, different residual blocks 50a-50h may operate in parallel with each other.In the example of FIG. 12, each block 50 (50a-50h) may repeat its convolutional layer twice. The first denormalization block 60a and the second denormalization block 60b may be used in series. The first denormalization block 60a may include an instance of a styling element 77 to apply adjusted feature parameters 74 and 75 to the first data 59(15) (or its residual version 59a). The first denormalization block 60a may include a normalization block 76. The normalization block 76 may perform normalization along the channels of the first data 59(15) (e.g., its residual version 59a). Thus, a normalized version c(76′) of the first data 59(15) (or its residual version 59a) may be obtained. Thus, the styling element 77 may be applied to the normalized version c(76′) to obtain a denormalized (adjusted) version of the first data 59(15) (or its residual version 59a). The denormalization at element 77 may be obtained, for example, by element-wise multiplication of a matrix (or more generally, a tensor) γ (which embodies condition 74) and signal 76′ (or another version of the signal between the input signal and the speech), and / or by element-wise addition of a matrix (or more generally, a tensor) β (which embodies condition 75) and signal 76′ (or another version of the signal between the input signal and the speech). Thus, a denormalized version 59b (adjusted by the adjusted feature parameters 74 and 75) of first data 59(15) (or its residual version 59a) may be obtained.

[0135] Next, a gated activation 900 may be performed on the denormalized version 59b of the first data 59 (e.g., its residual version 59a). Specifically, two convolutions 61b and 62b may be performed (e.g., using a 3x3 kernel and an expansion factor of 1, respectively). Different activation functions 63b and 64b may be applied to the results of the convolutions 61b and 62b, respectively. The activation 63b may be TanH. The activation 64b may be softmax. The outputs of the two activations 63b and 64b may be multiplied together to obtain a gated version 59c of the denormalized version 59b of the first data 59 (or its residual version 59a). Subsequently, a second denormalization 60b may be performed on the gated version 59c of the denormalized version 59b of the first data 59 (or its residual version 59a). The second denormalization 60b may be similar to the first denormalization and therefore will not be described here. Subsequently, a second activation 902 may be performed. Here, the kernel may be 3×3, but the expansion factor may be 2. In either case, the expansion factor of the second gated activation 902 may be larger than that of the first gated activation 900. A training set of learnable layers 71-73 and styling elements 77 (e.g., obtained from pre-trained learnable layers) may be applied to the signal 59a (e.g., twice for each block 50a, 50b, ...). Upsampling of the target data 12 may be performed in an upsampling block 70 to obtain an upsampled version 12′ of the target data 12. The upsampling may be obtained by nonlinear interpolation, for example, using a factor of 2, a power of 2, a multiple of 2, or another value greater than 2. Thus, in some examples, spectrogram (eg, mel-spectrogram) 12' may have the same dimensions as (eg, match) the signal (76, 76', c, 59, 59a, 59b, etc.) that is conditioned by the spectrogram.In an example, the first and second convolutions in 61b and 62b downstream of TADE blocks 60a or 60b, respectively, may be performed with the same number of elements in the kernel (e.g., 9, e.g., 3 x 3). However, the second convolution in block 902 may have an expansion factor of 2. In an embodiment, the maximum expansion factor of a convolution may be 2.

[0136] As described above, the target data 12 may be upsampled to match the input signal (also called the latent signal or activation signal, e.g., 59, 59a, 76′, the signal evolving therefrom). Convolutions 71, 72, 73 may then be performed (with intermediate values ​​of the target data 12 denoted by 71′) to obtain parameters γ (gamma, 74) and β (beta, 75). The convolutions in any of 71, 72, 73 may also involve a rectified linear unit (ReLu) or a leaky rectified linear unit (leaky ReLu). The parameters γ and β may have the same dimensions as the activation signal (the signal is processed to transform the audio signal 16 (e.g., 1724, 1824a, 1824b) generated from the input signal 14 and, when in normalized form, represented here as x, 59, 59a, or 76′). Therefore, if the activation signal (x, 59, 59a, 76') has two dimensions, then γ and β (74 and 75) also have two dimensions, and each of them can be superimposed on the activation signal (the length and width of γ and β may be the same as the length and width of the activation signal). In the styling element 77, the adjustment feature parameters 74 and 75 are applied to the activation signal (which may be the first data 59a or 59b output by the multiplier 65a). However, it should be noted that the activation signal 76' may be a normalized version (e.g., norm block 76) of the first data 59, 59a, 59b (15), where the normalization is in the channel dimension. Also, the styling element 77 (γc + β, in FIG. 12) TIFF2026502158000003.tif824 10 may be instantiated as block 50 in FIG. 12. Next, for example, convolution layer 44 reduces the number of channels to 1, after which TanH 46 is performed to obtain speech 16. Output 44′ of blocks 44 and 46 may have a reduced number of channels (e.g., 4 channels instead of 64) and / or the same number of channels (e.g., 40) as the previous block 50 or 42. A PQMF synthesis (see also below) 110 is performed on the signal 44' to obtain an audio signal 16 (eg, 1724, 1824a, 1824b) in one channel.

[0137] [Quantization and conversion from index to code] First, note that it is not strictly necessary that one single index is used to map one single code (e.g., tensor). Techniques such as the following are possible: -Split Tensor Quantization In the encoder (e.g., 1600a, 1600b), the quantizer 1608 converts a single tensor into multiple indices, for example, by: Splitting a tensor into multiple subtensors (e.g., subvectors) (e.g., at specific coordinates or locations within the tensor) Provide one index per subtensor To this end, different codebooks for different parts of the tensor may be defined. In some cases, a main part of a tensor (e.g., a main subtensor) and at least one lower-rank part of the tensor (e.g., a low-rank subtensor) may be defined. Therefore, the quantizer 1608 transforms each subtensor in each index using a respective codebook. In the audio signal representation decoder (e.g., 1710, 1810a, 1810b), the quantization index converter converts the indices for each tensor, for example by: Convert each index into its respective subtensor. · Combine sub-tensors into one single tensor. Like the encoder, it may use different codebooks.

[0138] -Residual quantization In the encoder (e.g., 1600a, 1600b), the quantizer 1608 converts a single tensor into multiple indices, for example, by: · Iteratively decompose the current tensor into a principal part and at least one residual part (e.g., error). For each part of a tensor, a transformation may be performed using a specific index. Even in this case, multiple codebooks (e.g., a primary codebook and a residual codebook) may be used. In the audio signal representation decoder (e.g., 1710, 1810a, 1810b), the quantization index converter converts the indices for each tensor, for example by: Convert each index into a part of the tensor (maybe the same high-rank codebook as the encoder is used). Compose all the parts together (e.g., by addition) In some instances, some of the tensors may be constituents (e.g., adducts).

[0139] In the following, reference will be made specifically to residual quantization, even though similar concepts may be used for split quantization. Here we describe the operation of the quantizer 1608 (e.g., in Figure 6a or 6b) and the quantization index converter 313 (inverse quantizer or inverse quantizer) when it is a quantization index. Note that the quantizer may be input with a scalar, vector, or more generally, a tensor, and the quantization index converter 313 (1818a, 1818b, 1718) converts the index to at least one code (obtained from a codebook). The codebooks used may be, for example, codebooks 1622 and 1624 (also 1122, 1124, 1124a, 1124b in FIG. 1a).

[0140] The following conventions are used here: ·x is speech (more than just a typical input signal 1602 to be coded). E(x) is the output of the audio signal generator 1604, which may be a vector or, more generally, a tensor. Code (e.g., i z , i r ,i q ) is a set of indices (e.g., z, r, q) that refer to (e.g., point to) at least one codebook (e.g., z e , r e , q e ) is within.

[0141] index (e.g., i z , i r ,i q ) is written to bitstream 3 (e.g., 1630 or 1830) by quantizer 1608 and read by quantization index converter 313 (1818a, 1818b, 1718). A primary code (e.g., z) is chosen to approximate the value E(x). The first (if present) residual code (e.g., r) is chosen to approximate the residual E(x)-z. The second (if present) residual code (e.g., q) is chosen to approximate the residual E(x)-zr. The decoder (e.g., at the quantization index converter 313, 1718, 1818a, 1818b) extracts the index (e.g., i z , i r ,i q ), obtain a code (e.g., z,r,q), and reconstruct a tensor (e.g., a tensor representing a frame in the first audio signal representation 220 of the first audio signal 1), for example by summing the code (e.g., z+r+q) as a tensor 112. Dithering can be added to avoid potential clustering effects.

[0142] The quantizer 1608 of Figure 6a or 6b may associate with each tensor of the first multi-dimensional audio signal representation of the input audio signal 1602 or the processed version of the first multi-dimensional audio signal representation a code that best approximates the tensor in the codebook (e.g., the code that minimizes the distance from the tensor), allowing the index in the codebook associated with the code that minimizes the distance to be written to the bitstream 3. As mentioned above, at least one codebook may be defined according to a residual technique. (1) Main (base) codebook z e (e.g., 1622, 1122) may be defined as having multiple codes, such that the particular code z∈z in the codebook that is relevant to approximating the main part of the frame E(x) (input vector) output by block 290. e is selected. (2) A specific code r∈r that best approximates the residual E(x)-z of the main part of the input vector E(x). e An optional first residual codebook r having multiple codes is selected such that e(e.g., 1624, 1124) may be defined. (3) an optional second residual codebook q having multiple codes; e (e.g., 1124a), so that the first-rank residual E(x)-z e -r e A specific code q∈q that approximates e is selected. (4) Possible optional lower-rank residual codebooks. The codes in each codebook may be indexed according to an index, and the association between each code in the codebook and the index may be obtained by training. The bitstream 3 (e.g., 1630 or 1830) is written with the index of each part (main part, first residual part, second residual part). For example, it may have:

[0143] (1)z∈z e The first index i points to z (2) First residual r∈r e The second index i points to r (3) Second residual q∈q e The third index i points to r The codes z, r, q may have the dimensions of the output E(x) of the audio signal representation generator 1604 for each frame, but the index i z , i r ,i q may be their encoded versions (e.g., strings of bits, such as 10 bits). Therefore, there may be multiple residual codebooks, resulting in the second residual codebook q e associates codes (e.g., scalars, vectors, or more generally, tensors) representing the second residual portion of the first multi-dimensional audio signal representation of the input audio signal with indices to be coded in the audio signal representation, and defines a first residual codebook r eassociates a code representing a first residual portion of a frame of the first multi-dimensional audio signal representation with an index to be coded in the audio signal representation, wherein the second residual portion of the frame is residual (e.g., low rank) with respect to the first residual portion of the frame.

[0144] Again, the audio generator 1700, 1800a, 1800b (or the audio signal representation decoder 1710, 1810a, 1810b, or in particular the quantization index converter 1718, 1818a, 1818b) may perform the inverse operation. z , i r ,i q ) from a code in the codebook to a code (e.g., z,r,q).

[0145] For example, in the case of the residual above, the bitstream may indicate for each frame (1630, 1830) of bitstream 3: (1) Index (code) i z to code z∈z e The primary index i z forms the principal part z of a tensor (e.g., a vector) that approximates E(x). (2) A code r∈r to convert from index ir to code r e The first residual index (second index) i r forms the first residual part of a tensor (eg, a vector) that approximates E(x). (3) Index i q code q∈r to convert from q The second residual index (third index) i q forms the second residual part of a tensor (eg, a vector) that approximates E(x). The coded version (tensor version) 212 of the frame may then be obtained, for example, as the sum z+r+q.

[0146] [GAN discriminator] 13 may be used, for example, during training to obtain parameters 74 and 75 (or processed and / or normalized versions thereof) to be applied to input signal 12. Training may be performed before inference, and the parameters (e.g., 74, 75, and / or at least one codebook) may be stored, for example, in non-transitory memory and used thereafter (although in some examples, parameters 74 or 75 may also be calculated on the fly).

[0147] The GAN discriminator 100 is responsible for learning how to recognize a generated audio signal (e.g., an audio signal 16 (e.g., 1724, 1824a, 1824b) synthesized as described above) from an actual input signal (e.g., actual speech) 104. Thus, the role of the GAN discriminator 100 is primarily exercised during a training session (e.g., for learning parameters 72 and 73) and is considered the opposite of the role of the GAN generator 11 (which may be considered an audio decoder 1700, 1800a, 1800b without the GAN discriminator 100).

[0148] Generally speaking, the GAN discriminator 100 receives as input both a synthesized audio signal 16 (e.g., 1724, 1824a, 1824b) generated by a GAN decoder 1700, 1800a, 1800b (obtained from a bitstream 3 (e.g., 1630 or 1830), which may be generated by an encoder 1600a or 1600b from an input audio signal 1602) and an actual audio signal (e.g., actual speech) 104 obtained, for example, from a microphone or another source, and may process the signals to obtain a metric (e.g., loss) to be minimized. The actual audio signal 104 may also be considered a reference audio signal. During training, the operations described above for synthesizing speech 16 may be repeated, for example, multiple times, to obtain, for example, parameters 74 and 75.

[0149] In an example, instead of analyzing the entire reference audio signal 104 and / or the entire generated audio signal 16 (e.g., 1724, 1824a, 1824b), it is possible to analyze only a portion thereof (e.g., a portion, slice, window, etc.). The generated signal portions are obtained in random windows (105a-105d) sampled from the generated audio signal 16 (e.g., 1724, 1824a, 1824b) and the reference audio signal 104. For example, a random window function can be used, so which windows 105a, 105b, 105c, 105d are used is not predefined. Also, the number of windows is not necessarily four and may vary.

[0150] Within the windows (105a-105d), a PQMF (Pseudo Quadrature Mirror Filter) bank 110 may be applied, resulting in subbands 120. Thus, a decomposition (110) of a representation of the generated audio signal (16) or of a representation of the reference audio signal (104) is obtained. An evaluation block 130 may be used to perform the evaluation. Multiple evaluators 132a, 132b, 132c, and 132d (collectively designated 132) may be used (although a different number may be used). In general, each window 105a, 105b, 105c, and 105d may be input to a respective evaluator 132a, 132b, 132c, and 132d. Sampling of the random window (105a-105d) may be repeated multiple times for each evaluator (132a-132d). In an example, the number of times the random window (105a-105d) is sampled for each evaluator (132a-132d) may be proportional to the length of the generated audio signal representation or the reference audio signal representation (104). Thus, each of the evaluators (132a-132d) may receive as input one or several parts (105a-105d) of a representation of the generated audio signal (16) or a representation of the reference audio signal (104).

[0151] Each evaluator 132a-132d may be a neural network itself. Each evaluator 132a-132d may, in particular, follow the convolutional neural network paradigm. Each evaluator 132a-132d may be a residual evaluator. Each evaluator 132a-132d may have parameters (e.g., weights) that are adapted during training (e.g., in a manner similar to one of those described above).

[0152] 13, each estimator 132a-132d also performs downsampling (e.g., by 4 or another downsampling ratio). The number of channels may increase for each estimator 132a-132d (e.g., by 4, or in some examples, by the same number as the downsampling ratio). Upstream and / or downstream of the estimator may be convolutional layers 131 and / or 134. The upstream convolutional layer 131 may, for example, have a kernel with dimensions 15 (e.g., 5x3 or 3x5). The downstream convolutional layer 134 may, for example, have a kernel with dimensions 3 (e.g., 3x3).

[0153] During training, a loss function (adversarial loss) 140 may be optimized. The loss function 140 may include a fixed metric (e.g., obtained during a pre-training step) between the generated audio signal (16) and the reference audio signal (104). The fixed metric may be obtained by calculating one or several spectral distortions between the generated audio signal (16) and the reference audio signal (104). The distortion may be measured by taking into account the following: the magnitude or logarithmic magnitude of the spectral representation of the generated audio signal (16) and the reference audio signal (104), and / or - Different time or frequency resolution.

[0154] In an example, the adversarial loss may be obtained by randomly feeding and evaluating representations of the generated audio signal (16) or representations of the reference audio signal (104) by one or more evaluators (132). The evaluation may include classifying the fed audio signals (16, 132) into a predetermined number of classes that indicate a pre-trained classification level of naturalness of the audio signals (14, 16). The predetermined number of classes may be, for example, "REAL" versus "FAKE."

[0155] An example of a loss may be obtained as follows. TIFF2026502158000004.tif1193 where x is the true speech 104, z is the latent input 14 (which may be noise obtained from the bitstream 3 or another input, e.g., 1630 or 1830), and s is a tensor representing x (or more generally the target signal 12). D(...) is the output of the estimator over the distribution of probabilities (D(...)=0 means "definitely false" and D(...)=1 means "definitely true"). Spectral reconstruction loss TIFF2026502158000005.tif610 is still used for regularization to prevent the appearance of adversarial artifacts. The final loss can be, for example, TIFF2026502158000006.tif1643 Here, each i is the contribution of each evaluator 132a to 132d (e.g., each evaluator 132a to 132d has a different D i (providing TIFF2026502158000007.tif712 is the pre-trained (fixed) loss. During the training session, There is a search for the minimum value of TIFF2026502158000008.tif66, which may be expressed, for example, as follows: TIFF2026502158000009.tif1653 Other types of minimization may be performed.

[0156] In general, the smallest adversarial loss 140 is associated with the best parameters (e.g., 74, 75) applied to the styling elements 77. (1) Note that during the training session, the encoder 1600a or 1600b (or at least the audio signal representation generator 1604) may be trained together with the decoders 1700, 1800a, 1800b (or even more in a typical audio generator 10). Thus, along with parameters for the decoders 1700, 1800a, 1800b (or even more in a typical audio generator 10), parameters for the encoder 1600a or 1600b (or at least the audio signal representation generator 1604) may also be obtained. In particular, at least one of the following may be obtained by training: Weights (e.g., kernels) of the learnable layers 230, 250. (2) Weights of the recurrent learnable layer 240. (3) The weights of the learnable block 290, including the weights (e.g., kernels) of layers 429, 440, and 460. (4) A codebook (e.g., z e , r e , q e at least one of the following: A common way to train the encoder 1600a or 1600b and the decoder 1700, 1800a, 1800b together is to use a GAN, where the discriminator 100 an audio signal 16 generated from frames in the bitstream 3 actually generated by the encoder 1; and an audio signal 16 generated from frames of the bitstream that were not generated by the encoder 1.

[0157] [Generate at least one codebook] The codebook (e.g., z) used by the quantizer 1608 and / or the quantization index converters 1818a, 1818b, 1718 (313) e , r e , q e With particular attention to at least one of the following, there may be different ways of defining the codebook: During a training session, multiple bitstreams 3 (1630, 1830) may be generated by the quantizer 1608 and obtained by the quantization index converter 313 (1818a, 1818b, 1718). To encode a known frame representing a known audio signal, the index (e.g., i z , i r , i q ) are written into the bitstream (3). The training session may include evaluation of the audio signals 16 generated in the audio signal representation decoders 1800a, 1800b, 1700 with respect to the known input audio signal 1602 provided to the audio signal representation generators 1610a, 1610b, and the association of at least one codebook index is adapted to the frames of the encoded bitstream (e.g., by minimizing the difference between the generated audio signals 16 (e.g., 1724, 1824a, 1824b) and the known audio signal 1602).

[0158] When a GAN is used, the discriminator 100 distinguishes between audio signals 16 (e.g., 1724, 1824a, 1824b) generated from frames of bitstream 3 (1630, 1830) actually generated by the encoders 1600a, 1600b and audio signals 16 generated from bitstreams not generated by the encoders 1600a, 1600b.

[0159] During the training session, it is possible to define the length of the index for each index (e.g., 10 bits instead of 15 bits). Thus, training may provide at least a number of first bitstreams having a first bitlength and having first candidate indices associated with first known frames representing the known audio signal, the first candidate indices forming a first candidate codebook, and a number of second bitstreams having a second bitlength and having second candidate indices associated with known frames representing the same first known audio signal, the second candidate indices forming a second candidate codebook.

[0160] The first bit length may be longer than the second bit length (and / or the first bit length has higher resolution but occupies more bandwidth than the second bit length). The training session may include evaluating the generated audio signals obtained from the multiple first bitstreams compared to the generated audio signals obtained from the multiple second bitstreams, thereby selecting a codebook (e.g., such that the selected trainable codebook is the selected codebook between the first candidate codebook and the second candidate codebook) (e.g., there may be an evaluation of a first ratio between a metric measuring the quality of the generated audio signals from the multiple first bitstreams in terms of bit length versus a second ratio between a metric measuring the quality of the generated audio signals from the multiple second bitstreams in terms of bit rate, and selecting the bit length that maximizes the ratio) (e.g., this may be repeated for each of the codebooks, e.g., primary, first residual, second residual, etc.). The discriminator 100 may determine whether the output signal 16 generated using the second candidate codebook with a low bitlength index is higher than the false bitstream 3 (e.g., TIFF2026502158000010.tif6150 and / or by evaluating the error rate in the discriminator 100), and if positive, the second candidate codebook with a low bit-length index will be selected, otherwise the first candidate codebook with a high bit-length index will be selected.

[0161] Additionally or alternatively, the training session may be performed by using a first number of first bitstreams having first indices associated with first known frames representing the known audio signal, the first indices being in a first maximum number, and the first number of first candidate indices forming a first candidate codebook, and a second number of second bitstreams having second indices associated with known frames representing the same first known audio signal, the second number of second candidate indices forming a second candidate codebook, and the second indices being in a second maximum number different from the first maximum number.

[0162] [Consideration] [Principle of the Invention] We propose a DNN-based autoregressive network (PLCNet) for PLC, which can be deeply integrated with the previously proposed codec NESC ([7]). NESC is an end-to-end speech codec consisting of a neural encoder and a neural decoder. The neural encoder learns a latent representation from the speech signal and vector quantizes it at a bit rate of 3.2 kbps. The neural decoder uses the quantized representation as a training feature to synthesize the original signal. Our PLCNet operates on the latent representation of a pre-trained NESC model and predicts future latent representations for concealment. PLCNet (primarily shown in Figures 1b and 2, and Figures 8a and 8b) has shown good results (see Figures 4 and 5).

[0163] We also propose an FEC mode for NESC (e.g., on the encoder side in Figures 6a and 6b, or on the decoder side in Figure 7) so that past latent representations can be transmitted along with the current frame for concealment, helping PLCNet with larger burst errors. FEC is self-contained and is a separate account of at least one past frame. FEC can be utilized when (de)jitter buffer management is performed at the receiver side. It avoids low-cost retransmissions or the silencing or concealment of lost frames, thus significantly improving system resilience. The FEC mode may involve an additional bitrate of 0.8 kbps to 3.2 kbps, depending on the number of past frames to be transmitted to the decoder (in our implementation, between 1 and 4), the desired quality in case of packet loss, and / or the desired total bitrate. [PLCNet and FEC] PLCNet NESC may include residual quantization with, for example, four codebooks. o The first codebook (e.g., 1622, 1222, etc.) may be the primary representation that is capable of producing an audio signal of acceptable quality, thus making only the first codebook suitable for concealment. ○PLCNet may be trained independently of the codec. The use of GRU-like memory elements in PLCNet facilitates autoregressive feature generation for burst error concealment.

[0164] Forward Error Correction (FEC) New FEC for neural codecs. In some instances, this is possible due to the availability of future frames in the jitter buffer (see, e.g., Figures 3a and 3b). o Redundant frames contain information from previous primary frames and are transmitted along with the current primary frame. o The selection of past frames to be considered in a redundant frame may depend on jitter buffer length and / or network conditions, and may be called the "FEC offset". The redundant frame can contain a single codebook stage index (e.g., 0.8 kbps) or all codebook stage indexes (e.g., 3.2 kbps), depending on the desired quality of the compensated packet. Alternatively, the redundant frame may contain information about multiple past frames (e.g., four different past frames with four different FEC offsets for a 3.2 kbps payload on the primary frame) to more efficiently compensate in very poor network conditions. Option to have a dedicated codebook for redundant information trained in the latent representation.

[0165] [Listening test results] Reference may be made to FIG. -Datasets are MS Challenge PLC dataset, maximum burst losses of 120ms and 320ms. -4% to 30% frame error rate. -NESC PLC is the proposed method for hiding used in deeply integrated estimation of latent representations. The results are shown in Figure 4. -NESC / PLCNet is a NESC with a hidden overhead from the conventional technique using a separate dedicated generation network used as post-processing for lost frames (high complexity overhead).

[0166] [Improved / New FEC] - A new method of FEC specifically for neural codecs that uses redundant information from previous frames. The new FEC operates in the latent space, does not require additional training layers, and involves minimal structural changes to the neural coder, allowing for simple yet powerful integration. FEC is performed on several stages of the codebook or on a completely newly trained codebook. - A new method of FEC specifically for neural codecs that uses current redundant information per frame. The new FEC operates in the latent space, does not require additional training layers, and involves minimal structural changes to the neural coder, allowing for simple yet powerful integration. FEC is performed on several stages of the codebook or on a completely newly trained codebook.

[0167] [Improved / New PLC] - An autoregressive method for PLC in the latent feature domain that predicts future codebook indices and is trained independently of the neural codec. -Good concealment with burst sizes up to 120ms or more and error rates up to 30%.

[0168] [Summary of some aspects] In the above examples, some aspects relate to an audio signal representation decoder configured to decode an audio signal representation from a bitstream, wherein the bitstream is divided into a sequence of packets, and the audio signal representation decoder: a bitstream reader configured to sequentially read the sequence of packets (e.g., to extract at least one index in at least one current packet); a packet loss controller configured to check whether the current packet has been well received (e.g., it has the correct format) or whether it should be considered lost; a quantization index converter configured to, if the packet loss controller determines that the current packet has been well received (e.g., has the correct format), convert at least one index extracted from the current packet into at least one current code (e.g., vector / tensor) from at least one codebook, thereby forming at least a portion of the audio signal representation; The audio signal representation decoder is configured to generate at least one current code by prediction (e.g., code prediction or index prediction) from at least one previous code or index (e.g., the current code may be obtained by prediction from a previously obtained index or code, or the current index may be obtained by prediction from a previously obtained index or code) via at least one learnable predictor layer when the packet loss controller determines that the current packet should be considered lost (the prediction may be based on a previously predicted code or index, or a code converted from a correctly received index or from a previously predicted index), thereby forming at least a part of the audio signal representation.

[0169] (For example, if the packet loss controller determines that the at least one current packet has the correct format, there may be a processing and / or rendering block configured to generate at least a portion of the audio signal by converting at least one transformed code (e.g., via at least one learnable processing layer, at least one deterministic layer, or at least one learnable processing layer and at least one deterministic layer) into at least a portion of the audio signal, and a code predictor, wherein the processing block is configured to generate at least a portion of the audio signal by converting at least one predicted code (e.g., via at least one learnable processing layer, at least one deterministic layer, or at least one learnable processing layer and at least one deterministic layer).

[0170] In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one codebook associates indices with codes or portions of codes, such that a quantization index converter converts at least one index extracted from the current packet into at least one transformed code or at least a portion of a transformed code.

[0171] In the above examples, some aspects relate to an audio signal representation decoder, comprising at least one codebook (e.g., z e , r e , q e )teeth, A base codebook (e.g., z) that associates indices with the main parts of the code. e )and, Associate an index with the residual part of the code (e.g., the lower the rank, the more residual part of the code) with at least one low-rank codebook (e.g., the first low-rank codebook (e.g., r e ), e.g., a second low-rank codebook having a lower rank than the first low-rank codebook, and possibly a third low-rank codebook having a lower rank than the second low-rank codebook, a fourth low-rank codebook having a lower rank than the third low-rank codebook, and additional codebooks are possible; the at least one index extracted from the current packet includes at least one high-rank index and at least one low-rank index; The quantization index converter is configured to convert at least one high-rank index into a major part of the current code and convert at least one low-rank index into at least one residual part of the current code; The quantization index transformer is further configured to reconstruct the current code by adding the dominant portion to the at least one residual portion.

[0172] In the above examples, some aspects relate to an audio signal representation decoder configured to predict at least one current code from at least one higher-ranked index of at least one preceding or subsequent packet, but not from a lowest-ranked index of at least one preceding or subsequent packet. In the above examples, some aspects relate to an audio signal representation decoder configured to predict a current code from at least a high-rank index and at least one medium-rank index of at least one previous packet, but not from a lowest-rank index of at least one previous packet.

[0173] In the above examples, some aspects relate to an audio signal representation decoder configured to store redundant information written to packets of a bitstream but referencing different packets, the audio signal representation decoder being configured to store the redundant information in a temporary storage unit; The audio signal representation decoder searches the temporary storage unit if at least one current packet is to be considered lost, and if redundant information referencing the at least one current packet is retrieved, Retrieving at least one index from the redundant information by referring to the current packet; causing a quantization index transformer to transform at least one index retrieved from the at least one codebook into a permutation code; The processing block is configured to generate at least a portion of the audio signal by transforming at least one substitution code into at least a portion of the audio signal.

[0174] In the above examples, some aspects relate to an audio signal representation decoder, wherein the redundant information provides at least a high-rank index of at least one preceding or subsequent packet, but does not provide at least one of a low-rank index of at least one preceding or subsequent packet. In the above examples, some aspects relate to an audio signal representation decoder further comprising at least one trainable predictor configured to perform prediction, the at least one trainable predictor having at least one trainable predictor layer.

[0175] In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one learnable predictor is trained by sequentially predicting a predicted current code or a current index, respectively, from preceding and / or subsequent packets, and by comparing the predicted current code or the current code obtained from the predicted index with a transformed code converted from a successfully received packet, to learn learnable parameters of at least one learnable predictor layer that minimize an error of the predicted current code relative to a transformed code converted from a packet having the correct format.

[0176] In the above examples, some aspects relate to an audio signal representation decoder, wherein the at least one trainable predictor layer includes at least one recurrent trainable layer. In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one learnable predictor layer includes at least one gated recurrent unit. In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one learnable predictor layer has at least one state, and the at least one learnable predictor layer is instantiated iteratively along a plurality of successive learnable predictor layer instantiations such that, to predict a current code, a current learnable predictor layer instantiation receives a state from at least one previous learnable predictor layer instantiation that predicted at least one previous code for at least one previous packet.

[0177] In the above examples, some aspects relate to an audio signal representation decoder, where to predict a current code, a current instantiation of a trainable predictor layer receives at input at least one previous transformed code when at least one previous packet is deemed to have been received well and at least one previous predicted code when at least one previous packet is deemed to have been lost. In the above examples, some aspects relate to an audio signal representation decoder, where, to predict a current code, a current learnable predictor layer instantiation receives state from at least one previous iteration both when at least one previous packet is considered to have been received successfully and when at least one previous packet is considered to have been lost.

[0178] In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one trainable predictor layer is configured to predict a current code and / or receive state from at least one previous trainable predictor layer instantiation both when at least one preceding packet is considered to be received well and when at least one preceding packet is considered to be lost, to provide a predicted code and / or output state to at least one subsequent trainable predictor layer instantiation.

[0179] In the above examples, some aspects relate to an audio signal representation decoder, where the current instantiation of a learnable predictor layer includes at least one learnable convolution unit. In the above examples, some aspects relate to an audio signal representation decoder, where a current learnable predictor layer instantiation includes at least one learnable recurrent unit. In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one recurrent unit of a current learnable layer receives a state from at least one corresponding recurrent unit from at least one preceding learnable predictor layer instantiation and outputs a state to at least one corresponding recurrent unit of at least one subsequent learnable predictor layer instantiation.

[0180] In the above examples, some aspects relate to an audio signal representation decoder, where the current learnable predictor layer instantiation has a series of learnable layers (e.g., each learnable layer in the series, apart from the last, outputs a processed code to the immediately succeeding layer in the series, and the last learnable layer in the series outputs a code to the immediately succeeding learnable predictor layer instantiation) (e.g., for each learnable predictor layer instantiation, apart from the last learnable predictor layer instantiation, each learnable layer in the series outputs its state to the corresponding learnable layer of the immediate learnable predictor layer instantiation).

[0181] In the above examples, some aspects relate to an audio signal representation decoder, where, for a current learnable predictor layer instantiation, the series of learnable layers includes at least one dimensionality-reduced learnable layer (e.g., GRU2) and at least one dimensionality-expanding learnable layer (e.g., FC) following the at least one dimensionality-reduced learnable layer (e.g., such that the output of the learnable predictor layer instantiation has the same dimensionality as the input of the learnable predictor layer instantiation). In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one reduced-dimensionality learnable layer (e.g., GRU2) includes at least one learnable layer having state (e.g., such that each learnable predictor layer instantiation provides state of at least one reduced-dimensionality learnable layer to at least one reduced-dimensionality learnable layer of the immediately succeeding learnable predictor layer instantiation, separate from the last learnable predictor layer instantiation).

[0182] In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one dimension-expanded learnable layer (e.g., FC) includes at least one learnable layer without state (e.g., such that an instantiation of a predictor layer does not provide the state of the at least one dimension-expanded learnable layer to the at least one dimension-expanded learnable layer of an immediately subsequent instantiation of the learnable predictor layer). In the above examples, some aspects relate to an audio signal representation decoder in which a series of learnable layers are gated. In the above examples, some aspects relate to an audio signal representation decoder in which a series of learnable layers are gated through a softmax activation function.

[0183] In the above examples, some aspects relate to an audio signal representation decoder configured to decode an audio signal representation from a bitstream, wherein the bitstream is divided into a sequence of packets, and the audio signal representation decoder: a bitstream reader (e.g., index extractor) configured to sequentially read a sequence of packets from at least one current packet, at least one index of the at least one current packet, and redundant information about at least one preceding or following packet; redundancy information that allows for reconstructing at least one index in at least one preceding or following packet; and a packet loss controller PLC configured to check whether at least one current packet has been successfully received (e.g., has a correct format) or should be considered lost (e.g., has an incorrect format); a quantization index converter (e.g., when the PLC determines that the at least one current packet has the correct format) configured to convert at least one index of the at least one current packet into at least one current transformed code (e.g., a tensor, or in certain cases a vector, which should preferably be multi-dimensional) from at least one codebook, thereby forming part of the audio signal representation; and a redundant information storage unit (e.g., via at least one learnable layer or deterministic layer) configured to store redundant information and provide the stored redundant information on the at least one current packet to form part of the audio signal representation via the redundant information (which may, for example, include an index or part of an index to be transformed by a quantization index converter, or a code or part of a code that has already been transformed) when the PLC determines that the at least one current packet should be considered lost.

[0184] (For example, as part of the audio generator, the PLC may comprise a processing and / or rendering block configured to generate at least a portion of an audio signal by converting at least one converted code into at least a portion of the audio signal when the PLC determines that at least one current packet has the correct format (e.g., via at least one learnable processing layer, at least one deterministic layer, or at least one learnable processing layer and at least one deterministic layer)); (The processing block is configured to generate at least a portion of the audio signal by transforming at least one stored redundant information about at least one current packet into at least a portion of the audio signal (e.g., via at least one learnable processing layer, at least one deterministic layer, or at least one learnable processing layer and at least one deterministic layer)).

[0185] In the above examples, some aspects relate to an audio signal representation decoder, wherein the redundant information storage unit is configured to store at least one index from a preceding or subsequent packet as redundant information, so as to provide the stored at least one index to a quantization index converter when the PLC determines that at least one current packet should be considered lost.

[0186] In the above examples, some aspects relate to an audio signal representation decoder, wherein the redundancy information storage unit is configured to store at least one code previously extracted from a preceding or subsequent packet as redundant information, to bypass the quantization index converter using the stored code when the PLC determines that at least one current packet should be considered lost.

[0187] In the above examples, some aspects relate to an audio signal representation decoder, wherein at least one codebook associates indices with codes or portions of codes, such that a quantization index converter converts at least one index extracted from the current packet into at least one transformed code or at least a portion of a transformed code.

[0188] In the above examples, some aspects relate to an audio signal representation decoder, comprising at least one codebook (e.g., z e , r e , q e ) is a base codebook (e.g., z ) that associates indices with the main parts of the code. e ) and at least one low-rank codebook (e.g., the first low-rank codebook (e.g., r) that associates an index with the residual part of the code (e.g., the lower the rank, the more residual part of the code) e), for example, a second low-rank codebook having a lower rank than the first low-rank codebook, and possibly a third low-rank codebook having a lower rank than the second low-rank codebook, a fourth low-rank codebook having a lower rank than the third low-rank codebook, and additional codebooks are possible), the at least one index extracted from the current packet includes at least one high-rank index and at least one low-rank index, the quantization index converter is configured to convert the at least one high-rank index into a prime part of the current code and convert the at least one low-rank index into at least one residual part of the current code, and the quantization index converter is further configured to reconstruct the current code by adding the prime part to the at least one residual part.

[0189] In the above examples, some aspects relate to an audio signal representation decoder configured to generate or retrieve at least one current code from at least one higher-ranked index of at least one preceding or subsequent packet, but not from the lowest-ranked index of at least one preceding or subsequent packet. In the above examples, some aspects relate to an audio signal representation decoder configured to generate or retrieve a current code from at least a high-rank index of at least one preceding or subsequent packet and from at least one intermediate-rank index, but not generate or retrieve a current code from a lowest-rank index of at least one preceding or subsequent packet.

[0190] In the above examples, some aspects relate to an audio generator for generating an audio signal from a bitstream, the audio generator comprising an audio signal representation decoder as described above and further configured to generate the audio signal by converting the audio signal representation into the audio signal. In the above examples, some aspects relate to an audio generator further configured to render the generated audio signal.

[0191] In the above example, some aspects relate to an audio generator, a first data provider configured to provide, for a given frame, first data derived from an input signal (e.g., from an external or internal source, or from an audio signal representation) (the first data may have one single channel or multiple channels, and the first data may, for example, be completely unrelated to the target data and / or the audio signal representation, although in other examples the first data may be obtained from the audio signal representation and therefore have some relationship to the audio signal representation); a first processing block configured to receive, for a given frame, first data and output first output data in the given frame (the first output data may include a single channel or multiple channels); (e.g., the audio generator further comprises a second processing block configured to receive, for a given frame, the first output data or data derived from the first output data as the second data), Here, the first processing block is: (possibly with at least one pre-trained trainable layer configured to receive the audio signal representation or a processed version thereof, and for a given frame, output target data representing the audio signal in the given frame (e.g., with multiple channels and multiple samples for the given frame); at least one training learnable layer configured to process, for a given frame, target data from the decoded audio signal representation to obtain training feature parameters for the given frame; a styling element configured to apply the adjustment feature parameters to the first data or the normalized first data.

[0192] (The second processing block, if present, may be configured to combine multiple channels of the second data to obtain an audio signal), (The at least one pre-trained learnable layer may include at least one recurrent learnable layer (e.g., a gated recurrent learnable layer such as a gated recurrent unit (GRU)).) (eg, configured to derive an audio signal from the first output data or a processed version of the first output data). In the above examples, some embodiments relate to an audio generator configured such that the bit rate of the audio signal is greater than the bit rate of both the target data and / or the first data and / or the second data.

[0193] In the above examples, some aspects relate to an audio generator, wherein the second processing block is configured to increase the bit rate of the second data to obtain the audio signal (and / or the second processing block is configured to reduce the number of channels of the second data to obtain the audio signal). In the above examples, some aspects relate to an audio generator, wherein the first processing block is configured to upsample first data from a number of samples of a given frame to a second number of samples that is greater than the initial number of samples of the given frame. In the above examples, some aspects relate to an audio generator, wherein the second processing block is configured to upsample the second data obtained from the first processing block from a second number of samples for a given frame to a third number of samples for the given frame that is greater than the second number.

[0194] In the above examples, some aspects relate to an audio generator configured to reduce the number of channels of the first data from a first number of channels to a second number of channels of the first output data that is less than the first number of channels. In the above examples, some aspects relate to an audio generator, wherein the second processing block is configured to reduce the number of channels of the first output data obtained from the first processing block from the second number of channels of the audio signal to a third number of channels, the third number of channels being less than the second number of channels.

[0195] In the above examples, some aspects relate to an audio generator, and the audio signal is a mono audio signal. In the above examples, some aspects relate to an audio generator configured to derive an input signal from an audio signal representation. In the above examples, some aspects relate to an audio generator configured to derive an input signal from noise. In the above examples, some aspects relate to an audio generator, where the training set of learnable layers includes one or at least two convolutional layers.

[0196] In the above examples, some aspects relate to an audio generator, the audio generator further comprising at least one pre-trained trainable layer configured to receive the audio signal representation, or a processed version thereof, and to output, for a given frame, target data representative of the audio signal in the given frame (e.g., with multiple channels and multiple samples for the given frame). In the above examples, some aspects relate to an audio generator, wherein at least one pre-trained learnable layer is configured to provide target data as a spectrogram or a decoded spectrogram. In the above examples, some aspects relate to an audio generator, wherein a first convolutional layer is configured to convolve target data or upsampled target data using a first activation function to obtain first convolutional data.

[0197] In the above examples, some aspects relate to an audio generator, where the adjustment learnable layer and styling elements are part of a weight layer within a residual block of a neural network that includes one or more residual blocks. In the above examples, some aspects relate to an audio generator, wherein the audio generator further comprises a normalization element configured to normalize the first data. In the above examples, some aspects relate to an audio generator, the audio generator further comprising a normalization element configured to normalize the first data in a channel dimension. In the above examples, some aspects relate to an audio generator, and the audio signal is a vocalized audio signal.

[0198] In the above example, some aspects relate to an audio generator, where the target data is upsampled by a factor that is a power of 2, or another factor such as 2.5 or a multiple of 2.5. In the above examples, some aspects relate to audio generators in which target data is upsampled by non-linear interpolation. In the above examples, some aspects relate to an audio generator, wherein the first processing block further comprises a further learnable layer configured to process data derived from the first data using a second activation function, the second activation function being a gated activation function. In the above examples, some aspects relate to an audio generator, where the further set of learnable layers comprises one or more convolutional layers.

[0199] In the above examples, some aspects relate to an audio generator in which the second activation function is a softmax-gated hyperbolic tangent (TanH) function. In the above examples, some aspects relate to an audio generator in which the first activation function is a leaky rectified linear unit (leaky ReLu) function. In the above example, some aspects relate to an audio generator, where the convolution operation is performed with a maximum expansion factor of 2. In the above example, some embodiments relate to an audio generator comprising eight first processing blocks and one second processing block. In the above example, some embodiments relate to an audio generator in which the first data has one dimension lower than the audio signal. In the above examples, some aspects relate to audio generators where the target data is a spectrogram.

[0200] In the above example, some aspects relate to an encoder, an audio signal representation generator configured to generate an audio signal representation (e.g., using at least one learnable layer, e.g., a combination of learnable and deterministic layers) from an input audio signal via at least one learnable layer, where the audio signal representation comprises a sequence of tensors (each tensor may be a vector, but if the tensor is a vector, it shall have at least two dimensions and each tensor / vector may be a code); a quantizer configured to convert each current tensor of the sequence of tensors into at least one index, each index being obtained from at least one codebook that associates a plurality of tensors with a plurality of indices; a bitstream writer configured to write packets to the bitstream such that the current packet includes at least one index of a current tensor of the sequence of tensors; The encoder is configured to write redundant information of the current tensor into at least one preceding or following packet of the bitstream that is different from the current packet.

[0201] In the above examples, some aspects relate to an encoder, where at least one codebook associates portions of a tensor with indices such that a quantizer transforms a current tensor into multiple indices. In the above example, some aspects relate to an encoder that includes at least one codebook (e.g., z e , r e , q e )teeth, A base codebook that associates the main part of a tensor with an index (e.g., z e )and, At least one low-rank codebook (e.g., the first low-rank codebook, e.g., r) that associates the residual part of the tensor with an index. e , a second low-rank codebook having a lower rank than the first low-rank codebook, a third low-rank codebook having a lower rank than the second low-rank codebook, and a fourth low-rank codebook having a lower rank than the third low-rank codebook, and additional codebooks are possible); At least one current tensor has at least one principal part and at least one residual part; The quantizer is configured to transform a principal portion of the at least one current tensor into at least one high-rank index and to transform at least one residual portion of the at least one tensor into at least one low-rank index; The bitstream writer thereby writes both the high-rank index and at least one low-rank index into the bitstream.

[0202] In the above examples, some aspects relate to an encoder, wherein the encoder is configured to provide redundant information having a high-rank index of at least one preceding or subsequent packet but not having a lowest-rank index of at least the same at least one preceding or subsequent packet. In the above examples, some aspects relate to an encoder configured to transmit a bitstream to a receiver (eg, an audio generator) over a communication channel. In the above examples, some aspects relate to an encoder configured to monitor payload conditions of a communication channel to increase the amount of redundant information when the payload conditions of the communication channel exceed a predetermined threshold.

[0203] In the above examples, some aspects relate to an encoder configured to: transmit, for each current packet, only a high-rank index of at least one preceding or subsequent packet as redundant information when the payload in the communication channel is below a predetermined threshold; and transmit, for each current packet, both a high-rank index of at least one preceding or subsequent packet and at least some low-rank indexes of at least one preceding or subsequent packet as redundant information when the payload of the communication channel is above the predetermined threshold.

[0204] In the above examples, some aspects relate to an encoder configured to calculate a packet offset between a current packet and at least one preceding or subsequent packet having redundant information depending at least on the payload of a communication channel. In the above examples, some aspects relate to an encoder configured to calculate a packet offset between a current packet and at least one preceding or following packet having redundant information depending on at least an envisioned application.

[0205] In the above examples, some aspects relate to an encoder configured to calculate a packet offset between a current packet and at least one preceding or subsequent packet having redundant information in response to at least input provided by an end user. In the above examples, some aspects relate to an encoder, wherein the at least one codebook includes a redundant codebook that associates a plurality of tensors with a plurality of indices, and the encoder is configured to write redundant information of a current tensor in at least one preceding or subsequent packet of the bitstream that is different from the current packet as an index received from the at least one quantization codebook.

[0206] [Further characterization of the figure] 1a and 1b may refer in some examples to the neural end-to-end speech codec (Fig. 1a) and the proposed PLCNet (Fig. 1b). FIG. 2 may refer to a detailed block diagram of PLCNet in some examples (tensor dimensions are shown in parentheses). FIG. 3 may refer in some instances to the NESC forward error correction method. Figure 4 shows the MUSHRA listening test for PLC using NESC. Figure 5 shows the P.800 listening test for PLC and FEC using NESC.

[0207] [Variation] Several variations and / or additional or alternative aspects are discussed herein. The hardware or software implementation may be performed using a digital storage medium, such as cloud storage, floppy disk, DVD, Blu-Ray, CD, ROM, PROM, EPROM, EEPROM or FLASH memory, on which electronically readable control signals are stored, which cooperates (or is capable of cooperating) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable. Some examples according to the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein. Generally, examples of the present invention may be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code may be stored on, for example, a machine-readable carrier.

[0208] Another example comprises the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an example of a method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further example of a method is therefore a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. A further example is therefore a data stream or a signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may for example be configured to be transferred via a data communication connection, for example via the Internet. A further example comprises processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further example comprises a computer having installed thereon a computer program for performing one of the methods described herein.

[0209] In some examples, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus. The above examples are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented by way of illustration and description of the examples herein.

Claims

1. an audio signal representation decoder (1810, 1810a, 1810b) configured to decode an audio signal representation (1820a, 1820b) from a bitstream (1830, 1630) divided into a sequence of packets, comprising: a bitstream reader (1802a, 1892b) configured to sequentially read the sequence of packets (1830, 1630); a packet loss controller (1806a, 1806b) configured to check whether the current packet (1830, 1630) has been successfully received or should be considered lost; a quantization index converter (1818a, 1818b) configured to convert at least one index (1804a, 1804b) extracted from the current packet (1830, 1630) into at least one current code (1820a, 1820b) from at least one codebook, if the packet loss controller (1806a, 1806b) determines that the current packet (1830, 1630) has been successfully received, thereby forming at least a part of the audio signal representation (1820a, 1820b); If the packet loss controller (1806a, 1806b) determines that the current packet should be considered lost, the audio signal representation decoder (1810, 1810a, 1810b) is configured to generate at least one current code by prediction (1810a, 1810b) from at least one previous code or index via at least one learnable predictor layer, thereby forming at least a part of the audio signal representation (1820a, 1820b).

2. The decoder of claim 1 , wherein the quantization index converter is configured to convert a plurality of indices into a respective plurality of subtensors, and then combine the subtensors to obtain the at least one current code.

3. 3. The decoder of claim 1, wherein the quantization index converter is configured to convert a plurality of indices into at least one main part of a code and at least one residual part of a code, and then compose the main part of a code and the at least one residual part of a code to obtain the at least one current code.

4. 4. An audio signal representation decoder according to claim 1, wherein the at least one codebook associates an index with a code or part of a code, such that the quantization index converter converts the at least one index extracted from the current packet into the at least one transformed code or at least part of a transformed code.

5. The at least one codebook a base codebook that associates indices with key parts of the code or high-ranking subcodes; at least one low-rank codebook that associates an index with a residual portion of the code or a low-rank subcode; the at least one index extracted from the current packet includes at least one high-ranked index and at least one low-ranked index; 5. The audio signal representation decoder of claim 1, wherein the quantization index converter is configured to convert the at least one high-ranked index into a dominant part or a high-ranked sub-code of the current code and to convert the at least one low-ranked index into at least one residual part or a low-ranked sub-code of the current code, and the quantization index converter is further configured to reconstruct the current code by adding the dominant part to the at least one residual part or by combining the high-ranked sub-code with the at least one low-ranked sub-code.

6. 6. An audio signal representation decoder according to claim 5, configured to predict the current code from at least one higher-ranked index of at least the at least one preceding or subsequent packet, but not from at least one lowest-ranked index of the at least one preceding or subsequent packet.

7. 7. An audio signal representation decoder according to claim 5 or 6, configured to predict the current code from at least a high-rank index and at least one medium-rank index of said at least one previous packet, but not from the lowest-rank index of said at least one previous packet.

8. configured to store redundancy information written to packets of the bitstream but referencing different packets, the audio signal representation decoder being configured to store the redundancy information in a storage unit; said audio signal representation decoder comprising: first searching the storage unit if the at least one current packet is to be considered lost, and retrieving at least one index from the redundant information referencing the current packet if the redundant information referencing the at least one current packet is to be retrieved; 8. An audio signal representation decoder according to claim 1, configured to cause the quantization index converter to convert the at least one index retrieved from the at least one codebook into a substitution code to become part of the audio signal representation.

9. 9. An audio signal representation decoder according to claim 8, configured to perform the prediction of the at least one current code only if the redundant information is not retrieved in the storage unit.

10. 10. An audio signal representation decoder according to claim 8 or 9 when dependent on any one of claims 5 to 7, wherein the redundant information provides at least the high-rank index of the at least one preceding or subsequent packet but not at least one of the low-rank indexes of the at least one preceding or subsequent packet.

11. 11. An audio signal representation decoder according to claim 1, further comprising at least one trainable predictor (1200, 1810a, 1810b) configured to perform the prediction, said at least one trainable predictor (1200, 1810a, 1810b) having at least one trainable predictor layer (1210, 1212, 1214, 1216).

12. 12. The audio signal representation decoder of claim 11, wherein the at least one trainable predictor (1200, 1810a, 1810b) is trained by sequentially predicting a predicted current code or a respective current index from a preceding and / or subsequent packet and comparing the predicted current code or the current code obtained from the predicted index with a transformed code converted from a successfully received packet, thereby learning trainable parameters of the at least one trainable predictor layer that minimize an error of the predicted current code relative to the transformed code converted from the packet having the correct format.

13. 13. An audio signal representation decoder according to any one of claims 11 to 12, wherein said at least one trainable predictor layer (1210) comprises at least one recurrent trainable layer (1212, 1214).

14. 14. An audio signal representation decoder according to any one of claims 11 to 13, wherein said at least one trainable predictor layer (1210) comprises at least one gated recurrent unit (1212, 1214).

15. 15. An audio signal representation decoder according to any one of claims 11 to 14, wherein the at least one trainable predictor layer (1210) comprises or is part of at least one neural network.

16. the at least one trainable predictor layer has at least one state (1222, 12221, 12222); 16. An audio signal representation decoder according to claim 11, wherein the at least one learnable predictor layer (1210) is instantiated iteratively along successive learnable predictor layer instantiations (1210) such that a current learnable predictor layer instantiation (1210n) receives state (1222, 12221, 12222) from at least one previous learnable predictor layer instantiation that predicted at least one previous code for at least one previous packet to predict the at least one current code (1811a, 1811an).

17. 17. An audio signal representation decoder according to claim 16, wherein, to predict the at least one current code (1204, 1811a, 1811an), the instantiation (1210n) of the trainable predictor layer receives at its input (1211) the at least one previous transformed code (1820'an) in case the at least one previous packet is considered to have been received well and the at least one previous predicted code (1220'(n-1)) in case the at least one previous packet is considered to have been lost.

18. 18. An audio signal representation decoder according to claim 17, wherein to predict the current code (1811an), the current learnable predictor layer instantiation (1210n) receives the state (1222, 12221, 12222) from the at least one previous iteration both when the at least one previous packet is considered to have been received successfully and when the at least one previous packet is considered to have been lost.

19. 19. An audio signal representation decoder according to any one of claims 16 to 18, wherein the at least one trainable predictor layer (1210) is configured to predict the current code and / or receive the state (1222) from the at least one previous trainable predictor layer instantiation both when the at least one previous packet is considered to be well received and when the at least one previous packet is considered to be lost, in order to provide the predicted code and / or output the state to at least one subsequent trainable predictor layer instantiation.

20. 20. An audio signal representation decoder according to any one of claims 16 to 19, wherein the current learnable predictor layer instantiation (1210n) comprises at least one learnable convolution unit (1216).

21. 21. An audio signal representation decoder according to any one of claims 16 to 20, wherein the current learnable predictor layer instantiation (1210n) comprises at least one learnable recurrent unit (1212, 1214).

22. 22. The audio signal representation decoder of claim 21, wherein the at least one recurrent unit (1212, 1214) of the current learnable layer (1210 n) receives a state from a corresponding at least one recurrent unit (1212, 1214) from the at least one preceding learnable predictor layer instantiation and outputs a state to a corresponding at least one recurrent unit (1212, 1214) of the at least one subsequent learnable predictor layer instantiation.

23. 23. An audio signal representation decoder according to any one of claims 16 to 22, wherein the current trainable predictor layer instantiation comprises a sequence of trainable layers.

24. 24. An audio signal representation decoder as described in claim 23, wherein for the current instantiation of the learnable predictor layer, the series of learnable layers includes at least one reduced-dimensionality learnable layer (1214) and at least one increased-dimensionality learnable layer (1216) following the at least one reduced-dimensionality learnable layer.

25. 25. An audio signal representation decoder according to claim 24, wherein the at least one reduced-dimensionality trainable layer (1214) comprises at least one trainable layer having states.

26. 26. An audio signal representation decoder according to any one of claims 24 to 25, wherein the at least one dimensionality expansion trainable layer (1216) comprises at least one stateless trainable layer.

27. 26. An audio signal representation decoder according to any one of claims 23 to 25, wherein the sequence of trainable layers is gated.

28. 28. An audio signal representation decoder according to any one of claims 23 to 27, wherein the sequence of learnable layers is gated via a softmax activation function.

29. An audio signal representation decoder (1700) configured to decode an audio signal representation (1720) from a bitstream (1630) divided into a sequence of packets, said audio signal representation decoder (1700) comprising: a bitstream reader (1702) sequentially reading said sequence of packets (1630), and from said at least one current packet: at least one index (1704) of said at least one current packet; redundancy information (1612b, 1714) relating to at least one preceding or subsequent packet, the redundancy information (1714) enabling at least one index (1704) in at least one packet of said at least one preceding or subsequent packet or information provided by said index to be reconstructed; a bitstream reader (1702) configured to extract a packet loss controller (PLC) (1706) configured to check whether the at least one current packet has been successfully received or should be considered lost; a quantization index converter (1718) configured to convert the at least one index (1704) of the at least one current packet (1630) into at least one current transformed code (1720) of at least one codebook, thereby forming part of the audio signal representation (1720); a redundancy information storage unit (1710) configured to store the redundancy information (1714) and, when the PLC (1706) determines (1708) that the at least one current packet (1630) should be considered lost, to provide the stored redundancy information (1714) on the at least one current packet to form part of the audio signal representation (1720) via the redundancy information (1712), An audio signal representation decoder (1700).

30. 30. An audio signal representation decoder according to claim 29, wherein the quantization index converter is configured to convert a plurality of indices into a respective plurality of subtensors, and then combine the subtensors to obtain the at least one current code.

31. 31. An audio signal representation decoder according to claim 29 or 30, wherein the quantization index converter is configured to convert a plurality of indices into at least one dominant part of a code and at least one residual part of a code, and then compose the dominant part of a code and the at least one residual part of a code to obtain the at least one current code.

32. 32. An audio signal representation decoder according to any one of claims 29 to 31, wherein the at least one codebook associates an index with a code or part of a code, such that the quantization index converter converts the at least one index extracted from the current packet into the at least one transformed code or at least part of a transformed code.

33. The at least one codebook a base codebook that associates indices with key parts of the code or high-ranking subcodes; at least one low-rank codebook that associates an index with a residual portion of the code or a low-rank subcode; the at least one index extracted from the current packet includes at least one high-ranked index and at least one low-ranked index; 33. An audio signal representation decoder according to any one of claims 29 to 32, wherein the quantization index converter is configured to convert the at least one high-ranked index into a dominant part or a high-ranked sub-code of the current code and to convert the at least one low-ranked index into at least one residual part or a low-ranked sub-code of the current code, and further configured to reconstruct the current code by adding the dominant part to the at least one residual part or by combining the high-ranked sub-code with the at least one low-ranked sub-code.

34. 34. An audio signal representation decoder as claimed in any one of claims 29 to 33, configured to read signaling indicating a packet offset between the current packet and at least one preceding or subsequent packet having the redundant information depending at least on the payload of the communication channel, so that the redundant information storage unit (17100), or another component of the audio signal representation decoder, reconstructs the packet to which the redundant information refers and stores the redundant information associated with the packet to which the redundant information refers.

35. 35. An audio signal representation decoder according to any one of claims 29 to 34, wherein the redundant information storage unit is configured to store at least one index from a previous or subsequent packet as redundant information so as to provide the stored at least one index to the quantization index converter if the PLC determines that the at least one current packet should be considered lost.

36. 36. An audio signal representation decoder according to any one of claims 29 to 35, wherein the redundancy information storage unit is configured to store at least one code or part of it previously extracted from a preceding or subsequent packet as redundant information, and to use the stored code to bypass the quantization index converter when the PLC determines that the at least one current packet should be considered lost.

37. 37. An audio signal representation decoder according to any one of claims 29 to 36, wherein the at least one codebook associates an index with a code or part of a code, such that the quantization index converter converts the at least one index extracted from the current packet into the at least one transformed code or at least part of a transformed code.

38. The at least one codebook a base codebook that associates indices with key parts of the code; at least one low-rank codebook; the at least one index extracted from the current packet includes at least one high-ranked index and at least one low-ranked index; the quantization index converter is configured to convert the at least one high-rank index into a dominant part of the current code or a high-rank sub-code, and to convert the at least one low-rank index into at least one residual part of the current code or a high-rank sub-code; the quantization index converter is further configured to reconstruct the current code by adding the dominant portion to the at least one residual portion, or by combining the at least one high-rank subcode with the at least one low-rank subcode, or by combining the at least one high-rank subcode with the at least one low-rank subcode, or by combining the high-rank subcode with the at least one low-rank subcode.

38. An audio signal representation decoder according to any one of claims 29 to 37.

39. 39. An audio signal representation decoder according to claim 38, configured to generate or retrieve said at least one current code from said at least one higher-ranked index of at least one preceding or subsequent packet, but not to generate or retrieve said at least one current code from said at least one lowest-ranked index of said at least one preceding or subsequent packet.

40. 40. An audio signal representation decoder according to claim 38 or 39, configured to generate or retrieve the current code from at least the high-ranked index and from at least one medium-ranked index of said at least one preceding or subsequent packet, but not from the lowest-ranked index of said at least one preceding or subsequent packet.

41. 41. An audio generator for generating an audio signal from a bitstream, comprising an audio signal representation decoder according to any one of claims 1 to 40, and further configured to generate said audio signal by converting said audio signal representation into said audio signal.

42. 42. The audio generator of claim 41, further configured to render the generated audio signal.

43. a first data provider (702) configured to provide, for a given frame, first data (15) derived from an input signal (14); a first processing block (40, 50, 50a-50h) configured to receive, for the given frame, the first data (15) and output first output data (69) within the given frame; at least one training learnable layer (71, 72, 73) configured to process target data (12) from the decoded audio signal representation for the given frame to obtain training feature parameters (74, 75); a styling element (77) configured to apply the adjustment feature parameters (74, 75) to the first data (15, 59a) or to the normalized first data (59, 76′), 43. An audio generator according to claim 41 or 42.

44. 44. An audio generator according to any one of claims 41 to 43, configured to derive the input signal from noise (14).

45. 45. The audio generator of claim 41, further comprising at least one pre-trained trainable layer (710) configured to receive the audio signal representation (1720, 1820a, 1820b) and to output target data (12) representing the audio signal.

46. 46. ​​The audio generator of claim 45, wherein the at least one pre-trained trainable layer (710) is configured to provide the target data (12) as a spectrogram or a decoded spectrogram.

47. 47. The audio generator of any one of claims 41 to 46, wherein a first convolutional layer (71-73) is configured to convolve the target data (12) or upsampled target data using a first activation function to obtain first convolved data (71′).

48. 48. An audio generator according to any one of claims 41 to 47, further comprising a normalisation element (76) configured to normalise the first data (59a, 15).

49. 49. An audio generator according to any one of claims 41 to 48, wherein the target data (12) comprises a spectrogram.

50. an audio signal representation generator (1604) configured to generate, via at least one learnable layer, an audio signal representation (1606) as a representation of the audio signal (1602), the audio signal representation (1606) comprising a sequence of tensors (1606); a quantizer (1608) configured to convert each current tensor (1606) of the sequence of tensors into at least one index (1626), each index being obtained from at least one codebook (1620) that associates a plurality of tensors with a plurality of indices; a bitstream writer (1628) configured to write a packet to the bitstream (1630) such that a current packet includes the at least one index (1626) for the current tensor (1606) of the sequence of tensors, The encoder (1600) is configured to write redundant information (1612) of the current tensor (1606) to at least one preceding or subsequent packet of the bitstream (1630) that is different from the current packet, and / or to write redundant information (1612) of a tensor that is different from the current packet to the current packet.

51. 51. The encoder of claim 50, wherein the at least one codebook (1620, 1622, 1624) associates portions of a tensor with indices such that the quantizer (1608) transforms the current tensor (1606) into a plurality of indices (1626, 1623, 1625).

52. The at least one codebook a base codebook (1622) that associates key parts of tensors with indices; at least one low-rank codebook (1624) that associates residual portions of the tensor with indices (1623); the at least one current tensor (1606) having at least one principal part and at least one residual part; the quantizer (1608) is configured to transform the dominant portion of the at least one current tensor into at least one high-rank index (1623) and to transform the at least one residual portion of the at least one tensor into at least one low-rank index (1625); 52. The encoder of claim 50, wherein the bitstream writer (1628) therefore writes both the high-rank index (1623) and the at least one low-rank index (1625) to the bitstream (1620).

53. 53. The encoder of claim 52, configured to provide in the redundant information (1121) at least the high-rank index (1623) of the at least one preceding or subsequent packet, but not at least the lowest-ranked low-rank index (1625) of the at least one preceding or subsequent packet.

54. 54. An encoder according to any one of claims 50 to 53, configured to split the current tensor into a number of subtensors for quantizing each subtensor.

55. 55. An encoder according to any one of claims 50 to 54, configured to decompose the current tensor between a main part and at least one residual part to quantize the main part and at least one residual part.

56. 56. An encoder according to any one of claims 50 to 55, configured to transmit the bitstream (1630) to a receiver over a communication channel.

57. 57. The encoder of claim 56, configured to monitor the payload condition (1643) of the communication channel (1644) so ​​as to increase the amount of redundant information if the payload condition (1643) of the communication channel (1640) exceeds a predetermined threshold.

58. transmitting, for each current packet, only the high-rank index of at least one preceding or subsequent packet as redundant information when the payload in the communication channel (1640) is below the predetermined threshold; and / or 58. An encoder as claimed in claim 57 when dependent at least on claim 52, configured to transmit, for each current packet, as redundant information, both the high-rank index of at least one preceding or subsequent packet and at least a portion of the low-rank index of at least one preceding or subsequent packet, when the payload (1643) of the communication channel (1640) exceeds the predetermined threshold.

59. 59. An encoder according to claim 57 or 58, configured to calculate a packet offset between the current packet and the at least one preceding or subsequent packet having the redundant information in accordance at least with the payload of the communication channel.

60. 60. An encoder according to any one of claims 57 to 59, configured to calculate a packet offset between the current packet and at least one preceding or subsequent packet having the redundant information in function of at least the intended application.

61. 61. An encoder as claimed in any one of claims 57 to 60, configured to calculate a packet offset between the current packet and at least one preceding or subsequent packet having the redundant information in response to at least input provided by the end user.

62. 60. An encoder according to claim 58 or 59, configured to calculate a packet offset between the current packet and at least one preceding or subsequent packet having the redundant information depending on at least the payload of the communication channel, such that the packet offset is higher the higher the payload in the communication channel or the higher the error rate in the communication channel.

63. 63. The encoder of claim 50, wherein the at least one codebook includes a redundant codebook that associates multiple tensors with multiple indices, and the encoder is configured to write the redundant information of the current tensor in at least one preceding or subsequent packet of the bitstream that is different from the current packet as an index received from the at least one quantization codebook.

64. 1. A method for decoding an audio signal representation from a bitstream, comprising: The method comprises: Read a sequence of packets contained in the bitstream, starting with the current packet: at least one index of the current packet; extracting redundant information about at least one preceding or subsequent packet, the redundant information enabling at least one index within the at least one preceding or subsequent packet to be reconstructed; checking whether the current packet has been successfully received or should be considered lost; - converting said at least one index of said current packet into at least one current transformed code from at least one codebook, thereby forming part of said audio signal representation; storing said redundant information to form part of said audio signal representation via said redundant information, and providing said stored redundant information for said at least one current packet if said checking determines that said at least one current packet should be considered lost. A method for decoding an audio signal representation.

65. A method for decoding an audio signal representation (1820a, 1820b) from a bitstream (1830, 1630), the bitstream (1830, 1630) being divided into a sequence of packets, the audio signal representation decoder (1810, 1810a, 1810b) comprising: reading said sequence of packets (1830, 1630) sequentially; checking whether the current packet (1830, 1630) was successfully received or should be considered lost; if said checking determines that said current packet (1830, 1630) has been successfully received, converting at least one index (1804a, 1804b) extracted from said current packet (1830, 1630) into at least one current code (1820a, 1820b) from at least one codebook, thereby forming at least a part of said audio signal representation (1820a, 1820b); and generating at least one current code by prediction (1810a, 1810b) from at least one previous code or index via at least one learnable predictor layer when the packet loss controller (1806a, 1806b) determines that the current packet should be considered lost. A method for decoding an audio signal representation.

66. generating an audio signal representation as a representation of the audio signal via at least one learnable layer, the audio signal representation comprising a sequence of tensors; converting each current tensor of the sequence of tensors into at least one index, each index being obtained from at least one codebook that associates a plurality of tensors with a plurality of indices; writing a packet in the bitstream such that the current packet includes the at least one index of the current tensor in the sequence of tensors; writing redundancy information of the current tensor to at least one preceding or subsequent packet of the bitstream different from the current packet, and / or writing redundancy information of at least one tensor to be written to at least one preceding or subsequent packet of the bitstream different from the current packet to the current packet; method.

67. When executed by a computer, the computer From the current packet, at least one index of the current packet; extracting redundant information about at least one preceding or subsequent packet, the redundant information enabling at least one index within the at least one preceding or subsequent packet to be reconstructed; checking whether the current packet has been successfully received or should be considered lost; transforming said at least one index of said current packet into at least one current transformed code from at least one codebook, thereby forming part of said audio signal representation; - controlling the storage of said redundant information and retrieving said stored redundant information on said at least one current packet if said checking determines that said at least one current packet should be considered lost, in order to form part of said audio signal representation via said redundant information. A non-transitory storage unit that stores instructions.

68. When executed by a computer, the computer Read the sequence of packets (1830, 1630) sequentially, Check whether the current packet (1830, 1630) has been received successfully or should be considered lost; if the check determines that the current packet (1830, 1630) has been successfully received, converting at least one index (1804a, 1804b) extracted from the current packet (1830, 1630) into at least one current code (1820a, 1820b) from at least one codebook, thereby forming at least a part of the audio signal representation (1820a, 1820b); if the check determines that the current packet should be considered lost, generating at least one current code by prediction (1810a, 1810b) from at least one previous code or index via at least one trainable predictor layer; A non-transitory storage unit that stores instructions.

69. When executed by a computer, the computer generating an audio signal representation as a representation of the audio signal via at least one learnable layer, the audio signal representation comprising a sequence of tensors; transforming each current tensor of the sequence of tensors into at least one index, each index being obtained from at least one codebook that associates a plurality of tensors with a plurality of indices; causing a packet to be written to the bitstream such that the current packet includes the at least one index of the current tensor of the sequence of tensors; causing at least one preceding or subsequent packet of the bitstream different from the current packet to write redundant information of the current tensor, and / or causing the current packet to write redundant information of at least one tensor to be written to at least one preceding or subsequent packet of the bitstream different from the current packet; A non-transitory storage unit that stores instructions.

Citation Information

Patent Citations

  • Systems and methods for communicating redundant frame information

    JP2016539536A

  • Trained generative model speech coding

    WO2022159247A1