Multi-lag format for audio coding

Through the encoding format of subband envelope and autocorrelation information based on auditory resolution, the problem of low encoding efficiency of existing audio encoding systems is solved, efficient audio signal encoding and decoding is realized, and high sound quality is maintained.

CN114258569BActive Publication Date: 2025-07-11DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080058713.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-20
Filing Date
2020-08-18
Publication Date
2025-07-11
Estimated Expiration
2040-08-18

AI Technical Summary

Technical Problem

The existing high-quality audio encoding system requires relatively large amounts of data, has low encoding efficiency, and it is difficult to effectively utilize auditory-related frequency resolution and perception characteristics.

Method used

Using an encoding format based on auditory resolution of subband envelope and additional information, the audio signal is decomposed into multiple subbands through a filter group, combined with spectral envelope and autocorrelation information for encoding, and decoding using iterative or machine learning methods to generate and reconstruct audio signals.

Benefits of technology

It improves encoding efficiency, maintains high sound quality, and reduces the bit rate required for encoding, achieving efficient audio signal encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114258569B_ABST
    Figure CN114258569B_ABST
Patent Text Reader

Abstract

The present description relates to a method for encoding an audio signal. The method includes: generating a plurality of sub-band audio signals based on the audio signal; determining a spectral envelope of the audio signal; for each sub-band audio signal, determining autocorrelation information of the sub-band audio signal based on an autocorrelation function of the sub-band audio signal; and generating an encoded representation of the audio signal, the encoded representation including a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information of the plurality of sub-band audio signals. Further described is a method for decoding the audio signal from the encoded representation, as well as corresponding encoder, decoder, computer program, and computer-readable recording medium.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to the following priority applications: U.S. Provisional Application 62 / 889,118, filed on August 20, 2019 (Reference No.: D19076USP1) and European Application 19192552.8, filed on August 20, 2019 (Reference No.: D19076EP), which are incorporated herein by reference. Technical field

[0003] The present disclosure generally relates to a method of encoding an audio signal into an encoded representation and a method of decoding an audio signal from the encoded representation.

[0004] Although some embodiments will be described herein with particular reference to this disclosure, it should be understood that the present disclosure is not limited to such fields of use and can be applied in a broader context. Background art

[0005] Any discussion of background art throughout the disclosure should in no way be taken as an admission that such technology is well - known in the art or forms part of the common general knowledge in the art.

[0006] In high - quality audio coding systems, it is common for the largest part of the information to describe the detailed waveform attributes of the signal. A small part of the information is used to describe more statistically - defined features (such as the energy in a frequency band), or control data aimed at shaping quantization noise according to known simultaneous masking properties of hearing (e.g., side information in an MDCT - based waveform encoder, which conveys the quantizer step size and range information necessary to correctly inverse - quantize the data representing the waveform at the decoder). However, these high - quality audio coding systems require a relatively large amount of data to encode audio content, i.e., they have a relatively low coding efficiency.

[0007] There is a need for audio coding methods and apparatuses that can encode audio data with improved coding efficiency. Summary of the invention

[0008] The present disclosure provides a method of encoding an audio signal, a method of decoding an audio signal, an encoder, a decoder, a computer program, and a computer - readable storage medium.

[0009] According to a first aspect of the present disclosure, a method for encoding an audio signal is provided. Encoding can be performed for each of a plurality of sequential portions (e.g., groups of samples, segments, frames) of the audio signal. In some embodiments, these portions may overlap with each other. An encoded representation can be generated for each such portion. The method may include generating a plurality of subband audio signals based on the audio signal. Generating a plurality of subband audio signals based on the audio signal may involve spectral decomposition of the audio signal, which may be performed by a filter bank of band-pass filters (BPFs). The frequency resolution of the filter bank may be related to the frequency resolution of the human auditory system. For example, the BPF may be a complex-valued BPF. Alternatively, generating a plurality of subband audio signals based on the audio signal may involve spectral and / or temporal flattening of the audio signal, optionally windowing the flattened audio signal by a window function, and spectrally decomposing the resulting signal into a plurality of subband audio signals. The method may further include determining a spectral envelope of the audio signal. The method may further include, for each subband audio signal, determining autocorrelation information of the subband audio signal based on the autocorrelation function (ACF) of the subband audio signal. The method may further include generating an encoded representation of the audio signal, the encoded representation including a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information of the plurality of subband audio signals. For example, the encoded representation may be related to a portion of a bitstream. In some embodiments, the encoded representation may further include waveform information related to the waveform of the audio signal and / or one or more waveforms of the subband audio signals. The method may further include outputting the encoded representation.

[0010] Configured as described above, the proposed method provides an encoded representation of the audio signal that has very high encoding efficiency (i.e., requires a very low bitrate to encode the audio), but at the same time includes appropriate information for achieving very good sound quality after reconstruction. This is achieved by providing autocorrelation information of multiple subbands of the audio signal in addition to the spectral envelope. Notably, it has been shown that two values per subband (one lag value and one autocorrelation value) are sufficient to achieve high sound quality.

[0011] In some embodiments, the autocorrelation information of a given subband audio signal may include a lag value of the corresponding subband audio signal and / or an autocorrelation value of the corresponding subband audio signal. Preferably, the autocorrelation information may include both a lag value of the corresponding subband audio signal and an autocorrelation value of the corresponding subband audio signal. Wherein, the lag value may correspond to the delay value (e.g., abscissa) at which the autocorrelation function reaches a local maximum, and the autocorrelation value may correspond to the local maximum (e.g., ordinate).

[0012] In some embodiments, the spectral envelope may be determined at a first update rate, and the autocorrelation information of the plurality of subband audio signals may be determined at a second update rate. In such a case, the first update rate and the second update rate may be different from each other. The update rate may also be referred to as the sampling rate. In one such embodiment, the first update rate may be higher than the second update rate. Further, different update rates may be applied to different subbands, i.e., the update rates of the autocorrelation information of different subband audio signals may be different from each other.

[0013] By reducing the update rate of the autocorrelation information compared to the update rate of the spectral envelope, the coding efficiency of the proposed method can be further improved without degrading the sound quality of the reconstructed audio signal.

[0014] In some embodiments, generating the plurality of subband audio signals may include applying spectral and / or temporal flattening to the audio signal. Generating the plurality of subband audio signals may further include windowing the flattened audio signal by a window function. Generating the plurality of subband audio signals may also further include spectrally decomposing the windowed flattened audio signal into a plurality of subband audio signals. In such a case, for example, applying spectral and / or temporal flattening to the audio signal may involve generating a perceptually weighted LPC residual of the audio signal.

[0015] In some embodiments, generating the plurality of subband audio signals may include spectrally decomposing the audio signal. Then, determining the autocorrelation function of a given subband audio signal may include determining the subband envelope of the subband audio signal. Determining the autocorrelation function may further include envelope flattening the subband audio signal based on the subband envelope. The subband envelope may be determined by taking the magnitude values of the windowed subband audio signal. Determining the autocorrelation function may further include windowing the envelope-flattened subband audio signal by a window function. Determining the autocorrelation function may also further include determining (e.g., calculating) the autocorrelation function of the envelope-flattened windowed subband audio signal. The autocorrelation function may be determined for a real-valued (envelope-flattened windowed) subband signal.

[0016] Another aspect of the present disclosure relates to a method of decoding an audio signal from an encoded representation of the audio signal. The encoded representation may include a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information of each of a plurality of subband audio signals of the audio signal (or generated from the audio signal). The autocorrelation information of a given subband audio signal may be based on the autocorrelation function of the subband audio signal. The method may include receiving the encoded representation of the audio signal. The method may further include extracting the spectral envelope and the (multiple) autocorrelation information from the encoded representation of the audio signal. The method may also further include determining a reconstructed audio signal based on the spectral envelope and the autocorrelation information. The reconstructed audio signal may be determined such that the autocorrelation function of each of the plurality of subband audio signals of the reconstructed audio signal (or generated from the reconstructed audio signal) will satisfy a condition derived from the autocorrelation information of the corresponding subband audio signal of the audio signal (or generated from the audio signal). For example, the reconstructed audio signal may be determined such that for each subband audio signal of the reconstructed audio signal, the value of the autocorrelation function of the subband audio signal of the reconstructed audio signal (or generated from the reconstructed audio signal) at the lag value (e.g., delay value) indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal (or generated from the audio signal) substantially matches the autocorrelation value indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal. This may mean that the decoder is able to determine the autocorrelation function of the subband audio signal in the same manner as done by the encoder. This may involve any one, some, or all of flattening, windowing, and normalizing. In some embodiments, the reconstructed audio signal may be determined such that the autocorrelation information of each of the plurality of subband signals of the reconstructed subband audio signal (or generated from the reconstructed subband audio signal) will substantially match the autocorrelation information of the corresponding subband audio signal of the audio signal (or generated from the audio signal). For example, the reconstructed audio signal may be determined such that for each subband audio signal of the reconstructed audio signal (or generated from the reconstructed subband audio signal), the autocorrelation value and the lag value (e.g., delay value) of the autocorrelation function of the subband signal of the reconstructed audio signal substantially match the autocorrelation value and the lag value indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal (or generated from the audio signal). This may mean that the decoder is able to determine the autocorrelation information (i.e., lag value and autocorrelation value) of each subband signal of the reconstructed audio signal in the same manner as done by the encoder. Here, for example, the term "substantially matches" may mean a match up to a predefined margin. In those embodiments where the encoded representation includes waveform information, the reconstructed audio signal may be further determined based on the waveform information. The subband audio signals may be obtained, for example, by spectral decomposition of the applicable audio signal (i.e., the original audio signal on the encoder side or the reconstructed audio signal on the decoder side), or they may be obtained by flattening, windowing, and then spectral decomposition of the applicable audio signal.

[0017] Thus, it can be considered that the decoder operates according to an analysis-by-synthesis approach, in that it attempts to find a reconstructed audio signal z that will satisfy at least one condition derived from the coded representation h(x) of the coded audio signal, or whose coded representation h(z) will match substantially the coded representation h(x) of the original audio signal x, where h is the coding mapping used by the encoder. In other words, it can be considered that the decoder finds a decoding mapping d such that As has been found, if the coded representation that the decoder attempts to reproduce includes spectral envelope and autocorrelation information as defined in the present disclosure, then this analysis-by-synthesis approach yields results that are perceptually very close to the original audio signal.

[0018] In some embodiments, the reconstructed audio signal can be determined in an iterative process that starts from an initial candidate for the reconstructed audio signal and generates a corresponding intermediate reconstructed audio signal in each iteration. In each iteration, an update mapping can be applied to the intermediate reconstructed audio signal to obtain an intermediate reconstructed audio signal for the next iteration. The update mapping can be configured such that the autocorrelation function of the subband audio signal of the intermediate reconstruction (or generated from the intermediate reconstruction of the audio signal) more closely satisfies the conditions derived from the autocorrelation information of the corresponding subband audio signal of the audio signal (or generated from the audio signal), and / or such that the difference between the measured signal power of the subband audio signal of the reconstructed audio signal (or generated from the reconstructed audio signal) and the signal power of the corresponding subband audio signal of the audio signal (or generated from the audio signal) indicated by the spectral envelope is decreased iteration by iteration. If both autocorrelation information and spectral envelope are considered, an appropriate difference metric can be defined for the degree of satisfaction of the conditions and the difference between the signal powers of the subband audio signals. In some implementations, the update mapping can be configured such that the difference between the coded representation of the intermediate reconstructed audio signal and the coded representation of the audio signal becomes gradually smaller iteration by iteration. To this end, an appropriate difference metric for the coded representation (including spectral envelope and / or autocorrelation information) can be defined and used. The autocorrelation function of the subband audio signal of the intermediate reconstructed audio signal (or generated from the intermediate reconstructed audio signal) can be determined in the same way as the encoder does for the subband audio signal of the audio signal (or generated from the audio signal). Similarly, the coded representation of the intermediate reconstructed audio signal can be the coded representation that would be obtained if the intermediate reconstructed audio signal had undergone the same coding technique that resulted in the coded representation of the audio signal.

[0019] This iterative method allows for a simple and efficient implementation of the above-described analysis-by-synthesis approach.

[0020] In some embodiments, determining a reconstructed audio signal based on spectral envelope and autocorrelation information may include: applying a machine learning-based generative model that receives the spectral envelope of an audio signal and the autocorrelation information of each of a plurality of subband audio signals of the audio signal as inputs and generates and outputs the reconstructed audio signal. In those embodiments where the encoded representation includes waveform information, the machine learning-based generative model may further receive the waveform information as an input. This means that the machine learning-based generative model can also be adjusted / trained using the waveform information.

[0021] This machine learning-based approach allows for a very efficient implementation of the above-described analysis-by-synthesis approach and enables a reconstructed audio signal that is perceptually very close to the original audio signal.

[0022] Another aspect of the present disclosure relates to an encoder for encoding an audio signal. The encoder may include a processor and a memory coupled to the processor, wherein the processor is adapted to perform the method steps of any one of the encoding methods described throughout the present disclosure.

[0023] Another aspect of the present disclosure relates to a decoder for decoding an audio signal from an encoded representation of the audio signal. The decoder may include a processor and a memory coupled to the processor, wherein the processor is adapted to perform the method steps of any one of the decoding methods described throughout the present disclosure.

[0024] Another aspect relates to a computer program including instructions for performing the method steps of any method described throughout the present disclosure when the instructions are executed.

[0025] Another aspect of the present disclosure relates to a computer-readable storage medium storing the computer program according to the foregoing aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Example embodiments of the present disclosure will now be described by way of example only with reference to the accompanying drawings, in which:

[0027] Figure 1 is a block diagram schematically illustrating an example of an encoder according to an embodiment of the present disclosure,

[0028] Figure 2 is a flowchart illustrating an example of an encoding method according to an embodiment of the present disclosure,

[0029] Figure 3 schematically illustrates an example of a waveform that may exist in Figure 2 the framework of the encoding method,

[0030] Figure 4is a block diagram schematically illustrating an example of a synthesis approach by analysis for determining a decoding function,

[0031] Figure 5 is a flowchart illustrating an example of a decoding method according to an embodiment of the present disclosure,

[0032] Figure 6 is an illustration of Figure 5 an example of steps in a decoding method of

[0033] Figure 7 is a block diagram schematically illustrating another example of an encoder according to an embodiment of the present disclosure, and

[0034] Figure 8 is a block diagram schematically illustrating an example of a decoder according to an embodiment of the present disclosure. Detailed Description

[0035] Introduction

[0036] High-quality audio coding systems typically require a relatively large amount of data to encode audio content, i.e., have a relatively low coding efficiency. Although the development of tools such as noise filling and high-frequency regeneration has shown that waveform descriptive data can be partially replaced by a smaller set of control data, no high-quality audio codec mainly relies on perceptually relevant features. However, the improvement of computing power and the latest progress in the field of machine learning have increased the feasibility of decoding audio mainly from arbitrary encoder formats. The present disclosure presents examples of such encoder formats.

[0037] Broadly speaking, the present disclosure presents a coding format based on auditory resolution-inspired subband envelopes and additional information. The additional information includes a single autocorrelation value and a single lag value per subband (and per update step). The envelope can be calculated at a first update rate and the additional information can be sampled at a second update rate. Decoding of the coding format can be performed using a synthesis approach by analysis, for example, which can be implemented by iterative or machine learning-based techniques.

[0038] Encoding

[0039] The coding format (coding representation) presented in the present disclosure can be referred to as a multi-lag format because it provides one lag per subband (and update step). Figure 1 is a block diagram schematically illustrating an example of an encoder 100 for generating a coding format according to an embodiment of the present disclosure.

[0040] The encoder 100 receives a target sound 10 corresponding to an audio signal to be encoded. The audio signal 10 may include a plurality of sequential or partially overlapping portions (e.g., groups of samples, segments, frames, etc.) processed by the encoder. The audio signal 10 is spectrally decomposed by a filter bank 15 into a plurality of sub-band audio signals 20 in corresponding frequency sub-bands. For example, the filter bank 15 may be a filter bank of band-pass filters (BPFs), which may be complex-valued BPFs. For audio, it is natural to use a BPF filter bank having a frequency resolution related to the human auditory system.

[0041] At an envelope extraction block 25, the spectral envelope 30 of the audio signal 10 is extracted. For each sub-band, the power is measured at a predetermined time step as a basic model of the cochlear auditory envelope or excitation pattern generated by the input sound signal, thereby determining the spectral envelope 30 of the audio signal 10. That is, the spectral envelope 30 can be determined based on the plurality of sub-band audio signals 20, for example, by measuring (e.g., estimating, calculating) the respective signal powers of each of the plurality of sub-band audio signals 20. However, the spectral envelope 30 can be determined by any suitable alternative means, such as, by way of example, a linear prediction coding (LPC) description. In particular, in some embodiments, the spectral envelope can be determined from the audio signal before spectral decomposition by the filter bank 15.

[0042] Optionally, the extracted spectral envelope 30 may be subjected to downsampling at a downsampling block 35, and the downsampled spectral envelope 40 (or spectral envelope 30) is output as part of an encoded format or encoded representation of the audio signal 10 (the applicable portion).

[0043] Only the reconstructed signal reconstructed from the spectral envelope may still lack sound quality. To address this issue, the present disclosure proposes to include a single value (i.e., the ordinate and abscissa) of the autocorrelation function of the (possibly envelope-flattened) signal of each subband, which results in a significant improvement in sound quality. To this end, the subband audio signal 20 is optionally flattened (envelope flattened) at the divider 45 and input to the autocorrelation block 55. The autocorrelation block 55 determines the autocorrelation function (ACF) of its input signal and outputs the corresponding number of autocorrelation information 50 for the corresponding subband audio signal 20 based on the ACF of each subband audio signal 20 (i.e., each subband). The autocorrelation information 50 for a given subband includes a representation 50 of the lag value T and the autocorrelation value ρ(T) (e.g., composed of it). That is, for each subband, a lag value T and the corresponding (possibly normalized) autocorrelation value (ACF value) ρ(T) are output (e.g., transmitted) as the autocorrelation information 50, which is part of the coded representation. Wherein, the lag value T corresponds to the delay value at which the ACF reaches a local maximum, and the autocorrelation value ρ(T) corresponds to the local maximum. In other words, the autocorrelation information for a given subband can include the delay value (i.e., the abscissa) and the autocorrelation value and (i.e., the ordinate) of the local maximum of the ACF.

[0044] The coded representation of the audio signal thus includes the spectral envelope of the audio signal and the autocorrelation information of each subband. The autocorrelation information for a given subband includes a representation of the lag value T and the autocorrelation value ρ(T). The coded representation corresponds to the output of the encoder. In some embodiments, the coded representation may additionally include waveform information related to the waveform of the audio signal and / or one or more waveforms of the subband audio signals.

[0045] Through the above process, a coding function (or coding mapping) h that maps the input audio signal to its coded representation is defined.

[0046] As described above, the spectral envelope and autocorrelation information of the subband audio signals can be determined and output at different update rates (sampling rates). For example, the spectral envelope can be determined at a first update rate, and the autocorrelation information of a plurality of subband audio signals can be determined at a second update rate different from the first update rate. The representation of the spectral envelope and the representation of the autocorrelation information (for all subbands) can be written into the bitstream at the corresponding update rates (sampling rates). In this case, the coded representation can relate to a part of the bitstream output by the encoder. In this regard, it should be noted that for each moment, the current spectral envelope and the current set of autocorrelation information (one piece of information per subband) are defined by the bitstream and can be regarded as the coded representation. Alternatively, the representation of the spectral envelope and the representation of the autocorrelation information (for all subbands) can be updated at the corresponding update rates in the corresponding output units of the encoder. In this case, each output unit of the encoder (e.g., coded frame) corresponds to an instance of the coded representation. Depending on the corresponding update rates, the representations of the spectral envelope and the autocorrelation information may be the same in a series of successive output units.

[0047] Preferably, the first update rate is higher than the second update rate. In one example, the first update rate R1 can be R1 = 1 / (2.5 ms) and the second update rate R2 can be R2 = 1 / (20 ms), so that the updated representation of the spectral envelope is output every 2.5 ms, while the updated representation of the autocorrelation information is output every 20 ms. In terms of the parts (e.g., frames) of the audio signal, the spectral envelope can be determined every n parts (e.g., each part), and conversely the autocorrelation information can be determined every m parts, where m > n.

[0048] (Multiple) coded representations can be output as a sequence of frames of a specific frame length. Among other factors, the frame length can also depend on the first update rate and / or the second update rate. Consider a frame having a length of a first period L1 corresponding to the first update rate R1 (e.g., 1 / (2.5 ms)) via L1 = 1 / R1 (e.g., 2.5 ms), which will include a representation of a spectral envelope and a representation of a set of autocorrelation information (one piece of information per subband audio signal). For the first update rate and the second update rate of 1 / (2.5 ms) and 1 / (20 ms) respectively, the autocorrelation information will be the same for eight consecutive frames of the coded representation. Generally, assuming that R1 and R2 are appropriately chosen to have an integer ratio, the autocorrelation information is the same for R1 / R2 consecutive frames of the coded representation. On the other hand, consider a frame having a length of a second period L2 corresponding to the second update rate R2 (e.g., 1 / (20 ms)) via L2 = 1 / R2 (e.g., 20 ms), which will include a representation of a set of autocorrelation information and R1 / R2 (e.g., eight) representations of the spectral envelope.

[0049] In some embodiments, different update rates can even be applied to different subbands, i.e., the autocorrelation information of different subband audio signals can be generated and output at different update rates.

[0050] Figure 2 FIG. is a flowchart illustrating an example of an encoding method 200 according to an embodiment of the present disclosure. The method (which can be implemented by the above encoder 100) receives an audio signal as input.

[0051] At Step S210 a plurality of subband audio signals are generated based on the audio signal. This can involve spectral decomposition of the audio signal, in which case this step can be performed according to the operation of the above filter bank 15. Alternatively, this can involve spectral and / or temporal flattening of the audio signal, optionally windowing the flattened audio signal by a window function, and spectrally decomposing the resulting signal into a plurality of subband audio signals.

[0052] At Step S220 the spectral envelope of the audio signal is determined (e.g., calculated). This step can be performed according to the operation of the above envelope extraction block 25.

[0053] At Step S230 for each subband audio signal, the autocorrelation information of the subband audio signal is determined based on the ACF of the subband audio signal. This step can be performed according to the operation of the above autocorrelation block 55.

[0054] At Step S240 an encoded representation of the audio signal is generated. The encoded representation includes a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information of each of the plurality of subband audio signals.

[0055] Next, an example of the implementation details of the steps of method 200 will be described.

[0056] For example, as described above, generating a plurality of subband audio signals can include (or be equivalent to) spectral decomposition of the audio signal by, for example, a filter bank. In this case, determining the autocorrelation function of a given subband audio signal can include determining the subband envelope of the subband audio signal. The subband envelope can be determined by taking the magnitude values of the subband audio signal. The ACF itself can be calculated for a real-valued (envelope-flattened windowed) subband signal.

[0057] Assuming that the sub-band filter response is complex-valued and the Fourier transform is substantially supported on positive frequencies, the sub-band signal becomes complex-valued. Then, the sub-band envelope can be determined by taking the magnitude of the complex-valued sub-band signal. The sub-band envelope has as many samples as the sub-band signal and may still be somewhat oscillatory. Optionally, the sub-band envelope can be downsampled, for example, for each shift along half of a particular length of the signal (e.g., 2.5 ms), by calculating the triangular window weighted sum of squares of the envelope in a segment of a particular length (e.g., length 5 ms, rising 2.5 ms, falling 2.5 ms), and then taking the square root of that sequence to obtain the downsampled sub-band envelope. This can be considered to correspond to the definition of an "rms envelope". The triangular window can be normalized such that a constant envelope with value 1 gives a sequence of 1s. Other ways of determining the sub-band envelope are also feasible, such as half-wave rectification followed by low-pass filtering in the case of a real-valued sub-band signal. In any case, the sub-band envelope can be considered to carry information about the energy in the sub-band signal (at the selected update rate).

[0058] Then, the sub-band audio signal can be envelope-flattened based on the sub-band envelope. For example, to obtain the fine-structure signal (carrier) from which the ACF data is calculated, a new full-sampling-rate envelope signal can be created by linearly interpolating the downsampled values and dividing the original (complex-valued) sub-band signal by this linearly interpolated envelope.

[0059] Then the envelope-flattened sub-band audio signal can be windowed by an appropriate window function. Finally, the ACF of the windowed envelope-flattened sub-band audio signal is determined (e.g., calculated). In some embodiments, determining the ACF of a given sub-band audio signal may further include normalizing the ACF of the windowed envelope-flattened sub-band audio signal by the autocorrelation function of the window function.

[0060] In Figure 3 the curve 310 in the upper half indicates the real value of the windowed envelope-flattened sub-band signal used for calculating the ACF. The solid curve 320 in the lower half indicates the real value of the complex ACF.

[0061] The main idea now is to find the maximum local maximum of the ACF of the subband signal among those local maxima above the ACF of the absolute value of the impulse response of the (complex-valued) subband filter (i.e., the corresponding BPF of the filter bank). For the ACF of a complex-valued subband signal, the real value of the ACF can be considered at this time. Finding the maximum local maximum above the ACF of the absolute value of the impulse response may be necessary to avoid picking a lag related to the subband center frequency rather than the input signal attributes. As a final adjustment, this maximum can be divided by the maximum of the ACF of the window function used for the subband ACF window (assuming the ACF of the subband signal itself has been normalized, e.g., such that the autocorrelation value at zero delay is normalized to one). This results in better utilization of the interval between 0 and 1, where ρ(T) = 1 is the maximum pitch.

[0062] Accordingly, determining the autocorrelation information of a given subband audio signal based on the ACF of the subband audio signal can further include comparing the ACF of the subband audio signal with the ACF of the absolute value of the impulse response of the corresponding bandpass filter associated with the subband audio signal. The ACF of the absolute value of the impulse response of the corresponding bandpass filter associated with the subband audio signal is indicated by Figure 3 the solid curve 330 in the lower half of Figure 3 Then, the autocorrelation information is determined based on the highest local maximum of the ACF of the subband signal above the ACF of the absolute value of the impulse response of the corresponding bandpass filter associated with the subband audio signal. In the lower half of Figure 3 the local maxima of the ACF are indicated by crosses, and the highest local maximum of the ACF of the subband signal selected above the ACF of the absolute value of the impulse response of the corresponding bandpass is indicated by a circle. Optionally, the selected local maximum of the ACF can be normalized by the ACF value of the window function (assuming the ACF itself has been normalized, e.g., such that the autocorrelation value at zero delay is normalized to 1). The highest local maximum of the normalized selected ACF is indicated by

[0063] the asterisk in the lower half, and the dashed curve 340 indicates the ACF of the window function.

[0064] A similar coding format can be defined within the framework of an LPC-based vocoder. Also in this case, the autocorrelation information is extracted from sub-band signals that are subject to at least some degree of spectral and / or temporal flattening. Different from the foregoing example, this is done by creating a (perceptually weighted) LPC residual, windowing it, and decomposing it into sub-bands to obtain a plurality of sub-band audio signals. After that, the ACF is computed and the lag values and autocorrelation values of each sub-band audio signal are extracted.

[0065] For example, generating a plurality of sub-band audio signals may include applying spectral and / or temporal flattening to the audio signal (e.g., generating a perceptually weighted LPC residual from the audio signal by using an LPC filter). After that, it may be windowing the flattened audio signal by a window function and spectrally decomposing the windowed flattened audio signal into a plurality of sub-band audio signals. As described above, the result of the temporal and / or spectral flattening may correspond to a perceptually weighted LPC residual, which is subsequently subject to windowing and spectral decomposition into sub-bands. For example, the perceptually weighted LPC residual may be a pink LPC residual.

[0066] Decoding

[0067] The present disclosure relates to audio decoding based on an analysis-by-synthesis approach. At the most abstract level, it is assumed that an encoding mapping h from the signal to the perceptual-motivation domain is given such that the original audio signal x is represented by y = h(x). In the best case, a simple distortion measure like least squares in the perceptual domain can well predict the subjective differences measured by a group of listeners.

[0068] One remaining problem is to design a decoder q that maps from y (the encoded and decoded version thereof) to the audio signal z = d(y). To this end, the concept of analysis-by-synthesis can be used, which involves "finding the waveform that is closest to generating a given image". The goal is that z and x should sound similar, so the decoder should solve the inverse problem h(z) = y = h(x). In terms of the composition of the mapping, d should approximate the left inverse of h, which means This inverse problem is generally ill-posed because it has many solutions. The opportunity to achieve significant bitrate savings lies in observing that a large number of different waveforms will produce the same sound impression.

[0069] Figure 4is a block diagram schematically illustrating an example of an analysis-by-synthesis approach for determining a decoding function (or decoding map) d given an encoding function (or encoding map) h. An encoding map h 415 is applied to an original audio signal x 410, resulting in an encoded representation y 420, where y = h(x). The encoded representation y can be defined in the perceptual domain. The aim is to find a decoding function (decoding map) d 425 that maps the encoded representation y to a reconstructed audio signal z 430, which has the property that applying the encoding map h 435 to the reconstructed audio signal z will produce an encoded representation h(z) 440 that matches the encoded representation y = h(x) substantially. Here, for example, "substantially matches" can mean matching up to a predefined margin. In other words, given the encoding map h, the aim is to find the decoding map d such that

[0070] Figure 5 is a flowchart illustrating an example of a decoding method 500 consistent with an analysis-by-synthesis approach according to an embodiment of the present disclosure. The method 500 is a method for decoding an audio signal from an encoded representation of the (original) audio signal. It is assumed that the encoded representation includes a representation of the spectral envelope of the original audio signal and a representation of the autocorrelation information of each of a plurality of subband audio signals of the original audio signal. The autocorrelation information for a given subband audio signal is based on the ACF of the subband audio signal.

[0071] At Step S510 an encoded representation of the audio signal is received.

[0072] At Step S520 the spectral envelope and the autocorrelation information are extracted from the encoded representation of the audio signal.

[0073] At Step S530At this point, a reconstructed audio signal is determined based on the spectral envelope and autocorrelation information. Wherein, the reconstructed audio signal is determined such that the autocorrelation function of each of the plurality of subband signals of the reconstructed subband audio signal will (substantially) satisfy the conditions derived from the autocorrelation information of the corresponding subband audio signal of the audio signal. For example, the condition may be that for each subband audio signal of the reconstructed audio signal, the value of the ACF of the subband audio signal of the reconstructed audio signal at the lag value (e.g., delay value) indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal substantially matches the autocorrelation value indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal. This may mean that the decoder can determine the ACF of the subband audio signal in the same way as the encoder did. This may involve any one, some, or all of flattening, windowing, and normalizing. In one embodiment, the reconstructed audio signal may be determined such that for each subband audio signal of the reconstructed audio signal, the autocorrelation value and lag value (e.g., delay value) of the ACF of the subband signal of the reconstructed audio signal substantially match the autocorrelation value and lag value indicated by the autocorrelation information of the corresponding subband audio signal of the original audio signal. This may mean that the decoder can determine the autocorrelation information of each subband signal of the reconstructed audio signal in the same way as the encoder did. In those embodiments where the encoded representation further includes waveform information, the reconstructed audio signal may be further determined based on the waveform information. The subband audio signals of the reconstructed audio signal may be generated in the same way as the encoder did. For example, this may involve spectral decomposition, or a series of flattening, windowing, and spectral decomposition.

[0074] Preferably, determining the reconstructed audio signal at step S530 also takes into account the spectral envelope of the original audio signal. Then, the reconstructed audio signal may be further determined such that for each subband audio signal of the reconstructed subband audio signal, the measured (e.g., estimated or calculated) signal power of the subband audio signal of the reconstructed audio signal substantially matches the signal power of the corresponding subband audio signal of the original audio signal indicated by the spectral envelope.

[0075] As can be seen from the above, the proposed method 500 can be considered to be inspired by the analysis-by-synthesis approach, because it attempts to find a reconstructed audio signal z that (substantially) satisfies at least one condition derived from the encoded representation y = h(x) of the original audio signal x, where h is the encoding mapping used by the encoder. In some embodiments, it can even be considered that the proposed method operates according to the analysis-by-synthesis approach, because it attempts to find a reconstructed audio signal z whose encoded representation h(z) will substantially match the encoded representation y = h(x) of the original audio signal x. In other words, it can be considered that the decoding method finds a decoding mapping d such that Next, two non-limiting example embodiments of method 500 will be described.

[0076] Embodiment Example 1: Parameter Synthesis or Signal Iteration

[0077] Given an update mapping z n = f(z n-1 , y), the inverse problem h(z) = y can be solved by an iterative method, where the update mapping modifies z n-1 such that h(z n ) is closer to y than h(z n-1 ). The starting point of the iteration (i.e., the initial candidate for the reconstructed audio signal) can be a random noise signal (e.g., white noise), or for example it can be determined based on the coded representation of the audio signal (e.g., as a manually crafted first guess). In the latter case, the initial candidate for the reconstructed audio signal can be related to an informed guess based on the spectral envelope and / or autocorrelation information of multiple sub-band audio signals. In those embodiments where the coded representation includes waveform information, an informed guess can be further made based on the waveform information.

[0078] More specifically, in this embodiment example, the reconstructed audio signal is determined during an iterative process that starts with an initial candidate for the reconstructed audio signal and generates a corresponding intermediate reconstructed audio signal in each iteration. In each iteration, the update mapping is applied to the intermediate reconstructed audio signal to obtain an intermediate reconstructed audio signal for the next iteration. The update mapping is chosen such that the difference between the coded representation of the intermediate reconstructed audio signal and the coded representation of the original audio signal gradually becomes smaller from one iteration to the next. To this end, an appropriate difference metric for the coded representation (e.g., spectral envelope, autocorrelation information) can be defined and used to evaluate the difference. The coded representation of the intermediate reconstructed audio signal can be the coded representation that would be obtained if the intermediate reconstructed audio signal had undergone the same coding scheme that results in the coded representation of the audio signal.

[0079] In the case where the process seeks a reconstructed audio signal that satisfies at least one condition derived from (multiple) autocorrelation information, the update mapping can be chosen such that the autocorrelation function of the sub-band audio signals of the intermediate reconstruction of the audio signal is closer to satisfying the corresponding condition derived from the autocorrelation information of the corresponding sub-band audio signals of the audio signal, and / or the difference between the measured signal power of the sub-band audio signals of the reconstructed audio signal and the signal power of the corresponding sub-band audio signals of the audio signal indicated by the spectral envelope decreases from one iteration to the next. If both autocorrelation information and spectral envelope are considered simultaneously, an appropriate difference metric can be defined for the degree of satisfaction of the condition and the difference between the signal powers of the sub-band audio signals.

[0080] Embodiment Example 2: Machine Learning-based Generative Model

[0081] Another option supported by modern machine learning methods is to train a machine learning-based generative model (or simply generative model for short) for audio x conditioned on data y. That is, given a large collection of examples of (x, y), where y = h(x), train a parameterized conditional distribution p(x|y) from y to x. Then, the decoding algorithm can consist of sampling from the distribution z ∼ p(x|y).

[0082] This option has been found to be particularly advantageous in cases where h(x) is a speech vocoder and p(x|y) is defined by a sequence generation model sample recurrent neural network (RNN). However, other generative models such as variational autoencoders or generative adversarial models are also relevant to this task. Thus, without being deliberately restrictive, the machine learning-based generative model can be one of a recurrent neural network, a variational autoencoder, or a generative adversarial model (e.g., generative adversarial network (GAN)).

[0083] In this example embodiment, determining the reconstructed audio signal based on the spectral envelope and autocorrelation information includes: applying a machine learning-based generative model that receives the spectral envelope of the audio signal and the autocorrelation information of each of a plurality of sub-band audio signals of the audio signal as inputs and generates and outputs the reconstructed audio signal. In those embodiments where the encoded representation further includes waveform information, the machine learning-based generative model can further receive the waveform information as an input.

[0084] As described above, the machine learning-based generative model can include a parameterized conditional distribution p(x|y) that relates the encoded representation y of the audio signal and the corresponding audio signal x to a respective probability p. Then, determining the reconstructed audio signal can include sampling from the parameterized conditional distribution p(x|y) for the encoded representation of the audio signal.

[0085] In the training phase, before decoding, the machine learning-based generative model can be adjusted / trained on a dataset of multiple audio signals and the corresponding encoded representations of the audio signals. If the encoded representation further includes waveform information, the waveform information can also be used to adjust / train the machine learning-based generative model.

[0086] Figure 6 is a flowchart of an example embodiment 600 that illustrates Figure 5 step S530 in the decoding method 500. In particular, embodiment 600 relates to a sub-band embodiment of step S530.

[0087] In Step 610There, a plurality of reconstructed subband audio signals are determined based on spectral envelope and autocorrelation information. Among them, the plurality of reconstructed subband audio signals are determined such that for each reconstructed subband audio signal, the autocorrelation function of the reconstructed subband audio signal will satisfy a condition derived from the autocorrelation information of the corresponding subband audio signal of the audio signal. In some embodiments, the plurality of reconstructed subband audio signals are determined such that for each reconstructed subband audio signal, the autocorrelation information of the reconstructed subband audio signal will substantially match the autocorrelation information of the corresponding subband audio signal.

[0088] Preferably, determining the plurality of reconstructed subband audio signals at step S610 also takes into account the spectral envelope of the original audio signal. Then, the plurality of reconstructed subband audio signals are further determined such that for each reconstructed subband audio signal, the measured (e.g., estimated, calculated) signal power of the reconstructed subband audio signal substantially matches the signal power of the corresponding subband audio signal indicated by the spectral envelope.

[0089] At Step S620 there, a reconstructed audio signal is determined based on the plurality of reconstructed subband audio signals by spectral synthesis.

[0090] The above-described Embodiment Examples 1 and 2 can also be applied to the subband implementation of step S530. For Embodiment Example 1, each reconstructed subband audio signal can be determined in an iterative process that starts from an initial candidate of the reconstructed subband audio signal and generates a corresponding intermediate reconstructed subband audio signal in each iteration. In each iteration, an update mapping can be applied to the intermediate reconstructed subband audio signal to obtain an intermediate reconstructed subband audio signal for the next iteration in such a way that the difference between the autocorrelation information of the intermediate reconstructed subband audio signal and the autocorrelation information of the corresponding subband audio signal gradually becomes smaller iteration by iteration, or such that the reconstructed subband audio signal better satisfies the respective conditions derived from the autocorrelation information of the respective corresponding subband audio signals of the audio signal.

[0091] Similarly, the spectral envelope can also be considered at this time. That is, the update mapping can make the (joint) difference between the corresponding signal powers of the subband audio signals and between the corresponding autocorrelation information terms gradually become smaller. This may mean defining an appropriate difference metric for evaluating the (joint) difference. In addition, the same explanations as in Embodiment Example 1 above can apply to this case.

[0092] Applying Example Embodiment 2 to the sub-band embodiment of step S530, determining a plurality of reconstructed sub-band audio signals based on the spectral envelope and autocorrelation information may include: applying a machine learning-based generative model that receives the spectral envelope of the audio signal and the autocorrelation information of each of the plurality of sub-band audio signals of the audio signal as inputs and generates and outputs a plurality of reconstructed sub-band audio signals. In addition, the same explanations as in Example Embodiment 2 above may apply to this case.

[0093] The present disclosure further relates to an encoder for encoding an audio signal, the encoder being capable of and adapted to perform the encoding methods described throughout the present disclosure. An example of such an encoder 700 is schematically illustrated in block diagram form in Figure 7 . The encoder 700 includes a processor 710 and a memory 720 coupled to the processor 710. The processor 710 is adapted to perform the method steps of any one of the encoding methods described throughout the present disclosure. To this end, the memory 720 may include corresponding instructions for the processor 710 to execute. The encoder 700 may further include an interface 730 for receiving an input audio signal 740 to be encoded and / or for outputting an encoded representation 750 of the audio signal.

[0094] The present disclosure further relates to a decoder for decoding an audio signal from an encoded representation of the audio signal, the decoder being capable of and adapted to perform the decoding methods described throughout the present disclosure. An example of such a decoder 800 is schematically illustrated in block diagram form in Figure 8 . The decoder 800 includes a processor 810 and a memory 820 coupled to the processor 810. The processor 810 is adapted to perform the method steps of any one of the decoding methods described throughout the present disclosure. To this end, the memory 820 may include corresponding instructions for the processor 810 to execute. The decoder 800 may further include an interface 830 for receiving an input encoded representation 840 of the audio signal to be decoded and / or for outputting a decoded (i.e., reconstructed) audio signal 850.

[0095] The present disclosure further relates to a computer program including instructions that, when executed, cause a computer to perform the encoding or decoding methods described throughout the present disclosure.

[0096] Finally, the present disclosure also relates to a computer-readable storage medium storing the computer program as described above.

[0097] Interpretation

[0098] Unless otherwise specifically stated, it will be apparent from the following discussion that, throughout the disclosed discussion, use of terms such as "processing", "computing", "calculating", "determining", "analyzing", etc., to refer to actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or transform data represented as physical (such as electronic) quantities into other data similarly represented as physical quantities.

[0099] In a similar manner, the term "processor" can refer to any device or portion of a device that processes electronic data, such as from registers and / or memory, to transform that electronic data into other electronic data, such as may be stored in registers and / or memory. A "computer" or "computing machine" or "computing platform" can include one or more processors.

[0100] In one exemplary embodiment, the methods described herein may be executed by one or more processors that receive computer-readable (also referred to as machine-readable) code that includes a set of instructions that, when executed by the one or more processors, perform at least one of the methods described herein. Any processor that includes a set of instructions (sequential or otherwise) that can execute the specified actions to be taken. Thus, one example is a typical processing system that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem that includes main RAM and / or static RAM and / or ROM. A bus subsystem may be included for communication between components. The processing system may further be a distributed processing system in which the processors are coupled together via a network. If the processing system requires a display, such a display may be included, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data input is required, the processing system also includes input devices such as one or more of an alphanumeric input unit (such as a keyboard), a pointing control device (such as a mouse), etc. The processing system may also encompass a storage system such as a disk drive unit. The processing system in some configurations may include a sound output device and a network interface device. The memory subsystem thus includes a computer-readable carrier medium that carries computer-readable code (e.g., software) that includes a set of instructions that, when executed by the one or more processors, causes one or more of the methods described herein to be executed. It should be noted that when the method includes several elements (e.g., several steps), no order of these elements is implied unless specifically stated. During the execution of software by a computer system, the software may reside on a hard disk, or it may also reside entirely or at least partially in RAM and / or in the processor. Thus, the memory and the processor also constitute a computer-readable carrier medium that carries computer-readable code. In addition, the computer-readable carrier medium may form or be included in a computer program product.

[0101] In alternative exemplary embodiments, one or more processors may operate as a stand-alone device or may be connected to (e.g., networked to) other processors in a networked deployment, and the one or more processors may operate as a server or a user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. One or more processors may form a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a cellular phone, a web appliance, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify the actions to be taken by the machine.

[0102] It should be noted that the term "machine" should also be considered to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.

[0103] Accordingly, an example embodiment of each method described herein takes the form of a computer-readable carrier medium carrying a set of instructions, such as a computer program for execution on one or more processors (e.g., one or more processors as part of a web server device). Thus, as will be understood by those skilled in the art, example embodiments of the present disclosure may be embodied as a method, an apparatus such as a special-purpose apparatus, an apparatus such as a data processing system, or a computer-readable carrier medium (e.g., a computer program product). The computer-readable carrier medium carries computer-readable code including a set of instructions that, when executed on one or more processors, cause the one or more processors to implement the method. Accordingly, aspects of the present disclosure may take the form of a method, a fully hardware example embodiment, a fully software example embodiment, or an example embodiment combining software and hardware aspects. Additionally, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) that carries computer-readable program code embodied in the medium.

[0104] Software may be further sent or received via a network interface device over a network. Although in the example embodiments the carrier medium is a single medium, the term "carrier medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store a set or multiple sets of instructions. The term "carrier medium" should also be considered to include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by one or more of the processors and that causes the one or more processors to execute any one or more of the methods of the present disclosure. The carrier medium may take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical discs, magnetic disks, and magneto-optical discs. Volatile media includes dynamic memory, such as main memory. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up a bus subsystem. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier medium" should thus be considered to include but not limited to solid-state memory, computer products embodied in optical and magnetic media; media that carry a propagated signal that can be detected by at least one processor or one or more processors and that represents a set of instructions that, when executed, implement the method; and transmission media in a network that carry a propagated signal that can be detected by at least one of the one or more processors and that represents the set of instructions.

[0105] It will be understood that, in one example embodiment, the steps of the methods discussed are performed by a suitable processor (or processors) in a processing (e.g., computer) system that executes instructions (computer-readable code) stored in a storage device. It will also be understood that the present disclosure is not limited to any particular implementation or programming technique, and the present disclosure can be implemented using any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.

[0106] References throughout this disclosure to "one example embodiment", "some example embodiments" or "example embodiments" mean that a particular feature, structure, or characteristic described in connection with the example embodiments is included in at least one example embodiment of the present disclosure. Thus, the phrases "in one example embodiment", "in some example embodiments" or "in example embodiments" appearing throughout this disclosure are not necessarily all referring to the same example embodiment. Additionally, in one or more example embodiments, the particular features, structures, or characteristics may be combined in any suitable manner, which will be apparent to those of ordinary skill in the art in light of the present disclosure.

[0107] As used herein, unless otherwise specified, the use of the ordinal adjectives "first", "second", "third", etc. to describe a common object merely indicates different instances of like objects and is not intended to imply that the objects so described must be in a given order in time, space, rank, or any other manner.

[0108] In the following claims and the description herein, any one of the terms comprising, comprised of, or which comprises is an open term, meaning it includes at least the subsequent element / feature, but does not exclude other elements / features. Thus, when the term "comprising" is used in a claim, it should not be construed as limited to the apparatus or elements or steps listed after it. For example, the expression for an apparatus comprising A and B should not be limited to an apparatus that only includes elements A and B. As used herein, any one of the terms including, which includes, or that includes is also an open term, and it also means it includes at least the element / feature after the said term, but does not exclude other elements / features. Thus, including is synonymous with comprising and means comprising.

[0109] It should be understood that in the above description of the exemplary embodiments of the present disclosure, various features of the present disclosure are sometimes combined in a single exemplary embodiment / figure or its description to simplify the present disclosure and help understand one or more of the inventive aspects. However, the method of the present disclosure should not be construed as reflecting an intention that the claims require more features than those expressly recited in each claim. On the contrary, as reflected in the following claims, the inventive aspects lie in less than all the features of a single previously disclosed exemplary embodiment. Accordingly, the claims following the specification are hereby expressly incorporated into this specification, where each claim stands on its own as a separate exemplary embodiment of the present disclosure.

[0110] In addition, although some of the exemplary embodiments described herein include some features included in other exemplary embodiments and do not include other features included in other exemplary embodiments, as will be understood by those skilled in the art, combinations of features of different exemplary embodiments are intended to be within the scope of the present disclosure and form different exemplary embodiments. For example, in the following claims, any of the exemplary embodiments recited in the claimed exemplary embodiments can be used in any combination.

[0111] In the description provided herein, numerous specific details are set forth. However, it should be understood that the exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail to avoid obscuring the understanding of this specification.

[0112] Accordingly, although a mode that is considered to be the best mode of the present disclosure has been described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of the present disclosure, and all such changes and modifications that fall within the scope of the present disclosure are intended to be claimed. For example, any of the formulas given above merely represent processes that can be used. Functions can be added to or deleted from the block diagrams, and operations can be interchanged between functional blocks. Steps can be added to or deleted from the methods described within the scope of the present disclosure.

[0113] The embodiments and various aspects of the present disclosure can be understood from the enumerated exemplary embodiments (EEEs) listed below.

[0114] EEE 1. A method for encoding an audio signal, the method comprising:

[0115] generating a plurality of subband audio signals based on the audio signal;

[0116] determining a spectral envelope of the audio signal;

[0117] For each subband audio signal, determine autocorrelation information of the subband audio signal based on the autocorrelation function of the subband audio signal; and

[0118] Generate an encoded representation of the audio signal, the encoded representation including a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information of the plurality of subband audio signals.

[0119] EEE 2. The method according to EEE 1, wherein the spectral envelope is determined based on the plurality of subband audio signals.

[0120] EEE 3. The method according to EEE 1 or 2, wherein the autocorrelation information of a given subband audio signal includes the lag value of the corresponding subband audio signal and / or the autocorrelation value of the corresponding subband audio signal.

[0121] EEE 4. The method according to the previous EEE, wherein the lag value corresponds to the delay value at which the autocorrelation function reaches a local maximum, and wherein the autocorrelation value corresponds to the local maximum.

[0122] EEE 5. The method according to any one of the previous EEEs, wherein the spectral envelope is determined at a first update rate, and the autocorrelation information of the plurality of subband audio signals is determined at a second update rate; and

[0123] wherein the first update rate and the second update rate are different from each other.

[0124] EEE 6. The method according to the previous EEE, wherein the first update rate is higher than the second update rate.

[0125] EEE 7. The method according to any one of the previous EEEs, wherein generating the plurality of subband audio signals includes:

[0126] Applying spectral and / or temporal flattening to the audio signal;

[0127] Windowing the flattened audio signal; and

[0128] Spectrally decomposing the windowed flattened audio signal into the plurality of subband audio signals.

[0129] EEE 8. The method according to any one of EEEs 1 to 6,

[0130] wherein generating the plurality of subband audio signals includes spectrally decomposing the audio signal; and

[0131] wherein determining the autocorrelation function of a given subband audio signal includes:

[0132] Determine the sub-band envelope of the sub-band audio signal;

[0133] Perform envelope flattening on the sub-band audio signal based on the sub-band envelope;

[0134] Window the envelope-flattened sub-band audio signal through a window function; and

[0135] Determine the autocorrelation function of the windowed envelope-flattened sub-band audio signal.

[0136] EEE 9. The method according to EEE 7 or 8, wherein determining the autocorrelation function of a given sub-band audio signal further comprises:

[0137] Normalize the autocorrelation function of the windowed envelope-flattened sub-band audio signal by the autocorrelation function of the window function.

[0138] EEE 10. The method according to any one of the foregoing EEEs, wherein determining the autocorrelation information of a given sub-band audio signal based on the autocorrelation function of the sub-band audio signal comprises:

[0139] Compare the autocorrelation function of the sub-band audio signal with the autocorrelation function of the absolute value of the impulse response of the corresponding band-pass filter associated with the sub-band audio signal; and

[0140] Determine the autocorrelation information based on the highest local maximum of the autocorrelation function of the sub-band signal above the autocorrelation function of the absolute value of the impulse response of the corresponding band-pass filter associated with the sub-band audio signal.

[0141] EEE 11. The method according to any one of the foregoing EEEs, wherein determining the spectral envelope comprises measuring the signal power of each of the plurality of sub-band audio signals.

[0142] EEE 12. A method for decoding an audio signal from an encoded representation of the audio signal, the encoded representation comprising a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information of each of a plurality of sub-band audio signals generated from the audio signal, wherein the autocorrelation information of a given sub-band audio signal is based on the autocorrelation function of the sub-band audio signal, the method comprising:

[0143] Receive the encoded representation of the audio signal;

[0144] Extract the spectral envelope and the autocorrelation information from the encoded representation of the audio signal; and

[0145] Determine a reconstructed audio signal based on the spectral envelope and the autocorrelation information,

[0146] Wherein, the reconstructed audio signal is determined such that the autocorrelation function of each of a plurality of subband signals generated from the reconstructed audio signal will satisfy a condition derived from the autocorrelation information of the corresponding subband audio signal generated from the audio signal.

[0147] EEE 13. The method according to the previous EEE, wherein the reconstructed audio signal is further determined such that for each subband audio signal of the reconstructed audio signal, the measured signal power of the subband audio signal of the reconstructed audio signal substantially matches the signal power of the corresponding subband audio signal of the audio signal indicated by the spectral envelope.

[0148] EEE 14. The method according to EEE 12 or 13,

[0149] wherein the reconstructed audio signal is determined in an iterative process that starts from an initial candidate of the reconstructed audio signal and generates a corresponding intermediate reconstructed audio signal in each iteration; and

[0150] wherein in each iteration, an update mapping is applied to the intermediate reconstructed audio signal in the following manner to obtain an intermediate reconstructed audio signal for the next iteration: such that the difference between the coded representation of the intermediate reconstructed audio signal and the coded representation of the audio signal gradually becomes smaller iteration by iteration.

[0151] EEE 15. The method according to EEE 14, wherein the initial candidate of the reconstructed audio signal is determined based on the coded representation of the audio signal.

[0152] EEE 16. The method according to EEE 14, wherein the initial candidate of the reconstructed audio signal is white noise.

[0153] EEE 17. The method according to EEE 12 or 13, wherein determining the reconstructed audio signal based on the spectral envelope and the autocorrelation information includes: applying a machine learning-based generative model that receives the spectral envelope of the audio signal and the autocorrelation information of each of a plurality of subband audio signals of the audio signal as inputs and generates and outputs the reconstructed audio signal.

[0154] EEE 18. The method according to the previous EEE, wherein the machine learning-based generative model includes a parametric conditional distribution that associates a coded representation of an audio signal and the corresponding audio signal with respective probabilities; and

[0155] wherein determining the reconstructed audio signal includes sampling from the parametric conditional distribution for the coded representation of the audio signal.

[0156] EEE 19. The method according to EEE 17 or 18 further includes, in a training phase, training the machine learning-based generative model on a dataset of a plurality of audio signals and corresponding encoded representations of the audio signals.

[0157] EEE 20. The method according to any one of EEE 17 to 19, wherein the machine learning-based generative model is one of a recurrent neural network, a variational autoencoder, or a generative adversarial model.

[0158] EEE 21. The method according to EEE 12, wherein determining the reconstructed audio signal based on the spectral envelope and the autocorrelation information includes:

[0159] determining a plurality of reconstructed sub-band audio signals based on the spectral envelope and the autocorrelation information; and

[0160] determining the reconstructed audio signal based on the plurality of reconstructed sub-band audio signals by spectral synthesis,

[0161] wherein the plurality of reconstructed sub-band audio signals are determined such that for each reconstructed sub-band audio signal, the autocorrelation function of the reconstructed sub-band audio signal will satisfy a condition derived from the autocorrelation information of the corresponding sub-band audio signal.

[0162] EEE 22. The method according to the previous EEE, wherein the plurality of reconstructed sub-band audio signals are further determined such that for each reconstructed sub-band audio signal, the measured signal power of the reconstructed sub-band audio signal substantially matches the signal power of the corresponding sub-band audio signal indicated by the spectral envelope.

[0163] EEE 23. The method according to EEE 21 or 22,

[0164] wherein each reconstructed sub-band audio signal is determined in an iterative process that starts from an initial candidate of the reconstructed sub-band audio signal and generates a corresponding intermediate reconstructed sub-band audio signal in each iteration; and

[0165] wherein in each iteration, an update mapping is applied to the intermediate reconstructed sub-band audio signal in the following manner to obtain an intermediate reconstructed sub-band audio signal for the next iteration: such that the difference between the autocorrelation information of the intermediate reconstructed sub-band audio signal and the autocorrelation information of the corresponding sub-band audio signal gradually becomes smaller iteration by iteration.

[0166] EEE 24. The method according to EEE 21 or 22, wherein determining the plurality of reconstructed subband audio signals based on the spectral envelope and the autocorrelation information includes: applying a machine learning-based generative model that receives the spectral envelope of the audio signal and the autocorrelation information of each of the plurality of subband audio signals of the audio signal as inputs and generates and outputs the plurality of reconstructed subband audio signals.

[0167] EEE 25. An encoder for encoding an audio signal, the encoder including a processor and a memory coupled to the processor, wherein the processor is adapted to perform the method steps according to any one of EEE 1 to 11.

[0168] EEE 26. A decoder for decoding the audio signal from an encoded representation of the audio signal, the decoder including a processor and a memory coupled to the processor, wherein the processor is adapted to perform the method steps according to any one of EEE 12 to 24.

[0169] EEE 27. A computer program including instructions for causing a computer to perform the method according to any one of EEE 1 to 24 when the instructions are executed.

[0170] EEE 28. A computer-readable storage medium storing the computer program according to the previous EEE.

Claims

1. A method for encoding an audio signal, the method comprising: generating a plurality of sub-band audio signals based on the audio signal; determining a spectral envelope of the audio signal; for each sub-band audio signal of the plurality of sub-band audio signals, determining autocorrelation information of the sub-band audio signal based on an autocorrelation function of the sub-band audio signal, wherein the autocorrelation information includes an autocorrelation value of the sub-band audio signal; and encoding the spectral envelope of the audio signal and the autocorrelation information of the plurality of sub-band audio signals into an encoded representation of the audio signal, wherein the autocorrelation information of a given sub-band audio signal further includes a lag value of the corresponding sub-band audio signal, wherein the spectral envelope is determined at a first sampling rate, and the autocorrelation information of the plurality of sub-band audio signals is determined at a second sampling rate, and the first sampling rate is higher than the second sampling rate.

2. The method according to claim 1, further comprising outputting a bitstream defining the encoded representation.

3. The method according to claim 1 or 2, wherein The spectral envelope is determined based on the plurality of sub-band audio signals.

4. The method according to claim 1 or 2, wherein The lag value corresponds to a delay value at which the autocorrelation function reaches a local maximum, and wherein the autocorrelation value corresponds to the local maximum.

5. The method according to claim 1 or 2, wherein Generating the plurality of sub-band audio signals includes: applying spectral and / or temporal flattening to the audio signal; windowing the flattened audio signal; and decomposing the windowed flattened audio signal spectrally into the plurality of sub-band audio signals.

6. The method according to claim 1 or 2, Among them, generating the plurality of sub-band audio signals includes decomposing the audio signal spectrally; and wherein determining the autocorrelation function of a given sub-band audio signal includes: determining a sub-band envelope of the given sub-band audio signal; envelope-flattening the given sub-band audio signal based on the sub-band envelope; windowing the envelope-flattened sub-band audio signal by a window function; and determining the autocorrelation function of the windowed envelope-flattened sub-band audio signal.

7. The method according to claim 6, wherein, Determining the autocorrelation function of the given sub-band audio signal further includes: normalizing the autocorrelation function of the windowed envelope-flattened sub-band audio signal by an autocorrelation function of the window function.

8. The method according to claim 1 or 2, wherein Determining the autocorrelation information of the sub-band audio signal based on the autocorrelation function of the sub-band audio signal includes: comparing the autocorrelation function of the sub-band audio signal with an autocorrelation function of an absolute value of an impulse response of a corresponding band-pass filter associated with the sub-band audio signal; and determining the autocorrelation information based on a highest local maximum of the autocorrelation function of the sub-band audio signal above the autocorrelation function of the absolute value of the impulse response of the corresponding band-pass filter associated with the sub-band audio signal.

9. The method according to claim 1 or 2, wherein Determining the spectral envelope of the audio signal includes measuring a signal power of each of the plurality of sub-band audio signals.

10. A method for decoding an audio signal from an encoded representation of the audio signal, the encoded representation including a spectral envelope of the audio signal and autocorrelation information for each of a plurality of subband audio signals generated from the audio signal, wherein, The autocorrelation information of a given sub-band audio signal is based on the autocorrelation function of the sub-band audio signal, the method comprising: receiving the encoded representation of the audio signal; extracting the spectral envelope and the autocorrelation information from the encoded representation of the audio signal; and Determine a reconstructed audio signal based on the spectral envelope and the autocorrelation information, wherein the autocorrelation information of a given subband audio signal includes the autocorrelation value of the subband audio signal and the lag value of the corresponding subband audio signal, wherein the spectral envelope is determined at a first sampling rate, and the autocorrelation information of the plurality of subband audio signals is determined at a second sampling rate, and the first sampling rate is higher than the second sampling rate.

11. The method according to claim 10, wherein, The reconstructed audio signal is determined such that the autocorrelation function of each of the plurality of subband signals generated from the reconstructed audio signal satisfies a condition derived from the autocorrelation information of the corresponding subband audio signal generated from the audio signal.

12. The method according to claim 10 or 11, wherein, The reconstructed audio signal is determined such that the autocorrelation information of each of the plurality of subband signals of the reconstructed audio signal matches the autocorrelation information of the corresponding subband audio signal of the audio signal up to a predefined margin.

13. The method according to claim 10 or 11, wherein The reconstructed audio signal is determined such that for each subband audio signal of the reconstructed audio signal, the value of the autocorrelation function of the subband audio signal of the reconstructed audio signal at the lag value indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal matches the autocorrelation value indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal up to a predefined margin.

14. The method according to claim 10 or 11, wherein, The reconstructed audio signal is further determined such that for each subband audio signal of the reconstructed audio signal, the measured signal power of the subband audio signal of the reconstructed audio signal matches the signal power of the corresponding subband audio signal of the audio signal indicated by the spectral envelope up to a predefined margin.

15. The method according to claim 10 or 11, Among them, wherein the reconstructed audio signal is determined in an iterative process that starts from an initial candidate of the reconstructed audio signal and generates a corresponding intermediate reconstructed audio signal in each iteration; and wherein in each iteration, an update mapping is applied to the intermediate reconstructed audio signal in the following manner to obtain an intermediate reconstructed audio signal for the next iteration: such that the difference between the coded representation of the intermediate reconstructed audio signal and the coded representation of the audio signal gradually becomes smaller iteration by iteration.

16. The method according to claim 15, wherein, The initial candidate of the reconstructed audio signal is determined based on the coded representation of the audio signal.

17. The method according to claim 15, wherein The initial candidate of the reconstructed audio signal is white noise.

18. The method according to claim 10 or 11, wherein Determining the reconstructed audio signal based on the spectral envelope and the autocorrelation information includes: applying a machine learning-based generative model that receives the spectral envelope of the audio signal and the autocorrelation information of each of the plurality of subband audio signals of the audio signal as inputs and generates and outputs the reconstructed audio signal.

19. The method according to claim 18, wherein, The machine learning-based generative model includes a parametric conditional distribution that associates a coded representation of an audio signal and the corresponding audio signal with respective probabilities; and wherein determining the reconstructed audio signal includes sampling from the parametric conditional distribution for the coded representation of the audio signal.

20. The method according to claim 18, further comprising: During the training phase, the machine learning-based generation model is trained on a dataset of multiple audio signals and corresponding encoded representations of the audio signals.

21. The method according to claim 18, wherein The machine learning-based generation model is one of a recurrent neural network, a variational autoencoder, or a generative adversarial model.

22. The method according to claim 11, wherein Determining the reconstructed audio signal based on the spectral envelope and the autocorrelation information includes: Determining a plurality of reconstructed subband audio signals based on the spectral envelope and the autocorrelation information; and Determining the reconstructed audio signal based on the plurality of reconstructed subband audio signals through spectral synthesis, wherein the plurality of reconstructed subband audio signals are determined such that for each reconstructed subband audio signal, the autocorrelation function of the reconstructed subband audio signal satisfies a condition derived from the autocorrelation information of the corresponding subband audio signal of the audio signal.

23. The method according to claim 22, wherein The plurality of reconstructed subband audio signals are determined such that the autocorrelation information of each reconstructed subband audio signal matches the autocorrelation information of the corresponding subband audio signal of the audio signal up to a predefined margin.

24. The method according to claim 22, wherein, The plurality of reconstructed subband audio signals are determined such that for each reconstructed subband audio signal, the value of the autocorrelation function of the reconstructed subband audio signal at the lag value indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal matches the autocorrelation value indicated by the autocorrelation information of the corresponding subband audio signal of the audio signal up to a predefined margin.

25. The method according to any one of claims 22 to 24, wherein The plurality of reconstructed subband audio signals are further determined such that for each reconstructed subband audio signal, the measured signal power of the reconstructed subband audio signal matches the signal power of the corresponding subband audio signal indicated by the spectral envelope up to a predefined margin.

26. The method according to any one of claims 22 to 24, Among them, each reconstructed subband audio signal is determined in an iterative process that starts from an initial candidate of the reconstructed subband audio signal and generates a corresponding intermediate reconstructed subband audio signal in each iteration; and wherein in each iteration, an update mapping is applied to the intermediate reconstructed subband audio signal in the following manner to obtain an intermediate reconstructed subband audio signal for the next iteration: such that the difference between the autocorrelation information of the intermediate reconstructed subband audio signal and the autocorrelation information of the corresponding subband audio signal gradually becomes smaller iteration by iteration.

27. The method according to any one of claims 22 to 24, wherein, Determining the plurality of reconstructed subband audio signals based on the spectral envelope and the autocorrelation information includes: applying a machine learning-based generation model that receives the spectral envelope of the audio signal and the autocorrelation information of each of the plurality of subband audio signals of the audio signal as inputs and generates and outputs the plurality of reconstructed subband audio signals.

28. An encoder for encoding an audio signal, the encoder comprising a processor and a memory coupled to the processor, wherein, The processor is adapted to execute the method steps according to any one of claims 1 to 9.

29. A decoder for decoding an audio signal from an encoded representation of the audio signal, the decoder comprising a processor and a memory coupled to the processor, wherein, The processor is adapted to execute the method steps according to any one of claims 10 to 27.

30. A computer program product comprising instructions that, when executed, cause a computer to perform the method according to any one of claims 1 to 27.

31. A computer-readable storage medium stores computer-executable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 27.

Citation Information

Patent Citations

  • Encoding device and encoding method

    CN106847295A