Multi-lag format for audio coding

The encoding method for audio signals using subband audio signals and autocorrelation information addresses the inefficiency of high-quality audio coding by achieving high coding efficiency and sound quality through a multi-lag format with different update rates.

JP7866661B2Active Publication Date: 2026-05-27DOLBY INTERNATIONAL AB

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2025-03-19
Publication Date
2026-05-27

AI Technical Summary

Technical Problem

High-quality audio coding systems require a large amount of data, resulting in low coding efficiency, and existing methods do not primarily rely on perceptually relevant features for improved coding efficiency.

Method used

An encoding method that generates subband audio signals based on spectral decomposition, determines spectral envelopes and autocorrelation information, and uses a synthesis by analysis approach for decoding, incorporating a multi-lag format with different update rates for spectral envelopes and autocorrelation information.

Benefits of technology

Achieves high coding efficiency with very low bitrate while maintaining excellent sound quality by using autocorrelation information and spectral envelopes, allowing for efficient decoding through iterative or machine learning-based techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007866661000005
    Figure 0007866661000005
  • Figure 0007866661000006
    Figure 0007866661000006
  • Figure 0007866661000007
    Figure 0007866661000007
Patent Text Reader

Abstract

To provide a method of coding an audio signal.SOLUTION: A method includes the steps of: generating a plurality of sub-band audio signals based upon an audio signal; determining a spectrum envelope of the audio signal; determining, for each sub-band audio signal, auto-correlation function information on the sub-band audio signal based upon an auto-correlation function of the sub-band audio signal; and generating coding representations of the audio signal, the coding representations including a representation of the spectrum envelope of the audio signal and representations of the auto-correlation information on the plurality of sub-band audio signals. There are further described a method for decoding the audio signal from the encoding representation, and a corresponding encoder, a decoder, a computer program, and a computer-readable recording medium.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Patent Application No. 62 / 889,118, filed on August 20, 2019, and European Patent Application No. 19192552.8, filed on August 20, 2019, and incorporates by reference the entire content of each application.

[0002] [Technical Field] This disclosure generally relates to methods for encoding an audio signal into an encoded representation and methods for decoding an audio signal from the encoded representation. [[ID=I5]]

[0003] Although some embodiments are described herein with particular reference to its disclosure, it is recognized that the present disclosure is not limited to such fields of use and is applicable in broader situations.

Background Art

[0004] No description of the background art through this disclosure should be construed as an admission that the technology is widely known or constitutes part of the common general knowledge in the technical field.

[0005] In high - quality audio coding systems, it is common to describe the detailed waveform characteristics of the signal in the largest part of the information. The smaller part of the information is used to describe more statistically defined features, such as control data intended to shape quantization noise according to the energy in the frequency band or known simultaneous masking characteristics of hearing (e.g., side information in an MDCT - based waveform coder that transmits the quantization step size and range information necessary to accurately inverse - quantize the data representing the waveform in the decoder). However, these high - quality audio coding systems require a relatively large amount of data to code audio content, i.e., they have a relatively low coding efficiency.

[0006] There is a need for an audio coding method and apparatus that can code audio data with improved coding efficiency. [Overview of the project]

[0007] This disclosure provides a method for encoding an audio signal, a method for decoding an audio signal, an encoder, a decoder, a computer program, and a computer-readable storage medium.

[0008] A first aspect of this disclosure provides a method for encoding an audio signal. The encoding may be performed on each of several consecutive parts of the audio signal (e.g., groups of samples, segments, or frames). In some implementations, these parts may overlap each other. Encoded representations may be generated for each such part. The method may include generating several subband audio signals based on the audio signal. Generating several subband audio signals based on the audio signal may include spectral decomposition of the audio signal, which may be performed by a filter bank of bandpass filters (BPFs). The frequency resolution of the filter bank may be related to the frequency resolution of the human auditory system. For example, the BPF may be a complex-valued BPF. Alternatively, generating several subband audio signals based on the audio signal may include spectrally and / or temporally flattening the audio signal, optionally windowing the flattened audio signal by a window function, and spectrally decomposing the resulting signal into several subband audio signals. The method may further include determining the spectral envelope of the audio signal. The method may further include determining autocorrelation information for each subband audio signal based on the autocorrelation function (ACF) of the subband audio signal. The method may also further include generating an encoded representation of the audio signal, which includes a representation of the spectral envelope of the audio signal and a representation of autocorrelation information for multiple subband audio signals. For example, the encoded representation may relate to a portion of the bitstream. In some implementations, the encoded representation may further include waveform information relating to the waveform of the audio signal and / or one or more waveforms of the subband audio signals. The method may further include outputting the encoded representation.

[0009] When configured as described above, the proposed method has very high coding efficiency (i.e., requires a very low bitrate to code the audio) while simultaneously providing an encoded representation of the audio signal that contains sufficient information to achieve very good sound quality after reconstruction. This is achieved by providing autocorrelation information for multiple subbands of the audio signal, in addition to the spectral envelope. Notably, it has been proven that two values ​​per subband, namely one lag value and one autocorrelation value, are sufficient to achieve high sound quality.

[0010] In some embodiments, the autocorrelation information for a given subband audio signal may include a lag value and / or an autocorrelation value for each subband audio signal. Preferably, the autocorrelation information may include both a lag value and an autocorrelation value for each subband audio signal. Here, the lag value may correspond to a delay value (e.g., x-coordinate) at which the autocorrelation function reaches a local maximum, and the autocorrelation value may correspond to this local maximum (e.g., y-coordinate).

[0011] In some embodiments, the spectral envelope may be determined at a first update rate, and the autocorrelation information for multiple subband audio signals may be determined at a second update rate. In this case, the first and second update rates may be different from each other. The update rate may also be called the sampling rate. In one such embodiment, the first update rate may be higher than the second update rate. Furthermore, different update rates may be applied to different subbands, i.e., the update rates for the autocorrelation information for different subband audio signals may be different from each other.

[0012] By reducing the update rate of autocorrelation information compared to the update rate of the spectral envelope, the coding efficiency of the proposed method can be further improved without affecting the sound quality of the reconstructed audio signal.

[0013] In some embodiments, generating multiple subband audio signals may include applying spectral and / or temporal flattening to the audio signals. Generating multiple subband audio signals may further include windowing the flattened audio signals using a window function. Generating multiple subband audio signals may further include spectrally decomposing the windowed, flattened audio signals into multiple subband audio signals. In this case, for example, spectrally and / or temporally flattening the audio signals may include generating perceptually weighted LPC residuals of the audio signals.

[0014] In some embodiments, generating multiple subband audio signals may include spectral decomposition of the audio signals. Then, determining the autocorrelation function for a given subband audio signal may include determining the subband envelope of the subband audio signal. Determining the autocorrelation function may further include envelope flattening the subband audio signal based on the subband envelope. The subband envelope may be determined by taking the magnitude values ​​of the windowed subband audio signal. Determining the autocorrelation function may also include windowing the envelope flattened subband audio signal by the window function. Furthermore, determining the autocorrelation function may further include determining (e.g., calculating) the autocorrelation function of the windowed subband audio signal after envelope flattening. The autocorrelation function may be determined for real-valued (windowed after envelope flattening) subband signals.

[0015] Another aspect of this disclosure relates to a method for decoding an audio signal from an encoded representation of the audio signal. The encoded representation may include a representation of the spectral envelope of the audio signal and a representation of autocorrelation information for each of several subband audio signals (or those generated from the audio signal) of the audio signal. The autocorrelation information for a given subband audio signal may be based on the autocorrelation function of the subband audio signal. The method may include receiving the encoded representation of the audio signal. The method may further include extracting the spectral envelope and (multiple) autocorrelation information from the encoded representation of the audio signal. The method may further include determining a reconstructed audio signal based on the spectral envelope and autocorrelation information. The reconstructed audio signal may be determined such that the autocorrelation function for each of the several subband audio signals (or those generated from the reconstructed audio signal) of the reconstructed audio signal satisfies conditions derived from the autocorrelation information for the corresponding subband audio signals (or those generated from the audio signal) of the audio signal. For example, the reconstructed audio signal may be determined such that, for each subband audio signal of the reconstructed audio signal, the value of the autocorrelation function of the subband audio signal (or one generated from the reconstructed audio signal) substantially matches the autocorrelation value indicated by the autocorrelation information for the corresponding subband audio signal of the audio signal, at a lag value (e.g., delay value) indicated by the autocorrelation information for the corresponding subband audio signal of the audio signal. This may mean that the decoder can determine the autocorrelation function of the subband audio signals in the same way that the encoder does. This may include, some or all, of flattening, windowing, and normalization.In some implementations, the reconstructed audio signal may be determined such that the autocorrelation information for each of the multiple subband signals (or those generated from the reconstructed subband audio signal) of the reconstructed subband audio signal substantially matches the autocorrelation information for the corresponding subband audio signal (or those generated from the audio signal) of the audio signal. For example, the reconstructed audio signal may be determined such that, for each subband audio signal (or those generated from the reconstructed audio signal) of the reconstructed audio signal, the autocorrelation value and lag value (e.g., delay value) of the autocorrelation function of the subband signal of the reconstructed audio signal substantially matches the autocorrelation value and lag value indicated, for example, by the autocorrelation information for the corresponding subband audio signal (or those generated from the audio signal) of the audio signal. This may mean that the decoder can determine the autocorrelation information (i.e., lag value and autocorrelation value) for each subband signal of the reconstructed audio signal in the same way that the encoder does. Here, the term substantially matches may mean, for example, matching up to a given margin. In these implementations where the encoded representation includes waveform information, the reconstructed audio signal may be determined further based on the waveform information. The subband audio signal may be obtained, for example, by spectral decomposition of the applicable audio signal (i.e., the original audio signal on the encoder side or the reconstructed audio signal on the decoder side), or by flattening and windowing the applicable audio signal and then spectrally decomposing it.

[0016] Therefore, it can also be said that the decoder operates according to a synthesis by analysis approach, which attempts to find a reconstructed audio signal z whose encoded representation h(z) substantially matches the encoded representation h(x) of the original audio signal x, or where h is the encoding map used by the encoder. In other words, the decoder operates according to a synthesis by analysis approach,

[0017]

number

[0018] In some embodiments, the reconstructed audio signal may be determined by an iterative procedure that starts from an initial candidate for the reconstructed audio signal and generates each intermediate reconstructed audio signal at each iteration. At each iteration, an update map may be applied to the intermediate reconstructed audio signal to obtain an intermediate reconstructed audio signal for the next iteration. The update map may be configured such that the autocorrelation function of the subband audio signal (or one generated from the intermediate reconstruction) of the intermediate reconstruction of the audio signal approaches a condition derived from the autocorrelation information for the corresponding subband audio signal (or one generated from the audio signal) of the audio signal, and / or the difference between the measured signal power of the subband audio signal (or one generated from the reconstructed audio signal) of the reconstructed audio signal and the signal power of the corresponding subband audio signal (or one generated from the audio signal) of the audio signal, as indicated by the spectral envelope, decreases from one iteration to the next. When both autocorrelation information and spectral envelopes are considered, an appropriate difference metric for the degree to which the condition is satisfied and the difference between the signal power for the subband audio signal may be defined. In some implementations, the update map may be configured such that the difference between the encoded representation of the intermediate reconstructed audio signal and the encoded representation of the audio signal decreases continuously from one iteration to the next. For this purpose, an appropriate difference metric for the encoded representation (including spectral envelope and / or autocorrelation information) may be defined and used. The autocorrelation function of the subband audio signals of the intermediate reconstructed audio signal (or those generated from the intermediate reconstructed audio signal) may be determined in the same way that the encoder does for the subband audio signals of the audio signal (or those generated from the audio signal). Similarly, the encoded representation of the intermediate reconstructed audio signal may be the encoded representation obtained when the intermediate reconstructed audio signal underwent the same encoding technique that resulted in the encoded representation of the audio signal.

[0019] Such an iterative method enables a simple yet efficient implementation of the synthesis technique according to the above analysis.

[0020] In some embodiments, determining a reconstructed audio signal based on spectral envelope and autocorrelation information may include applying a machine - learning - based generation model that receives, as input, the spectral envelope of the audio signal and the autocorrelation information for each of a plurality of sub - band audio signals of the audio signal, and generates and outputs a reconstructed audio signal. In these implementations where the encoded representation includes waveform information, the machine - learning - based generation model may further receive waveform information as input. This means that the machine - learning - based generation model may also be conditioned / trained using the waveform information.

[0021] Such a machine - learning - based method enables a very efficient implementation of the synthesis technique according to the above analysis and can achieve a reconstructed audio signal that is perceptually very close to the original audio signal.

[0022] Another aspect of the present disclosure relates to an encoder for encoding an audio signal. The encoder may include a processor and a memory coupled to the processor, and the processor is adapted to execute the steps of any one of the encoding methods described throughout the present disclosure.

[0023] Another aspect of the present disclosure relates to a decoder for decoding an audio signal from an encoded representation of the audio signal. The decoder may include a processor and a memory coupled to the processor, and the processor is adapted to execute the steps of any one of the decoding methods described throughout the present disclosure.

[0024] Another aspect relates to a computer program including instructions that, when executed, cause a computer to execute the steps of any of the methods described throughout the present disclosure.

[0025] Other aspects of the present disclosure relate to a computer-readable storage medium storing a computer program according to the above aspects.

Brief Description of the Drawings

[0026] Here, referring to the accompanying drawings, exemplary embodiments of the present disclosure will be described by way of example only. [Figure 1] It is a block diagram schematically showing an example of an encoder according to an embodiment of the present disclosure. [Figure 2] It is a flowchart showing an example of an encoding method according to an embodiment of the present disclosure. [Figure 3] An example of a waveform that may exist in the framework of the encoding method of FIG. 2 is schematically shown. [Figure 4] It is a block diagram schematically showing an example of a synthesis method by analysis for determining a decoding function. [Figure 5] It is a flowchart showing an example of a decoding method according to an embodiment of the present disclosure. [Figure 6] It is a flowchart showing an example of steps in the decoding method of FIG. 5. [Figure 7] It is a block diagram schematically showing another example of an encoder according to an embodiment of the present disclosure. [Figure 8] It is a block diagram schematically showing an example of a decoder according to an embodiment of the present disclosure.

Modes for Carrying Out the Invention

[0027] [First] High-quality audio coding systems generally require a relatively large amount of data to code audio content, i.e., they have relatively low coding efficiency. While the development of tools such as noise filling and high-frequency reproduction has shown that waveform description data can be partially replaced by a smaller set of control data, no high-quality audio codec relies primarily on perceptually relevant features. However, increased computing power and recent advances in the field of machine learning are increasing the feasibility of decoding audio from virtually arbitrary encoder formats. This disclosure proposes an example of such an encoder format.

[0028] Broadly speaking, this disclosure proposes an encoding format based on subband envelopes and additional information provided by auditory resolution. The additional information includes a single autocorrelation value and a single lag value per subband (and per update step). The envelope can be computed at a first update rate, and the additional information can be sampled at a second update rate. Decoding of the encoding format can be carried out using a synthesis by analysis approach, which can be implemented, for example, by iterative or machine learning-based techniques.

[0029] [Encoding] The encoding format (encoded representation) proposed in this disclosure provides one lag per subband (and update step) and may therefore be called a multi-lag format. Figure 1 is a schematic block diagram showing an example of an encoder 100 for generating an encoding format according to an embodiment of this disclosure.

[0030] The encoder 100 receives a target sound 10 corresponding to the audio signal to be encoded. The audio signal 10 may include multiple consecutive or partially overlapping portions (e.g., groups of samples, segments, frames, etc.) that are processed by the encoder. The audio signal 10 is spectrally decomposed by the filter bank 15 into multiple subband audio signals 20 within the corresponding frequency subbands. The filter bank 15 may be a filter bank of bandpass filters (BPFs), for example, a complex-value BPF. In audio, a filter bank of BPFs with frequency resolution relevant to the human auditory system is naturally used.

[0031] The spectral envelope 30 of the audio signal 10 is extracted by the envelope extraction block 25. For each subband, the power is measured at predetermined time steps as a basic model of the auditory envelope or excitation pattern for the cochlea resulting from the input sound signal, thereby determining the spectral envelope 30 of the audio signal 10. That is, the spectral envelope 30 may be determined based on a plurality of subband audio signals 20, for example by measuring (e.g., estimating, calculating) the respective signal power for each of the plurality of subband audio signals 20. However, the spectral envelope 30 may be determined by any suitable alternative tool, for example, linear predictive coding (LPC) description. In particular, in some implementations, the spectral envelope may be determined from the audio signal before spectral decomposition by the filter bank 15.

[0032] Optionally, the extracted spectral envelope 30 can be downsampled in the downsampling block 35, and the downsampled spectral envelope 40 (or spectral envelope 30) is output as part of the encoding format or encoding representation of the audio signal 10 (or the applicable portion thereof).

[0033] Reconstructed signals reconstructed solely from spectral envelopes may still lack sound quality. To address this problem, this disclosure proposes including a single value (i.e., y and x coordinates) of the autocorrelation function of the signal per subband (optionally envelope-flattened) which results in dramatically improved sound quality. For this purpose, the subband audio signals 20 are optionally flattened (envelope-flattened) by a divider 45 and input to an autocorrelation block 55. The autocorrelation block 55 determines the autocorrelation function (ACF) of its input signals and outputs autocorrelation information 50 for each subband audio signal 20 (i.e., for each subband) based on the ACF of each subband audio signal 20. The autocorrelation information 50 for a given subband includes (e.g., composed of) a lag value T and an autocorrelation value ρ(T). In other words, for each subband, one value of the lag T and the corresponding (potentially normalized) autocorrelation value (ACF value) ρ(T) are output (e.g., transmitted) as autocorrelation information 50, which is part of the encoded representation. Here, the lag value T corresponds to the delay value at which the ACF reaches a local maximum, and the autocorrelation value ρ(T) corresponds to this local maximum. In other words, the autocorrelation information for a given subband may include the delay value (i.e., x-coordinate) and autocorrelation value (i.e., y-coordinate) of the local maximum of the ACF.

[0034] Therefore, the encoded representation of an audio signal includes the spectral envelope of the audio signal and autocorrelation information for each of the subbands. The autocorrelation information for a given subband includes representations of the lag value T and the autocorrelation value ρ(T). The encoded representation corresponds to the output of the encoder. In some implementations, the encoded representation may further include waveform information related to the waveform of the audio signal and / or one or more waveforms of the subband audio signals.

[0035] The above procedure defines an encoding function (or encoding map) h that maps the input audio signal to its encoded representation.

[0036] As described above, the spectral envelope and autocorrelation information for subband audio signals may be determined and output at different update rates (sample rates). For example, the spectral envelope may be determined at a first update rate, and the autocorrelation information for multiple subband audio signals may be determined at a second update rate different from the first update rate. The representations of the spectral envelope and autocorrelation information (for all subbands) may be written to the bitstream at their respective update rates (sample rates). In this case, the encoded representation may be related to a portion of the bitstream output by the encoder. In this regard, it should be noted that at each point in time, the current spectral envelope and the current set of autocorrelation information (one for each subband) can be defined by the bitstream and become an encoded representation. Alternatively, the representations of the spectral envelope and autocorrelation information (for all subbands) may be updated at their respective update rates in each output unit of the encoder. In this case, each output unit of the encoder (e.g., an encoded frame) corresponds to an instance of the encoded representation. The representations of the spectral envelope and autocorrelation information may be the same across a series of consecutive output units, depending on their respective update rates.

[0037] Preferably, the first update rate is higher than the second update rate. For example, the first update rate R1 may be R1 = 1 / (2.5 ms) and the second update rate R2 may be R2 = 1 / (20 ms), resulting in the updated representation of the spectral envelope being output every 2.5 ms, while the updated representation of the autocorrelation information is output every 20 ms. With respect to a portion of the audio signal (e.g., a frame), the spectral envelope may be determined for every nth portion (e.g., each segment), while the autocorrelation information may be determined for every mth portion (m > n).

[0038] The encoded representation may be output as a sequence of frames of a specific frame length. Among other factors, the frame length may depend on the first and / or second update rates. Considering a frame with a length of a first period L1 (e.g., 2.5 ms) corresponding to a first update rate R1 (e.g., 1 / (2.5 ms)) via L1 = 1 / R1, this frame contains one representation of the spectral envelope and one set of autocorrelation information representations (one per subband audio signal). For the first and second update rates of 1 / (2.5 ms) and 1 / (20 ms), respectively, the autocorrelation information is the same for eight consecutive frames of the encoded representation, respectively. Generally, assuming that R1 and R2 are appropriately chosen to have an integer ratio, the autocorrelation information is the same for R1 / R2 consecutive frames of the encoded representation. On the other hand, considering a frame with a second period length L2 (e.g., 20ms) corresponding to a second update rate R2 (e.g., 1 / (20ms)) via L2 = 1 / R2, this frame contains one set of autocorrelation information representations and R1 / R2 (e.g., 8) representations of spectral envelopes.

[0039] In some implementations, different update rates may be applied to different subbands, meaning that autocorrelation information for different subband audio signals may be generated and output at different update rates.

[0040] Figure 2 is a flowchart showing an example of an encoding method 200 according to an embodiment of the present disclosure. This method may be implemented by the encoder 100 described above, which receives an audio signal as input.

[0041] In step S210, multiple subband audio signals are generated based on the audio signal. This may include spectrally decomposing the audio signal, in which case this step may be performed according to the operation of the filter bank 15 described above. Alternatively, this may include spectrally and / or temporally flattening the audio signal, optionally windowing the flattened audio signal by a window function, and spectrally decomposing the resulting signal into multiple subband audio signals.

[0042] In step S220, the spectral envelope of the audio signal is determined (e.g., calculated). This step may be performed according to the operation of the envelope extraction block 25 described above.

[0043] In step S230, for each subband audio signal, autocorrelation information for the subband audio signal is determined based on the ACF of the subband audio signal. This step may be performed according to the operation of the autocorrelation block 55 described above.

[0044] In step S240, an encoded representation of the audio signal is generated. The encoded representation includes a representation of the spectral envelope of the audio signal and a representation of autocorrelation information for each of the multiple subband audio signals.

[0045] Next, we will describe an example of the implementation details of the steps in Method 200.

[0046] For example, generating multiple subband audio signals as described above may include (or mean) spectrally decomposing the audio signals, for example, using a filter bank. In this case, determining the autocorrelation function for a given subband audio signal may include determining the subband envelope of the subband audio signal. The subband envelope may be determined by taking the magnitude value of the subband audio signal. The ACF itself may be calculated for real-valued (windowed after envelope flattening) subband signals.

[0047] Assuming that the subband filter response becomes complex by a Fourier transform supported in essence at positive frequencies, the subband signal becomes complex. The subband envelope can then be determined by taking the magnitude of the complex subband signal. This subband envelope has the same number of samples as the subband signal and is still oscillating to some extent. Optionally, the subband envelope can be downsampled by, for example, calculating the triangular window weighted sum of squares of the envelope in segments of a specific length (e.g., 5 ms, rising edge 2.5 ms, falling edge 2.5 ms) for every half-shift of a specific length (e.g., 2.5 ms) along the signal, and then taking the square root of this sequence to obtain the downsampled subband envelope. This may be said to correspond to the definition of an "rms envelope". The triangular window can be normalized such that a constant envelope of value 1 gives a sequence of 1s. Other methods for determining the subband envelope are similarly achievable, such as half-wave rectification and subsequent low-pass filtering in the case of a real-valued subband signal. In either case, the subband envelope can be said to transmit information about the energy within the subband signal (at a selected update rate).

[0048] Next, the subband audio signal may be envelope-flattened based on the subband envelope. For example, a new full sample-rate envelope signal may be created by linearly interpolating the downsampled values ​​to reach the fine-structured signal (carrier) from which the ACF data is calculated, and then dividing the original (complex-valued) subband signal with this linearly interpolated envelope.

[0049] Next, the envelope-flattened subband audio signal may be windowed by an appropriate window function. Finally, the ACF of the windowed envelope-flattened subband audio signal is determined (e.g., calculated). In some implementations, determining the ACF for a given subband audio signal may further include normalizing the ACF of the windowed envelope-flattened subband audio signal by the autocorrelation function of the window function.

[0050] In Figure 3, the upper curve 310 shows the real values ​​of the envelope-flattened subband signals after windowing, which are used to calculate the ACF. The lower solid curve 320 shows the real values ​​of the complex ACF.

[0051] The main concept here is to find the largest local maximum of the ACF of the subband signal from among the local maximums that are above the absolute value of the ACF of the absolute value of the impulse response of the (complex-valued) subband filter (i.e., the corresponding BPF in the filter bank). For the ACF of a complex-valued subband signal, the real value of the ACF may be considered at this point. To avoid picking lag, which is related to the center frequency of the subband rather than the characteristics of the input signal, it may be necessary to find the largest local maximum above the ACF of the absolute value of the impulse response. As a final adjustment, the maximum value may be divided by the value of the ACF of the window function used for the subband ACF window (assuming that the ACF of the subband signal itself is normalized so that, for example, the autocorrelation value at zero delay is normalized to 1). This results in better use of the interval between 0 and 1, where ρ(T)=1 is the maximum tonality.

[0052] Therefore, determining autocorrelation information for a given subband audio signal based on the ACF of the subband audio signal may further include comparing the ACF of the subband audio signal with the ACF of the absolute value of the impulse response of each bandpass filter associated with the subband audio signal. The ACF of the absolute value of the impulse response of each bandpass filter associated with the subband audio signal is shown by the solid curve 330 at the bottom of Figure 3. The autocorrelation information is then determined based on the highest local maximum value of the ACF of the subband signal that is above the ACF of the absolute value of the impulse response of each bandpass filter associated with the subband audio signal. At the bottom of Figure 3, the local maximum values ​​of the ACF are shown by crosses, and the selected highest local maximum value of the ACF of the subband signal that is above the ACF of the absolute value of the impulse response of each bandpass is shown by circles. Optionally, the selected local maximum value of the ACF may be normalized by the ACF value of the window function ACF (for example, assuming that the ACF itself is normalized such that the zero-delay autocorrelation value is normalized to 1). The selected best local maximum of the ACF after normalization is indicated by the lower asterisk in Figure 3, and the dashed curve 340 represents the ACF of the window function.

[0053] The autocorrelation information determined at this stage may include the autocorrelation and delay values ​​(i.e., y-coordinates and x-coordinates) of the selected (normalized) best local maximum of the ACF of the subband audio signal.

[0054] A similar encoding format may be defined in the framework of an LPC-based vocoder. In this case, the autocorrelation information is extracted from subband signals that are affected by at least some degree of spectral and / or temporal flattening. Unlike the example above, this is done by creating (perceptually weighted) LPC residuals, windowing them, and decomposing them into subbands to obtain multiple subband audio signals. Subsequently, the ACF is calculated and the lag and autocorrelation values ​​for each subband audio signal are extracted.

[0055] For example, generating multiple subband audio signals may involve applying spectral and / or temporal flattening to the audio signal (for example, by using an LPC filter to generate perceptually weighted LPC residuals from the audio signal). Subsequently, the flattened audio signal may be windowed using a window function, and the windowed flattened audio signal may be spectrally decomposed into multiple subband audio signals. As described above, the result of temporal and / or spectral flattening may correspond to perceptually weighted LPC residuals, which then undergo windowing and spectral decomposition into subbands. The perceptually weighted LPC residuals may be, for example, pink LPC residuals.

[0056] [Decrypt] This disclosure relates to audio decoding based on an analytical synthesis method. At its most abstract level, it is assumed that an encoded map h from the signal to a perceptually motivated domain is given such that the original audio signal x is represented by y = h(x). In the best case, a simple measure of distortion, such as least squares in the perceptual domain, would be a good predictor of the subjective difference measured by a group of listeners.

[0057] The remaining problem is to design a decoder q that maps y (in its encoded and decoded versions) to an audio signal z=d(y). For this, the concept of synthesis by analysis can be used, which involves "finding the waveform that comes closest to producing a given image." The goal is for z and x to sound the same, and as a result, the decoder should solve the inverse problem h(z)=y=h(x). Regarding the construction of the map, d should be approximated by the left reciprocal of h,

[0058]

number

[0059] Figure 4 is a schematic block diagram illustrating an example of synthesis by analysis to determine a decoding function (or decoding map) d given an encoding function (or encoding map) h. The original audio signal x, 410 receives an encoding map h, 415, producing an encoded representation y, 420, where y = h(x). The encoded representation y may be defined in the perceptual domain. The objective is to find a decoding function (decoding mapping) d, 425 that maps the encoded representation y to the reconstructed audio signal z, 430, which has the property that applying the encoding mapping h, 435 to the reconstructed audio signal z produces an encoded representation h(z), 440 that substantially matches the encoded representation y = h(x). Here, "substantially matching" may mean, for example, "matching to a given margin." In other words, given an encoding map h, the objective is,

[0060]

number

[0061] Figure 5 is a flowchart illustrating an example of a decoding method 500 according to an embodiment of the present disclosure, which follows an analytical synthesis method. Method 500 is a method for decoding an audio signal from an encoded representation of an (original) audio signal. The encoded representation is assumed to include a representation of the spectral envelope of the original audio signal and a representation of autocorrelation information for each of several subband audio signals of the original audio signal. The autocorrelation information for a given subband audio signal is based on the ACF of the subband audio signal.

[0062] In step S510, the encoded representation of the audio signal is received.

[0063] In step S520, the spectral envelope and autocorrelation information are extracted from the encoded representation of the audio signal.

[0064] In step S530, the reconstructed audio signal is determined based on the spectral envelope and autocorrelation information. Here, the reconstructed audio signal is determined such that the autocorrelation function of each of the multiple subband signals of the reconstructed subband audio signal (substantially) satisfies a condition derived from the autocorrelation information for the corresponding subband audio signals of the audio signal. This condition may be, for example, that for each subband audio signal of the reconstructed audio signal, at the lag value (e.g., delay value) indicated by the autocorrelation information for the corresponding subband audio signal of the audio signal, the ACF value of the subband audio signal of the reconstructed audio signal substantially matches the autocorrelation value indicated by the autocorrelation information for the corresponding subband audio signal of the audio signal. This may mean that the decoder can determine the ACF of the subband audio signals in the same way that the encoder does. This may include, in whole or in part, flattening, windowing, and normalization. In one implementation, the reconstructed audio signal may be determined such that, for each subband audio signal of the reconstructed audio signal, the autocorrelation and lag values ​​(e.g., delay values) of the ACF of the subband audio signal of the reconstructed audio signal substantially match the autocorrelation and lag values ​​indicated by the autocorrelation information of the corresponding subband audio signal of the original audio signal. This may mean that the decoder can determine the autocorrelation information for each subband signal of the reconstructed audio signal in the same way that the encoder does. In these implementations where the encoded representation also includes waveform information, the reconstructed audio signal may be determined further based on the waveform information. The subband audio signals of the reconstructed audio signal may be generated in the same way that the encoder does. For example, this may include a sequence of spectral decomposition, or flattening, windowing, and spectral decomposition.

[0065] Preferably, the determination of the reconstructed audio signal in step S530 also takes into account the spectral envelope of the original audio signal. The reconstructed audio signal may then be further determined such that, for each subband audio signal of the reconstructed subband audio signal, the measured (e.g., estimated or calculated) signal power of the subband audio signal of the reconstructed audio signal substantially matches the signal power for the corresponding subband audio signal of the original audio signal as indicated by its spectral envelope.

[0066] As can be seen from the above, the proposed method 500 can be said to be derived from an analytical synthesis technique in that it attempts to find a reconstructed audio signal z that (substantially) satisfies at least one condition derived from the encoded representation y=h(x) of the original audio signal x, where h is the encoding map used by the encoder. In some implementations, it can even be said that the proposed method operates according to an analytical synthesis technique in that it attempts to find a reconstructed audio signal z whose encoded representation h(z) substantially matches the encoded representation y=h(x) of the original audio signal x. In other words, the decoding method is

[0067]

number

[0068] [Implementation Example 1: Parametric synthesis or per signal iteration] The inverse problem h(z)=y is h(z n ) is h(z n-1 z becomes closer to y than ) n-1 Update map z to fix the problem. n =f(z n-1This can be solved by an iterative method assuming y). For example, the starting point of the iteration (i.e., the initial candidate for the reconstructed audio signal) may be a random noise signal (e.g., white noise), or it may be determined based on the encoded representation of the audio signal (e.g., as a manually made initial guess). In the latter case, the initial candidate for the reconstructed audio signal may be related to a learned guess made based on spectral envelope and / or autocorrelation information for multiple subband audio signals. In these implementations where the encoded representation includes waveform information, the learned guess may be made further based on the waveform information.

[0069] More specifically, the reconstructed audio signal in this implementation example is determined by an iterative procedure that starts from an initial candidate for the reconstructed audio signal and generates each intermediate reconstructed audio signal at each iteration. At each iteration, an update map is applied to the intermediate reconstructed audio signal to obtain the intermediate reconstructed audio signal for the next iteration. The update map is selected such that the difference between the encoded representation of the intermediate reconstructed audio signal and the encoded representation of the original audio signal decreases continuously from one iteration to the next. For this purpose, an appropriate difference metric for the encoded representation (e.g., spectral envelope, autocorrelation information) may be defined and used to evaluate the difference. The encoded representation of the intermediate reconstructed audio signal may also be the encoded representation obtained when the intermediate reconstructed audio signal is subjected to the same encoding scheme that yields the encoded representation of the audio signal.

[0070] If the procedure searches for a reconstructed audio signal that satisfies at least one condition derived from (multiple) autocorrelation information, the update map may be selected such that the autocorrelation function of the subband audio signals of the intermediate reconstruction of the audio signal satisfies the respective conditions derived from the autocorrelation information for the corresponding subband audio signals of the audio signal, and / or the difference between the measured signal power of the subband audio signals of the reconstructed audio signal and the signal power for the corresponding subband audio signals of the audio signal, as indicated by the spectral envelope, decreases from one iteration to the next. If both autocorrelation information and spectral envelopes are considered, an appropriate difference metric and the difference between the signal power for the subband audio signals may be defined for the degree to which the condition is satisfied.

[0071] [Implementation Example 2: Generative Model Based on Machine Learning] Another option made possible by modern machine learning methods is to train a machine learning-based generative model (or, simply a generative model) for audio x conditioned on data y. That is, given a large set of examples of (x,y) (where y=h(x)), a parametric conditional distribution p(x|y) from y to x is trained. The decoding algorithm may then consist of sampling from the distribution z~p(x|y).

[0072] This option has been shown to be particularly advantageous when h(x) is an audio vocoder and p(x|y) is defined by a sequential generative model sample recurrent neural network (RNN). However, other generative models such as variational autoencoders or generative adversarial models are equally relevant to this task. Therefore, without intent to limit, machine learning-based generative models can be any one of the following: recurrent neural networks, variational autoencoders, or generative adversarial models (e.g., Generative Adversarial Networks, GANs).

[0073] In this implementation example, determining the reconstructed audio signal based on the spectral envelope and autocorrelation information involves applying a machine learning-based generative model that receives the spectral envelope of the audio signal and autocorrelation information for each of the multiple subband audio signals of the audio signal as input, and generates and outputs the reconstructed audio signal. The encoded representation also includes waveform information. In these implementations, the machine learning-based generative model may further receive waveform information as input.

[0074] As described above, a machine learning-based generative model may include a parametric conditional distribution p(x|y) that associates an encoded representation y of an audio signal and the corresponding audio signal x with their respective probabilities p. Then, determining the reconstructed audio signal may include sampling from the parametric conditional distribution p(x|y) for the encoded representation of the audio signal.

[0075] During the training phase, prior to decoding, a machine learning-based generative model may be conditioned / trained on a dataset of multiple audio signals and their corresponding encoded representations. If the encoded representations also include waveform information, the machine learning-based generative model may also be conditioned / trained using the waveform information.

[0076] Figure 6 is a flowchart illustrating an exemplary implementation 600 for step S530 in the decoding method 500 of Figure 5. In particular, implementation 600 relates to the implementation of step S530 for each subband.

[0077] In step 610, multiple reconstructed subband audio signals are determined based on the spectral envelope and autocorrelation information. Here, for each reconstructed subband audio signal, the autocorrelation function of the reconstructed subband audio signal is determined to satisfy the conditions derived from the autocorrelation information for the corresponding subband audio signal of the audio signal. In some implementations, the multiple reconstructed subband audio signals are determined for each reconstructed subband audio signal such that the autocorrelation information for the reconstructed subband audio signal substantially matches the autocorrelation information for the corresponding subband audio signal.

[0078] Preferably, the determination of the multiple reconstructed subband audio signals in step S610 also takes into account the spectral envelope of the original audio signal. The multiple reconstructed subband audio signals are then further determined such that, for each reconstructed subband audio signal, the measured (e.g., estimated, calculated) signal power of the reconstructed subband audio signal substantially matches the signal power for the corresponding subband audio signal as indicated by the spectral envelope.

[0079] In step S620, the reconstructed audio signal is determined by spectral synthesis based on multiple reconstructed subband audio signals.

[0080] Implementation examples 1 and 2 described above may also be applied to the implementation of step S530 for each subband. In implementation example 1, each reconstructed subband audio signal may be determined by an iterative procedure that starts from an initial candidate for the reconstructed subband audio signal and generates each intermediate reconstructed subband audio signal in each iteration. In each iteration, an update map may be applied to the intermediate reconstructed subband audio signal to obtain an intermediate reconstructed subband audio signal for the next iteration such that the difference between the autocorrelation information for the intermediate reconstructed subband audio signal and the autocorrelation information for the corresponding subband audio signal becomes progressively smaller from one iteration to the next, or so that the reconstructed subband audio signal better satisfies the respective conditions derived from the autocorrelation information for each corresponding subband audio signal of the audio signal.

[0081] Furthermore, the spectral envelope may also be considered at this point. That is, the updated map may be such that the (congruent) differences between the respective signal powers of the subband audio signals and the (congruent) differences between the respective items of the autocorrelation information decrease continuously. This may mean defining an appropriate difference metric for evaluating the (congruent) differences. In addition, the same explanation given above for Implementation Example 1 may apply in this case.

[0082] When applying Implementation Example 2 to the implementation for each subband in step S530, determining multiple reconstructed subband audio signals based on spectral envelope and autocorrelation information may include applying a machine learning-based generative model that receives the spectral envelope of the audio signal and autocorrelation information for each of the multiple subband audio signals of the audio signal as input, and generates and outputs multiple reconstructed subband audio signals. In addition, the same explanation given above for Implementation Example 2 may apply in this case as well.

[0083] This disclosure further relates to an encoder for encoding an audio signal, which is capable of and adapted to perform the encoding methods described herein. An example of such an encoder 700 is schematically shown in Figure 7 in block diagram form. The encoder 700 includes a processor 710 and a memory 720 coupled to the processor 710. The processor 710 is adapted to perform a step of any one of the encoding methods described herein. For this purpose, the memory 720 may contain each instruction for the processor 710 to perform. The encoder 700 may further include an interface 730 for receiving an input audio signal 740 to be encoded and / or for outputting an encoded representation 750 of the audio signal.

[0084] This disclosure further relates to a decoder for decoding an audio signal from an encoded representation of an audio signal, which is capable of and adapted to perform the decoding methods described herein. An example of such a decoder 800 is schematically shown in Figure 8 in block diagram form. The decoder 800 includes a processor 810 and a memory 820 coupled to the processor 810. The processor 810 is adapted to perform a step of any one of the decoding methods described herein. For this purpose, the memory 820 may contain each instruction for the processor 810 to perform. The encoder 800 may further include an interface 830 for receiving an input encoded representation 840 of the audio signal to be decoded and / or for outputting a decoded (i.e., reconstructed) audio signal 850.

[0085] This disclosure further relates to computer programs that, when the instructions are executed, cause a computer to perform the encoding or decoding methods described through this disclosure.

[0086] Finally, this disclosure also relates to a computer-readable storage medium for storing the above-mentioned computer program.

[0087] [interpretation] Unless otherwise specified, as will be apparent from the following discussion, any discussion using terms such as “processing,” “calculation,” “operation,” “determination,” and “analysis” throughout this disclosure is understood to refer to the operation and / or process of a computer or computer system or similar electronic computing device that manipulates and / or converts data expressed as physical quantities, such as electron quantities, into other data similarly expressed as physical quantities.

[0088] Similarly, the term “processor” may refer to any device or part of a device that processes electronic data from, for example, registers and / or memory and converts it into other electronic data that can be stored, for example, in registers and / or memory. “Computer” or “calculator” or “computing platform” may include one or more processors.

[0089] The methods described herein are executable by one or more processors that accept computer-readable (also called machine-readable) code, which includes a set of instructions that, when executed by one or more processors, perform at least one of the methods described herein. This includes any processor capable of executing a set of instructions (sequential or otherwise) that specify an action to be performed. Thus, one example is a typical processing system comprising one or more processors. Each processor may include one or more of the following: a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem, which includes main RAM and / or static RAM and / or ROM. A bus subsystem for communication between components may be included. Furthermore, the processing system may be a distributed processing system having processors connected by a network. If the processing system requires a display, a display such as a liquid crystal display (LCD) or a cathode ray tube (CRT) display may be included. If manual data entry is required, the processing system also includes an input device such as one or more of the following: an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, etc. The processing system may also include a storage system such as a disk drive unit. In some configurations, the processing system may also include a sound output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier that carries computer-readable code (e.g., software) which, when executed by one or more processors, causes one or more of the methods described herein to be performed. Note that if the method involves several elements, for example, several steps, the order of such elements is not implied unless specifically mentioned.The software may reside permanently on the hard disk, or it may reside entirely or at least partially in RAM and / or the processor while it is being executed by the computer system. Thus, the memory and processor also constitute a computer-readable carrier for carrying computer-readable code. Furthermore, the computer-readable carrier may form a computer program product or be included in a computer program product.

[0090] In alternative exemplary embodiments, one or more processors may operate as standalone devices or, for example, be connected to other processors, or, for example, be networked, and in a networked deployment, one or more processors may operate at the capacity of a server or user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. One or more processors may form a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a mobile phone, a web device, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify the actions to be taken by the machine.

[0091] It should be noted that the term “machine” is also interpreted to include any set of machines that individually or collectively perform a set (or set) of instructions for performing one or more of the methodologies discussed herein.

[0092] Accordingly, each exemplary embodiment of the methods described herein is in the form of a computer-readable carrier that carries a set of instructions, for example, a computer program to run on one or more processors, for example, one or more processors that are part of a web server configuration. Accordingly, as will be recognized by those skilled in the art, exemplary embodiments of the disclosure may be embodied as a method, an apparatus such as a special-purpose apparatus, an apparatus such as a data processing system, or a computer-readable carrier, for example, a computer program product. The computer-readable carrier, when executed on one or more processors, carries computer-readable code that includes a set of instructions causing the processors to carry out the method. Accordingly, embodiments of the disclosure may take the form of a method, an exemplary embodiment entirely of hardware, an exemplary embodiment entirely of software, or an exemplary embodiment combining embodiments of software and hardware. Furthermore, the disclosure may take the form of a carrier that carries computer-readable program code embodied on a medium (for example, a computer program product on a computer-readable storage medium).

[0093] The software may be further transmitted or received over a network via a network interface device. In exemplary embodiments, the carrier is a single medium, but the term “carrier” should be interpreted to include a single or multiple mediums (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term “carrier” should also be interpreted to include any medium capable of storing, encoding, or carrying a set of instructions for execution by one or more processors, causing one or more processors to execute one or more of the methodologies of this disclosure. The carrier may take many forms, but is not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wires, and optical fibers, including wires that constitute a bus subsystem. The transmission media may also take the form of sound waves or light waves, such as those generated between radio waves and infrared data communications. For example, the term “carrier” includes, but is not limited to, solid memory, computer products embodied in optical and magnetic media, media carrying a propagation signal detectable by at least one processor or one or more processors and representing a set of instructions that, when executed, perform a method, and transmission media in a network carrying a propagation signal detectable by at least one of the one or more processors and representing a set of instructions.

[0094] It is understood that, in one exemplary embodiment, the steps of the described method are performed by a suitable processor (or more processors) of a processing system (e.g., a computer) that executes instructions (computer-readable code) stored in storage. It is also understood that this disclosure is not limited to any particular implementation or programming technique, and that this disclosure may be implemented using any suitable technique for implementing the functions described herein. This disclosure is not limited to any particular programming language or operating system.

[0095] Throughout this disclosure, any reference to “one exemplary embodiment,” “several exemplary embodiments,” or “exemplary embodiments” means that any particular feature, structure, or characteristic described in relation to an exemplary embodiment is included in at least one exemplary embodiment of this disclosure. Therefore, the appearance of the phrases “one exemplary embodiment,” “several exemplary embodiments,” or “exemplary embodiments” in various parts of this disclosure does not necessarily all refer to the same exemplary embodiment. Furthermore, any particular feature, structure, or characteristic may be combined in any appropriate manner in one or more exemplary embodiments, as will be apparent to those skilled in the art from this disclosure.

[0096] The use of ordinal adjectives such as “first,” “second,” and “third” used herein to describe common objects simply indicates that different instances of similar objects are being referred to, and is not intended to imply that the objects described in this manner must be in a given order, temporally, spatially, in rank, or otherwise.

[0097] In the following claims and descriptions herein, the terms “comprising,” “comprised of,” or “which comprises” are open terms meaning that at least an element / feature is included, but other elements are not excluded. Therefore, when used in the claims, the term “includes” should not be interpreted as being limited to the enumerated means, elements, or steps. For example, an apparatus including A and B should not be limited to a device consisting only of elements A and B. Similarly, when used herein, the terms “including” or “which includes” or “that includes” are also open terms meaning that at least an element / feature relating to the term is included, but other terms are not excluded. Therefore, “including” is synonymous with “comprising” and means the same thing.

[0098] In the above description of exemplary embodiments of the disclosure, it should be recognized that, for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various aspects of the invention, various features of the disclosure are sometimes summarized in a single exemplary embodiment, figure, or description thereof. However, this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than are expressly described in each claim. Rather, as reflected in the following claims, the aspects of the invention are fewer than all the features of a single exemplary embodiment disclosed above. Accordingly, the claims following the description are expressly incorporated herein so that each claim stands alone as a separate exemplary embodiment of the disclosure.

[0099] Furthermore, while some exemplary embodiments described herein include certain features but not others included in other exemplary embodiments, as will be understood by those skilled in the art, combinations of features from different exemplary embodiments are within the scope of this disclosure and constitute different exemplary embodiments. For example, any of the exemplary embodiments described in the following claims may be used in any combination.

[0100] Numerous specific details are given in the description provided herein. However, it should be understood that exemplary embodiments of this disclosure may be carried out without these specific details. In other cases, well-known methods, structures and techniques are not given in detail so as not to obscure the understanding of this description.

[0101] Therefore, while the best mode of this disclosure is described, those skilled in the art will recognize that other further modifications may be made without departing from the true intent of the disclosure and intend to request all such changes and modifications that fall within the scope of this disclosure. For example, the formulas above merely represent procedures that may be used. Functions may be added or removed from the block diagram, and operations may be swapped between function blocks. Steps may be added or removed to methods within the scope of this disclosure.

[0102] Various aspects and embodiments of this disclosure can be identified from the exemplary embodiments (EEE) listed below.

[0103] EEE1. A method for encoding an audio signal, A step of generating multiple subband audio signals based on an audio signal, The steps include determining the spectral envelope of the audio signal, For each subband audio signal, the autocorrelation information for the subband audio signal is determined based on the autocorrelation function of the subband audio signal. The step of generating an encoded representation of an audio signal, the encoded representation including a representation of the spectral envelope of the audio signal and a representation of autocorrelation information for multiple subband audio signals. A method that includes this.

[0104] EEE2. The spectral envelope is determined based on multiple subband audio signals, as described in EEE1.

[0105] EEE3. The method according to EEE1 or 2, wherein the autocorrelation information for a given subband audio signal includes the lag value and / or the autocorrelation value for each subband audio signal.

[0106] EEE4. The lag value corresponds to the delay value at which the autocorrelation function reaches a local maximum, and the autocorrelation value corresponds to this local maximum, as described in EEE3.

[0107] EEE5. The method according to any one of EEE1 to 4, wherein the spectral envelope is determined at a first update rate, and autocorrelation information for multiple subband audio signals is determined at a second update rate, and the first and second update rates are different from each other.

[0108] EEE6. The first renewal rate is higher than the second renewal rate, as described in EEE5.

[0109] EEE7. Generating multiple subband audio signals is possible. Apply spectral and / or temporal flattening to the audio signal. The flattened audio signal is then windowed. The method according to any one of EEE1 to 6, comprising spectrally decomposing a flattened audio signal after windowing into multiple subband audio signals.

[0110] EEE8. Generating multiple subband audio signals involves spectrally decomposing the audio signal. Determining the autocorrelation function for a given subband audio signal is: Determine the subband envelope of the subband audio signal. Based on the subband envelope, the subband audio signal is envelope-flattened. The subband audio signal, which has been envelope-flattened using a window function, is then windowed. The method according to any one of EEE1 to 6, comprising determining the autocorrelation function of the envelope-flattened subband audio signal after windowing.

[0111] EEE9. Determining the autocorrelation function for a given subband audio signal is: The method according to EEE7 or 8, further comprising normalizing the autocorrelation function of the envelope-flattened subband audio signal after windowing by the autocorrelation function of the window function.

[0112] EEE10. Determining the autocorrelation function for a subband audio signal based on the autocorrelation function of a given subband audio signal is possible. The autocorrelation function of the subband audio signal is compared with the autocorrelation function of the absolute values ​​of the impulse responses of each bandpass filter associated with the subband audio signal. The method according to any one of EEE1 to 9, comprising determining autocorrelation information based on the highest local maximum of the autocorrelation function of the subband signal that is above the autocorrelation function of the absolute value of the impulse response of each bandpass filter associated with the subband audio signal.

[0113] EEE11. Determining the spectral envelope is the method described in any one of EEE1 to 10, which includes measuring the signal power for each of several subband audio signals.

[0114] EEE12. A method for decoding an audio signal from an encoded representation of an audio signal, wherein the encoded representation includes a representation of the spectral envelope of the audio signal and a representation of autocorrelation information for each of a plurality of subband audio signals generated from the audio signal, and the autocorrelation information for a given subband audio signal is based on the autocorrelation function of the subband audio signal, and the method The steps include receiving an encoded representation of an audio signal, Steps include extracting spectral envelope and autocorrelation information from the encoded representation of an audio signal, The steps include determining the reconstructed audio signal based on the spectral envelope and autocorrelation information, and Includes, A method for reconstructing an audio signal, in which the autocorrelation function for each of the multiple subband audio signals generated from the reconstructed audio signal is determined such that it satisfies conditions derived from the autocorrelation information for the corresponding subband audio signals generated from the audio signal.

[0115] EEE13. The method according to EEE12, wherein the reconstructed audio signal is further determined such that, for each subband audio signal of the reconstructed audio signal, the measured signal power of the subband audio signal of the reconstructed audio signal substantially matches the signal power for the corresponding subband audio signal of the audio signal indicated by the spectral envelope.

[0116] EEE14. The reconstructed audio signal is determined by an iterative procedure that starts from an initial candidate for the reconstructed audio signal and generates each intermediate reconstructed audio signal at each iteration. The method according to EEE12 or 13, wherein in each iteration, an update map is applied to the intermediate reconstructed audio signal to obtain the intermediate reconstructed audio signal for the next iteration such that the difference between the encoded representation of the intermediate reconstructed audio signal and the encoded representation of the audio signal decreases sequentially from one iteration to the next iteration.

[0117] EEE15. The method described in EEE14, in which initial candidates for the reconstructed audio signal are determined based on the encoded representation of the audio signal.

[0118] EEE16. An initial candidate for the reconstructed audio signal is white noise, as described in EEE14.

[0119] EEE17. The method according to EEE12 or 13, wherein determining a reconstructed audio signal based on spectral envelope and autocorrelation information includes applying a machine learning-based generative model that receives the spectral envelope of an audio signal and autocorrelation information for each of several subband audio signals of the audio signal as input, and generates and outputs a reconstructed audio signal.

[0120] EEE18. A machine learning-based generative model includes an encoded representation of an audio signal and a parametric conditional distribution that associates the corresponding audio signals with their respective probabilities. Determining the reconstructed audio signal is a method according to EEE17, which includes sampling from a parametric conditional distribution of encoded representations of the audio signal.

[0121] EEE19. The method according to EEE17 or 18, further comprising the step of training a machine learning-based generative model on a dataset of multiple audio signals and corresponding encoded representations of the audio signals during the training phase.

[0122] EEE20. A machine learning-based generative model is one of a recurrent neural network, a variational autoencoder, or an adversarial generative model, as described in any one of EEE17 to 19.

[0123] EEE21. Determining a reconstructed audio signal based on spectral envelope and autocorrelation information is Multiple reconstructed subband audio signals are determined based on spectral envelope and autocorrelation information. This includes determining a reconstructed audio signal based on multiple reconstructed subband audio signals by spectral synthesis, The method according to EEE12, wherein multiple reconstructed subband audio signals are determined such that, for each reconstructed subband audio signal, the autocorrelation function of the reconstructed subband audio signal satisfies the conditions derived from the autocorrelation information for the corresponding subband audio signal.

[0124] EEE22. The method according to EEE21, wherein for each reconstructed subband audio signal, the measured signal power of the reconstructed subband audio signal is further determined so that it substantially matches the signal power of the corresponding subband audio signal indicated by the spectral envelope.

[0125] EEE23. Each reconstructed subband audio signal is determined by an iterative procedure that starts from an initial candidate for the reconstructed subband audio signal and generates each intermediate reconstructed subband audio signal in each iteration. The method according to EEE21 or 22, wherein in each iteration, an update map is applied to the intermediate reconstructed subband audio signal to obtain the intermediate reconstructed subband audio signal for the next iteration such that the difference between the autocorrelation information for the intermediate reconstructed subband audio signal and the autocorrelation information for the corresponding subband audio signal decreases continuously from one iteration to the next.

[0126] EEE24. The method according to EEE21 or 22, wherein determining multiple reconstructed subband audio signals based on spectral envelope and autocorrelation information includes applying a machine learning-based generative model that receives the spectral envelope of an audio signal and autocorrelation information for each of the multiple subband audio signals of the audio signal as input, and generates and outputs multiple reconstructed subband audio signals.

[0127] EEE25. An encoder for encoding an audio signal, comprising a processor and memory coupled to the processor, wherein the processor is adapted to perform the steps of the method described in any one of EEE1 to 11.

[0128] EEE26. A decoder for decoding an audio signal from an encoded representation of an audio signal, comprising a processor and memory coupled to the processor, wherein the processor is adapted to perform steps of any one of EEE12 to 24.

[0129] EEE27. A computer program that includes instructions that, when executed, cause a computer to perform any of the actions described in any one of EEE1 through EEE24.

[0130] A computer-readable storage medium containing computer programs as described in EEE28 and EEE27.

Claims

1. A method for encoding audio signals, The steps include generating a plurality of subband audio signals based on the aforementioned audio signal, The steps include determining the spectral envelope based on the plurality of subband audio signals, The steps include determining autocorrelation information for each subband audio signal, The steps include encoding the spectral envelope and the autocorrelation information into an encoded representation. A method that includes this.

2. The method according to claim 1, further comprising the step of outputting a bitstream based on the encoded representation.

3. The method according to claim 1, wherein the autocorrelation information includes the autocorrelation value for the subband audio signal.

4. The method according to claim 3, wherein the autocorrelation value corresponds to the local maximum of the autocorrelation function.

5. The method according to claim 4, wherein the autocorrelation information includes lag values.

6. The spectral envelope is determined by a first update rate. The aforementioned autocorrelation information is determined by a second update rate. The method according to claim 1, wherein the first update rate is different from the second update rate.

7. The method according to claim 6, wherein the first update rate is higher than the second update rate.

8. The method according to claim 1, wherein generating the plurality of subband audio signals includes flattening the audio signals.

9. The method of claim 8, wherein generating the plurality of subband audio signals includes decomposing the flattened audio signal into the plurality of subband audio signals.

10. A method for decoding an encoded representation of an audio signal, The steps include determining a plurality of reconstructed subband audio signals based on the spectral envelope and autocorrelation information of the encoded representation, The step of generating a reconstructed audio signal based on the plurality of reconstructed subband audio signals. Includes, A method wherein the spectral envelope is determined based on a plurality of original subband audio signals, the plurality of original subband audio signals are generated based on the audio signals, and the autocorrelation information is determined for each original subband audio signal.

11. The method according to claim 10, wherein the reconstructed audio signal is determined via a machine learning-based generative model and / or based on spectral synthesis.

12. The method according to claim 10, wherein the autocorrelation information includes an autocorrelation value for each of the original subband audios.

13. The method according to claim 12, wherein the autocorrelation value corresponds to the local maximum value of the autocorrelation function.

14. The method according to claim 13, wherein the autocorrelation information includes lag values.

15. The method according to claim 14, wherein the lag value corresponds to the delay value at which the autocorrelation function reaches the local maximum value.

16. The spectral envelope is determined by a first update rate. The aforementioned autocorrelation information is determined by a second update rate. The method according to claim 10, wherein the first update rate is different from the second update rate.

17. The method according to claim 16, wherein the first update rate is higher than the second update rate.

18. The method according to claim 10, wherein the plurality of original subband audio signals are generated by flattening the original audio signals.

19. The method according to claim 18, wherein the plurality of original subband audio signals are generated by decomposing the flattened original audio signal into the plurality of original subband audio signals.