Multi-lag format for audio coding
By generating subband audio signals and encoding spectral envelopes and autocorrelation information, the method enhances coding efficiency and sound quality in high-quality audio coding systems.
Patent Information
- Application Number
- JP2025044212
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-08-20
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-08-18
AI Technical Summary
High-quality audio coding systems require a large amount of data, resulting in low coding efficiency.
The method involves generating subband audio signals through spectral decomposition, determining the spectral envelope and autocorrelation information, and encoding these representations to achieve high coding efficiency.
This approach achieves high coding efficiency with low bit rates while maintaining excellent sound quality by incorporating spectral envelope and autocorrelation information in the encoding process.
Smart Images

Figure 2025083581000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims the priority of U.S. Provisional Patent Application No. 62 / 889,118, filed on August 20, 2019, and European Patent Application No. 19192552.8, filed on August 20, 2019, and incorporates the entire content of each application by reference.
[0002] [Technical Field] The present disclosure generally relates to a method of encoding an audio signal into an encoded representation and a method of decoding an audio signal from the encoded representation.
[0003] Although some embodiments are described herein with particular reference to its disclosure, it is recognized that the present disclosure is not limited to such fields of use and is applicable in a broader context.
Background Art
[0004] No description of the background art through the present disclosure should ever be construed as an admission that the technology is widely known or as part of the common general knowledge in the technical field.
[0005] In high - quality audio coding systems, it is common to describe the detailed waveform characteristics of the signal in the largest part of the information. The small part of the information is used to describe more statistically defined features, such as control data intended to shape quantization noise according to the energy in the frequency band or known simultaneous masking characteristics of hearing (e.g., side information in an MDCT - based waveform coder that transmits the quantization step size and range information necessary to accurately inverse - quantize the data representing the waveform in the decoder). However, these high - quality audio coding systems require a relatively large amount of data to code audio content, i.e., they have a relatively low coding efficiency.
[0006] There is a need for an audio coding method and apparatus capable of coding audio data with improved coding efficiency.
Summary of the Invention
[0007] The present disclosure provides a method for encoding an audio signal, a method for decoding an audio signal, an encoder, a decoder, a computer program, and a computer-readable storage medium.
[0008] According to a first aspect of the present disclosure, a method for encoding an audio signal is provided. The encoding may be performed for each of a plurality of consecutive portions (e.g., samples, segments, groups of frames) of the audio signal. In some implementations, these portions may overlap each other. An encoded representation may be generated for each such portion. The method may include generating a plurality of subband audio signals based on the audio signal. Generating a plurality of subband audio signals based on the audio signal may include spectral decomposition of the audio signal, which may be performed by a filter bank of bandpass filters (BPFs). The frequency resolution of the filter bank may be related to the frequency resolution of the human auditory system. For example, the BPF may be a complex-valued BPF. Alternatively, generating a plurality of subband audio signals based on the audio signal may include spectrally and / or temporally flattening the audio signal, optionally windowing the spectrally flattened audio signal by a window function, and spectrally decomposing the resulting signal into a plurality of subband audio signals. The method may further include determining a spectral envelope of the audio signal. The method may further include, for each subband audio signal, determining autocorrelation information about the subband audio signal based on the autocorrelation function (ACF) of the subband audio signal. The method may further include generating an encoded representation of the audio signal, the encoded representation including a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information about the plurality of subband audio signals. For example, the encoded representation may be related to a portion of a bitstream. In some implementations, the encoded representation may further include waveform information related to the waveform of the audio signal and / or the waveform of one or more of the subband audio signals. The method may further include outputting the encoded representation.
[0009] When configured as described above, the proposed method has a very high coding efficiency (i.e., requires a very low bit rate to code the audio), while at the same time providing an encoded representation of the audio signal that contains appropriate information to achieve a very good sound quality after reconstruction. This is done by providing, in addition to the spectral envelope, autocorrelation information for a plurality of subbands of the audio signal. Notably, it has been proven that two values per subband, namely one lag value and one autocorrelation value, are sufficient to achieve a high sound quality.
[0010] In some embodiments, the autocorrelation information for a given subband audio signal may include the lag value for each subband audio signal and / or the autocorrelation value for each subband audio signal. Preferably, the autocorrelation information may include both the lag value for each subband audio signal and the autocorrelation value for each subband audio signal. Here, the lag value may correspond to the delay value (e.g., the abscissa) at which the autocorrelation function reaches a local maximum, and the autocorrelation value may correspond to this local maximum (e.g., the ordinate).
[0011] In some embodiments, the spectral envelope may be determined at a first update rate, and the autocorrelation information for a plurality of subband audio signals may be determined at a second update rate. In this case, the first update rate and the second update rate may be different from each other. The update rate may also be referred to as the sampling rate. In such an embodiment, the first update rate may be higher than the second update rate. Furthermore, different update rates may be applied to different subbands, i.e., the update rates of the autocorrelation information for different subband audio signals may be different from each other.
[0012] By reducing the update rate of the autocorrelation information compared to the update rate of the spectral envelope, the coding efficiency of the proposed method can be further improved without affecting the sound quality of the reconstructed audio signal.
[0013] In some embodiments, generating a plurality of sub-band audio signals may include applying spectral and / or temporal flattening to the audio signal. Generating a plurality of sub-band audio signals may further include windowing the audio signal flattened by a window function. Also, generating a plurality of sub-band audio signals may further include spectrally decomposing the windowed flattened audio signal into a plurality of sub-band audio signals. In this case, for example, spectrally and / or temporally flattening the audio signal may include generating a perceptually weighted LPC residual of the audio signal.
[0014] In some embodiments, generating a plurality of sub-band audio signals may include spectrally decomposing the audio signal. Then, determining the autocorrelation function for a given sub-band audio signal may include determining the sub-band envelope of the sub-band audio signal. Determining the autocorrelation function may further include envelope flattening the sub-band audio signal based on the sub-band envelope. The sub-band envelope may be determined by taking the magnitude values of the windowed sub-band audio signal. Determining the autocorrelation function may include windowing the sub-band audio signal envelope-flattened by a window function. Also, determining the autocorrelation function may further include determining (e.g., calculating) the autocorrelation function of the windowed sub-band audio signal after envelope flattening. The autocorrelation function may be determined for a real-valued (windowed and envelope-flattened) sub-band signal.
[0015] Other aspects of the present disclosure relate to a method of decoding an audio signal from an encoded representation of the audio signal. The encoded representation may include a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information for each of a plurality of subband audio signals (or those generated from the audio signal). The autocorrelation information for a given subband audio signal may be based on the autocorrelation function of the subband audio signal. The method may include receiving the encoded representation of the audio signal. The method may further include extracting the spectral envelope and the (plural) autocorrelation information from the encoded representation of the audio signal. The method may further include determining a reconstructed audio signal based on the spectral envelope and the autocorrelation information. The reconstructed audio signal may be determined such that the autocorrelation function for each of the plurality of subband audio signals (or those generated from the reconstructed audio signal) of the reconstructed audio signal satisfies the condition derived from the autocorrelation information for the corresponding subband audio signal (or those generated from the audio signal) of the audio signal. For example, for each subband audio signal of the reconstructed audio signal, the reconstructed audio signal may be determined such that the value of the autocorrelation function of the subband audio signal (or those generated from the reconstructed audio signal) substantially matches the autocorrelation value indicated by the autocorrelation information for the corresponding subband audio signal (or those generated from the audio signal) of the audio signal at the lag value (e.g., delay value) indicated by the autocorrelation information for the corresponding subband audio signal (or those generated from the audio signal) of the audio signal. This may mean that the decoder can determine the autocorrelation function of the subband audio signal in the same way as performed by the encoder. This may include any, some, or all of flattening, windowing, and normalization.In some implementations, the reconstructed audio signal may be determined such that the autocorrelation information for each of a plurality of subband signals of the reconstructed subband audio signal (or those generated from the reconstructed subband audio signal) substantially matches the autocorrelation information for the corresponding subband audio signal of the audio signal (or those generated from the audio signal). For example, the reconstructed audio signal may be such that for each subband audio signal of the reconstructed audio signal (or those generated from the reconstructed audio signal), the autocorrelation values and lag values (e.g., delay values) of the autocorrelation function of the subband signal of the reconstructed audio signal substantially match the autocorrelation values and lag values indicated by the autocorrelation information for the corresponding subband audio signal of the audio signal (or those generated from the audio signal). This may mean that the decoder can determine the autocorrelation information (i.e., lag values and autocorrelation values) for each subband signal of the reconstructed audio signal in the same way as performed by the encoder. Here, the term "substantially match" may mean, for example, match up to a predetermined margin. In these implementations where the coded representation includes waveform information, the reconstructed audio signal may be further determined based on the waveform information. The subband audio signals may be obtained, for example, by spectral decomposition of the applicable audio signal (i.e., the original audio signal on the encoder side or the reconstructed audio signal on the decoder side), or may be obtained by flattening and windowing the applicable audio signal and then performing spectral decomposition.
[0016] Thus, the decoder may be said to operate according to a synthesis by analysis approach that attempts to detect a reconstructed audio signal z that satisfies at least one condition derived from the coded representation h(x) of the coded audio signal, or where the coded representation h(z) substantially matches the coded representation h(x) of the original audio signal x, where h is the coding map used by the encoder. In other words, the decoder
[0017]
Number
[0018] In some embodiments, the reconstructed audio signal may be determined by an iterative procedure that starts from an initial candidate for the reconstructed audio signal and generates respective intermediate reconstructed audio signals at each iteration. In each iteration, an update map may be applied to the intermediate reconstructed audio signal to obtain an intermediate reconstructed audio signal for the next iteration. The update map is such that the autocorrelation function of the sub-band audio signal of the intermediate reconstruction of the audio signal (or that generated from the intermediate reconstruction) approaches the condition derived from the autocorrelation information of the corresponding sub-band audio signal of the audio signal (or that generated from the audio signal), and / or the difference between the measured signal power of the sub-band audio signal of the reconstructed audio signal (or that generated from the reconstructed audio signal) and the signal power of the corresponding sub-band audio signal of the audio signal indicated by the spectral envelope (or that generated from the audio signal) is configured to be reduced from one iteration to the next. When both the autocorrelation information and the spectral envelope are considered, an appropriate difference metric for the degree to which the condition is satisfied and the difference between the signal powers for the sub-band audio signals may be defined. In some implementations, the update map may be configured such that the difference between the encoded representation of the intermediate reconstructed audio signal and the encoded representation of the audio signal continuously decreases from one iteration to the next. For this purpose, an appropriate difference metric for the encoded representation (including the spectral envelope and / or autocorrelation information) may be defined and used. The autocorrelation function of the sub-band audio signal of the intermediate reconstructed audio signal (or that generated from the intermediate reconstructed audio signal) may be determined in the same way as the encoder does for the sub-band audio signal of the audio signal. Similarly, the encoded representation of the intermediate reconstructed audio signal may be the encoded representation obtained when the intermediate reconstructed audio signal undergoes the same encoding technique that resulted in the encoded representation of the audio signal.
[0019] Such an iterative method enables a simple yet efficient implementation of the synthesis technique according to the above analysis.
[0020] In some embodiments, determining a reconstructed audio signal based on spectral envelope and autocorrelation information may include applying a machine learning-based generation model that receives, as input, the spectral envelope of the audio signal and the autocorrelation information for each of a plurality of sub-band audio signals of the audio signal, and generates and outputs a reconstructed audio signal. In these implementations where the encoded representation includes waveform information, the machine learning-based generation model may further receive waveform information as input. This means that the machine learning-based generation model may also be conditioned / trained using the waveform information.
[0021] Such a machine learning-based method enables a very efficient implementation of the synthesis technique according to the above analysis and can achieve a reconstructed audio signal that is very perceptually close to the original audio signal.
[0022] Another aspect of the present disclosure relates to an encoder for encoding an audio signal. The encoder may include a processor and a memory coupled to the processor, and the processor is adapted to execute the steps of any one of the encoding methods described throughout the present disclosure.
[0023] Another aspect of the present disclosure relates to a decoder for decoding an audio signal from an encoded representation of the audio signal. The decoder may include a processor and a memory coupled to the processor, and the processor is adapted to execute the steps of any one of the decoding methods described throughout the present disclosure.
[0024] Another aspect relates to a computer program including instructions that, when executed, cause a computer to execute the steps of any of the methods described throughout the present disclosure.
[0025] Other aspects of the present disclosure relate to a computer-readable storage medium storing a computer program according to the above aspects.
Brief Description of the Drawings
[0026] Here, with reference to the accompanying drawings, exemplary embodiments of the present disclosure will be described by way of example only.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0027] [First] High-quality audio coding systems generally require a relatively large amount of data to code audio content, i.e., they have a relatively low coding efficiency. The development of tools such as noise filling and high-frequency reproduction has shown that the waveform description data can be partially replaced by a smaller set of control data, but no high-quality audio codec mainly relies on perceptually relevant features. However, increased computing power and recent advances in the field of machine learning have mainly increased the feasibility of decoding audio from any encoder format. This disclosure proposes examples of such encoder formats.
[0028] Broadly speaking, this disclosure proposes a coding format based on the subband envelope and additional information brought about by auditory resolution. The additional information includes a single autocorrelation value and a single lag value per subband (and per update step). The envelope can be calculated at a first update rate, and the additional information can be sampled at a second update rate. Decoding of the coding format can proceed using an analysis-by-synthesis approach that can be implemented, for example, by techniques based on iteration or machine learning.
[0029] [Encoding] The coding format (coding representation) proposed in this disclosure provides one lag per subband (and update step), so it may be called a multi-lag format. FIG. 1 is a block diagram schematically showing an example of an encoder 100 for generating a coding format according to an embodiment of this disclosure.
[0030] Encoder 100 receives a target sound 10 corresponding to an audio signal to be encoded. The audio signal 10 may include a plurality of consecutive or partially overlapping portions (e.g., groups such as samples, segments, frames, etc.) to be processed by the encoder. The audio signal 10 is spectrally decomposed by a filter bank 15 into a plurality of sub-band audio signals 20 within corresponding frequency sub-bands. The filter bank 15 may be a filter bank of bandpass filters (BPFs), for example, this may be a complex-valued BPF. In audio, a filter bank of BPFs having a frequency resolution related to the human auditory system is naturally used.
[0031] The spectral envelope 30 of the audio signal 10 is extracted by an envelope extraction block 25. For each sub-band, power is measured at a predetermined time step as a basic model of the auditory envelope or excitation pattern for the cochlea resulting from the input sound signal, thereby determining the spectral envelope 30 of the audio signal 10. That is, the spectral envelope 30 may be determined based on the plurality of sub-band audio signals 20, for example, by measuring (e.g., estimating, calculating) the respective signal powers for each of the plurality of sub-band audio signals 20. However, the spectral envelope 30 may be determined by any suitable alternative tool such as, for example, a Linear Predictive Coding (LPC) description. In particular, in some implementations, the spectral envelope may be determined from the audio signal before spectral decomposition by the filter bank 15.
[0032] Optionally, the extracted spectral envelope 30 can be downsampled in a downsampling block 35, and the downsampled spectral envelope 40 (or spectral envelope 30) is output as part of the encoded format or encoded representation of the audio signal 10 (the applicable part).
[0033] A reconstructed signal reconstructed from only the spectral envelope may still lack sound quality. To address this problem, the present disclosure proposes including a single value (i.e., the vertical and horizontal coordinates) of the autocorrelation function of the signal (optionally envelope-flattened) per subband that results in dramatically improved sound quality. For this purpose, the subband audio signal 20 is optionally flattened (envelope-flattened) by a divider 45 and input to an autocorrelation block 55. The autocorrelation block 55 determines the autocorrelation function (ACF) of its input signal and outputs respective autocorrelation information 50 for each of the subband audio signals 20 (i.e., for each of the subbands) based on the ACF of each subband audio signal 20. The autocorrelation information 50 for a given subband includes (e.g., consists of) a representation 50 of the lag value T and the autocorrelation value ρ(T). That is, for each subband, as part of the encoded representation, one value of the lag T and the corresponding (optionally normalized) autocorrelation value (ACF value) ρ(T) are output (e.g., transmitted). Here, the lag value T corresponds to the delay value at which the ACF reaches a local maximum, and the autocorrelation value ρ(T) corresponds to this local maximum. In other words, the autocorrelation information for a given subband may include the delay value (i.e., the horizontal coordinate) and the autocorrelation value (i.e., the vertical coordinate) of the local maximum of the ACF.
[0034] Accordingly, the encoded representation of the audio signal includes the spectral envelope of the audio signal and the autocorrelation information for each of the subbands. The autocorrelation information for a given subband includes a representation of the lag value T and the autocorrelation value ρ(T). The encoded representation corresponds to the output of the encoder. In some implementations, the encoded representation may further include waveform information related to the waveform of the audio signal and / or one or more waveforms of the subband audio signals.
[0035] By the above procedure, an encoding function (or encoding map) h is defined that maps the input audio signal to its encoded representation.
[0036] As described above, the spectral envelope and autocorrelation information for the subband audio signal may be determined and output at different update rates (sample rates). For example, the spectral envelope can be determined at a first update rate, and the autocorrelation information for a plurality of subband audio signals can be determined at a second update rate different from the first update rate. The representations of the spectral envelope and the autocorrelation information (for all subbands) may be written to the bitstream at their respective update rates (sample rates). In this case, the encoded representation may be related to a part of the bitstream output by the encoder. In this regard, it should be noted that at each point in time, the current spectral envelope and the current set of autocorrelation information (one for each subband) are defined by the bitstream and can be used as the encoded representation. Alternatively, the representations of the spectral envelope and the autocorrelation information (for all subbands) may be updated at their respective update rates in each output unit of the encoder. In this case, each output unit of the encoder (e.g., an encoded frame) corresponds to an instance of the encoded representation. The representations of the spectral envelope and the autocorrelation information may be the same between a series of consecutive output units depending on their respective update rates.
[0037] Preferably, the first update rate is higher than the second update rate. In one example, the first update rate R 1 is R 1 = 1 / (2.5 ms) may be used, and the second update rate R 2 is R 2It may also be 1 / (20 ms), and as a result, the updated representation of the spectral envelope is output every 2.5 ms, while the updated representation of the autocorrelation information is output every 20 ms. For a portion (e.g., a frame) of the audio signal, the spectral envelope may be determined for every nth portion (e.g., every single portion), while the autocorrelation information may be determined for every mth portion (m > n).
[0038] The coded representation may be output as a sequence of frames of a specific frame length. Among other factors, the frame length may depend on the first and / or second update rate. L 1 = 1 / R 1 The first update rate R via 1 (e.g., 1 / (2.5 ms)) corresponds to the first period L 1 Considering a frame having a length of (e.g., 2.5 ms), this frame includes one representation of the spectral envelope and one set of representations of the autocorrelation information (one per subband audio signal). Regarding the first and second update rates of 1 / (2.5 ms) and 1 / (20 ms) respectively, the autocorrelation information is the same for each of eight consecutive frames of the coded representation. Generally, R 1 and R 2 Assuming that they are appropriately selected to have an integer ratio, the autocorrelation information is the same for R 1 / R 2 consecutive frames of the coded representation. On the other hand, L 2 = 1 / R 2 The second update rate R via 2 (e.g., 1 / (20 ms)) corresponds to the second period L 2 Considering a frame having a length of (e.g., 20 ms), this frame includes one set of representations of the autocorrelation information and R 1 / R 2 (e.g., eight) representations of the spectral envelope.
[0039] In some implementations, different update rates may also be applied to different sub-bands, i.e., the autocorrelation information for different sub-band audio signals may be generated and output at different update rates.
[0040] FIG. 2 is a flowchart showing an example of an encoding method 200 according to an embodiment of the present disclosure. The method may be implemented by the above-described encoder 100 and receives an audio signal as an input.
[0041] In step S210, a plurality of sub-band audio signals are generated based on the audio signal. This may include spectrally decomposing the audio signal, in which case this step may be performed according to the operation of the above-described filter bank 15. Alternatively, this may include spectrally and / or temporally flattening the audio signal, optionally windowing the audio signal flattened by a window function, and spectrally decomposing the resulting signal into a plurality of sub-band audio signals.
[0042] In step S220, the spectral envelope of the audio signal is determined (e.g., calculated). This step may be performed according to the operation of the above-described envelope extraction block 25.
[0043] In step S230, for each sub-band audio signal, autocorrelation information for the sub-band audio signal is determined based on the ACF of the sub-band audio signal. This step may be performed according to the operation of the above-described autocorrelation block 55.
[0044] In step S240, an encoded representation of the audio signal is generated. The encoded representation includes a representation of the spectral envelope of the audio signal and a representation of the autocorrelation information for each of the plurality of sub-band audio signals.
[0045] Next, an example of the details of the implementation of the steps of method 200 will be described.
[0046] For example, as described above, generating a plurality of sub-band audio signals may (or may mean) include spectrally decomposing an audio signal, for example, using a filter bank. In this case, determining an autocorrelation function for a given sub-band audio signal may include determining a sub-band envelope of the sub-band audio signal. The sub-band envelope may be determined by taking the magnitude value of the sub-band audio signal. The ACF itself may be calculated for a real-valued (windowed after envelope flattening) sub-band signal.
[0047] Assuming that the sub-band filter response becomes complex-valued by a Fourier transform supported at essentially positive frequencies, the sub-band signal becomes complex-valued. Then, the sub-band envelope can be determined by taking the magnitude of the complex-valued sub-band signal. This sub-band envelope has the same number of samples as the sub-band signal and is still somewhat oscillatory. Optionally, the sub-band envelope can be downsampled by, for example, calculating the sum of the squared triangular window weights of the envelope in segments of a specific length (e.g., length 5 ms, rise 2.5 ms, fall 2.5 ms) for every half shift of a specific length (e.g., 2.5 ms) along the signal, and then taking the square root of this sequence to obtain a downsampled sub-band envelope. This may be said to correspond to the definition of "rms envelope". The triangular window can be normalized so that a constant envelope of value 1 gives a sequence of 1. Other methods for determining the sub-band envelope are similarly feasible, such as half-wave rectification and subsequent low-pass filtering in the case of a real-valued sub-band signal. In any case, it can be said that the sub-band envelope conveys information about the energy within the sub-band signal (at a selected update rate).
[0048] Next, the sub-band audio signal may be envelope-flattened based on the sub-band envelope. For example, in order to reach the signal (carrier) of the fine structure for which the ACF data is calculated, the downsampled values are linearly interpolated, and the original (complex-valued) sub-band signal is divided by this linearly interpolated envelope, whereby a new envelope signal at the original sample rate may be created.
[0049] Next, the envelope-flattened sub-band audio signal may be windowed by an appropriate window function. Finally, the ACF of the windowed envelope-flattened sub-band audio signal is determined (e.g., calculated). In some implementations, determining the ACF for a given sub-band audio signal may further include normalizing the ACF of the windowed envelope-flattened sub-band audio signal by the autocorrelation function of the window function.
[0050] In FIG. 3, the upper curve 310 shows the real values of the windowed envelope-flattened sub-band signal used to calculate the ACF. The lower solid curve 320 shows the real values of the complex ACF.
[0051] The main concept here is to find the largest local maximum of the ACF of the subband signal from among the local maxima that are above the absolute value of the ACF of the impulse response of a subband filter (i.e., the corresponding BPF of a filter bank), which is a complex-valued one. For the ACF of a complex-valued subband signal, the real-valued ACF may be considered at this point. To avoid the picking lag associated with the center frequency of the subband, rather than the characteristics of the input signal, it may be necessary to find the largest local maximum above the ACF of the absolute value of the impulse response. As a final adjustment, the maximum value may be divided by the value of the ACF of the window function used for the subband ACF window (assuming, for example, that the ACF of the subband signal itself is normalized such that the autocorrelation value at zero delay is normalized to 1). This results in better use of the interval between 0 and 1, where ρ(T) = 1 is the maximum tonality.
[0052] Therefore, determining the autocorrelation information for a given subband audio signal based on the ACF of the subband audio signal may further include comparing the ACF of the subband audio signal with the ACF of the absolute value of the impulse response of each bandpass filter associated with the subband audio signal. The ACF of the absolute value of the impulse response of each bandpass filter associated with the subband audio signal is shown by the solid curve 330 in the lower part of FIG. 3. Then, the autocorrelation information is determined based on the highest local maximum of the ACF of the subband signal above the ACF of the absolute value of the impulse response of each bandpass filter associated with the subband audio signal. In the lower part of FIG. 3, the local maxima of the ACF are indicated by crosses, and the selected highest local maximum of the ACF of the subband signal above the ACF of the absolute value of the impulse response of each bandpass is indicated by a circle. Optionally, (assuming that the ACF itself is normalized such that the autocorrelation value at zero delay is normalized to 1) the selected local maximum of the ACF may be normalized by the ACF value of the ACF of the window function. The selected highest local maximum of the ACF after normalization is shown by the asterisk in the lower part of FIG. 3, and the dashed curve 340 shows the ACF of the window function.
[0053] The autocorrelation information determined at this stage may include the autocorrelation value and the delay value (i.e., the ordinate and abscissa) of the selected (normalized) highest local maximum of the ACF of the subband audio signal.
[0054] A similar coding format may be defined in the framework of an LPC-based vocoder. Also, in this case, the autocorrelation information is extracted from subband signals that are affected by at least some degree of spectral and / or temporal flattening. Different from the above example, this is done by creating a (perceptually weighted) LPC residual, windowing it, and decomposing it into subbands to obtain a plurality of subband audio signals. Subsequently, the ACF is calculated and the lag values and autocorrelation values of each subband audio signal are extracted.
[0055] For example, generating a plurality of sub-band audio signals may include applying spectral and / or temporal flattening to an audio signal (e.g., by generating perceptually weighted LPC residuals from the audio signal using, for example, an LPC filter). Subsequently, the audio signal flattened by the window function may be windowed, and the windowed flattened audio signal may be spectrally decomposed into a plurality of sub-band audio signals. As described above, the result of the temporal and / or spectral flattening may correspond to the perceptually weighted LPC residuals, which then undergo windowing and spectral decomposition into sub-bands. The perceptually weighted LPC residuals may be, for example, pink LPC residuals.
[0056] [Decoding] The present disclosure relates to audio decoding based on analysis-by-synthesis techniques. At the most abstract level, it is assumed that an encoding map h from the signal to a perceptually motivated domain is given such that the original audio signal x is represented by y = h(x). In the best case, a simple distortion measure such as least squares in the perceptual domain is a good predictor of the subjective differences measured by a population of listeners.
[0057] One remaining problem is to design a decoder q that maps from y to an audio signal z = d(y) (the encoded and decoded versions). For this, the concept of analysis-by-synthesis can be used, which includes "finding the waveform that comes closest to generating a given image". The goal is for z and x to sound the same, and as a result, the decoder should solve the inverse problem h(z) = y = h(x). Regarding the composition of the map, d should approximate the left inverse of h,
[0058] [Number] This means. This inverse problem is often a singular problem in the sense that it has many solutions. The opportunity to achieve a significant saving in bit rate lies in the observation that a large number of different waveforms produce the same sound impression.
[0059] Figure 4 is a block diagram schematically showing an example of synthesis by analysis for determining a decoding function (or decoding map) d when an encoding function (or encoding map) h is given. The original audio signal x, 410 receives the encoding map h, 415 and produces an encoded representation y, 420, where y = h(x). The encoded representation y may be defined in the perceptual domain. The objective is to find a decoding function (decoding mapping) d, 425 that maps the encoded representation y to the reconstructed audio signal z, 430, which has the property that applying the encoding mapping h, 435 to the reconstructed audio signal z produces an encoded representation h(z), 440 that substantially coincides with the encoded representation y = h(x). Here, "substantially coincides" may mean, for example, "coincides up to a predetermined margin". In other words, when the encoding map h is given, the objective is to
[0060] [Number] find a decoding map d like this.
[0061] Figure 5 is a flowchart showing an example of a decoding method 500 along the synthesis-by-analysis technique according to an embodiment of the present disclosure. The method 500 is a method for decoding an audio signal from an encoded representation of the (original) audio signal. The encoded representation is assumed to include a representation of the spectral envelope of the original audio signal and a representation of the autocorrelation information for each of a plurality of subband audio signals of the original audio signal. The autocorrelation information for a given subband audio signal is based on the ACF of the subband audio signal.
[0062] In step S510, an encoded representation of the audio signal is received.
[0063] In step S520, the spectral envelope and autocorrelation information are extracted from the encoded representation of the audio signal.
[0064] In step S530, the reconstructed audio signal is determined based on the spectral envelope and autocorrelation information. Here, the reconstructed audio signal is determined such that the autocorrelation function of each of the plurality of subband signals of the reconstructed subband audio signal (substantially) satisfies the condition derived from the autocorrelation information for the corresponding subband audio signal of the audio signal. This condition may be, for example, that for each subband audio signal of the reconstructed audio signal, the value of the ACF of the subband audio signal of the reconstructed audio signal substantially matches the autocorrelation value indicated by the autocorrelation information for the corresponding subband audio signal of the audio signal at the lag value (e.g., delay value) indicated by the autocorrelation information for the corresponding subband audio signal of the audio signal. This may mean that the decoder can determine the ACF of the subband audio signal in the same way as performed by the encoder. This may include any, some, or all of flattening, windowing, and normalization. In one implementation, the reconstructed audio signal is determined such that for each subband audio signal of the reconstructed audio signal, the autocorrelation value and lag value (e.g., delay value) of the ACF of the subband signal of the reconstructed audio signal substantially match the autocorrelation value and lag value indicated by the autocorrelation information for the corresponding subband audio signal of the original audio signal. This may mean that the decoder can determine the autocorrelation information for each subband signal of the reconstructed audio signal in the same way as performed by the encoder. In these implementations where the encoded representation also includes waveform information, the reconstructed audio signal may be further determined based on the waveform information. The subband audio signals of the reconstructed audio signal may be generated in the same way as performed by the encoder. For example, this may include spectral decomposition, or a sequence of flattening, windowing, and spectral decomposition.
[0065] Preferably, the determination of the reconstructed audio signal in step S530 also takes into account the spectral envelope of the original audio signal. The reconstructed audio signal is then further determined such that, for each sub-band audio signal of the reconstructed sub-band audio signals, the measured (e.g., estimated or calculated) signal power of the sub-band audio signal of the reconstructed audio signal substantially matches the signal power for the corresponding sub-band audio signal of the original audio signal indicated by the spectral envelope.
[0066] As can be seen from the above, the proposed method 500 can be said to be brought about by an analysis-by-synthesis approach in that it attempts to find a reconstructed audio signal z that (substantially) satisfies at least one condition derived from the encoded representation y = h(x) of the original audio signal x. Here, h is the encoding map used by the encoder. In some implementations, it can even be said that the proposed method operates according to an analysis-by-synthesis approach in that it attempts to find a reconstructed audio signal z such that the encoded representation h(z) substantially matches the encoded representation y = h(x) of the original audio signal x. In other words, the decoding method is said to find a decoding map d such as
[0067]
Number
[0068] [Implementation Example 1: For Each Parametric Synthesis or Signal Iteration] The inverse problem h(z) = y is to modify z n ) such that h(z n-1 ) is closer to y than h(z n-1 ) using an update map z n = f(z n-1It can be solved by an iterative method assuming (x, y). For example, the starting point of the iteration (i.e., the initial candidate for the reconstructed audio signal) may be a random noise signal (e.g., white noise), or it may be determined based on the coded representation of the audio signal (e.g., as an initially manually created guess). In the latter case, the initial candidate for the reconstructed audio signal may be related to a learned guess based on the spectral envelope and / or autocorrelation information for a plurality of subband audio signals. In these implementations where the coded representation includes waveform information, the learned guess may be further based on the waveform information.
[0069] More specifically, the reconstructed audio signal in this implementation example is determined by an iterative procedure that starts from an initial candidate for the reconstructed audio signal and generates respective intermediate reconstructed audio signals at each iteration. In each iteration, an update map is applied to the intermediate reconstructed audio signal in order to obtain the intermediate reconstructed audio signal for the next iteration. The update map is selected such that the difference between the coded representation of the intermediate reconstructed audio signal and the coded representation of the original audio signal continuously decreases from one iteration to the next. For this purpose, an appropriate difference metric for the coded representation (e.g., spectral envelope, autocorrelation information) may be defined and used to evaluate the difference. The coded representation of the intermediate reconstructed audio signal may be the coded representation obtained when the intermediate reconstructed audio signal undergoes the same coding scheme that yields the coded representation of the audio signal.
[0070] When searching for a reconstructed audio signal whose procedure satisfies at least one condition derived from the (plural) autocorrelation information, the update map is such that the autocorrelation function of the subband audio signal of the intermediate reconstruction of the audio signal satisfies each condition derived from the autocorrelation information for the corresponding subband audio signal of the audio signal, and / or the difference between the measured signal power of the subband audio signal of the reconstructed audio signal and the signal power for the corresponding subband audio signal of the audio signal indicated by the spectral envelope is selected to be reduced from one iteration to the next. When both the autocorrelation information and the spectral envelope are considered, an appropriate difference metric for the degree to which the condition is satisfied and the difference between the signal powers for the subband audio signals may be defined.
[0071] [Implementation Example 2: Generation Model Based on Machine Learning] Another option enabled by modern machine learning methods is to train a machine learning-based generative model (or, simply, generative model) for audio x conditioned on data y. That is, when a large set of examples of (x, y) (where y = h(x)) is given, the parametric conditional distribution p(x|y) from y to x is trained. Then, the decoding algorithm may consist of sampling from the distribution z~p(x|y).
[0072] This option has been found to be particularly advantageous when h(x) is a voice vocoder and p(x|y) is defined by a sequential generation model, a sample recurrent neural network (RNN). However, other generation models such as variational autoencoders or generative adversarial models are also relevant to this task. Therefore, without intending to limit, the machine learning-based generation model can be one of a recurrent neural network, a variational autoencoder, or a generative adversarial model (e.g., a Generative Adversarial Network (GAN)).
[0073] In this implementation example, determining the reconstructed audio signal based on the spectral envelope and autocorrelation information involves receiving, as inputs, the spectral envelope of the audio signal and the autocorrelation information for each of a plurality of subband audio signals of the audio signal, and applying a machine learning-based generation model that generates and outputs the reconstructed audio signal. The encoded representation also includes waveform information. In these implementations, the machine learning-based generation model may further receive waveform information as an input.
[0074] As described above, the machine learning-based generation model may include a parametric conditional distribution p(x|y) that associates the encoded representation y of the audio signal and the corresponding audio signal x with respective probabilities p. Then, determining the reconstructed audio signal may include sampling from the parametric conditional distribution p(x|y) for the encoded representation of the audio signal.
[0075] During the training phase, prior to decoding, a machine learning-based generative model may be conditioned / trained on a dataset of a plurality of audio signals and corresponding encoded representations of the audio signals. If the encoded representation also includes waveform information, the machine learning-based generative model may also be conditioned / trained using the waveform information.
[0076] FIG. 6 is a flowchart showing an exemplary implementation 600 of step S530 in the decoding method 500 of FIG. 5. In particular, implementation 600 relates to the implementation for each sub-band of step S530.
[0077] In step 610, a plurality of reconstructed sub-band audio signals are determined based on the spectral envelope and autocorrelation information. Here, for each reconstructed sub-band audio signal, the plurality of reconstructed sub-band audio signals are determined such that the autocorrelation function of the reconstructed sub-band audio signal satisfies the condition derived from the autocorrelation information about the corresponding sub-band audio signal of the audio signal. In some implementations, for each reconstructed sub-band audio signal, the plurality of reconstructed sub-band audio signals are determined such that the autocorrelation information about the reconstructed sub-band audio signal substantially matches the autocorrelation information about the corresponding sub-band audio signal.
[0078] Preferably, the determination of the plurality of reconstructed sub-band audio signals in step S610 also takes into account the spectral envelope of the original audio signal. Then, for each reconstructed sub-band audio signal, the plurality of reconstructed sub-band audio signals are further determined such that the measured (e.g., estimated, calculated) signal power of the reconstructed sub-band audio signal substantially matches the signal power of the corresponding sub-band audio signal indicated by the spectral envelope.
[0079] In step S620, a reconstructed audio signal is determined based on the plurality of reconstructed sub-band audio signals by spectral synthesis.
[0080] The above Implementation Examples 1 and 2 may also be applied to the implementation for each sub-band in step S530. In Implementation Example 1, each reconstructed sub-band audio signal may be determined by an iterative procedure that starts from an initial candidate for the reconstructed sub-band audio signal and generates an intermediate reconstructed sub-band audio signal in each iteration. In each iteration, an update map may be applied to the intermediate reconstructed sub-band audio signal to obtain the intermediate reconstructed sub-band audio signal for the next iteration such that the difference between the autocorrelation information about the intermediate reconstructed sub-band audio signal and the autocorrelation information about the corresponding sub-band audio signal continuously decreases from one iteration to the next, or such that the reconstructed sub-band audio signal better satisfies each condition derived from the autocorrelation information about the corresponding sub-band audio signal of the audio signal.
[0081] Also, the spectral envelope may be considered at this point. That is, the update map may be such that the (identical) differences between the respective signal powers of the sub-band audio signals and the (identical) differences between the respective terms of the autocorrelation information continuously decrease. This may mean the definition of an appropriate difference metric for evaluating the (identical) differences. Otherwise, a description similar to that given above for Implementation Example 1 may also apply in this case.
[0082] When applying Implementation Example 2 to the implementation for each sub-band in step S530, determining a plurality of reconstructed sub-band audio signals based on the spectral envelope and the autocorrelation information may include applying a machine learning-based generation model that receives, as input, the spectral envelope of the audio signal and the autocorrelation information about each of the plurality of sub-band audio signals of the audio signal, and generates and outputs a plurality of reconstructed sub-band audio signals. Otherwise, a description similar to that given above for Implementation Example 2 may also apply in this case.
[0083] The present disclosure further relates to an encoder for encoding an audio signal that is capable of executing and adapted to execute an encoding method described throughout the present disclosure. An example of such an encoder 700 is schematically shown in FIG. 7 in the form of a block diagram. Encoder 700 includes a processor 710 and a memory 720 coupled to processor 710. Processor 710 is adapted to execute the steps of any one of the encoding methods described throughout the present disclosure. For this purpose, memory 720 may include respective instructions for processor 710 to execute. Encoder 700 may further include an interface 730 for receiving an input audio signal 740 to be encoded and / or for outputting an encoded representation 750 of the audio signal.
[0084] The present disclosure further relates to a decoder for decoding an audio signal from an encoded representation of the audio signal that is capable of executing and adapted to execute a decoding method described throughout the present disclosure. An example of such a decoder 800 is schematically shown in FIG. 8 in the form of a block diagram. Decoder 800 includes a processor 810 and a memory 820 coupled to processor 810. Processor 810 is adapted to execute the steps of any one of the decoding methods described throughout the present disclosure. For this purpose, memory 820 may include respective instructions for processor 810 to execute. Encoder 800 may further include an interface 830 for receiving an input encoded representation 840 of the audio signal to be decoded and / or for outputting a decoded (i.e., reconstructed) audio signal 850.
[0085] The present disclosure further relates to a computer program including instructions that, when executed, cause a computer to execute an encoding or decoding method described throughout the present disclosure.
[0086] Finally, the present disclosure also relates to a computer-readable storage medium storing the above computer program.
[0087] [Interpretation] Unless otherwise specified, as will be apparent from the following discussion, throughout this disclosure, discussions using terms such as "processing," "calculating," "computing," "determining," "analyzing," etc. refer to operations and / or processes of a computer, computer system, or similar electronic computing device that manipulate and / or transform data represented as a physical quantity, such as an electronic quantity, into other data represented as a physical quantity as well.
[0088] Similarly, the term "processor" may refer to any device or portion of a device that processes electronic data from, for example, registers and / or memory and transforms it into other electronic data that may be stored in, for example, registers and / or memory. A "computer" or "computer system" or "computing platform" may include one or more processors.
[0089] The method described herein is executable by one or more processors that accept computer-readable (also referred to as machine-readable) code that includes a set of instructions that, when executed by one or more processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken is included. Thus, one example is a typical processing system that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem that includes main RAM and / or static RAM and / or ROM. A bus subsystem for communicating between components may be included. Further, the processing system may be a distributed processing system having processors coupled by a network. If the processing system requires a display, a display such as a liquid crystal display (LCD) or a cathode ray tube (CRT) display may be included. If manual data input is required, the processing system may also include an input device such as one or more of an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, etc. The processing system may also include a storage system such as a disk drive unit. In some configurations, the processing system may include a sound output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier carrying computer-readable code (e.g., software) that includes a set of instructions that, when executed by one or more processors, cause one or more of the methods described herein to be performed. Note that if a method includes several elements, e.g., several steps, the order of such elements is not implied unless specifically stated otherwise.The software may reside on a hard disk or, alternatively, may reside completely or at least partially in RAM and / or within the processor during its execution by a computer system. Thus, the memory and the processor also constitute a computer-readable medium carrying computer-readable code. Further, the computer-readable medium may form a computer program product or may be included in a computer program product.
[0090] In an alternative exemplary embodiment, one or more processors may operate as a stand-alone device or may be connected, for example, to other processors, such as, for example, network-connected. In a network-connected deployment, one or more processors may operate in the capacity of a server or user machine within a server-user network environment or may operate as a peer machine within a peer-to-peer or distributed network environment. One or more processors may form any machine that can execute a set of instructions (sequential or otherwise) that specify actions to be taken by that machine, such as a personal computer (PC), tablet PC, personal digital assistant (PDA), cellular phone, web appliance, network router, switch or bridge, or any other machine.
[0091] Note that the term "machine" is also to be construed to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0092] Accordingly, each one exemplary embodiment of the methods described herein is in the form of a computer-readable carrier carrying a computer program for execution on a set of instructions, e.g., on one or more processors, e.g., one or more processors that are part of a web server configuration. Thus, as will be recognized by those skilled in the art, the exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a special purpose device, an apparatus such as a data processing system, or a computer-readable carrier, e.g., a computer program product. The computer-readable carrier carries computer-readable code including a set of instructions that, when executed on one or more processors, cause the processors to implement the method. Thus, aspects of the present disclosure may take the form of a method, an exemplary embodiment that is entirely hardware, an exemplary embodiment that is entirely software, or an exemplary embodiment that combines aspects of software and hardware. Further, the present disclosure may take the form of a carrier carrying computer-readable program code embodied in a medium (e.g., a computer program product on a computer-readable storage medium).
[0093] The software may be further transmitted or received on a network via a network interface device. In an exemplary embodiment, the carrier is a single medium, but the term "carrier" should be construed to include a single medium or a plurality of media (e.g., a centralized or distributed database, and / or associated cache and server) that store one or more sets of instructions. The term "carrier" should also be construed to include any medium that can store, encode, or carry a set of instructions for execution by one or more processors and cause any one or more of the methodologies of the present disclosure to be executed by one or more processors. The carrier may take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media includes dynamic memory, such as main memory. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up a bus subsystem. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier" includes, but is not limited to, solid memory, computer products embodied in optical and magnetic media, media that carry a propagated signal detectable by at least one processor or one or more processors and represent a set of instructions that, when executed, implement a method, and transmission media within a network that carry a propagated signal detectable by at least one of one or more processors and represent a set of instructions.
[0094] It is understood that the steps of the methods described are, in one exemplary embodiment, executed by a suitable processor (or processors) of a processing (e.g., computer) system that executes instructions (computer-readable code) stored in a storage. It is also understood that the present disclosure is not limited to any particular implementation or programming technique, and the present disclosure may be implemented using any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.
[0095] Throughout this disclosure, references to "one exemplary embodiment", "some exemplary embodiments", or "an exemplary embodiment" mean that the particular features, structures, or characteristics described in connection with the exemplary embodiment are included in at least one exemplary embodiment of the disclosure. Thus, the appearances of the phrases "in one exemplary embodiment", "in some exemplary embodiments", or "in an exemplary embodiment" in various places throughout this disclosure are not necessarily all referring to the same exemplary embodiment. Further, the particular features, structures, or characteristics may be combined in any suitable manner in one or more exemplary embodiments, as will be apparent to those skilled in the art from this disclosure.
[0096] The use of ordinal adjectives such as "first", "second", "third", etc. in this specification to describe common objects is merely to indicate that different instances of like objects are being referred to, and is not intended to mean that the objects so described must be in a given order, whether temporally, spatially, in ranking, or otherwise.
[0097] In the following claims and the description herein, any one of the terms "comprising", "comprised of", or "which comprises" is an open term meaning including at least the element / feature but not excluding other elements. Thus, when used in a claim, the term "comprises" should not be construed as being limited to the recited means or elements or steps. For example, an apparatus comprising A and B should not be limited to a device consisting only of elements A and B. As used herein, any one of the terms "including" or "which includes or that includes" is also an open term meaning including at least the element / feature of the term but not excluding other terms. Thus, "including" is synonymous with "comprising" and means the same thing.
[0098] In the above description of the exemplary embodiments of the disclosure, for the purpose of streamlining the disclosure to assist in the understanding of one or more aspects of the various inventions, it should be recognized that the various features of the disclosure may, in some cases, be grouped together in a single exemplary embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claims require more features than are explicitly recited in each claim. Rather, as reflected by the following claims, aspects of the invention are in some cases smaller than all of the features of a single one of the exemplary embodiments disclosed above. Thus, the claims following the description are expressly incorporated herein such that each claim stands on its own as a separate exemplary embodiment of the present disclosure.
[0099] Furthermore, some of the exemplary embodiments described herein include some features but not other features included in other exemplary embodiments. As will be understood by those skilled in the art, combinations of features of different exemplary embodiments are within the scope of the present disclosure and are meant to form different exemplary embodiments. For example, in the following claims, any of the exemplary embodiments recited in the claims can be used in any combination.
[0100] In the description provided herein, numerous specific details are set forth. However, it is understood that the exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0101] Accordingly, while what is considered to be the best mode of the present disclosure is described, those skilled in the art will recognize that other further modifications may be made without departing from the spirit of the disclosure, and intend to claim all such changes and modifications that are included within the scope of the present disclosure. For example, the above formulas merely represent procedures that can be used. Functions may be added or removed from the block diagrams, and operations may be exchanged between the functional blocks. Steps may be added or removed to the methods described within the scope of the present disclosure.
[0102] Various aspects and embodiments of the present disclosure can be recognized from the exemplary embodiments (EEE) listed below.
[0103] EEE1. A method for encoding an audio signal, comprising: generating a plurality of sub-band audio signals based on the audio signal; determining a spectral envelope of the audio signal; for each sub-band audio signal, determining autocorrelation information about the sub-band audio signal based on the autocorrelation function of the sub-band audio signal; A step of generating an encoded representation of an audio signal, the encoded representation including a representation of the spectral envelope of the audio signal and a representation of autocorrelation information for a plurality of sub-band audio signals A method comprising
[0104] EEE2. The method according to EEE1, wherein the spectral envelope is determined based on a plurality of sub-band audio signals
[0105] EEE3. The method according to EEE1 or 2, wherein the autocorrelation information for a given sub-band audio signal includes a lag value for each sub-band audio signal and / or an autocorrelation value for each sub-band audio signal
[0106] EEE4. The method according to EEE3, wherein the lag value corresponds to a delay value at which the autocorrelation function reaches a local maximum, and the autocorrelation value corresponds to this local maximum
[0107] EEE5. The method according to any one of EEE1 to 4, wherein the spectral envelope is determined at a first update rate, the autocorrelation information for a plurality of sub-band audio signals is determined at a second update rate, and the first update rate and the second update rate are different from each other
[0108] EEE6. The method according to EEE5, wherein the first update rate is higher than the second update rate
[0109] EEE7. Generating a plurality of sub-band audio signals comprises applying spectral and / or temporal flattening to the audio signal windowing the flattened audio signal spectrally decomposing the windowed flattened audio signal into a plurality of sub-band audio signals, the method according to any one of EEE1 to 6
[0110] Generating a plurality of sub-band audio signals includes spectral decomposition of an audio signal, Determining an autocorrelation function for a given sub-band audio signal, determining a sub-band envelope of the sub-band audio signal, envelope flattening the sub-band audio signal based on the sub-band envelope, windowing the envelope-flattened sub-band audio signal by a window function, Determining an autocorrelation function of the envelope-flattened sub-band audio signal after windowing, including the method according to any one of EEE1 to 6.
[0111] EEE9. Determining an autocorrelation function for a given sub-band audio signal, The method according to EEE7 or 8, further comprising normalizing an autocorrelation function of the envelope-flattened sub-band audio signal after windowing by an autocorrelation function of the window function.
[0112] EEE10. Determining an autocorrelation function for a sub-band audio signal based on an autocorrelation function of a given sub-band audio signal, comparing the autocorrelation function of the sub-band audio signal with an autocorrelation function of an absolute value of an impulse response of each band-pass filter associated with the sub-band audio signal, Determining autocorrelation information based on a highest local maximum of the autocorrelation function of the sub-band signal above the autocorrelation function of the absolute value of the impulse response of each band-pass filter associated with the sub-band audio signal, including the method according to any one of EEE1 to 9.
[0113] EEE11. Determining a spectral envelope includes measuring signal power for each of a plurality of sub-band audio signals, including the method according to any one of EEE1 to 10.
[0114] EEE12. A method for decoding an audio signal from an encoded representation of the audio signal, the encoded representation including a representation of the spectral envelope of the audio signal and a representation of autocorrelation information for each of a plurality of sub-band audio signals generated from the audio signal, the autocorrelation information for a given sub-band audio signal being based on the autocorrelation function of the sub-band audio signal, the method comprising: receiving the encoded representation of the audio signal; extracting the spectral envelope and the autocorrelation information from the encoded representation of the audio signal; determining a reconstructed audio signal based on the spectral envelope and the autocorrelation information; and the reconstructed audio signal is determined such that the autocorrelation function for each of a plurality of sub-band audio signals generated from the reconstructed audio signal satisfies the condition derived from the autocorrelation information for the corresponding sub-band audio signal generated from the audio signal.
[0115] EEE13. The method according to EEE12, wherein the reconstructed audio signal is further determined such that, for each sub-band audio signal of the reconstructed audio signal, the measured signal power of the sub-band audio signal of the reconstructed audio signal substantially matches the signal power for the corresponding sub-band audio signal of the audio signal indicated by the spectral envelope.
[0116] EEE14. The reconstructed audio signal is determined by an iterative procedure that starts from an initial candidate for the reconstructed audio signal and generates an intermediate reconstructed audio signal at each iteration, wherein in each iteration, an update map is applied to the intermediate reconstructed audio signal to obtain the intermediate reconstructed audio signal for the next iteration such that the difference between the encoded representation of the intermediate reconstructed audio signal and the encoded representation of the audio signal continuously decreases from one iteration to the next, according to the method of EEE12 or 13.
[0117] EEE15. The initial candidate for reconstructing the audio signal is the method described in EEE14, which is determined based on the encoded representation of the audio signal.
[0118] EEE16. The initial candidate for reconstructing the audio signal is white noise, which is the method described in EEE14.
[0119] EEE17. Determining a reconstructed audio signal based on spectral envelope and autocorrelation information includes receiving, as input, the spectral envelope of the audio signal and the autocorrelation information for each of a plurality of subband audio signals of the audio signal, and applying a machine learning-based generation model that generates and outputs a reconstructed audio signal, which is the method described in EEE12 or 13.
[0120] EEE18. The machine learning-based generation model includes a parametric conditional distribution that associates the encoded representation of the audio signal and the corresponding audio signal with respective probabilities, Determining a reconstructed audio signal includes sampling from the parametric conditional distribution for the encoded representation of the audio signal, which is the method described in EEE17.
[0121] EEE19. The method described in EEE17 or 18 further includes, in a training phase, training a machine learning-based generation model on a dataset of a plurality of audio signals and corresponding encoded representations of the audio signals.
[0122] EEE20. The machine learning-based generation model is one of a recurrent neural network, a variational autoencoder, or an adversarial generation model, which is the method described in any one of EEE17 to 19.
[0123] EEE21. Determining a reconstructed audio signal based on spectral envelope and autocorrelation information is Determine a plurality of reconstructed sub-band audio signals based on spectral envelope and autocorrelation information, including determining a reconstructed audio signal based on the plurality of reconstructed sub-band audio signals by spectral synthesis, wherein the plurality of reconstructed sub-band audio signals are determined such that, for each reconstructed sub-band audio signal, the autocorrelation function of the reconstructed sub-band audio signal satisfies a condition derived from the autocorrelation information for the corresponding sub-band audio signal, the method according to EEE12.
[0124] EEE22. The plurality of reconstructed sub-band audio signals are further determined such that, for each reconstructed sub-band audio signal, the measured signal power of the reconstructed sub-band audio signal substantially matches the signal power for the corresponding sub-band audio signal indicated by the spectral envelope, the method according to EEE21.
[0125] EEE23. Each reconstructed sub-band audio signal is determined by an iterative procedure that starts from an initial candidate for the reconstructed sub-band audio signal and generates an intermediate reconstructed sub-band audio signal in each iteration, wherein in each iteration, an update map is applied to the intermediate reconstructed sub-band audio signal to obtain the intermediate reconstructed sub-band audio signal for the next iteration such that the difference between the autocorrelation information for the intermediate reconstructed sub-band audio signal and the autocorrelation information for the corresponding sub-band audio signal continuously decreases from one iteration to the next, the method according to EEE21 or 22.
[0126] EEE24. Determining a plurality of reconstructed sub-band audio signals based on spectral envelope and autocorrelation information includes applying a machine learning-based generation model that receives, as input, the spectral envelope of the audio signal and the autocorrelation information for each of a plurality of sub-band audio signals of the audio signal, and generates and outputs the plurality of reconstructed sub-band audio signals, the method according to EEE21 or 22.
[0127] EEE25. An encoder for encoding an audio signal, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to execute the steps of the method according to any one of EEE1 to EEE11.
[0128] EEE26. A decoder for decoding an audio signal from an encoded representation of the audio signal, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to execute the steps of the method according to any one of EEE12 to EEE24.
[0129] EEE27. A computer program including instructions which, when executed, cause a computer to execute the method according to any one of EEE1 to EEE24.
[0130] EEE28. A computer-readable storage medium storing the computer program according to EEE27.
Claims
1. 1. A method for encoding an audio signal, comprising the steps of: generating a plurality of subband audio signals based on the audio signal; determining a spectral envelope based on the plurality of subband audio signals; determining autocorrelation information for each subband audio signal; encoding said spectral envelope and said autocorrelation information into an encoded representation; The method includes:
2. The method of claim 1 , further comprising the step of outputting a bitstream based on the encoded representation.
3. The method of claim 1 , wherein the autocorrelation information comprises an autocorrelation value for the subband audio signal.
4. The method of claim 3 , wherein the autocorrelation values correspond to local maxima of the autocorrelation function.
5. The method of claim 4 , wherein the autocorrelation information includes a lag value.
6. The method of claim 5 , wherein the lag value corresponds to a delay value at which the autocorrelation value reaches the local maximum.
7. The spectral envelope is determined at a first update rate; The autocorrelation information is determined at a second update rate; The method of claim 1 , wherein the first update rate is different from the second update rate.
8. The method of claim 7 , wherein the first update rate is higher than the second update rate.
9. The method of claim 1 , wherein generating the plurality of sub-band audio signals comprises flattening the audio signal.
10. The method of claim 9 , wherein generating the plurality of sub-band audio signals comprises decomposing the flattened audio signal into the plurality of sub-band audio signals.
11. 1. A method of decoding an encoded representation of an audio signal, comprising the steps of: determining a plurality of reconstructed subband audio signals based on a spectral envelope and autocorrelation information of the encoded representation; generating a reconstructed audio signal based on the plurality of reconstructed subband audio signals; Including, 11. A method according to claim 10, wherein the spectral envelope is determined based on a plurality of original sub-band audio signals, the plurality of original sub-band audio signals being generated based on the audio signal, and the autocorrelation information is determined for each original sub-band audio signal.
12. The method of claim 11 , wherein the reconstructed audio signal is determined via a machine learning based generative model and / or based on spectral synthesis.
13. The method of claim 11 , wherein the autocorrelation information comprises an autocorrelation value for each original subband audio.
14. The method of claim 13 , wherein the autocorrelation values correspond to local maxima of an autocorrelation function.
15. The method of claim 14 , wherein the autocorrelation information includes a lag value.
16. The method of claim 15 , wherein the lag value corresponds to a delay value at which the autocorrelation value reaches the local maximum.
17. The spectral envelope is determined at a first update rate; The autocorrelation information is determined at a second update rate; The method of claim 11 , wherein the first update rate is different from the second update rate.
18. The method of claim 17 , wherein the first update rate is higher than the second update rate.
19. The method of claim 11 , wherein the plurality of original sub-band audio signals are generated by flattening the original audio signal.
20. The method of claim 19 , wherein the plurality of original sub-band audio signals are generated by decomposing the flattened original audio signal into the plurality of original sub-band audio signals.
Citation Information
Patent Citations
Method and device for coding / decoding voice
JP2001051698A
Multi-lag formats for audio coding
JP7654637B2