Audio signal processing during high-frequency reconstruction
The additional correction step in HFR processes adjusts the spectral envelope using frequency-dependent gain coefficients to address discontinuities in high-band signals, enhancing audio quality by reducing spectral discontinuities and level fluctuations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-25
AI Technical Summary
Conventional high-frequency reconstruction (HFR) technologies introduce artificial spectral envelope discontinuities and level fluctuations in high-band signals, particularly for audio signals with large fluctuations in energy levels, leading to perceived loss of high-frequency energy and audible discontinuities.
An additional correction step is introduced in the HFR process to adjust the spectral envelope of high-frequency signals, using spectral gain coefficients derived from a frequency-dependent curve fitted to the energies of low-frequency subband signals, and applying these coefficients to amplify and adjust the energy of high-frequency subband signals to match target energies while limiting noise introduction.
This approach improves the audio quality of high-frequency components by reducing spectral discontinuities and level fluctuations, resulting in more accurate and stable high-band signal regeneration.
Smart Images

Figure 2026053557000004 
Figure 2026053557000005 
Figure 2026053557000006
Abstract
Description
Technical Field
[0001] This application relates to High Frequency Reconstruction / Regeneration (HFR) of audio signals. In particular, this application relates to a method and system for performing HFR of an audio signal having large fluctuations in energy levels over a low frequency range used to reconstruct the high frequencies of the audio signal.
Background Art
[0002] HFR technologies, such as Spectral Band Replication (SBR) technology, allow for a significant improvement in the coding efficiency of traditional perceptual audio coders. In combination with MPEG-4 Advanced Audio Coding (AAC), HFR forms a very efficient audio coder, which is already used in the XM Satellite Radio system and Digital Radio Mondiale, and is also standardized in 3GPP (registered trademark), the DVD Forum, and others. The combination of AAC and SBR is called aacPlus. This is part of the MPEG-4 standard, which is called the High Efficiency AAC Profile (HE-AAC) in this standard. In general, HFR technologies can be combined with any perceptual audio coder in a way that is compatible with existing and future ones, thus providing the possibility to upgrade established broadcast systems such as MPEG Layer 2 used in the Eureka DAB system. The various methods of HFR can also be combined with a speech coder to allow for wideband speech at very low bitrates.
[0003] The fundamental idea behind HFR is the observation that there is usually a strong correlation between the characteristics of a signal in the high-frequency range and the characteristics of the same signal in the low-frequency range. Therefore, a good approximation of the representation of the original input high-frequency range of the signal can be achieved by a signal transposition from the low-frequency range to the high-frequency range.
[0004] This transfer concept was established in International Publication No. 98 / 57436, incorporated by reference, as a method for regenerating higher frequency bands from lower frequency bands of an audio signal. Substantial bitrate savings can be achieved by using this concept in audio coding and / or speech coding. While the following discussion will focus on audio coding, it should be noted that the methods and systems described are equally applicable to speech coding and unified speech and audio coding (USAC).
[0005] High-frequency reconstruction can be performed in the time domain or frequency domain using a selected filter bank or transform. This process typically involves several steps. Two main operations are first generating a high-frequency excitation signal, and then shaping the high-frequency excitation signal to approximate the spectral envelope of the original high-frequency spectrum. The step of generating the high-frequency excitation signal may be based on, for example, single-sideband modulation (SSB). In this case, a sine wave of frequency ω is mapped to a sine wave of frequency ω+Δω, with Δω as a fixed frequency shift. In other words, the high-frequency signal can be generated from the low-frequency signal by a "copy up" operation from the low-frequency subband to the high-frequency subband. A further technique for generating the high-frequency excitation signal may involve harmonic transposition of the low-frequency subband. Harmonic transposition of order T is typically designed to map a sine wave of frequency ω of the low-frequency signal to a sine wave of frequency Tω of the high-frequency signal, with T>1.
[0006] HFR technology may be used as part of a source coding system, where miscellaneous control information guiding the HFR process is transmitted from the encoder to the decoder along with a representation of the narrowband / low-frequency signal. For systems where additional control signals cannot be transmitted, the process may be applied on the decoder side using suitable control data inferred from available information on the decoder side.
[0007] The envelope tuning of the high-frequency excitation signal described above aims to achieve a spectral shape similar to the original high-band spectral shape. To do this, the spectral shape of the high-frequency signal needs to be modified. In other words, the tuning applied to the high band is a function of the existing spectral envelope and the desired target spectral envelope.
[0008] For systems operating in the frequency domain, such as HFR systems implemented in pseudo-QMF filter banks, conventional methods are not optimal in this respect. This is because the generation of high bands by combining several contributions from the source frequency range introduces an artificial spectral envelope in the high bands that should be envelope-tuned. In other words, typically, high-band or high-frequency signals generated from low-frequency signals during the HFR process exhibit an artificial spectral envelope (typically with spectral discontinuity). This presents a challenge for spectral envelope tuners, as they must not only be able to apply the desired spectral envelope with appropriate time and frequency resolution, but also be able to cancel out the spectral characteristics artificially introduced by the HFR signal generator. This presents a difficult design constraint for envelope tuners. As a result, these challenges tend to lead to perceived loss of high-frequency energy and, especially for speech-type signals, audible discontinuity in the spectral shape of high-band signals. In other words, conventional HFR signal generators tend to introduce discontinuities and level fluctuations into high-band signals for signals with large level variations across the low-band range, such as sibilants. Subsequently, when an envelope tuner encounters this high-band signal, it cannot rationally and consistently separate the newly introduced discontinuities from any natural spectral characteristics of the low-band signal. [Prior art documents] [Non-patent literature]
[0009] [Non-Patent Document 1] ISO / IEC14496-3 Information Technology―Coding of audio-visual objects―Part3: Audio [Non-Patent Document 2] MPEG-D USAC: ISO / IEC23003-3 United Speech and Audio Coding [Overview of the project] [Problems that the invention aims to solve]
[0010] This paper outlines solutions to the aforementioned problems that lead to improved perceived audio quality. In particular, this paper describes a solution to the problem of generating high-band signals from low-band signals, in which the spectral envelope of the high-band signal is effectively adjusted to resemble the original spectral envelope in the high band, without introducing undesirable artifacts. [Means for solving the problem]
[0011] This paper proposes an additional correction step as part of high-frequency reconstructed signal generation. As a result of this additional correction step, the audio quality of the high-frequency components or high-band signals is improved. This additional correction step can be applied to all source coding systems that use high-frequency reconstruction techniques, as well as to any single-ended post-processing method or system aimed at regenerating high frequencies in an audio signal.
[0012] In one aspect, a system is described that is configured to generate a plurality of high-frequency subband signals covering a high-frequency range. The system may be configured to generate the plurality of high-frequency subband signals from a plurality of low-frequency subband signals. The plurality of low-frequency subband signals may be subband signals of a low-band or narrow-band audio signal that can be determined using a decomposed filter bank or transform. In particular, the plurality of low-frequency subband signals may be determined from a low-band time-domain signal using a decomposed QMF (quadrature mirror filter) filter bank or an FFT (Fast Fourier Transform). The plurality of generated high-frequency subband signals may correspond to approximations of the high-frequency subband signals of the original audio signal from which the plurality of low-frequency subband signals were derived. In particular, the plurality of low-frequency subband signals and the plurality of (re)generated high-frequency subband signals may correspond to the subbands of a QMF filter bank and / or FFT transform.
[0013] The system has means for receiving the plurality of low-frequency subband signals. Therefore, the system may be placed downstream of the decomposition filter bank or transformer that generates the plurality of low-frequency subband signals from the low-band signal. The low-band signal may be an audio signal decoded in a core decoder from a received bitstream. The bitstream may be stored in a storage medium, such as a compact disc or DVD, or the bitstream may be received by the decoder through a transmission medium, such as an optical or radio wave transmission medium.
[0014] The system may have means for receiving a set of target energies, which may also be called scale factor energies. Each target energy may cover a different target interval, which may also be called a scale factor band. Typically, the set of target intervals corresponding to the set of target energies covers a complete high-frequency interval. The target energies included in the set of target energies typically represent the desired energy of one or more high-frequency subband signals within the corresponding target interval. In particular, the target energies may correspond to the desired average energy of the one or more high-frequency subband signals within the corresponding target interval. The target energies of a target interval are typically derived from the energy of the high-band signals of the original audio signal within that target interval. In other words, the set of target energies typically describes the spectral envelope of the high-band portion of the original audio signal.
[0015] The system may have means for generating the plurality of high-frequency subband signals from the plurality of low-frequency subband signals. For this purpose, the means for generating the plurality of high-frequency subband signals may be configured to perform a copy transfer onto the plurality of low-frequency subband signals and / or perform harmonic conversion of the plurality of low-frequency subband signals.
[0016] Furthermore, the means for generating the plurality of high-frequency subband signals may take into account a plurality of spectral gain coefficients during the generation process of the plurality of high-frequency subband signals. The plurality of spectral gain coefficients may be associated with each of the plurality of low-frequency subband signals. In other words, each of the plurality of low-frequency subband signals may have a corresponding spectral gain coefficient from the plurality of spectral gain coefficients. The spectral gain coefficients from the plurality of spectral gain coefficients may be applied to the corresponding low-frequency subband signal.
[0017] The plurality of spectral gain coefficients may be associated with the energy of each of the plurality of low-frequency subband signals. In particular, each spectral gain coefficient may be associated with the energy of its corresponding low-frequency subband signal. In one embodiment, the spectral gain coefficients are determined based on the energy of the corresponding low-frequency subband signal. For this purpose, a frequency-dependent curve may be determined based on the plurality of energy values of the plurality of low-frequency subband signals. In this case, the method for determining the plurality of gain coefficients may rely on the frequency-dependent curve determined from a representation (e.g., logarithmic representation) of the energy of the plurality of low-frequency subband signals.
[0018] In other words, the plurality of spectral gain coefficients may be derived from a frequency-dependent curve fitted to the energies of the plurality of low-frequency subband signals. In particular, the frequency-dependent curve may be a polynomial of a predetermined order. Alternatively, or in addition, the frequency-dependent curve may have various curve segments, which are fitted to the energies of the plurality of low-frequency subband signals in various frequency intervals. The various curve segments may be various polynomials of a predetermined order. In one embodiment, the various curve segments are polynomials of order 0, and the curve segments represent the average energy values of the energies of the plurality of low-frequency subband signals in the corresponding frequency intervals. According to one further embodiment, the frequency-dependent curve is fitted to the energies of the plurality of low-frequency subband signals by performing a moving average filtering operation along the various frequency intervals.
[0019] In one embodiment, the gain coefficients included in the plurality of gain coefficients are derived from the difference between the average energy of the plurality of low-frequency subband signals and the corresponding value of the frequency-dependent curve. The corresponding value of the frequency-dependent curve may be the value of the curve at frequencies within the frequency range of the low-frequency subband signals to which the gain coefficients correspond.
[0020] Typically, the energies of the plurality of low-frequency subband signals are determined on a time grid, for example, per frame. That is, the energy of a low-frequency subband signal within a time interval defined by the time grid corresponds to the average energy of a sample of the low-frequency subband signal within that time interval, for example, per frame. Thus, a plurality of different spectral gain coefficients may be determined on a selected time grid. For example, a plurality of different spectral gain coefficients may be determined for each frame of an audio signal. In one embodiment, the plurality of spectral gain coefficients may be determined sample by, for example, by determining the energy of the plurality of low-frequency subband signals using a floating window across samples of each low-frequency subband signal. It should be noted that the system may have means for determining the plurality of spectral gain coefficients from the plurality of low-frequency subband signals. These means may be configured to perform the methods described above for determining the plurality of spectral gain coefficients.
[0021] The means for generating the plurality of high-frequency subband signals may be configured to amplify the low-frequency subband signals using the respective spectral gain coefficients. Hereafter, the terms "amplify" or "amplify" will be used, but the "amplify" operation may be replaced by other operations such as "multiply," "rescaling," or "adjust." Amplification may be performed by multiplying a sample of the low-frequency subband signal by its corresponding spectral gain coefficient. In particular, the means for generating the plurality of high-frequency subband signals may be configured to determine a sample of the high-frequency subband signal at a given time from a sample of the low-frequency subband signal at that given time and at least one preceding time. Furthermore, the sample of the low-frequency subband signal may be amplified by the respective spectral gain coefficients of the plurality of spectral gain coefficients. In one embodiment, the means for generating the plurality of high-frequency subband signals is configured to generate the plurality of high-frequency subband signals from the plurality of low-frequency subband signals according to the "copy on top" algorithm specified in MPEG-4 SBR. The multiple low-frequency subband signals used in this “copy up” algorithm may be amplified using the multiple spectral gain coefficients, where the “amplification” operation may be performed as outlined above.
[0022] The system may have means for adjusting the energy of the plurality of high-frequency sub-band signals using the set of target energies. This operation is typically referred to as spectral envelope adjustment. Spectral envelope adjustment is performed by adjusting the energy of the plurality of high-frequency sub-band signals such that the average energy of the plurality of high-frequency sub-band signals within a target interval corresponds to the corresponding target energy. This may be achieved by determining an envelope adjustment value from the energy values of the plurality of high-frequency sub-band signals within the target interval and the corresponding target energy. In particular, the envelope adjustment value may be determined from the ratio of the target energy to the energy values of the plurality of high-frequency sub-band signals within the corresponding target interval. This envelope adjustment value may be used to adjust the energy of the plurality of high-frequency sub-band signals.
[0023] In one embodiment, the means for adjusting the energy has means for restricting the adjustment of the energy of the high-frequency sub-band signals within a limiter interval. Typically, the limiter interval covers two or more target intervals. The restricting means is typically used to avoid undesirable amplification of noise within certain high-frequency sub-band signals. For example, the restricting means may be configured to determine an average envelope adjustment value of the envelope adjustment values corresponding to the target intervals covered by or within the limiter interval. Further, the restricting means may be configured to limit the adjustment of the energy of the high-frequency sub-band signals within the limiter interval to a value proportional to the average envelope adjustment value.
[0024] Instead of, or in addition to, the means for adjusting the energy of the plurality of high-frequency sub-band signals may include means for ensuring that the adjusted high-frequency sub-band signals within the specific target interval have the same energy. This means is often referred to as "interpolation" means. In other words, the "interpolation" means ensures that the energy of each high-frequency sub-band signal within the specific target interval corresponds to the target energy. The "interpolation" means may be implemented by separately adjusting each high-frequency sub-band signal within the target interval such that the energy of the adjusted high-frequency sub-band signals corresponds to the target energy associated with the specific target interval. This may be achieved by determining different envelope adjustment values for each high-frequency sub-band signal within the specific target interval. The different envelope adjustment values may be determined based on the energy of the specific high-frequency sub-band signal and the target energy corresponding to the specific target interval. In one embodiment, the envelope adjustment value for a specific high-frequency sub-band signal is determined based on the ratio of the target energy to the energy of the specific high-frequency sub-band signal.
[0025] The system may further include means for receiving control data. The control data may indicate whether to apply the plurality of spectral gain factors to generate the plurality of high-frequency sub-band signals. In other words, the control data may indicate whether additional gain adjustment of the low-frequency sub-band signals should be performed. Instead of, or in addition to, this, the control data may indicate the method to be used for determining the plurality of spectral gain factors. As an example, the control data may indicate a predetermined degree of a polynomial to be used for determining the frequency-dependent curve to be applied to the energy of the plurality of low-frequency sub-band signals. The control data is typically received from a corresponding encoder that analyzes the original audio signal and notifies the corresponding decoder or HFR system how to decode the bitstream.
[0026] Another aspect describes an audio decoder configured to decode a bitstream containing a low-frequency audio signal and a set of target energies describing the spectral envelope of a high-frequency audio signal. In other words, an audio decoder is described configured to decode a bitstream containing a low-frequency audio signal and a set of target energies describing the spectral envelope of a high-frequency audio signal. The audio decoder may have a core decoder and / or conversion unit configured to determine a plurality of low-frequency subband signals associated with the low-frequency audio signal from the bitstream. Alternatively, or in addition, the audio decoder may have a high-frequency generation unit based on the system outlined herein, the system may be configured to determine a plurality of high-frequency subband signals from the plurality of low-frequency subband signals and the set of target energies. Alternatively, or in addition, the decoder may have a merge and / or inverse conversion unit configured to generate an audio signal from the plurality of low-frequency subband signals and the plurality of high-frequency subband signals. The merge and inverse transform unit may have a composite filter bank or transform, such as an inverse QMF filter bank or an inverse FFT.
[0027] In a further aspect, an encoder configured to generate control data from an audio signal is described. The audio encoder may have means for analyzing the spectral shape of the audio signal and determining the degree of spectral envelope discontinuity introduced when regenerating the high-frequency components of the audio signal from the low-frequency components of the audio signal. Thus, the encoder may have certain elements of a corresponding decoder. In particular, the encoder may have an HFR system outlined in this paper, which allows the encoder to determine the degree of discontinuity in the spectral envelope that may be introduced into the high-frequency components of the audio signal on the decoder side. Alternatively, or in addition, the encoder may have means for generating control data to control the regeneration of the high-frequency components based on the degree of discontinuity. In particular, the control data may correspond to control data received by a corresponding decoder or the HFR system. The control data may indicate whether to use the plurality of spectral gain coefficients during the HFR process and / or which of a predetermined polynomial order should be used to determine the plurality of spectral gain coefficients. To determine this information, the ratio of selected portions of the low-frequency section, i.e., the frequency range covered by the multiple low-frequency subband signals, can be determined. This ratio information can be determined, for example, by examining the lowest and highest frequencies of the low band. This provides access to the spectral variations of the low-band signals that will later be used in the decoder for high-frequency reconstruction. A large ratio may indicate an increased degree of discontinuity. Control data can also be determined using a signal type detector. For example, the detection of a speech signal may indicate an increased degree of discontinuity. On the other hand, the detection of a prominent sine wave in the original audio signal may lead to control data indicating that the multiple spectral gain coefficients should not be used during the HFR process.
[0028] Another aspect describes a method for generating multiple high-frequency subband signals covering a high-frequency interval from multiple low-frequency subband signals. The method may include the step of receiving the multiple low-frequency subband signals and / or a set of target energies. Each target energy may cover a different target interval within the high-frequency interval. Furthermore, each target energy may represent a desired energy for one or more high-frequency subband signals within the target interval. The method may include the step of generating the multiple high-frequency subband signals from the multiple low-frequency subband signals and a set of spectral gain coefficients associated with each of the multiple low-frequency subband signals. Alternatively, or in addition, the method may include the step of adjusting the energies of the multiple high-frequency subband signals using the set of target energies. The energy adjustment step may include the step of limiting the adjustment of the energies of the high-frequency subband signals within a limiter interval. Typically, the limiter interval covers two or more target intervals.
[0029] In a further aspect, a method is described for decoding a bitstream that represents or contains a set of target energies describing the spectral envelopes of a low-frequency audio signal and a corresponding high-frequency audio signal. Typically, the low-frequency and high-frequency audio signals correspond to the low-frequency and high-frequency components of the same original audio signal. The method may include the step of determining a plurality of low-frequency subband signals associated with the low-frequency audio signal from the bitstream. Alternatively, or in addition to this, the method may include the step of determining a plurality of high-frequency subband signals from the plurality of low-frequency subband signals and the set of target energies. This step is typically performed based on the HFR method outlined in this paper. The method then may include the step of generating an audio signal from the plurality of low-frequency subband signals and the plurality of high-frequency subband signals.
[0030] Another aspect describes a method for generating control data from an audio signal. This method may include the step of analyzing the spectral shape of the audio signal to determine the degree of discontinuity introduced when regenerating the high-frequency components of the audio signal from the low-frequency components of the audio signal. Furthermore, this method may include the step of generating control data that controls the regeneration of the high-frequency components based on the degree of discontinuity.
[0031] In a further aspect, a software program is described. This software program may be adapted for execution on a processor and for performing the method steps outlined in this paper when executed on a computing device.
[0032] Another aspect describes a storage medium. This storage medium may have a software program adapted for execution on a processor and for performing the method steps outlined in this paper when executed on a computing device.
[0033] In a further aspect, a computer program product is described. This computer program may have executable instructions for performing the method steps outlined in this paper when executed on a computer.
[0034] It should be noted that the preferred embodiments of the methods and systems outlined in this patent application may be used alone or in combination with other methods and systems described herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any way. In particular, the features of each claim may be combined with each other in any way. [Brief explanation of the drawing]
[0035] The present invention will be illustrated by examples with reference to the accompanying drawings. [Figure 1a] This figure shows the absolute spectrum of an exemplary high-band signal prior to spectral envelope adjustment. [Figure 1b] This figure illustrates an exemplary relationship between the time frame of audio data and the envelope time boundary of the spectral envelope. [Figure 1c] This figure shows the absolute spectrum of an exemplary high-band signal prior to spectral envelope adjustment, along with the corresponding scale factor band, limiter band, and HF (high frequency) patch. [Figure 2] This figure shows an embodiment of an HFR system in which the copy process is complemented by an additional gain adjustment step. [Figure 3] This figure shows an approximation of the coarse spectral envelope of an exemplary low-band signal. [Figure 4] This figure shows an embodiment with an additional gain tuner that operates based on arbitrary control data and QMF subband samples and outputs a gain curve. [Figure 5] This figure shows a more detailed embodiment of the additional gain adjuster shown in Figure 4. [Figure 6] This figure shows an embodiment of an HFR system that takes a narrowband signal as input and outputs a wideband signal. [Figure 7] This figure shows an embodiment of an HFR system incorporated within an SBR module of an audio decoder. [Figure 8] This figure shows an exemplary embodiment of a high-frequency reconstruction module for an audio decoder. [Figure 9] This figure shows an exemplary embodiment of an encoder. [Figure 10a] This is a spectrogram of an exemplary voice segment decoded using a conventional decoder. [Figure 10b] This is a spectrogram of an exemplary voice segment decoded using a decoder that applies additional gain adjustment processing. [Figure 10c]This is the spectrogram of the voice segment in Figure 10a for the original unencoded signal. [Modes for carrying out the invention]
[0036] The embodiments described below are merely illustrative of the principle of the present invention, "Audio Signal Processing during High-Frequency Reconstruction." It will be understood that modifications and variations of the configurations and details described herein will be obvious to those skilled in the art. Therefore, it is intended that the invention be limited only by the scope of the accompanying claims and not by the specific details presented in the description and explanation of the embodiments herein.
[0037] As outlined above, an audio decoder using the HFR technique typically has an HFR unit for generating a high-frequency audio signal and a subsequent spectral envelope adjustment unit for adjusting the spectral envelope of that high-frequency audio signal. When adjusting the spectral envelope of an audio signal, this is typically done by a filter bank implementation or by time-domain filtering. The adjustment can endeavor to correct the absolute spectral envelope, or it can be performed by filtering that also corrects the phase characteristics. In any case, the adjustment is typically a combination of two steps: removal of the current spectral envelope and application of a target spectral envelope.
[0038] It is important to note that the methods and systems outlined in this paper are not simply aimed at removing the spectral envelope of audio signals. The methods and systems in this paper, as part of the high-frequency regeneration step, strive to provide suitable spectral correction of the spectral envelope of low-band signals. This is to avoid introducing spectral envelope discontinuities in the high-frequency spectrum generated by combining different segments of low-band signals (i.e., low-frequency signals) that have been shifted or converted to different frequency ranges of high-band signals (i.e., high-frequency signals).
[0039] In Figure 1a, the stylized spectra 100 and 110 of the HFR unit output before entering the envelope tuner are displayed. In the upper panel, an up-copy method (having two patches) is used to generate the high-band signal 105 from the low-band signal 101, for example, the up-copy method used in MPEG-4 SBR (Split-Band Reproduction) outlined in Non-Patent Document 1, incorporated by reference. The up-copy method moves the lower frequency portions 101 to the higher frequency 105. In the lower panel, a harmonic conversion method (having two patches) is used to generate the high-band signal 115 from the low-band signal 111, for example, the harmonic conversion method of MPEG-D USAC described in Non-Patent Document 2, incorporated by reference.
[0040] In the subsequent envelope tuning stage, the target spectral envelope is applied to the high-frequency components 105 and 115. As can be seen from the spectra 105 and 115 entering the envelope tuner, discontinuities (particularly at patch boundaries) can be observed in the spectral shape of the high-band excitation signals 105 and 115, i.e., the high-band signals entering the envelope tuner. These discontinuities stem from the fact that some contributions from the low-frequency signals 101 and 111 are used to generate the high-band signals 105 and 115. As can be seen, the spectral shape of the high-band signals 105 and 115 is related to the spectral shape of the low-band signals 101 and 111. Consequently, a particular spectral shape of the low-band signals 101 and 111, for example the gradient shape shown in Figure 1a, can lead to discontinuities in the overall spectra 100 and 110.
[0041] In addition to spectra 100 and 110, Figure 1a shows an exemplary frequency band 130 of spectral envelope data representing the target spectral envelope. These frequency bands 130 are referred to as scale factor bands or target intervals. Typically, a target energy value, i.e., scale factor energy, is specified for each target interval, i.e., scale factor band. In other words, since there is typically only one target energy per target interval, the scale factor band defines the effective frequency resolution of the target spectral band. Using the scale factor or target energy specified for the scale factor band, the subsequent envelope tuner attempts to tune the high-band signals so that the energy of the high-band signals within the scale factor band is equal to the energy of the received spectral envelope data for each scale factor band, i.e., the target energy.
[0042] Figure 1c provides a more detailed description using an exemplary audio signal. This plot shows the spectrum of a real-world audio signal 121 entering the envelope tuner, along with the corresponding original signal 120. In this particular example, the SBR range, i.e., the high-frequency signal range, begins at 6.4 kHz and consists of three different copies of the low-band frequency range. These different copy frequency ranges are indicated by "Patch 1," "Patch 2," and "Patch 3." From the spectrogram, it is clear that this patch configuration introduces discontinuities in the spectral envelope at approximately 6.4 kHz, 7.4 kHz, and 10.8 kHz. In this example, these frequencies correspond to the patch boundaries.
[0043] Figure 1c further shows the scale factor band 130 and the limiter band 135, the functions of which are outlined in more detail below. In the illustrated embodiment, an envelope tuner for MPEG-4 SBR is used. This envelope tuner operates using a QMF filter bank. The main aspects of the operation of such an envelope tuner are as follows:
[0044] • Calculate the average energy of the input signal to the envelope tuner, i.e., the signal coming out of the HFR unit, through the scale factor band 130. In other words, the average energy of the regenerated high-band signal is calculated within each scale factor band / target interval 130.
[0045] For each scale factor band 130, a gain value, also called the envelope adjustment value, is determined. The envelope adjustment value is the square root of the energy ratio between the target energy (i.e., the energy target received from the encoder) and the average energy of the regenerated high-band signal 121 within each scale factor band 130.
[0046] Each envelope adjustment value is applied to the frequency band corresponding to the respective scale factor band 130 of the regenerated high-band signal 121.
[0047] Furthermore, the envelope adjuster may have additional steps and modifications, specifically as follows:
[0048] A limiting function that limits the maximum permissible envelope adjustment value applied to a certain frequency band, i.e., the limiting band 135. The maximum permissible envelope adjustment value is a function of the envelope adjustment values determined for various scale factor bands 130 that fall within the limiting band 135. Specifically, the maximum permissible envelope adjustment value is a function of the average of the envelope adjustment values determined for various scale factor bands 130 that fall within the limiting band 135. For example, the maximum permissible envelope adjustment value may be the average of the relevant envelope adjustment values multiplied by a limiting factor (e.g., 1.5). The limiting function is typically applied to limit the introduction of noise into the regenerated high-band signal 121. This is particularly important for audio signals containing prominent sine waves, i.e., audio signals with a spectrum that has a clear peak at a certain frequency. Without using the limiting function, the original audio signal would have a significant envelope adjustment value determined for scale factor bands 130 that contain such a clear peak. As a result, the spectrum of the entire scale factor band 130 (not just the distinct peak) is modified, thereby introducing noise.
[0049] • Interpolation function. This allows for the calculation of envelope adjustment values for each individual QMF subband within the scale factor band, rather than calculating a single envelope adjustment value for the entire scale factor band. Since the scale factor band typically contains two or more QMF subbands, the envelope adjustment value can be calculated as the ratio of the energy of a specific QMF subband within the scale factor band to the target energy received from the encoder, rather than calculating the ratio of the average energy of all QMF subbands within the scale factor band to the target energy received from the encoder. Thus, different envelope adjustment values may be determined for each QMF subband within the scale factor band. It should be noted that the received target energy value for a given scale factor band typically corresponds to the average energy of that frequency range in the original signal. How the received average target energy is applied to the corresponding frequency band of the regenerated high-band signal depends on the decoder operation. This can be done by applying an overall envelope adjustment value to the QMF subbands within the scale factor band of the regenerated high-band signal, or by applying individual envelope adjustment values to each QMF subband. The latter method can be thought of as if the received envelope information (i.e., one target energy per scale factor band) is "interpolated" through the QMF subbands within the scale factor band to provide higher frequency resolution. Therefore, this method is referred to as "interpolation" in MPEG-4 SBR.
[0050] Referring to Figure 1c, it can be seen that the envelope adjuster must apply a high envelope adjustment value to match the spectrum 121 of the signal entering the adjuster to the spectrum 120 of the original signal. It can also be seen that large fluctuations in the envelope adjustment value occur within the limiter bandwidth 135 due to the discontinuity. As a result of such large fluctuations, the envelope adjustment value corresponding to the minimum of the regenerated spectrum 121 is limited by the limiter function of the envelope adjustment value. Consequently, the discontinuity in the regenerated spectrum 121 remains even after the envelope adjustment operation has been performed. On the other hand, if the limiter function is not used, undesirable noise may be introduced as outlined above.
[0051] Therefore, any signal with large level fluctuations across the low-band range presents a problem in the regeneration of the high-band signal. This problem arises due to the discontinuity introduced during the high-frequency regeneration of the high band. When the envelope tuner then encounters this regenerated signal, it cannot reasonably and consistently separate the newly introduced discontinuity from any "real-world" spectral characteristics of the low-band signal. This problem has two aspects. First, spectral shapes that the envelope tuner cannot compensate for are introduced into the high-band signal. As a result, the output has an incorrect spectral shape. Second, an instability effect is perceived. This is due to the fact that this effect comes in and out depending on the low-band spectral characteristics.
[0052] This paper addresses the aforementioned problem by describing a method and system for providing an HFR high-band signal at the input of an envelope tuner that does not exhibit spectral discontinuities. For this purpose, it is proposed to remove or reduce the spectral envelope of the low-band signal when performing high-frequency regeneration. By doing so, any introduction of spectral discontinuities into the high-band signal before performing envelope tuning is also avoided. As a result, the envelope tuner does not need to handle such spectral discontinuities. In particular, a conventional envelope tuner may be used in which the limiter function of the envelope tuner is used to avoid introducing noise into the regenerated high-band signal. In other words, the method and system described can be used to regenerate an HFR high-band signal with little or no spectral discontinuities and low noise levels.
[0053] It should be noted that the time resolution of the envelope tuner may differ from the time resolution of the proposed processing of the spectral envelope during high-band signal generation. As mentioned above, the processing of the spectral envelope during high-band signal regeneration is intended to modify the spectral envelope of the low-band signal in order to reduce subsequent processing within the envelope tuner. This processing, i.e., modification of the spectral envelope of the low-band signal, may be performed, for example, once per audio frame. Here, the envelope tuner may adjust the spectral envelope over several time intervals, i.e., using several received spectral envelopes. This is outlined in Figure 1b, where a time grid 150 of spectral envelope data is shown in the upper panel, and a time grid 155 for processing the spectral envelope of the low-band signal during high-band signal regeneration is shown in the lower panel. As can be seen in the example in Figure 1b, the time boundaries of the spectral envelope data change with time, while the processing of the spectral envelope of the low-band signal acts on a fixed time grid. It can also be seen that several envelope adjustment cycles (represented by time boundary 150) may be performed during one cycle of processing the spectral envelope of the low-band signal. In the illustrated example, the processing of the spectral envelope of the low-band signal acts frame by frame; that is, multiple different spectral gain coefficients are determined for each frame of the signal. It should be noted that the processing of the low-band signal may act on any time grid, and such a processing time grid does not need to coincide with the time grid of the spectral envelope data.
[0054] Figure 2 depicts a filter bank-based HFR system 200. The HFR system 200 operates using a pseudo-QMF filter bank, and the system 200 can be used to generate the high-band and low-band signals 100 shown in the upper panel of Figure 1a. However, an additional gain adjustment step is added as part of the high-frequency generation process. The high-frequency generation process is a copy-and-paste process in the illustrated example. A low-frequency input signal is analyzed by a 32-subband QMF 201 to generate multiple low-frequency subband signals. Some or all of the low-frequency subband signals are patched to higher frequency positions based on the HF (high frequency) generation algorithm. Furthermore, these multiple low-frequency subbands are directly input to a composite filter bank 202. The composite filter bank 202 described above is a 64-subband inverse QMF 202. For the specific implementation shown in Figure 2, the use of the 32-subband QMF decomposition filter bank 201 and the 64-subband QMF composite filter bank 202 results in an output sampling rate of the output signal that is twice the input sampling rate of the input signal. However, the systems outlined in this paper are not limited to systems with different input and output sampling rates. A number of different sampling rate relationships can be conceivable by those skilled in the art.
[0055] As outlined in Figure 2, lower frequency subbands are mapped to higher frequency subbands. A gain adjustment stage 204 is introduced as part of this copy process. The resulting high-frequency signals, i.e., the multiple high-frequency subband signals, are input to an envelope tuner 203 (possibly having limiter and / or interpolation functions) prior to being combined with the multiple low-frequency subband signals in the composite filter bank 202. By using such an HFR system 200, and in particular by using the gain adjustment stage 204, the introduction of spectral envelope discontinuities shown in Figure 1 can be avoided. For this purpose, the gain adjustment stage 204 corrects the spectral envelope of the low-band signals, i.e., the spectral envelopes of the multiple low-frequency subband signals. This allows the corrected low-band signals to be used to generate high-band signals, i.e., multiple high-frequency subband signals, that do not exhibit discontinuities, particularly at patch boundaries. Referring to Figure 1c, the additional gain adjustment stage 204 ensures that the spectral envelopes 101 and 111 of the low-band signals are modified so that the generated high-band signals 105 and 115 have no discontinuities at all or only limited discontinuities.
[0056] Correction of the spectral envelope of a low-band signal can be achieved by applying a gain curve to the spectral envelope of the low-band signal. Such a gain curve can be determined by a gain curve determination unit 400 shown in Figure 4. Module 400 takes QMF data 402 as input, corresponding to the frequency range of the low-band signal used to regenerate the high-band signal. In other words, the multiple low-frequency subband signals are input to the gain curve determination unit 400. As previously mentioned, only a subset of the available QMF subbands of the low-band signal can be used to generate the high-band signal. That is, only a subset of the available QMF subbands can be input to the gain curve determination unit 400. Furthermore, module 400 may receive optional control data 404, for example, control data sent from a corresponding encoder. Module 400 outputs a gain curve 403 that is applied during the high-frequency regeneration process. In one embodiment, the gain curve 403 is applied to the QMF subbands of the low-band signal used to generate the high-band signal. That is, the gain curve 403 may be used in a copy process over the HFR process.
[0057] Optional control data 404 may include information about the resolution of the coarse spectral envelope estimated within module 400 and / or information about the suitability of applying the gain adjustment process. Thus, control data 404 can control the amount of additional processing involved during the gain adjustment process. Control data 404 may also trigger a bypass of additional gain adjustment processing if a signal that is not well suited to coarse spectral envelope estimation occurs, for example, a signal having a single sine wave.
[0058] Figure 5 provides a more detailed overview of module 400 in Figure 4. QMF data 402 of the low-band signal is input to an envelope estimation unit 501, which estimates the spectral envelope, for example, on a logarithmic energy scale. The spectral envelope is then input to module 502, which estimates a coarse spectral envelope from the high (frequency)-resolution spectral envelope received from envelope estimation unit 501. In one embodiment, this is done by fitting a low-order polynomial, i.e., a polynomial of order in the range of 1, 2, 3, or 4, to the spectral envelope data. The coarse spectral envelope may also be determined by performing a moving average calculation of the high-resolution spectral envelope along the frequency axis. The determination of the coarse spectral envelope 301 of the low-band signal is visualized in Figure 3. It can be seen that the absolute spectrum 302 of the low-band signal, i.e., the energy 302 of the various QMF bands, is approximated by a coarse spectral envelope 301, i.e., by a frequency-dependent curve fitted to the spectral envelope of the multiple low-frequency subband signals. Furthermore, it is shown that only 20 QMF subband signals are used to generate the high-band signal, i.e., only a portion of the 32 QMF subband signals are used in the HFR process.
[0059] The method used to determine a coarse spectral envelope from a high-resolution spectral envelope, particularly the degree of the polynomial fitted to the high-resolution spectral envelope, can be controlled by arbitrary control data 404. The degree of the polynomial may be a function of the size of the frequency range 302 of the low-band signal in which the coarse spectral envelope 301 is determined, and / or a function of other parameters important for the overall coarse spectral shape of the relevant frequency range 302 of the low-band signal. Polynomial fitting computes a polynomial that approximates the data in terms of least-squares error. A preferred embodiment is outlined below with Matlab code.
[0060] [Table 1] In the code above, the input is the spectral envelope (LowEnv) of the lowband signal, obtained by averaging QMF subband samples for each subband over a time interval corresponding to the current time frame of the data to be subsequently acted upon by the envelope tuner. As mentioned above, the gain adjustment process for the lowband signal may be performed on various other time grids. In the example above, the estimated absolute spectral envelope is represented in the logarithmic domain. A low-order polynomial, a cubic polynomial in the example above, is fitted to the data. Given the polynomial, the gain curve (GainVec) is calculated from the difference in average energy between the lowband signal and the curve obtained from the polynomial fitted to the data (lowBandEnvSlope). In the example above, the operation to determine the gain curve is performed in the logarithmic domain.
[0061] Gain curve calculation is performed by the gain curve calculation unit 503. As described above, the gain curve may be determined from the average energy of a portion of the low band signal used to regenerate the high band signal, and from the spectral envelope of the portion of the low band signal used to regenerate the high band signal. In particular, the gain curve may be determined from the difference between the average energy and a coarse spectral envelope represented, for example, by a polynomial. That is, the calculated polynomial may be used to determine the gain curve. The gain curve includes separate gain values for all significant QMF subbands of the low band signal. These gain values are also referred to as spectral gain coefficients. This gain curve, including these gain values, is then used in the HFR process.
[0062] As an example, the HFR generation process based on MPEG-4 SBR is described below. The HF-generated signal is derived by the following formula (see the document MPEG-4 Part 3 (ISO / IEC 14496-3), sub-part 4, section 4.6.18.6.2, incorporated here by reference).
[0063]
number
[0064]
number
[0065] Further details regarding the relationship between p and k in the copy process are specified in the aforementioned MPEG-4 Part 3 document. In the above formula, X Low (p,l) represents a sample of the low-frequency subband signal with subband index p at time l. This sample, combined with the preceding samples, represents the high-frequency subband signal X with subband index k. High Used to generate a sample of (k,l).
[0066] The gain adjustment aspect can be used in any filter bank-based high-frequency reconstruction system, as shown in Figure 6. Here, the present invention is part of a standalone HFR unit 601 that acts on a narrowband or lowband signal 602 and outputs a wideband or highband signal 604. Module 601 may receive additional control data 603 as input, which may specify, among other things, the amount of processing used for the described gain adjustment and information about, for example, the target spectral envelope of the highband signal. However, these parameters are merely examples of arbitrary control data 603. In some embodiments, the relevant information may be derived from the narrowband signal 602 input to module 601 or by other means. That is, the control data 603 may be determined within module 601 based on information available in module 601. It should be noted that the standalone HFR unit 601 may receive the plurality of low-frequency subband signals and output the plurality of high-frequency subband signals. In other words, the decomposition / synthesis filter bank or conversion may be located outside the HFR unit 601.
[0067] As already mentioned above, it can be beneficial to signal the activation of gain adjustment processing in the bitstream from encoder to decoder. For certain signal types, such as a single sine wave, gain adjustment processing may not be significant, and therefore, it can be beneficial to allow the encoder / decoder system to turn off the additional processing to avoid introducing undesirable behavior in such borderline cases. For this purpose, the encoder may be configured to analyze the audio signal and generate control data to turn the gain adjustment processing in the decoder on or off.
[0068] Figure 7 includes a proposed gain adjustment stage in a high-frequency reconstruction unit 703, which is part of an audio codec. An example of such an HFR unit 703 is an MPEG-4 spectral band replication tool used as part of a high-efficiency AAC codec or MPEG-D USAC (Unified Speech and Audio Codec). In this embodiment, a bitstream 704 is received by an audio decoder 700. The bitstream 704 is multiplexed and separated in a demultiplexer 701. The SBR-related portion 708 of the bitstream is fed to an SBR module or HFR unit 703, and the core decoder-related bitstream 707, for example AAC data or USAC core decoder data, is sent to the core decoder module 702. Furthermore, a lowband or narrowband signal 706 is passed from the core decoder 702 to the HFR unit 703. The present invention is incorporated as part of the SBR process in the HFR unit 703, based on a system outlined, for example, in Figure 2. The HFR unit 703 outputs a broadband or highband signal 705 using the processing outlined in this paper.
[0069] Figure 8 provides a more detailed overview of one embodiment of the high-frequency reconstruction module 703. Figure 8 shows that HF (high-frequency) signal generation may be derived from different HF generation modules at different points in time. HF generation may be based on a QMF-based copy-on-transitioner 803, or HF generation may be based on an FFT-based harmonic converter 804. For either HF signal generation module, the low-band signal is processed as part of the HF generation to determine the gain curve used in the copy-on-transitioner 803 or harmonic converter 804 process (801, 802). The outputs from the two transpositions are selectively input to the envelope tuner 805. The decision of which transposition signal to use is controlled by bitstreams 704 or 708. It should be noted that, due to the copy-on-transition nature of the QMF-based transposition, the shape of the spectral envelope of the low-band signal is more clearly preserved than when using a harmonic converter. This typically leads to a more pronounced discontinuity in the spectral envelope of the high-band signal when using a copy-on-transitioner. This is shown in the upper and lower panels of Figure 1a. As a result, it may be sufficient to incorporate gain adjustment only into the method of copying the QMF base performed in module 803. Nevertheless, it may also be beneficial to apply gain adjustment to the harmonic conversion performed in module 804.
[0070] Figure 9 outlines the corresponding encoder module. The encoder 901 may be configured to analyze a particular input signal 903 and determine a suitable amount of gain adjustment processing for a particular type of input signal 903. In particular, the encoder 901 may determine the degree of discontinuity in the high-frequency subband signal that will be caused in the decoder by the HFR unit 703. For this purpose, the encoder 901 may have the HFR unit 703 or at least a relevant portion of the HFR unit 703. Based on the analysis of the input signal 903, control data 905 can be generated for the corresponding decoder. The information 905 regarding the gain adjustment to be performed in the decoder is combined with the audio bitstream 906 in the multiplexer 902 to form a complete bitstream 904 that is transmitted to the corresponding decoder.
[0071] Figure 10 shows the output spectrum of a real-world signal. Figure 10a depicts the output of an MPEG USAC decoder decoding a 12kbps mono bitstream. This section of the real-world signal is the vocal portion of an a cappella recording. The horizontal axis corresponds to the time axis, and the vertical axis corresponds to the frequency axis. Comparing the spectrogram of Figure 10a with the corresponding spectrogram of the original signal in Figure 10c, it is clear that there is a gap (see reference numerals 1001 and 1002) in the spectrum of the fricative portion of the vocal segment. Figure 10b depicts the spectrogram of the output of an MPEG USAC decoder including the present invention. From this spectrogram, it can be seen that the gap in the spectrogram has disappeared (see reference numerals 1003 and 1004, corresponding to reference numerals 1001 and 1002).
[0072] The complexity of the proposed gain adjustment algorithm was calculated as weighted MOPS. Functions such as POW / DIV / TRIG (power / division / trigonometric functions) were weighted as 25 operations, while all other operations were weighted as 1 operation. Given these assumptions, the calculated complexity is approximately 0.1 WMOPS and negligible RAM / ROM usage. In other words, the processing and memory requirements of the proposed gain adjustment process are low.
[0073] This paper describes a method and system for generating high-band signals from low-band signals. The method and system are adapted to generate high-band signals with little to no spectral discontinuity, thereby improving the perceived performance of high-frequency reconstruction methods and systems. The method and system can be easily integrated into existing audio encoding / decoding systems. In particular, the method and system can be integrated without modifying the envelope adjustment processing of existing audio encoding / decoding systems. This applies especially to the limiter and interpolator functions of the envelope adjustment processing, which can perform their intended tasks. Thus, the described method and system can be used to regenerate high-band signals with little to no spectral discontinuity and low noise levels. Furthermore, the use of control data is described. Control data may be used to adapt the parameters (and computational complexity) of the described method and system to the type of audio signal.
[0074] The methods and systems described in this paper may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on, for example, a digital signal processor or microprocessor. Other components may be implemented as, for example, hardware and / or application-specific integrated circuits. Signals encountered in the methods and systems described may be stored on a medium such as random-access memory or optical storage media. Such signals may be transmitted over a network such as a radio network, satellite network, wireless network, or wired network, such as the Internet. Typical devices utilizing the methods and systems described in this paper are portable electronic devices or other consumer devices used to store and / or play audio signals. The methods and systems may be used on a computer system, such as an Internet web server, that stores and provides audio signals, such as music signals, for download.
[0075] Several aspects are described below. [Aspect 1] A system configured to generate multiple high-frequency subband signals covering a high-frequency range from multiple low-frequency subband signals: • Means for receiving the plurality of low-frequency subband signals; A means for receiving a set of target energies, each target energy covering different target intervals within the high-frequency interval and representing a desired energy of one or more high-frequency subband signals within the target interval; A means for generating a plurality of high-frequency subband signals from the plurality of low-frequency subband signals and a plurality of spectral gain coefficients associated with each of the plurality of low-frequency subband signals; The system includes means for adjusting the energy of the plurality of high-frequency subband signals using the set of target energies. system. [Aspect 2] The system according to embodiment 1, wherein the means for adjusting the energy includes means for limiting the adjustment of the energy of high-frequency subband signals within a limiter section (135), and the limiter section covers two or more target sections (130). [Aspect 3] The system according to embodiment 1 or 2, wherein the plurality of spectral gain coefficients are associated with the energy of each of the plurality of low-frequency subband signals. [Aspect 4] The system according to embodiment 3, wherein the plurality of spectral gain coefficients are derived from frequency-dependent curves fitted to the energies of the plurality of low-frequency subband signals. [Aspect 5] The system according to embodiment 4, wherein the frequency-dependent curve is a polynomial of a predetermined order. [Aspect 6] The system according to embodiment 4 or 5, wherein the spectral gain coefficients included in the plurality of spectral gain coefficients are derived from the difference between the average energy of the plurality of low-frequency subband signals and the corresponding value of the frequency-dependent curve. [Aspect 7] The system according to any one of embodiments 1 to 6, wherein the means for generating the plurality of high-frequency subband signals is configured to amplify the plurality of low-frequency subband signals using the respective plurality of spectral gain coefficients. [Aspect 8] The means for generating the plurality of high-frequency subband signals is, - Perform a copy transfer onto the aforementioned multiple low-frequency subband signals; and / or • It is configured to perform harmonic conversion of the plurality of low-frequency subband signals. A system as described in any one of the descriptions 1 to 7. [Aspect 9] The system according to embodiment 8, wherein the means for generating the plurality of high-frequency subband signals is - A sample of the low-frequency subband signal is multiplied by the respective spectral gain coefficient of the plurality of spectral gain coefficients, thereby obtaining a modified sample; The system is configured to determine a sample of the corresponding high-frequency subband signal at a specific time from a modified sample of the low-frequency subband signal at the specific time and at least one preceding time. system. [Aspect 10] The system according to embodiment 9, wherein a sample of the corresponding high-frequency subband signal at the specified time is determined from the modified sample of the low-frequency subband signal using a copy algorithm onto an MPEG-4 SBR. [Aspect 11] The system according to any one of embodiments 1 to 10, wherein the means for adjusting the energy of the plurality of high-frequency subband signals further includes means for ensuring that the adjusted high-frequency subband signals within a specific target interval have the same energy. [Aspect 12] The plurality of low-frequency subband signals and the plurality of high-frequency subband signals • QMF filter bank and / or FFT A system according to any one of the embodiments 1 to 11, corresponding to the subbands of the [Aspect 13] A system according to any one of embodiments 1 to 12, further comprising means for receiving control data, wherein the control data is Whether to apply the multiple spectral gain coefficients to generate the multiple high-frequency subband signals; and / or The method for determining the plurality of spectral gain coefficients is shown. system. [Aspect 14] The system according to embodiment 13, when referring to embodiment 5, wherein the control data indicates the predetermined degree of the polynomial. [Aspect 15] An audio decoder configured to decode a bitstream representing a set of target energies describing the spectral envelopes of a low-frequency audio signal and a corresponding high-frequency audio signal: A core decoder and conversion unit configured to determine a plurality of low-frequency subband signals associated with the low-frequency audio signal from the bitstream; A high-frequency generation unit based on the system according to any one of embodiments 1 to 14, configured to determine a plurality of high-frequency subband signals from the plurality of low-frequency subband signals and the set of target energies; The system includes a merge and inverse transformer configured to generate an audio signal from the plurality of low-frequency subband signals and the plurality of high-frequency subband signals, decoder. [Aspect 16] An encoder configured to generate control data from an audio signal, wherein the audio encoder: A means for analyzing the spectral shape of the audio signal and determining the degree of spectral envelope discontinuity introduced when regenerating the high-frequency components of the audio signal from the low-frequency components of the audio signal; The system includes means for generating control data to control the regeneration of the high-frequency component based on the degree of discontinuity. Encoder. [Aspect 17] A method for generating multiple high-frequency subband signals covering a high-frequency range from multiple low-frequency subband signals, wherein: The step of receiving the aforementioned multiple low-frequency subband signals; A step of receiving a set of target energies, each target energy covering a different target interval within the high-frequency interval and representing a desired energy of one or more high-frequency subband signals within the target interval; The step of generating the plurality of high-frequency subband signals from the plurality of low-frequency subband signals and the plurality of spectral gain coefficients associated with each of the plurality of low-frequency subband signals; The step includes adjusting the energy of the plurality of high-frequency subband signals using the set of target energies, method. [Aspect 18] A method for decoding a bitstream representing a low-frequency audio signal and a set of target energies describing the spectral envelope of the corresponding high-frequency audio signal: The steps include determining a plurality of low-frequency subband signals associated with the low-frequency audio signal from the bitstream; The steps include determining a plurality of high-frequency subband signals from the plurality of low-frequency subband signals and the set of target energies according to the method described in Embodiment 17; The step includes generating an audio signal from the plurality of low-frequency subband signals and the plurality of high-frequency subband signals. method. [Aspect 19] A method for generating control data from an audio signal: The steps include: analyzing the spectral shape of the audio signal and determining the degree of spectral envelope discontinuity introduced when regenerating the high-frequency components of the audio signal from the low-frequency components of the audio signal; The step includes generating control data to control the regeneration of the high-frequency component based on the degree of discontinuity, method. [Aspect 20] A software program adapted for execution on a processor and for performing a step of the method described in any one of embodiments 17 to 19 when executed on a computing device. [Aspect 21] A storage medium having a software program adapted for execution on a processor and for performing a step of the method described in any one of embodiments 17 to 19 when executed on a computing device. [Aspect 22] A computer program product having executable instructions for performing the method described in any one of embodiments 17 to 19 when executed on a computer.
Claims
1. A system (601, 703) for generating multiple high-frequency audio subband signals (604) covering a high-frequency range from multiple low-frequency audio subband signals (602), wherein the system (601, 703) is: - The step of receiving the plurality of low-frequency subband signals (602); - A step of receiving a set of target energies, each target energy covering a different target interval (130) within the high-frequency interval and representing a desired energy of one or more high-frequency subband signals within the target interval (130); - A step of generating a plurality of high-frequency subband signals (604) from a plurality of low-frequency subband signals (602) and a plurality of spectral gain coefficients associated with each of the plurality of low-frequency subband signals (602), wherein generating the plurality of high-frequency subband signals (604) includes scaling the plurality of low-frequency subband signals (602) using each of the plurality of spectral gain coefficients; - A step of adjusting the energy (203) of the plurality of high-frequency subband signals (604) using the set of target energies, wherein adjusting the energies includes determining a different envelope adjustment value for each of the high-frequency subband signals within each target interval (130), system.
2. A system for generating a bitstream (904), wherein the system is: - The stage of receiving the audio signal (903); - The step of generating an audio bitstream (906) from the audio signal (903); - In the step of generating control data (905) from the audio signal (903), generating the control data (905) is: - To analyze the spectral shape of the audio signal (903) and determine the degree of spectral envelope discontinuity introduced when regenerating the high-frequency components of the audio signal (903) from the low-frequency components of the audio signal (903); - Includes generating control data (905) for controlling the regeneration of the high-frequency component based on the degree of discontinuity, Determining the degree of spectral envelope discontinuity includes determining ratio information, which is determined by examining the lowest and highest frequencies of the low-frequency components. A high value of the determined ratio information indicates a high degree of spectral envelope discontinuity, and a low value of the determined ratio information indicates a low degree of spectral envelope discontinuity. Stages; - The step of combining the control data (905) with the audio bitstream (906) to form the bitstream (904) It is configured to perform system.
3. A method for generating multiple high-frequency audio subband signals (604) covering a high-frequency range from multiple low-frequency audio subband signals (602), the method being: - The step of receiving the plurality of low-frequency subband signals (602); - A step of receiving a set of target energies, each target energy covering a different target interval (130) within the high-frequency interval and representing a desired energy of one or more high-frequency subband signals (604) within the target interval (130); - A step of generating a plurality of high-frequency subband signals (604) from a plurality of low-frequency subband signals (602) and a plurality of spectral gain coefficients associated with each of the plurality of low-frequency subband signals (602), wherein generating the plurality of high-frequency subband signals (604) includes scaling the plurality of low-frequency subband signals (602) using each of the plurality of spectral gain coefficients; - A step of adjusting the energy (203) of the plurality of high-frequency subband signals (604) using the set of target energies, wherein adjusting the energies includes determining a different envelope adjustment value for each of the high-frequency subband signals within each target interval (130), for each target interval (130). method.
4. A method for generating a bitstream (904), the method being: - The stage of receiving the audio signal (903); - The step of generating an audio bitstream (906) from the audio signal (903); - In the step of generating control data (905) from the audio signal (903), generating the control data (905) is: - To analyze the spectral shape of the audio signal (903) and determine the degree of spectral envelope discontinuity introduced when regenerating the high-frequency components of the audio signal (903) from the low-frequency components of the audio signal (903); - Includes generating control data (905) for controlling the regeneration of the high-frequency component based on the degree of discontinuity, Determining the degree of spectral envelope discontinuity involves determining ratio information by examining the lowest and highest frequencies of the low-frequency components, where a high value of the determined ratio information indicates a high degree of spectral envelope discontinuity, and a low value of the determined ratio information indicates a low degree of spectral envelope discontinuity. Stages; - The step of combining the control data (905) with the audio bitstream (906) to form the bitstream (904) including, method.
5. A software program adapted for execution on a processor and for performing the steps of the method according to claim 3 or 4 when executed on a computing device.
6. A storage medium having an encoded bitstream, wherein the encoded bitstream includes control data generated by performing the method step of claim 4.
7. A computer program product having executable instructions for performing the method described in claim 3 or 4 when executed on a computer.
Citation Information
Patent Citations
IEC14496-3