Processing of audio stereo signals
By optimizing the upmixing parameters through a generator and coefficient processor, the problems of distortion and high complexity introduced by existing stereo encoding/decoding methods are solved, achieving efficient and low-complexity audio stereo signal reconstruction and improving the trade-off between audio quality and data rate.
Patent Information
- Application Number
- CN202480049238.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-26
- Filing Date
- 2024-07-17
- Publication Date
- 2026-02-24
AI Technical Summary
Existing parametric stereo encoding/decoding methods introduce distortion, variations, and artifacts in certain situations, leading to reduced audio quality, high data rates, high processing complexity, and imperfect reconstruction of multi-channel audio signals.
This device generates an output stereo audio signal by receiving a mono audio signal and upmixing parameters. It uses a coefficient generator to generate an upmixing matrix based on the upmixing parameters, and generates the output stereo signal through matrix multiplication. It prevents the signal cancellation metric from approaching zero or infinity, optimizes the upmixing parameters to match the downmixing matrix, and reduces numerical problems and artifacts.
It improves audio quality, reduces data rate and processing complexity, reduces computational load, and provides a better perceptual audio experience and improved spatial audio reconstruction.
Smart Images

Figure CN121569340A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the processing of audio stereo signals, such as encoding / decoding / downmixing / upmixing / generation, and specifically (but not exclusively) to generating audio stereo signals based on upmixing of a mono downmixed signal using upmixing parameter data. Background Technology
[0002] Spatial audio applications have become numerous and widespread, increasingly forming at least part of many audiovisual experiences. In fact, the continuous development of new and improved spatial experiences and applications has led to an increasing demand for audio processing and rendering.
[0003] For example, virtual reality (VR) and augmented reality (AR) have received increasing attention in recent years, and many implementations and applications are reaching the consumer market. In fact, devices are being developed for presenting experiences and for capturing or recording suitable data for such applications. For example, relatively low-cost devices are being developed to allow game consoles to deliver a full VR experience. This trend is expected to continue and will indeed grow rapidly as the VR and AR markets reach a considerable size in the short term. In the audio domain, significant field exploration involves the reproduction and synthesis of realistic and natural spatial audio. The ideal goal is to generate natural audio sources so that users cannot distinguish between the synthesized source and the original source.
[0004] Extensive research and development efforts have focused on providing efficient and high-quality audio coding and decoding for spatial audio. Frequently used spatial audio representations are multi-channel audio representations, including stereo representations, and efficient coding methods for such multi-channel audio have been developed based on downmixing the multi-channel audio signal into a downmixed channel with fewer channels. One of the major advances in low-bit-rate audio coding is the use of parametric multichannel coding, where the downmixed signal is generated along with parametric data, which can then be used to upmix the downmixed signal to recreate the multi-channel audio signal.
[0005] Specifically, instead of traditional mid-side coding or intensity coding, in parametric multichannel audio coding, the multichannel input signal is downmixed to a lower number of channels (e.g., two to one), and multichannel image (stereo) parameters are extracted. The downmixed signal is then encoded using a more conventional audio encoder (e.g., a mono audio encoder). The downmixed bitstream is multiplexed with the encoded multichannel image parameter bitstream. This bitstream is then transmitted to a decoder, where the process is reversed. First, the downmixed audio signal is decoded, and then the multichannel audio signal is reconstructed using the encoded multichannel image upmix parameters.
[0006] An example of stereo decoding is described in "Advances in Parametric Coding for High-Quality Audio" by E. Schuijers, W. Oomen, B. den Brinker, and J. Breebaart (114th AES Conference, Amsterdam, Netherlands, 2003, preprint 5852). In the described method, the downmixed mono signal is parameterized by utilizing the natural separation of the signal into three components (objects): transients, sine waves, and noise. Further details are provided in "Low Complex Parametric Stereo Coding" by E. Schuijers, J. Breebaart, and H. Purnhagen (116th AES Conference, Berlin, Germany, 2004, preprint 6073), which describes how parametric stereo can be implemented with low (decoder) complexity when combined with spectral band replication (SBR).
[0007] In the described method, decoding is based on the use of a so-called decorrelation process. The decorrelation process generates a decorrelated auxiliary signal from the mono signal. During stereo reconstruction, both the mono signal and the decorrelated auxiliary signal are used to generate an upmixed stereo signal based on upmixing parameters. Specifically, the two signals can be multiplied by a 2×2 matrix with time and frequency correlations having coefficients determined according to the upmixing parameters to provide the output stereo signal. When combined with Spectral Band Replication (SBR), the method allows for parametric stereo encoding / decoding with low (decoder) complexity. The decorrelation process generates a synthetic auxiliary signal d[n] from the mono signal m[n]. During stereo reconstruction, both signals m[n] and d[n] are mixed to form a stereo pair l[n], r[n]. To further reduce (decoder) complexity, how to move the decorrelation process into the subband domain is also described. This is also the form that has been standardized for HE-AACv2 (see, for example, “An Overview of the Coding Standard MPEG-4 Audio Amendments 1 and 2: HE-AAC, SSC, and HE-AAC v2” by ACden Brinker, J. Breebaart, P. Ekstrand, J. Engdegård, F. Henn, K. Kjörling, W. Oomen and H. Purnhagen, EURASIP Log on Audio Speech and Music Processing, January 2009, and ISO / IEC 14496-3:2005, Information Technology—Coding of Audiovisual Objects—Part 3: Audio).
[0008] However, while parametric stereo (PS) and similar downmixing encoding / decoding methods represent a leap forward from traditional stereo and multichannel decoding, they are not optimal in all situations. In particular, known encoding and decoding methods tend to introduce distortions, variations, artifacts, etc., which can introduce differences between the (original) stereo audio signal input to the encoder and the stereo audio signal reconstructed at the decoder. Typically, audio quality may degrade, and imperfect reconstruction of multiple channels may occur. Furthermore, the data rate may still be higher than desired, and / or the processing complexity / resource usage may be higher than preferred. The encoding and decoding process is generally suboptimal, and particularly for certain signals, the process may introduce undesirable effects, degradation, inaccuracies, and / or artifacts.
[0009] Therefore, improved methods will be advantageous. In particular, methods that allow for increased flexibility, improved adaptability, improved performance, prevention or mitigation of numerical problems in audio processing, including encoding and decoding, increased audio quality, improved trade-offs between audio quality and data rate, reduced complexity and / or resource usage, reduced computational load, facilitated implementation, and / or improved spatial audio experience will be advantageous. Summary of the Invention
[0010] Therefore, the present invention seeks to mitigate, alleviate or eliminate one or more of the above-mentioned disadvantages, preferably alone or in any combination.
[0011] According to one aspect of the present invention, an apparatus for generating an output audio stereo signal is provided, the apparatus comprising: a receiver arranged to receive an audio data signal including: a downmixed mono audio signal of two channel signals as a first audio stereo signal; upmixing parameters for the mono audio signal, the set of upmixing parameters including a first parameter indicating a level difference between the two channel signals, a second parameter indicating a correlation between the two channel signals, and a third parameter indicating a phase difference between the two channel signals; a coefficient generator arranged to generate coefficients for an upmixing matrix based on the upmixing parameters; and a generator arranged to generate the output audio stereo signal by applying the upmixing matrix to samples of the mono audio signal and an auxiliary mono audio signal; wherein the coefficient generator is arranged to: determine a signal cancellation metric based on the upmixing parameters, the signal cancellation metric indicating signal cancellation in the sum of the two channel signals; and determine coefficients for the upmixing matrix based on the signal cancellation metric.
[0012] In many embodiments and applications, this method can provide an improved audio experience. For many signals and scenarios, this method can provide improved generation / reconstruction of stereo audio signals with improved perceived audio quality.
[0013] This method can provide an efficient implementation and, in many embodiments, can allow for reduced complexity and / or resource usage. In many cases, this method can allow the use of downmixing to reduce the data rate of data representing multichannel audio signals.
[0014] This method can particularly mitigate and compensate for numerical problems and generally provides better parameter values and / or computations. Specifically, the processing and parameter determination can prevent the denominator of the equations evaluated to determine the overmixing coefficients from approaching zero. This method can prevent or reduce the risk of parameter values (whether final or intermediate) exceeding a suitable dynamic range, and specifically prevent or reduce the risk of these parameter values approaching infinity. This method can further achieve this effect while allowing for optimal (or improved) determination of the overmixing parameters for most scenarios. For example, modifications to prevent or mitigate numerical problems can be focused on these possible scenarios without significantly impacting operations in other scenarios.
[0015] As a specific example, the method for determining the overmixing parameters can closely follow the method of ISO / IEC 14496-3:2005 for many situations, while for some situations and signals, it can explicitly prevent or mitigate the numerical problems associated with this method.
[0016] When the values of the mixing parameters are determined, this method can reduce distortion and artifacts caused by numerical problems.
[0017] Furthermore, this method allows for coordination between the encoder and decoder sides. Specifically, in many embodiments, the means for generating the output stereo audio signal can determine the signal cancellation metric based solely on the received upmixing parameters, and these parameters can specifically reflect the characteristics of the channel signals of the input stereo signal at the encoder. Therefore, the same information is available on both the encoding and decoding sides, and the same signal cancellation metric can be determined on both sides. The upmixing parameters can be determined accordingly to match the applied downmixing parameters at the encoding side. Specifically, the upmixing parameters can be determined such that the upmixing matrix is closely complementary to the downmixing matrix. Specifically, the upmixing matrix can be determined as the inverse of the downmixing matrix, thereby generating a sequence of downmixing matrix multiplications and upmixing matrix multiplications, resulting in an identity matrix.
[0018] Samples of a mono audio signal can be frequency domain samples, or can span a specific time and frequency range (especially sub-band domain samples). Samples of an auxiliary audio signal can be time domain samples, frequency domain samples, or can span a specific time and frequency range (especially sub-band domain samples).
[0019] Upmixing parameter data may include data indicating the relative characteristics between the channel signals of the first stereo audio signal. Upmixing parameters may include data indicating characteristic differences between the channels of the stereo audio signal. Upmixing parameters include data perceptually relevant to the synthesis of the output stereo audio signal. Characteristics may be, for example, differences in phase and / or intensity and / or timing and / or correlation. In some embodiments and scenarios, upmixing parameters may represent abstract characteristics that are not directly understood by humans / experts (but can generally facilitate better reconstruction / lower data rates, etc.). Upmixing parameters may include data including at least one of the following: inter-channel intensity difference, inter-channel timing difference, inter-channel correlation, and / or inter-channel phase difference of the channel signals of the stereo audio signal.
[0020] Upmixing parameters can specifically include inter-ear intensity difference (IID), inter-ear level difference (ILD), inter-channel phase difference (IPD), total phase difference (OPD), inter-channel cross-correlation (ICC), and channel phase difference (CPD) parameters.
[0021] The generator can be arranged to generate an output stereo audio signal by applying matrix multiplication to the mono audio signal and the auxiliary audio signal, where the coefficients of the upmixing matrix are determined as a function of the parameters of the upmixing parameters. The upmixing matrix is time- and frequency-dependent. Equivalently, upmixing matrices can be provided for time and / or frequency segments, and different matrices can be provided for different time and / or frequency segments.
[0022] The auxiliary signal can be a decorrelation signal generated from the mono audio signal. The decorrelation signal can be generated to have the same level and / or frequency distribution as the mono audio signal. In some cases, the auxiliary signal can be a signal received together with the mono audio signal, and in particular, it can be a side signal or residual signal for the first audio stereo signal.
[0023] Signal cancellation metrics indicate the degree or level of signal cancellation in the sum of two channel signals. Specifically, they indicate the signal level / power / amplitude of the sum of the two channel signals relative to the sum of the signal levels / power / amplitude of the two channel signals.
[0024] Signal cancellation metrics can indicate the degree / level of signal cancellation in the sum of two channel signals and / or equivalently indicate the degree / level of signal cancellation in the difference / subtraction between two channel signals (which can be considered as negative signal cancellation of the signal).
[0025] In some embodiments, the signal cancellation metric may be a normalized signal cancellation metric, and specifically normalized relative to the level / power / energy of the first stereo signal. In some embodiments, the signal cancellation metric may be in the range of -1 to +1. In some embodiments, the sign of the signal cancellation metric may indicate whether signal cancellation occurs in the summation of the channel signals and / or in the difference / subtraction between the channel signals.
[0026] The signal cancellation metric can be a monotonic measure of the signal cancellation between the two channels of the (first audio stereo signal). The signal cancellation metric can be monotonicly increased (or decreased) to increase the signal cancellation between the two channels of the (first audio stereo signal).
[0027] In some embodiments, the coefficient generator may be arranged to determine the signal cancellation metric based solely on upmixing parameters. In some embodiments, the coefficient generator may be arranged to determine the signal cancellation metric based solely on a first parameter indicating the level difference between the two channel signals, a second parameter indicating the correlation between the two channel signals, and a third parameter indicating the phase difference between the two channel signals. In some embodiments, the coefficient generator may be arranged to determine the signal cancellation metric based on received parameters other than the first parameter indicating the level difference between the two channel signals, the second parameter indicating the correlation between the two channel signals, and the third parameter indicating the phase difference between the two channel signals. In some embodiments, the coefficient generator may be arranged to determine the signal cancellation metric based on parameters of the audio data signal other than the first parameter indicating the level difference between the two channel signals, the second parameter indicating the correlation between the two channel signals, and the third parameter indicating the phase difference between the two channel signals.
[0028] In some embodiments, the coefficient generator may be arranged to determine a signal cancellation metric based on parameters of the audio data signal, other than a first parameter indicating the level difference between the two channel signals, a second parameter indicating the correlation between the two channel signals, and a third parameter indicating the phase difference between the two channel signals.
[0029] The first parameter indicating the level difference between two channel signals can be the interaural intensity difference (IID) upmixing parameter. The second parameter indicating the correlation between two channel signals can be the interchannel cross-correlation (ICC) upmixing parameter. The third parameter indicating the phase difference between two channel signals can be the interchannel phase difference (IPD) upmixing parameter.
[0030] According to an optional feature of the invention, the coefficient processor is arranged to: adjust the upmixing coefficients to deviate from the coefficients used for the mono audio signal as a sum signal and channel signal, for a signal cancellation metric that satisfies the first signal cancellation requirement.
[0031] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0032] According to an optional feature of the invention, the coefficient processor is arranged to: increase the deviation of the upmixing coefficient from the coefficients of the mono audio signal used as the sum of the channel signals, for a signal cancellation metric indicating the increase in signal cancellation in the sum of the channel signals.
[0033] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0034] According to an optional feature of the invention, the coefficient processor is arranged to: for a signal cancellation metric indicating the increase in signal cancellation in the difference signal of the channel signal, increase the deviation of the upmixing coefficient from the coefficients of the mono audio signal used as the sum signal of the channel signal.
[0035] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0036] According to an optional feature of the invention, the coefficients of the mono audio signal used as the sum signal of the channel signal are given as follows: in Where IID is the interaural intensity difference upmixing parameter, ICC is the interchannel cross-correlation upmixing parameter, and IPD is the interchannel phase difference upmixing parameter.
[0037] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0038] According to an optional feature of the invention, the signal cancellation metric is essentially determined as follows: Where IID is the inter-aural intensity difference upmixing parameter, ICC is the inter-channel cross-correlation upmixing parameter, and IPD is the inter-channel phase difference upmixing parameter.
[0039] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0040] According to an optional feature of the invention, the coefficient processor is arranged to: generate a first intermediate parameter indicating a prediction of the difference signal between the channel signal and the mono audio signal; and generate upmixing coefficients in response to the first intermediate parameter, the first intermediate parameter depending on the signal cancellation metric.
[0041] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0042] According to an optional feature of the invention, the coefficient processor (107) is arranged to generate a second intermediate parameter indicating the residual signal for prediction, and to generate an overmixing coefficient in response to the intermediate parameter, the second intermediate parameter depending on the signal cancellation metric.
[0043] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0044] According to an optional feature of the invention, the coefficient processor (107) is arranged to generate the overmixing matrix as follows: Where c is the gain parameter and α and β are parameters that depend on the upmixing parameter and the signal cancellation metric, and parameter g 1,1 g 1,2 g 2,1 and g 2,2 It depends on the signal cancellation metric.
[0045] According to an optional feature of the invention, the coefficient processor (107) is arranged to generate the overmixing matrix as follows: Where c is the gain parameter, and in: Where IID is the interaural intensity difference upmixing parameter, ICC is the interchannel cross-correlation upmixing parameter, and IPD is the interchannel phase difference upmixing parameter; and where It depends on the signal cancellation metric.
[0046] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0047] In some embodiments, Where z depends on the signal cancellation metric.
[0048] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0049] According to optional features of the present invention, in and It depends on the signal cancellation metric.
[0050] This can provide a particularly advantageous implementation and / or performance, and can prevent or mitigate numerical problems, artifacts, and / or signal distortion, especially in many scenarios.
[0051] According to one aspect of the present invention, an apparatus for generating an audio data signal is provided, the apparatus comprising: a receiver arranged to receive an audio stereo signal including two channel signals; a downmixer arranged to generate a mono audio signal as a combination of the two channel signals according to a set of downmixing coefficients; a parameter generator arranged to generate an upmixing parameter set, the upmixing parameter set including a first parameter indicating a level difference between the two channel signals, a second parameter indicating a correlation between the two channel signals, and a third parameter indicating a phase difference between the two channel signals; a downmixing coefficient processor arranged to generate a downmixing coefficient set according to the set of upmixing parameters; a data signal generator arranged to generate the audio signal including a mono audio signal and the set of upmixing parameters; a signal cancellation estimator arranged to determine a signal cancellation metric according to the set of upmixing parameters, the signal cancellation metric indicating signal cancellation in the sum of the two channel signals; and wherein the downmixing coefficient processor is arranged to generate the downmixing coefficient set according to the signal cancellation metric.
[0052] According to one aspect of the present invention, a method for generating an output audio stereo signal is provided, the method comprising: receiving an audio data signal, the audio data signal including: a downmixed mono audio signal of two channel signals as a first audio stereo signal, and upmixing parameters for the mono audio signal, the upmixing parameter set including a first parameter indicating a level difference between the two channel signals, a second parameter indicating a correlation between the two channel signals, and a third parameter indicating a phase difference between the two channel signals; generating coefficients for an upmixing matrix based on the upmixing parameters; and generating an output audio stereo signal by applying the upmixing matrix to samples of the mono audio signal and an auxiliary mono audio signal; wherein generating the coefficients includes: determining a signal cancellation metric based on the upmixing parameters, the signal cancellation metric indicating signal cancellation in the sum of the two channel signals; and determining coefficients for the upmixing matrix based on the signal cancellation metric.
[0053] According to one aspect of the present invention, a method for generating an audio data signal is provided, the method comprising: receiving an audio stereo signal including two channel signals; generating a mono audio signal as a combination of the two channel signals according to a set of downmixing coefficients; generating a set of upmixing parameters, the set of upmixing parameters including a first parameter indicating a level difference between the two channel signals, a second parameter indicating a correlation between the two channel signals, and a third parameter indicating a phase difference between the two channel signals; generating a set of downmixing coefficients according to the set of upmixing parameters; generating the audio signal as including the mono audio signal and the set of upmixing parameters; determining a signal cancellation metric according to the set of upmixing parameters, the signal cancellation metric indicating signal cancellation in the sum of the two channel signals; and wherein the generation of the set of downmixing coefficients is based on the signal cancellation metric.
[0054] These and other aspects, features and advantages of the invention will become apparent from the embodiments described below, and will be illustrated with reference to the embodiments described below. Attached Figure Description
[0055] Embodiments of the invention will be described by way of example only with reference to the accompanying drawings, wherein...
[0056] Figure 1 Some elements of an example audio device according to some embodiments of the present invention are shown;
[0057] Figure 2 Some elements of an example audio device according to some embodiments of the present invention are shown;
[0058] Figure 3 It shows Figure 1 or Figure 2 The parameters of the audio device determine an example of foreign matter flowing through it;
[0059] Figure 4Examples of signal cancellation metrics according to some embodiments of the present invention are shown;
[0060] Figure 5 Examples of intermediate parameters as functions of overmixing parameters according to some embodiments of the present invention are shown;
[0061] Figure 6 Examples of intermediate parameters as functions of overmixing parameters according to some embodiments of the present invention are shown;
[0062] Figure 7 Examples of intermediate parameters as functions of overmixing parameters according to some embodiments of the present invention are shown;
[0063] Figure 8 Examples of intermediate parameters as functions of mixing parameters according to some embodiments of the present invention are shown; and
[0064] Figure 9 Some elements of a processor for implementing an apparatus are shown in some embodiments of the invention. Detailed Implementation
[0065] Figure 1 and Figure 2 Elements of an audio device according to some embodiments of the present invention are shown. Figure 1 Audio devices can generally be considered to perform decoding and upmixing functions / operations, and therefore will be referred to as decoders for the sake of brevity. Figure 2 Audio devices can generally be considered to perform encoding and downmixing functions / operations, and therefore will be referred to as encoders for the sake of brevity.
[0066] Figure 1 The audio device includes a receiver 101, which is arranged to receive a data signal / bitstream including a downmixed mono audio signal. The downmixed mono audio signal is a downmixed stereo audio signal that includes two channel signals, typically corresponding to a left channel signal and a right channel signal. In a particular example, the stereo signal has already been provided with... Figure 2 The encoder and the stereo signal downmixed to the mono audio signal by the encoder.
[0067] Additionally, the received data signal includes upmixing parameter data for upmixing the downmixed audio signal. The upmixing parameter data can specifically be a set of parameters indicating the relationship between the signals of two different audio channels of the stereo audio signal, i.e., the relationship between the channel signals combined into the downmixed mono audio signal. Typically, upmixing parameters can indicate measures of time difference, phase difference, level / intensity difference, and / or similarity, such as correlation, between the two channel signals (i.e., between the input left signal and the input right signal). Typically, upmixing parameters are provided on a per-time and per-frequency basis (time-frequency tiles). For example, new parameters can be provided periodically for a set of sub-bands. Parameters can specifically include inter-aural intensity difference (IID), inter-aural level difference (ILD), inter-channel phase difference (IPD), total phase difference (OPD), inter-channel cross-correlation (ICC), and channel phase difference (CPD) parameters, as known from the parameter stereo coding (and from the higher channel coding).
[0068] Typically, a mono audio signal is an encoded audio signal that has been encoded according to a suitable mono signal encoding standard or method, and the receiver 101 can use a decoding method corresponding to the encoding method of the encoder to decode the received encoded mono audio signal.
[0069] Receiver 101 is coupled to generator 103, which generates an output stereo audio signal corresponding to the stereo audio signal based on the downmixing signal. Generator 103 is arranged to generate the output stereo audio signal from the mono audio signal and the auxiliary audio signal based on upmixing parameter data. Specifically, the generator can generate the output stereo audio signal by applying a 2×2 matrix multiplication to samples of the mono audio signal and the auxiliary audio signal. The coefficients of the 2×2 matrix (also called the upmixing matrix) are typically determined based on the upmixing parameters of the upmixing parameter data, using time and frequency band as a basis.
[0070] Typically, upmixing involves generating an auxiliary audio signal in the form of a decorrelation signal from a mono audio signal. It has been found that by generating a decorrelation signal and mixing it with the mono audio signal, an improved quality of the upmixed signal is perceived, and decoders have therefore been developed to take advantage of this. The decorrelation signal is typically generated by a decorrelator 105, such as a full-phase filter applied to the mono audio signal. In some cases, the auxiliary signal can be a signal received along with the mono audio signal, specifically a signal generated based on the received residual or side signal generated on the encoder side and sent to the decoder side.
[0071] exist Figure 1 The device only receives mono audio signals. m The decorrelation unit 105 is used to generate the decorrelation signal.d As a decorrelated version of a mono audio signal (typically having the same energy / level and spectral shape as the mono audio signal). By using a mono audio signal m and go to related letters Number d (samples) multiplied by the above mixing matrix H To generate output stereo audio signals l',r' (Samples) are used to generate the output stereo audio signal. l',r' (in ' (This indicates that it is a decoder copy of the original input stereo audio signal provided to the encoder).
[0072] Figure 1 The decoder also includes a coefficient processor 107, which is configured to generate coefficients for the upmixing matrix H based on the received upmixing parameters, as will be described in more detail later. Specifically, the coefficients for the upmixing matrix H can be generated based on the received IID, ICC, and IPD parameters.
[0073] In some examples, the coefficients of the overmixing matrix can be generated for each sampling time of the signal, but typically at a much lower update rate. In this case, the same coefficients can be used, for example, for groups / blocks / segments of samples, or the coefficient processor 107 can be arranged, for example, to interpolate between defined values. For example, the overmixing matrix H can be defined at discrete time points where sampling is performed at a lower rate than the rate at which defined samples are taken, and time interpolation can be used to provide more appropriate time-varying coefficients.
[0074] Figure 2 It shows that it can be generated by Figure 1 An example of a device (hereafter referred to as an encoder) that receives audio data signals from a decoder.
[0075] In this example, the encoder includes a receiver 201 that receives the input stereo audio signal to be encoded and transmitted. The stereo audio signal includes two channel signals l and r fed to a downmixer 203, which is arranged to generate a mono audio signal comprising most of the signal energy of the channel signals l and r, as well as a typical residual signal or side signal s.
[0076] The encoder also includes an upmixing parameter generator 205, which is arranged to determine upmixing parameters characterizing the properties of the input channel signals l and r. Specifically, the upmixing parameter generator 205 is arranged to generate IID, ICC, and IPD parameters.
[0077] The encoder also includes a downmixing coefficient processor 207, which is arranged to determine the downmixing coefficients (which can therefore also be considered as downmixing parameters) for downmixing based on the upmixing parameters. The upmixing / downmixing parameters can specifically reflect how downmixing is performed in the encoder and how upmixing should be performed in the decoder.
[0078] The encoder also includes a data signal generator 209, which is arranged to receive at least the downmixed mono signal m and upmixing parameters, and to generate a data signal including these. The data signal generator 209 may be specifically arranged to generate suitable data representing these signals and parameters, and therefore may include suitable encoder functions, bitstream formatting functions, etc., as will be well known to those skilled in the art. In many embodiments, the data signal generator 209 is arranged to generate the data signal without including the residual / side signal s, but in some embodiments, this signal may also be encoded and included in the data signal. In such cases, the residual / side signal s is typically encoded at a data rate much lower than that of the mono audio signal m, thereby reflecting the reduced energy and reduced perceptual impact on the stereo signal generated at the decoder side.
[0079] The method for parametric stereo is defined by the Motion Picture Experts Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission in ISO / IEC 23003-3:2020, Information Technology - MPEG Audio Technology - Part 3: Unified speech and audio coding.
[0080] In this standard, upmixing (for each parameter stereo band) is described as a generalized 2×2 mixture of downmixing the mono signal m and the decorrelation signal d:
[0081] By analyzing the signal m Apply reverb type processing to derive the relevant signal d More specifically, supermixing can be described as (for convenience, the notation used is the opposite of that used in the standard specification). ):
[0082] This method uses intermediate parameters , and c The set consists of functions of the IID, ICC, and IPD parameters. Figure 3 The parameter flowchart is shown. Based on the IID, ICC, and IPD parameters, the intermediate parameter c is calculated. and Then, these parameters are used to calculate the entries of the H matrix.
[0083] c. and The parameters are specifically defined as follows:
[0084] The fundamental principle behind this upmixing can be seen by dissecting the inverse of the upmixing matrix (i.e., the downmixing matrix used by the encoder / downmixer 203):
[0085] The analysis shows the corresponding structure:
[0086] In the first step (the rightmost matrix multiplication), the traditional intermediate signal and side signal are formed as the sum signal l+r and the difference signal lr, respectively.
[0087] In the second step, the best possible (least squared) prediction based on the mono signal is achieved. Therefore, the parameter value α is a complex value that is determined to provide the best prediction based on the sum signal to the difference signal, and thus specifically provides the best prediction based on the generated mono audio signal m to the difference signal (note that the remaining matrix multiplication preserves the m signal as a direct sum signal l+r).
[0088] Then, in the third step, the residual signal of the difference signal is scaled to ensure that both signals m and d' have equal signal power. Therefore, the parameter value β is a gain parameter that adapts the level of the decorrelated signal d' to have a signal power corresponding to that of the intermediate signal and / or signal m. It should be noted that the residual signal is uncorrelated with the intermediate (mono) signal (due to the use of parameter...). (Prediction).
[0089] The final parameter c is a coefficient used to maintain the signal power in the downmixer, and specifically, it is set to ensure that c‧(l+r) has approximately the same power as the sum of the signal power of the left and right channel signals. This value is capped / limited to the value of cmax in order to maintain the practical range of the value.
[0090] Standardized methods for parametric stereo coding and decoding offer highly advantageous operation, particularly providing a high audio quality to data rate ratio / trade-off. However, the inventors have recognized that in some cases, and especially for some signals, the standardization methods result in less than ideal coding and may actually lead to significant degradation and distortion in certain specific situations. As will be described below, the inventors have also recognized that this effect and these scenarios can be mitigated or reduced by performing specific modifications to the operation.
[0091] Specifically, the inventors have recognized that problems can arise when the channel signals of the input stereo audio signal are identical or identical except for a 180° phase difference. In such cases, intermediate parameters may approach values that cause numerical and processing problems, resulting in degradation and distortion of the resulting decoded stereo audio signal. In particular, in these scenarios, the encoded and / or decoded values may approach infinite values that cannot be properly represented.
[0092] For example, when the channel signals are the same but 180° out of phase (l=-r), the upmixing parameters will have the following values: IID=1, ICC=1, IPD=π
[0093] However, this results in the following parameters:
[0094] Therefore, all parameters are difficult to process numerically, leading to problems in the encoder and decoder.
[0095] Therefore, the inventors have recognized that when the channel signals are essentially the same but 180° out of phase (l=-r), all parameters become numerically unstable. Furthermore, the downmix signal may begin to include (time-frequency) gaps, where the neutral signal may essentially have no energy (i.e., zero signal), and this can make reconstructing the stereo signal extremely difficult.
[0096] exist Figure 2 and Figure 1 The encoder and decoder devices employ methods that can mitigate and resolve these problems in many scenarios.
[0097] In this method, coefficient generator 107 is configured to determine intermediate parameters based on the upmixing parameters, and then determine the upmixing matrix coefficients based on the intermediate parameters. The intermediate parameters may closely correspond to the parameters applied in ISO / IEC 23003-3:2020, but may be specifically modified for certain signals.
[0098] In addition, Figure 1 In the decoder, coefficient processor 107 is arranged to determine a signal cancellation metric based on upmixing parameters, wherein the signal cancellation metric indicates the signal cancellation in the sum of the two channel signals of the original input stereo audio signal to the encoder. The characteristics of these channel signals of the original input stereo audio signal are represented by the upmixing parameters. In practice, upmixing parameters (such as specifically IID, IPD, and ICC) depend on the input channel signals, and specifically on the relative differences between the input channel signals. In general, upmixing parameters depend only on the characteristics of the channel signals, and specifically on the relative characteristics of the channel signals.
[0099] Therefore, the coefficient processor 107 can determine, based on information provided by upmixing parameters indicating the relative characteristics of the channel signals, how much signal cancellation will be generated when the channel signals are added together. The signal cancellation metric can indicate the sum / power / amplitude (square root of power) / signal level of the channel signals l+r relative to the sum / combined energy / power / amplitude (square root of power) / signal level of the two individual channel signals l and r.
[0100] For example, in the case of l=-r, i.e., the two signals are identical but with a phase shift of π, the summation signal l+r=0, meaning that if the two channel signals are added / summed, they will be completely canceled out. Furthermore, for this specific case, IID=1, ICC=1, IPD=π, and therefore this situation can be detected by evaluating these upmixing parameters.
[0101] When l=r, i.e., when the two signals are identical, the sum of the two signals l+r=2l=2r, thus there is negative signal cancellation. In fact, the signal energy of the sum is four times that of the left or right signal. Furthermore, for this specific case, IID=1, ICC=1, IPD=0, and therefore the situation can be detected by evaluating these upmixing parameters.
[0102] These two examples can be considered to correspond to extreme cases of signal cancellation.
[0103] As mentioned, the signal cancellation metric can indicate the sum signal energy metric determined according to the overmixing parameters, wherein the sum signal energy metric can indicate the energy level of the sum signal as the sum of the channel signals relative to the energy levels of the individual channel signals.
[0104] As a low-complexity example, a signal cancellation metric can be generated to reflect the difference between the received upmixing parameters and the upmixing parameters corresponding to the extreme cases of in-phase or out-of-phase signal cancellation. For example, the signal cancellation metric can be determined based on a comparison of the received upmixing parameters with the upmixing parameters corresponding to the maximum and / or minimum (inverse) cancellation (i.e., amplification) of the signal.
[0105] In many embodiments, the signal cancellation metric can be determined based on the upmixing parameters. For example, the signal cancellation metric can be determined as:
[0106] In the extreme case of complete signal cancellation, this value will reach values of 1 and -1, and will provide increasingly different values for other values of the upmixing parameter. Therefore, it can provide a suitable indication of how close the stereo signal is to the scenario where the channel signal is cancelled in the sum signal l+r or in the difference signal lr (corresponding to the maximum negative signal cancellation of the sum signal).
[0107] For a given signal cancellation metric, the closer the absolute value is to 1, the closer the input stereo audio signal is to the situation where the input channel signals are canceled out in the sum or difference signals. Furthermore, the sign of the signal cancellation metric indicates which of these signals the cancellation occurs in.
[0108] Another example of signal cancellation metrics is as follows:
[0109] This signal cancellation metric has some properties that are particularly advantageous in many scenarios and implementations. Because , ,and , Therefore, we can conclude that... .
[0110] Furthermore, R approaches -1 only for highly correlated out-of-phase channel signals, i.e., when the two channel signals cancel each other out in the sum signal. Therefore, R=-1 indicates that if two channel signals are added together, they will cancel each other out, resulting in a zero signal. Similarly, R approaches 1 only for highly correlated in-phase channel signals, i.e., when the two channel signals cancel each other out in the difference signal. Therefore, R=1 indicates that if one channel signal is subtracted from the other, they will cancel each other out, resulting in a zero signal.
[0111] The specific signal cancellation metric R is particularly advantageous in many scenarios. Consider the power ratio of the unadjusted intermediate and side signals: Two terms that can potentially cancel each other out can be identified: and The ratio of these two factors can provide a particularly advantageous signal cancellation metric for indicating how close the current scene is to the problematic in-phase and out-of-phase scenes, and in particular how close the current scene is to complete signal cancellation of the sum or difference of the channel signals. Therefore, the specific signal cancellation metric described above provides a particularly advantageous metric in many embodiments.
[0112] Figure 4 This shows how the R value mentioned above varies with the upmixing parameters. It can be seen that it provides a good indication of when signal cancellation might occur.
[0113] The following description will focus on the use of this particular signal cancellation metric, which has been found to provide particularly advantageous performance and allow for improved encoding, decoding, and rendering. However, it will be understood that other values and formulas for determining the signal cancellation metric based on the upmixing parameters may be used in other embodiments.
[0114] The coefficient processor 107 is arranged to determine the coefficients for the upmixing matrix based on a signal cancellation metric. Specifically, the coefficient processor 107 can be arranged to modify operations such that the operations are adapted to compensate / modify operations in scenarios where signal cancellation may occur in the sum and / or difference signals.
[0115] The coefficient processor 107 can be specifically configured to modify the determination of coefficients for near signal cancellation, thereby mitigating numerical problems and, in particular, ensuring that the determined intermediate parameters do not approach problematic values, and specifically, that they do not approach infinity. The coefficient processor 107 can be specifically configured to adapt the operations / equations used to determine the coefficients, allowing for greater constraint on the required dynamic range of intermediate calculations and parameters, thus enabling practical applications and reducing numerical challenges and problems.
[0116] In this example, the coefficient processor 107 is arranged to: adapt the upmixing coefficients for the mono audio signal to the signal cancellation metric that satisfies the first signal cancellation requirement. , The coefficients deviate from those used for the mono audio signal as a channel signal and the difference signal. In fact, when the mono audio signal is a channel signal and the difference signal, the optimal coefficients for determining the channel signals of the output stereo signal will be given as a function of the upmixing parameters. However, the coefficient processor 107 can be arranged to generate coefficients such that they differ from and deviate from such values if the signal cancellation metric satisfies the requirements. Specifically, if the signal cancellation metric indicates that the signal cancellation in the signal and / or difference signal is above a given threshold, the coefficient processor 107 can be arranged to differ from these values.
[0117] The cancellation requirement may require the signal cancellation metric and the signal cancellation of the signal to be above a threshold. Therefore, compared to existing systems where the coefficients in the encoder are calculated based on the received mono audio signal (intermediate signal) being the original input channel signal and the signal / downmixer, Figure 1 The coefficient processor 107 in the method continues to target at least some values that deviate from the optimal values for the mono audio signal as a signal.
[0118] In many embodiments, the degree of deviation from the coefficients used for the mono signal as a channel signal and signal can depend on a signal cancellation metric, and can specifically be a monotonically increasing function of the degree of signal cancellation in the channel signal and signal. Therefore, as signal cancellation in the and signal increases, the determination of the coefficients is modified to deviate increasingly from the coefficients to be determined for the mono audio signal as a channel signal and signal. In particular, for increased signal cancellation (and thus reduced signal level) in the and signal, the upmixing coefficients determined for the mono audio signal as a direct and signal signal may become increasingly large, and may indeed approach infinity or be undefined. However, in the described method, a signal cancellation metric is determined and used to control the coefficient determination, such that this is mitigated and prevented, and thus coefficients deviating from the potentially ideal coefficients used for the and signal but with reduced numerical problems are determined.
[0119] In many embodiments, the coefficient processor 107 is arranged to increase the deviation of the upmixing coefficients used for the mono audio signal from the coefficient values used for the mono audio signal as a sum of channel signals, for a signal cancellation metric indicating increased signal cancellation in the sum of channel signals. Therefore, for a signal cancellation metric indicating increased signal cancellation in the sum of channel signals, the deviation from a reference coefficient value increases, where the reference coefficient value is the (optimal) coefficient used for the mono audio signal as a sum of channel signals.
[0120] In many embodiments, the coefficient processor 107 is arranged to increase the deviation of the upmixing coefficients for the mono audio signal from the coefficient values of the mono audio signal used as the difference / subtraction signal, for a signal cancellation metric indicating increased signal cancellation in the difference / subtraction signal used for the channel signal. Therefore, for a signal cancellation metric indicating increased signal cancellation between the channel signals, the deviation from a reference coefficient value increases, where the reference coefficient value is the (optimal) coefficient for the mono audio signal used as the sum signal.
[0121] In many embodiments, coefficient processor 107 is arranged to determine, for a signal cancellation metric that satisfies a second signal cancellation requirement, the upmixing coefficients for the mono audio signal as coefficients for the mono audio signal as a sum signal of channel signals. Specifically, for a signal cancellation metric indicating that signal cancellation in the sum signal of channel signals is below a threshold, coefficient processor 107 may generate coefficients based on the mono audio signal being a sum signal, and in practice, in some embodiments, coefficient processor 107 may determine coefficients in this case substantially as defined, for example, in ISO / IEC 23003-3:2020.
[0122] In many embodiments, the coefficient processor 107 may be arranged to generate upmixing coefficients as (optimal) coefficients for the mono audio signal of the sum signal as the input channel signal when the signal cancellation metric indicates low signal cancellation in the sum signal and / or difference signal, and to generate upmixing coefficients that deviate from being optimal for the mono audio signal of the sum signal as the sum signal when the signal cancellation metric indicates high signal cancellation.
[0123] In most embodiments, the above method can be applied to all coefficients of the overmixing matrix, and in fact, for at least some values of the signal cancellation metric, all coefficients can be determined as the coefficients that will be applied to the signal. However, it should be understood that in some embodiments, the method may be applied to only a subset of one, two, or three coefficients. Specifically, in some embodiments, the method may be applied only to the coefficients used for the mono audio signal or the coefficients used for the auxiliary audio signal.
[0124] It can be used for the upmixing coefficients of the auxiliary signal ( , A similar method is used, but in this case, a deviation is introduced for signal cancellation in the difference signal of the vocal tract signal, corresponding to the maximum negative cancellation in the signal and (i.e., the maximum level increase in the signal).
[0125] The coefficients used for the mono audio signal as the input channel signal and the input channel signal, and are referred to as reference coefficients for brevity below, can specifically be the optimal coefficients used to generate the output stereo audio signal from the mono audio signal as the input channel signal and the input channel signal and the auxiliary signal, which can specifically be the decorrelated version of the mono audio signal.
[0126] The reference coefficients (used as coefficients for the mono audio signal l+r) can specifically be coefficients determined according to the method defined by ISO / IEC 23003-3:2020, that is, the reference coefficients can be specifically determined as follows: in Where IID is the inter-aural intensity difference upmixing parameter, ICC is the inter-channel cross-correlation upmixing parameter, and IPD is the inter-channel phase difference upmixing parameter.
[0127] The values IID, ICC, and IPD can be specifically determined according to ISO / IEC 23003-3:2020. Specifically, the mixing parameters IID, ICC, and IPD can be determined as follows: For complex-valued coefficients, the Hilbert inner product is defined as:
[0128] The summation on i can refer to a set of frequency domain coefficients, or it can refer to the summation over a window of time and frequency in the case of (complex-valued) subband representation.
[0129] It should be noted that the upmixing parameters IID, ICC, and IPD depend only on the input signal and are not modified based on any part of the downmixing, signal cancellation, or processing or encoding of the actual input stereo signal. This is highly advantageous in many scenarios and embodiments. In fact, a particular advantage is that the perceptual sensitivity of the parameters is well-known and understood.
[0130] Therefore, the coefficient processor 107 can be arranged to determine coefficients that deviate from the optimal coefficients, which will be applied to the mono audio signal as both a channel signal and a signal. However, the deviation depends on the signal cancellation metric and can therefore be specific to the particular cases in which signal cancellation will occur in both the channel and signal. Thus, although the method may deviate accordingly from the theoretical or optimal processing, in practice it can mitigate and often eliminate numerical problems and the associated difficulties and degradations. This can provide improved audio quality, improved robustness, and / or reduced degradation / artifacts in many scenarios.
[0131] In some embodiments, the method can be used with an encoder that consistently generates a mono audio signal as a sum signal, and thus can perform suboptimal upmixing. However, the benefits of mitigating and reducing numerical problems and issues may often far outweigh the effects of modifying the upmixing coefficients, especially since this may be limited to specific scenarios where numerical problems would be highly detrimental and cause significant distortion.
[0132] However, in many embodiments and systems, the encoder can be arranged to also determine downmixing coefficients to reflect differences in upmixing coefficients; that is, deviations in the upmixing coefficients can be compensated for by corresponding operations at the encoder, such that the generated mono audio signal (and possibly the residual signal) is modified in scenarios where signal cancellation may exceed a given level. Therefore, in many embodiments, downmixing at the encoder and upmixing at the decoder can be complementary, and both can depend on signal cancellation in the sum / difference of the two input channel signals.
[0133] For example, in some embodiments, the mono audio signal in the encoder can be generated as a sum of the input channel signals, i.e., m = l + r. However, in the specific scenario where l = -r, the signal can be modified so that it is not simply determined as a sum. In particular, if signal cancellation in the sum increases toward complete cancellation, the mono audio signal can be generated as a component including the difference signal s = lr. Therefore, the minimum level of the mono audio signal can be preserved.
[0134] However, since the generation of the mono audio signal has changed, the decoder can be configured to supplement and compensate for this change in operation. In particular, for cases where the encoder operation is modified to prevent signal cancellation, the upmixing on the decoder side can be modified accordingly.
[0135] A similar method can be used for signal cancellation in difference signals. In this case, when the input channel signal causes the difference signal to approach l=r such that lr approaches zero, the generation of the difference signal can be modified to include elements of the sum signal, thereby preventing the signal level from dropping below a given value.
[0136] Therefore, in many embodiments, the encoder can modify the downmixing based on signal cancellation in the sum and / or difference signals of the two channel signals, and in particular, can modify the downmixing coefficients of the downmixing matrix that generates the mono audio signal, and optionally modify the side signals or auxiliary signals.
[0137] Figure 2 The encoder therefore also includes a signal cancellation estimator 211, which receives the upmixing parameters determined by the upmixing parameter generator 205. The signal cancellation estimator 211 is then arranged to determine a signal cancellation metric based on the set of upmixing parameters, wherein the signal cancellation metric again indicates the signal cancellation in the sum of the two channel signals of the input stereo signal. Specifically, the signal cancellation estimator 211 can be arranged to determine the signal cancellation metric using the same algorithm, formula, and method as the coefficient processor 107 of the decoder. Therefore, the description of the signal cancellation metric generated by the coefficient processor 107 also applies (with necessary modifications) to determining the signal cancellation metric signal generated by the cancellation estimator 211.
[0138] The signal cancellation estimator 211 can accordingly generate the same signal cancellation metric as that generated by the coefficient processor 107 of the decoder. Therefore, in many embodiments, the encoder and decoder can generate the same signal cancellation metric, and thus can be arranged to use coordinated and complementary methods to generate coefficients for the encoder's downmixing matrix and upmixing matrix, respectively. In fact, in many embodiments, coefficients can be generated such that the two matrices are inverses of each other, resulting in a total downmixing and upmixing operation that recovers the original input stereo signal.
[0139] The following sections describe specific methods that can provide particularly advantageous implementations. These methods typically offer compatibility with existing standards and technical specifications, such as ISO / IEC 23003-3:2020.
[0140] As previously mentioned, the encoder matrix corresponding to the decoder matrix of ISO / IEC 23003-3:2020 can be represented by the following formula:
[0141] Parameters α, β, and c are determined based on the upmixing parameters to provide specific properties of the function / compensation channel signal. It should be noted that the signal d' is not typically explicitly calculated in the encoder, but the parameters α, β, and c involved in generating that signal are determined.
[0142] Specifically, the parameter value α is determined to generate a prediction of the difference signal lr based on the sum signal l+r. Therefore, it is a parameter that indicates the prediction of the difference signal based on the sum signal.
[0143] The parameter β is a gain parameter that is adapted to match the gain of the decorrelated signal d' to the gain of the mono audio signal m. Therefore, the parameter β is determined to indicate the relative difference (and more specifically, the ratio) between the energy / level / amplitude of the residual signal generated by the prediction and the generated mono audio signal.
[0144] Finally, determine the parameters to adjust the overall gain / energy of the mono audio signal.
[0145] exist Figure 1 and Figure 2 In some embodiments of the method of the apparatus, such a downmixing matrix can be modified to include additional matrix multiplication, such as by adding an additional gain matrix multiplied by the sum and difference signals generated by the first matrix multiplication, for example:
[0146] Then the gain / coefficients of the gain matrix can be determined. , , and This is to compensate for signal cancellation in the sum and difference signals, respectively. Therefore, the gain / coefficient can be determined based on a signal cancellation metric, which is determined in the encoder and reflects the signal cancellation in the sum and / or difference signals of the input channel signals.
[0147] Furthermore, the gain coefficients can be determined as a function of the upmix parameters / stereo parameters. These values depend only on the input signal and represent the properties of the input stereo audio signal. In particular, the upmix parameters do not depend on the output mono audio signal, but can be determined directly from the input stereo audio signal without considering any other signals.
[0148] Specifically, the encoder can be arranged to determine the upmixing parameters ICC, IID, and IPD based on the input stereo audio signal (i.e., the input channel signal). The encoder can then determine the gain / coefficients of the gain matrix based on the upmixing parameters. Specifically, the gain / coefficients can be determined such that for input channel signals that are substantially identical but out of phase and correspond to the high signal cancellation of the sum signal, gain matrix multiplication results in a portion of the difference signal being added to the intermediate signal; that is, gain matrix multiplication can cause the sum signal to be modified to include a portion of the difference signal, thereby preventing complete signal cancellation in the sum signal.
[0149] Similarly, the gain / coefficient can be determined such that for input channel signals that are substantially identical and in phase, corresponding to the high signal cancellation of the difference signal, gain matrix multiplication results in a portion of the sum signal being added to the difference signal. That is, gain matrix multiplication can cause the difference signal to be modified to include a portion of the sum signal, thereby preventing complete signal cancellation of the difference signal.
[0150] Gain can also be determined as a function of upmixing parameters / parameter stereo parameters, thus allowing them to be determined equally at the encoder and decoder sides.
[0151] A matrix can be compressed into a single matrix: , For the following situations This can be further simplified to:
[0152] In some embodiments, the downmixing coefficient processor 207 may be arranged to determine the gain such that, for scenarios where no significant signal cancellation occurs in the sum or difference signals (as indicated by a signal cancellation metric), the matrix may be determined to be very similar to the identity matrix. , ).
[0153] Then, for the scenario where significant signal cancellation occurs in the difference signal lr (IID) 1. ICC 1. IPD 0), the coefficient processor 107 can determine that the gain matrix has the following characteristics:
[0154] In this case, the downmixing operation is modified so that a portion of the sum signal is mixed into the difference signal.
[0155] Then, for the scenario where significant signal cancellation occurs in the difference signal lr (IID) 1. ICC 1. IPD (π), the coefficient processor 107 can determine that the gain matrix has the following characteristics:
[0156] In this case, the downmixing operation is modified so that a portion of the difference signal is mixed into the sum signal to generate a mono audio signal.
[0157] Alternatively or additionally, for highly correlated out-of-phase signals, the encoder can modify the downmixing coefficients so that a portion of the difference signal is leaked (added) to the sum signal. This ensures that attenuation of the sum signal / mono audio signal is prevented.
[0158] For other scenarios, downmixing can maintain a method close to, for example, the ISO / IEC 23003-3:2020 specification.
[0159] For the decoder, the upmixing matrix can be achieved by inverting each of the downmixing matrices described above: The reversal resulted in: It can be written as a single matrix: According to the above definition , This simplifies to: For the identity matrix G= This reduces the upmixing to the traditional PS prediction upmixing.
[0160] The following is given and The generalized equations. Note that these are simplified for the different gain matrices G and G-1 as described above. And among them:
[0161] In this method, parameter c is a gain parameter / coefficient, which in many embodiments can be set to a suitable value by the decoder, and specifically, it can be a design parameter that can be set according to any suitable algorithm or standard.
[0162] For example, in some embodiments, the gain parameter c can be determined using a similar and detailed formula based on the received upmixing parameters, and the gain coefficient c may further depend on the signal cancellation metric. However, in other cases (e.g., when backward compatibility for mono is not considered relevant), c can simply be set to a constant value, such as c = 0.5.
[0163] Gain matrix The value depends on the signal cancellation metric, but it should be understood that the exact dependency will depend on the specific preferences and requirements of each embodiment and application.
[0164] Figure 5 and Figure 6 The intermediate parameters determined according to the above equation and the equation according to ISO / IEC 14496-3:2005 are shown. Examples of the absolute difference between them.
[0165] Figure 7 and Figure 8 The intermediate parameters determined according to the above equation and the equation according to ISO / IEC 14496-3:2005 are shown. Examples of the absolute difference between them.
[0166] It can be seen that the deviations allowed by this method from the parameter values of ISO / IEC 14496-3:2005 are mainly limited to scenarios where the overmixing parameter indicates significant signal cancellation.
[0167] In some embodiments, the gain matrix can advantageously be given by the following equation: Wherein, the decoder inverse matrix is:
[0168] Then you can define a function. To construct the parameter-dependent gain matrix G. For example: in This represents the maximum allowed mixing value, for example... ,and m It is an appropriate value, usually a high value (e.g., 4) that ensures that modifications to traditional methods only begin to occur near the critical scenario.
[0169] In some embodiments, the gain matrix can advantageously be given by the following equation: Wherein, the inverse matrix is:
[0170] Similarly, ,For example For example, and m =4.
[0171] The above definition of the gain matrix is symmetric, making it easy to generate an invertible matrix. However, this has a minor drawback: unnecessary power loss for both in-phase and out-of-phase cases. For example, if z is small in option A2, this also means that the intermediate signal is (unnecessarily) scaled by a factor less than 1. An alternative asymmetric matrix can be defined as: Wherein, the inverse matrix is:
[0172] Weight and It can be defined, for example, as a function with mixed parameters, as follows:
[0173] Note that in this case, the inverse matrix is further simplified because the determinant of matrix G is always equal to 1:
[0174] It should be understood that other methods and gain matrices may be used in other embodiments.
[0175] In the described method, coefficient processor 107 can accordingly continue to generate intermediate parameters α based on the upmixing parameters and the signal cancellation metric (since the gain g depends on the signal cancellation metric). The intermediate parameter α indicates the prediction of the difference signal of the channel signal based on the mono audio signal, where the difference signal can specifically be the subtraction signal lr (or rl). Coefficient processor 107 can then generate coefficients for the upmixing matrix based on the first intermediate parameter.
[0176] Furthermore, the coefficient processor 107 can be configured to generate a second intermediate parameter β, which indicates the residual signal generated after the prediction based on the first intermediate parameter α. The second intermediate parameter β is determined based on the upmixing parameters and the signal cancellation metric. The coefficient processor 107 can then continue to generate upmixing matrix coefficients based on these intermediate parameters.
[0177] As another example, the default encoder matrix can be extended by adding two additional gain matrices, one before prediction and one after prediction:
[0178] This can be written as a single matrix operation:
[0179] The decoder upmixing matrix can be determined as the inverse of the encoder downmixing matrix:
[0180] This can be written as a single matrix operation:
[0181] In this case, the above equation can be used again to determine and parameter.
[0182] Similarly, the same signal cancellation metric can be used. R To control the first additional gain matrix ( ):
[0183] Weight and It can be defined, for example, as a function of the signal cancellation metric, and therefore the upmixing parameters are as follows:
[0184] It should be noted that the combined encoder downmixing matrix in the example is given as follows: as well as
[0185] If we ignore the normalization factor of the determinant (g11g22-g12g21), we can see that: g 11 =1 g 12 = g 21 = g 22 =1+
[0186] The normalization factor of the determinant can be included in or compensated by the gain factor c, so it can be seen that the example method is equivalent.
[0187] Audio devices can be specifically implemented in one or more appropriately programmed processors. In particular, artificial neural networks can be implemented in one or more such appropriately programmed processors. Different functional blocks, especially artificial neural networks, can be implemented in separate processors and / or, for example, in the same processor. Examples of suitable processors are provided below.
[0188] Figure 9 This is a block diagram illustrating an example processor 900 according to an embodiment of the present disclosure. Processor 900 can be used to implement one or more processors, which implement the apparatus or elements thereof as described above (particularly including one or more artificial neural networks). Processor 900 can be any suitable processor type, including but not limited to microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable arrays (FPGAs) (wherein the FPGA is programmed to form a processor), graphics processing units (GPUs), application-specific integrated circuits (ASICs) (wherein the ASIC is designed to form a processor), or combinations thereof.
[0189] Processor 900 may include one or more cores 902. Core 902 may include one or more arithmetic logic units (ALUs) 904. In some embodiments, in addition to or in place of ALU 904, core 902 may include a floating-point logic unit (FPLU) 906 and / or a digital signal processing unit (DSPU) 908.
[0190] Processor 900 may include one or more registers 312 communicatively coupled to core 902. Registers 312 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 312 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 902.
[0191] In some embodiments, processor 900 may include one or more levels of cache memory 910 communicatively coupled to core 902. Cache memory 910 may provide computer-readable instructions to core 902 for execution. Cache memory 910 may provide data for processing by core 902. In some embodiments, the computer-readable instructions may have already been provided to cache memory 910 by local memory (e.g., local memory attached to external bus 916). Cache memory 910 may be implemented using any suitable cache memory type, such as metal-oxide-semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.
[0192] Processor 900 may include controller 914, which can control inputs to processor 900 from other processors and / or components included in the system and / or outputs from processor 900 to other processors and / or components included in the system. Controller 914 can control data paths in ALU 904, FPLU 906, and / or DSPU 908. Controller 914 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 914 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.
[0193] Register 912 and cache 910 can communicate with controller 914 and core 902 via internal connections 920A, 920B, 920C, and 920D. These internal connections can be implemented as buses, multiplexers, cross switches, and / or any other suitable connection technology.
[0194] Inputs and outputs for processor 900 may be provided via bus 916, which may include one or more conductive lines. Bus 916 may be communicatively coupled to one or more components of processor 900, such as controller 914, cache 910, and / or register 912. Bus 916 may be coupled to one or more components of the system.
[0195] Bus 916 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 932. ROM 932 may be a mask ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 933. RAM 933 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 935. The external memory may include flash memory 934. The external memory may include a magnetic storage device such as a disk 936. In some embodiments, the external memory may be included within the system.
[0196] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination of these. The invention can optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Therefore, the invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.
[0197] Although the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is defined only by the appended claims. Furthermore, while features may appear to have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0198] Furthermore, although listed separately, multiple units, elements, circuits, or method steps can be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features can be advantageously combined, and inclusion in different claims does not imply that such a combination of features is not feasible and / or advantageous. Furthermore, including a feature in a class of claims does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other claim classes. Moreover, the order of features in a claim does not imply any particular order in which the features must operate, and in particular, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Furthermore, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided as clarifying examples only and should not be construed as limiting the scope of the claims in any way.
Claims
1. An apparatus for generating an output stereo audio signal, the apparatus comprising: Receiver (101), which is arranged to receive audio data signals, said audio data signals including: A mono audio signal, which is a downmixing of the two channels of the first stereo audio signal; Upmixing parameters for the mono audio signal, the set of upmixing parameters including a first parameter indicating the level difference between the two channel signals, a second parameter indicating the correlation between the two channel signals, and a third parameter indicating the phase difference between the two channel signals; A coefficient generator (107) is arranged to generate coefficients for the overmixing matrix based on the overmixing parameters; A generator (103) is arranged to generate the output audio stereo signal by applying the upmixing matrix to samples of the mono audio signal and the auxiliary mono audio signal. in The coefficient generator (107) is arranged as follows: A signal cancellation metric is determined based on the upmixing parameters, the signal cancellation metric indicating the signal cancellation in the sum of the two channel signals; and The coefficients used for the upmixing matrix are determined based on the signal cancellation metric.
2. The apparatus according to claim 1, wherein, The coefficient processor (107) is arranged to: for the signal cancellation metric that satisfies the first signal cancellation requirement, adapt the overmixing coefficients, which are used as coefficients for the overmixing matrix, to deviate from the coefficients used for the mono audio signal, which is the sum signal of the channel signal.
3. The apparatus according to claim 1 or 2, wherein, The coefficient processor (107) is arranged to: for the signal cancellation metric indicating the addition of signal cancellation in the sum signal of the channel signal, increase the overmixing coefficients relative to the coefficients of the mono audio signal used as the sum signal of the channel signal.
4. The apparatus according to any of the preceding claims, wherein, The coefficient processor (107) is arranged to: for the signal cancellation metric indicating the increase in signal cancellation in the difference signal of the channel signal, increase the overmixing coefficients relative to the coefficients of the mono audio signal used as the sum signal of the channel signal.
5. The apparatus according to any one of claims 2 to 4, wherein, The coefficients of the mono audio signal used as the sum of the channel signals are given as follows: in Wherein, IID is the inter-aural intensity difference upmixing parameter, ICC is the inter-channel cross-correlation upmixing parameter, and IPD is the inter-channel phase difference upmixing parameter.
6. The apparatus according to any of the preceding claims, wherein, The signal cancellation metric is essentially determined as follows: Wherein, IID is the inter-aural intensity difference upmixing parameter, ICC is the inter-channel cross-correlation upmixing parameter, and IPD is the inter-channel phase difference upmixing parameter.
7. The apparatus according to any of the preceding claims, wherein, The coefficient processor (107) is configured to generate a first intermediate parameter based on the upmixing parameter and the signal cancellation metric, the first intermediate parameter indicating a prediction of the difference signal for the channel signal based on the mono audio signal, and to generate the coefficients based on the first intermediate parameter.
8. The apparatus according to claim 7, wherein, The coefficient processor (107) is configured to generate a second intermediate parameter indicating the residual signal for the prediction based on the upmixing parameter and the signal cancellation metric, and to generate the coefficients based on the intermediate parameter.
9. The apparatus according to any of the preceding claims, wherein, The coefficient processor (107) is arranged to generate the overmixing matrix as follows: Where c is the gain parameter and α and β are parameters that depend on the upmixing parameter and the signal cancellation metric, and parameter g 1,1 g 1,2 g 2,1 and g 2,2 It depends on the signal cancellation metric.
10. The apparatus according to any of the preceding claims, wherein, The coefficient processor (107) is arranged to generate the overmixing matrix as follows: Where c is the gain parameter, and in: Wherein, IID is the interaural intensity difference upmixing parameter, ICC is the interchannel cross-correlation upmixing parameter, and IPD is the interchannel phase difference upmixing parameter; and where, It depends on the signal cancellation metric.
11. The apparatus according to claim 9, wherein: in, and It depends on the signal cancellation metric.
12. An apparatus for generating audio data signals, the apparatus comprising: Receiver (201) is configured to receive an audio stereo signal comprising two channel signals; A downmixer (203) is arranged to generate a mono audio signal as a combination of the two channel signals according to a set of downmixing coefficients; A parameter generator (205) is arranged to generate an upmixing parameter set, the upmixing parameter set including a first parameter indicating the level difference between the two channel signals, a second parameter indicating the correlation between the two channel signals, and a third parameter indicating the phase difference between the two channel signals; A downmixing coefficient processor (207) is arranged to generate the downmixing coefficient set based on the upmixing parameter set; A data signal generator (209) is arranged to generate an audio signal comprising the mono audio signal and the set of upmixing parameters; A signal cancellation estimator (211) is arranged to determine a signal cancellation metric based on the set of upmixing parameters, the signal cancellation metric indicating signal cancellation in the sum of the two channel signals; as well as The downmixing coefficient processor (207) is configured to generate the downmixing coefficient set based on the signal cancellation metric.
13. A method for generating an output stereo audio signal, the method comprising: Receive audio data signals, the audio data signals including: A mono audio signal, which is a downmixing of the two channels of the first stereo audio signal; Upmixing parameters for the mono audio signal, the set of upmixing parameters including a first parameter indicating the level difference between the two channel signals, a second parameter indicating the correlation between the two channel signals, and a third parameter indicating the phase difference between the two channel signals; Generate coefficients for the upmixing matrix based on the upmixing parameters; The output stereo audio signal is generated by applying the upmixing matrix to samples of the mono audio signal and the auxiliary mono audio signal. in Generating the coefficients includes: A signal cancellation metric is determined based on the upmixing parameters, the signal cancellation metric indicating the signal cancellation in the sum of the two channel signals; and The coefficients used for the overmixing matrix are determined based on the signal cancellation metric.
14. A method for generating an audio data signal, the method comprising: Receives stereo audio signals including signals from two channels; A mono audio signal is generated as a combination of the two channel signals based on the set of downmixing coefficients. Generate an upmixing parameter set, which includes a first parameter indicating the level difference between the two channel signals, a second parameter indicating the correlation between the two channel signals, and a third parameter indicating the phase difference between the two channel signals; The set of lower mixing coefficients is generated based on the set of upper mixing parameters; The audio signal is generated as a set of upmixing parameters, including the mono audio signal and the upmixing parameter set. A signal cancellation metric is determined based on the set of upmixing parameters, the signal cancellation metric indicating the signal cancellation in the sum of the two channel signals; as well as The set of downmixing coefficients is generated based on the signal cancellation metric.
15. A computer program product comprising computer program code units adapted to perform all the steps of claim 14 when the program is run on a computer.