Generation of multi-channel audio signal and audio data signal representing multi-channel audio signal

By combining the parameters of horizontal difference, coherence, and phase difference with the receiver and upmixer, the complexity and audio quality issues in the encoding and decoding of multi-channel audio signals are solved, achieving efficient multi-channel audio signal reconstruction and improved spatial audio experience.

CN121925701APending Publication Date: 2026-04-24KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-09-12
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing multi-channel audio signal encoding and decoding methods are insufficient in reducing complexity and computational load, leading to problems such as audio quality degradation and high data rates. This is particularly true in virtual reality and augmented reality applications, where it is difficult to achieve an efficient audio experience and reconstruction.

Method used

The receiver receives the downmixed audio signal and upmixing parameters, and generates an output multi-channel signal through the first and second upmixers. Upmixing is performed using parameters of horizontal difference, coherence and phase difference, and combined with the upmixing of transient audio components, to generate a reconstruction that is closer to the original multi-channel audio signal.

Benefits of technology

It improves the perceived quality of multi-channel audio signals, reduces data rate and decoder complexity, provides a transient representation and channel relationships closer to the original audio signal, and enhances the spatial audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925701A_ABST
    Figure CN121925701A_ABST
Patent Text Reader

Abstract

A decoder audio device comprises: a receiver (101) that receives an audio data signal, the audio data signal comprising: a downmix audio signal that is a downmix of a multi-channel audio signal; the upmix parameter set comprises a level difference parameter, a related parameter and a phase difference parameter; and at least one transient upmix parameter indicative of an inter-channel level difference for a transient. A first upmixer (103) generates an upmix multichannel signal by upmixing a downmix audio signal according to an upmix parameter, and a second upmixer (105) generates a transient audio component by upmixing a transient audio component according to a transient upmix parameter. A generator (107) generates an output multi-channel signal from a combination of the upmix multi-channel signal and the transient audio component. The encoder audio device may generate an audio data signal including a downmix audio signal, a set of upmix parameters, and transient parameters from a received multi-channel audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the generation of multi-channel audio signals and / or audio data signals representing multi-channel audio signals, and specifically, but not exclusively, to the encoding and / or decoding of stereo signals. Background Technology

[0002] Spatial audio applications have become numerous and widespread, gradually forming at least a part of many audiovisual experiences. In fact, the continuous development of new and improved spatial experiences and applications has led to an increasing demand for audio processing and rendering.

[0003] For example, in recent years, virtual reality (VR) and augmented reality (AR) have received increasing interest, and multiple implementations and applications are reaching the consumer market. In fact, devices are being developed for rendering experiences and for capturing or recording suitable data for such applications. For example, relatively low-cost equipment is being developed to allow game consoles to provide a full VR experience. This trend is expected to continue and will actually accelerate as the VR and AR market reaches a considerable size in the near future. In the audio domain, a prominent area of ​​exploration is the reproduction and synthesis of audio from real and natural spaces. The ideal goal is to produce natural audio sources so that users cannot distinguish between synthesized and original audio sources.

[0004] Significant research and development efforts have focused on providing efficient and high-quality audio coding and decoding for spatial audio. Frequently used spatial audio representations are multi-channel audio representations, including stereo representations, and efficient coding for such multi-channel audio has been developed based on downmixing the multi-channel audio signal into a downmixed channel with fewer channels. One of the major advances in low-bit-rate audio decoding is the use of parametric multi-channel decoding, where the downmixed signal is generated along with parametric data, which can be used to upmix the downmixed signal to recreate the multi-channel audio signal.

[0005] Specifically, instead of traditional mid-side or intensity decoding, in parametric multichannel audio decoding, the multichannel input signal is downmixed to a lower number of channels (e.g., two to one), and multichannel image (stereo) parameters are extracted. The downmixed signal is then encoded using a more conventional audio decoder (e.g., a single-channel audio encoder). The downmixed bitstream is multiplexed with the encoded multichannel image parameter bitstream. This bitstream is then transmitted to the decoder, where the process is reversed. First, the downmixed audio signal is decoded, and then the multichannel audio signal is reconstructed using the encoded multichannel image / upmix parameters.

[0006] An example of stereo decoding is described in E. Schuijers, W. Oomen, B. den Brinker, J. Breebaart, “Advances in Parametric Coding for High-Quality Audio” (114th AES Convention, Amsterdam, The Netherlands, 2003, Preprint 5852). In the described method, a downmixed single-channel signal is parameterized by utilizing the natural separation of the signal into three components (objects): transients, sine waves, and noise. Further details are provided in E. Schuijers, J. Breebaart, H. Purnhagen, J. Engdegård, “Low Complexity Parametric Stereo Coding” (116th AES, Berlin, Germany, 2004, Preprint 6073), which describes how parametric stereo can be implemented with low (decoder) complexity when combined with spectral band replication (SBR).

[0007] In the described method, decoding is based on the use of a so-called decorrelation process. This decorrelation process generates a decorrelation helper signal from the single-channel signal. During stereo reconstruction, both the single-channel signal and the decorrelation helper signal are used to generate an upmixed stereo signal based on upmixing parameters. Specifically, the two signals can be multiplied by a time- and frequency-dependent 2x2 matrix with coefficients determined from the upmixing parameters to provide the output stereo signal.

[0008] However, while parametric stereo (PS) and similar downmixing encoding / decoding methods represent a leap forward from traditional stereo and multichannel decoding, they are not optimal in all scenarios. In particular, known encoding and decoding methods tend to introduce distortions, alterations, artifacts, etc., which can introduce differences between the (original) multichannel audio signal supplied to the encoder and the multichannel audio signal reconstructed at the decoder. Typically, audio quality may degrade, and imperfect reconstruction of the multichannel audio signal may occur. Furthermore, data rates may still be higher than desired, and / or processing complexity / resource usage may be higher than preferred.

[0009] Another issue is that there is often a strong desire to reduce complexity and computational load, especially on the decoder side.

[0010] Therefore, improved methods will be advantageous. In particular, methods that allow for increased flexibility, improved adaptability, improved performance, increased audio quality, improved audio quality with data rate trade-offs, reduced complexity and / or resource usage, improved decoder-side operation / processing over encoder-side input, reduced computational load, facilitated implementation methods, and / or improved spatial audio experience will be advantageous. Summary of the Invention

[0011] Therefore, the present invention seeks to mitigate, alleviate or eliminate one or more of the above-mentioned disadvantages, preferably alone or in any combination.

[0012] According to one aspect of the present invention, an audio apparatus for generating an output multi-channel signal is provided, the audio apparatus comprising: a receiver arranged to receive an audio data signal, the audio data signal comprising (data describing the following items): a downmixed audio signal, which is a downmixed version of a first multi-channel signal; an upmixing parameter set for the downmixed audio signal, each upmixing parameter set including at least: a level difference parameter indicating a level difference between channels of the first multi-channel signal; a correlation parameter indicating coherence between channels of the first multi-channel signal; a phase difference parameter indicating a phase difference between channels of the first multi-channel signal; and a parameter indicating a level difference between the first multi-channel signal; and a parameter indicating a level difference between channels of the first multi-channel signal; At least one transient upmixing parameter for at least one transient audio component of the first multichannel signal, the transient upmixing parameter indicating the level difference between the at least one transient audio component between channels of the first multichannel signal; a first upmixer arranged to generate an upmixed multichannel signal by upmixing the downmixed audio signal and according to the set of upmixing parameters; a second upmixer arranged to generate an upmixed multichannel transient audio component by upmixing the at least one transient audio component according to the at least one transient upmixing parameter; and a generator arranged to generate the output multichannel signal based on a combination of the upmixed multichannel signal and the upmixed multichannel transient audio component.

[0013] In many embodiments, the method can provide an improved audio experience. For many signals and scenarios, the method can provide improved generation / reconstruction of multi-channel audio signals with improved perceived audio quality. The method can provide not only a significantly improved representation of the transients of the original first multi-channel audio signal in the generated output multi-channel audio signal, but also improved upmixing with parameters that more closely represent the relationships between the channels of the multi-channel audio signal, as the effects of transient behavior can be compensated for.

[0014] The method can provide a particularly advantageous arrangement that, in many embodiments and scenarios, allows for the representation of both longer-term audio properties and shorter-term properties (such as both background applause and foreground applause).

[0015] In many embodiments, the method can allow the generation of improved multi-channel audio signals by allowing the encoder to adjust / modify the means of generating the multi-channel audio signal based on transient properties.

[0016] The method described can provide an efficient implementation and, in many embodiments, can allow for reduced complexity and / or resource usage. In many scenarios, the method can allow for the use of downmixing to reduce the data rate of data representing multi-channel audio signals. In fact, in many embodiments, significant improvements in audio quality can be achieved with only a small increase in the overall data rate.

[0017] In many embodiments, the method can allow for reduced complexity and / or resource usage on the decoder / generation / reconstruction side.

[0018] The samples of the downmixed audio signal can be time-domain samples or frequency-domain samples (specifically, sub-band samples). The samples can span a specific time and frequency range.

[0019] The first upmixer can be arranged to generate an upmixed multichannel signal by applying matrix multiplication to the downmixed signal and the auxiliary audio signal, wherein the coefficients of the matrix are determined based on parameters of the upmixing parameter set. The matrix can be time- and frequency-dependent.

[0020] The audio device can specifically be an audio decoder device.

[0021] The processing can be time-frequency segments or blocks, which can be different time intervals and frequency intervals. Each time-frequency segment / block can represent a frequency interval within a time interval. In many embodiments, the first multi-channel audio signal can be divided into time periods / intervals, and the frequency representation of the signal within a time period / interval can be provided by signal values ​​representing different frequency segments of the signal within the time period / interval.

[0022] Transient parameters may have a different time-frequency resolution than the set of upmixing parameters; they may have different time resolutions and / or different frequency resolutions. In many embodiments, transient parameters and upmixing parameters may have the same time resolution but different frequency resolutions. In particular, transient parameters may typically have a coarser frequency resolution than upmixing parameters. For at least some frequency intervals where upmixing parameters provide individual values, transient parameters may provide only a single parameter value. In many cases, transient parameters may provide only a single value for the entire frequency band. Therefore, in some embodiments, the frequency spectrum is not divided, and time-frequency segments / blocks may be time periods / intervals.

[0023] The first upmixer can be specifically arranged to generate an output multichannel audio signal by applying matrix multiplication to samples of the downmixed audio signal and the decorrelated signal, wherein the matrix coefficients are determined based on the set of upmixing parameters.

[0024] Upmixing multi-channel audio signals, upmixing multi-channel transient audio components, and outputting multi-channel audio signals can all have the same number of channels.

[0025] According to an optional feature of the invention, the receiver is arranged to extract a first transient audio component from the downmixed audio signal, and the second upmixer is arranged to upmix the first transient audio component according to transient upmixing parameters.

[0026] This can provide a favorable approach for many scenarios, including, for example, offering a favorable trade-off between complexity, computational resources, and / or the perceived audio quality of the generated multi-channel audio signal.

[0027] According to an optional feature of the invention, the receiver is arranged to generate a residual downmixed audio signal, which is generated by extracting a set of transient audio components from the downmixed audio signal comprising audio data for the at least one transient audio component, and the first upmixer is arranged to upmix the residual downmixed audio signal according to the set of upmixing parameters.

[0028] This can provide a favorable approach for many scenarios and offer a good trade-off between performance, data rate, and computational resources / complexity.

[0029] According to an optional feature of the invention, the first upmixer is arranged to decorrelate the residual downmixed audio signal to generate a decorrelated residual downmixed audio signal, and to generate an upmixed multichannel signal by upmixing the residual downmixed audio signal and the decorrelated residual downmixed audio signal according to an upmixing parameter set.

[0030] This can provide a useful approach in many scenarios.

[0031] According to an optional feature of the invention, the audio data signal includes audio data for the at least one transient audio component, and the second upmixer is arranged to upmix the audio data for the at least one transient audio component according to the transient upmixing parameters.

[0032] This can provide a useful approach in many scenarios and often results in improved audio quality.

[0033] In some embodiments, the downmixed audio signal is a residual downmixed audio signal representing a first multichannel signal after extracting a set of transient audio components, the set of transient audio components including at least one transient audio component.

[0034] According to an optional feature of the invention, the second upmixer is arranged to perform a translation of at least one transient audio component between two channels of the upmixed multichannel transient audio component based on the transient upmixing parameters.

[0035] This can provide a useful approach in many scenarios and often results in improved audio quality.

[0036] In some embodiments, the upmixing of at least one transient audio component by the second upmixer does not depend on other inter-channel property parameters other than the transient upmixing parameters.

[0037] In some embodiments, the audio data signal does not include any other upmixing parameters for at least one transient audio component besides the transient upmixing parameters.

[0038] In some embodiments, the audio data signal does not include correlation data of at least one transient audio component, does not include coherent data of at least one transient audio component, and does not include phase data of at least one transient audio component.

[0039] According to an optional feature of the invention, the first upmixer is arranged to perform sub-band domain upmixing of the downmixed audio signal; and the second upmixer is arranged to perform time domain upmixing of the at least one transient audio component.

[0040] This can provide particularly advantageous performance, operation, reduced complexity, and / or implementation methods in many embodiments and scenarios.

[0041] In some embodiments, the audio data signal includes transient indications for time periods used to downmix the audio signal, each transient indication indicating the number of transients within the time period.

[0042] According to an optional feature of the invention, the audio data signal includes a timing indication for at least one transient audio component, and the combination depends on the timing indication.

[0043] This can provide particularly advantageous performance, operation, and / or implementation methods in many embodiments and scenarios.

[0044] According to an optional feature of the invention, the frequency resolution for at least one transient upmixing parameter is coarser than the frequency resolution of the set of upmixing parameters.

[0045] This can provide advantageous operation and / or implementation and / or performance in many embodiments.

[0046] According to one aspect of the present invention, an audio apparatus for generating an audio data signal is provided, the audio apparatus comprising: a receiver arranged to receive a first multi-channel signal; a downmixer arranged to generate a single-channel downmixed audio signal based on the first multi-channel signal and determine an upmixing parameter set for the downmixed audio signal, each upmixing parameter set including at least: a level difference parameter indicating a level difference between channels of the first multi-channel signal; a correlation parameter indicating coherence between channels of the first multi-channel signal; a phase difference parameter indicating a phase difference between channels of the first multi-channel signal; a transient detector arranged to detect at least one transient audio component of the first multi-channel signal and generate at least one transient upmixing parameter for the at least one transient audio component of the first multi-channel signal, the at least one transient upmixing parameter indicating a level difference between the at least one transient audio component between channels of the first multi-channel signal; and a data generator arranged to generate the audio data signal including the single-channel downmixed audio signal, the upmixing parameter set, and the transient upmixing parameter.

[0047] According to an optional feature of the invention, the transient detector is arranged to detect at least one transient audio component by applying transient detection to the downmixed audio signal.

[0048] This can provide advantageous operation and / or implementation and / or performance in many embodiments.

[0049] According to an optional feature of the invention, the transient detector is arranged to detect at least one transient audio component by applying transient detection to a channel of the first multi-channel signal.

[0050] This can provide advantageous operation and / or implementation and / or performance in many embodiments.

[0051] In some embodiments, the data generator is arranged to include audio data describing at least one transient audio component in the audio data signal.

[0052] In some embodiments, the transient detector is arranged to remove at least one transient audio component from the downmixed audio signal before the downmixed audio signal is included in the audio data signal.

[0053] According to one aspect of the present invention, a method for generating an output multichannel signal is provided, the method comprising: receiving an audio data signal, the audio data signal including (data describing the following items): a downmixed audio signal, which is a downmix of a first multichannel signal; an upmixing parameter set for the downmixed audio signal, each upmixing parameter set including at least: a level difference parameter indicating a level difference between channels of the first multichannel signal; a correlation parameter indicating coherence between channels of the first multichannel signal; a phase difference parameter indicating a phase difference between channels of the first multichannel signal; and at least one transient upmixing parameter for at least one transient audio component of the first multichannel signal, the transient upmixing parameter indicating a level difference between the at least one transient audio component between channels of the first multichannel signal; generating an upmixed multichannel signal by upmixing the downmixed audio signal and according to the upmixing parameter set; generating an upmixed multichannel transient audio component by upmixing the at least one transient audio component according to the at least one transient upmixing parameter; and generating the output multichannel signal based on a combination of the upmixed multichannel signal and the upmixed multichannel transient audio component.

[0054] According to one aspect of the present invention, a method for generating an output audio signal is provided, the method comprising: receiving a first multi-channel signal; generating a single-channel downmixed audio signal based on the first multi-channel signal and determining an upmixing parameter set for the downmixed audio signal, each upmixing parameter set comprising at least: a level difference parameter indicating a level difference between channels of the first multi-channel signal; a correlation parameter indicating coherence between channels of the first multi-channel signal; and a phase difference parameter indicating a phase difference between channels of the first multi-channel signal; detecting at least one transient audio component of the first multi-channel signal and generating at least one transient upmixing parameter for the at least one transient audio component of the first multi-channel signal, the at least one transient upmixing parameter indicating a level difference between the at least one transient audio component between channels of the first multi-channel signal; and generating the audio data signal comprising the single-channel downmixed audio signal, the upmixing parameter set, and the transient upmixing parameter.

[0055] These and other aspects, features and advantages of the invention will be apparent from and set forth with reference to the embodiments described below. Attached Figure Description

[0056] Embodiments of the invention will be described by way of example only with reference to the accompanying drawings, wherein...

[0057] Figure 1 The illustration shows some elements of an example audio device according to some embodiments of the present invention;

[0058] Figure 2 The illustration shows some elements of an example audio device according to some embodiments of the present invention;

[0059] Figure 3 An example of the transient representation of an audio signal frame is illustrated;

[0060] Figure 4 An example illustrating the transient representation of an audio signal in frames;

[0061] Figure 5 Examples of elements of a transient detector according to some embodiments of the present invention are illustrated;

[0062] Figure 6 An example of the transient representation of an audio signal frame is illustrated;

[0063] Figure 7 An example of the transient representation of an audio signal frame is illustrated;

[0064] Figure 8 Examples of stereo audio signals, stereo transient audio signals, and stereo residual audio signals are illustrated.

[0065] Figure 9 The illustration shows an example of the arrangement of elements used to train a neural network;

[0066] Figure 10 The illustration shows an example of the time-frequency representation (spectral graph) of a stereo applause signal;

[0067] Figure 11 The illustration shows an example of the time-frequency representation (spectral graph) of a stereo applause signal;

[0068] Figure 12 The illustration shows an example of a time period that includes a transient stereo signal;

[0069] Figure 13 The illustration shows some elements of an example audio device according to some embodiments of the present invention;

[0070] Figure 14 The illustration shows some elements of an example audio device according to some embodiments of the present invention;

[0071] Figure 15 The illustration shows some elements of an example audio device according to some embodiments of the present invention; and

[0072] Figure 16 The illustration shows some elements of a processor for implementing an audio device according to some embodiments of the present invention. Detailed Implementation

[0073] Figure 1 The illustration shows some elements of an audio apparatus for generating multi-channel audio signals according to some embodiments of the present invention. Figure 2 The illustration shows an example of an audio device arranged to generate an audio data signal representing a multi-channel audio signal, hereinafter referred to as a first multi-channel audio signal. Figure 2 The audio data signal generated by the audio device can be specifically fed to Figure 1 An audio device that can be arranged to generate an output multichannel audio signal as a copy of a first multichannel audio signal. Figure 1 The audio device will also be referred to as a decoder audio device (or simply as a decoder), and Figure 2 The audio device will also be referred to as an encoder audio device (or simply as an encoder).

[0074] The decoder audio device includes a receiver 101, which is arranged to receive a data signal / bitstream comprising a downmixed audio signal as a multi-channel audio signal. Specifically, the data signal / bitstream may be a data signal / bitstream generated by the encoder audio device to represent the first multi-channel audio signal.

[0075] The following description will focus on the case where the multichannel audio signal is a stereo signal and the downmix signal is a single-channel signal. However, it should be understood that the methods and principles described are equally applicable to multichannel audio signals with more than two channels and downmix signals with more than one channel (although fewer channels than multichannel audio signals).

[0076] In addition to the downmixed audio signal, the received data signal also includes upmixing parameter data, which comprises a set of upmixing parameters used to upmix the downmixed audio signal. Upmixing parameters can specifically be parameters indicating the relationship between signals of different audio channels of a multi-channel audio signal (specifically a stereo signal) and / or between the downmixed signal and the audio channels of the multi-channel audio signal. Typically, upmixing parameters can indicate measures of time difference, phase difference, level / intensity difference, and / or similarity, such as correlation.

[0077] The set of upmixing parameters should include at least the following items: The level difference parameter indicates the level difference between channels (and specifically two channels) of the first multi-channel audio signal. The level difference parameter can be specifically, for example, the binaural intensity difference (IID) and / or binaural level difference (ILD) known from ISO / IEC 23003-3:2020 Information technology —MPEG audio technologies — Part 3: Unified speech and audio coding. The correlation parameter indicates the coherence between channels (specifically, two channels) of the first multi-channel audio signal. The correlation parameter can be, for example, the inter-channel cross-correlation (ICC) parameter known from ISO / IEC 23003-3:2020 Information technology — MPEG audio technologies — Part 3: Unified speech and audio coding. The phase difference parameter indicates the phase difference between channels (and specifically two channels) of the first multi-channel audio signal. The phase difference parameter can be, for example, the inter-channel phase difference (IPD), total phase difference (OPD), or channel phase difference (CPD) parameter known from ISO / IEC 23003-3:2020 Information technology —MPEG audio technologies — Part 3: Unified speech and audio coding.

[0078] The parameter set can specifically include the IID, ICC, and IPD parameters as defined in ISO / IEC 23003-3:2020 Information technology—MPEG audio technologies—Part 3: Unified speech and audio coding. In particular, the upmixing parameters IID, ICC, and IPD can be defined as follows: Where l and r represent the signal values ​​of the two channels / signals of the first multi-channel audio signal (specifically, the left and right channel signals of the stereo signal), and<a,b> This represents the complex inner product between vectors a and b.

[0079] Typically, upmixing parameters are provided based on each time and each frequency (time-frequency block). For example, new parameters can be provided periodically for each subband in the subband set.

[0080] The encoder audio device is arranged to receive a first multi-channel audio signal and generate an audio data signal representing the first multi-channel audio signal, having a representation including a downmixed audio signal and a set of upmixing parameters. Specifically, the encoder audio device may be a parametric stereo (PS) encoder that receives a stereo signal and encodes it into a single-channel audio signal with associated upmixing parameterized data.

[0081] Typically, the downmixed audio signal is encoded, and the receiver 101 is arranged to decode the downmixed audio signal to provide the downmixed audio signal, i.e., the single-channel signal in this particular example, as well as the set of upmixing parameters and any other desired data.

[0082] The decoder audio device includes a first upmixer 103 and a second upmixer 105, which are arranged to generate an upmixed multichannel signal. The upmixed multichannel signal is fed to a generator 107, which is arranged to combine the signals to generate an output multichannel signal.

[0083] The first upmixer 103 is configured to generate an upmixed multichannel signal by upmixing the downmixed audio signal based on a set of upmixing parameters, specifically based on level difference parameters, correlation parameters, and phase difference parameters (e.g., IID, ICC, IPD parameters). In some embodiments, conventional upmixing of the downmixed signal can be performed. For example, in a stereo case, the first upmixer 103 can be configured to perform parameterized stereo upmixing.

[0084] In this method, the audio data signal further includes transient data indicating one or more properties of one or more transients in the first multi-channel signal represented by the audio data signal. Therefore, the audio data signal includes data reflecting one or more properties of transients in the original multi-channel signal that the decoder audio device attempts to reproduce. A second upmixer 105 is arranged to generate upmixed multi-channel transient audio components by upmixing one or more transient audio components according to at least one transient upmixing parameter. The second upmixer 105 can accordingly generate transient components for the output multi-channel signal based on the transient data provided in the audio data signal.

[0085] The first upmixer 103 and the second upmixer 105 are coupled to the generator 107, which is arranged to receive the upmixed multichannel signal and the upmixed multichannel transient audio components and combine them into an output multichannel signal.

[0086] In this method, multiple separate upmixing operations are therefore employed by using different upmixed data extracted from the audio data signal and upmixing different elements of the audio signal individually.

[0087] The encoder audio device includes a receiver 201 that receives a first multi-channel audio signal from an internal or external source. The receiver 201 is coupled to a downmixer 203, which is arranged to downmix the first multi-channel audio signal to generate a downmixed audio signal having fewer channels than the first multi-channel audio signal. In addition to the downmixed audio signal, the downmixer 203 continues to generate sets of upmixing parameters, each set of upmixing parameters as previously described with respect to the audio data signal including at least: a level difference parameter indicating the level difference between channels of the multi-channel audio signal; a correlation parameter indicating the coherence between channels of the multi-channel audio signal; and a phase difference parameter indicating the phase difference between channels of the multi-channel audio signal.

[0088] It should be understood that various methods are known for generating such downmixed audio signals and associated upmixing parameters, and any method may be used where appropriate without diminishing the invention.

[0089] In many embodiments, the first multi-channel audio signal may specifically be a stereo signal, and the upmixing parameters can be generated based on samples of the left and right channel signals of the input stereo signal. In this case, the downmixed audio signal is a single-channel downmixed audio signal.

[0090] Additionally, the encoder audio device includes a transient detector 205 arranged to detect transients in the first multi-channel audio signal (and / or equivalently, in the downmixed audio signal). The transient detector 205 is arranged to detect transient audio components in the first multi-channel signal and generate at least one transient upmixing parameter for each detected transient audio component. The upmixing parameter indicates the level difference of transient audio components between channels of the first multi-channel signal. Specifically, the transient upmixing parameter can be an inter-channel intensity difference (IID) specifically determined for the transient audio component (e.g., determined for a small duration corresponding to the duration of the transient audio component).

[0091] In many embodiments, the transient upmixing parameter can be a composite transient upmixing parameter, which includes multiple parameter / property values ​​for the transient (or equivalently, multiple transient upmixing parameters can be provided for a given transient audio component). Specifically, in many embodiments, the transient upmixing parameter can also indicate the timing properties of the transient audio component, such as the timing when the transient specifically occurs and / or the duration of the transient.

[0092] Typically, in different embodiments, the specific transient parameters and properties transmitted may differ. For example, in many cases, transient data may include an indication of the presence, number, time, duration, amplitude, inter-channel level difference, inter-channel phase difference, and inter-channel coherence of one or more transients present in the first multi-channel audio signal (and typically in the downmixed audio signal).

[0093] The encoder audio device also includes a data signal generator 207, which generates audio data signals including data representing downmixed audio signals, upmixed parameters, and transient parameters.

[0094] In many embodiments, the encoder audio device and the decoder audio device are arranged to perform sub-band processing. Specifically, upmixing parameters can be generated for different (frequency) sub-bands of the downmixed audio signal and the first multi-channel audio signal.

[0095] Specifically, receiver 201 or downmixer 203 may include a filter bank arranged to generate a frequency sub-band representation of the downmixed audio signal. Typically, receiver 201 or downmixer 203 may include a filter bank applied to all channels of the first multi-channel audio signal, such that each channel signal is divided into sub-bands. Downmixing can then be performed on a per-sub-band basis, wherein upmixing parameters are determined for each sub-band and a sub-band downmixed signal is generated. Then, in some cases, the sub-band downmixed audio signal may be directly included in the audio data signal as a sub-band downmixed audio signal, or it may be transformed to the time domain to provide a time-domain signal.

[0096] The filter bank can be a group of quadrature mirror filters (QMFs), or it can be implemented, for example, by a fast Fourier transform (FFT). However, it should be understood that many other filter banks and methods for dividing an audio signal into multiple sub-band signals are known and can be used. Specifically, the filter bank can be a group of complex-valued pseudo-QMFs, thereby producing, for example, 32 or 64 complex-valued sub-band signals.

[0097] Furthermore, this processing is typically performed within time intervals. In most embodiments, the first multi-channel audio signal is divided into time intervals / segments by applying, for example, FFT or QMF filtering to samples of each signal, utilizing a transformation to the frequency / subband domain. For example, each channel of the multi-channel audio signal can be divided into time intervals of 2048, 1024, or 512 samples, and said time intervals can then be processed to generate, for example, multiple (e.g., 32, 16, 8) subband signals with 64, 32, or 16 subband samples. Thus, a sample set can be determined for each subband of the downmixed audio signal. Furthermore, for each time interval / interval and frequency interval / subband, an upmixing parameter set can be generated.

[0098] It should be noted that the number of time-domain samples is not directly coupled to the number of subbands. Typically, for a so-called critical sampling filter bank with N frequency bands, every N input samples will yield N subband samples (one for each subband). An oversampling filter bank will produce more output samples. For example, for every N input samples, it will generate k*N output samples, i.e., k consecutive samples for each frequency band.

[0099] Therefore, a set of mixing parameters can be generated, where each set is provided for a given time interval and a given frequency interval (also known as a given time-frequency block or segment).

[0100] Each set of upmixing parameters may include IID, ICC, and IPD values ​​as previously described in detail, and therefore these parameters are provided with a given time resolution and a given frequency resolution. The time interval may vary, but in many embodiments it typically has a fixed duration. In some embodiments, the subband size / frequency resolution may also be fixed / constant for all subbands, but in many embodiments, subbands may have different resolutions / sizes. In many embodiments, the filter bank may be arranged to generate subband signals with subbands having equal bandwidth, and in many other embodiments, the filter bank may be arranged to generate subband signals with subbands having different bandwidths. For example, a higher frequency subband may have a higher bandwidth compared to a lower frequency subband. Furthermore, subbands may be grouped together to form a higher bandwidth subband.

[0101] Typically, subbands can have bandwidths ranging from 10 Hz to 10,000 Hz.

[0102] However, in many embodiments, the time-frequency resolution of the transient upmixing parameters may differ from the time-frequency resolution of the upmixing parameter set.

[0103] Transient data / parameters are typically provided with a different time-frequency resolution than the overmixed parameter set, and in many cases can be provided with a finer timing resolution or a coarser frequency resolution than the overmixed parameter set, and in many cases both a finer time resolution and a coarser frequency resolution. For example, a finer timing granularity than the (processing) time period / interval can be used to indicate the timing of transients, and, for example, non-frequency-dependent inter-channel level differences can be provided.

[0104] A set of upmixing parameters can be provided for time-frequency blocks corresponding to subbands and time periods as previously described in the processing (e.g., for a fixed number of samples per time period and subband). However, in some cases, a finer temporal resolution can be provided for the transient parameter values. For example, the timing or duration of the transient can be provided at a higher temporal resolution than the sampling time of the subband samples.

[0105] Furthermore, in many embodiments, transient properties can be provided with a coarser frequency resolution compared to a set of upmixing parameters. Specifically, in some embodiments, a set of upmixing parameters may be provided for each subband, while a single common transient parameter value may be provided for multiple, and possibly all, subbands.

[0106] In many embodiments, a set of one or more parameters can be provided for each transient in the set of detected transients. Parameters may specifically include an indication of the channel level difference between channels of a first multi-channel audio signal.

[0107] Figure 3 The diagram illustrates a specific example of a time period / frame FRM that detects three transients. Here, for a given time period / frame, three transients are detected at positions p0, p1, and p2. These positions / moments can be encoded and included in the audio data signal and transmitted accordingly to the decoder audio device. Furthermore, for each transient parameter position p, the amplitude level difference of the transient is determined and included in the audio data signal. For example, the IID (Inter-channel Intensity Difference) can be determined and included in the audio data signal. In this case, a positive IID can correspond to a left shift of the transient signal, and a negative IID can correspond to a right shift of the transient signal. In some embodiments, individual amplitudes a0, a1, and a2 can be additionally or alternatively transmitted. Thus, transient data can be used to encode a representation of the detected transients.

[0108] In some embodiments, such as Figure 4 As illustrated, the duration of a transient can be determined, encoded, and transmitted in the audio data signal to the decoder audio device.

[0109] In many embodiments, parameter values ​​can be quantized into a relatively small number of levels, and therefore a relatively small number of bits can be used for each value. In many embodiments, the word length can be no more than 1, 2, 3, or 4 bits. For example, amplitude values ​​can be quantized into several levels (e.g., 5 or 7 discrete levels) using only a few bits.

[0110] The encoding of parameter values ​​in audio data signals can, for example, use absolute or differential encoding. Specifically, the three IID values ​​corresponding to positions p0, p1, and p2 can be differentially encoded into an IID (averaged over the frequency band) transmitted for the entire frame (each frequency band).

[0111] The encoder audio device can accordingly generate transient data indicating transients in the first multi-channel audio signal and provide it to the decoder audio device, where it is used to perform separate upmixing on the upmixing of the set of upmixing parameters, and thus generate additional transient upmixing signals. The encoder audio device can provide this transient data with a different time-frequency resolution than the upmixing parameters, thereby allowing independent optimization of the transient data. In particular, a coarser frequency resolution can be used, and in many scenarios, the transient data may not include any frequency dependence, but can provide the same parameter values ​​for all sub-bands of the downmixed audio signal. In many embodiments, the parameter values ​​can also be coarsely quantized to several discrete levels. Therefore, very low data overhead can be achieved in many embodiments. However, it has been found that providing this transient data / information and using it for a second separate transient upmixing operation can achieve significantly improved perceived audio quality, and in particular, can significantly improve the perceived audio realism for some scenarios and environments.

[0112] It should be understood that different methods can be used to detect transients in the first multi-channel audio signal and / or the downmixed audio signal (in fact, these operations can be considered equivalent, since transients in the first multi-channel audio signal also exist in the downmix, and therefore detecting transients in the first multi-channel audio signal also detects transients in the downmixed audio signal, and vice versa).

[0113] Figure 5 An example of an element that can be used in a transient detector 205 in an encoder audio device is illustrated. In this example, the transient detector 205 can be arranged to detect transients in a stereo signal.

[0114] Transient detection can be performed independently for different channels, and in this example, specifically for the left and right channels, the detections are then combined. In other embodiments, it can be based directly on information from both channels, for example, by taking into account the downmixed audio signal. This approach can be advantageous because it can utilize inter-channel level (e.g., IID) parameters. The following examples will primarily consider such an approach.

[0115] In many embodiments, as will be described below, the transient detector 205 may detect a transient in response to detecting that a first level difference measure (specifically, a first IID measure) indicating the level difference between channels of a first multi-channel audio signal differs from a second level difference measure (specifically, a second IID measure for the same channel) indicating the level difference between channels differs by more than a threshold, wherein the first level difference measure is determined over a shorter time interval than the second level difference. In many embodiments, the shorter time interval may not exceed 10%, 20%, 30%, or 50% of the time interval for the second level difference.

[0116] exist Figure 5 In the example, transient detector 205 includes an analysis filter bank 501 that typically decomposes the left and right stereo channels into time-frequency (TF) representations, where the distribution of center frequencies follows, for example, the log-interval critical bands of the human auditory system (inner ear). Spectral decomposition can be performed using a hybrid quadrature mirror filter bank (QMF), which produces fine resolution at low frequencies, where the resolution decreases (bandwidth increases) with increasing frequency. It should be understood that in many embodiments, such an analysis filter bank 501 can be equivalently part of receiver 201, and the generated sub-band representation can also be used for downmixing, upmixing parameter estimation, etc.

[0117] QMF decomposition can produce complex outputs, and the transient detector 205 includes an envelope circuit 503 that determines the real envelope of both the left and right channels. The simple envelope is given by the following equation: in, and yes Frequency slab and time slices The real and imaginary parts of the time-frequency samples. The square root operation can be omitted to reduce complexity.

[0118] It should be noted that, depending on the implementation, it is not always necessary to calculate the real envelope, because the IID calculation already includes calculating the squared amplitude of the complex signal.

[0119] The transient detector 205 includes detection circuitry 505, which is coupled to envelope circuitry 503 and, in this example, processes the current and previous frames of time-frequency samples. In this example, the IID on the windowed frame is determined (e.g., using a method similar to that of a conventional PS encoder containing a symmetrical Hanning window). This IID value serves as a baseline for predicting the perceptual effect of transients detected later in processing and will be determined by... express.

[0120] Positive baseline value Left shift above the indicator frame, values ​​close to zero indicate center shift (no shift), and negative baseline IID indicates right shift.

[0121] Corresponding to a short sliding window of approximately 10ms Used to calculate the IID between the left and right channels:

[0122] Deviation of The values ​​can be considered transient candidates, and these values ​​can have the same characteristics as the baseline IID. The IID values ​​are encoded separately. It should be noted that the IID values ​​can be calculated for each frequency band at each QMF time, or aggregated across frequencies either globally or according to a custom binning scheme that may or may not omit certain frequencies that are irrelevant to the detection of certain transients.

[0123] Next, the sensing excitation step can evaluate the (PS) decoder's sensing performance on the detected transients. It can... set and The comparison is performed, and transients that are not affected by the baseline (traditional) IID reconstruction in the decoder are filtered out. This helps reduce the number of parameters that must be sent in the bitstream, thus keeping the resulting bit rate under control.

[0124] It is known that humans perceive slowly changing stereo parameters as moving sources, but can only detect rapidly changing parameters as an increase or decrease in the width of a stereo image.

[0125] The perceptual filtering step has two objectives: 1. Decide whether the IID parameters should be calculated and transmitted separately for a given transient. 2. Group transients together according to their IID properties, that is, group transients originating from the same source / location.

[0126] The primary objective is to perceive the stimulus and to determine the deviation between the transient IID parameters and the overall estimated frame parameters. If this deviation exceeds a certain threshold, the transient is included as a component of the stereo transient parameters.

[0127] The second objective could be to further reduce the bit rate, but by binding transients to the stereo properties of transients and assigning them to virtual sources (objects) in the stereo image. Thus, if the IID value for a given source is stable over time, only timing information, rather than both timing and IID characteristics, must be transmitted for the same source.

[0128] Figure 6 The diagram illustrates the relationship with Figure 3 and Figure 4 The same stereo transient representation is shown, but the average IID of the frame is also shown. ) and surrounding The IID range of the given average IID is specified. In this case, transients falling outside the indicated range can be represented by parameters included in the audio data signal.

[0129] Tuning can be based on perception (listening) tests or models. The value of this test or model can, for example, determine the informable difference (JND) between the transient frame IID and the average frame IID.

[0130] However, since JND is also known to be a function of the IID level—generally, the larger the IID, the larger the JND— The area between can be utilized Replace the exclusion area around the horizontal line itself, such as Figure 7 As shown. If the average IID of the frame falls within this range, the transient can be ignored and will not be encoded (as in the transient). (in the case of)

[0131] The perceptual filtering step may also include a masking model. Similar to excluded regions, for example, a masking model can indicate to the user that a given transient cannot be fully perceived based on background noise.

[0132] In this example, short-term properties are compared with long-term properties accordingly, and transients are detected based on these sufficient differences.

[0133] Many suitable transient detection methods are based on tracking the changes in the signal envelope (wideband or per spectrum) relative to slowly changing or rapidly changing residuals to detect the start (and end (offset)) of the transient.

[0134] In another method, the amplitude of the signal (e.g., a channel signal of a first multi-channel audio signal or a downmixed audio signal) is determined in a time-frequency representation, and the resulting frequency envelope is summed across the frequencies. Two smoothed versions of this envelope can then be created, one tracking the envelope more slowly than the other. Thus, the slowly changing residual envelope value is determined, and the other value is determined using the time constant of the faster-tracking envelope. A first-order exponential smoothing or, for example, a smoothed moving average filter can be applied to create the smoothed envelope. In the case of a fast-tracking envelope, an instantaneous envelope can also be used. in, and It corresponds to and The slow envelope and the fast envelope. Then, the ratio between the fast tracking envelope and the slow tracking envelope can be used to indicate the abrupt change in the transient relative to the residual signal, and is used as a time-domain (wideband) gain function, where , It is transient and Indication of transient existence. Such a method is described, for example, in Adami, A., Herzog, A., Disch, S., and Herre, J., (October 2017) "Transient-to-noise ratio restoration of coded applause-like signals" (2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) (pp. 349-353). IEEE).

[0135] This metric can be used to detect transients and the timing of these transients. Specifically, if g(n) exceeds a threshold, a transient can be considered to have been detected. In some cases, the metric g(n) can then be compared to a second threshold (which can specifically be the same as the first threshold), and if it falls below the second threshold, the end of the transient can be considered to have been detected. Thus, this method can detect both the beginning and end of a transient, and therefore also its duration.

[0136] The interval between the start and end of a transient in the first multi-channel audio signal and / or downmixed audio signal can also be used to filter out non-transient components with longer durations, and thus transient detection can detect only relatively short transients rather than step changes with longer durations.

[0137] In some embodiments, the encoder audio device may, in certain cases, divide the downmixed audio signal into a set of transients and a residual signal after removing these transients. For example, a portion of the downmixed audio signal between the detection of the start of a transition and the detection of the end of a transition may be extracted and represented as a separate transient, wherein the resulting downmixed audio signal represents the residual signal.

[0138] In some cases, a softer separation of transient and residual signals can be performed by using weighted selection instead of binary selection. For example, the transient signal t(n) can be generated by multiplying the first multi-channel audio signal and / or the downmixed audio signal by the detection signal g(n), i.e., Where m(n) represents, for example, the channel signal of the first multi-channel audio signal or the downmixed audio signal.

[0139] Similarly, residual signals can be generated, for example: .

[0140] Figure 8 The diagram illustrates the stereo input signal. Divided into transient signals and residual signals Examples.

[0141] Another approach to separating transients is to use minimum tracking of the envelope for each frequency band (based on minimum statistical methods for tracking stationary noise, such as those described in Martin, Rainer. “Noise power spectral density estimation based on optimal smoothing and minimum statistics.” (IEEE Transactions on Speech and Audio Processing 9.5 (2001): 504-512)) to track the residual signal (r(n)). In this case, the transient signal can be written in the frequency domain as follows: , in, It is a frequency domain estimation using the residual signal with minimum tracking, and It is the oversubtraction factor used to explain the underestimation of the residual signal. ).

[0142] As another example, source separation techniques employing neural networks can be used. Examples of such techniques can be found in Daniel Stoller, Sebastian Ewert, Simon Dixon, “Wave-U-Net: A Multi-Scale Neural Network for End-to-End Audio Source Separation” (http: / / arxiv.org / abs / 1806.03185, 2018).

[0143] Transient detection can be performed using a neural network model that has been trained to detect the location of transients using various neural network implementations, such as fully connected layers, convolutional layers, or recursive layers. Figure 9 The diagram below illustrates a neural network and its training.

[0144] In this example, the input corresponds to the current frame or a combination of the current (F) and previous (F-1) frames, where a time block from the previous frame can be used as additional padding if a transient occurs at the beginning of the current block F. Furthermore, the samples can correspond to all time-frequency samples or a range of frequency samples relevant to transient detection (e.g., for clapping, most of the energy is between 1 kHz and 3 kHz). Assuming a 2×f×n block of stereo data is used as input, the training data can include 2×f×n input blocks, and labels corresponding to 1×n vectors of 1 or 0 indicate the transient locations in the training data. Additionally, a transient presence flag indicating whether a frame includes at least a single transient can be added to the training labels, resulting in a 1x(n+1) vector label.

[0145] The loss function used can correspond to aggregate cross-entropy loss (assuming a frame length of n samples): in, The baseline truth label corresponding to the i-th time step of a given frame, and It is the probability predicted by the neural network that a transient exists at time i.

[0146] Alternative loss functions can be based on the mean squared error between the true baseline and the predicted label.

[0147] Manually annotating the applause signal using foreground clapping is possible, although time-consuming. Therefore, for the purposes of this invention, several available transient synthesizer models (e.g., clapping / applause) can be used to train the synthesized data. These models can generate realistic clapping sound signals using signal processing techniques. During inference, the model estimates the location of the detected stereo transient, which can then be used to calculate the corresponding IID value.

[0148] In the described method, the encoder audio device is thus arranged to generate an audio data signal that includes not only a set of downmixed and associated upmixed data, but also transient parameters reflecting the inter-channel level differences of the transient audio components of the multi-channel audio signal. The decoder audio device uses this data in the upmixing to generate an output multi-channel audio signal by employing two different upmixers and generating two different upmixed signals that are then combined.

[0149] The first upmixer 103 downmixes the downmixed audio signal based on the upmixing parameter set, and can specifically apply matrix multiplication to samples of the single-channel downmixed audio signal and the decorrelated signal, such as those known from traditional PS upmixing, for example, for a single-channel downmixed signal that has been upmixed into a stereo signal. Where m represents the single-channel downmixed audio signal and d represents the decorrelation signal.

[0150] This upmixing process typically operates in a time- and frequency-dependent manner, where the upmixing parameters are time- and frequency-dependent. For example, the upmixing coefficient h is usually determined for each time period / frame and each subband. xy The set of upmixing parameters. The upmixing coefficients are determined based on the received set of upmixing parameters. The exact dependence of the upmixing coefficients on the set of upmixing parameters will depend on the specific implementation. For example, in many implementations, the relationships defined for a conventional PS may be used as defined in ISO / IEC 23003-3:2020 Information technology — MPEG audio technologies — Part 3: Unified speech and audio coding.

[0151] The second upmixer 105 is arranged to generate upmixed multichannel transient audio components for each transient of the audio data signal, including transient parameters.

[0152] In some embodiments, the second upmixer 105 may be arranged to operate similarly to the first upmixer 105, wherein it may also use matrix multiplication of the downmixed audio signal and possibly one or more auxiliary signals to generate upmixed multichannel transient audio components. However, the coefficients are determined not based on the set of upmixing parameters, but based on the transient parameter coefficients.

[0153] For example, similar to the method used in conventional PS upmixing, upmixed multi-channel transient audio components can be generated using a 2×2 matrix multiplication of a single-channel downmixed audio signal and a decorrelation signal. In practice, in some embodiments, the same upmixing method as the first upmixer 105 can be used, but the upmixing coefficients are determined based on specific transient parameters provided for the transient audio components. In some embodiments, even IID, IPD, and ICC can be specified for the transient data of a given transient audio component, and coefficients for transient upmixing can be determined using formulas similar to those used in the first upmixer 103. However, instead of applying such upmixing to the entire downmixed audio signal, the transient upmixing operation is applied only to the corresponding transient audio components.

[0154] In some embodiments, transient audio components can be defined, for example, as the time interval of the downmixed audio signal in which a transient has been detected. Therefore, in some embodiments, the second upmixer 105 can be arranged to specifically upmix the time interval (and typically a very short time interval) of the downmixed audio signal using dedicated upmixing coefficients determined according to dedicated transient parameters. Thus, the upmixing specifically reflects the transient nature, rather than the nature of the entire audio segment for which it provides the set of upmixing parameters.

[0155] Therefore, even in such cases, a separate and dedicated upmixing process is used to generate upmixed multichannel transient audio components that specifically reflect and reproduce the transient properties of the first multichannel audio signal.

[0156] However, in most embodiments, the type of transient parameters also differs from those used in conventional upmixing for the first upmixer 103. Specifically, transient parameters typically provide less information than the set of upmixing parameters, where they are generally less frequency-dependent and typically represent fewer inter-channel signal properties.

[0157] In most embodiments, transient parameters can have a coarse frequency dependence, and in many embodiments, they may have no frequency dependence at all. Typically, the same parameter value is provided for the entire audio band. Therefore, a single value can be provided for the entire frequency range, rather than separate values ​​for each frequency subband (as is often the case with upmixing sets of parameters). Thus, the data overhead of transient data can be substantially reduced compared to the data overhead of upmixing sets of parameters.

[0158] In many embodiments, the type of information / parameters provided for transient upmixing may also be reduced relative to the set of upmixing parameters. In fact, in many embodiments, the transient parameters for a given transient audio component may not include any correlation / coherence and / or phase difference information for the channels of a multi-channel audio signal. In fact, in many embodiments, the only relative inter-channel property represented by the transient parameters(s) for a given transient audio component may be the inter-channel level difference. In many embodiments, for a given transient audio component, the transient parameters(s) may include an indication of the inter-channel level difference(s) and optionally, the timing properties of the transient audio component. For example, for a given transient audio component, a single full-band IID value and an indication of the timing (and typically the duration) of the given transient audio component may be provided.

[0159] In this case, the upmixing performed by the second upmixer 105 can simply scale the transient audio components to different channels of the upmixed multichannel transient audio components, such that the resulting upmixed multichannel transient audio components have an inter-channel level difference that matches the inter-channel level difference indicated by the transient parameters.

[0160] In many embodiments, the second upmixer 105 may be arranged to perform a translation of at least one transient audio component between two channels of the upmixed multichannel transient audio component according to transient upmixing parameters.

[0161] In a translation operation, individual channel signals of the upmixed multi-channel transient audio components can be generated by scaling the transient audio components with weights, where the weights, or different channels, depend on the level difference between channels. Therefore, in such an operation, the channel signals are scaled copies of the transient audio components, but the scaling depends on the transient parameters. Effectively, such an operation can generate upmixed multi-channel transient audio components that directly correspond to the transient audio components but are spatially positioned (by relative scaling) to have the desired location as indicated by the transient parameters.

[0162] For example, in the case of stereo, transient audio components in the form of defined segments of a single-channel downmixed audio signal can be shifted between the left and right signals, and are thus located in the stereo image by determining the scaling / weighting of the left and right signals to correspond to the inter-channel level difference indicated by the transient parameters.

[0163] As a specific example, the second upmixer 105 can be arranged to perform simple matrix multiplication on samples of a single-channel downmixed audio signal, for example: Here, h1 and h2 are scalar values ​​determined based on the inter-channel level difference indicated by transient parameters (e.g., as IID), and in many cases may depend solely on the inter-channel level difference.

[0164] As a specific example, the second upmixer 105 can perform the following translation operation: Wherein, the iid value represents the linear inter-channel intensity difference between the left and right transient signals, and t is the transient signal component of the single-channel downmixed audio signal.

[0165] This method leverages the inventors' understanding that transients are typically well-localized spatially compared to diffuse background signals. Furthermore, due to the usually short duration of transients, it is generally unnecessary to reconstruct a complete (diffuse) spatial image of the transient. Instead, translating the transient signal to the correct location is usually sufficient.

[0166] Therefore, in many embodiments, transient parameters can provide substantially less inter-channel information than the set of upmixing parameters, and thus the overhead of providing transient data can be very low, and is generally insignificant compared to the overhead imposed by the set of upmixing parameters. Furthermore, it has been found that, despite the low overhead, significantly improved audio experience and perceived quality can be achieved by upmixing transient audio components in parallel and separately based on the provided transient data.

[0167] It has been found that the described method can be used to achieve particularly advantageous generation of upmixed multichannel signals in many scenarios and for many signals (and embodiments). It has been found that including transient data and information allows the decoder audio device to consider information that might otherwise be lost (due to insufficient representation by the upmixing parameters). Furthermore, it has been found that representing transient information using a time-frequency resolution different from that used for the upmixing parameters allows for significantly improved upmixing and audio quality, while introducing only small overhead. In particular, it has been found that even a very coarse frequency resolution, including frequency dependence without transient parameter values, can still allow for very accurate and significantly improved upmixing including transient components. Moreover, it has been found that different time resolutions (including those particularly allowing finer time resolutions than the segments / time intervals used for the upmixing parameters) allow for improved audio quality and, in particular, allow for better representation of some audio components and sounds.

[0168] Furthermore, the inventors have recognized that such improved audio quality and user experience can be achieved compared to conventional upmixing, while providing only reduced information. In particular, they have recognized that substantial improvements can often be achieved by providing only information about the inter-channel horizontal differences, without having to provide information about the inter-channel coherence or phase differences.

[0169] To illustrate these considerations, consider an audio signal representing the applause of a group of people. Such an applause signal tends to consist of a seemingly random (spatial) superposition of individual clapping, i.e., short bursts of energy in time. On the one hand, estimating the set of stereo parameters and interpolating these between frames does not result in an accurate reconstruction of the stereo signal at the decoder. On the other hand, fine-grained estimation of the parameters can significantly increase the total bit rate.

[0170] To further understand the impact of a low (per frame) parameter update rate, we can consider Figure 10 The image illustrates an example of the time-frequency decomposition of a clapping stereo signal sampled at 44.1 kHz (spectral graph). The signal comprises background and foreground clapping, with the background clapping dominating below 3 kHz. The foreground clapping is noticeably shifted to the left, as it does not appear dominant in the right channel. The time-frequency energy of the background clapping is fairly random and resembles noise.

[0171] Figure 11 The output of a conventional parametric stereo decoder for the same segment is shown. It should be noted that this approach results in smeared background applause and foreground clapping. Conventional methods of stereo parametric interpolation are not very effective here. For the background applause, the smearing effect results in a musical pitch-like effect while clearly generating harmonic components. The foreground clapping is also slightly blurred temporally and loses its translational characteristics (IID): some foreground clapping now appears more dominant in the right channel. The latter effect is a result of the frame-by-frame update rate of the stereo parameter estimator.

[0172] To better understand the impact on stereo parameters, consider Figure 12 The left channel spectrograms of two data frames are depicted. The foreground clapping 1201 appears in the middle of the previous frame. Stereo parameters are estimated by first windowing both frames to allow energy attenuation near the beginning of the previous frame and the end of the current frame.

[0173] If IID, IPD, and ICC are calculated over these two frames, the distribution of phase, intensity, and coherence of the background applause from the left and right channels will largely determine the estimated stereo parameters. Even for the foreground clapping, the parameters should be different because, at least for IID, it is clear that the signal is shifted to the left.

[0174] Therefore, providing transient data with higher temporal resolution allows additional information to be included in the upmix, thereby enabling the generation of an improved output audio signal. In the described method, such information is used to perform additional and separate upmixing of the transient parameters of one or more transient audio components to generate upmixed multichannel transient audio components that are combined with a more conventionally generated upmixed signal.

[0175] This method allows transient information to be specifically adapted to the particular importance of the transient information. In fact, for transients, it has been found that upmixing is less sensitive to frequency dependence than to timing accuracy, and in the described method, transient data can be adjusted / generated accordingly, and is not limited to following the same resolution as the upmixing data.

[0176] Furthermore, in many embodiments, in addition to timing information indicating the timing of transients, the inter-channel horizontal difference may be significant, and for example, allows for an accurate representation of the spatial location of transients (e.g., in a stereoscopic image).

[0177] In many embodiments, the only inter-channel information provided for transients may be an inter-channel level difference indication. Specifically, in many embodiments, transient data may not include any inter-channel phase difference or inter-channel coherence. In practice, such information provides substantially irrelevant information and typically has a significantly smaller impact on the resulting audio quality of the generated output multichannel signal.

[0178] The second upmixer 105 is arranged to upmix the transient audio components based on transient parameter values ​​provided for the transient audio components. The exact signal components that are upmixed to generate the upmixed multi-channel transient audio components will depend on the individual scenario and implementation.

[0179] In many embodiments, transient audio components can be determined by a decoder audio device based on parameter data provided by an encoder audio device, and specifically, transient audio components can be generated as specific time intervals of a downmixed audio signal indicated by received transient timing values. For example, transient detection at the encoder audio device can detect the occurrence of a transient and determine the start and end times of the corresponding transient time interval. These times can be passed to a decoder audio device, which can be arranged to generate transient audio components by extracting segments from the received downmixed audio signal corresponding to the indicated time intervals. This transient audio component (i.e., the identified time interval of the downmixed audio signal) can then be upmixed based on transient parameters to generate corresponding upmixed multichannel transient audio components. A fully first upmixer 103 can upmix the fully downmixed audio signal to generate an upmixed multichannel signal, which is then combined with the generated upmixed multichannel transient audio components in a generator 107. In many embodiments, this combination can simply add the two multichannel signals / components together, potentially utilizing scaling / weighting of each signal.

[0180] In many embodiments, the downmixed audio signal is used to perform transient upmixing and is also used for upmixing by the first upmixer 103. In such a case, two contributions can be generated for a given portion of the downmixed audio signal corresponding to the transient; that is, in the previous example, the indicated segment of the downmixed audio signal is upmixed based on the set of upmixing parameters (by the first upmixer 103) and based on transient parameters (by the second upmixer 105). This may lead to some degradation in some cases, but may be acceptable (or even desirable) in others. In some embodiments, the parameters may also be adapted to include compensation for such combined upmixing. For example, the inter-channel level difference provided for the transient may not directly indicate the actual level difference, but may indicate a relative level difference relative to the level difference indicated by the set of upmixing parameters including the transient segment. Specifically, the IID value provided for the transient may be relative to the IID value of the corresponding set of upmixing parameters, and thus include both a contribution that eliminates the contribution from the first upmixer 103 and a contribution representing the desired level difference of the transient.

[0181] However, in other embodiments, the first upmixer 103 may be performed on the residual downmixed audio signal generated by extracting transient audio components (or at least one transient audio component) from the downmix. For example, in the previous example, the residual downmixed audio signal may simply be generated when time intervals indicating transients are removed from the downmixed audio signal. In such a case, the first upmixer 103 may then upmix the downmixed audio signal based on a set of upmixing parameters other than the transient segments upmixed by the second upmixer 105 based on transient parameters.

[0182] As another example, the separation of the downmixed audio signal into transient signals / transient audio components and residual downmixed audio signals can be performed by analyzing the downmixed audio signal (or analyzing the first multi-channel audio signal at the encoder audio device), as previously described in conjunction with transient detection. For example, the transient signal can be determined as: And the residual signal is: . Where m(n) represents, for example, a channel signal of a first multi-channel audio signal or a downmixed audio signal, and g(n) is the detection function described previously.

[0183] In some embodiments, the decoder audio device may be arranged to perform transient detection on the downmixed audio signal to detect transients. In some embodiments, the previous description of transient detection at the encoder audio device may be applied equivalently (with necessary modifications) to the decoder audio device, and specifically to the receiver 101 of the decoder audio device capable of identifying transients in the downmixed audio signal.

[0184] The second upmixer 105 can then be configured to upmix the identified portion of the downmixed audio signal using the received transient parameters.

[0185] In many embodiments, receiver 101 can be arranged to generate transient signals / transient audio components as described above, as well as residual downmixed audio signals. Therefore, as described above, the decoder audio device can be arranged to receive downmixed audio signals (such as single-channel downmixed audio signals) and perform transient detection to generate multiple transient audio components and residual downmixed audio signals, wherein these transient audio components are extracted. The transient audio components are then upmixed by a second upmixer 105 based on the received transient parameters, and the residual downmixed audio signals are upmixed by a first upmixer 103, wherein the resulting upmixed signals are combined by generator 107 to generate an output multi-channel audio signal.

[0186] In such embodiments, the decoder audio device itself can detect transients from transient data provided by the encoder audio device. The link between received transient parameters and detected transients can be accomplished in different ways. For example, the same transient detection can be performed in both the encoder and decoder based on the same signal (e.g., transient detection at the encoder audio device can also be based on the downmixed audio signal (e.g., after the encoding and decoding processes performed at the matching decoder audio device)). In such cases, a strong correspondence will exist between the detected transients, and simple assignment can be performed (e.g., assigning a first received transient parameter to a first detected transient, assigning a second received transient parameter to a second detected transient, etc.). In some cases, the audio data signal can, for example, indicate the number of transient parameters for each set of upmixed parameter time periods, and the receiver 101 can be arranged to detect the corresponding number of transients and assign transient parameters provided for the segment. Therefore, in many embodiments, the decoder audio device also includes functionality for generating one or more transient audio components based on the received downmixed audio signal.

[0187] Providing a fully downmixed audio signal that allows the decoder audio device to detect and extract transients offers several advantages. It typically reduces the data rate, and often significantly so. For example, it can also promote backward compatibility, as full downmixing can be used with upmixing performed by conventional decoders that do not include a dedicated transient upmixing path.

[0188] In some embodiments, receiver 101 may also be configured to extract one or more transient audio components based on provided transient parameters. As previously indicated, if the transient parameters include a timing indication, the extraction can be performed by extracting a segment of the downmixed audio signal corresponding to the indicated time interval.

[0189] In some embodiments, the encoder audio device may be arranged to generate an audio data signal to include audio data for one or more transient audio components. For example, dedicated audio data describing the transients may be provided. In such a case, the receiver 301 of the decoder audio device may also be arranged to recreate a local copy of the transient audio components based on the received audio signal and upmix them in a second upmixer 105 using the received transient parameters.

[0190] Such methods can increase the required data rate and bandwidth in many cases, but usually by only a relatively small amount. However, they can often allow for a significant reduction in complexity and computational resources at the decoder audio device, which is a significant benefit in many practical applications. Furthermore, they can allow for improved audio quality because transients can be represented more accurately.

[0191] In many embodiments, the encoder downmixed audio signal can still provide a fully downmixed audio signal including transients. However, in other embodiments, the encoder can be arranged to generate a residual downmixed audio signal and transmit it as a downmixed audio signal. In such cases, the decoder audio device can simply decode the residual downmixed audio signal and feed it to the first upmixer 103 for upmixing, and decode the transient audio components and feed these to the second upmixer 105 for upmixing. Therefore, the decoder audio device needs to operate efficiently with relatively low complexity and resource requirements.

[0192] Therefore, in some embodiments, the transient detector 205 may be arranged to remove one or more transient audio components from the downmixed audio signal before the downmixed audio signal is included in the audio data signal. The transient detector 205 may, for example, perform the operations described above to generate the transient signal and the residual downmixed signal based on the transient detection function g(n).

[0193] As a specific example, in some embodiments, the decoder audio device may specifically be a parametric stereo decoder that receives a single-channel downmixed audio signal and an upmixed parameter set including parameterized stereo parameters. The single-channel downmixed audio signal may be separated into a transient signal t and a residual signal r, wherein the transient signal is upmixed into a stereo signal (l) by a second upmixer 105. t r t The residual signal is upmixed into a stereo signal by the first upmixer 103. r r r Then, the output stereo pair (l', r') is constructed based on the sum of the stereo pair signals.

[0194] The residual upmixer (first upmixer 103) will include a decorrelation module to generate a signal substantially decorrelated with the residual signal r, and the residual stereo pair consists of a mixture of the residual signal r and its decorrelation signal, which is controlled by a set of PS parameters.

[0195] However, in a specific example, the transient upmixer (second upmixer 105) includes only a translation module, which is controlled by a set of PS parameters to translate the transient signal to the stereo pair (l t r t ).

[0196] Figure 13 The image shows a specific example of the corresponding encoder audio device. This encoder can generate a downmixed signal and estimate transient (foreground) and residual (background) parameters. For both the input left and right signals (l, r), the signals are separated into a left transient (t) by a transient separator 1301. l ), left residual (r) l ), right transient (t)r ) and right residual (r r The left and right transient signals are fed to transient parameter estimator 1303 to generate transient parameters. The left and right residual signals are fed to residual parameter estimator 1305 to generate residual parameters. Residual parameter estimator 1305 can be a conventional PS parameter estimator that estimates IID, ICC, and IPD values. Transient parameter estimator 1303 can be a simplified PS parameter estimator that only estimates the intensity difference (translation / IID) between the left and right transient signals. The transient signals, residual signals, and transient and residual parameters are fed to downmixing module 1307, which downmixes the input signal. Furthermore, the transient and residual parameters can be used to normalize the power of the downmixed signal relative to the input signal. The resulting downmixed signal and transient and residual parameters are also encoded ( Figure 13 (Not shown in the image) and transmitted to the decoder.

[0197] In some embodiments, to ensure similarity to the decoder's transient separation, the encoder may include a conventional PS encoder that generates downmixing and then runs the same transient separation module as the decoder to generate single-channel transient and residual signals. The transient signal can then be used to control the detection of stereo transients in the encoder, since it is clear that the decoder will generate the transient signal based on the single-channel signal.

[0198] Since not all frames can contain transient information, it can be beneficial to signal the decoder whether a frame contains transient information. In this case, the overhead remains low on average because transient parameters do not need to be transmitted for every frame, and the decoder's transient separation only needs to operate when this flag is present. Instead of this binary signaling, the signaling can also include the transmission of multiple transients within a frame, possibly entropy-coded to provide shorter word lengths for a lower number of transients. In this case, the transient parameter data can include individual parameters, such as an IID for each transient.

[0199] As previously mentioned, at least some of the processing typically takes place on the subband signal and is performed in the subband domain. Specifically, for a decoder audio device, the downmixed audio signal can be converted to a subband representation (or received in subband representation), and all processing can take place in the subband domain until the left and right output signals are generated and converted to the time domain. Figure 14 An example of such a decoder audio device is illustrated.

[0200] In other embodiments, the downmixed audio signal can be received in the time domain, and all processing can be performed in the time domain.

[0201] In some embodiments, some of the processing can advantageously be performed on subband samples / signals in the subband domain, and other processing can advantageously be performed on time-domain samples / signals in the time domain. Specifically, advantageously, in many scenarios, the first upmixer 103 can be arranged to perform subband domain upmixing of the downmixed audio signal, while the second upmixer 105 can be arranged to perform time-domain upmixing of at least one transient audio component.

[0202] In some embodiments, the downmixed audio signal / residual downmixed audio signal fed to the first upmixer 103 may be a sub-band domain signal, and upmixing can be performed in the sub-band using a set of upmixing parameters provided for different sub-bands. Conversely, the transient audio component fed to the second upmixer 105 may be a time-domain signal, and upmixing of the time-domain sample can be performed using transient parameters that are also provided in the time domain and may not be time-dependent. Figure 15 An example of such a decoder audio device is illustrated.

[0203] Operating in different frequency bands is generally advantageous for the described processing. An efficient means of operation in the frequency domain can be, for example, subband processing using (hybrid) quadrature mirror filtering (QMF) groups. From an efficiency standpoint, all processing blocks operate in the same domain (e.g., Figure 14 As shown, this is generally beneficial. Therefore, the incoming single-channel signal m is processed, for example, by a mixing QMF group to produce a set of sub-band domain signals representing the single-channel signal m. Then, all subsequent processing (transient separation, upmixing) operates in the same domain. Finally, after the transient and residual stereo (sub-band domain) signals have been reconstructed and added, they are synthesized back into the time domain.

[0204] In some embodiments, it may be desirable to process transient signals, especially in the time domain, separately (e.g. Figure 15 (As shown). This has the advantage of even lower complexity in transient upmixing. Note that transient separation can still be performed (partially) in the frequency domain.

[0205] One or more audio devices can be specifically implemented in one or more appropriately programmed processors. In particular, artificial neural networks can be implemented in one or more such appropriately programmed processors. Different functional blocks, especially artificial neural networks, can be implemented in separate processors and / or can be implemented, for example, in the same processor. Examples of suitable processors are provided below.

[0206] Figure 16This is a block diagram illustrating an example processor 1600 according to an embodiment of the present disclosure. Processor 1600 can be used to implement one or more processors that implement the apparatus or elements thereof as previously described (particularly including one or more artificial neural networks). Processor 1600 can be any suitable processor type, including but not limited to microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs) where the FPGA has been programmed to form a processor, graphics processing units (GPUs), application-specific integrated circuits (ASICs) where the ASIC has been designed to form a processor or a combination thereof.

[0207] Processor 1600 may include one or more cores 1602. Core 1602 may include one or more arithmetic logic units (ALUs) 1604. In some embodiments, in addition to or in place of ALU 1604, core 1602 may include a floating-point logic unit (FPLU) 1606 and / or a digital signal processing unit (DSPU) 1608.

[0208] Processor 1600 may include one or more registers 1612 communicatively coupled to core 1602. Registers 1612 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 1612 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 1602.

[0209] In some embodiments, processor 1600 may include one or more levels of cache memory 1610 communicatively coupled to core 1602. Cache memory 1610 may provide computer-readable instructions to core 1602 for execution. Cache memory 1610 may provide data for core 1602 to process. In some embodiments, computer-readable instructions may have been provided to cache memory 1610 from local memory (e.g., local memory attached to external bus 1616). Cache memory 1610 may be implemented using any suitable cache memory type, such as metal-oxide-semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0210] Processor 1600 may include controller 1614, which can control inputs to processor 1600 from other processors and / or components included in the system and / or outputs from processor 1600 to other processors and / or components included in the system. Controller 1614 can control data paths in ALU 1604, FPLU 1606, and / or DSPU 1608. Controller 1614 may be implemented as one or more state machines, data paths, and / or dedicated control logic units. The gates of controller 1614 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0211] Register 1612 and cache 1610 can communicate with controller 1614 and core 1602 via internal connections 1620A, 1620B, 1620C and 1620D. Internal connections can be implemented as buses, multiplexers, cross switches and / or any other suitable connection technology.

[0212] Inputs and outputs to processor 1600 may be provided via bus 1616, which may include one or more wires. Bus 1616 may be communicatively coupled to one or more components of processor 1600, such as controller 1614, cache 1610, and / or register 1612. Bus 1616 may be coupled to one or more components of the system.

[0213] Bus 1616 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 1632. ROM 1632 may be a masked ROM, electronically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 1633. RAM 1633 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 1635. The external memory may include flash memory 1634. The external memory may include a magnetic storage device, such as a disk 1636. In some embodiments, the external memory may be included in the system.

[0214] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The invention can optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.

[0215] Although the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is limited only by the appended claims. Furthermore, although features may appear to be described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, terms include those that do not exclude the presence of other elements or steps.

[0216] Furthermore, although listed individually, multiple units, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features may be advantageously combined, and their inclusion in different claims does not imply that such a combination of features is infeasible and / or disadvantageous. Moreover, including a feature in a class of claims does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other claim classes. Furthermore, the order of features in a claim does not imply any particular order in which the features must be performed, and in particular, the order of individual steps in a method claim does not imply that the steps must be performed in that order. Rather, these steps may be performed in any suitable order. Furthermore, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided merely as clarifying examples and should not be construed as limiting the scope of the claims in any way.

Claims

1. An audio device for generating and outputting multi-channel signals, the audio device comprising: Receiver (101), which is arranged to receive audio data signals, said audio data signals including: The downmixed audio signal is the downmixing of the first multi-channel signal; For the set of upmixing parameters of the downmixed audio signal, each set of upmixing parameters includes at least: The level difference parameter indicates the level difference between the channels of the first multi-channel signal; Relevant parameters, which indicate the coherence between channels of the first multi-channel signal; and A phase difference parameter, which indicates the phase difference between the channels of the first multi-channel signal; and At least one transient upmixing parameter for at least one transient audio component of the first multi-channel signal, the transient upmixing parameter indicating the level difference between the at least one transient audio component between channels of the first multi-channel signal; A first upmixer (103) is arranged to upmix the downmixed audio signal and generate an upmixed multichannel signal according to the set of upmixing parameters; A second upmixer (105) is arranged to generate upmixed multichannel transient audio components by upmixing the at least one transient audio component according to the at least one transient upmixing parameter; and A generator (107) is arranged to generate the output multichannel signal based on a combination of the upmixed multichannel signal and the upmixed multichannel transient audio components.

2. The audio device according to claim 1, wherein, The receiver (101) is configured to extract a first transient audio component from the downmixed audio signal, and the second upmixer is configured to upmix the first transient audio component according to the transient upmixing parameters.

3. The audio device according to claim 2, wherein, The receiver (101) is arranged to generate a residual downmixed audio signal, which is generated by extracting a set of transient audio components from the downmixed audio signal, which includes audio data for the at least one transient audio component, and the first upmixer (103) is arranged to upmix the residual downmixed audio signal according to the set of upmixing parameters.

4. The apparatus according to claim 3, wherein, The first upmixer (103) is configured to decorrelate the residual downmixed audio signal to generate a decorrelated residual downmixed audio signal, and to generate the upmixed multichannel signal by upmixing the residual downmixed audio signal and the decorrelated residual downmixed audio signal according to the upmixing parameter set.

5. The audio device according to any of the preceding claims, wherein, The audio data signal includes audio data for the at least one transient audio component, and the second upmixer (105) is arranged to upmix the audio data for the at least one transient audio component according to the transient upmixing parameters.

6. The audio device according to any of the preceding claims, wherein, The second upmixer (105) is arranged to perform translation of at least one transient audio component between two channels of the upmixed multichannel transient audio component according to the transient upmixing parameters.

7. The audio device according to any of the preceding claims, wherein, The first upmixer (103) is arranged to perform sub-band domain upmixing of the downmixed audio signal; and the second upmixer (105) is arranged to perform time domain upmixing of the at least one transient audio component.

8. The audio device according to any of the preceding claims, wherein, The audio data signal includes a timing indication for the at least one transient audio component, and the combination depends on the timing indication.

9. The audio device according to any of the preceding claims, wherein, The frequency resolution for the at least one transient upmixing parameter is coarser than the frequency resolution of the set of upmixing parameters.

10. An audio apparatus for generating audio data signals, the audio apparatus comprising: A receiver (201) is configured to receive a first multichannel signal; Downmixer (203), configured to generate a single-channel downmixed audio signal based on the first multi-channel signal and determine an upmixing parameter set for the downmixed audio signal, each upmixing parameter set including at least: The level difference parameter indicates the level difference between the channels of the first multi-channel signal; Relevant parameters, which indicate the coherence between channels of the first multi-channel signal; and The phase difference parameter indicates the phase difference between the channels of the first multi-channel signal; A transient detector (205) is arranged to detect at least one transient audio component of the first multi-channel signal and generate at least one transient upmixing parameter for the at least one transient audio component of the first multi-channel signal, the at least one transient upmixing parameter indicating the level difference of the at least one transient audio component between channels of the first multi-channel signal; A data generator (207) is arranged to generate the audio data signal to include the single-channel downmixed audio signal, the set of upmixing parameters, and the transient upmixing parameters.

11. The audio device according to claim 10, wherein, The transient detector (205) is arranged to detect the at least one transient audio component by applying transient detection to the downmixed audio signal.

12. The audio device according to claim 10 or 11, wherein, The transient detector (205) is arranged to detect the at least one transient audio component by applying transient detection to the channels of the first multi-channel signal.

13. A method for generating an output multi-channel signal, the method comprising: Receive audio data signals, the audio data signals including: The downmixed audio signal is the downmixing of the first multi-channel signal; For the set of upmixing parameters of the downmixed audio signal, each set of upmixing parameters includes at least: The level difference parameter indicates the level difference between the channels of the first multi-channel signal; Relevant parameters, which indicate the coherence between channels of the first multi-channel signal; and A phase difference parameter, which indicates the phase difference between the channels of the first multi-channel signal; and At least one transient upmixing parameter for at least one transient audio component of the first multi-channel signal, the transient upmixing parameter indicating the level difference between the at least one transient audio component between channels of the first multi-channel signal; The upmixed audio signal is upmixed and an upmixed multichannel signal is generated based on the upmixing parameter set. Upmixed multichannel transient audio components are generated by upmixing the at least one transient audio component according to the at least one transient upmixing parameter; and The output multichannel signal is generated by combining the upmixed multichannel signal with the upmixed multichannel transient audio components.

14. A method for generating an output audio signal, the method comprising: Receive the first multi-channel signal; A single-channel downmixed audio signal is generated based on the first multi-channel signal, and a set of upmixing parameters for the downmixed audio signal is determined, wherein each set of upmixing parameters includes at least: The level difference parameter indicates the level difference between the channels of the first multi-channel signal; Relevant parameters, which indicate the coherence between channels of the first multi-channel signal; and The phase difference parameter indicates the phase difference between the channels of the first multi-channel signal; At least one transient audio component of the first multi-channel signal is detected and at least one transient upmixing parameter is generated for the at least one transient audio component of the first multi-channel signal, the at least one transient upmixing parameter indicating the level difference of the at least one transient audio component between channels of the first multi-channel signal; and The audio data signal is generated to include the single-channel downmixed audio signal, the upmixing parameter set, and the transient upmixing parameters.

15. A computer program product comprising a computer program code module, the computer program code module being adapted to perform all the steps of claim 13 or 14 when the program is run on a computer.