Multichannel audio encoder, decoder, method and computer program for switching between parametric multichannel operation and individual channel operation

By switching between parametric multichannel coding and single-channel coding, and utilizing inter-channel relationships and speaker interference detection, the problem of stereo image reproduction in multi-speaker scenarios is solved, thus improving the perception quality.

CN113874937BActive Publication Date: 2026-01-27FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080032830.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-04
Filing Date
2020-04-02
Publication Date
2026-01-27
Estimated Expiration
2040-04-02

Smart Images

  • Figure CN113874937B_ABST
    Figure CN113874937B_ABST
Patent Text Reader

Abstract

A multi-channel audio encoder (100) for providing an encoded audio representation (112) based on an input audio representation (110) is provided. The multi-channel audio encoder (100) is configured to switch (140) between a parametric multi-channel encoding (120) of a plurality of channels and a separate encoding (130) of the plurality of channels depending on a characteristic of the input audio representation (110).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to multichannel audio encoding and decoding for stereo, two-channel, or more-than-two-channel applications. More specifically, it relates to general audio encoding / decoding, or speech encoding / decoding, or encoding / decoding using transform-domain encoding / decoding with scaling factors and / or encoding / decoding based on linear prediction coefficients. Background Technology

[0002] For transmitting stereo speech signals captured by a microphone arrangement with two or more microphones (at a distance between them), parametric stereo techniques can be used when a low bit rate is required. An exemplary parametric stereo technique is described in [1]. Parametric stereo systems can perform adequately in most cases where there are two or more speakers around the microphone arrangement and more than one speaker is speaking simultaneously in the same time period. However, there are some cases where the parametric model may fail to reproduce the stereo image and deliver an intelligible speech output for scenarios with interfering speakers. This can occur, for example, when each of the two or more speakers is captured with a different ITD (inter-channel time difference), a large ITD value (large distance between microphones), and / or when the speakers are sitting in opposite positions around the microphone arrangement axis.

[0003] Furthermore, in the parametric stereo scheme described in [1], some parameters are extracted to reproduce the spatial stereo scene, and the stereo signal is derived into a mono-channel downmixed signal that is further encoded. In the case of interfering speakers, the downmixed signal can be encoded using a speech encoder such as CELP described in [2]. However, such an encoding scheme is a source filter model of speech generation, designed to represent the speech of a single speaker. For interfering speakers, this may violate the core encoding model and result in a decrease in perceptual quality.

[0004] The purpose of this invention is to overcome, at least in part, the shortcomings of conventional methods. Summary of the Invention

[0005] This objective is achieved by a multi-channel audio encoder according to claim 1, a multi-channel audio decoder according to claim 26, an encoded multi-channel audio representation according to claim 26, a method for multi-channel audio encoding according to claim 30, a method for multi-channel audio decoding according to claim 31, and a computer program according to claim 32.

[0006] A multichannel audio encoder is provided. The multichannel audio encoder can be stereo, or a two-channel or more-than-two-channel audio encoder. The audio encoder can be a general audio encoder, a speech encoder, or an encoder that switches between transform-domain encoding using scaling factors and encoding based on linear prediction coefficients. The encoder is configured to provide an encoded audio representation based on an input audio representation. The encoder is configured to switch between parametric multichannel encoding of multiple channels (e.g., multiple channels of the input audio representation) and individual encoding of multiple channels (e.g., multiple channels of the input audio representation) according to the characteristics of the input audio representation.

[0007] Parametric multichannel coding can encode a combined signal that incorporates signals from multiple channels, and encode the relationships between two or more channels in the form of parameters. Parameters may include inter-channel time difference parameters, and / or inter-channel level difference parameters, and / or inter-channel phase parameters and / or inter-channel correlation parameters.

[0008] Switching between parametric multichannel coding and individual coding, based on the characteristics of the input audio representation, advantageously allows for the adaptation of coding to the characteristics of the input audio representation. The selective switching between parametric multichannel coding and individual coding can lead to the selection of a coding more suitable for encoding the potential input audio representation, resulting in an encoded audio representation that possesses advantageous properties regarding, for example, perceptual performance.

[0009] In other words, the present invention relates to a trade-off between the effort required to obtain the characteristics of an input audio representation and then take action (e.g., switching) based on those characteristics, and the benefit of encoding the input audio representation using, for example, encoding that may be advantageous to a particular input audio representation (or a portion thereof) in terms of performance standards.

[0010] According to an embodiment, a multichannel encoder can be configured to determine whether the input audio representation satisfies the assumptions of a model under parametric multichannel coding, and to switch accordingly. Assumptions may include the presence of a single loudspeaker, for example, the presence of a single significant interchannel time difference / interaural time difference (ITD) in each time-frequency portion. For example, characteristics of the input audio representation may provide an indication of interference from two or more speakers, thus potentially violating the assumption of a single loudspeaker in the model under parametric multichannel coding.

[0011] According to an embodiment, a multichannel encoder can be configured to switch to individual encoding if the assumptions of the model under parametric multichannel encoding are not satisfied. For example, for some input audio representations, the assumptions of the model under parametric multichannel encoding regarding the number of speakers and the ITD / multiple ITDs of these speakers may not be satisfied. However, the assumptions of the model under individual encoding can be satisfied. Therefore, switching to individual encoding can lead to advantageous performance.

[0012] According to an embodiment, a multichannel encoder can be configured to determine whether an input audio representation corresponds to a dominant source (e.g., a single dominant source). In this case, other sources (e.g., all other sources) may be weaker, for example, differing by at least a predetermined intensity difference. The encoder can be configured to switch based on this determination. The presence or absence of a dominant source can provide an indication of whether parametric encoding or individual encoding may be advantageous in terms of performance.

[0013] According to embodiments, a multichannel encoder can be configured to determine whether a single dominant source exists in multiple time-frequency portions, and / or to determine whether two or more sources exist in a given time-frequency portion, wherein the multichannel coding parameters of the two or more sources differ by at least a predetermined deviation or more than a predetermined deviation. The multichannel encoder can be configured to switch based on this determination. The multiple time-frequency portions may alternatively include all time-frequency portions. The two or more sources may satisfy a source saliency condition, for example, being relevant and / or significant and / or noteworthy sources located at different positions. The multichannel coding parameters may be ITDs. Determining a single source may allow selection of coding under which a model is suitable for processing the single source, for example, parametric coding. Determining a single source in one or more time-frequency portions may allow selection of coding for one or more portions that satisfy assumptions about the model under the coding (e.g., a parametric model). Determining two or more sources in a given time-frequency portion may indicate that coding with a potential model based on a single source may not provide the expected performance for the time-frequency portion, and therefore switching the coding of the given portion may result in favorable performance. Determining whether multichannel parameters differ by at least a predetermined deviation (or more) can allow for the determination of whether two or more sources may cause the assumptions of the model under the encoding to be violated, and thus can be an indication to switch to a different encoding.

[0014] In an embodiment, a multichannel encoder can be configured to determine the parameters of a model under parameterized multichannel encoding and to switch based on the model's parameters. For example, the model's parameters could be inter-channel time difference (ITD) or inter-ear time difference (ITD). The parameters can describe the relationship between two or more channels of the input audio representation. Determining the parameters of the model under parameterized multichannel encoding allows for the evaluation of the parameterized model's ability to deliver desired performance for a given relationship between two or more channels of the input audio representation, and the execution of switching to achieve favorable performance.

[0015] In embodiments, a multichannel encoder can be configured to determine whether a property defining the relationship between channels in an input audio representation allows for explicit determination of multichannel coding parameters, or indicates two or more distinct possible values ​​for the multichannel coding parameters, and to switch accordingly. For example, the property defining the relationship between channels could be the evolution of a generalized cross-correlation phase transform (GCC-PHAT) over a hysteresis parameter, or the evolution of a cross-correlation function between two or more channels over a hysteresis parameter. The multichannel coding parameter could be an ITD. Two or more distinct possible (e.g., meaningful) values ​​could differ by at least a predetermined value and could be distinguished from the noise floor. The property could include two or more values ​​(e.g., peak values, or values ​​satisfying a saliency condition) that differ by at most one (predetermined or signal-adaptive) difference regarding their significance, or the property could include only a single value satisfying a saliency condition. By using the evolution of the generalized cross-correlation phase transform or the evolution of the cross-correlation function to determine the relationship between channels in an input audio representation, the relationship between channels can be quantized to obtain the property. Determining whether two or more distinct values ​​of a multichannel coding parameter differ by at least a predetermined value and whether these two or more distinct values ​​can be distinguished from the background noise allows for a favorable and reliable determination of whether an explicit determination of the multichannel coding parameter is possible, or whether two or more distinct meaningful values ​​of the multichannel coding parameter can be determined. Alternatively or additionally, for example by using a saliency criterion to determine whether a characteristic includes two or more values ​​that differ by at most one difference in saliency with respect to their determined significance allows for a favorable and reliable determination of whether an explicit determination of the multichannel coding parameter is possible, or whether two or more distinct meaningful values ​​of the multichannel coding parameter can be determined.

[0016] In embodiments, a multichannel encoder can be configured to determine whether a characteristic defining the relationship between channels of an input audio representation includes only a single saliency value satisfying a saliency condition, or whether the characteristic defining the relationship between channels of an input audio representation includes two or more (e.g., different) saliency values ​​satisfying a saliency condition, and to switch between parameterized multichannel encoding and individual encoding based on this determination, for example. The characteristic defining the relationship between channels can be the evolution of GCC-PHAT on the hysteresis parameter, or the evolution of the cross-correlation function between two or more channels on the hysteresis parameter. A single saliency value can involve a single saliency peak, representing a single ITD value. A saliency condition can include the amplitude relationship between two or more local peaks or maxima, and / or the distance relationship between two local peaks or maxima, and / or the distance from the noise floor. A saliency condition can be predetermined or signal-adaptive, for example, it can be adaptive based on the characteristics of the input audio representation. Two or more saliency values ​​can include at least two saliency peaks, representing two or more different ITD values. The satisfaction of a saliency condition can be determined in a single time-frequency portion. Determining the relationships between channels in an input audio representation using the evolution of GCC-PHAT or cross-correlation functions advantageously allows for the quantization of these relationships to obtain characteristics. Determining whether a characteristic includes only a single saliency value or two or more values ​​advantageously allows for the determination of which coding method (e.g., parametric multichannel coding or individual coding) might be more suitable for a given input audio representation. Saliency conditions advantageously allow for the use of one or more criteria to evaluate values, such as the amplitude between two local peaks or maxima in the time domain (e.g., time lag), the distance between two local peaks or maxima, and / or the distance to the noise floor, in order to determine which values ​​can be considered evolutionarily when determining whether a characteristic includes only a single saliency value or two or more saliency values.

[0017] In an embodiment, the multi-channel encoder can be configured to determine (e.g., in the form of encoded audio representation) parameters of a previous frame and switch according to those parameters. The parameters of the previous frame may be a SAD flag. Determining the parameters of the previous frame can advantageously be used, for example, to determine whether the previous frame includes an active signal, allowing for selective avoidance of switching at the first frame of the signal portion.

[0018] In embodiments, a multichannel encoder can be configured to determine whether an interference source is present in the input audio representation and to switch based on that determination. The interference source may include two or more interfering sound sources, or two or more interfering speakers, or two or more interfering speakers. The interference source (or speaker or speaker) in the input audio representation may be determined, for example, in a time-frequency portion or, for example, in an overlapping time-frequency resource or portion. Determining the presence of an interference source can advantageously allow switching between parametric multichannel coding and individual coding, for example, based on the determination that the input audio representation includes an interference source. This can, for example, lead to a performance degradation in parametric multichannel coding and, for example, lead to the advantageous performance of individual coding.

[0019] In embodiments, a multichannel encoder can be configured to determine whether there exist two or more values ​​describing the relationship between two or more channels of an input audio representation, and to switch based on this determination, wherein the two or more values ​​satisfy a saliency condition and are associated with a single time-frequency component. The two or more values ​​may include correlated values ​​or saliency values. Determining whether there exist two or more values ​​that satisfy a saliency condition and are associated with a single time-frequency component can advantageously allow the determination of whether the input audio representation may, for example, lead to a performance degradation in parameterized multichannel coding, and, for example, lead to advantageous performance in individual coding.

[0020] In an embodiment, a multichannel encoder can be configured to determine whether two or more peaks exist in the cross-correlation (e.g., GCC-PHAT) between two or more channels of the input audio representation, and switch accordingly. The cross-correlation may be related to a given time-frequency portion. Determining whether two or more peaks exist in the cross-correlation between two or more channels can advantageously allow for a quantitative determination of whether there may be speaker interference in the input audio representation, which may degrade the performance of, for example, parameterized multichannel coding, and, after determination, switch to, for example, individual coding.

[0021] In embodiments, a multichannel encoder may include an estimator configured to estimate the relationship between two or more channels of an input audio representation based on cross-correlation. The estimator may be configured to estimate the relationship separately for multiple time-frequency portions. The estimator may be an ITD estimator. The cross-correlation may be GCC-PHAT or smoothed cross-correlation. The cross-correlation may be performed in the time domain or in the frequency domain. The multichannel encoder may also be configured to determine whether the difference between two peaks (e.g., correlation values ​​and / or significance values ​​estimated by the estimator) associated with different cross-correlation hysteresis is greater than a value (e.g., a predetermined value or a signal adaptive value), and to switch accordingly. The estimator (e.g., an ITD estimator) may exist within the encoder, such as an encoder using parameterized multichannel coding, so using an estimator to determine whether the difference between two peaks associated with different cross-correlation hysteresis is greater than a threshold may not introduce significant additional complexity.

[0022] In an embodiment, a multichannel encoder can be configured to determine whether the distance between two or more values ​​(e.g., correlation values ​​or saliency values) describing the relationship between two or more channels of an input audio representation is greater than a value (e.g., a predetermined value or a signal adaptation value), and to switch based on this determination, wherein the two or more values ​​satisfy a saliency condition and are associated with the same time-frequency component. The distance can be determined, for example, in the time domain relative to a time lag or cross-correlation lag. The two or more values ​​can be peaks of cross-correlation between two or more channels of the input audio representation and can be provided by an estimator (e.g., an ITD estimator). Peak values ​​can be values ​​that satisfy a saliency condition. Determining whether the distance between two or more values ​​that satisfy a saliency condition and are associated with the same time-frequency component is greater than a threshold allows an advantageous distinction between: for example, two or more peaks located at small distances that may be attributed to a single source, and two or more peaks located at significant (e.g., larger) distances that may be attributed to more than a single source.

[0023] In embodiments, a multichannel encoder can be configured to determine a first characteristic value based on the evolution of cross-correlation (e.g., on hysteresis parameters) and to switch based on this determination. The first characteristic value can be a dominant peak or a primary peak. The cross-correlation can include GCC-PHAT. The first characteristic value can satisfy a significance condition. The peak value can be the maximum (e.g., absolute) value in the evolution. The determination can include evaluating the evolution over one or more frames (including, for example, one or more previous frames). The determination can also include determining whether the value satisfies a stability condition. For example, if the value is within a range (e.g., a predetermined range or a signal-adaptive range) for multiple previous frames (e.g., a predetermined number of previous frames or a signal-adaptive number of previous frames), then the stability condition is satisfied. Alternatively or additionally, the satisfaction of the stability criterion can be determined based on a hysteresis mechanism that takes the value for multiple frames (e.g., a predetermined number of previous frames or a signal-adaptive number of previous frames) as input. Determining the first characteristic value (e.g., the dominant peak) can allow, individually or in combination with one or more other values, to advantageously evaluate whether the determined value (in many cases, the maximum value in the evolution of cross-correlation) causes a switching of encoding between parameterized multichannel encoding and individual encoding. In addition, optional consideration of significance and / or stability conditions can advantageously allow for determining, for example, whether to selectively avoid switching if the detected value is not stable enough over time and / or not far enough away from, for example, background noise.

[0024] In embodiments, a multichannel encoder can be configured to determine one or more dependent characteristic values ​​based on the evolution of cross-correlation, and to switch based on this determination. The one or more dependent characteristic values ​​can be secondary peaks or second peaks. Dependent values ​​can be determined based on a portion of the cross-correlation evolution. For example, the distance of each element of this portion to the first characteristic value (e.g., in the time domain, as relative to time lag) can exceed (e.g., a predetermined or signal-adaptive) threshold. One or more dependent characteristic values ​​can satisfy a significance condition. One or more dependent characteristic values ​​can be one or more maximum (e.g., absolute) values ​​in the portion of the evolution. One or more dependent characteristic values ​​can satisfy a stability condition. Determining one or more dependent characteristic values ​​can advantageously allow for the evaluation of whether the determined values ​​(e.g., the first characteristic value and / or one or more dependent characteristic values) cause a switching of encoding between parameterized multichannel encoding and individual encoding. Furthermore, optionally evaluating one or more dependent values ​​at a distance from the first characteristic value in a portion of the cross-correlation evolution can advantageously allow for the reliable attribution of the input audio representation to a single source or multiple sources. Alternatively or additionally, the multichannel encoder can be configured to determine the existence of one or more dependent feature values ​​based on the evolution of cross-correlation, and to switch accordingly. In other words, the existence of only one or more dependent feature values ​​can be determined, for example, based on a pattern recognition algorithm.

[0025] In embodiments, a multichannel encoder can be configured to determine that a primary peak and one or more subordinate peaks satisfy a significance criterion, and switch accordingly. For example, for multiple frames satisfying a stability criterion, a significance criterion is satisfied if the difference (e.g., relative difference) between the primary peak and one or more subordinate peaks is greater than a threshold (e.g., a predetermined threshold, or a signal adaptive threshold). For example, the difference between peaks can be determined relative to the amplitude of the peak, or relative to the phase of the peak, or relative to the time lag of the peak. Alternatively or additionally, the multichannel encoder can be configured to determine whether one or more subordinate peaks satisfy a correlation criterion in the cross-correlation, and switch accordingly. For example, the correlation criterion can be defined relative to the primary peak and / or relative to the background noise of the cross-correlation. Determining a significant difference between the primary peak and one or more subordinate peaks advantageously allows for reliable determination that more than one source exists in the input audio representation, and, for example, switching to separate encoding based on this determination.

[0026] In embodiments, a multichannel encoder can be configured to selectively consider a subordinate peak in a given frame of the input audio representation if one or more corresponding subordinate peaks already exist in one or more frames preceding the given frame. For example, the one or more corresponding subordinate peaks may be located at the same autocorrelation hysteresis as the considered subordinate peak, or within a predetermined range of autocorrelation hysteresis around the considered subordinate peak's autocorrelation hysteresis. Considering one or more corresponding subordinate peaks in one or more previous frames to selectively consider subordinate peaks in a given frame advantageously allows for the determination of whether certain spatial and / or level / phase / frequency stability can be attributed to one or more sources prior to the switching encoding. Stability can encompass one or more frames and therefore can be related to the environment of one or more sources, without being limited by frame length.

[0027] In embodiments, a multichannel encoder can be configured to determine whether one or more characteristic values ​​describing the relationship between two or more channels of an input audio representation satisfy a stability condition, and to switch based on this determination. The characteristic value can be a dominant peak and / or one or more subordinate peaks. For example, if the value is within a range (e.g., a predetermined range or a signal-adaptive range) or greater than a threshold (e.g., a predetermined threshold or a signal-adaptive threshold), the stability condition can be satisfied for multiple previous frames (e.g., a predetermined number of previous frames, or a signal-adaptive number of previous frames). Alternatively or additionally, the satisfaction of the stability condition can be determined based on a hysteresis mechanism that takes values ​​for multiple (e.g., a predetermined number of previous frames, or a signal-adaptive number of previous frames) frames as input. Determining that the stability condition is satisfied can advantageously allow for avoiding switching on noisy input audio representations or portions thereof (e.g., on noisy frames).

[0028] In embodiments, a multichannel encoder can be configured to determine whether a noise condition is met for multiple frames (e.g., a predetermined number of frames or a signal-adaptive number of frames), and selectively avoid switching if the noise condition is met. These frames may include the current frame. For example, a noise condition may be met if the noise characteristics (e.g., noise floor) of a frame (or multiple frames) are greater than a threshold (e.g., a predetermined threshold or a signal-adaptive threshold). Determining that the noise condition is met can advantageously allow avoiding switching on noisy input audio representations or portions thereof (e.g., on noisy frames).

[0029] In an embodiment, the multichannel encoder can be configured to determine, for multiple frames, whether a saliency condition and / or stability condition for a characteristic value are met, and to switch accordingly. The characteristic value can be a dominant peak and / or one or more subordinate peaks. The number of frames can be predetermined or signal-adaptive. These frames can include one or more previous frames and / or the current frame. Determining the satisfaction of the saliency condition and / or stability condition for multiple frames advantageously allows selective avoidance of switching on unstable signals (e.g., unstable and / or noisy portions of the input audio representation).

[0030] In an embodiment, a multichannel encoder can be configured to determine whether the distance of one or more subordinate peaks is within a predetermined range, and to switch and / or selectively avoid switching based on this determination. For example, one or more subordinate peaks may have a maximum value (e.g., a maximum absolute value) and may be referred to as peaks (2). The distance may be determined relative to a time lag (e.g., absolute time lag or relative time lag), and / or may be determined in the time domain or frequency domain. The distance may be determined for multiple frames (e.g., a predetermined number of frames or a signal-adaptive number of frames). These frames may include one or more previous frames and / or the current frame. Determining whether the distance of one or more peaks is within a predetermined range and switching and / or selectively avoiding switching based on this can advantageously allow selective avoidance of switching on unstable signals (e.g., unstable and / or noisy portions of the input audio representation).

[0031] In embodiments, the multichannel encoder can be configured to selectively avoid switching at or after a first frame following an inactive frame of the input audio representation. An inactive frame may include a noisy frame. Alternatively or additionally, the multichannel encoder can be configured to determine whether a given flag in a frame has changed relative to one or more previous frames, and selectively avoid switching based on that determination. For example, the flag may indicate an active signal and may be a SAD flag. Selectively avoiding switching may include avoiding switching at or after the first frame where the flag has an active value. Thus, it is advantageous to selectively avoid switching at the first frame of the signal portion.

[0032] In an embodiment, a multichannel encoder can be configured to selectively switch to individual encoding in response to detecting a change in a characteristic of the input audio representation greater than a threshold (e.g., a predetermined threshold or a signal adaptive threshold). The characteristic of the input audio representation can be, for example, ITD, or a main peak, or a peak (1). Selectively switching to individual encoding in response to detecting a change in a characteristic greater than a threshold can advantageously allow action to be taken on sudden changes without evaluating additional characteristics / parameters.

[0033] In an embodiment, a multichannel encoder can be configured to determine whether a parameter describing the direction of a sound source has changed by at least one value (e.g., a threshold) (e.g., relative to a previous / last frame), and to switch based on this determination. The parameter may be the position of the dominant peak in the time-frequency portion of a cross-correlation (e.g., in GCC-PHAT). Switching may include switching to separate encoding. Determining whether the parameter describing the direction of the sound source has changed by at least a threshold can advantageously allow switching to a particular encoding, such as separate encoding, if the sound source moves rapidly relative to a microphone, or if an additional sound source suddenly appears and interferes with the existing sound source in the time-frequency portion.

[0034] Furthermore, a multichannel audio decoder is provided. The multichannel audio decoder can be stereo, or a two-channel or more-than-two-channel audio decoder. The audio decoder can be a general audio decoder, a speech decoder, or a decoder that switches between transform-domain decoding using scaling factors and decoding based on linear prediction coefficients. The decoder is configured to provide a decoded audio representation based on an encoded audio representation. The decoder is configured to switch between parameterized multichannel decoding of multiple channels (e.g., multiple channels of the input audio representation) and individual decoding of multiple channels (e.g., channels of the input audio representation).

[0035] Parametric multichannel decoding can encode combined signals from multiple channels, and can encode the relationships between two or more channels in the form of parameters. Parameters may include inter-channel time difference parameters, and / or inter-channel level difference parameters, and / or inter-channel phase parameters and / or inter-channel correlation parameters.

[0036] Switching between parametric multichannel decoding and individual decoding advantageously allows the decoding (and therefore the encoding) to be adapted to the characteristics of the input audio representation. Selective switching between parametric multichannel decoding and individual decoding allows the selection of an encoding more suitable for encoding the potential input audio representation, resulting in an encoded audio representation that possesses advantageous properties regarding, for example, perceptual performance.

[0037] In other words, the present invention relates to a trade-off between the effort required to obtain the characteristics of an input audio representation and then to act (e.g., switch) based on those characteristics, and the benefit of encoding the input audio representation (and thus making it available for decoding) using, for example, a code that may be advantageous to a particular input audio representation (or a portion thereof) in terms of performance standards.

[0038] In an embodiment, the multichannel audio decoder can be configured to switch between parameterized multichannel decoding and individual decoding based on signaling included in the encoded audio representation. Signaling included in the encoded audio representation simplifies the decoder compared to a decoder that, for example, infers a potential encoding scheme based on the context of the obtained encoded audio representation.

[0039] In addition, an encoded multichannel audio representation is provided. The multichannel audio representation can be stereo, or a two-channel or more-than-two-channel audio representation. The encoded multichannel audio representation includes (e.g., an input audio representation) an encoded parametric multichannel representation of multiple channels and (e.g., an input audio representation) an encoded individual representation of multiple channels.

[0040] Parametric multichannel coding can encode a combined signal that incorporates signals from multiple channels, and encode the relationships between two or more channels in the form of parameters. Parameters may include inter-channel time difference parameters, and / or inter-channel level difference parameters, and / or inter-channel phase parameters and / or inter-channel correlation parameters.

[0041] In other words, the multichannel audio representation of the present invention advantageously allows for the selective use of encodings more suitable for encoding potential input audio representations, such that the resulting encoded audio representation can have advantageous properties, for example, with respect to perceptual performance or any other standard.

[0042] In an embodiment, the encoded multichannel audio representation may further include (e.g., to the decoder) signaling to switch between a parameterized multichannel representation and a single representation. This signaling may indicate the switching while, for example, the encoded multichannel audio representation is being decoded.

[0043] Furthermore, a method for multichannel audio coding is provided. Multichannel coding can include stereo, or stereo or more than stereo audio coding. Audio coding can be performed by a general audio encoder, a speech encoder, or an encoder that switches between transform-domain coding using scaling factors and coding based on linear prediction coefficients. Coding provides an encoded audio representation based on an input audio representation. The method includes switching between parametric multichannel coding of multiple channels (e.g., multiple channels of the input audio representation) and individual coding of multiple channels (e.g., multiple channels of the input audio representation) based on the characteristics of the input audio representation.

[0044] Parametric multichannel coding can encode a combined signal that incorporates signals from multiple channels, and encode the relationships between two or more channels in the form of parameters. Parameters may include inter-channel time difference parameters, and / or inter-channel level difference parameters, and / or inter-channel phase parameters and / or inter-channel correlation parameters.

[0045] Switching between parametric multichannel coding and individual coding, based on the characteristics of the input audio representation, advantageously allows for the adaptation of the coding to the characteristics of the input audio representation. Selective switching between parametric multichannel coding and individual coding can lead to the selection of a coding scheme more suitable for encoding the potential input audio representation, resulting in an encoded audio representation that possesses advantageous properties with respect to, for example, perceptual performance or any other performance criterion.

[0046] Furthermore, a method for multichannel audio decoding is provided. Multichannel audio decoding may include stereo, or two-channel or more-than-two-channel audio decoding. Audio decoding may be performed by a general audio decoder, a speech decoder, or a decoder that switches between transform-domain decoding using scaling factors and decoding based on linear prediction coefficients. Decoding provides a decoded audio representation based on an encoded audio representation. The method includes switching between parametric multichannel decoding of multiple channels (e.g., multiple channels of the input audio representation) and individual decoding of multiple channels (e.g., multiple channels of the input audio representation).

[0047] Parametric multichannel decoding can encode combined signals from multiple channels, and can encode the relationships between two or more channels in the form of parameters. Parameters may include inter-channel time difference parameters, and / or inter-channel level difference parameters, and / or inter-channel phase parameters and / or inter-channel correlation parameters.

[0048] Switching between parametric multichannel decoding and individual decoding advantageously allows the decoding (and therefore the encoding) to be adapted to the characteristics of the input audio representation. The selective switching between parametric multichannel decoding and individual decoding allows for the selection of an encoding more suitable for encoding the potential input audio representation, resulting in an encoded audio representation that possesses advantageous properties regarding, for example, perceptual performance.

[0049] The method may optionally be supplemented by any features, functions, and details of the apparatus also disclosed herein. The method may optionally be supplemented by these features, functions, and details, individually and in combination.

[0050] In addition, a computer program is provided that, when run on a computer, performs one of the methods described above.

[0051] Embodiments of the invention will be discussed below with reference to the accompanying drawings. Attached Figure Description

[0052] Embodiments of the present invention will then be described with reference to the accompanying drawings, wherein:

[0053] Figure 1 A schematic block diagram of an audio encoder according to an embodiment is shown;

[0054] Figure 2 A schematic block diagram of an audio decoder according to an embodiment is shown;

[0055] Figure 3 A flowchart of a method for providing an encoded audio representation according to an embodiment is shown;

[0056] Figure 4 A flowchart of a method for providing a decoded audio representation according to an embodiment is shown;

[0057] Figure 5 A schematic block diagram of an audio encoder according to an embodiment is shown;

[0058] Figure 6 The audio signal and its associated peaks are shown.

[0059] Figure 7 The representation of the relevant function is shown; and

[0060] Figure 8 A schematic block diagram of an audio encoder according to an embodiment is shown. Detailed Implementation

[0061] 1. According to Figure 1 audio encoder

[0062] Figure 1 A multi-channel audio encoder 100 is schematically illustrated. An input audio representation 110 is provided to the multi-channel audio encoder 100 as input. For example, the input audio representation 110 may include multiple channels. The multi-channel audio encoder 100 provides an encoded audio representation 112 as output.

[0063] The multichannel audio encoder 100 includes a function block 120 for performing parametric multichannel encoding and a function block 130 for performing individual encoding of multiple channels. An input audio representation 110 is provided to each of the function blocks 120 and 130. The output of each of the function blocks 120 and 130 is selectively switched by a switching element 140 such that the multichannel audio encoder 100 provides an encoded audio representation 112.

[0064] The multi-channel audio encoder 100 controls the switching element 140 by using a switching control signal 145, based on the characteristics of the input audio representation 110. The control signal 145 may be provided by an optional function block for performing switching control 150, which is included in the multi-channel audio encoder 100 or any other suitable device.

[0065] Alternatively or additionally, the switching control signal 145 can be provided to either function block 120 or 130 such that blocks 120 and 130 can be selectively disabled (e.g., turned off). For example, if the switching control signal 145 instructs a function block used to perform separate encoding of multiple channels 130 to be used to encode the input audio representation 110, then function block 120 used to perform parametric multichannel encoding can be disabled based on the switching control signal 145.

[0066] Alternatively, if the switching control signal 145 indicates that the function block 120 for performing parameterized multichannel encoding is to be used to encode the input audio representation 110, then the function block 130 for performing individual encoding of multiple channels can be disabled based on the switching control signal 145.

[0067] The audio encoder 100 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.

[0068] 2. According to Figure 2 audio decoder

[0069] Figure 2 A multi-channel audio decoder 200 is schematically illustrated. An encoded audio representation 210 is provided as input to the multi-channel audio decoder 200. The multi-channel audio decoder 200 provides a decoded audio representation 212. For example, the decoded audio representation 212 may include multiple channels.

[0070] The multichannel decoder 200 includes a function block 220 for performing parameterized multichannel decoding and a function block 230 for performing individual decoding of multiple channels. An encoded audio representation 210 is provided to each of the function blocks 220 and 230. The output of each of the function blocks 220 and 230 is selectively switched by a switching element 240 such that the multichannel audio decoder 200 provides a decoded audio representation 212.

[0071] The switching element 240 is a controller, for example, via implicit or explicit signaling (not shown) included in the encoded audio representation 210.

[0072] The audio decoder 200 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.

[0073] 3. According to Figure 3 Methods for providing encoded audio representations

[0074] Figure 3 A method 300 for multichannel audio coding is schematically illustrated. Method 300 includes step 310, switching between parametric multichannel coding of multiple channels and individual coding of multiple channels based on the characteristics of the input audio representation. Furthermore, method 300 includes step 320, wherein an encoded audio representation is provided.

[0075] Please note that method 300 may optionally perform other suitable actions disclosed in conjunction with any device (e.g., a multi-channel encoder according to this aspect).

[0076] 4. According to Figure 4 Methods for providing encoded audio representations

[0077] Figure 4 A method 400 for multichannel audio decoding is schematically illustrated. Method 400 includes step 410, switching between parameterized multichannel decoding of multiple channels and individual decoding of multiple channels. Furthermore, method 400 includes step 420, wherein a decoded audio representation is provided.

[0078] Please note that method 400 may optionally perform other suitable actions disclosed in conjunction with any device (e.g., a multi-channel decoder according to this aspect).

[0079] 5. According to Figure 5 audio encoder

[0080] Figure 5 An embodiment of a multi-channel audio encoder 500 is schematically illustrated. Two input audio representation signals, namely audio representation signal 510a and audio representation signal 510b, are provided to the multi-channel audio encoder 500. Audio representation signal 510a corresponds to the left channel and is designated by L, while audio representation signal 510b corresponds to the right channel and is designated by R.

[0081] Each of the input audio representation signals 510a and 510b undergoes optional frequency domain analysis in function blocks 520a and 520b, respectively. Each of function blocks 520a and 520b obtains the signal in the time domain, i.e., the signal's evolution over time, and provides information about the signal's amplitude and / or phase relative to a given frequency band within the frequency range. Function blocks 520a and 520b provide output signals 522a and 522b, respectively. Alternatively, function blocks 520a and 520b may be absent, and signal 522a may be equivalent to signal 510a, and signal 522b may be equivalent to signal 510b.

[0082] Signals 522a and 522b can be provided to function block 530. Block 530 performs a cross-correlation operation on signals 522a and 522b and provides a detection signal 532 indicating whether an interfering speaker is detected in the input audio representation signals 510a and 510b. More specifically, block 530 performs a generalized cross-correlation phase transform, also known as GCC-PHAT, on signals 522a and 522b. GCC-PHAT performs the cross-correlation operation using a weighting function that normalizes the spectral density of the signals to obtain peaks that are advantageously distinguishable relative to, for example, background noise. GCC-PHAT provides a value as a parameter indicating a similarity metric of its input signals, which have a time lag between the two signals. Therefore, by analyzing the peaks in the result of the GCC-PHAT operation, block 530 determines the inter-channel time difference (also known as inter-aural time difference or ITD) and concludes whether an interfering speaker is present in the audio representation signals 510a and 510b. To determine whether there is an interfering speaker in signals 510a and 510b, block 530 may optionally use saliency conditions, stability conditions, and / or noise conditions discussed in conjunction with other embodiments of the invention. Signal 532 may also include an estimate of ITD.

[0083] Signal 532 is provided to controller 540. Controller 540 also receives signals 522a and 522b as inputs. Based on the detection signals provided by block 530, controller 540 selectively provides signals 522a, 522b and an estimated ITD to either parametric stereo encoder 550 (i.e., a function block for parametric multichannel encoding) or LR encoding block 560 (i.e., a function block for encoding individual channels). More specifically, in response to an indication that there is no speaker interference in signals 510a and 510b, controller 540 provides the ITD estimate and signals 522a and 522b to parametric stereo encoder 550. In response, encoder 550 provides an encoded audio representation 552 based on parametric multichannel encoding as the output of multichannel audio encoder 500. Alternatively, in response to an indication that there is speaker interference in signals 510a and 510b, controller 540 provides signals 522a and 522b to LR encoding block 560. In response, the coding block 560 provides a coded audio representation 562 based on individual coding (e.g., left-right, LR coding).

[0084] The parametric stereo encoder 550 can be implemented by the encoding described in [1] or [2]. It should be understood that appropriate standards (or more, rule sets) defining parametric stereo encoding (e.g., in MPEG-4 Part 3 or HE-AAC v2) can be used by the encoder 550. The encoder block 560 can be implemented as described in [4]. It should be understood that appropriate standards (or rule sets) defining individual encodings for multiple channels can be used by the encoding block 560. The encoding block 560 can also implement joint stereo encoding, M / S stereo encoding, etc.

[0085] Figure 6 The exemplary operation of the GCC-PHAT functional unit is visualized, for example, as included in the above combination. Figure 5 The GCC-PHAT functional unit in block 530 is discussed. More specifically, Figure 6 It is a two-dimensional representation of the GCC-PHAT values ​​and an analysis of these values ​​in determining one or more peaks and, based on them, detecting interfering speakers. Figure 6 The horizontal axis shown relates to the progression of time expressed in frames. For the purposes of the following explanation, different time ranges are defined by identifying exemplary time points, such as t1, t2, etc., which are the endpoints of each range. Figure 6 The vertical axis shown is related to the parameters of GCC-PHAT, namely the time lag (e.g., denoted as ITD) between two signals provided to the functional unit that performs GCC-PHAT. Figure 6 The color on the two-dimensional plane corresponds to the GCC-PHAT value for a given frame and a given time lag.

[0086] Within the exemplary time range (i.e., frame range) between t1 and t2, multiple main peaks determined by the GCC-PHAT functional unit are shown (in... Figure 6 In the legend, each main peak is represented by a cross and designated as "peak 1". According to one or more embodiments of the invention, the GCC-PHAT functional unit can determine the main peak. Multiple subordinate peaks determined by the GCC-PHAT functional unit are also shown in the range t1 to t2 (in...). Figure 6 In the legend, each subordinate peak is represented by a circle and designated as "Peak 2". According to one or more embodiments of the present invention, the GCC-PHAT functional unit can determine subordinate peaks.

[0087] Within the range t1 to t2, the GCC-PHAT function can determine that multiple main peaks 610 satisfy stability conditions, for example, given that the positions of peaks 610 (in terms of time lag) (across consecutive frames) differ from each other by at most a certain threshold. Furthermore, for example, although the position of peak 620 indicates some scattering in at least a series of consecutive frames within the range t1 to t2 adjacent to t2, the GCC-PHAT function can determine that multiple subordinate peaks 615 included in the range t1 to t2 satisfy stability conditions (either the same or different parameterizations as the main peak 610). Therefore, given that peaks 610 and 615 satisfy stability conditions, the GCC-PHAT function (or, for example, different functional units included in block 530) can determine the presence of interfering speakers.

[0088] In another exemplary range t3 to t4, the main peak 620 exhibits a similar pattern to that in range t1 to t2. Therefore, the stability condition can be determined by the GCC-PHAT function. For the multiple subordinate peaks 625, the GCC-PHAT function can determine that, considering the scattering pattern (i.e., positions that differ significantly in terms of time lag for at least some subranges of consecutive frames), at least some of the peaks 625 do not satisfy the stability condition. Therefore, given that only one of the two evaluated stability conditions is satisfied, it can be determined that there is no interfering speaker.

[0089] For the exemplary ranges t5 to t6 and t6 to t7, given the stability of the main peak and the scattering of the subordinate peaks, this determination can correspond to the determination made for the range t3 to t4. For the exemplary range t8 to t9, given the stability of the main peak and the subordinate peaks, this determination can correspond to the determination made for the range t1 to t2.

[0090] Figure 7 This illustrates an example of a single frame (e.g.) Figure 6 The evolution of GCC-PHAT (one of the frames shown). Figure 7 In the figure, the horizontal axis is related to the time lag parameter, and is related to... Figure 6 The vertical axis corresponds to this. Figure 7 The ordinate is related to the cross-correlation value, for example, to the value provided by the GCC-PHAT function. For Figure 7 The evolution of the peaks, the main peak (represented as peak 1, 710) and the subordinate peaks (represented as peak 2, 720), are determined by the GCC-PHAT function. According to one or more embodiments of the invention, given that the distance between the amplitudes (i.e., cross-correlation values) of the main peak 710 and the subordinate peak 720 and the cross-correlation value of the background noise 730 is greater than a threshold (e.g., defined according to one or more embodiments of the invention), it can be determined that both the main peak 710 and the subordinate peak 720 satisfy a noise condition.

[0091] Furthermore, according to one or more embodiments of the present invention, given that the distance between peak 710 and dependent peak 720 in terms of time lag (i.e., along the horizontal axis) is greater than a threshold (e.g., as defined by one or more embodiments of the present invention), (e.g., via the GCC-PHAT function or Figure 5 Block 530) can determine that peak 710 and subordinate peak 720 can satisfy the significance condition.

[0092] Furthermore, according to one or more embodiments of the present invention, given that the cross-correlation value of each of peaks 710 and dependent peaks 720 is greater than a threshold (e.g., a threshold defined according to one or more embodiments of the present invention, specifically, for example, greater than the value of 0.15 defined for peak (1) in option 1 below), (e.g., via the GCC-PHAT function or Figure 5 Block 530) can determine that peak 710 and dependent peak 720 satisfy different significance conditions shown.

[0093] Furthermore, according to one or more embodiments of the invention, given that the relationship between the cross-correlation values ​​of peaks 710 and 720 has a proportion below a threshold (e.g., a threshold defined according to one or more embodiments of the invention, and illustrated below by using an example with a constant c = 0.8), (e.g., via the GCC-PHAT function or Figure 5 Block 530) can determine that peak 710 and dependent peak 720 satisfy different significance conditions shown.

[0094] Please note that this invention is not limited to the use of GCC-PHAT, but rather to any technique capable of providing an indication of cross-correlation values, i.e., any suitable cross-correlation technique, and suitable pattern recognition techniques, such as pattern recognition techniques involving neural networks, can be used.

[0095] Other embodiments of the invention are described below. In addition to the aspects disclosed above, the embodiments described below may constitute alternatives or may be considered. The embodiments described below relate to detecting interfering speakers captured using a stereo microphone setup. The embodiments described below are useful tools, for example, for stereo speech codecs that can be used in communication applications.

[0096] Referring to the above description, for certain specific situations, discrete coding of the two stereo channels may be preferred for better performance. In cases of speaker interference, advantageous embodiments may involve switching between a parametric model (Mode A) and a discrete model (Mode B). On the other hand, it involves the ability to automatically detect when to switch from Mode A to Mode B and when to switch from Mode B to Mode A. The following considerations generally apply to the first case, i.e., when to switch from Mode A to Mode B.

[0097] When two speakers have different ITDs (Interaural Time Difference) and the difference between the two ITDs is large (significant), the exemplary solution considers it an important case (e.g., only the most critical case).

[0098] In some embodiments, such as those described in [3], it can be assumed that the codec already has an ITD estimator and that the ITD estimator is based on GCC-PHAT (Generalized Cross-Correlation Phase Transform). The basic principle of such an estimator is to detect a peak in the GCC-PHAT that corresponds to the ITD of the stereo signal. However, when two speakers are speaking simultaneously and they have two different ITDs, there are in most cases two peaks in the GCC-PHAT. Some embodiments detect whether there is only one peak in the GCC-PHAT (Mode A) or two peaks that are far apart from each other (Mode B).

[0099] In one embodiment, the starting point can be mode A. The GCC-PHAT of the stereo signal can be calculated, possibly using a smoothed version of the cross-spectrum or any other processing. The main peak of the GCC-PHAT can be estimated. In most cases, this can correspond to the maximum value of the absolute value of the GCC-PHAT. Alternatively or additionally, some hysteresis mechanism can be applied to obtain a more stable ITD estimate. A portion of the GCC-PHAT that is sufficiently far from the main peak can be selected. The distance between the main peak and the boundary of this portion can be above a certain threshold. A second peak in the selected portion can be found: for example, this can be the maximum value of the absolute value of the GCC-PHAT. If the value of the second peak is above a certain threshold, for example, if peak (2) > c * peak (1), where peak (1) and peak (2) are the values ​​of the first and second peaks, respectively, and c can be a constant (e.g., c = 0.8) or a signal adaptive variable, then the GCC-PHAT can be considered to contain two significant peaks and a switch to mode B can occur. Otherwise, there is no significant second peak, and mode A is still in use.

[0100] In addition, the following embodiments / options are disclosed:

[0101] In option 1, a check can be performed to ensure that the peak (1) is above a certain threshold (e.g., 0.15) to avoid switching on noisy frames.

[0102] In Option 2, it may be necessary to verify the two conditions of the two embodiments described above on two consecutive frames. This avoids switching on unstable signals.

[0103] In option 3, it may be necessary for the peaks (2) of two consecutive frames to be close to each other (e.g., their difference may be less than 4). This avoids switching on unstable signals.

[0104] In option 4, the SAD flag of the previous frame must be 1 (meaning it is an active signal). This avoids switching at the first frame of the signal section.

[0105] In option 5, peak (1) can change drastically from one frame to the next. In this case, it may not be necessary to check for the second peak, and it can be assumed that the second speaker has started speaking, and a switch to mode B may occur.

[0106] In some embodiments, after the GCC-PHAT detector determines the presence of a distracting speaker as described in one or more of the above embodiments: if no distracting speaker is detected, the system remains in its default parameterization mode and may forward the estimated ITD value to parameterization processing, for example, as described in [1]. If a distracting speaker is detected, the system may switch to an LR coding scheme, for example, encoding each channel individually using an EVS codec [4].

[0107] The described embodiments enable the detection of interfering speech segments in stereo speech signals under certain conditions, for which a switch from a parametric stereo coding system to a discrete coding system is preferable. In this way, the perceptual quality of the codec can be improved. For parametric coding schemes, inter-channel time difference (ITD) detectors may be present in some codecs. Therefore, the additional complexity overhead or additional latency may be acceptable.

[0108] The following aspects are further disclosed and may be used alone or optionally in combination with any features, functions, and details disclosed herein:

[0109] Aspect 1: A stereo speech coding system in which the codec can switch from parametric coding mode (mode A) to discrete LR coding mode (mode B) once the classifier / signal analyzer determines that the conditions for switching from parametric coding mode (mode A) to discrete LR coding mode (mode B) are met.

[0110] Aspect 2: A stereo speech coding system in which the codec can switch from parametric coding mode (mode A) to discrete LR coding mode (mode B) once the classifier / signal analyzer detects that the signal violates the model under the parametric coding scheme.

[0111] Aspect 3: A stereo speech coding system in which the codec switches from a parametric coding mode (mode A) to a discrete LR coding mode (mode B) once the system detects an interfering speaker.

[0112] Aspect 4: For stereo speech coding, PHAT generalized cross-correlation is used to detect the first maximum absolute value (peak) and the second highest absolute value, and interference speech segments are detected according to the conditions applicable to the second highest absolute value.

[0113] The above discussion Figure 6 This is a visualization of the steps / aspects / implementations described above, in which a scatter plot of the signal is drawn, and... Figure 7 The image shows the scaling of a single frame representation.

[0114] 6. According to Figure 8 audio encoder

[0115] Figure 8 A schematic block diagram of an audio encoder 800 according to an embodiment of the present invention is shown.

[0116] The audio encoder 800 receives an input audio representation 810, which may include, for example, multiple channels (e.g., channels L, R). The audio encoder 800 provides an encoded audio representation 812, which may, for example, represent the audio content of the input audio representation.

[0117] The audio encoder 800 optionally includes a first frequency domain analysis 820, which receives, for example, a first channel 810a of the input audio representation, and provides a frequency domain representation 822 of the first channel 810a based on the first channel 810a of the input audio representation. The audio encoder 800 optionally includes a second frequency domain analysis 824, which receives, for example, a second channel 810b of the input audio representation, and provides a frequency domain representation 826 of the second channel 810b based on the second channel 810b of the input audio representation. For example, the first and second frequency domain analyses may use, for example, short-term Fourier transform, MDCT transform, filter banks, etc., to provide frequency domain or spectral domain representations 822, 826 of the channels of the input audio representation.

[0118] The audio decoder 800 also includes parametric multichannel encoding 830 and individual encoding 834 for multiple channels. For example, multichannel encoding 830 may receive channels 810a, 810b of the input audio representation, or alternatively, receive frequency domain representations 822, 826 provided by frequency domain analyses 820, 824. Alternatively, multichannel encoding may receive different representations of the channels of the input audio representation. Parametric multichannel encoding provides encoded representations of two or more input channels to parametric multichannel representation 832, wherein the channels of the input signal representation may be represented, for example, using a combined signal (e.g., a downmixed signal) and parametric auxiliary information, such as representing similar signal components in all channels (or at least some channels, e.g., two or more channels) of the input signal representation, and parametric auxiliary information describing, for example, the similarity and / or difference between the two or more channels of the input audio representation in the form of parameter values. For example, parametric auxiliary information may include inter-channel level differences and / or inter-channel phase differences and / or inter-channel time differences and / or inter-channel correlation values ​​and / or any other parameters describing the relationships between the channels of the input audio representation. Parametric auxiliary information may preferably be available on the audio decoder side to reconstruct the channels of the input audio representation at least approximately based on the combined signal. For example, parameter values ​​for the parametric auxiliary information may be provided individually for different time-frequency ranges or for different spectral bins. For example, parametric multichannel coding may consider the concept of "parametric stereo," which, for example, is used as an extension of MPEG4 High Efficiency Advanced Audio Coding (HE-AAC), and may provide corresponding representations of the channels of the input audio representation.

[0119] The audio encoder 800 also includes separate encoding 834 for multiple channels, wherein, for example, different channels of the input audio representation are separately encoded, for example, using separate encoding of spectral values. Thus, the separate encoding 834 provides separately encoded information 836 associated with different channels of the input audio representation, which, for example, allows for separate decoding of the channels of the input audio representation on the audio decoder side.

[0120] Furthermore, the audio encoder is configured to switch between parametric multichannel encoding 830 and individual encoding 834, allowing the control block of the audio encoder to select whether to include information from the parametric multichannel representation 832 or the individually encoded representation in the encoded audio representation 812. Regarding this, the following is irrelevant: whether both parametric multichannel encoding 830 and individual encoding 834 were performed for a given frame and whether the encoded representation 832 provided by parametric multichannel encoding or the encoded representation 836 provided by individual encoding is actually included in the encoded audio representation 812, or whether only parametric multichannel encoding or individual encoding was selected for a given frame (where the latter solution is generally more efficient but may introduce additional latency).

[0121] The following will describe how the selection of parametric multichannel encoding 830 or individual encoding 834 (or equivalently, information 836 regarding parametric multichannel representation 832 or individual encoding associated with different channels of the input audio representation) should be included in the encoded audio representation 812.

[0122] For this purpose, the audio encoder 800 includes a decorrelation determination 840, which can determine the correlation (e.g., cross-correlation) between two or more channels of the input audio representation, for example, based on the frequency domain representations 822, 826 of the channels of the input audio representation. However, it should be noted that the correlation determination 840 can operate, for example, based on the time domain representations of the channels of the input audio representation. Furthermore, it should be noted that the correlation determination can provide separate correlation information 842 for different frequency ranges or time-frequency portions of the input audio representation. Therefore, not only can there be separate correlation information 842 for subsequent frames of the input audio representation, but there can even be separate correlation information 842 for individual frequency ranges or frequency blocks. Furthermore, it should be noted that the correlation information 842 can be in the form of a correlation function (e.g., for each time-frequency portion) that includes different correlation values ​​for different correlation hysteresis values ​​(also specified as hysteresis or time hysteresis).

[0123] For example, the so-called "GCC-PHAT" technique can be used to obtain relevant information, which has been found to yield particularly meaningful results. However, different concepts can also be used to determine (mutual)related information.

[0124] The audio decoder 800 also includes a dominant peak determination 850, which can be configured to determine the dominant peak (e.g., the maximum absolute value of GCC_PHAT) of the cross-correlation between two or more channels of the input audio representation based on cross-correlation information, and to provide information 852 describing the dominant peak (e.g., including the inter-channel time difference, peak value, or peak intensity). For example, the dominant peak determination 850 can determine that the cross-correlation information (or the cross-correlation function represented by the cross-correlation information) includes a (global) maximum value for which correlation lag (or equivalently for which time lag, or equivalently for which inter-channel time difference). Optionally, the dominant peak determiner can also determine the peak value (or peak intensity) itself. However, it should be noted that the dominant peak determiner does not necessarily need to identify the maximum value of the cross-correlation function as the dominant peak. Instead, the dominant peak determiner can, for example, disregard “sporadic” or “unstable” peaks and identify stable peaks (e.g., peaks that are stable across multiple frames and can be classified as “significant,” such as those greater than a threshold or exceeding a predetermined value of the noise floor) as dominant peaks (wherein, for example, a lag mechanism can be used to have a more stable ITD estimate). It should be noted that different algorithms known to those skilled in the art for identifying peaks or main peaks of correlation functions can be used.

[0125] Optionally, the audio decoder also includes a peak checker 852 that receives the main peak information 852 and checks the reliability of the main peak information. For example, the peak checker can identify unreliable main peak information that includes large fluctuations over time (e.g., peak ITD and / or peak intensity), and / or that indicates an excessively low peak intensity. For example, the value of the main peak can be checked to avoid switching on noisy frames. Optionally, it can also be determined whether the main peak meets one or more conditions on multiple frames (e.g., regarding peak value). In summary, such unreliable main peak information can be suppressed and / or substituted, and / or signaled by default information.

[0126] Furthermore, the audio decoder may include a second peak determination 860, which may be configured to determine a second peak of cross-correlation between two or more channels of the input audio representation based on cross-correlation information 842, and provide information 862 describing the second peak (e.g., including inter-channel time difference, peak value, or peak intensity). For example, the second peak may be a local maximum of the cross-correlation function described by the cross-correlation information 842, which includes the second largest peak after the peak value of the main peak. Alternatively, it may be optional to identify the local maximum of the cross-correlation information as the second peak, which satisfies one or more predetermined conditions regarding the main peak of the cross-correlation function and / or regarding the noise floor. For example, the second peak determination may receive information about the main peak from the main peak determination 850 and take that information into account when identifying the second peak. For example, the second peak determination 860 may check whether the distance of the second peak candidate (e.g., a local maximum of the cross-correlation function) includes a predetermined distance condition from the main peak (e.g., in terms of correlation hysteresis or ITD), where, for example, it may be necessary for the second peak to include a predetermined minimum distance from the main peak. Alternatively, the determination of the second peak can be performed based on a (selected) portion of GCC-PHAT that is “far from the main peak,” for example, at a predetermined distance from the main peak in terms of ITD, wherein, for example, the (absolute) maximum value of the absolute value of GCC-PHAT in the selected portion of GCC-PHAT can be identified as the second peak.

[0127] Alternatively or additionally, the second peak determination can examine whether a second peak candidate (e.g., in terms of the relationship between the peak values ​​of the primary peak and the second peak) meets predetermined peak conditions. For example, it may be necessary for the value of the second peak to be above a certain threshold, which can be defined relative to the value of the primary peak.

[0128] Furthermore, the second peak determination can check whether the peak value of the second peak candidate is sufficient to be above the background noise of the cross-correlation information.

[0129] Therefore, the second peak determination 860 can determine whether there exists a second peak that meets the requirements to be identified as a second peak, and provide second peak information 862 describing the second peak (e.g., in terms of relevant hysteresis and / or ITD and / or peak value and / or peak intensity). Optionally, the second peak information can indicate that there is no second peak that meets the conditions.

[0130] Optionally, the audio decoder may also include a second peak saliency evaluation 864, which may, for example, receive second peak information 862 and determine whether the second peak described by the second peak information 862 is significant and / or reliable. For example, the second peak saliency evaluation may check whether the second peak satisfies one or more conditions over multiple frames. For example, the second peak saliency evaluation may determine whether the second peak is above a certain threshold (e.g., relative to the main peak) for multiple frames. Alternatively or additionally, the second peak saliency evaluation may check whether the relevant hysteresis value or ITD value of the second peak is sufficiently close over two or more (subsequent) frames. However, other conditions of the second peak may also be optionally checked.

[0131] It should be noted that the functions described in the main peak check 854 can be optionally integrated into the main peak determination 850. Furthermore, the function for assessing the significance of the second peak can be optionally included in the second peak determination 860. Additionally, it should be noted that when determining the information 856 describing the main peak and the information 866 describing the second peak, the above conditions may not be checked, some or all of the above conditions may be checked, or additional conditions may be checked.

[0132] Furthermore, it should be noted that information 856 describing the primary peak may optionally only indicate whether a valid primary peak has been found. Similarly, information 866 describing the secondary peak may optionally only indicate whether a valid secondary peak has been found. However, information 856 and 866 may also optionally describe details about the peaks, such as relevant hysteresis and / or ITD and / or peak value.

[0133] The audio encoder 800 may optionally include detection 870 of detecting a change in the correlation hysteresis or ITD of the main peak that is greater than a threshold, and providing information 872 describing whether such a change exists.

[0134] The audio encoder 800 also includes a switching decision 880, which is configured to determine whether a parameterized multichannel representation 832 or separate encoded information 836 associated with different channels of the input audio representation should be included in the encoded audio representation.

[0135] In a simple case, the switching decision 880 can simply check whether a significant (or valid) second peak is available. If only a single peak (i.e., the main peak) exists, parametric multichannel encoding 830 can be used (or parametric multichannel representation 832 can be included in the encoded audio representation). If the information 866 describing the second peak indicates the presence of a significant (or valid) second peak, the switching decision can decide to use separate encoding 834 (or include separate encoding information 836 associated with different channels of the input audio representation in the encoded audio representation).

[0136] However, the switching decision can optionally use one or more additional criteria to determine which information should be included in the encoded audio representation.

[0137] For example, the switching decision may optionally consider whether there is a change in the main peak greater than a (predetermined or variable) threshold, wherein, in response to the discovery that the change in the main peak is greater than the threshold (e.g., signaled by information 872), the switching decision may switch to using separate encoding 834 (or include separate encoding information 836 associated with different channels of the input audio representation into the encoded audio representation).

[0138] As another example, the handover decision may optionally consider indicators that indicate whether a previous frame was already active (such as the SAD flag). For instance, if the handover decision finds that a previous frame was already inactive, the handover decision may selectively suppress the handover.

[0139] However, the switching decision may also optionally evaluate information about other signal characteristics of the input audio representation and, based on that information, determine which information should be included in the encoded audio representation.

[0140] In summary, the audio encoder 800 determines, for example on a frame-by-frame basis, whether to include the parameterized multichannel representation 832 or separate encoded information 836 associated with different channels of the input audio representation in the encoded audio representation, based on an analysis of the characteristics of the input audio representation (e.g., based on the determination of how many “significant” or “valid” peaks exist in the cross-correlation function).

[0141] However, it should be noted that it is not necessary to specifically assign functions to different functional blocks. Instead, some or all functions can be combined into a single functional block if needed.

[0142] Furthermore, it should be noted that the audio encoder 800 may optionally be supplemented individually and in combination by any features, functions, and details disclosed herein.

[0143] Furthermore, any features, functionalities, and details disclosed herein may be optionally incorporated, individually or in combination, into any of the embodiments disclosed herein.

[0144] 7. Implement alternative solutions

[0145] Although some aspects have been described in the context of the apparatus, it will be clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of the features of the corresponding block or item or the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices (such as microprocessors, programmable computers, or electronic circuits). In some embodiments, such devices may be used to perform one or more of the most important method steps.

[0146] Novel coded audio signals can be stored on digital storage media or transmitted on transmission media such as wireless or wired transmission media (e.g., the Internet).

[0147] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or software. Implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) on which electronically readable control signals are stored, which cooperate with (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.

[0148] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0149] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable medium.

[0150] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.

[0151] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.

[0152] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded, the computer program being used to perform one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.

[0153] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence of a computer program used to perform one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0154] Another embodiment includes a processing means, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.

[0155] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0156] Another embodiment of the invention includes an apparatus or system configured to transmit a computer program to a receiver (e.g., electronically or optically), the computer program being used to perform one of the methods described herein. The receiver may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.

[0157] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0158] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.

[0159] The apparatus described herein or any component thereof may be implemented, at least in part, in hardware and / or in software.

[0160] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.

[0161] The methods or any components of the apparatus described herein may be performed, at least in part, by hardware and / or by software.

[0162] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the invention is intended to be limited only by the scope of the appended claims and not by the specific details given by way of the description and explanation of the embodiments herein.

[0163] References[1]S. Bayer, M. Dietz, S. Doehla, E. Fotopoulou, G. Fuchs, W. Jaegers, G. Markovic, M. Multrus, E. Ravelli and M. Schnell, "APPARATUSES AND METHODS FOR ENCODING OR DECODING A MULTI-CHANNEL AUDIO SIGNAL USING FRAME CONTROL SYNCHRONIZATION," WO17125562, July 27, 2017.

[0164] [2]M. Schroeder and B. Atal, "Code-excited linear prediction (CELP): High-quality speech at very low bit rates," ICASSP'85. IEEE International Conference on Acoustics, Speech, and Signal Processing, Tampa, FL, USA, 1985.

[0165] [3]S. Bayer, M. Dietz, S. Doehla, E. Fotopoulou, G. Fuchs, W. Jaegers, G. Markovic, M. Multrus, E. Ravelli and M. Schnell, "APPARATUS AND METHOD FOR ENCODING OR DECODING A MULTI-CHANNEL SIGNAL USING A BROADBAND ALIGNMENT PARAMETER AND A PLURALITY OF NARROWBAND ALIGNMENT PARAMETERS", WO17125558, July 27, 2017.

[0166] [4]3GPP TS 26.445, Codec for Enhanced Voice Services (EVS); Detailed algorithmic description.

Claims

1. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multichannel audio encoder (100, 500, 800) is configured to determine whether a single dominant source exists in multiple time-frequency portions, or whether two or more sources exist in a given time-frequency portion, wherein the multichannel coding parameters of the two or more sources differ by at least a predetermined deviation or more than a predetermined deviation, and to switch based on the determination of whether the multichannel coding parameters differ by at least a predetermined deviation or more than a predetermined deviation; Wherein, the multichannel coding parameters are based on the relationship between the channels represented by the input audio; and The multi-channel audio encoder is configured to switch to the parameterized multi-channel encoding in the case of a single source.

2. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to determine whether the input audio representations (110, 510a, 510b, 810) satisfy the assumptions of the model under the parameterized multi-channel encoding (120, 550, 830), and switch accordingly.

3. The multi-channel audio encoder (100, 500, 800) according to claim 2, wherein, The multi-channel audio encoders (100, 500, 800) are configured to switch to the individual encoding (130, 560, 834) if the assumptions of the model under the parameterized multi-channel encoding (120, 550, 830) are not met.

4. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to determine whether the input audio representation (110, 510a, 510b, 810) corresponds to the dominant source, and switch accordingly based on the determination.

5. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multichannel audio encoders (100, 500, 800) are configured to determine whether a single dominant source exists in multiple time-frequency portions, and / or to determine whether two or more sources exist in a given time-frequency portion, and to switch based on the determination, wherein the multichannel encoding parameters of the two or more sources differ by at least a predetermined deviation or more than a predetermined deviation.

6. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to determine the parameters of the model under the parameterized multi-channel encoding (120, 550, 830) and switch according to the parameters of the model.

7. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multichannel audio encoders (100, 500, 800) are configured to determine whether the characteristics defining the relationship between the channels of the input audio representation (110, 510a, 510b, 810) allow explicit determination of the multichannel coding parameters or indicate two or more different possible values ​​of the multichannel coding parameters, and switch according to the determination.

8. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multichannel audio encoder (100, 500, 800) is configured to determine whether the property defining the relationship between the channels of the input audio representation (110, 510a, 510b, 810) includes only a single salience value that satisfies the salience condition, or whether the property defining the relationship between the channels of the input audio representation (110, 510a, 510b, 810) includes two or more salience values ​​that satisfy the salience condition, and to switch according to the determination.

9. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to determine parameters of the previous frame and switch according to the parameters of the previous frame.

10. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to determine whether there is an interference source in the input audio representations (110, 510a, 510b, 810) and switch accordingly.

11. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multichannel audio encoder (100, 500, 800) is configured to determine whether there are two or more values ​​describing the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810), and to switch according to the determination, wherein the two or more values ​​satisfy a significance condition and are associated with a single time-frequency component.

12. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to determine whether there are two or more peaks (610, 615, 620, 625, 710, 720) in the cross-correlation between two or more channels of the input audio representation, and to switch according to the determination.

13. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) include estimators (530, 840) configured to estimate the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810) based on cross-correlation. The multi-channel audio encoder (100, 500, 800) is configured to determine whether the difference between two peaks (610, 615, 620, 625, 710, 720) associated with different cross-correlation hysteresis is greater than a value, and to switch according to the determination.

14. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multichannel audio encoder (100, 500, 800) is configured to determine whether the distance between two or more values ​​describing the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810) is greater than a value, and to switch according to the determination, wherein the two or more values ​​satisfy a significance condition and are associated with the same time-frequency component.

15. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to determine a first characteristic value based on the evolution of cross-correlation, and to switch according to the determination.

16. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to determine one or more dependent characteristic values ​​based on cross-correlation evolution, and to switch according to the determination, and / or The multi-channel audio encoders (100, 500, 800) are configured to determine whether one or more dependent characteristic values ​​exist based on the evolution of the cross-correlation, and to switch accordingly.

17. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to determine whether the main peak (610, 620, 710) and one or more subordinate peaks (615, 625, 720) satisfy a significance condition, and to switch based on the determination, and / or The multi-channel audio encoder (100, 500, 800) is configured to determine whether there are one or more dependent peaks (615, 625, 720) in the cross-correlation that meet the correlation criteria, and to switch according to the determination.

18. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to selectively consider the dependent peaks (615, 625, 720) in the given frame represented by the input audio if one or more corresponding dependent peaks (615, 625, 720) exist in one or more frames preceding the given frame.

19. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to determine whether one or more characteristic values ​​describing the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810) satisfy a stability condition, and to switch based on the determination.

20. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to determine whether a noise condition is met for multiple frames, and if the noise condition is met, selectively avoid switching.

21. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to determine, for multiple frames, whether a saliency condition and / or stability condition for a characteristic value is met, and to switch accordingly.

22. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to determine whether the distance between one or more subordinate peaks (615, 625, 720) is within a predetermined range, and to switch and / or selectively avoid switching based on the determination.

23. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to selectively avoid switching at or after the first frame following an inactive frame of the input audio representation. The multi-channel audio encoder (100, 500, 800) is configured to determine whether a given flag in a frame has changed relative to one or more previous frames, and selectively avoid switching based on the determination.

24. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoders (100, 500, 800) are configured to selectively switch to the individual encoding (130, 560, 834) in response to detecting a change in the characteristics of the input audio representation (110, 510a, 510b) greater than a threshold.

25. The multi-channel audio encoder (100, 500, 800) according to claim 1, wherein, The multi-channel audio encoder (100, 500, 800) is configured to determine whether a parameter describing the direction of a sound source has changed by at least one value, and to switch accordingly.

26. A method (300) for multichannel audio coding, for providing (320) encoded audio representation based on an input audio representation, the method comprising: Based on the characteristics of the input audio representation, switching is performed between parameterized multichannel encoding of multiple channels and individual encoding of multiple channels (310). The method includes: determining whether a single dominant source exists in multiple time-frequency portions, or whether two or more sources exist in a given time-frequency portion, wherein the multi-channel coding parameters of the two or more sources differ by at least a predetermined deviation or more than a predetermined deviation, and switching based on the determination of whether the multi-channel coding parameters differ by at least a predetermined deviation or more than a predetermined deviation; The multichannel coding parameters are based on the relationship between the channels represented by the input audio; and The switch to the parameterized multichannel encoding is performed in the case of a single source.

27. A computer-readable storage medium having instructions stored thereon that, when executed on a computer, perform the method according to claim 26.

28. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) is configured to determine whether the characteristic defining the relationship between the channels of the input audio representation (110, 510a, 510b, 810) includes only a single significant value that satisfies a saliency condition, or whether the characteristic defining the relationship between the channels of the input audio representation (110, 510a, 510b, 810) includes two or more significant values ​​that satisfy a saliency condition, and to switch according to the determination.

29. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) is configured to determine whether there are two or more values ​​describing the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810), and to switch according to the determination, wherein the two or more values ​​satisfy a saliency condition and are associated with a single time-frequency component.

30. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) is configured to determine whether there are two or more peaks (610, 615, 620, 625, 710, 720) in the cross-correlation between two or more channels of the input audio representation, and to switch according to the determination. The cross-correlation is related to a given time-frequency component.

31. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) includes an estimator (530, 840) configured to estimate the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810) based on cross-correlation. The multi-channel audio encoder (100, 500, 800) is configured to determine whether the difference between two peaks (610, 615, 620, 625, 710, 720) associated with different cross-correlation hysteresis is greater than a value, and to switch according to the determination.

32. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) is configured to determine whether the distance between two or more values ​​describing the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810) is greater than a value, and to switch according to the determination, wherein the two or more values ​​satisfy a significance condition and are associated with the same time-frequency component.

33. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) is configured to determine whether the main peak (610, 620, 710) and one or more subordinate peaks (615, 625, 720) satisfy a significance condition, and to switch based on the determination, and / or The multi-channel audio encoder (100, 500, 800) is configured to determine whether there are one or more dependent peaks (615, 625, 720) in the cross-correlation that meet the correlation criteria, and to switch according to the determination.

34. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) is configured to determine whether one or more characteristic values ​​describing the relationship between two or more channels of the input audio representation (110, 510a, 510b, 810) satisfy a stability condition, and to switch according to the determination.

35. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoders (100, 500, 800) are configured to determine whether a noise condition is met for multiple frames, and if the noise condition is met, selectively avoid switching.

36. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoders (100, 500, 800) are configured to selectively avoid switching at or after the first frame following an inactive frame of the input audio representation. The multi-channel audio encoder (100, 500, 800) is configured to determine whether a given flag in a frame has changed relative to one or more previous frames, and selectively avoid switching based on the determination.

37. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoders (100, 500, 800) are configured to selectively switch to the individual encoding (130, 560, 834) in response to detecting a change in the characteristics of the input audio representation (110, 510a, 510b) greater than a threshold. The characteristic represented by the input audio is the time difference between channels or the main peak of the cross-correlation between two or more channels represented by the input audio.

38. A multi-channel audio encoder (100, 500, 800) for providing an encoded audio representation (112, 552, 562, 812) based on an input audio representation (110, 510a, 510b, 810). in, The multi-channel audio encoders (100, 500, 800) are configured to switch between parametric multi-channel encoding (120, 550, 830) and individual encoding (130, 560, 834) for multiple channels based on the characteristics of the input audio representations (110, 510a, 510b, 810). The multi-channel audio encoder (100, 500, 800) is configured to determine whether a parameter describing the direction of a sound source in the input audio representation has changed by at least one value, and to switch accordingly.

Citation Information

Patent Citations

  • Stereo audio signal transmitter

    JP1994252863A

  • Audio coding selection based on device operating conditions

    JP2012516471A

  • Coding device and coding method

    WO2018221138A1