Audio encoding and audio decoding
By separating the multi-channel audio signal into the first and second subsets and performing differential encoding processing, the problem of low encoding efficiency of the multi-channel audio signal under different bandwidth conditions is solved, and efficient audio signal transmission and quality preservation are achieved.
Patent Information
- Application Number
- CN202080067697.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-26
- Filing Date
- 2020-09-16
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2040-09-16
Smart Images

Figure BDA0003564927550000121 
Figure BDA0003564927550000161 
Figure BDA0003564927550000162
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to audio encoding and audio decoding. In particular, embodiments of the present disclosure relate to encoding a multi-channel audio signal and also decoding the same to obtain the multi-channel audio signal. Background Art
[0002] The multi-channel audio signal includes a plurality of audio signals.
[0003] In order to store or transmit a multi-channel audio signal, it is desirable to compress the multi-channel audio signal through encoding. Summary of the Invention
[0004] According to various, but not all, embodiments, an apparatus is provided that includes means for:
[0005] Receive multi-channel audio signals;
[0006] identifying at least one audio signal to separate from the multi-channel audio signal;
[0007] Based on the identified at least one audio signal, separate the plurality of audio signals into at least a first subset of audio signals and a second subset of audio signals, wherein the first subset includes the identified at least one audio signal and the second subset includes remaining audio signals in the received multi-channel audio signal;
[0008] analyzing remaining audio signals in the second subset of audio signals to determine one or more transmitted audio signals and metadata; and
[0009] The at least one audio signal, the one or more transmission audio signals, and the metadata are encoded.
[0010] In some but not all examples, the first subset of audio signals is a fixed subset of the plurality of audio signals and the second subset of audio signals is a fixed subset of the plurality of audio signals.
[0011] In some but not all examples, the first subset includes a center speaker channel signal and / or a pair of stereo channel signals, and / or the first subset of audio channels includes one or more dominant speech audio channel signals.
[0012] In some but not all examples, the first subset of audio signals is a variable subset of the plurality of audio signals and the second subset of audio signals is a variable subset of the plurality of audio signals.
[0013] In some but not all examples, the count of the first audio signal subset is variable, and / or the composition of the first audio signal subset is variable.
[0014] In some but not all examples, the first subset of audio signals are signals determined to satisfy a first criterion and the second subset of audio signals are signals determined not to satisfy the first criterion.
[0015] In some but not all examples, the first criterion depends on one or more first audio characteristics of the audio signals, the first subset of audio signals having and sharing the one or more first audio characteristics, and the second subset of audio signals not having the one or more first audio characteristics.
[0016] In some but not all examples, the first criterion depends on one or more spectral characteristics of the audio signals, at least some of the audio signals in the first subset of audio signals share the one or more spectral characteristics, and the second subset of audio signals does not share the one or more spectral characteristics.
[0017] In some but not all examples, the one or more first audio characteristics include energy levels of the audio signals, and each audio signal in the first subset of audio signals has an energy level greater than any audio signal in the second subset of audio signals.
[0018] In some but not all examples, the one or more first audio characteristics include audio signal correlation, and each audio signal in the first subset of audio signals has a greater cross-correlation with audio signals in the first subset than with audio signals in the second subset.
[0019] In some but not all examples, the one or more first audio characteristics include audio signal decorrelation, and at least some audio signals in the first subset of audio signals have low cross-correlation with other audio signals in the first subset and with audio signals in the second subset.
[0020] In some but not all examples, the one or more first audio characteristics include audio characteristics defined by an audio classifier, and at least some of the audio signals in the first subset of audio signals convey speech and the audio signals in the second subset do not convey speech.
[0021] In some but not all examples, the multi-channel audio signal includes multiple audio signals, where each audio signal is used to render audio via a different output channel.
[0022] In some but not all examples, the count of the first subset depends on the available bandwidth.
[0023] In some but not all examples, analyzing the remaining audio signals in the second subset of audio signals to determine the transmit audio signal and the metadata includes analyzing the second subset of audio signals instead of the first subset of audio signals.
[0024] In some but not all examples, the metadata parameterizes a time-frequency portion of the second audio signal subset.
[0025] In some but not all examples, the metadata encodes at least a spatial energy distribution of a sound field defined by the second subset of the audio signal.
[0026] In some examples, the analysis is a parametric spatial analysis that produces both parametric and spatialized metadata, wherein the parametric spatial analysis parameterizes the time-frequency portion of the second audio signal subset and at least partially encodes a spatial energy distribution of a sound field defined at least by the second audio signal subset.
[0027] In some but not all examples, the metadata encodes at least a spatial energy distribution of a sound field defined by the second subset of the audio signal.
[0028] In some but not all examples, the apparatus includes means for providing control information that identifies at least an audio signal of the plurality of audio signals that is included in the first subset of audio signals.
[0029] In some but not all examples, the control information identifies at least a processed audio signal produced by the analysis.
[0030] In some, but not all, examples, analysis of the second audio signal subset provides one or more processed audio signals and metadata, wherein the one or more processed audio signals and metadata are jointly encoded with the first audio signal subset or the one or more processed audio signals and metadata are jointly encoded but separately from the first audio signal subset.
[0031] According to various, but not all, embodiments, there is provided a method of encoding a multi-channel audio signal, comprising:
[0032] identifying at least one audio signal to separate from the multi-channel audio signal;
[0033] Based on the identified at least one audio signal, separate the plurality of audio signals into at least a first subset of audio signals and a second subset of audio signals, wherein the first subset includes the identified at least one audio signal and the second subset includes remaining audio signals in the received multi-channel audio signal;
[0034] analyzing remaining audio signals in the second subset of audio signals to determine one or more transmitted audio signals and metadata; and
[0035] The at least one audio signal, the one or more transmission audio signals, and the metadata are encoded.
[0036] According to various, but not all, embodiments, there is provided a computer program comprising program instructions for causing an apparatus to at least:
[0037] identifying at least one audio signal to separate from the multi-channel audio signal;
[0038] Based on the identified at least one audio signal, separate the plurality of audio signals into at least a first subset of audio signals and a second subset of audio signals, wherein the first subset includes the identified at least one audio signal and the second subset includes remaining audio signals in the received multi-channel audio signal;
[0039] analyzing remaining audio signals in the second subset of audio signals to determine one or more transmitted audio signals and metadata; and
[0040] Encoding of the at least one audio signal, the one or more transmission audio signals, and the metadata is enabled.
[0041] According to various, but not all, embodiments, there is provided an apparatus comprising means for:
[0042] receiving encoded data for decoding, the encoded data comprising at least one audio signal, one or more transport audio signals, and metadata;
[0043] decoding the received encoded data to decode the at least one audio signal, the one or more transport audio signals, and the metadata;
[0044] synthesizing the decoded one or more transmitted audio signals and the decoded metadata to provide a set of audio signals;
[0045] identifying a multi-channel index of the at least one audio signal and / or the group of audio signals; and
[0046] At least the decoded at least one audio signal and the group of audio signals are combined using the index to provide a multi-channel audio signal.
[0047] According to various, but not all, embodiments, there is provided a method comprising:
[0048] receiving encoded data for decoding, the encoded data comprising at least one audio signal, one or more transport audio signals, and metadata;
[0049] decoding the received encoded data to decode the at least one audio signal, the one or more transport audio signals, and the metadata;
[0050] synthesizing the decoded one or more transmitted audio signals and the decoded metadata to provide a set of audio signals;
[0051] identifying a multi-channel index of the at least one audio signal and / or the group of audio signals; and
[0052] At least the decoded at least one audio signal and the group of audio signals are combined using the index to provide a multi-channel audio signal.
[0053] According to various, but not all, embodiments, there is provided a computer program comprising program instructions for causing an apparatus to at least:
[0054] decoding the received encoded data comprising at least one audio signal, one or more transmission audio signals, and metadata to decode the at least one audio signal, one or more transmission audio signals, and metadata;
[0055] synthesizing the decoded one or more transmitted audio signals and the decoded metadata to provide a set of audio signals;
[0056] identifying a multi-channel index of the at least one audio signal and / or the group of audio signals; and
[0057] At least the decoded at least one audio signal and the group of audio signals are combined to provide a multi-channel audio signal.
[0058] According to various, but not all, embodiments, there is provided an apparatus comprising means for:
[0059] receiving a multi-channel audio signal for rendering spatial audio via a plurality of output channels, the multi-channel audio signal comprising a plurality of audio signals, wherein each audio signal is for rendering audio via a different output channel;
[0060] separating the plurality of audio signals into at least a first subset of audio signals and a second subset of audio signals;
[0061] performing analysis on the second audio signal subset instead of the first audio signal subset to provide a spatially encoded second audio signal subset; and
[0062] At least the first audio signal subset is encoded to provide an encoded first audio signal subset.
[0063] According to various, but not all, embodiments, there is provided a method comprising:
[0064] Changing audio encoding of a multi-channel audio signal for rendering spatial audio via a plurality of output channels, wherein the multi-channel audio signal comprises a plurality of audio signals, wherein each audio signal is used to render audio via a spatial output channel, comprising:
[0065] selecting a first subset of the plurality of audio signals and selecting a second subset of the plurality of audio signals;
[0066] performing analysis on the second subset of audio signals instead of the first subset of audio signals; and
[0067] A first subset of the plurality of audio signals is individually encoded.
[0068] According to various, but not all, embodiments, there is provided a computer program comprising program instructions for causing an apparatus to at least:
[0069] selecting a first subset and a second subset of the plurality of audio signals for use in rendering spatial audio via a plurality of output channels, wherein the multi-channel audio signal comprises a plurality of audio signals, wherein each audio signal is used to render audio via a spatial output channel;
[0070] performing analysis on the second subset of audio signals instead of the first subset of audio signals;
[0071] Encoding of a first subset of the plurality of audio signals is enabled.
[0072] According to various, but not all, embodiments, there is provided an apparatus comprising means for:
[0073] decoding the encoded first audio signal subset to generate a first audio signal subset;
[0074] decoding the spatially encoded second audio signal subset to generate a second audio signal subset;
[0075] The first subset of audio signals and the second subset of audio signals are combined to synthesize a plurality of audio signals for rendering spatial audio via a plurality of output channels, wherein each audio signal is used to render audio via a different output channel.
[0076] According to various, but not all, embodiments, there is provided a method comprising:
[0077] decoding the encoded first audio signal subset to generate a first audio signal subset;
[0078] decoding the spatially encoded second audio signal subset to generate a second audio signal subset;
[0079] The first subset of audio signals and the second subset of audio signals are combined to synthesize a plurality of audio signals for rendering spatial audio via a plurality of output channels, wherein each audio signal is used to render audio via a different output channel.
[0080] According to various, but not all, embodiments, there is provided a computer program comprising program instructions for causing an apparatus to at least:
[0081] decoding the encoded first audio signal subset to generate a first audio signal subset;
[0082] decoding the spatially encoded second audio signal subset to generate a second audio signal subset;
[0083] The first subset of audio signals and the second subset of audio signals are combined to synthesize a plurality of audio signals for rendering spatial audio via a plurality of output channels, wherein each audio signal is used to render audio via a different output channel.
[0084] According to various, but not all, embodiments, there is provided an apparatus comprising means for:
[0085] receiving a multi-channel audio signal for rendering spatial audio via a plurality of output channels, the multi-channel audio signal comprising a plurality of audio signals, wherein each audio signal is for rendering audio via a different output channel;
[0086] separating the plurality of audio signals into at least a first subset of audio signals and a second subset of audio signals;
[0087] providing a first coding path for encoding the first subset of audio signals and a second, different coding path for encoding the second subset of audio signals,
[0088] Wherein, the second encoding path, rather than the first encoding path, includes performing analysis.
[0089] According to various, but not all, embodiments, there is provided a method comprising:
[0090] Audio encoding a multi-channel audio signal for rendering spatial audio via a plurality of output channels, wherein the multi-channel audio signal comprises a plurality of audio signals, wherein each audio signal is used for rendering audio via a spatial output channel, comprising:
[0091] selecting a first subset of the plurality of audio signals and selecting a second subset of the plurality of audio signals;
[0092] providing a first coding path for encoding the first subset of audio signals and a second, different coding path for encoding the second subset of audio signals,
[0093] Wherein, the second encoding path, rather than the first encoding path, includes performing analysis.
[0094] According to various, but not all, embodiments, there is provided a computer program comprising program instructions for causing an apparatus to at least:
[0095] selecting a first subset and a second subset of the plurality of audio signals for use in rendering spatial audio via a plurality of output channels, wherein the multi-channel audio signal comprises a plurality of audio signals, wherein each audio signal is used to render audio via a spatial output channel;
[0096] providing a first coding path for encoding the first subset of audio signals and a second, different coding path for encoding the second subset of audio signals,
[0097] Wherein, the second encoding path, rather than the first encoding path, includes performing analysis.
[0098] According to various, but not all, embodiments, an apparatus is provided that includes means for:
[0099] receiving a multi-channel audio signal for rendering spatial audio via a plurality of output channels, the multi-channel audio signal comprising a plurality of audio signals, wherein each audio signal is for rendering audio via a different output channel;
[0100] separating the plurality of audio signals into at least a first subset of audio signals and a second subset of audio signals;
[0101] providing a first coding path for encoding the first subset of audio signals and a second, different coding path for encoding the second subset of audio signals,
[0102] wherein the second encoding path, rather than the first encoding path, comprises performing analysis, wherein the first encoding path (after analysis) and the second encoding path use a joint encoder, or wherein the first encoding path (after analysis) and the second encoding path use separate encoders.
[0103] According to various, but not all, embodiments, examples are provided as claimed in the following claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] Some examples will now be described with reference to the accompanying drawings, in which:
[0105] Figure 1 An example of the subject matter described herein is shown;
[0106] Figure 2 Another example of the subject matter described herein is shown;
[0107] Figure 3 Another example of the subject matter described herein is shown;
[0108] Figure 4 Another example of the subject matter described herein is shown;
[0109] Figure 5 Another example of the subject matter described herein is shown;
[0110] Figure 6 Another example of the subject matter described herein is shown;
[0111] Figure 7 Another example of the subject matter described herein is shown;
[0112] Figure 8 Another example of the subject matter described herein is shown;
[0113] Figure 9 Another example of the subject matter described herein is shown;
[0114] Figure 10 Another example of the subject matter described herein is shown;
[0115] Figure 11 Another example of the subject matter described herein is shown;
[0116] Figure 12 Another example of the subject matter described herein is shown;
[0117] Figure 13 Another example of the subject matter described herein is shown;
[0118] Figure 15 Another example of the subject matter described herein is shown;
[0119] Figure 16 Another example of the subject matter described herein is shown. DETAILED DESCRIPTION
[0120] Figure 1 An example of an apparatus 100 is shown. The apparatus 100 is an audio encoder apparatus configured to encode a multi-channel audio signal 110.
[0121] The apparatus 100 is configured to receive a multi-channel audio signal 110. In at least some examples, the received multi-channel audio signal 110 is a multi-channel audio signal 110 for rendering spatial audio via multiple output channels. In at least some examples, the multi-channel audio signal 110 includes multiple audio signals 110, and each audio signal 110 is used to render audio via a different output channel.
[0122] The device 100 includes circuitry for performing functions. These functions include:
[0123] At block 130 , the plurality of audio signals 110 are separated into at least a first subset 111 of the audio signals 110 and a second subset 112 of the audio signals 110 ;
[0124] At block 150 , before providing subsequent encoding of the second subset 122 of the encoded audio signal 110 , performing analysis 152 on the second subset 112 of the audio signal 110 instead of the first subset 111 of the audio signal 110 ; and
[0125] At block 140 , at least a first subset 111 of the audio signal 110 is encoded to provide a first encoded subset 121 of the audio signal 110 .
[0126] The apparatus 100 provides a first encoding path 101 for encoding a first subset 111 of the audio signal 110 and a second, different encoding path 103 for encoding a second subset 112 of the audio signal 110. The second encoding path 103, instead of the first encoding path 101, comprises performing an analysis 152.
[0127] Although in this example, the encoding of the first subset 111 of the audio signal 110 is shown as being performed separately from the second subset 112 of the audio signal 110, in other examples, as will be described subsequently, after the analysis 152 of the second subset 112 of the audio signal, joint encoding of the analyzed second subset 112 of the audio signal 110 and the first subset 111 of the audio signal 110 may occur.
[0128] Multi-channel audio signal
[0129] In some but not all examples, the multi-channel audio signal 110 includes multiple audio signals 110, and each audio signal 110 is configured to render audio via a different speaker channel. Examples of these multi-channel audio signals 110 include 5.1, 5.1+2, 5.1+4, 7.1, 7.1+4, etc.
[0130] In some but not all examples, the multi-channel audio signal 110 includes a plurality of audio signals 110, and each audio signal 110 represents a virtual microphone. Examples of these multi-channel audio signals 110 may include higher-order ambisonics.
[0131] For example, the multi-channel audio signal 110 may be received after being converted from a different spatial audio format, such as an object-based audio format.
[0132] For example, the multi-channel audio signal 110 may be received after being accessed from a memory storage by the apparatus 100 or after being transmitted to the apparatus 100 .
[0133] fixed
[0134] In some examples, the apparatus 100 has fixed (non-adaptive) operation and is configured to separate 130 the plurality of audio signals 110 in the same manner over time. The separation can be permanently fixed or temporarily fixed. If temporarily fixed, it can be fixed by the user. It does not adapt based on the content of the plurality of audio signals 110.
[0135] For example, in some but not all examples, the apparatus 100 separates 130 the plurality of audio signals 110 into at least a first subset 111 of the audio signals 110 and a second subset 112 of the audio signals 110 in a fixed manner, i.e., the first subset 111 of the audio signals 110 is a fixed subset of the plurality of audio signals 110, and the second subset 112 of the audio signals 110 is a fixed subset of the plurality of audio signals 110.
[0136] The first subset 111 may include a single audio signal, for example, a center speaker channel signal. The first subset may include a pair of audio signals, for example, a pair of stereo channel signals.
[0137] The first subset 111 may include one or more dominant speech audio channel signals, or other source-dominated audio signals that are dominated by one or more audio sources and best capture the one or more sources (which may be, for example, a lead instrument, vocals, or some other type of audio source).
[0138] Adaptive
[0139] In some examples, apparatus 100 has adaptive operation and is configured to dynamically separate 130 multiple audio signals 110, i.e., to separate them in different ways over time. The separation is adaptive because apparatus 100 controls the adaptation itself. For example, apparatus 100 can adapt separation 130 of multiple audio signals 110 based on the content of multiple audio signals 110.
[0140] For example, in some but not all examples, the apparatus 100 adaptively (over time) separates 130 the plurality of audio signals into at least a first subset 111 of the audio signals 110 and a second subset 112 of the audio signals 110 , wherein the first subset 111 of the audio signals 110 is a variable subset of the plurality of audio signals 110 and the second subset 112 of the audio signals 110 is a variable subset of the plurality of audio signals 110 .
[0141] The subset 111 of the audio signals 110 may be changed by changing the count (number) of the first subset 111 of the audio signals 110. The first subset 111 may include a single audio signal 110, a pair of audio signals 110, or more audio signals 110.
[0142] The subset 111 of the audio signals 110 may be changed by changing the composition (identity) of the first subset 111 of the audio signals 110. The first subset 111 may, for example, be mapped to different combinations of the plurality of audio signals 110.
[0143] In some, but not all, examples, separation 130 of audio signal 110 depends on available bandwidth. For example, the number of first subset 111 of audio channels and / or the composition of first subset 111 of audio channels 110 can depend on available bandwidth. Apparatus 100 can adapt to changes in available bandwidth, for example, by adapting separation 130 of audio signal 110.
[0144] As an example, the multi-channel audio signal 110 may have a 7.1 surround sound format. There are seven audio signals 110, of which one audio channel is a center audio signal 110. The following table shows some examples of how the count of the first subset 111 can be changed. The following table shows how the bandwidth allocated to the first subset 111 of audio channels 110 can be changed. The table shows how the division of the available bandwidth between the first subset 111 of audio signals 110 and the second subset 112 of audio signals 110 can be changed.
[0145]
[0146] In some examples, there may be a minimum bandwidth for each audio signal 110 in the first subset 111. In some examples, a suitable minimum bandwidth may be 9.6 kbps or 10 kbps.
[0147] In some examples, there may be a minimum bandwidth for the second subset 112 of the audio signals 110. In some examples, a suitable minimum bandwidth may be 20 kbps.
[0148] The first subset 111 of the audio signal 110 may be encoded at a variable bit rate per audio signal. Alternatively or additionally, the second subset 112 of the audio signal 110 may be encoded at a variable bit rate. The bit rate distribution between the first subset 111 and the second subset 112 may be controlled so as to achieve optimal perceptual quality.
[0149] Figure 2An example of a method 300 that may be performed by the apparatus 100 is shown. The method 300 alters audio coding of a multi-channel audio signal 110 for rendering spatial audio via a plurality of output channels. The multi-channel audio signal 110 comprises a plurality of audio signals 110, each for rendering audio via a spatial output channel.
[0150] The method includes, at block 302 , selecting 302 a first subset 111 of the plurality of audio signals 110 and selecting 302 a second subset 112 of the plurality of audio signals 110 .
[0151] The method includes, at block 306 , performing analysis on the second subset 112 of the audio signal 110 instead of the first subset 111 of the spatial audio signal 110 .
[0152] The method includes, at block 304 , encoding at least a first subset 111 of the plurality of audio signals 110 .
[0153] In some examples, the first subset 111 of the plurality of audio signals 110 is encoded separately from the second subset 112 of the plurality of audio signals 110. In some examples, after the second subset 112 of the audio signals 110 is analyzed, the first subset 111 of the plurality of audio signals 110 is jointly encoded with the second subset 112 of the plurality of audio signals 110.
[0154] Figure 3 Shown is an example of an apparatus 200. The apparatus 200 is an audio decoder apparatus configured to decode a first subset 121 of the encoded audio signal 110 and a second subset 122 of the encoded audio signal 110 to synthesize a multi-channel audio signal 110'.
[0155] The apparatus 200 includes circuitry for performing functions.
[0156] The apparatus 200 decodes 240 the first subset 121 of the encoded audio signal 110 to generate a first subset 111 ′ of the audio signal 110 .
[0157] The apparatus 200 decodes 250 the second subset 122 of the encoded audio signal 110 to generate a second subset 112 ′ of the audio signal 110 .
[0158] The first subset 111 ′ and the second subset 112 ′ of the audio signals 110 are combined to synthesize a plurality of audio signals 110 ′ for rendering spatial audio via a plurality of output channels, wherein each audio signal 110 ′ is used to render audio via a different output channel.
[0159] Figure 4 An example of a method 310 that may be performed by apparatus 200 is shown.
[0160] The method 310 comprises, at block 312 , decoding the first subset 121 of the encoded audio signal 110 to produce a first subset 111 ′ of the audio signal 110 .
[0161] The method 310 includes, at block 314 , decoding the second subset 122 ′ of the encoded audio signal 110 to produce the second subset 112 ′ of the audio signal 110 .
[0162] The method 310 includes, at block 316 , combining a first subset 111 ′ of the audio signal 110 and a second subset 112 ′ of the audio signal 110 to synthesize a plurality of audio signals 110 ′ for rendering spatial audio via a plurality of output channels, wherein each audio signal 110 ′ is for rendering audio via a different output channel.
[0163] As about Figure 1 and 2 As described, separating 130 the audio signal 110 into the first subset 111 and the second subset 112 can be based on evaluating a criterion. The criterion can be, for example, a simple single criterion, or a logical criterion that uses Boolean logic to define a more complex conditional statement as a criterion. Thus, the criterion can depend on one or more parameters.
[0164] exist Figure 5 In the example shown, a first subset 111 of the audio signals 110 are signals determined at block 132 to satisfy the criterion, while a second subset 112 of the audio signals 110 are signals determined at block 132 to not satisfy the criterion.
[0165] In some examples, the evaluation of the audio signal 110 at block 132 is frequency independent (wideband). In other examples, the evaluation of the audio signal 110 at block 132 is frequency dependent, and the audio signal 110 is transformed 134 from the time domain to the frequency domain prior to the evaluation of the criterion at block 132.
[0166] The first criterion may, for example, depend on one or more audio characteristics of the audio signals 110. Thus, in some examples, the first subset 111 of audio signals 110 shares one or more audio characteristics, while the second subset 112 of audio signals 110 does not share the one or more first audio characteristics.
[0167] The first criterion may depend on one or more spectral characteristics of the audio signals 110. Thus, in some examples, at least some of the first subset 111 of audio signals 110 share one or more spectral characteristics, while the second subset 112 of audio signals 110 does not share the one or more spectral characteristics.
[0168] The first criterion may depend on both audio characteristics and spectral characteristics.For example, the first subset 111 of audio signals may share audio characteristics within a first frequency range, while the second subset 112 of audio signals 110 does not share audio characteristics within the first frequency range.
[0169] In some examples, the one or more audio characteristics include an energy level of audio signal 110. Thus, in some examples, each audio signal in first subset 111 of audio signals 110 has an energy level greater than any audio signal in second subset 112 of audio signals 110. In some examples, each audio signal in first subset 111 of audio signals 110 has an energy level greater than any audio signal in second subset 112 of audio signals 110 and, in addition, greater than a threshold. In some examples, the energy level is determined only within one or more defined frequency bands. For example, the defined frequency bands may correspond to human speech.
[0170] In some examples, the one or more audio characteristics identify dialogue or other prominent audio, such that the first subset 111 includes dialogue / most prominent audio signals 110 .
[0171] In some examples, the one or more first audio characteristics include audio signal correlation. Thus, in some examples, each audio signal in the first subset 111 of audio signals 110 has a greater cross-correlation with the audio signals 110 in the first subset than with the audio signals 110 in the second subset. This can occur, for example, when a prominent audio content is simultaneously present on multiple channels. Thus, prominence is caused by a wider spatial distribution compared to other audio content.
[0172] In some examples, the one or more first audio characteristics include audio signal decorrelation. Thus, in some examples, at least some of the first subset 111 of audio signals 110 have low cross-correlation with both other audio signals 110 in the first subset and with audio signals 110 in the second subset. This can occur, for example, when prominent audio content is confined to a single channel. Thus, the prominence arises from a narrower spatial distribution compared to the other audio content.
[0173] In some examples, the one or more first audio characteristics include audio characteristics defined by an audio classifier. The audio classifier can, for example, be configured to classify sound sources. Thus, the audio classifier can identify audio signals 110 that include (dominant) human speech, or musical instruments, or speech, or singing, or some other type of audio source. Thus, at least some of the first subset 111 of audio signals 110 can convey a specific sound source, wherein the audio signals 110 of the second subset 112 do not convey the specific sound source.
[0174] Figure 6 An example of a more detailed method for evaluating a criterion for separating 130 the audio signal 110 into the first subset 111 and the second subset 112 is shown.
[0175] The input of the method is a multichannel signal s(i, m), where i is the index of the audio signal 110 for the channel and m is the time index. First, at block 171, the signal 110 is transformed from the time domain to the time-frequency domain. This can be performed using, for example, a short-time Fourier transform (STFT) or a complex quadrature mirror filter bank (QMF). The resulting time-frequency domain signal is denoted as S(i, b, n), where b is the frequency bin index and n is the time frame index.
[0176] At block 172, the energy E(i,k,n) of the time-frequency domain input signal S(i,b,n) is estimated in the frequency band:
[0177]
[0178] Where k is the frequency band index, b k,low is the lowest bin of the band, b k,high It is the highest warehouse.
[0179] At optional block 173, the energy E(i, k, n) can be weighted with frequency-dependent weighting, for example, to focus more on certain frequencies, such as the speech frequency range. As another example, weighting can be applied to simulate the loudness perception of human hearing. The weighting can be performed by the following formula:
[0180] E w (i,k,n)=E(i,k,n)w(k)
[0181] Among them, w(k) is the weighting function.
[0182] At block 174, the weighted energies are summed across the frequency band to obtain a wideband estimate:
[0183]
[0184] E w,sm (i, n) = aE w (i, n) + bE w,sm (i, n-1)
[0185] Where a and b are smoothing coefficients (eg, a=0.01 and b=1-a).
[0186] Next, at block 176, the ratio of the energy of the audio signal 110 of some channel i to the total energy of all channels is calculated:
[0187]
[0188] Finally, at block 178, r(i, n) is used to select the index i of the audio signal 110 to be separated into the first subset 111. These indices can be provided as control information 180 for use in separating the plurality of audio signals 110 into the first subset 111 of audio signals 110 and the second subset 112 of audio signals 110. As an example, the audio signal 110 with the largest ratio r(i, n) can be selected. As another option, if the largest ratio r(i, n) is above a certain threshold τ (e.g., τ=0.2), then the audio signal 110 with the largest ratio can be selected. If the largest audio signal 110 is below the threshold, no channels will be separated into the first subset 111.
[0189] As another example, more than one audio signal 110 may be selected to be separated into the first subset 111. For example, the two audio signals with the largest ratio r(i, n) may be selected. The selection may also be "paired," so that the audio signals 110 for symmetric channels (e.g., left front and right front) are considered together (so as not to interfere with the stereo image). In this case, both audio signals 110 for the symmetric channels may need to have a ratio r(i, n) above a threshold τ.
[0190] As another example, if the audio signal 110 of the center channel has a ratio r(i,n) above a threshold, it is separated into the first subset 111 .
[0191] Therefore, the audio signal 110 to be separated into the first subset 111 can be flexibly selected, and there are multiple selection methods.
[0192] The selection made (whether fixed or flexible) needs to be known at the decoder, which identifies the multichannel index of the first subset 111 and / or the second subset 112 or otherwise defines the relationship or mapping of the first subset 111 and / or the second subset 112 of the audio signal 110 to the multichannel audio signal.
[0193] The selection may depend on the available bit rate. For example, when a higher bit rate is available, on average more of the audio signal 110 can be separated into the first subset.
[0194] Figure 7 An example of the previously described apparatus 100 is shown. Like reference numerals are used to describe like components and functions.
[0195] The device 100 includes circuitry for performing functions. These functions include:
[0196] identifying 132 at least one audio signal to separate from the multi-channel audio signal;
[0197] Based on the identified at least one audio signal, separate 130 the plurality of audio signals 110 into at least a first subset 111 of the audio signals 110 and a second subset 112 of the audio signals 110, wherein the first subset 111 includes the identified at least one audio signal and the second subset includes remaining audio signals in the received multi-channel audio signal 110;
[0198] analyzing 152 the remaining audio signals in the second subset 112 of the audio signals 110 to determine one or more transmitted audio signals 151 and metadata 153 ; and
[0199] The at least one identified audio signal in the first subset 111 , the one or more transmission audio signals 151 and the metadata 152 are encoded 140 , 154 .
[0200] exist Figure 7 Features shown in include: blocks 132, 133 within block 130 show blocks for logical separation 132 and physical separation 133 of the audio signal 110; blocks 152, 154 within block 150 show analysis 152 and encoding 154 of the second subset 112 of the audio signal 110; a multiplexer 160 combines not only the encoded first audio signal subset 121 and the encoded second audio signal subset 122, but also combines control information 180 from block 132 to form a data stream 161.
[0201] Block 152 performs analysis on the second subset 112 of the audio signal 110 instead of the first subset 111 of the audio signal 110 to provide one or more processed (transmitted) audio signals 151 and metadata 153. The provided one or more processed (transmitted) audio signals 151 and metadata 153 are encoded at block 154 to provide a second subset 122 of the encoded audio signal 110.
[0202] Processing 152 the audio signal 110 to form the processed audio signal 151 may, for example, include down-condensing or selecting. The processed audio signal 151 for transmission may, for example, be a down-condensing of some or all of the audio signals in the second subset 112 of the audio signals 110. Alternatively, the processed audio signal 151 for transmission may, for example, be a selected subset of the audio signal 110 in the second subset 112 of the audio signals 110.
[0203] In some but not all examples, block 152 performs spatial audio encoding. For example, block 152 may include one or more metadata assisted spatial audio (MASA) codecs, or analyzers, or processors or pre-processors. The MASA codec generates two processed audio signals 151 for transmission.
[0204] In some, but not all, examples, the metadata 153 parameterizes the time-frequency portion of the second subset 112 of the audio signal 110. For example, in some examples, the metadata 153 encodes at least the spatial energy distribution of the sound field defined by the second subset 112 of the audio signal 110.
[0205] Metadata 153 may, for example, encode one or more of the following parameters:
[0206] Direction index defining the direction of the sound;
[0207] Provides a direction / energy (ratio) of the energy ratio for the direction specified by the direction index (e.g., energy in the direction / total energy);
[0208] Sound field information;
[0209] Coherence information (such as spread coherence and surrounding coherence);
[0210] Diffuseness information;
[0211] distance.
[0212] These parameters may be provided in the time-frequency domain.
[0213] The metadata 153 for metadata-assisted spatial audio may use one or more of the following parameters:
[0214] i) Direction index: The direction of arrival of the sound at the time-frequency parameter interval. Expressed in a sphere with approximately 1 degree accuracy;
[0215] ii) Direct-to-total energy ratio: The energy ratio for a direction index (i.e., time-frequency subframe). It is calculated as energy in the direction / total energy;
[0216] iii) Spread coherence: Energy spread for a directional index (ie, time-frequency subframe), defining the direction to be reproduced as a point source or the direction to be coherently reproduced around it.
[0217] iv) Diffuse-to-total energy ratio: The energy ratio of non-directional sound in the surrounding directions. It is calculated as the energy of non-directional sound / total energy
[0218] v) Surround coherence: the coherence of non-directional sound in the surrounding directions;
[0219] vi) Remainder-to-total energy ratio: The energy ratio of the remaining (such as microphone noise) sound energy, calculated as the energy of the remaining sound / the total energy in order to meet the requirement of the sum of energy ratios;
[0220] vii) Distance: The distance of the sound from the direction index (ie, time-frequency subframe) in meters on a logarithmic scale.
[0221] The function of separating 130 the audio channels 110 includes: a sub-block 132 for determining a logical separation of the audio channels 110 into a first subset 111 and a second subset 112; and a sub-block 133 for physically separating the audio channels 110 into a first encoding path 101 for the first subset 111 of the audio signal 110 and a second encoding path 103 for the second subset 112 of the audio signal 110.
[0222] In some, but not all, examples, sub-block 132 analyzes multiple audio signals 110. For example, as previously described, it determines whether the received audio signals 110 meet a criterion. Sub-block 133 can logically separate the audio signals 110 into a first subset 111 and a second subset 112. For example, first subset 111 of audio signals 110 is determined to meet the criterion, while second subset 112 of audio signals 110 is determined (explicitly or implicitly) to not meet the criterion.
[0223] Sub-block 132 generates control information 180 that identifies at least a logical separation of audio signals 110 into first and second subsets 111 and 112. Control information 180 identifies at least audio signals of the plurality of audio signals 110 that are included in first subset 111 of audio signals 110.
[0224] In some examples, control information 180 identifies at least processed audio signal 151 produced by analysis 152 .
[0225] In some examples, control information 180 identifies at least metadata, eg, identifying a type of analysis or parameters to use for the analysis.
[0226] Figure 8 Shown for Figure 7The decoder device 200 is used together with the encoder device 100 shown in FIG. Figure 8 An example of the previously described apparatus 200 is shown. Like reference numerals are used to describe like components and functions.
[0227] The apparatus 200 is an audio decoder apparatus configured to decode a first subset 121 of the encoded audio signal 110 and a second subset 122 of the encoded audio signal 110 to synthesize a multi-channel audio signal 110 ′.
[0228] The apparatus 200 includes circuitry for performing functions. These functions include:
[0229] receiving, for decoding, encoded data 161 comprising at least one audio signal 111 , one or more transport audio signals 151 , and metadata 153 ;
[0230] decoding 240 , 250 the received encoded data 161 to provide a decoded at least one audio signal 111 ′ as a first subset 111 ′ of the audio signal 110 ′, decoded one or more transmitted audio signals 151 ′ and decoded metadata 153 ′;
[0231] synthesizing 254 the decoded one or more transmitted audio signals 151 ′ and the decoded metadata 153 ′ to provide a second audio signal subset 112 ′;
[0232] identifying a multi-channel index of the at least one audio signal and / or the set of audio signals; and
[0233] At least the decoded at least one audio signal 111 ′ (the first subset) and the second subset of audio signals 112 ′ are combined 230 to provide a multi-channel audio signal 110 ′.
[0234] exist Figure 8 The features shown in the include:
[0235] The demultiplexer 210 recovers the first subset 121 of the encoded audio signal, the second subset 122 of the encoded audio signal 110 and the control information 180 from the received data stream 161 ;
[0236] decoding 240 the first subset 121 of the encoded audio signal to provide at least one audio signal as a first subset 111 ′ of the audio signal 110 ′;
[0237] Blocks 252 , 254 within block 250 illustrate: decoding 252 and synthesis 254 of the second subset 122 of the encoded audio signal 110 to recover the second subset 112 ′ of the audio signal 110 ; and combining 230 the first subset 111 ′ of the audio signal 110 and the second subset 112 ′ of the audio signal 110 to synthesize a plurality of audio signals 110 ′ depending on the received control information 180 .
[0238] The second subset 122 of the encoded audio signal 110 is decoded at block 252 to provide one or more processed (transmitted) audio signals 151 ′ and metadata 153 ′.
[0239] The block 254 performs synthesis on the processed (transmitted) audio signal 151 ′ and the metadata 153 ′ to synthesize the second subset 112 ′ of the audio signal 110 .
[0240] In some but not all examples, block 254 includes one or more Metadata Assisted Spatial Audio (MASA) codecs, or synthesizers, or renderers, or processors.The MASA codec decodes both the processed audio signal 151 and the metadata 153 for transmission.
[0241] The function of combining 230 the first subset 111′ of the audio signals 110 and the second subset 112′ of the audio signals 110 to synthesize the plurality of audio signals 110′ may depend on the received control information 180. The control information 180 defines a logical separation of the audio channels 110 into the first subset 111 and the second subset 112. The control information may, for example, identify a multi-channel index of at least one audio signal and / or a set of audio signals.
[0242] In some examples, the control information 180 identifies at least the processed audio signal 151 produced by the analysis 152. In this example, the control information 180 is provided to block 254.
[0243] In some examples, the control information 180 identifies at least the metadata 153 , for example, identifying a type of analysis or parameters for the analysis. In this example, the control information 180 is provided to block 254 .
[0244] exist Figure 7 In the example of , the analysis 152 of the second subset 112 of the audio signal 110 instead of the first subset 111 of the audio signal 110 provides one or more processed audio signals 151 and metadata 153. Figure 7In the example of , the one or more processed audio signals 151 and the metadata 153 are not jointly encoded with the first subset 111 of the audio signal 110. The first encoding path 101 for the first subset 111 of the audio signal 110 and the second encoding path 103 for the second subset 112 of the audio signal 110 are rejoined at the multiplexer 160.
[0245] Figure 9 The device 100 shown in FIG. Figure 7 The apparatus 100 shown in FIG is similar to FIG. However, in FIG. Figure 9 In FIG, one or more processed audio signals 151 and metadata 153 are jointly encoded with a first subset 111 of the audio signal 110 at a joint encoder 190. The first encoding path 101 for the first subset 111 of the audio signal 110 and the second encoding path 103 for the second subset 112 of the audio signal 110 are rejoined at the joint encoder 190. The joint encoder 190 replaces Figure 7 Blocks 140, 154 in.
[0246] Figure 10 An example of a joint encoder 190 is shown. In the joint encoder 190, possible interdependencies between the first set 111 of audio signals 110 and the processed (transmitted) audio signal 151 may be taken into account when encoding them.
[0247] The signals in the first set 111 of audio signals 110 and one or more transmitted audio signals 151 are forwarded to a computation block 191. Block 191 combines these signals 111, 151 into one or more downmix signals 194 and a residual signal 192. In addition, prediction coefficients 196 are output. In the decoder, the prediction coefficients 196 and the residual signal 192 can be used to retrieve the original signals 111, 151 from the downmix signal 194. Details of the prediction and residual processing can be found in publicly available literature.
[0248] The residual signal 192 is forwarded to block 193 for encoding. The downmix signal 194 is forwarded to block 195 for encoding. The residual coefficients 196 are forwarded to block 197 for encoding. The metadata 153 is encoded at block 198.
[0249] The encoded residual signal, the encoded downmix signal, the encoded residual coefficients and the encoded metadata 153 are provided to a multiplexer 199 which outputs a data stream comprising a first set 121 of encoded audio signals 110 and a second set 122 of encoded audio signals.
[0250] Figure 11 Shown for Figure 9The decoder device 200 is used together with the encoder device 100 shown in FIG. Figure 11 The device 200 shown in FIG. Figure 8 The apparatus 200 shown in FIG is similar. However, in Figure 11 In the embodiment, the received jointly encoded data stream 121 , 122 comprises a first subset 121 of the encoded audio signal 110 and a second subset 122 of the encoded audio signal 110 .
[0251] The joint decoder 280 decodes the joint coded data stream and creates a first decoding path for the first subset 111' of the audio signal 110 and a second decoding path for the second subset 112' of the audio signal 110. The one or more processed audio signals 151' and metadata 153' are provided by the joint decoder 280 to the block 254 in the second decoding path. The joint decoder 280 replaces Figure 8 Blocks 240, 252 in.
[0252] Figure 12 Shown with Figure 10 An example of a joint decoder 280 corresponding to the joint encoder 190 is shown in Using the joint decoder 280 , the first subset 111 of the audio signal 110 , one or more transmission audio signals 151 and metadata 153 are generated.
[0253] The data stream comprising the first set 121 of encoded audio signals 110 and the second set 122 of encoded audio signals is demultiplexed at block 270 to provide an encoded residual signal 271 , an encoded downmix signal 273 , encoded residual coefficients 275 and encoded metadata 277 .
[0254] The coded residual signal 271 is forwarded to a block 272 for decoding. This reproduces the residual signal 192.
[0255] The encoded downmix signal 273 is forwarded to block 274 for decoding. This reproduces the downmix signal 194.
[0256] The coded residual coefficients 275 are forwarded to a block 276 for decoding. This reproduces the residual coefficients 196.
[0257] The encoded metadata 277 is forwarded to block 278 for decoding. This reproduces the metadata 153.
[0258] The block 279 processes the downmix signal 194 using the prediction coefficients 196 and the residual signal 192 to reproduce the first set 111 of audio signals 110 and the one or more transmitted audio signals 151 .
[0259] One or more transmission audio signals 151 and metadata 153 are output together with the metadata 153 to Figure 11 Block 254 in .
[0260] Figure 13 The device 200 shown in FIG. Figure 7 Possible interdependencies between the first set 111 of audio signals 110 and the processed (transmitted) audio signal 151 may be considered. In this example, joint processing occurs at block 133 before separation of the audio signals 110.
[0261] Pre-processing begins by determining a first subset 111 of the audio signals 110 at block 132. Control information 180 is provided to block 133. Block 133 first performs pre-processing on the audio signals 110 in the first subset 111 and at least some of the remaining audio signals 110 in the second subset 112.
[0262] For example, if it is determined that the center channel audio signal 110 in the first subset 111 is also coherently present in the left front channel audio signal 110 and the right front channel audio signal 110 , the center channel audio signal 110 may be subtracted from the left front channel audio signal 110 and the right front channel audio signal 110 .
[0263] As another example, prediction and residual processing may be applied between the center channel audio signal 110 and the left front channel audio signal 110 and the right front channel audio signal 110, as shown in FIG. Figure 10 described.
[0264] The preprocessing results in a modified multi-channel audio signal 110 and preprocessing coefficients 181 containing information about which preprocessing was applied.
[0265] The block 133 outputs the preprocessing coefficients 181 , the first set 111 of audio signals 110 as one stream, and the second set 112 of audio signals as a second stream.
[0266] The pre-processing coefficients 181 may be provided separately from the control information 180 , or may be provided together with the control information 180 , or may be provided as part of the control information 180 .
[0267] Figure 14 Shown for Figure 13 The decoder device 200 is used together with the encoder device 100 shown in FIG. Figure 14 The device 200 shown in FIG. Figure 8 The apparatus 200 shown in FIG is similar. However, in Figure 14In the embodiment of the present invention, the combination 230 of the first set 111' of audio signals 110 and the second set 112' of audio signals 110 uses coefficients 181 to combine and recover the synthesized original multi-channel signal 110'. The first subset 111 of audio signals and the second subset 112 of audio signals are post-processed before they are combined. The post-processing reverses the pre-processing applied in the encoder. For example, if the pre-processing coefficients 181 indicate that such pre-processing was applied in the encoder, the center channel audio signal 110 can be added back to the left front channel audio signal 110 and the right front channel audio signal 110.
[0268] Figure 15 An example of a controller 500 is shown. The controller may provide the functionality of the encoding apparatus 100 and / or the decoding apparatus 200.
[0269] The controller 500 may be implemented as a controller circuit. The controller 500 may be implemented solely in hardware, with certain aspects of software including only firmware, or may be a combination of hardware and software (including firmware).
[0270] like Figure 15 As shown in , the controller 500 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 506 in a general-purpose or special-purpose processor 502, which can be stored on a computer-readable storage medium (disk, memory, etc.) for execution by such processor 502.
[0271] The processor 502 is configured to read from and write to the memory 504. The processor 502 may also include an output interface through which the processor 502 outputs data and / or commands and an input interface through which data and / or commands are input to the processor 502.
[0272] The memory 504 stores a computer program 506 comprising computer program instructions (computer program code) which, when loaded into the processor 502, controls the operation of the apparatus 100, 200. The computer program instructions of the computer program 506 provide the logic and routines which enable the apparatus to perform Figures 1 to 14 By reading the memory 504 , the processor 502 can load and execute the computer program 506 .
[0273] Thus, the apparatus 100 may include:
[0274] at least one processor 502; and
[0275] at least one memory 504 comprising computer program code,
[0276] The at least one memory 504 and the computer program code are configured to, together with the at least one processor 502, cause the apparatus 100, 200 to at least perform:
[0277] identifying at least one audio signal to separate from the multi-channel audio signal 110;
[0278] Based on the identified at least one audio signal, separate the plurality of audio signals into at least a first subset 111 of the plurality of audio signals and a second subset 112 of the plurality of audio signals, wherein the first subset 111 includes the identified at least one audio signal and the second subset 112 includes remaining audio signals in the received multi-channel audio signal 110;
[0279] analyzing the remaining audio signals in the second subset 112 of audio signals to determine one or more transmitted audio signals 151 and metadata 153; and
[0280] The at least one audio signal, the transmission audio signal 151 and the metadata 153 are enabled to be encoded.
[0281] Thus, the apparatus 200 may include:
[0282] at least one processor 502; and
[0283] at least one memory 504 comprising computer program code,
[0284] The at least one memory 504 and the computer program code are configured to, together with the at least one processor 502, cause the apparatus 100, 200 to at least perform:
[0285] decoding 240 , 250 the received encoded data 160 comprising the at least one audio signal 111 , the one or more transport audio signals 151 , and the metadata 153 to provide a decoded at least one audio signal 111 ′ as the first subset 111 ′ of the audio signal 110 ′, the decoded one or more transport audio signals 151 ′, and the decoded metadata 153 ′;
[0286] synthesizing 254 the decoded one or more transmitted audio signals 151 ′ and the decoded metadata 153 ′ to provide a second audio signal subset 112 ′;
[0287] identifying a multi-channel index of the at least one audio signal and / or the set of audio signals; and
[0288] At least the decoded at least one audio signal 111 ′ (the first subset) and the second subset of audio signals 112 ′ are combined 230 to provide a multi-channel audio signal 110 ′.
[0289] like Figure 16 As shown in FIG, the computer program 506 may arrive at the apparatus 100, 200 via any suitable delivery mechanism 508. The delivery mechanism 508 may be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a storage device, a recording medium such as a compact disc read-only memory (CD-ROM) or a digital versatile disc (DVD) or a solid-state memory, or an article of manufacture that includes or tangibly embodies the computer program 506. The delivery mechanism may be a signal configured to reliably deliver the computer program 506. The apparatus 100, 200 may propagate or transmit the computer program 506 as a computer data signal.
[0290] The computer program 506 may include computer program instructions for causing the apparatus to perform at least the following operations or computer program instructions for performing at least the following operations:
[0291] identifying at least one audio signal to separate from the multi-channel audio signal 110;
[0292] Based on the identified at least one audio signal, separate the plurality of audio signals 110 into at least a first subset 111 of the plurality of audio signals and a second subset 112 of the plurality of audio signals, wherein the first subset 111 includes the identified at least one audio signal and the second subset 112 includes remaining audio signals in the received multi-channel audio signal 110;
[0293] Analyzing the remaining audio signals in the second audio signal subset 112 to determine one or more transmitted audio signals 151 and metadata 153; and
[0294] The at least one audio signal, the transmission audio signal 151 and the metadata 153 are enabled to be encoded.
[0295] The computer program 506 may include program instructions for causing the apparatus to at least perform the following operations:
[0296] decoding 240 , 250 the received encoded data 160 comprising the at least one audio signal 111 , the one or more transport audio signals 151 , and the metadata 153 to provide a decoded at least one audio signal 111 ′ as the first subset 111 ′ of the audio signal 110 ′, the decoded one or more transport audio signals 151 ′, and the decoded metadata 153 ′;
[0297] synthesizing 254 the decoded one or more transmitted audio signals 151 ′ and the decoded metadata 153 ′ to provide a second audio signal subset 112 ′;
[0298] identifying a multi-channel index of the at least one audio signal and / or the group of audio signals; and
[0299] At least the decoded at least one audio signal 111 ′ (the first subset) and the second subset of audio signals 112 ′ are combined 230 to provide a multi-channel audio signal 110 ′.
[0300] The computer program instructions may be included in a computer program, a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not all examples, the computer program instructions may be distributed over more than one computer program.
[0301] Although memory 504 is shown as a single component / circuit, it may be implemented as one or more separate components / circuits, some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cache storage.
[0302] Although processor 502 is shown as a single component / circuit, it may be implemented as one or more separate components / circuits, some or all of which may be integrated / removable.Processor 502 may be a single-core or multi-core processor.
[0303] References to "computer-readable storage medium," "computer program product," "tangibly embodied computer program," etc., or "controller," "computer," "processor," etc., should be understood to encompass not only computers having different architectures, such as single / multi-processor architectures and serial (von Neumann) / parallel architectures, but also specialized circuits, such as field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc., should be understood to encompass software for a programmable processor, or firmware, such as the programmable content of a hardware device, that may include instructions for a processor, or configuration settings for a fixed-function device, gate array, or programmable logic device, etc.
[0304] As used in this application, the term "circuitry" may refer to one or more or all of the following:
[0305] (a) hardware circuit implementation only (such as implementation of analog and / or digital circuits only);
[0306] (b) a combination of hardware circuitry and software such as (if applicable):
[0307] (i) a combination of analog and / or digital hardware circuitry and software / firmware; and
[0308] (ii) any portion of a hardware processor with software (including a digital signal processor, software, and memory that work together to enable a device such as a mobile phone or server to perform various functions); and
[0309] (c) Hardware circuits and / or processors, such as a microprocessor or portion of a microprocessor, that require software (e.g., firmware) to operate, but may not be present when the software is not required for operation.
[0310] This definition of "circuitry" applies to all uses of this term in this application, including in any claims. As another example, as used in this application, the term "circuitry" also covers an implementation of merely a hardware circuit or processor and its accompanying software and / or firmware. The term "circuitry" also covers (for example, and if applicable to the specifically claimed element) a baseband integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or networking equipment.
[0311] Figures 1 to 14 The blocks shown in the figure can represent steps in a method and / or code segments in a computer program 506. The illustration of a specific order of blocks does not imply that there is a required or preferred order for these blocks, but the order and arrangement of the blocks can be varied. In addition, it is possible to omit some blocks.
[0312] At moderate bit rates (e.g., around 128 kbps), the above approach can yield significant perceptible audio quality advantages. This is particularly true for channel-based multi-channel audio with a large number of channels. Separate encoding of one or a few channels, for example, provides a more "stable" image for the main dialog, while simultaneously making the spatial image "wider" because the spatial parameters do not have to "waste" a large portion of the parameter space used to represent the main dialog. The increase in bit rate, if any, is manageable.
[0313] Where a structural feature has been described, it may be replaced by a component that performs the function or functions of the structural feature, whether that function or functions are explicitly described or implicitly described.
[0314] As used herein, a "module" refers to a unit or device excluding certain parts / components added by a terminal manufacturer or a user.
[0315] The device 100 may be a module. The device 200 may be a module.
[0316] The component blocks of the apparatus 100 may be modules. The component blocks of the apparatus 200 may be modules. The controller 500 may be a module.
[0317] As used herein, the term "comprising" has an inclusive, rather than exclusive, meaning. That is, any expression "X comprises Y" means that X may comprise only one Y or may comprise more than one Y. If the exclusive meaning of "comprising" is intended, this will be made clear in the context by reference to "comprising only one..." or by the use of "consisting of..."
[0318] Reference has been made to various examples in this description. The description of features or functions for an example indicates that these features or functions are present in that example. Whether explicitly stated or not, the use of the term "example" or "for example" or "may" or "could" in the text indicates that such feature or function is present in at least the example being described, whether or not described as an example, and that such feature or function may, but need not, be present in some or all other examples. Thus, an "example," "for example," or "may" or "could" refers to a specific instance within a class of examples. A property of an instance may be only a property of that instance or a property of a class of instances or a subclass of that class of instances that includes some but not all of the class of instances. Thus, features described for one example but not for another example are implicitly disclosed to be available for use in other examples as part of a working combination, but are not required to be used in other examples.
[0319] Although the examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims.
[0320] Features described in the preceding description may be used in combinations other than those explicitly described above.
[0321] Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not.
[0322] Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not.
[0323] As used herein, the terms "a," "an," or "the" are intended to be inclusive, rather than exclusive. That is, any reference to "X includes one or the Y" indicates "X may include only one Y" or "X may include more than one Y," unless the context clearly indicates otherwise. If the exclusive meaning of "a," "an," or "the" is intended, the context will clearly indicate this. In some contexts, "at least one" or "one or more" may be used to emphasize the inclusive meaning, but the absence of these terms should not be construed as implying a non-exclusive meaning.
[0324] The presence of a feature (or combination of features) in a claim is a reference to that feature (or combination of features) itself, and also a reference to features that achieve substantially the same technical effect (equivalent features). Equivalent features include, for example, features that are variations and achieve substantially the same result in substantially the same manner. Equivalent features include, for example, features that perform substantially the same function in substantially the same manner to achieve substantially the same result.
[0325] In this specification, reference has been made to various examples using adjectives or adjective phrases to describe characteristics of the examples. Such description of characteristics with respect to the examples means that the characteristics are identical to the described characteristics in some examples and substantially the same as the described characteristics in other examples.
[0326] While an attempt has been made in the foregoing description to identify those features regarded as important, it will be understood that the applicant may seek protection by way of the claims for any patentable feature or combination of features herein before referenced and / or shown in the drawings, whether emphasized or not.
Claims
1. An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: Receive multi-channel audio signals; identifying at least one audio signal to separate from the multi-channel audio signal; Based on the identified at least one audio signal, separate the multi-channel audio signal into at least a first audio signal subset and a second audio signal subset, wherein the first audio signal subset includes the identified at least one audio signal and the second audio signal subset includes remaining audio signals in the received multi-channel audio signal; analyzing the remaining audio signals in the second subset of audio signals to determine one or more transmitted audio signals and metadata; and The at least one audio signal, the one or more transmission audio signals, and the metadata are encoded.
2. The device according to claim 1, wherein The first subset of audio signals is a fixed subset of the multi-channel audio signal, and the second subset of audio signals is a fixed subset of the multi-channel audio signal.
3. The device according to claim 2, wherein The first subset of audio signals includes a center speaker channel signal and / or a pair of stereo channel signals, and / or the first subset of audio signals includes one or more dominant speech audio channel signals.
4. The device according to claim 1, wherein The first audio signal subset is a variable subset of the multi-channel audio signal, and the second audio signal subset is a variable subset of the multi-channel audio signal.
5. The device according to claim 4, wherein The count of the first audio signal subset is variable, and / or wherein the composition of the first audio signal subset is variable.
6. The device according to claim 1, wherein The first subset of audio signals are signals determined to satisfy a first criterion, and the second subset of audio signals are signals determined not to satisfy the first criterion.
7. The device according to claim 6, wherein The first criterion depends on one or more first audio characteristics of the audio signals, wherein the first subset of audio signals share the one or more first audio characteristics and the second subset of audio signals do not share the one or more first audio characteristics.
8. The device according to claim 6, wherein The first criterion depends on one or more spectral characteristics of the audio signals, wherein at least some of the audio signals in the first subset of audio signals share the one or more spectral characteristics and the second subset of audio signals does not share the one or more spectral characteristics.
9. The device according to claim 7, wherein The one or more first audio characteristics include energy levels of audio signals, wherein each audio signal in the first subset of audio signals has an energy level greater than any audio signal in the second subset of audio signals.
10. The device according to claim 7, wherein The one or more first audio characteristics include audio signal correlation, wherein each audio signal in the first audio signal subset has a greater cross-correlation with audio signals in the first audio signal subset than with audio signals in the second audio signal subset, or wherein the one or more first audio characteristics include audio signal decorrelation, wherein at least some audio signals in the first subset of audio signals have low cross-correlation with other audio signals in the first subset of audio signals and with audio signals in the second subset of audio signals, or The one or more first audio characteristics include audio characteristics defined by an audio classifier, wherein at least some of the audio signals in the first subset of audio signals convey speech and audio signals in the second subset of audio signals do not convey speech.
11. The device according to claim 1, wherein The multi-channel audio signal includes a plurality of audio signals, wherein each audio signal is used to render audio via a different output channel.
12. The device according to claim 5, wherein The count of the first subset of audio signals depends on available bandwidth.
13. The device according to claim 1, wherein The apparatus being caused to analyze the remaining audio signals in the second subset of audio signals to determine the transmission audio signal and metadata comprises analyzing the second subset of audio signals instead of the first subset of audio signals.
14. The device according to claim 13, wherein The metadata is configured as at least one of the following: parameterizing the time-frequency portion of the second audio signal subset; and At least a spatial energy distribution of a sound field defined by the second audio signal subset is encoded.
15. The device according to claim 1, wherein The apparatus is caused to provide control information identifying at least at least one of: audio signals of the multi-channel audio signals included in the first audio signal subset; and A processed audio signal is generated by the analysis.
16. The device according to claim 1, wherein The analysis of the second audio signal subset provides one or more processed audio signals and metadata, wherein the one or more processed audio signals and metadata are jointly encoded with the first audio signal subset or the one or more processed audio signals and metadata are jointly encoded but separately from the first audio signal subset.
17. A method for encoding a multi-channel audio signal, comprising: identifying at least one audio signal to separate from the multi-channel audio signal; Based on the identified at least one audio signal, separate the multi-channel audio signal into at least a first audio signal subset and a second audio signal subset, wherein the first audio signal subset includes the identified at least one audio signal and the second audio signal subset includes remaining audio signals in the received multi-channel audio signal; analyzing the remaining audio signals in the second subset of audio signals to determine one or more transmitted audio signals and metadata; and The at least one audio signal, the one or more transmission audio signals, and the metadata are encoded.
18. An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: receiving encoded data for decoding, the encoded data comprising at least one audio signal, one or more transport audio signals, and metadata; decoding the received encoded data to decode the at least one audio signal, the one or more transmission audio signals, and the metadata; synthesizing the decoded one or more transmitted audio signals and the decoded metadata to provide a set of audio signals; identifying a multi-channel index of the at least one audio signal and / or the set of audio signals; as well as At least the decoded at least one audio signal and the set of audio signals are combined using the index to provide a multi-channel audio signal.
19. The apparatus according to claim 18, comprising: a joint decoder for decoding the received encoded data to decode the at least one audio signal, the one or more transport audio signals and the metadata, or The invention also provides a method for decoding a first audio signal and a second audio signal processing unit (hereinafter referred to as "the method"). The method further comprises: a first decoder for decoding at least a first subset of the received encoded data to provide the at least one audio signal; and a second, different decoder for decoding at least a second subset of the received encoded data to provide the one or more transmitted audio signals and the metadata.
20. A method comprising: receiving encoded data for decoding, the encoded data comprising at least one audio signal, one or more transport audio signals, and metadata; decoding the received encoded data to decode the at least one audio signal, the one or more transport audio signals, and the metadata; synthesizing the decoded one or more transmitted audio signals and the decoded metadata to provide a set of audio signals; as well as At least the decoded at least one audio signal and the set of audio signals are combined to provide a multi-channel audio signal.
Citation Information
Patent Citations
Synchronization and switchover methods and systems for an adaptive audio system
CN103621101A
System aspects of an audio codec
CN105531928A