Encoding of a down-mixed multi-channel audio signal comprising a primary input channel and two or more scaled non-primary input channels
By processing multi-channel audio signals in the encoder premixer to form dominant channels and reduce channel correlation, the problem of low coding efficiency of multi-channel audio signals is solved, achieving efficient coding and preservation of audio quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2021-06-10
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to efficiently encode multi-channel audio signals, especially when there is high correlation between channels and no clear dominant channel, resulting in low coding efficiency.
The input multichannel audio signal is processed by an encoder premixer to form an output multichannel signal with dominant audio channels and less correlation, and then encoded with a simple encoder, discarding or reducing the bit allocation for non-dominant channels.
It achieves the maintenance of audio quality at lower bit rates by forming dominant channels and reducing channel correlation, thereby improving coding efficiency and audio signal reconstruction quality.
Smart Images

Figure CN116406471B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 037,635, filed June 11, 2020, and U.S. Provisional Patent Application No. 63 / 193,926, filed May 27, 2021, each of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates generally to audio coding, and particularly to the coding of multichannel audio signals. Background Technology
[0004] When storing or transmitting input audio signals for later use (e.g., playback to listeners), it is generally desirable to encode the audio signals with a reduced amount of data. The data reduction process applied to the input audio signal is commonly referred to as "audio encoding" (or "encoding"), and the device used for encoding is commonly referred to as an "audio encoder" (or "encoder"). The process of regenerating the output audio signal from the reduced data is commonly referred to as "audio decoding" (or "decoding"), and the device used for decoding is commonly referred to as an "audio decoder" (or "decoder"). Audio encoders and decoders can be adapted to operate on input signals consisting of a single audio channel or multiple audio channels. When the input signal consists of multiple audio channels, the audio encoder and audio decoder are respectively called a multichannel audio encoder and a multichannel audio decoder. Summary of the Invention
[0005] An implementation of adaptive downmixing for audio signals with improved continuity is disclosed.
[0006] In some embodiments, an audio encoding method includes receiving, with at least one processor, an input multi-channel audio signal including a primary input audio channel and L non-primary input audio channels, where L is a positive integer greater than 1; determining, with the at least one processor, a set of L input gains; for each of the L non-primary input audio channels and the L input gains, forming a respective scaled non-primary input audio channel from a respective non-primary input audio channel scaled according to the input gain; forming a primary output audio channel from a sum of the primary input audio channel and the scaled non-primary input audio channels; determining, with the at least one processor, a set of L prediction gains; for each of the L prediction gains, forming, with the at least one processor, a prediction channel from the primary output audio channel scaled according to the prediction gain; forming, with the at least one processor, L non-primary output audio channels from a difference between a respective non-primary input audio channel and a respective prediction signal; forming, with the at least one processor, an output multi-channel audio signal from the primary output audio channel and the L non-primary output audio channels; encoding the output multi-channel audio signal with an audio encoder; and transmitting or storing, with the at least one processor, the encoded output multi-channel audio signal.
[0007] In some embodiments, wherein determining the set of L input gains includes determining a set of L mixing coefficients; determining an input mixing intensity coefficient; and determining the L input gains by scaling the L mixing coefficients by the input mixing intensity coefficient.
[0008] In some embodiments, determining the set of L prediction gains includes determining a set of L mixing coefficients; determining a prediction mixing intensity coefficient; and determining the L prediction gains by scaling the L mixing coefficients by the prediction mixing intensity coefficient.
[0009] In some embodiments, the input mixing intensity coefficient h is determined by a pre-prediction constraint equation h = fg, where f is a predetermined constant value greater than zero and less than or equal to 1, and g is the prediction mixing intensity coefficient.
[0010] In some embodiments, the prediction mixing intensity coefficient g is a maximum real value solution of the equation βf 2 g 3 + 2αfg 2 - βfg - α + gw = 0, where β = u H x E x u, and the quantities w, column vector v, and matrix E are components of a covariance matrix of mid-signals of a dominant channel.
[0011] In some embodiments, the covariance matrix of the mid signal is computed from a covariance matrix of the multi-channel input audio signal.
[0012] In some embodiments, two or more input multi-channel audio channels are processed in accordance with a mixing matrix to produce the primary input audio channel and the L non-primary input audio channels.
[0013] In some embodiments, the primary input audio channel is determined by a dominant eigenvector of an expected covariance of a typical input multi-channel audio signal.
[0014] In some embodiments, each of the L mixing coefficients is determined based on a correlation of a respective one of the non-primary input audio channels and the primary input audio channel.
[0015] In some embodiments, the encoding includes allocating more bits to the primary output audio channel than to the L non-primary output audio channels, or discarding one or more of the L non-primary output audio channels.
[0016] Other implementations disclosed herein relate to a system, apparatus, and computer readable medium. The details of the disclosed implementations are set forth in the accompanying drawings and the description below. Other features, objects, and advantages are apparent from the description, drawings, and claims.
[0017] Particular implementations disclosed herein provide one or more of the following advantages. An input multi-channel audio signal is processed by an audio encoder premixer to form an output multi-channel audio signal having two desirable properties for efficient encoding. The first property is that at least one dominant audio channel of the output multi-channel audio signal contains most or all of the sound elements of the input multi-channel audio signal. The second property is that each audio channel of the output multi-channel audio signal is largely uncorrelated with each other audio channel. A simple encoder can provide data to a simple decoder to help regenerate audio channels that are discarded by the simple encoder.
[0018] The two properties described above allow the simple encoder to efficiently encode the output multi-channel audio signal by allocating fewer bits to the encoding of less dominant channels or by choosing to discard less dominant audio channels entirely. BRIEF DESCRIPTION OF DRAWINGS
[0019] In the drawings, specific arrangements or orderings of illustrative elements are shown as specific examples. However, it should be understood that such elements can be arranged or ordered differently, and / or that specific arrangements or orderings can not be required. Additionally, some elements can be included or omitted from the figures, and / or some elements can be included or omitted from the description of the figures, as needed for clarity and conciseness. Further, it should be understood that reference numerals can be repeated among the figures to refer to elements with similar or related functions.
[0020] Further, in the drawings, connections between two or more other illustrative elements can be represented by a connection element, such as a wire, a bus, a pipe, a channel, and the like, which can be understood as a person skilled in the art would understand it. However, it should be understood that no connection is necessarily implied, unless explicitly described as such. Further, connection elements can represent a single connection or multiple connections between two or more elements, unless explicitly described as a single connection. Further, connections between elements in the drawings can be used to represent one or more connections between the elements, unless explicitly described as a single connection or multiple connections. Further, connections drawn between elements can also represent element ordering or a sequence and not a necessary connection or association between the elements. Additionally, it should be understood that statements that one element is "connected" or "associated" with another element can represent that the elements are directly connected or associated with each other, or that additional one or more elements can be found between the two elements.
[0021] Figure 1 is a block diagram of an arrangement of a simple audio encoder and a simple audio decoder aimed at forming a facsimile of an input multi-channel audio signal according to some embodiments.
[0022] Figure 2 is a block diagram of an audio codec system comprising an audio encoder, an audio decoder 106, an encoder pre-mixer, and a decoder post-mixer according to some embodiments.
[0023] Figure 3 is shown an arrangement of processing elements according to some embodiments, where an input multi-channel audio signal is split into sub-band signals by a filter bank, where each sub-band is processed by a mixing matrix to produce a remixed sub-band signal.
[0024] Figure 4 is a block diagram of an arrangement of two mixing operations aimed at implementing the functionality of an encoder pre-mixer of Figure 2 or Figure 3 an encoder pre-mixer according to some embodiments.
[0025] Figure 5 is a block diagram of a prediction mixer according to some embodiments.
[0026] Figure 6 is shown an arrangement of processing elements according to some embodiments, where an input multi-channel audio signal is split into sub-band signals by a filter bank, where each sub-band is processed by a mixing matrix to produce a remixed sub-band signal. Figure 2arrangement of processing elements of a decoder post-mixer.
[0027] Figure 7 is a flowchart of an adaptive downmix process of an audio signal with improved continuity according to some embodiments.
[0028] Figure 8 is a flowchart of an adaptive downmix process of an audio signal with improved continuity according to some embodiments. Figures 1-7 Block diagram of a system describing features and processes.
[0029] The same reference numbers in the various drawings represent the same elements. DETAILED DESCRIPTION
[0030] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described implementations. Those skilled in the art will understand that the various described implementations can be practiced without these specific details. In other instances, well-known methods, procedures, components and circuits have not been described in detail so as not to unnecessarily obscure aspects of the implementations. A number of features are described below, each of which can be used independently of the others and in any combination with other features.
[0031] Nomenclature
[0032] As used herein, the term “includes” and its variants are to be read as open terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless expressly indicated otherwise. The term “based on” is to be understood as “based, at least in part, on.” The terms “one example implementation” and “an example implementation” are to be understood as “at least one example implementation.” The term “another implementation” is to be understood as “at least one other implementation.” The term “determining” is to be understood as obtaining, receiving, calculating, computing, estimating, predicting or deriving. Also, in the description and the claims, unless otherwise specified, all technical and scientific terms used herein are to be interpreted as is customary in the art to which this disclosure pertains.
[0033] Figure 1is a block diagram of an arrangement 10 of a simple audio encoder and a simple audio decoder, aimed at forming a multi-channel audio signal 17 (Z') which is a copy of a multi-channel audio signal 13 (Z). The multi-channel audio signal 13 is processed by a simple audio encoder 14 to produce an encoded representation 15, which can be stored and / or transmitted 20 to a simple audio decoder 16 which produces the multi-channel audio signal 17. Preferably, the data size of the encoded representation 15 is minimized while minimizing the difference between the multi-channel audio signal 13 and the multi-channel audio signal 17. Furthermore, the difference between the multi-channel audio signal 13 and the multi-channel audio signal 17 can be measured in terms of similarity as perceived by a human listener. The measure of human perceived similarity between the audio signal 13 and the audio signal 17 is based on a reference playback method (i.e. the assumption of a hypothetical default means by which the audio channels of the multi-channel audio signals 13, 17 are presented as an auditory experience to a listener).
[0034] The efficiency of the simple audio encoder 14 and decoder 16 can be defined in terms of the data rate (measured in bits per second) of the encoded representation 15 required to provide the multi-channel audio signal 17 which a listener will judge to match the multi-channel audio signal 13 with a certain level of perceived quality. The simple audio encoder 14 and decoder 16 can achieve higher efficiency (i.e. lower data rate) when the multi-channel audio signal 13 is known to have certain properties. In particular, higher efficiency can be achieved when the multi-channel audio signal 13 is known to have the following properties (DD1 and DD2):
[0035] DD1 : One or more channels of the multi-channel audio signal are generally more dominant than the other channels, where a more dominant audio channel is an audio channel which will contain the substantial elements of most (or all) of the sound elements in the scene. That is, when the multi-channel audio signal is presented to a listener by the reference playback method, the dominant audio signal, when presented to the listener as a single audio channel, will contain most (or all) of the sound elements of the multi-channel signal.
[0036] DD2: Each audio channel of the multi-channel audio signal is largely uncorrelated with each other audio channel.
[0037] In the case that the multi-channel audio signal 13 has the attributes DD1 and DD2, the simple audio encoder 14 can achieve improved efficiency using several techniques, including but not limited to: allocating fewer bits to the encoding of less dominant channels or selecting to discard less dominant channels altogether. The simple audio encoder 14 can provide data to the simple audio decoder 16 to assist in regenerating channels discarded by the simple encoder audio encoder 14. Preferably, the multi-channel audio signal without the attributes DD1 and DD2 can be processed by an encoder pre-mixer to form (e.g., compute, determine, construct, or generate) a multi-channel audio signal with the attributes DD1 and DD2, as described further with reference to Figure 2 A corresponding decoder post-mixer can be applied to the simple decoder output to form an output multi-channel audio signal, such that the decoder post-mixer performs an approximate inverse operation to that of the encoder pre-mixer.
[0038] Figure 2 is a block diagram of an audio codec system 100, which includes an audio encoder 104 and an audio decoder 106, an encoder pre-mixer 102, and a decoder post-mixer 108. The audio encoder 104 and the audio decoder 106 form a multi-channel audio signal 109 (X') that is a copy of the multi-channel audio signal 101 (X). Preferably, the data size of the encoded representation 105 is minimized while the difference between the multi-channel audio signal 101 and the multi-channel audio signal 109 is minimized. Furthermore, the difference between the multi-channel audio signal 101 and the multi-channel audio signal 109 can be measured in terms of similarity as perceived by a human listener.
[0039] The measurement of human-perceived similarity between the multi-channel audio signal 101 and the multi-channel audio signal 109 is based on a reference playback method (i.e., the assumed default means by which the audio channels of the audio signal 101, 109 are presented to a listener as an auditory experience). The efficiency of the multi-channel audio encoder 104 and the multi-channel audio decoder 106 can be defined in terms of the data rate (measured in bits per second) of the encoded representation 105 of the multi-channel audio signal 109 that a listener would judge to match the multi-channel audio signal 101 with a particular level of perceived quality.
[0040] Reference Figure 2, the input multi-channel audio signal 101 is mixed according to an encoder pre-mixer 102 (R) to produce an output multi-channel audio signal 103 (Z), which is processed by a simple audio encoder 104 to produce an encoded representation 105, which can be stored and / or transmitted 110 to a simple audio decoder 106, which produces a multi-channel audio signal 107 (Z'). The multi-channel audio signal 107 is processed by a decoder post-mixer 108 (R') to produce a decoded multi-channel audio signal 109. The encoder pre-mixer 102 provides metadata 112 (Q) comprising information needed to determine the behavior of the decoder post-mixer 108. The metadata 112 can be stored and / or transmitted 110 together with the encoded representation 105. As will be appreciated by the skilled person, a measure of the efficiency of the multi-channel audio encoder 104 and the multi-channel audio decoder 106 can comprise the size of the metadata 112, typically measured in bits per second.
[0041] The multi-channel audio signal 101 can consist of N audio channels, wherein there can be a significant correlation between some pairs of channels, and wherein no single channel can be considered as a dominant channel. That is, the multi-channel audio signal 101 can not have the properties DD1 and DD2, and thus the multi-channel audio signal 101 can not be a suitable signal for encoding and decoding using the simple audio encoder 104 and the decoder 106, respectively.
[0042] Preferably, the encoder pre-mixer 102 is adapted to process the input multi-channel audio signal 101 to produce an output multi-channel audio signal 103, wherein the output multi-channel audio signal 103 has the properties DD1 and DD2. Given that the input multi-channel audio signal X consists of N channels:
[0043]
[0044] The output multi-channel audio signal Z is computed as follows:
[0045]
[0046] = R(t) x X(t). [3]
[0047] The coefficients of the encoder premixer matrix R can vary with time, so R can be considered as a function of time. The values of the elements of R can be calculated at regular intervals (for example, where the intervals can be 20ms, or a value between 1ms and 100ms) or at irregular intervals. When the values of the elements of R change, the changes can be smoothly interpolated. In the following discussion, references to R should be taken as references to the time-varying encoder premixer R(t), and references to R' should be taken as references to the time-varying decoder premixer R'(t).
[0048] In embodiments, the encoder premixer 102 can process components of the audio signal in a frequency band b, where 1 < b < B, using mixing coefficients R b (t). Figure 4 An arrangement of processing elements 150 is shown, whereby a multi-channel audio signal 151 (X) is split into B sub-band signals, X [1] (t), X [2] (t),... X [,] (t), each sub-band signal (e.g. 153 (X [1] (t))) is processed by a mixing matrix (e.g. 154 (R1)) to produce a re-mixed sub-band signal (e.g. 155 (Z [1] (t)). [1] (t), Z [2] (t),... Z [,] (t), which are recombined by a combiner 156 to form a multi-channel audio signal 157 (Z).
[0049] For the purposes of the following discussion, references to the matrix R(t) can be interpreted as references to R b (t), where b refers to a sub-band. It can be appreciated that the following discussion can apply to signals processed in sub-bands, or can apply to signals processed without sub-band processing. Those skilled in the art will appreciate that many methods can be used to process audio signals according to sub-bands, and the discussion of the matrix R will apply to these methods.
[0050] With reference to Figure 2 , R mixes the channels of the multi-channel audio signal 101 to produce a multi-channel audio signal 103 having properties DD1 and DD2, as described above, thereby enabling the encoder 104 to achieve improved data efficiency. The decoder post-mixer 108 (R') provides a mixing operation that is the inverse of the mixer R, such that:
[0051] X'(t) = R'(t) x Z'(t) [4]
[0052] Figure 3 is intended to achieve Figure 2encoder pre-mixer 102(R) of Fig. 1 or Figure 4 encoder pre-mixer R of Fig. 1 b Fig. 2 is a block diagram of an arrangement 200 of two mixing operations of the functionality of Fig. 1. An N-channel multi-channel input signal 201 (X) is mixed by a mixing matrix 202 (M) to produce an N-channel intermediate signal 203 (Y), which is then processed by a mixer 204 (P) to produce an N-channel signal 205 (Z). Figure 3 The signals 201 (X) and 205 (Z) in Fig. 2 are intended to correspond to Figure 2 the input signals 101 (X) and 103 (Z) in Fig. 1, respectively, or to Figure 4 the sub-band signals 153 (X / (t)) and 155 (Z / (t)) in Fig. 1.
[0053] An analysis block 210 (A) obtains input from the signal 201 and computes coefficients 212 which will be used to adjust the operation of the mixer 204. The analysis block 210 also produces metadata 211 (Q) corresponding to the metadata 112 of Fig. 1, which will be provided to the decoder 113 (Q) for use by the decoder post-mixer 108. Figure 2
[0054] From the arrangement of mixers 202 and 204 in Fig. 2, it will be understood that the matrix R will be: Figure 4
[0055] R(t) = P(t) x M [5]
[0056] where the matrix P(t) can vary over time.
[0057] Thus:
[0058]
[0059] The matrix M is adapted to ensure that the intermediate signal 203 (Y) has the property DD1. That is, the N-channel signal 203 (Y) contains one channel which can be considered a dominant channel. Without loss of generality, the matrix M is adapted to ensure that the first channel Y1(t) is the dominant channel. In the following, when the first channel of a multi-channel signal is the dominant channel, this first channel will be referred to as the primary channel. In some contexts, the primary channel can also be referred to as the "eigencannel".
[0060] The [N x N] matrix M can be determined from the [N x N] desired covariance matrix Cov of the N-channel input signal X(t):
[0061]
[0062] where X(t) H The operation represents the Hermitian transpose of a column vector X(t) of length N, and the operation E() represents the expected value of a variable.
[0063] The expected value used in Equation
[10] can be estimated based on the assumed characteristics of a typical input multichannel audio signal, or it can be estimated by statistical analysis of a set of typical input multichannel audio signals.
[0064] The covariance matrix Cov can be factored using eigenvalue analysis, as is familiar to those skilled in the art:
[0065] Cov=V×D×V H
[12]
[0066] Here, matrix V is a unitary matrix, and matrix D is a diagonal matrix, where the diagonal elements are non-negative real values sorted in descending order.
[0067] Matrix M can be chosen as:
[0068] M = V H
[13]
[0069] Those skilled in the art will understand that the covariance matrix Cov will depend on the panning method used to form the original input signal X(t), and the typical use of the panning method by the creator of the typical signal.
[0070] For example, when the original input signal is a 2-channel stereo signal intended for playback on stereo speakers, the typical translation rules used by content creators will result in some audio objects being translated to the first channel (often referred to as the left channel in this context), some audio objects being translated to the second channel (often referred to as the right channel in this context), and some objects being translated to both channels simultaneously. In this case, the covariance matrix can be similar to:
[0071] For L / R stereo: And according to equations
[12] and
[13] :
[0072] For L / R stereo:
[0073] The matrix M in Equation
[15] will be familiar to those skilled in the art as a mixing matrix suitable for converting the original input audio signal X in L / R stereo format into an intermediate signal Z in Mid / Side format. Those skilled in the art will also understand that the first channel of Z (often referred to as the Mid signal in this case) is the dominant audio signal (main channel) and has the property that most of the audio elements in the stereo mix will appear in the Mid signal.
[0074] As an alternative example, when the original input signal is a 5-channel surround signal intended for playback on a common five-speaker setup, the typical translation rules used by content creators will result in some audio objects being translated to one of the five channels, while others are translated to two or more channels simultaneously. In this case, the covariance matrix can resemble:
[0075] For 5 channels: And according to equations
[12] and
[13] :
[0076] For 5 channels:
[0077] It will be understood that the top row of matrix M in equation
[17] consists of similar (or identical) positive values. This means that, according to equation [6], the first channel of the intermediate signal Y(t) will be formed by the sum of the five channels of the original input audio signal X(t), and this ensures that all sound elements translated in the original input audio signal will appear in Y1(t) (the first channel of the N-channel signal Y(t)). Therefore, this choice of matrix M ensures that the intermediate signal Y has the property DD1 (Y1(t) is the primary channel).
[0078] In another alternative example, when the input multichannel audio signal X(t) already contains the dominant channel (and without loss of generality, it is assumed that the first channel X1(t) is the dominant channel), the matrix M can be an [N×N] identity matrix. In a more specific example of an input multichannel audio signal with a dominant / primary first channel, the input multichannel audio signal can represent an acoustic scene encoded in Ambisonic format (means for encoding acoustic scenes that will be familiar to those skilled in the art).
[0079] At time t, by Figure 4 Analysis block 210(A) in the middle calculates matrix 212(P(t)) according to the following procedure:
[0080] 1. Determine the covariance of the intermediate signal Y(t) at time t. An example of a method for calculating the covariance is:
[0081]
[0082] Alternatively, the covariance of the intermediate signal Y(t) can be calculated from the covariance of the input multi-channel audio signal X(t), as follows:
[0083] Cov Y (t)=M×Cov X (t)×M H
[19]
[0084] in
[0085]
[0086] 2. Extract the scalar \(w = [Cov Y (t)]\) from the \([L\times L]\) covariance matrix \(Cov Y (t)\), 1,1 ,
[0087] the \([N\times1]\) column vector \(v = [Cov Y (t)]\) 2..L,1 and the \([N\times N]\) matrix \(E == [Cov Y (t)]\) 2..L,2..L , where \(N = L - 1\), and:
[0088]
[0089] 3. Determine the quantities \(\alpha\), \(\beta\) of the mixing coefficient \(u\) and the \([N\times1]\) vector:
[0090]
[0091] \(\beta=u\) H \times E\times u
[24]
[0092] 4. Given the quantities \(w\), \(\alpha\) and \(\beta\), solve the equation
[25] to determine the input mixing intensity coefficient \(h\) and the predicted mixing intensity coefficient \(g\):
[0093] \(\beta h\) 2 \(g + 2\alpha hg-\beta h-\alpha+gw = 0
[25]
[0094] where the solution of this equation will also satisfy the pre - prediction constraint equation. An example of the pre - prediction constraint equation is:
[0095] PPC1: \(h = fg
[26]
[0096] where \(f\) is a predetermined constant value satisfying \(0\lt f\leq1\).
[0097] When using the pre - prediction constraint PPC1, the equation
[25] can be modified to:
[0098] \(\beta f\) 2 \(g\) 3 \(+ 2\alpha fg\) 2 \(-\beta fg-\alpha+gw = 0
[27]
[0099] And for the maximum real value of \(g\), the equation
[27] can be solved, and thus the value of \(h\) can be determined using the equation
[26] .
[0100] 5. Form the \([L\times L]\) matrix \(Q\) as:
[0101]
[0102] 6. The [L×L] matrix P(t) is formed as follows:
[0103] P(t=(I L -gQ)×(I L +hQ H
[29]
[0104] Where I L It is an [L×L] identity matrix.
[0105] Figure 4 Metadata 211(Q) in the document can convey information that allows [the following to be communicated]: Figure 2 The decoder post-mixer 113 determines the unit vector u and the information of the coefficients g and h.
[0106] The solution to g in equation
[27] can be approximated by choosing an initial estimate g1 = 1 and iterating (according to Newton's method, as is known in the art) multiple times:
[0107]
[0108] This allows a reasonable approximation of the solution to be found from g = g5. It will be understood that other methods are known in the art for finding approximate solutions to cubic equations
[27] .
[0109] According to an alternative embodiment, at time t, a [N×1] vector u representing the correlation between the primary channel and the remaining N non-primary channels of the intermediate signal Y(t) can be determined, and the input mixing intensity coefficient h and the predicted mixing intensity coefficient g can be determined according to equation
[28] to form P(t), thereby determining the [L×L] matrix P(t) such that the signal Z(t) = P(t) × Y(t) will have attributes DD1 and DD2.
[0110] The determination of the coefficients g and h can be governed by the predictive constraint equations. An example of the predictive constraint equations (PPC1) is given in equation
[26] . The preferred choice for the coefficient f is f = 0.5, but a range of f values of 0.2 ≤ f ≤ 1 may be suitable.
[0111] In an alternative embodiment, the following pre-prediction constraints may be used:
[0112]
[0113] Where c is a predetermined constant. A typical value can be c = 1, but the value of c can be chosen within the range of 0.25 ≤ c ≤ 4.
[0114] According to the constraint PPC2 in equation
[31] , the solution to equation
[25] is:
[0115] when:
[0116]
[0117] h = 0
[0118] otherwise:
[0119]
[0120] Figure 5 This is a block diagram of a predictive mixer 300 according to some embodiments. The matrix term (I) of equation
[29] L -gQ) and (I L +hQ H This can be implemented by a predictive mixer 300, where, in this example, the signal Y(t) consists of four channels (L=4), the first channel 301 (Y1) is the primary channel, and the remaining three non-primary channels 302 (e.g., Y2, Y3, Y4) are scaled according to three input gains 312 (H2, H3, and H4) to form scaled input signal components (e.g., 304). The scaled input signal components are added 305 to the primary input channel 301 (Y1) to form the primary output 306 (Z1). The primary output 306 (Z1) is scaled by three prediction gains 313 (G2, G3, and G4) to form three prediction signals (e.g., 311). Each prediction signal (e.g., 308 and 309) is subtracted from the corresponding input (e.g., Y2 302) to form the corresponding non-dominant output 310 (Z2).
[0121] The three input gains 312 (H2, H3, and H4) can be determined from the mixing coefficient u (determined according to equation
[23] ) and the input mixing intensity coefficient h (determined according to the solution of equation
[25] ), where:
[0122]
[0123] The three prediction gains 313 (G2, G3, and G4) can be determined from the mixing coefficient u (determined according to equation
[23] ) and the prediction mixing intensity coefficient g (determined according to the solution of equation
[25] ), where:
[0124]
[0125] Those skilled in the art will understand that Figure 4 The arrangement of linear matrix operations M 202 and P 204 can be achieved using a single matrix R = P × M.
[0126] Those skilled in the art will understand that Figure 2 The decoder matrix R′ can be formed from matrices M′ (the inverse of M) and P′ (the inverse of P):
[0127] R′(t)=M′×P′(t)
[38]
[0128] Furthermore, M′ can be pre-calculated (not changing over time), and P′ can be formed using this method:
[0129] P′=(I L -hQ H )×(I L +gQ)
[39]
[0130] Figure 6 Implementation shown Figure 2 The processing elements of the decoder post-mixer 108 are arranged in an arrangement 400. Metadata 402(Q) provides information to the inverse prediction determination block 403(B), which calculates the coefficients 404 required to determine the operation of the inverse predictor 405(P′). The signal 401(Z′) is processed by the inverse predictor 405(P′) to produce an intermediate signal 406(Y′), which is then processed by the matrix 407(M′) to produce an output signal 408X′.
[0131] Example process
[0132] Figure 7 This is a flowchart of an adaptive downmixing process 700 for audio signals, which improves continuity according to some embodiments. Process 700 can be, for example... Figure 8 The system shown is implemented using system 800.
[0133] Process 700 includes the following steps: receiving an input multichannel audio signal comprising a main input audio channel and L non-main input audio channels (701); determining a set of L input gains, where L is a positive integer greater than 1 (702); for each of the L non-main input audio channels and L input gains, forming a corresponding scaled non-main input audio channel from the corresponding non-main input audio channel scaled according to the input gain (703); forming a main output audio channel from the sum of the main input audio channel and the scaled non-main input audio channels (704). 04); For each of the L prediction gains, determine a set of L prediction gains (705); form prediction channels from the main output audio channels scaled according to the prediction gains (706); form L non-main output audio channels from the difference between the corresponding non-main input audio channels and the corresponding prediction signals (706); form an output multi-channel audio signal from the main output audio channels and the L non-main output audio channels (707); encode the output multi-channel audio signal (708); and transmit or store the encoded output multi-channel audio signal (709). Reference Figures 1-6 To describe each of these steps more comprehensively.
[0134] Example System Architecture
[0135] Figure 8 A reference for implementation according to an embodiment is shown. Figures 1-7 A block diagram of an example system 800 describing the features and processes. System 800 includes any device capable of playing audio, including but not limited to: smartphones, tablet computers, wearable computers, in-vehicle computers, game consoles, surround sound systems, and self-service terminals.
[0136] As shown in the figure, system 800 includes a central processing unit (CPU) 801, which is capable of executing various processes based on a program stored, for example, in a read-only memory (ROM) 802 or a program loaded from, for example, a storage unit 808 into a random access memory (RAM) 803. The RAM 803 also stores data required as needed when the CPU 801 executes various processes. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0137] The following components are connected to I / O interface 805: input unit 806, which may include a keyboard, mouse, etc.; output unit 807, which may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 808, including a hard disk or other suitable storage device; and communication unit 809, including a network interface card such as a network card (e.g., wired or wireless).
[0138] In some implementations, the input unit 806 includes one or more microphones at different locations (depending on the host device), enabling the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0139] In some implementations, the output unit 807 includes a system with a variety of numbers of speakers. For example... Figure 8 As shown, the output unit 807 (depending on the capabilities of the host device) can present audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0140] Communication unit 809 is configured to communicate with other devices (e.g., via a network). Drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media, is mounted on drive 810 such that computer programs read from it can be installed into storage unit 808 as needed. Those skilled in the art will understand that although system 800 is described as including the components described above, in practice, it is possible to add, remove, and / or replace some of these components, and all such modifications or alterations fall within the scope of this disclosure.
[0141] The aspects of the system described herein can be implemented in a suitable computer-based sound processing network environment for processing digital or digitized audio files. Parts of the adaptive audio system may include one or more networks comprising any desired number of individual machines, including one or more routers (not shown) for buffering and routing data transmitted between computers. Such networks can be built on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0142] According to exemplary embodiments of this disclosure, the above processes can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include computer program products comprising computer programs tangibly implemented on machine-readable media, the computer programs including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 809, and / or installed from removable media 811, such as... Figure 8 As shown.
[0143] Generally, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be implemented by control circuitry (e.g., with...) Figure 8 The control circuitry, in conjunction with other components of the CPU, executes the actions described in this disclosure. Some aspects may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). Although various aspects of exemplary embodiments of this disclosure are shown and described as block diagrams, flowcharts, or other illustrated representations, it is understood that, as non-limiting examples, the blocks, apparatuses, systems, techniques, or methods described herein may be implemented as hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0144] Furthermore, the various blocks shown in the flowchart can be considered as method steps, and / or operations generated by the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform (multiple) related functions. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly implemented on a machine-readable medium, the computer program containing program code configured to perform the methods described above.
[0145] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program used by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transient and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0146] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device with control circuitry, such that when executed by the processor of a computer or other programmable data processing device, the program code enables the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0147] While this document contains numerous specific implementation details, these details should not be construed as limiting the scope of possible claims, but rather as descriptions of features that may be specific to particular embodiments. Certain features described in this specification within the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed to be so, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof. The logical flow described in the accompanying drawings does not require the specific order or sequential order shown to achieve the desired result. Furthermore, additional steps or steps may be provided or removed from the described flow, and additional components may be added to or removed from the described system. Therefore, other implementations are within the scope of the appended claims.
Claims
1. A method of audio encoding, comprising: receiving, with at least one processor, an input multi-channel audio signal comprising a primary input audio channel and L non-primary input audio channels; determining, with the at least one processor, a set of L input gains, wherein L is a positive integer greater than one, and wherein the set of L input gains is determined by scaling a set of L mixing coefficients by input mixing intensity coefficients; for each of the L non-primary input audio channels and the L input gains, forming a respective scaled non-primary input audio channel from the respective non-primary input audio channel scaled by the input gain; forming a primary output audio channel from a sum of the primary input audio channel and the scaled non-primary input audio channels; determining, with the at least one processor, a set of L prediction gains, wherein the set of L prediction gains is determined by scaling the set of L mixing coefficients by prediction mixing intensity coefficients; for each of the L prediction gains, forming, with the at least one processor, a prediction channel from the primary output audio channel scaled by the prediction gain; forming, with the at least one processor, L non-primary output audio channels from a difference between the respective non-primary input audio channels and the respective prediction channels; forming, with the at least one processor, an output multi-channel audio signal from the primary output audio channel and the L non-primary output audio channels; encoding, with an audio encoder, the output multi-channel audio signal; and transmitting or storing, with the at least one processor, the encoded output multi-channel audio signal.
2. A system, comprising: one or more computer processors; and a non-transitory computer-readable medium storing instructions that, when executed by the one or more computer processors, cause the one or more computer processors to perform the following operations: receiving, with at least one processor, an input multi-channel audio signal comprising a primary input audio channel and L non-primary input audio channels; determining, with the at least one processor, a set of L input gains, wherein L is a positive integer greater than one, and wherein the set of L input gains is determined by scaling a set of L mixing coefficients by input mixing intensity coefficients; for each of the L non-primary input audio channels and the L input gains, forming a respective scaled non-primary input audio channel from the respective non-primary input audio channel scaled by the input gain; forming a primary output audio channel from a sum of the primary input audio channel and the scaled non-primary input audio channels; determining, with the at least one processor, a set of L prediction gains, wherein the set of L prediction gains is determined by scaling the set of L mixing coefficients by prediction mixing intensity coefficients; for each of the L prediction gains, forming, with the at least one processor, a prediction channel from the primary output audio channel scaled by the prediction gain; forming, with the at least one processor, L non-primary output audio channels from a difference between the respective non-primary input audio channels and the respective prediction channels; forming, with the at least one processor, an output multi-channel audio signal from the primary output audio channel and the L non-primary output audio channels; encoding, with an audio encoder, the output multi-channel audio signal; and transmitting or storing, with the at least one processor, the encoded output multi-channel audio signal.
3. A non-transitory computer-readable medium storing instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the following operations: receiving, with at least one processor, an input multi-channel audio signal comprising a primary input audio channel and L non-primary input audio channels; determining, with the at least one processor, a set of L input gains, where L is a positive integer greater than one, and where the set of L input gains is determined by scaling a set of L mixing coefficients by input mixing intensity coefficients; for each of the L non-primary input audio channels and the L input gains, forming, with the at least one processor, a respective scaled non-primary input audio channel from the respective non-primary input audio channel scaled by the input gain; forming a primary output audio channel from a sum of the primary input audio channel and the scaled non-primary input audio channels; determining, with the at least one processor, a set of L prediction gains, where the set of L prediction gains is determined by scaling the set of L mixing coefficients by prediction mixing intensity coefficients; for each of the L prediction gains, forming, with the at least one processor, a prediction channel from the primary output audio channel scaled by the prediction gain; forming, with the at least one processor, L non-primary output audio channels from a difference between the respective non-primary input audio channels and the respective prediction channels; forming, with the at least one processor, an output multi-channel audio signal from the primary output audio channel and the L non-primary output audio channels; encoding, with an audio encoder, the output multi-channel audio signal; and transmitting or storing, with the at least one processor, the encoded output multi-channel audio signal.
Citation Information
Patent Citations
Method and an Apparatus for Decoding an Audio Signal
US20080192941A1