ADAPTIVE DOWNWARD MIXING OF AUDIO SIGNALS WITH ENHANCED CONTINUITY
Patent Information
- Application Number
- MX2022015325
- Authority / Receiving Office
- MX · MX
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-27
- Filing Date
- 2022-12-02
- Publication Date
- 2026-06-12
- Estimated Expiration
- 2041-06-10
AI Technical Summary
Existing audio coding technologies struggle to efficiently encode multi-channel audio signals while maintaining audio continuity and reducing data requirements, particularly when dominant channels are not clearly defined and channels are highly correlated.
An adaptive downmixing process that involves forming a primary output channel from a sum of scaled non-primary channels and forming non-primary output channels from differences, using mixing and prediction gains to create an output multichannel audio signal with a dominant primary channel and uncorrelated channels, allowing for efficient encoding by allocating fewer bits or discarding less dominant channels.
The solution enables efficient encoding by minimizing data usage while preserving audio quality, as the dominant channel contains most sonic elements and channels are largely uncorrelated, facilitating effective regeneration by a simple decoder.
Smart Images

Figure MX435487B0
Abstract
Description
ADAPTIVE DOWNWARD MIXING OF AUDIO SIGNALS WITH ENHANCED CONTINUITY Cross-reference to Related Applications
[001] This application claims priority to U.S. Provisional Patent Application No. 63 / 037,635, filed June 11, 2020, and U.S. Provisional Patent Application No. 63 / 193,926, filed May 27, 2021, each of which is hereby incorporated by reference in its entirety. Field of Invention
[002] This disclosure relates generally to audio coding and, in particular, to multi-channel audio signal coding. Background of the Invention
[003] When an input audio signal is to be stored or transmitted for later use (for example, for playback to a listener), it is often desirable to encode the audio signal with a reduced amount of data. The process of data reduction, as applied to an input audio signal, is commonly called audio coding (or encoding), and the apparatus used for coding is commonly called an audio encoder (or encoder). The process of regenerating an output audio signal from the reduced data is commonly called audio decoding (or decoding), and the apparatus used for decoding is commonly called an audio decoder (or decoder). Audio encoders and decoders can be adapted to operate on input signals that are composed of a single audio channel or multiple audio channels.When an input signal is composed of multiple audio channels, the audio encoder and audio decoder are called a multi-channel audio encoder and a multi-channel audio decoder, respectively. Brief Description of the Invention
[004] Implementations for adaptive downmixing of audio signals with improved continuity are disclosed.
[005] In some embodiments, an audio coding method comprises: receiving, with at least one processor, a multichannel input audio signal comprising one primary input audio channel and L non-primary input audio channels; determining, with the at least one processor, a set of L input gains, where L is a positive integer greater than one; for each of the L non-primary input audio channels and L input gains, forming a respective scaled non-primary input audio channel from the respective scaled non-primary input audio channel according to the input gain; forming an output primary audio channel from the sum of the input primary audio channel and the scaled non-primary audio channels;Determine, using at least one processor, a set of L prediction gains; for each of the L prediction gains, form, using at least one processor, a prediction channel of the primary output audio channel scaled according to the prediction gain; form, using at least one processor, L non-primary output audio channels from the difference between the respective non-primary input audio channel and the respective prediction signal; form, using at least one processor, a multichannel output audio signal from the primary output audio channel and the L non-primary output audio channels; encode, using an audio encoder, the multichannel output audio signal; and transmit or store, using at least one processor, the encoded multichannel output audio signal.
[006] In some modalities, where the determination of the set of L input gains comprises: determining a set of L mixing coefficients; determining an input mixing intensity coefficient; and determining the L input gains by modifying the scale of the L mixing coefficients by the input mixing intensity coefficient.
[007] In some modalities, determining the set of L prediction gains comprises: determining a set of L mixing coefficients; determining a prediction mixing intensity coefficient; and determining the L prediction gains by scaling the L mixing coefficients by the prediction mixing intensity coefficient.
[008] In some modes, the input mixing intensity coefficient, h, is determined by a pre-prediction constraint equation, h=fg, where f is a predetermined constant value greater than zero and less than or equal to one, and g is the prediction mixing intensity coefficient.
[009] In some modalities, the prediction mixing intensity coefficient, g, is a larger real-valued solution for: Pf2g3+ 2afg2- Pfg - a + gw = O, where β = uHx E xu, u =^v,a = |v|2=vn' Ylacidad w, the column vector vy, and the matrix E are components of a covariance matrix for an intermediate signal having a dominant channel.
[0010] In some modes, the intermediate signal covariance matrix is calculated from a multichannel input audio signal covariance matrix.
[0011] In some modes, two or more multichannel input audio channels are processed according to a mixing matrix to produce the primary input audio channel and the L non-primary input audio channels.
[0012] In some modalities, the primary input audio channel is determined by a dominant eigenvector of an expected covariance of a typical input multichannel audio signal.
[0013] In some modes, each of the L mixing coefficients is determined based on a correlation of a respective channel of the non-primary input audio channels and the primary input audio channel.
[0014] In some modes, encoding includes assigning more bits to the primary output audio channel than to the L non-primary output audio channels, or discarding one or more of the L non-primary output audio channels.
[0015] Other implementations disclosed herein relate to a computer-readable system, apparatus, and medium. Details of the disclosed implementations are set forth in the accompanying figures and the description below. Other features, objects, and advantages are evident from the description, drawings, and claims.
[0016] The particular implementations disclosed herein provide one or more of the following advantages. zQQRnn / Qznz / q / υιλι An input multichannel audio signal is processed by an audio encoder premixer to form an output multichannel audio signal that has two desirable attributes for efficient encoding. The first attribute is that at least one dominant audio channel in the output multichannel audio signal contains most or all of the sonic elements of the input multichannel audio signal. The second attribute is that each of the audio channels in the output multichannel audio signal is not highly correlated with each of the other audio channels. The simple encoder can provide data to a simple decoder to assist in the regeneration of audio channels that were discarded by the simple encoder.
[0017] The two attributes described above allow the output multichannel audio signal to be efficiently encoded by a simple encoder by allocating fewer bits to encoding less dominant channels or by choosing to discard less dominant audio channels altogether. Brief Description of the Figures
[0018] The figures show specific arrangements or orderings of schematic elements, such as those representing devices, units, instruction blocks, and data elements, to facilitate description. However, it should be understood by those skilled in the art that the specific order or arrangement of schematic elements in the figures is not intended to imply that a particular order or sequence of processing, or separation of processes, is required. Furthermore, the inclusion of a schematic element in a drawing is not intended to imply that this element is required in all modalities or that the features represented by this element may not be included or combined with other elements in some implementations.
[0019] Furthermore, in drawings where connecting elements, such as solid or dashed lines or arrows, are used to illustrate a connection, relationship, or association between two or more schematic elements, the absence of these connecting elements is not intended to imply that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings to avoid complicating disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, where a connecting element represents communication of signals, data, or instructions, it should be understood by those skilled in the art that this element represents one or more signal paths, as necessary, to affect the communication.
[0020] Figure 1 is a block diagram of an arrangement of a simple audio encoder and a simple audio decoder intended to form an output multichannel audio signal that is a facsimile of an input multichannel audio signal, according to some modalities.
[0021] Figure 2 is a block diagram of an audio codec system that includes an audio encoder, audio decoder 106, encoder premixer, and decoder postmixer, according to some modalities.
[0022] Figure 3 illustrates an arrangement of processing elements whereby an input multichannel audio signal is divided by a filter bank into subband signals, where each subband is processed by a mixing matrix to produce a remixed subband signal, according to some modalities.
[0023] Figure 4 is a block diagram of a two-operation mixing arrangement intended to implement the function of the encoder premixer of Figure 2 or the encoder premixer of Figure 3 according to some modalities.
[0024] Figure 5 is a block diagram of a prediction mixer, according to some modalities.
[0025] Figure 6 shows an arrangement of processing elements that implement the decoder postmixer of Figure 2 according to some modalities.
[0026] Figure 7 is a flow diagram of an adaptive downmixing process of audio signals with enhanced continuity, according to some modalities.
[0027] Figure 8 is a block diagram of a system to implement the features and processes described with reference to Figures 1-7 according to some modalities.
[0028] The same reference symbol used in several drawings indicates similar elements. Detailed Description of the Invention
[0029] In the following detailed description, several specific details are established to provide a complete understanding of the various modalities described. It will be evident to someone skilled in the art that the various implementations described can be practiced without these specific details. In other cases, well-known methods, procedures, components, and circuits were not described in detail so as not to unnecessarily complicate aspects of the modalities. Several features are described below, which can be used individually or in any combination with other features. Nomenclature
[0030] As used herein, the term “includes” and its variants should be read as open terms meaning “includes, but is not limited to.” The term “or” should be read as “and” unless the context clearly indicates otherwise. The term “based on” should be read as “based at least in part on.” The terms “an example implementation” and “an example implementation” should be read as “at least one example implementation.” The term “another implementation” should be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” should be read as “obtain,” “receive,” “compute,” “calculate,” “estimate,” “predict,” or “derive.” Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by a person skilled in the art to which this disclosure pertains.
[0031] Figure 1 is a block diagram of an arrangement 10 of a simple audio encoder and a simple audio decoder, intended to form a multichannel audio signal 17 (Z) that is a facsimile of the multichannel audio signal 13 (Z). The multichannel audio signal 13 is processed by the simple audio encoder 14 to produce the encoded representation 15, which can be stored and / or transmitted 20 to the simple audio decoder 16 that produces the multichannel audio signal 17. Preferably, the data size of the encoded representation 15 is minimized zQQRnn / Qznz / q / υιλι while minimizing the difference between the multichannel audio signal 13 and the multichannel audio signal 17. Furthermore, the difference between the multichannel audio signal 13 and the multichannel audio signal 17 can be measured according to the similarity as perceived by a human listener.The measurement of the human-perceived similarity between audio signal 13 and audio signal 17 is based on a reference playback method (i.e., the assumed default means by which the audio channels of multichannel audio signals 13,17 are presented as an auditory experience for the listener).
[0032] The efficiency of the simple audio encoder 14 and decoder 16 can be defined in terms of the data rate (measured in bits per second) of the encoded representation 15 required to provide the multichannel audio signal 17 that will be judged by a listener to match the multichannel audio signal 13 with a particular perceived quality level. The simple audio encoder 14 and decoder 16 can achieve greater efficiency (i.e., a lower data rate) when the multichannel audio signal 13 is known to possess particular attributes. In particular, greater efficiency can be achieved when the multichannel audio signal 13 is known to possess the following attributes (DD1 and DD2):
[0033] DD1: One or more channels of the multichannel audio signal are generally more dominant than others, where a more dominant audio channel is one that will contain substantial elements of most (or all) of the sonic elements in the scene. That is, a dominant audio signal, when presented as a single audio channel to a listener, will contain most (or all) of the sonic elements of the multichannel signal when the multichannel audio signal is presented to a listener through a reference playback method.
[0034] DD2: Each of the audio channels of the multichannel audio signal are not highly correlated with each of the other audio channels.
[0035] Given the knowledge that the multichannel audio signal 13 possesses DD1 and DD2 attributes, the simple audio encoder 14 can achieve improved efficiency using various techniques, including but not limited to: allocating fewer bits to the encoding of less dominant channels or choosing to discard less dominant channels altogether. The simple audio encoder 14 can provide data to the simple audio decoder 16 to assist in the regeneration of channels that were discarded by the simple audio encoder 14. Preferably, a multichannel audio signal that does not possess DD1 and DD2 attributes can be processed by an encoder premixer to form, for example, to calculate, determine, construct, or generate, a multichannel audio signal that does possess DD1 and DD2 attributes, as described later with reference to Figure 2.A corresponding decoder postmixer can be applied to the output of the simple decoder to form a multichannel output audio signal, so that the decoder postmixer performs an approximate inverse operation with respect to the operation of the encoder premixer.
[0036] Figure 2 is a block diagram of an audio codec system 100 comprising an audio encoder 104 and an audio decoder 106, an encoder premixer 102 and a decoder postmixer 108. The audio encoder 104 and the audio decoder 106 form a multichannel audio signal 109 (X) which is a facsimile zQQRnn / Qznz / q / υιλι of the multichannel audio signal 101 (X). Preferably, the size of the data of the encoded representation 105 is minimized while minimizing the difference between the multichannel audio signal 101 and the multichannel audio signal 109. Furthermore, the difference between the multichannel audio signal 101 and the multichannel audio signal 109 can be measured according to the similarity as perceived by a human listener.
[0037] The measurement of the human-perceived similarity between multichannel audio signal 101 and multichannel audio signal 109 is based on a reference playback method (i.e., the assumed default means by which the audio channels of audio signals 101, 109 are presented as an auditory experience to the listener). The efficiency of multichannel audio encoder 104 and multichannel audio decoder 106 can be defined in terms of the data rate (measured in bits per second) of the encoded representation 105 that provides a multichannel audio signal 109 that will be judged by a listener to match multichannel audio signal 101 with a particular perceived quality level.
[0038] With reference to Figure 2, the input multichannel audio signal 101 is mixed according to the encoder premixer 102 (R) to produce the output multichannel audio signal 103 (Z) which is processed by the simple audio encoder 104 to produce the encoded representation 105, which can be stored and / or transmitted 110 to the simple audio decoder 106, which produces the multichannel audio signal 107 (Z'). The multichannel audio signal 107 is processed by the decoder postmixer 108 (R') to produce the decoded multichannel audio signal 109. The encoder premixer 102 provides metadata 112 (Q) which includes information necessary to determine the behavior of the decoder postmixer 108. The metadata 112 can be stored and / or transmitted 110 with encoded representation 105.Measuring the efficiency of the multichannel audio encoder 104 and the multichannel audio decoder 106 may include the size of the metadata 112 (commonly measured in bits per second), as will be appreciated by those skilled in the art.
[0039] The multichannel audio signal 101 may consist of N audio channels where significant correlations may exist between some pairs of channels, and where no individual channel can be considered a dominant channel. That is, the multichannel audio signal 101 may not possess attributes DD1 and DD2, and therefore the multichannel audio signal 101 might not be a suitable signal for encoding and decoding using a simple audio encoder 104 and a decoder 106, respectively.
[0040] Preferably, the encoder premixer 102 is adapted to process the input multichannel audio signal 101 to produce the output multichannel audio signal 103, wherein the output multichannel audio signal 103 has attributes DD1 and DD2. Given the input multichannel audio signal X composed of / V channels: / *i(O\ X(t) = ( j [1] the multichannel output audio signal Z is calculated as: zQQRnn / Qznz / q / υιλι CO\ Z(t) =[2]\ZN(t) / = R(t)^X(tY [3]
[0041] The coefficients of the encoder premixer matrix R can vary with time, and therefore R can be considered a function of time. The values of the elements of R can be calculated at regular intervals (e.g., where the interval can be 20 ms, or a value between 1 ms and 100 ms) or at irregular intervals. When the values of the elements of R are changed, the change can be interpolated without problems.In the following analysis, references to R are to be treated as references to a time-variable encoder premixer R(t) and references to R' are to be treated as references to a time-variable decoder premixer R'(t\
[0042] In one mode, the encoder premixer 102 can make use of mixing coefficients, Rb(t) to process the audio signal components in a band b, where 1 < b < B: Figure 3 illustrates an arrangement of processing elements 150 whereby the multichannel audio signal 151 (X) is divided by the filter bank 152 into B subband signals, X1!(t),(t), ...X[B1(t), with each subband signal (e.g., 153 (xW(t))) processed by a mixing matrix (e.g., 154 (R^) to produce a subband signal remixed (for example, 155 [Z^ (t))). The remixed subband signals, Z[1\t\ Z[2\t), ...Z[B](t), are recombined by combiner 156 to form the multichannel audio signal 157 (Z).
[0043] For the purpose of the following analysis, references to the matrix R(t) may be interpreted as references to β^ζί), where b refers to a subband. It will be appreciated that the analysis that follows can be applied to signals that are processed in subbands, or to signals that are processed without subband processing. It will be appreciated by those skilled in the art that many methods can be used to process audio signals according to subbands, and the analysis of the matrix R will be applied to those methods.
[0044] With reference to Figure 2, R mixes the channels of the multichannel audio signal 101 to produce the multichannel audio signal 103, which has attributes DD1 and DD2, as described above, thus enabling the encoder 106 to achieve improved data efficiency. The decoder's premixer 108 (R') provides a mixing operation that is the inverse of mixer R, such that: X'(t) = R'(t)xZ'(t). [4]
[0045] Figure 4 is a block diagram of a 200 arrangement of two mixing operations intended to implement the function of the encoder premixer 102 (R) of Figure 2 or the encoder premixer Rb of Figure 3. The N-channel multichannel input signal 201 (X) is mixed by the mixing matrix 202 (M) to produce the W-channel intermediate signal 203 (Y), which is then processed by the mixer 204 (P) to produce the N-channel signal 205 (Z). Signals 201 (X) and 205 (Z) in Figure 4 are intended to correspond respectively to input signals 101 (X) and 103 (Z) in Figure 2, or to subband signals 153 (¾ (O) and 155 (¾ (t)) in Figure 3.
[0046] Analysis block 210 (A) takes the signal input 201 and calculates the coefficients 212 to be used to adapt the operation of the mixer 204. Analysis block 210 also produces the metadata 211 (Q), corresponding to the metadata 112 in Figure 2, which will be provided to the decoder, as 113 (Q), for use by the decoder postmixer 108.
[0047] It will be seen from the arrangement of mixers 202 and 204 in figure 4, that the matrix R will be: R(t) = xm [5] where the matrix P(t) can vary with time.
[0048] Therefore: Y(t) = Mx X(t) Z(t) = P(t)x Y(t) [6]-[9] = P(t) x M x X(t) = R(t) XX(t)
[0049] The matrix M is adapted to ensure that the intermediate signal 203 (Y) possesses the attribute DD1. That is, the N-channel signal 203 (Y) contains a channel that can be considered a dominant channel. Without loss of generality, the matrix M is adapted to ensure that the first channel, Y^t), is a dominant channel. Hereafter, when the first channel of a multichannel signal is a dominant channel, this first channel will be referred to as a primary channel. The primary channel may also be referred to as a proper channel in some contexts.
[0050] The matrix M of [N x / V] can be determined from the expected covariance matrix [N x N] Cov of the N-channel input signal, X(t'): Cov = E(X(t) x X(t)H) where the operation X(t)H indicates the Hermitian transpose of the column vector of length- / VX(ty and the operation E() indicates the expected value of a variable quantity.
[0051] Expected values, as used in Equation
[10] , can be estimated based on assumed characteristics of typical multichannel input audio signals, or they can be estimated by statistical analysis of a set of typical multichannel input audio signals.
[0052] The covariance matrix, Cov, can be factored according to self-analysis, as will be familiar to those proficient in the technique: Cov = V x DxVH
[12] where the l matrix is a unitary matrix and the D matrix is a diagonal matrix with the diagonal elements being non-negative real values arranged in descending order. zoQAnn / eznz / q / υιλι
[0053] The matrix M can be chosen to be: M = Vh
[13]
[0054] It will be appreciated by those skilled in the technique that the covariance matrix, Cov, will depend on the panning methods used to form the original input signal X(t), as well as the typical use of panning methods used by typical signal creators.
[0055] As an example, when the original input signal is a 2-channel stereo signal intended for playback on stereo speakers, typical panning rules used by content creators will result in some audio objects being panned to the first channel (in this context, this is often referred to as the left channel), some audio objects being panned to the second channel (in this context, this is often referred to as the right channel), and some objects being panned simultaneously to both channels. In this case, the covariance matrix might look something like this: for stereo L / R : Cov = (θ'θ θ θ)Π^] and according to equations
[12] and
[13] : (ii \ j
[15] 72 /
[0056] The matrix M in equation
[15] will be familiar to those skilled in the art as a mixing matrix suitable for converting the original input audio signal X in stereo L / R format into an intermediate signal Z in Mid / Side format. It will also be appreciated by those skilled in the art that the first channel of Z (often referred to as the Mid signal in this case) is a dominant audio signal (the primary channel), which has the property that most of the audio elements in a stereo mix will be present in the Mid signal.
[0057] As an alternative example, when the original input signal is a 5-channel surround signal intended for playback on a common five-speaker array, typical panning rules used by content creators will result in some audio objects being panned to one of the five channels, and some objects being panned simultaneously to two or more channels. In this case, the covariance matrix might look something like this: / 1.500 0.595 1.155 1.155 0.595\ 0.595 1.500 1.155 0.595 1.155 for 5 channels: Cov = 1.155 1.155 1.500 0.595 0.595 ,
[16] 1.155 0.595 0.595 1.500 1.155 \0.595 1.155 0.595 1.155 1.500 / and according to equations
[12] and
[13] : / 0.447 -0.195 0.447 -0.195 0.447 -0.632 0.447 0.447 \ 0.512 0.512 for 5 channels: M = 0.602 -0.602 0.000 0.372 -0.372 •
[17] -0.512 -0.512 0.632 0.195 0.195 \—0.372 0.372 0.000 0.602 -0.602 /
[0058] It will be noticed that the top row of the matrix M in equation
[17] is composed of similar (or identical) positive values. This means that, according to Equation [6], the first channel of the intermediate signal Y(t) will be formed by the sum of the five channels of the original input audio signal, Y(t), and this ensures that all the sonic elements that are displaced in the original input audio signal will be present in ^(t) (the first channel of the N-channel signal 7(t)). Therefore, this choice of matrix Ai ensures that the intermediate signal Y possesses the attribute DD1 (^(t) is a primary channel).
[0059] In a further alternative example, when the input multichannel audio signal, X(t), already contains a dominant channel (and, without loss of generality, the first channel, X^t), is assumed to be dominant), the matrix M can be an identity matrix [IVxlV]. In a more specific example of an input multichannel audio signal with a dominant / primary first channel, the input multichannel audio signal can represent an acoustic scene encoded in an ambisonic format (a means of encoding acoustic scenes that will be familiar to those skilled in the art).
[0060] The matrix 212 (P(t)) is computed by the analysis block 210 (A) in Figure 4, at time t, according to the following process: 1. Determine the covariance of the intermediate signal 7(t) at time t. An example of a method for calculating covariance is: CovY^ =^%Y(t) xY^H
[18]
[0061] Alternatively, the covariance of the intermediate signal 7(t) can be calculated from the covariance of the input multichannel audio signal X(t), as: Covy(t) = M x Covx(t) x Mh,
[19] where Covx(V =1-f^X(t)xX(tr.
[20] 2. From the covariance matrix of [L x L], CovY(t), extract the scalar quantity w = [Covr(t)]1;1, the column vector of [N x 1] v = [CovY(t)]2..L,i and the matrix of [N x N] E == [Covr(t)]2ii2 where N = L - 1, and: Cov^t} = (wv\
[21] Vp E' 3. Determine the quantities a, β and the vector of [ / V x 1] of the mixing coefficients u: a = |v|2=
[22] u=ϊνt23] β = uHx E xu
[24]
[0062] 4. Given the quantities w, ay, and β, solve equation
[25] to determine the input mixture intensity coefficient h and the prediction mixture intensity coefficient g\ βh2g + 2ahg — ph — a + gw = 0
[25] where the solutions to this equation will also satisfy a pre-prediction constraint equation. An example of a pre-prediction constraint equation is: PPC1: h = fg,
[26] where f is a predetermined constant value that satisfies 0 ≤ β < 1.
[0063] When using the pre-prediction constraint PPC1, equation
[25] can be modified to be: Pf2g3+ 2afg2- βfg - a + gw = 0
[27] and equation
[27] can be solved for the largest actual value of and thus the value of h can be determined using equation
[26] , 5. Form the matrix of [¿ x L] Q as: / 0 0...0 \ Q = L0- Ί.
[28] V 0 ... o / 6. Form the matrix of [L x ¿] P(t) as: P^ = (IL-gQ)x^L+ hQH)
[29] where ILes the identity matrix [L x ¿].
[0064] The metadata 211 (Q) in Figure 4 can transmit information that will allow the unit vector uy and the coefficients g and h to be determined by the postmixer 113 of the decoder in Figure 2.
[0065] The solution for g of equation
[27] can be approximated by choosing an initial estimate g1= le and iterating (according to Newton's method, as known in the technique) several times: „ _ „ _ rom ^+19k 3f2gk+4afgk-pf+w ' so that a reasonable approximation for the solution can be found starting from g = g5. It will be appreciated that other methods are known in the technique for finding approximate solutions to the cubic equation
[27] ,
[0066] According to an alternative modality, the matrix P(f) of [L x ¿] can be determined, at time t, by determining a vector of pV x 1] u indicative of the correlation between the primary channel of the intermediate signal Y(t) and the remaining N non-primary channels, and determining the input mixing intensity coefficient h and the prediction mixing intensity coefficient g to form P(t) according to Equation
[28] , so that the signal Z(t) = P(t) x Y(t) will possess attributes DD1 and DD2.
[0067] The determination of coefficients gyh can be governed by a pre-prediction constraint equation. An example of a pre-prediction constraint equation is given (PPC1) in Equation
[26] . A preferred choice for the coefficient f may be f = 0.5, but values of f in the range 0.2 < / < 1 may be appropriate for use. zQQRnn / Qznz / q / υιλι
[0068] In an alternative mode, the following pre-prediction restrictions can be used: PPC2: when: — < cw otherwise
[31] zoQAnn / eznz / q / υιλι where c is a predetermined constant. A typical value might be c = 1, but values of c can be chosen in the range 0.25 < c < 4.
[0069] According to the PPC2 restriction in equation
[31] , the solution to equation
[25] is: when: — < c wq= £,
[32] -
[33] h =0 otherwise: β-2ca+gβ __
[34] -
[35] 2+4a2c2-4c2^w 2c β
[0070] Figure 5 is a block diagram of a 300 predictive mixer, according to some modalities. The matrix terms, (IL- gQ} and (JL+ hQH) of Equation
[29] can be implemented by the prediction mixer 300, wherein, in this example, the signal r(t) is composed of 4 channels (L = 4), the first channel 301 (PJ) being a primary channel and the remaining 3 non-primary channels 302 (e.g., Y2, Y3, Y4) are scaled according to the three input gains 312 (H2, H3, and Y4) to form the scaled input signal components (e.g., 304). The scaled input signal components are summed 305 with the primary input channel 301 (Y1) to form the primary output 306 (Zj). The primary output 306 is scaled by the three prediction gains 313 (G2, G3, and G^) to form three prediction signals (e.g., 311).Each prediction signal is subtracted (e.g., 308 and 309) from the respective input (e.g., Y2302) to form the respective non-dominant output 310 (Z2}.
[0071] The three inlet gains 312 (H2, H3y) can be determined from the mixing coefficients u (determined according to equation
[23] ) and the inlet mixing intensity coefficient h (according to the solution to equation
[25] ), where: / H2\ I H3j = hu.
[36] \h4J
[0072] The three prediction gains 313 (G2, G3 and G4) can be determined from the mixing coefficients u (determined according to equation
[23] ) and the prediction mixing intensity coefficient g (according to the solution to equation
[25] ), where: í&2\ G3= gu.
[37] \Gj
[0073] It will be appreciated by those skilled in the art that the arrangement of the linear matrix operations M 202 and P 204 of figure 4 can be implemented using a single matrix R = P x M.
[0074] It will be appreciated by those skilled in the art that the decoder matrix R' of Figure 2 can be formed from the matrices M' (the inverse of M) and P' (the inverse of P): R'(f) = M' xP'(t),
[38] and M' can be precalculated (without varying as a function of time) and P' can be formed by the method: P' = (IL- hQHj x (IL+ gQ).
[39]
[0075] Figure 6 shows an arrangement 400 of processing elements implementing a decoder postmixer 108 in Figure 2. The metadata 402 (Q) provides information to the inverse prediction determination block 403 (B), which calculates the coefficients needed to determine the operation of the inverse predictor 405 (P'). The signal 401 (Z') is processed by the inverse predictor 405 (P') to produce the intermediate signal 406 (Y'), which is then processed by the matrix 407 (M1) to produce the output signal 408 X'. Example process
[0076] Figure 7 is a flowchart of an adaptive downmixing process 700 for audio signals with enhanced continuity, according to several modalities. Process 700 can be implemented, for example, using the system 800 shown in Figure 8.
[0077] Process 700 includes the steps of: receiving a multichannel input audio signal comprising one primary input audio channel and L non-primary input audio channels (701); determining a set of L input gains, where L is a positive integer greater than one (702); for each of the L non-primary input audio channels and L input gains, forming a respective scaled non-primary input audio channel from the respective scaled non-primary input audio channel according to the input gain (703); forming an output primary audio channel from the sum of the input primary audio channel and the scaled non-primary audio channels (704); determining a set of L prediction gains for each of the L prediction gains (705); forming a prediction channel from the scaled primary output audio channel according to the prediction gain (706);forming L non-primary output audio channels from the difference between the respective non-primary input audio channel and the respective prediction signal (707); forming a multi-channel output audio signal from the primary output audio channel and the L non-primary output audio channels (708); encoding the multi-channel output audio signal (709); and transmitting or storing the encoded multi-channel output audio signal (710). Each of these steps is described more fully with reference to Figures 1-6. Example system architecture
[0078] Figure 8 shows a block diagram of an example 800 system for implementing the features and processes described with reference to Figures 1-7 according to one modality. The 800 system includes any device capable of playing audio, including but not limited to: smartphones, tablets, wearables, vehicle computers, game consoles, surround sound systems, and kiosks.
[0079] As shown, the 800 system includes a central processing unit (CPU) 801 that is capable of performing various processes according to a program stored in, for example, a read-only memory (ROM) 802 or a program loaded from, for example, a storage unit 808 into random-access memory (RAM) 803. In RAM 803, the data required when the CPU 801 performs the various processes is also stored, as needed. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 809. An input / output (I / O) interface 805 is also connected to bus 804.
[0080] The following components connect to the I / O interface 805: an input unit 806, which may include a keyboard, mouse, or the like; an output unit 807, which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 808, which includes a hard disk or other suitable storage device; and a communication unit 809, which includes a network interface card such as a network card (for example, wired or wireless).
[0081] In some implementations, the 806 input unit includes one or more microphones in different positions (depending on the host device) that allow the capture of audio signals in various formats (e.g., monophonic, stereo, spatial, immersive, and other suitable formats).
[0082] In some implementations, the 807 output unit includes systems with varying numbers of speakers. As shown in Figure 8, the 807 output unit (depending on the capabilities of the host device) can render audio signals in various formats (e.g., monophonic, stereo, immersive, binaural, and other suitable formats).
[0083] The communication unit 809 is configured to communicate with other devices (for example, over a network). A unit 810 is also connected to the I / O interface 805, as required. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media, is mounted on unit 810 so that a computer program read from it is installed on storage unit 808, as required. A person skilled in the art will understand that although the 800 system is described as including the components described above, in actual applications it is possible to add, remove, and / or replace some of these components, and all such modifications or alterations are within the scope of this disclosure.
[0084] The aspects of the systems described herein may be implemented in a computer-based sound processing network environment suitable for processing digital or digitized audio files. Portions of the adaptive audio system may include one or more networks comprising any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route transmitted data between the computers. This network may be built upon various different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0085] According to example embodiments in this disclosure, the processes described above can be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments in this disclosure include a computer program product that includes a computer program tangibly embedded on a machine-readable medium, the computer program that includes program code for performing methods. In these embodiments, the computer program can be downloaded and mounted from the network through the communication unit 809, and / or installed from removable media 811, as shown in Figure 8.
[0086] Generally, several example modes of this disclosure can be implemented in hardware or special-purpose circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be run by control circuitry (e.g., a CPU in combination with other components in Figure 8); therefore, the control circuitry can perform the actions described in this disclosure. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be run by a controller, microprocessor, or other computing device (e.g., control circuitry).While various aspects of the example modalities in this disclosure are illustrated and described using block diagrams, flowcharts, or some other illustrated representation, it will be appreciated that the blocks, devices, systems, techniques, or methods described herein may be implemented in, as non-limiting examples, special-purpose hardware, software, firmware, circuits or logic, general-purpose hardware or controller, or other computer devices, or some combination thereof.
[0087] In addition, several blocks shown in the flowcharts can be viewed as method steps and / or as operations resulting from the operation of the computer program code and / or as a plurality of coupled logic circuit elements constructed to carry out the associated functions. For example, modalities of the present disclosure include a computer program product comprising a computer program tangibly embodied in a machine-readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0088] In the context of disclosure, a machine-readable medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction-executing system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transient and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof.More specific examples of machine-readable storage media would include an electrical connection having one or more wires, a laptop floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc (CD-ROM) read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0089] The computer program code for carrying out the methods of this disclosure may be written in any combination of one or more programming languages. Such computer program codes may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data-processing apparatus having control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data-processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented.The program code can be executed entirely on one computer, partially on the computer as a standalone software package, partially on the computer and partially on a remote computer, or entirely on the remote computer or server, or distributed across one or more remote computers and / or servers.
[0090] While this document contains many specific implementation details, these should not be interpreted as limitations on the scope of what can be claimed, but rather as descriptions of features that may be specific to particular modalities. Certain features described in this specification in the context of separate modalities may also be implemented in combination in a single modality. Conversely, several features described in the context of a single modality may also be implemented in multiple separate modalities or in any suitable subcombination.Furthermore, although the features may be described above as acting in certain combinations and even initially claimed as such, one or more features of a claimed combination may, in some cases, be removed from the combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination. The logical flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desired results. Moreover, other steps may be provided, or steps may be removed, from the described flows, and other components may be added to or removed from the described systems. Accordingly, other implementations are within the scope of the following claims.
Claims
1. An audio coding method, comprising: receiving, with at least one processor, a multichannel input audio signal comprising one primary input audio channel and L non-primary input audio channels; determining, with the at least one processor, a set of L input gains, where L is a positive integer greater than one; for each of the L non-primary input audio channels and L input gains, forming a respective scaled non-primary input audio channel from the respective scaled non-primary input audio channel according to the input gain; forming a primary output audio channel from the sum of the primary input audio channel and the scaled non-primary input audio channels;Determine, using at least one processor, a set of L prediction gains; for each of the L prediction gains, form, using at least one processor, a prediction channel of the primary output audio channel scaled according to the prediction gain; form, using at least one processor, L non-primary output audio channels from the difference between the respective non-primary input audio channel and the respective prediction signal; form, using at least one processor, a multichannel output audio signal from the primary output audio channel and the L non-primary output audio channels; encode, using an audio encoder, the multichannel output audio signal; and transmit or store, using at least one processor, the encoded multichannel output audio signal.
2. The method according to claim 1, wherein the determination of the set of L inlet gains comprises: determining a set of L mixing coefficients; determining an inlet mixing intensity coefficient; and determining the L inlet gains by scaling the L mixing coefficients by the inlet mixing intensity coefficient.
3. The method according to claim 2, wherein the determination of the set of L prediction gains comprises: determining a set of L mixing coefficients; determining a prediction mixing intensity coefficient; and determining the L prediction gains by scaling the L mixing coefficients by the prediction mixing intensity coefficient.
4. The method according to claim 3, wherein the input mixing intensity coefficient, h, is determined by a pre-prediction constraint equation, h=fg, where f is a predetermined constant value greater than zero and less than or equal to one, and g is the prediction mixing intensity coefficient.
5. Method according to claim 4, wherein the prediction mixing intensity coefficient, g, is a larger real-valued solution for: Pf2g3 + Zafg2 - βfg - a + gw = 0, where β = uH x E xu, u =^v,a = |v|2 = y 'a quantity w, the column vector vy, and the matrix E are components of a covariance matrix for an intermediate signal having a dominant channel.
6. The method according to claim 5, wherein the intermediate signal covariance matrix is calculated from a multichannel input audio signal covariance matrix.
7. The method according to claim 2 or 3, wherein two or more multichannel input audio channels are processed according to a mixing matrix to produce the primary input audio channel and the L non-primary input audio channels.
8. The method according to claim 7, wherein the primary input audio channel is determined by a dominant eigenvector of an expected covariance of a typical input multichannel audio signal.
9. The method according to claim 2 or 3, wherein each of the L mixing coefficients is determined based on a correlation of a respective channel of the non-primary input audio channels and the primary input audio channel.
10. The method according to claim 1, wherein the encoding includes assigning more bits to the primary output audio channel than to the L non-primary output audio channels, or discarding one or more of the L non-primary output audio channels.
11. A system, comprising: one or more computer processors; and a non-transient computer-readable means for causing the one or more computer processors to perform the method according to any one of claims 1-10.
12. A non-transient, computer-readable audio encoding medium comprising the method according to any of claims 1-10 above.