OPTIMIZED CODING OF INFORMATION REPRESENTATIVE OF A SPATIAL IMAGE OF A MULTI-CHANNEL AUDIO SIGNAL
Patent Information
- Application Number
- DE602021036148
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-30
- Filing Date
- 2021-06-23
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Existing coding methods for ambisonic signals, such as multi-mono and parametric coding, result in spatial distortions and degradation of the sound scene due to the lack of consideration for channel correlations, leading to artifacts like phantom sound sources and diffuse noise.
A method involving the decomposition of covariance matrices into eigenvalues and eigenvectors, followed by quantization of these parameters in the Euler angle or quaternion domain, to optimize the coding rate and reduce spatialization distortions.
This approach allows for a significant reduction in coding rate while maintaining a spatial image close to the original sound scene, achieving improved spatial accuracy and reduced distortion in the decoded sound.
Description
[0001] The present invention relates to the coding / decoding of spatialized sound data, in particular in an ambiophonic context (hereinafter also referred to as “ambisonic”).
[0002] The codecs currently used in mobile telephony are mono (a single signal channel for playback on a single speaker). The 3GPP EVS codec (for "Enhanced Voice Services") provides "Super-HD" quality (also called "High Definition Plus" or HD+ voice) with a super-wideband (SWB) audio band for signals sampled at 32 or 48 kHz or fullband (FB) for signals sampled at 48 kHz; the audio bandwidth is 14.4 to 16 kHz in SWB mode (9.6 to 128 kbit / s) and 20 kHz in FB mode (16.4 to 128 kbit / s).
[0003] The next quality evolution in the conversational services offered by operators should be immersive services, using terminals such as smartphones equipped with several microphones or spatial audio conferencing equipment or videoconferencing such as telepresence or 360° video, or even equipment for sharing “live” audio content, with a 3D spatial sound rendering that is much more immersive than a simple 2D stereo reproduction. With the increasingly widespread use of listening on mobile phones with an audio headset and the appearance of advanced audio equipment (accessories such as a 3D microphone, voice assistants with acoustic antennas, virtual reality headsets, etc.) the capture and rendering of spatial sound scenes are now widespread enough to offer an immersive communication experience.
[0004] As such, the future 3GPP “IVAS” standard (for “Immersive Voice And Audio Services”) proposes the extension of the EVS codec to immersive audio by accepting as codec input format at least the spatial sound formats listed below (and their combinations): Multichannel format (channel-based in English) of stereo or 5.1 type where each channel feeds a loudspeaker (for example L and R in stereo or L, R, Ls, Rs and C in 5.1); Object format (object-based in English) where sound objects are described as an audio signal (generally mono) associated with metadata describing the attributes of this object (position in space, spatial width of the source, etc.), Ambisonic format (scene-based in English) which describes the sound field at a given point, generally captured by a spherical microphone or synthesized in the domain of spherical harmonics.
[0005] The following is typically concerned with the coding of a sound in Ambisonic format, as an example of an embodiment (at least certain aspects presented in connection with the invention below can also be applied to formats other than Ambisonic).
[0006] Ambisonics is a method of recording ("coding" in the acoustic sense) spatialized sound and a system of reproducing ("decoding" in the acoustic sense). An ambisonic microphone (of order 1) comprises at least four capsules (typically cardioid or subcardioid) arranged on a spherical grid, for example the vertices of a regular tetrahedron. The audio channels associated with these capsules are called the "A-format". This format is converted into a "B-format", in which the sound field is decomposed into four components (spherical harmonics) denoted W, X, Y, Z, which correspond to four coincident virtual microphones. The W component corresponds to an omnidirectional capture of the sound field while the X, Y and Z components, more directional, are similar to pressure gradient microphones oriented along the three orthogonal axes of space.An Ambisonics system is flexible in that recording and playback are separate and decoupled. It allows decoding (in the acoustic sense) on any loudspeaker configuration (e.g., binaural, 5.1 surround sound, or 7.1.4 periphery (with elevation). The Ambisonics approach can be generalized to more than four channels in B-format, and this generalized representation is commonly called "HOA" (for "Higher-Order Ambisonics"). Decomposing the sound into more spherical harmonics improves the spatial accuracy of reproduction when rendered on loudspeakers.
[0007] An ambisonic signal of order M includes K=(M+1) 2< components and, at order 1 (if M=1), we find the four components W, X, Y, and Z, commonly called FOA (for First-Order Ambisonics). There is also a so-called "planar" variant of ambisonics (W, X, Y) which decomposes the sound defined in a plane which is generally the horizontal plane (where Z=0). In this case, the number of components is K =2M+1 channels. First-order ambisonics (4 channels: W, X, Y, Z), planar first-order ambisonics (3 channels: W, X, Y), and higher-order ambisonics are all referred to hereinafter as "ambisonics" indiscriminately to facilitate reading, the treatments presented being applicable regardless of the planar type or not and the number of ambisonic components. From now on, we will call an "ambisonic signal" a B-format signal in a predetermined order with a certain number of ambisonic components.This also includes hybrid cases, where for example at order 2 we only have 8 channels (instead of 9) - more precisely, at order 2 we find the 4 channels of order 1 (W, X, Y, Z) to which we normally add 5 channels (usually denoted R, S, T, U, V), and we can for example ignore one of the higher order channels (for example R).
[0008] The signals to be processed by the coder / decoder are presented as successions of blocks of sound samples called “frames” or “sub-frames” below.
[0009] Furthermore, hereinafter, mathematical notations follow the following convention: Scalar: s or N (lower case for variables or upper case for constants) the operator Re(.) designates the real part of a complex number Vector: u (minuscule, gras) Matrix: A (majuscule, gras)
[0010] The ratings A T< And A H< indicates respectively the transposition and the Hermitian transposition (transposed and conjugate) of A. A one-dimensional discrete-time signal, s(i), defined over a time interval i=0, ..., L-1 of length L is represented by a row vector s = s 0 , … , s L − 1 .
[0011] We can also write: s = [s 0 ,..., S L-1 ] to avoid the use of parentheses. A discrete-time multidimensional signal, b (i), defined on a time interval i=0, ..., L-1 of length L and with K dimensions is represented by a matrix of size LxK: B = b 0 0 ⋯ b 0 L − 1 ⋮ ⋯ ⋮ b K − 1 0 ⋯ b K − 1 L − 1 .
[0012] We can also note: B = [B ij ], i=0,...K-1, j=0...L-1, to avoid the use of parentheses. The Cartesian coordinates (x,y,z) of a 3D point can be converted into spherical coordinates (r, Θ,φ), where r is the distance from the origin, Θ is the azimuth and φ the elevation. We use here, without loss of generality, the mathematical convention where the elevation is defined with respect to the horizontal plane (0xy); the invention can be easily adapted to other definitions, including the convention used in physics where the azimuth is defined with respect to the Oz axis.
[0013] Furthermore, we do not recall here the known conventions of the state of the art in Ambisonics concerning the order of Ambisonic components (including ACN for Ambisonic Channel Number, SID for Single Index Designation, FuMA for Furse-Malham) and the normalization of Ambisonic components (SN3D, N3D, maxN). More details can be found for example in the resource available online: https: / / en.wikipedia.ora / wiki / Ambisonic data exchange formats By convention, the first component of an ambisonic signal generally corresponds to the omnidirectional component W.
[0014] The simplest approach to encoding an Ambisonics signal is to use a mono encoder and apply it in parallel to all channels, possibly with different bit allocation for each channel. This approach is called "multi-mono" here. The multi-mono approach can be extended to multi-stereo encoding (where pairs of channels are encoded separately by a stereo codec) or more generally to the use of multiple parallel instances of the same core codec.
[0015] Since the multi-mono coding approach does not take into account the correlation between channels, it produces spatial distortions with the addition of various artifacts such as the appearance of phantom sound sources, diffuse noise or displacements of sound source trajectories. Thus, coding an ambisonic signal using this approach causes degradation of spatialization.
[0016] An alternative approach to the separate coding of all channels is given, for a stereo or multichannel signal, by parametric coding. For this type of coding, the input multichannel signal is reduced into a smaller number of channels, after a process called "downmix", these channels are encoded and transmitted and additional spatialization information is also encoded. Parametric decoding consists of increasing the number of channels after decoding the transmitted channels, using a process called "upmix" (typically implemented by decorrelation) and spatial synthesis based on the decoded additional spatialization information.
[0017] An example of stereo parametric coding is given by the 3GPP e-AAC+ codec.
[0018] An example of parametric coding for ambisonics is given by the DirAC codec (for "Directional Audio Coding" in English) of which there are several variants for coding at order 1 or higher orders. At order 1 (4 channels W, X, Y, Z) the DirAC method can take the W signal as a "downmix" signal and applies to the input ambisonic signal a time / frequency analysis to estimate two parameters per sub-band: the direction of the main source and the diffuse character of the scene. To do this, the active intensity vector is calculated at the time / frequency index interval ( n,f ) to within a normalization constant: I n f = Re W n f X * n f Re W n f Y * n f Re W n f Z * n f where * is the Hermitian conjugate, Re (.) corresponds to the real part. From the intensity vector, we estimate the direction of arrival (DoA) of the source:
[0019] Or ∠ gives the angle of the 3D vector and mathematical expectation, and we estimate the diffuse character of the scene by a “diffuseness” parameter defined for example as: ψ n = 1 − E I E I where ∥ . ∥ is the complex modulus.
[0020] In the case of higher ambisonic orders, the DirAC method divides the sound space into sectors S m , corresponding to a part of the sphere (unit). For each sector S m , a directional beamforming process extracts 3 channels X m , Y m , Z m , and an “omni” channel corresponding to the sum of the 3 channels called W m . Similar to the first-order DirAC method, for each sector the signal is coded with the sub-band spatialization parameters (DoA and diffuseness). For more details, see V. Pulkki et al., Parametric time-frequency domain spatial audio, Wiley, 2017, pp. 89-159. The “downmix” operation of existing parametric coding methods causes degradation of the spatialization and modifications of the spatial image of the original signal.
[0021] The DirAC approach described above seeks to re-spatialize one or more sources in space, with a limitation on the maximum number of sources. The reproduction of the sound scene during decoding is then not always optimal.
[0022] There is therefore a need to find a sound scene close to the original sound scene during decoding while optimizing the coding rate.
[0023] Here, the term "spatial image" is understood to mean a distribution of the sound energy of the ambisonic sound scene at different directions in space; the spatial image describes the sound scene and generally corresponds to positive quantities evaluated at different predetermined directions in space - these positive quantities can be interpreted as energies and are seen as such hereinafter.
[0024] A spatial image associated with an ambisonic sound scene therefore represents the sound energy (or more generally a positive quantity) as a function of different directions in space. Information representative of a spatial image can be, for example, a covariance matrix calculated between the channels of the multichannel signal or energy information associated with directions of origin of the sound (associated with directions of virtual loudspeakers distributed on a unit sphere). The energy information can be obtained according to different directions (associated with directions of virtual loudspeakers distributed on a unit sphere). For this, different spatial image calculation methods known to those skilled in the art can be used: SRP type method (for "Steered-Response Power" in English), MUSIC pseudo-spectrum, histogram of directions of arrival, etc.
[0025] In general, the representation of a spatial image in the form of a covariance matrix involves coding a matrix of size KxK with K(K+1) / 2 non-redundant coefficients. The energy information requires coding at least the energy in N=K discrete points distributed on a sphere; in practice a higher number of points (N>>K) must be defined to have a sufficiently precise and usable representation.
[0026] We are therefore particularly interested here in the problem of coding a covariance matrix. The known state-of-the-art approach is for example described in the article by Dai Yang et al., High-Fidelity Multichannel Audio Coding with Karhunen-Loève Transform, IEEE Trans. Speech and Audio Processing, vol. 11, no. 4, July 2003.
[0027] Coding a KxK covariance matrix is done by encoding K(K+1) / 2 values (corresponding to the lower or upper triangle - the matrix being symmetric) with a 16-bit floating point representation (per coefficient). For example, if a single 4x4 matrix (K=4 for the FOA) is coded per 20 ms frame, this corresponds to a data rate of 16 x 10 bits / 20 ms = 8 kbit / s. If several covariance matrices are transmitted per frame, this data rate becomes very significant.
[0028] US 2016 / 155448 A1 relates to multi-channel audio coding, and more specifically to discrete multi-channel audio coding and decoding techniques. In particular, this document relates to systems and a method for coding acoustic fields.
[0029] An audio encoder configured to encode a frame of a sound field signal comprising a plurality of audio signals is disclosed. The audio encoder comprises a transformation determination unit configured to determine an energy-condensing orthogonal transformation based on the frame of the sound field signal. Further, the encoder comprises a transformation unit configured to apply the energy-condensing orthogonal transformation to the frame of the sound field signal, and configured to provide a frame of a rotated sound field signal comprising a plurality of rotated audio signals.The audio encoder comprises a waveform encoding unit configured to encode a first rotated audio signal of the plurality of rotated audio signals, and a parametric encoding unit configured to determine a set of spatial parameters for determining a second rotated audio signal of the plurality of rotated audio signals based on the first rotated audio signal.
[0030] The invention improves the state of the art.
[0031] To this end, the invention relates to a method for coding a multi-channel sound signal, comprising the following steps: coding of at least one audio signal channel of the original multichannel signal or of a downmix of the original multichannel signal; division of the original multichannel signal into frequency sub-bands; determination of a covariance matrix per frequency sub-band, representative of a spatial image of the original multichannel signal; decomposition of the determined covariance matrices into eigenvalues; coding by quantization of the parameters resulting from the decomposition into eigenvalues comprising both eigenvalues and eigenvectors.
[0032] Thus, the coding of the covariance matrices by frequency band of the original multichannel signal will allow the decoder to reconstruct the sound scene as close as possible to that of the original signal by applying corrections to the transmitted signals.
[0033] The decomposition into eigenvalues of the covariance matrices and the coding of parameters resulting from this decomposition makes it possible to restrict the number of information items to be transmitted to the decoder and thus to optimize the coding rate of these parameters and to reduce distortion for a given budget.
[0034] According to one embodiment of the invention, the coding method makes it possible to decompose the K(K+1) / 2 degrees of freedom into two parts on which more efficient coding is possible: K(K-1) / 2 degrees of freedom (eigenvectors in the form of a rotation matrix in dimension K) + K degrees of freedom (eigenvalues).
[0035] Typically, we obtain a flow rate of around 2.5 kbit / s for the example of a 4x4 matrix and 20 ms frames.
[0036] In one embodiment, the eigenvalues are ordered before quantization and the quantization is performed according to differential scalar quantization.
[0037] Thus, the coding rate is further reduced to quantize these eigenvalues.
[0038] According to a first embodiment, the decomposition of a covariance matrix into eigenvalues is carried out according to the following steps: obtaining a matrix of eigenvectors Q such that C = Q HAS Q T< with C, the covariance matrix and Λ=diag(λ 1 ,..., λ K ), a diagonal matrix of eigenvalues; modification of the eigenvector matrix as a function of a determinant value of the eigenvector matrix Q; conversion of the eigenvector matrix Q into the domain of generalized Euler angles; the generalized Euler angles obtained being part of the parameters to be quantified.
[0039] The conversion in the domain of Euler angles allows to quantize the angles resulting from this conversion to encode the eigenvector matrix, which allows to reduce the coding rate for a given distortion or to reduce the distortion for a given rate. The quantization, for this first embodiment, is of low complexity.
[0040] Also, in a particular embodiment, the quantization of generalized Euler angles is carried out by uniform quantization.
[0041] According to a second embodiment, the decomposition of a covariance matrix into eigenvalues is carried out according to the following steps: obtaining a matrix of eigenvectors Q such that C = Q Λ Q T< with C, the covariance matrix and Λ=diag(λ 1 ,..., λ K ), a diagonal matrix of eigenvalues; modification of the eigenvector matrix as a function of a determinant value of the eigenvector matrix Q; conversion of the eigenvector matrix Q into the quaternion domain; at least one quaternion obtained being part of the parameters to be quantified.
[0042] The conversion into the quaternion domain allows the quaternions resulting from this conversion to be quantized to encode the eigenvector matrix, which makes it possible to reduce the coding rate for a given distortion or to reduce the distortion for a given rate. For this second embodiment, the quantization has a greater complexity but the spherical vector quantization used to quantize these parameters is more efficient than a scalar quantization.
[0043] Also, in a particular embodiment, the quantization of the quaternions is carried out by spherical vector quantization.
[0044] The invention also relates to a method for decoding a multi-channel sound signal, comprising the following steps: decoding at least one coded channel and obtaining a decoded multichannel signal; division of the decoded multichannel signal into frequency sub-bands; decoding of parameters resulting from a decomposition into eigenvalues of covariance matrices of the original multichannel signal; determination of the covariance matrices of the original multichannel signal from the decoded parameters: determination of a covariance matrix, by frequency sub-bands, of the decoded multichannel signal; determination of a set of corrections to be made to the decoded signal from the covariance matrices of the original multichannel signal (Inf. B ) and covariance matrices of the decoded multichannel signal (Inf. B̂ ); correction of the decoded multichannel signal by the determined set of corrections; Thus, the decoder can receive and decode the covariance matrices of the original multichannel signal with reduced distortion for a given rate compared to conventional methods using direct coding of the covariance matrix. These decoded covariance matrices then make it possible to determine corrections to be made to the decoded multichannel signal so that the spatial image of the decoded multichannel signal is as close as possible to the spatial image of the original multichannel signal.
[0045] The invention also relates to a coding device comprising a processing circuit for implementing the coding method as described above. The invention also relates to a decoding device comprising a processing circuit for implementing the decoding method as described above. The invention relates to a computer program comprising instructions for implementing the coding or decoding methods as described above, when they are executed by a processor.
[0046] Finally, the invention relates to a storage medium, readable by a processor, storing a computer program comprising instructions for executing the coding or decoding methods described above.
[0047] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which: [ Fig 1 ] There figure 1 illustrates an embodiment of an encoder and a decoder, an encoding method and a decoding method according to the invention; [ Fig 2 ] There figure 2 illustrates a detailed embodiment of the correction set determination block; [ Fig 3a ] There figure 3a illustrates in flowchart form an embodiment of the covariance matrix coding block according to an embodiment of the invention; [ Fig 3b ] There figure 3b illustrates in flowchart form an embodiment of the covariance matrix decoding block according to an embodiment of the invention; [ Fig 4 ] There figure 4 illustrates exemplary structural embodiments of an encoder and a decoder according to one embodiment of the invention.
[0048] We recall here the known technique of encoding (in the acoustic sense) a sound source in ambisonic format. A mono sound source can be artificially spatialized by multiplying the associated signal by the values of the spherical harmonics associated with its direction of origin (assuming the signal is carried by a plane wave) to obtain as many ambisonic components. To do this, we calculate the coefficients for each spherical harmonic for a given position in azimuth Θ and elevation φ at the desired order: B = Y θ φ . s where s is the mono signal to be spatialized and Y(Θ,φ) is the encoding vector defining the coefficients of the spherical harmonics associated with the direction (Θ,φ) for order M. An example of an encoding vector is given below for order 1 with the SN3D convention and the SID or FuMa channel order: Y θ φ = 1 cos θ cos φ sin θ cos φ sin φ
[0049] Other normalization conventions (e.g. maxN, N3D) and channel order (e.g. ACN) exist and the different implementation modes are then adapted according to the convention used for the order or normalization of the ambisonic components (FOA or HOA). This amounts to modifying the order of the lines Y ( Θ ,φ) or multiply these lines by predefined constants.
[0050] For higher orders, the coefficients Y( Θ ,φ) of the spherical harmonics can be found in the book by B.Rafaely, Fundamentals of Spherical Array Processing, Springer, 2015. Generally speaking, for an order M, the ambisonic signals are K=(M+1) 2< .
[0051] Similarly, we recall here some notions on the rendering or restitution of ambisonics by loudspeakers. An ambisonic sound is not made to be listened to as is; for immersive listening on loudspeakers or headphones, a "decoding" step in the acoustic sense also called rendering ("renderer" in English) must be done. We consider the case of N loudspeakers (virtual or physical) distributed on a sphere - typically of unit radius - and whose directions ( Θ n , φ n ), n=0, ..., N-1, in terms of azimuth and elevation are known. Decoding, as considered here, is a linear operation which consists of applying a matrix D to ambisonic signalsB to obtain the sn signals from the speakers, which can be gathered into a matrix S =[ s 0 , ... s N- 1 ], S=D.B Or S = s 0 ⋮ s N − 1 .
[0052] We can decompose the matrix D in line vectors d n , either D = d 0 ⋮ d N − 1 d n can be seen as a weighting vector for the nth speaker, used to recombine the components of the ambisonic signal and calculate the signal played on the nth speaker: s n =d n .B.
[0053] There are multiple methods of "decoding" in the acoustic sense. The so-called "basic decoding" method, also called "mode-matching," is based on the encoding matrix E associated with all virtual speaker directions: E = Y θ 0 φ 0 … Y θ N − 1 φ N − 1
[0054] According to this method, the matrix D is typically defined as the pseudo-inverse of E: D = pinv E = D T D . D T − 1
[0055] Alternatively, the method that can be called "projection" gives similar results for certain regular distributions of directions, and is described by the equation: D = 1 N E T
[0056] In this latter case, we see that for each direction of index n, d n = 1 N Y θ n φ n T
[0057] In the context of this invention, such matrices will serve as a beamforming matrix describing how to obtain signals characteristic of directions in space for the purpose of performing spatial analysis and / or transformations.
[0058] In the context of the present invention, it is useful to describe the reciprocal conversion for passing from the loudspeaker domain to the ambisonic domain. It is appropriate that the successive application of the two conversions exactly reproduces the original ambisonic signals if no intermediate modification is applied in the loudspeaker domain. The reciprocal conversion is therefore defined as involving the pseudo-inverse of D : pinv (D).S=D T< (D.D T< ) -1< .S
[0059] When K=(M+1) 2< , the matrix D of size KxK is invertible under certain conditions and in this case: B=D -1< .S
[0060] In the case of the "mode-matching" method, it appears that pinv( D )= E . In variants, other decoding methods by D can be used, with the reverse conversion E corresponding; the only condition to check is that the combination of decoding by Dand the reverse conversion by E must give a perfect reconstruction (when no intermediate processing is carried out between acoustic decoding and acoustic encoding).
[0061] Such variants are for example given by: “mode-matching” decoding with a regulation term in the form D T< (D.D T< +ε D I) -1< where ε D is a low value (for example 0.01), The “in phase” or “max-rE” decodings known from the state of the art or variants where the distribution of the directions of the loudspeakers is not regular on the sphere.
[0062] The method described below is based on the transmission of a spatial image representation in the form of a covariance matrix and the correction of spatial degradations, in particular to ensure that the spatial image of the decoded signal is as close as possible to the original signal. Unlike known parametric coding approaches for stereo or multichannel signals, where perceptual cues are coded, the invention does not rely on a perceptual interpretation of the spatial image information because the ambisonic domain is not directly "listenable".
[0063] In the embodiment described below, coding is performed, with optionally, a downmix / upmix using a mapping, hereinafter called spatial image of the original ambisonic sound scene. During coding, a certain number of channels (preferably less than the number of input channels) are transmitted to the decoder. These channels can be a subset of the original channels (e.g.: W or X channel, 4 channels of the 3D FOA, 3 channels of the planar FOA, etc.) or a rematrixing of the input channels (e.g.: stereo downmix from a FOA input). In addition to these channels, the coder transmits information from a mapping of the original sound scene.This information can be defined on a signal on a single frequency band (e.g.: 0-16000 or 100-14000 Hz for a signal sampled at 32 kHz) but in the preferred embodiment the spectrum is divided by sub-bands (which can be derived from existing cuts of the Bark, Mel, or other types as described later). According to an embodiment of the invention, the spatial image of the original sound scene is a covariance matrix as defined later. An optimized coding of this covariance matrix is provided in order to optimize the coding rate of this representation of a spatial image, especially when it is defined by sub-bands.
[0064] During decoding, the received and decoded signals are optionally extended by "upmix" (by decorrelation) described below. Depending on the type of information received, a mapping of the "degraded" sound scene is carried out. A transformation operation is determined to recreate the original sound scene. This transformation is determined at the decoder from the original mapping received and decoded (information of the spatial image of the original signal) and the degraded mapping (information of the spatial image of the decoded multichannel signal).
[0065] Alternatively, downmix / upmix can be replaced by direct channel coding, e.g. multi-mono or multi-stereo.
[0066] There figure 1 represents an exemplary embodiment of an encoder and a decoder according to the invention for the implementation respectively of the encoding and decoding methods according to an embodiment of the invention.
[0067] The original multi-channel signal B of dimension KxL (i.e. K components of L time or frequency samples) is at the input of the encoder.
[0068] We are interested here in the case of a multichannel signal in ambisonic representation, as described previously. The invention can also be applied to other types of multichannel signal such as a B-format signal with modifications, such as for example the suppression of certain components (e.g.: suppression of the R component at order 2 to keep only 8 channels) or the matrixing of the B-format to move into an equivalent domain (called "Equivalent Spatial Domain") as described in the 3GPP TS 26.260 specification - another example of matrixing is given by the "channel mapping 3" of the IETF Opus codec and in the 3GPP TS 26.918 specification (clause 6.1.6.3).
[0069] In the embodiment thus described, the input signal is sampled at 32 kHz. The encoder operates in frames which are preferably 20 ms long, i.e. L=640 samples per frame at 32 kHz. In variants, other frame lengths and sampling frequencies are possible (e.g. L=480 samples per frame of 10 ms at 48 kHz).
[0070] In a preferred embodiment, the coding of the spatial image is carried out in sub-bands in the frequency domain after temporal short-term discrete Fourier transform (STFT) (on one or more bands), however in variants, the invention can be implemented in sub-bands by applying a real or complex filter bank to process sub-bands in time, or by using another type of transform such as the modified discrete cosine transform (MDCT) or the Modulated Complex Lapped Transform (MCLT).
[0071] A block 110 for reducing the number of channels (DMX) is implemented optionally. For example, for an ambisonic input signal of order 1, it consists of keeping only the W channel and for an ambisonic input signal of order >1, of keeping only the first 4 ambisonic components W, X, Y, Z (therefore truncating the signal to order 1). Other types of downmix (selection of a subset of channels and / or matrixing, use of a “delay-sum” type “beamforming”) can be implemented without this modifying the method according to the invention.
[0072] Block 111 encodes the audio signal b' k , k=1, ..., K dmx (where K dmx ≤K) of B' at the exit of block 110.
[0073] In a preferred embodiment, block 111 uses multi-mono coding (MDC) with variable allocation, where the core codec is the 3GPP EVS standardized codec. In this multi-mono approach, each channel b'k is coded separately by an instance of the codec; however, in variants other coding methods are possible, for example multi-stereo coding or joint multichannel coding. At the output of this coding block 111, at least one coded channel of an audio signal from the original multichannel signal is therefore obtained, in the form of a bitstream which is sent to the multiplexer 140.
[0074] Block 120 extracts a given frequency band (which may correspond to the full-band signal or in a restricted band) or performs a division into several frequency sub-bands. In variants, the extraction of a given band or the division into sub-bands may reuse equivalent processing carried out in blocks 110 or 111.
[0075] Generally, the subbanding may be uniform or non-uniform. In a preferred embodiment, when the signal is not encoded into a frequency band, the channels of the original multi-channel audio signal are divided into frequencies using frequency intervals defined according to the Bark scale.
[0076] The Bark scale is defined over the following 24 intervals (in Hz) for a signal sampled at 32 kHz: [20, 100], [100, 200], [200, 300], [300, 400], [400, 510], [510, 630], [630, 770], [770, 920], [920, 1080], [1080, 1270], [1270, 1480], [1480, 1720], [1720, 2000], [200, 2320], [2320, 2700], [2700, 3150], [3150, 3700], [3700, 4400], [4400, 5300], [5300, 6400], [6400, 7700], [7700, 9500], [9500, 12000], [12000, 16000]
[0077] This predefined split can be modified for the case of a sampling frequency to use a different number of bands, for example by keeping only 21 bands at 16 kHz and changing the last interval to [6400, 8000] or by adding a band [16000, 20000] at 48 kHz.
[0078] This sub-band splitting, which is implemented in the domain of the short-term discrete Fourier transform (STFT) calculated on 20 ms frames with a windowing over 30 ms (10 ms of signal passed), amounts to band-pass filtering in the Fourier domain. In variants, a filter bank may be applied with or without critical sampling to obtain real or complex signals corresponding to the sub-bands. It should be noted that the sub-band splitting operation generally involves a processing delay which is a function of the type of filter bank implemented; according to the invention, a time alignment may be applied before or after coding-decoding and / or before the extraction of spatial image information, so that the spatial image information is well synchronized temporally with the corrected signal.
[0079] In the following description, the individual steps of coding and decoding are described as if they were processing in the complex frequency domain. The real case is also described as a variant.
[0080] Block 121 determines (Inf. B ) information representative of a spatial image of the original multichannel signal.
[0081] In the embodiment described herein, the information representative of the spatial image of the original multichannel signal is a covariance matrix of the input channels B in each frequency band predetermined by block 120. It will be noted here that for simplicity, the description does not distinguish here the sub-band index for the matrix C. In the preferred embodiment, the invention is implemented in a domain by complex-valued transform, the covariance is calculated as follows: C = Re B . B H within a normalization factor.
[0082] This matrix is calculated as follows in the real case: C = B . B T within a normalization factor.
[0083] In cases of a multichannel signal in the time domain, the covariance can be estimated recursively (sample by sample) in the form: Cij n = n / n + 1 Cij n − 1 + 1 / n + 1 bi n bj n .
[0084] In variants, temporal smoothing operations of the covariance matrix may be used.
[0085] In variants, the covariance matrix C can be regularized before quantification in the form C+ ε I or by applying a threshold on the diagonal coefficients of C to ensure a minimum value ε (for example ε=10 -9< if the input ambisonic signals are normalized in amplitude over the interval + / -1).
[0086] The covariance matrix C (of size KxK) is, by definition, symmetric, K being the number of ambisonic components.
[0087] Block 130 performs a quantization of the coefficients of the matrix.
[0088] There figure 3a illustrates the steps implemented by block 130 to quantize the coefficients of a covariance matrix according to an embodiment of the invention. Thus, according to the invention, the coding of the covariance matrix is carried out according to the following steps: It is assumed at this step that the covariance matrix has been estimated and that it has been modified (regularized) to ensure that no eigenvalue is zero. This can be achieved by replacing the values Cii of the diagonal of C by Cii=max(Cii, ε), where ε is a small value fixed for example at 10 -9< (if the values of the ambisonic signal in time are defined in the interval + / -1). In variants, the matrix can be modified C=C+ ε I Or I is the identity matrix.
[0089] The covariance matrix C(thus regularized) is decomposed into eigenvalues at step S1, in the form: C = Q Λ Q T< Or Q is an orthogonal matrix (with in particular det Q = + / -1) and A=diag (λ 1 ,..., λ K ) is a diagonal matrix of eigenvalues. Without loss of generality we assume that λ 1 ≥ ... ≥ λ K ≥0. We note that the regularization of C, if applied, guarantees that the eigenvalues are strictly positive.
[0090] Several methods are known in the state of the art to perform this factorization: QR decomposition (iterative), Householder transformation, Givens rotations or variants of these methods such as the “sorted QR” decomposition. It does not matter which method is chosen if the eigenvalues λ i (i=1..K) are not positive and are not ordered in descending order, according to the invention, if λ i < 0, the sign of λ i and of the associated eigenvector will be reversed; the eigenvalues will also be permuted if necessary to respect the constraint λ 1 ≥ ... ≥ λ K ≥0 by applying the same permutations to the eigenvectors (columns) in Q.
[0091] In step S2, we calculate the determinant of the eigenvector matrix Q and we determine if det Q = -1. If this is the case (O at step S2), we modify Qin step S3, preferably by inverting the sign of the eigenvector associated with the lowest eigenvalue so as to obtain a rotation matrix (orthogonal, unitary matrix with det Q = +1). The vector matrix Q is therefore called below the “rotation matrix” after step S2.
[0092] For the case of the planar FOA (three channels) step S2 is adapted to a calculation of a determinant of size 3x3 and for the ambisonics of the FOA (4 channels) we use a determinant of a 4x4 matrix. An example of implementation is given for the 4x4 case in APPENDIX 3.
[0093] In step S4, we convert the matrix Q resulting from step S3 or step S1, depending on the value of the determinant in step S2. This conversion is done either in the domain of Euler angles (K=3) or generalized Euler angles (K>4), or in the domain of quaternions in the case of FOA (K=3 or 4).
[0094] The conversion into Euler angles (for K=3) is for example given in Appendix I of the article by K. Shoemake, Animating Rotation with Quaternion Curves Proc. SIG-GRAPH 1985, pp. 245 - 254. It is recalled that there are variants of definition of the Euler angles according to the chosen axes of rotation (X, Y, Z) and according to whether the axes are fixed or not. In variants of the invention, other variants of definition of the Euler angles than that adopted in the article by K. Shoemake for the conversion may be used.
[0095] The conversion to generalized Euler angles (for K>3) is for example detailed in the article DK Hoffman, RC Raffenetti, and K. Ruedenberg, "Generalization of Euler Angles to N-Dimensional Orthogonal Matrices," Journal of Mathematical Physics, vol. 13, no. 4, pp. 528-533, 1972. This parameterization by generalized Euler angles is general and applies to any dimension.
[0096] Alternatively, in the case K=3, one can convert the rotation matrix Q (after steps S1 and S2) into a single unit quaternion. An example of this is given in Appendix I of Ken Shoemake's article, Animating Rotation with Quaternion Curves Proc. SIG-GRAPH 1985, pp. 245–254.
[0097] In the case K=4, a parameterization of Q by double unit quaternion is also possible; the conversion to double quaternion is given for example in the article P. Mahé, S. Ragot, S. Marchand, "First-Order Ambisonic Coding with PCA Matrixing and Quaternion-Based Interpolation," Proc. DAFx, Birmingham, UK, Sept. 2019.
[0098] In step S5, the parameters obtained in step S4 are quantized. For the Euler angles (K=3) or generalized Euler angles (K>3), noted in APPENDIX 1 angles [i] (i=1,...,6 for the example K=4), in the preferred embodiment, we apply for example a scalar quantization with a quantization step (noted in APPENDIX 1 “stepSize”) identical for each angle. We define for example a budget of 5 and 6 bits for intervals of length π and 2π which gives a budget of 33 bits for 6 generalized Euler angles. A pseudo-code performing this quantization operation is given in APPENDIX 1. In the case K=3, with 3 Euler angles we would have for example a budget of 17 bits (6+6+5 bits for 2 angles defined on an interval of length 2π and one angle on an interval of length π). In variants other methods of quantization of Euler angles can be used.
[0099] In the case K=3, if the rotation matrix Qis converted into a single unit quaternion, this quaternion is coded preferentially with a hemispherical dictionary of spherical vector quantization in dimension 4. In an exemplary implementation, we can take the vertices of a polytope of dimension 4, using preferentially the vertices of a truncated 600-cell (7200 vertices) or omnitruncated (14400 vertices) as defined in the literature or the 7200 vertices of a “120-cell snub” whose code words (coordinates in dimension 4) are available for example in: http: / / paulbourke.net / geometry / hyperspace / 120cell_snub.ascii.qz (start of file, lines 2-7201).
[0100] Spherical vector quantization is done by simple comparison by scalar product in dimension 4 with code words (typically normalized to a unit norm equal to 1). The exhaustive search of the nearest neighbor can be carried out efficiently by taking into account the possible permutations of the same code word in the dictionary. According to the invention, the dictionary can be truncated into one hemisphere to retain in the nearest neighbor search only the code words whose last (fourth) component is positive (or negative according to the alternative convention which can be used in variants). In variants, the truncation by the sign can be done on one of the three other components of the unit quaternion. In variants, the quantization dictionary can not be truncated to one hemisphere.We do not recall here the known principles of spherical vector quantization with the use of "leaders", which are for example defined in the article by C. Lamblin and J.-P. Adoul. Algorithm for spherical vector quantization from the Gosset network of order 8. Ann. Télécommun., vol. 43, no. 3-4, pp. 172-186, 1988 (Lamblin, 1988). Here, the scalar product is calculated for all the elements of the dictionary (with or without restriction to the hemisphere) and the number of calculations can be equivalently reduced to a subset by listing the signed or unsigned "leaders" in a pre-calculated table. The calculation of the quantization index is given either by the index in the exhaustive table, or by the addition of a permutation index and a cardinality offset, according to approaches known to those skilled in the art. An example of spherical quantization (which can be easily adapted) can be found in clause 6.6.9 of ITU-T recommendation G.729.1.
[0101] In the case K=4, in the case of double quaternions, the pair of unit quaternions q 1 and q 2 is quantified by a spherical quantification dictionary in dimension 4; by convention we quantify q 1 with a hemispherical dictionary (because q 1 and - q 1 correspond to the same 3D rotation) and q 2 is quantified with a spherical dictionary. Examples of dictionaries can be given by predefined points in 4-dimensional polyhedra. The quantification dictionaries for q 1 and q 2 can be swapped for quantization. Quantization is implemented as explained above by repeating the case K=3 for q 1 and q 2 with a hemispherical and a spherical dictionary.
[0102] In step S5, the eigenvalue matrix is also coded. According to the invention, the eigenvalues are ordered so that λ 1 ≥ ... ≥ λ K ≥ 0
[0103] In an exemplary embodiment, differential scalar quantization on a logarithmic scale is used.
[0104] An example of quantization consists of coding λ 1 in absolute on 5 bits, then coding the difference (in dB) between λ k and λ k-1 on 3 bits, i.e. a budget of 17 bits for K=4. An example of implementation is given in APPENDIX 4 using a base 2 logarithm - in variants a base 10 (or other) can be used. In variants, other realizations of logarithmic scalar quantization can be used. One can also use a vector quantization after having converted the eigenvalues in the logarithmic domain, for example by using a Pyramidal Vector Quantization (PVQ) type quantization described in the article T. Fischer, "A pyramid vector quantizer," IEEE transactions on information theory, vol. 32, no. 4, pp. 568-583, 1986, or in variants (as in the Opus codec defined in IETF RFC 6716).Vector quantization uses only one quadrant of the possible codewords because the eigenvalues are positive and ordered, so the indexing of the codewords can be simplified to take these two constraints into account. For the PVQ case, a preferred implementation example scales the eigenvalues before applying the search to a 4-dimensional pyramid face.
[0105] In variants, the eigenvalues can be normalized to code only K-1 normalized eigenvalues λ 2 / λ 1 , ..., λ K / λ 1 . We then use a scalar quantization on a 14-bit logarithmic scale for K=4. In this case, the same normalization constraint must be applied to the decoding on the covariance matrix calculated on the decoded signal. The exemplary implementation can be adapted to directly code differential indices.
[0106] In some variants, the eigenvalues resulting from the decomposition of the matrix C can be quantized predictively using inter-frame or intra-frame prediction. In other variants, if the coding uses a division into several sub-bands, a joint quantization of the eigenvalues of all the sub-bands can be used.
[0107] The quantization indices of the rotation matrix and the eigenvalue matrix are sent to the multiplexer (block 140).
[0108] The quantized values (index_angle[i], ...) are sent to the multiplexer 140.
[0109] In the implementation example for the 4-channel FOA case (with 6 generalized Euler angles coded on 33 bits and 4 eigenvalues coded on 17 bits), we therefore obtain a budget of 50 bits (or 2.5 kbit / s) to code a covariance matrix of size 4x4 in each sub-band. For example, if a division by sub-bands is defined with respectively 4, 6, 12 or 24 sub-bands and if a covariance matrix is transmitted for each of the sub-bands, we obtain a "meta-data" rate describing the spatial image of 10, 15, 30, or 60 kbit / s.
[0110] The decoder illustrated in the figure 1 , receives in the demultiplexer block 150, a binary stream comprising at least one coded channel of an audio signal from the original multichannel signal and the information representative of a spatial image in at least one frequency band (a sub-band or a single band which can cover up to the Nyquist band) of the original multichannel signal.
[0111] Block 160 decodes (Q -1< ) the covariance matrix in each band or sub-band defined by the encoder or other information representative of the spatial image of the original signal. To keep the notations simple, the decoded covariance matrix is also denoted C as to the coder.
[0112] Block 160 implements the steps illustrated in figure 3b to decode the covariance matrix. The steps depend on the parameterization used in the encoder.
[0113] If the matrix Q has been encoded in the domain of generalized Euler angles, block 160 can decode in S'1 the quantization indices of the generalized Euler angles. In the FOA case (with 4 channels) the following pseudo-area is given in APPENDIX 2. The same approach is easily adapted to the case of three Euler angles for K=3 or in the general case K>3. If the matrix Qhas been coded in the quaternion domain, we decode the quantification index(es) corresponding for example to a code word in a quantification dictionary in dimension 4 (possibly restricted to a hemisphere by restriction of the sign of one of the components of the unit quaternion in the dictionary).
[0114] At step S'2, block 160 reconstructs the decoded matrix Q by applying the conversion of generalized Euler angles or quaternion(s) to a rotation matrix, for example according to the aforementioned articles for the encoder part.
[0115] We also decode the eigenvalues in S'1, to obtain Λ=diag(λ 1 ,..., λ K ) then we calculate in step S'3 the covariance matrix: C = Q Λ Q T< .
[0116] Block 170 of the figure 1 , decodes (DEC) the audio signal as represented by the bitstream.
[0117] The decoding implemented in block 170 makes it possible to obtain a decoded audio signal B̂ ' which is sent as input to upmix block 171. Thus, block 171 implements a step (UPMIX) of increasing the number of channels. In one embodiment of this step, for the channel of a mono signal B̂ ', it consists of convolving the signal B̂ ' by different spatial impulse responses which implement all-pass decorrelating filters normalized in energy to the different signal channels B̂ ' .
[0118] In variants, the signal can be convolved B̂ ' by room impulse responses (SRIR for "Spatial Room Impulse Response"); these SRIRs are defined at the original ambisonic order of B. In other variants, decorrelation will be implemented in a transformed domain or in sub-bands (by applying a real or complex filter bank).
[0119] The upmix will add a number of channels K up to obtain K dmx + K up = K , with K the number of channels of the original signal. In a particular embodiment, with a FOA downmix signal, K dmx = 1 (the W channel) and K up = 3.
[0120] Block 172 implements a step (SB) of division into sub-bands in a transformed domain. In variants, a filter bank may be applied to obtain signals in the time or frequency domain. A reverse step, in block 191, recombines the sub-bands to reconstruct a decoded output signal.
[0121] In the preferred embodiment, the signal decorrelation (block 171) is implemented before the division into sub-bands (block 172) but it is entirely possible in variants to invert these two blocks. The only condition to be checked is to ensure that the decorrelation is adapted to the predefined band or sub-bands.
[0122] Block 175 determines (Inf B̂ ) information representative of a spatial image of the decoded multichannel signal in a similar manner to that described for block 121 (for the original multichannel signal), this time applied to the decoded multichannel signal B̂ obtained at the output of block 171.
[0123] In the same way as described for block 121, in one embodiment, this information is a covariance matrix of the channels of the decoded multichannel signal. In one embodiment, in the STFT domain, the complex case will be used where Ĉ=Re(B̂.B̂ H< ) to within a normalization factor.
[0124] This covariance matrix is obtained as follows in the real case: Ĉ=B̂.B̂ T< within a normalization factor.
[0125] We can eventually normalize the matrices Ĉ by the term Ĉ 11 associated with the W channel, if a similar normalization is applied on the matrix C.
[0126] Alternatively, temporal smoothing operations of the covariance matrix may be used. In cases of a multichannel signal in the time domain, the covariance can be estimated recursively (sample by sample).
[0127] In variants, the covariance matrix Ĉ of the decoded signal can be decomposed into eigenvalues (ordered as at the encoder) and the eigenvalues can be normalized by the largest eigenvalue.
[0128] From the information representative of the spatial images respectively of the original multichannel signal (Inf. B ) and the decoded multichannel signal (Inf. B̂ ) , for example, the covariance matrices Cand Ĉ, block 180 implements a step of determining (Det.Corr) a set of corrections by sub-bands (in at least one band).
[0129] For this, a transformation matrix T to be applied to the decoded signal is determined, so that the modified spatial image after application of the transformation matrix T to the decoded signal B̂ be the same as that of the original signal B.
[0130] There figure 2 illustrates this determination step implemented by block 180. In this embodiment, it is considered that the information representative of the spatial image of the original multichannel signal and of the decoded multichannel signal are the respective covariance matrices C etc.
[0131] So we are looking for a matrix T which verifies the following equation: T.Ĉ.T T< =C Or C=B.B T< is the covariance matrix of B And Ĉ = B̂.B̂ T< is the covariance matrix of B̂ , in the current frame.
[0132] In this embodiment, a factorization called Cholesky factorization is used to solve this equation.
[0133] Given a matrix A of size nxn, Cholesky factorization consists of determining a matrix L triangular (lower or upper) such that A=LL T< (real case) and A=LL H< (complex case). For the decomposition to be possible, the matrix A must be a symmetric positive definite matrix (real case) or Hermitian positive definite matrix (complex case); in the real case, the diagonal coefficients of L are strictly positive.
[0134] In the real case, a matrix M size nxn is said to be symmetrical positive definite if it is symmetrical (M T< =M) and positive definite (x T< Mx>0 for everything x∈R n< { 0} ).
[0135] For a symmetric matrix M, it is possible to verify that the matrix is positive definite if all its eigenvalues are strictly positive ( λ i >0 ). If the eigenvalues are positive ( λ i ≥0 ), the matrix is said to be positive semi-definite.
[0136] A matrix M size nxn is said to be positive definite symmetric Hermitian if it is Hermitian (M H< =M) and positive definite (z H< Mz is a real >0 for all z∈C n< { 0} ).
[0137] Cholesky factorization is for example used to find a solution to a system of linear equations of the type Ax=b. For example, in the complex case, it is possible to transform A in LL H< by Cholesky factorization, to solve Ly=b then to solve L H< x=y.
[0138] Equivalently, the Cholesky factorization can be written as A=U T< U (real case) and A=U H< U (complex case), where U is an upper triangular matrix.
[0139] In the embodiment described here, without loss of generality, only the case of a Cholesky factorization by triangular matrix is treated. L.
[0140] Thus, Cholesky factorization allows us to decompose a matrix C=L.L T< into two triangular matrices provided that the matrix C is symmetrical and positive definite. This gives the following equation: T . L ^ . L ^ T T T = L . L T .
[0141] By identification, we find: T . L ^ = L
[0142] Either : T = L . L ^ − 1
[0143] Covariance matrices C And Ĉ being in general positive semi-definite matrices, the Cholesky factorization cannot be used as is.
[0144] We note here when the matrices L And L̂ are lower (respectively upper) triangular, the transformation matrix T is also lower (respectively upper) triangular.
[0145] So, block 210 forces the covariance matrix C to be positive definite. This modification of the matrix C can be omitted for the decoded covariance matrix if the quantization guarantees that the eigenvalues are indeed non-zero. If it is used, the values of the diagonal Cii can be replaced by max(Cii, ε) where ε is a small value set for example to 10 -9< (if the values of the ambisonic signal in time are defined in the interval + / -1). In variants, ε is added (Fact. C for factorization of C ) on the coefficients of the diagonal of the matrix to ensure that the matrix is positive definite: C = C +ε I , And I is the identity matrix.
[0146] Similarly, block 220 forces the covariance matrix Ĉ to be defined positive, by replacing the values of the diagonal Cii by max(Cii, ε) where ε is a low value fixed for example at 10 -9< (if the values of the ambisonic signal in time are defined in the interval + / -1) or by modifying this matrix in the form Ĉ= Ĉ +ε I . In the preferred embodiment this conditioning of the covariance matrices is preferably integrated into the blocks 121 (at the encoder) for the matrix C and 175 (at the decoder) for the matrix Ĉ .
[0147] Once the two covariance matrices C And Ĉ are conditioned (regularized) to be positive definite, block 230 calculates the associated Cholesky factorizations and finds (Det.T) the optimal transformation matrix T in the form T = L . L ^ − 1 .
[0148] In this embodiment, it is possible that the relative energy difference between the decoded Ambisonic signal and the corrected Ambisonic signal is very significant, particularly at the high frequencies which can be significantly deteriorated by coders such as multi-mono EVS coding. To avoid amplifying certain frequency areas too much, a regularization term can be added. Block 240 optionally normalizes (Norm. T ) this correction. In the preferred embodiment, a normalization factor is therefore calculated so as not to amplify frequency zones.
[0149] From the covariance matrix Ĉ of the coded then decoded multichannel signal and the transformation matrix T we can calculate the covariance matrix of the corrected signal as: R = T . C ^ . T T
[0150] Only the value of the first coefficient R 00 of the matrixR, corresponding to the omnidirectional component (W channel), is kept to be applied as a normalization factor to T and avoid an increase in the overall gain due to the correction matrix T: B ^ corr = T norm . B ^ T norm = g norm . T with g norm = C ^ 00 / R 00 where Ĉ 00 corresponds to the first coefficient of the covariance matrix of the decoded multichannel signal.
[0151] In variants, the normalization factor g norm can be determined without calculating the entire matrix R, because it is sufficient to calculate only a subset of matrix elements to determine R 00 (and therefore g norm ).
[0152] The Matrix T Or T norm thus obtained in each band or sub-band corresponds to the corrections to be made to the multichannel signal decoded in block 190 of the figure 1 .
[0153] Block 190 performs the step of correcting the decoded multichannel signal by applying in each band or sub-band of the transformation matrix T Or T norm directly to the decoded multichannel signal, in the ambisonic domain (preferably in the transformed domain), to obtain the corrected output ambisonic signal ( B̂ corr).
[0154] Even if the invention applies to the Ambisonic case, in variants it will be possible to convert other formats (multichannel, object, etc.) into Ambisonic to apply the methods implemented according to the different embodiments described. An example of carrying out such a conversion from a multichannel or object format to an Ambisonic format is described in figure 2 of the 3GPP TS 26.259 specification (v15.0.0).
[0155] It has been illustrated on the figure 4 a DCOD coding device and a DDEC decoding device, within the meaning of the invention, these devices being dual to each other (in the sense of “reversible”) and connected to each other by a RES communication network.
[0156] The DCOD coding device comprises a processing circuit typically including: a memory MEM1 for storing instruction data of a computer program within the meaning of the invention (these instructions can be distributed between the coder DCOD and the decoder DDEC); an interface INT1 for receiving an original multichannel signal B,for example an ambisonic signal distributed over different channels (for example four channels W, Y, Z, X in order 1) for the purpose of its compression coding within the meaning of the invention; a processor PROC1 for receiving this signal and processing it by executing the computer program instructions stored in the memory MEM1, for the purpose of its coding; and a communication interface COM 1 for transmitting the coded signals via the network.
[0157] The DDEC decoding device has its own processing circuitry, typically including: a memory MEM2 for storing instruction data of a computer program within the meaning of the invention (these instructions can be distributed between the coder DCOD and the decoder DDEC as indicated previously); a COM2 interface for receiving the coded signals from the network RES with a view to their compression decoding within the meaning of the invention; a processor PROC2 for processing these signals by executing the computer program instructions stored in the memory MEM2, with a view to their decoding; and an output interface INT2 for delivering the corrected decoded signals ( B̂ Corr) for example in the form of Ambisonic W...X channels, for their restitution.
[0158] Of course, this figure 4 illustrates an example of a structural embodiment of a codec (coder or decoder) within the meaning of the invention. The figures 1 à 3 commented above describe in detail rather functional implementations of these codecs. ANNEXE 1
[0159] min_angle 6 = − PI_ 2 , − PI_ 2 , − PI , − PI_ 2 , − PI , − PI max_angle 6 = PI_ 2 , PI_ 2 , PI , PI_ 2 , PI , PI excess_bit 6 = 0 , 0 , 1 , 0 , 1 , 1 bits = 5 + v_excess_bit i stepSize = max_angle i − min_angle i / 1 ≪ bits index_angle i = int angles i − min_angle i / stepSize + 0.5 index_angle i = index_angle i % 1 ≪ bits ANNEXE 2
[0160] min_angle 6 = − PI_ 2 , − PI_ 2 , − PI , − PI_ 2 , − PI , − PI max_angle 6 = PI_ 2 , PI_ 2 , PI , PI_ 2 , PI , PI excess_bit 6 = 0 , 0 , 1 , 0 , 1 , 1 bits = 5 + v _ excess _ bit i stepSize = max _ angle i − min _ angle i / 1 ≪ bits angles _ q i = index * stepSize + min _ angle i ANNEXE 3
[0161] Calculation of the determinant d = det M in literal form for a matrix M = [aij] of size 4x4: d = a 11 * a 22 * a 33 * a 44 + a 11 * a 24 * a 32 * a 43 + a 11 * a 23 * a 34 * a 42 − a 11 * a 24 * a 33 * a 42 − a 11 * a 22 * a 34 * a 43 − a 11 * a 23 * a 32 * a 44 − a 12 * a 21 * a 33 * a 44 − a 12 * a 23 * a 34 * a 41 − a 12 * a 24 * a 31 * a 43 + a 12 * a 24 * a 33 * a 41 + a 12 * a 21 * a 34 * a 43 + a 12 * a 23 * a 31 * a 44 + a 13 * a 21 * a 32 * a 44 + a 13 * a 22 * a 34 * a 41 + a 13 * a 24 * a 31 * a 42 − a 13 * a 24 * a 32 * a 41 − a 13 * a 21 * a 34 * a 42 − a 13 * a 22 * a 31 * a 44 − a 14 * a 21 * a 32 * a 43 − a 14 * a 22 * a 33 * a 41 − a 14 * a 23 * a 31 * a 42 + a 14 * a 23 * a 32 * a 41 + a 14 * a 21 * a 33 * a 42 + a 14 * a 22 * a 31 * a 43 ANNEXE 4
[0162] Assuming matrix conditioning C by ε=10 -9< (the interval of indices is adapted according to this value): Quantification: index _ val i = round 1 / 2 log 2 λi , i = 1 , … , K − 1 index _ val i = clip index _ val i , − 15 , 37 # saturation dans l ′ intervalle − 15 , 37 diff _ index _ val i = index _ val i − index i − 1 , i = 2 .. K − 1 diff _ index _ val i = clip diff _ index _ val i , 0 7 # saturation dans l ′ intervalle 0 7 Decoding: index _ val i = index _ val i − 1 + diff _ index i , i = 2 … K − 1 λi = 2 1 2 index _ val i
Claims
1. Method for coding a multichannel sound signal, comprising the following steps: - coding (111) at least one audio signal channel of the original multichannel signal or a downmix of the original multichannel signal; - dividing (120) the original multichannel signal into frequency sub-bands; - determining (121) a covariance matrix per frequency sub-band, representative of a spatial image of the original multichannel signal; - decomposing (130) the determined covariance matrices into eigenvalues; - coding (130) by quantizing the parameters resulting from the decomposition into eigenvalues comprising both eigenvalues and eigenvectors.
2. Method according to Claim 1, wherein the eigenvalues are ordered before quantization and the quantization is performed by a differential scalar quantization.
3. Method according to either of Claims 1 and 2, wherein a covariance matrix is decomposed into eigenvalues using the following steps: - obtaining a matrix of eigenvectors Q such that C = Q A QT, where C is the covariance matrix and Λ=diag(λ1,..., λK) is a diagonal matrix of eigenvalues; - modifying the matrix of eigenvectors as a function of a determinant value of the matrix of eigenvectors Q; - converting the matrix of eigenvectors Q into the domain of generalized Euler angles; the generalized Euler angles that are obtained forming part of the parameters to be quantized.
4. Method according to Claim 3, wherein the generalized Euler angles are quantized by uniform quantization.
5. Method according to either of Claims 1 and 2, wherein a covariance matrix is decomposed into eigenvalues using the following steps: - obtaining a matrix of eigenvectors Q such that C = Q A QT, where C is the covariance matrix and Λ=diag(λ1,..., λK) is a diagonal matrix of eigenvalues; - modifying the matrix of eigenvectors as a function of a determinant value of the matrix of eigenvectors Q; - converting the matrix of eigenvectors Q into the domain of quaternions; at least one quaternion that is obtained forming part of the parameters to be quantized.
6. Method according to Claim 5, wherein the quaternions are quantized by spherical vector quantization.
7. Method for decoding a multichannel sound signal, comprising the following steps: - decoding (170) at least one coded channel or a downmix of the original multichannel signal and obtaining a decoded multichannel signal; - dividing (172) the decoded multichannel signal into frequency sub-bands; - decoding (160) parameters resulting from a decomposition of covariance matrices of the original multichannel signal into eigenvalues; - determining (160) the covariance matrices of the original multichannel signal from the decoded parameters; - determining (175) a covariance matrix, per frequency sub-band, of the decoded multichannel signal; - determining (180) a set of corrections to be made to the decoded signal based on the covariance matrices of the original multichannel signal (Inf. B) and the covariance matrices of the decoded multichannel signal (Inf. B̂ ); - correcting (190) the decoded multichannel signal using the determined set of corrections.
8. Coding device comprising a processing circuit for implementing the steps of the coding method according to one of Claims 1 to 6.
9. Decoding device comprising a processing circuit for implementing the steps of the decoding method according to Claim 7.
10. Storage medium, able to be read by a processor, storing a computer program comprising instructions for executing the coding method according to one of Claims 1 to 6 or the decoding method according to Claim 7.