Spatial audio coding having a configuration of decorrelation processing operation

The adaptive encoding method for ambisonic signals dynamically adjusts decorrelation processes based on signal characteristics, addressing spatial distortion issues and enhancing encoding quality for immersive audio experiences.

JP2025519179APending Publication Date: 2025-06-24オランジュ
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024570378
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-30
Filing Date
2023-05-30
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing methods for encoding ambisonic signals, such as multi-monaural encoding, fail to effectively capture the spatial correlations between channels, leading to spatial distortion and artifacts like false sound sources.

Method used

An adaptive method for encoding ambisonic signals that determines the active or inactive mode of decorrelation based on gain criteria, inter-frame distance between rotation matrices, and distance between the current rotation matrix and the identity matrix, allowing for dynamic adjustment of decorrelation processes.

Benefits of technology

This approach improves the encoding quality by adapting decorrelation processes to the characteristics of the input signal, reducing spatial artifacts and enhancing the immersive experience in spatial audio applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025519179000001_ABST
    Figure 2025519179000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method for decoding an audio signal that forms a series of frames (t-1, t) of samples within each of n channels of an ambisonic representation of degree greater than 0 over time. The method comprises: determining a binary value indicating an active or inactive mode of an uncorrelating processing operation applied to the signal of the current frame with respect to the current frame to be encoded, and encoding this value into a bitstream; encoding into the bitstream uncorrelating processing information if the mode is determined to be active; generating an output signal encoded into the bitstream depending on the mode determined for the current frame and the mode determined for the previous frame. The present invention also relates not only to a corresponding decoding method, but also to an encoding device and a decoding device that implement each encoding method and decoding method, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the encoding / decoding of spatial audio data, particularly related to ambiophonic (also hereinafter referred to as "ambisonic").

Background Art

[0002] The encoders / decoders currently used in mobile phones (hereinafter referred to as "codecs") are monaural (a single signal channel for rendering on a single loudspeaker). The codec 3GPP (registered trademark) EVS ("Enhanced Voice Services") enables the presentation of "Super HD" (also called "High Definition Plus" or HD+ voice) quality with the SWB ("super-wideband") audio band of signals sampled at 32 or 48 kHz or the FB ("Fullband") of signals sampled at 48 kHz; the audio bandwidth is 14.4 to 16 kHz in the SWB mode (9.6 to 128 kbit / s) and 20 kHz and above in the FB mode (16.4 to 128 kbit / s).

[0003] The next development in the quality of the conversation services presented by operators should consist of immersive services using terminals such as smartphones equipped with several microphones, or telepresence or 360° video type spatial audio conferencing or videoconferencing equipment, or otherwise "live" audio content sharing equipment with 3D spatial sound that is much more immersive than simple 2D stereo rendering. The increasingly widespread use of audio headsets on mobile phones and the emergence of state-of-the-art audio equipment (accessories such as 3D microphones, voice assistants with acoustic antennas, virtual reality headsets, etc.) have made the capture and rendering of spatial acoustic scenes now common enough to present an immersive communication experience.

[0004] In this regard, the future standard 3GPP ( "IVAS" ( "Immersive Voice and Audio Services")) includes an extension of the codec EVS to immersive by accepting at least the spatial audio formats (and combinations thereof) listed below as input formats for the codec: - Stereo or 5.1 type channel-based formats: Here, each channel feeds a loudspeaker (e.g., L and R in stereo) or L, R, Ls, Rs, and C in 5.1; - Object-based formats: Here, a sound object is described as an audio signal (generally monaural) associated with metadata that describes the attributes of this object (position in space, spatial width of the source, etc.); - Scene-based formats that describe the sound field at a given point (generally captured by a spherical microphone or synthesized in the spherical harmonic domain).

[0005] The emphasis below is usually on the encoding of audio in a scene-based (or ambisonic) format according to an exemplary embodiment (where at least some of the aspects presented below regarding the present invention may also be applicable to formats other than scene-based formats).

[0006] Ambisonics is a method of recording (encoding in the acoustic sense) and reproducing (decoding in the acoustic sense) spatial sound. An Ambisonic microphone (order 1) includes at least four capsules (usually of the cardioid or sub-cardioid type) arranged on a spherical lattice (e.g., the vertices of a regular tetrahedron). The audio channels associated with these capsules are called the "A format". This format is converted to the "B format", where the sound field is decomposed into four components (spherical harmonics) represented by W, X, Y, and Z corresponding to four simultaneous virtual microphones. Component W corresponds to the omnidirectional capture of the sound field while the more directional components X, Y, and Z can be considered as pressure gradient microphones oriented along three mutually perpendicular spatial axes. The Ambisonic system is a flexible system in the sense that recording and rendering are separated and decoupled. The Ambisonic system enables decoding (in the acoustic sense) for any given configuration of loudspeakers (e.g., 5.1 type binaural "surround" sound, or 7.1.4 type periphonic (with elevation)). The Ambisonic technique can be generalized to more than 5 channels in the B format, and this generalized representation is commonly called "HOA" ("Higher-Order Ambisonics"). By decomposing the sound over a greater number of spherical harmonics in total, the spatial accuracy during rendering onto loudspeakers is improved.

[0007] An Ambisonic signal of order M has K = (M + 1) 2It contains components, and at order 1 (M = 1), four components W, X, Y, and Z, which are commonly referred to as First-Order Ambisonic (FOA), are recovered. There also exists a deformed form called the "plane" of Ambisonic (W, X, Y) that decomposes sounds defined within a plane, which is generally a horizontal plane. In this case, the number of components is K = 2M + 1 channels. For order 1 (4 channels: W, X, Y, Z) of Ambisonic, order 1 (3 channels: W, X, Y) of planar Ambisonic, and all higher orders of Ambisonic, for ease of reading, they will hereafter be represented as "Ambisonic". The presented processing operations are applicable regardless of the type of Ambisonic component, plane, or otherwise the number thereof. Hereinafter, an "Ambisonic signal" will refer to a signal in B-format of a given order having a fixed number of Ambisonic components. This also includes hybrid cases where, for example, at order 2, only 8 (instead of 9) channels exist. - More precisely, at order 2, there are 4 channels (W, X, Y, Z) of order 1, to which 5 channels (usually denoted as R, S, T, U, V) are typically added, and one of the higher order channels (e.g., R) can be ignored, for example. This also includes cases where the Ambisonic signal has been pre-processed to be converted to pre-processed channels prior to encoding.

[0008] The signal processed by the encoder / decoder takes the form of a series of blocks of audio samples, hereafter referred to as "frames" or "sub-frames".

[0009] Furthermore, hereafter, mathematical notations follow the following conventions: - Scalars: s or N (lowercase for variables or uppercase for constants) - The operator Re(.) represents the real part of a complex number - Vectors: u (lowercase, bold) - Matrices: A (uppercase, bold) The notations A T and A H respectively indicate the transpose of A and the Hermitian transpose (transpose and conjugate) of A. - A one-dimensional signal with discrete times \(s(i)\) defined over a time interval of length \(L\), \(i = 0,\ldots,L - 1\), is represented by a row vector. \(s=[s(0),\ldots,s(L - 1)]\)

[0010] This can also be written as follows to avoid using parentheses: \(s=[s_0,\ldots,s_{L - 1}]\). L-1 \) - A multi-dimensional signal with discrete times \(b(i)\) defined over a time interval of length \(L\), \(i = 0,\ldots,L - 1\), and having \(K\) dimensions is represented by a matrix of size \(L\times K\):

Number

[0011] This can also be written as follows to avoid using parentheses: \(B = [B_{i,j}], i = 0,\ldots,K - 1, j = 0,\ldots,L - 1\). ij \)

[0012] Furthermore, the known conventions in ambisonics related to the order of ambisonic components (including ACN: “Ambisonic Channel Number”, SID: “Single Index Designation”, FuMA: “Furse - Malham”) and the normalization of ambisonic components (SN3D, N3D, maxN) are not recalled here. Further details can be found, for example, online at: https: / / en.wikipedia.org / wiki / Ambisonic data exchange formats in the resources available there.

[0013] By convention, the first component of an ambisonic signal generally corresponds to the omnidirectional component \(W\).

[0014] The essence of the simplest approach for encoding an ambisonic signal lies in using a monaural encoder and applying the monaural encoder separately to each of the individual channels, which may, perhaps, involve different bit allocations according to the channels. This approach is herein called "multi-monaural: multi-mono". The multi-monaural approach can be extended to multi-stereo encoding (where pairs of channels are encoded separately by a stereo codec), or more generally, to the use of several parallel instances of the same core codec. The input signal is split into channels (one monaural channel or several channels). These channels are encoded separately depending on a given variance and binary allocation. At decoding, the decoded channels are recombined according to the convention of the input signal.

[0015] The quality of multi-monaural or multi-stereo encoding varies depending on the core encoding and decoding used and is generally only satisfactory at very high rates. For example, in the case of multi-monaural, EVS encoding can be judged to be quasi-transparent (from a perceptual point of view) at a rate of at least 48 kbit / s per channel (mono) (thus, for a first-order ambisonic signal, a minimum rate of 4×48 = 192 kbit / s). Since the multi-monaural encoding approach does not take into account the correlation between channels, it causes spatial distortion due to the addition of various artifacts (such as the appearance of false sound sources) of the diffuse noise of the sound source or the displacement of the sound source path. Therefore, the encoding of an ambisonic signal by this approach leads to a deterioration of the spatialization.

[0016] Alternative approaches to the separate encoding of channels are given, for example, by parameter encoding (such as DiRAC encoding) as described in the article (Non-Patent Document 1). In this document, the directionality analysis of an ambisonic signal is performed by frames and subbands for determining the direction of arrival (DoA). The DoA is completed by "diffusion" parameters that give a parametric description of the audio scene. The multi-channel input signal is encoded in the form of a downmix channel (usually a monaural or stereo signal obtained by reducing a plurality of captured channels) and spatial metadata (DoA and "diffusivity" by subbands).

[0017] The present invention also relates to another specific ambisonic encoding method described in the following publications: - (Non-Patent Document 2) - (Non-Patent Document 3) 。

[0018] Hereinafter, this method, which is encoding by principal component analysis or simply called PCA (principal component analysis) encoding, uses quantization and interpolation of the rotation matrix related to the eigenvectors of the PCA analysis (such as those also described in (Patent Document 1)). The strategy of this type of ambisonic encoding is to decorrelate the channels of the ambisonic signal and then separately encode the transformed channels by a core (e.g., multi-monaural) codec. This strategy makes it possible to limit the spatial artifacts in the decoded ambisonic signal.

[0019] In this method, for a first-order ambisonic signal, a 4×4 rotation matrix in 3D (resulting from PCA / KLT analysis as described in the example of the aforementioned patent application) is converted into parameters to be encoded (e.g., 6 generalized Euler angles or 2 unitary quaternions).

[0020] Without loss of generality, the quaternion domain is here held more specifically, enabling the transformation matrix computed for PCA / KLT analysis to be efficiently interpolated; i.e., since the transformation matrix is a rotation matrix at decoding time, the inverse matrix operation is performed simply by transposing the matrix applied at encoding time.

[0021] Figure 1 shows this method of encoding in the case where the quaternion representation is used for both encoding and interpolation of the rotation matrix. Encoding occurs in several steps.

[0022] The original multi-channel signal A of dimension K×L (i.e., K components of L time samples or frequency samples) is at the input. In block 100, PCA analysis is performed and split into several steps: - The signals of the channels (e.g., W, Y, Z, X for the FOA case) are assumed to take the form of a matrix A with an n×L matrix (n (here 4) ambisonic channels per frame and L samples). These channels can optionally be preprocessed, e.g., by a high-pass filter.

[0023] For example, the covariance matrix A of the multi-channel signal is obtained as follows: C = A.A T : within the normalization factor (in the real number case), or C = Re(A.A H ): within the normalization factor (in the complex number case).

[0024] Operations for time smoothing of the covariance matrix can be used. In the case of a multi-channel signal in the time domain, the covariance can be estimated recursively (per sample). The frame can also be split into sub-frames, and one covariance matrix is determined per sub-frame and smoothed later.

[0025] The diagonal elements of C represent in particular the energy of the i-th input channel of the PCA process

Number

[0026] In block 110, the new matrix V (which is a rotation matrix) of eigenvalues of the current frame t is transformed into the appropriate domain of quantization parameters. The corresponding matrix of eigenvalues here is Λ = diag(λ1,…,λ n ) where, in this case, this can be thought of as the transformation of a 4×4 matrix into two unitary quaternions; in the planar ambisonic case there would be a single unitary quaternion for a 3×3 matrix.

[0027] For a dimension of 4 (n = 4), the rotation matrix V can be parameterized by the product of two unitary quaternions q1 and q2 in matrix form:

Number

Number

[0028] Conversely, given a 4×4 rotation matrix, it is possible to find the associated dual quaternion (q1,q2) and the corresponding matrix. In other words, this matrix can be of the form by a method known as, for example, "Cayley factorization"

Number

[0029] These parameters q1, q2 are encoded according to prior art encoding methods over the total number of bits allocated to quantization of the parameters (block 120). For example, 19 bits may be used for q1 and 18 bits for q2, which gives a budget of N Q = 37 bits per frame.

[0030] The current frame is divided into sub - frames, where the number is assumed to be fixed. The representation by the encoded quaternion is interpolated by sequential sub - frames of index t' from the end of the previous frame t - 1 to the end of the current frame t in order to smooth the difference between inter - frame matrixing over time (block 130). The interpolated quaternion within each sub - frame is converted to a rotation matrix [Number] (block 140), and then the resultant rotation matrix, decoded and interpolated within each sub - frame (block 150), is applied.

[0031] At the output of block 150, a matrix representing each sub - frame of the signals of the ambisonic channels is obtained to decorrelate these signals and obtain the transformed signal B. A binary allocation to separate channels is also done based on the total number of bits (block 160), and N Q bits used within block 120 are subtracted from the total number of bits.

[0032] Figure 2 shows the corresponding decoding. The quantization index of the quantization parameter of the rotation matrix within the current frame is demultiplexed (block 200) according to the decoding method corresponding to the encoding (block 120) and decoded at block 230. The transformed channels are also decoded (block 220) based on the same binary allocation (block 210) as the encoder (block 160).

[0033] The conversion process and the interpolation process (blocks 240 and 250) of the decoder are the same as those performed in the encoder (blocks 130 and 140).

[0034] Block 260 applies the inverse matrix transformation resulting from block 250 to the decoded signal of the ambisonic channel by means of the subframe (recall that the inverse matrix of the rotation matrix is its transpose). Note that the algorithm delay associated with the encoding and decoding (blocks 170 and 220) must be corrected by storing the inverse matrix values in memory in an appropriate manner.

[0035] Ambisonic encoding as implemented in Figures 1 and 2 assumes that the input channels are (sufficiently) correlated. In particular, it is assumed that the decorrelation by block 150 provides an encoding gain; furthermore, it is assumed that the matrixing is stable from one frame to another so as not to generate audio artifacts within the transformed signal B. Note also that the encoding of the metadata (block 120) typically uses a rate of about 2 kbit / s (e.g., 1.85 kbit / s for N Q = 37 bits per 20 ms frame).

[0036] However, for some signals such as the recording of applause in a place where the sound field is relatively spread out, the decorrelation gain may be low. For spatially unstable signals (e.g., impact sounds whose localization rapidly alternates in each frame within the acoustic space), the PCA analysis (block 100) is

Number

Prior Art Documents

Patent Documents

[0037]

Patent Document 1

Non-Patent Documents

[0038]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0039] The present invention improves this situation.

[0040] For this purpose, the present invention provides a method for encoding an audio signal that forms a series of frames (t-1,t) of samples in each of n channels in an ambisonic representation of degree greater than 0 over time, the method comprising: - Determining a binary value indicating an active mode (ON) or an inactive mode (OFF) of a decorrelation process applied to the signal of the current frame with respect to the current frame to be encoded, and encoding this value into a bitstream; - Encoding decorrelation process information into the bitstream if the mode is determined to be active; - Generating an output signal to be encoded into the bitstream according to the mode determined for the current frame and the mode determined for the previous frame.

[0041] Thus, the present invention enables the use of decorrelation between n channels to be adapted according to the characteristics of the input signal.

Means for Solving the Problem

[0042] In one embodiment, the determination of the binary value indicating the active mode or the inactive mode is performed according to at least one gain criterion of the encoded signal before and after the decorrelation process.

[0043] Thus, this criterion makes it possible to ensure that "the decorrelation process provides a sufficient gain to be activated".

[0044] According to one specific embodiment, the encoding gain is defined by the following logarithmic value:

Equation

Equation

[0045] In one embodiment, the binary determination of the active or inactive mode is made according to a criterion of the frame - to - frame distance between the rotation matrices to which the decorrelation process is applied.

[0046] Thus, depending on the value of this distance, the generation of the signal to be encoded is adapted to avoid excessive variations in the transformation matrix to which the decorrelation process is applied.

[0047] According to a particular embodiment in which the rotation matrix is represented as a dual quaternion, the frame - to - frame distance between the rotation matrices is expressed by using the scalar product between the quaternion in the current frame and the quaternion of the previous frame.

[0048] In one embodiment, the binary determination of the active or inactive mode is made according to a criterion of the distance between the rotation matrix of the current frame and the identity matrix (by applying the decorrelation process).

[0049] Thus, here again, depending on the value of this distance, the generation of the signal to be encoded is adapted to avoid excessive variations in the transformation matrix to which the decorrelation process is applied to the direct encoding of the input.

[0050] In a particular embodiment in which the rotation matrix is represented as a dual quaternion, the distance between the rotation matrix of the current frame and the identity matrix is expressed in the form of the scalar product between the quaternion in the current frame and the unit quaternion.

[0051] The present invention is applicable to a method for decoding an audio signal that forms a series of frames (t - 1,t) of samples within each channel of an n - channel as an ambisonic representation of degree greater than 0 over time, the method comprising: - For the current frame (t), in addition to the signals of the n channels of this current frame, receiving a binary value indicating the active mode or the inactive mode of the decorrelation process applied to the signals of the current frame; - When it is determined that the mode is active, decoding the decorrelation process information received in the bit stream; - Generating an output signal according to the mode determined for the current frame and the mode determined for the previous frame.

[0052] The decoding method has the same advantages as the corresponding encoding method.

[0053] The present invention also aims at an encoding device including a processing circuit for implementing the previously presented encoding method.

[0054] The present invention also aims at a decoding device including a processing circuit for implementing the aforementioned decoding method.

[0055] The present invention also aims at a computer program including instructions for implementing the aforementioned method, and these instructions are executed by a processor of a processing circuit.

[0056] The present invention also aims at a non-volatile memory medium storing instructions of such a computer program.

[0057] Other advantages and features of the present invention will become apparent by reading the exemplary embodiments presented in the following detailed description and examining the accompanying drawings.

Brief Description of the Drawings

[0058]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0059] Without loss of generality, the input signal is assumed to be an ambisonic signal of first-order FOA in ACN format and by SN3D normalization. In a variant form, the input signal may be subjected to a preprocessing operation so as to obtain 4 channels derived from the original ambisonic signal (FOA).

[0060] FIG. 3 shows an encoding method according to the present invention, where decorrelation by PCA is applied in an adaptive manner (active or inactive PCA) within each frame assumed herein to be of length 20 ms (e.g., L = 960 samples at 48 kHz). It is assumed that R bits are allocated to the current frame for encoding R = 5120 bits, for example, at 256 kbit / s. In a variant form, this budget R may be reduced according to the bits already used for the signal by a potential preprocessing operation prior to encoding by PCA.

[0061] An adaptive determination is given by block 300 to determine the value of the activation indicator applied to the current frame or otherwise the value of the decorrelation process (PCA) indicator, which corresponds to mode = ON (active or activated PCA) or OFF (inactive or deactivated PCA). This determination criterion will be explained below.

[0062] Therefore, a binary value indicating the active mode (ON) or inactive mode (OFF) of the PCA type whitening process is determined by block 300. The operation of the encoder depends on the mode in the current frame of the index t(mode) and the mode in the previous frame of the index t-1(prev_mode). Without loss of generality, it is assumed that the initial state of the previous frame at the start of the encoder is prev_mode = OFF.

[0063] The following four possible combinations can be identified: ● If mode = ON and prev_mode = ON, the same operation as that in FIG. 1 is obtained. In other words, blocks 100, 110, 120, 320, 140, 170 in FIG. 3 apply the same processing as blocks 100, 110, 120, 130, 140, 170 described in FIG. 1, respectively. Branch 2 where the input signal undergoes conversion is selected by block 330. In the allocation block 340, since PCA-based encoding uses 1 bit to indicate the mode (ON) of the current frame, N Q bits for encoding the metadata and R - 1 - N Q bits for encoding the channel (block 170) are considered. ● If mode = OFF and prev_mode = OFF, the whitening by PCA is "short-circuited" and branch 1 is selected by the selection block 330. The encoded signal B is the same as the input signal A. The allocation block 340 takes into account the fact that "1 bit is used to indicate the mode (OFF) of the current frame and R - 1 bits are used to encode the channel (block 170)". In the default setting, the eigenvalue

number

number

Number

Number

Number

[0064] Following the processing of the previous frame, the eigenvalue

Number

Number

[0065] Therefore, the multiplexing module 350 inserts the encoded data into the bitstream depending on the activation indicators determined for the current and previous frames according to the allocation defined in block 340.

[0066] Next, an embodiment of the interpolation (block 320) will be described. This block depends on the encoded quaternion in the current frame t

Number

Number

Number

Number

Number

[0067] In a deformed form, other interpolations of quaternions (e.g., spherical linear interpolation or SLERP of other interpolations) are possible. Here, the NLERP method is used because it is less complex and more stable due to limited digital precision. -

Number

Number

[0068] Therefore, this interpolation leads to a constant value of piece - wise matrixization by normal division into N sub - frames.

[0069] In the deformation mode, other interpolation methods become possible. For example, interpolation is applied to each sample at the start of a frame (across the first N samples in total), and then a constant value

Number

[0070] In the deformation mode, other embodiments of interpolation are possible with different divisions into sub - frames.

[0071] An exemplary embodiment of the decision block 300 will now be described.

[0072] In one exemplary embodiment, the determination of the indicator for activation of the decorrelation process uses some of the following criteria together with a threshold - based determination: 1. Encoding gain of PCA 2. Inter - frame distance between rotation matrices (which can be regarded as a measure of spatial stability) 3. Distance between the rotation matrix in the current frame and the identity matrix

[0073] The motivation for the various criteria is as follows respectively: 1. To ensure that decorrelation by PCA provides sufficient gain to justify its activation with respect to direct encoding. 2. To ensure that the interpolation between the rotation matrix in the current frame and the rotation matrix in the previous frame does not cause very large variations in the matrix over time. 3. To ensure that the interpolation between the rotation matrix in the current frame and the potential invalidation of PCA does not lead to very large variations in the matrix over time.

[0074] In one possible embodiment, another criterion can be based on the correlation matrix of the input signal and its values outside the diagonal values. The correlation matrix is equal to the covariance matrix except that "the ambisonic components are each normalized by their standard deviation before calculating the correlation relationship". Therefore, the criterion can be defined independently of the signal of the input signal, and a predetermined threshold value is applied, for example, to the maximum or average non-diagonal value (absolute value) within the correlation matrix.

[0075] This makes it possible to verify that there is a minimum correlation between the input signals and that the decorrelation process is useful.

[0076] Exemplary embodiments of these criteria are given below.

[0077] The definition of the coding gain of the PCA / KLT transform in the case of a Gaussian source with n channels is recalled here:

Equation

[0078] This corresponds to the ratio of the arithmetic mean of the components to be encoded to the geometric mean of the variance (energy) (in the Gaussian case).

[0079] According to the present invention, rather, it is the energy of the input channel A of the molecule

Equation

Equation

[0080] It is assumed that the eigenvalues are in descending order and positive. The term ε is fixed, for example, at ε = 10 -8 for adjusting the calculation of the logarithm.

[0081] The gain G actually corresponds to the sum (in the logarithmic domain) of the coding (or decorrelation) gains between the individual (separate) channels taken before and after PCA.

[0082] In the deformed form, the eigenvalues may also be normalized; in this case, the normalization coefficient is also applied to the value [Number] is applied.

[0083] An example of a criterion for determining the distance between matrices is defined below. The preferred embodiment depends on the representation as a dual quaternion, and the angular distance is determined as follows by the scalar product between the quaternion in the current frame t and the quaternion in the previous frame t - 1: P1 = q1(t - 1).q1(t) P2 = q2(t - 1).q2(t)

[0084] The distance between the rotation matrix associated with frame t and the rotation matrix associated with frame t - 1 is evaluated as follows: minP = min(P1, P2)

[0085] In a similar way, this distance can be defined between the rotation matrix in frame t and the identity matrix (represented by the unitary quaternion q1 = q2 = (1, 0, 0, 0)) as follows: (1, 0, 0, 0).q1(t) = a1 (1, 0, 0, 0).q2(t) = a2 and minP2 = min(a1, a2)

[0086] Therefore, this determination (block 300) can be implemented as follows in one example: If the activation indicator is positive in the default setting (in other words, such that mode = ON (active mode)), the indicator switches to the negative mode so that mode = OFF (inactive mode) if the coding gain G is lower than a predetermined threshold (for example, 6), or the inter-frame distance between the rotation matrices is smaller than a threshold (for example, 0.8), or otherwise the distance between the rotation matrix of the current frame and the identity matrix is smaller than a threshold (for example, 0).

[0087] At this time, this means the following: If G < 6 mode = OFF, If minP < 0.8 mode = OFF, If minP2 < 0 mode = OFF.

[0088] In a variant form, the values of the thresholds can be different (6, 0.8, 0 respectively).

[0089] In a variant form, the activation indicator is negative in the default setting (in other words, such that mode = OFF), and the indicator switches to the positive mode so that mode = ON if it verifies that all the criteria are "the opposite of those defined above". Therefore, this simply reverses the decision logic of the same result.

[0090] In other variant forms, at least one of the three criteria defined according to the present invention may be used, and the others may not be used or may be replaced by other criteria.

[0091] Variant forms of the criteria are described next.

[0092] In a variant form, other definitions of the criteria may be adopted. For example, the coding gain may be as follows:

Number

Number

Number

[0093] In the deformed form, other measurements of the distance between rotation matrices (e.g., Frobenius distance, or another distance between the rotation matrix

Number

Number

[0094] In the case of the criterion based on the correlation matrix of the input signal, the correlation matrix corresponds to the covariance matrix applied to the normalized components of signal A. Such a covariance matrix (such as those described with reference to FIG. 1) is obtained, for example, as follows: C = A.A T : within the normalization factor (in the case of real numbers), or C = Re(A.A H ): within the normalization factor (in the case of complex numbers).

[0095] In particular, the diagonal element of C representing the energy of the i-th input channel

Number

[0096] The elements of the covariance matrix considered in this specification are the matrix terms C ij and C ji(i ≠ j). The maximum value of these terms is determined and compared with a threshold value (e.g., a value of 0.1). If the value is greater than this threshold value, the mode of the current frame is determined to be active (ON mode). Otherwise, the mode is determined to be inactive (OFF mode).

[0097] In a variant form, the determination block may perform the determination "in closed-loop mode"; that is, this involves applying the PCA processing of blocks 100 to 150 to obtain an initial version of the transformed signal B prior to confirming that the determination mode within the current frame of the index t is ON. In this case, the eigenvalue λ i can be replaced by the energy of each channel of the initial version B. If the determination is finally OFF, the initial version of the transformed signal B must be replaced.

[0098] In a variant form where closed-loop mode determination is used, additional activation criteria (such as detection of the maximum absolute value within each individual channel before and after PCA (in A and B)) may be added; that is, if this absolute value (within the current frame) within one of the channels of B exceeds that of the corresponding channel in the input signal A, the mode is set to OFF within the current frame.

[0099] FIG. 4 shows a decoder that implements the decoding method according to an embodiment of the present invention.

[0100] The bitstream is demultiplexed at 400, and the decoder 220 receives the channels of the multiplexed channel signal that is decoded according to the binary allocation determined at 420.

[0101] Module 410 receives an indicator for activation of the decorrelation processing of the current frame and applies decoding and conversion processing operations adapted to this indicator (in the same way as performed in encoding).

[0102] Depending on the value of the indicator of the previous frame, the following four combinations are identified: ● If mode = ON and prev_mode = ON, the same operations as those in FIG. 2 are obtained again: In other words, blocks 220, 230, 430, 250, and 260 in FIG. 4 apply the same processing operations as blocks 220, 230, 240, 250, and 260 described in FIG. 2, respectively. Branch 2, where the decoded signal undergoes conversion by block 260, is selected by block 440. The module for decoding the conversion information received in the bit stream is implemented by connection block 430. Allocation block 420 uses R - 1 - N Q bits for decoding the channel at 220. In the default setting, the

Number

Number

Number

Number

Number

Number

[0103] Note that the algorithm delays associated with encoding and decoding (blocks 170 and 220) must be corrected by storing the values of the inverse matrix V in memory for each sub - frame not only within the current frame but also within the previous frame. For example, if the encoding / decoding (blocks 170 - 220) is multi - monaural EVS encoding, there will usually be a 12 - ms delay to be corrected. Thus, if interpolation uses N = 40 sub - frames (i.e., 0.5 - ms sub - frames) within a 20 - ms frame, this will require a memory of 40 + 24 = 64 matrix values (4×4 matrices), and block 260 will apply the time - shifted matrices of the past 24 sub - frames.

[0104] After processing the previous frame, the eigenvalue

Number

Number

[0105] Figure 5 shows the encoding device DCOD and the decoding device DDEC in the context of the present invention, which are dual to each other (in the sense of "reversible") and are connected together via a communication network RES.

[0106] The encoding device DCOD includes a processing circuit, which typically includes the following: - A memory MEM1 for storing data of instructions of a computer program in the context of the present invention (where these instructions can be distributed between the encoder DCOD and the decoder DDEC); - An interface INT1 for receiving an ambisonic signal distributed over various channels (e.g., primarily four channels W, Y, Z, X) for the purpose of their compression encoding in the context of the present invention; - A processor PROC1 for processing these signals by executing computer program instructions stored in the memory MEM1 for the purpose of their encoding; and - A communication interface COM1 for transmitting the encoded signal via a network.

[0107] The decoding device DDEC includes its own processing circuit, which typically includes the following: - A memory MEM2 for storing data of computer program instructions in the sense of the present invention, where these instructions can be distributed between the encoder DCOD and the decoder DDEC as previously indicated; - An interface COM2 for receiving an encoded signal from the network RES for the purpose of decompressing them in the sense of the present invention; - A processor PROC2 for processing these signals by executing computer program instructions stored in the memory MEM2 for the purpose of decoding them; and - An output interface INT2 for delivering the decoded signals in the form of ambisonic channels W’, Y’, Z’, X’ for purposes such as their playback.

[0108] It goes without saying that this Figure 5 shows an example of a structural embodiment of a codec (encoder or decoder) in the sense of the present invention. Figures 3 to 4 explain the functional embodiments of these codecs in more detail.

Claims

1. A method of encoding an audio signal that forms a series of frames (t-1, t) of samples within each of n channels as an ambisonic representation of degree greater than 0 over time, comprising: - Determining a binary value indicating an active mode (ON) or an inactive mode (OFF) of a whitening process applied to the signal of the current frame with respect to the current frame to be encoded, and encoding this value into a bitstream (300); - Encoding whitening process information into the bitstream (310) when it is determined that the mode is active; - Generating an output signal to be encoded into the bitstream according to the mode determined for the current frame and the mode determined for the previous frame (330). A method comprising the steps of:

2. The method according to claim 1, wherein the determination of the binary value indicating the active mode or the inactive mode is performed according to at least one gain determination criterion for the encoded signal before and after the whitening process.

3. The encoding gain is defined by the following logarithmic value: 【Number 1】 where 【Number 2】 is the energy of the input channel before the decorrelation process, λ i is the eigenvalue of the input channel, and the mode is determined to be inactive with respect to a predetermined value of the gain G, the method according to claim 2.

4. The method according to claim 1, wherein the determination of the binary value indicating the active mode or the inactive mode is performed according to a determination criterion for the inter-frame distance between rotation matrices to which the whitening process is applied.

5. The method according to claim 4, wherein the rotation matrix is represented as a dual quaternion, and the inter-frame distance between the rotation matrices is expressed by using the scalar product between the quaternion in the current frame and the quaternion in the previous frame.

6. The method according to claim 1, wherein the determination of the binary value indicating the active mode or the inactive mode is performed according to a distance determination criterion between the rotation matrix of the current frame and the identity matrix after applying the whitening process.

7. The method according to claim 6, wherein the rotation matrix is represented as a dual quaternion, and the distance between the rotation matrix of the current frame and the identity matrix is expressed in the form of the scalar product between the quaternion in the current frame and the unit quaternion.

8. A method of decoding an audio signal that forms a series of frames (t-1, t) of samples within each of n channels as an ambisonic representation of degree greater than 0 over time, comprising: - For the current frame (t), in addition to the signals of the n channels of this current frame, receiving a binary value indicating an active mode (ON) or an inactive mode (OFF) of the decorrelation process applied to the signals of the current frame (410); - When it is determined that the mode is active, decoding the decorrelation process information received in the bit stream (430); - A method including generating an output signal (440) according to the mode determined for the current frame and the mode determined for the previous frame.

9. An encoding device including a processing circuit for implementing the steps of the encoding method according to any one of Claims 1 to 7.

10. A decoding device including a processing circuit for implementing the steps of the decoding method according to Claim 8.

11. A storage medium storing in a memory a computer program including instructions for executing the encoding method according to any one of Claims 1 to 7 or the decoding method according to Claim 8, the storage medium being readable by a processor.

Citation Information

Patent Citations

  • Spatialized audio coding with interpolation and quantification of rotations

    WO2020177981A1