Ambisonics Audio Encoding Phase Discontinuity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for encoding and decoding three-dimensional stereophonic audio signals suffer from phase discontinuity issues, leading to spatial discontinuity and artefacts when sound sources are placed or move in certain directions, particularly due to the use of generic panning laws and separation of signals into ambient and directional components.
Innovation Solution
A method that converts complex frequency coefficients into spatial correspondence in azimuth and elevation coordinates, using a frequency representation of First-Order Ambisonics (FOA) signals, allowing for continuous phase representation without requiring non-directional components or matrix encoding, thereby maintaining stability and location accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If generic panning laws and separation of signals into ambient and directional components are used, then encoding and decoding of three-dimensional audio signals can be achieved, but phase discontinuity occurs leading to spatial discontinuity and artefacts
Solution Approach 1:
The patent extracts the problematic generic panning law separation into a specific directional component model based on First-Order Ambisonics (FOA). By taking out the ambient/directional separation and replacing it with FOA's explicit spatial representation, the patent eliminates phase discontinuity while maintaining encoding/decoding functionality.
Solution Approach 2:
The patent changes the parameter representation from generic panning law coefficients to FOA-based spatial coordinates (azimuth and elevation). This parameter transformation ensures continuous phase representation by using physically meaningful spatial parameters that maintain stability across all directions, including extreme positions.
2Measurement precision
If sound sources are placed or moved in certain directions using conventional methods, then spatial positioning is achieved, but phase discontinuity and spatial discontinuity artefacts occur
Solution Approach 1:
The patent implements dynamic spatial positioning using FOA direction vectors that continuously adapt to sound source movement. The direction vectors are updated based on current azimuth and elevation coordinates, ensuring smooth transitions and continuous phase representation as sources move through any spatial path, eliminating discontinuity artefacts.
Solution Approach 2:
The patent transitions from two-dimensional panning law representations to three-dimensional FOA spatial coordinates (azimuth, elevation, and radial distance). This dimensional enhancement provides continuous phase representation in all spatial directions, eliminating the discontinuities that occur in conventional 2D panning approaches at certain angles.
3Quantity of substance
If matrix encoding is used to encode three-dimensional audio, then spatial information can be compressed, but phase discontinuity and instability occur at extreme positions
Solution Approach 1:
The patent applies preliminary spatial decomposition using FOA direction vectors before encoding. By pre-calculating the spatial representation in terms of continuous azimuth and elevation coordinates, the system ensures phase stability is maintained throughout the encoding process, particularly at extreme positions where conventional matrix encoding fails.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method for the conversion, encoding, decoding and transcoding of a sound field, especially a first-order Ambisonics three-dimensional sound field, including at least one method of converting said sound field into a spherical field, a method of encoding said spherical field into a stereophonic signal, a method of decoding a stereophonic signal to a spherical field, or a method of transcoding said spherical field to a randomly chosen audio format. According to the method of encoding the Ambisonics sound field into a spherical field, said sound field is separated, in the frequency domain, into three components, or optionally two components, and these components are recombined into a total spherical field. According to the method of encoding the spherical field into a stereophonic signal, in the frequency domain, the panning values and phase-difference values are determined, the singularity of the phase difference in the interchannel domain is determined, the phase correspondence function in the interchannel domain is determined, and the left-hand and right-hand components of the signal encoded into stereophonic form are calculated. The spherical coordinates are optionally subjected to affine modification such that they correspond to the standard geometric disposition of the left-hand and right-hand channels. The method of decoding into a spherical field is applied to any stereophonic signal, in particular a stereophonic signal obtained by the above encoding method. According to the aforementioned method of decoding into a spherical field, in the frequency domain, the panning and phase difference are determined, the new position of the phase difference singularity in the interchannel domain is determined, said position varying temporally, the phase correspondence function in the interchannel domain is determined, a complex coefficient corresponding to the desired spherical field is determined, and the direction of provenance in the spherical field is determined, said direction being optionally subject to affine modification in order to correspond to the standard geometric disposition of the left-hand and right-hand channels. The method of transcoding on the basis of a stereophonic signal comprises the above method of decoding to a spherical field, followed by a method ensuring that the spherical field is projected on a specified audio panning law or a binauralisation method.