6DOF rendering of audio captured by microphone array
Through high-order panoramic surround sound audio source and short-time Fourier transform technology, combined with cross fade method, the problem of 6-degree-of-free audio rendering in virtual reality and augmented reality systems is solved, high-quality spatial audio output is achieved, and the realism and immersion of the audio environment is enhanced.
Patent Information
- Application Number
- CN202380090819.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-09
- Filing Date
- 2023-12-15
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art is difficult to effectively capture and render spatial audio of 6 degrees of freedom in virtual reality and augmented reality systems, especially when the listener moves and rotates freely in an audio environment, resulting in poor audio rendering.
High-order panoramic surround sound audio source and short-time Fourier transform technology are used, combined with the cross fade method, spatial audio output is generated. By analyzing the listener position and audio source position, the audio signal processing is dynamically adjusted to achieve 6 degrees of freedom audio rendering.
High-quality spatial audio rendering when listeners move and rotate freely in virtual reality and augmented reality systems is achieved, improving the realism and immersion of the audio environment.
Smart Images

Figure CN120476615A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to apparatus and methods for audio rendering of a 6-DOF system using audio captured by a microphone array. Background Art
[0002] Spatial audio capture methods attempt to capture an audio environment or audio scene so that it can be perceptually recreated for the listener in an efficient manner, and also allow the listener to move and / or rotate within the recreated audio environment. For spatial audio capture and to linearly record spatial sound at a single location in the recording space, a high-end microphone array is required. One such microphone is the spherical 32-microphone Eigenmike. High-order ambisonics (HOA) signals can be obtained from high-end microphone arrays and used for rendering. Using the HOA audio signals, spatial audio can be rendered so that sounds arriving from different directions are well separated within a reasonable hearing bandwidth. In some systems, multiple microphone positions enable multi-point HOA (MPHOA) capture systems, in which multiple HOA audio signals are present at multiple locations within the audio scene. In some embodiments, even basic microphone arrays for approximating first-order ambisonics can be used to record the audio scene. In some other embodiments, the audio scene can include two or more synthesized FOA or HOA sources.
[0003] The audio rendering (wherein the captured audio signals are presented to the listener) can be part of a virtual reality (VR) or augmented reality (AR) system. Furthermore, the audio rendering can be performed as part of a VR or AR where the listener can freely move and rotate their head within the environment or audio scene, which is referred to as a 6 degrees of freedom (6DoF) configuration. Furthermore, the audio rendering can be a multi-point HOA (MPHOA) audio rendering where the audio scene includes multiple HOA audio signal recordings that are rendered to the user in a 6DoF manner. That is, the user is able to listen to the recorded scene from locations other than the location of the recorded HOA source. Summary of the Invention
[0004] According to a first aspect, a method for generating a spatialized audio output is provided, comprising: obtaining at least two higher-order ambisonics audio sources, wherein the at least two higher-order ambisonics audio sources are associated with corresponding audio source positions within an audio environment; obtaining a listener position within the audio environment, wherein the listener position is freely movable within the audio environment; determining at least one source as an active source; determining at least one current higher-order ambisonics audio source from at least one determined active source based on a current listener position; determining at least one previous higher-order ambisonics audio source associated with a signal interpolation determination; performing processing on all channel signals of the at least one current higher-order ambisonics audio source; performing processing on at least one channel signal of another active source in the at least one determined active source; determining a difference between the at least one current higher-order ambisonics audio source and the at least one previous higher-order ambisonics audio source associated with the signal interpolation determination; continuing processing on all channel signals of the at least one previous active higher-order ambisonics audio source; processing at least one channel signal of the at least one previous higher-order ambisonics audio source associated with the signal interpolation determination; the processing of at least one channel signal of at least one other active source among the at least one determined active source; processing all channel signals of at least one current higher-order panoramic surround sound audio source among the current higher-order panoramic surround sound audio sources; performing signal interpolation based on: within a first time period after determining the difference, continuing the processing of channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation; crossfading between continuing the processing of all channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; and ceasing to continue processing all channel signals of the at least one previous active higher-order panoramic surround sound audio source after the crossfading ends; and generating a spatialized audio output based on the determined at least one signal interpolation.
[0005] The processing may be a short-time Fourier transform.
[0006] The cross-fade may be during a second time period.
[0007] The first time period may be a first number of time frames, the first number of time frames being based on a short time Fourier transform priming delay, and the second time period may be a second defined number of processing frames.
[0008] The first number of time frames may be 6 frames, and the second number of time frames may be 12 frames.
[0009] The method may further include determining the spatial metadata by: analyzing processing of all channel signals of at least one current high-order panoramic surround sound audio source and at least one channel signal of another active source among at least one determined active source; and analyzing continued processing of at least one channel signal of at least one previously active high-order panoramic surround sound audio source.
[0010] Generating the spatialized audio output may be further based on the determined spatial metadata.
[0011] The at least one corresponding channel signal of the at least one previous higher-order panoramic surround sound audio source may be a subset of channels. The at least one channel signal of the at least one previous higher-order panoramic surround sound audio source may be first four channel signals of the at least one previous higher-order panoramic surround sound audio source.
[0012] Based on the listener position, determining at least one current higher-order panoramic surround sound audio source from at least two higher-order panoramic surround sound audio sources may include: determining an area in which the current listener position is located, the area being defined by vertex positions of at least three higher-order panoramic surround sound audio sources; and selecting at least one current higher-order panoramic surround sound audio source from the at least two higher-order panoramic surround sound audio sources, the at least one current higher-order panoramic surround sound audio source being the current higher-order panoramic surround sound audio source at the vertex of the area defined by its position.
[0013] According to a second aspect, a device for generating a spatialized audio output is provided, the device comprising components configured to: obtain at least two higher-order panoramic surround sound audio sources, wherein the at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within an audio environment; obtain a listener position within the audio environment, wherein the listener position is freely movable within the audio environment; determine at least one source as an active source; determine at least one current higher-order panoramic surround sound audio source from at least one determined active source based on a current listener position; determine at least one previous higher-order panoramic surround sound audio source associated with a signal interpolation determination; perform processing of all channel signals of the at least one current higher-order panoramic surround sound audio source; perform processing of at least one channel signal of other active sources in the at least one determined active source; determine a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continue processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source; perform processing of at least one channel signal of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination The invention relates to a method for processing a channel signal of at least one of the at least two determined active sources; continuing processing of at least one channel signal of the other active source in at least one determined active source; processing all channel signals of at least one current higher-order panoramic surround sound audio source in the current higher-order panoramic surround sound audio sources; performing signal interpolation based on the following items: within a first time period after determining the difference, continuing processing of channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation; after the first time period after determining the difference, cross-fading between continuing processing of all channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source in at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources; and stopping continuing processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source after the cross-fade ends; and generating a spatialized audio output based on the determined at least one signal interpolation.
[0014] The processing may be a short-time Fourier transform.
[0015] The cross-fade may be during a second time period.
[0016] The first time period may be a first number of time frames based on a short time Fourier transform start delay, and the second time period may be a second defined number of processing frames.
[0017] The first number of time frames may be 6 frames, and the second number of time frames may be 12 frames.
[0018] The component can be further configured to determine the spatial metadata by: analyzing the processing of all channel signals of at least one current high-order panoramic surround sound audio source and at least one channel signal of other active sources among at least one determined active source; analyzing the continued processing of at least one channel signal of at least one previously active high-order panoramic surround sound audio source.
[0019] Generating the spatialized audio output may be further based on the determined spatial metadata.
[0020] The at least one corresponding channel signal of the at least one previous higher-order panoramic surround sound audio source may be a subset of channels. The at least one channel signal of the at least one previous higher-order panoramic surround sound audio source may be first four channel signals of the at least one previous higher-order panoramic surround sound audio source.
[0021] The component configured to determine at least one current higher-order panoramic surround sound audio source from at least two higher-order panoramic surround sound audio sources based on a listener position can be configured to: determine an area in which the current listener position is located, the area being defined by vertex positions of at least three higher-order panoramic surround sound audio sources; and select at least one current higher-order panoramic surround sound audio source from the at least two higher-order panoramic surround sound audio sources, the at least one current higher-order panoramic surround sound audio source being the current higher-order panoramic surround sound audio source at the vertex of the area defined by its position.
[0022] According to a third aspect, a device for generating a spatialized audio output is provided, the device comprising at least one processor and at least one memory storing instructions, which, when executed by the at least one processor, causes the system to perform at least the following operations: obtain at least two higher-order panoramic surround sound audio sources, wherein the at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within an audio environment; obtain a listener position within the audio environment, wherein the listener position is freely movable within the audio environment; determine at least one source as an active source; determine at least one current higher-order panoramic surround sound audio source from at least one determined active source based on a current listener position; determine at least one previous higher-order panoramic surround sound audio source associated with a signal interpolation determination; perform processing of all channel signals of the at least one current higher-order panoramic surround sound audio source; perform processing of at least one channel signal of other active sources in the at least one determined active source; determine a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continue processing of all channel signals of the at least one previously active higher-order panoramic surround sound audio source; perform processing of at least one channel signal of other active sources in the at least one determined active source; processing at least one channel signal of a previous higher-order panoramic surround sound audio source; continuing processing of at least one channel signal of another of the at least one determined active source; processing all channel signals of at least one current higher-order panoramic surround sound audio source of the current higher-order panoramic surround sound audio source; performing signal interpolation based on: continuing processing of channel signals of at least one of the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation within a first time period after determining the difference; cross-fading between continuing processing of all channel signals of at least one of the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; and ceasing to continue processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source after the cross-fade ends; and generating a spatialized audio output based on the determined at least one signal interpolation.
[0023] The processing may be a short-time Fourier transform.
[0024] The cross-fade may be during a second time period.
[0025] The first time period may be a first number of time frames based on a short time Fourier transform start delay, and the second time period may be a second defined number of processing frames.
[0026] The first number of time frames may be 6 frames, and the second number of time frames may be 12 frames.
[0027] The apparatus may be further configured to determine the spatial metadata by: analyzing processing of all channel signals of at least one current high-order panoramic surround sound audio source and at least one channel signal of another active source among at least one determined active source; and analyzing continued processing of at least one channel signal of at least one previously active high-order panoramic surround sound audio source.
[0028] The apparatus being caused to perform generating the spatialized audio output may be further based on the determined spatial metadata.
[0029] The at least one corresponding channel signal of the at least one previous higher-order panoramic surround sound audio source may be a subset of channels. The at least one channel signal of the at least one previous higher-order panoramic surround sound audio source may be first four channel signals of the at least one previous higher-order panoramic surround sound audio source.
[0030] The device, which is configured to determine at least one current higher-order panoramic surround sound audio source from at least two higher-order panoramic surround sound audio sources based on a listener position, may be configured to perform: determining an area in which the current listener position is located, the area being defined by vertex positions of at least three higher-order panoramic surround sound audio sources; and selecting at least one current higher-order panoramic surround sound audio source from the at least two higher-order panoramic surround sound audio sources, the at least one current higher-order panoramic surround sound audio source being the current higher-order panoramic surround sound audio source at the vertex of the area defined by its position.
[0031] According to a fourth aspect, a device for generating a spatialized audio output is provided, the device comprising: an obtaining circuit configured to obtain at least two high-order panoramic surround sound audio sources, wherein the at least two high-order panoramic surround sound audio sources are associated with corresponding audio source positions within an audio environment; an obtaining circuit configured to obtain a listener position within the audio environment, wherein the listener position is freely movable within the audio environment; a determining circuit configured to determine at least one source as an active source; a determining circuit configured to determine at least one current high-order panoramic surround sound audio source from at least one determined active source based on a current listener position; and a determining circuit configured to determine at least one previous high-order panoramic surround sound audio source associated with the signal interpolation determination; an execution circuit configured to execute processing of all channel signals of at least one current high-order panoramic surround sound audio source; an execution circuit configured to execute processing of at least one channel signal of other active sources in at least one determined active source; a determination circuit configured to determine a difference between at least one current high-order panoramic surround sound audio source and at least one previous high-order panoramic surround sound audio source associated with the signal interpolation determination; a continuation circuit configured to continue processing of all channel signals of at least one previous active high-order panoramic surround sound audio source; and a processing circuit configured to execute processing of at least one channel signal of at least one other active source in at least one determined active source. a continuing circuit configured to continue processing at least one channel signal of at least one previous high-order panoramic surround sound audio source among the at least one determined active source; processing all channel signals of at least one current high-order panoramic surround sound audio source among the current high-order panoramic surround sound audio sources; an executing circuit configured to perform signal interpolation based on the following items: within a first time period after determining the difference, continuing processing of the channel signal of at least one previous high-order panoramic surround sound audio source among the at least one previous high-order panoramic surround sound audio source associated with the signal interpolation; after the first time period after determining the difference, A cross-fade between continued processing of all channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from at least two higher-order panoramic surround sound audio sources; and a stopping circuit configured to stop continuing processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source after the cross-fade ends; and a generating circuit configured to generate a spatialized audio output based on the determined at least one signal interpolation.
[0032] According to a fifth aspect, there is provided a computer program [or a computer-readable medium comprising instructions], the instructions being for causing an apparatus to generate a spatialized audio output, the apparatus being caused to perform at least the following operations: obtaining at least two higher-order panoramic surround sound audio sources, wherein the at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within an audio environment; obtaining a listener position within the audio environment, wherein the listener position is free to move within the audio environment; determining at least one source as an active source; determining at least one current higher-order panoramic surround sound audio source from at least one determined active source based on a current listener position; determining at least one previous higher-order panoramic surround sound audio source associated with a signal interpolation determination; performing processing of all channel signals of the at least one current higher-order panoramic surround sound audio source; performing processing of at least one channel signal of other active sources in the at least one determined active source; determining a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continuing processing of all channel signals of the at least one previously active higher-order panoramic surround sound audio source; performing processing of at least one channel signal of other active sources in the at least one determined active source; processing at least one channel signal of a higher-order panoramic surround sound audio source; continuing processing of at least one channel signal of another active source among at least one determined active source; processing all channel signals of at least one current higher-order panoramic surround sound audio source among the current higher-order panoramic surround sound audio sources; performing signal interpolation based on: continuing processing of channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation within a first time period after determining the difference; cross-fading between continuing processing of all channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; and ceasing to continue processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source after the cross-fade ends; and generating a spatialized audio output based on the determined at least one signal interpolation.
[0033] According to a sixth aspect, a non-transitory computer-readable medium comprising program instructions is provided, the program instructions being configured to cause an apparatus for generating a spatialized audio output to at least perform the following operations: obtain at least two higher-order panoramic surround sound audio sources, wherein the at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within an audio environment; obtain a listener position within the audio environment, wherein the listener position is freely movable within the audio environment; determine at least one source as an active source; determine at least one current higher-order panoramic surround sound audio source from at least one determined active source based on a current listener position; determine at least one previous higher-order panoramic surround sound audio source associated with a signal interpolation determination; perform processing of all channel signals of the at least one current higher-order panoramic surround sound audio source; perform processing of at least one channel signal of other active sources in the at least one determined active source; determine a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continue processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source; perform processing of at least one channel signal of other active sources in the at least one determined active source; The invention relates to a method for processing at least one channel signal of a surround sound audio source; continuing processing of at least one channel signal of another active source among at least one determined active source; processing all channel signals of at least one current higher-order panoramic surround sound audio source among the current higher-order panoramic surround sound audio sources; performing signal interpolation based on: continuing processing of channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation within a first time period after determining the difference; cross-fading between continuing processing of all channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; and ceasing to continue processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source after the cross-fade ends; and generating a spatialized audio output based on the determined at least one signal interpolation.
[0034] According to a seventh aspect, there is provided an apparatus for generating a spatialized audio output, comprising: a component for obtaining at least two higher-order panoramic surround sound audio sources, wherein the at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within an audio environment; a component for obtaining a listener position within the audio environment, wherein the listener position is freely movable within the audio environment; a component for determining at least one source as an active source; a component for determining at least one current higher-order panoramic surround sound audio source from at least one determined active source based on a current listener position; a component for determining at least one prior associated with a signal interpolation determination; Components for performing processing of all channel signals of at least one current higher-order panoramic surround sound audio source; components for performing processing of at least one channel signal of another active source among at least one determined active source; components for determining a difference between at least one current higher-order panoramic surround sound audio source and at least one previous higher-order panoramic surround sound audio source associated with a signal interpolation determination; components for continuing processing of all channel signals of at least one previous active higher-order panoramic surround sound audio source; components for processing at least one channel signal of at least one previous higher-order panoramic surround sound audio source associated with a signal interpolation determination; components for continuing processing of at least one channel signal of another active source among at least one determined active source; components for processing all channel signals of at least one current higher-order panoramic surround sound audio source among the current higher-order panoramic surround sound audio source; components for performing signal interpolation based on: continued processing of channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation within a first time period after determining the difference; processing of channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation after a first time period after determining the difference A cross-fade between continued processing of all channel signals of at least one previous higher-order panoramic surround sound audio source of the interpolation-associated at least one previous higher-order panoramic surround sound audio source and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from at least two higher-order panoramic surround sound audio sources; and means for stopping continued processing of all channel signals of at least one previously active higher-order panoramic surround sound audio source after the cross-fade ends; and means for generating a spatialized audio output based on the determined at least one signal interpolation.
[0035] According to an eighth aspect, a computer-readable medium comprising instructions is provided, the instructions being for causing an apparatus for generating a spatialized audio output to perform at least the following operations: obtain at least two higher-order panoramic surround sound audio sources, wherein the at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within an audio environment; obtain a listener position within the audio environment, wherein the listener position is freely movable within the audio environment; determine at least one source as an active source; determine at least one current higher-order panoramic surround sound audio source from at least one determined active source based on a current listener position; determine at least one previous higher-order panoramic surround sound audio source associated with a signal interpolation determination; perform processing of all channel signals of the at least one current higher-order panoramic surround sound audio source; perform processing of at least one channel signal of other active sources in the at least one determined active source; determine a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continue processing of all channel signals of the at least one previously active higher-order panoramic surround sound audio source; perform processing of at least one channel signal of other active sources in the at least one determined active source; the processing of at least one channel signal of at least one of the at least one determined active sources; continuing the processing of at least one channel signal of the other active sources among the at least one determined active sources; processing all channel signals of at least one current higher-order panoramic surround sound audio source among the current higher-order panoramic surround sound audio sources; performing signal interpolation based on: continuing the processing of channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation within a first time period after determining the difference; cross-fading between continuing the processing of all channel signals of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; and ceasing to continue the processing of all channel signals of the at least one previous active higher-order panoramic surround sound audio source after the cross-fade ends; and generating a spatialized audio output based on the determined at least one signal interpolation.
[0036] An apparatus comprises means for performing the actions of the method as described above.
[0037] An apparatus is configured to perform the actions of the method described above.
[0038] A computer program includes program instructions for causing a computer to execute the method described above.
[0039] A computer program product stored on a medium may cause an apparatus to execute the method as described herein.
[0040] An electronic device may include an apparatus as described herein.
[0041] A chipset may include the apparatus as described herein.
[0042] The embodiments of the present application are intended to solve the problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings, in which:
[0044] Figure 1 schematically illustrates an apparatus system according to some embodiments showing an audio rendering or reproduction of an example audio scene and within which a user can move;
[0045] Figure 2 schematically illustrates an example audio scene comprising a reproduction of an audio scene in which a user moves within an area determined by a high-order panoramic surround sound audio signal source;
[0046] Figure 3 Schematically shown as Figure 2 , and wherein the source can be identified as a higher-order panoramic surround sound audio source or a full-order panoramic surround sound audio source;
[0047] Figure 4 Schematically shown as Figure 2 or Figure 3 The example audio scene shown in , where the listener moves and a crossfade is implemented;
[0048] Figure 5 Schematically shown as Figure 4 The example audio scene shown in follows a crossfade;
[0049] Figure 6 illustrates an apparatus suitable for implementing some embodiments, wherein the capture apparatus may be separate from the rendering apparatus element;
[0050] Figure 7 an example flow chart illustrating the operation of an example apparatus according to some embodiments; and
[0051] Figure 8 An exemplary device suitable for implementing the illustrated means is schematically illustrated. DETAILED DESCRIPTION
[0052] The concepts discussed in further detail herein with respect to the following embodiments relate to audio scene rendering, wherein the audio scene is captured based on linear or parametric spatial audio methods using two or more microphone arrays corresponding to different locations in the recording space (or in other words, using audio signal sets captured at corresponding signal set locations in the recording space). Furthermore, the concepts also involve attempts to reduce the computational complexity of the spatial analysis required for MPHOA processing.
[0053] As discussed above, 6DoF is currently common in virtual reality (such as VR games), where movement in the audio scene is easy to render because all spatial information is readily available (i.e., the position of each sound source and the audio signal of each source, respectively).
[0054] In the following examples, the audio signal set is generated by a microphone (or microphone array). For example, the microphone arrangement may include one or more microphones and generate one or more audio signals for the audio signal set. In some embodiments, the audio signal set includes audio signals that are virtual audio signals or generated audio signals (e.g., virtual speaker audio signals with associated virtual speaker positions). In addition, in some embodiments, the microphone array is separated from or physically remote from any processing device, however, this does not exclude examples in which the microphone is located on or physically connected to the processing device.
[0055] about Figure 1 , shows an example apparatus that can be configured to implement MPHOA processing according to some embodiments. In some embodiments, the apparatus is part of a suitable MPEG-I Audio Reference Audio Renderer.
[0056] In some embodiments, the apparatus 101 comprises a pre-processor 103. The pre-processor 103 is configured to receive the head-related impulse response (HRIR) 100 and the high-order Ambisonic microphone position p i 104 and generates a head related transfer function (HRTF) 110 and the determined microphone triangle T108.
[0057] In some embodiments, the apparatus 101 comprises a position pre-processor 105 configured to receive the position p of the listener. i 104. High-order Ambisonic microphone position p L (j) The position of 106 and T108, and thus generate T A (j) 112 (active triangle for frame j), 128 (interpolation weight selected for subframe k of frame j), m c(j) 126 (HOA source selected for frame j).
[0058] In some embodiments, the apparatus 101 includes a spatial analyzer 107 configured to receive from an input s ESD (i, j) 102 (input time domain HOA signal in equivalent spatial domain representation) and receives T from position preprocessor 105 A The spatial analyzer 107 is configured to generate metadata θ(i,j,k,b) 116 (azimuth angle for HOA source i, frame j, subframe k, and frequency bin b), θ(i,j,k,b) 118 (elevation angle for HOA source i, frame j, subframe k, and frequency bin b), r(i,j,k,b) 120 (direct-to-total energy ratio for HOA source i, frame j, subframe k, and frequency bin b), e(i,j,k,b) 122 (energy for HOA source i, frame j, subframe k, and frequency bin b), and signal S(i,j,k,b) 114.
[0059] In some embodiments, the apparatus 101 includes a spatial metadata interpolator 111 configured to receive θ(i,j,k,b) 116 from the spatial analyzer 107. 118r(i,j,k,b)120e(i,j,k,b)122 and received from the position preprocessor 105 and thus generate interpolated metadata 134 136 138 132.
[0060] In some embodiments, the apparatus 101 includes a signal interpolator 109 configured to receive m from the position preprocessor 105. c (j) 126 receives S(i,j,k,b) 114 and e(i,j,k,b) 122 from the spatial analyzer 107 and receives 132, thereby generating an interpolation signal 130.
[0061] In some embodiments, the apparatus 101 includes a mixer 113 configured to receive the interpolated signal from the signal interpolator 109. 130, and receives interpolation metadata from the spatial metadata interpolator 111 134 136 138 132. The mixer thus generates output audio O(j,k,b) 142.
[0062] In some embodiments, the apparatus 101 includes an output processor 115 configured to receive the output audio O(j, k, b) 142 (binaural output time-frequency domain signal) from the mixer 113 and generate an output audio signal S out (j) 144 (binaural output time domain audio signal).
[0063] The operation of the spatial analyzer 107 as described above is configured to receive from the input s ESD (i, j) 102, and receives T from the position preprocessor 105 A (j) 112, and generate metadata θ(i,j,k,b) 116, 118, r(i,j,k,b) 120, e(i,j,k,b) 122, and signal S(i,j,k,b) 114. This will be described in more detail herein. The spatial analyzer 107 and renderer will be further described with reference to GB2007710.8, EP21201766.9, and Section 6.6.18 of the MPEG-I Immersive Audio Standard Working Draft (ISO / IEC 23090-4WD).
[0064] The spatial analysis block uses the equivalent spatial domain (ESD) representation of the audio input signal s ESD (i, j) 102 as input and provides spatial metadata as output for determining interpolated spatial metadata (i, j, k, b) 116 at the listener's position 118r(i,j,k,b)120e(i,j,k,b)122 and the time-frequency domain signal S(i,j,k,b)114 to be used for signal interpolation.
[0065] First, these signals are converted into High-Order Ambisonics (HOA) signals as follows:
[0066] s HOA (i,j)=M ESDtoHOA s ESD (i,j),
[0067] Among them, M ESDtoHOA Yes N ch ×N ch ESD to HOA conversion matrix, i is the HOA source index, and j is the frame index. Then, the output HOA signal is divided into N equal length sf Subframes:
[0068] [s HOA (i,j,1) … s HOA(i,j,N sf )]=s HOA (i,j)
[0069] Furthermore, a time-frequency domain transformation is applied to all active HOA sources i. This transformation can be performed using a suitable function such as the afSTFT function, which can be found, for example, at https: / / github.com / jvilkamo / afSTFT.
[0070]
[0071] Among them, S(i,j,k) is N ch ×N b Matrix, containing a length of N for each HOA channel b time-frequency domain signals.
[0072] The afSTFT transformation is run separately for each channel ch. Therefore, the more channels there are to process, the more computationally expensive the processing becomes.
[0073] For each frequency bin b of the signal S(i,j,k), the spatial analysis block computes spatial metadata, including direction, spread, and energy information. These are then passed to the spatial metadata interpolator 111, which, in some embodiments, can be configured to perform interpolation in a manner similar to that described in Section 6.6.18.3.4.1, “Metadata interpolation,” of ISO / IEC 23090-4 WD1, as discussed in further detail below.
[0074] In addition, spatial metadata can be obtained from the signal covariance matrix C FOA The signal covariance matrix is obtained from the signal as follows:
[0075] C FOA (i,j,k,b)=s(i,j,k,b)s H (i,j,k,b),
[0076] in:
[0077]
[0078] Among them, S b,ch (i,j,k) is the value in the matrix S(i,j,k) corresponding to channel ch and frequency bin b.
[0079] Signal S(i c ,j,k)(where i cis the index of the selected HOA source for signal interpolation) is passed to the signal interpolation block, where a prototype binaural signal is calculated from the signal. In some embodiments, the prototype signal creation involves applying EQ gain to the signal, rotating the signal according to the listener head orientation, and then multiplying it with the HOA to binaural transformation matrix. For example, in some embodiments, this can be implemented in a form similar to that described in Section 6.6.18.3.4.2 "Signal Interpolation" in ISO / IEC 23090-4 WD1, as discussed in further detail below.
[0080] In some embodiments, to reduce computational complexity, spatial analysis (and the signal processing described above) can be performed on sources that require spatial analysis (i.e., active sources). The determination of which HOA sources are active can be performed based on the listener's position in the scene relative to the HOA sources in the scene and the triangulation of the HOA source positions that has been performed in the pre-processing stage.
[0081] For example, in some embodiments, an active source is a source that includes a triangle within it.
[0082] exist Figure 2 An example scene 201 is shown in FIG. The audio scene 201 includes microphones m1 203, m2 205, m3 207, and m4 209, which are arranged such that two triangles are defined by the positions of the microphones. Thus, the scene includes a first triangle defined by a "connection" 204 between m1 203 and m2 205, a "connection" 208 between m2 205 and m4 209, and a "connection" 206 between m1 203 and m4 209. Furthermore, the scene includes a second triangle defined by a "connection" 210 between m3 207 and m2 205, a "connection" 208 between m2 205 and m4 209, and a "connection" 212 between m4 209 and m3 207. In this example, the listener p l 211 is located within the second triangle. Therefore, the active sources are the microphones that form the vertices or corners of the second triangle m2205, m3207 and m4209.
[0083] The signal interpolator 109 is further configured to select an audio signal associated with the HOA source closest to the listener as the input audio signal. In turn, the signal interpolator 109 is configured to create an interpolated time-frequency domain audio signal. In some embodiments, this may be achieved by processing the selected input signal corresponding to the HOA source closest to the user, for example, by applying an equalization (EQ) gain to the selected audio signal. For example, in Figure 2In the case shown in , the selected source is m2205. In some embodiments, this selection can be implemented in a manner as described in ISO / IEC 23090-4 WD, section 6.6.18.3.2.4, "Determining HOA Source for Signal Interpolation," as discussed in further detail below.
[0084] In the event that the listener moves so that the HOA source selected for signal interpolation is no longer his closest HOA source, the signal interpolator 109 can be configured to implement a cross-fade process to smoothly transition to using a new HOA source for signal interpolation. During the cross-fade (which lasts for 12 frames), the signal interpolator 109 creates a cross-faded frequency domain signal from the previous closest HOA source and the new HOA closest source by selecting frequency bands from the two signals. As the cross-fade progresses, more frequency bands are selected from the new closest signal and fewer frequency bands are selected from the previous closest signal. In some embodiments, cross-fading can be implemented using the example shown in Section 6.6.18.3.2.5 "Cross-fading" of ISO / IEC 23090-4 WD1, as discussed in further detail below.
[0085] To support the listener moving to a new triangle, the HOA sources comprising the new triangle will be added to the active HOA source list (i.e., the list of HOA sources for which the STFT is run). Since only the three HOA sources belonging to the previous triangle are active for the STFT before moving to the new triangle, and since the STFT used here takes 6 frames of input for the output to be meaningful, the new HOA sources in the new triangle are not yet ready for processing. Therefore, the system uses delayed switching of triangles. For 6 frames, or the time it takes for the STFT to produce a meaningful output, the spatial metadata interpolator 111 is configured to perform interpolation using the HOA sources of the previous triangle. Once the output of the STFT indicates the HOA sources of the new triangle, the process will resume normal operation using the HOA sources of the new triangle. The HOA sources that are not part of the new triangle will be discarded from the active HOA source list.
[0086] Thus, as discussed above, for 6DoF HOA rendering, the spatial analyzer 107 may be configured to calculate or determine spatial metadata for all HOA sources required for spatial metadata interpolation, and the signal interpolator uses the most suitable HOA source (e.g., the closest HOA source) for signal interpolation.
[0087] Typically, the signal is transformed and processed in the frequency domain (based on STFT processing), and this processing is applied to a few selected channels for spatial metadata calculation. On the other hand, for signal interpolation, it is beneficial to have the maximum amount of information available, so the STFT processing is performed on all channels of the HOA source.
[0088] Furthermore, as the listener (or HOA source) moves, the most suitable or determined HOA source currently used for signal interpolation may change. This will cause perceptible discontinuities or errors and thus adversely affect the subjective audio consumption experience. This is because the processing or rendering involves interpolating the HOA source (HOA) from the previous signal. Prev-Sigint ) to the new signal interpolation HOA source (HOA Current-Sigint ) changes. This requires the HOA Prev-Sigint Switch from STFT calculation for all channels to calculation for first-order channels (first four channels), and HOA Current-Sigint Switch from STFT calculation for the first-order channels (the first four channels) to STFT calculation for all channels. This audible error or discontinuity occurs due to the delay in obtaining meaningful STFT results after initiating STFT processing. This delay is equal to a predefined number of audio frames.
[0089] The concepts discussed herein in embodiments herein relate to 6DoF rendering of an audio scene comprising two or more HOA sources, wherein an apparatus is provided that is configured to implement a method for switching between HOA sources for signal interpolation to achieve seamless switching without audible errors. In some embodiments, this can be achieved by initiating STFT processing of all channels of a new signal interpolation HOA source; initiating STFT processing of first-order channels of a previous signal interpolation HOA source; delaying the switch until the STFT processing provides meaningful output for all channels of the new signal interpolation HOA source, and, after the switch, stopping STFT processing of all channels of the previous signal interpolation HOA source (but continuing STFT processing of first-order channels).
[0090] In some embodiments, the method steps may be:
[0091] οGet listener position
[0092] o Get active HOA sources based on listener location
[0093] o Obtain the previous HOA source (HOA) used for signal interpolation Prev-Sigint ), which is used for signal interpolation of the previous audio frame
[0094] o Determine the current HOA source (HOA) for signal interpolation from the active HOA sourcesCurrent-Sigint )
[0095] o If the previous HOA source used for signal interpolation is different from the current HOA source used for signal interpolation, then
[0096] Set the crossfade in progress flag to true, and set the number of frames before the crossfade starts (e.g., the number of audio frames required for STFT priming)
[0097] Initiate STFT calculations for all available channels of the HOA source for the current HOA source used for signal interpolation, while retaining separate STFT calculations for the first four channels
[0098] Initiate STFT calculations for the first four channels of the previous HOA source for signal interpolation, while retaining separate STFT calculations for all available channels of the HOA source
[0099] o If the number of audio frames before the crossfade starts is greater than zero, delay the crossfade until the audio frame count before the crossfade is equal to zero.
[0100] o After the number of frames before the cross-fade is zero, a cross-fade is performed between the previous HOA source for signal interpolation and the current HOA source for signal interpolation
[0101] In some embodiments, determining the difference between the at least one current high-order Ambisonics audio source and the at least one previous high-order Ambisonics audio source associated with the signal interpolation determination may be accomplished by a comparison based on a comparison between the at least one current high-order Ambisonics audio source and the at least one previous high-order Ambisonics audio source associated with the signal interpolation determination.
[0102] Additionally, in some embodiments, if, based on the comparison, it is determined that at least one current high-order Ambisonics audio source is the same as at least one previous high-order Ambisonics audio source associated with the signal interpolation determination, the process continues.
[0103] In some embodiments, STFT enable refers to the number of audio frames that need to be processed before STFT processing provides meaningful results.
[0104] In some embodiments, from the HOA Prev-Sigint To HOA Current-Sigint The switching is performed as follows:
[0105] οInitiate a campaign against HOA Prev-Sigint STFT calculation for first-order channels or any other subset of channels
[0106] οContinue to target HOA Current-Sigint STFT calculation for first-order channels or any other subset of channels
[0107] o Perform signal interpolation based on first-order or any other low-order channel subset
[0108] οInitiate a campaign against HOA Current-Sigint STFT calculation for all channels
[0109] o after meaningful STFT results start to appear (e.g., the delay is defined in terms of number of audio frames),
[0110] Signal interpolation is performed based on the STFT information from all channels.
[0111] In some embodiments, the Prev-Sigint To HOA Current-Sigint During the conversion, signal interpolation is performed so that the candidate HOA Current-Sigint is added to the list of active HOA sources and is determined to be an HOA by predicting the listener's movement based on past location information or any other suitable information. Current-Sigint Previously, the STFT calculation for all channels was initiated.
[0112] In some other embodiments of the present invention, Prev-Sigint To HOA Current-Sigint During conversion, the signal is interpolated with any order between all channels and the first order channel.
[0113] As discussed above, Figure 1 The device for implementing MPHOA processing shown in is configured to render an audio scene comprising a microphone array (high-order Ambisonics) audio signal as input to a listener. In some embodiments, the rendering can be implemented when the listener has 6DoF in their movement. That is, the listener is allowed to move in the audio scene or environment, and its position is inconsistent with the position of the microphone. In such an embodiment, the device renders binaural (or other multi-channel format) audio signals to the listener, which sounds like the expected audio scene is heard from the listener's position. This is not an easy task because there is no direct recording of the audio scene at the listener's position, and therefore the device is configured to implement a method for inferring what the scene would sound like at the listener's position. Therefore, the device is configured to analyze (HOA) microphone signals near the listener's position to estimate the (binaural or multi-channel) microphone audio signals at the listener's position.
[0114] Therefore, some embodiments will be further described herein. Figure 1 The apparatus shown in and briefly described above.
[0115] As discussed above, the pre-processor 103 uses the position p of the HOA microphone array i 104 and a set of HRIR filters 100 as input. Then, the audio scene can be divided into triangular parts T by performing Delauney triangulation. Figure 2 An example is shown and described in the example audio scene shown in . The triangulation can then be used in processing to determine which HOA sources surround the listener and used to generate a binaural signal at the listener's location.
[0116] In addition, the pre-processor 103 can be configured to sample the HRIR filter 100 in a uniform directional grid and convert it into a frequency domain HRTF 110. This is performed when implementing MPHOA processing in the frequency domain. The pre-processor is configured to perform these operations during initialization of the device. Therefore, in some embodiments, the pre-processor 103 is used only once for each audio scene.
[0117] In some embodiments, the position pre-processor 105 may be configured to determine, for each frame of audio, an active triangle T that is used for processing at the spatial analysis block. A 112. Active Triangle T A 112 is a triangle surrounding the listener (or in other words, the triangle in which the listener is located or positioned) from the available triangle portion T108.
[0118] Furthermore, in some embodiments, the position preprocessor 105 is configured to determine or select a "chosen" HOA source for signal interpolation. In some embodiments, the "chosen" HOA source is the source determined to be closest to the listener's position. Furthermore, in some embodiments, the position preprocessor 105 is further configured to determine the interpolation weights 128, these refer to the weights of the HOA sources in the active triangle. The closer the HOA source is to the user, the higher the weighting factor. In some embodiments, the weighting factor can be obtained by calculating the barycentric coordinates of the triangle by solving the following equation:
[0119] T xy,n b n =p L,xy
[0120] Among them, p L,xy =[p x p y 1] is the listener position on the xy plane. T xy,nThe value contains the coordinates of the HOA source in the active triangle on the xy plane. Furthermore, in some embodiments, the barycentric coordinates may be used as a weighting factor.
[0121] In some embodiments, when no crossfading is in progress or the listener is not traveling to a new triangle, the selected HOA source may be marked or indicated by setting a "FullOrder" indicator value, with the remaining HOA sources including the triangle the listener is in being indicated by setting a "FOA" indicator value. For example, Figure 3 As shown in , it shows that Figure 2 , but with microphone m2 identified as the “Full Order” microphone 305 , and microphone m3 identified as the “FOA” 307 and m4 identified as the “FOA” 309 microphone.
[0122] The spatial analyzer 107 receives an input signal s represented in the equivalent spatial domain (ESD). ESD (i, j) 102, and provide spatial metadata θ(i, j, k, b) 116 118r(i,j,k,b)120e(i,j,k,b)122 as output for determining the interpolated spatial metadata interpolation signal at the listener position 134 136 138 132 and the time-frequency domain signal S(i,j,k,b) 114 to be used for signal interpolation.
[0123] In some embodiments, the input signal s is represented in the equivalent spatial domain (ESD) ESD (i,j) 102 is first converted into a high-order ambisonics (HOA) signal as follows:
[0124] s HOA (i,j)=M ESDtoHOA s ESD (i,j),
[0125] Among them, M ESDtoHOA Yes N ch ×N ch ESD to HOA conversion matrix, i is the HOA source index, and j is the frame index. Then, the output HOA signal is divided into N equal length sf Subframes:
[0126] [s HOA (i,j,1) … s HOA (i,j,N sf )]=s HOA(i,j)
[0127] Then, a time-frequency domain transformation is applied to all active HOA sources i. For sources otherwise marked as "FOA", this transformation is performed only on the first four channels:
[0128]
[0129] Among them, S(i,j,k) is 4×N b The matrix containing the first four HOA channels is of length N b time-frequency domain signals.
[0130] For sources that are additionally marked as "FullOrder", the conversion is performed on all channels. In both cases, the function afSTFT[] can be used to perform the conversion.
[0131]
[0132] Among them, S(i,j,k) is N ch ×N b Matrix, containing a length of N for each HOA channel b The afSTFT process is performed separately for each channel ch. Therefore, the more channels there are to process, the more computationally expensive the process becomes.
[0133] For each frequency bin b of the signal S(i,j,k), the spatial analyzer 107 may be configured to calculate spatial metadata, including direction, diffuseness and energy information. These are then passed to the spatial metadata interpolator 111. The spatial metadata is obtained from the signal covariance matrix C FOA Calculated, the signal covariance matrix is obtained from the signal as follows:
[0134] C FOA (i,j,k,b)=s(i,j,k,b)s H (i,j,k,b),
[0135] in:
[0136]
[0137] Among them, S b,ch (i,j,k) is the value in the matrix S(i,j,k) corresponding to channel ch and frequency bin b.
[0138] Then, for each frequency bin of each active HOA source, spatial metadata is calculated from the covariance matrix. This includes direction information, spread information, and energy:
[0139]
[0140] where θ(i,j,k,b) 116 is the azimuth angle, 118 is the elevation angle, r(i,j,k,b) 120 is the direct to total energy ratio, and e(i,j,k,b) 122 is the energy for HOA source i for frame j (subframe k) and frequency bin b. These are obtained as follows. First, the intensity vector is calculated from the covariance matrix:
[0141]
[0142] Then, energy:
[0143]
[0144] And the rest of the spatial metadata:
[0145]
[0146] Signal S(i c ,j,k)(where i c is the index of the HOA source selected for signal interpolation) is passed to the signal interpolation block, where the prototype binaural signal is calculated from this signal.
[0147] As previously indicated, the spatial metadata interpolator 111 is configured to take the metadata associated with the HOA sources of the active triangle (calculated by the spatial analyzer 107) and create interpolated metadata (i.e., metadata at the listener's position). Thus, the spatial metadata interpolator 111 is configured to describe the sound field at the listener's position (what it should sound like at the listener's position, which frequencies come from which direction, how much energy, etc.).
[0148] The output of the spatial metadata interpolator 111 is the interpolated metadata, which is a weighted sum of the spatial metadata of the HOA sources of the active triangle. The weights used for weighted interpolation can be the weights calculated in the position preprocessor 105 128.
[0149] The signal interpolator 109 is configured to take the selected HOA source frequency domain signal S(i c ,j,k) as input, where i c is the index of the HOA source selected for signal interpolation as part of the signal S(i,j,k,b) 114 and provides the prototype frequency domain signal 130 as output. In summary, prototype signal creation involves applying EQ gains to the signal (based on the interpolated signal energy), rotating the signal according to the listener head orientation, and then multiplying it with the HOA to binaural transformation matrix. In some embodiments, the EQ gains are calculated as follows:
[0150]
[0151] Among them, m c (j) is the index of the selected HOA source for frame j.
[0152] The interpolated signal is calculated as follows:
[0153]
[0154] In some embodiments, where the listener moves such that the selected source for signal interpolation is no longer the original closest source, the apparatus is configured to perform a crossfade process to smoothly transition to using the new source for signal interpolation. During the crossfade, which can be configured to last for a determined number of frames (e.g., 12 frames), the system creates a crossfaded frequency domain signal from the previous closest source and the new closest source by selecting frequency bands from the two signals. As the crossfade progresses, more frequency bands are selected from the new closest signal and fewer frequency bands are selected from the previous closest signal.
[0155] For example, in some embodiments, when the crossfade process begins, the new closest source is determined to be the "FOA" source, which means that signal processing and frequency domain transform (STFT) have been performed on the first four channels of the source. However, in order for the crossfade operation to be performed correctly, the frequency domain transform (STFT) of all orders for the new closest source needs to be determined or calculated.
[0156] In some embodiments, the new closest source is also labeled or indicated as a "Full Order" source. Thus, during the crossfade, there are two sources indicated as "Full Order" and a single "FOA" source.
[0157] For example, Figure 4 Shown in Figure 2 and Figure 3 An example of an audio scenario is shown in , where the listener position is moved closer to microphone m4, Figure 3 Thus, the cross-fade operation 421 is configured with microphone m2 identified as the first or current "Full Order" microphone 305, m4 identified as the second or new "Full Order" microphone 409, and microphone m3 identified as the "FOA" 307 microphone.
[0158] Since the frequency domain transform STFT requires a certain number of time domain audio frames (for example, 6 in afSTFT) as input to generate meaningful output values, and since the STFT has not yet been run for all channels of the new closest source, it is necessary to delay the start of the crossfade operation until the correct output of all channels of the new closest source is available. To this end, when it is determined that a crossfade is required, a counter is added that is incremented for each audio frame, and once it reaches the number of frames required for the correct STFT output, the crossfade is started. During the period of waiting for the crossfade to start, the signal is interpolated using the previous closest source.
[0159] When the crossfade is complete, normal operation is resumed, where the new closest HOA source will be marked as "FullOrder" and the other two HOA sources in the triangle where the listener is located are marked as "FOA". This can be done by Figure 5 , where completion of the cross-fade operation results in microphone m2 being identified as the “FOA” microphone 505 , m4 being identified as the “Full Order” microphone 409 , and microphone m3 being identified as the “FOA” microphone 307 .
[0160] As described above, the mixer 113 is configured to receive or obtain interpolated spatial metadata at the listener position. 134 136 138 132 and the prototype signal S(i,j,k,b) 114 as input. Thus, at this stage we have a description of the sound field at the listener's position (interpolated spatial metadata) and a binaural signal that is an approximation of the output we want (the signal of the HOA source closest to the listener, which has been EQed based on the energy of the interpolated signal). To obtain the final output, the mixing stage creates a binaural signal from the interpolated signal so that it has the same characteristics as the interpolated metadata. For this purpose, an optimal mixing algorithm is used, which can be described as in Vilkamo, J., The present invention is implemented by the method described in T. and Kuntz, A., “Optimized covariance domain framework for time-frequency processing of spatial audio” (Journal of the Audio Engineering Society, Vol. 61, No. 6, pp. 403-411, 2013), and is summarized as follows.
[0161] First, the prototype binaural signal B(j,b) is calculated from the interpolated signal:
[0162]
[0163] Among them, R sh (j) is the spherical harmonic function rotation matrix calculated based on the listener's head and sound source orientation, and M HOA2bin (b) is the Ambisonics to binaural matrix for frequency bin b.
[0164] Compute the covariance matrix from the prototype signal:
[0165] C x (j,b)=B(j,b)B(j,b) H
[0166] The target covariance matrix is calculated from the interpolated metadata and the HRTF function calculated in the preprocessing step. The direct part of the covariance matrix:
[0167]
[0168] Diffusion part:
[0169]
[0170] in,
[0171]
[0172] And the final target covariance matrix:
[0173]
[0174] Then, the mixing matrix is obtained through the optimal mixing algorithm [MPHOA_optimal_mixing]. The output of the optimal mixing algorithm is the mixing matrix ( and ), which when applied to the binaural prototype signal will produce the output binaural signal O(j,k,b), where the covariance matrix is equal to C y :
[0175]
[0176] where D(j,b) is the decorrelated time-frequency domain signal obtained from the buffer of the previous binaural signal B.
[0177] Furthermore, the output processor 115 is configured to perform an inverse frequency domain transform (eg, inverse STFT) on the output frequency domain signal O(j,k,b) 142 to provide a final time domain output signal 144 .
[0178] The computational complexity savings that can be achieved using this approach are due to the amount of STFT processing required per frame. For example, conventional frequency domain (STFT) would be performed for (3 x 16 =) 48 channels (assuming a 3rd order HOA input signal). In some embodiments, frequency domain processing (STFT) is performed for (4 + 4 + 16 =) 24 channels, thereby reducing the STFT computation by half. For a 4th order HOA input signal, the savings are even greater (75 channels vs 33 channels).
[0179] Furthermore, in some embodiments, during a crossfade, a "FOA" STFT can be calculated instead of calculating a "Full Order" STFT for the HOA source during which the crossfade is performed. In these embodiments, signal interpolation is performed on the first-order signal during the crossfade. This will result in a decrease in quality for the duration of the crossfade, but will have the advantage of not delaying the start of the crossfade.
[0180] In an embodiment, the number of channels for which to run the STFT may be determined as follows.
[0181] In some embodiments, the HOA format signal is converted into the time-frequency domain using an alias-free short-time Fourier transform. For each subframe k of audio frame j and for each HOA source i (which is a triangular track recording (TRR, see Section 6.6.18.3.2.3) s HOA (i, j, k) and use the afSTFT forward transform to determine the time-frequency domain signal matrix S(i, j, k). S(i, j, k) is N ch ×N b A matrix of length N containing the values for each HOA channel b For the HOA source that has been selected for signal interpolation, afSTFT is run for all channels, and S(i,j,k) is N ch ×N b A matrix of length N containing the values for each HOA channel b For all other HOA sources found in the triangular orbit record, afSTFT is run only for the first four channels (FOA), and S(i,j,k) is 4×N b A matrix of length N containing the first four channels b During crossfade (6.6.18.3.2.5), the HOA source is the target of the crossfade and afSTFT shall also be run for all channels.
[0182] Furthermore, in some embodiments, if the selected HOA source changes, a cross-fade should be initiated. To this end, the fade-in and fade-out weights (w fi (j,k) and w fo (h,k)) and fade-out band B fo (j) Fade-in weight, fade-out weight and fade-out band are used in metadata and signal interpolation during cross-fading (see Sections 6.6.18.3.4.1 and 6.6.18.3.4.2).
[0183]
[0184] get_fade_out_bands() is used to get the fade-out band B fo (j) The size of this list decreases as the crossfade progresses.
[0185]
[0186]
[0187] When crossFadeProgressIndx == cfLen, cross fading is stopped by setting crossFadeInProgress = False.
[0188] about Figure 6 , a flowchart showing the method steps of some embodiments:
[0189] For example, Figure 6 In the example, the listener position is first obtained, as shown in 601. The renderer receives this information from the listener position and orientation interface in audio scene coordinates.
[0190] Then, if Figure 6 In the example, as shown by 603, based on the listener's location, the active HOA sources are obtained. This can be evaluated periodically, in other words, for each scene state update. For the HOA sources in which the user is located, all HOA sources are classified as active.
[0191] In addition, if Figure 6 As shown in 605, the previous HOA source is obtained for signal interpolation and spatial metadata calculation during the previous scene state update or audio frame.
[0192] After the above operations have been performed, if Figure 6 As shown in 607, the HOA source for signal interpolation is determined and the spatial metadata is calculated based on the proximity of the HOA source to the listener. For example, the HOA source closest to the listener is selected because it has the most representative audio compared to the listener's position.
[0193] like Figure 6 As shown by 609, if the previous signal interpolation HOA source is different from the HOA source determined for signal interpolation in the current audio frame time or scene state update, in other words, the signal interpolation HOA source changes:
[0194] like Figure 6 As shown in 621, the new signal interpolation HOA source (HOA Current-Sigint ) STFT calculation for all channels of the HOA source while continuing with separate STFT calculation for the first-order channels of the HOA source; and
[0195] like Figure 6 As shown in 623, a separate STFT calculation is initiated for the first order channel of the previous signal interpolated HOA source.
[0196] like Figure 6 As shown by 611, the double STFT calculation for the previous signal interpolation HOA source and the new signal interpolation HOA source is continued.
[0197] In addition, if Figure 6 As shown by 613, there is a delay in switching the signal interpolation HOA source until a meaningful output is available from the new STFT processing.
[0198] Then, if Figure 6 As shown by 615 , a cross-fade is performed between the previous signal interpolated HOA source and the new signal interpolated HOA source.
[0199] Finally, if Figure 6 As shown by 617, after a predetermined number of audio frames, after the STFT is started and ready, the HOA source (HOA) for the previous and new signals is stopped. Prev-Sigint and HOA Current-Sigint ) is calculated using the first four channels of the STFT.
[0200] exist Figure 7 An example system employing some embodiments is described in FIG. This figure illustrates an end-to-end system overview of an audio scene comprising multiple HOA sources rendered in accordance with the present invention. The present invention modifies the rendering that occurs at an MPEG-I audio renderer located in a playback device. The renderer receives a scene description and an audio bitstream and performs rendering accordingly. Whenever a scene comprises multiple HOA sources, Figure 1 The MPHOA processing described in and presented in this invention is performed in an MPEG-I audio renderer.
[0201] about Figure 7, schematically illustrating an example system in which embodiments may be implemented.
[0202] The system may include a content creator 701, which may be implemented on any suitable computer or processing device. The content creator 701 includes an (MPEG-I) encoder 711, which is configured to receive an audio scene description 700 and an audio signal or data 702. The audio scene description 700 may be provided in an MPEG-I encoder input format (EIF) or in other suitable formats. Typically, the audio scene description includes an acoustically relevant description of the content of the audio scene, and may include, for example, scene geometry (such as a grid or voxels), acoustic materials, an acoustic environment with reverberation parameters, sound source locations, and other audio element related parameters (such as whether reverberation is to be rendered for the audio element). The MPEG-I encoder 711 is configured to output coded data 712.
[0203] Furthermore, in some embodiments, the content creator 701 includes a bitstream encoder 713 that is configured to receive the output 712 of the MPEG-I encoder 711 and the encoded audio signal from the MPEG-H encoder 619 and generate a bitstream 714 that can be passed to the bitstream decoder 631. In some embodiments, the bitstream 714 can be streamed to an end-user device or made available for download or storage.
[0204] Additionally, the system includes a server configured to obtain the bitstream 714, store it, and provide it to the player 705. In some embodiments, this is achieved by a streaming server 721 configured to provide audio data 722 and MPEG-I Audio 6DoF metadata bitstream 724.
[0205] The player 705 retrieves the associated bitstream 724 and audio data 722. In some embodiments, other implementation options such as broadcast, multicast, etc. are also possible.
[0206] In some embodiments, the player 705 includes a playback device 731 that is configured to obtain or receive audio data 722 and an MPEG-I audio 6DoF metadata bitstream 724, and may also be configured to receive or otherwise obtain 6DoF tracking information (listener orientation or position information) 734 from a suitable listener user interface, such as from a head-mounted device (HMD) 741. This may be generated, for example, by sensors within the HMD 741, or from sensors that sense the listener's orientation or position in the environment.
[0207] In some embodiments, the playback device 731 includes a bitstream parser 733 that is configured to obtain the encoded metadata bitstream 724 and decode it in a reverse or inverse operation to the bitstream encoder 713 and the MPEG-I encoder 711 to generate audio scene description information 732, which can be passed to the MPEG-I audio renderer 735.
[0208] In some embodiments, the playback device 731 includes an MPEG-I audio renderer 735 that is configured to implement the rendering operations described above and generate an audio output signal that can be output to the head mounted device 741 .
[0209] The playback device 731 can be implemented in different form factors depending on the application. In some embodiments, the playback device is equipped with its own listener position tracking device, or receives listener position information from an external device. In some embodiments, the playback device can also be equipped with a headphone connector to deliver the rendered binaural audio output to headphones.
[0210] about Figure 8 , shows an example electronic device that can be used as a computer, an encoder processor, a decoder processor, or any functional block described herein. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1600 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0211] In some embodiments, device 1600 includes at least one processor or central processing unit 1607. Processor 1607 may be configured to execute various program codes, such as the methods described herein.
[0212] In some embodiments, the device 1600 includes a memory 1611. In some embodiments, at least one processor 1607 is coupled to the memory 1611. The memory 1611 can be any suitable storage component. In some embodiments, the memory 1611 includes a program code portion for storing program code that can be implemented on the processor 1607. In addition, in some embodiments, the memory 1611 can also include a stored data portion for storing data (e.g., data that has been processed or will be processed according to the embodiments described herein). The implemented program code stored in the program code portion and the data stored in the stored data portion can be retrieved by the processor 1607 via the memory-processor coupling when needed.
[0213] In some embodiments, device 1600 includes a user interface 1605. In some embodiments, user interface 1605 can be coupled to processor 1607. In some embodiments, processor 1607 can control the operation of user interface 1605 and receive input from user interface 1605. In some embodiments, user interface 1605 can enable a user to enter commands to device 1600, for example, via a keyboard. In some embodiments, user interface 1605 can enable a user to obtain information from device 1600. For example, user interface 1605 can include a display configured to display information from device 1600 to the user. In some embodiments, user interface 1605 can include a touch screen or touch interface that can enable information to be entered into device 1600 and can also display information to a user of device 1600.
[0214] In some embodiments, device 1600 includes input / output port 1609. In some embodiments, input / output port 1609 comprises a transceiver. In such embodiments, the transceiver may be coupled to processor 1607 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components may be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.
[0215] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol (such as, for example, IEEE 802.x), a suitable short-range radio frequency communication protocol (such as Bluetooth), or an infrared data communication path (IRDA).
[0216] The transceiver input / output port 1609 may be configured to transmit / receive audio signals and bitstreams, and in some embodiments, perform the operations and methods described above by executing appropriate code using the processor 1607 .
[0217] In general, various embodiments of the present invention may be implemented using hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented using hardware, while other aspects may be implemented using firmware or software that can be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein may be implemented using hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof, as non-limiting examples.
[0218] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. In addition, it should be noted in this regard that any block of the logic flow as in the accompanying drawings may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media, and optical media.
[0219] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and a processor.
[0220] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0221] Programs, such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design, Inc. of San Jose, Calif., use well-established design rules and a library of pre-stored design blocks to automatically route conductors and position components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design, in a standardized electronic format (e.g., Opus, GDSII, etc.), can be transferred to a semiconductor fabrication facility, or "fab," for fabrication.
[0222] The foregoing description has provided by way of exemplary and non-limiting examples a complete and informative description of the exemplary embodiments of the present invention. However, various modifications and adaptations will become apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar modifications of the teachings of this invention will still fall within the scope of the invention as defined by the appended claims.
Claims
1. A method for generating spatialized audio output, comprising: Obtain at least two higher-order panoramic surround sound audio sources, wherein the at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within the audio environment; obtaining a listener position within the audio environment, wherein the listener position is free to move within the audio environment; determining at least one source as an active source; determining at least one current higher-order panoramic surround sound audio source from the at least one determined active source based on the current listener position; determining at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; Performing processing on all channel signals of the at least one current high-order panoramic surround sound audio source; performing processing on at least one channel signal of other activity sources among the at least one determined activity source; determining a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continuing processing of all channel signals of the at least one previously active higher-order panoramic surround sound audio source; processing at least one channel signal of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continuing processing of at least one channel signal of the other activity source among the at least one determined activity source; processing all channel signals of at least one current high-order panoramic surround sound audio source among the current high-order panoramic surround sound audio sources; Perform signal interpolation based on: continuing processing of a channel signal of at least one of the at least one previous higher-order panoramic surround sound audio source associated with signal interpolation for a first time period after determining the difference; cross-fading between continued processing of all channel signals of the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; and After the cross-fade ends, stopping processing of all channel signals of the at least one previously active high-order panoramic surround sound audio source; and The spatialized audio output is generated based on the determined at least one signal interpolation.
2. The method according to claim 1, wherein The processing is a short-time Fourier transform.
3. The method according to any one of claims 1 or 2, wherein The crossfade is in the second time period.
4. The method according to claim 1, wherein The first time period is a first number of time frames based on a short time Fourier transform start delay, and the second time period is a second defined number of processing frames.
5. The method according to any one of claims 1 to 4, further comprising: Determine spatial metadata by doing the following: analyzing processing of all channel signals of the at least one current higher-order panoramic surround sound audio source and at least one channel signal of another active source among the at least one determined active source; as well as Continued processing of at least one channel signal of the at least one previously active higher-order panoramic surround sound audio source is analyzed.
6. The method according to claim 5, wherein: Generating the spatialized audio output is further based on the determined spatial metadata.
7. The method according to any one of claims 1 to 6, wherein The determined at least one corresponding channel signal of the at least one previous higher-order panoramic surround sound audio source is a subset of the channels.
8. The method according to any one of claims 1 to 7, wherein The at least one channel signal of the determined at least one previous higher-order panoramic surround sound audio source is first four channel signals of the determined at least one previous higher-order panoramic surround sound audio source.
9. The method according to any one of claims 1 to 7, wherein Determining at least one current higher-order panoramic surround sound audio source from the at least two higher-order panoramic surround sound audio sources based on the listener position includes: Determining an area in which the current listener position is located, the area being defined by vertex positions of at least three higher-order panoramic surround sound audio sources; and The at least one current higher-order panoramic surround sound audio source is selected from the at least two higher-order panoramic surround sound audio sources, the at least one current higher-order panoramic surround sound audio source being a current higher-order panoramic surround sound audio source whose position defines a vertex of the area.
10. An apparatus comprising means for performing the method according to any one of claims 1 to 9.
11. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform the method according to any one of claims 1 to 9.
12. An apparatus for generating spatialized audio output, the apparatus comprising components configured to: At least two high-order panoramic surround sound audio sources are obtained, wherein: The at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within the audio environment; obtaining a listener position within the audio environment, wherein the listener position is free to move within the audio environment; determining at least one source as an active source; determining at least one current higher-order panoramic surround sound audio source from the at least one determined active source based on the current listener position; determining at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; Performing processing on all channel signals of the at least one current high-order panoramic surround sound audio source; performing processing on at least one channel signal of other activity sources among the at least one determined activity source; determining a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continuing processing of all channel signals of the at least one previously active higher-order panoramic surround sound audio source; processing at least one channel signal of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continuing processing of at least one channel signal of the other activity source among the at least one determined activity source; processing all channel signals of at least one current high-order panoramic surround sound audio source among the current high-order panoramic surround sound audio sources; Perform signal interpolation based on: continuing processing of a channel signal of at least one of the at least one previous higher-order panoramic surround sound audio source associated with signal interpolation for a first time period after determining the difference; cross-fading between continued processing of all channel signals of the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; After the cross-fade ends, stopping processing of all channel signals of the at least one previously active high-order panoramic surround sound audio source; and The spatialized audio output is generated based on the determined at least one signal interpolation.
13. An apparatus for generating spatialized audio output, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: At least two high-order panoramic surround sound audio sources are obtained, wherein: The at least two higher-order panoramic surround sound audio sources are associated with corresponding audio source positions within the audio environment; obtaining a listener position within the audio environment, wherein the listener position is free to move within the audio environment; determining at least one source as an active source; determining at least one current higher-order panoramic surround sound audio source from the at least one determined active source based on the current listener position; determining at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; Performing processing on all channel signals of the at least one current high-order panoramic surround sound audio source; performing processing on at least one channel signal of other activity sources among the at least one determined activity source; determining a difference between the at least one current higher-order panoramic surround sound audio source and the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continuing processing of all channel signals of the at least one previously active higher-order panoramic surround sound audio source; processing at least one channel signal of at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation determination; continuing processing of at least one channel signal of the other activity source among the at least one determined activity source; processing all channel signals of at least one current high-order panoramic surround sound audio source among the current high-order panoramic surround sound audio sources; Perform signal interpolation based on: continuing processing of a channel signal of at least one of the at least one previous higher-order panoramic surround sound audio source associated with signal interpolation for a first time period after determining the difference; cross-fading between continued processing of all channel signals of the at least one previous higher-order panoramic surround sound audio source associated with the signal interpolation and processing of all channel signals of at least one current higher-order panoramic surround sound audio source determined from the at least two higher-order panoramic surround sound audio sources after the first time period after determining the difference; After the cross-fade ends, stopping processing of all channel signals of the at least one previously active high-order panoramic surround sound audio source; and The spatialized audio output is generated based on the determined at least one signal interpolation.
14. The device according to claim 13, wherein The processing is a short-time Fourier transform.
15. The device according to any one of claims 13 or 14, wherein The crossfade is in the second time period.
16. The device according to claim 13, wherein The first time period is a first number of time frames based on a short time Fourier transform start delay, and the second time period is a second defined number of processing frames.
17. The apparatus according to any one of claims 13 to 16, further configured to determine the spatial metadata by: analyzing processing of all channel signals of the at least one current higher-order panoramic surround sound audio source and at least one channel signal of another active source among the at least one determined active source; and Continued processing of at least one channel signal of the at least one previously active higher-order panoramic surround sound audio source is analyzed.
18. The device according to claim 17, wherein Causing the apparatus to generate the spatialized audio output is further based on the determined spatial metadata.
19. The device according to any one of claims 13 to 18, wherein The determined at least one corresponding channel signal of the at least one previous higher-order panoramic surround sound audio source is a subset of the channels.
20. The device according to any one of claims 13 to 19, wherein The at least one channel signal of the determined at least one previous higher-order panoramic surround sound audio source is first four channel signals of the determined at least one previous higher-order panoramic surround sound audio source.
21. The device according to any one of claims 13 to 19, wherein The device is configured to determine, based on the listener position, at least one current higher-order panoramic surround sound audio source from the at least two higher-order panoramic surround sound audio sources. The device is configured to: Determining an area in which the current listener position is located, the area being defined by vertex positions of at least three higher-order panoramic surround sound audio sources; as well as The at least one current higher-order panoramic surround sound audio source is selected from the at least two higher-order panoramic surround sound audio sources, the at least one current higher-order panoramic surround sound audio source being a current higher-order panoramic surround sound audio source whose position defines a vertex of the area.
Citation Information
Patent Citations
Hydraulic drive system
GB202007710D0