6DOF rendering of audio captured by a microphone array at a location outside of the microphone array
By using at least two microphone arrays to capture audio signals and perform interpolation processing, the problem of 6-DOF audio rendering outside the microphone array area is solved, improving the user experience and providing spatial navigation prompts.
Patent Information
- Application Number
- CN202211224290.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-08
- Filing Date
- 2022-10-08
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-10-08
AI Technical Summary
Existing technologies struggle to provide high-quality 6-DOF audio rendering outside of microphone arrays, resulting in a poor listening experience when moving outside the microphone array area and a lack of spatial navigation cues.
By capturing audio signals using at least two microphone arrays, determining the listener's location, and interpolating based on metadata, a modified audio signal and spatial metadata are generated to provide reasonable audio rendering outside the microphone array area.
It enables reasonable audio rendering at the listener's location outside the microphone array area, improves the user experience, provides spatial navigation cues, and avoids orientation errors and confusion.
Smart Images

Figure CN115955622B_ABST
Abstract
Description
Technical Field
[0001] This application relates to an apparatus and method for audio rendering using a 6-DOF system for audio captured by a microphone array located outside the microphone array. Background Technology
[0002] Spatial audio capture methods attempt to capture the audio environment so that it can be effectively recreated perceptibly for the listener, and further, allow the listener to move and / or rotate within the recreated audio environment. For example, in some systems (3DoF, 3 degrees of freedom), the listener can rotate their head, and the rendered audio signal reflects this rotational motion. In some systems (3DoF+, 3DoF+), the listener can slightly “move” within the environment and rotate their head, while in other systems (6DoF, 6 degrees of freedom), the listener can freely move and rotate their head within the environment.
[0003] Linear spatial audio capture refers to an audio capture method in which the processing is not adapted to the characteristics of the captured audio. Instead, the output is a predetermined linear combination of the captured audio signals.
[0004] To linearly record spatial sound at a location within a recording space, a high-end microphone array is required. One such microphone is the Eigenmike 32-spherical microphone. High-order Ambisonics (HOA) signals can be obtained from a high-end microphone array and used for linear rendering. Using HOA signals, spatial audio can be linearly rendered so that sounds arriving from different directions are satisfactorily separated within a reasonable audible bandwidth.
[0005] The problem with linear spatial audio capture technology is the requirements for the microphone array. Short wavelengths (higher frequency audio signals) require small microphone spacing, while long wavelengths (lower frequencies) require large array size, and it is difficult to meet both conditions within a single microphone array.
[0006] Most practical capture devices (e.g., virtual reality cameras, SLR cameras, mobile phones) are not equipped with microphone arrays like those provided by Eigenmike, and there is insufficient microphone arrangement for linear spatial audio capture. Furthermore, achieving linear spatial audio capture for a capture device results in spatial audio being obtained only for a single location.
[0007] Parametric spatial audio capture refers to a system that estimates perceptually relevant parameters based on audio signals captured by a microphone, and that spatial sound can be synthesized based on these parameters and the audio signals. Analysis and synthesis typically occur in frequency bands approaching the resolution of human spatial hearing.
[0008] As is well known, for most compact microphone setups (e.g., VR cameras, multi-microphone arrays, mobile phones with microphones, SLR cameras with microphones), parametric spatial audio capture can produce perceptually accurate spatial audio rendering, while linear methods typically do not produce feasible results in terms of the spatial aspects of sound. For high-end microphone arrays, such as Eigenmike, parametric methods can further provide spatial sound perception with an average quality better than linear methods. Summary of the Invention
[0009] According to a first aspect, an apparatus is provided, comprising components configured to perform the following operations: obtaining two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; obtaining a listener location within an audio environment, wherein the audio environment includes one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and, for at least two of the two or more audio signal sets, based on the at least two audio signal sets in the two or more audio signal sets... The process involves processing two audio signals to obtain metadata; determining a second listener position for a listener position in an audio environment outside an inner region, wherein the second listener position is located in an outer region and closer to, or on, the boundary of one or more inner and outer regions, or within one or more inner regions; determining modified metadata for the second listener position based on the metadata; determining at least two modified audio signals for the second listener position based on at least two audio signals; determining spatial metadata for the listener position based on the modified metadata for the second listener position; and outputting at least two modified audio signals and spatial metadata.
[0010] A component configured to determine spatial metadata for a listener location based on modified metadata for a second listener location can be configured to: determine at least one audio location relative to the second listener location based on the modified metadata for the second listener location, wherein the modified metadata for the second listener location includes a direction parameter representing the direction from the second listener location to one of the at least one audio location; and determine spatial metadata for the listener location based on at least one audio signal set location relative to the second listener location, wherein the spatial metadata includes a spatial direction parameter representing the direction from the listener location to the at least one of the at least one audio location.
[0011] A component configured to obtain two or more audio signal sets can be configured to obtain two or more audio signal sets from a microphone arrangement, wherein each microphone arrangement may be at a corresponding location and includes one or more microphones.
[0012] A component configured to obtain the listener's location can be configured to obtain the listener's location from another device.
[0013] A component configured to obtain metadata based on the processing of at least two audio signals from at least two audio signal sets in two or more audio signal sets can be configured to: determine orientation parameters based on the processing of at least two audio signals.
[0014] The component configured to determine a second listener position for a listener position in an audio environment outside an internal region can be configured to determine the second listener position at one of the following locations: within a plane or volume at least partially defined by an edge or surface connecting two or more audio signal set positions and the listener position; within a plane or volume within an associated internal region at least partially defined by an edge or surface connecting two or more audio signal set positions; on an edge or surface defined by two or more audio signal set positions; and at the nearest audio signal set position among the two or more audio signal set positions.
[0015] The component configured to determine modified metadata for the second listener location based on metadata can be configured to: generate at least two interpolation weights based on the audio signal set location and the second listener location; apply the at least two interpolation weights to the corresponding audio signal set audio metadata to generate interpolated audio metadata; and combine the interpolated audio metadata to generate modified metadata for the second listener location.
[0016] The component configured to determine the spatial metadata for the listener location based on the modified metadata for the second listener location can be configured to map the modified metadata to a Cartesian coordinate system based on the second listener location.
[0017] A component configured to determine at least two modified audio signals for a second listener's location based on at least two audio signals can be configured to generate an interpolated audio signal from the at least two audio signals.
[0018] A component configured to determine spatial metadata for a listener position based on at least one audio position relative to a second listener position, wherein the spatial metadata includes a component representing a spatial direction parameter indicating the direction from the listener position to one of the at least one audio positions, can be configured to determine the spatial direction parameter based on one of the following: an interpolation difference between the at least one audio position relative to the second listener position and the listener position; and a difference between the listener position and the at least one audio position relative to the second listener position.
[0019] A component configured to determine spatial metadata for a listener's location based on modified metadata for a second listener's location can be configured to modify at least one direct-to-total energy ratio based on the difference between at least one audio location relative to the second listener's location and the listener's location.
[0020] The component can be further configured to process at least two modified audio signals to generate spatial audio output based on spatial metadata for the listener's location.
[0021] The component configured to generate spatial audio output can be configured to generate at least one of the following: binaural audio output, which includes two audio signals for headphones and / or earphones; Ambisonic audio output, which includes multiple audio signals for headphones or an Ambisonic renderer for a multi-channel speaker group; and multi-channel audio output, which includes at least two audio signals for a multi-channel speaker group.
[0022] According to a second aspect, a method is provided for an apparatus for generating spatialized audio output based on a listener's location, the method comprising: obtaining two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; obtaining a listener's location within an audio environment, wherein the audio environment includes one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and, for at least two of the two or more audio signal sets, generating spatialized audio output based on at least two audio signals in the two or more audio signal sets. Processing at least two audio signals from a signal set to obtain metadata; determining a second listener position for a listener position in an audio environment outside an inner region, the second listener position being located in an outer region and closer to, or on, the boundary of one or more inner and outer regions, or within one or more inner regions; determining modified metadata for the second listener position based on the metadata; determining at least two modified audio signals for the second listener position based on the at least two audio signals; determining spatial metadata for the listener position based on the modified metadata for the second listener position; and outputting at least two modified audio signals and spatial metadata.
[0023] Determining spatial metadata for a listener location based on modified metadata for a second listener location may include: determining at least one audio location relative to the second listener location based on the modified metadata for the second listener location, wherein the modified metadata for the second listener location includes a direction parameter representing the direction from the second listener location to one of the at least one audio location; and determining spatial metadata for a listener location based on at least one audio signal set location relative to the second listener location, wherein the spatial metadata includes a spatial direction parameter representing the direction from the listener location to the at least one of the at least one audio location.
[0024] Obtaining two or more audio signal sets may include obtaining two or more audio signal sets from a microphone arrangement, wherein each microphone arrangement may be at a corresponding location and includes one or more microphones.
[0025] Obtaining the listener's location can include obtaining the listener's location from another device.
[0026] For at least two audio signal sets in two or more audio signal sets, obtaining metadata based on the processing of at least two audio signals in the two or more audio signal sets may include: determining orientation parameters based on the processing of at least two audio signals.
[0027] Determining a second listener location for a listener location in an audio environment outside an internal region may include determining the second listener location at one of the following locations: within a plane or volume at least partially defined by an edge or surface connecting two or more audio signal set locations and the listener location; within a plane or volume within an associated internal region at least partially defined by an edge or surface connecting two or more audio signal set locations; on an edge or surface defined by two or more audio signal set locations; and at the nearest audio signal set location among the two or more audio signal set locations.
[0028] Determining modified metadata for the second listener's location based on metadata may include: generating at least two interpolation weights based on the audio signal set location and the second listener's location; applying the at least two interpolation weights to the corresponding audio signal set audio metadata to generate interpolated audio metadata; and combining the interpolated audio metadata to generate modified metadata for the second listener's location.
[0029] Determining the spatial metadata for the listener location based on the modified metadata for the second listener location may include mapping the modified metadata to a Cartesian coordinate system based on the second listener location.
[0030] Determining at least two modified audio signals for the second listener's location based on at least two audio signals may include generating an interpolated audio signal from the at least two audio signals.
[0031] Spatial metadata for a listener position is determined based on at least one audio position relative to a second listener position, wherein the spatial metadata includes a spatial direction parameter representing the direction from the listener position to one of the at least one audio positions. The spatial direction parameter may be determined based on one of the following: an interpolation difference between the at least one audio position relative to the second listener position and the listener position; and the difference between the listener position and the at least one audio position relative to the second listener position.
[0032] Determining the spatial metadata for the listener location based on the modified metadata for the second listener location may include: modifying at least one direct-to-total energy ratio based on the difference between at least one audio location relative to the second listener location and the listener location.
[0033] The method may further include: processing at least two modified audio signals to generate a spatial audio output based on spatial metadata used for the listener's location.
[0034] Generating spatial audio output may include generating at least one of the following: binaural audio output, which includes two audio signals for headphones and / or earbuds; Ambisonic audio output, which includes multiple audio signals for headphones or an Ambisonic renderer for a multi-channel speaker group; and multi-channel audio output, which includes at least two audio signals for a multi-channel speaker group.
[0035] According to a third aspect, an apparatus is provided, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, to cause the apparatus to at least: obtain two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; obtain a listener location within an audio environment, wherein the audio environment comprises one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and, for at least two audio signal sets in the two or more audio signal sets, based on the two... Processing at least two audio signals from at least two audio signal sets in one or more audio signal sets to obtain metadata; determining a second listener position for a listener position in an audio environment outside an inner region, the second listener position being located in an outer region and closer to, or on, the boundary of one or more inner and outer regions, or within one or more inner regions; determining modified metadata for the second listener position based on the metadata; determining at least two modified audio signals for the second listener position based on at least two audio signals; determining spatial metadata for the listener position based on the modified metadata for the second listener position; and outputting at least two modified audio signals and spatial metadata.
[0036] An apparatus for determining spatial metadata of a listener's location based on modified metadata for a second listener's location can be made to: determine at least one audio location relative to the second listener's location based on the modified metadata for the second listener's location, wherein the modified metadata for the second listener's location includes a direction parameter representing the direction from the second listener's location to one of the at least one audio location; and determine spatial metadata for the listener's location based on at least one audio signal set location relative to the second listener's location, wherein the spatial metadata includes a spatial direction parameter representing the direction from the listener's location to the at least one of the at least one audio location.
[0037] The means for obtaining two or more audio signal sets can be made to obtain two or more audio signal sets from a microphone arrangement, wherein each microphone arrangement may be at a corresponding position and includes one or more microphones.
[0038] The device that is made to obtain the listener's location can be made to obtain the listener's location from another device.
[0039] An apparatus for obtaining metadata based on processing at least two audio signals from at least two audio signal sets in two or more audio signal sets may be configured to: determine orientation parameters based on processing at least two audio signals.
[0040] The means for determining a second listener position for a listener position in an audio environment outside an internal region can be made to determine the second listener position at one of the following locations: within a plane or volume at least partially defined by an edge or surface connecting two of the two or more audio signal set positions and the listener position; within a plane or volume within an associated internal region at least partially defined by an edge or surface connecting two of the two or more audio signal set positions; on an edge or surface defining two of the two or more audio signal set positions; and at the closest audio signal set position among the two or more audio signal set positions.
[0041] The apparatus for determining modified metadata based on metadata for a second listener's location can be configured to: generate at least two interpolation weights based on the location of the audio signal set and the second listener's location; apply the at least two interpolation weights to the corresponding audio signal set audio metadata to generate interpolated audio metadata; and combine the interpolated audio metadata to generate modified metadata for the second listener's location.
[0042] An apparatus that enables the determination of spatial metadata for a listener's location based on modified metadata for a second listener's location can be made to: map the modified metadata to a Cartesian coordinate system based on the second listener's location.
[0043] An apparatus that enables the determination of at least two modified audio signals for the location of a second listener based on at least two audio signals can be configured to generate an interpolated audio signal from the at least two audio signals.
[0044] A means for determining spatial metadata for a listener position based on at least one audio position relative to a second listener position, wherein the spatial metadata includes a spatial direction parameter representing the direction from the listener position to one of the at least one audio positions, can be made to determine the spatial direction parameter based on one of: an interpolation difference between the at least one audio position relative to the second listener position and the listener position; and a difference between the listener position and the at least one audio position relative to the second listener position.
[0045] The means for determining spatial metadata for a listener's location based on modified metadata for a second listener's location can be made to: modify at least one direct-to-total energy ratio based on the difference between at least one audio location relative to the second listener's location and the listener's location.
[0046] The device can be further configured to process at least two modified audio signals to generate a spatial audio output based on spatial metadata for the listener's location.
[0047] The means for generating spatial audio output can be made to generate at least one of the following: binaural audio output, which includes two audio signals for headphones and / or earphones; Ambisonic audio output, which includes multiple audio signals for an Ambisonic renderer for headphones or a multi-channel speaker group; and multi-channel audio output, which includes at least two audio signals for a multi-channel speaker group.
[0048] According to a fourth aspect, an apparatus is provided, comprising: means for acquiring two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; means for acquiring a listener location within an audio environment, wherein the audio environment includes one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and means for acquiring, based on processing at least two audio signals from at least two of the two or more audio signal sets, a meta-analysis of the audio signal set location. The data includes: components for determining a second listener location for a listener location in an audio environment outside an internal region, the second listener location being located in an external region and closer to, or on, the boundary of one or more internal and external regions, or within one or more internal regions; components for determining modified metadata for the second listener location based on metadata; components for determining at least two modified audio signals for the second listener location based on at least two audio signals; components for determining spatial metadata for the listener location based on the modified metadata for the second listener location; and components for outputting at least two modified audio signals and spatial metadata.
[0049] According to a fifth aspect, a computer program [or a computer-readable medium including program instructions] is provided, the instructions being configured to cause a device to perform at least the following operations: obtaining two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; obtaining a listener location within an audio environment, wherein the audio environment comprises one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and, for at least two of the two or more audio signal sets, based on the two or more audio signal sets... The method involves processing at least two audio signals from at least two audio signal sets to obtain metadata; determining a second listener location for a listener location in an audio environment outside an inner region, wherein the second listener location is located in an outer region and closer to, or on, the boundary of one or more inner and outer regions, or within one or more inner regions; determining modified metadata for the second listener location based on the metadata; determining at least two modified audio signals for the second listener location based on the at least two audio signals; determining spatial metadata for the listener location based on the modified metadata for the second listener location; and outputting at least two modified audio signals and spatial metadata.
[0050] According to a sixth aspect, a non-transitory computer-readable medium is provided comprising program instructions for causing a device to perform at least the following operations: obtaining two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; obtaining a listener location within an audio environment, wherein the audio environment comprises one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and, for at least two of the two or more audio signal sets, based on at least one of the two or more audio signal sets... Processing at least two audio signals from two audio signal sets to obtain metadata; determining a second listener position for a listener position in an audio environment outside an inner region, the second listener position being located in an outer region and closer to, or on, the boundary of one or more inner and outer regions, or within one or more inner regions; determining modified metadata for the second listener position based on the metadata; determining at least two modified audio signals for the second listener position based on at least two audio signals; determining spatial metadata for the listener position based on the modified metadata for the second listener position; and outputting at least two modified audio signals and spatial metadata.
[0051] According to a seventh aspect, an apparatus is provided comprising: acquisition circuitry configured to acquire two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; acquisition circuitry configured to acquire a listener location within an audio environment, wherein the audio environment includes one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and acquisition of metadata based on processing of at least two audio signals from at least two of the two or more audio signal sets. The circuitry includes: a determining circuitry configured to determine a second listener position for a listener position in an audio environment outside an internal region, the second listener position being located in an external region and closer to, or on, the boundary of one or more internal and external regions, or within one or more internal regions; a determining circuitry configured to determine modified metadata for the second listener position based on metadata; a determining circuitry configured to determine at least two modified audio signals for the second listener position based on at least two audio signals; a determining circuitry configured to determine spatial metadata for the listener position based on the modified metadata for the second listener position; and an output circuitry configured to output the at least two modified audio signals and the spatial metadata.
[0052] According to an eighth aspect, a computer-readable medium including program instructions is provided for causing a device to perform at least the following operations: obtaining two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; obtaining a listener location within an audio environment, wherein the audio environment includes one or more regions having one or more internal and external regions relative to the corresponding audio signal set location, wherein the internal region is defined by the corresponding audio signal set location; and, for at least two audio signal sets in the two or more audio signal sets, based on at least two audio signal sets in the two or more audio signal sets... Processing at least two audio signals from a signal set to obtain metadata; determining a second listener position for a listener position in an audio environment outside an inner region, the second listener position being located in an outer region and closer to, or on, the boundary of one or more inner and outer regions, or within one or more inner regions; determining modified metadata for the second listener position based on the metadata; determining at least two modified audio signals for the second listener position based on the at least two audio signals; determining spatial metadata for the listener position based on the modified metadata for the second listener position; and outputting at least two modified audio signals and spatial metadata.
[0053] An apparatus comprising components for performing the actions described above.
[0054] An apparatus configured to perform the actions of the above-described method.
[0055] A computer program comprising program instructions for causing a computer to perform the methods described above.
[0056] A computer program product stored on a medium can enable a device to perform the methods described herein.
[0057] An electronic device may include the means as described herein.
[0058] A chipset may include devices as described herein.
[0059] The embodiments of this application are intended to solve problems associated with the prior art. Attached Figure Description
[0060] To better understand this application, reference will now be made to the accompanying drawings by way of example, wherein:
[0061] Figure 1 The illustration schematically depicts a device system that captures and reproduces an example sound scene, and allows a user to move within the reproduced scene.
[0062] Figure 2 An example reproduction of an audio scene is shown schematically, in which the user moves outside the area defined by the microphone array;
[0063] Figure 3 An example planar microphone array arrangement is schematically shown in which a user can move in and out of an area defined by the microphone array;
[0064] Figure 4 The illustration schematically depicts an apparatus, according to some embodiments, suitable for rendering audio signals to a user capable of moving within and outside an area defined by a microphone array;
[0065] Figure 5 Illustrations based on some embodiments Figure 4 A flowchart illustrating the operation of the device shown;
[0066] Figure 6 The listener positions for example edge rendering and vertex rendering scenes are schematically shown according to some embodiments;
[0067] Figure 7 An example area is schematically shown, covered by an example edge-rendered scene and a vertex-rendered scene according to some embodiments;
[0068] Figure 8 The diagram illustrates the determination of the normal vector for the listener's position, for example, edge rendering, according to some embodiments.
[0069] Figure 9 The interpolation of the original and projected parameters for listener position for example edge rendering is illustrated according to some embodiments.
[0070] Figure 10 An example normal for omitted edges in a non-convex shape arrangement of a microphone array is schematically shown according to some embodiments;
[0071] Figure 11 An example edge / vertex selection for a non-convex shape arrangement of a microphone array is schematically shown according to some embodiments;
[0072] Figures 12a to 12c Example scenarios according to some embodiments are shown respectively, wherein the user is in an area defined by the microphone array, the user is outside an area defined by the microphone array, and the user is outside an area defined by the microphone array;
[0073] Figure 13 An apparatus suitable for implementing some embodiments is shown, wherein the capturing means can be separated from the rendering means elements;
[0074] Figure 14 Suitable means for implementing some embodiments are schematically shown; and
[0075] Figure 15 An example device suitable for implementing the illustrated apparatus is shown schematically. Detailed Implementation
[0076] The concepts discussed in further detail herein relate to the rendering of audio scenes, where an audio scene (or, in other words, a set of audio signals captured at corresponding signal locations in the recording space) is captured using a parametric spatial audio method and two or more microphone arrays corresponding to different locations in the recording space. Furthermore, this concept relates to the rendering of audio scenes where the user (or listener) can move to different locations both within and outside the area defined by the microphone arrays.
[0077] 6DoF is currently common in virtual reality (such as VR games), where motion in the audio scene is easy to render because all spatial information is readily available (i.e., the location of each sound source and the audio signal of each sound source).
[0078] In the following examples, an audio signal set is generated by a microphone (or microphone array). For example, a microphone arrangement may include one or more microphones and generate one or more audio signals for the audio signal set. In some embodiments, the audio signal set includes audio signals that are virtual or generated audio signals (e.g., virtual speaker audio signals with associated virtual speaker locations). In some embodiments, the microphone array is further separated from or physically located away from any processing device; however, this does not preclude examples where the microphones are located on or physically connected to the processing device.
[0079] Before discussing this concept in more detail, some aspects of spatial capture and reproduction will first be described in more detail. For example, Figure 1 An example of spatial capture and playback is shown. Therefore, for example... Figure 1 The spatial audio signal capture environment is shown on the left. The environment or audio scene includes sound sources, namely source 1102 and source 2104, which can be actual audio signal sources or abstract representations of sounds or audio sources. In other words, a sound source can represent an actual sound source, such as a musical instrument, or an abstract sound source, such as the distributed sound of wind blowing through trees. Furthermore, Figure 1 Section 106, representing a non-directional or non-location-specific environment for an audio scene, is shown. These can be captured by at least two microphone arrangements / arrays, each of which may include two or more microphones.
[0080] Audio signals can be captured as described above, and can further be encoded, transmitted, received, and reproduced, such as... Figure 1 As shown by arrow 110 in the image.
[0081] exist Figure 1 The right side shows an example reproduction. The reproduction of the spatial audio signal results in presenting a reproduced audio environment to user 150 (who is shown in this example wearing a head-tracking headset) in the form of a 6DoF spatial rendering 118, which includes a perceived source 1 112 (which is a copy of source 1102), a perceived source 2 114 (which is a copy of source 2 104), and a perceived surrounding environment 116 (which is a copy of surrounding environment 106).
[0082] Traditionally, parametric capture methods have been used only for single-point reproduction, but recently, a 6DoF reproduction method that allows for free movement has been proposed. The method proposed in UK patent application GB2002710.8 uses at least two microphone arrays and analyzes spatial metadata for each array (to determine parameters for more than one frequency band, such as orientation and energy ratio). At the renderer, 6DoF audio is rendered using the microphone array signals and spatial metadata based on the listener's position and orientation.
[0083] The method proposed in GB2002710.8 can be used for, for example Figure 1 In the scenario shown, the audio scene can be captured using a relatively small number of microphone arrays (e.g., six arrays), and the listener can move freely within the space. Furthermore, the method employed is completely blind, meaning no information about the source location is required.
[0084] However, while this method can be used when the listener is able to move within the area spanned by the microphone array, the consistency of audio spatialization can be significantly degraded when the listener moves outside that area.
[0085] As proposed by the method shown in GB2002710.8, for locations outside the area spanned by the microphone array, a rendering is generated based on the location determined by projecting the listener onto the nearest edge of the area spanned by the microphone array.
[0086] In the following discussion, the terms location and position are used interchangeably.
[0087] Therefore, if the sound source is located within the area spanned by the microphone array, this can produce relatively high-quality audio rendering when the listener moves outside that area, because the projection onto the edge maintains the sound source position relative to the correct side of the listener, although the exact orientation may be slightly off.
[0088] However, if the sound source is located outside the area defined by the microphone array, the mentioned method can produce significant orientation errors.
[0089] This situation is as follows Figure 2 As shown. In this example, as illustrated in the diagram on the left (201), the listener is located at a first position 209 at the outer edge of the area defined by microphone arrays 203, 205, and 207, and the sound source 213 is located outside this area. Figure 2 As shown on the right side 251, if the listener moves from the first position 209 to the second position 257 and passes the position of the source 213, where the second position 257 is away from the area defined by the microphone arrays 203, 205, 207, then even after the listener moves past the source, the perceived sound source maintains the same orientation relative to the listener, because the rendering is based on the projected position (and the direction indicated by the arrow marker 261, which is a straight line from the earlier listener position 209 to the source 213).
[0090] This can lead to a confusing experience for listeners because they cannot perceive the actual source direction (and the perceived source direction is incorrect). Furthermore, when a listener moves away from the area of the microphone array, any movement within that area results in spatial audio rendering corresponding to the user moving at the edge of the area defined by the microphone array. Therefore, there is no auditory cue to help the listener navigate back to the main listening area (i.e., the area defined by the microphone array position).
[0091] The methods discussed above suggest making the rendered directionality worse outside the area spanned by the microphone array. This prevents the sound source from being perceived as being in a completely incorrect direction, since the sound source is rendered with a "fuzzy" direction when outside the area. However, this can still confuse the listener, as the listener can no longer navigate solely by the sound source and may not be able to navigate back to the main listening area without assistance.
[0092] Therefore, 6DoF rendering outside the area spanned by the microphone array suffers from significant orientation errors, resulting in a poor user experience where the user perceives the sound source location incorrectly and does not receive spatial cues to be able to perceive that the area spanned by the microphone array will be able to return there.
[0093] Therefore, the embodiments described herein relate to 6-DOF (i.e., the listener can move within the scene and the listener's position is tracked) binaural (and other spatial output formats) rendering of audio captured at a known location using at least two microphone arrays, wherein apparatus and methods are described to provide spatially plausible binaural (and other spatial output formats) audio rendering for listening locations outside the area spanned by the microphone arrays.
[0094] As described in this article, this can be achieved in the following ways:
[0095] Determine the user's location relative to the area defined by the microphone array;
[0096] Based on the user's location and audio captured using at least two microphone arrays, determine the directional parameters (spatial metadata);
[0097] When the user's location is determined to be outside the area defined by the microphone array, identify or select a microphone (and its associated parameters) corresponding to the user's location and orientation parameters;
[0098] A single set of parameters is determined using the parameters associated with the selected microphone;
[0099] Modified (orientation) parameters are obtained by applying spatial modification rules to (orientation) parameters to modify the value of at least one parameter by at least one amount. This amount may depend on the location of the determined position relative to the area defined by the microphone array (e.g., modifying additional orientation parameters corresponding to positions outside the area defined by the microphone array); and
[0100] Based on modified directional parameters and audio signals from one or more microphone arrays, render spatial audio signals (e.g., binaural audio signals).
[0101] The term "spatially plausible binaural audio rendering" can be understood as sound sources within the area spanned by the microphone array (at the listening position) being rendered as "points" from roughly the correct direction, thus allowing navigation into that area. Since the location of the sources is assumed to be unknown, sound sources outside the area are rendered in a way that does not conflict with spatial cues from sources within the area, thus avoiding confusion and aiding navigation. Furthermore, the assumption of a certain distance from these external sources helps to make their rendering geometrically more consistent and believable as the listener moves, rather than having an unnatural, fixed orientation.
[0102] In some embodiments, when the parameters correspond to a sound source outside the area defined by the microphone array, at least one parameter is modified to a greater extent than when they correspond to a sound source within the area defined by the microphone array.
[0103] In some embodiments, the determination of whether the directional parameter corresponds to a sound source outside or inside the area defined by the microphone array is achieved by comparing whether the directional parameter associated with the directional parameter is a first directional parameter that is closer to or farther away from the area defined by the microphone array or a second directional parameter that is closer to or toward the area defined by the microphone array.
[0104] For example, Figure 3The diagram illustrates a microphone arrangement on a plane where the microphone arrays (shown as circular arrays 1 301, 2 303, 3 305, 4 307, and 5 309) are located. Spatial metadata has been determined at the array locations. This arrangement has five microphone arrays on the plane. For example, the plane can be divided into interpolation triangles using Delaunay triangulation. When a user moves to a location within a triangle (e.g., location 1 311), three microphone arrays forming the triangle containing that location are selected for interpolation (arrays 1 301, 3 305, and 4 307 in this example case). When a user moves outside the area spanned by the microphone arrays (e.g., location 2 313), the user's location can be projected to the nearest location within the area spanned by the microphone arrays (e.g., projected location 2 314), and the array triangle containing the projected location is selected for interpolation (in this example, with respect to location 2 and projected location 2, these microphone arrays are arrays 2 303, 3 305, and 5 309).
[0105] about Figure 4 Example apparatuses suitable for implementing some of the embodiments described herein are shown.
[0106] In this example, the system input is a plurality of signal sets based on microphone array signals 400. These plurality of signal sets can be, for example, a plurality of high-order Ambisonics (HOA) signal sets. In some embodiments, the plurality of signal sets based on microphone array signals may include J multi-channel signal sets. The signals can be the microphone array signals themselves, or array signals in some converted form, such as Ambisonic signals. These signals can be represented as s j (m,i), where j is the index of the microphone array from which the signal originates (i.e., the signal set index), m is the sampling time, and i is the channel index of the signal set.
[0107] Furthermore, further inputs to the system may include microphone array positions 404. The microphone array positions 404 (for each array j) can be defined as a position column vector p. j,arr This can be a 3×1 vector containing Cartesian coordinates of x, y, and z in meters. In the following example, only a 2×1 column vector containing x and y coordinates is shown, where it is assumed that the source, microphone, and listener have the same elevation angle (z-axis). However, the method described herein can be directly extended to include the z-axis as well. Further inputs are listener position 418 and listener orientation 416.
[0108] Figure 4The example shown illustrates a spatial analyzer 401 configured to receive multiple sets of signals from microphone array signals 400 determined based on (spatial) metadata for each array. These spatial / parameterized audio parameters can be determined based on any known mechanism, such as those described in GB2002710.8. The method for determining the spatial metadata can be similar to that implemented in Directional Audio Coding (DirAC). DirAC can employ a method that provides directional values and ratio values in the frequency band indicating how the sound is directional or non-directional based on a first-order capture signal. This is also an example spatial metadata dataset derived for each array. The spatial analyzer 401 is then configured to output the generated (for each array) metadata 402 to a spatial metadata and audio signal determiner 407 for projected listener positions. The projected listener positions can also be referred to as second listener positions.
[0109] In the example shown here, the second listener position may be located on the boundary of one of the “inner” regions; in other words, on the edge of a plane defined by the locations of the two (closest) audio signal sets (or on a surface of a volume defined at least partially by the locations of the two audio signal sets), where the signal sets are shown in the following example as the locations of the capture microphone array. However, in some embodiments, the second listener position (or the projected listener position) may be a location in an “outer” region, but closer to the “inner” region than the determined listener position. Furthermore, as described later, the second listener position may be located within an “inner” region (which may still be outside a different “inner” region). Furthermore, the modified metadata for these locations outside the “inner” regions may be determined in a manner similar to that defined below. For example, modified metadata from an edge or surface boundary (or some other point in the inner region) may be used for a second listener position slightly outside the “inner” region.
[0110] In some embodiments, the spatial analyzer 401 may include a suitable time-frequency converter configured to receive multiple signal sets based on the microphone array signal 400. The time-frequency converter is configured to convert the input signal s j (m,i) can be converted to the time-frequency domain, for example, using a Short-Time Fourier Transform (STFT) or a group of complex modulated quadrature mirror filters (QMF). As an example, STFT is typically configured such that for a frame length of N samples, the current frame and the previous frame are windowed and processed using a Fast Fourier Transform (FFT). The result is represented as S... j The time-frequency domain signal is represented by (b,n,i), where b is the frequency bin and n is the time frame index. The audio signal from the time-frequency microphone array can then be output to various estimators.
[0111] Spatial analysis can be based on any suitable technique, and there are known suitable methods for various input types. For example, if the input signal is in Ambisonic or Ambisonic-related form (e.g., derived from B-format microphones), or if the array can be reasonably converted to Ambisonic form (e.g., Eigenmike), then directional audio coding (DirAC) analysis can be performed. First-order DirAC has been described in Pulkki and Ville, “Spatial Sound Reproduction with Directional Audio Coding” (Journal of the Audio Engineering Society, Vol. 55, No. 6 (2007): 503-516), which specifies a method for estimating a set of spatial metadata from B-format signals (variants of first-order Ambisonics), consisting of directional parameters in the frequency band and environment-to-total-energy ratio parameters.
[0112] When higher-order Ambisonics are available, Archontis Politis, Juha Vilkamo, and VillePulkki's "Sector-Based Parametric Reconstruction of Sound Fields in the Spherical Harmonic Domain" (IEEE Transactions on Signal Processing, Vol. 9, No. 5 (2015): 852-866) provides a method for obtaining multiple simultaneous directional parameters. Other methods that can be implemented in some embodiments include estimating spatial metadata from planar devices such as mobile phones and tablets, as described in PCT Publication Patent Application WO2018 / 091776, and similar delay-based analysis methods for non-planar devices, as described in GB Publication Patent Application GB2572368.
[0113] In other words, there are multiple methods to obtain spatial metadata, and the method chosen may depend on the array type and / or audio signal format. In some embodiments, one method is applied in one frequency range and another method is applied in another frequency range.
[0114] In some embodiments, the apparatus includes a listener position projector 405. The listener position projector 405 is configured to receive a microphone array position 404 and a listener position 418, and determine a projected listener position 406. The projected listener position 406 is passed to a spatial metadata and audio signal determiner 407 for use in determining the projected listener position.
[0115] As is known in the prior art, a key objective of parametric spatial audio capture and rendering is to obtain spatial audio reproduction that is perceptually accurate for the listener. Therefore, the listener position projector 405 is configured to determine the position or interpolation data of the projection for any location (since the listener can move to any location), allowing modification of metadata based on the microphone array position 404 and the listener position 418.
[0116] In the example here, the microphone array is located on a plane. In other words, the array has no z-axis displacement component. However, in some embodiments, it is possible to extend the embodiment to the z-axis, and to have the microphone array located on a line (in other words, only one axis displacement).
[0117] In some embodiments, the listener position projector 405 may, for example, determine the projected listener position vector p. L (In this example, it is a 2x1 vector containing x and y coordinates);
[0118] Therefore, the spatial metadata and audio signal determiner 407 for the projected listener position is configured to obtain multiple signal sets based on the microphone array signal 400, metadata 402 for each array, microphone array position 404, and the projected listener position 406. The spatial metadata and audio signal determiner 407 for the projected listener position is configured to determine the spatial metadata and audio signal corresponding to the projected listener position.
[0119] The determination of the spatial metadata and audio signal corresponding to the listener position block of the projection can be achieved in a manner similar to that described in GB2002710.8.
[0120] For example, the spatial metadata of the listener's position used for projection and the audio signal determiner 407 can be configured to specify interpolation weights w1, w2, w3. These weights can be specified, for example, using the following known transformation between barycentric coordinates and Cartesian coordinates. First, based on the microphone array position vector, a uniform value is appended to each vector and the resulting vectors are combined into a matrix. To determine the 3×3 matrix:
[0121]
[0122] Microphone array position vector and The microphone arrays j1, j2, and j3 correspond to the triangle within which the listener's position forms the projection.
[0123] Then, the weights are determined using the inverse of the matrix and a 3×1 vector, which is derived from the (projected) listener position vector p. L Obtained by appending a uniform value:
[0124]
[0125] Then, the interpolation weights (w1, w2, and w3) and the position vector (p) can be used. L , and Together with the microphone placement indices (j1, j2, and j3), they are used to determine the spatial metadata and audio signal used for the listener's location in the projection.
[0126] The spatial metadata for the listener's location determined for projection can be an interpolation of metadata using interpolation weights w1, w2, and w3. In some embodiments, this can be achieved by first interpolating the azimuth angle θ used for the frequency band k and time index n. j (k,n), elevation angle and the ratio of direct energy to total energy r j This is achieved by converting the spatial metadata of (k,n) into vector form:
[0127]
[0128] Then, these vectors are averaged using the following formula:
[0129]
[0130] Then it is represented as
[0131] v(k,n)=[v1(k,n) v2(k,n) v3(k,n)] T
[0132] The interpolated metadata is obtained using the following formula:
[0133] θ(k,n)=atan2(v2(k,n),v1(k,n))
[0134]
[0135]
[0136] Then, the interpolated spatial metadata 410 is output to the metadata direction-to-location mapper 411 and the modified spatial metadata determiner 413.
[0137] The above provided an example of metadata interpolation. Other interpolation rules can be designed and implemented in other embodiments. For example, the interpolation ratio parameter can also be determined as a weighted average of the input ratios (based on w1, w2, w3). Furthermore, in some embodiments, averaging may also involve weighting based on the energy of the array signal.
[0138] The audio signal used to determine the listener's position for projection can be an interpolation of the input audio signal 400. Therefore, multiple sets of audio signals (or their time-frequency domain transformed versions) can be used to determine the total energy for each audio signal and each frequency band. In the example where the multiple sets of signals based on the microphone array signal 400 are in the form of FOA signals, the total energy can be determined as the omnidirectional energy, i.e., the energy of the first channel of the FOA signal.
[0139]
[0140] Among them, b k,low It is the first cell of frequency band k, b k,high This is the last warehouse.
[0141] The spatial metadata and audio signal determiner 407 used for the listener's location in the projection can be configured to determine distance values for indices j1, j2, and j3. And the index with the minimum distance is denoted as j. minD .
[0142] Then, the spatial metadata of the listener's location used for projection and the audio signal determiner 407 are configured to determine the selected index j. sel For the first frame (or, when processing begins), the spatial metadata of the listener's position for projection and the audio signal determiner 407 can be set. sel =j minD .
[0143] For the next or subsequent frame (or any desired temporal resolution), when the user position has potentially changed, the spatial metadata of the listener's position used for projection and the audio signal determiner 407 are configured to determine whether a change in selection j is necessary. sel If j sel If a value is not contained within j1, j2, or j3, then a change is needed. This condition means the user has moved to a location that does not contain j1, j2, or j3. sel Another region. If d jsel >d jminD α, where α is a threshold, also needs to be changed. For example, α = 1.2. This condition means that, in relation to j sel Compared to the array position, the user has clearly moved closer to j. minD The array position. This threshold is needed so that when the user is in the middle of two positions, the selection does not change back and forth erratically (in other words, a hysteresis threshold is provided to prevent rapid switching between arrays).
[0144] If any of the above conditions are met, then j sel =j minD Otherwise, keep j. sel The previous value.
[0145] The intermediate interpolation signal was determined as
[0146]
[0147] Using this processing, when j sel When changes are made, it follows a simultaneous change selection for all frequency bands. In some embodiments, the selection is set to change in a frequency-dependent manner. For example, when j sel When a change is made, some frequency bands are updated immediately, while others are changed in subsequent frames, until all frequency bands have been changed. It may be necessary to change the signal in this frequency-dependent manner to reduce the frequency fluctuations in signal S′. interp Possible switching artifacts at (b,n,i). In this configuration, when a switch occurs, for a short transition period, the signal S′... interp Some frequencies of (b,n,i) may come from one microphone array, while others may come from another.
[0148] Then, the spatial metadata of the listener's location for projection and the audio signal determiner 407 are configured to determine the energy-corrected intermediate signal S′. interp (b,n,i). Represents the equalization gain in the frequency band.
[0149]
[0150] g max Values that limit excessive amplification, for example, g max =4. Then, balance is achieved through multiplication.
[0151] S(b,n,i)=g(k,n)S′ interp (b,n,i)
[0152] Where k is the frequency band index of the bin b. The spatial metadata of the listener's location for projection and the audio signal determiner 407 are configured to output the signal S(b,n,i) as an audio signal 408 to the synthesis processor 415.
[0153] In this example embodiment, the spatial metadata 410 (projected location) includes direction (azimuth θ(k,n) and elevation φ(k,n)) parameters in the time-frequency domain and a parameter directly related to the total energy ratio r(k,n) (where k is the frequency band index and n is the time frame index). In other embodiments, other parameters may be used additionally or alternatively.
[0154] In some embodiments, the apparatus 499 includes a metadata direction-to-position mapper 411. The metadata direction-to-position mapper 411 is configured to receive spatial metadata 410 from the listener position for projection and the spatial metadata 410 from the audio signal determiner 407, the projected listener position 406, and to map the direction [θ(k,n), φ(k,n)] to spatial positions x(k,n), y(k,n), and z(k,n) in a Cartesian coordinate system, in this example, on the surface of a shape. The shape can be any suitable shape, and it can be fixed or adaptive. The mapped position in Cartesian coordinates is the location where a line from the projected listener position toward the direction [θ(k,n), φ(k,n)] intersects the determined shape. In other words, the shape in this example is determined by distance parameters d(θ(k,n), φ(k,n)). The projected listener position 406 at time index n is represented as x P (n), y P (n), z P (n), and the mapping is performed by the following formula:
[0155] x(k,n)=cos(θ(k,n)) cos(φ(k,n)) d(θ(k,n),φ(k,n))+x P (n)
[0156] y(k,n)=sin(θ(k,n)) cos(φ(k,n)) d(θ(k,n),φ(k,n))+y P (n)
[0157] z(k,n)=sin(φ(k,n)) d(θ(k,n),φ(k,n))+z P (n)
[0158] In some embodiments, the shapes in different directions, i.e., the distances d(θ(k,n), φ(k,n)), will make the distances of the sound source from the projected position in the corresponding directions more than equal to the distances of the sound source. For example, multi-array source localization techniques or visual analysis methods can be used to determine the general area where the source is located, and an approximate function for d(θ(k,n), φ(k,n)) can be determined accordingly.
[0159] If this information is unavailable or cannot be reliably estimated, it can also be set to a predefined fixed distance value, or it can use geometric information to limit possible source distances in different directions. For example, in the simplest case, a sphere with a specific radius (e.g., 2 meters) in meters can be globally set. Alternatively, if there are room boundaries around the array, or specific known boundaries (e.g., walls) in different directions, the distance from the array edge to these boundaries can be used as the assumed maximum source distance.
[0160] Therefore, the directions [θ(k,n),φ(k,n)] are mapped to the mapped metadata positions 412x(k,n), y(k,n) and z(k,n), which are output and can then be passed to the modified spatial metadata determiner 413.
[0161] In some embodiments, the apparatus 499 includes a modified spatial metadata determiner 413. The modified spatial metadata determiner 413 is configured to receive mapped metadata location 412, spatial metadata 410, listener location 418, and microphone array location 404, wherein the microphone array location 404 is configured to determine appropriate metadata for the actual listener location, while the original spatial metadata 410 is determined for the projected listener location 406. In other words, the modified spatial metadata determiner 413 is configured to determine a modified orientation [θ]. mod (k,n),φ mod [(k,n)] and the modified direct-to-total energy ratio r mod (k,n). When the projected listener position 406 is the same as the listener position 408, i.e., when the user is within the area defined by the microphone array, the modified direction and ratio can be the same as the direction and ratio of the original spatial metadata 410. Otherwise, the following procedure can be applied.
[0162] Therefore, in some embodiments, the modified spatial metadata determiner 413 may first convert the mapped position (mapped metadata position 412) into direction [θ′(k,n), φ′(k,n)] based on the listener position 418. x L (n), y L (n), z L (n) represents the listener's position; these directions can be determined by the following formula.
[0163] θ′(k,n)=atan2((y(k,n)-y L (n)),(x(k,n)-x L (n)))
[0164]
[0165] In some embodiments, these directions can be used directly as modified directions (i.e., θ). mod (k,n)=θ′(k,n) and φ mod(k,n) = φ′(k,n)). Alternatively, in some embodiments, the modified spatial metadata determiner 413 is configured to (adaptively) interpolate between the original [θ(k,n), φ(k,n)] and the mapped direction [θ′(k,n), φ′(k,n)]. For example, the original direction can be used for directions pointing "inside" the area traversed by the microphone array, while the mapped direction can be used for directions pointing "outside".
[0166] Modified direction [θ] mod (k,n),φ mod [k,n] is a fair estimate of the possible directions at the listener's position. However, it should be noted that these estimates are only "seemingly reasonable estimates" and they do not have to be accurate estimates (e.g., if the directions are simply mapped onto the surface of a sphere at a fixed distance).
[0167] In some embodiments, the modified spatial metadata determiner 413 is thus configured to modify the direct-to-total energy ratios so that the smaller they are modified, the greater the uncertainty. This modification mitigates the effects of uncertain directions because they are at least partially rendered as diffuse, while more determined directions are rendered normally.
[0168] The direct-to-total-energy ratio can be modified in any suitable manner. For example, the distances between the mapped locations (mapping metadata location 412) x(k,n), y(k,n), and z(k,n) and the listener location 418 can be determined, and the closer the listener is to the mapped location, the greater the reduction in the direct-to-total-energy ratio r(k,n) for that time-frequency tile. For example, the reduction operation could be based on a function...
[0169]
[0170] in,
[0171]
[0172]
[0173] In some embodiments, the modified spatial metadata determiner 413 is configured not to modify the direct-to-total energy ratio r(k,n) corresponding to the direction pointing to the “inside” of the area traversed by the microphone array.
[0174] Modifying the total energy ratio directly can have the following effects.
[0175] First, as the listener approaches the assumed location, the "direction" of sound sources outside the area (for which there is no accurate information about their actual direction) becomes worse. Therefore, the listener will not arrive at the erroneous assumption that the sound source is located in a precise (and potentially incorrect) location.
[0176] Secondly, sound sources within the area should be rendered as point sources. These sources are fairly directional, so for quality reasons, it's best to render them as point sources. This helps the listener navigate the sound scene and makes the rendered audio scene more natural, since only some sound sources are non-directional (when outside the area).
[0177] Third, if the listener is very far from the area, all sound sources will be reoriented (both inside and outside the area). This is because it can be assumed that the sound sources captured by the microphone array may not be very far from the microphone array.
[0178] Furthermore, the synthesis processor 415 is configured to receive audio signal 408, modified spatial metadata 414, and listener orientation 416. The synthesis processor 415 is configured to perform spatial rendering of the audio signal 408 to generate spatialized audio output 420. Spatialized audio output 420 can be any suitable format, such as binaural, surround sound, or ambisonics.
[0179] Spatial processing can be any suitable synthetic processing. For example, suitable spatial processing is described in GB2002710.8.
[0180] Therefore, for example, the synthesis processor can be configured to determine the vector rotation function to be used in the following formula. Based on the principles of Laitinen and MV in their 2008 Master's thesis at Helsinki University of Technology, pages 54-55, "Binaural Reproduction for Directional Audio Coding," the rotation function can be defined as...
[0181]
[0182] Here, yaw, pitch, and roll are the head orientation parameters, and x, y, and z are the values of the unit vector being rotated. The result is x′, y′, and z′, which are the rotated unit vectors. The mapping function performs the following steps:
[0183] 1. Yaw rotation
[0184] x1=cos(yaw)x+sin(yaw)y
[0185] y1 = -sin(yaw)x + cos(yaw)y
[0186] z1=z
[0187] 2. Pitch and Rotation
[0188]
[0189] y2=y1
[0190]
[0191] 3. Finally, roll and rotate.
[0192] x′=x2
[0193]
[0194]
[0195] Once these parameters are determined, the compositing processor 415 can implement any suitable spatial rendering. For example, in some embodiments, the compositing processor 415 can implement 3DOF rendering, for instance, based on the principles described in PCT Publication WO2019086757. It should be noted that "3DOF rendering" actually means 6DOF rendering, because position processing has already been considered in the audio signal 408 and the modified spatial metadata 414, and the compositing processor only needs to consider head rotation (the remaining 3 degrees of freedom out of 6).
[0196] In other words, the operation of the synthesis processor 415 can be summarized as follows:
[0197] 1) Convert the “audio signal” to a time-frequency representation using any known filter bank suitable for audio processing (unless it has already been done).
[0198] 2) Based on spatial metadata, process time-frequency audio signals in the frequency band, and
[0199] 3) Convert the processed audio back to a time domain signal to obtain spatial audio output 420.
[0200] In some embodiments, the synthesis processor 415 is configured to, if rendering binaural output signals, first rotate the orientation parameter [θ] based on the head orientation. mod (k,n),φ mod (k,n). This is achieved by converting the directions into unit vectors [xyz] pointing in the corresponding directions. T Use the function rotate([xyz]) T (yaw, pitch, roll) to obtain the rotated unit vector [x′y′z′] T Then, the unit vector is converted into rotated azimuth and elevation parameters [θ]. modR (k,n),φmodR [k,n] is then implemented. The synthesis processor 415 is then configured to use the head-related transfer function (HRTF) in the frequency band to reduce the direct energy ratio r of the audio signal. mod (k,n) guides to the direction [θ] modR (k,n),φ modR [(k,n)], and uses a decorrelationer configured to provide appropriate diffusion field interaural correlation to reduce the ambient energy ratio 1-r of the audio signal. mod (k,n) is guided as a spatially non-localizable sound. This processing is adapted for each frequency and time interval (k,n) determined by spatial metadata. Similarly, for speaker output, a translation function can be used to render the direct portion to make the target speaker layout and environment discontinuous between speakers. In speaker playback, metadata rotation is not required because it is considered during the listening time when the sound is reproduced from the speaker. Similarly, for Ambisonic output, the translation function can be an Ambisonic translation function, and the environment can also be discontinuous between output channels, but with levels depending on the Ambisonic normalization scheme used. In Ambisonic rendering, rotation is generally not required because if the Ambisonic sound is ultimately rendered as a binaural output, head orientation is assumed to be considered in the Ambisonic renderer.
[0201] about Figure 5 This shows, as Figure 4 The flowchart of the example device shown.
[0202] Multiple signal sets are obtained based on the audio signals from the microphone array, such as Figure 5 As shown in step 501.
[0203] Spatial analysis is performed on multiple signal sets to determine metadata for each microphone array, such as... Figure 5 Step 511 is shown in the diagram.
[0204] Obtain the microphone array position, such as Figure 5 As shown in step 503.
[0205] Additionally, obtaining the listener's location, such as... Figure 5 The step 507 in the middle is shown.
[0206] After obtaining the listener's location and the microphone array location, determine the listener's location for the projection, such as... Figure 5 As shown in step 509.
[0207] Then, after obtaining the listener location and spatial metadata for the projection (and having already obtained the microphone array location), the spatial metadata and audio signal used for the listener location for the projection are determined, such as... Figure 5 Step 513 is shown in the diagram.
[0208] Then, after determining the spatial metadata of the listener's location for projection, the metadata orientation is mapped to the location, such as... Figure 5 As shown in step 515.
[0209] Furthermore, after determining the location of the mapping, the modified spatial metadata is then determined, such as... Figure 5 Step 517 is shown in the diagram.
[0210] After obtaining the listener's location and modified spatial metadata (as well as the audio signal), the generation of spatialized audio signals (e.g., binaural, surround sound, ambisonics) is performed, such as... Figure 5 Step 519 is shown in the diagram.
[0211] Then, the spatialized audio signal is output (to an output device, such as headphones), as... Figure 5 As shown in step 521.
[0212] In some embodiments, portions of the environment can be rendered based on the distance between the listener and the area spanned by the microphone array. For example, when the listener is near the area (or within the area), the target orientation distribution for environment rendering may follow the orientation distribution of the audio signal captured by the nearest microphone array, while when the listener is far from the area, the target orientation distribution may be more omnidirectional. This can help avoid misperceptions of the environment's orientation when the listener is far from the microphone array.
[0213] In some embodiments, the direct and ambient portions are not rendered separately as described above, because improved processing quality can be achieved by using a blending technique that renders the direct and ambient portions in the same processing step. The advantage is that it minimizes the need for decorrelation, which can be detrimental to perceived audio quality. This optimized audio processing procedure is further described in detail in GB2002710.8.
[0214] In the example embodiments described above, the listener position requires spatial parameters determined from the external microphones forming the microphone array arrangement. If the listener position can be projected onto the outer edge of the array (edge rendering), then when the listener position is on the edge, parameter interpolation is performed from the two microphone pairs forming the edge, similar to GB2002710.8. In this embodiment, a smooth transition from the internal rendering method of GB2002710.8 to the external rendering as described in the embodiments herein can be achieved when the listener crosses the boundary via the edge. Valid edges can be found by projecting the listener onto the nearest edge and determining whether the projection point is on or outside the edge. One method for determining the nearest edge is to maintain a list of external edges and find the two edges connected to it based on the nearest microphone.
[0215] Therefore, for example, such as Figure 6 As shown in edge rendering 601, example microphone array positions (shown as circular arrays 1 603, 2 611, 3 609, 4 605, and 5 607) are illustrated. Furthermore, the listener at position 626 has a first projection 612 that intersects the (vector) line from the positions connecting arrays 1 603 and 2 611 at a point P 616 between the positions of arrays 1 603 and 2 611. There is also a second projection 614 that intersects the (vector) line from the positions connecting arrays 1 603 and 4 605 outside the positions of arrays 1 603 and 4 605.
[0216] However, when the listener is outside the array, there are areas at corners where no projection exists on the edge segments. In this case (vertex rendering), the spatial metadata to be used comes directly from the nearest microphone forming the corner. This strategy enables a smooth transition from internal rendering of GB2002710.8 to external rendering as described in the embodiments herein, as the listener crosses the boundary through the microphone in the corner.
[0217] Therefore, for example, such as Figure 6 The vertex rendering configuration 651 shows example microphone array positions (displayed as circular arrays 1 603, 2 611, 3 609, 4 605, and 5 607). Furthermore, the listener at position 661 has a first projection 662 that intersects a (vector) line from the positions connecting arrays 1 603 and 2 611, outside the positions of arrays 1 603 and 2 611. There is also a second projection 664 that intersects a (vector) line from the positions connecting arrays 1 603 and 4 605, outside the positions of arrays 1 603 and 4 605. A third projection 666 extends directly to the nearest array microphone position, array 1 603.
[0218] Therefore, in some embodiments, a geometric check can be performed to determine whether edge rendering or vertex rendering should be applied. The geometric check can be based on identifying the two edges adjacent to the nearest microphone and projecting the listener onto these two edges. If either projection falls within an edge segment, edge rendering is considered, and if no projection falls within an edge segment, vertex rendering is considered.
[0219] Therefore, for example, such as Figure 7The edge rendering configuration 701 shows the locations of example microphone arrays (displayed as circles: array 1 603, array 2 611, array 3 609, array 4 605, and array 5 607). Furthermore, an edge rendering region 711 is shown, defined by (vector) lines connecting the locations of array 1 603 and array 2 611, (vector) lines connecting the locations of array 1 603 and array 4 605, and (vector) lines connecting the locations of array 2 611 and array 3 609.
[0220] The vertex rendering configuration 751 is further illustrated, in which there is a vertex rendering region 761 defined by (vector) lines connecting the positions of array 1 603 and array 2 611 and (vector) lines connecting the positions of array 1 603 and array 4 605.
[0221] In some embodiments, the spatial parameters can be modified by a modified spatial metadata determiner 413 based on an angular weighting between the original spatial parameters of the edges or vertices and the spatial parameters due to projection. In this embodiment, the modified spatial metadata determiner 413 uses information from the array geometry and the estimated DOA to allow modification primarily of parameters from sources that appear to originate outside the array, while leaving spatial parameters originating from the array region largely unaffected. In this way, external sounds become “blurred” when the listener leaves the microphone array area, but sounds emanating from the microphone array area can maintain their directional clarity, thus providing an acoustic anchor to the array when the listener moves outside of it.
[0222] In some embodiments, the modified spatial metadata determiner 413 is configured to determine orientation weights as follows:
[0223] Calculate the vertex normal pointing outward for each microphone on the outer array boundary. Each vertex normal is the average of the two normals from the two edges connected to that vertex. These normals can then be used in both vertex and edge rendering modes to indicate the direction of maximum "outward" from the inside of the array. If the listener is on vertex rendering, the normal vector from the nearest microphone is used. If the listener is on edge rendering, the normal vector is determined by inserting two vertex normals at the ends of the edges based on the listener's projected position.
[0224]
[0225] Where, d AB The distance between points A and B is indicated by unit{}, which is a function that normalizes a vector into a unit vector with the same direction.
[0226] Therefore, for example, Figure 8As shown, example microphone array positions are illustrated (displayed as circular arrays 1603, 2611, 3609, 4605, and 5607). Additionally, the listener at listener position 813 is also shown.
[0227] There exists a first vertex normal. 811 is a combination of the (vector) line connecting the positions of array 1603 and array 2611 and the (vector) line connecting the positions of array 1603 and array 4605.
[0228] There exists a second vertex normal. 815 is a combination of the (vector) line connecting the positions of array 1603 and array 2611 and the (vector) line connecting the positions of array 2611 and array 3609.
[0229] In addition, the edge "normals" are also shown. 819 is a combination of the first and second vertex normals from the projection point 817. In this example, point P is the listener position of the projection, and there exists an edge normal n_p based on n_1 and n_2, as described above. Therefore, as the listener moves along the edge, point P changes with the listener position to modulate from one vertex normal to the other vertex normal on one side of the edge.
[0230] Therefore, in some embodiments, it can be based on the analysis The weighting function is determined using the listener's position relative to the projection and the normal vector.
[0231]
[0232] w2(k,n) = 1 - w1(k,n)
[0233] Here, N is the power factor, which determines the magnitude by which the directional weights are increased outwards from the array. For example, for N=1, the weights have a cardioid pattern with the peak at the outward-pointing normal; for N=2, it has a second-order cardioid pattern, and so on.
[0234] Therefore, when the listener is far from the edge or vertex, the mapped DOA is determined as described above using vector notation:
[0235]
[0236] here, It is the mapped DOA. It is the listener's location. d is the listener position projected onto the vertex or edge, and d is the distance to the mapping boundary.
[0237] In some embodiments, the mapping effect is (primarily) applied to the external DOA; therefore, the directional weighting can be determined as
[0238]
[0239] From the final revised From The direction is determined by the modified azimuth and elevation angle θ. mod ,
[0240] Furthermore, in some embodiments, to increase diffuseness as we move away from the edge, the maximum effect at distance R can be achieved by reducing the direct-to-total energy ratio in a manner similar to reducing the direct-to-total energy ratio as described above:
[0241]
[0242] in, and Similar to the previous embodiments.
[0243] In some embodiments, the external DOA modification is primarily targeted at the total energy ratio, while the rendering of internal sources remains largely unaffected. Therefore, in some embodiments, the directional weighting can be determined as follows:
[0244] r mod (k,n)=min[w1(k,n)k′ mod (k,n)+w2(k,n)r(k,n),1]
[0245] This is for example in Figure 9 The image shows edge rendering at 900 on the left. In this example, the microphone array positions are shown as circles: array 1 901, array 2 903, and array 3 905. The listener is located at position 919, outside the area defined by the microphone array positions and within the area extending from the (vector) line between array 1 901 and array 2 903. The normal is also shown. 913, DOA 917 and the mapped DOA 921, Directional Weighting Function w1915, and DOA The product of the distance d 923.
[0246] Indicates the outer edge normal of the array 913 is displayed vertically for ease of visualization, but in practice, it can be tilted more towards the vertex normal, depending on the listener's position.
[0247] The right side shows vertex rendering at 950. In this example, the microphone array position or location is shown as a circular array 1 901, and the listener is located at listener position 969, outside the area defined by the microphone array position and outside the area extending from the (vector) line between array 1 901 (and any other array). Furthermore, normals are also shown. 963, DOA 967 and the mapped DOA 971. Directed weighting function w1965 and DOA Product with distance d 973. Range 907 / 957 shows a surface whose direction is mapped to the location mapper 411 by the metadata direction. In this example, the surface is a simple sphere, and therefore has a constant radius.
[0248] While it is always possible to construct an outer boundary (which is convex (convex hull)) between all microphone array locations, sometimes the resulting edges are not effective; for example, the edges may be too long for the effective space interpolation between connected microphone arrays. In some embodiments, the outer edges can be removed, resulting in a non-convex hull arrangement. In this case, the derived normals may lose their effectiveness because they do not necessarily point from the inside to the outside. Therefore, in some embodiments, the non-convex edge normals and connecting vertices can be replaced with the normals of the omitted edges.
[0249] For example, this is in Figure 10 The diagram illustrates an example arrangement 1000 with microphone array positions (shown as circular arrays 1603, 2611, 3609, 4605, and 5607) and a listener at listener position 1003. Additionally, an example "long" side 1001 between arrays 1603 and 4605 is shown. Furthermore, there is a vertex normal 1013 associated with array 1603 and a vertex normal 1015 associated with array 4, as well as example interpolated or weighted edge normals 1021, 1023, and 1025 positioned along the "long" side 1001.
[0250] A modified arrangement 1050 is further shown, wherein the example arrangement 1000 is modified by removing the example "long" side 1001. This results in a non-convex arrangement, and the listener position 1003 is now located outside the area defined by the microphone array positions. Furthermore, the new non-convex normals (not shown) along the two new short sides, the first "short" side defined by the line between array 1 603 and array 5 607 positions, and the second "short" side defined by the line between array 5 607 and array 4 605 positions do not point outwards. Therefore, as... Figure 10As shown, the original discarded edge normal 1023 is copied to all microphones on the new outer edge (replacing the discarded or deleted edge). In this embodiment, the listener is projected onto one of the new outer edges, and the new vector pointing outward is determined, as before, by interpolating the microphone normals around the edge (which can be either copied microphone normals or original microphone normals).
[0251] Furthermore, in some embodiments, besides determining the modified outward vector pointing outward, the process of projecting the listener onto the edge or microphone is handled differently for non-convex boundaries. After omitting an edge, if the listener is projected perpendicularly onto a new edge below the omitted edge, then, as is typically done for the convex exterior of the array region, there will be positions where the listener is projected onto both edges simultaneously, instead of the preferred position. To avoid this, the listener is always projected onto the new edge, not perpendicular to them, but perpendicular to the originally discarded edge (see...). Figure 11 ), thus projecting onto the unique non-convex edge.
[0252] For example, this is in Figure 11 The diagram illustrates an exemplary "clear" arrangement 1103 having microphone array positions (shown as circles, arrays 1 603, 2 611, 3 609, 4 605, and 5 607) and a listener position 1101 outside the microphone array area. In this example, an "effective" projection 1123 to a first "short" side defined by the line between arrays 1 603 and 5 607 and an "ineffective" projection 1121 relative to a second "short" side defined by the line between arrays 5 607 and 4 605 are shown.
[0253] An example “fuzzy” arrangement 1113 of a listener with the same microphone array position and a listener position 1111 outside the microphone array area is shown, wherein there are two “effective” projections, namely, a first projection 1133 to a first “short” side defined by the line between positions 1603 and 5607 and a second projection 1131 relative to a second “short” side defined by the line between positions 5607 and 4605.
[0254] Based on the above embodiments, this can be achieved by implementing a vertical projection from the omitted edges to the listener that will intersect with one of the new edges, as shown in example arrangement 1123. In other words, the listeners are projected 1141 onto new edges that are not perpendicular to them, but perpendicular to the original discarded or deleted edges, resulting in projection onto a single non-convex edge.
[0255] Reference Figures 12a to 12c Describe the actual effects of the embodiments.
[0256] For example, Figure 12a The listener 1209 is shown within a region 1201 spanned by a microphone array (not shown), where sound sources (shown as sound sources 1203 and 1205 within the region and sound source 1207 outside the region) are reproduced in the correct direction.
[0257] Figure 12b This illustrates that the listener 1219 has now moved outside the area 1201 spanned by the microphone array (not shown). The conventional rendering method involves the sound sources—namely, the rendered sources 1213, 1215, and 1217 marked with solid boxes—being moved from the original source positions 1203, 1205, and 1207, respectively, as the listener moves. In this example, the sound sources are reproduced in the wrong direction, and the listener has difficulty understanding their location.
[0258] For example, although the rendered source 1213 is roughly in the correct direction relative to the listener 1219 when compared to the direction of source 1203 relative to the listener 1219 position, the direction of source 1217 is roughly opposite to the direction of source 1207 relative to the listener 1219 position. Furthermore, although a sound source can be rendered with poor directionality, it does not help navigation and may even make navigation more difficult.
[0259] Figure 12c The embodiment described above illustrates that, using the listener position 1219, is first projected 1235 onto the edge 1231 of region 1201. Then, sound sources outside the region (sound source 1207) are mapped onto the surface of sphere 1233. These mapped sources are then rendered in those directions from the listener's position. Furthermore, as described above, the mapped sound sources near the listener's position have poorer orientation. Sound sources within regions (1203, 1205) are rendered with less modification (based on the projected position). As a result, sound sources within the region are rendered as point sources from approximately the correct direction, thus allowing them to be used for navigation within the region. Sound sources outside the region are rendered in seemingly reasonable locations (even if not entirely accurate), with poorer orientation as the listener approaches them. Therefore, they do not confuse the listener but rather provide some seemingly reasonable positioning.
[0260] although Figure 4 The example device shown is depicted as implemented in a single device; however, the capture and processing / rendering components may be physically separate or implemented at different times. For example, regarding... Figure 13 , showed Figure 4 A variation of the illustrated embodiment. In this embodiment, the difference between the two examples is the addition of an encoder / multiplexer 1305 and a decoder / demultiplexer 1307.
[0261] The encoder / multiplexer 1305 is configured to receive multiple signal sets based on microphone array signals 400, metadata 402 for each array, and microphone array positions 404, and apply a suitable encoding scheme to the audio signals, such as any method of encoding Ambisonic signals described in the context of MPEG-H, i.e., ISO / IEC 23008-3:2019 Information Technology—Efficient Encoding and Media Delivery in Heterogeneous Environments—Part 3: 3D Audio. In some embodiments, the encoder / multiplexer 1305 may also downmix or otherwise reduce the number of audio channels to be encoded. Furthermore, in some embodiments, the encoder / multiplexer 1305 may quantize and encode the spatial metadata 402 and array position 404 information, and embed the encoded result along with the encoded audio signal into a bitstream 1399. The bitstream 1399 may further be provided as an encoded video signal at the same media container. The encoder / multiplexer 1305 may then be configured to output (e.g., transmit or store) the bitstream 1399.
[0262] In some embodiments, the encoder / multiplexer 1305 may be configured to omit the encoding of some signal sets based on the bit rate used, and if so, also omit the array positions and metadata corresponding to the encoding.
[0263] The decoder / demultiplexer 1307 can be configured to receive (or retrieve or otherwise obtain) bitstream 1399 and decode and demultiplex multiple signal sets based on microphone array 1300 (and provide them to the spatial metadata and audio signal determiner 407 for the listener location used for projection), microphone array locations 1304 (and provide them to the listener location projector 405 and the spatial metadata and audio signal determiner 407 for the listener location used for projection), and metadata 1302 for each array (and provide them to the spatial metadata and audio signal determiner 407 for the listener location used for projection).
[0264] about Figure 14 , showed Figure 13 Encoder and decoder embodiments (and Figure 4 Example application of the embodiments.
[0265] In this example, there are three microphone arrays, which could be, for example, a spherical array with a sufficient number of microphones (e.g., 30 or more), or a VR camera with microphones mounted on a surface (e.g., an OZO or similar product from Nokia). Thus, microphone array 1 1401, microphone array 2 1411, and microphone array 3 1421 are shown, configured to output audio signals to computer 1 1405 (in this example, an FOA / HOA converter 1415).
[0266] In addition, each array is equipped with a locator that provides position information for the corresponding array. Thus, microphone array 1 locator 1403, microphone array 2 locator 1413, and microphone array 3 locator 1423 are shown, which are configured to output position information to computer 1 1405 (encoder processor 1305 in this example).
[0267] Figure 14 The system further includes a computer, namely computer 11405, which includes an FOA / HOA converter 1415 configured to convert array signals into first-order Ambisonic (FOA) or high-order Ambisonic (HOA) signals. Converting microphone array signals to Ambisonic signals is known and not described in detail herein; however, if the array is, for example, an Eigenmikes array, there are available components for converting microphone signals to Ambisonic form.
[0268] The FOA / HOA converter 1415 outputs the converted Ambisonic signal to the encoder processor 1305 in the form of multiple signal sets based on the microphone array signal 400. The encoder processor 1305 can operate as an encoder processor as described above.
[0269] Microphone array positioners 1403, 1413, and 1423 are configured to provide microphone array position information to an encoder processor in computer 11405 via a suitable interface (e.g., via Bluetooth connection). In some embodiments, the array positioner also provides rotation alignment information, which can be provided to rotate and align the FOA / HOA signals at computer 11405.
[0270] The encoder processor 1445 at computer 11405 is configured as follows: Figure 13 (or Figure 4In the context of the description, the encoder processor 1445 processes multiple signal sets and microphone array positions based on microphone array signals and provides an encoded bitstream 1399 as output. In other words, in some embodiments, the encoder processor 1445 may include a spatial analyzer (per array) 401 and an encoder / multiplexer 1305.
[0271] The bitstream 1399 can be stored and / or transmitted, and then the decoder processor 1447 of computer 2 1407 is configured to receive or obtain the bitstream 1399 from storage. The decoder processor 1447 can also obtain listener position and orientation information from a position / orientation tracker worn by the user on an HMD (head-mounted display) 1431. Therefore, in some embodiments, the decoder processor 1447 includes a DEMUX / decoder 1307 and, as in [other embodiments], [other components]. Figure 13 The other remaining blocks are shown.
[0272] Based on bitstream 1399 and listener location and orientation information 1430, the decoder processor 1447 of computer 2 1407 is configured to generate binaural spatialized audio output signals 1432 and provide them via a suitable audio interface for reproduction through headphones 1433 worn by the user.
[0273] In some embodiments, computer 2 1407 is the same device as computer 1 1405; however, in typical cases, they are different devices or computers. In this context, a computer can refer to a desktop / laptop computer, a cloud processing device, a game console, a mobile device, or any other device capable of performing the processes described in this disclosure.
[0274] In some embodiments, bitstream 1399 is an MPEG-I bitstream. In some other embodiments, it can be any suitable bitstream.
[0275] In some embodiments, the listener's position can be tracked relative to the captured audio environment and / or the captured audio scene. For example, the listener may be equipped with a tracker that provides the position and orientation of the listener's head. Audio can then be rendered to the listener based on this position and orientation information, making it appear as if he / she is moving within the captured audio environment. It should be noted that the listener typically does not actually move within the captured audio environment, but rather within his / her physical environment. Therefore, movement can be relative, and the listener's motion can be scaled (up / down) to represent movement within the captured environment according to the scene. Furthermore, it should be noted that the captured audio environment can also be virtual rather than a real environment. In other words, instead of reflecting physical space, the captured audio environment is a simulated, generated, or augmented space. Additionally, it should be noted that the listener's movement can be virtual. For example, the listener can use appropriate user input such as a keyboard, mouse, or any suitable input device to indicate movement.
[0276] about Figure 15 This illustrates an example electronic device that can be used as a computer, encoder processor, decoder processor, or any functional block described herein. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.
[0277] In some embodiments, device 1600 includes at least one processor or central processing unit 1607. Processor 1607 may be configured to execute various program codes, such as those described herein.
[0278] In some embodiments, device 1600 includes memory 1611. In some embodiments, at least one processor 1607 is coupled to memory 1611. Memory 1611 can be any suitable storage device. In some embodiments, memory 1611 includes a program code portion for storing program code that can be implemented on processor 1607. Furthermore, in some embodiments, memory 1611 may further include a storage data portion for storing data, such as data that has been processed or will be processed according to the embodiments described herein. The implemented program code stored in the program code portion and the data stored in the storage data portion can be retrieved by processor 1607 via memory-processor coupling when needed.
[0279] In some embodiments, device 1600 includes a user interface 1605. In some embodiments, user interface 1605 may be coupled to processor 1607. In some embodiments, processor 1607 may control the operation of user interface 1605 and receive input from user interface 1605. In some embodiments, user interface 1605 enables a user to input commands to device 1600, for example, via a keypad. In some embodiments, user interface 1605 enables a user to obtain information from device 1600. For example, user interface 1605 may include a display configured to show information from device 1600 to a user. In some embodiments, user interface 1605 may include a touchscreen or touch interface capable of allowing information to be input to device 1600 and further displaying the information to a user of device 1600.
[0280] In some embodiments, device 1600 includes an input / output port 1609. In some embodiments, input / output port 1609 includes a transceiver. In this embodiment, the transceiver may be coupled to processor 1607 and configured to communicate with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver device may be configured to communicate with other electronic devices or devices via wired or wired coupling.
[0281] The transceiver can communicate with other devices using any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol (such as IEEE 802.X), a suitable short-range radio frequency communication protocol (such as Bluetooth), or an Infrared Data Communication Path (IRDA).
[0282] The transceiver input / output port 1609 can be configured to send / receive audio signals, bit streams, and in some embodiments, the operations and methods described above are performed by using a processor 1607 that executes appropriate code.
[0283] Generally, various embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while others may be implemented in firmware or software executable by a controller, microprocessor, or other computing device, although the invention is not limited thereto. Although various aspects of the invention may be illustrated and described as block diagrams, flowcharts, or represented using certain other graphical representations, it is well understood that such blocks, apparatuses, systems, techniques, or methods described herein may be implemented (by way of non-limiting example) in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0284] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Further, in this respect, it should be noted that any block in the logical flow of the figures may represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip, or a memory block implemented within a processor, magnetic media, or optical media.
[0285] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic storage devices and systems, optical storage devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include one or more of the following as non-limiting examples: general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.
[0286] Embodiments of the present invention can be practiced in a variety of elements, such as integrated circuit modules. The design of integrated circuits is largely a highly automated process. Sophisticated and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.
[0287] Programs (such as those provided by Synopsys in Mountain View, California, and Cadence Design in San Jose, California) use well-established design rules and pre-stored libraries of design modules to automatically route conductors and position components on semiconductor chips. Once the semiconductor circuit design is complete, the final design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility or “fab” for fabrication.
[0288] The foregoing description has provided a complete and informative description of exemplary embodiments of the invention by way of exemplary and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and appended claims, given the foregoing description. Nevertheless, all such modifications and similar alterations to the teachings of the invention will still fall within the scope of the invention as defined by the appended claims.
Claims
1. An apparatus for generating spatialized audio output based on the listener's location, comprising: At least one processor; as well as At least one memory for storing computer program code; The at least one memory and the computer program code are configured, together with the at least one processor, to make the device at least: Obtain two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; Obtain the listener's location within an audio environment, wherein the audio environment includes one or more regions, the one or more regions having one or more internal regions and external regions relative to the corresponding audio signal set location, wherein the one or more internal regions are defined by the corresponding audio signal set location; For at least two audio signal sets in the two or more audio signal sets, metadata is obtained based on the processing of at least two audio signals in the two or more audio signal sets. For the listener's location in an audio environment outside the one or more inner regions, a second listener location is determined, the second listener location being located in the one or more outer regions and closer to the boundary between the one or more inner regions and the outer regions, or located on the boundary, or located within the one or more inner regions; Based on the metadata, the modified metadata for the location of the second listener is determined; Based on the at least two audio signals, at least two modified audio signals are determined for the second listener's location; Based on the modified metadata for the second listener's location, determine the modified spatial metadata for the listener's location; and Output the at least two modified audio signals and the modified spatial metadata.
2. The apparatus according to claim 1, wherein, The modified spatial metadata is determined to make the device: Based on the modified metadata for the second listener's location, at least one audio location is determined relative to the second listener's location, wherein the modified metadata for the second listener's location includes a direction parameter representing the direction from the second listener's location to one of the at least one audio location; and Based on the location of the at least one audio signal set relative to the second listener's location, the modified spatial metadata for the listener's location is determined, wherein the modified spatial metadata includes a spatial direction parameter representing the direction from the listener's location to one of the at least one audio locations.
3. The apparatus according to claim 1, wherein, The device acquires two or more audio signal sets from a microphone arrangement, wherein each microphone is arranged at a corresponding location and includes one or more microphones.
4. The apparatus according to claim 1, wherein, The device obtains the listener's location from another device.
5. The apparatus according to claim 1, wherein, For at least two audio signal sets in the two or more audio signal sets, the metadata obtained enables the device to determine directional parameters based on the processing of the at least two audio signals.
6. The apparatus according to claim 1, wherein, The device determines the location of the second listener at a location that is: Within a plane or volume at least partially defined by an edge or surface connecting the two or more audio signal set locations and the listener location; Within at least a portion of the associated internal region, in a plane or volume defined by an edge or surface connecting two or more audio signal set locations; On the edge or surface defined by the two or more audio signal set locations; as well as At the closest audio signal set location among the two or more audio signal set locations.
7. The apparatus according to claim 1, wherein, The device determines the modified metadata for the location of the second listener. Based on the location of the audio signal set and the location of the second listener, at least two interpolation weights are generated; The at least two interpolation weights are applied to the corresponding audio signal set audio metadata to generate interpolated audio metadata; as well as The interpolated audio metadata is combined to generate the modified metadata for the second listener's location.
8. The apparatus according to claim 7, wherein, Based on the modified metadata for the second listener's location, the device determines modified spatial metadata for the listener's location, enabling the device to map the modified metadata to a Cartesian coordinate system based on the second listener's location.
9. The apparatus according to claim 1, wherein, The device determines at least two modified audio signals for the second listener's location, causing the device to generate an interpolated audio signal from the at least two audio signals.
10. The apparatus according to claim 2, wherein, The modified spatial metadata for the listener's location is determined based on at least one audio location relative to the second listener's location, wherein the modified spatial metadata includes a spatial direction parameter representing the direction from the listener's location to one of the at least one audio locations, such that the device determines the spatial direction parameter based on one of the following: The interpolation difference between the at least one audio position relative to the second listener position and the listener position; and The difference between the listener's position and the at least one audio position relative to the second listener's position.
11. The apparatus according to claim 10, wherein, Based on the modified metadata for the second listener's location, the device determines the modified spatial metadata for the listener's location such that it modifies at least one direct-to-total energy ratio based on the difference between the at least one audio location relative to the second listener's location and the listener's location.
12. The apparatus of claim 1 is further configured to: process the at least two modified audio signals to generate a spatial audio output based on the modified spatial metadata for the listener's location.
13. The apparatus according to claim 12, wherein, Generating spatial audio output causes the device to generate at least one of the following: Binaural audio output, which includes two audio signals for headphones and / or earbuds; Stereo mixed sound audio output, which includes multiple audio signals for a stereo mixed sound renderer for headphones or multi-channel speaker groups; as well as Multi-channel audio output, which includes at least two audio signals for a multi-channel speaker array.
14. A method for generating spatialized audio output based on listener location, the method comprising: Obtain two or more audio signal sets, wherein each of the two or more audio signal sets is associated with a corresponding audio signal set location; Obtain the listener's location within an audio environment, wherein the audio environment includes one or more regions, the one or more regions having one or more internal regions and external regions relative to the corresponding audio signal set location, wherein the one or more internal regions are defined by the corresponding audio signal set location; For at least two audio signal sets in the two or more audio signal sets, metadata is obtained based on the processing of at least two audio signals in the two or more audio signal sets. For the listener's location in an audio environment outside the one or more inner regions, a second listener location is determined, the second listener location being located in the one or more outer regions and closer to the boundary between the one or more inner regions and the outer regions, or located on the boundary, or located within the one or more inner regions; Based on the metadata, the modified metadata for the location of the second listener is determined; Based on the at least two audio signals, at least two modified audio signals are determined for the second listener's location; Based on the modified metadata for the second listener's location, determine the modified spatial metadata for the listener's location; and Output the at least two modified audio signals and the modified spatial metadata.
15. The method according to claim 14, wherein, Determining the modified spatial metadata for the listener's location based on the modified metadata for the second listener's location includes: Based on the modified metadata for the second listener's location, at least one audio location is determined relative to the second listener's location, wherein the modified metadata for the second listener's location includes a direction parameter representing the direction from the second listener's location to one of the at least one audio location; and Based on the location of the at least one audio signal set relative to the second listener's location, the modified spatial metadata for the listener's location is determined, wherein the modified spatial metadata includes a spatial direction parameter representing the direction from the listener's location to one of the at least one audio locations.
16. The method of claim 14, wherein, Obtaining the two or more audio signal sets includes: obtaining the two or more audio signal sets from a microphone arrangement, wherein each microphone is arranged at a corresponding location and includes one or more microphones.
17. The method of claim 14, wherein, Obtaining the listener's location includes obtaining the listener's location from another device.
18. The method according to claim 14, wherein, For at least two audio signal sets in the two or more audio signal sets, obtaining the metadata includes: determining directional parameters based on the processing of the at least two audio signals.
19. The method of claim 14, wherein, Determining the location of the second listener includes: determining the location of the second listener at one of the following locations: Within a plane or volume at least partially defined by an edge or surface connecting the two or more audio signal set locations and the listener location; Within at least a portion of the associated internal region, in a plane or volume defined by an edge or surface connecting two or more audio signal set locations; On the edge or surface defined by two of the two or more audio signal set locations; and At the closest audio signal set location among the two or more audio signal set locations.
20. The method of claim 14, wherein, The modified metadata used to determine the location of the second listener includes: Based on the location of the audio signal set and the location of the second listener, at least two interpolation weights are generated; The at least two interpolation weights are applied to the corresponding audio signal set audio metadata to generate interpolated audio metadata; and The interpolated audio metadata is combined to generate the modified metadata for the second listener's location.
Citation Information
Patent Citations
Spatial audio capture
GB2572368A
Audio rendering with spatial metadata interpolation
GB2592388A
Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices
WO2018091776A1
Determination of targeted spatial audio parameters and associated spatial audio playback
WO2019086757A1
Audio rendering with spatial metadata interpolation
WO2021170900A1