Audio Processing in Immersive Audio Services
By encoding and decoding directional audio to compensate for device motion, the apparatus stabilizes immersive audio experiences, addressing motion-related issues and improving user experience.
Patent Information
- Application Number
- JP2024076517
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-01-28
- Filing Date
- 2024-05-09
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2039-11-12
AI Technical Summary
Existing audio codecs lack the capability to handle immersive audio experiences efficiently, particularly in scenarios where the capturing device moves relative to the acoustic scene, leading to undesirable spatial rotations and potential motion sickness.
An apparatus and method for capturing and encoding directional audio that compensates for unintended motion by modifying audio characteristics based on spatial data, and a corresponding decoder and renderer that adjust audio characteristics using metadata to stabilize the audio scene.
This approach stabilizes the audio scene by compensating for device motion at the capture end, reducing the need for high bitrates and minimizing spatial data transmission, thereby enhancing user experience and preventing motion sickness.
Smart Images

Figure 0007815321000001 
Figure 0007815321000002 
Figure 0007815321000003
Abstract
Description
[Technical Field]
[0001] The present disclosure generally relates to capturing, acoustically preprocessing, encoding, decoding, and rendering directional audio of an audio scene. In particular, the present disclosure relates to an apparatus adapted to modify directional characteristics of captured directional audio in response to spatial data of a microphone system capturing the directional audio. The present disclosure further relates to a rendering apparatus configured to modify directional characteristics of received directional audio in response to the received spatial data. [Background technology]
[0002] The introduction of 4G / 5G high-speed wireless access into communications networks, combined with the availability of increasingly powerful hardware platforms, is providing the foundation for advanced communications and multimedia services to be deployed faster and easier than ever before.
[0003] The Third Generation Partnership Project (3GPP) Enhanced Voice Services (EVS) codecs have brought very significant improvements in user experience with the introduction of super-wideband (SWB) and full-band (FB) speech and audio coding, along with improved packet loss resilience. However, expanded audio bandwidth is only one dimension required for a truly immersive experience. Immersing users in convincing virtual worlds in a resource-efficient manner ideally requires support beyond the mono and multi-mono currently offered by EVS.
[0004] Furthermore, while the audio codecs currently specified in 3GPP provide suitable quality and compression for stereo content, they lack the conversational features (e.g., sufficiently low latency) required for conversational voice and videoconferencing. These coders also lack the multi-channel capabilities required for immersive services such as live and user-generated content streaming, virtual reality (VR), and immersive videoconferencing.
[0005] Extensions to the EVS codec have been proposed for Immersive Voice and Audio Services (IVAS) to bridge this technology gap and meet the growing demand for rich multimedia services. Furthermore, videoconferencing applications in 4G / 5G will benefit from the IVAS codec being used as an improved speech coder supporting multi-stream coding (e.g., channel-, object-, and scene-based audio). Use cases for this next-generation codec include, but are not limited to, speech voice, multi-stream videoconferencing, VR conversations, and user-generated live and non-live content streaming. Summary of the Invention [Problem to be solved by the invention]
[0006] IVAS is thus expected to provide immersive as well as VR, AR, and / or XR user experiences. In many of these applications, the device capturing directional (immersive) audio (e.g., a mobile phone) often moves relative to the acoustic scene during the session, which can cause spatial rotation and / or translation of the captured audio scene. Depending on the type of experience being provided (e.g., immersive, VR, AR, or XR) and depending on the specific use case, this behavior may or may not be desirable. For example, if the rendered scene constantly rotates every time the capturing device is rotated, it may be annoying to the listener. In the worst case, motion sickness may occur.
[0007] Thus, improvements are needed in this context. [Brief explanation of the drawings]
[0008] Exemplary embodiments will now be described with reference to the accompanying drawings. [Figure 1] 1 illustrates a method for encoding directional audio, according to an embodiment. [Figure 2] 1 illustrates a method for rendering directional audio, according to an embodiment. [Figure 3] 2 shows an encoder device configured to perform the method of FIG. 1 according to an embodiment; [Figure 4] 3 illustrates a rendering device configured to perform the method of FIG. 2 according to an embodiment. [Figure 5] 5 illustrates a system comprising the apparatus of FIGS. 3 and 4, according to an embodiment. [Figure 6] 1 illustrates a physical VR conferencing scenario, according to an embodiment. [Figure 7] 1 illustrates a virtual meeting space, according to an embodiment.
[0009] All figures are schematic and generally show only parts necessary to explain the present disclosure, while other parts may be omitted or merely suggested. Unless otherwise noted, like reference numerals refer to like parts in different figures. DETAILED DESCRIPTION OF THE INVENTION
[0010] In view of the above, it is an object to provide an apparatus and associated methods for capturing, acoustically preprocessing, and / or encoding directional audio to compensate for undesired motion in a spatial sound scene that may result from unintended motion of a microphone system capturing the directional audio. It is also an object to provide a corresponding decoder and / or rendering apparatus and associated methods for decoding and rendering the directional audio. For example, a system including the encoder apparatus and the rendering apparatus is also provided.
[0011] I. Overview - Sender According to a first aspect, there is provided an apparatus comprising or connected to a microphone system including one or more microphones for capturing audio. The device (also referred to herein as the sender or capture device): receiving directional audio captured by a microphone system; receiving metadata related to the microphone system, the metadata including spatial data of the microphone system, the spatial data indicating a spatial orientation and / or a spatial location of the microphone system, including at least one from a list of azimuth angle, pitch angle, roll angle(s), and spatial coordinates of the microphone system. It has a receiving unit.
[0012] In this disclosure, the term "directional audio" (directional sound) generally refers to immersive audio, i.e., audio captured by a directional microphone system that can pick up sound including from the direction of arrival. The reproduction of directional audio allows for a natural three-dimensional sound experience (binaural rendering). Audio, which may include audio objects and / or channels (e.g., representing scene-based audio or channel-based audio in Ambisonics B format), is thus associated with the direction from which it is received. In other words, directional audio originates from a directional source and is incident from a direction of arrival (DOA), e.g., represented by an azimuth and elevation angle. In contrast, diffuse ambient sound is assumed to be omnidirectional, i.e., spatially invariant or spatially uniform. Other terms that may be used for the characteristic of "directional audio" include "spatial audio," "spatial sound," "immersive audio," "immersive sound," "stereo," and "surround audio."
[0013] In this disclosure, the term "spatial coordinates" generally refers to the spatial location of a microphone system or capture device in space. Cartesian coordinates are one realization of spatial coordinates. Other examples include cylindrical coordinates or spherical coordinates. It should be noted that the location in space may be relative (e.g., coordinates within a room or coordinates relative to other devices / units) or absolute (e.g., GPS coordinates, etc.).
[0014] In this disclosure, "spatial data" generally refers to either the current rotational orientation and / or spatial position of a microphone system, or a change in rotational orientation and / or spatial position compared to a previous orientation / position of the microphone system.
[0015] The device thus receives metadata including spatial data indicating the spatial orientation and / or spatial location of a microphone system capturing directional audio.
[0016] The apparatus further has a computing unit configured to modify at least a portion of the directional audio to generate modified directional audio, whereby directional characteristics of the audio are modified in response to a spatial orientation and / or a spatial position of the microphone system.
[0017] The modification can be performed using any suitable means, for example by defining a rotation / translation matrix based on the spatial data and multiplying the directional audio with this matrix to achieve modified directional audio. Matrix multiplication is suitable for non-parametric spatial audio. Parametric spatial audio may also be modified by adjusting spatial metadata, such as the directional parameters of sound objects.
[0018] The modified directional audio is then encoded into digital audio data, which is transmitted by a transmission unit of the apparatus.
[0019] The inventors have come to realize that the rotational / translational motion of the sound capture device (microphone system) is best compensated for at the transmitting end, i.e., at the end capturing the audio. This may likely allow for the best possible stabilization of the captured audio scene, e.g., with respect to unintended motion. Such compensation may be part of the capture process, i.e., during acoustic preprocessing or as part of the IVAS encoding stage. Furthermore, performing compensation at the transmitting end alleviates the need to transmit spatial data from the transmitting end to the receiving end. If compensation for the rotational / translational motion of the sound capture device were performed at the audio receiver, all spatial data would have to be transmitted to the receiving end. Assuming that the rotational coordinates for all three axes are each represented by 8 bits and estimated and transmitted at a rate of 50 Hz, the resulting bit rate is 1.2 kbps. A similar assumption can be made for the spatial coordinates of the microphone system.
[0020] According to some embodiments, the spatial orientation of the microphone system is represented in the spatial data by parameters describing the rotational movement / orientation of one degree of freedom DoF, e.g., for a conference call it may be sufficient to consider only the azimuth angle.
[0021] According to some embodiments, the spatial orientation of the microphone system is represented by parameters describing the rotational orientation / motion with three degrees of freedom DoF in the spatial data.
[0022] According to some embodiments, the spatial data of the microphone system is represented in 6 DoF, capturing the changed position of the microphone system as translations (referred to herein as spatial coordinates) of the microphone system in three perpendicular axes: forward / back (surge), up / down (heave), and left / right (sway), often combined with the change in orientation (or current rotational orientation) of the microphone system through rotations about three perpendicular axes, often referred to as yaw or azimuth (normal / vertical axis), pitch (horizontal axis), and roll (vertical axis).
[0023] According to some embodiments, the received directional audio includes audio that includes directional metadata. For example, such audio may include audio objects, i.e., object-based audio (OBA). OBA is a parametric form of spatial / directional audio with spatial metadata. One specific form of parametric spatial audio is metadata-assisted spatial audio (MASA).
[0024] According to some embodiments, the computing unit is further configured to encode at least a portion of the metadata, including spatial data of the microphone system, into the digital audio data. Advantageously, this allows compensation for directional adjustments made to the captured audio at the receiving end. Depending on the definition of a suitable rotating reference system, e.g., one in which the z-axis corresponds to the vertical direction, it may often be sufficient to transmit only the azimuth angle (e.g., at 400 bps). The pitch and roll angles of the capture device within the rotating reference system may only be required for certain VR applications. By compensating for the spatial data of the microphone system at the transmitting end and conditionally including at least a portion of the spatial data in the encoded digital audio data, cases in which the rendered sound scene should remain unchanged from the position of the capture device and other cases in which the rendered sound scene should rotate with the corresponding movement of the capture device are advantageously supported.
[0025] According to some embodiments, the receiving unit is further configured to receive a first instruction indicating to the computing unit whether to include the at least part of the metadata containing the spatial data of the microphone system in the digital audio data, so that the computing unit operates accordingly. As a result, the sender conditionally includes part of the spatial data in the digital audio data to save bitrate when possible. The instruction may be received multiple times during a session so that whether (part of) the spatial data should be included in the digital audio data changes over time. In other words, there may be intra-session adaptation, where the first instruction can be received by the device both continuously and discontinuously. Continuously would be, for example, once per frame. Discretely could be only once when a new instruction should be given. It is also possible to receive the first instruction only once during session setup.
[0026] According to some embodiments, the receiving unit is further configured to receive second instructions indicating to the computing unit which parameter(s) of the spatial data of the microphone system to include in the digital audio data, so that the computing unit acts accordingly. As mentioned above, the sender may be instructed to include only the azimuth angle or to include all data defining the spatial orientation of the microphone system. The command may be received multiple times during a session so that the number of parameters included in the digital audio data varies over time. In other words, there may be intra-session adaptation, and the second command may be received by the device both continuously and discontinuously. Continuously would be, for example, once per frame. Discretely could be only once when a new command is to be given. It is also possible to receive the second command only once during session setup.
[0027] According to some embodiments, the sending unit is configured to send the digital audio data to a further device, and the indication regarding the first and / or second instructions is received from the further device. In other words, the receiving side (including a renderer for rendering the received decoded audio) may instruct the sending side, depending on the context, whether to include part of the spatial data in the digital audio data and / or which parameters to include. In other embodiments, the indication regarding the first and / or second instructions may be received from, for example, a coordination unit (call server) for multi-user immersive audio / video conferencing or any other unit not directly involved in rendering directional audio.
[0028] According to some embodiments, the receiving unit is further configured to receive metadata including a timestamp indicating a capture time of the directional audio, and the computing unit is configured to encode said timestamp into said digital audio data. Advantageously, this timestamp can be used for synchronization at the receiving end, for example to synchronize audio rendering with video rendering or to synchronize multiple digital audio data received from different capture devices.
[0029] According to some embodiments, encoding the modified audio signal includes downmixing the modified directional audio, the downmixing being performed taking into account the spatial orientation of the microphone system, and encoding the downmix and a downmix matrix used in the downmixing into the digital audio data, for example, acoustic beamforming towards a particular directional source of the directional audio is advantageously adapted based on the directional modification made to the directional audio.
[0030] According to some embodiments, the device is implemented in virtual reality (VR) gear or augmented reality (AR) gear having the microphone system and a head tracking device configured to determine spatial data for the device with 3-6 degrees of freedom. In other embodiments, the device is implemented in a mobile phone having a microphone system.
[0031] II. Overview - Receiving Side According to a second aspect, there is provided an apparatus for rendering an audio signal. The apparatus (also referred to herein as a receiving side or a rendering apparatus) has a receiving unit configured to receive digital audio data. The apparatus further includes a decoding unit configured to decode the received digital audio data into directional audio and metadata, the metadata including spatial data including at least one from a list of azimuth, pitch, roll angle(s), and spatial coordinates. The spatial data may be received, for example, in the form of parameters, such as 3DoF angles. In other embodiments, the spatial data may be received as a rotation / translation matrix.
[0032] The apparatus further comprises: Correcting the directional characteristics of directional audio using rotational spatial data; Configured to render corrected directional audio With rendering units.
[0033] Advantageously, devices according to this aspect can modify directional audio as indicated in the metadata, for example, movement of the device capturing the audio may be taken into account during rendering.
[0034] According to some embodiments, the spatial data indicates the spatial orientation and / or spatial location of a microphone system including one or more microphones capturing directional audio, and the rendering unit modifies the directional characteristics of the directional audio to at least partially recreate the audio environment of the microphone system. In this embodiment, the device applies the acoustic scene rotation by reapplying at least a portion of the acoustic scene rotation compensated for by the capture device (relative acoustic scene rotation, i.e., scene rotation with respect to the moving microphone system).
[0035] According to some embodiments, the spatial data includes parameters describing rotational movement / orientation of one degree of freedom DoF.
[0036] According to some embodiments, the spatial data includes parameters describing rotational movement / orientation in three degrees of freedom DoF.
[0037] According to some embodiments, the decoded directional audio includes audio that includes directional metadata. For example, the decoded directional audio may include audio objects, i.e., object-based audio (OBA). In other embodiments, the decoded directional audio may be channel-based, representing, for example, scene-based audio or channel-based audio in Ambisonics B format.
[0038] According to some embodiments, the device includes a transmitting unit configured to transmit instructions to a further device from which the digital audio is received, the instructions indicating to the further device which parameter(s) (if any) the rotation data should include. Consequently, the rendering device may instruct the capture device to transmit, for example, only rotation parameters, only azimuth parameters, or all 6DoF parameters, depending on the use case and / or available bandwidth. Furthermore, the rendering device may make this decision based on the available computational resources in the renderer for applying the sound scene rotation or the level of complexity of the rendering unit. The instructions may be transmitted more than once during a session and thus may vary over time, i.e., based on the above. In other words, there may be intra-session adaptation, where the device can transmit the instructions both continuously and discontinuously. Continuously would be, for example, once per frame. Discretely could be only once when a new instruction should be given. There is also the possibility of transmitting the instructions only once during session setup.
[0039] According to some embodiments, the decoding unit is further configured to extract a timestamp from the digital audio data indicating the capture time of the directional audio, which timestamp may be used for synchronization reasons as discussed above.
[0040] According to some embodiments, decoding of the received digital audio data into directional audio by the decoding unit comprises: Decoding the received digital audio data into downmixed audio; and upmixing the downmixed audio into directional audio by a decoding unit using a downmix matrix included in the received digital audio data.
[0041] According to some embodiments, the spatial data includes spatial coordinates, and the rendering unit is further configured to adjust the volume of the rendered audio based on the spatial coordinates. In this embodiment, the volume of audio received from "far" locations may be attenuated relative to audio received from closer locations. It should be noted that the relative proximity of the received audio may be determined by applying a suitable distance metric, e.g., a Euclidean measure, based on a virtual space, and the position of the capturing device relative to the receiving device in this space is determined based on the spatial coordinates of the devices. A further step may include using any arbitrary mapping scheme to determine audio rendering parameters, such as sound level, from the distance metric. Advantageously, in this embodiment, the immersive experience of the rendered audio may be improved.
[0042] According to some embodiments, the device is implemented in virtual reality (VR) gear or augmented reality (AR) gear having a head-tracking device configured to measure the spatial orientation and spatial position of the device with 6 DoF. In this embodiment, the spatial data of the rendering device may also be used when modifying the directional characteristics of the directional audio. For example, the received rotation / translation matrix may be multiplied by a similar matrix defining the rotational state of the rendering device, and the resulting matrix may then be used to modify the directional characteristics of the directional audio. Advantageously, this embodiment may improve the immersive experience of the rendered audio. In other embodiments, the device is implemented in a teleconferencing device or similar device that is assumed to be stationary, and the rotational state of the device is ignored altogether.
[0043] According to some embodiments, the rendering unit is configured for binaural audio rendering.
[0044] III. Overview - System According to the third aspect: 1. A system having a first device according to a first aspect configured to transmit digital audio data to a second device according to a second aspect, the system being configured for audio and / or video conferencing. is provided.
[0045] According to some embodiments, the first device further comprises a video recording unit and is configured to encode the recorded video into digital video data and transmit the digital video data to the second device, the second device further comprising a display for displaying the decoded digital video data.
[0046] According to the fourth aspect: 1. A system comprising a first device according to a first aspect configured to transmit digital audio data to a second device, the second device comprising: a receiving unit configured to receive digital audio data; a decoding unit configured to decode the received digital audio data into directional audio and metadata, the metadata including spatial data including at least one from a list of azimuth angle, pitch, roll angle(s), and spatial coordinates; It has a rendering unit for rendering audio and The rendering unit, when the second device further receives encoded video data from the first device: modifying directional characteristics of directional audio using the spatial data; configured to render modified directional audio, The rendering unit, when the second device does not receive encoded video data from the first device: configured to render the directional audio; A system is provided.
[0047] Advantageously, the decision of whether to recreate the audio environment of a microphone system by compensating for the spatial orientation and / or spatial position of the microphone system is made based on whether video is being transmitted. In this embodiment, the transmitting device may not always know when that motion compensation is necessary or desirable. For example, consider a situation where audio is rendered alongside video. In that case, it may be advantageous to be able to rotate the audio scene along with a moving visual scene or keep the audio scene stable, at least when video capture is performed on the same device as capturing the audio. If video is not being consumed, keeping the audio scene stable by compensating for the motion of the capture device may be the preferred option.
[0048] According to a fifth aspect, there is provided a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the operations of aspects 1 to 4.
[0049] IV. Overview - General The second through fifth aspects may generally have the same or corresponding features and advantages as the first aspect. Other objects, features and advantages of the present invention will become apparent from the following detailed disclosure, the attached dependent claims and the drawings. The steps of any method, or apparatus implementing a sequence of steps, disclosed herein do not have to be performed in the exact order disclosed, unless explicitly stated.
[0050] V. Illustrative Embodiments Immersive audio and sound services are expected to provide immersive and virtual reality (VR) user experiences. Augmented reality (AR) and extended reality (XR) experiences may also be provided. This disclosure addresses the fact that mobile devices, such as handheld UEs, capturing immersive or AR / VR / XR scenes may often be moving relative to the audio scene during a session. This highlights cases where the rotational movement of the capturing device should be avoided from being reproduced as a corresponding rendered scene rotation by the receiving device. This disclosure relates to how this can be efficiently handled to meet the requirements that users have for immersive audio, depending on the context.
[0051] Although some examples herein are described in the context of an IVAS encoder, decoder, and / or renderer, it should be noted that this is merely one type of encoder / decoder / renderer to which the general principles of the present invention are applicable, and that there may be many other types of encoders, decoders, and renderers that may be used in conjunction with the various embodiments described herein.
[0052] It should also be noted that although the terms "upmix" and "downmix" are used throughout this document, they do not necessarily mean an increase and a decrease in the number of channels, respectively. While this is often the case, both terms can also mean either a decrease or an increase in the number of channels. As such, both terms fall under the more general concept of "mix."
[0053] Referring now to FIG. 1, a method 1 for encoding and transmitting a representation of directional audio is described according to one embodiment.
[0054] A device 300 configured to perform Method 1 is shown in FIG. 3. Device 300 may generally be a mobile phone (smartphone), but may also be part of a VR / AR / XR installation or any other type of device having or connected to a microphone system 302 having one or more microphones for capturing directional audio. Thus, device 300 may have microphone system 302 or may be connected (wired or wirelessly) to a remote microphone system 302. In some embodiments, device 300 is implemented in VR or AR gear having microphone system 302 and a head-tracking device configured to determine spatial data for the device with 1 to 6 DoF.
[0055] In some audio capture scenarios, the position and / or spatial orientation of the microphone system 302 may change during directional audio capture.
[0056] Two example scenarios are now described.
[0057] Changes in the position and / or spatial orientation of the microphone system 302 during audio capture can cause a spatial rotation / translation of the rendered scene at the receiving device. Depending on the type of experience being provided (e.g., immersive, VR, AR, or XR), and depending on the particular use case, this behavior may or may not be desirable. One example where this may be desirable is when a service additionally provides a visual component, such as when a capture camera (e.g., 360-degree video capture, not shown in FIG. 1) and microphone 302 are integrated into the same device. In that case, rotation of the capture device should be expected to result in a corresponding rotation of the rendered audiovisual scene.
[0058] On the other hand, if audiovisual capture is not performed by the same physical device, or if there is no video component, the rotation of the rendered scene with each rotation of the capture device can be distracting to the listener. In the worst case, it can cause motion sickness. Therefore, it is desirable to compensate for position changes (translation and / or rotation) of the capture device. Examples include immersive telephony and immersive conferencing applications that use a smartphone as the capture device (i.e., one that includes a set of microphones 302). In these use cases, it can frequently happen that the set of microphones is unintentionally moved due to handholding or contact by the user during operation. Capture device users may not be aware that moving the capture device can cause instabilities in the spatial audio rendered at the receiving device. In general, users cannot be expected to hold the phone stationary during a conversation.
[0059] The methods and apparatus described below are directed to some or all of the above scenarios.
[0060] Thus, device 300 has or is connected to a microphone system 302 that includes one or more microphones for capturing audio. Thus, the microphone system may include 1, 2, 3, 5, 10, etc. microphones. In some embodiments, the microphone system includes multiple microphones. Device 300 has multiple functional units. These units may be implemented in hardware and / or software and may have one or more processors to handle their functionality.
[0061] The apparatus 300 comprises a receiving unit 304 configured to receive (S13) directional audio 320 captured by the microphone system 302. The directional audio 320 is preferably an audio representation that easily allows for rotation and / or translation of the audio scene. The directional audio 320 may for example include audio objects and / or channels that allow for rotation and / or translation of the audio scene. The directional audio may include: Channel-based audio (CBA), e.g., stereo, multichannel / surround, 5.1, 7.1, etc. Scene-based audio (SBA), e.g., 1st and higher-order Ambisonics Object-based audio (OBA).
[0062] CBA and SBA are non-parametric forms of spatial / directional audio, while OBA is parametric with spatial metadata. One specific form of parametric spatial audio is Metadata-Assisted Spatial Audio (MASA).
[0063] The receiving unit 304 is further configured to receive (S14) metadata 322 associated with the microphone system 302. The metadata 322 includes spatial data of the microphone system 302. The spatial data indicates the spatial orientation and / or spatial location of the microphone system 302. The spatial data of the microphone system includes the azimuth, pitch, roll angle(s) of the microphone system, and at least one from a list of spatial coordinates. The spatial data may be expressed in one degree of freedom, DoF (e.g., only the azimuth angle of the microphone system), 3DoF (e.g., the spatial orientation of the microphone system with 3DoF), or 6DoF (both the spatial orientation with 3DoF and the spatial location with 3DoF). The spatial data may, of course, be expressed in any number of DoFs, from 1 to 6.
[0064] The apparatus 300 further comprises a computing unit 306 that receives the directional audio 320 and the metadata 322 from the receiving unit 304 and modifies (S15) at least a portion of the directional audio 320 (e.g., at least some of the audio objects of the directional audio) to generate modified directional audio, which results in a modified directional characteristic of the audio depending on the spatial orientation and / or spatial position of the microphone system.
[0065] The computing unit 306 then encodes (S16) the digital data by encoding (S17) the modified directional audio into digital audio data 328. The apparatus 300 further comprises a transmission unit 310 configured to transmit (wired or wirelessly) the digital audio data 328, for example as a bitstream.
[0066] Compensating for the rotational and / or translational motion of microphone system 302 already at encoding device 300 (sometimes referred to as source device, capture device, transmitting device, or sender) reduces the requirements for transmitting the spatial data of microphone system 302. If such compensation were performed by the device receiving the encoded directional audio (e.g., an immersive audio renderer), all required metadata would need to be included in digital audio data 328 at all times. Assuming that the rotational coordinates of microphone system 302 in all three axes are represented by 8 bits each and estimated and transmitted at a rate of 50 Hz, the resulting increase in bit rate of signal 332 would be 1.2 kbps. Furthermore, auditory scene variations in the absence of motion compensation at the capture side can make spatial audio coding more demanding and potentially less efficient.
[0067] Furthermore, since the information on which the correction decisions are based is readily available in the device 300, it is now appropriate and can be done efficiently to compensate for the rotational / translational movements of the microphone system 302. Thus, the maximum algorithmic delay for this operation can be reduced.
[0068] Yet another advantage is that by always compensating for rotational / translational motion in the capture device 300 (rather than conditionally on demand) and conditionally providing spatial orientation data of the capture system to the receiving end, potential collisions are avoided when multiple endpoints with different rendering needs are served, such as in multi-party conferencing use cases.
[0069] The above covers all cases where the rendered sound scene should be invariant with the position and rotation of the microphone system 302 capturing directional audio. To address the remaining cases where the rendered sound scene should rotate with the corresponding movement of the microphone system 302, the computing unit 306 may optionally be configured to encode (S18) at least a portion of the metadata 322 containing spatial data of the microphone system into the digital audio data 328. Depending on the definition of a preferred rotating reference frame, e.g., where the z-axis corresponds to the vertical direction, it may often be sufficient to simply transmit the azimuth angle (e.g., at 400 bps). The pitch and roll angles of the microphone system 302 within the rotating reference frame may only be required in certain VR applications.
[0070] Conditionally provided rotation / translation parameters may typically be transmitted as a single conditional element in the IVAS RTP payload format, and thus these parameters require a small fraction of the allocated bandwidth.
[0071] To accommodate these different scenarios, the receiving unit 304 may optionally be configured to receive (S10) instructions on how to handle the metadata 322 when the computing unit 306 is encoding the digital audio data 328. The instructions may be received (S10) from a rendering device (e.g., another part of the audio conference) or from a coordinating device such as a call server.
[0072] In some embodiments, the receiving unit 304 is further configured to receive (S11) a first instruction indicating to the computing unit 306 whether to include in the digital audio data the at least part of the metadata 322, including spatial data of a microphone system. In other words, the first instruction informs the device 300 whether any of the metadata or no metadata should be included in the digital audio data 328. For example, if the device 300 is transmitting the digital audio data 328 as part of an audio conference, the first instruction may specify that no part of the metadata 322 should be included.
[0073] Alternatively or additionally, in some embodiments, the receiving unit 304 is further configured to receive second instructions indicating to the computing unit which parameter(s) of the microphone system's spatial data to include in the digital audio data 328, so that the computing unit operates accordingly. For example, for bandwidth reasons or other reasons, the second instructions may specify to the computing unit 306 to include only the azimuth angle in the digital audio data 328.
[0074] The first and / or second instructions may typically be the subject of session setup negotiation, so none of these instructions would require transmission during the session, e.g., none of the allocated bandwidth for immersive audio / video conferencing.
[0075] As mentioned above, the device 300 may be part of a video conference. To this end, the receiving unit 304 may be further configured to receive metadata (not shown in FIG. 1 ) including a timestamp indicating the capture time of the directional audio, and the computing unit 306 is configured to encode said timestamp into said digital audio data. Advantageously, the modified directional audio may then be synchronized with the captured video at the rendering side.
[0076] In some embodiments, encoding S17 the modified directional audio includes downmixing the modified directional audio, which is performed by taking into account the spatial orientation of the microphone system 302 and encoding the downmix and a downmix matrix used in the downmix into the digital audio data 328. Downmixing may include, for example, adjusting the beamforming operation of the directional audio 320 based on the spatial orientation of the microphone system 302.
[0077] Thus, the digital audio data is transmitted (S19) from device 300, e.g., as the transmitting portion of an immersive audio / video conferencing scenario. The digital audio data is then received by a device for rendering an audio signal, e.g., as the receiving portion of an immersive audio / video conferencing scenario. Rendering device 400 will now be described with reference to Figures 2 and 4.
[0078] The device 400 for rendering an audio signal comprises a receiving unit 402 configured to receive (S21) (wired or wireless) digital audio data 328.
[0079] The device 400 further has a decoding unit 404 configured to decode (S22) the received digital audio data 328 into directional audio 420 and metadata 422, the metadata 422 including spatial data including azimuth angle, pitch, roll angle(s), and at least one from a list of spatial coordinates.
[0080] In some embodiments, the upmixing is performed by the decode unit 404. In these embodiments, the decoding of the received digital audio data 328 into directional audio 420 by the decode unit 404 includes: decoding the received digital audio data 328 into downmixed audio, and upmixing the downmixed audio into directional audio 420 by the decode unit 404 using a downmix matrix included in the received digital audio data 328.
[0081] The apparatus further includes a rendering unit 406 configured to modify directional characteristics of the directional audio using the spatial data (S23) and render the modified directional audio 424 using speakers or headphones (S24).
[0082] Thus, the apparatus 400 (its rendering unit 406) is configured to apply an audio scene rotation / translation based on the received spatial data.
[0083] In some embodiments, the spatial data indicates the spatial configuration and / or spatial location of a microphone system including one or more microphones capturing directional audio, and the rendering unit modifies the directional characteristics of the directional audio to at least partially recreate the audio environment of the microphone system (S23). In this embodiment, the apparatus 400 reapplies at least a portion of the acoustic scene rotation compensated at the capture end by the apparatus 300 of FIG.
[0084] The spatial data may include spatial data including rotational data representing movement in three degrees of freedom DoF. Alternatively or additionally, the spatial data may include spatial coordinates.
[0085] The decoded directional audio may, in some embodiments, include audio objects, or more generally, audio associated with spatial metadata, as described above.
[0086] The decoding S22 of the received digital audio data by the decoding unit 404 into directional audio may, in some embodiments, include decoding the received digital audio data into downmixed audio and upmixing the downmixed audio into directional audio by the decoding unit 404 using a downmix matrix included in the received digital audio data 328. To provide increased flexibility and / or to meet bandwidth requirements, the apparatus 400 may include a sending unit 306 configured to send (S20) instructions to a further device from which the digital audio data 328 is received, the instructions indicating to the further device which parameter(s), if any, the rotation or translation data should include. This functionality may thus facilitate meeting potential user preferences or preferences related to the type of rendering and / or service used.
[0087] In some embodiments, device 400 may be configured to send instructions to the further device indicating whether metadata including spatial data is to be included in digital audio data 328. In these embodiments, if the received S21 digital audio data 328 does not include such metadata, the rendering unit renders the decoded directional audio as received (possibly upmixed as described above) without any modification of the directional characteristics of the directional audio due to compensation made in capture device 300. However, in some embodiments, the received directional audio is modified in response to head tracking information of the renderer (described below).
[0088] The device 400 may, in some embodiments, be implemented in VR or AR gear with a head tracking device configured to measure the spatial orientation of the device with 6 DoF. The rendering unit 406 may be configured for binaural audio rendering.
[0089] In some embodiments, the rendering unit 406 is configured to adjust (S25) the volume of the rendered audio based on the spatial coordinates received in the metadata. This functionality is further described below in connection with FIGS. 6-7. 5 illustrates a system including a capture device 300 (described in connection with FIG. 3) and a rendering device 400 (described in connection with FIG. 4). In some embodiments, the capture device 300 may receive (S10) instructions 334 transmitted (S20) from the rendering device 400 indicating whether and to what extent the capture device 300 should include spatial data of the capture device's microphone system in the digital audio data 328.
[0090] In some embodiments, the capture device 300 further comprises a video recording unit and is configured to encode the recorded video into digital video data 502 and transmit the digital video data to the rendering device 400, which further comprises a display for displaying the decoded digital video data.
[0091] As mentioned above, changes in the position and / or spatial orientation of the microphone system of the capture device 300 during audio capture can cause a spatial rotation / translation of the rendered scene at the rendering device 400. Depending on the type of experience being provided (e.g., immersive, VR, AR, or XR), and depending on the particular use case, this behavior may or may not be desirable. One example where this may be desirable is when a service additionally provides a visual component 502, where the capture camera and the one or more microphones 302 are integrated into the same device. In that case, rotation of the capture device 300 would be expected to result in a corresponding rotation of the rendered audiovisual scene at the rendering device 400.
[0092] On the other hand, if the audiovisual capture is not performed by the same physical device, or if there is no video component, the rotation of the rendered scene every time the capture device 300 rotates can be distracting to the listener, and in the worst case scenario, can cause motion sickness.
[0093] For this reason, according to some embodiments, the rendering unit of the rendering device 400 may be configured to use the spatial data to modify the directional characteristics of the directional audio (received in the digital audio data 328) and render the modified directional audio when the rendering device 400 further receives the encoded video data 502 from the capture device 300.
[0094] However, when rendering device 400 does not receive encoded video data from capture device 300, the rendering unit of rendering device 400 may be configured to render directional audio without direction modification.
[0095] In another embodiment, rendering device 400 is informed prior to the conference that the data received from capture device 300 will not include a video component. In this case, rendering device 400 may indicate in instruction 334 that spatial data of the microphone system of capture device 300 does not need to be included in digital audio data 328, thereby configuring the rendering unit of rendering device 400 to render directional audio received in digital audio data 328 without direction correction.
[0096] The above briefly outlines downmixing and / or encoding of directional audio on a capture device, which will now be described in more detail.
[0097] In many cases, the capture device 300 does not have information about whether the decoded presentation (at the rendering device) is intended for a single mono speaker, stereo speakers, or headphones. The actual rendering scenario may also vary during a service session, which may change, for example, with connected playback equipment, such as connecting or disconnecting headphones to a mobile phone. Yet another scenario in which the capabilities of the rendering device are unknown is when a single capture device 300 needs to support multiple endpoints (rendering devices 400). For example, in an IVAS conferencing or VR content distribution use case, one endpoint may be using a headset and another may be rendering to stereo speakers, and yet it is advantageous to be able to feed both endpoints with a single encode. This reduces the complexity on the encoding side and may also reduce the overall network bandwidth required.
[0098] A less desirable but straightforward way to support these cases would be to always assume the lowest receiving device capability, i.e., mono, and select the corresponding audio operating mode. However, a more reasonable approach would be to require that the codec used (e.g., the IVAS codec) always be capable of producing a decoded audio signal that can be presented on a device 400 with the respective lower audio capability, even if it is operated in a presentation mode that supports spatial, binaural, or stereo audio. In some embodiments, a signal encoded as a spatial audio signal may be decodable for binaural, stereo, and / or mono rendering. Similarly, a signal encoded as binaural may be decodable as stereo or mono, and a signal encoded as stereo may be decodable for mono presentation. Illustratively, the capture device 300 may implement a single encoding (digital audio data 328) and simply transmit the same encoding to multiple endpoints 400. Some of the multiple endpoints 400 may support binaural presentation, and some may be stereo only.
[0099] It should be noted that the codecs discussed above can be implemented in a capture device or in a call server. In the case of a call server, the call server receives digital audio data 328 from a capture device, transcodes the digital audio data to meet the above requirements, and then transmits the transcoded digital audio data to the one or more rendering devices 400. Such a scenario is now illustrated with reference to FIG. 6.
[0100] A physical VR conference scenario 600 is shown in Figure 6. Five VR / AR conference users 602a-e from different sites are virtually meeting. The VR / AR conference users 602a-e may be IVAS-enabled. Each user uses VR / AR gear, including binaural playback and video playback, for example, using an HMD. All users' equipment supports 6DOF movement with corresponding head tracking. The users' user equipment (UE) 602 exchanges encoded audio with a conference call server 604 via uplink and downlink. Visually, the users may be represented through their respective avatars, which can be rendered based on information related to their relative position parameters and rotational orientation.
[0101] To further improve the immersive user experience, the rotational and / or translational movement of the listener's head is also taken into account when rendering audio received from other participant(s) in a conference scenario. As a result, head tracking informs the rendering unit of the user's rendering device (reference numeral 400 in FIGS. 4-5 ) of the user's current spatial data (6 DOF) of the VR / AR gear. This spatial data is combined with spatial data received in digital audio data received from another user 602 (e.g., through matrix multiplication or modification of metadata associated with the directional audio), causing the rendering unit to modify the directional characteristics of the directional audio received from the other user 602 based on the combination of spatial data. The modified directional audio is then rendered to the user.
[0102] Additionally, the volume of rendered audio received from a particular user may be adjusted based on the spatial coordinates received in the digital audio data. Based on the virtual (or real) distance between two users (as calculated by the rendering device or the call server 604), the volume may be increased or decreased to further improve the immersive user experience.
[0103] 7 shows an example of a virtual conference space 700 created by a conference call server. Initially, the server assigns conference users Ui, i=1...5 (also referred to as 702a-e) to virtual location coordinates K i =(x i ,y i ,z i ) The virtual conference space is shared between users. Thus, the audiovisual rendering for each user is done in that space. For example, from the perspective of user U5 (corresponding to user 602d in Figure 6), the rendering effectively places the other conference participants at relative positions K i -K5, i≠5. For example, user U5 places user U2 at a distance |K i -K5|, the vector (K i -K5) / |K i -K5|, and thus directional rendering is done for U5's rotational position. Figure 2 also shows U5's movement toward U4. This movement affects U5's position relative to the other users, which is taken into account during rendering. Simultaneously, U5's UE transmits its changing position to the conference server 604, which updates the virtual conference space with U5's new coordinates. Because the virtual conference space is shared, users U1-U4 are aware of the moving user U5 and can adapt their respective renderings accordingly. The simultaneous movement of user U2 works according to a corresponding principle. The call server 604 is configured to maintain position data for participants 702a-e in the shared conference space.
[0104] In the scenario of Figures 6-7, with regard to audio, one or more of the following 6DOF requirements may be applied to the coding framework: Providing a metadata framework for the representation and upstream transmission of receiving endpoint location information, including spatial and / or rotational coordinates (as described above in connection with Figures 1-4). Ability to associate input audio elements (e.g. objects) with 6DOF attributes including spatial coordinates, rotational coordinates and directionality. · Capability of simultaneous spatial rendering of multiple received audio elements, each with associated 6DOF attributes. · Sufficient adjustment of the rendered scene to the rotation and translation of the listener's head.
[0105] It should be noted that the above also applies to XR meetings, which are a blend of physical and virtual meetings. Physical participants see and hear avatars representing remote participants through AR glasses and headphones. Participants interact with those avatars in the discussion as if they were physically present participants. For them, interactions with other physical and virtual participants occur in mixed reality. The locations of the real and virtual participants are merged (e.g., by the call server 604) into a combined shared virtual meeting space that matches the locations of the real participants in the physical meeting space and is mapped into the virtual meeting space using absolute and relative physical / real location data.
[0106] In VR / AR / XR scenarios, virtual conference subgroups may be formed. These subgroups may be used to inform the call server 604 of which users should receive a higher quality of service (QoS), for example, and which users may receive a lower QoS. In some embodiments, the virtual environment provided to these subgroups via VR / AR / XR gear includes only participants from the same subgroup. For example, a scenario in which subgroups may be formed is a poster session offering virtual participation from remote locations. Remote participants are equipped with an HMD and headphones. They virtually exist and can walk from poster to poster. They can listen to the ongoing poster presentation and approach a presentation if they find the topic or ongoing discussion interesting. To improve the potential for immersive interaction between virtual and physical participants, subgroups may be formed based, for example, on which of the multiple posters participants are currently interested in.
[0107] An embodiment of this scenario includes: · Receiving topics from virtual meeting participants via remote conferencing systems; · Grouping participants into subgroups for virtual meetings based on topics through remote conferencing systems; receiving, by the teleconferencing system, a request from a new participant's device to join the virtual conference, the request being associated with an indicator of a preferred topic; selecting, by a teleconferencing system, a subgroup from among the subgroups based on the preferred topic and the topics of the subgroups; The teleconferencing system provides a virtual environment for the virtual conference on a device of the new participant, the virtual environment indicating at least one of a visual virtual proximity and an auditory virtual proximity between the new participant and one or more participants of a selected subgroup.
[0108] In some embodiments, the virtual environment indicates visual or auditory virtual proximity by providing a virtual reality display or a virtual reality sound field in which at least the avatar of the new participant and one or more avatars of the participants of the selected subgroup are in close proximity to each other.
[0109] In some embodiments, each participant is connected by open headphones and AR glasses.
[0110] VI. Equivalents, Extensions, Substitutions and Other Further embodiments of the present disclosure will be apparent to those skilled in the art after reviewing the above description. While the description and drawings disclose embodiments and examples, the present disclosure is not limited to such specific examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure, which is defined by the appended claims. Any reference signs appearing in the claims should not be construed as limiting the scope thereof.
[0111] Moreover, variations to the disclosed embodiments can be understood and implemented by those skilled in the art in practicing the present disclosure, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprises" does not exclude other elements or steps, and the word "a" or "an" does not exclude a plurality. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
[0112] The systems and methods disclosed above may be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks among functional units referred to in the above description does not necessarily correspond to a division into physical units. Conversely, a single physical component may have multiple functions, and a single task may be performed by several cooperating physical components. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or application-specific integrated circuits. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, those skilled in the art will appreciate that communication media typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0113] All figures are schematic and generally show only parts necessary to explain the present disclosure, while other parts may be omitted or may only be suggested. Unless otherwise noted, like reference signs refer to like parts in different figures.
[0114] [Aspect 1] 1. A device having or connected to a microphone system including one or more microphones for capturing audio, the device comprising: A receiving unit comprising: receiving directional audio captured by the microphone system; a receiving unit configured to perform the steps of: receiving metadata related to the microphone system, the metadata including spatial data of the microphone system, the spatial data indicating a spatial orientation and / or a spatial position of the microphone system, the spatial data including at least one from a list of an azimuth angle, a pitch angle, a roll angle, and a spatial coordinate of the microphone system; modifying at least a portion of the directional audio to generate modified directional audio, whereby directional characteristics of the audio are modified in response to a spatial orientation and / or spatial position of the microphone system; encoding the modified directional audio into digital audio data; and a transmitting unit configured to transmit the digital audio data; Device. [Aspect 2] 2. The apparatus of claim 1, wherein the spatial orientation of the microphone system is represented in the spatial data by parameters describing a rotational movement / orientation of one degree of freedom (DoF). Aspect 3 2. The apparatus of claim 1, wherein the spatial orientation of the microphone system is represented in the spatial data by parameters describing 3DoF rotational movement / orientation. Aspect 4 4. The apparatus of any one of aspects 1 to 3, wherein the spatial data of the microphone system is represented in 6 DoF. Aspect 5 5. The apparatus of any one of aspects 1-4, wherein the received directional audio includes audio that includes directional metadata. Aspect 6 6. The apparatus of any one of aspects 1 to 5, wherein the computing unit is further configured to encode at least a portion of the metadata, including spatial data of the microphone system, into the digital audio data. Aspect 7 The apparatus of aspect 6, wherein the receiving unit is further configured to receive a first instruction indicating to the computing unit whether to include at least a portion of the metadata including spatial data of the microphone system in the digital audio data, and the computing unit operates accordingly. Aspect 8 The apparatus of aspect 6 or 7, wherein the receiving unit is further configured to receive a second instruction indicating to the computing unit which parameter(s) of the spatial data of the microphone system to include in the digital audio data, and the computing unit operates accordingly. Aspect 9 The apparatus of aspect 7 or 8, wherein the transmitting unit is configured to transmit the digital audio data to a further device (400), and an indication regarding the first instruction and / or the second instruction is received from the further device. Aspect 10 10. The apparatus of any one of aspects 1 to 9, wherein the receiving unit is further configured to receive metadata including a timestamp indicating a capture time of the directional audio, and the computing unit is configured to encode the timestamp into the digital audio data. Aspect 11 11. The apparatus of any one of aspects 1 to 10, wherein encoding the modified directional audio includes downmixing the modified directional audio, the downmix being performed by encoding a downmix and a downmix matrix used in the downmixing into the digital audio data, taking into account the spatial orientation of the microphone system. Aspect 12 12. The apparatus of embodiment 11, wherein the downmixing includes beamforming. Aspect 13 13. The device of any one of aspects 1 to 12, implemented in virtual reality (VR) gear or augmented reality (AR) gear having the microphone system and a head tracking device configured to determine spatial data of the device with 3 to 6 DoF. Aspect 14 1. An apparatus for rendering an audio signal, the apparatus comprising: a receiving unit configured to receive digital audio data; a decoding unit configured to decode the received digital audio data into directional audio and metadata, the metadata including spatial data including at least one from a list of an azimuth angle, a pitch angle, a roll angle, and a spatial coordinate; modifying directional characteristics of the directional audio using the spatial data; Configured to render corrected directional audio Having a rendering unit and Device. Aspect 15 The device of claim 14, wherein the spatial data indicates the spatial orientation and / or spatial position of a microphone system including one or more microphones that capture the directional audio, and the rendering unit modifies the directional characteristics of the directional audio to at least partially reproduce the audio environment of the microphone system. Aspect 16 16. The apparatus of claim 14 or 15, wherein the spatial data includes parameters describing rotational movement / orientation in one degree of freedom DoF. Aspect 17 16. The apparatus of claim 14 or 15, wherein the spatial data includes parameters describing rotational movement / orientation in 3 DoF. Aspect 18 18. The apparatus of any one of aspects 14-17, wherein the decoded directional audio includes audio that includes directional metadata. Aspect 19 19. The apparatus of any one of aspects 14 to 18, further comprising a transmitting unit configured to transmit instructions to a further device (300) from which the digital audio is received, the instructions indicating to the further device which parameter(s) the rotation data should include. Aspect 20 20. The apparatus of any one of aspects 14 to 19, wherein the decoding unit is further configured to extract a timestamp from the digital audio data indicating a capture time of the directional audio. Aspect 21 The decoding unit may decode the received digital audio data into directional audio by: Decoding the received digital audio data into downmixed audio; upmixing, by the decoding unit, the downmixed audio into the directional audio using a downmix matrix included in the received digital audio data. 21. The apparatus of any one of aspects 14 to 20. Aspect 22 22. The apparatus of any one of aspects 14 to 21, wherein the spatial data includes spatial coordinates, and the rendering unit is further configured to adjust a volume of rendered audio based on the spatial coordinates. Aspect 23 23. The device of any one of aspects 14 to 22, implemented in virtual reality (VR) gear or augmented reality (AR) gear having a head tracking device configured to measure the spatial orientation and spatial position of the device in 6 DoF. Aspect 24 24. The apparatus of any one of aspects 14 to 23, wherein the rendering unit is configured for binaural audio rendering. Aspect 25 A system having a first device of any one of aspects 1 to 13 configured to transmit digital audio data to a second device of any one of aspects 14 to 24, the system being configured for audio and / or video conferencing. Aspect 26 The system further includes a video recording unit configured to encode the recorded video into digital video data and transmit the digital video data to the second device, the second device further including a display for displaying the decoded digital video data. Aspect 27 14. A system comprising a first device of any one of aspects 1 to 13 configured to transmit digital audio data to a second device, wherein the second device: a receiving unit configured to receive digital audio data; a decoding unit configured to decode the received digital audio data into directional audio and metadata, the metadata including spatial data including at least one from a list of an azimuth angle, a pitch angle, a roll angle, and a spatial coordinate; It has a rendering unit for rendering audio and The rendering unit, when the second device further receives encoded video data from the first device: modifying the directional characteristics of the directional audio using the spatial data; Configured to render corrected directional audio; The rendering unit, when the second device does not receive encoded video data from the first device: configured to render the directional audio. system. Aspect 28 28. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations recited in any one of aspects 1-27.
Claims
1. 1. A method for encoding spatial audio, the method comprising: obtaining directional audio captured by one or more microphones of a microphone system; acquiring metadata associated with the microphone system, the metadata including spatial data indicating a spatial orientation and / or a spatial location of the microphone system, the spatial data including at least one from a list of an azimuth angle, a pitch angle, a roll angle, and a spatial coordinate of the microphone system; modifying the directional audio to generate modified directional audio, wherein directional characteristics of the directional audio are modified in response to the metadata to generate the modified directional audio, and modifying the directional characteristics of the directional audio in response to the metadata to at least partially compensate for rotational or translational movement of the microphone system; encoding the modified directional audio and at least a portion of the metadata into digital audio data. method.
2. The method of claim 1 , wherein encoding comprises encoding the digital audio data for Immersive Voice and Sound Services (IVAS) compliance.
3. 2. The method of claim 1, further comprising the step of outputting said digital audio data by storing said digital audio data.
4. The method of claim 1 , wherein the spatial orientation of the microphone system is represented in the spatial data using parameters describing rotational movement / orientation in 3 DoF.
5. The method according to claim 1 or 4, wherein the spatial data of the microphone system is expressed in 6 DoF.
6. 6. The method of claim 1, further comprising receiving instructions on how to handle the metadata for the encoding.
7. 7. The method of claim 2, further comprising receiving a first instruction indicating whether to include the at least part of the metadata comprising spatial data of the microphone system in the digital audio data.
8. 8. The method of claim 2, further comprising receiving a second instruction indicating which parameter(s) of spatial data of the microphone system to include in the digital audio data.
9. 8. The method of claim 7, further comprising transmitting the digital audio data to a further device, wherein an indication regarding the first command is received from the further device.
10. 9. The method of claim 8, further comprising transmitting the digital audio data to a further device, wherein the indication regarding the second command is received from the further device.
11. 11. The method of claim 2, wherein receiving metadata comprises receiving metadata including a timestamp indicating a capture time of the directional audio, and wherein encoding comprises encoding the timestamp in the digital audio data.
12. 12. The method according to any one of claims 2 to 11, wherein the method is performed by a device including a microphone system with one or more microphones and a head tracking system, the device including a virtual reality (VR) gear (602a-e) or an augmented reality (AR) gear (602a-e), and the head tracking system is configured to determine spatial data of the device in 3 to 6 DoF.
Citation Information
Patent Citations
Method and apparatus for improving the rendering of multichannel audio signals
JP2015527610A
Encoding Higher Order Ambisonic Audio Data Using Motion Stabilization
JP2018511070A
Encoding device and method, decoding device and method, and program
WO2014192602A1
Merging audio signals with spatial metadata
WO2017182714A1