Audio signal capture
By introducing a control data processing device for adjusting the sound capture beam in the audio capture device, the problem of poor experience of users when listening to the output audio signals of multiple physical speakers is solved, and a higher sensitivity to the audio signals of multiple speaker directions is achieved.
Patent Information
- Application Number
- CN202411750055.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-24
AI Technical Summary
Users wearing specific audio capture devices may not have the best user experience when listening to audio signals output by two or more physical speakers.
An apparatus is provided, including means for receiving audio data, means for determining the direction of sound source in an audio signal, and means for sending control data to an audio capture device. The device is able to adjust the sound capture beam when the audio capture device is in directivity mode to improve the sensitivity of the audio signal to a particular physical speaker direction.
By adjusting the sound capture beam, the audio capture device is increased sensitivity to audio signals from the directions of two or more physical speakers, thereby improving the user experience.
Smart Images

Figure CN120201346A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments relate to audio signal capture, such as in the case where an audio capture device captures an audio signal output or intended to be output using two or more physical speakers. Background Art
[0002] Certain audio signal formats are suitable for output by two or more physical speakers. Such audio signal formats may include stereo, multi-channel, and immersive formats. By using two or more physical speakers to output an audio signal, a listening user can perceive one or more sound objects as coming from a specific direction different from the direction of the physical speakers.
[0003] A user wearing a particular audio capture device may not obtain an optimal user experience when listening to an audio signal output by two or more physical speakers. Summary of the Invention
[0004] The scope of protection sought for various embodiments of the present invention is defined by the independent claims. Embodiments and features (if any) described in this specification that do not fall within the scope of the independent claims will be construed as examples useful for understanding various embodiments of the present invention.
[0005] A first aspect provides an apparatus, comprising: means for receiving audio data representing an audio signal for output by two or more physical speakers; means for determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers with respect to the user; and means for, in response to the determination, sending control data to an audio capture device of the user, the audio capture device operating in a directional mode for steering a sound capture beam to the first direction, wherein the control data is for causing the audio capture device to disable its directional mode or modify the sound capture beam such that the audio capture device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
[0006] In some example embodiments, the apparatus may further comprise: means for receiving from the audio capture device a notification message for indicating that the audio capture device is operating in the directional mode, and wherein, further in response to receiving the notification message, the control data is sent to the audio capture device.
[0007] In some example embodiments, the control data may be used to cause the audio capture device to widen the sound capture beam such that the audio capture device has higher sensitivity to audio signals from a wider range of directions relative to the user, the wider range of directions including the direction of at least one of the two or more specific physical speakers.
[0008] In some example embodiments, the control data may be used to cause the audio capture device to widen the sound capture beam such that the audio capture device has higher sensitivity to audio signals from the respective directions of the two or more specific physical speakers.
[0009] In some example embodiments, the control data may be used to cause the audio capture device to turn the sound capture beam from the first direction to the direction of one of the two or more specific physical speakers.
[0010] In some example embodiments, the control data may include data indicating the spatial position of at least one of the two or more specific physical speakers for enabling the audio capture device to estimate the direction or respective directions of at least one of the two or more specific physical speakers.
[0011] In some example embodiments, the apparatus may further include: a component for receiving position data indicating the spatial position of the audio capture device and the direction of the sound capture beam from the audio capture device; and a component for using the position data and the known position of at least one of the two or more specific physical speakers to determine a modification to be applied to the sound capture beam of the audio capture device, wherein the control data includes the determined modification to be applied by the audio capture device.
[0012] In some example embodiments, the modification may include the amount of widening of the sound capture beam.
[0013] In some example embodiments, the modification may include the direction and amount of turning the sound capture beam from the first direction to the direction of one of the two or more specific physical speakers.
[0014] In some example embodiments, the apparatus may further include: a component for receiving spatial metadata associated with the audio data, the spatial metadata indicating spatial characteristics of an audio scene including at least the first sound source, wherein the component for determining is configured to: determine, based on the spatial metadata, that the first sound source will be perceived as having a first direction relative to the user that is different from the physical speaker direction.
[0015] In some example embodiments, the audio data and the spatial metadata may be received in an Immersive Voice and Audio Service (IVAS) bitstream.
[0016] In some example embodiments, a data format including one of the following may be used to provide the IVAS bitstream: Metadata-Assisted Spatial Audio (MASA); Object with Metadata-Assisted Spatial Audio (OMASA); and Independent Stream with Metadata (ISM).
[0017] In some example embodiments, the apparatus may further include: a component for identifying, in response to detecting that the audio data and the spatial metadata are received in an IVAS bitstream, one or more of the MASA, OMASA, and ISM data formats supported by the IVAS bitstream; and a component for selecting one of the MASA, OMASA, and ISM data formats or a priority order for decoding the IVAS bitstream and obtaining the spatial metadata.
[0018] In some example embodiments, the apparatus may include a mobile terminal.
[0019] A second aspect provides an apparatus, including: a component for capturing an audio signal output by two or more physical speakers, the audio signal including an audio signal representing a first sound source output by two or more specific physical speakers, such that the first sound source will be perceived as having a first direction relative to the user that is different from the physical speaker direction; a component for operating in a directional mode for steering a sound capture beam to the first direction; and a component for receiving control data from a control device, wherein the control data causes the directional mode to be disabled or the sound capture beam to be modified such that the apparatus has a higher sensitivity to an audio signal from the direction of at least one of the two or more specific physical speakers.
[0020] In some example embodiments, the apparatus may further include: a component for sending a notification message to the control device for indicating that the apparatus is operating in the directional mode, wherein the control data is received from the control device in response to sending the notification message.
[0021] In some example embodiments, the control data may cause the sound capture beam to be widened so that the device has higher sensitivity to audio signals from a wider range of directions, the wider range of directions including the direction of at least one of the two or more specific physical speakers.
[0022] In some example embodiments, the control data may cause the sound capture beam to be widened so that the device has higher sensitivity to audio signals from the respective directions of the two or more specific physical speakers.
[0023] In some example embodiments, the control data may cause the sound capture beam to turn from the first direction to the direction of one of the two or more specific physical speakers.
[0024] In some example embodiments, the control data may include data indicating the spatial position of at least one of the two or more physical speakers, and the device may further include: components for estimating the direction or respective directions of at least one of the two or more specific physical speakers.
[0025] In some example embodiments, the device may further include: components for sending position data indicating the spatial position of the device and the direction of the sound capture beam to the control device; wherein the control data includes modifications to be applied to the sound capture beam determined based on the position data and the known positions of at least one of the two or more specific physical speakers.
[0026] In some example embodiments, the modification may include the amount by which the sound capture beam is widened.
[0027] In some example embodiments, the modification may include the direction and amount by which the sound capture beam is turned from the first direction to the direction of one of the two or more specific physical speakers.
[0028] In some example embodiments, the device may include a head-mounted or ear-mounted user device.
[0029] A third aspect provides an apparatus, comprising: means for receiving audio data representing an audio signal for output by two or more physical speakers; means for determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to the user and for determining that the user's audio capture device operates in a directional mode for steering a sound capture beam towards the first direction; and means for presenting, in response to the determination, at least some of the audio signals of the first sound source from a specific physical speaker selected from the two or more specific physical speakers and not from other specific physical speakers, such that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered towards the selected physical speaker.
[0030] A fourth aspect provides an apparatus, comprising: means for receiving audio data representing an audio signal for output by two or more physical speakers; means for determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to the user and for determining that the user's audio capture device operates in a directional mode for steering a sound capture beam towards the first direction; means for receiving, from the audio capture device, a notification message indicating that one or more other real-world sound sources have been captured by the sound capture beam; and means for presenting, in response to receiving the notification message, at least some of the audio signals of the first sound source such that the first sound source will be perceived as having a second direction different from the first direction relative to the user.
[0031] A fifth aspect provides a method, comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to the user; and in response to the determination, sending control data to the user's audio capture device, the audio capture device operating in a directional mode for steering a sound capture beam towards the first direction, wherein the control data is for causing the audio capture device to disable its directional mode or modify the sound capture beam such that the audio capture device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
[0032] In some example embodiments, the method may further include: receiving, from the audio capture device, a notification message indicating that the audio capture device is operating in the directional mode, wherein, further in response to receiving the notification message, the control data is sent to the audio capture device.
[0033] In some example embodiments, the control data may be used to cause the audio capture device to widen the sound capture beam, such that the audio capture device has a higher sensitivity to audio signals from a wider range of directions relative to the user, the wider range of directions including the direction of at least one of the two or more specific physical speakers.
[0034] In some example embodiments, the control data may be used to cause the audio capture device to widen the sound capture beam, such that the audio capture device has a higher sensitivity to audio signals from the respective directions of the two or more specific physical speakers.
[0035] In some example embodiments, the control data may be used to cause the audio capture device to turn the sound capture beam from the first direction to the direction of one of the two or more specific physical speakers.
[0036] In some example embodiments, the control data may include data indicating the spatial position of at least one of the two or more specific physical speakers, for enabling the audio capture device to estimate the direction or respective directions of at least one of the two or more specific physical speakers.
[0037] In some example embodiments, the method may further include: receiving, from the audio capture device, position data indicating the spatial position of the audio capture device and the direction of the sound capture beam; and using the position data and the known position of at least one of the two or more specific physical speakers, determining a modification to be applied to the sound capture beam of the audio capture device, wherein the control data includes the determined modification to be applied by the audio capture device.
[0038] In some example embodiments, the modification may include the amount by which the sound capture beam is widened.
[0039] In some example embodiments, the modification may include the direction and amount by which the sound capture beam is turned from the first direction to the direction of one of the two or more specific physical speakers.
[0040] In some example embodiments, the method may further include: receiving spatial metadata associated with the audio data, the spatial metadata indicating spatial characteristics of an audio scene including at least the first sound source, wherein, based on the spatial metadata, it is determined that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speaker.
[0041] In some example embodiments, the audio data and the spatial metadata may be received in an Immersive Voice and Audio Service (IVAS) bitstream.
[0042] In some example embodiments, a data format including one of the following may be used to provide the IVAS bitstream: Metadata-Assisted Spatial Audio (MASA); Object with Metadata-Assisted Spatial Audio (OMASA); and Independent Stream with Metadata (ISM).
[0043] In some example embodiments, the method may further include: in response to detecting that the audio data and the spatial metadata are received in an IVAS bitstream, identifying one or more of the MASA, OMASA, and ISM data formats supported by the IVAS bitstream; and selecting one or a priority order of the MASA, OMASA, and ISM data formats for decoding the IVAS bitstream and obtaining the spatial metadata.
[0044] In some example embodiments, the method may be executed at a mobile terminal.
[0045] A sixth aspect provides a method, including: capturing an audio signal output by two or more physical speakers, the audio signal including an audio signal representing a first sound source output by two or more specific physical speakers, such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speaker; operating in a directional mode for steering a sound capture beam towards the first direction; and receiving control data from a control device, wherein the control data causes the directional mode to be disabled or the sound capture beam to be modified such that the device has a higher sensitivity to an audio signal from at least one of the two or more specific physical speakers.
[0046] In some example embodiments, the method may further include: sending a notification message to the control device for indicating that the device is operating in the directional mode, wherein the control data is received from the control device in response to sending the notification message.
[0047] In some example embodiments, the control data may cause the sound capture beam to be widened such that the device has a higher sensitivity to audio signals from a wider range of directions, the wider range of directions including the direction of at least one of the two or more specific physical speakers.
[0048] In some example embodiments, the control data may cause the sound capture beam to be widened such that the device has a higher sensitivity to audio signals from the respective directions of the two or more specific physical speakers.
[0049] In some example embodiments, the control data may cause the sound capture beam to turn from the first direction to the direction of one of the two or more specific physical speakers.
[0050] In some example embodiments, the control data may include data indicating the spatial position of at least one of the two or more physical speakers, and the method may further include: estimating the direction or respective directions of at least one of the two or more specific physical speakers.
[0051] In some example embodiments, the method may further include: sending position data indicating the spatial position of the device and the direction of the sound capture beam to the control device, wherein the control data includes modifications to be applied to the sound capture beam determined based on the position data and the known positions of at least one of the two or more specific physical speakers.
[0052] In some example embodiments, the modification may include the amount by which the sound capture beam is widened.
[0053] In some example embodiments, the modification may include the direction and amount by which the sound capture beam is turned from the first direction to the direction of one of the two or more specific physical speakers.
[0054] In some example embodiments, the method may be performed by a head-mounted or ear-mounted user device.
[0055] A seventh aspect provides a method, comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to the user, and determining that the user's audio capture device operates in a directional mode for steering a sound capture beam towards the first direction; and in response to the determination, presenting at least some of the audio signals of the first sound source from a specific physical speaker selected from the two or more specific physical speakers and not from other specific physical speakers, such that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered towards the selected physical speaker.
[0056] An eighth aspect provides a method, comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to the user, and determining that the user's audio capture device operates in a directional mode for steering a sound capture beam towards the first direction; receiving a notification message from the audio capture device indicating that one or more other real-world sound sources are captured by the sound capture beam; and in response to receiving the notification message, presenting at least some of the audio signals of the first sound source such that the first sound source will be perceived as having a second direction different from the first direction relative to the user.
[0057] A ninth aspect provides a computer program comprising an instruction set which, when executed on a device, is configured to cause the device to perform a method, the method comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to the user; and in response to the determination, sending control data to the user's audio capture device, the audio capture device operating in a directional mode for steering a sound capture beam towards the first direction, wherein the control data is for causing the audio capture device to disable its directional mode or modify the sound capture beam such that the audio capture device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
[0058] In some example embodiments, the ninth aspect may include any other features mentioned for the method of the fifth aspect.
[0059] The tenth aspect provides a computer program comprising an instruction set which, when executed on a device, is configured to cause the device to perform a method comprising: capturing an audio signal output by two or more physical speakers, the audio signal comprising an audio signal representing a first sound source output by two or more particular physical speakers such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speakers; operating in a directional mode for steering a sound capture beam towards the first direction; and receiving control data from a control device, wherein the control data causes the directional mode to be disabled or the sound capture beam to be modified such that the device has a higher sensitivity to an audio signal from the direction of at least one of the two or more particular physical speakers.
[0060] In some example embodiments, the tenth aspect may include any other features mentioned for the method of the sixth aspect.
[0061] The eleventh aspect provides a computer program comprising an instruction set which, when executed on a device, is configured to cause the device to perform a method comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more particular physical speakers such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speakers, and determining that the user's audio capture device is operating in a directional mode for steering a sound capture beam towards the first direction; and in response to the determination, presenting the at least some audio signals of the first sound source from a particular physical speaker selected from the two or more particular physical speakers and not from other particular physical speakers such that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered towards the selected physical speaker.
[0062] A twelfth aspect provides a computer program including an instruction set that, when executed on a device, is configured to cause the device to perform a method including: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to a user; and determining that an audio capture device of the user is operating in a directional mode for steering a sound capture beam towards the first direction; receiving, from the audio capture device, a notification message indicating that one or more other real-world sound sources have been captured by the sound capture beam; and in response to receiving the notification message, presenting the at least some of the audio signals of the first sound source such that the first sound source will be perceived as having a second direction different from the first direction relative to the user.
[0063] A thirteenth aspect of the present invention provides a non-transitory computer-readable medium having computer-readable code stored thereon that, when executed by at least one processor, causes the at least one processor to perform a method including: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to a user; and in response to the determination, sending control data to an audio capture device of the user, the audio capture device operating in a directional mode for steering a sound capture beam towards the first direction, wherein the control data is for causing the audio capture device to disable its directional mode or modify the sound capture beam such that the audio capture device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
[0064] In some example embodiments, the thirteenth aspect may include any other features mentioned for the method of the fifth aspect.
[0065] A fourteenth aspect of the present invention provides a non-transitory computer-readable medium having computer-readable code stored thereon, the computer-readable code causing the at least one processor to perform a method when executed by the at least one processor, the method comprising: capturing an audio signal output by two or more physical speakers, the audio signal including an audio signal representing a first sound source output by two or more particular physical speakers such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speakers; operating in a directional mode for steering a sound capture beam towards the first direction; and receiving control data from a control device, wherein the control data causes the directional mode to be disabled or the sound capture beam to be modified such that the device has a higher sensitivity to audio signals from the direction of at least one of the two or more particular physical speakers.
[0066] In some example embodiments, the fourteenth aspect may include any of the other features mentioned for the method of the sixth aspect.
[0067] A fifteenth aspect of the present invention provides a non-transitory computer-readable medium having computer-readable code stored thereon, the computer-readable code causing the at least one processor to perform a method when executed by the at least one processor, the method comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more particular physical speakers such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speakers, and determining that the user's audio capture device is operating in a directional mode for steering a sound capture beam towards the first direction; and in response to the determination, presenting the at least some audio signals of the first sound source from a particular physical speaker selected from the two or more particular physical speakers and not from other particular physical speakers such that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered towards the selected physical speaker.
[0068] A sixteenth aspect of the present invention provides a non-transitory computer-readable medium having computer-readable code stored thereon, the computer-readable code causing at least one processor to perform a method when executed by the at least one processor, the method comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to a user; and determining that the user's audio capture device is operating in a directional mode for steering a sound capture beam towards the first direction; receiving from the audio capture device a notification message indicating that one or more other real-world sound sources have been captured by the sound capture beam; and in response to receiving the notification message, presenting the at least some of the audio signals of the first sound source such that the first sound source will be perceived as having a second direction different from the first direction relative to the user.
[0069] A seventeenth aspect of the present invention provides an apparatus having at least one processor and at least one memory having computer-readable code stored thereon, the computer-readable code controlling the at least one processor when executed to: receive audio data representing an audio signal for output by two or more physical speakers; determine that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to a user; and in response to the determination, send control data to the user's audio capture device, the audio capture device operating in a directional mode for steering a sound capture beam towards the first direction, wherein the control data is for causing the audio capture device to disable its directional mode or modify the sound capture beam such that the audio capture device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
[0070] In some example embodiments, the seventeenth aspect may include any other features mentioned for the method of the fifth aspect.
[0071] The eighteenth aspect of the present invention provides a device having at least one processor and at least one memory storing computer-readable code thereon, the computer-readable code when executed controlling the at least one processor to: capture an audio signal output by two or more physical speakers, the audio signal including an audio signal representing a first sound source output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers with respect to a user; operate in a directional mode for steering a sound capture beam to the first direction; and receive control data from a control device, wherein the control data causes the directional mode to be disabled or the sound capture beam to be modified such that the device has a higher sensitivity to an audio signal from the direction of at least one of the two or more specific physical speakers.
[0072] In some example embodiments, the eighteenth aspect may include any other features mentioned for the method of the sixth aspect.
[0073] The nineteenth aspect of the present invention provides a device having at least one processor and at least one memory storing computer-readable code thereon, the computer-readable code when executed controlling the at least one processor to: receive audio data representing an audio signal for output by two or more physical speakers; determine that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers with respect to a user, and determine that an audio capture device of the user is operating in a directional mode for steering a sound capture beam to the first direction; and in response to the determination, present at least some of the audio signals of the first sound source from a specific physical speaker selected from the two or more specific physical speakers and not from other specific physical speakers such that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered to the selected physical speaker.
[0074] The twentieth aspect of the present invention provides an apparatus having at least one processor and at least one memory storing computer-readable code thereon, the computer-readable code when executed controlling the at least one processor to: receive audio data representing an audio signal for output by two or more physical speakers; determine that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers with respect to the user, and determine that the user's audio capture device is operating in a directional mode for steering a sound capture beam towards the first direction; receive from the audio capture device a notification message indicating that one or more other real-world sound sources have been captured by the sound capture beam; and in response to receiving the notification message, present the at least some of the audio signals of the first sound source such that the first sound source will be perceived as having a second direction different from the first direction with respect to the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The present invention will now be described by way of non-limiting examples with reference to the accompanying drawings, in which:
[0076] Figure 1 A system for audio presentation is shown;
[0077] Figure 2 A system having a sound source direction indication is shown Figure 1 is shown;
[0078] Figure 3 An audio capture device is shown;
[0079] Figure 4 is a flowchart showing operations according to one or more example embodiments;
[0080] Figure 5 A system for audio presentation that can be used to understand one or more example embodiments is shown;
[0081] Figure 6 A system for audio presentation according to one or more example embodiments is shown;
[0082] Figure 7 A system for audio presentation according to one or more other example embodiments is shown;
[0083] Figure 8 A system for audio presentation according to one or more other example embodiments is shown;
[0084] Figure 9 is a flowchart showing operations according to another example embodiment;
[0085] Figure 10 is a flowchart showing operations according to another exemplary embodiment;
[0086] Figure 11 shows a system for audio rendering according to another exemplary embodiment;
[0087] Figure 12 is a flowchart showing operations according to another exemplary embodiment;
[0088] Figure 13 shows an audio field that can be used to understand one or more other exemplary embodiments;
[0089] Figure 14 shows, when being modified, according to one or more other exemplary embodiments Figure 13 of the audio field;
[0090] Figure 15 is a block diagram of a device that can be configured according to one or more exemplary embodiments; and
[0091] Figure 16 is a non - transitory computer - readable medium according to one or more exemplary embodiments. Detailed Description
[0092] Exemplary embodiments relate to audio signal capture, for example, in a case where an audio capture device can capture an audio signal output or intended to be output using two or more physical speakers.
[0093] Exemplary embodiments focus on immersive audio, but it should be understood that other audio formats (including but not limited to stereo and multi - channel audio formats) for output by two or more physical speakers are also applicable.
[0094] In this context, immersive audio can refer to any of the following techniques: which presents sound objects in space such that a listening user in that space can perceive one or more sound objects as coming from corresponding directions in that space. The user can also perceive a sense of depth.
[0095] In this context, immersive audio can include any techniques, such as surround sound and different types of spatial audio techniques, which utilize two or more physical speakers with corresponding spaced - apart positions to provide an immersive audio experience. 3GPP Immersive Voice and Audio Service (IVAS) and MPEG - I Audio are example immersive audio formats or codecs, but exemplary embodiments are not limited to such examples.
[0096] Figure 1Shown is a system 100 for outputting immersive audio, which includes an audio processor 102 (sometimes referred to as an audio receiver or audio amplifier) and first to fifth physical speakers 104A - 104E (hereinafter referred to as "speakers"), which are spaced apart in a listening space 105 (which can be a room) and have corresponding positions. The first, second, and third speakers 104A, 104B, 104C can be referred to as front left, front right, and front center speakers based on their corresponding positions relative to a typical listening position (indicated by reference numeral 106). Similarly, the fourth and fifth speakers 104D, 104E can be referred to as rear left and rear right speakers based on their corresponding positions relative to the listening position 106. There may also be another speaker (not shown) for outputting lower frequency audio signals, and this can be referred to as a subwoofer, bass speaker, etc. In some example embodiments, there may be fewer speakers. Thus, system 100 can represent a 5.1 surround sound setup, but it will be understood that there are many other setups, such as but not limited to 2.0, 2.1, 3.1, 4.0, 4.1, 5.1, 5.1.2, 5.1.4, 6.1, 7.1, 7.1.2, 7.1.4, 7.2, 9.1, 9.1.2, 10.2, 13.1, and 22.2.
[0097] The audio processor 102 can be configured to store audio data representing immersive audio content for output via all or specific ones of the first to fifth speakers 104A - 104E. The audio processor 102 can include an amplifier, signal processing functions, one or more memories, such as a hard disk drive (HDD) and / or solid state drive (SSD) for storing audio data. The audio processor 102 can be provided in any suitable form, such as a set - top box, a mobile terminal (such as a mobile phone), a tablet, etc. The audio processor 102 can be a pure digital processor, in which case it may not include an amplifier. For example, the audio data can be received from a remote source 108 via a network 110 and stored on one or more memories. The network 110 can include the Internet. The audio data can be received via a wired or wireless connection to the network 110 (such as via a home router or hub). Alternatively, the audio data can be streamed from the remote source 108 using a suitable streaming protocol (such as the Real - Time Streaming Protocol (RTSP), etc.). Alternatively, the audio data can be provided on a non - transitory computer - readable medium (such as a compact disc, memory card, memory stick, or removable hard disk drive inserted or connected to a suitable component of the audio processor 102).
[0098] Audio data can represent an audio signal for any form of audio, whether it is speech, singing, music, background sound, or a combination thereof. The audio data can include data as part of a voice call or a conference. The audio data can be associated with video data, such as part of a video call, a video conference, a video clip, a video game, or a movie. The audio data can represent an audio scene that includes one or more sound objects.
[0099] The audio processor 102 can be configured to present the audio data by outputting the audio signal using a specific speaker among the first to fifth speakers 104A - 104E. Thus, the audio processor 102 can include hardware, software, and / or firmware that is configured to process the audio signal and output (or present) the audio signal to these specific speakers among the first to fifth speakers 104A - 104E. The audio processor 102 can also provide other signal processing functions, such as modifying the overall volume, modifying the respective volumes of different frequency ranges, and / or performing specific effects, such as modifying the reverb and / or performing panning, such as vector base amplitude panning (VBAP). VBAP is a method of placing a sound source in any direction using the current speaker setup; the number of speakers is arbitrary as they can be placed in a 2D or 3D setup. VBAP produces a virtual source that is localized for a relatively narrow area. VBAP processing can involve finding a speaker triplet (i.e., three speakers) that encloses the desired sound source panning position and then calculating the gain of the audio signal to be applied to the sound source such that the sound source will be reproduced using these three speakers. The audio processor 102 can implement VBAP, for example. An alternative method is speaker placement correction amplitude panning (SPCAP). Another alternative method is edge fade amplitude panning (EFAP).
[0100] The audio data may include metadata or other computer-readable indications, and the audio processor 102 processes these metadata and indications to determine how the audio signal is to be presented, such as by which one of the first to fifth speakers 104A - 104E and in what signal proportion. For example, if the audio format is an IVAS bitstream or the like, the audio data may have associated spatial metadata. The spatial metadata may indicate the spatial characteristics of the audio scene, such as by indicating direction and direct-to-total ratio parameters, which together control how much signal energy is to be reproduced by a particular one of the first to fifth speakers 104A - 104E. The spatial metadata may also indicate parameters such as diffuseness coherence, diffuseness energy to total energy ratio, surround coherence, and residual energy to total energy ratio. For example, a sound with a direction pointing forward and a direct-to-total ratio of "1" will be reproduced only from the front (i.e., the third speaker 104C), while if the direct-to-total ratio is "0", the sound will be reproduced diffusely from each of the first to fifth speakers 104A - 104E.
[0101] In this case, the IVAS bitstream may have a specific format, including but not limited to Metadata-Assisted Spatial Audio (MASA), Object with Metadata-Assisted Spatial Audio (OMASA), and / or Independent Stream with Metadata (ISM). In some cases, the audio processor 102 may determine which audio format to decode by negotiating with the remote source 108. The remote source 108 may indicate in the initial data which audio formats are supported in the IVAS bitstream, and then the audio processor 102 may select, for example, one or more audio formats to use in a preferred order (this selection may be based on the availability of such formats in the audio processor 102's own decoder), and thus configure its decoding function. The audio signal may be arranged into channels, for example, each of the first to fifth speakers 104A - 104E has one channel.
[0102] In some cases, based on the metadata or other computer-readable indications, only a subset of the first to fifth speakers 104A - 104E may be used.
[0103] The audio processor 102 may present a sound source by outputting an audio signal from two or more specific ones of the first to fifth speakers 104A - 104E such that the user perceives the sound source as coming from a direction different from the direction of any of the first to fifth speakers relative to the user. This may be referred to as a phantom sound source.
[0104] Figure 2 There is shown a first sound source 200 indicated at a position between the first and third speakers 104A, 104C Figure 1A system such that the first sound source 200 will be perceived by a user at position 106 as coming from a first direction 202 relative to the user. The first sound source 200 is an example of a virtual sound source.
[0105] In this example, the audio processor 102 can use the first and third speakers 104A, 104C to present the first sound source 200.
[0106] The same process can be performed for one or more other sound sources not shown such that these sound sources will be perceived by the user as coming from corresponding directions relative to the user position 106.
[0107] A user wearing a particular audio capture device may not obtain an optimal user experience when experiencing immersive audio, such as as Figure 2 shown. This is especially true for audio capture devices (such as hearing aids or headphone devices) that can operate in a directional or unoccluded mode for hearing assistance. In this context, such an audio capture device can not only capture sound but also process and reproduce the captured sound.
[0108] Figure 3 is a schematic diagram of an example audio capture device including headphones 300. In other examples, the audio capture device can include any ear-worn or head-worn device that includes one or more microphones and one or more speakers, such as a beamforming hearing aid. Although not shown, the headphones 300 can include one of a pair of headphones. The headphones 300 can include a speaker 302 (to be placed over or within the user's ear in use) and a microphone array 304. The headphones 300 can be configured to provide hearing assistance when operating in a so-called directional (or unoccluded) mode in use, which can be the default mode or a mode enabled by a control input of the headphones or by another device (such as the user device 306) paired with the headphones.
[0109] In some example embodiments, the user device 306 can include Figure 1 the audio processor 102 as shown. The control input can be provided by any suitable means (such as touch input, gesture, or voice input).
[0110] The microphone array 304 can be configured to steer the sound capture beam 308 towards the perceived direction of a particular sound (such as a particular sound object) or towards a direction relative to the headphones (such as the front direction).
[0111] More specifically, the earphone 300 may include a signal processing function 310 that spatially filters the surrounding audio field such that sound from one or more specific directions (which may adaptively change) or from within a predetermined range of directions is amplified more than sound from other directions. In other words, the earphone 300 (or more precisely, its microphone array 304) is more sensitive to sound from one or more specific directions or a range of directions than sound outside of one or more specific directions or a range of directions. These directions effectively form the aforementioned sound capture beam 308, which is used to visualize the sensitivity of the microphone array 304 at different times. It can be seen that the direction of the sound capture beam 308 can be steered under the control of the signal processing function 310, which amplifies the sound captured within the sound capture beam and delivers the captured sound to the speaker 302.
[0112] Known methods can be used to configure the signal processing function 310 to widen the sound capture beam 308 and / or steer the sound capture beam towards the direction of one or more specific sound objects or relative to the direction of the earphone 300.
[0113] Specific sound objects may include predetermined types of sound objects, such as speech sound objects and / or sound objects in a specific direction relative to the earphone (e.g., towards the front side of the earphone). The audio processor 102 may infer the importance of a sound object to the user based on the predetermined type or corresponding direction of the sound object.
[0114] Returning to Figure 2 , if the user at position 106 is wearing an audio capture device (such as the earphone 300) operating in a directional mode, then Figure 3 the sound capture beam 308 of can be directed by the signal processing function 310 to a first direction 202 because the first direction 202 is the perceived direction of the first sound source 200. However, the amplification may be suboptimal and may affect the clarity of the first sound source 200. The amplification may be suboptimal because the sound capture beam 308 is directed to a location without a speaker and may perform attenuation on audio signals outside of the sound capture beam (such as speaker audio signals). Additionally, the size of the sound capture beam 308 and / or the steering of the sound capture beam 308 performed by the signal processing function 310 may be affected. Overall, the user experience may be negatively impacted.
[0115] Figure 4is a flowchart showing operations 400 that can be performed by one or more example embodiments. Operations 400 can be performed by hardware, software, firmware, or a combination thereof. Operations 400 can be performed by one or corresponding components, which are any suitable components, such as one or more processors or controllers in combination with computer-readable instructions provided on one or more memories. Operations 400 can be performed, for example, by the audio processor 102 described in the example for Figure 2 to execute.
[0116] The first operation 401 can include: receiving audio data representing an audio signal for output by two or more physical speakers.
[0117] The second operation 402 can include: determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction different from the direction of the physical speakers relative to the user.
[0118] The third operation 403 can include: in response to the determination, sending control data to the user's audio capture device, which operates in a directional mode for steering a sound capture beam to the first direction, wherein the control data is for causing the audio capture device to disable its directional mode or modify the sound capture beam such that the audio capture device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
[0119] In this way, the audio capture device operating in the directional mode can be controlled so as to overcome or at least mitigate the problems described above. The audio capture device can be configured to capture sound and also process and reproduce the sound for output via one or more speakers of the audio capture device.
[0120] For ease of explanation, it will be assumed hereinafter that the audio capture device includes the headset 300 and the control device includes an audio processor, which can be part of, for example, a mobile phone.
[0121] Figure 5 A system 500 for outputting immersive audio according to one or more example embodiments is shown.
[0122] System 500 is similar to the system shown in Figure 2 . System 500 includes an audio processor 502, which includes a processing module 504 configured to perform operations 400 described with reference to Figure 4 to execute.
[0123] The processing module 504 may receive audio data from the remote source 108 according to the first operation 401, for example, in an immersive audio data format (such as the IVASMASA format).
[0124] The processing module 504 may determine according to the second operation 402 that the audio signal representing the first sound source 200 is output or will be output from the first and third speakers 104A, 104C as shown in Figure 2 . Accordingly, the processing module 504 may determine that the first sound source 200 is or is intended to be perceived as coming from the first direction 202 relative to the user at the position 106. This determination may be based on the spatial metadata associated with the audio data, such as MASA spatial metadata.
[0125] Then, the processing module 504 may send control data to the headset 300 via the control channel 510 according to the third operation 403.
[0126] As shown in the figure, the headset 300 may operate in a directional mode for steering the sound capture beam 506 to the first direction 202.
[0127] The fact that the headset 300 is operating in a directional mode may be unknown or known.
[0128] For example, the processing module 504 may send control data to the headset 300 without knowing that the headset 300 is operating in a directional mode. In this case, the control channel 510 may be a broadcast channel. The same control data may also be received by one or more other audio capture devices within the receiving range of the processing module 504, such that the one or more other audio capture devices will operate in the same manner as the headset 300.
[0129] In other examples, the processing module 504 may receive a notification message from the headset 300 for indicating that the headset is operating in a directional mode. The notification message may be sent by the headset 300 in response to a discovery signal sent (such as broadcast) by the processing module 504. Alternatively, the notification message may be sent by the headset 300 in response to enabling the directional mode at the headset. The processing module 504 may further send control data in response to receiving the notification message. The control channel 510 may be a point-to-point channel.
[0130] This signal communication between the audio processor 502 and the headset 300 may be facilitated by any suitable wireless protocol, such as WiFi, Bluetooth, Zigbee, or any variant thereof. For example, there may be a pairing relationship between the audio processor 502 and the headset 300, and when the headset 300 is within the communication range of the audio processor 502, this pairing relationship automatically establishes a link and performs signaling between these devices.
[0131] The control data may cause the earphone 300 (or more specifically, the signal processing function 310 of the earphone 300) to disable its directional pattern, in which case the microphone array 304 becomes sensitive to sounds from all possible directions (and thus includes the first and third speakers 104A, 104C).
[0132] Alternatively, the control data may cause the earphone 300 (or more specifically, the signal processing function 310 of the earphone 300) to modify the sound capture beam 506 such that the earphone 300 has a higher sensitivity to audio signals from the direction of at least one of the first and third speakers 104A, 104C.
[0133] For example, as Figure 6 shown, the control data may cause the earphone 300 to configure its signal processing function 310 to generate a (spatially) wider sound capture beam 606. Compared with Figure 5 the case of, the wider sound capture beam 606 has a higher sensitivity to audio signals from a wider range of directions (in this case, including the direction of the first speaker 104A).
[0134] For example, as Figure 7 shown, the control data may cause the earphone 300 to configure its signal processing function 310 to generate a (spatially) wider sound capture beam 706 that includes the directions of both the first and third speakers 104A, 104C.
[0135] For example, as Figure 8 shown, the control data may cause the earphone 300 to configure its signal processing function 310 to turn the sound capture beam 506 from the first direction 202 to the direction of one of the first and third speakers 104A, 104C. In Figure 8 this case, the sound capture beam 506 is turned from the first direction 202 to the direction 806 of the first speaker 104A. In other examples, the sound capture beam 506 may be turned from the first direction 202 to the direction of the third speaker 104C.
[0136] In some example embodiments, the control data may include data indicating the spatial position of at least one of the specific speakers (in this case, the spatial position of one or both of the first and third speakers 104A, 104C).
[0137] The earphone 300 may estimate the direction or corresponding directions of the first and / or third speakers 104A, 104C in order to modify the sound capture beam 506 according to the above examples.
[0138] For example, the earphone 300 can determine its own spatial position (more precisely, the position 106 of the user) using known methods (e.g., by using ranging signals sent from or sent to a reference position and multilateration processing). The earphone 300 knows that its sound capture beam 506 has a specific direction or orientation relative to the user position 106.
[0139] Then, the earphone 300 can use the spatial positions of the first and / or third speakers 104A, 104C relative to its own position to determine how to modify the width of the sound capture beam 506 so that the microphone array 304 has higher sensitivity in the direction of the first and / or third speakers 104A, 104C.
[0140] In the case where the control data is used to cause the earphone 300 to turn the sound capture beam 506 from the first direction 202 to the direction of one of the first and third speakers 104A, 104C, the earphone 300 can determine that direction and the amount of rotation required to turn the sound capture beam.
[0141] In some example embodiments, the processing module 504 can be configured to receive position data from the earphone 300 indicating the spatial position of the earphone and the direction of the sound capture beam 506.
[0142] Then, the processing module 504 can use the position data of the earphone and the direction of the sound capture beam to determine the modification to be applied to the sound capture beam 506.
[0143] For example, the processing module 504 can determine the amount by which to widen the sound capture beam 506 so that the microphone array 304 has higher sensitivity in the direction of the first and / or third speakers 104A, 104C.
[0144] For example, the processing module 504 can determine the direction and the amount of rotation to turn the sound capture beam 506 from the first direction 202 to the direction of one of the first and third speakers 104A, 104C.
[0145] The control data sent by the processing module 504 to the earphone 300 can include the determined modifications to be applied by the earphone. In response to receiving the control data from the processing module 504, the earphone 300 can perform the determined modifications.
[0146] Figure 9is a flowchart showing operations 900 that may be performed by one or more example embodiments. Operations 900 may be performed by hardware, software, firmware, or a combination thereof. Operations 900 may be performed by one or corresponding components, where a component is any suitable component, such as one or more processors or controllers in combination with computer-readable instructions provided on one or more memories. Operations 900 may be performed, for example, by the audio capture device (such as the headset 300) described for the above examples.
[0147] A first operation 901 may include: capturing audio signals output by two or more physical speakers, where the audio signals include audio signals representing a first sound source output by two or more specific physical speakers, such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speakers.
[0148] Assuming that a directional mode is enabled for steering a sound capture beam to be directed towards the first direction, a second operation 902 may include: receiving control data from a control device, where the control data causes the directional mode to be disabled or the sound capture beam to be modified such that the device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
[0149] As will be appreciated, the control device in the second operation 902 may include the audio processor 502 described for Figures 5 - 8 the description.
[0150] In some example embodiments, other operations may include: sending a notification message to the control device for indicating that the device is operating in the directional mode, where the control data is received from the control device in response to sending the notification message.
[0151] In some example embodiments, the control data may cause the sound capture beam to be widened such that the device has a higher sensitivity to audio signals from a wider range of directions, where the wider range of directions includes the direction of at least one of the two or more specific physical speakers. For example, the control data may cause the sound capture beam to be widened such that the device has a higher sensitivity to audio signals from the respective directions of the two or more specific physical speakers.
[0152] In some example embodiments, the control data may cause the sound capture beam to be turned from the first direction to the direction of one of the two or more specific physical speakers.
[0153] In some example embodiments, the control data may include data indicating the spatial location of at least one of two or more physical speakers, and other operations may include: estimating the direction or corresponding orientation of the at least one particular physical speaker among the two or more particular physical speakers.
[0154] In some example embodiments, other operations may include: sending location data indicating the spatial location of the audio capture device and the direction of the sound capture beam to a control device, wherein the control data includes a modification to be applied to the sound capture beam determined based on the location data and the known locations of at least one of two or more particular physical speakers. The modification may include an amount by which to widen the sound capture beam. Alternatively, the modification may include a direction and an amount by which to turn the sound capture beam from a first direction to the direction of one of the two or more particular physical speakers.
[0155] From the above it will be understood that by disabling the directional pattern or modifying the sound capture beam, the user of the audio capture device will have an improved perception of the sound source.
[0156] Other embodiments will now be described, which may incorporate the specific features and considerations described above.
[0157] Figure 10 is a flowchart showing operations 1000 that may be performed by one or more other example embodiments. Operations 1000 may be performed by hardware, software, firmware, or a combination thereof. Operations 1000 may be performed by one or corresponding components, where a component is any suitable component, such as one or more processors or controllers in combination with computer-readable instructions provided on one or more memories. Operations 1000 may be performed, for example, by the audio processor 502 already described for the above examples.
[0158] A first operation 1001 may include: receiving audio data representing an audio signal for output by two or more physical speakers.
[0159] A second operation 1002 may include: determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more particular physical speakers such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speakers.
[0160] A third operation 1003 may include: determining that the user's audio capture device is operating in a directional mode for turning a sound capture beam to the first direction.
[0161] The fourth operation 1004 may include: in response to the second and third determination operations 1002, 1003, presenting at least some audio signals of the first sound source from a specific physical speaker selected from two or more specific physical speakers and not from other specific physical speakers, so that the first sound source is perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to turn to the selected physical speaker.
[0162] According to this specific example, the audio processor 502 may present the audio signal of the first sound source in a manner different from expected based on the received audio data. This may include, for example, modifying the spatial metadata received together with the audio data to effectively move the first sound source to the selected physical speaker.
[0163] Referring again to Figure 5 , for example, according to the first operation 1001, the audio processor 502 may receive audio data in an IVAS bitstream having a specific format (including but not limited to MASA, OMASA, and / or ISM).
[0164] According to the second operation 1002, the audio processor 502 may analyze the spatial metadata included in one of these formats to determine that at least some audio signals representing the first sound source 200 in the audio signal are for output by the first and third speakers 104A, 104C, so that the first sound source is perceived as having a first direction 202 different from the direction of the physical speaker relative to the user.
[0165] According to the third operation 1003, the audio processor 502 may determine that the headset 300 is operating in a directional mode, for example, based on a notification message received from the headset 300, to turn the sound capture beam 506 to the first direction 202.
[0166] According to the fourth operation 1004, the audio processor 502 may present at least some audio signals in the audio signal of the first sound source 200 from the first speaker 104A and not from the third speaker 104C, so that the first sound source is perceived from the direction of the first speaker. Alternatively, the audio signal of the first sound source 200 may be presented from the third speaker 104C and not from the first speaker 104A.
[0167] Referring to Figure 11 , this will cause the sound capture beam 506 of the headset 300 to turn to the first speaker 104A.
[0168] From the above, it will be understood that by presenting only the audio signal of the first sound source 200 from the first speaker 104A, the user of the headset 300 will have an improved perception of the first sound source.
[0169] Figure 12is a flowchart showing operations 1200 that can be performed by one or more other example embodiments. Operations 1200 can be performed by hardware, software, firmware, or a combination thereof. Operations 1200 can be performed by one or corresponding components, where a component is any suitable component, such as one or more processors or controllers in combination with computer-readable instructions provided on one or more memories. Operations 1200 can be performed, for example, by the audio processor 502 described for the above examples.
[0170] A first operation 1201 can include: receiving audio data representing an audio signal for output by two or more physical speakers.
[0171] A second operation 1202 can include: determining that at least some of the audio signals in the audio signal representing a first sound source are for output by two or more specific physical speakers such that the first sound source will be perceived as having a first direction relative to the user that is different from the direction of the physical speakers.
[0172] A third operation 1203 can include: determining that the user's audio capture device is operating in a directional mode for steering a sound capture beam to the first direction.
[0173] A fourth operation 1204 can include: receiving from the audio capture device a notification message indicating that one or more other real-world sound sources have been captured by the sound capture beam.
[0174] In some example embodiments, the notification message can be received in response to user feedback indicating that the first sound source is masked or interfered with by real-world sound sources. The user feedback can be received as a voice notification or by the user selecting a specific option on the audio capture device or the audio processor.
[0175] A fifth operation 1205 can include: in response to receiving the notification message, presenting at least some of the audio signals of the first sound source such that the first sound source will be perceived as having a second direction relative to the user that is different from the first direction.
[0176] The second direction can be at least a predetermined angle relative to (i.e., away from) the first direction, such as at least 25 degrees relative to the first direction.
[0177] This example embodiment can be applicable to a situation where the audio capture device is a pair of headphones or a headset and the audio data is for binaural presentation, possibly with head tracking capabilities such that when the user turns their head, the audio source remains stationary in the audio field represented by the audio data. The audio capture device can operate in a so-called transparent mode, thereby also capturing sounds from the environment.
[0178] Reference Figure 13, the user at position 106 is shown wearing a pair of head-tracking headphones 1300 that can operate in a directional mode and a transparent mode. For clarity, Figure 13 the audio processor 502 and the speakers 104A - 104E are omitted in Figure 13 An example audio scene including a first sound source 200 is shown. There are also first, second, and third real-world sound sources 1302, 1304, 1306 within the user's environment.
[0179] According to the first operation 1201, audio data can be received by the audio processor 502 in an IVAS bitstream having a specific format (including but not limited to MASA, OMASA, and / or ISM).
[0180] According to the second operation 1202, the spatial metadata included in these formats can be analyzed by the audio processor 502 to determine that at least some of the audio signals representing the first sound source 200 in the audio signal are for output by the first and third speakers 104A, 104C, such that the first sound source will be perceived as having a first direction 202 relative to the user that is different from the physical speaker direction.
[0181] According to the third operation 1203, the audio processor 502 can determine that the head-tracking headphones 1300 are operating in the directional mode, for example, based on a notification message received from the head-tracking headphones 1300, to steer the sound capture beam 506 to the first direction 202.
[0182] According to the fourth operation 1204, the audio processor 502 can receive another notification message from the head-tracking headphones 1300 or another user device indicating that the sound capture beam 506 is capturing a real-world sound source (in this case, the first real-world sound source 1302). For example, the user can select an option on the head-tracking headphones 1300 or the audio processor 502 to indicate that they are experiencing a masking effect due to the sound from the first real-world sound source 1302.
[0183] According to the fifth operation 1205 and as Figure 14 shown, the audio processor 502 can present the audio signal of the first sound source 200 such that the first sound source 200 will be perceived as having a second direction 1402 relative to the user.
[0184] The audio processor 502 can, for example, modify the spatial metadata received together with the audio data, such as rotating the direction in which the first sound source 200 is perceived by 25 degrees. If the first sound source 200 is part of an audio scene containing multiple sound sources, all the sound sources can be rotated by the same amount in the same direction.
[0185] In this manner, the voice capture beam 506 will be steered by the head-tracking headset to the second direction 1402, and the masking is reduced or eliminated.
[0186] In the above embodiments, it will be noted that audio data may be received in the IVAS bitstream. In some examples, this may involve negotiating an IVAS session with the remote source 108, for example, before starting to process the audio data, such as when starting an audio call.
[0187] As part of this process, the audio processor 502 may preferentially negotiate a specific IVAS subformat or a specific order of IVAS subformats based on, for example, the rendering capabilities of the audio processor 502 and possibly also based on determining that the audio capture device is operating in a directional mode. Specific IVAS subformats may include but are not limited to MASA, OMASA, and / or ISM.
[0188] For example, the audio processor 502 may receive a Session Description Protocol (SDP) message from the remote source 108, and the SDP message may be as follows:
[0189] m = audio 49152 RTP / AVP 96
[0190] a = rtpmap:96 IVAS / 16000
[0191] a = fmtp:96 inf = 9,21-24,10-13;
[0192] a = ptime:20
[0193] a = maxptime:240
[0194] a = sendonly
[0195] where inf indicates the IVAS input format capabilities.
[0196] The inf parameter may have values from a set including 1-24. If a range of input formats is supported, the range is indicated by the first input format in the range and the last input format in the range, with the two input formats separated by a hyphen (inf1-inf2).
[0197] If multiple input formats are not a continuous range but are individual formats, these formats may be listed as comma-separated values (inf1,inf2). Comma-separated values are also used when the input formats are in a certain range but the preferred order of the formats is not the default continuous range.
[0198] In both cases (i.e., hyphen-separated list or comma-separated list), the input formats are listed in the preferred order of the input format from the most preferred to the least preferred. In the case of using different input formats in the sending direction and the receiving direction respectively, the parameters inf-send and inf-recv are used. If the inf parameter does not exist, all possible IVAS input formats are supported for the session.
[0199] The IVAS input formats and their assigned inf attribute values are as follows:
[0200]
[0201] Thus, with respect to the embodiments described above for the audio processor 502, other operations may include: in response to detecting that audio data and spatial metadata are received in an IVAS bitstream, identifying one or more of the MASA, OMASA, and ISM data formats supported by the IVAS bitstream; and selecting one or the order of preference among the MASA, OMASA, and ISM data formats for decoding the IVAS bitstream and obtaining spatial metadata for decoding using an appropriate decoder. The selection may be based on which data formats the audio processor 502 supports.
[0202] Example apparatus
[0203] Figure 15 An apparatus according to some example embodiments is shown. The apparatus may be configured to perform the operations described herein, such as the operations described with reference to any disclosed process. The apparatus includes at least one processor 1500 and at least one memory 1501 directly or closely connected to the processor. The memory 1501 includes at least one random access memory (RAM) 1501a and at least one read-only memory (ROM) 1501b. Computer program code (software) 1506 is stored in the ROM 1501b. The apparatus may be connected to a transmitter (TX) and a receiver (RX). Optionally, the apparatus may be connected to a user interface (UI) for indicating the apparatus and / or for outputting data. At least one processor 1500, at least one memory 1501, and computer program code 1506 are arranged such that the apparatus at least performs at least a method according to any previous process (such as those disclosed for any flowchart and its related features described herein).
[0204] Figure 16 A non-transitory medium 1600 according to some embodiments is shown. The non-transitory medium 1600 is a computer-readable storage medium. It may be, for example, a CD, DVD, USB flash drive, Blu-ray disc, etc. The non-transitory medium 1600 stores computer program instructions such that an apparatus performs a method according to any previous process disclosed, such as those described for any flowchart and its related features described.
[0205] The names of network elements, protocols, and methods are based on current standards. In other versions or other technologies, the names of these network elements and / or protocols and / or methods may be different as long as they provide corresponding functions. For example, embodiments can be deployed in 2G / 3G / 4G / 5G networks and other generations of 3GPP, and also in non-3GPP radio networks such as WiFi.
[0206] The memory can be volatile or non-volatile. The memory can be, for example, RAM, SRAM, flash memory, FPGA block RAM, DCD, CD, USB flash drive, and Blu-ray disc.
[0207] If not otherwise stated or the context does not clearly indicate, a statement that two entities are different means that they perform different functions. This does not necessarily mean that they are based on different hardware. That is, each entity described in this specification can be based on different hardware, or some or all of the entities can be based on the same hardware. This does not necessarily mean that they are based on different software. That is, each entity described in this specification can be based on different software, or some or all of the entities can be based on the same software. Each entity described in this specification can be embodied in the cloud.
[0208] As a non-limiting example, the implementation of any of the blocks, devices, systems, technologies, or methods described above includes implementation as hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof. Some embodiments can be implemented in the cloud.
[0209] It will be understood that the above description is currently considered to be a preferred embodiment. However, it should be noted that the description of the preferred embodiment is given only by way of example, and various modifications can be made without departing from the scope defined by the appended claims.
Claims
1. A device comprising: means for receiving audio data representing an audio signal for output by two or more physical speakers; for determining that at least some of the audio signals representing a first sound source are intended to be output by two or more specific physical speakers so that the first sound source will be perceived as having a first direction relative to a user that is different from the direction of the physical speakers; as well as means for sending control data to an audio capture device of the user in response to the determination, the audio capture device operating in a directional mode for steering a sound capture beam toward the first direction, wherein the control data is used to cause the audio capture device to disable its directional mode or to modify the sound capture beam so that the audio capture device has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.
2. The apparatus according to claim 1, further comprising: means for receiving a notification message from the audio capture device indicating that the audio capture device is operating in the directional mode, and Wherein, further in response to receiving the notification message, the control data is sent to the audio capture device.
3. The device according to claim 1 or claim 2, wherein: The control data is used to cause the audio capture device to widen the sound capture beam so that the audio capture device has a higher sensitivity to audio signals from a wider range of directions relative to the user, the wider range of directions including the direction of at least one of the two or more specific physical speakers.
4. The device according to claim 3, wherein: The control data is used to cause the audio capture device to widen the sound capture beam so that the audio capture device has a higher sensitivity to audio signals from corresponding directions of the two or more specific physical speakers.
5. The device according to claim 1 or claim 2, wherein: The control data is used to cause the audio capturing device to turn the sound capturing beam from the first direction to the direction of a specific physical speaker among the two or more specific physical speakers.
6. The device according to any one of claims 3 to 5, wherein: The control data comprises data indicating a spatial position of at least one of the two or more specific physical speakers for enabling the audio capturing device to estimate a direction or a corresponding direction of the at least one of the two or more specific physical speakers.
7. The device according to any one of claims 3 to 5, further comprising: means for receiving, from the audio capture device, position data indicative of the spatial position of the audio capture device and the direction of the sound capture beam; as well as means for determining a modification to be applied to the sound capturing beam of the audio capturing device using the position data and the known position of the at least one specific physical speaker of the two or more specific physical speakers, Therein, the control data comprises the determined modification to be applied by the audio capture device.
8. The device according to claim 7 as appended to claim 3 or claim 4, wherein: The modification includes widening an amount of the sound capturing beam.
9. The device according to claim 7 as appended to claim 5, wherein: The modification includes a direction and an amount of turning the sound capturing beam from the first direction to the direction of the one of the two or more specific physical speakers.
10. The apparatus according to any preceding claim, further comprising: means for receiving spatial metadata associated with the audio data, the spatial metadata indicating spatial characteristics of an audio scene including at least the first sound source, The component for determining is configured to: determine, based on the spatial metadata, that the first sound source will be perceived as having the first direction relative to the user that is different from a physical speaker direction.
11. The device according to claim 10, wherein: The audio data and the spatial metadata are received in an Immersive Voice and Audio Services (IVAS) bitstream.
12. The device according to claim 11, wherein The IVAS bitstream includes one of the following: Metadata-assisted spatial audio MASA; OMASA object with metadata-assisted spatial audio; and Independent stream ISM with metadata.
13. Apparatus according to any preceding claim, comprising: Mobile terminal.
14. An apparatus comprising: means for capturing audio signals output by two or more physical speakers, the audio signals including an audio signal representative of a first sound source output by the two or more particular physical speakers such that the first sound source will be perceived as having a first direction relative to a user that is different than a direction of the physical speakers; means for operating in a directional mode for steering the sound capturing beam in said first direction; as well as A component for receiving control data from a control device, wherein the control data causes disabling the directivity pattern or modifying the sound capture beam to make the apparatus more sensitive to audio signals from the direction of at least one of the two or more specific physical speakers.
15. The apparatus according to claim 14, further comprising: means for sending a notification message to the control device indicating that the apparatus is operating in the directional mode, and The control data is received from the control device in response to sending the notification message.
16. The device according to claim 14 or claim 15, wherein: The control data causes the sound capturing beam to be widened so that the apparatus has a higher sensitivity to audio signals from a wider range of directions, the wider range of directions including the direction of the at least one specific physical speaker of the two or more specific physical speakers.
17. The device according to claim 16, wherein: The control data causes the sound capturing beam to be widened so that the apparatus has a higher sensitivity to audio signals from corresponding directions of the two or more specific physical speakers.
18. The apparatus of claim 14 or claim 15, wherein: The control data causes the sound capturing beam to turn from the first direction to the direction of a specific physical speaker among the two or more specific physical speakers.
19. The device according to any one of claims 16 to 18, wherein: The control data comprises data indicative of a spatial position of the at least one physical speaker of the two or more physical speakers, and The apparatus further comprises means for estimating a direction or a corresponding direction of the at least one specific physical speaker of the two or more specific physical speakers.
20. The device according to any one of claims 16 to 18, further comprising: means for sending position data indicative of the spatial position of the apparatus and the direction of the sound capturing beam to the control device; Wherein the control data comprises a modification to be applied to the sound capturing beam determined based on the position data and a known position of the at least one specific physical speaker of the two or more specific physical speakers.
21. Apparatus according to claim 20 when appended to claim 16 or claim 17, wherein: The modification includes widening an amount of the sound capturing beam.
22. The device according to claim 20 as appended to claim 18, wherein: The modification includes a direction and an amount of turning the sound capturing beam from the first direction to the direction of the one of the two or more specific physical speakers.
23. A method comprising: receiving audio data representing an audio signal for output by two or more physical speakers; determining that at least some of the audio signals representing a first sound source are intended to be output by two or more specific physical speakers so that the first sound source will be perceived as having a first direction relative to a user that is different from a direction of the physical speakers; as well as In response to the determination, control data is sent to an audio capture device of the user, the audio capture device operating in a directional mode for steering a sound capture beam toward the first direction, wherein the control data is used to cause the audio capture device to disable its directional mode or modify the sound capture beam so that the audio capture device has a higher sensitivity to audio signals from the direction of at least one specific physical speaker among the two or more specific physical speakers.
24. A method comprising: capturing audio signals output by two or more physical speakers, the audio signals including an audio signal representative of a first sound source output by the two or more specific physical speakers such that the first sound source will be perceived as having a first direction relative to a user that is different from a direction of the physical speakers; operating in a directional mode for steering a sound capturing beam in said first direction; as well as Control data is received from a control device, wherein the control data causes the directivity pattern to be disabled or the sound capture beam to be modified so that the apparatus has a higher sensitivity to audio signals from the direction of at least one of the two or more specific physical speakers.