Output of audio signal

By detecting the operating status of the audio capture device and selecting the closest speaker to output the audio signal, the problem of the user not being able to get the best experience in directional mode is solved, and more accurate sound capture and rendering is achieved.

CN119946545APending Publication Date: 2025-05-06NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411506849.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-10-28
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Users wearing certain audio capture devices may not have the best user experience when listening to audio signals output by two or more physical speakers, especially in directional mode.

Method used

By detecting that the user's audio capture device operates in a directional mode, select the physical speaker closest to the user, output the audio signal of the first sound source to the speaker, and then adjust the direction of the sound capture beam to improve the user experience.

Benefits of technology

This approach can effectively improve the user experience and ensure that the sound capture beam is accurately directed to the selected speaker, thereby optimizing the rendering and capture of the audio signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946545A_ABST
    Figure CN119946545A_ABST
Patent Text Reader

Abstract

Example embodiments relate to the output of an audio signal, and in particular to the use of two or more physical speakers to render one or more sound sources represented by such an audio signal. In an example method, a method is disclosed that includes rendering at least a first sound source by outputs of audio signals from two or more physical speakers having different respective locations such that the first sound source is intended to be perceived as having a first direction different from the physical speaker direction relative to a user. The method may also include detecting that an audio capture device of the user is operating in a directionality mode for steering the sound capture beam in a first direction. The method may also include, in response to the detection, performing a modified rendering by outputting an audio signal of the first sound source from a selected physical speaker from the two or more physical speakers rather than from other physical speakers such that the first sound source will be perceived from a direction of the selected physical speaker, thus, a sound capture beam of the audio capture device will be diverted to the selected physical speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments relate to the output of audio signals, and in particular to the use of two or more physical speakers to render one or more sound sources represented by such audio signals. Background Art

[0002] Some audio signal formats are suitable for being output by two or more physical speakers. Such audio signal formats may include stereo, multi-channel, and immersive formats. By using two or more physical speakers to output audio signals, a listening user may perceive one or more sound objects from a specific direction different from the direction of the physical speakers.

[0003] Users wearing certain audio capture devices may not have an optimal user experience when listening to audio signals output by two or more physical speakers. Summary of the invention

[0004] The scope of protection sought by various embodiments of the present invention is set out in the independent claims. Embodiments and features described in this specification that do not fall under the scope of the independent claims, if any, should be interpreted as examples that aid in understanding the various embodiments of the present invention.

[0005] A first aspect provides an apparatus comprising: a component for rendering at least a first sound source by output of audio signals from two or more physical speakers having different corresponding positions, so that the first sound source is intended to be perceived as having a first direction relative to a user that is different from the direction of the physical speakers; and a component for detecting that the user's audio capture device is operating in a directional mode for steering a sound capture beam toward the first direction; wherein the component for rendering is configured to: in response to the detection, perform modified rendering by outputting an audio signal of the first sound source from a physical speaker selected from the two or more physical speakers but not from other physical speakers, so that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered toward the selected physical speaker.

[0006] The selected physical speaker may have an orientation closest to the first direction relative to the user.

[0007] The apparatus may further include means for detecting that the first sound source is of user interest, wherein the means for rendering is further configured to perform modified rendering in response to detecting that the audio capture device is operating in a directional mode only when the first sound source is detected as of user interest.

[0008] The means for detecting that the first sound source is of interest to the user may be configured to detect that the first sound source is a predetermined type of sound. The predetermined type of sound may be a speech type sound.

[0009] The apparatus may further include: a component for determining a head direction of a user, wherein the component for detecting that the first sound source is of interest to the user is configured to: detect that the first direction is within a predetermined angular range of the head direction of the user.

[0010] The component for rendering can be configured to: render one or more other sound sources through the output of other audio signals from two or more physical speakers so that they are intended to be perceived as coming from corresponding directions relative to the user, and wherein, for only the first sound source and not for the other sound sources, a modified rendering is performed so that the first sound source will be perceived from the direction of the selected physical speaker.

[0011] The apparatus may further comprise: a component for determining that, for the other audio signals of the one or more other sound sources, a first group of the other audio signals are outputted by or intended to be outputted by only the selected physical speakers, and a second group of the other audio signals are outputted by or intended to be outputted by one or more other physical speakers, wherein the component for rendering is configured to: in response to the determination, perform other modified rendering of the first group of other audio signals and / or the second group of other audio signals of one or more other sound sources.

[0012] The further modified rendering may include rendering the first set of audio signals of one or more other sound sources with reduced reverberation; and / or rendering the second set of audio signals of one or more other sound sources with increased reverberation.

[0013] The other modified rendering may include outputting the first set of audio signals of one or more other sound sources from different physical speakers to the selected physical speakers.

[0014] For a particular other sound source, different physical speakers may have directions relative to the user that are closest to the direction of the particular other sound source relative to the user.

[0015] The further modified rendering may include rendering one or more other sound sources by rendering the first set of audio signals thereof at a reduced volume.

[0016] The apparatus may further include: means for determining a respective type of audio content comprising the first sound source and one or more other sound sources; and means for determining an amount of the further modified rendering to be performed based on the determined respective type of audio content.

[0017] The apparatus may further include means for receiving metadata associated with audio content including the first sound source and one or more other sound sources; and means for determining an amount of the other modified rendering to be performed based on the received metadata.

[0018] The means for rendering of at least the first audio source may comprise an MPEG-I renderer.

[0019] A second aspect provides a method comprising: rendering at least a first sound source by output of audio signals from two or more physical speakers having different corresponding positions so that the first sound source is intended to be perceived as having a first direction relative to a user that is different from the direction of the physical speakers; detecting that the user's audio capture device is operating in a directional mode for steering a sound capture beam toward the first direction; and in response to the detection, performing a modified rendering by outputting an audio signal of the first sound source from a selected physical speaker among the two or more physical speakers but not from other physical speakers so that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered toward the selected physical speaker.

[0020] The selected physical speaker may have an orientation closest to the first direction relative to the user.

[0021] The method may further comprise detecting that the first sound source is of user interest, wherein the modified rendering is performed in response to detecting that the audio capture device is operating in the directional mode only when the first sound source is detected as of user interest.

[0022] Detecting that the first sound source is of interest to the user may include detecting that the first sound source is a predetermined type of sound. The predetermined type of sound may be a speech type sound.

[0023] The method may further include determining a head direction of the user, wherein detecting that the first sound source is of interest to the user may include detecting that the first direction is within a predetermined angular range of the head direction of the user.

[0024] Rendering may include: rendering one or more other sound sources through output of other audio signals from two or more physical speakers so that they are intended to be perceived as coming from corresponding directions relative to the user, and wherein, for only the first sound source and not for the other sound sources, a modified rendering is performed so that the first sound source will be perceived from the direction of the selected physical speaker.

[0025] The method may further include determining that, for the other audio signals of the one or more other sound sources, a first group of the other audio signals are output only by the selected physical speakers or are intended to be output only by the selected physical speakers, and a second group of the other audio signals are output by one or more other physical speakers or are intended to be output by one or more other physical speakers, and wherein the rendering may include performing other modified rendering of the first group of other audio signals and / or the second group of other audio signals of one or more other sound sources in response to the determination.

[0026] The further modified rendering may include rendering the first set of audio signals of one or more other sound sources with reduced reverberation; and / or rendering the second set of audio signals of one or more other sound sources with increased reverberation.

[0027] The other modified rendering may include outputting the first set of audio signals of one or more other sound sources from different physical speakers to the selected physical speakers.

[0028] For a particular other sound source, different physical speakers may have directions relative to the user that are closest to the direction of the particular other sound source relative to the user.

[0029] The further modified rendering may include rendering one or more other sound sources by rendering the first set of audio signals thereof at a reduced volume.

[0030] The method may further include determining a respective type of audio content comprising the first sound source and one or more other sound sources; and determining an amount of the further modified rendering to be performed based on the determined respective type of audio content.

[0031] The method may further comprise: receiving metadata associated with audio content comprising the first sound source and one or more other sound sources; and determining an amount of the other modified rendering to be performed based on the received metadata.

[0032] The rendering of at least the first audio source may be performed by an MPEG-I renderer.

[0033] A third aspect provides a computer program comprising a set of instructions which, when executed on a device, are configured to cause the device to perform a method comprising: rendering at least a first sound source by output of audio signals from two or more physical speakers having different corresponding positions so that the first sound source is intended to be perceived as having a first direction relative to a user that is different from the direction of the physical speakers; detecting that the user's audio capture device is operating in a directional mode for steering a sound capture beam toward the first direction; and in response to the detection, performing a modified rendering by outputting an audio signal of the first sound source from a selected physical speaker among the two or more physical speakers but not from other physical speakers so that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered toward the selected physical speaker.

[0034] In some example embodiments, the third aspect may include any other features mentioned with respect to the method of the second aspect.

[0035] A fourth aspect of the present invention provides a non-transitory computer-readable medium having computer-readable code stored thereon, which, when executed by at least one processor, causes the at least one processor to perform a method comprising: rendering at least a first sound source by outputting audio signals from two or more physical speakers having different corresponding positions so that the first sound source is intended to be perceived as having a first direction relative to a user that is different from the direction of the physical speakers; detecting that the user's audio capture device is operating in a directional mode for steering a sound capture beam toward the first direction; and in response to the detection, performing a modified rendering by outputting an audio signal of the first sound source from a physical speaker selected from the two or more physical speakers but not from other physical speakers so that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered toward the selected physical speaker.

[0036] The fourth aspect may include any other features mentioned in relation to the method of the second aspect.

[0037] A fifth aspect of the present invention provides an apparatus having at least one processor and at least one memory having computer-readable code stored thereon, which, when executed, controls the at least one processor to: render at least a first sound source by outputting audio signals from two or more physical speakers having different corresponding positions, so that the first sound source is intended to be perceived as having a first direction relative to a user that is different from the direction of the physical speakers; detect that the user's audio capture device is operating in a directional mode for steering a sound capture beam toward the first direction; and in response to the detection, perform modified rendering by outputting an audio signal of the first sound source from a physical speaker selected from the two or more physical speakers but not from other physical speakers, so that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capture beam of the audio capture device to be steered toward the selected physical speaker.

[0038] The fifth aspect may include any other features mentioned in relation to the method of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The invention will now be described by way of non-limiting example with reference to the accompanying drawings, in which:

[0040] Figure 1 is a schematic diagram of a system for rendering that is useful for understanding one or more example embodiments;

[0041] Figure 2 Indicates the direction of a sound source relative to the user Figure 1 Schematic diagram of the system;

[0042] Figure 3 is a schematic diagram of an audio capture device useful for understanding one or more example embodiments;

[0043] Figure 4 is a flowchart illustrating operations according to one or more example embodiments;

[0044] Figure 5 is a schematic diagram of a system for rendering audio according to one or more example embodiments;

[0045] Figure 6 is a schematic diagram of a system for rendering audio according to one or more other example embodiments;

[0046] Figure 7 is a schematic diagram of a system for rendering audio according to one or more other example embodiments;

[0047] Figure 8 is a schematic diagram of a system for rendering audio according to one or more other example embodiments;

[0048] Fig. 9 is a block diagram of an apparatus that may be configured according to one or more example embodiments; and

[0049] Fig.10 is a non-transitory computer-readable medium according to one or more example embodiments. DETAILED DESCRIPTION

[0050] Example embodiments relate to using two or more physical speakers to render one or more sound sources.

[0051] The example embodiments focus on immersive audio, but it will be appreciated that other audio formats for output by two or more physical speakers, including but not limited to stereo and multi-channel audio formats, are also applicable.

[0052] Immersive audio in this context can refer to any technology that renders sound objects in a space so that a listening user in that space can perceive one or more sound objects as coming from corresponding directions in that space. The user can also perceive a sense of depth.

[0053] Immersive audio in this context may include any technology such as surround sound and different types of spatial audio technologies that utilize two or more physical speakers with corresponding spaced-apart positions to provide an immersive audio experience. Ambisonics and MPEG-I are example immersive audio formats, but example embodiments are not limited to such examples.

[0054] Figure 1A system 100 for outputting immersive audio is shown, the system including an audio processor 102 (sometimes referred to as an audio receiver or audio amplifier) ​​and first to fifth physical speakers 104A-104E (hereinafter referred to as "speakers"), which are spaced apart from each other and have corresponding positions in a listening space 105, which may be a room. The first, second and third speakers 104A, 104B, 104C may be referred to as left front, right front and center front speakers based on their respective positions relative to a typical listening position (indicated by reference numeral 106). Similarly, the fourth and fifth speakers 104D, 104E may be referred to as left rear and right rear speakers based on their respective positions relative to the listening position 106. There may also be another speaker, not shown, for outputting lower frequency audio signals, and it may be referred to as a subwoofer, a woofer, etc. In some example embodiments, there may be fewer speakers. Thus, system 100 may represent a 5.1 surround sound setup, but it will be appreciated that many other setups are possible, such as but not limited to 2.0, 2.1, 3.1, 4.0, 4.1, 5.1, 6.1, 7.1, 7.1.2, 7.2, 9.1, 9.1.2, 10.2, and 13.1.

[0055] The audio processor 102 may be configured to store audio data representing immersive audio content for output via the first to fifth speakers 104A-104E. The audio processor 102 may include an amplifier, a signal processing function, one or more memories, for example, a hard disk drive (HDD) and / or a solid state drive (SSD) for storing audio data. The audio processor 102 may be provided in any suitable form (such as a set-top box, a mobile phone, a tablet computer, etc.). The audio processor 102 may be a digital processor only, in which case it may not include an amplifier. For example, audio data may be received from a remote source 108 via a network 110 and stored on one or more memories. The network 110 may include the Internet. The audio data may be received via a wired or wireless connection to the network 110 (such as via a home router or hub). Alternatively, the audio data may be streamed from a remote source 108 using a suitable streaming protocol (e.g., a real-time streaming protocol (RTSP) etc.). Alternatively, the audio data may be provided on a non-transitory computer readable medium, such as a compact disc, memory card, memory stick or removable hard disk inserted into or connected to a suitable portion of the audio processor 102 .

[0056] The audio data may represent an audio signal in any audio form, whether speech, singing, music, environment / ambience or a combination thereof. The audio data may be associated with video data, for example, as part of a video clip, video game or movie.

[0057] The audio processor 102 may be configured to render audio data by outputting an audio signal using a suitable speaker among the first to fifth speakers 104A-104E. Therefore, the audio processor 102 may include a rendering component, which may include hardware, software and / or firmware configured to process (or render) an audio signal and output the audio signal to the suitable speaker among the first to fifth speakers 104A-104E. The audio processor 102 may also provide other signal processing functions, such as modifying the overall volume, modifying the corresponding volume for different frequency ranges and / or performing certain effects, such as modifying reverberation and / or performing translation such as vector basis amplitude translation (VBAP). VBAP is a method for localizing a sound source to an arbitrary direction using the current speaker setting; the number of speakers is arbitrary because they can be positioned in a 2-dimensional or 3-dimensional setting. VBAP produces a virtual source confined to a relatively narrow area. VBAP processing may involve finding a speaker triplet (i.e., three speakers) surrounding the desired sound source translation position, and then calculating the gain of the audio signal to be applied to the sound source so that three speakers will be used for reproduction. The audio processor 102 may implement VBAP, for example. An alternative approach is Speaker Placement Correction Amplitude Panning (SPCAP).

[0058] The audio data may include metadata or other computer-readable instructions that the audio processor 102 processes to determine how to render the audio signal, for example, by which of the first to fifth speakers 104A-104E to render it. The audio signal may be arranged into channels, for example, one for each of the first to fifth speakers 104A-104E.

[0059] In some cases, only a subset of the first through fifth speakers 104A- 104E may be used.

[0060] In some cases, metadata or other computer-readable indications may determine certain effects to be applied to the audio signal at certain times during the output of the audio data. For example, the audio data may accompany a movie, in which certain channels or sound sources may be amplified, attenuated, or have certain effects, such as panning and / or reverb modifications, used at specific times.

[0061] Through the output of audio signals from two or more of the first to fifth speakers 104A- 104E, the audio processor 102 can render the sound source so that it will be perceived by the user as coming from a direction different from the direction of (any one of) the first to fifth speakers relative to the user.

[0062] Figure 2 The first sound source 200 is shown. Figure 1system, a first sound source 200 is indicated to be located between the first and third speakers 104A, 104C so that it will be perceived by a user at position 106 as coming from a first direction 202 relative to the user.

[0063] In this example, the audio processor 102 may render the first sound source 200 using the first and third speakers 104A, 104C via VBAP or the like.

[0064] The same process may be performed for one or more other sound sources (not shown) so that they will be perceived by the user as coming from corresponding directions relative to the user.

[0065] Users wearing certain audio capture devices may not have the best user experience when experiencing immersive audio, for example Figure 2 This is particularly true for audio capture devices that can operate in directional or accessibility modes for hearing assistance, such as hearing aids or headphone devices.

[0066] Figure 3 300 is a schematic diagram of an example audio capture device (including earphone 300). Although not shown, earphone 300 may include one earphone in a pair of earphones. Earphone 300 may include speaker 302 and microphone array 304, and in use, speaker 302 is placed on or in the ear of the user. Earphone 300 may be configured to provide hearing assistance when operating in so-called directional (or accessibility) mode (which may be a default mode) in use, or enabled by means of a control input to the earphone or by another device (such as a user device 306 that communicates with the earphone pairing). The control input may be provided by any suitable means / components, for example, touch input, gesture, or voice input.

[0067] The microphone array 304 may be configured to steer the sound capturing beam 308 toward the perceived direction of a particular sound (such as a particular sound object) or toward a direction relative to the headset, such as a frontal direction.

[0068] More specifically, the earphone 300 may include a signal processing function 310 that spatially filters the surrounding audio field so that sounds from one or more specific directions or from a predetermined range of directions are amplified over sounds from other directions. These directions effectively form the referenced sound capture beam 308. It will be seen that the direction and size of the sound capture beam 308 can be steered / directed under the control of the signal processing function 310, which amplifies the sound captured within the sound capture beam and passes it to the speaker 302.

[0069] The signal processing functionality 310 may be configured using known methods to steer the sound capturing beam 308 in the direction of one or more specific sound objects or relative to the earphones.

[0070] The specific sound object may include a predetermined type of sound object, such as a speech sound object and / or a sound object in a specific direction relative to the headset, for example, toward its front side. The audio processor 102 may infer that the sound object is important to the user based on the predetermined type or corresponding direction of the sound object.

[0071] Back to Figure 2 , if the user at location 106 is wearing an audio capture device (e.g., headphones 300) operating in a directional mode, the sound capture beam 308 can be directed by the signal processing function 310 in the first direction 202 because it is the perceived direction of the first sound source 200. However, the amplification will likely be suboptimal and the clarity / intelligibility of the first sound source 200 may be affected. The amplification may be suboptimal because the sound capture beam 308 is directed to a location without a speaker, and attenuation may be performed on audio signals outside the sound capture beam (e.g., speaker audio signals). In addition, the size and / or steering / steering of the sound capture beam 308 caused by the signal processing function 310 may be affected. Overall, the user experience may be negatively affected.

[0072] According to one or more example embodiments, the rendering of one or more sound sources may be modified to alleviate this problem.

[0073] Figure 4 4 is a flow chart illustrating operation 400 that may be performed by one or more example embodiments. Operation 400 may be performed by hardware, software, firmware, or a combination thereof. Operation 400 may be performed by one or a corresponding component, which is any suitable component, such as a combination of one or more processors or controllers and computer readable instructions provided on one or more memories. Operation 400 may be performed by, for example, a processor or controller that has been described in detail with respect to FIG. Figure 2 The example described is performed by the audio processor 102.

[0074] The first operation 401 may include rendering at least a first sound source through output of audio signals from two or more physical speakers having different respective positions such that the first sound source is intended to be perceived as having a first direction relative to a user different from a direction of the physical speakers.

[0075] A second operation 402 may include detecting that an audio capturing device of a user is operating in a directional mode for steering a sound capturing beam toward a first direction.

[0076] The third operation 403 may include: in response to the detection, performing modified rendering by outputting the audio signal of the first sound source from a selected physical speaker among the two or more physical speakers instead of from other physical speakers, so that the first sound source will be perceived from the direction of the selected physical speaker, so that the sound capturing beam of the audio capturing device will be steered toward the selected physical speaker.

[0077] In this way, an audio capture device operating in directional mode will direct its sound capturing beam towards the selected physical speaker, alleviating the above-mentioned problems.

[0078] Figure 5 A system 500 for outputting immersive audio according to one or more example embodiments is shown.

[0079] System 500 and Figure 2 The system 500 includes an audio processor 502, which includes a rendering component 504, which is configured to perform reference Figure 4 Operation 400 is described.

[0080] The rendering component 504 may be configured to operate at a first time according to the first operation 401 .

[0081] Thus, the rendering component 504 may output or is intended to output the audio signal of the first sound source 200 from the first and third speakers 104A, 104C. The first sound source 200 is perceived or is intended to be perceived as coming from the first direction 202.

[0082] The rendering component 504 or another component or function of the audio processor 502 may be configured to operate according to the second operation 402 .

[0083] That is, the rendering component 504 or other component or function can detect that the user at the location 106 is wearing an audio capture device (in this case, Figure 3 ), the audio capturing device operates in a directional mode for steering a sound capturing beam 505 (shown in dashed lines) toward the first direction 202.

[0084] The second operation 402 may involve the rendering component 504 or other component or function receiving a signal indicating that the headset 300 is operating in a directional mode.

[0085] The received signal may be sent by the headset 300 or an associated device, such as the user device 306 .

[0086] The received signal may be sent in response to a discovery signal sent by the rendering component 504 or other components or functions of the audio processor 502 to the headset 300 or the user device 306. Alternatively, the signal may be sent in response to the user enabling the directional mode at the headset 300 during the execution of the first operation 401. The signal communication between the audio processor 502 and the headset 300 or the user device 306 may be by means of any suitable wireless protocol, such as by WiFi, Bluetooth, Zigbee or any variant thereof. For example, there may be a pairing relationship between the audio processor 502 and the headset 300, and the link and signaling between the devices are automatically established when the latter is in the communication range of the former.

[0087] The second operation 402 may be performed in response to determining that the headset 300 is in proximity to the audio processor 502 .

[0088] Then, the rendering component 504 may responsively perform the third operation 403 .

[0089] That is, the rendering component 504 modifies the rendering of the audio signal of the first sound source 200 by outputting it from the first speaker 104A (in this case, the first speaker 104A) rather than from the third speaker 104C.

[0090] Thus, the earphone 300 steers its sound capturing beam 505 towards a second direction 508 aligned with the first speaker 104 A. This avoids or mitigates the above mentioned disadvantages.

[0091] In some example embodiments, the audio processor 502 may be configured such that the selected speaker has a direction closest to the first direction relative to the user.

[0092] In this regard, the audio processor 502 can be configured to determine the direction of at least the first and third speakers 104A, 104C relative to the user position 106 (e.g., based on knowing or determining their respective positions), and can select the third speaker based on knowing the first direction relative to the user. In another approach, the audio processor 502 can be further configured to determine the expected spatial position 106 of the first sound source 200 relative to at least the first and third speakers 104A, 104C based on the first direction 202. Figure 6 As shown in FIG. 1 , the audio processor 502 may determine that the third speaker 104C has a Figure 2 The earphone 300 directs its sound capturing beam 505 towards the third speaker 104C, which may further improve clarity / intelligibility.

[0093] In some example embodiments, the third operation 403 is further performed in response to the audio processor 502 detecting that the first sound source is of interest to the user. For example, if the first sound source is a predetermined type of sound (such as a speech type sound), the third operation 403 may be performed. The type of sound may be indicated with the audio data, for example, in metadata, or may be determined using a signal processing method (such as determining the audio type by applying the audio data to one or more classifier models).

[0094] Alternatively or additionally, if the expected direction of the first sound source 200 (i.e., the first direction 202) corresponds to the user's head direction, the third operation 403 may be performed by the audio processor 502. In this regard, the audio processor 502 may determine the user's head direction with knowledge of the user's position 106, and if the first direction 202 is within a predetermined angular range of the user's head direction, the first sound source 200 is detected to be of interest. The user's position 106 may be determined by the audio processor 502 using known methods (such as by using ranging signals sent from or to a reference position and multilateration processing), or determined by another device using known methods and sent to the audio processor.

[0095] The user's head orientation may be determined using conventional methods, such as based on the orientation of the headset 300 when worn. It may be assumed that the front portion of the headset 300 corresponds to the user's head orientation.

[0096] In some example embodiments, the first operation 401 may include rendering or intending to render one or more other sound sources so that they, or at least some, are intended to be perceived as coming from respective directions relative to the user's position 106. Again, this may be accomplished by the audio processor 502 using two or more of the first to fifth speakers 104A-104E to output other audio signals of the one or more other sound sources.

[0097] In this case, the audio processor 502 may perform the third operation 403 only for the first sound source 200 (based on it being of interest to the user). The rendering of other sound sources may remain unaffected, or may be modified in one or more other ways, as will be explained below.

[0098] Figure 6 Shows Figure 5 System 500, wherein second through fifth sound sources 611-614 are shown as being rendered in respective directions relative to the user.

[0099] It will be seen that only the first sound source 200 undergoes a modified rendering, which effectively moves it to the direction of the third speaker 104C. This may be because the first sound source 200 is detected to be of interest to the user, for example because it is speech and / or it is within the direction of the user's head.

[0100] The third speaker 104C may be selected because it is closest to the first direction relative to the direction of the user and / or because it is the closest expected spatial location of the first sound source. Thus, the sound capture beam 604 of the headset 300 is steered in the direction of the third speaker 104C.

[0101] Some example embodiments may comprise applying different forms of modified rendering to audio signals of at least some other sound sources to further enhance the user experience.

[0102] Figure 7 Shows Figure 5 System 500, wherein the second to fifth sound sources 611-614 are shown as being rendered in respective directions relative to the user. Figure 6 Modified rendering described (movement of the first sound source 200).

[0103] It will be seen that the second sound source 611 is rendered using the first set of audio signals 701 from the third speaker 104C and the second set of audio signals 702 from the second speaker 104B. It will also be seen that the fifth sound source 614 is rendered using the third set of audio signals 703 from the third speaker 104C and the fourth set of audio signals 704 from the first speaker.

[0104] Since the third speaker 104C is the selected speaker for the first sound source 200 in this case, the following modification may be performed.

[0105] In one example, the first set of audio signals 701, 703 of the second and fifth sound sources 611, 614 may be rendered with reduced reverberation (clean / reverb ratio) using one or more known methods (e.g., as described in the MPEG-I standard).

[0106] As an alternative or in addition to the above, the second set of audio signals 702, 704 of the second and fifth sound sources 611, 614 may be rendered with increased reverberation (clean / reverberation ratio) using one or more known methods (e.g., as described in the MPEG-I standard). If this method is used, this may be used to compensate for the reduced reverberation of the first set of audio signals 701, 703.

[0107] As an alternative or in addition to the above, the first set of audio signals 701, 703 of the second and fifth sound sources 611, 614 may be rendered with a reduced (or muted) output volume. In this way, the audio signal of the first sound source 200 will tend to mask the audio signals 701, 703 of the second and fifth sound sources 611, 614, especially when they share a reasonable amount of common frequencies.

[0108] As an alternative or in addition to the above, a smaller number of loudspeakers may be used to render at least a portion of the audio signal without panning.

[0109] As an alternative or in addition to the above, at least a portion of the audio signal may be rendered using a smaller number of loudspeakers.

[0110] Alternatively, at least the first set of audio signals 701, 703 of the second and fifth sound sources 611, 614 may be output by one or more different speakers (i.e., speakers different from the third speaker 104C) so that the second and fifth sound sources 611, 614 are perceived as coming from different directions, i.e., the directions of the one or more other speakers.

[0111] Figure 8 The audio signals of the second and fifth sound sources 611, 614 are shown to be outputted by the second and first speakers 104B, 104A, respectively, and the third speaker 104C makes an audio signal contribution. In this way, the respective perceived directions of the second and fifth sound sources 611, 614 are changed, making the first sound source 200 more perceptible, while keeping the second and fifth sound sources in the overall audio scene.

[0112] In some example embodiments, different speakers 104B, 104A are selected based on which speaker is closest to the expected spatial location of the second and fifth sound sources 611 , 614 .

[0113] Example embodiments may be performed using object rendering with a powerful renderer such as an MPEG-I renderer.

[0114] In some example embodiments, the amount of sound source modification (such as the amount of change in the perceived direction of one or more sound sources) may depend on the type of audio content that includes the sound sources.

[0115] For example, an example embodiment may include determining a type of audio content, and determining an amount of sound source modification to perform based on the determined type.

[0116] For example, certain types of audio content may be processed differently than other types of audio content. For example, music may be processed differently than ambient / ambience. For example, if audio content accompanies video content, the amount of sound source modification may be different than if it is not accompanied by video content. If there is accompanying video content, for example, it may be assumed (or indicated in the accompanying metadata) that the audio direction of one or more sound sources is critical; for example, speech may be considered critical to be rendered from appropriately positioned speakers or speakers corresponding to the direction of the user's head, while other audio sources (e.g., ambient / ambience sounds) may be less important, and one or more other effects (e.g., moving to other speakers) may be used.

[0117] The content creator may indicate one or more preferences, for example via metadata associated with the audio content, which indicate which modifications are permitted for which sound sources and / or when during the rendering process. For example, the metadata may indicate the degree of deviation from the original sound source direction that is permitted at certain times compared to the improved clarity / intelligibility due to the modifications. The metadata may be embedded in the scene data, for example, in the accessibility model of MPEG-I.

[0118] Alternatively or additionally, the user may determine which modifications are permitted for which sound sources and / or when during the rendering process. The user may provide input to the renderer via a suitable user interface to set one or more preferences in this regard, for example, via the audio processor 502.

[0119] Example embodiments are applicable to both object-based and non-object-based audio rendering methods. Ambisonics is an example of a non-object-based audio rendering method, whose rendering may include beamforming at the signal level to focus on important sound sources (such as speech), and adjusting the direction of the beam towards the physical speakers in the output rendering (Ambisonics panning) to achieve an experience similar to having an object. Therefore, the Ambisonics signal can be rotated during panning so that the position of one or more sound sources of interest coincides with the speaker position, resulting in a clearer reproduction. Ambisonics beamforming can be used to enhance sound sources. Speaker channel-based methods (such as 5.1) are examples of non-object-based audio rendering methods. All channels can be modified so that fewer speakers are used to render channel-based signals.

[0120] Example Device

[0121] Fig. 9An apparatus according to some example embodiments is shown. The apparatus may be configured to perform the operations described herein, for example, the operations described with reference to any disclosed process. The apparatus includes at least one processor 900 and at least one memory 901 directly or closely connected to the processor. The memory 901 includes at least one random access memory (RAM) 901a and at least one read-only memory (ROM) 901b. Computer program code (software) 906 is stored in the ROM 901b. The apparatus may be connected to a transmitter (TX) and a receiver (RX). The apparatus may optionally be connected to a user interface (UI) for instructing the apparatus and / or for outputting data. At least one processor 900 together with at least one memory 901 and computer program code 906 are configured to cause the apparatus to at least perform operations according to any previous process (for example, as described with respect to Figure 4 The method disclosed by the flowchart and its related features).

[0122] Fig.10 A non-transitory medium 1000 according to some embodiments is shown. The non-transitory medium 1000 is a computer readable storage medium. It may be, for example, a CD, a DVD, a USB stick, a Blu-ray disc, etc. The non-transitory medium 1000 stores computer program instructions so that the device performs any of the previous processes (e.g., as described in relation to Figure 4 The method disclosed by the flowchart and its related features).

[0123] The names of network elements / components, protocols, and methods are based on current standards. In other versions or other technologies, the names of these network elements / components and / or protocols and / or methods may be different, as long as they provide corresponding functions. For example, embodiments may be deployed in 2G / 3G / 4G / 5G networks and subsequent generations of 3GPP, and may also be deployed in non-3GPP radio networks such as WiFi.

[0124] The memory may be volatile or non-volatile. It may be, for example, RAM, SRAM, flash memory, FPGA block RAM, DCD, CD, USB stick, and Blu-ray disc.

[0125] If not otherwise stated or clear from the context, a statement that two entities are different means that they perform different functions. This does not mean that they are based on different hardware. That is, each entity described in this specification can be based on different hardware, or some or all entities can be based on the same hardware. This does not mean that they are based on different software. That is, each entity described in this specification can be based on different software, or some or all entities can be based on the same software. Each entity described in this specification can be embodied in the cloud.

[0126] As non-limiting examples, implementation of any of the above boxes / blocks, devices, systems, techniques or methods includes implementation as hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof. Some embodiments may be implemented in the cloud.

[0127] It will be appreciated that the foregoing is what is presently considered to be the preferred embodiment. It should be noted, however, that the description of the preferred embodiment is given by way of example only and that various modifications may be made without departing from the scope as defined in the appended claims.

Claims

1. A device comprising: for rendering at least a first sound source by output of audio signals from two or more physical speakers having different respective positions such that the first sound source is intended to be perceived as having a first direction relative to a user different from the direction of the physical speakers; as well as means for detecting that the user's audio capturing device is operating in a directional mode for steering a sound capturing beam towards the first direction; Wherein, the component for rendering is configured to: in response to the detection, perform modified rendering by outputting the audio signal of the first sound source from a physical speaker selected from the two or more physical speakers instead of from other physical speakers, so that the first sound source will be perceived from the direction of the selected physical speaker, so that the sound capture beam of the audio capture device will be directed to the selected physical speaker.

2. The device according to any one of the preceding claims, wherein: The selected physical speaker has an orientation closest to the first orientation relative to the user.

3. The device according to claim 1 or claim 2, further comprising: for detecting that the first sound source is a component of interest to the user, Wherein the means for rendering is further configured to: in response to detecting that the audio capture device is operating in a directional mode only when the first sound source is detected as being of interest to the user, perform the modified rendering.

4. The device according to claim 3, wherein: The means for detecting that the first sound source is of interest to the user is configured to detect that the first sound source is a predetermined type of sound.

5. The device according to claim 4, wherein: The predetermined type of sound includes a voice type sound.

6. The device according to any one of claims 3 to 5, further comprising: means for determining the orientation of the user's head, The component for detecting that the first sound source is of interest to the user is configured to detect that the first direction is within a predetermined angle range of the head direction of the user.

7. The device according to any one of claims 3 to 6, in, the means for rendering being configured to render one or more other sound sources through output of other audio signals from the two or more physical speakers so that they are intended to be perceived as coming from respective directions relative to the user, and Therein, the modified rendering is performed only for the first sound source and not for the other sound sources, so that the first sound source will be perceived from the direction of the selected physical speaker.

8. The apparatus according to claim 7, further comprising: means for determining that, for said one or more other sound sources, a first set of said other audio signals is output or intended to be output only by the selected physical speaker, and a second set of said other audio signals is output or intended to be output by one or more other physical speakers; Therein, the means for rendering is configured to: in response to the determination, perform further modified rendering of the first set of further audio signals and / or the second set of further audio signals of the one or more further sound sources.

9. The device according to claim 8, wherein: Said other modified renderings include: rendering the first set of audio signals of the one or more other sound sources with reduced reverberation; and / or The second set of audio signals of the one or more other sound sources are rendered with increased reverberation.

10. The device according to claim 8 or claim 9, wherein: The other modified rendering includes outputting the first set of audio signals of the one or more other sound sources from different physical speakers to the selected physical speakers.

11. The device according to claim 10, wherein: For a specific other sound source, the different physical speaker has a direction relative to the user that is closest to the direction of the specific other sound source relative to the user.

12. The device according to claim 8 or claim 9, wherein: The further modified rendering comprises rendering the first set of audio signals of the one or more other sound sources by rendering them at a reduced volume.

13. The device according to any one of claims 8 to 12, further comprising: means for determining a respective type of audio content comprising the first sound source and the one or more other sound sources; as well as Means for determining an amount of said other modified rendering to be performed based on the determined respective type of audio content.

14. The device according to any one of claims 8 to 12, further comprising: means for receiving metadata associated with audio content including the first sound source and the one or more other sound sources; as well as Means for determining an amount of said other modified rendering to perform based on the received metadata.

15. A device according to any one of the preceding claims, wherein: Said means for rendering of said at least first audio source comprises an MPEG-I renderer.

16. A method comprising: rendering at least a first sound source via audio signal output from two or more physical speakers having different respective positions such that the first sound source is intended to be perceived as having a first direction relative to a user that is different from a direction of the physical speakers; as well as detecting that the user's audio capture device is operating in a directional mode for steering a sound capture beam toward the first direction; In response to the detection, modified rendering is performed by outputting the audio signal of the first sound source from a physical speaker selected from the two or more physical speakers but not from other physical speakers, so that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capturing beam of the audio capturing device to be steered toward the selected physical speaker.

17. A computer program comprising a set of instructions which, when executed on a device, are configured to cause the device to perform a method comprising: rendering at least a first sound source by output of audio signals from two or more physical speakers having different respective positions such that the first sound source is intended to be perceived as having a first direction relative to a user that is different from a direction of the physical speakers; as well as detecting that the user's audio capture device is operating in a directional mode for steering a sound capture beam toward the first direction; In response to the detection, modified rendering is performed by outputting the audio signal of the first sound source from a physical speaker selected from the two or more physical speakers but not from other physical speakers, so that the first sound source will be perceived from the direction of the selected physical speaker, thereby causing the sound capturing beam of the audio capturing device to be steered toward the selected physical speaker.