Audio processing
By selecting the main sound direction that deviates from the center of the image area, the problems of poor audio focus selection and quality in the prior art are solved, and clearer sound source distinction and user experience improvement are achieved.
Patent Information
- Application Number
- CN202080039259.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2020-05-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-05-19
AI Technical Summary
When realizing audio focus, the prior art is limited by the processing capability of multi-channel audio signals and the limitations of beamforming technology, resulting in poor audio focus selectivity and audio quality. Especially in multi-sound source scenarios, it is difficult to clearly distinguish different sound sources.
By receiving indications of the multi-channel audio signal and the audio focus direction, the main sound direction is selected so that it corresponds to the second position of the image area, which deviates from the center point of the image area relative to the first position, so that the main sound direction is further away from the center of the focus area, thereby obtaining the output audio signal according to the main sound direction based on the multi-channel audio signal, emphasizing the sound in the main sound direction.
Improved the selectivity and audio quality of audio focus, can distinguish different sound sources more clearly, and improve user experience, especially in multi-sound source scenarios.
Smart Images

Figure CN113906769B_ABST
Abstract
Description
Technical Field
[0001] Examples and non - limiting embodiments of the present invention relate to the processing of multi - channel audio signals. In particular, various embodiments of the present invention relate to obtaining beamformed audio signals based on multi - channel audio signals. Background Art
[0002] For many years, mobile devices such as mobile phones and tablet computers have been equipped with camera and microphone arrangements that enable users of the devices to capture audio and video simultaneously. With the development of microphone technology and the increasing processing power and storage capacity available in mobile devices, it is becoming increasingly common to provide multi - microphone arrangements in such mobile devices that are capable of capturing multi - channel audio, which in turn can process the captured multi - channel audio into spatial audio to accompany the simultaneously captured video.
[0003] Generally, the process of using a mobile device to capture multi - channel audio signals includes: operating a microphone array arranged in the mobile device to capture a plurality of microphone signals; and processing the captured microphone signals into a recorded multi - channel audio signal for further processing in the mobile device, for storage in the mobile device together with the associated video and / or for transmission to one or more other devices. In a typical scenario, a user of a mobile device aims to record a multi - channel audio signal that represents an audio scene corresponding to the field of view (FOV) of a camera, thereby enabling a comprehensive presentation of the audiovisual scene at the time of capture.
[0004] When capturing or rendering an audiovisual scene, a user may wish to apply audio focusing to emphasize sounds in certain directions of the audio scene and / or to fade out sounds in certain other directions of the audio scene. Audio focusing schemes based on beamforming techniques known in the art enable, for example, amplifying sounds arriving from a selected direction that may also correspond to a respective sub - part of the FOV of the video, thereby providing audio in which sounds arriving from directions of the audio scene corresponding to the selected sub - part of the FOV that can depict an object of interest are emphasized.
[0005] However, in practical implementations, the number of available microphone signals in a mobile device, the corresponding positions of the microphones, and the limitations of available beamforming techniques impose limitations on the selectivity of audio focusing and / or the audio quality of the resulting audio signal. In particular, due to limitations in generating arbitrary spatially selective beam patterns, the microphone signals available at a mobile device typically only enable beamforming that results in relatively wide beams, where, relative to sounds from sound sources located in regions where the beam pattern has a smaller magnitude, a single beam pattern can amplify sounds from multiple sound sources located in regions where the beam pattern has a larger magnitude. This characteristic of beamforming or spatial filtering can be conceptualized as a focus region, where the focus region consists of directions in which the magnitude of the beam pattern is relatively high. In practice, the beam pattern can vary with frequency (and time, depending on the beamforming technique), and the beam pattern can have side lobes, and thus it can be understood that the term "focus region" is a conceptual term herein that illustrates the main capture region of focus processing. Known beamforming techniques generally do not allow for a clear boundary between sounds arriving within the focus region and sounds arriving from directions outside the focus region, and thus in practical scenarios, the attenuation of sounds residing outside the focus region gradually increases as the distance from the focus region increases. Thus, sounds from sound sources outside but relatively close to the focus region are generally not attenuated to a sufficient extent.
[0006] Accordingly, in practical implementations, in a scenario where the captured multi-channel audio signal represents two or more sound sources that are close to each other in corresponding spatial positions, even if a user sets or focuses the audio focus on a single sound source of interest, the audio focus generally emphasizes sounds from all of these sound sources. Additionally, in such a scenario, moving the center of the audio focus from one sound source to another by the user may have only a negligible effect (if any) on the resulting processed audio. These two aspects both limit the applicability of audio focus schemes and, in many cases, result in a degraded user experience. SUMMARY OF THE INVENTION
[0007] According to an example embodiment, there is provided a method for audio focus, the method comprising: receiving a multi-channel audio signal that represents sounds in sound directions corresponding to respective positions in an image region of an image; receiving an indication of an audio focus direction corresponding to a first position in the image region; selecting a main sound direction such that it corresponds to a second position in the image region, the second position deviating from the first position in a direction that is farther from the center point of the image region; and obtaining an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sounds in the sound direction defined by the selected main sound direction are emphasized relative to sounds in sound directions other than the sound direction defined by the selected main sound direction.
[0008] According to another exemplary embodiment, a method for audio focusing is provided, the method comprising: receiving a multi-channel audio signal representing sound in sound directions corresponding to respective positions in an image region of an image; receiving an indication of an audio focus direction corresponding to a first position in the image region; selecting a main sound direction from a plurality of different available candidate directions, wherein the plurality of different available candidate directions includes the audio focus direction and one or more offset candidate directions, and wherein each offset candidate direction corresponds to a respective candidate offset deviating from the first position in the image region; and obtaining an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sound in the sound direction defined by the selected main sound direction is emphasized relative to sound in sound directions other than the sound direction defined by the selected main sound direction.
[0009] According to another preferred embodiment, a device for audio focusing is provided, the device being configured to: receive a multi-channel audio signal representing sound in sound directions corresponding to respective positions in an image region of an image; receive an indication of an audio focus direction corresponding to a first position in the image region; select a main sound direction such that it corresponds to a second position in the image region, the second position deviating from the first position in a direction further away from the center point of the image region; and obtain an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sound in the sound direction defined by the selected main sound direction is emphasized relative to sound in sound directions other than the sound direction defined by the selected main sound direction.
[0010] According to another preferred embodiment, a device for audio focusing is provided, the device being configured to: receive a multi-channel audio signal representing sound in sound directions corresponding to respective positions in an image region of an image; receive an indication of an audio focus direction corresponding to a first position in the image region; select a main sound direction from a plurality of different available candidate directions, wherein the plurality of different available candidate directions includes the audio focus direction and one or more offset candidate directions, and wherein each offset candidate direction corresponds to a respective candidate offset deviating from the first position in the image region; and obtain an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sound in the sound direction defined by the selected main sound direction is emphasized relative to sound in sound directions other than the sound direction defined by the selected main sound direction.
[0011] According to another exemplary embodiment, there is provided an apparatus for audio focusing, the apparatus comprising: means for receiving a multi-channel audio signal representing sound in sound directions corresponding to respective positions in an image region of an image; means for receiving an indication of an audio focus direction corresponding to a first position in the image region; means for selecting a main sound direction such that it corresponds to a second position in the image region, the second position being offset from the first position in a direction further away from the center point of the image region; and means for obtaining an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sound in the sound direction defined by the selected main sound direction is emphasized relative to sound in sound directions other than the sound direction defined by the selected main sound direction.
[0012] According to another exemplary embodiment, there is provided an apparatus for audio focusing, the apparatus comprising: means for receiving a multi-channel audio signal representing sound in sound directions corresponding to respective positions in an image region of an image; means for receiving an indication of an audio focus direction corresponding to a first position in the image region; means for selecting a main sound direction from a plurality of different available candidate directions, wherein the plurality of different available candidate directions includes the audio focus direction and one or more offset candidate directions, and wherein each offset candidate direction corresponds to a respective candidate offset that is offset from the first position in the image region; and means for obtaining an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sound in the sound direction defined by the selected main sound direction is emphasized relative to sound in sound directions other than the sound direction defined by the selected main sound direction.
[0013] According to another exemplary embodiment, there is provided an apparatus for audio focusing, the apparatus comprising at least one processor and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to: receive a multi-channel audio signal representing sound in sound directions corresponding to respective positions in an image region of an image; receive an indication of an audio focus direction corresponding to a first position in the image region; select a main sound direction such that it corresponds to a second position in the image region, the second position being offset from the first position in a direction further away from the center point of the image region; and obtain an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sound in the sound direction defined by the selected main sound direction is emphasized relative to sound in sound directions other than the sound direction defined by the selected main sound direction.
[0014] According to another exemplary embodiment, there is provided an apparatus for audio focusing, the apparatus including at least one processor and at least one memory including computer program code, the computer program code causing the apparatus, when executed by the at least one processor, to: receive a multi-channel audio signal representing sounds in sound directions corresponding to respective positions in an image region of an image; receive an indication of an audio focus direction corresponding to a first position in the image region; select a main sound direction from a plurality of different available candidate directions, wherein the plurality of different available candidate directions includes the audio focus direction and one or more offset candidate directions, and wherein each offset candidate direction corresponds to a respective candidate offset deviating from the first position in the image region; and obtain an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sounds in the sound direction defined by the selected main sound direction are emphasized relative to sounds in sound directions other than the sound direction defined by the selected main sound direction.
[0015] According to another example embodiment, there is provided a computer program for audio focusing, the computer program including computer-readable program code configured to cause at least the method according to the example embodiment described above to be performed when the computer program is executed on a computing device.
[0016] According to an example embodiment, the computer program may be embodied on a volatile or non-volatile computer-readable recording medium, such as a computer program product including at least one computer-readable non-transitory medium having program code stored thereon, the program causing the device to perform at least the operations described above for the computer program according to the example embodiments of the present invention when executed by the device.
[0017] The exemplary embodiments of the present invention presented in this patent application should not be construed as limiting the applicability of the appended claims. The verb "comprising" and its derivatives are used in this patent application as open limitations, which do not exclude the existence of also unrecited features. Unless otherwise explicitly stated, the features described below may be combined with each other arbitrarily.
[0018] Some features of the present invention are set forth in the appended claims. However, the present invention will be best understood from the following description of some example embodiments, when read in conjunction with the accompanying drawings, with respect to its construction and method of operation, as well as its additional objects and advantages, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Embodiments of the present invention are illustrated in the drawings by way of example and not limitation, wherein
[0020] Figure 1A A block diagram showing some components and / or entities of a media capture device according to an example;
[0021] Figure 1B A block diagram showing some components and / or entities of a media rendering device according to an example;
[0022] Figure 2A A diagram showing an arrangement for implementing a media capture device and a media rendering device according to an example;
[0023] Figure 2B A diagram showing an arrangement for implementing a media capture device and a media rendering device according to an example;
[0024] Figure 2C A diagram showing an arrangement for implementing a media capture device and a media rendering device according to an example;
[0025] Figure 3A A block diagram showing some components and / or entities of a media capture device according to an example;
[0026] Figure 3B A block diagram showing some components and / or entities of a media rendering device according to an example;
[0027] Figure 4 Schematically shows the mapping of an audio focus area to two sound sources in an image area according to an example;
[0028] Figure 5 A flowchart showing a method according to an example;
[0029] Figure 6A Schematically shows the offset of a focus position in an image area according to an example;
[0030] Figure 6B Schematically shows the offset of a focus position in an image area according to an example;
[0031] Figure 6C Schematically shows the offset of a focus position in an image area according to an example;
[0032] Figure 7 Schematically shows the division of an image area into image parts and the shift of an audio focus area according to an example;
[0033] Figure 8 A flowchart depicting a method according to an example;
[0034] Figure 9 Schematically shows the mapping of an audio focus area to two sound sources in an image area corresponding to multiple candidate sound directions according to an example;
[0035] Figure 10 Schematically shows the mapping of a plurality of analysis regions according to an example to two sound source positions in an image region;
[0036] Figure 11 Shows a block diagram of some units of a device according to an example. Detailed implementation
[0037] Figure 1A Shows a block diagram of some components and / or entities of a media capture device 100 according to an example. The media capture device 100 includes a media capture entity 110, which includes an audio capture entity 111, a video capture entity 112, and a media processing entity 115. Figure 1B Shows a block diagram of some components and / or entities of a media rendering device 200 according to an example. The media rendering device 200 includes a media rendering entity 210, which includes an audio rendering entity 211, a video rendering entity 212, and a media processing entity 215.
[0038] The audio capture entity 111 is coupled to a microphone array 121, and it is configured to receive respective microphone signals from a plurality of microphones 121-1, 121-2, …, 121-K, and record the captured multi-channel audio signal based on the received microphone signals. The microphones 121-1, 121-2, ..., 121-K represent a plurality (i.e., two or more) microphones, where a single one of these microphones can be referred to as microphone 121-k. In this document, the concept of the microphone array 121 will be interpreted broadly to include any arrangement of two or more microphones 121-k provided in or coupled to the device implementing the media capture device 100. The video capture entity 112 is coupled to a camera entity 122, and it is configured to receive images from the camera entity 122 and record these images as a captured video stream. The camera entity 122 can include, for example, a digital video camera device or a digital video camera module. The media processing entity 115 can be configured to control at least some aspects of the operations of the audio capture entity 111 and the video capture entity 112.
[0039] Each microphone signal provides a different representation of the captured sound, the difference depending on the positions of microphones 121-k relative to each other. For a sound source at a certain spatial position relative to the microphone array 121, this results in different sound representations of the sound source in each microphone signal: A microphone 121-k closer to the certain sound source captures the sound from the certain sound source with a higher amplitude and earlier than a microphone 121-j farther away from the certain sound source. Together with the knowledge of the positions of microphones 121-k relative to each other, this difference in amplitude and / or time delay enables the use of the microphone signals as a basis for extracting or amplifying an audio signal representing the sound arriving from a desired direction relative to the microphone array 121 and / or converting the microphone signals into a spatial audio signal providing a spatial representation of the captured audio, wherein the sound of a sound source in the environment of the microphone array 121 at the time of capture is perceived as arriving in its corresponding direction relative to the microphone array 121. Audio processing techniques for extracting or amplifying an audio signal representing the sound arriving from a desired direction relative to the microphone array 121 and for converting the microphone signals into a spatial audio signal are well known in the art and are further described in detail herein only to the extent necessary for understanding certain aspects of the audio focus processing disclosed herein.
[0040] Accordingly, the microphone signals from the microphone array 121 are used as a multi-channel audio signal representing the sound captured within a sound direction range relative to the microphone array. Hereinafter, the sound direction range represented by the microphone signals or the spatial audio signal obtained therefrom is in most cases referred to as the spatial audio image captured at the position of the microphone array 121, while the audio signal obtained from the microphone signals and representing the sound arriving from a desired direction relative to the microphone array 121 can be regarded as representing the corresponding sound direction within the spatial audio image. Since the microphone array 121 and the camera entity 122 are operated at the same physical location, the multi-channel audio signal formed by or obtained from the microphone signals represents the sound in the sound directions corresponding to the positions in the image region of the image obtained from the camera entity 122. Based on the known characteristics of the image sensor of the camera entity 122 and its position and orientation relative to the microphone array 121, there can be at least an approximate predefined mapping between the spatial positions in the image region of the image obtained from the camera entity 122 and the corresponding sound directions within the spatial audio image represented by the microphone signals received from the microphone array 121, and thus, each position in the image region can be mapped to the corresponding sound direction in the spatial audio image represented by the microphone signals and vice versa. Accordingly, the correspondence between the sound direction and the position in the image region can be defined, for example, via a mapping function.
[0041] The media processing entity 115 may further be configured to provide the captured multi-channel audio signal and the captured video stream to the media rendering device 200. In this regard, the media capture device 100 may be implemented in the first device 101, while the media rendering device 200 may be implemented in the second device 201, as shown by the block diagram of Figure 2A . The providing may include sending the captured multi-channel audio signal and the captured video stream from the first device 101 to the second device 201 via a communication network, for example, as corresponding audio and video packet streams. In this example, the processing in the media processing entity 115 may include encoding the captured multi-channel signal and encoding the captured video stream for transmission to the second device 102 in the corresponding audio and video packet streams, while the processing in the media processing entity 215 may include, for example, decoding the reconstructed multi-channel audio signal based on the received audio packet stream and providing the reconstructed multi-channel audio signal for further audio processing in the audio rendering entity 211, and decoding the reconstructed video stream based on the received video packet stream and providing the reconstructed video stream for further video processing in the video rendering entity 212.
[0042] In other examples, the media capture device 100 and the media rendering device 200 may be implemented in the first device 101, as shown by the corresponding block diagrams of Figure 2B and Figure 2C . In the example of Figure 2B , providing the multi-channel audio signal and the captured video stream may include: the media capture device 100 storing the captured multi-channel audio signal and the captured video stream in the memory 102, and the media rendering device 200 reading the captured multi-channel audio signal and the captured video stream from the memory 102. In the example of Figure 2C , the media rendering device 200 directly receives the captured multi-channel audio signal and the captured video stream from the media capture device 100. In this example, the media capture device 100 and the media rendering device 200 may be implemented as a single logical entity, which may be referred to as the media processing device 103. In the examples of Figure 2B and Figure 2C , the corresponding encoding and decoding of the captured multi-channel audio signal and the captured video stream may not be necessary, and thus, the media processing entity 215 may provide the captured audio signal to the audio rendering entity 211, and provide the captured video stream directly ( Figure 2C ) or via the memory 102 ( Figure 2B ) to the video rendering entity 212.
[0043] The audio rendering entity 211 may be configured to apply audio focus processing to the multi-channel audio signal received there, so as to extract or emphasize the sound in the desired audio focus direction of the spatial audio image represented by the received multi-channel audio signal. In this regard, the audio focus processing may, for example, result in a single-channel audio signal (at least) representing the sound in the desired audio focus direction or a multi-channel audio signal having a focused audio component, wherein the sound in the desired audio focus direction is emphasized relative to the sound in other sound directions of the audio image. If the output includes a multi-channel audio signal having a focused audio component, the audio rendering entity 211 may further be configured to process the multi-channel audio signal having a focused audio component into a predefined or selected spatial audio format suitable for audio playback by the audio playback entity 221 (e.g., a speaker system or headphones). The video rendering entity 212 may process the video stream received there into a format suitable for video rendering by the video playback entity 222 (e.g., a display device).
[0044] If the processing in the media processing entities 115, 215 includes the respective steps of encoding and decoding the captured multi-channel audio signal into a reconstructed multi-channel audio signal and encoding and decoding the captured video stream into a reconstructed video stream, then in this regard, the media processing may be performed by using techniques known in the art, and thus, no further details in this regard are provided in the present disclosure. In addition, some aspects of the audio processing performed by the audio rendering entity 211 (such as processing the reconstructed audio stream into the desired spatial audio format) may likewise be performed by using techniques known in the art, and thus, no further details in this regard are provided in the present disclosure.
[0045] Figure 3A A block diagram showing some components and / or entities of a media capture device 100' according to an example is shown, while Figure 3B A block diagram showing some components and / or entities of a media rendering device 200' according to an example is shown. The media capture device 100' includes a media capture entity 110', which includes an audio capture entity 111', a video capture entity 112, and a media processing entity 115'. The media rendering device 200' includes a media rendering entity 210', which includes an audio rendering entity 211', a video rendering entity 212, and a media processing entity 215'. The system including the media capture device 100' and the media rendering device 200' is different from the system including the media capture device 100 and the media rendering device 200 in that audio focus processing for extracting or emphasizing the sound arriving from the desired audio focus direction described above with reference to the audio rendering entity 211 is applied in the audio capture entity 111', and no audio focusing occurs in the audio rendering entity 211.
[0046] The audio focus processing in the audio capture entity 111' may, for example, result in a single-channel audio signal representing (at least) sounds in a desired audio focus direction of the spatial audio image or a multi-channel audio signal having a focused audio component, wherein the sounds in the desired audio focus direction are emphasized relative to sounds located in other sound directions of the audio image. In the latter case, the media processing entity 115' may further process the multi-channel audio signal having the focused audio component into a predefined or selected spatial audio format that makes it readily suitable for audio playback by an audio playback entity (e.g., the audio playback entity 221). Regardless of the format of the audio signal resulting from the processing applied in the audio capture entity 111' and the media processing entity 115', the audio output from the media capture entity 110' is referred to as a captured audio signal, which may be presented in a manner similar to that described above with reference to Figure 2A , Figure 2B and Figure 2C The described approach for capturing a multi-channel audio signal is transferred (with appropriate modifications) from the media capturing entity 110' to the media rendering entity 210'.
[0047] Along the lines described above, audio focus processing (e.g. in the audio capturing entity 111' or in the audio rendering entity 211) is intended to emphasize sounds in a sound direction of interest relative to sounds in other sound directions, based on an audio focus indication provided as input to the audio capturing entity 111' or the respective one of the audio rendering entity 211. The audio focus indication defines at least an audio focus direction of interest within a spatial audio image represented by a multi-channel audio signal, and the audio focus indication may also define an audio focus amount indicating a desired intensity of emphasis to be applied to the sounds in the audio focus direction. In the following, the audio focus processing is described via a non-limiting example, which refers to the audio focus processing performed in the audio rendering entity 211, while it is easily generalized to audio focus processing performed in the audio capturing entity 111' (e.g. based on a multi-channel signal composed of or obtained from microphone signals) or by another entity.
[0048] As described above, the sound direction within the spatial audio image is associated with the position in the image area of the image accompanying the video stream, and conversely, the position in the image area is associated with the sound direction within the spatial audio image, for example via the mapping described above. Thus, the audio focus direction can be mapped to the corresponding position in the image area, and vice versa, the position in the image area showing the sound source of interest can be mapped to the audio focus direction within the spatial audio image.
[0049] In one example, the audio focus direction received at the audio capture entity 111' corresponds to a single (fixed or static) sound direction that remains the same or substantially the same over time (e.g., from one image to another in the images of a video stream), and it can be selected by the user or by another unit of the media capture entity 110'. In another example, the audio focus direction received at the audio capture entity 111' corresponds to a sound direction that varies over time (e.g., from one image to another in the images of a video stream), and it can be obtained by another unit of the media capture entity 110' by, for example, tracking the image region location of an object of interest (e.g., an object selected by the user) over time. Similar considerations (with appropriate modifications) also apply to the reception of the audio focus direction in the audio rendering entity 211. The audio focus processing described in the present disclosure can be performed at capture time (e.g., in the audio capture entity 111'), or as a post-processing stage after capture time (e.g., in the audio capture entity 111' or in the audio rendering entity 121).
[0050] The audio focus processing can include: applying a predefined beamforming technique to a multi-channel audio signal received at the audio rendering entity 211 to extract a beamformed (single-channel or multi-channel) audio signal that represents sound in the desired audio focus direction of the spatial audio image represented by the multi-channel audio signal. In some examples, the beamformed audio signal can further be used as a basis for: creating a focused (multi-channel) audio component, where the beamformed audio signal is repositioned to its original spatial position in the spatial audio image; and combining the focused audio component with the multi-channel audio signal, in view of the desired amount of audio focus (or, if no desired amount of audio focus is specified, in view of a predefined amount of audio focus), to create a multi-channel audio signal with the focused audio component. In this regard, combining the focused audio component with the multi-channel audio signal can include: amplifying (e.g., multiplying) the focused audio component by a first scaling factor representing the desired or predefined amount of audio focus; or, attenuating (e.g., multiplying) the multi-channel audio signal by a second scaling factor representing the desired or predefined amount of audio focus. In a further example, combining the focused audio component with the multi-channel audio signal can include: amplifying (e.g., multiplying) the focused audio component by a first scaling factor and attenuating (e.g., multiplying) the multi-channel audio signal by a second scaling factor, where the first scaling factor and the second scaling factor together represent the desired or predefined amount of audio focus.
[0051] The beamforming technique applied by the audio rendering entity 211 when creating a beamformed audio signal may include using a suitable beamformer known in the art. Due to the limited spatial selectivity of beamforming techniques known in the art, in practical implementations, the beamformed audio signal represents not only the sound strictly located in the desired audio focus direction in the spatial audio image, but also the sound within the audio focus area around the desired audio focus direction in the spatial audio image, thereby representing the sound in the desired audio focus direction together with the sound in the sound directions related to the beamforming technique around the desired audio focus direction. Generally, in addition to the sidelobes and fluctuations in the beam pattern, the attenuation (or suppression) of sound sources in the sound directions around the desired audio focus direction usually increases with the increase in the distance from the desired audio focus direction, where the degree of attenuation depends on the applied beamforming technique and / or the positioning of the microphones 121-k (relative to each other and relative to the desired audio focus direction) when capturing potential multi-channel audio signals. In this regard, the audio focus area can be regarded as containing those sound directions where the sound is substantially not attenuated, while the sound in the sound directions outside the audio focus area is substantially attenuated.
[0052] Beamformers known in the art can be classified as dynamic beamformers and static beamformers. An example of a dynamic beamformer is the minimum variance distortionless response (MVDR) beamformer, while an example of a static beamformer is the phase shift (PS) beamformer. Generally, compared with static beamformers such as PS, dynamic beamformers such as MVDR can achieve a smaller audio focus area and can especially better suppress discrete sound sources in the sound directions outside the audio focus area. However, due to the increased possibility of audio distortion in the beamformed audio signal, this advantage of dynamic beamformers usually comes at the cost of reducing the quality of the beamformed audio signal compared to the quality obtained via using static beamformers. The computational complexity of dynamic beamformers is also generally higher than that of static beamformers. The trade-off between the size of the resulting audio focus area and / or the degree or probability of distortion in the resulting beamformed audio signal can be further adjusted to some extent by selecting the parameters of the applied beamformer (e.g., the white noise gain of the beamformer). Therefore, the spatial part of the spatial audio image represented by the multi-channel audio signal covered by a certain audio focus area is at least partially defined by the main sound direction of the audio focus area, the characteristics of the applied beamformer, and possibly also by the parameters of the applied beamformer.
[0053] Along the lines discussed above, the actual shape and size of the audio focus region set in view of a given desired audio focus direction can depend, for example, on the beamforming technique applied, the relative positions of the microphones 121-k in the microphone array 121 that are applied to capture the potential multi-channel audio signals and / or the position of the desired audio focus direction within the spatial audio image. Additionally, the shape and size of the audio focus region can be different at different frequencies (e.g., in different frequency sub-bands). Thus, while some of the figures of the present disclosure show the audio focus region as circular for graphical clarity in illustration, in actual implementation the audio focus region can have a somewhat arbitrary shape having an "envelope" that is similar to but not strictly circular (or elliptical).
[0054] Hereinafter, the position of the audio focus region relative to the spatial audio image is described via the main sound direction of the audio focus region, such that setting or selecting a certain sound direction of the spatial audio image as the main direction results in an audio focus region positioned around the main direction. In other words, the main amplification direction of the beam pattern is around the main sound direction. Thus, beamforming based on the main sound direction results in a beamformed audio signal, wherein the sound in the sound direction defined via the main sound direction is emphasized relative to the sound in the sound directions other than those defined via the main sound direction. The main sound direction can be regarded as the conceptual center point of the audio focus region, even though it may not be the geometric center point of the audio focus region due to the somewhat arbitrary shape of the audio focus region and the differences in size and shape across frequencies. However, conceptually, the main sound direction can be regarded as representing the center point of the audio focus region. In an example, the main sound direction of the audio focus region includes the sound direction in which the sound is amplified to the maximum extent compared to other directions. In some examples, the main sound direction of the audio focus region includes the sound direction in which the sound is amplified to the maximum extent relative to other directions within the image region, i.e., there may be stronger amplification in some sound directions mapped outside the image region, but these are not considered. Nevertheless, in the context of the present disclosure, the relative position of the audio focus region obtained from selecting the main sound direction within the spatial audio image plays a more important role than its absolute position, and thus, the concept of "main sound direction" serves as a sufficient position reference for the purposes of the present disclosure.
[0055] In the following description, the expression that the main sound direction of a proposed audio focus area is arranged / set / located at a certain position in an image area may be applied. Obviously, the meaning of such an expression itself is limited. However, for the purpose of improving the readability of the present disclosure, such a concise expression is applied to convey the following meaning: the main sound direction is arranged / set / located in the spatial audio image in the sound direction that is mapped to a certain position in the image area. Similarly, the following text may adopt the expression that a proposed audio focus area overlaps / covers a certain spatial position or part of an image area as a concise version of a complete expression, the meaning of which is that the audio focus area includes one or more sound directions in the spatial audio image that are mapped to a certain spatial position or part of the image area.
[0056] In a scenario where multi-channel audio is accompanied by a video stream, the user is typically mainly interested in the sound arriving within an image area of an image that constitutes an associated video stream (e.g., defined by the FOV of camera entity 122), while the sound arriving outside the image area can be ignored without significantly affecting the perceived quality of the resulting audiovisual representation of the scenario. On the other hand, the spatial audio image represented by the multi-channel audio signal can extend to also cover sound directions outside the image area. In this regard, the audio rendering entity 211 may be set to suppress or attenuate the sound in the sound directions in the spatial audio image that originate from sound sources outside the image area of the video stream's image. As described previously, due to the limited spatial selectivity of beamforming techniques known in the art, in a practical implementation, by significantly attenuating (or even suppressing) the sound in the sound directions outside the audio focus area while not significantly attenuating the sound in the sound directions within the audio focus area, the beamformed audio signal necessarily represents the sound within the audio focus area around the desired audio focus direction (rather than only strictly representing the sound of the desired audio focus direction). Therefore, the beamformed audio signal represents not only the sound of an object shown at a desired point in the image area of the video stream but also the sound of objects within a part of the image area around that desired point.
[0057] During the operation of the audio rendering entity 211, a user may select a direction of an audio focus of interest via a user interface (UI) of the devices 101, 201 that implement the audio rendering entity 211. As an example in this regard, the audio rendering entity 211 may receive, via the UI, a selection of a position of an image region depicting a desired audio focus direction and map the selected position of the image region to a corresponding sound direction in the spatial audio image. In another example, the audio rendering entity 211 may receive, via the UI, a selection of an object depicted in an image region, apply appropriate image analysis techniques to identify a position of the object in the image region in an image of a video stream, and, in each considered image, map the identified position of the object in the image region to a corresponding sound direction in the spatial audio image. In previously known solutions, beamforming is performed using this sound direction as the main sound direction, which results in an audio focus area in the sound direction in the spatial audio image to which the selected position of the image region is mapped. In this document, even if a user does not directly select an audio focus direction, the sound direction of the spatial audio image selected in response to the position of the image region selected by the user or in response to the tracked image region position of the object selected by the user shown in the image may be referred to as the user-selected (or received) audio focus direction. Accordingly, the audio rendering entity 211 performs beamforming based on the user-selected audio focus direction, which results in an audio focus area that includes sound directions around the user-selected audio focus direction and, thus, results in audio focusing that emphasizes the sound in all sound directions within the obtained audio focus area in the spatial audio image relative to the sound in the sound directions outside the audio focus area obtained from the beamforming performed based on the user-selected audio focus direction in the spatial audio image.
[0058] Along the lines discussed above, the above method is used to provide audio focusing of the sound in the user-selected audio focus direction, while it may inadvertently also provide audio focusing of the sound in the sound directions around the desired sound direction. Figure 4An example in this regard is schematically shown, where a first object is depicted at position A in the image region 312, and a second object is depicted at position B in the image region 312, where the first and second objects represent respective sound sources within a spatial audio image. Assuming that the user wants to set the audio focus to the sound originating from the first object, the resulting audio focus region 311 covers a part of the image region around position A. However, due to the limitations of the spatial selectivity of the applied beamforming technique, the audio focus region 311 also includes the second object at position B in the image region. Therefore, instead of emphasizing the sound originating from the first object depicted at position A relative to the sound originating from the second object depicted at position B, beamforming is performed using the audio focus region 311, resulting in the emphasis of both the sound originating from the first object and the sound originating from the second object, which in many cases leads to a degraded user experience in terms of audio focus processing.
[0059] For example, via operations of method 400 as shown in the flowchart depicted by Figure 5 an improved audio focus can be obtained. Method 400 can be performed, for example, by the audio capture entity 111' or the audio rendering entity 211. The operations described with reference to blocks 402 to 408 of method 400 can be varied or supplemented in various ways without departing from the scope of audio focus processing according to the present disclosure (e.g., according to the examples described hereinbefore and hereinafter).
[0060] As indicated in block 402, method 400 begins with receiving a multi-channel audio signal that represents sounds in sound directions corresponding to respective positions in an image region of an image. Herein, the image includes a video stream received at the media processing entity 215 or an image obtained therefrom. As indicated in block 404, method 400 further includes receiving an indication of an audio focus direction corresponding to a first position in the image region.
[0061] Method 400 further includes: as indicated in block 406, selecting a main sound direction corresponding to a second position in the image region, the second position being offset from the first position in the image region in a direction that is farther from the center point of the image region; and as indicated in block 408, obtaining an output audio signal based on the multi-channel audio signal and according to the main sound direction, wherein sounds in a sound direction other than the sound direction defined by the selected main sound direction are de-emphasized relative to sounds in the sound direction defined by the selected main sound direction.
[0062] In the following, non-limiting examples of operations related to block 406 are provided. In this regard, the selection of the primary sound direction for obtaining the output audio signal and the arrangement of the resulting audio focus regions relative to their (mapped) positions in the image region are described in more detail. In an example with respect to method 400, the primary sound direction is selected such that (in addition to the primary sound direction) the received audio focus directions (also) are included in the audio focus region around the primary sound direction. In the following description, the term "received focus position" is applied to refer to the position in the image region to which the received audio focus direction is mapped (i.e., the "first position" mentioned above in the context of blocks 404 and 406), while the term "shifted focus position" is applied to refer to the position to which the selected primary sound direction is mapped (i.e., the "second position" mentioned above in the context of block 406). Thus, the shifted focus position is arranged according to the image region position such that the distance between the shifted focus position and the center point of the image region is greater than the distance between the received focus position and the center point of the image region, thereby shifting the provided audio focus to include the sound directions mapped to the image region positions that are farther from the center of the image region than the sound directions mapped to the received focus position.
[0063] Furthermore, the mention below of shifting or offsetting the received focus position to the shifted focus position implies adjusting the received audio focus directions within the spatial audio image mapped to the received focus position to the selected primary sound direction within the spatial audio image mapped to the shifted focus position. Thus, shifting or offsetting the received focus position in the image region to the shifted focus position in the image region is essentially the result of shifting or offsetting the audio focus directions to the primary sound direction within the spatial audio image, but for the sake of brevity and clarity of description, the following examples mostly refer to shifting the audio focus directions within the spatial audio image such as a movement or deviation occurring in the image plane.
[0064] According to a first example, the shifted focus position is offset from the received focus position in one or both of the horizontal direction of the image plane and the vertical direction of the image plane such that the point in the image region to which the primary sound direction is mapped is farther from the center point of the image region. The terms "horizontal direction" and "vertical direction" are used herein in a non-limiting manner to include any pair of a first direction and a second direction that are perpendicular to each other. The degree of offset is selected such that the audio focus region (also) resulting from the use of the applied beamformer includes the received audio focus directions.
[0065] From Figure 4 the example continues and further assumes that the received audio focus direction is mapped to position A in the image region (where a first object is shown here), Figure 6AAn example in this regard is schematically shown, in which the shifted focus position is deviated from the received focus position in the vertical direction of the image plane (represented by the y-axis in the illustration of Figure 6A ). In Figure 6A , the solid circle represents the offset audio focus area 311' obtained by shifting the audio focus direction from the user-selected audio focus direction, while the dashed circle represents the audio focus area 311 according to the example of Figure 4 .
[0066] In the example of Figure 6A , the main sound direction is selected such that it causes the shifted focus position to be mapped to the position of the image area indicated by the cross mark in the illustration of Figure 6A , thereby providing a shifted focus position whose distance to the center point of the image area (indicated by C in the illustration of Figure 6A ) is greater than the distance from position A to the center point of the image area. Thus, with sufficient deviation, the obtained offset audio focus area 311' is shifted such that the sound from the direction mapped to position B (showing a second object) in the image area is not included in the offset audio focus area 311', while the audio focus area 311 includes the sound in the received audio focus direction mapped to position A in the image area. Therefore, beamforming using the offset audio focus area 311' enables obtaining a beamformed audio signal in which the sound in the sound direction mapped to position A in the image is also emphasized relative to the sound in the sound direction mapped to position B in the image area.
[0067] Figure 6B An example is schematically shown, in which the shifted focus position is deviated from the received focus position in the horizontal direction of the image plane (indicated by the x-axis in the illustration of Figure 6B ), such that it is farther from the center point of the image area (at position C). Similarly, the solid circle represents the offset audio focus area 311' obtained by deviating the audio focus direction from the received audio focus direction, while the dashed circle represents the audio focus area 311 according to the example of Figure 4 . As shown in Figure 6B , with sufficient deviation, the obtained offset audio focus area 311' is shifted such that the sound from the direction mapped to position B (showing a second object) in the image area is not included in the offset audio focus area 311', while the audio focus area 311 includes the sound in the received audio focus direction mapped to position A in the image area.
[0068] Figure 6C Another example is schematically shown, in which the shifted focus position is in both the horizontal and vertical directions of the image plane (in the illustration of Figure 6Cis offset from the received focus position on the x-axis and y-axis respectively (indicated in the illustration). In this example, the focus position is shifted along the (conceptual) line that intersects both the center point of the image region (at position C) and the received focus position (at position A) such that the focus position is further from the center of the image region. Again, the solid circle represents the offset audio focus area 311' obtained from deviating the audio focus direction from the received audio focus direction, and the dashed circle represents the audio focus area 311 according to the Figure 4 example. As Figure 6C shown, with sufficient deviation, the resulting offset audio focus area 311' is shifted such that sound from the direction mapped to position B (showing a second object here) in the image region is not included in the offset audio focus area 311', while the audio focus area 311 includes sound in the received audio focus direction mapped to position A in the image region.
[0069] Both the degree and direction of deviation can be predefined. Under the conditions described above, the deviation direction results in the shifted focus being further from the center point of the image region compared to the received focus. Even if the predefined degree and direction of deviation do not guarantee providing an offset audio focus area 311' that does not include sound from important sound sources in the direction of the sound source of interest relatively close to and mapped to positions residing within the image region, nevertheless, it still increases the likelihood of excluding such sound sources from the beamformed audio signal, thereby enabling improved audio focusing.
[0070] In one example, a predefined degree of deviation can be applied that is independent of the position of the received focus position in the image region. In other words, the same predefined degree of deviation can be applied to all received focus positions. In another example, the degree of deviation depends on the position of the received focus position in the image region such that the degree of deviation increases as the distance between the received focus position and the center point of the image region increases. In another example, the image region can be (at least conceptually) divided into multiple non-overlapping image parts, and the corresponding predefined degree of deviation is applied according to the image part in which the received focus position is located. As an example in this regard, the degree of deviation in the image part further from the center point of the image region can be greater than that in the image part closer to the center point of the image region.
[0071] As a further example of the degree of deviation, the deviation can be applied only to those received focus positions that are further than a (first) predefined distance from the center point of the image region (in other words, for received focus positions that are within the (first) predefined distance from the center point of the image region, the degree of deviation can be "zero"), the degree of deviation can be limited such that it remains within the image region, and / or the degree of deviation can be limited such that it does not extend beyond the image region by more than a predefined threshold distance.
[0072] In one example, a predefined deviation direction can be applied that is independent of the position of the received focus position within the image region. In other words, the same predefined deviation direction can be applied to all received focus positions. In another example, the deviation direction can be selected based on the position of the received focus position within the image region such that the image region can be (at least conceptually) divided into a plurality of non-overlapping image portions, and the predefined deviation direction is applied based on the image portion in which the received focus position is located. As an example in this regard, in an image portion bounded by a single edge of the image region (e.g., an image portion adjacent to one of the top, bottom, left, and right edges of the image region), the deviation direction can be in the vertical or horizontal direction in the image plane such that the shifted focus is closer to one side of the image portion bounded by the edge of the image region rather than closer to the opposite side of the image portion bounded by another image portion, in an image portion bounded by two non-opposite edges of the image region (e.g., an image portion at a corner of the image region), for example, a deviation direction can be provided in both the horizontal and vertical directions along a (conceptual) line that intersects both the center point of the image region and the received focus, and / or, in an image portion not bounded by any edge of the image region (e.g., an image portion all of whose sides are bounded by adjacent image portions), the deviation direction can be in one or both of the horizontal and vertical directions in the image plane, or alternatively, no deviation can be applied in such an image portion.
[0073] The sound direction included by the offset audio focus area 311' further depends on the choice of beamforming technique applied to create the beamformed audio signal according to the shifted focus position. As an example in this regard, a predefined beamformer can be applied when obtaining the beamformed audio signal. In another example, the operations associated with block 406 can further include selecting the beamformer or the type of beamformer to be applied when obtaining the beamformed audio signal. In one example, the same beamformer and / or the same or similar type of beamformer can be applied regardless of the position of the reception focus position in the image area, where the applied beamformer can be a static beamformer such as PS or a dynamic beamformer such as MVDR. In another example, the beamformer or the type of beamformer to be applied can be selected according to the position of the reception focus position in the image area, for example such that a dynamic beamformer is applied to reception focus positions closer to the center point of the image area than a (second) predefined distance, while a static beamformer is applied to reception focus positions farther from the center point of the image area than the (second) predefined distance. In another example, the beamformer or the type of beamformer to be applied can be selected according to the reception focus position in the image area such that the image area can be (at least conceptually) divided into a plurality of non-overlapping image parts, and the beamformer or the type of beamformer assigned to the image part where the reception focus position is located is applied. As an example in this regard, a dynamic beamformer can be assigned to an image part bounded by a single edge of the image area (e.g., an image part adjacent to one of the top edge, bottom edge, left edge, and right edge of the image area) and to an image part not bounded by any edge of the image area (e.g., an image part all sides of which are bounded by adjacent image parts), and / or, a static beamformer can be assigned to an image part area bounded by two adjacent edges of the image (e.g., an image part at the corner of the image area).
[0074] Depending on the details of the selected method, the selection of the beamformer or the type of beamformer according to the position of the reception focus position as described above results in: using a dynamic beamformer near the center of the image area (usually enabling a smaller-sized audio focus area with an increased risk of audio distortion) and using a static beamformer closer to the edges and / or corners of the image area (usually resulting in a larger-sized audio focus area with a reduced risk of audio distortion), thereby (further) reducing the possibility of providing the offset audio focus area 311' such that it does not include the sound from an important sound source in the sound direction relatively close to the sound source of interest and mapped to a position residing within the image area.
[0075] In Figure 7is schematically shown a non-limiting example of partitioning an image region into a set of non-overlapping rectangular image parts, while in other examples, image parts of some other shape (e.g., hexagon) may alternatively be applied. In Figure 7 , the image region 312 is partitioned into eight image parts labeled 312-1 to 312-8, and each image part is shown to have a corresponding exemplary shifted audio focus region 311-1' to 311-8'. It should be noted that Figure 7 the illustration of does not depict the absolute position of the shifted audio focus region 311-j' relative to the corresponding image part 312-j, but is used to indicate the corresponding direction to which the received focus position is shifted relative to the center point of the image region 312 to define the corresponding shifted focus position (see the arrows extending outward from the circles representing the audio focus regions 311-j'). In addition, the corresponding size of the audio focus region 311-j' is used to indicate the type of beamformer assigned to the corresponding image part 312-j: a larger circle represents a static beamformer (such as PS), and a smaller circle represents a dynamic beamformer (such as MVDR). Thus, in Figure 7 the example of, it can be assumed that dynamic beamformers are assigned to the image parts 312-2, 312-3, 312-6, 312-7, while the deviation direction is in the vertical direction of the image plane towards the closer one of the top and bottom edges of the image region 312, and static beamformers are assigned to the image parts 312-1, 312-4, 312-5, 312-8, while the deviation direction is in both the horizontal and vertical directions of the image plane in the general direction towards the corresponding corners of the image region.
[0076] Now referring to the operations related to block 408, obtaining the output audio signal may include, for example: extracting a beamformed audio signal from the received multi-channel audio signal using a predefined or selected beamformer, the beamformed audio signal representing the sound in the sound direction within the audio focus region 311' around the selected main sound direction of the spatial audio image, wherein the beamformed audio signal may include a single-channel audio signal or a multi-channel audio signal. As described above, the obtained offset audio focus region 311' also contains the sound in the received audio focus direction, whereby the beamformed audio signal serves as an audio signal that emphasizes the sound in the received audio focus direction relative to the sound in the sound directions outside the audio focus region 311'.
[0077] In one example, the beamformed audio signal is provided as the output audio signal. In another example, the operations associated with block 408 may further include or be followed by: constructing a multi-channel output audio signal having a focused audio component based on the received multi-channel audio signal and the beamformed audio signal, wherein sounds in the sound directions within the audio focus region 311' around the selected primary sound direction of the spatial audio image are emphasized relative to sounds in the sound directions outside the audio focus region 311'. Generally, only the sound directions mapped to positions within the image region are considered, while amplification and / or attenuation of sounds in the sound directions mapped to positions outside the image region are ignored.
[0078] Obtaining such a multi-channel output audio signal may include: obtaining a focused (multi-channel) audio component, wherein the beamformed audio signal is repositioned at its original spatial position in the spatial audio image; and combining the focused audio component with the received multi-channel audio signal to create a multi-channel output audio signal having a focused audio component in view of the desired amount of audio focus (or in view of a predefined amount of audio focus if no desired amount of audio focus is specified). As an example in this regard, combining the focused audio component with the multi-channel audio signal may include: amplifying (e.g., multiplying) the focused audio component by a first scaling factor representing the desired or predefined amount of audio focus; or, attenuating (e.g., multiplying) the received multi-channel audio signal by a second scaling factor representing the desired or predefined amount of audio focus. In a further example, combining the focused audio component with the multi-channel audio signal may include: amplifying (e.g., multiplying) the focused audio component by a first scaling factor and attenuating (e.g., multiplying) the multi-channel audio signal by a second scaling factor, wherein the first scaling factor and the second scaling factor together represent the desired or predefined amount of audio focus. According to a predefined channel configuration (such as 5.1 channel surround sound or 7.1 channel surround sound), the multi-channel output audio signal may be provided as or (further) processed into, for example, a two-channel binaural audio signal or a multi-channel surround signal.
[0079] Still referring to the first example, the degree of deviation, the direction of deviation, and / or the applied beamformer or beamformer type may be differently selected or defined in different frequency sub-bands. In an example, the degree of deviation, the direction of deviation, and / or the applied beamformer or beamformer type may be selected or defined as described above for one or more first frequency sub-bands, while for one or more second frequency sub-bands, no deviation (or a smaller deviation) may be applied and / or a predefined beamformer or beamformer type may be applied.
[0080] According to a second example, assume that two or more respective microphones 121-k of the microphone array 121 are located on both sides of the image sensor of the camera entity 122, which typically results in the fact that even when the same or a similar beamformer or beamforming type is applied to each of the audio focus regions 311, 311', the size of the audio focus regions 311, 311' is smaller when (more) close to the center of the image region compared to the size of its side (more) close to the image region (e.g., those edges closer to the image region corresponding to the respective edges of the image sensor adjacent to the two or more microphones 121-k). In this regard, in the second example, the beamformer can be a predefined beamformer, e.g., a static beamformer such as PS or a dynamic beamformer such as MVDR. Thus, in the context of the second example, the selection of the main sound direction (see block 406) and the obtaining of the output audio signal (see block 408) can be performed in the manner described above for the first example, except that the beamformer or beamformer type may be selected according to the position of the reception focus position in the image region (the position to which the received audio focus direction is mapped).
[0081] Still referring to the second example, the degree of deviation and / or the direction of deviation can be selected or defined differently in different frequency sub-bands. In the example, the degree of deviation and / or the direction of deviation can be selected or defined as described above for one or more first frequency sub-bands (e.g., for frequency sub-bands below a predefined frequency threshold), while no deviation (or a smaller deviation) may be applied for one or more second frequency sub-bands (e.g., for frequency sub-bands above a predefined frequency threshold).
[0082] According to a third example, in a manner slightly different from method 400 and / or the examples regarding Figure 6A , Figure 6B , Figure 6C and Figure 7 , the problems of the previously known audio focusing methods discussed with reference to Figure 4 are solved. In this regard, an improved audio focus can be provided, for example, according to method 500 shown in the flowchart depicted in Figure 8 . The operations described with reference to blocks 502 to 508 of method 500 can be varied or supplemented in various ways without departing from the scope of audio focus processing according to the present disclosure (e.g., according to the examples described above and below).
[0083] As indicated in block 502, method 500 begins with receiving a multi-channel audio signal that represents sounds in sound directions corresponding to respective positions in an image region of an image. As indicated in block 504, method 500 also includes receiving an indication of an audio focus direction corresponding to a first position in the image region. Here, the operations associated with blocks 502 and 504 are respectively similar to the operations described with reference to blocks 402 and 404 in the context of method 400.
[0084] As indicated in block 506, method 500 also includes selecting a main sound direction from a plurality of different available candidate directions, where each candidate direction corresponds to a respective candidate offset that is offset from the first position. In this regard, the offset can be in any direction on the image plane. As indicated in block 508, method 500 also includes obtaining an output audio signal based on the multi-channel audio signal and according to the main sound direction, where sounds in the sound direction defined by the selected main sound direction are emphasized relative to sounds in sound directions other than those defined by the main sound direction. In an example regarding method 500, the main sound direction is selected such that the received audio focus direction (also) is included in an audio focus region around the main sound direction. Non-limiting examples of the operations associated with blocks 506 and 508 are described below.
[0085] Now referring to the operations associated with block 506 of method 500, as described above, the main sound direction can be selected from a plurality of different available candidate sound directions (i.e., two or more different available candidate sound directions), which include the received audio focus direction and one or more offset candidate directions, and each offset candidate direction can be described, for example, via a corresponding candidate offset relative to the position of the image region mapped to by the received audio focus direction. In this regard, each candidate offset can define a corresponding pair of deviation directions and degrees of deviation in the image plane, in other words, the direction and distance of the corresponding candidate shifted focus position relative to the received focus position, and the deviation direction can be in any direction in the image plane. The same or similar beamformer is applicable to obtain a corresponding candidate beamformed audio signal using each candidate sound direction, so as to enable obtaining a corresponding candidate beamformed audio signal based on the corresponding candidate audio focus region around the corresponding candidate sound direction. Due to using the same or similar beamformer, each candidate audio focus region has substantially the same size in terms of the sound directions included in the corresponding candidate audio focus region. The degree of deviation is selected for each offset candidate sound direction such that, given the characteristics of the applied beamformer, the corresponding candidate audio focus region includes the received audio focus direction. Since each candidate audio focus region includes the received audio focus direction, they necessarily overlap partially with each other. On the other hand, each candidate audio focus region also includes a range of directions around an audio focus direction different from that included in other candidate audio focus regions.
[0086] As a non-limiting example in this regard, Figure 9 schematically shows the following scenario: corresponding candidate audio focus regions 311, 311a, 311b, and 311c obtained from a scenario in which three different offset candidate sound directions are available in addition to the received audio focus region: the first offset candidate audio focus region 311a is obtained by shifting the received focus position in the direction of the vertical axis of the image plane (towards the upper edge of the image region) according to the first candidate offset, the second offset candidate audio focus region 311b is obtained by shifting the received focus position in the direction of the horizontal axis of the image plane (towards the right edge of the image region) according to the second candidate offset, and the third offset candidate audio focus region 311c is obtained by shifting the received focus position in the direction of the vertical axis of the image plane (towards the lower edge of the image region) according to the third candidate offset. In Figure 9 the example, sounds from both the direction mapped to position A (showing the first object) in the image region and the direction mapped to position B (showing the second object) in the image region are included in the audio focus regions 311, 311b, and 311c, while the audio focus region 311a only includes the direction mapped to position A and does not include the direction of position B.
[0087] In a third example, selecting the primary sound direction (see block 506) can include: for each of a plurality of different available candidate directions, estimating the energy of the corresponding candidate beamformed audio signal that can be obtained via use of the applied beamformer; and selecting one of the candidate sound directions as the primary sound direction based on the corresponding energies of the candidate beamformed audio signals. In one example, the energy of a candidate beamformed audio signal resulting from beamforming according to a certain candidate direction can be obtained by performing the following operations: performing beamforming using the applied beamformer to obtain the corresponding candidate beamformed audio signal, and calculating the energy of the corresponding candidate beamformed audio signal. In another example, the energy of a candidate beamformed audio signal resulting from beamforming according to a certain candidate direction via use of the applied beamformer can be obtained via use of a directional energy estimation method associated with the applied beamformer, thereby avoiding the calculations required for the actual obtaining of the candidate beamformed audio signal. Such directional energy estimation methods are known in the art.
[0088] As a specific example in this regard, selecting one of the candidate sound directions as the primary sound direction can include: selecting the candidate sound direction that results in the candidate beamformed audio signal having the lowest energy as the primary sound direction. In another example, the energy-based selection of the primary sound direction can be performed separately for a plurality of frequency subbands. Thus, different ones of the candidate sound directions can be selected as the primary sound direction in different frequency subbands. In an example, the same energy-based criterion for selecting one of the candidate sound directions as the primary sound direction can be applied to these frequency subbands. In another example, the energy-based criterion for selecting one of the candidate sound directions as the primary sound direction can vary from one frequency subband to another. As an example of the latter, in a frequency subband below a predefined frequency threshold, the candidate sound direction that provides the candidate beamformed audio signal having the lowest energy can be selected as the primary sound direction, while in a frequency subband above the predefined frequency threshold, the candidate sound direction that provides the candidate beamformed audio signal having the highest energy can be selected as the primary sound direction.
[0089] Now referring to block 508, in an example, based on the selected main sound direction via the operations of block 506 described above, an output audio signal can be obtained from the received multi-channel audio signal by applying a predefined beamformer to extract a beamformed audio signal from the received multi-channel audio signal, where the beamformed audio signal represents the sound in the main sound direction of the spatial audio image represented by the received multi-channel audio signal. In another example, if the energy estimation described above involves obtaining a candidate beamformed audio signal, the candidate beamformed audio signal obtained by beamforming based on the candidate sound direction (via the operations of block 506) selected as the main sound direction can be applied as the beamformed audio signal.
[0090] Along the lines described above in the context of the examples related to method 400, in an example, the beamformed audio signal can be provided as the output audio signal. In another example, the operations related to block 508 can also include or be followed by: constructing a multi-channel output audio signal with focused audio components based on the received multi-channel audio signal and the beamformed audio signal, where, relative to the sound in the sound directions outside the audio focus region 311', the sound in the sound directions within the audio focus region 311' around the selected main sound direction of the spatial audio image is emphasized. The obtaining of such a multi-channel output audio signal can be performed as described above. According to a predefined channel configuration (such as 5.1-channel surround sound or 7.1-channel surround sound), the multi-channel output audio signal can be provided as or (further) processed into, for example, a two-channel binaural audio signal or a multi-channel surround signal.
[0091] According to a fourth example provided within the framework of method 500, the selection of the main sound direction (see block 506) includes: performing an analysis process to attempt to identify the respective sound directions of one or more (directional) sound sources included in the spatial audio image represented by the received multi-channel audio signal; and selecting the main sound direction at least partially based on the identified sound directions.
[0092] The analysis process includes applying a set of analysis zones, the respective main sound directions of which are set such that these analysis zones together cover or substantially cover the sound directions of the spatial audio image corresponding to the entire image region, thereby enabling the identification of the sound directions of those audio sources (if any) depicted in the image region. Hereinafter, we will refer to the main sound direction of the analysis zone as the analysis direction to avoid confusion with the main sound direction (to be) selected for obtaining the output audio signal via the application of the analysis zone. The analysis directions can include respective predefined sound directions of the spatial audio image represented by the received multi-channel audio signal, so that these predefined sound directions are mapped to respective predefined positions of the image region.
[0093] Figure 10 Schematically shows a plurality of analysis regions 313 overlaid on an image region, and image region positions A and B, which are again used to indicate the respective image region positions of a first object and a second object depicting respective sound sources representing a spatial audio image. In Figure 10 the example of, each of the analysis regions 313 overlaps with two or more adjacent analysis regions 313, while in other examples, the overlap between the analysis regions 313 can be greater than the overlap depicted in the Figure 10 example of, or the analysis regions 313 can be non-overlapping. A dynamic beamformer such as MVDR can be used to provide the analysis regions 313, and the applied beamformer can consider only a sub-part of the frequency range so as to be able to keep the analysis regions 313 as small as possible. Conversely, a static beamformer such as PS can be used to perform obtaining an output audio signal according to a selected main sound direction, thereby obtaining a larger (shifted) audio focus region compared to the analysis regions 313, as will be described hereinafter.
[0094] The analysis process can include: for each of the said analysis directions, estimating the energy of the corresponding preliminary beamformed audio signal obtainable via the applied dynamic beamformer; and identifying those analysis directions of the corresponding preliminary beamformed audio signals that result in an energy exceeding an energy threshold. In this regard, along the lines described above in the context of the third example, the energy estimation can be performed by obtaining the corresponding preliminary beamformed audio signals and calculating their energy or by applying a directional energy estimation method (appropriately modified) associated with the applied dynamic beamformer. The energy threshold can be a predefined energy threshold, or can be defined, for example, based on the average audio signal energy over a predefined-duration time window. The identified analysis directions are regarded as the analysis directions representing the respective (different) sound sources. Thus, the selection of the main sound direction for obtaining the output signal is partly based on the knowledge of the identified analysis directions representing the respective (different) sound sources.
[0095] As an example, selecting the main sound direction according to the identified analysis directions can apply a plurality of candidate sound directions described above in the context of the third example to identify candidate sound directions, thereby resulting in respective candidate audio focus regions that contain the minimum contribution in the identified analysis directions, and selecting the identified candidate sound direction as the main sound direction. Referring to the Figure 10 example of, and assuming that the received audio focus direction is mapped to the position A in the image region (showing the first object here) and the available candidate sound directions include those that result in Figure 9Those sound directions of the candidate audio focus regions 311, 311a, 311b, 311c shown in [description], the analysis process will result in identifying the analysis directions of the candidate analysis regions 313a and 313b as the analysis directions representing the corresponding (different) sound sources. Since in this example, the candidate audio focus region 311a contains the identified analysis direction leading to the analysis region 313, while both candidate audio focus regions 311b and 311c contain the identified analysis directions leading to the analysis regions 313a and 313b, thus, identifying the candidate sound direction of the candidate audio focus region that results in the smallest contribution among the identified analysis directions will result in identifying the candidate sound direction that gives rise to the audio focus region 311a, and thus selecting the identified candidate sound direction as the main sound direction.
[0096] In the example, identifying the candidate sound direction of the candidate audio focus region that results in the smallest contribution among the identified audio directions may include identifying the candidate sound direction of the candidate audio focus region that results in the smallest number of the identified audio directions. In another example, identifying the candidate sound direction of the candidate audio focus region that results in the smallest contribution among the identified audio directions may include identifying the candidate sound direction of the candidate beamformed audio signal that results in the smallest energy contribution from the identified audio directions.
[0097] Therefore, the analysis process applied in the fourth example enables avoiding emphasizing at least some sound sources that are in the sound directions near the received audio focus direction but are preferably excluded from the output audio signal, thus achieving improved selectivity due to avoiding the known spatial positions of the undesired sound sources, and enabling an improved user experience for audio focusing.
[0098] Still referring to the fourth example, the analysis of the analysis regions generated from the corresponding analysis directions and the subsequent selection of one of the available candidate focus directions as the main sound direction can be performed separately for multiple frequency sub-bands. Therefore, different candidate sound directions among the available candidate sound directions can be selected as the main focus directions in different frequency bands.
[0099] Figure 11 A block diagram shows some components of an exemplary device 900. The device 900 may include other components, units, or parts not shown in [description]. For example, the device 900 can be used when implementing one or more components described above in the context of the media capture entity 110 and / or the media rendering entity 210. Figure 11 For example, the device 900 can be used when implementing one or more components described above in the context of the media capture entity 110 and / or the media rendering entity 210.
[0100] Device 900 includes a processor 916 and a memory 915 for storing data and computer program code 917. A portion of the memory 915 and the computer program code 917 stored therein may further be configured to, together with the processor 916, implement at least some of the operations, processes, and / or functions described hereinabove in the context of the media capture entity 110 and / or the media rendering entity 210 or one or more of their components.
[0101] Device 900 includes a communication section 912 for communicating with other devices. The communication section 912 includes at least one communication device enabling wired or wireless communication with other devices. The communication devices of the communication section 912 may also be referred to as corresponding communication components.
[0102] Device 900 may further include a user I / O (input / output) component 918, which may be configured to, possibly together with the processor 916 and a portion of the computer program code 917, provide a user interface for receiving input from a user of the device 900 and / or providing output to a user of the device 900 to control at least some aspects of the operation of the media capture entity 110 and / or the media rendering entity 210 or one or more of their components implemented by the device 900. The user I / O component 918 may include hardware components such as a display, a touch screen, a touchpad, a mouse, a keyboard, and / or an arrangement of one or more keys or buttons, etc. The user I / O component 918 may also be referred to as a peripheral device. The processor 916 may be configured to control the operation of the device 900, for example, in accordance with a portion of the computer program code 917 and possibly further in accordance with user input received via the user I / O component 918 and / or in accordance with information received via the communication section 912.
[0103] Although the processor 916 is depicted as a single component, it may be implemented as one or more separate processing components. Similarly, although the memory 915 is depicted as a single component, it may be implemented as one or more separate components, some or all of which may be integrated / removable, and / or may provide permanent / semi-permanent / dynamic / cache storage.
[0104] The computer program code 917 stored in the memory 915 may include computer-executable instructions that, when loaded into the processor 916, control one or more aspects of the operation of the apparatus 900. As an example, the computer-executable instructions may be provided as one or more sequences of one or more instructions. By reading one or more sequences of one or more instructions contained therein from the memory 915, the processor 916 is able to load and execute the computer program code 917. One or more sequences of one or more instructions may be configured to, when executed by the processor 916, cause the apparatus 900 to perform at least some of the operations, processes, and / or functions described hereinabove in the context of the media capture entity 110 and / or the media rendering entity 210 or one or more of its components.
[0105] Accordingly, the apparatus 900 may include at least one processor 916 and at least one memory 915 including computer program code 917 for one or more programs, the at least one memory 915 and the computer program code 917 being configured to, with the at least one processor 916, cause the apparatus 900 to perform at least some of the operations, processes, and / or functions described hereinabove in the context of the media capture entity 110 and / or the media rendering entity 210 or one or more of its components.
[0106] The computer program stored in the memory 915 may be provided as, for example, a corresponding computer program product that includes at least one computer-readable non-transitory medium having stored thereon the computer program code 917, the computer program code causing the apparatus 900 to perform at least some of the operations, processes, and / or functions described hereinabove in the context of the media capture entity 110 and / or the media rendering entity 210 or one or more of its components when executed by the apparatus 900. The computer-readable non-transitory medium may include a memory device or a recording medium such as a CD-ROM, a DVD, a Blu-ray disc, or another article of manufacture tangibly embodying the computer program. As another example, the computer program may be provided as a signal configured to reliably convey the computer program.
[0107] References to a processor should not be construed as limited to programmable processors only, but also include dedicated circuits such as field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), signal processors. The features described in the foregoing description may be used in combinations other than the explicitly described combinations.
[0108] Although functions have been described with reference to certain features, those functions may be performed by other features whether described or not. Although features have been described with reference to certain embodiments, those features may also be present in other embodiments whether described or not.
Claims
1. An apparatus for audio focusing, the apparatus comprising at least one processor and at least one memory including computer program code, the computer program code, when executed by the at least one processor, causing the apparatus to: Receive a multi-channel audio signal, the multi-channel audio signal representing sounds in sound directions corresponding to respective positions in an image area of an image; Receive an indication of an audio focus direction corresponding to a first position in the image area; Select a main sound direction such that it corresponds to a second position in the image area, the second position being defined by shifting the first position to the second position such that the second position deviates from the first position in a direction that is farther from the center point of the image area; And Based on the multi-channel audio signal, obtain an output audio signal according to the selected main sound direction, wherein sounds in the sound direction defined by the selected main sound direction are emphasized relative to sounds in sound directions other than the sound direction defined by the selected main sound direction.
2. The apparatus according to claim 1, Wherein, The degree of deviation is at least one of the following: Depending on the position of the first position within the image area; and Increasing as the distance from the center point of the image area increases.
3. The apparatus according to claim 2, Wherein, The image area is divided into a plurality of non-overlapping image parts, and the degree of the deviation depends on the image part in which the first position is located.
4. The apparatus according to claim 1, Wherein, The direction of the deviation is at least one of the following: Depending on the position of the first position within the image area; and / or Along a conceptual line intersecting both the first position and the center point of the image area.
5. The apparatus according to claim 4, Wherein, The image area is divided into a plurality of non-overlapping image parts, and the direction of the deviation depends on the image part in which the first position is located.
6. The apparatus according to claim 1, Wherein, Obtaining the output audio signal further causes the apparatus to apply a beamformer to extract a beamformed audio signal representing the sound in the main sound direction from the multi-channel audio signal, and wherein the apparatus is caused to select the beamformer according to the position of the first position within the image area for obtaining the output audio signal.
7. The apparatus according to claim 6, Wherein, Selecting the beamformer further causes the apparatus to: In response to the first position being within a predefined distance from the center point of the image area, select a dynamic beamformer; and In response to the first position being farther from the center point of the image area than the predefined distance, select a static beamformer.
8. The apparatus according to claim 6, Wherein, The image area is divided into a plurality of non-overlapping image parts, and wherein the apparatus is caused to select the beamformer according to the image part in which the first position is located.
9. The apparatus according to claim 8, Wherein, The apparatus is caused to perform at least one of the following: select a dynamic beamformer for an image portion bounded by a single edge of the image region and / or an image portion not bounded by an edge of the image region; and select a static beamformer for an image portion bounded by two non-opposite edges of the image region.
10. The apparatus according to claim 7, wherein the static beamformer includes a phase-shift beamformer, and wherein the dynamic beamformer includes a minimum variance distortionless response beamformer.
11. The apparatus according to claim 1, wherein the main sound direction is selected such that the apparatus selects the main sound direction for at least two frequency sub-bands respectively.
12. The apparatus according to claim 1, wherein an output audio signal is obtained such that the apparatus applies a beamformer to extract a beamformed audio signal from the multi-channel audio signal, the beamformed audio signal representing sound in a sound direction within an audio focus area around the selected main sound direction, and wherein the main sound direction is selected such that the apparatus selects the main sound direction in view of the characteristics of the beamformer so that the audio focus area includes the received audio focus direction.
13. A method for audio focusing, the method comprising: receiving a multi-channel audio signal representing sound in sound directions corresponding to respective positions in an image region of an image; receiving an indication of an audio focus direction corresponding to a first position in the image region; selecting a main sound direction such that it corresponds to a second position in the image region, the second position being defined by shifting the first position to the second position such that the second position deviates from the first position in a direction further away from the center point of the image region; and obtaining an output audio signal based on the multi-channel audio signal according to the selected main sound direction, wherein sound in the sound direction defined by the selected main sound direction is emphasized relative to sound in sound directions other than the sound direction defined via the selected main sound direction.
14. The method according to claim 13, wherein obtaining the output audio signal further comprises: applying a beamformer to extract a beamformed audio signal representing sound in the main sound direction from the multi-channel audio signal, and wherein the method further comprises: selecting the beamformer for obtaining the output audio signal according to the position of the first position within the image region.
15. The method according to claim 14, wherein selecting the beamformer further comprises: selecting a dynamic beamformer in response to the first position being within a predefined distance from the center point of the image region; and selecting a static beamformer in response to the first position being further from the center point of the image region than the predefined distance.
Citation Information
Patent Citations
Audio Signal Beam Forming
US20160249134A1