Apparatus and related methods for spatial audio capture
By receiving spatial audio data and video images, and controlling audio capture attributes using direction information and user input, beamforming technology is used to solve the difficulties of spatial audio capture and control, achieving efficient audio capture and rich user experience.
Patent Information
- Application Number
- CN202080037691.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-20
- Filing Date
- 2020-05-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-05-11
AI Technical Summary
Capture and control of spatial audio are difficult in the prior art, especially under the limitations of the capture device, to provide efficient control of the audio capture attributes of spatial audio content.
By receiving spatial audio data and video images, using direction information to correlate the audio source with the video image area or out-of-view graphics, providing user input to select the audio source and controlling its capture attributes, and beamforming technology is used to focus audio capture.
It realizes efficient capture and control of spatial audio data, and can accurately select and adjust the capture attributes of audio sources according to user input, overcomes device limitations and technical obstacles, and provides a richer user experience.
Smart Images

Figure CN113853529B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of capture of spatial audio. In particular, the present disclosure relates to the presentation of a user interface that provides modification of one or more audio capture attributes of spatial audio, associated apparatus, methods, and computer programs. Background Art
[0002] The capture of spatial audio may be useful, and the control of such capture may be difficult.
[0003] A list or discussion of prior published documents or any background in this specification should not be taken as an admission that the document or background is part of the prior art or common general knowledge. One or more aspects / examples of the present disclosure may or may not address one or more of the background problems. Summary of the Invention
[0004] In a first example aspect, there is provided an apparatus including components configured to:
[0005] Receive spatial audio data including audio captured from one or more audio sources in a space extending around a capture device and direction information indicating at least a direction towards the one or more audio sources, wherein the spatial audio data is captured by the capture device;
[0006] Receive video images captured by a camera of the capture device, the video images having a field of view, wherein the extent of the space in which the spatial audio data is captured is greater than the field of view;
[0007] For an audio source within the field of view, associate each of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction towards the audio source, and for an audio source outside the field of view, associate each of the one or more audio sources determined according to the direction information with a part of an off-view graphic indicating the spatial extent of the space outside the field of view corresponding to the direction towards the audio source;
[0008] Provide a display on a display of the video images together with the off-view graphic;
[0009] Receive a user input for selecting a region of the video image or a part of the off-view graphic;
[0010] Provide control of at least one audio capture attribute of a selected one of the one or more audio sources, wherein the selected audio source among the one or more audio sources includes an audio source among the one or more audio sources associated with the region or part selected by the user input.
[0011] In one or more examples, the component is configured to provide a display of a marker at one or more of the following (e.g., both of these):
[0012] One or more portions of off-view graphics corresponding to one or more directions toward one or more audio sources; and
[0013] One or more regions of a video image corresponding to one or more directions toward one or more audio sources.
[0014] In one or more examples, control of at least one audio capture attribute includes the component being configured to provide signaling to cause capture or recording of a selected one of the audio sources using beamforming techniques.
[0015] In one or more examples, control of at least one audio capture attribute includes the component being configured to perform at least one of the following:
[0016] Capture or record a selected one of the audio sources with a greater volume gain relative to the volume gain of other audio applied to the spatial audio data;
[0017] Capture or record a selected one of the audio sources with a higher quality relative to the quality of other audio applied to the spatial audio data; or
[0018] Capture or record the audio of a selected one of the audio sources as an audio stream separate from the other audio of the spatial audio data.
[0019] In one or more examples, the component is configured to determine one or more audio sources by using direction information to determine from which direction audio having a volume above a predetermined threshold is received.
[0020] In one or more examples, off-view graphics indicating a spatial extent of space outside the field of view includes at least one of the following:
[0021] A line, where positioning along the line from one end to the other represents a direction of receiving audio of an audio source from a direction corresponding to at least a first edge of the field of view to a direction corresponding to at least a second edge of the field of view opposite the first edge; or
[0022] A sector of an ellipse, where positioning within the sector represents a direction of receiving audio of an audio source from a direction corresponding to at least a first edge of the field of view to a direction corresponding to at least a second edge of the field of view opposite the first edge.
[0023] In one or more examples, an off-view graphic indicating a spatial extent of space outside the field of view surrounds a plane of a capture device, where the positioning of a presented marker relative to the off-view graphic represents an azimuthal direction from which audio of an audio source is received, and where the positioning of a marker depicted at a distance above or below the off-view graphic corresponds to an elevation direction from which the audio of the audio source is received above or below the plane.
[0024] In one or more examples, an off-view graphic indicating a spatial extent of space outside the field of view includes lines, where the positioning along a line from one end to the other represents an azimuthal direction from which audio of an audio source is received from an azimuthal direction corresponding at least to a first edge of the field of view to an azimuthal direction corresponding at least to a second edge of the field of view opposite the first edge, and where a distance above or below the line corresponds to an elevation direction from which the audio of the audio source is received.
[0025] In one or more examples, the component is configured to: provide control over at least one audio capture attribute by modifying the audio capture attribute based on a user input including a tap that selects a region of a video image or a portion of the off-view graphic at a location on a touch-sensitive input device, the beamforming technique focusing on a region of space corresponding to the selected region or portion.
[0026] In one or more examples, the component is configured to: provide control over at least one audio capture attribute by modifying the audio capture attribute based on a user input including a pinch gesture that selects a region of a video image or a portion of the off-view graphic at a location on a touch-sensitive input device, the beamforming technique having a degree related to the size of the pinch gesture.
[0027] In one or more examples, the component is configured to: provide a display of a second marker indicating the absence of an audio source in a direction corresponding to a selected region of the video image or a portion of the off-view graphic based on a user input that selects a region of the video image or a portion of the off-view graphic for which there is no associated audio source.
[0028] In one or more examples, the beamforming technique includes at least one of the following: delay-and-sum beamforming technique or parametric spatial audio processing technique in which audio of a selected audio source is emphasized.
[0029] In one or more examples, the component is configured to provide one or more (e.g., both) of presentation and recording of spatial audio data using a selected audio source having controlled audio capture attributes.
[0030] In one or more examples, the component of the device includes at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the device to perform the functions of the first aspect.
[0031] In a second example aspect, there is provided an electronic device including the device of the first aspect, a camera configured to capture video images, a plurality of microphones configured to capture spatial audio data, and a display for use by the device to display video images and off-view graphics.
[0032] In a third example aspect, there is provided a method including:
[0033] Receiving spatial audio data including audio captured from one or more audio sources in a space extending around a capture device and direction information at least indicating a direction towards the one or more audio sources, wherein the spatial audio data is captured by the capture device;
[0034] Receiving video images captured by a camera of the capture device, the video images having a field of view, wherein a range of the space in which the spatial audio data is captured is greater than the field of view;
[0035] For an audio source within the field of view, associating each of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction towards the audio source, and for an audio source outside the field of view, associating each of the one or more audio sources determined according to the direction information with a part of the off-view graphics corresponding to the direction towards the audio source, the off-view graphics indicating a spatial extent of the space outside the field of view;
[0036] Providing a display on a display of the video images together with the off-view graphics;
[0037] Receiving a user input for selecting a region of the video image or a part of the off-view graphics;
[0038] Providing control of at least one audio capture attribute of a selected one of the one or more audio sources, wherein the selected one of the one or more audio sources includes an audio source among the one or more audio sources associated with the region or part selected by the user input.
[0039] In one or more examples, the method includes providing a display of a marker at one or both of the following:
[0040] One or more parts of the off-view graphics corresponding to one or more directions towards one or more audio sources; and
[0041] One or more regions of a video image corresponding to one or more directions towards one or more audio sources.
[0042] In one or more examples, controlling at least one audio capture attribute includes providing signaling to cause capture or recording of a selected one of the audio sources using beamforming techniques.
[0043] In one or more examples, controlling at least one audio capture attribute includes a method of performing at least one of the following:
[0044] Capture or record a selected one of the audio sources with a greater volume gain relative to the volume gain of other audio applied to the spatial audio data;
[0045] Capture or record a selected one of the audio sources with a higher quality relative to the quality of other audio applied to the spatial audio data; or
[0046] Capture or record the audio of a selected one of the audio sources as an audio stream separate from the other audio of the spatial audio data.
[0047] In one or more examples, the method includes determining one or more audio sources by using direction information to determine from which direction audio having a volume above a predetermined threshold is received.
[0048] In one or more examples, the method includes receiving user input including a tap selecting a region of the video image or a portion of an out-of-view graphic at a location on a touch-sensitive input device, and providing control over at least one audio capture attribute by modifying the audio capture attribute by applying beamforming techniques that focus on a region of space corresponding to the selected region or portion.
[0049] In one or more examples, the method receives user input including a pinch gesture selecting a region of the video image or a portion of an out-of-view graphic at a location on a touch-sensitive input device, and provides control over at least one audio capture attribute by modifying the audio capture attribute by applying beamforming techniques having a degree related to the size of the pinch gesture.
[0050] In one or more examples, the method includes receiving user input selecting a region of the video image or a portion of an out-of-view graphic for which there is no associated audio source, and providing a display of a second marker for indicating that there is no audio source in the direction corresponding to the selected region or portion of the out-of-view graphic of the video image.
[0051] In one or more examples, the method includes providing one or both of rendering and recording of spatial audio data using the selected audio source having a controlled audio capture attribute.
[0052] In a fourth example aspect, there is provided a computer-readable medium including computer program code stored thereon, the computer-readable medium and the computer program code being configured to perform a method when run on at least one processor, the method including:
[0053] Receiving spatial audio data, the spatial audio data including audio captured from one or more audio sources in a space extending around a capture device and direction information indicating at least a direction towards the one or more audio sources, wherein the spatial audio data is captured by the capture device;
[0054] Receiving video images captured by a camera of the capture device, the video images having a field of view, wherein a range of the space in which the spatial audio data is captured is larger than the field of view;
[0055] For an audio source within the field of view, associating each audio source of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction towards the audio source, and for an audio source outside the field of view, associating each audio source of the one or more audio sources determined according to the direction information with a part of a view-outside graphic corresponding to the direction towards the audio source, the view-outside graphic indicating a spatial extent of the space outside the field of view;
[0056] Providing a display on a display of the video images together with the view-outside graphic;
[0057] Receiving a user input for selecting a region of the video images or a part of the view-outside graphic;
[0058] Providing control of at least one audio capture attribute of a selected audio source of the one or more audio sources, wherein the selected audio source of the one or more audio sources includes an audio source among the one or more audio sources associated with the region or part selected by the user input.
[0059] In a fourth example aspect, there is provided an apparatus, the apparatus including:
[0060] At least one processor; and
[0061] At least one memory including computer program code,
[0062] The at least one memory and the computer program code being configured to, together with the at least one processor, cause the apparatus to at least perform the following operations:
[0063] Receive spatial audio data, the spatial audio data including audio captured from one or more audio sources in a space extending around a capture device and direction information indicating at least a direction towards the one or more audio sources, wherein the spatial audio data is captured by the capture device;
[0064] Receive a video image captured by a camera of a capture device, the video image having a field of view, wherein a range of the space in which the spatial audio data is captured is greater than the field of view;
[0065] For an audio source within the field of view, associate each audio source of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction towards the audio source, and for an audio source outside the field of view, associate each audio source of the one or more audio sources determined according to the direction information with a part of an out-of-view graphic corresponding to the direction towards the audio source, the out-of-view graphic indicating a spatial extent of the space outside the field of view;
[0066] Provide a display of the video image together with the out-of-view graphic on a display;
[0067] Receive a user input for selecting a region of the video image or a part of the out-of-view graphic;
[0068] Provide control of at least one audio capture attribute of a selected audio source of the one or more audio sources, wherein the selected audio source of the one or more audio sources includes an audio source among the one or more audio sources associated with the region or part selected by the user input.
[0069] Optional features of the first aspect are equally applicable to the apparatus of the fourth aspect. In addition, the functions provided by the optional features of the first aspect can be performed by the method of the second aspect and the code of the computer-readable medium of the third aspect.
[0070] This disclosure includes one or more corresponding aspects, examples, or features, either alone or in various combinations, whether or not specifically stated (including claimed) in that combination or alone. Corresponding components and corresponding functional units for performing one or more of the discussed functions (e.g., function enablers, AR / VR graphics renderers, display devices) are also within this disclosure.
[0071] Corresponding computer programs for implementing one or more of the disclosed methods are also within this disclosure and are covered by one or more of the described examples.
[0072] The above overview is intended to be merely exemplary and not restrictive. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Now described only by way of example with reference to the accompanying drawings, in which:
[0074] Figure 1 Illustrates an example apparatus for controlling at least one audio capture attribute, the apparatus being shown as part of an electronic device or "capture device" in a space having an audio source;
[0075] Figure 2 Shows a first example view of a display showing an interface based on signaling from the apparatus;
[0076] Figure 3 Shows a second example view of a display showing an interface based on signaling from the apparatus;
[0077] Figure 4 Shows a third example view of a display showing an interface based on signaling from the apparatus;
[0078] Figure 5 Shows a fourth example view of a display showing an interface based on signaling from the apparatus, where the user provides user input including a pinch gesture;
[0079] Figure 6 Shows a fifth example view of a display showing an interface based on signaling from the apparatus, where the user provides user input for selecting a portion of a graphic outside the view;
[0080] Figure 7 Shows a flowchart illustrating an example method; and
[0081] Figure 8 Shows a computer-readable medium. DETAILED DESCRIPTION
[0082] The capture of spatial audio can be used to provide rich and diverse user experiences, such as in the fields of virtual reality, augmented reality, communication, and video capture. As a result, the number of devices capable of capturing spatial audio may increase. Given that spatial audio includes capturing audio with direction information that indicates the direction towards one or more audio sources, or in other words, the direction from which the audio arrives from one or more audio sources, the effective capture of such audio can be complex. It may be desirable to provide control over one or more audio capture attributes of spatial audio content in an efficient manner, despite the limitations of the devices capturing the spatial audio.
[0083] Spatial audio includes audio with direction information such as that captured by a spatial audio capture device. Thus, the captured spatial audio can have information representing the audio itself and information indicating the spatial arrangement of the source of the audio in the space around the spatial audio capture device. The spatial audio can be presented to a user in such a way that each audio source can be perceived as originating from a specific location, as if the individual sources of the audio were located at those specific locations. Spatial audio data includes audio for presentation as spatial audio and thus typically includes audio and direction information, such as explicitly specified as metadata or inherently present in the way the audio is captured. The spatial audio data can be presented such that its component audio (e.g., audio sources captured in space) is perceived as originating from one or more points or one or more directions according to the direction information. Audio rendering can take into account early reflections and reverberation, which can be modeled, for example, according to the virtual or real space in which the audio presentation occurs.
[0084] The captured spatial audio can be parametric spatial audio, such as DirAC or first- or higher-order Ambisonics (FOA, HOA). The capture of spatial audio data can be provided by using a number of microphones (such as at least three). In one or more examples, parametric spatial audio capture processing can be used. As is known to those skilled in the art, parametric spatial audio capture can include: for each time-frequency bin of the captured multi-microphone signal, analyzing sufficient spatial parameters to represent the perceptually relevant attributes of the signal. For example, these can include direction of arrival and ratio parameters, such as diffuseness for each time-frequency bin. The spatial audio signal can then be represented with direction information (e.g., spatial metadata), which can include a transfer signal formed from the multi-microphone input signal. During rendering, the transferred audio signal together with the direction information is used to synthesize a sound field that produces an auditory perception similar to that of a listener when the listener's head is located at the position of the microphone arrangement.
[0085] Spatial localization of spatial audio can be provided by 3D audio effects, such as 3D audio effects that utilize head-related transfer functions to create a spatial audio space (aligned with the real-world space in the case of augmented reality), and audio can be localized in this spatial audio space for presentation to the user. Spatial audio can be presented by headphones using head-related transfer function (HRTF) filtering techniques, or by speakers by using vector base amplitude panning techniques to localize the perceived auditory source of the audio content. Spatial audio can use one or more of the volume difference, time difference, and pitch difference between the audible presentations presented to each ear of the user to create a perception of a specific location or specific direction of the audio source in space (e.g., not necessarily aligned with the speakers). The perceived distance to the perceived source of the audio can be rendered by controlling the amount of reverb and gain to indicate proximity or distance from the perceived source of the spatial audio. It should be understood that spatial audio presentation as described herein can involve audio presentations that only have a perceived direction towards their origin and audio presentations that give the origin of the audio a perceived location (e.g., including the perceived distance from the user).
[0086] Virtual reality (VR) content can be provided with spatial audio having a direction attribute such that the audio is perceived to originate from a point in the VR space, which can be linked to an image of the VR content. Augmented or mixed reality content can be provided with spatial audio such that the spatial audio is perceived to originate from real-world objects visible to the user and / or from augmented reality graphics overlaid on the user's view. Communication between electronic devices can use spatial audio to present an auditory scene perceived by a first user to a second user who is remote from the first user.
[0087] Figure 1 An example apparatus 100 is shown that is configured to provide control of at least one audio capture attribute of an audio source for selection from one or more audio sources. Apparatus 100 includes components such as a processor 101 and a memory 102 for receiving spatial audio data and providing control of the audio capture attributes. In this example and one or more examples, apparatus 100 can be part of an electronic device 103 such as a smart phone or a tablet. Electronic device 103 can include an embodiment of a capture device configured to receive spatial audio data and / or video images.
[0088] Device 100 is configured to receive spatial audio data from one or more microphones 104. In one or more examples, the microphone 104 may be part of the electronic device 103, but in other examples may be separate from the electronic device 103. The one or more microphones 104 may include, for example, at least three microphones arranged as a microphone array for capturing spatial audio data. The device 100 or the electronic device 103 may be configured to process the audio captured from the microphone 104 to generate associated direction information. In one or more examples, tracking of an audio source in the space 105 around the electronic device 103 may be used to generate the direction information.
[0089] Device 100 is configured to receive video images from the camera 106. In one or more examples, the camera may be part of the electronic device 103, but in other examples may be separate from the electronic device 103. The camera has a field of view 107 of the space 105, and the field of view 107 is represented by an arrow between a first edge 108 of the field of view and a second edge 109 of the field of view. The field of view 107 of the camera 106 is smaller than the spatial extent of the space 105 from which the spatial audio data is captured by the microphone 104. Thus, there is a region 110 of the space 105 outside the field of view 107 of the camera 106. The electronic device 103 may be referred to as a "capture device" because it is used to capture spatial audio data and video images. However, if the camera 106 and the microphone 104 are separate from or independent of the electronic device 103, the camera 106 and the microphone 104 may be collectively regarded as including the capture device.
[0090] Device 100 may be configured to provide a display by signaling to the display 111. The display 111 may be associated with a touch-sensitive user input device 112 that provides touchscreen input to be provided to a user interface presented on the display 111. It should be understood that other user input functions may be provided by the device 100 or by the electronic device 103 for use by the device 100.
[0091] Although in this example, the device 100 is shown as part of the electronic device 103 and may share hardware resources such as the processor 101, the memory 102, the camera 106, the display 111, and the microphone 104 with the electronic device 103, in other embodiments, the device 100 may include a part of a server (not shown) that communicates with the electronic device 103 or with the camera 106, the microphone 104, and the display 111, regardless of whether the camera 106, the microphone 104, and the display 111 are part of the electronic device 103. Thus, the device 100 may utilize communication elements to receive spatial audio data and video images and provide signaling to cause an image to be displayed on the display.
[0092] Regardless of how the apparatus 100 is embodied, such as in the form of part of a server or an electronic device 103, the apparatus 100 can include or be connected to a processor 101 and a memory 102, and can be configured to execute computer program code. The apparatus 100 can have only one processor 101 and one memory 102, but it should be understood that other embodiments can utilize more than one processor and / or more than one memory (e.g., the same or different processor / memory types). Additionally, the apparatus 100 can be an application specific integrated circuit (ASIC).
[0093] The processor can be a general-purpose processor dedicated to executing / processing information received from other components (such as from a microphone 104, a camera 106, and a touch-sensitive user input device 112) according to instructions stored in the memory in the form of computer program code. The output signaling generated by such operation of the processor is provided to other components, such as a display 111 or an audio processing module configured to process spatial audio data according to the instructions of the apparatus 100. In other examples, the apparatus 100 can include components for processing spatial audio data and can modify spatial audio capture attributes.
[0094] The memory 102 (not necessarily a single memory unit) is a computer-readable medium (in this example, a solid-state memory, but can be other types of memory, such as a hard disk drive, ROM, RAM, flash memory, etc.) that stores computer program code. When the program code runs on the processor, the computer program code stores instructions executable by the processor. In one or more example embodiments, the internal connection between the memory and the processor can be understood to provide an active coupling between the processor and the memory to allow the processor to access the computer program code stored on the memory.
[0095] In this example, the corresponding processor and memory are electrically connected to each other internally to allow electrical communication between the corresponding components. In this example, all components are placed close to each other so as to form an ASIC together, in other words, so as to be integrated together as a single chip / circuit that can be installed into an electronic device. In some examples, one or more or all components can be placed separately from each other.
[0096] In one or more examples, the apparatus 100 is configured to receive spatial audio data that includes audio captured from one or more audio sources in a space 105 extending around the electronic device 103. In Figure 1In the example, there are four audio sources, including a first audio source 113 and a second audio source 114 within the field of view 107 of the camera 106, and a third audio source 115 and a fourth audio source 116 outside the field of view 107 of the camera 106 (i.e., region 110). The apparatus 100 may be configured to identify the first audio source 113 to the fourth audio source 116 as audio sources when they are currently generating audio. In other examples, the first audio source 113 to the fourth audio source 116 may be considered audio sources when less than a predetermined silence time has elapsed since the audio source last generated audio. Depending on user preferences, the silence time may include up to 5, 10, 20, 30, 40, 50 or 60 seconds or more. Thus, the apparatus 100 may be configured to analyze the captured audio and determine one or more of the audio sources based on whether the audio is currently audible or has been audible within the silence time. In other examples, the apparatus may receive information identifying where in the spatial audio data the audio source is present. The spatial audio data also includes at least direction information indicating the direction towards the one or more audio sources. Thus, the direction information may indicate a first direction 117 of the first audio source 113; a second direction 118 of the second audio source 114; a third direction 119 of the third audio source 115; and a fourth direction 120 of the fourth audio source 116. It should be understood that the spatial audio data may be encoded in a variety of different ways and the directions 117 - 120 may be recorded as metadata or the audio itself may be encoded to indicate the directions 117 - 120 and other techniques.
[0097] As described above, the apparatus 100 may be configured to receive video images captured by the camera 106 of the electronic device 103, where the spatial extent of the space 105 in which the spatial audio data is captured is greater than the field of view 107. Thus, the audio from the third audio source 115 and the fourth audio source 116 has characteristics in the spatial audio data, but the images of the third audio source 115 and the fourth audio source 116 do not have characteristics in the video image at a given time. It should be understood that the electronic device 103 may move around the space 105 while the video image and the spatial audio data are being captured, such that the field of view falls on other audio sources over time. Thus, the audio sources within the field of view 107 may change over time.
[0098] Example Figure 2An electronic device 103 and its display 111 are shown, with a user interface presented on the display 111. The apparatus 100 is configured to provide a display of video images from the camera 106. Thus, the apparatus 100 can provide signaling such that video images captured within the field of view 107 of the camera 106 are presented on the display 111. It should be understood that the range of content captured by the camera may not be exactly the same as the content presented on the display 111. For example, the camera 106 can default to cropping regions to match the resolution or aspect of the video images to the display 111. Thus, the field of view 107 of the camera 106 can be considered to include the field of view for presentation on the display 111. In Figure 2 the example of, the first audio source 113 and the second audio source 114 are visible in the video images provided for presentation on the display 111.
[0099] Example Figure 2 A first example of an off-view graphic 200 is shown. The off-view graphic 200 includes graphic elements or images that are displayed to represent the spatial extent of the space 105 outside the field of view 107. In particular, it can represent the extent of the space 105 outside the field of view 107 from which spatial audio data is captured, such as only that portion of the space 105 outside the field of view. Thus, audio sources that appear in the video images are not represented on the off-view graphic 200. In one or more examples, the off-view graphic 200 can not only represent the space 105 outside the field of view 107, and can include portions that represent portions of the space 105 within the field of view 107.
[0100] In this example and other examples, the off-view graphic 200 includes a sector of an ellipse, such as a half-ellipse. Thus, an ellipse or a circle can be used to represent 360 degrees of the space 105 around the electronic device 103, while a half-ellipse or other sector portion can represent the region 110 of the space 105 outside the field of view 107. In one or more examples, the off-view graphic 200 has a first radial portion 201 and a second radial portion 202, the first radial portion 201 representing a direction corresponding to at least a first edge 108 of the field of view 107, and the second radial portion 202 representing a direction corresponding to at least a second edge 109 of the field of view 107, the second edge 109 being opposite the first edge 108. Assuming the off-view graphic 200 represents the portion of the space 105 outside the field of view 107. The positioning within the off-view graphic can be used to represent the direction from which the audio of the audio sources 115, 116 is received.
[0101] To provide control over the audio capture properties of an audio source selected based on its location shown on the display 111, the device 100 may associate a region / portion of the displayed video image or off-view graphics 200 with each of one or more audio sources, which audio sources themselves may be determined from orientation information. Thus, for audio sources 113, 114 within the field of view 107, the device 100 may associate regions 203, 204 of the video image with audio sources 113, 114 or the directions towards the audio sources. For a third audio source 115 and a fourth audio source 116 outside the field of view 107, the device 100 may associate a portion of the off-view graphics 200 shown as markers 215 and 216 corresponding to the directions towards the audio sources 115, 116. Thus, marker 215 represents the location or direction towards the third audio source 115, and marker 216 represents the location or direction towards the fourth audio source 116.
[0102] The device 100 may be configured to receive user input. In this example, the user input may be provided by user input at the touch-sensitive user input device 112. It should be understood that other user input methods may be used, such as eye gaze position or movement of a cursor or pointer via a controller. The location of the user input on the display 111 may select a region of the video image or a portion of the off-view graphics 200, such as one of regions 203, 204 or one of markers 215 or 216. Given the associations made for these regions 203, 204 and markers 215, 216, the device 100 is provided with a selection of audio from one of the first audio source 113 to the fourth audio source 116. It should be understood that in other examples, multiple selections may be made. In the example Figure 2 shown, the user indicated by finger 206 has selected marker 216 and thus selected the audio of the fourth audio source 116 in the spatial audio data.
[0103] The device 100 may be configured to provide control over at least one audio capture property, where the control is specific to the selected audio source among the one or more audio sources 113-116.
[0104] Accordingly, the apparatus 100 may be configured to receive spatial audio data that represents audio captured from a direction range that is larger than the spatial range of a jointly received video image. Thus, there are technical limitations with the electronic device 103 or more generally with the incoming spatial audio data and video image, namely that when the audio of the audio sources 113-116 is captured, without a multi-camera such as a spherical camera arrangement, a visual image of the same degree cannot be captured. Such a multi-camera arrangement may limit the situations in which spatial audio data can be captured because such a multi-camera arrangement is generally large and cumbersome. Accordingly, the apparatus 100 provides the processing of the spatial audio data and video image and the presentation of the interface in such a way that the problems associated with the control of spatial audio capture and the technical limitation of the smaller field of view of the camera 106 combined with the larger capture area of the spatial audio data can be overcome.
[0105] The out-of-view graphic 200 is shown as having a number of arrows that may be labeled to assist the user in understanding what is represented. For example, the arrow 207 may be labeled with 180° to indicate that it represents a direction that is 180° from the forward direction of the video image. Similarly, the other arrows 208 and 209 may be labeled with 135° and 225° to indicate the directions represented by these portions of the out-of-view graphic 200.
[0106] In one or more examples, the apparatus 100 may provide the presentation of the spatial audio data via an audio presentation device (not shown) such as headphones. In other examples, given that the user of the electronic device 103 will be able to hear firsthand the audio captured as spatial audio data, the presentation of the spatial audio data may not be necessary. However, the presentation of the spatial audio data may be advantageous so that the user can understand the impact of the changes they have indicated via user input on the audio capture attributes. Accordingly, in one or more examples, the apparatus may be configured to provide the presentation of only the audio from the audio sources for which the audio capture attributes have been modified.
[0107] The positions of markers 215, 216 on off-view graphic 200 can be updated in real time or periodically to represent the current localization of audio sources outside the field of view 107. If an audio source moves from region 110 into the field of view 107, its associated marker can be removed from the display. Similarly, if an audio source moves into region 110, device 100 can add a marker to off-view graphic 200. In one or more examples, markers that are visually similar to markers 215, 216 or different can be presented at one or more regions 203, 204 of the video image that correspond to one or more directions towards one or more audio sources within the field of view 107. Thus, the device can provide the presentation of markers to indicate that device 100 views a person in the video image as the first audio source 113 or the second audio source 114 at the current time, and shows the localization of the audio sources (in addition to the audio sources that have features in the video image). In one or more examples, markers for audio sources in the video image can include outlines or translucent shadows to mark the associated regions 203, 204.
[0108] In one or more examples, device 100 is configured to provide control over live audio capture attributes when spatial audio data and video images are captured. In one or more examples, the spatial audio data and video images are captured and recorded simultaneously, and the device is provided with pre-recorded spatial audio data and pre-recorded video images.
[0109] Control over the audio capture attributes can be provided in various ways. In one or more examples, device 100 can be configured to control how the spatial audio data is captured, such as by modifying the arrangement of microphones 104 or the parameters of microphones 104, such as the gain applied to the audio source or the directional focus of the (multiple) microphones. In one or more examples, device 100 can be configured to control how the spatial audio data is recorded, and thus can provide audio processing of the spatial audio data and record the spatial audio data and the modifications to the audio capture attributes applied thereto. The purpose of controlling the audio capture attributes can be to provide enhancement of the audio from a particular audio source 113 - 116 or direction.
[0110] In one or more examples, control of audio capture attributes is provided by using beamforming techniques. Beamforming techniques can be used to capture a mono audio stream of a selected audio source. The mono audio stream can be specific to the selected audio source, while other audio sources can be recorded together in a common stream. Beamforming techniques can provide spatial audio data in which the audio source in the selected direction is relatively enhanced and / or the audio sources in other directions are relatively attenuated. An example of a beamforming technique is the delay and sum beamforming technique, which utilizes a microphone array of at least three microphones to focus the microphones on audio capture from the selected audio source or direction. Alternatively, the beamforming technique can include parametric spatial audio processing for forming a beamforming output, where a certain region or direction of the spatial audio is enhanced or "extracted" from the spatial audio field representing the audio received from space 105.
[0111] Thus, user input identifying a location on either the video image or the off-view graphic 200 can cause the device 100 to control the audio capture attributes of the selected audio source, such as by beamforming.
[0112] In one or more examples, control of at least one audio capture attribute includes components configured to capture or record a selected one of the audio sources 116 at a greater volume gain relative to the volume gain of the audio of other audio sources 113, 114, 115 applied to the spatial audio data. Thus, audio processing for increasing the volume level can be selectively applied to the audio from the direction of the fourth audio source 116. It should be understood that in other examples, rather than including a volume gain, the audio capture attribute can include the quality of capturing the audio from the selected audio source / direction. Thus, a higher bit rate can be used to record the audio from the selected audio source 116 compared to the bit rate of the other audio for the spatial audio data.
[0113] In the examples described, reference is made to being able to apply control of audio capture attributes to the audio received from the selected direction or to the audio from the selected audio source. In one or more embodiments, these two can be considered interchangeable. However, the device 100 can be configured to identify the direction towards the audio source and thus identify the presence of the audio source in the spatial audio data. The device 100 can be configured to use the direction information to determine from which direction the audio with a volume higher than a predetermined threshold is received. If audio higher than the threshold is received from a particular direction, it can be determined that the direction points to the audio source.
[0114] Different techniques can be used to locate the primary sound source in space 105. One example is the Steering Response Power Phase Transform (SRP-PHAT). This algorithm can be understood as a beamforming-based method that searches for candidate localizations or directions of the audio source and maximizes the output of a controlled delay and sum beamformer for "scanning" space 105. In one or more examples, to limit the computational burden of the method, the front, back, and / or sides of the electronic device 103 can be divided into sectors of a fixed size, and a fixed beamformer designed for each sector, which is formed by a microphone or microphone array, is shown together at 104. It can be understood that filtering can be applied to identify only the audio sources that meet a desired threshold. In one or more examples, the device can be configured to apply deep learning methods or voice activity detection to determine when the audio source is active and then search for its location / direction via beamforming means such as SRP-PHAT. When the location is determined, an association and / or markers 215, 216 can be placed at appropriate points on the out-of-view graphic 200. In one or more examples, the device 100 can apply a threshold duration during which the audio source needs to be active to be detected as an audio source. This can help filter out sounds of shorter duration that may be considered unwanted noise.
[0115] Example Figure 3 An alternative embodiment of the out-of-view graphic 300 is shown. Example Figure 3 Similar to the example Figure 2 , and thus uses the same reference numerals, except for the out-of-view graphic 300. Thus, in this example, the out-of-view graphic includes a line, where the orientation along the line from one end 301 to the other end 302 represents the direction of receiving audio from an audio source from a direction corresponding at least to the first edge 108 of the field of view 107 to a direction corresponding at least to the second edge 109 of the field of view 107 opposite the first edge 108.
[0116] As previously described, markers 215 and 216 are provided to show on the line to represent the location of the audio source in space 105 outside the field of view 107. In this example, only the azimuthal directions around the microphone 104 are represented. However, in other examples, the spherical localization of the audio sources 113 - 116 can be depicted, i.e., having a height above or below the horizontal plane extending around the electronic device 103.
[0117] Therefore, referring to the example Figure 4, in one or more examples, the out-of-view graphic 400 includes a line, wherein the positioning along the line from one end to the other end thereof represents an azimuth direction from which audio from the audio source is received from an azimuth direction corresponding to at least a first edge 108 of the field of view to an azimuth direction corresponding to at least a second edge 109 of the field of view opposite the first edge, and wherein distances 401, 402 above or below the out-of-view graphic 400 (e.g., the line) correspond to elevation directions from which audio from the audio sources 115, 116 is received. Thus, the marker 215 representing the third audio source 115 is shown as being centered between the ends of the line and at a distance 401 above the line to indicate that the audio is received from behind and above the microphone 104. Additionally, the marker 216 representing the fourth audio source 116 is shown as being at a distance 402 below the line to indicate that the audio is received from to the right and above the microphone 104. It will be appreciated that in some examples, elevations may be provided only for audio sources that are above the horizontal plane or only for audio sources that are below the horizontal plane.
[0118] User input has been described as causing control of audio capture properties, which may be provided by applying a predetermined audio focus to an audio source, such as through audio processing or beamforming. In one or more examples, the degree to which the audio processing or beamforming or other control is applied is controllable. Apparatus 100 may be configured to provide an efficient method for selecting and controlling what audio controls are applied to and how the audio capture properties are controlled.
[0119] Example Figure 5 With example Figure 4 Substantially the same, and like reference numerals are applied.In one or more examples, apparatus 100 may be configured to receive a pinch gesture. Figure 5 Two fingers 501 and 502 of a user are shown performing a pinch gesture on one of the markers 216. It should be understood that the pinch gesture may be applied to other markers 215 or to audio sources 113, 114 visible in the video image.
[0120] The apparatus 100 may be configured to provide modification of the audio capture properties by applying a beamforming technique having a degree associated with the size 503 of the pinch gesture. It should be understood that the pinch gesture may be used to select one of the audio sources and to control the degree of change to the audio capture properties. In one or more examples, control of the audio capture properties may be performed during application of the pinch gesture to "preview" the final effect. In one or more examples, the size 503 may be determined when the pinch gesture is completed and the user's finger is removed from the display 111. In summary, the user input may include a pinch gesture that selects an area of a video image or a portion of an out-of-view graphic at a location on the touch-sensitive user input device 112 (e.g., the location of the marker 216) and controls the degree of modification of the audio capture properties.
[0121] In terms of the application of beamforming technology, the size 503 of a pinch gesture can determine the dominance of an audio focus in spatial audio data relative to other spatial audio data. For example, the pinch gesture can control the beam width or maximum gain provided by the beamforming technology. In one or more examples, the beam width is the width of the beamforming technology, which is measured in degrees of a sector in which the audio source is amplified and outside of which the audio source is attenuated (e.g., relative to the amplified source). In one or more examples, the maximum gain is the decibel difference between the maximum amplified sound source and the maximum attenuated sound source.
[0122] The device can be configured to provide user feedback to the pinch gesture by controlling the size of a marker representing the audio source selected by the pinch gesture. Thus, marker 216 is shown as larger than marker 215 because it has caused the audio of the associated fourth audio source 116 to be focused by beamforming or otherwise modified.
[0123] Example Figure 6 A user input applied to a portion 600 of an off-screen graphic 300 is shown, where there is no marker and thus no audio source at that portion 600. It should be understood that the user input may have been applied to a video image region that has no associated audio source. In response to such user input, device 100 can be configured to provide a display of a second marker 601 to indicate the absence of an audio source in a direction corresponding to the selected region of the video image or the portion of the off-screen graphic 300. The color or pattern or appearance of the second marker 601 can be different from that of markers 215, 216.
[0124] Figure 7 A flowchart illustrating the following steps is shown: receiving 701 spatial audio data that includes audio captured from one or more audio sources in a space extending around a capture device and direction information indicating at least the direction toward the one or more audio sources, where the spatial audio data is captured by the capture device;
[0125] Receiving 702 a video image captured by a camera of the capture device, the video image having a field of view, where the extent of the space in which the spatial audio data is captured is larger than the field of view;
[0126] For audio sources within the field of view, associating 703 each of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction toward the audio source, and for audio sources outside the field of view, associating 703 each of the one or more audio sources determined according to the direction information with a portion of an off-screen graphic corresponding to the direction toward the audio source, the off-screen graphic indicating the spatial extent of the space outside the field of view;
[0127] Provide the display of 704 video images together with out-of-view graphics on a display;
[0128] Receive 705 user input for selecting a region of a video image or a part of an out-of-view graphic;
[0129] Provide 706 control of at least one audio capture attribute of a selected audio source among the one or more audio sources, where the selected audio source among the one or more audio sources includes an audio source among the one or more audio sources associated with the region or part selected by the user input.
[0130] The method may be characterized by any of the features described above with respect to the apparatus.
[0131] Figure 8 Schematically shows a computer / processor-readable medium 800 of a provider according to an example. In this example, the computer / processor-readable medium is a disc such as a digital versatile disc (DVD) or a compact disc (CD). In some examples, the computer-readable medium may be any medium that has been programmed in a manner to perform the inventive functions. The computer program code may be distributed among multiple memories of the same type or among multiple different types of memories such as ROM, RAM, flash memory, hard disk, solid state, etc.
[0132] The user input may be a gesture, which includes one or more of tapping, swiping, sliding, pressing, holding, rotating gesture, static hovering gesture near the user interface of the device, moving hovering gesture near the device, bending at least a part of the device, squeezing at least a part of the device, multi-finger gesture, tilting the device or flipping the control device. Additionally, the gesture may be any free-space user gesture using the user's body, such as their arm, or a stylus or other element suitable for performing free-space user gestures.
[0133] The apparatus shown in the above example may be a portable electronic device, laptop computer, mobile phone, smart phone, tablet computer, personal digital assistant, digital camera, smart watch, smart glasses, pen computer, non-portable electronic device, desktop computer, display, smart TV, server, wearable device, virtual reality device, or a module / circuit system for one or more of the above.
[0134] Any device mentioned and / or other features of a specifically mentioned apparatus may be provided by means arranged such that they are configured to perform the desired operations only when enabled (e.g., powered on, etc.). In such cases, they may not have to load appropriate software into active memory in a non-enabled state (e.g., off state), but only load appropriate software in an enabled state (e.g., on state). The means may include hardware circuitry and / or firmware. The means may include software loaded onto a memory. Such software / computer program may be recorded on the same memory / processor / functional unit and / or one or more memories / processors / functional units.
[0135] In some examples, a specifically mentioned apparatus may be pre-programmed with appropriate software to perform the desired operations, and wherein the appropriate software may be enabled for use by a user downloading a "key", e.g., to unlock / enable the software and its associated functions. Advantages associated with such examples may include reducing the requirement for downloading data when the device requires additional functionality, and this is useful in examples where the device is considered to have sufficient capacity to store such pre-programmed software to implement functions that the user may not have enabled.
[0136] Any apparatus / circuitry / component / processor mentioned may have other functions in addition to the functions mentioned, and these functions may be performed by the same apparatus / circuitry / component / processor. One or more of the disclosed aspects may include the electronic distribution of relevant computer programs and computer programs (which may be source / transmission encoded) recorded on an appropriate carrier (e.g., memory, signal).
[0137] Any "computer" described herein may include a collection of one or more individual processors / processing elements, which may or may not be located on the same circuit board or in the same area / location of a circuit board or even the same device. In some examples, one or more of any mentioned processors may be distributed across multiple devices. The same or different processors / processing elements may perform one or more of the functions described herein.
[0138] The term "signaling" may refer to one or more signals transmitted as a series of transmitted and / or received electrical / optical signals. The series of signals may include one, two, three, four or even more individual signal components or different signals that make up the said signaling. Some or all of these individual signals may be transmitted / received simultaneously, sequentially and / or in such a way that they overlap with each other in time, either wirelessly or by wire communication.
[0139] Refer to any discussion of any computers and / or processors and memories mentioned (e.g., including ROM, CD-ROM, etc.), which may include computer processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and / or other hardware components that have been programmed in a manner enabling the functions of the present invention to be performed.
[0140] The applicant hereby separately discloses each individual feature described herein and any combination of two or more such features, such that such features or combinations can be implemented based on the present specification as a whole in view of the general knowledge of those skilled in the art, regardless of whether such features or combinations of features solve any of the problems disclosed herein, and without limiting the scope of the claims. The applicant points out that the disclosed aspects / examples can consist of any such individual feature or combination of features. In view of the foregoing description, it will be apparent to those skilled in the art that various modifications can be made within the scope of the present disclosure.
[0141] While the basic novel features have been shown and described and pointed out as applied to examples thereof, it should be understood that various omissions and substitutions and changes in the form and detail of the described apparatus and methods can be made by those skilled in the art without departing from the scope of the present disclosure. For example, all combinations of those elements and / or method steps that perform substantially the same function in substantially the same way to achieve the same result are within the scope of the present disclosure. In addition, it should be recognized that the structures and / or elements and / or method steps shown and / or described in connection with any disclosed form or example can be incorporated, as a general design matter, into any other disclosed or described or suggested form or example. Further, in the claims, means-plus-function clauses are intended to cover the structures described herein as performing the recited function, and not only structural equivalents but also equivalent structures. Thus, while a nail and a screw may not be structural equivalents in that a nail uses a cylindrical surface to fasten wooden parts together while a screw uses a helical surface, the nail and the screw can be equivalent structures in the environment of fastening wooden parts.
Claims
1. An apparatus for processing spatial audio, comprising components configured to perform the following operations: Receive spatial audio data, the spatial audio data including audio captured from one or more audio sources in a space extending around a capture device and direction information indicating at least a direction towards the one or more audio sources, wherein the spatial audio data is captured by the capture device; Receive video images captured by a camera of the capture device, the video images having a field of view, wherein a range of the space in which the spatial audio data is captured is larger than the field of view; For an audio source within the field of view, associate each audio source of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction towards the audio source, and for an audio source outside the field of view, associate each audio source of the one or more audio sources determined according to the direction information with a part of an out-of-view graphic corresponding to the direction towards the audio source, the out-of-view graphic indicating a spatial extent of the space outside the field of view, wherein the out-of-view graphic includes a line, wherein a positioning along the line from one end to the other end represents a direction in which the audio of the audio source associated with the out-of-view graphic is received, wherein the one end of the line represents at least a direction corresponding to a first edge of the field of view, and the other end of the line represents at least a direction corresponding to a second edge of the field of view opposite to the first edge; Provide a display on a display of the video images together with the out-of-view graphic; Receive a user input for selecting a part of the out-of-view graphic; Provide control of at least one audio capture attribute of a selected one of the one or more audio sources, wherein the selected audio source of the one or more audio sources includes one of the one or more audio sources associated with the part selected by the user input.
2. The apparatus according to claim 1, wherein the components are configured to provide a display of a marker at one or more of the following: One or more parts of the out-of-view graphic corresponding to one or more of the directions towards the one or more audio sources; and One or more regions of the video images corresponding to one or more of the directions towards the one or more audio sources.
3. The apparatus according to claim 1 or claim 2, wherein the control of at least one audio capture attribute Comprises: The components are configured to provide signaling to cause capture or recording of the selected one audio source using beamforming techniques.
4. The apparatus according to claim 1 or claim 2, wherein the control of at least one audio capture attribute includes the components being configured to perform at least one of the following operations: Capture or record the selected one audio source with a greater volume gain relative to a volume gain applied to other audio of the spatial audio data; Capture or record the selected one audio source with a higher quality relative to the quality of other audio applied to the spatial audio data; Or Capture or record the audio of the selected one audio source as an audio stream separated from other audio of the spatial audio data.
5. The apparatus according to claim 1 or claim 2, wherein the component is configured to determine the one or more audio sources by using the direction information to determine from which direction the audio having a volume higher than a predetermined threshold is received.
6. The apparatus according to claim 1 or claim 2, wherein the out-of-view graphic indicating the spatial range of the space outside the field of view includes a sector of an ellipse, and the sector of the ellipse is the line, wherein the positioning within the sector represents the direction from which the audio of the audio source associated with the out-of-view graphic is received, wherein the first part of the sector represents the direction corresponding to at least the first edge of the field of view, and the second part of the sector represents the direction corresponding to at least the second edge of the field of view opposite to the first edge.
7. The apparatus according to claim 1 or claim 2, wherein the out-of-view graphic indicating the spatial range of the space outside the field of view represents a plane around the capture device, wherein the positioning of the presented marker relative to the out-of-view graphic represents the azimuth direction from which the audio of the audio source is received, and wherein the positioning of the presented marker depicted at a distance above or below the out-of-view graphic corresponds to the elevation direction from which the audio of the audio source is received above or below the plane.
8. The apparatus according to claim 1 or claim 2, wherein the component is configured to: based on the user input including a tap that selects the region of the video image or the part of the out-of-view graphic at a position on the touch-sensitive input device, modify the audio capture attribute by applying beamforming technology to provide control over at least one audio capture attribute, and the beamforming technology focuses on the region of the space corresponding to the selected region or part.
9. The apparatus according to claim 1 or claim 2, wherein the component is configured to: based on the user input including a pinch gesture that selects the region of the video image or the part of the out-of-view graphic at a position on the touch-sensitive input device, modify the audio capture attribute by applying beamforming technology to provide control over at least one audio capture attribute, and the beamforming technology has an angle degree related to the size of the pinch gesture.
10. The apparatus according to claim 1 or claim 2, wherein the component is configured to: based on the received user input selecting a region of the video image or a part of the out-of-view graphic for which there is no associated audio source, provide a display of a second marker to indicate that there is no audio source in the direction corresponding to the selected region of the video image or the part of the out-of-view graphic.
11. The apparatus according to claim 3, wherein the beamforming technique includes at least one of the following: a delay-and-sum beamforming technique or a parametric spatial audio processing technique in which the audio of the selected audio source is enhanced.
12. The apparatus according to claim 1 or claim 2, wherein the component is configured to provide one or both of the rendering and recording of the spatial audio data using the selected audio source having controlled audio capture attributes.
13. An electronic device, comprising the apparatus according to any one of the preceding claims, a camera configured to capture the video image, a plurality of microphones configured to capture the spatial audio data, and a display for use by the apparatus to display the video image together with the off-view graphic.
14. A method of processing spatial audio, the method comprising: receiving spatial audio data, the spatial audio data including audio captured from one or more audio sources in a space extending around a capture device and direction information indicating at least a direction towards the one or more audio sources, wherein the spatial audio data is captured by the capture device; receiving a video image captured by a camera of the capture device, the video image having a field of view, wherein a range of the space in which the spatial audio data is captured is greater than the field of view; for an audio source within the field of view, associating each audio source of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction towards the audio source, and for an audio source outside the field of view, associating each audio source of the one or more audio sources determined according to the direction information with a part of an off-view graphic corresponding to the direction towards the audio source, the off-view graphic indicating a spatial extent of the space outside the field of view, wherein the off-view graphic includes a line, wherein a positioning along the line from one end to the other end represents a direction in which the audio of the audio source associated with the off-view graphic is received, wherein one end of the line represents at least a direction corresponding to a first edge of the field of view, and the other end of the line represents at least a direction corresponding to a second edge of the field of view opposite to the first edge; providing a display on a display of the video image together with the off-view graphic; receiving a user input for selecting a part of the off-view graphic; providing control of at least one audio capture attribute of a selected audio source of the one or more audio sources, wherein the selected audio source of the one or more audio sources includes an audio source of the one or more audio sources associated with the part selected by the user input.
15. A computer-readable medium including computer program code stored thereon, the computer-readable medium and the computer program code being configured to perform a method when run on at least one processor, the method comprising: Receive spatial audio data, the spatial audio data including audio captured from one or more audio sources in a space extending around a capture device and direction information indicating at least a direction towards the one or more audio sources, wherein the spatial audio data is captured by the capture device; Receive a video image captured by a camera of the capture device, the video image having a field of view, wherein a range of the space in which the spatial audio data is captured is greater than the field of view; For an audio source within the field of view, associate each of the one or more audio sources determined according to the direction information with a region of the video image corresponding to the direction towards the audio source, and for an audio source outside the field of view, associate each of the one or more audio sources determined according to the direction information with a part of an out-of-view graphic corresponding to the direction towards the audio source, the out-of-view graphic indicating a spatial extent of the space outside the field of view, wherein the out-of-view graphic includes a line, wherein a positioning along the line from one end to the other end represents a direction in which the audio of the audio source associated with the out-of-view graphic is received, wherein the one end of the line represents at least a direction corresponding to a first edge of the field of view, and the other end of the line represents at least a direction corresponding to a second edge of the field of view opposite to the first edge; Provide a display on a display of the video image together with the out-of-view graphic; Receive a user input for selecting a part of the out-of-view graphic; Provide control of at least one audio capture attribute of a selected one of the one or more audio sources, wherein the selected audio source among the one or more audio sources includes one of the one or more audio sources associated with the part selected by the user input.
Citation Information
Patent Citations
Audio processing apparatus
EP2824663A2
Visual Audio Processing Apparatus
US20160299738A1