Computing audio engine

CN122603322APending Publication Date: 2026-08-18DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202580010509.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-13
Filing Date
2025-01-14
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]在移动设备上捕获音视频内容期间,用户通常可以看到正在捕获的视频,但(例如,不使用一个或多个附加设备)无法轻易监测正在捕获的音频信号

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122603322A_ABST
    Figure CN122603322A_ABST
Patent Text Reader

Abstract

Some methods disclosed involve identifying two or more audio sources in an audio scene based at least in part on audio data and video data, estimating at least one audio characteristic of each of the two or more audio sources based on the audio data, and storing the audio data and video data received during a capture phase. Some methods involve controlling a display of a device to display an image corresponding to the video data and to display a graphical user interface (GUI) overlaid on the image prior to and during the capture phase, the GUI including an audio source image corresponding to the at least one audio characteristic and including one or more user input areas, receiving user input via the one or more user input areas prior to the capture phase, and causing the audio data received during the capture phase to be modified in accordance with the user input.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 621,960, filed January 17, 2024, and U.S. Provisional Application No. 63 / 744,455, filed January 13, 2025, each of which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure generally relates to audio capture and related user feedback and audio processing. Background Technology

[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims of this application and are not acknowledged as prior art by virtue of their inclusion in this section.

[0004] During the capture of audio and video content on mobile devices, users can typically see the video being captured, but (e.g., without the use of one or more additional devices) cannot easily monitor the audio signal being captured. While audio capture processes implemented on mobile devices (also known as “audio capture stacking” or “audio stacking”) are generally designed to maximize the quality of the captured audio according to some assumed artistic intent, mobile device users are often unaware of whether the techniques operating within the audio capture stack are satisfactory. Other issues will be discussed in the detailed implementation below. Improvements to methods, devices, and systems are desired. Summary of the Invention

[0005] Techniques for audio capture and for related user feedback and audio signal processing are described. In some example embodiments, the method may involve receiving audio data from a microphone system and video data from a camera system by a device's control system. In some examples, the audio data may be received from the device's microphone system, and the video data may be received from the device's camera system. Some methods may involve identifying two or more audio sources in an audio scene by a control system, at least in part, based on the audio data and video data. Some methods may involve estimating at least one audio characteristic of each of the two or more audio sources by a control system based on the audio data. Some methods may involve storing the audio and video data received during the capture phase by a control system. In some examples, this storage may involve storing unmodified audio data received during the capture phase.

[0006] Some methods may involve a control system controlling a device's display to show images corresponding to video data, and displaying a graphical user interface (GUI) overlaid on these images before and during the capture phase. In some examples, the GUI may include audio source images corresponding to at least one audio characteristic of each of two or more audio sources. In some examples, the GUI may include one or more user input areas for receiving user input. Some methods may involve a control system receiving user input via one or more user input areas before the capture phase. Some methods may involve a control system causing the audio data received during the capture phase to be modified based on user input.

[0007] Some methods may involve the control system receiving user input via a user input area after the capture phase has begun. Other methods may involve the control system modifying the audio data received during the duration of the capture phase based on the user input.

[0008] Some methods may involve classifying two or more audio sources into two or more audio source categories by a control system. In some examples, the GUI may include user input area portions corresponding to each of the two or more audio source categories. In some examples, the two or more audio source categories may include a background category and a foreground category. According to some examples, the two or more audio source categories may include at least one user input area configured to receive user input regarding a selected level or level ratio for each of the two or more audio source categories.

[0009] Some methods may involve creating a sound source inventory by a control system. According to some examples, the sound source inventory may include actual sound sources and potential sound sources. In some examples, classifying two or more audio sources into two or more audio source categories may be at least partially based on the sound source inventory.

[0010] Some approaches may involve determining one or more actionable feedback types related to the audio scene. In some examples, the GUI may be based in part on one or more actionable feedback types.

[0011] Some methods may involve storing unmodified audio data received during the capture phase. Other methods may involve storing modified audio data that has been modified based on user input.

[0012] Some methods may involve a control system creating and storing user input metadata corresponding to user input received via a user input area. Some methods may involve modifying audio data received during the capture phase based on user input, involving post-capture audio processing at least in part based on the user input metadata. In some examples, the control system may be configured to perform at least a portion of the post-capture audio processing. According to some examples, another control system (such as a cloud-based service control system, e.g., a server control system) may be configured to perform at least a portion of the post-capture audio processing. In some examples, identification may involve the control system performing a first sound source separation process. In some such examples, post-capture audio processing may involve performing a second sound source separation process.

[0013] In some examples, identification may involve the control system detecting one or more potential sound sources, at least in part, based on video data. According to some examples, at least one of the one or more potential sound sources may not be indicated by audio data.

[0014] Some methods may involve a control system detecting one or more candidate sound sources for enhanced audio capture. In some examples, enhanced audio capture may involve replacing candidate sound sources with external or synthesized audio. According to some examples, the GUI may include at least one user input area configured to receive user selections for a selected potential sound source or a selected candidate sound source. In some examples, the GUI may include at least one user input area configured to receive user selections for enhanced audio capture. According to some examples, enhanced audio capture may include at least one of external or synthesized audio from a selected potential sound source or a selected candidate sound source. In some examples, the GUI may include at least one user input area configured to receive user selections for a ratio between enhanced audio capture and real-world audio capture.

[0015] According to some examples, modifying the audio data received during the capture phase based on user input can involve modifying the audio data corresponding to a selected audio source or a selected audio source category. In some examples, modifying the audio data received during the capture phase based on user input can involve a beamforming process corresponding to a selected audio source.

[0016] Other example embodiments describe an apparatus. In some example embodiments, the apparatus may include an interface system comprising an input / output (I / O) system. According to some implementations, the apparatus may include a control system comprising one or more processors. In some example embodiments, the apparatus may include a display system and a touch sensor system, the display system comprising one or more displays, the touch sensor system being proximate to at least one of the one or more displays. According to some example embodiments, the control system may be configured to receive audio data from a microphone system and video data from a camera system via the interface system. In some example embodiments, the control system may be configured to identify two or more audio sources in an audio scene, at least partially based on the audio and video data. According to some example embodiments, the control system may be configured to estimate at least one audio characteristic of each of the two or more audio sources. In some example embodiments, the control system may be configured to store the audio and video data received during a capture phase. According to some example embodiments, the control system may be configured to control the display of the display system to display an image corresponding to the video data and to display a graphical user interface (GUI) overlaid on these images before and during the capture phase. In some examples, the GUI may include an audio source image corresponding to at least one audio characteristic of each of the two or more audio sources. In some examples, the GUI may include one or more user input areas for receiving user input via a touch sensor system. According to some example embodiments, the control system may be configured to receive user input via one or more user input areas before the capture phase through an interface system. In some example embodiments, the control system may be configured to modify audio data received during the capture phase based on user input.

[0017] In some example embodiments, the control system may be configured to receive user input via a user input area after the start of the capture phase, and to modify the audio data received during the duration of the capture phase according to the user input.

[0018] According to some example embodiments, the control system can be configured to classify two or more audio sources into two or more audio source categories. In some examples, the GUI may include user input area portions corresponding to each of the two or more audio source categories. According to some examples, the two or more audio source categories may include a background category and a foreground category. In some examples, one or more user input areas may include at least one user input area configured to receive user input regarding a selected level or level ratio for each of the two or more audio source categories.

[0019] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory computer-readable media, which control one or more devices to perform one or more methods. Such non-transitory computer-readable media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc.

[0020] Some such methods may involve the control system of a device receiving audio data from a microphone system and video data from a camera system. In some examples, the audio data may be received from the device's microphone system, and the video data may be received from the device's camera system. Some methods may involve the control system identifying two or more audio sources in an audio scene, at least in part, based on the audio and video data. Some methods may involve the control system estimating at least one audio characteristic of each of the two or more audio sources based on the audio data. Some methods may involve the control system storing the audio and video data received during the capture phase. In some examples, this storage may involve storing unmodified audio data received during the capture phase.

[0021] Some methods may involve a control system controlling a device's display to show images corresponding to video data, and displaying a graphical user interface (GUI) overlaid on these images before and during the capture phase. In some examples, the GUI may include audio source images corresponding to at least one audio characteristic of each of two or more audio sources. In some examples, the GUI may include one or more user input areas for receiving user input. Some methods may involve a control system receiving user input via one or more user input areas before the capture phase. Some methods may involve a control system causing the audio data received during the capture phase to be modified based on user input.

[0022] Some methods may involve the control system receiving user input via a user input area after the capture phase has begun. Other methods may involve the control system modifying the audio data received during the duration of the capture phase based on the user input.

[0023] Some methods may involve classifying two or more audio sources into two or more audio source categories by a control system. In some examples, the GUI may include user input area portions corresponding to each of the two or more audio source categories. In some examples, the two or more audio source categories may include a background category and a foreground category. According to some examples, the two or more audio source categories may include at least one user input area configured to receive user input regarding a selected level or level ratio for each of the two or more audio source categories.

[0024] Some methods may involve creating a sound source inventory by a control system. According to some examples, the sound source inventory may include actual sound sources and potential sound sources. In some examples, classifying two or more audio sources into two or more audio source categories may be at least partially based on the sound source inventory.

[0025] Some approaches may involve determining one or more actionable feedback types related to the audio scene. In some examples, the GUI may be based in part on one or more actionable feedback types.

[0026] Some methods may involve storing unmodified audio data received during the capture phase. Other methods may involve storing modified audio data that has been modified based on user input.

[0027] Some methods may involve a control system creating and storing user input metadata corresponding to user input received via a user input area. Some methods may involve post-capture audio processing, at least in part based on the user input, modifying audio data received during the capture phase according to the user input. In some examples, the control system may be configured to perform at least a portion of the post-capture audio processing. According to some examples, another control system (such as a cloud-based service control system, e.g., a server control system) may be configured to perform at least a portion of the post-capture audio processing. In some examples, identification may involve the control system performing a first sound source separation process. In some such examples, post-capture audio processing may involve performing a second sound source separation process.

[0028] In some examples, identification may involve the control system detecting one or more potential sound sources, at least in part, based on video data. According to some examples, at least one of the one or more potential sound sources may not be indicated by audio data.

[0029] Some methods may involve a control system detecting one or more candidate sound sources for enhanced audio capture. In some examples, enhanced audio capture may involve replacing candidate sound sources with external or synthesized audio. According to some examples, the GUI may include at least one user input area configured to receive user selections for a selected potential sound source or a selected candidate sound source. In some examples, the GUI may include at least one user input area configured to receive user selections for enhanced audio capture. According to some examples, enhanced audio capture may include at least one of external or synthesized audio from a selected potential sound source or a selected candidate sound source. In some examples, the GUI may include at least one user input area configured to receive user selections for a ratio between enhanced audio capture and real-world audio capture.

[0030] According to some examples, modifying the audio data received during the capture phase based on user input can involve modifying the audio data corresponding to a selected audio source or a selected audio source category. In some examples, modifying the audio data received during the capture phase based on user input can involve a beamforming process corresponding to a selected audio source.

[0031] The embodiments described herein can generally be described as technology, wherein the term “technology” can refer to a system, apparatus, method, computer-readable instructions, module, component, hardware logic and / or operation as suggested in the context to which it is applied herein.

[0032] Features and technical benefits beyond those explicitly described above will become apparent upon reading the following detailed description and consulting the accompanying drawings. This summary is provided to illustrate the choice of techniques in a simplified form and is not intended to identify key or essential features of the claimed subject matter as defined by the appended claims. Attached Figure Description

[0033] Figure 1A A schematic block diagram of an example device architecture that can be used to implement various aspects of this disclosure is shown.

[0034] Figure 1B The illustration shows that Figure 1A A schematic block diagram of an example central processing unit (CPU) implemented in a device architecture that can be used to implement various aspects of this disclosure.

[0035] Figure 1C Examples of timelines and time intervals that may occur for some of the disclosed processes are shown.

[0036] Figure 1DThis is a flowchart outlining various example methods according to some publicly available implementations.

[0037] Figure 2A Examples of timelines and time intervals are shown when some of the disclosed processes may occur.

[0038] Figure 2B This is a flowchart outlining various example methods according to some publicly available implementations.

[0039] Figure 3A The illustration shows an example of a GUI that can be presented to indicate one or more audio scene preferences and one or more current audio scene characteristics.

[0040] Figure 3B This is a block diagram illustrating an example of a custom layer that can be rendered via one or more GUIs.

[0041] Figure 4A The diagram shows that... Figure 3B An example of a GUI rendered in the first custom layer.

[0042] Figure 4B The diagram shows that... Figure 3B Another example of a GUI rendered by the first custom layer.

[0043] Figure 4C The diagram shows that... Figure 3B An example of a GUI rendered by a second custom layer.

[0044] Figure 4D and Figure 4E The diagram shows that... Figure 3B Example elements of the GUI rendered by the third or fourth custom layer.

[0045] Figure 5A This is a flowchart outlining various example methods according to some publicly available implementations.

[0046] Figure 5B This is a flowchart outlining various example methods according to some publicly available implementations.

[0047] Figure 5C A table is shown representing example elements of a video object data structure according to some publicly disclosed implementations.

[0048] Figure 5D A table showing example elements of an audio source manifest data structure according to some publicly available implementations is shown.

[0049] Figure 6 Examples of modules that can be implemented based on some publicly available examples are shown.

[0050] Figure 7 Example components of an analysis module configured for contextual hypothesis evaluation are shown, based on some publicly available examples.

[0051] Figure 8 This is a flowchart outlining various example methods according to some publicly available implementations.

[0052] Figure 9 This is a flowchart outlining various example methods according to some publicly available implementations.

[0053] Figure 10A An element representing a list of audio sources according to some publicly available implementations.

[0054] Figure 10B This is a flowchart outlining various example methods according to some publicly available implementations.

[0055] Figure 10C This is a flowchart outlining additional example methods according to some disclosed implementations.

[0056] Figure 11A Examples of media assets and hybrid modules are shown, based on some publicly available examples.

[0057] Figure 11B This is a flowchart outlining various example methods 1150 according to some disclosed implementations.

[0058] Figure 12A Examples of media assets and interpolators are shown based on some publicly available examples.

[0059] Figure 12B , Figure 12C and Figure 12D The illustration shows sample elements of the GUI that can be rendered during post-editing processes that include interpolation.

[0060] Figure 12E This is a flowchart outlining various example methods according to some publicly available implementations.

[0061] Figure 13 Examples of media assets and interpolators are shown based on some publicly available examples.

[0062] Figure 14 This is a flowchart outlining various example methods according to some publicly available implementations.

[0063] Figure 15 This is a block diagram of an example IVAS encoder / decoder (“coder / decoder”) framework for encoding and decoding Immersive Speech and Audio Services (IVAS) bitstreams according to one or more embodiments.

[0064] In the accompanying drawings, for ease of description, a specific arrangement or order of schematic elements is shown, such as those representing devices, units, instruction blocks, and data elements. However, those skilled in the art will understand that the specific order or arrangement of schematic elements in the drawings does not imply a requirement for a particular processing sequence or order, or process separation. Furthermore, the inclusion of schematic elements in the drawings does not imply that such elements are required in all embodiments, or that in some embodiments, features represented by such elements may not be included in or combined with other elements.

[0065] Furthermore, in the accompanying drawings, where connecting elements such as solid or dashed lines or arrows are used to illustrate connections, relationships, or associations between two or more other schematic elements, the absence of any such connecting element does not imply the impossibility of such connections, relationships, or associations. In other words, some connections, relationships, or associations between elements are not shown in the drawings to avoid obscuring this disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, where a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such an element represents one or more signal paths that may be necessary to achieve the communication.

[0066] The same reference numerals used in the various figures indicate similar elements. Detailed Implementation

[0067] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the various embodiments described with reference to the accompanying drawings. The illustrative embodiments in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of this disclosure. In view of this disclosure, it will be apparent to those skilled in the art that the various features and implementations described can be practiced without many of the details in these specific details. In some instances, well-known methods, processes, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Several features are described below, each of which can be used independently of each other or in any combination with other features. Therefore, these features can be arranged, replaced, combined, separated, or designed into other configurations as contemplated in light of this disclosure.

[0068] Terminology Explanation As used herein, the term “comprising” and variations thereof should be understood as open-ended terms meaning “including but not limited to”. Unless the context explicitly states otherwise, the term “or” should be understood as “and / or”. The term “based on” should be understood as “at least partially based on”. The terms “one example implementation” and “an example implementation” should be understood as “at least one example implementation”. The term “another implementation” should be understood as “at least one other implementation”. The term “determine” should be understood as obtaining, receiving, calculating, estimating, predicting, or obtaining. Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0069] definition An audio bed is a mono or multi-channel audio waveform associated with a specific channel representation of content.

[0070] An audio object is a mono or stereo audio waveform associated with some location metadata, which facilitates rendering of the audio object for any configuration of the playback system.

[0071] A microphone feed is a signal captured or being captured by a microphone or microphone array connected to a mobile device performing capture. A “microphone feed” can also be handled by enhancement tools. A microphone feed can be created at any layer of the audio capture stack implemented on a mobile device. In other words, the term “microphone feed” is used to refer to any signal (mono, multi-channel, sound field, etc.) that is available for use with the disclosed computational capture system.

[0072] A video feed is a signal captured or being captured by one or more cameras on a mobile device. A video feed is any video signal that can be used with the disclosed computational capture system.

[0073] A standard media container is an interchangeable media format supported by a software ecosystem associated with mobile devices.

[0074] acronym AR - Augmented Reality ASIC — Application-Specific Integrated Circuit BS - Bitstream CD-ROM—Compact Disc Read-Only Memory CNN - Convolutional Neural Network 3D-CNN — 3D CNN CNN-RNN—CNN Recurrent Neural Network CPU — Central Processing Unit DIRAC - Directional Audio Coding DSP—Digital Signal Processor EPROM—Erasable Programmable Read-Only Memory EVS – Enhanced Voice Service FOA - First-order high-fidelity stereo FPGA - Field Programmable Gate Array GUI — Graphical User Interface HOA - High-end High-fidelity stereo I / O — Input / Output IVAS – Immersive Voice and Audio Services MD - Metadata MP4 - Moving Image Experts Group (MPEG)-4 Part 14 MViTv2 - Multi-scale Visual Transformer PCM—Pulse Code Modulation QMF – Quadrature Mirror Filter RAM—Random Access Memory ROM - Read-Only Memory SPAR – Space Reconstruction SNR – Signal-to-Noise Ratio VR - Virtual Reality VSC - Video Scene Classification YAMNet—yet another multi-scale convolutional neural network This disclosure describes various audio capture methods, devices, and systems configured to operate in the context of audio and video capture. Some embodiments presented in this disclosure describe the use of mobile devices, such as mobile phones, to perform mobile audio and video capture, the mobile devices being equipped with a built-in microphone or microphone array and a camera.

[0075] As described above, during the capture of audio and video content on a mobile device, the user can typically see the video being captured, but (e.g., without the use of one or more additional devices) cannot easily monitor the audio signal being captured. Some embodiments presented in this disclosure describe a situation where mobile audio capture is performed on a mobile phone equipped with a built-in microphone or microphone array, and the user performing the audio capture cannot monitor the audio material being recorded, for example, by playing it through headphones. Although some embodiments will be described in this context, this disclosure is not limited to this field of use and can be applied to a broader context.

[0076] While the audio capture process implemented on mobile devices (also known as "audio capture stacking" or "audio stacking") is typically designed to maximize the quality of captured audio according to some assumed artistic intent, mobile device users are often unaware of whether the techniques operating within the audio capture stack are satisfactory. This can lead to scenarios where users might abandon audio capture, for example, under challenging audio conditions, because they don't know if any usable audio can be captured under those circumstances. Furthermore, users may be unaware of solvable audio problems that arise during audio capture that require corrective action. For example, audio capture performance can often be improved by moving closer to the audio source, changing how the mobile device is held, etc. Additionally, audio stacks can include powerful audio enhancement tools, but these tools typically have limitations in their performance (e.g., in terms of minimum signal-to-noise ratio (SNR)). The performance of these audio enhancement tools generally degrades as audio capture conditions worsen. However, users are often unaware of these limitations. Moreover, users are often unaware of how far their current operating point is from these limitations.

[0077] Some of the disclosed methods involve generating feedback on the composition of the audio scene, which adapts to the context, allows the user to address the aforementioned problems, and facilitates corrective actions from the user. In some examples, this feedback can be real-time adaptive, allowing the user to correlate the feedback with changes in the acoustic scene as perceived by the user. According to some examples, the feedback can be available at the start of the pre-capture phase and can continue throughout the actual capture phase. In some examples, the feedback can include graphical feedback on the composition of the audio scene superimposed on the currently acquired video. Some disclosed examples allow the user to introduce enhancements to the audio capture (before, during, or after capture). Some disclosed examples provide means for falling back from enhanced audio capture to real-world audio capture to interpolate between enhanced audio capture and real-world audio capture, or both.

[0078] To implement the desired user experience that can be provided according to some implementation methods, various technical problems may need to be solved. For example, it may be necessary to analyze the audio-visual scene (also referred to herein as the audio scene) to identify the acoustic components of the audio scene, and preferably identify one or more audio characteristics (such as sound level information) together. In some instances, both actual acoustic components and potential acoustic components can be determined. Audio scene analysis can be available at any time after the capture app has been launched and can be updated during audio capture. In some examples, the visualization of the audio scene can include not only the presence of actual or potential audio sources, but also estimates of the contribution of the respective audio sources to the audio scene. The estimate of the contribution of the audio sources to the audio scene can be obtained, for example, based on the estimated level of the respective audio sources. In some examples, information about the components of the audio scene can be filtered by estimating which audio sources are most likely to be relevant to the user's artistic intent.

[0079] Some aspects of this disclosure relate to providing real-time graphical feedback on the composition of an audio scene. In some such examples, the graphical feedback can be overlaid on a video image being acquired by a mobile device. In some examples, the graphical feedback can be provided via a graphical user interface (GUI) including a representation of one or more audio sources in the audio scene and one or more user input areas for receiving user input. In some examples, audio data received during the capture phase can be modified based on user input.

[0080] Some examples involve providing or facilitating one or more types of audio data enhancement that can occur before, during, or after the capture process. Audio data enhancement can involve replacing candidate sound sources with external or synthesized audio, beamforming processes, or both. Some disclosed examples may facilitate inserting audio signals not present in the microphone feed, removing one or more components of the audio scene, etc. According to some examples, the control system can select one or more candidate sound sources for modification, to perform enhanced audio capture, or both. In some such examples, the user can select an audio source for modification, enhancement, or both via a GUI.

[0081] Some of the disclosed examples provide apparatus for falling back from enhanced capture to “real-world” or unmodified audio capture. Some of these examples involve storing unmodified audio data received during the capture phase, as well as enhanced audio data. According to some examples, a GUI (which may be provided during a post-capture editing process) may include at least one user input area configured to receive a user selection of the ratio between the enhanced audio capture and the real-world audio capture. The user input area may include a virtual slider, a virtual dial, etc.

[0082] Multimodal analysis of video and audio can be used in the context of audio source identification by leveraging the correlation between objects present in the video feed of a capture system and objects present in the microphone feed. Typically, a set of detected audio objects can contain many elements, and its composition may be unstable or ambiguous. Some of the technical problems addressed by this disclosure involve generating context-relevant feedback on the composition of a scene, where only important audio sources are selected and their contribution levels are estimated, while other sources can, for example, be classified as background components. In other words, while users may easily be overwhelmed by the many details associated with the composition of an audio scene, some of the disclosed examples involve filtering the details of the audio scene and presenting information about one or more sound sources estimated to be the most relevant components of the audio scene.

[0083] In some instances, users can indicate one or more preferences regarding audio capture, such as the desired level of a particular type of source (e.g., speech, background sound, or a component not present in the scene but integrated into the capture), by interacting with graphical feedback (e.g., by interacting with a GUI). For example, a user can focus on a specific sound source, suppress a specific source, or adjust a specific source. By acquiring and recording this user input, some of the disclosed examples provide apparatus for estimating and implementing a user's artistic intent.

[0084] Figure 1AA schematic block diagram of an example device architecture 101 (in this example, apparatus 101) that can be used to implement various aspects of the present disclosure is illustrated. Architecture 101 includes, but is not limited to, server and client devices, systems, etc., which can be configured to perform any or all of the methods described in the accompanying drawings of the reference disclosure. As shown, architecture 101 includes a central processing unit (CPU) 141 capable of executing various processes according to a program stored, for example, in read-only memory (ROM) 142 or loaded from, for example, storage unit 148 into random access memory (RAM) 143. CPU 141 may be, for example, an electronic processor 141. CPU 141 is an example of the contents of an element of a control system, which may be referred to herein as a "control system". A control system may, for example, include a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. ROM 142 and RAM 143 are examples of the contents of an element of a memory system, which may be referred to herein as a "memory system". RAM 143 also stores data required by CPU 141 when performing various processes. In this example, CPU 141, ROM 142, and RAM 143 are interconnected via bus 144. Input / output (I / O) interface 145 is also connected to bus 144. Bus 144 and I / O interface 145 are examples of elements of the term "interface system" as used in this disclosure.

[0085] According to this example, the following components are connected to I / O interface 145: input unit 146, which may include a keyboard, mouse, etc.; output unit 147, which may include a display system including one or more displays, a loudspeaker system including one or more loudspeakers, etc.; storage unit 148, which includes a hard disk or other suitable storage device; and communication unit 149, which includes a network interface card such as a network card (e.g., wired or wireless). Communication unit 149 may be referred to herein as part of the interface system.

[0086] In some implementations, the input unit 146 may include a microphone system comprising one or more microphones. In some examples, the microphone system may include two, three, or more microphones located at different locations, which enable the capture of audio signals in various formats, such as mono, stereo, spatial, immersive, and other suitable formats.

[0087] According to some implementations, output unit 147 may include a system with a variety of amplifiers. Output unit 147 (depending on the capabilities of the host device) may be able to render audio signals in various formats, such as mono, stereo, immersive, binaural, and other suitable formats.

[0088] In some embodiments, communication unit 149 is configured to communicate with other devices (e.g., via a network). Drive 150 is also connected to I / O interface 145 as needed. Removable media 151 (such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media) is mounted on drive 150 such that computer programs read therefrom can be installed into storage unit 148 as needed. Those skilled in the art will understand that while apparatus 101 is described as including the components described above, in practice, some of these components may be added, removed, and / or replaced, and all such modifications or alterations fall within the scope of this disclosure.

[0089] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 149, and / or from, for example... Figure 1A The removable medium 151 shown is installed.

[0090] Figure 1AA schematic block diagram of an example device architecture 101 (in this example, apparatus 101) that can be used to implement various aspects of the present disclosure is illustrated. Architecture 101 includes, but is not limited to, server and client devices, systems, etc., which can be configured to perform any or all of the methods described in the accompanying drawings of the reference disclosure. In some examples, apparatus 101 may be a mobile display device, such as a cellular phone. As shown, architecture 101 includes a central processing unit (CPU) 141 capable of executing various processes according to a program stored, for example, in read-only memory (ROM) 142 or loaded from, for example, storage unit 148 into random access memory (RAM) 143. CPU 141 may be, for example, an electronic processor 141. CPU 141 is an example of the contents of an element that may be referred to as a "control system" or a control system. ROM 142 and RAM 143 are examples of the contents of an element that may be referred to as a "memory system" or a memory system. In RAM 143, data required by CPU 141 when executing various processes is also stored as needed. In this example, CPU 141, ROM 142, and RAM 143 are interconnected via bus 144. Input / output (I / O) interface 145 is also connected to bus 144. Bus 144 and I / O interface 145 are examples of elements of the term "interface system" as used in this disclosure.

[0091] According to this example, the following components are connected to I / O interface 145: input unit 146, which may include a keyboard, mouse, etc.; output unit 147, which may include a display system including one or more displays, a loudspeaker system including one or more loudspeakers, etc.; storage unit 148, which includes a hard disk or other suitable storage device; and communication unit 149, which includes a network interface card such as a network card (e.g., wired or wireless). Communication unit 149 may be referred to herein as part of the interface system.

[0092] In some implementations, the input unit 146 may include a microphone system comprising one or more microphones. In some examples, the microphone system may include two, three, or more microphones located at different locations, which enable the capture of audio signals in various formats, such as mono, stereo, spatial, immersive, and other suitable formats.

[0093] According to some implementations, output unit 147 may include a system with a variety of amplifiers. Output unit 147 (depending on the capabilities of the host device) may be able to render audio signals in various formats, such as mono, stereo, immersive, binaural, and other suitable formats.

[0094] In some embodiments, communication unit 149 is configured to communicate with other devices (e.g., via a network). Drive 150 is also connected to I / O interface 145 as needed. Removable media 151 (such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media) is mounted on drive 150 such that computer programs read therefrom can be installed into storage unit 148 as needed. Those skilled in the art will understand that while apparatus 101 is described as including the components described above, in practice, some of these components may be added, removed, and / or replaced, and all such modifications or alterations fall within the scope of this disclosure.

[0095] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 149, and / or from, for example... Figure 1A The removable medium 151 shown is installed.

[0096] According to some examples, CPU 141 may be a control system configured to perform some or all of the methods disclosed herein, or may be part of such a control system. In some examples, the control system may be configured to receive audio data from a microphone system and video data from a camera system via an interface system. According to some examples, the control system may be configured to identify two or more audio sources in an audio scene, at least in part, based on the audio and video data, and to estimate at least one audio characteristic of each of the two or more audio sources. In some examples, the control system may be configured to store the audio and video data received during the capture phase. According to some examples, the control system may be configured to control the display of a display system to display an image corresponding to the video data, and to display a graphical user interface (GUI) overlaid on these images before and during the capture phase. In some examples, the GUI may include an audio source image corresponding to at least one audio characteristic of each of the two or more audio sources and one or more user input areas for receiving user input via a touch sensor system. According to some examples, the control system may be configured to receive user input via one or more user input areas before the capture phase through the interface system, and to modify the audio data received during the capture phase based on the user input.

[0097] Figure 1B The illustration shows that Figure 1AA schematic block diagram of an example CPU 141 implemented in device architecture 101 that can be used to implement various aspects of this disclosure. CPU 141 includes an electronic processor 160 and a memory 161. The electronic processor 160 is electrically and / or communicatively connected to the memory 161 for bidirectional communication. The memory 161 stores encoding software 162 and decoding software 163. The memory 161 may be, for example, ROM, RAM, or another non-transitory computer-readable medium. The electronic processor 160 may implement the encoding software 162 stored in the memory 161 to perform one, some, or all of the disclosed methods. Additionally, the electronic processor 160 may implement the decoding software 163 stored in the memory 161 to perform one, some, or all of the disclosed methods.

[0098] Typically, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the unit discussed above can be implemented by control circuitry (e.g., CPU 141 and...). Figure 1A The control circuitry can perform the actions described in this disclosure, as executed by other components. Some aspects may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). Although various aspects of exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, other computing devices, or some combination thereof, as non-limiting examples.

[0099] Furthermore, the various blocks shown in the flowchart can be viewed as method steps, and / or operations resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform associated functions. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code configured to perform the methods described above.

[0100] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0101] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that when executed by the processor of the computer or the processor of the other programmable data processing apparatus, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0102] Figure 1C Examples of timelines and time intervals that may occur during some of the disclosed processes are shown. These processes may, for example, be at least in part caused by Figure 1A The device 101 or a similar device performs the operation. In these examples, timeline 102 indicates the capture application start time 104, capture start time 106, capture end time 108, post-capture editing start time 110, and post-capture editing end time 112. According to these examples, timeline 102 breaks between capture end time 108 and post-capture editing start time 110 to indicate that the time interval can be variable.

[0103] The capture application launch time 104 indicates when one or more software applications related to audio and video capture launch, while the capture start time 106 indicates when the actual audio and video capture process begins. In other words, the capture start time 106 indicates when actual audio and video recording begins. In these examples, the time interval between the capture application launch time 104 and the capture start time 106 is referred to as the pre-capture phase 114, the time interval between the capture start time 106 and the capture end time 108 is referred to as the capture phase 116, and the time interval between the post-capture editing start time 110 and the post-capture editing end time 112 is referred to as the post-capture phase 118, during which the post-capture editing process 124 may occur. The post-capture editing process 124 may be performed at least partially by the device used during the capture phase 116. However, in some examples, the post-capture editing process 124 may be performed at least partially by one or more other devices, such as by one or more servers of a cloud-based audio processing service.

[0104] According to some examples, analysis of the composition of an audio scene can begin during the pre-capture phase 114. The analysis may involve, for example, identifying audio sources, classifying and labeling audio sources, associating audio sources with components of the video scene (e.g., active speakers present in the video scene), associating audio sources with components not present in the video scene (e.g., the person performing the capture), etc. During the audio scene analysis, audio sources may be associated with the spatial locations of objects present in the video (e.g., the locations of active speakers present in the video scene). The analysis may involve estimating the level of each audio source. At least some results of the audio scene analysis may be provided to components or modules that are configured or generate user interfaces, such as graphical user interfaces (GUIs) including graphical feedback on the composition of the audio scene. Some examples of such GUIs are provided in this disclosure.

[0105] Accordingly, during the pre-capture phase 114 and the capture phase 116, the user can interact with a GUI that provides feedback on the current composition of the audio scene. Various types of user interaction are possible. One type involves altering the audio scene through interaction with the audio scene itself (e.g., by approaching or moving away from the audio source). Another type involves interaction with the GUI. Examples of user interaction with the GUI may include adjusting the audio levels between the acoustic background and acoustic foreground, adjusting the audio levels of one or more speakers (e.g., by touching a displayed audio source indicator associated with a speaker), etc. In some examples, the user may interact with the GUI to start or end an instance of capture phase 116, for example, by touching a virtual recording button presented by the capture application, to begin or end the capture of a specific video segment and associated audio data.

[0106] exist Figure 1C In the diagram, arrow 120 corresponds to a time interval during which the process 121 for registering user input can be performed, and arrow 122 corresponds to a time interval during which the process 123 for providing feedback on the composition of the audio scene can be performed. In some examples, the user input used for the process 121 for registering user input can be obtained via touch sensor signals corresponding to one or more user input areas of the GUI.

[0107] According to some examples, the process 123 of providing feedback on the composition of the audio scene may also involve a GUI. The GUI may, for example, be overlaid on an image corresponding to video and audio data obtained by device 101 or a similar device. The image corresponding to the audio data may, for example, include shapes, outlines, etc., corresponding to one or more audio sources. In some examples, the GUI may include an image of an audio source corresponding to at least one audio characteristic (e.g., volume or level) of each audio source. In some such examples, user input may indicate that the user desires to modify at least one audio characteristic, for example, the user desires to increase or decrease the volume of one or more audio sources. Some disclosed methods may involve, for example, a control system causing the audio data received during the capture phase 116 to be modified based on at least some of the user input received via the GUI. Depending on the implementation and the type of modification, audio data modification may or may not occur during the capture phase 116. Some such methods may involve storing the modified audio data. Some disclosed methods may involve creating and storing user input metadata corresponding to user input received via a user input area.

[0108] exist Figure 1CIn the example shown, it can be observed that during some or all of the pre-capture phase 114 and during some or all of the capture phase 116, a process 121 for registering user input can potentially be performed, and a process 123 for providing feedback on the composition of the audio scene can potentially be provided. Accordingly, in some examples, the GUI can be displayed before capture phase 116 (during pre-capture phase 114) and during capture phase 116. Such examples offer several potential advantages. For example, before the start of capture phase 116, the user can be able to evaluate information about the audio scene (including, but not limited to, level information corresponding to one or more audio sources) and provide user input about that audio scene information. During capture phase 116, the user can be able to focus relatively more attention on other aspects of capture phase 116, such as aspects of video capture. Furthermore, in some implementations, information corresponding to user input obtained before and during capture phase 116 can be “registered,” for example, stored as user input metadata. In some such examples, post-capture audio processing can be based at least in part on the registered user input, for example, as described below with reference to process 127.

[0109] Typically, it is not necessary to create the final audio result of the capture during capture phase 116. For example, some audio manipulations disclosed herein may require operating a high-quality audio source separation algorithm, which could be computationally limiting during capture phase 116. Instead, some of the disclosed examples involve using a relatively low-quality audio source separation method during capture phase 116 and a relatively high-quality audio source separation method during post-capture editing process 124. The relatively low-quality audio source separation method may have relatively low computational requirements and may be more suitable for providing real-time feedback during capture phase 116. In some examples, capture phase 116 may involve estimating the contribution of each audio component to the audio scene, presenting the estimated contribution of each audio component to the user performing the capture, and registering the user's intent based on user input, which may include some manipulation the user expects to be applied to the source (e.g., adjusting the source level, introducing a source not present in the microphone feed, etc.). Information corresponding to the user input (e.g., metadata) may be appended to the captured content, and a final version of the audio data may be generated during post-capture editing process 124 and may be based at least in part on the metadata. While the final version of the audio data can be obtained during the post-capture editing process 124, according to some examples, changes made to the audio scene during this process can also be undone if the user wishes. For example, the user can preview the automatic rendering of the scene based on user input registered during the pre-capture phase 114 and capture phase 116. During or after the preview, the user can decide to revert to the default capture or introduce further adjustments to the scene composition.

[0110] Accordingly, in Figure 1C In the example shown, the post-capture editing process 124 may involve implementing a process 125 to revert to a real-world capture. In some such examples, both modified and unmodified versions of the captured audio data may be stored. According to some examples, implementing the process 125 to revert to a real-world capture may involve, for example, allowing a user during the post-capture editing process 124 to select the degree of modification or unmodification of the final version of the audio corresponding to at least one sound source. In some such examples, a GUI may be presented to the user, comprising a virtual slider, dial, or other such virtual tool that the user can interact with to indicate the degree of audio modification in the final version of the audio.

[0111] exist Figure 1CIn the example shown, the post-capture editing process 124 may involve a process 127 that implements post-processing guided by registered user input. According to some examples, process 127 may involve referencing stored input metadata corresponding to user input received via a user input area during the pre-capture phase 114, during the capture phase, or both. In some examples, the post-capture editing process 124 may involve one or more video editing processes.

[0112] After the post-capture editing process 124 is completed, in this example, process 128 involves creating one or more multimedia files, including a final version of audio and video data. Alternatively or additionally, process 128 may involve creating and transmitting a bitstream corresponding to the final version of the audio and video data. According to some examples, the bitstream may be a bitstream encoded with Immersive Speech and Audio Services (IVAS). This document provides some examples of encoding and decoding IVAS bitstreams.

[0113] Figure 1D This is a flowchart outlining various example methods 160 according to some disclosed embodiments. Example method 160 can be divided into blocks, such as blocks 161, 163, 165, 166, 167, 168, 169, 170, 171, 173, and 175. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 160 need not be performed in the indicated order. In some embodiments, one or more blocks of method 160 can be performed simultaneously. Furthermore, some embodiments of method 160 may include more or fewer blocks than those shown and / or described. The blocks of method 160 can be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0114] In these examples, boxes 161, 163, 165, and 167 are executed during the pre-capture phase 114. According to these examples, boxes 168, 169, and 171 are executed during the capture phase 116. In these examples, boxes 173 and 175 are executed during the post-capture phase 118. The pre-capture phase 114, capture phase 116, and post-capture phase 118 can be, for example, as referenced herein. Figure 1C The functions described in boxes 163, 165, 166, 167, 168, 169, and 171 may also be referred to herein as the “audio interaction framework.” Processing can begin at box 161.

[0115] Box 161 relates to “capture application launch”. In some examples, the capture application may be a software application used to capture audio and video data. In some examples, the capture application may be stored in the memory system of a mobile device (such as a cellular phone). According to some examples, Box 161 may relate to, for example, initializing or launching the capture application in response to user input received by the device. Processing may continue to Box 163.

[0116] Box 163 relates to “analyzing an audio scene and generating a GUI with customizable options.” According to some examples, Box 163 may involve receiving and analyzing audio data from a microphone system. According to some examples, Box 163 may involve receiving and analyzing video data from a camera system. According to some examples, Box 163 may involve identifying one or more audio sources in an audio scene, at least in part, based on audio and video data. Audio sources may also be referred to herein as sound sources. According to some examples, analyzing an audio scene may involve creating a list of sound sources in the audio scene. The list of sound sources may include one or more actual sound sources and one or more potential sound sources. For example, potential sound sources may be identified based on video data even if no audio data from a potential sound source is detected or the audio data is below a threshold level. According to some examples, Box 163 may involve classifying two or more audio sources into two or more audio source categories, which may include a background category and a foreground category. According to some examples, Box 163 may involve presenting a GUI that includes a user input area portion corresponding to each of the two or more audio source categories. According to some examples, Box 163 may involve obtaining feedback options and customizable options. In some examples, box 163 may involve presenting at least one feedback option, at least one custom option, or both on the GUI. In some examples, box 163 may involve overlaying one or more images corresponding to at least one feedback option, at least one custom option, or both onto a displayed image corresponding to video data from a camera system. In some examples, the GUI includes one or more areas for receiving user input. Processing may continue to box 165.

[0117] Box 165 relates to determining whether user input is received from the GUI. According to some examples, box 165 may relate to determining whether user input is received from an area of ​​the touch sensor system corresponding to one or more areas of the GUI configured to receive user input. Processing may proceed to box 163 or box 166. If it is determined in box 165 that no user input has been received from the GUI, in this example, processing returns to box 163. If it is determined in box 165 that user input has been received from the GUI, in this example, processing proceeds to box 166.

[0118] Box 166 relates to determining whether a “capture start” event has occurred. In some examples, box 166 may relate to determining whether user input received via the aforementioned GUI corresponds to the initiation of capture phase 116, during which audio data received by the microphone system and video data received by the camera system are stored in memory. In some examples, method 160 may relate to storing unmodified audio data received during the capture phase, storing modified audio data received during the capture phase (which may have been modified based on user input), or both. Processing may proceed to box 168 or box 167. If it is determined in box 166 that user input from the GUI corresponds to the initiation of capture phase 116, in this example, the process proceeds to box 168. If it is determined in box 166 that user input from the GUI does not correspond to the initiation of capture phase 116, in this example, the process proceeds to box 167.

[0119] Box 167 relates to “Registering User Input”. In some examples, user input received via the GUI may respond to custom options provided by the GUI, such as increasing the volume of a sound source, decreasing the volume of a sound source, selecting a potential sound source, etc. According to some examples, registering user input may involve creating and storing metadata corresponding to the user input. Registered user input can be used to guide modifications to the audio data during capture phase 116, post-capture phase 118, or both. Processing can continue to box 163.

[0120] Box 168 relates to “analyzing the audio scene and generating a GUI with custom options.” According to some examples, Box 168 may simply be a continuation of an “audio interaction framework” that includes Box 163 and continues from Box 163 through the pre-capture stage 114 and the capture stage 116. Processing can continue to Box 169.

[0121] Box 169 relates to determining whether "user input was received from [the] GUI". In this example, the GUI is rendered during capture phase 116. Processing may proceed to either box 168 or box 170. If it is determined in box 169 that user input was received from the GUI, then in this example, processing proceeds to box 170.

[0122] Box 170 relates to determining whether a "capture end" event has occurred. In some examples, box 170 may respond to user input, which may or may not be received via the GUI mentioned above. In this example, box 170 may relate to determining whether user input received via the GUI mentioned above corresponds to the end of capture phase 116. If it is determined in box 170 that the user input from the GUI corresponds to the end of capture phase 116, in some examples, the process continues to box 173. However, as noted elsewhere in this document, multiple instances of starting and ending capture may exist before box 173. Accordingly, in Figure 1D In this example, the arrows connecting boxes 170 and 173 are broken. If it is determined in box 170 that the user input from the GUI does not correspond to the end of capture phase 116, then in this example, the process continues to box 171.

[0123] Box 171 relates to “Registering User Input.” In some examples, user input may respond to custom options provided by the GUI, such as increasing the volume of a sound source, decreasing the volume of a sound source, selecting a potential sound source, etc. As noted elsewhere in this document, registering user input may involve creating and storing metadata corresponding to the user input, and the registered user input may be used to guide modifications to the audio data during capture phase 116, post-capture phase 118, or both. For example, if user input indicating an intention to increase the volume of a selected sound source is received, a beamforming process can be provided during capture phase 116 to enhance the sound from the selected sound source. This is an example of what may be referred to as “enhancing audio capture” in this document. Processing can continue to Box 168.

[0124] Box 173 relates to "post-capture editing." According to some examples, Box 173 may be performed by a control system of a device participating in audio and video capture performed by one or more remote devices (such as one or more servers) or a combination thereof. In some examples, Box 173 may involve modifying audio data received during the capture phase based on registered user input (such as metadata corresponding to the user input). According to some examples, Box 173 may involve presenting a GUI including at least one user input area allowing selection of potential or candidate sound sources. In some examples, user input corresponding to the selection of potential or candidate sound sources may have been previously received. According to some examples, Box 173 may involve replacing potential or candidate sound sources with external or synthesized audio. For example, Box 173 may involve replacing music detected in audio data from a microphone system, an example of a "candidate sound source," with downloaded music. In another example, Box 173 may involve adding external or synthesized audio corresponding to a selected potential sound source, for example, adding external or synthesized audio corresponding to a ticking clock when the selected potential sound source is a clock detected via video data.

[0125] As noted elsewhere in this document, method 160 may involve storing unmodified audio data received during the capture phase. According to some examples, box 173 may involve presenting a GUI including at least one user input area (e.g., a virtual slider, virtual dial, etc.) configured to receive a user selection of a ratio between modified audio data corresponding to enhanced audio capture and unmodified audio data corresponding to “real-world” audio capture.

[0126] According to some examples, pre-capture phase 114, capture phase 116, or both may involve performing a first sound source separation process. In some examples, box 173 may involve, for example, a second sound source separation process performed by one or more servers. This is potentially advantageous because the second sound source separation process may be more accurate, but it is computationally too intensive for the capture device to perform during pre-capture phase 114 or capture phase 116. Processing can continue to box 175.

[0127] Box 175 relates to preparing one or more media files, storing one or more media files, transmitting one or more media files, or a combination thereof. In this example, the one or more media files include a final version of audio data produced by the post-capture editing process of Box 173. In some examples, Box 175 may relate to providing a final version of audio and video data in a standard media container, such as Moving Picture Experts Group (MPEG)-4 Part 14 (MP4) . According to some examples, Box 175 may relate to transmitting a bitstream comprising the final version of audio and video data. According to some examples, the bitstream may be a bitstream encoded with Immersive Speech and Audio Services (IVAS). This document provides some examples of encoding and decoding IVAS bitstreams.

[0128] Figure 2A Examples of timelines and time intervals are shown when some of the disclosed processes may occur. These processes may, for example, be at least in part composed of Figure 1A The device 101 or a similar device performs the operation. In these examples, timeline 202 indicates the capture application start time 204, capture start time 206a, capture end time 208a, capture start time 206b, and capture end time 208b. According to these examples, the first capture instance of fragment capture #1 occurs during the time interval between capture start time 206a and capture end time 208a. Here, the second capture instance of fragment capture #2 occurs during the time interval between capture start time 206b and capture end time 208b. In these examples, timeline 202 breaks between capture end time 208a and capture start time 206b to indicate that the time interval can be variable.

[0129] In these examples, capturing the application startup time 204 is... Figure 1C An instance of capturing application startup time 104 can be basically as described in the reference. Figure 1C As described. Similarly, the capture start times 206a and 206b are... Figure 1C The instance with capture start time 106, and capture end times 208a and 208b are... Figure 1C The capture end time of 108 instances can be basically as referenced. Figure 1C As described. Like other published examples, Figure 2A The types, quantities, and arrangements of the elements shown are provided as examples only. For example, although Figure 2A The capture phase 116 includes two capture instances, but other examples may include more or fewer capture instances.

[0130] Depending on the specific diagram, changes to the composition of the audio scene can be made in several ways. Figure 2AThe diagram illustrates some options. According to these examples, user interaction #1 is detected during pre-capture phase 114, and user interaction #2 is detected during capture phase 116. User interaction #1 and user interaction #2 can be detected, for example, based on user interaction with a GUI (such as one of the interactions disclosed herein). For example, the GUI can be overlaid on an image corresponding to video data acquired and displayed by a capture device (such as a cellular phone or other mobile device). The GUI may include one or more audio source images corresponding to at least one audio feature and may include one or more user input areas.

[0131] Based on some examples, user interaction #1 could correspond to adjusting the level of the sound source during the pre-capture phase 114. In some examples, this adjustment can immediately affect the graphical feedback provided to the user. For example, a virtual level or volume control slider, dial, etc., could change its size, position, etc. Figure 2A In the example shown, user interaction effect #1 is applied to the audio captured during the subsequent capture phase 116, and is therefore referred to as "causal interaction".

[0132] In some examples, user interaction #2 (during capture phase 116) can correspond to user input used to adjust the ratio between the acoustic background and acoustic foreground, adjust the level of the audio source, etc. Figure 2A In the example shown, the user interaction effect #2 is applied in a "reverse causal" manner by applying it from the start of segment capture #1 at capture start time 206a. For example, the capture device may automatically apply the reverse causal effect during capture phase 116, at the end of capture phase 116, or during post-capture phase 118.

[0133] Figure 2B This is a flowchart outlining various example methods 250 according to some disclosed embodiments. Example methods 250 can be divided into blocks, such as blocks 251, 252, 253, 255, 256, 257, 259, 261, 263, and 265. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 250 need not be performed in the indicated order. In some embodiments, one or more blocks of method 250 can be performed simultaneously. Furthermore, some embodiments of method 250 may include more or fewer blocks than those shown and / or described. The blocks of method 250 can be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0134] According to these examples, boxes 252, 253, 255, 256, 257, and 259 are executed during capture phase 116. In these examples, boxes 263 and 265 are executed during post-capture phase 118. Capture phase 116 and post-capture phase 118 can be, for example, as referenced herein. Figure 1C As described. Processing can begin at box 251.

[0135] Box 251 relates to "capturing application startup". In this example, box 251 corresponds to Figure 1D Box 161. Processing can continue to box 252.

[0136] Box 252 relates to “obtaining capture preferences”. In this example, capture preferences are or include audio capture preferences. Such audio capture preferences can be obtained in various ways. According to some examples, one or more audio capture preferences can be obtained in box 252 based on explicit user actions (e.g., when the user explicitly sets audio capture preferences by interacting with the GUI). Accordingly, box 252 can relate to receiving user input from the GUI corresponding to one or more capture preferences.

[0137] In some examples, box 252 may involve retrieving previously acquired user preference data from memory. According to some examples, one or more audio capture preferences can be obtained by accessing a stored data structure that includes audio capture preferences. According to some implementations, audio capture preferences may be represented as one or more lookup tables, which include a list of audio sources and corresponding signal level characteristics for each audio source. Such lookup tables may take the form of a data structure that includes a list of detected audio sources present in the audio scene. In some such examples, the list of detected audio sources may include foreground audio objects and one or more background components of the audio scene. Background components may include multiple context-dependent sound sources, such as multiple vehicle sound sources, background speakers, other street noise sources, etc., in a street scene, background music, background speakers, food-related sounds, and food-related sounds, etc., in a coffee shop scene. Within the data structure, each audio source may be associated with a semantic tag, such as "human speaker #1", "human speaker #2", etc., for foreground audio objects, and "street noise", "noisy speech noise", "wind noise", etc., for background components. In some examples, each audio source in the data structure may be associated with a timestamp sequence and a binary flag indicating the presence or absence of the audio source, a signal level measurement of the audio source, and the detection probability of the audio source.

[0138] In some examples, semantic labels for objects can be generated from instances of “audio taggers,” which are implemented by the control system of a device used for audio capture. Such audio taggers can be implemented in many ways depending on the specific implementation. According to some examples, the audio tagger can be implemented as a trained neural network built according to a so-called YAMNet (yet another multi-scale convolutional neural network) architecture, which can employ the MobileNet architecture that operates on the Mel spectrogram of the audio signal. An example of MobileNet is described in Howard, Andrew G. et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications” (arXiv preprint arXiv:1704.04861 (2017)), which is incorporated herein by reference. This architecture facilitates the assignment of tags from a large set of labels to segments of a waveform. The assigned tags are examples of semantic labels.

[0139] More generally, an audio tagger can be constructed using a convolutional neural network (CNN). Such an audio tagger operates on a spectral representation of the audio signal (e.g., a Mel spectrogram) and can update a set of tags and their confidence intervals at a predefined time resolution (e.g., 1 second). The CNN can be trained to provide an output that includes an indication of the tags (within a large set of all possible tags) corresponding to audio objects present in the segment, and another output representing the confidence level of those tags. The output representing the confidence level can be, for example, an object detection probability, such as a value between zero and one, where values ​​closer to 1.0 indicate high confidence, while values ​​closer to zero indicate low confidence. This confidence level can be determined based on an estimated object detection probability, which can be computed by the same CNN.

[0140] Based on some examples, the signal level of an individual audio object can be determined as follows. An example of what is referred to herein as a “first source separator” can be used to find the distribution of signal levels (or signal energy) among all foreground components of an audio scene and the overall signal level of the background components of the audio scene. Then, knowing the foreground components of the audio scene present in the segments of the audio signal and the signal levels of that group of foreground components, the control system can estimate the weighted signal levels of individual foreground components of the audio scene, given that the detection probability has been known by weighting the total signal level of the foreground components of the audio scene relative to the detection probability of each foreground component. The detection probability can be provided via analysis of the video signal associated with the audio signal, for example, by detecting potentially active speakers in the audio signal and determining the corresponding confidence level for each potentially active speaker (e.g., based on mouth movements corresponding to time intervals of speech).

[0141] In some examples, object detection in video associated with an audio scene can be used to adjust the distribution of signal energy. For instance, if multiple speakers are present in the scene simultaneously, video analytics can be used to identify active speakers. The instantaneous signal energy associated with all current speakers can then be distributed among the active speakers.

[0142] According to some examples, previously acquired user preference data may have already been obtained during a previous capture phase 116 or during a previous segmentation of the current capture phase 116 (e.g., when a previous segment is acquired during the current capture phase 116). In some examples, box 252 may involve obtaining one or more audio capture preferences by analyzing the audio segments obtained during the previous capture phase 116 (e.g., using analysis tools of the capture engine) and using the analysis results to populate the audio capture settings of the current capture phase 116. The one or more audio capture preferences obtained in box 252 may include (a) a preferred ratio between the total audio signal level of all foreground audio components and (b) the total audio signal level of the background audio components. Alternatively or additionally, the user may prefer that the audio level associated with the background components in a scene involving a human speaker should not exceed a certain level. Accordingly, the user's preference for foreground and background audio levels may be indicated by a preferred minimum foreground to background ratio, a preferred maximum background audio level, or both. In some embodiments, obtaining one or more audio capture preferences may involve receiving user input indicating the classification of audio sources to focus on or emphasize. For example, users can specify that the captured audio signal should be emphasized to correspond to a human speaker, an audio signal should correspond to a musical instrument, and so on.

[0143] In some examples, after one or more audio scene preferences are known or set, these preferences can be used to facilitate a comparison of the current composition of the audio scene with respect to these preferences. Accordingly, in some examples, box 252 (or another box of method 250, such as box 253) may involve mapping one or more capture preferences to one or more elements of the current audio scene. Processing can continue to box 253.

[0144] Box 253 relates to “Capture and Feedback”. In this example, box 253 relates to acquiring and storing audio and video data during capture phase 116, and providing feedback, for example, via a GUI. According to some examples, the feedback provided in box 253 may relate to signaling deviations from a typical scene, audio capture preferences, or both captured audio scenes. For example, the feedback provided in box 253 may indicate that the ratio between the acoustic background level and the acoustic foreground level of the current audio scene differs from a preferred level. In some examples, the feedback provided in box 253 may indicate that the levels of one or more sound sources (such as wind, traffic noise, etc.) may exceed the desired levels. Processing may continue to box 255.

[0145] Box 255 relates to determining whether an adjustment has been made. According to some examples, box 255 may relate to determining whether input indicating that the user expects an adjustment has been received from the user during capture phase 116 (e.g., through user interaction with the GUI presented in box 253). The interaction may correspond, for example, to adjusting the desired level of the audio source, adjusting the ratio between the acoustic background and acoustic foreground, etc. If it is determined in box 255 that no adjustment will currently be made, in this example, the process returns to box 253. In many cases, boxes 253 and 255 involve concurrent operation, as the capture process will typically continue during box 255. However, if it is determined in box 255 that an adjustment will currently be made, in this example, the process proceeds to box 257.

[0146] Box 257 relates to “applying changes including feedback.” In this example, box 257 relates to processing received user input indicating that the user expects adjustments, and providing feedback, for example via a GUI from which the input was received, indicating that one or more changes have been applied or will be applied. In some examples, the changes may involve some type of enhancement capture, such as a microphone beamforming process for enhancing audio received from a selected sound source. According to some examples, enhancing audio capture may involve replacing the audio corresponding to a selected potential sound source or a selected candidate sound source with external or synthesized audio. In some examples, such adjustments can be performed in a “reverse causal” manner by applying the adjustments from the beginning of the captured segment. However, in some examples, the changes can be performed in a “causal” manner by applying the adjustments from the time the corresponding user input was received until the end of the captured segment. Alternatively or additionally, in some examples, the adjustments may be performed at least partially after capture stage 116, for example during a post-capture editing process.

[0147] Box 259 relates to determining whether “[Yes] capture is complete.” In some examples, box 259 may relate to determining whether a “capture end” event has occurred. In some examples, box 259 may respond to user input, which may or may not be received via the GUI mentioned above. In this example, box 259 may relate to determining whether user input received via the GUI mentioned above corresponds to the end of capture phase 116. If it is determined in box 259 that user input from the GUI corresponds to the end of capture phase 116, then in some examples, the process continues to box 263. However, as noted elsewhere in this document, multiple instances of starting and ending capture may exist before box 263. Accordingly, in Figure 2B In this example, the arrows connecting boxes 259 and 263 are broken. If it is determined in box 259 that the user input from the GUI does not correspond to the end of capture phase 116, then in this example, the process returns to box 253.

[0148] Box 263 relates to "post-capture editing." According to some examples, box 263 can be executed by a control system of a device participating in audio and video capture performed by one or more remote devices (such as one or more servers) or a combination thereof. In some examples, box 263 can relate to modifying audio data received during the capture phase based on registered user input (such as metadata corresponding to the user input). In some examples, box 263 can be as referenced above. Figure 1DThe operation is performed as described in box 173. In some examples, for example, during capture phase 116, user input corresponding to the selection of a potential or candidate sound source may have been received previously. According to some examples, box 263 may involve replacing the potential or candidate sound source with external or synthesized audio. For example, box 263 may involve replacing music detected in the audio data from a microphone system, which is an example of a “candidate sound source,” with downloaded music. In another example, box 263 may involve adding external or synthesized audio corresponding to the selected potential sound source, for example, adding external or synthesized audio corresponding to a ticking clock when the selected potential sound source is a clock detected via video data.

[0149] As noted elsewhere in this document, method 160 may involve storing unmodified audio data received during the capture phase. In some examples, box 263 may involve allowing a user to interpolate between modified audio data corresponding to the enhanced audio capture and unmodified audio data corresponding to the “real-world” audio capture. According to some examples, box 263 may involve presenting a GUI including at least one user input area (e.g., a virtual slider, virtual dial, etc.) configured to receive a user selection of the ratio between the modified audio data corresponding to the enhanced audio capture and the unmodified audio data corresponding to the “real-world” audio capture. In some examples, box 263 may involve allowing a user to edit the modified audio data by replacing another instance of external or synthesized audio, for example, by allowing a user to select a different type of background music, selecting a different audio segment corresponding to a potential sound source, such as a different “tickling clock” audio segment, etc. Processing may continue to box 265.

[0150] Box 265 relates to preparing one or more media files, storing one or more media files, transmitting one or more media files, or a combination thereof. In this example, the one or more media files include a final version of audio data produced by the post-capture editing process of Box 263. In some examples, Box 265 may relate to providing a final version of audio and video data in a standard media container, such as Moving Picture Experts Group (MPEG)-4 Part 14 (MP4) . According to some examples, Box 265 may relate to transmitting a bitstream comprising the final version of audio and video data. According to some examples, the bitstream may be a bitstream encoded with Immersive Speech and Audio Services (IVAS).

[0151] In some examples, the visualization of the audio scene can allow for customization of the audio capture. According to some such examples, the audio customization interface can be implemented as a layer overlaid on the video scene. For example, the GUI can provide multiple customization layers, where a basic layer involves feedback on the composition of the audio scene, and more advanced customization options can be presented as one or more subsequent or optional layers. In this context, the terms "customization layer" and "layer of customization" refer to one or more graphical features that can be presented on the GUI. In some examples, one or more graphical features can be presented in an accumulated or additive manner. Some examples of this advanced customization could involve adjusting the audio source levels between background and foreground audio sources, adjusting the speech levels from multiple speakers to a consistent level, and so on.

[0152] Figure 3A The illustration shows an example of a GUI that can be presented to indicate one or more audio scene preferences and one or more current audio scene characteristics. In this example, Figure 3A The diagram illustrates a GUI 300, an image corresponding to a video frame of video clip 301, a person 305 in the current video clip 301, an audio source summary region 315 of the GUI 300, audio source information regions 330a and 330b within the audio source summary region 315, audio source level information regions 332a and 332b within the audio source information regions 330a and 330b respectively, and an audio source level preference indicator 333 in the audio source level information region 332b. In this example, the GUI 300 is overlaid on the video clip 301. The GUI 300 and video clip 301 can be presented as follows: Figure 1A On devices such as device 101, for example, presented in relation to Figure 1A The output unit 147 shown corresponds to the display. In this example, device 101 is a cellular phone.

[0153] In this example, audio source-level information region 332a provides information about the speaker's audio source, which in this instance corresponds to the speech of person 305. Here, audio source-level information region 332b provides information about the audio source category, which in this example is cafeteria background noise, and may include multiple individual audio sources. In some examples, device 101 may be configured to determine or estimate the audio source category, for example, by implementing an audio source classifier as disclosed herein, to assess the background audio currently being received by device 101's microphone system.

[0154] According to this example, audio source level information areas 332a and 332b respectively indicate the estimated current audio level corresponding to the speaker and cafeteria noise. In this example, audio source level preference indication 333 is obtained from a template corresponding to a specific type of background noise, in this instance, cafeteria noise. According to some such examples, audio source level preference indication 333 may correspond to one or more audio levels previously selected by the current user of device 101 for cafeteria noise.

[0155] In this example, audio source level information areas 332a and 332b each include a dotted portion 337a and a portion 337b with diagonal hash markers. Portion 337b indicates the degree to which the current sound level has been modified according to user input. The modification indicated by portion 337b can be an increase or a decrease. For example, portion 337b in audio source level information area 332a could indicate the degree to which the current sound level corresponding to a speaker has increased according to user input, while portion 337b in audio source level information area 332b could indicate the degree to which the current sound level corresponding to cafeteria noise has decreased according to user input. In some examples, user input may involve touching audio source level information areas 332a and 332b, for example, interacting with a virtual slider provided by audio source level information areas 332a and 332b.

[0156] As with other published examples, Figure 3A The types, quantities, and arrangements of the elements shown are provided as examples only. For example, although Figure 3A The audio source summary area 315 provides information about a single sound source and a sound source category, but other examples can provide information about multiple individual sound sources. Alternatively or additionally, some alternative examples can provide different graphical depictions of sound source characteristics. Some such examples are provided in this disclosure.

[0157] Figure 3B This is a block diagram illustrating an example of a custom layer that can be rendered via one or more GUIs. One or more GUIs can be rendered as... Figure 1A On devices such as device 101, for example, presented in relation to Figure 1A The output unit 147 shown corresponds to the display. In some examples, the device can be a mobile device, such as a cellular phone. In this example, Figure 3BBoxes 355, 360, 365, and 370 are shown, each corresponding to a custom layer that can be rendered via one or more GUIs. According to this example, each of boxes 355, 360, 365, and 370 is interconnected with a bidirectional arrow labeled "User Interaction," indicating that the user can interact with the GUI corresponding to any of boxes 355, 360, 365, and 370 to add a custom layer to the currently rendered GUI or return to a different GUI corresponding to another custom layer.

[0158] In these examples, box 355 corresponds to a first custom layer, which corresponds to a variety of possible graphical feature sets for providing basic feedback levels about the composition of the audio scene. For example, the GUI corresponding to the first custom layer can provide a graphical depiction of one or more audio sources or groups of audio sources in the current audio scene. According to some examples, the GUI corresponding to the first custom layer can provide an estimate of the level of the content of the most relevant audio sources estimated to be present in the audio scene. Alternatively or additionally, the GUI corresponding to the first custom layer can display images corresponding to one or more background audio sources and one or more foreground audio sources. In some such examples, the GUI corresponding to the first custom layer can provide an estimate of the ratio of the level of the background audio sources to the level of the foreground audio sources.

[0159] Based on these examples, box 360 corresponds to a second custom layer, which corresponds to various possible sets of graphical features to facilitate basic customization of one or more sound sources in an audio scene by the user. In some examples, one or more additional graphical features corresponding to the second custom layer may include one or more areas for receiving user input, such as adjusting the ratio between the audio levels of the background audio components and the foreground audio components, adjusting the level of one or more foreground components, etc. In some such examples, audio data corresponding to one or more sound sources may be modified (e.g., suppressed or enhanced) based on the received user input. Based on these examples, the user may be able to interact with the user input area of ​​the GUI corresponding to box 355 to add one or more additional graphical features corresponding to the second custom layer of box 360.

[0160] In these examples, box 365 corresponds to a third custom layer, which corresponds to a variety of possible sets of graphical features to facilitate advanced customization of one or more sound sources in the audio scene by the user. According to these examples, the user can interact with user input areas of the GUI corresponding to box 355 or box 360 to add one or more additional graphical features corresponding to the third custom layer of box 365. In some examples, one or more additional graphical features corresponding to the third custom layer may include one or more areas for receiving user input to introduce audio components not currently present in the audio data captured by the microphone system of the capture device (e.g., device 101), or audio components whose acoustic signals are below a threshold level. For example, video data obtained by the camera system of the capture device may indicate one or more potential sound sources whose audio data captured by the microphone system is below a threshold level. Potential sound sources may include distant people or animals, clocks, audio devices that are not currently in use or are producing low-level sounds, etc. One or more additional graphical features corresponding to the third custom layer may indicate one or more such potential sound sources and one or more user input areas for selecting potential sound sources. After a user selects a potential sound source, one or more additional graphical features can be presented to the user, corresponding to the options for selecting audio data that corresponds to the selected potential sound source.

[0161] According to these examples, box 370 corresponds to a fourth custom layer, which corresponds to various possible sets of graphical features to facilitate post-capture editing of one or more sound sources in the captured audio scene. According to these examples, a user can interact with user input areas of the GUI corresponding to boxes 355, 360, or 365 to add one or more additional graphical features corresponding to the fourth custom layer of box 370. In some examples, the one or more additional graphical features corresponding to the fourth custom layer may include one or more areas for receiving user input indicating a ratio between enhanced audio capture and real-world audio capture. These one or more additional graphical features may include, for example, virtual sliders, virtual dials, etc.

[0162] As with other published examples, Figure 3B The types, quantities, and arrangements of the elements shown are provided as examples only. For example, although Figure 3B Four examples of custom layers are shown, but other implementations may involve more or fewer than four custom layers, one or more different types of custom layers, etc.

[0163] Figure 4A The diagram shows that... Figure 3B An example of a GUI rendered in the first custom layer. In this example, Figure 4AThe diagram illustrates a GUI 400a, an image corresponding to a video frame of video clip 401a, a person 405a in the current video clip 401a, audio source level representations 410a and 410b, an audio source summary region 415a, and audio source information regions 420a and 420b within the audio source summary region 415a. In this example, the GUI 400a is overlaid on the video clip 401a. The GUI 400a and video clip 401a can be presented as follows: Figure 1A On devices such as device 101, for example, presented in relation to Figure 1A The output unit 147 shown corresponds to the display. In this example, device 101 is a cellular phone.

[0164] In this example, audio source level information area 410a provides information about the speaker's audio source, which in this instance corresponds to the speech of person 405a. Here, audio source level information area 410b provides information about another speaker's audio source, which corresponds to the speech of the person operating device 101. According to this example, audio source level information areas 410a and 410b, based on their sizes, indicate the estimated current audio level corresponding to the speech of person 405a and the speech of the person operating device 101, respectively.

[0165] Here, audio source information areas 420a and 420b provide information about two audio source categories: street noise and speech, and may include multiple individual audio sources. In this example, the relative sizes of audio source information areas 420a and 420b correspond to relative audio levels. In some examples, the total size of audio source information areas 420a and 420b may correspond to a total instantaneous audio level, which may be scaled relative to the maximum instantaneous audio level that can be captured by the audio capture system of device 101. In some such examples, the length or volume of audio source summary area 415a may correspond to the maximum instantaneous audio level that can be captured by the audio capture system of device 101. According to this example, audio source information area 420a provides information about the overall sound levels of various sound sources identified by device 101 (or one or more devices communicating with device 101, such as devices of a cloud service) as belonging to the audio source category of street noise. In some examples, device 101 may be configured to determine or estimate the audio source category by implementing an audio source classifier as disclosed herein, to assess background audio currently being received by the microphone system of device 101 as street noise. Depending on the specific implementation, audio source information area 420b may provide information about the overall speech level of person 405a and the person operating device 101, or information only about the speech level of person 405a.

[0166] Figure 4B The diagram shows that... Figure 3B Another example of a GUI rendered in the first custom layer. In this example, Figure 4B It shows in Figure 4A The GUI 400a, located at a time following the time depicted in the image, simultaneously displays the same video clip 401a. Accordingly, in addition to the following details, Figure 4B The elements of GUI 400a shown are... Figure 4A The elements shown are identical: audio source level representation 410a is not shown because person 405a is not speaking at this time; vehicle 425 (a bus in this example) is now part of the audio scene; GUI 400a now includes audio source level representation 410c corresponding to vehicle 425; and audio source summary area 415a now includes audio source information area 420c, which also corresponds to vehicle 425. Figure 4A and Figure 4B The audio source information areas 420a and 420b (and 420c when present) shown in the figure provide a graphical depiction of the total instantaneous audio level of the audio scene and the instantaneous distribution of the total audio level among the components of the audio scene.

[0167] In this example, the audio source-level information region 410c is greater than Figure 4A The audio source level information areas 410a and 410b shown are... Figure 4B The audio source level information area 410b indicates that the bus sound is relatively louder than any other audio source in any audio scene, and therefore results in a higher received audio signal level. Furthermore, the audio source level information area 410c shows a triangle with an exclamation mark, indicating that the estimated current audio level corresponding to the bus may be problematicly high.

[0168] For example, the sound source separator implemented by device 101 will typically have performance limitations beyond which the components of the audio signal cannot be reliably estimated. In this example, if the level of the background audio signal (such as the background audio signal corresponding to a bus) is too high, the instance of the speech separator may fail to separate the speech. Therefore, it is potentially advantageous to notify the user performing the capture of this situation, as this can prompt the user to take corrective actions, such as moving closer to the audio source of interest, pausing the audio segment until the background audio level decreases (e.g., waiting until the bus has continued moving), etc.

[0169] GUI (e.g.) Figure 4A or Figure 4BThe GUI 400a offers a variety of other potential advantages. For example, the device 101 presenting the GUI 400a has reduced the number of possible audio sources providing information about it and the types of information provided to a manageable number of user feedback elements presented in the GUI 400a. Accordingly, the user is not overwhelmed by a large amount of information, but can instead focus on the information categories that the device 101 estimates to be most relevant. Thus, the user is able to focus on engaging in the task of capturing the desired type of video footage and the associated audio.

[0170] As with other published examples, Figure 4A and Figure 4B The types, quantities, and arrangements of the elements shown are provided as examples only. For example, although Figure 4A and Figure 4B The audio source summary area 415a provides information about the two sound source categories, and Figure 4B The audio source summary area 415a provides information about the sound from a single vehicle, but other examples may provide information about more or fewer individual sound sources or groups of sound sources. Alternatively or additionally, some alternative examples may provide different graphical depictions of the sound source characteristics. Some such examples are provided in this disclosure.

[0171] Figure 4C The diagram shows that... Figure 3B An example of a GUI rendered in a second custom layer. In this example, Figure 4C It shows something similar to Figure 4A and Figure 4B GUI 400a and GUI 400b are shown. For example, GUI 400b correspondingly shows circular audio source level representations 410a, 410b, and 410c for person 405b, person operating device 101, and vehicle 425. In this example, vehicle 425 is a truck. Furthermore, the size of the audio source level representations 410a, 410b, and 410c shown in GUI 400b corresponds to the level of each corresponding audio source.

[0172] However, GUI 400b includes two significant differences. One difference is that the audio source summary area 415 includes audio source information areas 420d and 420e, which correspond to the background audio source and the foreground audio source, respectively. The “background” audio source information area 420d may, for example, correspond to street noise and vehicle noise 425. In some examples, the “foreground” audio source information area 420e may correspond only to the speech of a person 405b. Another significant difference is that the audio source summary area 415 includes virtual sliders 440a and 440b at the edges of the audio source information areas 420d and 420e, respectively. Users can interact with the virtual sliders 440a and 440b to indicate desired increases or decreases in the level of the background or foreground audio, respectively. Figure 4C In the image, the user's finger 435 is shown interacting with a virtual slider 440b to indicate desired modifications to the foreground sound level.

[0173] GUIs (such as GUI 400b) offer a variety of potential advantages. For example, the device 101 presenting GUI 400b has reduced the number of possible audio sources providing information about them, as well as the types of information provided, to a manageable number of user feedback elements. The audio source summary area 415 of GUI 400b is even simpler than that of GUI 400a, because all sound sources are categorized as foreground or background. Users can provide feedback to modify the level of foreground or background audio, which is both convenient and not overly complex. Accordingly, users are not overwhelmed by a large amount of information, but can instead focus on engaging in the task of capturing the desired type of video clip and the associated audio. As with other disclosed examples, Figure 4C The types, quantities, and arrangements of the elements shown are provided as examples only.

[0174] Figure 4D and Figure 4E The diagram shows that... Figure 3B The example elements of the GUI rendered by the third or fourth custom layer. In other words, Figure 4D and Figure 4E The GUI 400c can be used for advanced customization, post-editing processes, or both. The GUI 400c and video clip 401c are presented in... Figure 1A In instances of device 101, the device is, in these examples, a cellular phone. In these examples, the GUI 400c is overlaid on video footage 401c depicting a restaurant scene. As with other disclosed examples, Figure 4D and Figure 4E The types, quantities, and arrangements of the elements shown are provided as examples only.

[0175] In these examples, Figure 4D and Figure 4E The diagram shows an image corresponding to a video frame in a video segment 401c at two different times, a person 405a in the video segment 401c at two different times, an audio source summary region 415c, audio source information regions 430a, 430b, 430c and 430d within the audio source summary region 415c, audio source tags and audio source level information regions 432a, 432b, 432c and 432d within the audio source information regions 430a, 430b, 430c and 430d respectively, an average audio source level indicator 433 within the audio source information regions 430a, 430b, 430c and 430d, a region 437a within the audio source level information regions 432a, 432b, 432c and 432d, and a region 437b within at least some of the audio source level information regions 432a, 432b, 432c and 432d. Figure 4E It also includes a user input area 460, which is configured to receive user selections for enhanced audio capture. In this example, enhanced audio capture involves replacing candidate audio sources with external audio.

[0176] According to these examples, the audio source labels for audio source information areas 430a, 430b, 430c, and 430d are "cafeteria noise," "speaker," "off-screen speech," and "music playing," respectively. In these examples, the speaker is person 405a, the off-screen speech corresponds to the speech of the person using device 101, and "music playing" corresponds to background music currently playing in the restaurant. According to these examples, "cafeteria noise" corresponds to background noise in the restaurant, which has been classified into the cafeteria noise category by the audio classifier implemented by device 101. In these examples, audio source level information areas 432a, 432b, 432c, and 432d indicate the estimated current audio level corresponding to the background noise in the restaurant, the speech of person 405a, the speech of the person using device 101, and the background music, respectively. According to these examples, a user can interact (e.g., touch) with a plus or minus sign in any of the audio source information areas 430a, 430b, 430c, and 430d to indicate a desired increase or decrease in the relative signal level of the corresponding audio source relative to other audio sources. In some examples, the control system of device 101 can be configured to create and store user input metadata corresponding to user input received via the plus or minus sign. Depending on the specific implementation, audio source level changes may or may not occur during the capture process. In some examples, audio source level changes may occur during a post-capture editing process and may be based at least in part on the user input metadata corresponding to user input received via the plus or minus sign.

[0177] According to these examples, regions 437a and 437b correspond to the modified and unmodified audio levels within audio source level information regions 432a, 432b, 432c, and 432d, respectively. In these examples, the average audio source level indicator 433 indicates the audio source signal levels of background noise in the restaurant, the speech of person 405a, the speech of the person using device 101, and background music. The average audio source level indicator 433 may, for example, indicate the audio source signal of each audio source or a group of audio sources within a time interval.

[0178] exist Figure 4D In the example shown, text feedback area 450 displays "Background music detected here." Figure 4E In the example shown, text feedback area 450 displays "Background music can now be replaced with a streaming version." According to some examples, this type of instruction can be in response to, for example, the detection of a candidate sound source for enhanced audio capture by the control system of device 101, which involves replacing the candidate sound source with external audio. For example, the control system can be configured to detect and identify background music and suggest replacing the currently detected background music with another version of the same music or with other music of user selection. In this example, Figure 4E The user input area 460 displays an image corresponding to another version of the same song being detected, and this user input area is configured to receive the user's selection of the song for enhanced audio capture.

[0179] In some implementations, a GUI (such as GUI 400c) can be used for post-capture editing. For example, virtual sliders, virtual dials, etc., can be presented on the GUI (such as GUI 400c) to allow a user to select the ratio between enhanced audio capture and real-world audio capture. In one such example, a slider can be presented in one or more of the audio source-level information areas 432a, 432b, and 432c, allowing a user to modify the enhanced audio to include relatively more or relatively less unmodified audio.

[0180] GUIs (such as the GUI 400c) offer a variety of potential advantages. While the GUI 400c is not as simplified as the GUI 400a or GUI400b, it allows for more advanced customization that can occur before, during, or after audio capture.

[0181] Figure 5AThis is a flowchart outlining various example methods 500 according to some disclosed embodiments. Example methods 500 can be divided into blocks, such as blocks 505, 510, 515, 520, 525, 530, 535, and 540. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 500 need not be performed in the indicated order. In some embodiments, one or more blocks of method 500 can be performed simultaneously. Furthermore, some embodiments of method 500 may include more or fewer blocks than those shown and / or described. The blocks of method 500 can be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0182] Box 505 relates to "receiving audio data from a microphone system by the device's control system". The control system may be or may include Figure 1A CPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. In some examples, the microphone system may be a microphone system of a mobile device that includes a control system. Processing may continue to box 510.

[0183] Box 510 relates to "receiving video data from a camera system by a control system". In some examples, the camera system may include one or more cameras of a mobile device that includes a control system. Processing may continue to box 515.

[0184] Box 515 relates to “identifying two or more audio sources in an audio scene by a control system based at least in part on audio data and video data.” Audio sources may also be referred to herein as sound sources. In some examples, identification may involve creating a list of sound sources in the audio scene. The list of sound sources may include one or more actual sound sources and one or more potential sound sources. According to some examples, identification may involve a first sound source separation process performed by a control system. According to some such examples, post-capture audio processing may involve performing a second sound source separation process. Depending on the specific implementation, the second sound source separation process may be performed by the control system or by one or more other devices, such as one or more servers. In some examples, identification may involve the detection of one or more potential sound sources by a control system based at least in part on video data. In some such examples, at least one potential sound source may not be indicated by audio data. Processing may continue to box 520.

[0185] Box 520 relates to "estimating at least one audio characteristic of each of two or more audio sources based on audio data by a control system." In some examples, box 520 may relate to estimating the level of each of the two or more audio sources. Processing may continue to box 525.

[0186] Box 525 relates to "audio and video data received during the capture phase and stored by the control system". Box 525 may, for example, relate to storing the audio and video data in the memory of the device receiving the audio and video data. The capture phase may correspond to, for example, the references herein. Figure 1C The capture phase 116 is described. In some examples, boxes 505, 510, 515, and 520 may be executed at least partially during the capture phase. According to some examples, boxes 505, 510, 515, and 520 may be executed at least partially during a pre-capture phase, which may correspond to, for example, the references herein. Figure 1C The pre-capture phase 114 is described. Processing can continue to box 530.

[0187] Box 530 relates to "a display controlled by a control system to display images corresponding to video data and to display a graphical user interface (GUI) overlaid on these images before and during the capture phase, the GUI including audio source images corresponding to at least one audio characteristic of each of two or more audio sources, the GUI including one or more user input areas for receiving user input." In some examples, box 530 may relate to a display of control device 101 to present, as Figure 3A The GUI shown or Figures 4C to 4E One of the GUIs shown. Figure 4C The virtual sliders 440a and 440b are examples of "one or more user input areas for receiving user input". Processing can continue to box 535.

[0188] Box 535 relates to “receiving user input by the control system via one or more user input areas prior to the capture phase.” Figure 2A User interaction #1 provides an example of this type of user input. Box 535 may, for example, involve via... Figure 4C User input is received via one of the virtual sliders 440a and 440b (e.g., via the user's finger 435). In other examples, box 535 may involve receiving input via another type of virtual slider, via a virtual knob, a virtual dial, etc. In some alternative examples, user input may be received via one or more voice commands, one or more gestures, etc. Processing may continue to box 540.

[0189] Box 540 relates to "the control system modifying the audio data received during the capture phase according to user input." (See reference...) Figure 2A The described "causal" effect provides examples. Box 540 may, for example, involve modifying audio data corresponding to a selected audio source or a selected audio source category. Box 540 may, for example, involve modifying audio data received during the capture phase. In some such examples, box 540 may involve a beamforming process corresponding to a selected audio source. In some examples, method 500 may involve receiving user input via a user input area by a control system after the start of the capture phase. According to some such examples, box 540 may involve modifying audio data received during the duration of the capture phase based on user input. Reference Figure 2A The description of the "reverse causality" effect provides an example.

[0190] Alternatively or additionally, box 540 may involve modifying audio data received during the capture phase during a post-capture editing phase. In some such examples, box 525 may involve storing modified audio data that has been modified according to user input. According to some examples, method 500 may involve storing unmodified audio data received during the capture phase.

[0191] In some examples, method 500 may involve the control system creating and storing user input metadata corresponding to user input received via a user input area. As described above, block 540 may involve modifying audio data received during the capture phase during a post-capture editing phase. In some examples, modifying audio data received during the capture phase based on user input may involve post-capture audio processing based at least in part on the user input metadata. According to some examples, the control system may be configured to perform at least a portion of the post-capture audio processing. Alternatively or additionally, another control system (such as a server's control system) may be configured to perform at least a portion of the post-capture audio processing.

[0192] According to some examples, method 500 may involve a control system classifying two or more audio sources into two or more audio source categories. In some such examples, the GUI may include user input area portions corresponding to each of the two or more audio source categories. One such category could be "cafeteria noise," for example, as referenced. Figure 3A ,or Figure 4D or Figure 4E As described. Another such category could be "street noise," for example, as referenced. Figure 4A and Figure 4BAs described. In some examples, classifying two or more audio sources into two or more audio source categories may be based on a list of audio sources. According to some examples, the two or more audio source categories may include a background category and a foreground category. In some examples, one or more user input areas of the GUI may include at least one user input area configured to receive user input regarding a selected level or level ratio for each of the two or more audio source categories. Figure 4C The virtual sliders 440a and 440b are examples.

[0193] In some examples, method 500 may involve determining one or more actionable feedback types regarding the audio scene. Based on some such examples, the GUI may be based in part on one or more actionable feedback types. Figure 4B An example of actionable feedback is provided: the audio source level information area 410c shows a triangle with an exclamation mark, indicating that the estimated current audio level corresponding to the bus may be problematicly high. In this example, if the level of the background audio signal (such as the background audio signal corresponding to the bus) is too high, the speech separator instance may be unable to separate the speech. Notifying the user performing the capture of this situation can prompt the user to take corrective actions, such as moving closer to the audio source of interest, pausing the audio segment until the background audio level decreases (e.g., waiting until the bus has resumed moving), etc.

[0194] According to some examples, method 500 may involve a control system detecting one or more candidate sound sources for enhancing audio capture, which involves replacing the candidate sound sources with external or synthesized audio. (Reference) Figure 4D and Figure 4E An example is described where the control system detects background music and indicates that the background music is a candidate sound source for enhancing audio capture. In this example, enhancing audio capture involves replacing the candidate sound source with external audio. Foley effects are an example of synthesized audio.

[0195] In some examples, the GUI may include at least one user input area configured to receive user selections for a selected potential sound source or a selected candidate sound source. For example, in some implementations, GUI 400c may allow a user to select another sound source or sound source category, such as cafeteria noise, for example, by touching the corresponding audio source information area 430a.

[0196] According to some examples, the GUI may include at least one user input area configured to receive user selections for enhanced audio capture. Figure 4EUser input area 460 is an example of such a user input area. Enhanced audio capture can involve external or synthesized audio for a selected potential or candidate sound source. In some alternative examples, GUI 400c or another publicly disclosed GUI may provide the user with options for synthesized audio, such as possible onomatopoeia effects, for enhanced audio capture corresponding to the selected potential or candidate sound source.

[0197] In some examples, the GUI may include at least one user input area configured to receive user selection of the ratio between enhanced audio capture and real-world audio capture. As described above, in some examples, GUI 400c or a similar GUI may allow user selection of the ratio between enhanced audio capture and real-world audio capture.

[0198] Figure 5B This is a flowchart outlining various example methods 550 according to some disclosed embodiments. Example methods 550 can be divided into blocks, such as blocks 555a, 555b, 558a, 558b, 560, 562, 564, and 565. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 550 need not be performed in the indicated order. In some embodiments, one or more blocks of method 550 can be performed simultaneously. Furthermore, some embodiments of method 550 may include more or fewer blocks than those shown and / or described. The blocks of method 550 can be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0199] Box 555a relates to "receiving audio data". (And...) Figure 5A Similar to box 505, box 555a may relate to the receiving of audio data from the microphone system by the device's control system. The control system may be or may include Figure 1A The CPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. In some examples, the microphone system may be a microphone system of a mobile device that includes a control system. Processing may continue from block 555a to block 558a.

[0200] Box 555b relates to "receiving video data". (And...) Figure 5ASimilar to box 510, box 555a may involve the receiving of video data from a camera system by the control system of the device. In some examples, the camera system may be a camera system of a mobile device that includes a control system. According to some examples, boxes 555a and 555b may be executed simultaneously. Processing can continue from box 555b to box 558b.

[0201] Box 558a relates to “classifying audio scenes.” In this example, box 558a involves the control system selecting the most probable environmental context from a set of possible environmental contexts, such as “city park,” “airport,” “subway station,” “coffee shop,” and “living room,” based at least in part on the audio data received in box 555a. Box 558a may, for example, involve the implementation of an audio scene classification (ASC) method by the control system. The ASC method may, for example, involve using a trained neural network implemented by the control system. In some instances, the ASC method may be one of the methods described by B. Ding et al. in “Acoustic scene classification: acomprehensive survey” in “Expert Systems and Applications,” Volume 238, Part B, page 121902 (Elsevier, March 15, 2024), which is incorporated herein by reference. Processing may continue from box 558a to box 564.

[0202] Box 558b relates to “classifying video scenes.” In this example, box 558b involves the control system selecting the most probable environmental context from a set of possible environmental contexts based at least in part on the video data received in box 555b. Box 558b may, for example, involve the control system implementing a video scene classification (VSC) method, such as a multi-scale visual Transformer (MViTv2), a 3D CNN-based VSC method, a CNN-RNN-based VSC method, or another suitable VCS method. Processing can continue from box 558b to box 560.

[0203] Box 560 relates to “classifying video objects.” In this example, box 560 relates to the selection of objects in a currently acquired video image by a control system based at least in part on the video data received in box 555b and the audio scene classification in box 558b. Box 560 may, for example, relate to a video-based object detection method implemented by the control system, such as one described by J. Redmon et al. in “You Only Look Once: Unified, Real-Time Object Detection” (arXiv:1506.02640v5 [cs.CV] May 9, 2016), which is incorporated herein by reference, or other applicable video-based object detection methods. Processing may continue from box 560 to box 562.

[0204] Box 562 relates to “identifying potential audio sources.” In this example, box 562 relates to identifying potential audio sources by a control system based at least in part on the video object classification of box 560. In some examples, the operation of box 562 may be based at least in part on video data received in box 555b. For example, if the video object classification of box 560 identifies people in the scene, determining whether a person is a potential audio source may be based at least in part on one or more of the person’s activities, such as whether the person’s mouth is moving, whether the person is playing a musical instrument, whether the person is tapping on a table, etc. In some examples, certain types of objects (such as an empty chair or table) may not be classified as potential audio sources. According to some examples, box 562 may relate to identifying candidate sound sources for enhancing audio capture. In some such examples, enhancing audio capture may involve replacing candidate sound sources with external or synthesized audio. An example of the output of box 562 is shown in [the image / description]. Figure 5C As shown in the diagram and described below, the processing can continue from box 562 to box 564.

[0205] Box 564 relates to “identifying and classifying audio sources, and estimating audio source levels.” In this example, box 564 relates to the identification and classification of audio sources, and the estimation of audio source levels, by a control system based at least in part on potential audio sources identified in box 562, the audio scene classification in box 558a, and the audio data received in box 555a. In some examples, the operation of box 564 may involve creating a data structure (such as a lookup table) that includes a list of detected audio sources present in the audio scene. In some examples, the data structure may indicate whether each audio source present in the audio scene is a “foreground” audio source or a “background” audio source, wherein a foreground audio source is estimated to be an audio source of interest, and a background audio source is estimated to be noise or at least less significant than a foreground audio source. Box 564 may involve selecting semantic labels for each audio source in the foreground audio source (e.g., "human speaker #1", "human speaker #2", "musician #1", etc.) and selecting semantic labels for each background audio source (e.g., "street noise", "voice noise", "restaurant noise", "wind noise", "traffic noise", etc.) and including the semantic labels in the data structure. In some instances, multiple background audio sources existing in the audio scene can be grouped into a single background category in the data structure, such as "background noise".

[0206] According to some examples, box 564 may involve generating a timestamp sequence associated with binary flags indicating the presence or absence of a particular audio source. In some examples, box 564 may involve a signal level measurement for each audio source, which may also be associated with a timestamp sequence. According to some examples, box 564 may involve generating a detection probability for that source, which may also be associated with a timestamp sequence. An example of the output of box 564 is shown in... Figure 5D As shown in the figure, and described below.

[0207] In some implementations of box 564, semantic labels for an audio source can be generated using an instance of an audio tagger implemented by a control system. Such an audio tagger can be implemented in many ways. For example, it can be constructed using a convolutional neural network (CNN). This audio tagger can operate on a spectral representation of the audio signal (e.g., a Mel spectrogram) and can be configured to update a set of semantic labels or “tags” and confidence intervals for these labels at a selected or predefined time resolution (e.g., 1 second). The CNN can be trained to provide an output including an indication of the label corresponding to an audio object present in the segment (within a large set of all possible labels) and another output representing the confidence level of these labels, indicating the probability of object detection. For example, a value close to 1.0 can indicate a high confidence level, while a value close to zero can indicate a low confidence level. This confidence level can be determined based on an estimated object detection probability, which can be computed by the same CNN. In some examples, the audio tagger can be based on a so-called YAMNet architecture that operates on a Mel spectrogram of the audio signal. YAMNet is a deep neural network capable of classifying over 500 different types of sound sources, including horns, barks, laughter, and more. The YAMNet architecture can be derived using the MobileNet-v1 architecture, disclosed in "MobileNets: Efficient convolutional neural networks for mobile vision applications" (arXiv preprint arXiv:1704.04861 (2017) https: / / arxiv.org / pdf / 1704.04861), which is incorporated herein by reference. This architecture facilitates the assignment of semantic labels from a large set of semantic tags to segments of the waveform.

[0208] The signal levels of individual audio sources (which may be referred to herein as audio objects) can be determined as follows. An instance of an audio source separator, which may be called a "first source separator," can be implemented by a control system (such as the control system of a device for audio and video capture). In some examples, the first source separator can be used to find the distribution of signal levels (or signal energy) between all foreground and background components of an audio scene. Then, knowing the foreground audio sources present in the segmentation of the audio signal and the signal levels of a set of foreground audio sources, the control system can calculate the level of each foreground audio source by weighting the total level relative to the detection probability of these components, thus knowing their detection probability to estimate the signal level of each audio source. The detection probability can be provided by analyzing the video signal associated with the audio signal, for example, by detecting an active speaker based on whether their mouth is moving when apparent speech is detected.

[0209] In some embodiments, object detection in the video associated with the audio scene can be used to perform adjustments to the signal energy distribution. For example, if multiple speakers are present in the scene simultaneously, video analytics can be used to identify active speakers, and the instantaneous signal energy associated with each speaker can be distributed among the active speakers. Processing can continue from box 564 to box 565.

[0210] Box 565 relates to "a list of output audio sources". In this example, the list of audio sources corresponds to at least a portion of the data structure generated in Box 564. See below for reference. Figure 5D This describes a list of audio sources. In some examples, box 565 may involve a control system that displays at least a portion of the list of audio sources, such as in a GUI. Figures 4A to 4D The audio source summary areas 415a, 415b, and 415c provide examples of such an audio source listing display. According to some examples, box 565 may relate to the data structure generated in box 564.

[0211] Figure 5C A table is shown illustrating example elements of a video object data structure according to some disclosed implementations. In some examples, it is possible to... Figure 5B Table 570 is generated in box 562. In this example, table 570 includes fields 572, 574, 576, and 568. The number of fields, field types, and field order shown in table 570 are merely examples. For instance, in some implementations, the video object data structure may include an estimated probability of each estimated video object type, timestamps, etc. In some examples, the video object data structure may not include predictions of whether a video object is a background or foreground object. Furthermore, field titles (“Video Object ID”, etc.) are presented only for viewer convenience: the actual video object data structure may not necessarily include such field titles.

[0212] In this example, Table 570 shows information about the environment of a coffee shop or cafeteria (e.g., Figure 4D and Figure 4E Information about video objects in the environment of a café or cafeteria (as shown). Figure 5C In Table 570, the Video Object ID field 572 includes an identifier for each estimated video object. Here, the Estimated Video Object Type field 574 is the estimated video object type for each identified video object in Table 570. Based on some examples, the video scene classification in box 558b may correspond to a set of possible video object types. For example, the video scene classification in box 558b might indicate that the video scene corresponds to a coffee shop or cafeteria environment, and this set of possible video object types could include objects common in a coffee shop or cafeteria environment.

[0213] According to this example, field 576 indicates whether the control system estimates whether each video object type is a potential audio source at a specific time corresponding to table 570. In this example, potted plants, empty chairs, and empty tables are not considered potential audio sources, while people and clocks are considered potential audio sources.

[0214] In this example, field 578 instructs the control system to estimate, at a specific time corresponding to table 570, whether each actual or potential audio source is likely a foreground or background audio source. (Reference) Figure 4D The person shown near the counter in the background of video clip 401a can be considered as an actual or potential background audio source, while the person 405a near the center of video clip 401a can be considered as an actual or potential foreground audio source.

[0215] Figure 5D A table is shown illustrating example elements of an audio source manifest data structure according to some disclosed implementations. In some examples, Table 580 may be... Figure 5B The data is generated in boxes 564 or 565. In these examples, table 580 includes fields 582, 584, 586, 568, 590, and 592. The number, type, and order of fields shown in table 580 are merely examples. For instance, in some implementations, the audio source inventory data structure may include an estimated probability for each estimated audio source type, timestamps, etc. Other examples of table 580 may not include fields indicating how actual or potential audio sources are detected. Furthermore, field headers (“Audio Source ID”, etc.) are presented only for viewer convenience: the actual audio source inventory data structure may not necessarily include such field headers.

[0216] In these examples, Table 580 shows information about coffee shop or cafeteria environments (e.g.) Figure 4D and Figure 4EInformation about the audio source in the environment shown (such as a coffee shop or cafeteria). Figure 5D In Table 580, the Audio Source ID field 582 includes an identifier for each estimated audio source. Here, the Estimated Audio Source Type field 584 is the estimated audio source type for each identified audio source in Table 580. Based on some examples, the audio scene classification in box 558a can correspond to a set of possible audio source types. In these examples, the audio scene classification in box 558b has indicated that the video scene corresponds to a coffee shop or cafeteria environment, and the set of possible audio source types includes objects common in a coffee shop or cafeteria environment.

[0217] Based on these examples, field 586 indicates how each actual or potential audio source is detected. For example, the "cafeteria noise" of audio source 1 and the "background dialogue" of audio source 5 are primarily detected based on the audio data received in box 555a, but are expected due to the video scene and video object classification in boxes 558b and 560. In these examples, audio source 3 is detected only in the audio data, and audio source 6 is detected only in the video data.

[0218] In these examples, field 588 instructs the control system to estimate, at the time corresponding to Table 580, whether each actual or potential audio source is likely to be a foreground or background audio source. According to these examples, field 590 instructs the control system to estimate, at a specific time corresponding to Table 580, whether each actual audio source is likely to be heard by a human listener. In these examples, field 592 instructs the control system to estimate the audio signal level of each actual or potential audio source at the time corresponding to Table 580. According to these examples, the audio signal level is indicated with a rating from zero to ten, where the estimated audio value is rounded up or down to an integer. For example, an estimated audio signal level of 0.45 would be rounded down to 0, and an estimated audio signal level of 0.55 would be rounded up to 1. In some examples, a threshold may be applied to the estimated audio signal level to estimate whether each actual audio source is likely to be heard by a human listener. For example, the threshold could be 0.3, 0.5, 0.8, 1.0, etc. According to some examples, the threshold may vary based on the levels of one or more other audio sources, such as one or more background noise levels. Other examples could involve different ranges of audio signal levels, such as zero to one, zero to one hundred, etc.

[0219] Figure 6An example of a module that can be implemented based on some publicly available examples is shown. In this example, the control system section 606 is configured to implement a first audio source separation module 605, a video object classifier 608 including a video object detection module 610 and a video object confidence estimation module 620, an audio source level estimation module 615, an analysis module 625, a GUI module 630, and a file generation module 635. According to some examples, Figure 6 The inputs and outputs of the module shown can be accessed during the capture phase (e.g., reference). Figure 1C and Figure 2A Provided during the described capture phase 116). In some such examples, the control system section 606 may be part of the capture device or a control system comprising a capture system including more than one device.

[0220] In this example, the first audio source separation module 605 is configured to determine and output based at least in part on audio data 601 received from the microphone system and received auxiliary information 604. n The separate audio source data 607a-607n for each of the separate audio sources. In some examples, n It can be an integer of two or greater. As noted elsewhere herein, the term "audio source" has the same meaning as the term "sound source" in this disclosure. According to some examples, the first audio source separation module 605 can be configured to perform what is referred to herein as the "first audio source separation process" or "first sound source separation process". In some examples, the first sound source separation process can be implemented by a neural network (e.g., a neural network trained to separate a particular type of audio signal from a mixture containing a particular type of audio signal and some arbitrary background audio). In some examples, such a neural network can be trained using a regression-based training objective. In some examples, the architecture of the first source separator can depend on the category of the audio scene. For scenes involving human speakers, the first source separator can be, for example, an example of the speech-to-background separator disclosed in U.S. Patent Publication No. 2023368807A1, which is incorporated herein by reference.

[0221] More generally, the first sound source separator can be implemented by imposing a typical set of foreground objects (e.g., human speakers, musical instruments, etc.) by a control system configured to implement the first audio source separation module 605, and designing a dedicated foreground and background audio source separator for this set. Instances of such separators for a specific signal category may include a neural network model operating on a time-frequency representation of an audio signal, trained to generate separated audio from unseparated audio in that time-frequency representation. In some examples, such a network can be implemented using convolutional neural networks developed for image segmentation, such as the U-Net autoencoding framework. Instances of such sound source separators for a specific signal category can be configured to estimate a pair of separated signals, where one signal contains all foreground objects belonging to the category, and the other signal contains all background. The estimation of the audio source levels can be performed, for example, by measuring the signal levels of the separated audio components of the mixed signal in the input of the first audio source separation module 605.

[0222] According to some examples, the source separation module 605 can be configured to select a specific instance of the sound source separator by defining one or more sound source separation targets using auxiliary information 604. The auxiliary information 604 can be provided by the user or automatically provided by the control system based on the analysis of the input video and audio data.

[0223] In some such examples, the post-capture editing process (e.g.) Figure 1C The post-capture editing process 124 may involve, for example, a second source separation process performed by one or more servers. This is potentially advantageous because the second source separation process may be more accurate, but it is computationally too intensive for the capture device to perform during the pre-capture phase 114 or the capture phase 116.

[0224] According to some examples, auxiliary information 604 can be or may include audio source labels assigned by the audio source classifier, video object source labels assigned by the video object classifier, or both. In some instances, the audio source classifier and the video object classifier can be references. Figure 5B One or more of those described. For example, in some instances, an audio source classifier can be configured to perform... Figure 5B Frame 558a Figure 5B Frame 562 Figure 5B The bounding box 564 or a combination thereof. In some instances, the video object classifier can be configured to perform... Figure 5B Frame 558b Figure 5B The box 560 or both. In some examples, the separated audio source data 607a-607n may include separated audio components of the received audio data 601, as well as "tags" or other identification data, annotations, etc.

[0225] According to some implementations, the first audio source separation module 605 can be configured to identify audio scene foreground components and audio scene background components. In some such examples, the separated audio source data 607a-607n may include “tags” or other identifying data, annotations, etc., indicating which separated audio components of the received audio data 601 are estimated as audio scene foreground components and which are estimated as audio scene background components. Accordingly, in some instances, the first audio source separation module 605 can be configured to implement what is referred to herein as an “audio tagger.”

[0226] As disclosed elsewhere in this document, in some examples, the audio tagger can be or can include a convolutional neural network (CNN) implemented by a control system. Some such audio taggers can operate on a spectral representation of the audio signal (e.g., a Mel spectrogram) and can update a set of tags and the confidence intervals of these tags at a predefined time resolution (e.g., 1 second). The CNN can be trained to provide an output including an indication of the tags corresponding to audio objects present in the segment (within a large set of all possible tags) and another output representing the confidence level of these tags. The output representing the confidence level can be, for example, an object detection probability, such as a value between zero and one, where a value closer to 1.0 would indicate a high confidence level, while a value closer to zero would indicate a low confidence level. This confidence level can be determined based on an estimated object detection probability, which can be computed by the same CNN. According to some examples, the audio tagger can be implemented as a trained neural network built according to the so-called YAMNet architecture, as described in more detail elsewhere in this document.

[0227] In this example, the audio source level estimation module 615 is configured to estimate the audio signal level of each audio component of the separated audio source data 607a-607n and output the corresponding audio signal levels 617a-617n to the analysis module 625. In some examples, the separated audio source data 607a-607n may also be provided to the analysis module 625. According to some examples, the audio signal levels 617a-617n may be or may include instantaneous audio signal levels. In some examples, the audio signal levels 617a-617n may be or may include averaged audio signal levels over a certain time interval, which may be a fixed time interval or a variable time interval. According to some examples, the time interval may be on the order of tens of milliseconds, hundreds of milliseconds, seconds, tens of seconds, etc. According to some implementations, the audio source level estimation module 615 may be configured to find the distribution of signal levels (or signal energy) among all foreground components of the audio scene and the overall signal level of the background components of the audio scene. In some implementations, the first audio source separation module 605 may be configured for both audio source separation and audio source level estimation.

[0228] According to this example, the video object classifier 608 is configured to detect objects in the currently acquired video data. The video object classifier 608 can be implemented, for example, by a CNN. In some examples, the video object classifier 608 can be one or more of the methods described by J. Redmon et al. in “You Only Look Once: Unified, Real-Time Object Detection” (arXiv:1506.02640v5 [cs.CV] May 9, 2016), which is incorporated herein by reference, or other applicable video-based object detection methods.

[0229] In this example, the video object detection module 610 is configured to determine and output video objects based at least in part on video data 602 received from the camera system and auxiliary information 604 received. v The corresponding estimated video object data 612a-612v for each of the separated video objects. In some examples, v It can be an integer of two or greater. As noted elsewhere in this document, in some examples, the auxiliary information 604 can be or can include the video object source label assigned by the video object classifier. In some instances, the video object classifier can be configured to perform... Figure 5B Frame 558b Figure 5B The frame is 560 or both.

[0230] According to this example, the video object confidence estimation module 620 is configured to estimate the confidence level of each of the video objects corresponding to the video object data 612a-612v output by the video object detection module 610, and output the corresponding video object confidence levels 622a-622v to the analysis module 625. In some examples, the video object confidence levels 622a-622v can vary from zero to one, where zero is the lowest confidence level and one is the highest confidence level. Other examples may use other numerical ranges. In some examples, the video object data 612a-612r may also be provided to the analysis module 625.

[0231] In this example, the analysis module 625 is configured to output audio scene analysis information 627 to the GUI module 630 based at least in part on the received video object confidence levels 622a-622v and audio signal levels 617a-617n. In some examples, the audio scene analysis information 627 may be based at least in part on auxiliary information 604. According to some examples, the audio scene analysis information 627 may be based at least in part on video data 602 received from the camera system, video object data 612a-612r, separate audio source data 607a-607n, or a combination thereof. The analysis module 625 may be implemented, for example, by one or more CNNs.

[0232] Audio scene analysis information 627 may include, for example, audio source inventory data and audio scene analysis metadata. In some examples, the audio source inventory data may include something similar to a reference... Figure 5D The data structure described is the audio source inventory data structure. Audio scene analysis metadata may include, for example, audio source tags, audio source level data, audio source location data, and combinations thereof. In some examples, audio source level data may include data indicating the audio source signal of each audio source or group of audio sources within a certain time interval (see [link to documentation]). Figure 4D and Figure 4E The average audio source level indicator 433 describes this. Audio source location data may include, for example, location data relative to the current video scene, allowing audio object locations to be associated with one or more corresponding people, objects, etc., within the video scene. In some examples, audio scene analysis information 627 may include audio source category information grouping audio sources in the audio source list into two or more audio source categories, which may include a foreground category and a background category. See below for further details. Figure 7 Further details are provided for an example of the analysis module 625.

[0233] According to this example, GUI module 630 is configured to perform audio scene visualization and user input collection via one or more types of GUIs. Accordingly, in this example, GUI module 630 is configured to control one or more displays of a display system to present one or more types of GUIs. In some examples, one or more types of GUIs may be overlaid on an image corresponding to the video data 602 received during the capture phase or pre-capture phase, for example, as shown in Reference 1. Figures 3A to 5A As described, at least one of the GUIs may include an audio source image corresponding to at least one audio characteristic of one or more audio sources indicated by the audio scene analysis information 627.

[0234] In this example, at least one GUI provided by GUI module 630 includes one or more user input areas for receiving user input. Accordingly, in this example, GUI module 630 is configured to receive user input 628, for example, via one or more touchscreen locations corresponding to one or more user input areas. In some examples, user input 628 may indicate that the user expects to increase or decrease the level of one or more audio sources, such as the level of a speaker whose image is presented simultaneously with the GUI. In some instances, user input 628 may indicate a selected potential sound source or a selected candidate sound source for enhancing audio capture. Enhancing audio capture may, for example, involve replacing a selected candidate sound source with external or synthesized audio, adding external or synthesized audio to a selected potential sound source, or a combination thereof. In some examples, enhancing audio capture may involve a beamforming process for enhancing the audio of a selected candidate sound source.

[0235] According to this example, GUI module 630 (or another module implemented by control system section 606) is configured to generate and store metadata corresponding to one or more types of received user input. In some examples, at least one of the one or more types of GUIs can be similar to Figure 3A and Figures 4A to 4E One or more of the GUIs shown.

[0236] According to this example, the file generation module 635 is configured to generate an audio asset stream 637 that includes audio data 601 and video data 602 received during the capture phase. According to some embodiments, the file generation module 635 (or another component of the control system section 606) may be configured to store audio assets corresponding to the audio asset stream 637. In some embodiments, the audio asset stream 637 may include user input metadata, audio scene analysis metadata, or both. According to some embodiments, the audio asset stream 637 may include modified audio data, modified video data, or both modified during the capture phase.

[0237] The context of an audio scene can be associated with one or more assumptions about the user's intent regarding a particular scene (e.g., about which audio sources(s) the user considers most important). This intent can be conveyed through the user's interaction with the GUI (e.g., by...). Figure 6 The user's intent regarding a particular scene can be indicated through interaction with the GUI module 630. For example, the intent of the capturing device user regarding a particular scene can be indicated by received user input instructing the user to focus on a detected human speaker in the scene. In some examples, the user input may instruct the user to pay relatively more attention to another type of sound source or potential sound source (such as a musician, animal, vehicle, aircraft, fountain, or waterfall). In some instances, user input instructing the user's attention may be received via the GUI, such as input regarding changes in the audio source (e.g., input regarding a desired increase in signal level). Alternatively or additionally, the user input instructing the user's attention may be, or may correspond to, user control of the camera system, such as zooming in on a particular video object (such as a person), centering a particular video object in the current video frame, etc.

[0238] In some implementations, where user input is absent or, in addition to indicating a primary or potential sound source of user attention, the control system (such as the control system of a capture device or capture system) can be configured to select the most likely assumption about the context of the audio scene as perceived by the control system. Alternatively or additionally, in some examples, even if the control system has received one or more previous indications of user intent, it can assess the most likely assumption about the current context of the audio scene and whether the user's focus of attention may have changed. For example, the control system can provide such a contextual assumption assessment when a defined time interval has elapsed since the last explicit indication of user intent, or when the indication of user intent is ambiguous.

[0239] Figure 7 Example components of an analysis module configured for contextual hypothesis evaluation are shown, based on some publicly available examples. In this example, analysis module 625 is... Figure 6 An example of analysis module 625 is implemented by control system section 606 and configured to generate audio scene analysis information 627. According to this example, analysis module 625 includes audio scene analysis module 705, context hypothesis evaluation module 710, and audio scene list filtering module 715.

[0240] In this example, the audio scene analysis module 705 is configured to generate scene analysis data 707 based at least in part on auxiliary information 604, the audio signal levels 617a-617n corresponding to the audio source an, and the video object confidence levels 622a-622v corresponding to the video object av. In some examples, the audio scene analysis module 705 may be implemented via a CNN. According to some examples, the audio scene analysis module 705 may be configured to generate audio scene analysis information 627 based at least in part on video data 602 received from the camera system. The auxiliary information 604 may include, for example, audio source labels assigned by an audio source classifier, video object source labels assigned by a video object classifier, or both. The scene analysis data 707 may include, for example, a set of audio scene components, which may be referred to as the "first set of audio scene components". In some examples, the scene analysis data 707 may include audio source labels, audio source levels, and audio source coordinates (e.g., relative to an image in the input video data).

[0241] According to some examples, the audio scene analysis module 705 can be configured to estimate the correspondence between an audio source an and a video object av, or the absence of such a correspondence. For example, the audio scene analysis module 705 can be configured (e.g., based on video data 602) to detect that the mouth of one or more individuals in the audio scene foreground is currently moving when one or more audio sources corresponding to speech are detected in audio data 601 received from the microphone system. The audio scene analysis module 705 can be configured to estimate which current “speaker” audio source corresponds to each of the one or more individuals whose mouth is currently moving. Similarly, in some instances, two audio scene foreground speakers may be detected in audio data 601, but only one individual in the audio scene foreground may currently have their mouth moving. In some examples, the audio scene analysis module 705 can estimate that the other audio scene foreground speaker is the interviewer and / or the person currently operating the capture device. However, in some alternative examples, the audio scene analysis module 705 can simply classify unseen foreground speakers as people speaking outside the video scene, such as… Figure 5D As shown. This type of situation is similar to... Figure 4D and Figure 4E This corresponds to the "Speaking Off-Screen" audio source tag example in GUI 400.

[0242] In some implementations, the audio scene analysis module 705 can be configured to determine whether one or more aspects of the current audio are inconsistent with video-based analysis. In some such examples, scene analysis data 707 can indicate whether one or more aspects of the current audio are inconsistent with video-based analysis. According to some such examples, the audio scene analysis module 705 can be configured to compare an audio source level with a video object confidence level. For example, if the video object confidence level of a particular object is greater than a video object confidence threshold and the corresponding audio source signal level is less than an audio source signal threshold, then the scene analysis data 707 can indicate (e.g., in metadata, by setting a flag, by setting a value, etc.) that there is currently an inconsistency between the audio-based analysis and the video-based analysis of a particular audio source or potential audio source. This inconsistency can be an indication that one or more aspects of the audio capture system are operating close to or below a desired performance threshold.

[0243] According to this example, the context hypothesis evaluation module 710 is configured to estimate the current scene context hypothesis 712 based at least in part on scene analysis data 707 and auxiliary information 604. According to some examples, the context hypothesis evaluation module 710 may be configured to access a memory in which a list of context hypotheses is stored.

[0244] The list of context hypotheses may include, for example, one or more contexts relating to capturing audio and video corresponding to a human speaker, one or more contexts relating to capturing audio and video corresponding to a musical instrument, one or more contexts relating to capturing audio and video corresponding to a natural scene, and so on. In some such examples, each hypothesis may be associated with the presence of a set of audio sources (also referred to herein as audio objects or sound sources) typically associated with the corresponding context. For example, if the context is “city sidewalk” and involves capturing audio and video corresponding to a human speaker, the set of audio sources may include foreground speakers, background speakers, cars, buses, trucks, dogs, sirens, etc. In some such examples, the context hypothesis evaluation module 710 may be configured to score each hypothesis based on the presence of the corresponding audio or video object and the confidence level of its detection.

[0245] In some such examples, the context hypothesis evaluation module 710 can be configured to score each hypothesis using a predefined utility function. According to some such examples, the context hypothesis evaluation module 710 can be configured to select the hypothesis whose utility function maximizes its value. For example, a scene can be evaluated given audio scene analysis data 707 and auxiliary information data 604. utility function In this context, audio scene analysis data is used to obtain a set of detection probabilities for (typically) audio sources associated with the scene (expected). Furthermore, auxiliary information data is used to obtain a set of (typically) expected non-audio objects (e.g., objects detected in the video feed) associated with this scenario. Typically, sets and It is predefined for possible scenario categories (by listing typical objects expected for such scenarios). However, it includes... and The number of objects in the dataset may vary depending on the scenario. (The last part, "will use," appears to be a fragment and doesn't translate directly. It's left as is.) and This indicates the number of objects in the corresponding collection, and uses... The objects are represented as being included in the set X. Then, the utility function can be computed as... + ,in, Represents the detection probability of an object If in the scene No object was detected in = 0 Then, the detection probability can be estimated in the audio scene analysis module 705. The context evaluation module 701 may contain several predefined functions. Each predefined function and set and Different numbers of objects are associated within. However, due to the... and Normalization allows for comparison of different functions, and box 701 can classify scenarios based on the achievement of the maximum value. The scene is represented by a practical function. For example, a scene associated with capturing in a restaurant environment may include the expected objects. (e.g., [human speakers, restaurant noise, background music, etc.]), and scenes associated with street capture can be linked to the expected objects. (For example, [street noise, car horns, human speakers, wind, etc.]) are associated. Similarly, for scenes captured and associated with restaurant environments, this set... This can include objects such as [tables, people, etc.], and for scenes associated with street capture, this collection... This can include objects such as [cars, buses, streetlights, etc.]. A predefined set of typical objects can be defined for the most likely scenario categories where audio capture will be performed (e.g., "outdoor nature," "street," "indoor restaurant," "sports," "default," etc.).

[0246] In this example, the audio scene list filtering module 715 is configured to generate audio scene analysis information 627 based at least in part on auxiliary information 604, scene analysis data 707, and the current scene context hypothesis 712. In some examples, the audio scene list filtering module 715 may be implemented via a CNN. The audio scene analysis information 627 may be or may include content referred to herein as the "second set of audio scene components." In some examples, the audio scene analysis information 627 may include audio source labels, audio source levels, and audio source coordinates (e.g., relative to images in the input video data) for the second set of audio scene components.

[0247] In some instances, the second set of audio scene components may be a subset of the first set of audio scene components in the scene analysis data 707. According to some examples, the audio scene list filtering module 715 may be configured to prioritize and / or rank the first set of audio scene components to determine the second set of audio scene components based on current or recent audio scene activity, current or recent video scene activity, and current or recent user indications. User indications may include user input, the current frame of the video scene, or both. User indications may, for example, correspond to a person or object at the center of the video scene, a person or object currently focused on by the camera system, etc.

[0248] In some instances, at least some of the second set of audio scene components can be represented in the GUI provided by the GUI module 630. Figure 4D and Figure 4E The audio source information areas 430a, 430b, 430c, and 430d are examples. In some instances, the second set of audio scene components (or a subset of the second set of audio scene components currently shown in the GUI) can vary over time based on whether the audio source is currently providing audio estimated by the control system to be audible, or whether the audio source has provided audio estimated by the control system to be audible during the most recent time interval, etc. For example, Figure 4B GUI 401a indicates that the bus is one of the audio scene components. If no bus sound is detected within a certain time interval, the control system can be configured to remove the audio source information region 420c from GUI 401a, and in some examples, remove the audio source information region from the current "second group of audio scene components".

[0249] Figure 8This is a flowchart outlining various example methods 800 according to some disclosed embodiments. Example method 800 can be divided into blocks, such as blocks 805, 810, 815, and 820. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 800 need not be performed in the indicated order. In some embodiments, one or more blocks of method 800 can be performed simultaneously. Furthermore, some embodiments of method 800 may include more or fewer blocks than those shown and / or described. The blocks of method 800 can be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0250] Box 805 relates to "the control system retrieving an audio data asset file saved from a previous capture stage from a memory system." The control system may be, or may include, [the following]. Figure 1A The CPU 141. The control system may, for example, include a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. According to some examples, the control system may be a server's control system. In some examples, audio data asset files may correspond to those generated by… Figure 6 The file generation module 635 generates and stores the audio asset stream 637 in the storage system. Accordingly, the audio data asset file may include audio and video data received during the capture phase, user input metadata, audio scene analysis metadata, etc. Processing can continue to box 810.

[0251] Box 810 relates to “implementing a second audio source separation process on audio data from an audio data asset file by a control system.” In some examples, the second audio source separation process can be relatively more precise than the audio source separation process used during the capture phase. According to some examples, the second audio source separation process can be relatively more computationally intensive than the audio source separation process used during the capture phase. In some instances, the second audio source separation process can be implemented by a neural network trained via a regression-based process, which can be similar to a reference process. Figure 6The described first audio source separation process is a regression-based process, representing some publicly known examples. In some alternative examples, the second audio source separation process may be implemented by a generative neural network. In some embodiments, the second audio source separation process may be implemented by a control system (e.g., a server) based on an algorithm comprising a set of dedicated generative separators. In some examples, each of these dedicated generative separators may be optimized for a specific audio signal category. In some examples, metadata about audio objects created by the capture system (such as scene analysis metadata, user input metadata, or both) may be used in the second audio source separation process to select an instance of a dedicated audio separation algorithm from the set of dedicated generative separators. Processing may continue to box 815.

[0252] Box 815 relates to “remixing an audio scene created by a control system.” According to this example, the remixing will include audio sources and associated audio data generated by the second audio source separation process in Box 810. In some examples, the remixing process in Box 815 (or the process preceding the remixing process in Box 815) may be at least partially based on user input metadata. For example, the user input metadata may include metadata corresponding to a desired level increase for a particular audio source. This level increase may be implemented in Box 815 (or another box in method 800). In some examples, Box 815 (or another box in method 800) may involve replacing candidate sound sources (e.g., replacing the candidate sound source indicated by the user input metadata with external or synthesized audio). Processing may continue to Box 820.

[0253] Box 820 relates to "storing the remix of an audio scene as an updated audio asset file by the control system." According to this example, box 820 relates to storing the actual output of remix box 815 or a modified version of the output of box 815. In some examples, the updated audio asset file can be a backward-compatible audio asset file, which may include the output of remix block 815, context metadata, and a copy of the audio data received in box 805 from the audio data asset file. According to some examples, the updated audio asset file can be a standard media container, such as an MP4 file, an IVAS file, or a Dolby AC-4 file.

[0254] Figure 9This is a flowchart outlining various example methods 900 according to some disclosed embodiments. Example method 900 can be divided into blocks, such as blocks 905, 910, 915, 920, 925, 930, and 935. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 900 need not be performed in the indicated order. In some embodiments, one or more blocks of method 900 can be performed simultaneously. Furthermore, some embodiments of method 900 may include more or fewer blocks than those shown and / or described. The blocks of method 900 can be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0255] Box 905 relates to "receiving audio data from a microphone system by the control system of the device". The control system may be or may include Figure 1A CPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. In some examples, the microphone system may be a microphone system of a mobile device that includes a control system. Processing may continue to box 910.

[0256] Box 910 relates to "receiving video data from a camera system by a control system". In some examples, the camera system may include one or more cameras of a mobile device that includes a control system. Processing may continue to box 915.

[0257] Box 915 relates to “creating an audio source inventory by a control system based at least in part on audio data, video data, or both.” Audio sources may also be referred to herein as sound sources. In some examples, the audio source inventory may be generated via a method similar to [reference needed]. Figure 5D The described audio source list data structure is provided or stored (at least temporarily) as such a data structure. In some examples, the audio source list may serve as a reference. Figure 6 The audio scene analysis information described in section 627 is provided, or is used as a reference. Figure 7 The scene analysis data described in section 707 is provided. The audio source list may include one or more actual audio sources and one or more potential audio sources. Processing can continue to box 920.

[0258] Box 920 relates to “selecting one or more subsets of audio sources from a list of audio sources by a control system.” In some examples, box 920 may relate to selecting one or more audio sources currently receiving audio data indicating that one or more audio sources are currently emitting sound. According to some examples, box 920 may relate to selecting one or more audio sources based on user input. In some examples, selection may involve the control system estimating which audio sources in the list of audio sources are the most important sound sources. The subset of one or more selected audio sources may include the audio sources estimated to be the most important sound sources. According to some examples, selection may involve estimating which audio sources in the list of audio sources correspond to a speaker. The subset of one or more selected audio sources may include the audio sources estimated to be the speaker. Processing may continue to box 925.

[0259] Box 925 relates to “estimating at least one audio characteristic of at least one or more selected audio sources based on audio data by a control system.” In some examples, box 925 may relate to estimating the current level of each of the one or more selected audio sources, the average level of each of the one or more selected audio sources, one or more other audio characteristics of the at least one or more selected audio sources, or a combination thereof. Processing may continue to box 930.

[0260] Box 930 relates to "audio and video data received during the capture phase and stored by the control system". Processing can proceed to box 935.

[0261] Box 935 relates to "a display controlled by a control system to display images corresponding to video data, and to display a graphical user interface (GUI) overlaid on these images before and during the capture phase, wherein the GUI includes audio source images corresponding to at least one audio characteristic of a subset of one or more selected audio sources." In some examples, box 935 may relate to a display of control device 101 to present, for example... Figure 3A The GUI shown or Figures 4C to 4E One of the GUIs shown. Audio source level information regions 432a, 432b, 432c, and 432d within audio source information regions 430a, 430b, 430c, and 430d respectively, and the average audio source level indicator 433 are examples of audio source images corresponding to at least one audio characteristic.

[0262] In some examples, the GUI may include one or more user input areas configured to receive user input. Figure 4C The virtual sliders 440a and 440b are examples of GUI user input areas configured to receive user input. In other examples, one or more user input areas may include another type of virtual slider, virtual knob, virtual dial, etc.

[0263] According to some examples, method 900 may involve a control system classifying audio sources in an audio source list into two or more audio source categories. In some such examples, the GUI may include a user input area portion corresponding to at least one of the two or more audio source categories. In some examples, one of the audio source categories may be a foreground category corresponding to one or more selected audio sources. According to some examples, one of the audio source categories may be a background category corresponding to one or more audio sources in the audio source list that are not in a subset of one or more selected audio sources.

[0264] In some examples, method 900 may involve a first sound source separation process performed by a control system. In some such examples, method 900 may involve the control system updating the audio scene state at least in part based on the first sound source separation process, and the control system causing the GUI to be updated according to the updated audio scene state. According to some examples, the creation process of box 915 may be at least in part based on the first sound source separation process. Some disclosed methods may involve performing post-capture audio processing on audio data received during the capture phase. In some examples, post-capture audio processing may involve a second sound source separation process that is more complex than the first sound source separation process.

[0265] According to some examples, method 900 may involve a control system detecting one or more potential sound sources based at least in part on video data. In some such examples, at least one of the one or more potential sound sources may not be indicated by audio data. For example, at least one of the one or more potential sound sources may be indicated by video data. According to some examples, the list of audio sources may include one or more potential sound sources.

[0266] In some examples, method 900 may involve a control system detecting one or more candidate sound sources for enhanced audio capture. In some such examples, enhanced audio capture may involve replacing the candidate sound sources with external or synthesized audio. According to some examples, the GUI may include at least one user input area configured to receive user selection for a selected potential sound source or a selected candidate sound source. In some examples, the GUI may include at least one user input area configured to receive user selection for enhanced audio capture. Enhanced audio capture may, for example, include external audio, synthesized audio, or both of the selected potential sound source or the selected candidate sound source.

[0267] Some of the disclosed methods involve providing a GUI including at least one user input area configured to receive user selection of a ratio between enhanced audio capture and real-world audio capture. In some instances, the GUI may be provided during a post-capture editing process.

[0268] According to some examples, method 900 may involve a control system causing a display to show an audio source label in a GUI. Figure 4D and Figure 4E Examples of audio source labels are provided in audio source information areas 430a, 430b, 430c, and 430d. In some examples, at least one of the audio source labels may correspond to an audio source identified by the control system based on audio data, video data, or both.

[0269] In some examples, method 900 may involve updating the estimate of the current audio scene by the control system, and causing the GUI to be updated based on the updated estimate of the current audio scene by the control system. Updating the estimate of the current audio scene may involve the control system implementing an audio classifier, a video classifier, or both. In some examples, updating the estimate of the current audio scene may involve the control system implementing an audiovisual classifier. According to some examples, the updated estimate of the current audio scene may include updated level estimates of one or more audio sources.

[0270] Figure 10A This represents an element that indicates an audio source list according to some publicly available implementations. In some examples, audio source list 1000 can be... Figure 5B In box 564 or box 565, or in Figure 9 The audio source list 1000 is generated in box 915, as shown in some examples. According to some references, the audio source list 1000 can be generated via a similar method. Figure 5D The described audio source list data structure is provided or stored (at least temporarily) as such a data structure. In some examples, audio source list 1000 can be used as a reference. Figure 6 The audio scene analysis information described in section 627 is provided, or is used as a reference. Figure 7 A portion of the scene analysis data 707 described is provided. In these examples, the audio source list 1000 includes audio sources 1001 currently present in the received audio data and potential audio sources 1004 not currently present in the received audio data (or present but the corresponding audio data is below a threshold level, but has been detected in the video feed). Audio sources 1001 include a subset of selected audio sources 1002, which have been chosen for possible enhancement or replacement.

[0271] In some instances, the selected audio source 1002 may have been selected at least in part based on user actions, such as user input via a GUI, user zooming in on a specific video object, or user framing (e.g., centering) a specific video object. According to some examples, the selected audio source 1002 may have been selected by a control system (e.g., the control system of a capture device). In some such examples, the control system may have selected one or more audio sources whose corresponding audio data has low quality (e.g., background music partially masked by background noise), which is estimated to be close to or below human audibility levels. Alternatively or additionally, the control system may select one or more potential or actual audio sources at least in part based on the estimated audio or video scene context. For example, if the context is a “natural scene”, animals currently not producing sound or producing sound below a threshold in the video feed may be selected for possible enhancement or replacement.

[0272] The number, type, and order of elements shown in the audio source list 1000 are merely examples. For instance, in some instances, the potential audio source 1004 may include a subset of potential audio sources selected for association with synthetic audio data or external audio data. The subset of potential audio sources may be selected, for example, based on estimated audio scene context, user preferences, or a combination thereof.

[0273] Figure 10B This is a flowchart outlining various example methods 1005 according to some disclosed implementations. Example method 1005 can be divided into blocks, such as blocks 1010, 1015, 1020, 1025, 1030, 1035, and 1040. Each block can be described as an operation, process, method, step, action, or function. According to some examples, Figure 10B The frames are executed during the capture phase. As with other methods described herein, the frames of method 1005 do not necessarily need to be executed in the indicated order. In some implementations, one or more frames of method 1005 may be executed simultaneously. Furthermore, some implementations of method 1005 may include more or fewer frames than those shown and / or described. The frames of method 1005 may be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0274] Box 1010 relates to "receiving audio data from a microphone system and video data from a camera system by the device's control system." The control system may be or may include... Figure 1ACPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. In some examples, the microphone system may be a microphone system of a mobile device that includes a control system. According to some examples, the camera system may include one or more cameras of a mobile device that includes a control system. In other examples, the microphone system, camera system, or both may reside in one or more devices other than the device that includes the control system. Processing may continue to box 1015.

[0275] Box 1015 relates to "creating an audio source inventory by a control system based at least in part on audio data, video data, or both." In some examples, the audio source inventory can be... Figure 10A The audio source list corresponds to 1000. Accordingly, in some examples, the audio source list can be derived via a similar reference. Figure 5D The described audio source list data structure is provided or stored (at least temporarily) as such a data structure, and in some examples, it can be used as a reference. Figure 6 The audio scene analysis information described in section 627 is provided, or, in some examples, may be used as a reference. Figure 7 The scene analysis data described is provided in part 707. Processing can continue to box 1020.

[0276] Box 1020 relates to “a display of a device controlled by a control system to provide a graphical user interface (GUI), which includes representations of at least some audio sources in an audio source list.” In some examples, box 1020 may relate to a display of control device 101 to present, for example… Figure 3A The GUI shown or Figures 4C to 4E One of the GUIs shown. Audio source information areas 430a, 430b, 430c, and 430d are examples of "representations of at least some audio sources in the audio source list." In some examples, box 1020 or another box of method 1005 may involve selecting a subset of audio sources in the audio source list. According to some such examples, Figure 7 The audio scene list filtering module 715 can select a subset of audio sources. Processing can continue to box 1025.

[0277] Box 1025 relates to "receiving user input via a GUI from a control system regarding enhancement or replacement of audio data corresponding to one or more selected audio sources". In some examples, the GUI may include one or more user input areas configured to receive user input. According to some examples, box 1025 may relate to receiving user input via a touch person 405a, one of other video objects, or... Figure 4A and Figure 4B The audio source information area 420a, 420b, or 420c shown is used to receive user input. In some examples, box 1025 may involve receiving user input via a person touching 405b, another video object, or... Figure 4C The audio source information area shown is either 420d or 420e, used to receive user input. According to some examples, box 1025 may involve receiving input via a touch person 405a, one of other video objects, or... Figure 4D and Figure 4E The audio source information area shown is one of 430a, 430b, 430c, or 430d to receive user input. Processing can continue to box 1030.

[0278] Box 1030 relates to "the creation of metadata corresponding to user input by the control system". Processing can continue to box 1035.

[0279] Box 1035 relates to "the creation of media assets by a control system, which includes metadata, audio data received during the capture phase, and video data." In some examples, Box 1035 may involve... Figure 6 The operation of the file generation module 635. Processing can continue to box 1040.

[0280] Box 1040 relates to “the storage of media assets in memory by a control system”. According to some examples, box 1040 may relate to storing media assets in the memory of a capture device, storing media assets in the memory of another device (such as a storage device for a cloud-based service), or both.

[0281] Figure 10C This is a flowchart outlining an additional example method 1070 according to some disclosed implementations. Example method 1070 can be divided into blocks, such as blocks 1075, 1080, 1085, 1090, and 1095. Each block can be described as an operation, process, method, step, action, or function. According to some examples, Figure 10CThe boxes are executed during the post-capture editing phase. As with other methods described herein, the boxes of method 1070 do not necessarily have to be executed in the indicated order. In some implementations, one or more boxes of method 1070 may be executed simultaneously. Furthermore, some implementations of method 1070 may include more or fewer boxes than those shown and / or described. The boxes of method 1070 may be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0282] Box 1075 relates to “the acquisition of media assets by a control system, the media assets including metadata, audio data received during the capture phase, and video data, the metadata corresponding to enhancement or replacement of audio data corresponding to one or more user-selected audio sources.” The control system may be or may include… Figure 1A The CPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof.

[0283] In some examples, box 1075 may involve obtaining metadata corresponding to a label generated by the control system indicating that one or more audio sources or potential audio sources have been selected for possible enhancement or replacement. In some such examples, whether during the capture phase or during a previous post-capture editing phase, the metadata may correspond to user input indicating the user's selection of one or more audio sources or potential audio sources. According to some examples, the metadata may correspond to user input indicating that the user expects to change at least one audio characteristic of the audio source (such as the audio source's level).

[0284] According to some examples, box 1075 or another box of method 1070 may involve obtaining unmodified audio data corresponding to at least one audio source. In some such examples, box 1075 may involve obtaining unmodified audio data corresponding to all audio sources for which audio data was received during the capture phase. In some examples, box 1075 may involve obtaining modified audio data corresponding to at least one audio source. In some such examples, box 1075 may involve obtaining modified audio data corresponding to at least one audio source whose audio data was previously at least partially enhanced or replaced during the capture phase or a previous post-capture editing phase. Processing may continue to box 1080.

[0285] Box 1080 relates to “obtaining synthesized audio data, external audio data, or both by a control system to enhance or replace audio data corresponding to an audio source selected by a user.” In some examples, box 1080 may relate to obtaining one or more types of replacement of audio data from memory via download or streaming. For example, box 1080 may relate to obtaining audio data corresponding to music, audio data corresponding to pre-recorded sound effects (which may be referred to herein as foley effects), such as a ticking clock effect, a door closing effect, a glass breaking effect, etc. According to some examples, box 1080 may relate to generating synthesized audio data or obtaining the generated synthesized audio data. In some such examples, the synthesized audio data may be generated by a neural network. Processing may continue to box 1085.

[0286] Box 1085 relates to “enhancing or replacing audio data corresponding to a selected audio source by a control system to produce modified audio data including enhanced audio data, replaced audio data, or both.” In some examples, box 1085 may involve increasing or decreasing the level of an audio source, for example, based on metadata corresponding to a user’s desired increase or decrease in the level of the audio source. According to some examples, box 1085 may involve completely replacing the audio data corresponding to the selected audio source, for example, replacing recorded background music with stored, streamed, or downloaded music. In some instances, box 1085 or another box of method 1070 may involve associating the generated, stored, streamed, or downloaded audio with a selected potential audio source. In some such examples where the selected potential audio source is a video object (such as a door, clock, animal, fountain, etc.), the associated audio can be placed in the audio scene such that the apparent position of the associated audio matches the position of the video object. Processing may continue to box 1090.

[0287] Box 1090 relates to "the control system mixing modified audio data with audio data corresponding to other audio sources in a media asset to produce a modified media asset." According to some examples, another associated process of the mixing process or method 1070 may be an audio source mixing process known to those skilled in the art, which may include equalization, balancing, compression, adding or reducing reverberation effects, etc. In some examples, the mixing may be as described with reference to box 815. Processing may continue to box 1095.

[0288] Box 1095 relates to “the storage of modified media assets in memory by a control system.” According to some examples, Box 1095 may relate to storing media assets in the memory of a capture device, storing media assets in the memory of another device (such as a storage device for a cloud-based service), or both. In some examples, Box 1095 may relate to storing modified and unmodified audio data. According to some examples, Box 1095 may relate to storing one or more types of metadata, such as metadata generated during the capture process, metadata corresponding to aspects of the post-capture editing process, etc. Storing such metadata, along with video data and unmodified audio data, can produce content referred to herein as “backward-compatible media assets.” The audio portion of backward-compatible media assets may be referred to herein as “backward-compatible audio assets.”

[0289] Figure 11A Examples of media assets and hybrid modules are shown, based on some publicly available examples. Figure 11A Media asset 1105, mixing module 1120, and backward-compatible media asset 1130 are illustrated. According to these examples, media asset 1105 includes unmodified audio data 1105, modified unmixed audio data 1110, and context metadata 1115. In these examples, control system portion 1106 is configured to implement mixing module 1120. According to these examples, backward-compatible media asset 1130 includes a copy of the unmodified audio data 1105, modified mixed audio data 1125 output by mixing module 1120, and a copy of context metadata 1115. In this example, the modified mixed audio data 1125 is in a format playable on various consumer devices, such as IVAS, Dolby AC-4, MP-4, etc.

[0290] In some examples, media asset 1105 may have been stored after the capture phase, while in other examples, it may have been stored after the post-capture editing phase. Accordingly, the modified mixed audio data 1125 may have been modified (at least partially) during the capture phase, during the post-capture editing phase, or both. Modifications may involve enhancement, replacement, or both.

[0291] According to these examples, mixing module 1120 is configured to mix modified unmixed audio data 1110 based on input 1117, which may include user input, default settings, or both. User input may correspond, for example, to user input received during a post-capture editing process. In some instances, user input may be received in response to a user preview of one or more audio files in the modified unmixed audio data, a previous mix produced by the mixing module, etc. According to some examples, N audio components may be provided to mixing module 1120. In some such examples, mixing module 1120 may output M components obtained through a linear transformation of the N input components. This linear transformation typically has N×M degrees of freedom (e.g., corresponding to a rectangular transformation matrix). The coefficients of the transformation matrix may be adjusted based on user input, or default values ​​may be used. In some examples, mixing module 1120 may perform the transformation directly on time samples (e.g., pulse code modulation (PCM) samples), while in other examples, mixing module 1120 may perform the transformation on time-frequency slots provided by a filter bank (e.g., quadrature mirror filter (QMF)).

[0292] In these examples, the mixing module 1120 is configured to output modified mixed audio data 1125 and store the modified mixed audio data 1125 as part of a backward-compatible media asset 1130.

[0293] In some examples, the mixing module 1120 can be configured to mix modified unmixed audio data 1110 based on at least a portion of the context metadata 1115. In some examples, at least a portion of the context metadata 1115 can correspond to... Figure 7 The context hypothesis evaluation module 710 outputs the scenario context hypothesis 712, context metadata generated based on user input, or both.

[0294] Based on these examples, the control system section 1106 is also configured to copy unmodified audio data 1105 and context metadata 1115, and store copies of the unmodified audio data 1105 and context metadata 1115 as components of backward-compatible media assets 1130.

[0295] Figure 11BThis is a flowchart outlining various example methods 1150 according to some disclosed embodiments. Example method 1150 can be divided into blocks, such as blocks 1155, 1160, 1165, 1170, 1175, and 1180. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 1150 need not be performed in the indicated order. In some embodiments, one or more blocks of method 1150 can be performed simultaneously. Furthermore, some embodiments of method 1150 may include more or fewer blocks than those shown and / or described. The blocks of method 1150 can be generated by one or more devices (e.g., Figure 1A The device shown will be used for execution.

[0296] Box 1155 relates to "receiving audio data from a microphone system by the control system of the device". The control system may be or may include Figure 1A CPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. In some examples, the microphone system may be a microphone system of a mobile device that includes a control system. Processing may continue to box 1160.

[0297] Box 1160 relates to "receiving video data from a camera system by a control system". In some examples, the camera system may include one or more cameras of a mobile device that includes a control system. In other examples, a microphone system, a camera system, or both may reside in one or more devices other than the device that includes the control system. Processing may continue to box 1165.

[0298] Box 1165 relates to “creating an audio source inventory by a control system based at least in part on audio data, video data, or both.” In some examples, the audio source inventory may be as referenced herein. Figure 10A As described in this article, as referenced Figure 10B The audio source list may be as described in box 1015, or both. Accordingly, in some examples, the audio source list may include one or more actual audio sources and one or more potential audio sources. Processing may continue to box 1170.

[0299] Box 1170 relates to “the control system selecting one or more selected audio sources from a list of audio sources, wherein the one or more selected audio sources are selected for possible enhancement or replacement.” In some examples, the one or more selected audio sources may have been selected at least in part based on user actions, such as user input previously received via the GUI, user zooming in on a specific video object, user framing of a specific video object (e.g., centering), etc.

[0300] In some examples, selection may involve estimating which audio sources in the list correspond to the speaker. In some such examples, one or more selected audio sources do not include those estimated to be the speaker.

[0301] Alternatively or additionally, in some examples, at least one audio source may have been selected by a control system (such as the control system of a capture device). In some such examples, the control system may have selected one or more audio sources whose corresponding audio data has low quality (such as background music partially masked by background noise), which is estimated to be close to or below human audibility levels, etc. According to some examples, the control system may select one or more potential or actual audio sources at least in part based on the estimated audio or video scene context. In some such examples, at least one audio source may be selected at least in part based on video data, for example, by selecting video objects as actual or potential audio sources. According to some such examples, Figure 7 The audio scene list filtering module 715 can be configured to at least partially perform box 1170. Processing can then proceed to box 1175.

[0302] Box 1175 relates to "audio and video data received during the capture phase and stored by the control system". Processing can continue to box 1180.

[0303] Box 1180 relates to "a display of a device controlled by a control system to display images corresponding to video data and to display a graphical user interface (GUI) overlaid on these images, wherein the GUI indicates one or more selected audio sources." The GUI may be displayed before the capture phase, during the capture phase, or in both cases. In some instances, the GUI may include audio source labels. At least one of the audio source labels may correspond to an audio source or potential audio source identified by the control system based on audio data, video data, or both.

[0304] In some examples, box 1180 may relate to a display of control device 101 to present, for example... Figure 3A The GUI shown or Figures 4C to 4EOne of the GUIs shown. Audio source information areas 430a, 430b, 430c, and 430d are examples of GUIs indicating "one or more selected audio sources." In some examples, the GUI may present a representation of a potential audio source whose audio data is not currently being received (or whose audio data is currently being received at a level below a threshold), but which is still a video object that has been identified as a potential audio source by the control system. According to some examples, the GUI may represent one or more audio sources that are selected differently from other audio sources (e.g., in a different color) for possible enhancement or replacement.

[0305] In some examples, the GUI may include one or more user input areas configured to receive user input. In some such examples, an audio source information area may be configured to receive user input, for example, to allow a user to select one or more audio sources for enhancement or replacement. According to some examples, the GUI may display one or more audio sources selected for possible enhancement or replacement, along with text prompts associated with the enhancement or replacement of audio data corresponding to the one or more selected audio sources. For example, the GUI may display one or more selected audio sources and prompts such as "Modify?" or "Enhance or Replace?".

[0306] In some instances of method 1150, the GUI may display one or more selected audio sources, along with prompts regarding the enhancement of audio data corresponding to the selected audio sources, such as indicating the proposed enhancement type. In some such examples, the proposed enhancement type may involve a microphone beamforming process for enhancing the audio data corresponding to the selected audio sources. According to some examples, the GUI may display one or more selected audio sources, along with prompts regarding the replacement of audio data corresponding to the selected audio sources. For example, the prompts may indicate the proposed enhancement type, such as replacing the audio data corresponding to the first selected audio source with synthesized audio data or with external audio data.

[0307] Some examples of method 1150 may involve the control system receiving user input via a GUI instructing the enhancement or replacement of audio data corresponding to a selected audio source. Some such examples may involve the control system enhancing or replacing the audio data corresponding to the selected audio source to produce modified audio data. The modified audio data may be enhanced audio data or replaced audio data. Some examples may involve the control system tagging the enhanced or replaced audio data and storing the tag along with the enhanced or replaced audio data. The tag may, for example, be a type of audio metadata.

[0308] Some examples of method 1150 may involve a control system causing a GUI to indicate that the audio data corresponding to the selected audio source is enhanced or replaced audio data; in other words, indicating that the audio data corresponding to the selected audio source has been enhanced or replaced. Some examples of method 1150 may involve a control system causing a GUI to indicate one or more audio sources corresponding to unmodified audio data.

[0309] Some examples of method 1150 may involve the control system storing unmodified audio data corresponding to at least one audio source that has been selected for enhancement or replacement, for example, after the audio data corresponding to the selected audio source has been enhanced or replaced.

[0310] Some examples of Method 1150 allow a user to interpolate between enhanced or replaced audio data and “real-world” or unmodified audio data. Some such examples may involve a control system causing the GUI to indicate an audio source with corresponding enhanced or replaced audio data. Some such examples may involve a control system causing the GUI to include one or more user input areas for receiving user input to modify the enhanced or replaced audio data based on the unmodified audio data. Modification may, for example, involve interpolating between the enhanced or replaced audio data and the unmodified audio data. In some examples, modification may involve replacing the enhanced or replaced audio data with the unmodified audio data. According to some examples, the GUI may be displayed after the capture phase and during the post-capture review process.

[0311] Figure 12A Examples of media assets and interpolators based on some publicly available examples are shown. Figure 12A Media asset 1208, interpolator 1220, and media asset 1218 with a set of artistic intents are illustrated. According to these examples, media asset 1208 includes unmodified audio data 1205, modified unmixed audio data 1210, and contextual metadata 1215. In these examples, control system portion 1206 is configured to implement interpolator 1220. According to these examples, media asset 1218 with a set of artistic intents includes a copy of the unmodified audio data 1205, adjusted modified unmixed audio data 1212 output by interpolator 1220, and a copy of contextual metadata 1215.

[0312] In some examples, media asset 1208 may have been stored after the capture phase, while in other examples, it may have been stored after the post-capture editing phase. Accordingly, the modified unmixed audio data 1210 may (at least partially) have been modified during the capture phase, during the post-capture editing phase, or both. Modifications may involve enhancement, replacement, or both.

[0313] However, in this example, if the post-capture editing phase has already begun, it is not yet complete. Instead, the user provides user input 1203 to the interpolator 1220 to adjust the unmodified audio data 1205, the modified unmixed audio data 1210, or both, to the user's satisfaction. After the user has adjusted the unmodified audio data 1205, the modified unmixed audio data 1210, or both to satisfactorily represent the user's artistic intent, the user can terminate the interpolation process and allow the adjusted modified unmixed audio data 1212 to be output by the interpolator 1220 and stored as part of the media asset 1218 with the artistic intent set.

[0314] According to some examples, interpolator 1220 can be configured to interpolate between modified unmixed audio data 1210 and unmodified audio data 1205 based on user input 1203. Alternatively or additionally, interpolator 1220 can be configured to interpolate between unmodified audio data 1205 and synthesized or external audio data 1207 based on user input 1203. Synthesized or external audio data 1207 may include, for example, music, sound effects, etc., which may be pre-recorded or generated. In some instances, interpolator 1220 can be configured to interpolate at least in part based on context metadata 1215. For example, interpolator 1220 can be configured to propose an initial interpolation amount based on context metadata 1215, which can be modified based on user input 1203 if desired by the user.

[0315] Based on these examples, the control system section 1206 is also configured to copy unmodified audio data 1205 and context metadata 1215, and store copies of unmodified audio data 1205 and context metadata 1215 as components of media assets 1218 with artistic intent sets.

[0316] Figure 12B , Figure 12C and Figure 12D The illustrations show sample elements of the GUI that can be rendered during post-editing processes that include interpolation. In these examples, Figure 12B , Figure 12C and Figure 12DThe GUIs 1260a, 1260b, and 1260c are respectively mounted on the display 1255 of the device 1251, which is a cellular phone and in these examples is... Figure 1A Examples of device 101. In these examples, the control system (not shown) of device 1251 controls the display 1255 to present GUIs 1260a, 1260b, and 1260c. Similar to other disclosed examples, Figures 12B to 12D The types, quantities, and arrangements of the elements shown are provided as examples only.

[0317] Figure 12B The GUI 1260a includes audio source information areas 1230a, 1230b, 1230c, 1230d, and 1230e, and text prompts 1262a and 1262b. In this example, audio source information areas 1230b, 1230d, and 1230e have been selected by the control system as candidates for possible modification or further modification, and are displayed differently from audio source information areas 1230a and 1230c. Furthermore, audio source information areas 1230b, 1230d, and 1230e have associated text prompts 1262b that ask the user if they want to modify the corresponding audio source. According to this example, if the user wants to select the corresponding audio source for modification or further modification, text prompt 1262a encourages the user to touch the audio source area (meaning one of audio source information areas 1230a-1230e).

[0318] Based on this example, it has already responded to detecting that the user is... Figure 12B The audio source information area 1230e is presented by touch. Figure 12C The GUI 1260b. In this example, Figure 12C The GUI 1260b includes a text prompt 1262c and an interpolation control 1264a, which in this example includes a virtual slider 1266. According to this example, the interpolation control 1264a controls the interpolator based on an indicated ratio of modified audio to unmodified audio (e.g.,...). Figure 12A The interpolator 1220 provides a ratio ranging from 0 / 1 (meaning completely unmodified) to 1 / 1 (meaning completely modified). In this example, text prompt 1262c encourages the user to move a virtual slider 1266 to select the desired ratio of modified to unmodified audio. According to some examples, the audio corresponding to the selected ratio may be provided by the speaker of device 1251. In some examples, video may also be presented, such as a video showing a scene including video objects corresponding to the selected audio source. According to some examples, the user may be able to select different ratios for different portions of the audio corresponding to the audio source (e.g., for various time intervals).

[0319] In some examples, this can be done in response to detecting a user in Figure 12B The audio source information area 1230e is presented by touch. Figure 12D The alternative GUI 1260c. Based on this example, Figure 12D The GUI 1260c includes a text prompt 1262d and an interpolation control 1264b, which in this example also includes a virtual slider 1266. According to this example, the interpolation control 1264b controls the interpolator based on an indicated percentage of modified audio to unmodified audio, ranging from 0% (meaning no modification) to 100% (meaning full modification). In this example, the text prompt 1262d encourages the user to move the virtual slider 1266 to select the desired percentage of modified to unmodified audio. According to some examples, the audio corresponding to the selected percentage may be provided by the speaker of device 1251. In some examples, the user may be able to select different percentages for different time intervals of audio corresponding to the audio source. According to some examples, video may also be presented, such as a video showing a scene including video objects corresponding to the selected audio source.

[0320] Figure 12E This is a flowchart outlining various example methods 1270 according to some disclosed embodiments. Example method 1270 can be divided into blocks, such as blocks 1272, 1275, 1277, 1280, 1282, 1285, 1287, 1290, and 1292. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 1270 need not be performed in the indicated order. In some embodiments, one or more blocks of method 1270 can be performed simultaneously. Furthermore, some embodiments of method 1270 may include more or fewer blocks than those shown and / or described. The blocks of method 1270 can be generated by one or more devices (e.g., Figure 1A The method 1270 is performed on the device shown. Some blocks of method 1270 may be performed during the capture phase, and other blocks of method 1270 may be performed during the post-capture editing phase. Depending on the specific implementation, the post-capture editing phase may or may not be performed on the same device used for the capture phase.

[0321] Box 1272 relates to "receiving audio data from a microphone system by the control system of the device". The control system may be or may include Figure 1ACPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. In some examples, the microphone system may be a microphone system of a mobile device that includes a control system. Processing may continue to box 1275.

[0322] Box 1275 relates to "receiving video data from a camera system by a control system". In some examples, the camera system may include one or more cameras of a mobile device that includes a control system. In other examples, a microphone system, a camera system, or both may reside in one or more devices other than the device that includes the control system. Processing may continue to box 1277.

[0323] Box 1277 relates to "creating an audio source inventory by a control system based at least in part on audio data, video data, or both." In some examples, the audio source inventory may be as referenced herein. Figure 10A As described in this article, as referenced Figure 10B The process may involve either or both described in box 1015. Some examples may involve the detection of one or more potential sound sources by a control system based at least in part on video data. Accordingly, in some examples, the list of audio sources may include one or more actual audio sources and one or more potential audio sources. In some examples, at least one potential sound source may not be indicated by audio data. Processing may continue to box 1280.

[0324] Box 1280 relates to "the control system selecting at least a first selected audio source from a list of audio sources, the first selected audio source being chosen for enhancement or replacement." Figure 4D and Figure 4E In the example shown, the selected audio source is detected background music, which is replaced with a streaming version of the same music. According to some examples, method 1270 may involve controlling a display to present an audio data modification GUI before or during the capture phase, the audio data modification GUI including user prompts associated with enhancement or replacement of audio data corresponding to one or more selected audio sources in a list of audio sources. In some examples, the audio data modification GUI may include user prompts associated with replacing audio data corresponding to the selected audio source. Replacement may be associated with replacing the audio data corresponding to the selected audio source with synthesized audio data or with external audio data. Figure 4DThe text prompt provides an example of this type of user prompt. In some examples, an audio data modification GUI may include user prompts, virtual controls, or a combination thereof associated with enhancing or reducing audio data corresponding to the selected audio source. Figure 4D and Figure 4E The plus and minus signs in the audio source information areas 430a, 430b, 430c, and 430d are examples of such virtual controls. According to some examples, enhancement can be associated with a microphone beamforming process used to enhance audio data corresponding to the selected audio source. For example, if a user touches... Figure 4D and Figure 4E One of the plus signs in the text is that, in some implementations, the control system can initiate the microphone beamforming process.

[0325] In some examples, one or more selected audio sources may have been selected at least in part based on user actions, such as user input previously received via the GUI, user zooming in on a specific video object, user framing of a specific video object (e.g., centering), etc. In some examples, selection may involve estimating which audio sources in the list correspond to the speaker. In some such examples, one or more selected audio sources do not include audio sources estimated to be speakers. Alternatively or additionally, in some examples, at least one audio source may have been selected by a control system (e.g., the control system of a capture device). In some such examples, the control system may have selected one or more audio sources whose corresponding audio data has low quality (e.g., background music partially masked by background noise), which is estimated to be close to or below human audibility levels, etc. According to some examples, the control system may select one or more potential or actual audio sources at least in part based on the estimated audio or video scene context. In some such examples, at least one audio source may be selected at least in part based on video data, for example, by selecting a video object as an actual or potential audio source. Processing may continue to box 1282.

[0326] Box 1282 relates to "enhancing or replacing audio data corresponding to a first selected audio source by a control system to generate first modified audio data, the first modified audio data including at least one of first enhanced audio data or first replaced audio data." Figure 4E In the example shown, the selected audio source is detected background music, which is replaced with a streaming version of the same music. This replacement is an example of box 1282. In some examples, box 1282 may involve audio data enhancement, such as a microphone beamforming process for enhancing the audio data corresponding to the selected audio source. Processing can continue to box 1285.

[0327] Box 1285 relates to "the storage of first modified audio data by the control system". Figure 12A The modified unmixed audio data 1210 is an example of stored modified audio data. Processing can continue to box 1287.

[0328] Box 1287 relates to “audio and video data received during the capture phase and stored by the control system, the audio data including first unmodified audio data corresponding to at least a first selected audio source”. Figure 12A Unmodified audio data 1205 is an example of stored unmodified audio data. Processing can continue to box 1290.

[0329] Box 1290 relates to "a display of a device controlled by a control system to present images corresponding to video data and to display a post-capture graphical user interface (GUI) overlaid on these images, wherein the post-capture GUI at least indicates a first selected audio source and one or more user input areas to receive user input." The GUI may be similar to... Figure 12B , Figure 12C or Figure 12D The GUI shown includes one or more additional images of the video data, which at least include the selected audio source. In some instances, the GUI may include audio source tags, such as... Figure 12B The audio source information areas 1230a-1230e contain audio source labels. At least one of the audio source labels may correspond to an audio source or potential audio source identified by the control system based on audio data, video data, or both. In some examples, the post-capture GUI may include at least one user input area configured to receive user selection of a ratio between modified and unmodified audio data (e.g., a ratio between first modified audio data and first unmodified audio data). This at least one user input area may be or may include a slider. The slider may be configured to allow the user to select a ratio between modified and unmodified audio data, a modification percentage, etc. In some examples, the slider may be configured to allow the user to select a ratio from 0% to 100%. Processing may continue to box 1292.

[0330] Box 1292 relates to "editing first modified audio data to include at least a portion of first unmodified audio data based on user input received by the post-capture GUI during a post-capture review process." According to some examples, box 1292 may relate to an interpolation process, such as interpolating between the first modified audio data and the first unmodified audio data. In some examples, box 1292 may relate to a GUI (such as...) Figure 12B , Figure 12C or Figure 12D (GUI) interaction.

[0331] Some examples of method 1270 may involve a control system receiving audio data modification user input via an audio data modification GUI, the audio data modification user input indicating enhancement or replacement of audio data corresponding to a first selected audio source. The editing can efficiently provide the first modified audio data in response to the audio data modification user input. Some examples may involve a control system tagging the enhanced or replaced audio data, and storing the tag along with the enhanced or replaced audio data. The tag may, for example, be a type of audio metadata.

[0332] Some examples of method 1270 may involve a control system causing a GUI to indicate that the audio data corresponding to the selected audio source is enhanced or replaced audio data; in other words, indicating that the audio data corresponding to the selected audio source has been enhanced or replaced. Some examples of method 1270 may involve a control system causing a GUI to indicate one or more audio sources corresponding to unmodified audio data.

[0333] Some examples of method 1270 may involve a control system causing a post-capture GUI to indicate that the audio data corresponding to the first selected audio source is modified audio data. Some such examples may involve a control system causing a post-capture GUI to indicate one or more audio sources corresponding to unmodified audio data.

[0334] Some examples of method 1270 may involve the control system causing a display to show audio source labels in an audio data modification GUI or a post-capture GUI. In some such examples, at least one of the audio source labels corresponds to an audio source or potential audio source identified by the control system based on audio data, video data, or both.

[0335] Figure 13 Examples of media assets and interpolators based on some publicly available examples are shown. Figure 13 Media asset 1308, cloud processing system 1325, and backward-compatible media asset 1330 are illustrated. According to these examples, media asset 1308 includes unmodified audio data 1305, modified audio data 1310, and contextual metadata 1315. In these examples, cloud processing system 1325 includes a toolset selection module 1302, an editing toolbox 1304, and a processing module 1312. Cloud processing system 1325 may be implemented, for example, by one or more servers. According to these examples, backward-compatible media asset 1330 includes a copy of unmodified audio data 1305, processed and modified audio data 1325 (which includes processed audio data 1320 output by cloud processing system 1325), and a copy of contextual metadata 1315.

[0336] In some examples, media asset 1308 may have been stored after the capture phase, while in other examples, media asset 1308 may have been stored after the post-capture editing phase. Accordingly, modified audio data 1310 may (at least partially) have been modified during the capture phase, during the post-capture editing phase, or both. Modifications may involve enhancement, replacement, or both.

[0337] However, in this example, if the post-capture editing phase has already begun, it is not yet complete. Instead, in this example, unmodified audio data 1305 is provided to cloud processing system 1325 for processing audio corresponding to one or more audio sources of unmodified audio data 1305. In some examples, at least some of the modified audio data 1310 may also be provided to cloud processing system 1325 for processing.

[0338] According to this example, the processing would involve applying one or more selected audio processing tools 1308 chosen by the toolset selection module 1302 from the editing toolbox 1304. The audio processing tools 1308 may be implemented, for example, via software stored on one or more non-transitory computer-readable storage media. The editing toolbox 1304 may include various types of audio processing tools 1308, such as one or more generation tools for a specific audio category, one or more signal generators (e.g., sound effects, also known as foley effects generators), one or more audio warping tools (e.g., tools for warping audio corresponding to human speech), one or more speech enhancement tools, one or more tools for automatically mixing audio corresponding to different audio sources in an audio scene, etc.

[0339] In some examples, the processing may involve what is referred to herein as a "second audio source separation process," which may be relatively more accurate than a previous audio source separation process applied, for example, during the capture phase. In some such examples, the unmodified audio data 1305 may include a copy of the original audio data acquired by the microphone system during the capture phase. In some examples, the individual audio sources output by the second audio source separation process may be further processed according to one or more selected audio processing tools 1308.

[0340] According to some examples, the toolset selection module 1302 can be controlled at least in part based on the input context metadata 1315. In some examples, the context metadata 1315 can be generated by the control system during the capture process. According to some such examples, the context metadata 1315 can be based on the existence of an estimate of human speech in audio captured by a microphone system, for example, during the capture phase. At least some of the context metadata 1315 can, for example, correspond to user input regarding one or more audio sources (e.g., obtained during the capture phase), such as the desire to improve the signal level of a speaker's audio, reduce street noise levels, enhance audio corresponding to one or more musical performers, etc. For example, if the user provides input during the capture phase indicating a desire to improve the signal level of a speaker's audio, the toolset selection module 1302 can be configured to automatically select one or more tools to enhance the speaker's speech to improve its intelligibility, etc., without requiring further input during the post-capture editing process. This aspect demonstrates improvements and technical advantages compared to previously deployed methods.

[0341] In some examples, the cloud processing system 1325 can operate automatically without human input. In some such examples, the cloud processing system 1325 can be configured to determine when audio processing is complete and cause the processing module 1312 to store the processed audio data 1320 as at least a portion of the processed and modified audio data 1325.

[0342] However, in some examples, processing module 1312 may perform at least some processing based on optional user input 1317. User input 1317 may control processing module 1312, toolset selection module 1302, or both. According to some such examples, after a user is satisfied with the audio processing provided by cloud processing system 1325, the user may provide user input 1317, thereby causing processed audio data 1320 to be stored as at least a portion of processed and modified audio data 1325.

[0343] Based on these examples, the cloud processing system 1325 is also configured to replicate unmodified audio data 1305 and context metadata 1315, and store copies of the unmodified audio data 1305 and context metadata 1315 as components of backward-compatible media assets 1330.

[0344] Figure 14This is a flowchart outlining various example methods 1400 according to some disclosed embodiments. Example method 1400 can be divided into blocks, such as blocks 1402, 1405, 1407, 1410, 1412, 1415, 1417, and 1420. Each block can be described as an operation, process, method, step, action, or function. As with other methods described herein, the blocks of method 1400 do not necessarily need to be performed in the indicated order. In some embodiments, one or more blocks of method 1400 can be performed simultaneously. Furthermore, some embodiments of method 1400 may include more or fewer blocks than those shown and / or described. The blocks of method 1400 can be generated by one or more devices (e.g., Figure 1A The method 1400 is performed on the device shown. Some blocks of method 1400 may be performed during the capture phase, and other blocks of method 1400 may be performed during the post-capture editing phase. Depending on the specific implementation, the post-capture editing phase may or may not be performed on the same device used for the capture phase.

[0345] Box 1402 relates to "receiving audio data from a microphone system by the control system of the device". The control system may be or may include Figure 1A CPU 141. The control system may include, for example, a general-purpose single-chip or multi-chip processor, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, one or more discrete hardware components, or combinations thereof. In some examples, the microphone system may be a microphone system of a mobile device that includes a control system. Processing may continue to box 1405.

[0346] Box 1405 relates to "receiving video data from a camera system by a control system". In some examples, the camera system may include one or more cameras of a mobile device that includes a control system. In other examples, a microphone system, a camera system, or both may reside in one or more devices other than the device that includes the control system. Processing may continue to box 1407.

[0347] Box 1407 relates to “identifying two or more audio sources in an audio scene by a control system, at least in part, based on audio data and video data.” Box 1407 may relate to an audio source separation process or be performed after such separation process. According to some examples, Box 1407 may relate to creating a list of audio sources in the current audio scene. In some examples, the audio source list may be as referenced herein. Figure 10A As described in this article, as referenced Figure 10BThe process may involve either or both described in box 1015. Some examples may involve the detection of one or more potential sound sources by a control system based at least in part on video data. Accordingly, in some examples, the list of audio sources may include one or more actual audio sources and one or more potential audio sources. In some examples, at least one potential sound source may not be indicated by audio data. Processing may continue to box 1410.

[0348] Box 1410 relates to “audio and video data received during the capture phase and stored by the control system”. Figure 13 The unmodified audio data 1305 is an example of stored audio data received during the capture phase. Processing can continue to box 1412.

[0349] Box 1412 relates to “a display controlled by a control system to display images corresponding to video data and to display a graphical user interface (GUI) overlaid on these images before and during the capture phase, wherein the GUI includes audio source images corresponding to each of two or more audio sources, and wherein the GUI includes one or more user input areas for receiving user input.” In some examples, box 1412 may relate to a display of control device 101 to present, as Figure 3A The GUI shown or Figures 4C to 4E One of the GUIs shown. For example, Figure 4D and Figure 4E A GUI is shown that provides information about four audio sources and includes user input areas, such as plus and minus icons for each of the four audio sources in audio source information areas 430a, 430b, 430c, and 430d. Accordingly, in some examples, one or more user input areas may include at least one user input area configured to receive user input about the selected level. Processing can continue to box 1415.

[0350] Box 1415 relates to "receiving user input by a control system via one or more user input areas, the user input corresponding to at least one of two or more audio sources." User input may include, for example, user touch. Figure 4D or Figure 4E The plus icon in the "speaker" audio source information area 430b, and the minus icon in the "cafeteria noise" audio source information area 430a, when the user touches them, can be processed up to box 1417.

[0351] Box 1417 relates to "the generation of revision metadata corresponding to user input by the control system." The revision metadata may include data related to the aforementioned user touch. Figure 4D or Figure 4EThe corresponding metadata includes the plus icon in the "speaker" audio source information area 430b and the minus icon in the "cafeteria noise" audio source information area 430a, which were touched by the user. In some examples, revised metadata can be stored as content referred to herein as "contextual metadata" (e.g., Figure 13 This is part of the context metadata (1315). In some examples, the control system can be configured to generate at least some context metadata corresponding to user actions other than GUI input, such as zooming in on a video object as an actual or potential audio source, centering a video frame on the video object as an actual or potential audio source, etc. According to some examples, the control system can be configured to generate at least some context metadata that may not directly correspond to user input, such as context metadata indicating the presence of one or more human speakers, context metadata indicating the presence of one or more performing musicians, etc. Processing can continue to box 1420.

[0352] Box 1420 relates to "the storage of revision metadata and audio data received at least during the capture phase by the control system." Some examples of method 1400 may involve storing other types of metadata created by the control system during the capture phase, such as other types of context metadata.

[0353] Some examples of method 1400 may involve modifying audio data corresponding to the revision metadata based on the revision metadata. As noted elsewhere in this document, in some examples, Figure 13 The context metadata 1315 may include revision metadata. In some examples, the cloud processing system 1325 may cause audio data corresponding to the revision metadata to be modified at least partially based on the context metadata 1315 including the revision metadata. In other examples, another device or system (such as a capture device) may cause audio data corresponding to the revision metadata to be modified at least partially based on the revision metadata. Accordingly, in some examples, causing the audio data to be modified may involve the control system modifying the audio data corresponding to the revision metadata. In other examples, causing the audio data received during the capture phase to be modified may involve the control system sending the revision metadata and the audio data received during the capture phase to one or more other devices. For example, causing the audio data received during the capture phase to be modified may involve the control system sending the revision metadata and the audio data received during the capture phase to one or more servers.

[0354] According to some examples, the audio data corresponding to the revised metadata may include unmodified audio data received during the capture phase. Figure 13 An example of unmodified audio data 1305 is provided.

[0355] In some examples, the audio data corresponding to the revised metadata may include modified audio data. Modified audio data may include, for example, enhanced audio data, replaced audio data, or both. Figure 13 The modified audio data 1310 provides an example.

[0356] According to some examples, modifying audio data received during the capture phase can involve applying an audio enhancement tool to the audio data corresponding to the revised metadata. In some such examples, the audio data corresponding to the revised metadata may include speech audio data corresponding to speech from at least one person, and the audio enhancement tool may be or may include a speech enhancement tool. According to some examples, the audio enhancement tool may be or may include a sound source separation tool.

[0357] Some examples of method 1400 may involve the control system receiving modified user input via one or more user input areas after the start of the capture phase. Some such examples may involve the control system causing audio data received during the capture phase to be modified according to the modified user input. In some instances, the audio data received during the capture phase can be modified during the capture phase itself. For example, the audio data received during the capture phase can be modified according to a microphone beamforming process performed during the capture phase. In some examples, the stored procedure of block 1410 or another aspect of method 1400 may involve storing the modified audio data that has been modified according to the user input.

[0358] According to some examples, the identification process of another aspect of block 1407 or method 1400 may involve the execution of a first sound source separation process by a control system. In some examples, causing the audio data to be modified may involve performing a second sound source separation process, or causing the execution of a second sound source separation process.

[0359] Example IVAS codec framework Figure 15 This is a block diagram of an example IVAS encoder / decoder (“codec”) framework 1500 for encoding and decoding Immersive Voice and Audio Services (IVAS) bitstreams according to one or more embodiments. IVAS is expected to support a range of audio service capabilities, including but not limited to mono-to-stereo upmixing and fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a variety of devices, endpoints, and network nodes, including but not limited to: mobile and smartphones, tablet computers, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater devices, and other suitable devices.

[0360] IVAS codec 1500 includes IVAS encoder 1501 and IVAS decoder 1504. IVAS encoder 1501 includes a receiver. N A spatial encoder 1502 for input spatial audio (e.g., FOA, HOA) of each channel. In some embodiments, the spatial encoder 1502 implements SPAR and DirAC to... N_dmx The spatial audio channels are analyzed / downmixed, as described in further detail below. The output of the spatial encoder 1502 includes the spatial metadata (MD) bitstream (BS) and the spatial downmixed data. N_dmx Each channel. Spatial MD is quantized and entropy-coded. In some implementations, quantization may include fine, medium, coarse, and very coarse quantization strategies, and entropy coding may include Huffman or arithmetic coding. In some implementations, the framework may allow no more than three quantization levels in a given operating mode; however, as the bit rate decreases, in some such implementations, these three levels generally become increasingly coarse to meet bit rate requirements. The core audio encoder 1503 (e.g., based on a mono Enhanced Voice Services (EVS) coding unit) will spatially undermix the... N_dmx Each channel ( N_ dmx = 1-16 channels) are encoded into an audio bitstream, which is combined with the spatial MD bitstream to form an IVAS encoded bitstream transmitted to the IVAS decoder 1504. As described below, given the bit rate constraint of low bit rate scene-based audio (SBA), in some implementations, the number of channels will be limited to a single channel.

[0361] The IVAS decoder 1504 includes a core audio decoder 1505 (e.g., an EVS decoder) that decodes the audio bitstream extracted from the IVAS bitstream to recover the original audio data. N_dmx Each audio channel. The spatial decoder / renderer 1506 (e.g., SPAR / DirAC) decodes the spatial MD bitstream extracted from the IVAS bitstream to recover the spatial MD, and uses the spatial MD and spatial upmixing to synthesize / render the output audio channels for playback on various audio systems with different speaker configurations and capabilities.

[0362] The various features and aspects will be understood from the following enumerated example embodiments (EEE): EEE1A. A method comprising: receiving audio data from a microphone system by a control system of a device; receiving video data from a camera system by the control system; creating an audio source list by the control system based at least in part on the audio data, the video data, or both; selecting one or more subsets of selected audio sources from the audio source list by the control system; estimating at least one audio characteristic of at least the one or more selected audio sources by the control system based on the audio data; storing the audio data and video data received during a capture phase by the control system; and controlling a display of the device by the control system to display an image corresponding to the video data, and displaying a graphical user interface (GUI) overlaid on the image before and during the capture phase, wherein the GUI includes an audio source image corresponding to the at least one audio characteristic of the subset of the one or more selected audio sources.

[0363] EEE2A. The method of claim EEE1A, wherein the GUI includes one or more user input areas configured to receive user input.

[0364] EEE3A. The method of claim EEE1A or claim EEE2A further includes the control system classifying the audio sources in the audio source list into two or more audio source categories, wherein the GUI includes a user input area portion corresponding to at least one of the two or more audio source categories.

[0365] EEE4A. The method of claim EEE3A, wherein one of the audio source categories is a foreground category corresponding to the one or more selected audio sources.

[0366] EEE5A. The method of claim EEE3A or claim EEE4A, wherein one of the audio source categories is a background category corresponding to one or more audio sources in the audio source list that are not in a subset of the one or more selected audio sources.

[0367] EEE6A. The method of any one of claims EEE1A to EEE5A, wherein the selection includes estimating which audio sources in the list of audio sources are the most important sound sources, and wherein a subset of the one or more selected audio sources includes the audio sources estimated to be the most important sound sources.

[0368] EEE7A. The method of any one of claims EEE1A to EEE6A, wherein the selection includes estimating which audio sources in the list of audio sources correspond to a speaker, and wherein a subset of the one or more selected audio sources includes audio sources estimated to be speakers.

[0369] EEE8A. The method of any one of claims EEE1A to EEE7A, further comprising the first sound source separation process performed by the control system.

[0370] EEE9A. The method of claim EEE8A further includes updating the audio scene state by the control system at least in part based on the first sound source separation process, and causing the GUI to be updated according to the updated audio scene state by the control system.

[0371] EEE10A. The method of claim EEE8A or claim EEE9A, wherein the creation is at least partially based on the first sound source separation process.

[0372] EEE11A. The method of any one of claims EEE8A to EEE10A, further comprising performing post-capture audio processing on the audio data received during the capture phase, wherein the post-capture audio processing includes a second source separation process that is more complex than the first source separation process.

[0373] EEE12A. The method of any one of claims EEE1A to EEE11A, further comprising detecting one or more potential sound sources by the control system based at least in part on the video data.

[0374] EEE13A. The method of claim EEE12A, wherein at least one of the one or more potential sound sources is not indicated by the audio data.

[0375] EEE14A. The method of claim EEE12A or claim EEE13A, wherein the audio source list includes the one or more potential sound sources.

[0376] EEE15A. The method of any one of claims EEE1A to EEE14A, further comprising detecting one or more candidate sound sources for enhanced audio capture by the control system, wherein the enhanced audio capture comprises replacing the candidate sound sources with external audio or synthesized audio.

[0377] EEE16A. The method of claim EEE15A, wherein the GUI includes at least one user input area configured to receive user selection of a selected potential sound source or a selected candidate sound source.

[0378] EEE17A. The method of claim EEE16A, wherein the GUI includes at least one user input area configured to receive user selection of enhanced audio capture, wherein the enhanced audio capture includes at least one of external audio or synthesized audio from a selected potential sound source or a selected candidate sound source.

[0379] EEE18A. The method of claim EEE17A, wherein the GUI includes at least one user input area configured to receive user selection of a ratio between enhanced audio capture and real-world audio capture.

[0380] EEE19A. The method of any one of claims EEE1A to EEE18A, further comprising the control system causing the display to show an audio source label in the GUI.

[0381] EEE20A. The method of claim EEE19A, wherein at least one of the audio source tags corresponds to an audio source identified by the control system based on the audio data, the video data, or both.

[0382] EEE21A. The method of any one of claims EEE1A to EEE20A, wherein the audio data is received from the microphone system of the device, and the video data is received from the camera system of the device.

[0383] EEE22A. The method of any one of claims EEE1A to EEE21A, further comprising updating the estimate of the current audio scene by the control system, and causing the GUI to be updated based on the updated estimate of the current audio scene by the control system.

[0384] EEE23A. The method of claim EEE22A, wherein updating the estimate of the current audio scene includes the control system implementing an audio classifier and a video classifier.

[0385] EEE24A. The method of claim EEE22A, wherein updating the estimate of the current audio scene includes the control system implementing an audiovisual classifier.

[0386] EEE25A. The method of any one of claims EEE22A to EEE24A, wherein the updated estimation of the current audio scene includes updated level estimates of one or more audio sources.

[0387] EEE26A. One or more non-transient media storing instructions for controlling one or more devices to perform a method comprising: receiving audio data from a microphone system by a control system of the devices; receiving video data from a camera system by the control system; creating an audio source list by the control system based at least in part on the audio data, the video data, or both; selecting one or more subsets of selected audio sources from the audio source list by the control system; estimating at least one audio characteristic of at least the one or more selected audio sources by the control system based on the audio data; storing the audio data and video data received during a capture phase by the control system; and controlling a display of the devices by the control system to display an image corresponding to the video data, and displaying a graphical user interface (GUI) overlaid on the image before and during the capture phase, wherein the GUI includes an audio source image corresponding to the at least one audio characteristic of the subset of the one or more selected audio sources.

[0388] EEE27A. One or more non-transient media as claimed in claim EEE26A, wherein the GUI includes one or more user input areas configured to receive user input.

[0389] EEE28A. One or more non-transient media as claimed in claim EEE26A or claim EEE27A, further comprising classifying the audio sources in the audio source list into two or more audio source categories by the control system, wherein the GUI includes a user input area portion corresponding to at least one of the two or more audio source categories.

[0390] EEE29A. One or more non-transient media as claimed in EEE28A, wherein one of the audio source categories is a foreground category corresponding to the one or more selected audio sources, and one of the audio source categories is a background category corresponding to one or more audio sources in the audio source list that are not in a subset of the one or more selected audio sources.

[0391] EEE30A. An apparatus comprising: an interface system; a memory system; a display system including at least one display; and a control system configured to: receive audio data from a microphone system via the interface system; receive video data from a camera system via the interface system; create an audio source list based at least in part on the audio data, the video data, or both; select one or more subsets of selected audio sources from the audio source list; estimate at least one audio characteristic of at least the one or more selected audio sources based on the audio data; store the audio data and video data received during a capture phase in the memory system; and control the display of the display system by the control system to display an image corresponding to the video data, and display a graphical user interface (GUI) overlaid on the image before and during the capture phase, wherein the GUI includes an audio source image corresponding to the at least one audio characteristic of the subset of the one or more selected audio sources.

[0392] EEE31A. The apparatus of claim EEE30A, wherein the GUI includes one or more user input areas configured to receive user input.

[0393] EEE32A. The apparatus of claim EEE30A or claim EEE31A further includes classifying the audio sources in the audio source list into two or more audio source categories by the control system, wherein the GUI includes a user input area portion corresponding to at least one of the two or more audio source categories.

[0394] EEE33A. The apparatus of claim EEE32A, wherein one of the audio source categories is a foreground category corresponding to the one or more selected audio sources, and one of the audio source categories is a background category corresponding to one or more audio sources in the audio source list that are not in a subset of the one or more selected audio sources.

[0395] EEE34A. The apparatus of any one of claims EEE30A to EEE33A, wherein the apparatus comprises the microphone system and the camera system.

[0396] EEE1B. A method comprising: receiving audio data from a microphone system by a control system of a device; receiving video data from a camera system by the control system; creating an audio source list by the control system based at least in part on the audio data, the video data, or both; selecting one or more selected audio sources from the audio source list by the control system, wherein the one or more selected audio sources are selected for possible enhancement or replacement; storing the audio data and video data received during a capture phase by the control system; and controlling a display of the device by the control system to display an image corresponding to the video data and display a graphical user interface (GUI) overlaid on the image, wherein the GUI indicates the one or more selected audio sources.

[0397] EEE2B. The method of claim EEE1B, wherein the GUI includes one or more displayed user input areas for receiving user input.

[0398] EEE3B. The method of claim EEE2B, wherein the GUI includes user prompts associated with enhancements or replacements of audio data corresponding to the one or more selected audio sources.

[0399] EEE4B. The method of claim EEE2B or claim EEE3B, wherein the selection is based at least in part on previously received user input.

[0400] EEE5B. The method of any one of claims EEE1B to EEE4B, wherein at least a first selected audio source among the one or more selected audio sources is selected based at least in part on the video data.

[0401] EEE6B. The method of claim EEE5B, wherein the audio data corresponding to the first selected audio source is below a threshold level.

[0402] EEE7B. The method of any one of claims EEE4B to EEE6B, wherein the GUI includes user prompts associated with enhancement or replacement of audio data corresponding to the first selected audio source.

[0403] EEE8B. The method of claim EEE7B, wherein the GUI includes user prompts associated with enhancing the audio data corresponding to the first selected audio source, and wherein the enhancement involves a microphone beamforming process for enhancing the audio data corresponding to the first selected audio source.

[0404] EEE9B. The method of claim EEE7B or claim EEE8B, wherein the GUI includes a user prompt associated with replacing audio data corresponding to the first selected audio source, and wherein the replacement involves replacing the audio data corresponding to the first selected audio source with synthesized audio data or with external audio data.

[0405] EEE10B. The method of any one of claims EEE7B to EEE9B, further comprising: receiving user input by the control system via the GUI, wherein the received user input indicates enhancement or replacement of audio data corresponding to the first selected audio source; and enhancing or replacing the audio data corresponding to the first selected audio source by the control system to generate enhanced audio data or replaced audio data.

[0406] EEE11B. The method of claim EEE10B further includes: tagging the enhanced audio data or the replacement audio data by the control system; and storing the tag together with the enhanced audio data or the replacement audio data.

[0407] EEE12B. The method of claim EEE11B, wherein the tag includes audio metadata.

[0408] EEE13B. The method of any one of claims EEE10B to EEE12B, further comprising the control system causing the GUI to indicate that the audio data corresponding to the first selected audio source is enhanced audio data or replacement audio data.

[0409] EEE14B. The method of claim EEE13B further includes the control system causing the GUI to indicate one or more audio sources corresponding to unmodified audio data.

[0410] EEE15B. The method of any one of claims EEE10B to EEE14B, further comprising storing, by the control system, unmodified audio data corresponding to at least the first selected audio source.

[0411] EEE16B. The method of claim EEE15B, further comprising: the control system causing the GUI to indicate an audio source having corresponding enhanced audio data or replacement audio data; the control system causing the GUI to include one or more user input areas for receiving user input for modifying the enhanced audio data or replacement audio data based on the unmodified audio data.

[0412] EEE17B. The method of claim EEE16B, wherein the modification includes interpolating between the enhanced audio data or replacement audio data and the unmodified audio data.

[0413] EEE18B. The method of claim EEE16B, wherein the modification includes replacing the enhanced audio data with the unmodified audio data or replacing the audio data.

[0414] EEE19B. The method of any one of claims EEE1B to EEE18B, wherein the selection includes estimating which audio sources in the list of audio sources correspond to a speaker, and wherein the one or more selected audio sources do not include audio sources estimated to be speakers.

[0415] EEE20B. The method of any one of claims EEE1B to EEE19B, further comprising detecting one or more potential sound sources by the control system at least in part based on the video data, wherein at least one of the one or more potential sound sources is not indicated by the audio data, and wherein the list of audio sources includes the one or more potential sound sources.

[0416] EEE21B. The method of any one of claims EEE1B to EEE20B, further comprising the control system causing the display to show an audio source label in the GUI.

[0417] EEE22B. The method of claim EEE21B, wherein at least one of the audio source tags corresponds to an audio source or potential audio source identified by the control system based on the audio data, the video data, or both.

[0418] EEE23B. The method of any one of claims EEE1B to EEE22B, wherein the audio data is received from the microphone system of the device, and the video data is received from the camera system of the device.

[0419] EEE24B. The method of any one of claims EEE1B to EEE23B, wherein the GUI is displayed before the capture phase, during the capture phase, or in both cases.

[0420] EEE25B. The method of any one of claims EEE1B to EEE24B, wherein the GUI is displayed after the capture phase and during the post-capture review process.

[0421] EEE26B. One or more non-transient media storing instructions for controlling one or more devices to perform a method comprising: receiving audio data from a microphone system by a control system of the devices; receiving video data from a camera system by the control system; creating an audio source list by the control system based at least in part on the audio data, the video data, or both; selecting one or more selected audio sources from the audio source list by the control system, wherein the one or more selected audio sources are selected for possible enhancement or replacement; storing the audio data and video data received during a capture phase by the control system; and controlling a display of the devices by the control system to display an image corresponding to the video data and display a graphical user interface (GUI) overlaid on the image, wherein the GUI indicates the one or more selected audio sources.

[0422] EEE27B. One or more non-transient media as claimed in claim EEE26B, wherein the GUI includes one or more displayed user input areas for receiving user input.

[0423] EEE28B. One or more non-transient media as claimed in claim EEE27B, wherein the GUI includes user prompts associated with enhancement or replacement of audio data corresponding to the one or more selected audio sources.

[0424] EEE29B. One or more non-transient media as claimed in claim EEE27B or claim EEE28B, wherein the selection is based at least in part on previously received user input.

[0425] EEE30B. One or more non-transient media as claimed in any one of EEE26B to EEE28B, wherein at least a first selected audio source among the one or more selected audio sources is selected based at least in part on the video data.

[0426] EEE31B. An apparatus comprising: an interface system; a display system including one or more displays; a memory system; and a control system configured to: receive audio data from a microphone system via the interface system; receive video data from a camera system via the interface system; create an audio source list based at least in part on the audio data, the video data, or both; select one or more selected audio sources from the audio source list, wherein the one or more selected audio sources are selected for possible enhancement or replacement; store the audio data and video data received during a capture phase in the memory system; and control the displays of the display system to display an image corresponding to the video data and to display a graphical user interface (GUI) overlaid on the image, wherein the GUI indicates the one or more selected audio sources.

[0427] EEE32B. The apparatus of claim EEE31B, wherein the GUI includes one or more displayed user input areas for receiving user input.

[0428] EEE33B. The apparatus of claim EEE32B, wherein the GUI includes user prompts associated with enhancement or replacement of audio data corresponding to the one or more selected audio sources.

[0429] EEE34B. The apparatus of claim EEE32B or claim EEE33B, wherein the selection is based at least in part on previously received user input.

[0430] EEE35B. The apparatus of any one of claims EEE31B to EEE34B, wherein at least a first selected audio source among the one or more selected audio sources is selected based at least in part on the video data.

[0431] EEE1C. A method comprising: receiving audio data from a microphone system by a control system of a device; receiving video data from a camera system by the control system; creating an audio source list by the control system based at least in part on the audio data, the video data, or both; selecting at least a first selected audio source from the audio source list by the control system, the first selected audio source being selected for enhancement or replacement; enhancing or replacing audio data corresponding to the first selected audio source by the control system to produce first modified audio data, the first modified audio data including at least one of first enhanced audio data or first replaced audio data; and storing the first modified audio data by the control system. The control system stores audio and video data received during the capture phase, the audio data including first unmodified audio data corresponding to at least the first selected audio source; the control system controls the device's display to present an image corresponding to the video data and displays a post-capture graphical user interface (GUI) overlaid on the image, wherein the post-capture GUI at least indicates the first selected audio source and one or more user input areas to receive user input; and during the post-capture phase review process, the first modified audio data is edited based on the user input received by the post-capture GUI to include at least a portion of the first unmodified audio data.

[0432] EEE2C. The method of claim EEE1C, wherein the editing includes interpolating between the first modified audio data and the first unmodified audio data.

[0433] EEE3C. The method of claim EEE1C or claim EEE2C, wherein the post-capture GUI includes at least one user input area configured to receive user selection of a ratio between the first modified audio data and the first unmodified audio data.

[0434] EEE4C. The method of claim EEE3C, wherein the at least one user input area includes a slider.

[0435] EEE5C. The method of claim EEE4C, wherein the slider is configured to allow a user to select a ratio from 0% to 100%.

[0436] EEE6C. The method of any one of claims EEE1C to EEE5C, wherein the selection is based at least in part on user input.

[0437] EEE7C. The method of any one of claims EEE1C to EEE6C, wherein the selection is based at least in part on the video data.

[0438] EEE8C. The method of claim EEE7C, wherein the audio data corresponding to the first selected audio source is below a threshold level.

[0439] EEE9C. The method of any one of claims EEE1C to EEE8C, wherein the control further includes adjusting the display to present an audio data modification GUI before or during the capture phase, the audio data modification GUI including user prompts associated with enhancement or replacement of audio data corresponding to one or more selected audio sources in the audio source list.

[0440] EEE10C. The method of claim EEE9C, wherein the audio data modification GUI includes a user prompt associated with enhancing the audio data corresponding to a selected audio source, and wherein the enhancement is associated with a microphone beamforming process for enhancing the audio data corresponding to the selected audio source.

[0441] EEE11C. The method of claim EEE9C or claim EEE10C, wherein the audio data modification GUI includes a user prompt associated with replacing audio data corresponding to a selected audio source, and wherein the replacement is associated with replacing the audio data corresponding to the selected audio source with synthesized audio data or with external audio data.

[0442] EEE12C. The method of any one of EEE9C to EEE11C, further comprising receiving audio data modification user input by the control system via the audio data modification GUI, the audio data modification user input indicating enhancement or replacement of audio data corresponding to a first selected audio source; and wherein the editing effectively responds to the audio data modification user input by providing the first modified audio data.

[0443] EEE13C. The method of claim EEE12C further includes: tagging the modified audio data by the control system; and storing the tag together with the modified audio data.

[0444] EEE14C. The method of claim EEE13C, wherein the tag includes audio metadata.

[0445] EEE15C. The method of any one of claims EEE1C to EEE14C, further comprising the control system causing the post-capture GUI to indicate that the audio data corresponding to the first selected audio source is modified audio data.

[0446] EEE16C. The method of claim EEE15C further includes the control system causing the post-capture GUI to indicate one or more audio sources corresponding to the unmodified audio data.

[0447] EEE17C. The method of any one of claims EEE1C to EEE16C, wherein the selection involves estimating which audio sources in the list of audio sources correspond to a speaker, and wherein the one or more selected audio sources do not include audio sources estimated to be speakers.

[0448] EEE18C. The method of any one of claims EEE1C to EEE17C, further comprising detecting one or more potential sound sources by the control system at least in part based on the video data, wherein at least one of the one or more potential sound sources is not indicated by the audio data, and wherein the list of audio sources includes the one or more potential sound sources.

[0449] EEE19C. The method of any one of claims EEE1C to EEE18C, further comprising the control system causing the display to show an audio source label in the audio data modification GUI or the post-capture GUI.

[0450] EEE20C. The method of claim EEE19C, wherein at least one of the audio source tags corresponds to an audio source or potential audio source identified by the control system based on the audio data, the video data, or both.

[0451] EEE21C. The method of any one of claims EEE1C to EEE20C, wherein the audio data is received from the microphone system of the device, and the video data is received from the camera system of the device.

[0452] EEE22C. One or more non-transient media storing instructions for controlling one or more devices to perform a method, the method comprising: receiving audio data from a microphone system by a control system of the devices; receiving video data from a camera system by the control system; creating an audio source list by the control system based at least in part on the audio data, the video data, or both; selecting at least a first selected audio source from the audio source list by the control system, the first selected audio source being selected for enhancement or replacement; enhancing or replacing audio data corresponding to the first selected audio source by the control system to produce first modified audio data, the first modified audio data including at least one of first enhanced audio data or first replaced audio data; and ...; and enhancing or replacing audio data based on the first selected audio source to produce first modified audio data, the first modified audio data including at least one of first enhanced audio data or first replaced audio data; and enhancing or replacing audio data based on the first selected audio source; and enhancing or replacing audio data based on the first selected audio source to produce first modified audio data, the first modified audio data including at least one of first enhanced audio data or first replaced audio data; and enhancing or replacing audio data based on the first selected audio source; and enhancing or replacing audio data based on the first selected audio source to produce first modified audio data, the first modified audio data including at least one of first enhanced audio data or first replaced audio data; and enhancing or replacing audio data based on the first selected audio The system stores the first modified audio data; the control system stores audio and video data received during the capture phase, the audio data including first unmodified audio data corresponding to at least the first selected audio source; the control system controls the device's display to present an image corresponding to the video data and display a post-capture graphical user interface (GUI) overlaid on the image, wherein the post-capture GUI at least indicates the first selected audio source and one or more user input areas to receive user input; and during the post-capture phase review process, the first modified audio data is edited based on the user input received by the post-capture GUI to include at least a portion of the first unmodified audio data.

[0453] EEE23C. One or more non-transient media as claimed in claim EEE22C, wherein the editing includes interpolating between the first modified audio data and the first unmodified audio data.

[0454] EEE24C. One or more non-transient media as claimed in claim EEE22C or claim EEE23C, wherein the post-capture GUI includes at least one user input area configured to receive a user selection of a ratio between the first modified audio data and the first unmodified audio data.

[0455] EEE25C. One or more non-transient media as claimed in claim EEE24C, wherein the at least one user input area includes a slider.

[0456] EEE26C. An apparatus comprising: an interface system; a memory system; a display system including at least one display; and a control system configured to: receive audio data from a microphone system; receive video data from a camera system; create an audio source list based at least in part on the audio data, the video data, or both; select at least a first selected audio source from the audio source list, the first selected audio source being selected for enhancement or replacement; enhance or replace audio data corresponding to the first selected audio source to produce first modified audio data, the first modified audio data including at least one of first enhanced audio data or first replaced audio data; and store the first modified audio data. Modifying audio data; storing audio and video data received during the capture phase, the audio data including first unmodified audio data corresponding to at least the first selected audio source; controlling the display of the device by the control system to present an image corresponding to the video data and displaying a post-capture graphical user interface (GUI) overlaid on the image, wherein the post-capture GUI at least indicates the first selected audio source and one or more user input areas to receive user input; and during the post-capture phase review process, editing the first modified audio data based on the user input received by the post-capture GUI to include at least a portion of the first unmodified audio data.

[0457] EEE27C. The apparatus of claim EEE26C, wherein the editing includes interpolating between the first modified audio data and the first unmodified audio data.

[0458] EEE28C. The apparatus of claim EEE26C or claim EEE27C, wherein the post-capture GUI includes at least one user input area configured to receive a user selection of a ratio between the first modified audio data and the first unmodified audio data.

[0459] EEE29C. The apparatus of claim EEE28C, wherein the at least one user input area includes a slider.

[0460] EEE1D. A method comprising: receiving audio data from a microphone system by a control system of a device; receiving video data from a camera system by the control system; identifying two or more audio sources in an audio scene by the control system at least in part based on the audio data and the video data; storing the audio data and video data received during a capture phase by the control system; controlling a display of the device by the control system to display an image corresponding to the video data, and displaying a graphical user interface (GUI) overlaid on the image before and during the capture phase, wherein the GUI includes an audio source image corresponding to each of the two or more audio sources, and wherein the GUI includes one or more user input areas for receiving user input; receiving user input via the one or more user input areas by the control system, the user input corresponding to at least one of the two or more audio sources; generating revision metadata corresponding to the user input by the control system; and storing the revision metadata and the audio data received at least during the capture phase by the control system.

[0461] EEE2D. The method of claim EEE1D further includes modifying the audio data corresponding to the revised metadata according to the revised metadata.

[0462] EEE3D. The method of claim EEE2D, wherein the audio data corresponding to the revised metadata includes unmodified audio data received during the capture phase.

[0463] EEE4D. The method of claim EEE2D or claim EEE3D, wherein the audio data corresponding to the revised metadata includes modified audio data, and wherein the modified audio data includes enhanced audio data or replacement audio data.

[0464] EEE5D. The method of any one of claims EEE2D to EEE4D, wherein modifying the audio data includes modifying the audio data corresponding to the revised metadata by the control system.

[0465] EEE6D. The method of any one of claims EEE2D to EEE4D, wherein modifying the audio data received during the capture phase includes the control system sending the revised metadata and the audio data received during the capture phase to one or more other devices.

[0466] EEE7D. The method of any one of EEE2D, EEE3D, EEE4D or EEE6D, wherein modifying the audio data received during the capture phase includes the control system sending the revised metadata and the audio data received during the capture phase to one or more servers.

[0467] EEE8D. The method of any one of claims EEE2D to EEE7D, wherein modifying the audio data received during the capture phase includes applying an audio enhancement tool to the audio data corresponding to the revised metadata.

[0468] EEE9D. The method of claim EEE8D, wherein the audio data corresponding to the revised metadata includes speech audio data corresponding to speech from at least one person, and wherein the audio enhancement tool includes a speech enhancement tool.

[0469] EEE10D. The method of claim EEE8D, wherein the audio enhancement tool includes a sound source separation process.

[0470] EEE11D. The method of any one of claims EEE1D to EEE10D, further comprising: receiving modified user input via the one or more user input areas after the start of the capture phase by the control system; and causing the control system to modify the audio data received during the capture phase according to the modified user input.

[0471] EEE12D. The method of any one of claims EEE1D to EEE11D, wherein the one or more user input areas include at least one user input area configured to receive user input regarding a selected level.

[0472] EEE13D. The method of any one of claims EEE1D to EEE12D, wherein the identification includes creating a sound source list by the control system.

[0473] EEE14D. The method of claim EEE13D, wherein the list of sound sources includes actual sound sources and potential sound sources.

[0474] EEE15D. The method of any one of claims EEE1D to EEE14D, wherein the storage includes storing modified audio data that has been modified according to the user input.

[0475] EEE16D. The method of any one of claims EEE1D to EEE15D, wherein the identification includes performing a first sound source separation process by the control system, and wherein modifying the audio data includes performing a second sound source separation process.

[0476] EEE17D. One or more non-transient media storing instructions for controlling one or more devices to perform a method comprising: receiving audio data from a microphone system by a control system of the devices; receiving video data from a camera system by the control system; identifying two or more audio sources in an audio scene by the control system at least in part based on the audio data and the video data; storing the audio data and video data received during a capture phase by the control system; controlling a display of the devices by the control system to display an image corresponding to the video data, and displaying a graphical user interface (GUI) overlaid on the image before and during the capture phase, wherein the GUI includes an audio source image corresponding to each of the two or more audio sources, and wherein the GUI includes one or more user input areas for receiving user input; receiving user input via the one or more user input areas by the control system, the user input corresponding to at least one of the two or more audio sources; generating revision metadata corresponding to the user input by the control system; and storing the revision metadata and the audio data received at least during the capture phase by the control system.

[0477] EEE18D. One or more non-transient media as claimed in EEE17D, further comprising modifying audio data corresponding to the revised metadata according to the revised metadata.

[0478] EEE19D. One or more non-transient media as claimed in EEE18D, wherein the audio data corresponding to the revised metadata includes unmodified audio data received during the capture phase.

[0479] EEE20D. One or more non-transient media as claimed in claim EEE18D or claim EEE19D, wherein the audio data corresponding to the revised metadata includes modified audio data, and wherein the modified audio data includes enhanced audio data or replacement audio data.

[0480] EEE21D. One or more non-transient media as claimed in any one of EEE18D to EEE20D, wherein modifying the audio data includes modifying the audio data corresponding to the revised metadata by the control system.

[0481] EEE22D. An apparatus comprising: an interface system; a display system including one or more displays; a memory system; and a control system configured to: receive audio data from a microphone system via the interface system; receive video data from a camera system via the interface system; identify two or more audio sources in an audio scene based at least in part on the audio data and the video data; store the audio data and video data received during a capture phase in the memory system; control the displays of the display system to display an image corresponding to the video data, and display a graphical user interface (GUI) overlaid on the image before and during the capture phase, wherein the GUI includes an audio source image corresponding to each of the two or more audio sources, and wherein the GUI includes one or more user input areas for receiving user input; receive user input via the one or more user input areas, the user input corresponding to at least one of the two or more audio sources; generate revision metadata corresponding to the user input; and store the revision metadata and the audio data received at least during the capture phase in the memory system.

[0482] EEE23D. The apparatus of claim EEE22D further includes causing audio data corresponding to the revised metadata to be modified according to the revised metadata.

[0483] EEE24D. The apparatus of claim EEE23D, wherein the audio data corresponding to the revised metadata includes unmodified audio data received during the capture phase.

[0484] EEE25D. The apparatus of claim EEE23D or claim EEE24D, wherein the audio data corresponding to the revised metadata includes modified audio data, and wherein the modified audio data includes enhanced audio data or replacement audio data.

[0485] EEE26D. The apparatus of any one of claims EEE23D to EEE25D, wherein modifying the audio data includes modifying the audio data corresponding to the revised metadata by the control system.

[0486] According to exemplary embodiments of this disclosure, the processes disclosed above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 709, and / or from a removable medium (e.g., Figure 1A The removable media 151 shown is installed.

[0487] Generally, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be executed by a control circuitry system, which can therefore perform or be configured to perform the actions described in this disclosure. Some aspects can be implemented in hardware, while others can be implemented in firmware or software (e.g., control circuitry) that can be executed by a controller, microprocessor, or other computing device. Although various aspects of the exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, other computing devices, or some combination thereof, as non-limiting examples.

[0488] Furthermore, the various blocks shown in the flowchart can be viewed as method steps, and / or operations resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform associated functions. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code configured to perform the methods described above.

[0489] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0490] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that when executed by the processor of the computer or the processor of the other programmable data processing apparatus, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0491] While this document contains numerous details of specific implementation, these details should not be construed as limiting the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Specific features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in this way, in some cases one or more features of a claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof. The logical flow depicted in the drawings does not require the specific order or ordered sequence shown to achieve the desired result. Additionally, other steps may be provided from the described flow, or steps may be removed, and other components ...

Claims

1. A method comprising: The device's control system receives audio data from the microphone system; The control system receives video data from the camera system; The control system identifies two or more audio sources in an audio scene, at least in part, based on the audio data and the video data. The control system estimates at least one audio characteristic of each of the two or more audio sources based on the audio data; The audio and video data received during the capture phase are stored by the control system. The control system controls the device's display to display an image corresponding to the video data, and displays a graphical user interface (GUI) overlaid on the image before and during the capture phase. The GUI includes an audio source image corresponding to at least one audio characteristic of each of the two or more audio sources, and includes one or more user input areas for receiving user input. The control system receives user input via one or more user input areas prior to the capture phase. as well as The control system causes the audio data received during the capture phase to be modified based on the user input.

2. The method of claim 1, further comprising: After the capture phase begins, the control system receives user input via the user input area. as well as The control system causes the audio data received during the duration of the capture phase to be modified according to the user input.

3. The method of claim 1 or claim 2, further comprising classifying the two or more audio sources into two or more audio source categories by the control system, wherein, The GUI includes a user input area portion corresponding to each of the two or more audio source categories.

4. The method of claim 3, wherein, The two or more audio source categories include a background category and a foreground category.

5. The method as claimed in claim 3 or claim 4, wherein, The one or more user input areas include at least one user input area configured to receive user input regarding a selected level or level ratio for each of the two or more audio source categories.

6. The method according to any one of claims 3 to 5, wherein, The storage includes unmodified audio data received during the capture phase.

7. The method according to any one of claims 3 to 6, wherein, The identification involves the control system creating a list of sound sources.

8. The method of claim 7, wherein, The list of sound sources includes both actual and potential sound sources.

9. The method of claim 7 or claim 8, wherein, Classifying the two or more audio sources into two or more audio source categories is based on the audio source list.

10. The method of any one of claims 7 to 9, further comprising determining one or more operable feedback types regarding the audio scene, wherein, The GUI is based in part on one or more of the operable feedback types.

11. The method according to any one of claims 1 to 10, wherein, The storage includes unmodified audio data received during the capture phase.

12. The method according to any one of claims 1 to 11, wherein, The storage includes storing modified audio data that has been modified based on the user input.

13. The method of any one of claims 1 to 12, further comprising creating and storing user input metadata corresponding to user input received via the user input area by the control system.

14. The method of claim 13, wherein, Modifying the audio data received during the capture phase based on the user input involves post-capture audio processing based at least in part on the user input metadata.

15. The method of claim 14, wherein, The control system is configured to perform at least a portion of the post-capture audio processing.

16. The method of claim 14 or claim 15, wherein, Another control system is configured to perform at least a portion of the captured audio processing.

17. The method according to any one of claims 14 to 16, wherein, The identification involves the control system performing a first sound source separation process, and the post-capture audio processing involves performing a second sound source separation process.

18. The method according to any one of claims 1 to 17, wherein, The identification involves the control system detecting one or more potential sound sources, at least in part, based on the video data.

19. The method of claim 18, wherein, At least one of the one or more potential sound sources is not indicated by the audio data.

20. The method of any one of claims 1 to 19, further comprising detecting one or more candidate sound sources for enhanced audio capture by the control system, the enhanced audio capture involving replacing the candidate sound sources with external audio or synthesized audio.

21. The method of claim 20, wherein, The GUI includes at least one user input area configured to receive user selections for a selected potential sound source or a selected candidate sound source.

22. The method of claim 21, wherein, The GUI includes at least one user input area configured to receive user selections for enhanced audio capture, which includes at least one of external or synthesized audio from a selected potential or candidate sound source.

23. The method of claim 22, wherein, The GUI includes at least one user input area configured to receive user selections regarding the ratio between enhanced audio capture and real-world audio capture.

24. The method according to any one of claims 1 to 23, wherein, The audio data is received from the device's microphone system, and the video data is received from the device's camera system.

25. The method according to any one of claims 1 to 24, wherein, Modifying the audio data received during the capture phase based on the user input involves modifying the audio data corresponding to the selected audio source or the selected audio source category.

26. The method of claim 25, wherein, This allows the modification of the audio data received during the capture phase based on the user input to involve a beamforming process corresponding to the selected audio source.

27. One or more non-transient media having instructions encoded thereon for controlling one or more devices to perform a method, the method comprising: The device's control system receives audio data from the microphone system; The control system receives video data from the camera system; The control system identifies two or more audio sources in an audio scene, at least in part, based on the audio data and the video data. The control system estimates at least one audio characteristic of each of the two or more audio sources based on the audio data; The audio and video data received during the capture phase are stored by the control system. The control system controls the device's display to display an image corresponding to the video data, and displays a graphical user interface (GUI) overlaid on the image before and during the capture phase. The GUI includes an audio source image corresponding to at least one audio characteristic of each of the two or more audio sources, and includes one or more user input areas for receiving user input. The control system receives user input via one or more user input areas prior to the capture phase. as well as The control system causes the audio data received during the capture phase to be modified based on the user input.

28. One or more non-transient media as described in claim 27, wherein, The method further includes: The control system receives user input via the user input area after the capture phase begins; and The control system causes the audio data received during the duration of the capture phase to be modified according to the user input.

29. One or more non-transient media as described in claim 27 or claim 28, wherein, The method further includes classifying the two or more audio sources into two or more audio source categories by the control system, wherein the GUI includes a user input area portion corresponding to each of the two or more audio source categories.

30. One or more non-transient media as described in claim 29, wherein, The two or more audio source categories include a background category and a foreground category.

31. One or more non-transient media as described in claim 29 or claim 30, wherein, The one or more user input areas include at least one user input area configured to receive user input regarding a selected level or level ratio for each of the two or more audio source categories.

32. An apparatus comprising: Interface system; A display system, the display system including one or more displays; A touch sensor system, the touch sensor system being located near at least one of the one or more displays; as well as The control system includes one or more processors and is configured to: The interface system receives audio data from the microphone system. The interface system receives video data from the camera system. Identify two or more audio sources in an audio scene, at least in part, based on the audio data and the video data; Estimate at least one audio characteristic of each of the two or more audio sources; The audio and video data received during the capture phase are stored. The display system is controlled to display an image corresponding to the video data, and a graphical user interface (GUI) overlaid on the image is displayed before and during the capture phase. The GUI includes an audio source image corresponding to at least one audio characteristic of each of the two or more audio sources, and includes one or more user input areas for receiving user input via the touch sensor system. The interface system receives user input via one or more user input areas prior to the capture phase. as well as This allows the audio data received during the capture phase to be modified based on the user input.

33. The apparatus of claim 32, wherein, The control system is further configured to: After the capture phase begins, user input is received via the user input area; and This allows the audio data received during the duration of the capture phase to be modified based on the user input.

34. The apparatus of claim 32 or claim 33, wherein, The control system is further configured to classify the two or more audio sources into two or more audio source categories, wherein the GUI includes a user input area portion corresponding to each of the two or more audio source categories.

35. The apparatus of claim 34, wherein, The two or more audio source categories include a background category and a foreground category.

36. The apparatus of claim 34 or claim 35, wherein, The one or more user input areas include at least one user input area configured to receive user input regarding a selected level or level ratio for each of the two or more audio source categories.

Citation Information

Patent Citations

  • Deep-learning based speech enhancement

    US20230368807A1