Depth-enhanced audio optimization system

WO2026183192A3PCT designated stage Publication Date: 2026-10-01DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016610
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2026-01-30
Filing Date
2026-02-25
Publication Date
2026-10-01

AI Technical Summary

Technical Problem

Existing audio capture technologies face challenges in handling environmental obstacles such as wind, noise, low light conditions, and bright light conditions, particularly in user-generated content like vlogs, and struggle with accurately positioning audio sources in a 3D video scene relative to visual elements.

Method used

Integration of real-time depth-enhanced audio optimization techniques using multi-microphone arrays and advanced signal processing, which involve capturing 3D depth data, adjusting audio source spatialization parameters based on target object positions, and maintaining synchronization between audio and video data.

Benefits of technology

Ensures accurate audio-visual synchronization and optimal sound quality, even in dynamic environments, by dynamically adjusting audio parameters and reducing noise, thus enhancing the immersive experience in vlogs and professional-generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026016610_01102026_PF_FP_ABST
    Figure US2026016610_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Methods for dynamic audio positioning of target objects corresponding to audio sources located in a three-dimensional (3D) space may involve obtaining audio data and video data via a capture device, identifying, based on the video data, one or more of the target objects, determining parameters of a bounding box around each of the one or more target objects, and estimating, based on the bounding box parameters, target object position data. Some methods involve adjusting audio source spatialization parameters of audio sources corresponding to each of the one or more target objects according to the target object position data and producing an output content stream that includes output video data and output audio data. The identifying, determining, estimating and adjusting may all occur during capture of the audio data and the video data.
Need to check novelty before this filing date? Find Prior Art

Description

D25022W001DEPTH-ENHANCED AUDIO OPTIMIZATION SYSTEM CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from International Patent Application No. PCT / CN2025 / 079958, filed on 18 February 2025, and United States Provisional Patent Application No. 63 / 971,867, filed on 30 January 2026, the contents of each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] This application relates generally to visual and multi-modal processing. In particular, the present disclosure is directed to systems, devices, methods, and / or computer program products configured for depth-enhanced audio optimization.BACKGROUND

[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.

[0004] Recent developments in personal mobile devices (e.g., smart phone, tablet device, wearable device) have resulted in a number of new formats and standards for communicating and processing captured audio. For example, voice and video encoder / decoder (“codec”) standards have been developed, including codecs for Immersive Voice and Audio Services (IVAS). The IVAS standard is expected to support a range of service capabilities with audio processing that may be in any variety of formats such as mono, stereo or fully spatial audio signals (e.g., stereo, multi-channel, Ambisonics, etc.). The enhanced audio processing that is supported by IVAS leads to improved listener enjoyment due to the high quality of audio with a fully immersive spatial sound experience.

[0005] User-generated content (UGC) is often created by consumers in the form of a mixedmode format including a variety of images, videos, text and audio, typically using a personal mobile device. A so called video log (vlog) is one form of UGC where a user creates a multimodal recording including both audio and video, to capture their personal thoughts and moments in their lives to share through some form of recorded media outlet (e.g., SMS, Facebook, YouTube, Instagram, Twitter, X, etc.)D25022W001

[0006] It is with respect to these and other considerations that the disclosure made herein is presented.BRIEF SUMMARY OF THE DISCLOSURE

[0007] The techniques described herein include visual and multi-modal processing. In particular, the present disclosure is directed to techniques for real-time depth-enhanced audio optimization. The present disclosure appreciates that capture of multi-modal UGC may experience any variety of environmental obstacles, including but not limited to wind, noise, low light conditions, bright light conditions, to name a few. Additional challenges discussed herein, which are peculiar to vlog audio / video capture, involve processing the spatial audio so as to position a narrator’s voice in a desired location of the audio scene relative to the video capture. The disclosed techniques are suitable for UGC, such as vlog recordings, as well as for professional-generated content (PGC).

[0008] Briefly stated, methods, systems, devices and computer-readable media for enhancing audio quality by integrating real-time depth information are disclosed. Example methods may utilize an electronic processor configured for: real-time capture of three-dimensional (3D) depth data for target objects or individuals located in a 3D video scene; real-time adjustment of spatial parameters of audio sources in the video scene based on the position of the audio sources relative to the target objects in the 3D video scene; and continual synchronization between captured audio and video.

[0009] Methods for dynamic audio positioning of target objects corresponding to audio sources located in a 3D space may involve obtaining audio data and video data via a capture device, identifying, based on the video data, one or more of the target objects, determining parameters of a bounding box around each of the one or more target objects, and estimating, based on the bounding box parameters, target object position data. Some methods involve adjusting audio source spatialization parameters of audio sources corresponding to each of the one or more target objects according to the target object position data and producing an output content stream that includes output video data and output audio data. The identifying, determining, estimating and adjusting may all occur during capture of the audio data and the video data.

[0010] Some methods may involve maintaining real-time synchronization between the audio data and the video data. The real-time synchronization may be performed during capture ofD25022W001the audio data and the video data and may be based, at least in part, on repeated estimation of the target object position data to produce updated target object position data. The real-time synchronization may be based, at least in part, on repeated adjustment of the audio source spatialization parameters according to the updated target object position data.

[0011] Some disclosed methods may involve creating audio source metadata corresponding to the audio source spatialization parameters and / or the target object position data and embedding the audio source metadata in the output content stream. Some methods may involve modifying an intensity of audio corresponding to one or more audio sources based, at least in part, on the target object position data. Some such methods may involve creating audio source metadata corresponding to this intensity modification of audio and embedding the audio source metadata in the output content stream.

[0012] Some methods may involve capturing, with a multi-microphone array of the capture device, audio data associated with an audio scene. The audio scene may include ambient sound signals and target sounds associated with audio data from the one or more target objects. Some such methods may involve applying an emphasis process to emphasize target sounds in the audio scene while reducing interference from other sounds in the audio scene. The emphasis process may include a beam forming process and / or an adaptive filtering process. Some such methods may involve separating and suppressing background noise from the audio scene by applying an Independent Component Analysis (ICA) process.

[0013] Various aspects of the present disclosure provide for processing of video and audio signals, and effect improvements in the technical fields of audio and video processing including at least spatial audio and video capture, spatial audio and video encoding, spatial audio and video decoding, and the like.

[0014] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.

[0015] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplifiedD25022W001form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims.DESCRIPTION OF DRAWINGS

[0016] Figure 1 shows an illustration of the vertical field of view of a camera.

[0017] Figure 2 shows an example of the geometrical relationship between the actual size of a target object and a bounding box in a captured image.

[0018] Figure 3 is a graph that shows the relationship between the bounding box size and the distance to a person.

[0019] Figure 4 shows examples of pixel distance corrections according to the abovedescribed example.

[0020] Figure 5 shows an example of multiple devices being used for audio and video capture.

[0021] Figure 6 is a top view that shows examples of two target objects and a capture device.

[0022] Figure 7 illustrates a block diagram of a system for content generation in accordance with aspects of the present disclosure.

[0023] Figure 8 is a detailed flow diagram that outlines various example methods 800 according to some disclosed implementations.

[0024] Figure 9 illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.

[0025] Figure 10 is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure.

[0026] In the drawings, specific arrangements or orderings of schematic elements, such as those representing devices, units, instruction blocks and data elements, are shown for ease of description. However, it should be understood by those skilled in the art that the specific ordering or arrangement of the schematic elements in the drawings is not meant to imply that a particular order or sequence of processing, or separation of processes, is required. Further, the inclusion of a schematic element in a drawing is not meant to imply that such element is required in all embodiments or that the features represented by such element may not be included in or combined with other elements in some implementations.

[0027] Further, in the drawings, where connecting elements, such as solid or dashed lines or arrows, are used to illustrate a connection, relationship, or association between or among two or more other schematic elements, the absence of any such connecting elements is not meant to imply that no connection, relationship, or association can exist. In other words,D25022W001some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the disclosure. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships or associations between elements. For example, where a connecting element represents a communication of signals, data, or instructions, it should be understood by those skilled in the art that such element represents one or multiple signal paths, as may be needed, to effect the communication.

[0028] The same reference symbol used in various drawings indicates like elements.D25022W001DETAILED DESCRIPTION

[0029] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of various described embodiments with reference to the accompanying drawings. The illustrative embodiments in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes made, without departing from the spirit or scope of the present disclosure. In light of the present disclosure, it will be apparent to one of ordinary skill in the art that the various described features and implementations may be practiced without many of these specific details. In some instances, well-known methods, procedures, components, and circuits, have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described hereafter that can each be used independently of one another or with any combination of other features. Thus, the features may be arranged, substituted, combined, separated, or designed into other configurations, which is contemplated in light of the present disclosure.

[0030] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0031] AcronymsAI - Artificial IntelligenceD25022W001BRIR - Binaural Room Impulse ResponseCNN - Convolutional Neural NetworkDSP - Digital Signal ProcessorDRC - Dynamic Range CompressionDOA - Direction of ArrivalEPROM - Erasable Programmable Read-Only MemoryFPGA - Field-Programmable Gate ArrayGUI - Graphical User InterfaceICA - Independent Component AnalysisI / O - Input / OutputIVAS - Immersive Voice and Audio ServicesPGC - Professional-Generated ContentRAM - Random Access MemoryR-CNN - Region-Based Convolutional Neural NetworkROM - Read-Only MemorySNR - Signal-to-Noise RatioToF – Time-of-FlightUGC - User-Generated Contentvlog - Video LogYOLO - You Only Look Once

[0032] The present disclosure generally describes a comprehensive system for enhancing audio quality in vlog recordings and other audiovisual recordings by integrating real-time depth information, multi-microphone arrays, and advanced signal processing techniques. In some examples, the system dynamically adjusts audio parameters such as spatial position, gain control, and noise reduction based on the target object position data that indicates the distance between the capture device (e.g., a camera device, video device, mobile device, etc.) and the target object (e.g., a person or other sound source of interest in the video scene). The distance to a target object may be estimated through image analysis. For example, the distance to a target object may be estimated by detecting the target object in each frame using an object detection algorithm, calculating the size of a bounding box around a target object image, and referencing against known dimensions to compute an actual distance. Some examples involve adjusting audio source spatialization parameters of an audio sourceD25022W001corresponding to the target object according to the target object position data during capture of audio data and video data. The described approaches can ensure accurate audio- visual synchronization and optimal sound quality, even in dynamic recording environments, such as recording environments in which one or more audio sources and / or a capture device are moving.1. Depth Estimation

[0033] The present disclosure proposes methods that utilize target object position data that indicates depth information based on video data captured by a capture device (e.g., a camera device, video device, mobile device, etc.) to real-time adjust the position of sound sources in a three-dimensional (3D) audio scene, ensuring synchronization with objects or individuals in the video frame.

[0034] Some disclosed methods may include one or more of the following elements.• Depth Map Acquisition: Target object position data including depth information, particularly target object position data corresponding to audio sources in a video scene, may be captured or estimated for each video frame. In some examples, the depth information for target objects — such as target objects corresponding to audio sources — may be obtained using capture devices such as structured light sensors or time-of-flight (ToF) cameras. In other examples, the distance to target objects may be estimated by detecting the target objects in a video frame using an object detection algorithm, calculating the size and location of a bounding box around each detected target object image, and computing target object according to the bounding box size and location. Detailed examples are provided below. The captured or estimated target object position data, including depth information, may be used to construct a three-dimensional spatial model. In some examples, target object position data may be obtained for each video frame.• Sound Localization Calculation: Based on the positions of target objects (e.g., a person and / or other sound source of interest) in the three-dimensional space of the video scene, some methods involve processing their corresponding sound sources to achieve a spatial audio representation. Some such methods involve repeatedly obtaining target object position data, e.g., for some or all sound sources in each videoD25022W001frame, ensuring consistency between the visual presentation and the spatial audio in the video scene.• Dynamic Update: Some methods involve continuously adjusting audio source spatialization parameters in real time when either the target object or the capture device changes in position relative to one another, ensuring the accuracy and continuity of sound localization. The audio source spatialization parameters may correspond to audio source / target object locations relative to the capture device. For some examples, creating audio source metadata includes at least one of the audio source spatialization parameters or the target object position data, and embedding the audio source metadata in the output content stream. The audio source metadata may include audio source location data, audio source size data, etc. The audio source metadata may, for example, be Atmos™ audio object metadata. The audio source metadata may be used to facilitate accurate rendering of audio objects in a playback environment.

[0035] Another alternative price-friendly solution is proposed for use with some mobile phone devices that are without ToF or other depth sensing capabilities. In some examples a novel pixel distance correction algorithm is applied, which may also use a bounding box based distance estimation method. In some such examples, the distance between a video camera and a person in a vlog is estimated by analyzing the bounding box of the person within the image frame.

[0036] Some steps in example methods are as follows:• Detect the target object(s) (e.g., person or persons, or other audio sources of interest) in each video frame using an object detection algorithm. In some examples, a control system may perform a target object detection process by implementing a “you only look once” (YOLO) object detection model. YOLO is a real-time computer vision process that identifies and localizes objects in an image by running a single neural network pass over the entire image. YOLO treats the object detection task as a regression problem, predicting bounding boxes and class probabilities directly from the input image to achieve high processing speeds. In other examples, a control system may perform a target object detection process by implementing a Faster R- CNN (region-based convolutional neural network) object detection model. Unlike R- CNN and Fast R-CNN, which relied on external algorithms such as selective searchD25022W001for generating target object / region proposals, Faster R-CNN uses a small convolutional network (the region proposal network (RPN)) to directly make target object proposals. This RPN shares convolutional layers with the main target object detection neural network, drastically reducing the time spent on proposal generation and making the overall target object detection process much faster than previously- deployed methods.• Calculate the size of the bounding box around the detected target object(s). For example, a control system may determine the height and width of the bounding box in a particular video frame. The bounding box size may be determined in units of image pixels. In some examples, the position of the bounding box in a video frame may be determined. For example, the position of the bottom of the bounding box — e.g., the position of the midpoint of the bottom edge of the bounding box — may be determined with reference to the bottom of the video frame.• Estimate the distance(s) based on the reference size(s) associated with the target object(s), using geometric calculations such as triangulation or camera calibration parameters (e.g., focal length, sensor size). Some examples of estimating target object distance with reference to camera focal length are described below with reference to Figures 1-4. In some examples, the average dimensions of a typical person may be used for the reference size. In another example, a person’s known dimensions may be used for the reference size. For example, the author of a vlog may provide the author’s known dimensions (e.g., height and width), the dimensions of another person or object in a video scene, etc. In some examples, a person’s known dimensions may be provided in response to a prompt caused by the control system. The prompt may be a visual prompt, e.g., on a graphical user interface (GUI), an audio prompt, or both.• Optionally, the reference sizes can be adjusted based on estimations for different sizes of people: e.g., an estimated adult being one reference size, and an estimated child being another reference size. Some examples are provided below.Optionally, the reference sizes can be adjusted based on a posture estimation. For example, the position of a target person may be estimated as standing, bending, sitting, squatting, kneeling, or lying down, and the reference size may be adjustedD25022W001based on the estimated position. Some such examples involve estimating the posture of a person by obtaining the aspect ratio of the box boundary, calculating a head-to- body ratio, and then selecting reference body and / or body part dimensions. The specific dimensions may be determined based on the average height and head-to-body ratio of a country, the average height and head-to-body ratio of a region within the country, etc.• Optionally, the reference sizes can be adjusted based on an estimated body type of a target person. For example, the body type may be estimated based on the heightwidth ratio or head-body ratio, and the reference size may be adjusted based on the estimated body type. In some examples, body types may be classified as ectomorph, endomorph and mesomorph, and mesomorphs may be assumed to have the reference size (height and width). In some such examples, bodies classified as ectomorph types will be assigned a higher height / width ratio than the reference size, and bodies classified as endomorph types will be assigned a lower height / width ratio than the reference size.

[0037] The above-described methods allow for accurate distance estimation without requiring the capture device to include any type of dedicated depth sensor, enhancing the system's versatility and practicality in various recording environments.

[0038] In various example implementations, a person may be a target object and the reference size of a person may be based a worldwide average height of an adult of approximately 1.7 m, and an average width of 0.4 m. This reference size is merely one example, and it is expected that the reference size can be set based on any appropriate criteria. For example, several reference sizes can be determined by combining posture detection, adult vs. child detection or other body analysis algorithms. In some examples, the reference size may be based on a person engaged in an activity that would augment or diminish the person’s natural height, such as when operating a bicycle, operating a motorcycle, operating a wheelchair, operating a golf cart, operating a snowmobile, operating a jet ski or wave runner, skateboarding, skiing or snowboarding, to name a few. For simplicity, in some examples we may assume that we have an average body type of an average height and width person.D25022W001

[0039] With the knowledge of the various characteristics of a digital camera or other capture device, the focal length can be obtained. Naturally, the type of capture device as well as other specific features can be selected based on any desired image capture criteria.

[0040] Figure 1 shows an illustration of the vertical field of view of a camera. As with other disclosed examples, the types and numbers of elements that are shown in Figure 1 and described herein are merely presented by way of example. Other examples may include one or more different types of elements, different numbers of elements, or both. In the simple example shown in Figure 1, a captured image 125 is shown within a camera 105. In this example, the camera 105 includes a lens 110 having a focal length / and a vertical field of view of theta (Θ). Light 115 that passes through the focal point is focused by the lens 110 onto an image sensor 130 of a camera 105. The image sensor 130 is configured to convert photons into electrical signals. The image sensor 130 may, for example, be an active pixel sensor or a charge-coupled device. The camera 105 may, for example, include an image processor configured to convert the electrical signals into digital data corresponding to the captured image 125.

[0041] A simple image capture device such as a pinhole camera with a vertical field of view of 6 - 39°, with an image capture size of 640 x 640 pixels, thus having a width w = 640, has a focal length f that can be determined as:a>f=- 7o\(1)2 tan ( 2 J

[0042] Thermal and RGB type cameras have isotropic lenses and can be modeled with a pinhole camera model to approximate the same focal length, where f = fx= fyfor both x and y axes.

[0043] Figure 2 shows an example of the geometrical relationship between the actual size of a target object and a bounding box in a captured image. In this example, the target object 205 is a person having a height ph and a width pw. Here, the bounding box 210 in the captured image 125 corresponds to the target object 205. In this example, the bounding box 210 is an ithbounding box having a height hi and a width Wi. According to this example, the y coordinate of the upper left comer of the bounding box 210 is yi.

[0044] The apparent size of the target object 205 in the captured image 125 is related to the actual size of the target object 205 by the focal length f and the distance d of the target objectD25022W001205 from the camera 105. As shown in Figure 2, a triangle similarity property can be applied to demonstrate the relationship between apparent and actual target object size based on focal length f from Equation (1) and distance d to the target object (in this example, a person). For a given focal length f, the distance d to a person may be given by the following equation:d = f (2)Js*

[0045] In Equation 2, pwrepresents a person’s width, phrepresents the person’s height, and sbrepresents the size of the bounding box measured in pixels. For a given bounding box width of w and a bounding box height of h, the size of the bounding box sbmay be estimated as:sb= wh (3)

[0046] By applying Equation (1) and Equation (3) to Equation (2), the distance d to a person can be calculated by<*> fPwPh(4)ztant?-] wh

[0047] Based on a known or assumed reference person size, for a selected bounding box size sb, the distance d from the camera to a person in the field of view can be estimated.

[0048] Figure 3 is a graph that shows the relationship between the bounding box size and the distance to a person. In the example shown in Figure 3, the captured image has a size of 640 x 640 pixels. The graph 300 is based on an underlying assumption that the person has a reference person size corresponding to a height of 1.7 m and a width of 0.4 m. As illustrated in graph 300 of Figure 3, the relationship between the distance to a target object 205 and the size of the bounding box 210 is a non-linear relationship. All target objects 205 for which the corresponding bounding boxes 210 have a size sbexceeding approximately 5000 pixels have a corresponding distance d that is closer than 10 m to the camera. In Figure 3, the distance d is shown to grow exponentially for smaller-sized bounding boxes sb, thereby showing the importance (e.g., challenges in accuracy in distance measurement) when detecting small target objects.2. Pixel distance correction algorithmD25022W001

[0049] To make the algorithm robust for target objects for people of different body shapes and ages, a pixel distance correction algorithm is implemented in some examples disclosed herein. This pixel distance correction algorithm is based on an assumption that the lower edge of the captured image is closer to the capture device, e.g., to the camera, than the upper edge of the captured image. That is, if the midpoint of the lower edge of the target object's bounding box is farther away from the lower edge of the captured image, then the target object is assumed to be farther away from the user. Although the assumption may not be true for all target objects, the assumption can be considered approximately true when the object is a person.

[0050] The true distance to a target object is linearly related to the “pixel distance’ — which is the distance in units of pixels from the lower edge of the captured image to the lower edge of a bounding box corresponding to the target object — raised to the power of 1 / 2. So, for a captured image containing N target objects, a control system may calculate the distance to the Ithtarget object, d,, i e [1,2,... A] using formula (4), which corresponds to the pixel distanced,mts' fromjower ecjge of,hebounding box to the lower edge of the captured image. In some examples, the control system may then calculate the average ratio ratioavebetween d, and d‘ntsfor N target objects, e.g., as follows:

[0051] In some examples, -ntscan be obtained by:d-nts= height — yt—L(6)

[0052] In Equation 6, height represents the height of the captured image, y, corresponds to the y coordinate of the upper left corner of the ithbounding box, and hi represents the height of the i'hbounding box. By applying equation (5), the corrected distance d1° of target object i can be calculated by:

[0053] Figure 4 shows examples of pixel distance corrections according to the above- described example. Figure 4 shows bounding boxes 410A, 410B, 410C, 410D and 410E, corresponding to a total of 5 target objects (here, 5 different people of varying size and location), in the video scene of captured image 425. In these examples, a box is outliningD25022W001each of the bounding boxes 410A-410E to outline their respective identified locations in the video frame of captured image 425. The numbers above each bounding box denote the original estimated distance of the corresponding target object to the capture device using Equation 4. The numbers shown below each bounding box denote the corrected distances calculated according to Equations 5, 6 and 7.

[0054] Figure 4 shows an example of a “pixel distance” 415B for bounding box 410B that may be used in these calculations. The pixel distance 415B is the distance, measured in pixels, from the lower edge 411 of the bounding box 410B to the lower edge 413 of the captured image 425. The pixel distance 415B corresponds to the pixel distance d-ntsof Equations 5, 6 and 7.

[0055] In Figure 4, target object corresponding to the bounding box 410A has an original distance estimated to be 3.78m and a corrected distance of 3.35m. The target object corresponding to the bounding box 410B has an original estimated distance of 5.08m and a corrected distance of 6.13m. The target object corresponding to the bounding box 410C has an original estimated distance of 9.35m and a corrected distance of 7.71m. The target object corresponding to the bounding box 410D has an original estimated distance of 7.47m and a corrected distance of 8.34m. The target object corresponding to the bounding box 410E has an original estimated distance of 7.98m and a corrected distance of 8.54m. Because the target object corresponding to the bounding box 410C is a child in this example, Figure 4 shows how a mismatch between a reference size of an adult and a child may lead to an incorrect distance measurement. The original error in the child’s distance measurement is ameliorated by the presently disclosed methods.

[0056] According to some examples, a single capture device will be used to capture all of the audio data and video data that will be obtained for a content stream. However, in some use cases, multiple devices may be used for audio and video capture. The disclosed methods also apply to such use cases.

[0057] Figure 5 shows an example of multiple devices being used for audio and video capture. According to this example, the capture device 1001 is being used to capture video data corresponding to target objects 505 A, 505B, 505C and 505D, which are musicians in this example. In this example, the capture device 1001 is a cellular telephone mounted on a tripod 535. The captured image 525 displayed on the capture device 1001 includes target objects 505A-505D. Audio data is being captured by microphones 515A, 515B, and 515C, as well as by at least two “pickup” transducers attached to the guitars shown in Figure 5, as evidenced by the cords 530 A and 530B.D25022W001

[0058] In a capture implementation such as that shown in Figure 5, the primary device capturing the visual content (the capture device 1001) may primarily record ambient environmental sounds, while other devices (such as the microphones 515A-515C and pickups) may capture audio for specific objects, individuals, or instruments. This strategic approach ensures a harmonious blend of all audio elements, creating an immersive and realistic sound experience that synchronizes with the visual narrative.

[0059] The disclosed methods may be applied to vlogs, reality TV programs, or other high-quality User-Generated Content (UGC) and Professional Generated Content (PGC) where multiple individuals are present in a video scene. Some disclosed implementations are configured for managing multi-microphone configurations, providing precise synchronization by implementing one or more of the disclosed techniques. In such cases, it is important to capture each person’s voice clearly while synchronizing the audio with the video, creating a more immersive and realistic audio-visual experience for audiences. The control system may analyze the distance and position of each individual within the video frame and may adjust the audio data corresponding to each person within the video frame accordingly, including volume and spatial orientation. The control system may be configured to synchronize the audio data with the video data, ensuring that the sound matches the visuals, providing a sense of depth and directionality.

[0060] In some examples, a control system may be configured to provide some or all of the functionality described in Sections 3-6, below.3. Distance-Related Automatic Volume Adjustment

[0061] Some disclosed examples involve measuring real-time distances between target objects (e.g., the shooting subjects in a video scene) and the capture device (e.g., camera device, video device, mobile device, etc.). Figure 6 is a top view that shows examples of two target objects and a capture device. In this example, based on image analysis and implementation of the disclosed methods, the capture device 1001 has determined that the target object 505e is at a distance dl and an angle cpl relative to the capture device 1001, and that the target object 505f is at a distance d2 and an angle tp2 relative to the capture device 1001. In these examples, the angles cp 1 and cp2 are determined according to the positions of the centroid 640a and the centroid 640b, respectively.

[0062] This disclosure provides methods for dynamically controlling the intensity of sounds based on the distance between objects or people and the camera, including:D25022W001• Distance and Angle Measurement: In some examples, a control system may be configured to obtain target object position data, including target object distance data and angle data for each of one or more target objects. According to some examples, the control system may be configured to create audio source metadata corresponding to the target object position data and to embed the audio source metadata in an output content stream. In some examples, the control system may be configured to construct a 3D spatial model of a video scene, including “depth map data” indicating the distances or “depths” to each of the target objects.• Audio Gain Adjustment: Based on predefined distance-gain curves or through adaptive determination using machine learning models, a control system may be configured to adjust the intensity of sound signals to match real-world audio source characteristics, including target object position data corresponding to audio sources. For example, if two people are talking in a video scene, the audio corresponding to the speech of the closer person may be made louder than the audio corresponding to the speech of the person farther away. According to some examples, the control system may be configured to create audio source metadata corresponding to the target object position data and to embed the audio source metadata in an output content stream.D25022W0014.Noise Control

[0063] Some aspects of this disclosure involve a microphone array control methods, advanced signal processing algorithms, or both, to detect and suppress unnecessary environmental noise in real-time.• Beamforming Technology: Microphone arrays are utilized with adaptive filtering techniques to actively focus on target sound sources while suppressing noise from other directions. For example, a control system may implement receiver-side beamforming techniques to enhance sound from a target sound source, such as a person who is speaking in a video scene, while suppressing audio from other sound sources.• Adaptive Noise Elimination: In some implementations, a control system may implement advanced noise elimination algorithms such as Independent Component Analysis (ICA) to separate and reduce background noise components in real-time, enhancing the clarity of primary sounds. For example, ICA may be used to identify audio corresponding to a person of interest who is speaking in a video scene, such as the author of a vlog. Audio corresponding to the person of interest may be enhanced relative to audio from other sources.5. Dynamic Microphone Signal Fusion

[0064] In some examples, the control system may process differences in signal-to-noise ratio (SNR), reverberation, and other attributes between microphones to align the corresponding audio properly. This disclosure proposes methods for automatically selecting and fusing optimal microphone signals based in part on distance information:• Microphone Array Management: According to some examples, a control system may be configured to dynamically determine and evaluate one or more quality indicators (such as SNR, interference levels, etc.) of each microphone signal. The microphone signal evaluation may be based on part on the position and movement of target objects (e.g., the subject(s) of interest in a video scene).• Adaptive Weighted Fusion: In some examples, a control system may be configured to perform a real-time weighted overlay on the audio signals captured by different microphones. Weights may, for example, be determined by the position and distanceD25022W001of target sound sources relative to each microphone (e.g., with the highest weight being given to the microphone nearest to a target sound source), enhancing the quality of sound output.6. Real-Time and Post-Production Metadata Management

[0065] Some aspects of this disclosure involve a metadata recording and application system:• Real-Time Embedding: During the process of capturing audio data and video data, metadata may be created corresponding to key parameters such as depth information and audio adjustments. This metadata may be embedded in the output content stream.• Post-Production Optimization: Some aspects of this disclosure involve using these metadata in a post-production editing process to fine-tune sound localization, volume adjustment, and noise control, ensuring the final audio output is optimized for an immersive experience.

[0066] Figure 7 illustrates a block diagram of a system for content generation in accordance with aspects of the present disclosure. As with other disclosed examples, the types and numbers of elements that are shown in Figure 7 and described herein are merely presented by way of example. Other examples may include one or more different types of elements, different numbers of elements, or both.

[0067] fhe illustrated system 700 of Figure 7 includes a set of blocks to illustrate various functional partitions. The system 700 may, for example, be implemented by an instance of the control system 1010 of Figure 10. According to this example, system 700 includes a video analysis block 705, an audio processing block 710 and an audio / video synchronization block 715.

[0068] In this example, the video analysis block 705 is configured to analyze frames of captured video data 702 and to provide video analysis data 707 to the audio processing block 710 and the audio / video synchronization block 715. The video analysis data 707 may include target object position data for each target object of interest in a video frame.

[0069] The video analysis block 705 may be configured to analyze video data for the detection and spatial tracking of particular audio sources of interest. In some examples, people and musical instruments may be the primary audio sources of interest. This focus is based on their prevalence and significance in typical applications such as concerts, interviews, and video conferences, where human voices and musical instruments are the mostD25022W001common and critical sound sources. The video analysis block 705 may employ advanced artificial intelligence (Al) models that excel in recognizing human faces, facial expressions, lip movements, and body language to identify individual people as likely audio sources. Similarly, in order to detect musical instruments, the video analysis block 705 may use deep learning algorithms trained on extensive datasets of various musical instruments, enabling the video analysis block 705 to detect visual cues such as the shape, texture, and movement associated with instrument play. By prioritizing “person” and “instrument,” the video analysis block 705 enhances detection accuracy and efficiency, ensuring that the most relevant audio sources are identified and tracked effectively. This prioritization not only improves the overall performance of the module but also elevates the user experience by delivering more precise and meaningful results in real-world scenarios.

[0070] According to this example, the audio processing block 710 is configured to process captured audio data 701 and to output processed audio data 712 to the audio / video synchronization block 715. In this example, the audio processing block 710 is configured for audio signal processing, which may include source separation, noise reduction, acoustic feature extraction, or combinations thereof. The audio processing block 710 may be configured for performing a wide array of audio signal processing tasks, ensuring that the processed audio data 712 meets high-quality standards. The audio processing block 710 may perform processes such as audio source separation and noise reduction to isolate individual audio elements and eliminate unwanted background interference. The audio processing block 710 may perform acoustic feature extraction to analyze and enhance specific characteristics of the audio signals. To further refine the audio quality, the audio processing block 710 may incorporate dynamic range compression (DRC) in order to balance the loudness and softness of sounds, ensuring a consistent listening experience. The audio processing block 710 may align the SNR across different audio sources to maintain uniformity in sound clarity. The audio processing block 710 may be configured for sibilance removal, ameliorating the high-frequency harshness often present in speech and resulting in smoother and more pleasant audio. The audio processing block 710 may be configured for reverberation consistency adjustment in order to ensure that the spatial effects of sounds are coherent, preventing an unnatural listening experience. The audio processing block 710 may be configured to make equalization (EQ) adjustments in order to fine-tune frequency responses, optimizing the overall tonal balance. The audio processing block 710 may be configured to detect whether an audio signal is speech or music and to apply tailored processing techniques accordingly. This differentiation allows for specific enhancements, such as noise reduction for speech andD25022W001preservation of musical nuances. The audio processing block 710 may be configured to remix the processed audio sources in order to complement the video footage by implementing a binaural room impulse response (BRIR), for which the distance and direction of arrival (DOA) may be obtained from the video analysis data 707.

[0071] In this example, the audio / video synchronization block 715 is configured to produce an output content stream of audio and video data 720 based on the video analysis data 707 and the processed audio data 712. In this example, the audio / video synchronization block 715 is configured to ensure precise alignment between the processed audio data 712 and the video data, based at least in part on the video analysis data 707. This alignment, which is also referred to herein as “real-time synchronization,” is important for providing immersive experiences such as those offered by Dolby Atmos™. In some examples, the audio / video synchronization block 715 may be configured for one or more of the following:• Precise synchronization between audio and video using timestamp embedding and phase-correlation algorithms;• Dynamic control of DOA: As characters or objects move in the video scene, the sound sources should adapt accordingly;• Management of distance-related audio cues, ensuring that the perceived proximity of sounds matches the visual motion;• Maintaining consistent reverberation and spatial effects, creating a unified auditory environment;• Maintaining control over SNR, thereby providing audio clarity in various scenes.

[0072] In some examples, the audio / video synchronization block 715 may be configured to support Dolby Atmos’s object-based audio, enabling precise 3D sound placement and dynamic metadata adaptation. By harmonizing these advanced audio features with visual elements, the audio / video synchronization block 715 may be configured to enhance immersion, making the system 700 capable of providing content for premium audiovisual experiences.

[0073] Figure 8 is a detailed flow diagram that outlines various example methods 800 according to some disclosed implementations. The example methods 800 may be partitioned into blocks, such as blocks 805, 810, 815, 820, 825 and 830. The various blocks may be described as operations, processes, methods, steps, acts or functions. The blocks of methods 800, like other methods described herein, are not necessarily performed in the order indicated. In some implementations, one or more of the blocks of methods 800 may beD25022W001performed concurrently. Moreover, some implementations of methods 800 may include more or fewer blocks than shown and / or described. The blocks of methods 800 may be performed by one or more devices, for example, the apparatus 1001 of Figure 10.

[0074] Processing may commence at block 805. Block 805 involves “obtaining, by a capture device, audio data and video data, wherein the capture device includes a video recording module and a microphone system including one or more microphones.” In some examples, block 805 may involve obtaining the audio data and the video data via the capture system 1020 of Figure 10. Processing may continue to block 810.

[0075] Block 810 involves “identifying, by a control system and based at least in part on the video data, one or more of the target objects.” In some examples, the control system may be a capture device control system. In other examples, the control system may reside in a device other than the capture device. In some examples, block 810 may involve implementing a YOLO object detection module or a Faster R-CNN object detection module. In some examples, one or more of blocks 810-830 may be performed by a control system of the capture device. In other examples, one or more of blocks 810-830 may be performed by a control system of another device, which may or may not include a capture system 1020. Processing may continue from block 810 to block 815.

[0076] Block 815 involves “determining, by the control system, one or more bounding box parameters of a bounding box around each of the one or more target objects.” According to some examples, block 815 may involve determining a width and a height of a bounding box around each of the one or more target objects. In some examples, block 815 may involve determining a pixel distance from the bottom of a bounding box to the bottom of a captured image. Processing may continue to block 820.

[0077] Block 820 involves “estimating, by the control system and based at least in part on the one or more bounding box parameters, target object position data for each of the one or more target objects, wherein the target object position data indicates a distance from the capture device to each of the one or more target objects.” In some examples, block 820 may involve using an assumed target object size, such as an assumed size of a person. According to some examples, block 820 may involve implementing an algorithm such as Equation 4 of this disclosure. Processing may continue to block 825.

[0078] Block 825 involves “adjusting, by the control system, audio source spatialization parameters of one or more audio sources corresponding to each of the one or more target objects according to the target object position data, wherein the identifying, the determining, the estimating and the adjusting all occur during capture of the audio data and the videoD25022W001data.” The audio source spatialization parameters may include audio source position and / or audio source size. According to some examples, block 825 may involve adjusting an audio source position based on the target object position data. Processing may continue to block 830.

[0079] Block 830 involves “producing, by the control system, an output content stream that includes output video data and output audio data." According to some examples, block 830 may involve producing the output audio and video data 720 of Figure 7.

[0080] In some examples, method 800 may involve maintaining real-time synchronization between the audio data and the video data. In some examples, the real-time synchronization may be performed during capture of the audio data and the video data and may be based, at least in part, on repeated estimation of the target object position data to produce updated target object position data. The real-time synchronization may be based, at least in part, on repeated adjustment of the audio source spatialization parameters according to the updated target object position data.

[0081] According to some examples, method 800 may involve creating audio source metadata corresponding to the audio source spatialization parameters and / or the target object position data and embedding the audio source metadata in the output content stream. In some examples, method 800 may involve modifying an intensity of audio corresponding to one or more audio sources based, at least in part, on the target object position data. Some such examples may involve creating audio source metadata corresponding to this intensity modification of audio and embedding the audio source metadata in the output content stream.

[0082] In some examples, method 800 may involve capturing, with a multi -microphone array of the capture device, audio data associated with an audio scene. The audio scene may include ambient sound signals and target sounds associated with audio data from the one or more target objects. Some such examples may involve applying an emphasis process to emphasize target sounds in the audio scene while reducing interference from other sounds in the audio scene. The emphasis process may include a beam forming process and / or an adaptive filtering process. Some such examples may involve separating and suppressing background noise from the audio scene by applying an Independent Component Analysis (ICA) process. Some such examples may involve evaluating one or more microphone signal quality metrics of microphone signals from each microphone of the multi-microphone array and adjusting weighting parameters for the microphone signals from each microphone of the multi-microphone array based on evaluated quality metrics.D25022W001

[0083] According to some examples, method 800 may involve constructing a 3D spatial model of a video scene based, at least in part, on the target object position data. In some examples, method 800 may involve performing a sound localization calculation based, at least in part, on the target object position data.

[0084] Estimating the target object position data may involve calculating the size of each bounding box based on a reference size associated with an identified target object. In some examples, method 800 may involve adjusting the reference size based on one or more estimated characteristics of the identified target object in the bounding box. The one or more estimated characteristics may include: (a) an estimated person type of either a child or an adult: (b) an estimated body type based on a height-width ratio or head-body ratio of a person; (c) an estimated gesture type of standing, bending, sitting, squatting, kneeling, or lying down; or (d) combinations thereof.

[0085] Estimating the target object position data may involve using geometric calculations, including triangulation, camera calibration parameters, or both triangulation and camera calibration parameters. The camera calibration parameters may include focal length, sensor size, or both focal length and sensor size.

[0086] Estimating the target object position data may involve identifying a midpoint between an upper edge of the bounding box and a lower edge of the bounding box of an identified target object, evaluating the target object position data to determine a distance to the midpoint, and estimating a distance to the identified target object based a determined distance to the midpoint of the bounding box. Identifying the midpoint may involve determining a centroid of the identified target object.

[0087] Estimating the target object position data may involve identifying a first midpoint along a lower edge of a first bounding box of a first identified target object, calculating a first pixel distance corresponding to a number of pixels between an image bottom edge and the first midpoint, and determining a first estimated target object distance based, at least in part, on the first pixel distance. Determining the first estimated target object distance may be based on an assumption that the image bottom edge is closer to the capture device than an image top edge. The first identified target object may be a person. Determining the first estimated target object distance may be based, in part, on a square root of the first pixel distance.

[0088] Estimating the target object position data may involve determining N estimated target object distances for N identified target objects based, at least in part, on N pixel distances. The N identified target objects may correspond to N people. Estimating the target objectD25022W001position data for the N identified target objects may involve a distance estimate correction process. The distance estimate correction process may involve computing a calibration ratio based on averaging a ratio of each of the N estimated target object distances and square roots of each of the N pixel distances and calculating corrected target object distance estimates by multiplying a square root of each of the N pixel distances by the calibration ratio.

[0089] Figure 9 illustrates a schematic block diagram of an example device or system architecture that may be used to implement various aspects of the present disclosure.Architecture 901 may include, but is not limited to, servers and client devices, stand-alone devices, systems, etc., which may be configured to perform the methods that are described with reference to any or all of the disclosed figures. As shown, the architecture 901 includes central processing unit (CPU) 941, which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 942 or a program loaded from, for example, storage unit 948 to random access memory (RAM) 943. The CPU 941 may be, for example, an electronic processor 941. The CPU 941 is an instance of what may be referred to herein as a “control system” or an element of a control system. The ROM 942 and RAM 943 are instances of what may be referred to herein as a “memory system” or an element of a memory system. In RAM 943, the data required when CPU 941 performs the various processes is also stored, as required. In this example, CPU 941, ROM 942, and RAM 943 are connected to one another via bus 944. Input / output (I / O) interface 945 is also connected to bus 944. The bus 944 and the I / O interface 945 are instances of “interface system” elements as that term is used in this disclosure.

[0090] According to this example, the following components are connected to I / O interface 945: input unit 946, which may include a keyboard, a mouse, or the like; output unit 947, which may include a display system including one or more displays, a loudspeaker system including one or more loudspeakers, etc.; storage unit 948 including a hard disk, or another suitable storage device; and communication unit 949 including a network interface card such as a network card (e.g., wired or wireless). The communication unit 949 may be referred to herein as being part of an interface system.

[0091] In some implementations, the input unit 946 may include a microphone system that includes one or more microphones. In some examples, the microphone system may include two, three or more microphones in different positions, enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0092] According to some implementations, the output unit 947 may include systems with various numbers of loudspeakers. Output unit 947 — and / or another component of theD25022W001apparatus 901, such as the CPU 941 — may be capable of rendering audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0093] In some embodiments, communication unit 949 may be configured to communicate with other devices (e.g., via a network). In this example, drive 950 is also connected to I / O interface 945, as required. Removable medium 951, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 950, so that a computer program read therefrom may be installed into storage unit 948, as required. A person skilled in the art would understand that although apparatus 901 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0094] Figure 10 is a block diagram that shows examples of components of an apparatus or system capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 10 are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus or system 1001 may be, or may include, a device that is configured for performing at least some of the methods disclosed herein, such as a cellular telephone, a video camera, a tablet device, a smart audio device, a laptop computer, etc.

[0095] In this example, the apparatus or system 1001 includes an interface system 1005, a control system 1010 and a capture system 1020. The interface system 1005 may, in some implementations, be configured for providing audio data and video data from the capture system 1020 to the control system 1010. In some examples, interface system 1005 may be configured for outputting an output content stream that includes output video data and output audio data to the memory system 1015, to another device, etc. In some alternative implementations, the apparatus or system 1001 may not include a capture system 1020. In some such implementations, the apparatus or system 1001 may be a server that is configured for performing at least some of the methods disclosed herein.

[0096] The interface system 1005 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 1005 may include one or more wireless interfaces. The interface system 1005 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, aD25022W001display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 1005 may include one or more interfaces between the control system 1010 and a memory system, such as the optional memory system 1015 shown in Figure 10.However, the control system 1010 may include a memory system in some instances.

[0097] The control system 1010 may, for example, include a general purpose single- or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0098] In some implementations, the control system 1010 may reside in more than one device. For example, a portion of the control system 1010 may reside in a capture device (such as a cellular telephone, a video camera, etc.) and another portion of the control system 1010 may reside in another device, such as a server.

[0099] In some implementations, the control system 1010 may be configured for performing, at least in part, the methods disclosed herein. Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 1015 shown in Figure 10 and / or in the control system 1010. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, include instructions for controlling at least one device to process audio data. The software may, for example, be executable by one or more components of a control system such as the control system 1010 of Figure 10.

[0100] In this example, the apparatus or system 1001 includes a capture system 1020. The capture system 1020 may include a microphone system including one or more microphones. In some implementations, the microphone system may include a multimicrophone array. The capture system 1020 may include a video recording module configured for capturing video data. The video recording module may include one or more video cameras.D25022W001

[0101] According to some implementations, the apparatus or system 1001 may include the optional loudspeaker system 1025 shown in Figure 10. The optional loudspeaker system 1025 may include one or more loudspeakers. Loudspeakers may sometimes be referred to herein as “speakers.”

[0102] In some implementations, the apparatus or system 1001 may include the optional sensor system 1030 shown in Figure 10. The optional sensor system 1030 may include a depth sensor system, which also may be referred to as a distance sensor system, a touch sensor system, a gesture sensor system, an optical sensor system, or combinations thereof.

[0103] In some implementations, the apparatus or system 1001 may include the optional display system 1035 shown in Figure 10. The optional display system 1035 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some instances, the optional display system 1035 may include one or more organic lightemitting diode (OLED) displays. In some examples wherein the apparatus or system 1001 includes the display system 1035, the sensor system 1030 may include a touch sensor system and / or a gesture sensor system proximate one or more displays of the display system 1035. According to some such implementations, the control system 1010 may be configured for controlling the display system 1035 to present a graphical user interface (GUI), such as a GUI related to implementing one of the methods disclosed herein.

[0104] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units and modules discussed above can be executed by control circuitry (e.g., CPU 601 in combination with other components of FIG.6A), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as nonlimiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.D25022W001

[0105] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0106] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine -readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0107] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0108] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):D25022W001

[0109] EEE 1. A method for dynamic audio positioning of target objects corresponding to audio sources located in a three-dimensional (3D) space, the method comprising: obtaining, by a capture device, audio data and video data, wherein the capture device includes a video recording module and a microphone system including one or more microphones; identifying, by a control system and based at least in part on the video data, one or more of the target objects; determining, by the control system, one or more bounding box parameters of a bounding box around each of the one or more target objects; estimating, by the control system and based at least in part on the one or more bounding box parameters, target object position data for each of the one or more target objects, wherein the target object position data indicates a distance from the capture device to each of the one or more target objects; adjusting, by the control system, audio source spatialization parameters of one or more audio sources corresponding to each of the one or more target objects according to the target object position data, wherein the identifying, the determining, the estimating and the adjusting all occur during capture of the audio data and the video data; and producing, by the control system, an output content stream that includes output video data and output audio data.

[0110] EEE 2. The method of EEE 1, further comprising maintaining real-time synchronization between the audio data and the video data, wherein the real-time synchronization is performed during capture of the audio data and the video data and is based, at least in part, on repeated estimation of the target object position data to produce updated target object position data.

[0111] EEE 3. The method of EEE 2, wherein the real-time synchronization is based, at least in part, on repeated adjustment of the audio source spatialization parameters according to the updated target object position data.

[0112] EEE 4. The method of EEE 2 or EEE 3, further comprising creating audio source metadata corresponding to at least one of the audio source spatialization parameters or the target object position data and embedding the audio source metadata in the output content stream.

[0113] EEE 5. The method of any one of EEEs 1-4, further comprising modifying an intensity of audio corresponding to one or more audio sources based, at least in part, on the target object position data.D25022W001

[0114] EEE 6. The method of EEE 5, further comprising creating audio source metadata corresponding to intensity modification of audio corresponding to one or more audio sources based, at least in part, on the target object position data and embedding the audio source metadata in the output content stream.

[0115] EEE 7. The method of any one of EEEs 1-6, further comprising: capturing audio data associated with an audio scene with a multi-microphone array of the capture device, wherein the audio scene includes ambient sound signals and target sounds associated with audio data from the one or more target objects; and applying an emphasis process to emphasize target sounds in the audio scene while reducing interference, wherein the emphasis process includes a beam forming process and / or an adaptive filtering process.

[0116] EEE 8. The method of EEE 7, further comprising separating and suppressing background noise from the audio scene by applying an Independent Component Analysis (ICA) process.

[0117] EEE 9. The method of EEE 7 or EEE 8, further comprising: evaluating one or more microphone signal quality metrics of microphone signals from each microphone of the multi-microphone array; and adjusting weighting parameters for the microphone signals from each microphone of the multi-microphone array based on evaluated quality metrics.

[0118] EEE 10. The method of any one of EEEs 1-9, wherein the audio source spatialization parameters comprise audio source position, audio source size, or both.

[0119] EEE 11. The method of any one of EEEs 1-10, further comprising constructing a 3D spatial model of a video scene based, at least in part, on the target object position data.

[0120] EEE 12. The method of any one of EEEs 1-11, further comprising performing a sound localization calculation based, at least in part, on the target object position data.

[0121] EEE 13. The method of any one of EEEs 1-12, wherein estimating the target object position data involves calculating a size of each bounding box based on a reference size associated with an identified target object.

[0122] EEE 14. The method of EEE 13, further comprising adjusting the reference size based on one or more estimated characteristics of the identified target object inD25022W001the bounding box and wherein the one or more estimated characteristics include one or more of the following: an estimated person type of either a child or an adult; an estimated body type based on a height-width ratio or head-body ratio of a person; and an estimated gesture type of standing, bending, sitting, squatting, kneeling, or lying down.

[0123] EEE 15. The method of any one of EEEs 1-14, wherein estimating the target object position data comprises using geometric calculations, including triangulation, camera calibration parameters, or both triangulation and camera calibration parameters, wherein the camera calibration parameters include focal length, sensor size, or both focal length and sensor size.

[0124] EEE 16. The method of any one of EEEs 1-15, wherein estimating the target object position data comprises: identifying a midpoint between an upper edge of the bounding box and a lower edge of the bounding box of an identified target object; evaluating the target object position data to determine a distance to the midpoint; and estimating a distance to the identified target object based a determined distance to the midpoint of the bounding box.

[0125] EEE 17. The method of EEE 16, wherein identifying the midpoint comprises determining a centroid of the identified target object.

[0126] EEE 18. The method of any one of EEEs 1-17, wherein estimating the target object position data comprises: identifying a first midpoint along a lower edge of a first bounding box of a first identified target object; calculating a first pixel distance corresponding to a number of pixels between an image bottom edge and the first midpoint; and determining a first estimated target object distance based, at least in part, on the first pixel distance.

[0127] EEE 19. The method of EEE 18, wherein determining the first estimated target object distance is based on an assumption that the image bottom edge is closer to the capture device than an image top edge.

[0128] EEE 20. The method of EEE 18 or EEE 19, wherein the first identified target object is a person.D25022W001

[0129] EEE 21. The method of any one of EEEs 18-20, wherein determining the first estimated target object distance is based, in part, on a square root of the first pixel distance.

[0130] EEE 22. The method of any one of EEEs 18-21, wherein estimating the target object position data further comprises determining N estimated target object distances for N identified target objects based, at least in part, on N pixel distances, wherein the N identified target objects correspond to N people.

[0131] EEE 23. The method of EEE 22, wherein estimating the target object position data for the N identified target objects further comprises a distance estimate correction process, comprising: computing a calibration ratio based on averaging a ratio of each of the N estimated target object distances and square roots of each of the N pixel distances; and calculating corrected target object distance estimates by multiplying a square root of each of the N pixel distances by the calibration ratio.

[0132] EEE 24. The method of any one of EEEs 1-23, wherein the control system is a capture device control system.

[0133] EEE 25. The method of any one of EEEs 1-23, wherein the control system resides in a device other than the capture device.

[0134] EEE 26. An apparatus, comprising: an interface system; and a control system configured to perform operations including the method of any one of EEEs 1-23.

[0135] EEE 27. A non-transitory computer-readable storage medium recording a program of instructions that is executable by a device to perform the method of any one of EEEs 1-23.

[0136] EEE 28. A system comprising: one or more processors, and a non-transitory computer-readable storage medium storing instructions that, when executed by the one or more processors, perform the method of any one of EEEs 1-23.

[0137] A person skilled in the art realizes that the present disclosure by no means is limited to the examples described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims.D25022W001

[0138] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0139] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many examples and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future examples. In sum, it should be understood that the application is capable of modification and variation.

[0140] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

Claims

D25022W001CLAIMSWhat is claimed is:

1. A method for dynamic audio positioning of target objects corresponding to audio sources located in a three-dimensional (3D) space, the method comprising:obtaining, by a capture device, audio data and video data, wherein the capture device includes a video recording module and a microphone system including one or more microphones;identifying, by a control system and based at least in part on the video data, one or more of the target objects;determining, by the control system, one or more bounding box parameters of a bounding box around each of the one or more target objects;estimating, by the control system and based at least in part on the one or more bounding box parameters, target object position data for each of the one or more target objects, wherein the target object position data indicates a distance from the capture device to each of the one or more target objects;adjusting, by the control system, audio source spatialization parameters of one or more audio sources corresponding to each of the one or more target objects according to the target object position data, wherein the identifying, the determining, the estimating and the adjusting all occur during capture of the audio data and the video data; andproducing, by the control system, an output content stream that includes output video data and output audio data.

2. The method of claim 1, further comprising maintaining real-time synchronization between the audio data and the video data, wherein the real-time synchronization is performed during capture of the audio data and the video data and is based, at least in part, on repeated estimation of the target object position data to produce updated target object position data.

3. The method of claim 2, wherein the real-time synchronization is based, at least in part, on repeated adjustment of the audio source spatialization parameters according to the updated target object position data.D25022W0014. The method of claim 2 or claim 3, further comprising creating audio source metadata corresponding to at least one of the audio source spatialization parameters or the target object position data and embedding the audio source metadata in the output content stream.

5. The method of any one of claims 1-4, further comprising modifying an intensity of audio corresponding to one or more audio sources based, at least in part, on the target object position data.

6. The method of claim 5, further comprising creating audio source metadata corresponding to intensity modification of audio corresponding to one or more audio sources based, at least in part, on the target object position data and embedding the audio source metadata in the output content stream.

7. The method of any one of claims 1-6, further comprising:capturing audio data associated with an audio scene with a multi-microphone array of the capture device, wherein the audio scene includes ambient sound signals and target sounds associated with audio data from the one or more target objects; andapplying an emphasis process to emphasize target sounds in the audio scene while reducing interference, wherein the emphasis process includes a beam forming process and / or an adaptive filtering process.

8. The method of claim 7, further comprising separating and suppressing background noise from the audio scene by applying an Independent Component Analysis (ICA) process.

9. The method of claim 7 or claim 8, further comprising:evaluating one or more microphone signal quality metrics of microphone signals from each microphone of the multi-microphone array; andadjusting weighting parameters for the microphone signals from each microphone of the multi-microphone array based on evaluated quality metrics.

10. The method of any one of claims 1-9, wherein the audio source spatialization parameters comprise audio source position, audio source size, or both.D25022W00111. The method of any one of claims 1-10, further comprising constructing a 3D spatial model of a video scene based, at least in part, on the target object position data.

12. The method of any one of claims 1-11, further comprising performing a sound localization calculation based, at least in part, on the target object position data.

13. The method of any one of claims 1-12, wherein estimating the target object position data involves calculating a size of each bounding box based on a reference size associated with an identified target object.

14. The method of claim 13, further comprising adjusting the reference size based on one or more estimated characteristics of the identified target object in the bounding box and wherein the one or more estimated characteristics include one or more of the following:an estimated person type of either a child or an adult;an estimated body type based on a height- width ratio or head-body ratio of a person; andan estimated gesture type of standing, bending, sitting, squatting, kneeling, or lying down.

15. The method of any one of claims 1-14, wherein estimating the target object position data comprises using geometric calculations, including triangulation, camera calibration parameters, or both triangulation and camera calibration parameters, wherein the camera calibration parameters include focal length, sensor size, or both focal length and sensor size.

16. The method of any one of claims 1-15, wherein estimating the target object position data comprises:identifying a midpoint between an upper edge of the bounding box and a lower edge of the bounding box of an identified target object;evaluating the target object position data to determine a distance to the midpoint; and estimating a distance to the identified target object based a determined distance to the midpoint of the bounding box.

17. The method of claim 16, wherein identifying the midpoint comprises determining a centroid of the identified target object.D25022W00118. The method of any one of claims 1-17, wherein estimating the target object position data comprises:identifying a first midpoint along a lower edge of a first bounding box of a first identified target object;calculating a first pixel distance corresponding to a number of pixels between an image bottom edge and the first midpoint; anddetermining a first estimated target object distance based, at least in part, on the first pixel distance.

19. The method of claim 18, wherein determining the first estimated target object distance is based on an assumption that the image bottom edge is closer to the capture device than an image top edge.

20. The method of claim 18 or claim 19, wherein the first identified target object is a person.

21. The method of any one of claims 18-20, wherein determining the first estimated target object distance is based, in part, on a square root of the first pixel distance.

22. The method of any one of claims 18-21, wherein estimating the target object position data further comprises determining N estimated target object distances for N identified target objects based, at least in part, on N pixel distances, wherein the N identified target objects correspond to N people.

23. The method of claim 22, wherein estimating the target object position data for the N identified target objects further comprises a distance estimate correction process, comprising:computing a calibration ratio based on averaging a ratio of each of the N estimated target object distances and square roots of each of the N pixel distances; andcalculating corrected target object distance estimates by multiplying a square root of each of the N pixel distances by the calibration ratio.

24. The method of any one of claims 1-23, wherein the control system is a capture device control system.D25022W00125. The method of any one of claims 1-23, wherein the control system resides in a device other than the capture device.

26. An apparatus, comprising:an interface system; anda control system configured to perform operations including the method of any one of claims 1-23.

27. A non-transitory computer-readable storage medium recording a program of instructions that is executable by a device to perform the method of any one of claims 1-23.

28. A system comprising:one or more processors, anda non-transitory computer-readable storage medium storing instructions that, when executed by the one or more processors, perform the method of any one of claims 1-23.