Rendering audio captured by multiple devices

The system processes audio from multiple sources to create immersive experiences by integrating binaural audio with listener head movement adjustments, addressing the lack of complete audio objects in user-generated content.

JP2025529877APending Publication Date: 2025-09-09DOLBY LABORATORIES LICENSING CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025511550
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-20
Filing Date
2023-08-21
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

User-generated content (UGC) often lacks complete audio objects due to consumer capture devices' limitations, and there is no effective way to integrate audio from binaural microphones and mobile devices, preventing immersive audio experiences.

Method used

A system that processes audio from multiple sources, including binaural microphones and mobile devices, adjusts audio based on listener head movements using head-related transfer functions and rebalancing techniques to create immersive experiences without requiring complete audio objects.

Benefits of technology

Enables immersive audio experiences by adapting audio output to listener head movements, enhancing user-generated content with binaural audio integration and interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529877000001_ABST
    Figure 2025529877000001_ABST
Patent Text Reader

Abstract

A method of audio processing includes receiving user-generated content having two audio sources, extracting an audio object and a residual signal, adjusting the audio object and the residual signal according to a listener's head movement, and generating a binaural audio signal by mixing the adjusted audio signals. In this way, the binaural signal is adjusted according to the listener's head movement without requiring a complete audio object.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to International Patent Application No. PCT / CN2022 / 114596, filed August 24, 2022, U.S. Provisional Patent Application No. 63 / 432,385, filed December 14, 2022, and U.S. Provisional Patent Application No. 63 / 509,121, filed June 20, 2023, each of which is incorporated by reference in its entirety.

[0002] The present disclosure relates to audio processing, and in particular to processing audio captured by binaural microphones and additional microphones. [Background technology]

[0003] Unless otherwise specified, the approaches described in this section are not prior art to the claims herein, and their inclusion in this section is not an admission that they are prior art.

[0004] Audiovisual capture devices are becoming increasingly popular among consumers. These devices include portable cameras, such as Sony Action Cam® cameras and GoPro® cameras, as well as mobile phones with integrated camera functionality. Devices typically capture audio simultaneously with video, for example, using mono or stereo microphones. Audiovisual content sharing systems, such as the YouTube® service and Twitch.tv® service, are also becoming increasingly popular. Users upload captured audiovisual content to content sharing systems or broadcast it simultaneously. Because this content is generated by users, it is called user-generated content (UGC), which distinguishes it from professional-generated content (PGC), which is typically created by professionals. UGC often differs from PGC in that it is created using consumer equipment, which may be less expensive and have fewer features than professional equipment. Another difference between UGC and PGC is that UGC is often captured in uncontrolled environments, such as outdoors, while PGC is often captured in controlled environments, such as recording studios.

[0005] Another difference between UGC and PGC is that PGC may use complete audio objects, while UGC may not. For example, a PGC content creator can place high-resolution audio at a specific object location, and the PGC system can generate an audio object that exactly corresponds to the creator's intent. This audio object is called a complete audio object. In contrast, UGC content creators generally cannot use complete audio objects.

[0006] Binaural audio involves audio recorded using two microphones positioned at the user's ears. Captured binaural audio is sometimes called immersive audio and provides an immersive listening experience when played through headphones. Compared to stereo audio, binaural audio also includes the user's head and ear shadows, resulting in time and level differences between the two ears when the binaural audio is captured. Binaural audio also differs from stereo in that stereo audio may contain crosstalk between speakers. PGC binaural audio may be captured in a studio environment with controllable sound sources and acoustics. UGC binaural audio may be captured with earphones and may contain unwanted sounds from the surrounding environment.

[0007] Head tracking (or head tracking) generally refers to tracking the orientation of a user's head in order to adjust the input or output to a system. In the case of audio, head tracking refers to altering the audio signal depending on the orientation of the listener's head. Summary of the Invention [Problem to be solved by the invention]

[0008] Existing audiovisual capture systems for UGC have many problems. One problem is that UGC often does not use the complete audio object, because consumer capture devices often cannot capture the complete audio object. Another problem is that, although UGC content creators can typically capture audiovisual content using a mobile phone or binaural audio content using binaural earphones, there is no good way for UGC content creators to integrate the output of these two devices.

[0009] In view of these issues, embodiments relate to processing audio from multiple sources and adjusting the audio based on the listener's head movements. [Means for solving the problem]

[0010] According to one embodiment, a computer-implemented method for audio processing includes receiving, by one or more playback devices, user-generated content (UGC) captured by a group of interconnected capture devices. Each audio source of the UGC corresponds to a respective characteristic in an audio scene. The method further includes receiving, by one or more playback devices, information indicative of listener behavior of a user of the one or more playback devices from one or more sensors of the one or more playback devices. The listener behavior may include head movement of the listener. The method further includes adapting the UGC according to the listener behavior, including compensating for characteristics of the audio source according to the listener behavior. The method further includes providing the listener with an interactive experience related to the audio scene by rendering the adapted UGC.

[0011] As a result, the output audio responds to the listener's head movements, even for user-generated content, without the need for complete audio objects or professionally generated content.

[0012] According to another embodiment, an apparatus includes a processor configured to control the apparatus to perform one or more of the methods described herein. The apparatus may further include similar details to one or more of the methods described herein.

[0013] According to another embodiment, a non-transitory computer-readable medium stores a computer program that, when executed by a processor, controls an apparatus to perform processes including one or more of the methods described herein.

[0014] The following detailed description and accompanying drawings provide a further understanding of the nature and advantages of the various embodiments. [Brief explanation of the drawings]

[0015] [Figure 1A] FIG. 1A is a diagram of a user with a UGC capture device. [Figure 1B] FIG. 1B is an illustration of a user with a UGC capture device.

[0016] [Figure 2] FIG. 2 is a block diagram of a system 200 for interactively rendering UGC captured on multiple devices.

[0017] [Figure 3] FIG. 3 is a block diagram illustrating additional details of the HRTF adjuster 220 (see FIG. 2).

[0018] [Figure 4] FIG. 4 is a block diagram illustrating additional details of rebalancer 230 (see FIG. 2).

[0019] [Figure 5] FIG. 5 is a block diagram illustrating additional details of mixer 240 (see FIG. 2).

[0020] [Figure 6] FIG. 6 is a device architecture 600 for implementing the features and processes described herein, according to an embodiment.

[0021] [Figure 7] FIG. 7 is a flow chart of a method 700 of audio processing. DETAILED DESCRIPTION OF THE INVENTION

[0022]

[0013] This specification describes technologies related to audio processing. In the following description, for purposes of explanation, numerous examples and specific details are set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure, as defined by the claims, may include some or all of the features in these examples, alone or in combination with other features described below, and may also include variations and equivalents of the features and concepts described herein.

[0023] In the following description, various methods, processes, and procedures are detailed. While certain steps may be described in a certain order, such order is primarily for convenience and clarity. Certain steps may be repeated multiple times, may occur before or after other steps, or may occur in parallel with other steps, even if those steps are otherwise described in a different order. A second step is required to follow a first step only if the first step must be completed before the second step can begin. If this situation is not clear from the context, it will be specifically noted.

[0024] The terms "and," "or," and "and / or" are used herein. Such terms are to be read as having an inclusive meaning. For example, "A and B" may mean at least "both A and B" or "at least both A and B." As another example, "A or B" may mean at least "at least A," "at least B," "both A and B," or "at least both A and B." As another example, "A and / or B" may mean at least "A and B" or "A or B." When an exclusive or is intended, this is specifically noted (e.g., "either A or B," "at most one of A and B")

[0025] This specification describes various processing functions in terms of structures such as blocks, elements, components, circuits, etc. Generally, these structures may be implemented by a processor controlled by one or more computer programs.

[0026] 1A-1B are diagrams of a user holding a UGC capture device. FIG. 1A is a side perspective view, and FIG. 1B is an overhead view. FIGS. 1A-1B show a user 102 holding a mobile phone 104 and wearing earphones 106a and 106b (collectively 106). The mobile phone 104 typically includes a camera, a microphone, a screen, a speaker, a processor, volatile and non-volatile memory and storage, a radio, and other components. Examples of mobile phones 104 include the Apple iPhone® mobile phone and the Samsung Galaxy® mobile phone. The earphones 106 can wirelessly connect to the mobile phone 104 via an IEEE 802.15.1 standard protocol, such as the Bluetooth® protocol. The earphones 106 typically include a speaker, a microphone, a processor, volatile and non-volatile memory and storage, a radio, and other components.

[0027] The user 102 uses these devices to capture UGC of their surroundings (referred to as audiovisual scene 110). To capture UGC, the user 102 may hold the mobile phone 104 in their hand or on a selfie stick. For example, the phone's screen (on the front, facing the user) is used to frame the video scene in front of the user, the phone's camera (on the back, facing the video scene) is used to capture video, and the phone's microphones are used to capture audio (e.g., a single microphone captures mono audio, two microphones capture stereo audio, etc.). The user may use earphones 106 to capture binaural audio of the audiovisual scene 110 while simultaneously capturing audio and video using the mobile phone 104.

[0028] As discussed in the background, existing UGC devices do not provide an easy way to integrate the output of two devices, especially when both devices are capturing audio. For example, when playing content, a listener may have to choose between the audio captured by the mobile phone 104 (which is not binaural audio) and the binaural audio captured by the earphones 106. The following sections describe how the systems described herein integrate these two audio sources. As another example, the listener may have earphones with head tracking capabilities, but there is no easy way to adjust the captured UGC binaural audio to account for the listener's head movements. The following sections also describe how the systems described herein address this issue.

[0029] 2 is a block diagram of a system 200 for interactively rendering UGC captured by multiple devices. System 200 can be implemented using multiple devices, including two or more capture devices (e.g., earphones and a mobile phone), a server device, a playback device, etc. A device implementing system 200 can include circuitry such as a microprocessor that executes a computer program that implements the functionality of system 200. System 200 includes an object extractor 210, a head-related transfer function (HRTF) adjuster 220, a rebalancer 230, a mixer 240, and a remixer 250.

[0030] The object extractor 210 receives the audio signal 262 and the binaural audio signal 264, performs object extraction, and generates one or more binaural objects 266 and a residual signal 268. The audio signal 262 is captured by a UGC audiovisual capture device, such as a mobile phone, and the audio signal 262 is captured simultaneously with the video data. The audio signal 262 typically has N channels corresponding to the number of microphones in the UGC audiovisual capture device. For example, the audio signal 262 may have one channel for captured mono audio, two channels for captured stereo audio, etc. The mobile phone used to capture the audio signal 262 may have two microphones (e.g., located below and above, left and right, rear and front, etc.), three microphones (located below, above, left, right, front, rear, or an omnidirectional microphone, etc.), etc. The binaural audio signal 264 is captured by a UGC audio capture device, such as binaural earphones. The binaural audio signal 264 typically has two channels.

[0031] The binaural objects 266 generally correspond to audio data that the object extractor 210 has localized to a particular location in the audiovisual scene. For example, bird songs or airplane noises can be extracted to generate a height object. Similarly, identified sounds originating from the left side of the capture device can be extracted to generate a second audio object, and identified sounds originating from the right side of the capture device can be extracted to generate a third audio object.

[0032] The residual signal 268 corresponds to the audio signal 262 and the binaural audio signal 264 excluding the binaural object 266. The residual signal 268 has N+2 channels and corresponds to the N-channel residual of the audio signal 262 (excluding the binaural object 266) and the 2-channel residual of the binaural audio signal 264 (excluding the binaural object 266).

[0033] The object extractor 210 may implement a machine learning system for extracting binaural objects 266 from audio input. Typically, a machine learning system has a model that is learned in a training phase using training data. During the operation phase, the machine learning system uses this model as part of processing the input data to generate the machine learning system's output.

[0034] According to one embodiment, a machine learning system implements a trained model based on the signal-to-noise ratio (SNR) of a training audio dataset. The model can have multiple submodels or layers and can be configured to handle sparse objects. In operation, the machine learning system performs feature extraction on the audio input (e.g., including SNR features), performs classification on the extracted features, and uses the model as part of generating binaural objects 266 based on the extracted features. The machine learning system may reduce or eliminate leakage from the classified objects by processing the extracted objects based on at least one of the SNR or audiovisual context to generate binaural objects 266. Additional details of this embodiment of the machine learning system are described in International Patent Application No. PCT / CN2022 / 114613.

[0035] The object extractor 210 may be implemented by a capture device of a UGC content creator, such as a mobile phone (e.g., mobile phone 104 of FIG. 1). For example, the mobile phone may capture an audio signal 262 using its microphone, receive a binaural signal 264 from earphones (e.g., earphones 106) connected to the mobile phone, for example, via a Bluetooth wireless connection, and locally generate a binaural object 266.

[0036] Alternatively, the object extractor 210 may be implemented by a server device. Instead of generating the binaural object 266 locally, the capture device (e.g., the mobile phone 104) transmits the audio signal 262 and the binaural signal 264 to a server, which generates the binaural object 266. The UGC content creator can then receive the binaural object 266 from the server and play it locally on the capture device or another device. Furthermore, other users can receive the binaural object 266 (along with the captured video, if the captured video was also transmitted to the server device) from the server and play it using their own playback devices.

[0037] Alternatively, the object extractor 210 may be implemented by a computer. Instead of generating the binaural objects 266 locally or uploading the captured signals to a server, the UGC content creator may connect a mobile phone to a personal computer that generates the binaural objects 266. The UGC content creator may then play back the binaural objects 266 using a computer or other device for local playback. Additionally, the UGC content creator may upload the captured video and received binaural objects 266 to a server for other users to play back using their own playback devices.

[0038] The HRTF adjuster 220 receives the binaural object 266 and head orientation information 270 and adjusts the binaural object 266 according to the head orientation information 270 to generate an adjusted binaural object 272. The head orientation information 270 may be generated by a playback device of the listener, for example, by a binaural headset having a gyroscope that tracks the movement of the headset as the listener's head moves. Thus, the adjusted binaural object 272 corresponds to the binaural object 266 that has been adjusted using HRTFs based on the head orientation information 270. Further details of the HRTF adjuster 220 are provided with reference to FIG. 3.

[0039] The rebalancer 230 receives the adjusted binaural objects 272 and the head orientation information 270 and rebalances the adjusted binaural objects 272 according to the head orientation information 270 to generate rebalanced binaural objects 274. In general, the rebalancer 230 makes level and timbre adjustments based on the listener's head movement as indicated by the head orientation information 270. Further details of the rebalancer 230 are described with reference to FIG.

[0040] Mixer 240 receives residual signal 268 and head orientation information 270 and mixes residual signal 268 according to head orientation information 270 to generate residual signal 276. In contrast to residual signal 268, which has N+2 channels, residual signal 276 has 2 channels. Further details of mixer 240 are described with reference to FIG. 5.

[0041] The remixer 250 receives the binaural object 274 and the residual signal 276 and mixes the audio corresponding to the binaural object 274 with the residual signal 276 to generate a modified binaural signal 278. In general, the remixer 250 renders the binaural object 274 into an intermediate binaural signal having two channels and adds this to the residual signal 276, which also has two channels, to generate the modified binaural signal 278 having two channels.

[0042] As described above, the functionality of system 200 may be realized by multiple devices. As one example, a UGC content creator may use their own capture device as a playback device. In such an embodiment, the UGC content creator's mobile phone may perform object extraction and the UGC content creator's earphones may be used to capture head movements during playback. As another example, the UGC content creator may provide the captured video and processed audio to a listener, and the listener's device may play back the audio modified by the listener's current head movements. In such an embodiment, the UGC content creator's mobile phone may be used to perform object extraction and the listener's earphones may be used to capture the listener's current head movements. As another example, a server may perform object extraction, and the UGC content creator or another listener may play back the audio modified by their current head movements.

[0043] Figure 3 is a block diagram illustrating additional details of the HRTF adjuster 220 (see Figure 2). The HRTF adjuster 220 includes a direction estimator 302, a delta HRTF generator 304, a delta HRTF calculator 306, and an object adjuster 308. In general, the HRTF adjuster 220 adjusts primarily for changes in azimuth (left and right) of the listener's head orientation, but also for changes in elevation (up and down).

[0044] The direction estimator 302 receives the binaural object 266, estimates the direction of arrival (DOA) of the sound represented by the object, and generates a weighting vector 320. The weighting vector 320 may correspond to the time-delay-of-arrival (TDOA) of the sound represented by the object. The direction estimator 302 may implement one or more techniques for estimating the DOA, including digital signal processor (DSP)-based techniques, machine learning (ML)-based techniques, etc. The DSP-based technique may operate based on the level difference and time difference of the sound represented by the object. An example of the DSP-based technique is described in Kwon, Byoungho, Youngjin Park, and Youn-sik Park, "Analysis of the GCC-PHAT technique for multiple sources", ICCAS 2010, 2010) <doi:10.1109 / ICCAS.2010.<5670137>. Another example of the DSP-based approach is described in Dmochowski, Jacek P., Jacob Benesty, and Sofiene Affes, “A generalized steered response power method for computationally viable source localization”, IEEE Transactions on Audio, Speech, and Language Processing 15, no. 8 (2007): 2510-2526 <doi: 10.1109 / TASL.2007.906694>. As an example of the ML-based approach, there is an adaptive boosting technique such as AdaBoost. Another example of the ML-based technique is a neural network having multiple layers between an input layer and an output layer, such as a deep neural network (DNN).

[0045] The delta HRTF generator 304 receives the head orientation information 270 and adjusts the HRTFs calculated for several predefined positions according to the head orientation information 270 to generate delta HRTFs 322. The number of predefined positions may be four, corresponding to front, back, left, and right. The delta HRTFs 320 then correspond to the HRTFs for the predefined positions adjusted according to the head orientation information 270. The use of predefined positions reduces the computational complexity of generating HRTFs based on the listener's head movements. The number of predefined positions can be adjusted as desired.

[0046] Equation (1) describes the operation of the delta HRTF generator 304. JPEG2025529877000002.jpg28148

[0047] In equation (1), ω is the angular frequency, and the HRTF is frequency dependent. i and φ i are the azimuth and elevation angles, respectively, of the i-th predefined position. Δθ and Δφ are the azimuth and elevation angles, respectively, changed due to head rotation and are based on head rotation information 270. In other words, the delta HRTF group 322 for a given predefined position is proportional to a function of at least one of the azimuth and elevation angles of the given predefined position after head rotation, and inversely proportional to a function of at least one of the azimuth and elevation angles of the given predefined position.

[0048] Delta HRTF calculator 306 receives weighting vector 320 and delta HRTFs 322 and applies weighting vector 320 to delta HRTFs 322 for predefined locations to generate weighted delta HRTFs 324. Equation (2) describes the operation of delta HRTF calculator 306. JPEG2025529877000003.jpg25149

[0049] In equation (2), W iis the weighting vector for the i-th predefined location. In other words, the weighted delta HRTF group 324 is the sum over the set of predefined locations of the weighting vectors 320 applied to the delta HRTF group 322 for each predefined location.

[0050] The object adjuster 308 receives the binaural objects 266 and the weighted delta HRTFs 324 and applies the weighted delta HRTFs to each object to generate adjusted binaural objects 272. The adjusted binaural objects 272 thus correspond to the binaural objects 266 rotated according to the head orientation information 270. Equation (3) describes the operation of the object adjuster 308. JPEG2025529877000004.jpg19141

[0051] In equation (3), X obj (ω) is the frequency domain representation of a given object obj, dHRTF(ω) is the weighted delta HRTF set 324, and Y obj_HA (ω) is the frequency domain representation of the object after HRTF adjustment according to head orientation information 270. In other words, the adjusted binaural object 272 is proportional to the weighted delta HRTFs 324.

[0052] In summary, the HRTF adjuster 220 calculates the DOA of each object generated by the object extractor 210 (using the direction estimator 302), and then uses the object adjuster 308 to weight the HRTFs of each object between two of the predefined positions.

[0053] A playback device (e.g., a listener's mobile phone) may implement all components of the HRTF adjuster 220. Alternatively, a capture device (e.g., a UGC content creator's mobile phone) may implement the direction estimator 302, and the playback device may implement the other components. In such an embodiment, the capture device may provide the weighting vector 320 to the playback device, for example, as metadata, along with the binaural object 266.

[0054] 4 is a block diagram illustrating additional details of rebalancer 230 (see FIG. 2). Rebalancer 230 includes a direction estimator 402, a rebalancing coefficient calculator 404, a level adjuster 406, and a timbre adjuster 408. In general, rebalancer 230 adjusts for elevational (upward and downward) changes in the listener's head orientation, for example, associated with elevational objects (planes, birds singing, etc.).

[0055] The direction estimator 402 receives the adjusted binaural object 272, estimates the direction of arrival (DOA) of the sound represented by the object, and generates a weighting vector 420. The direction estimator 402 may be the same component as the direction estimator 302 (see FIG. 3), in which case the weighting vector 420 corresponds to the weighting vector 320 calculated based on the binaural object 266. Alternatively, the direction estimator 402 may be a different component from the direction estimator 302, in which case the weighting vector 420 is calculated based on the adjusted binaural object 272. In either case, the direction estimator 402 may perform direction estimation using a technique similar to that of the direction estimator 302, such as a DSP-based technique or an ML-based technique.

[0056] Rebalancing coefficient calculator 404 receives weighting vector 420 and head orientation information 270 and generates steering coefficients 422. Equation (4) describes the operation of rebalancing coefficient calculator 404. JPEG2025529877000005.jpg17143

[0057] In equation (4), g d is the direction coefficient of a given object calculated by the direction estimator 402 and corresponds to the weighting vector 420. Δφ is the elevation change due to the listener's head movement and corresponds to the head orientation information 270. σ(·) is an activation function. The rebalancer 230 may use the activation function to rebalance (e.g., apply a gain) the elevation direction object when the listener moves their head upward. g Rebalance corresponds to the steering coefficient 422. In other words, the steering coefficient 422 is proportional to the weighting vector 420 and the weights of the activation function (related to the head orientation information 270).

[0058] The level adjuster 406 receives the adjusted binaural object 272 and the steering coefficients 422 and adjusts the level of the adjusted binaural object 272 according to the steering coefficients 422 to generate a level-adjusted binaural object 424. The level adjuster 406 may perform the level adjustment using dynamic range control (DRC) applied to the object. For example, if the height object (before adjustment) is already large, there is no need to apply much additional gain when the listener moves their head upward. In such a case, the steering coefficients 422 control the aggressiveness of the DRC. The level adjuster 406 may adjust the amount of DRC based on psychoacoustic principles.

[0059] The timbre adjuster 408 receives the level-adjusted binaural object 424 and the steering coefficients 422 and adjusts the timbre of the level-adjusted binaural object 424 according to the steering coefficients 422 to generate the rebalanced binaural object 274. The timbre adjuster 408 may perform timbre adjustment using equalization applied to specific bands. For example, when a listener moves their head upward, the timbre of the sound changes, and timbre adjustment may be performed to boost specific bands. That is, if a listener perceives a height object and moves their head upward, the timbre adjustment may make the listener perceive the height object as looking directly at it, rather than perceiving it as directly above the listener. In such a case, the steering coefficients 422 control the aggressiveness of the EQ. The bands adjusted by the timbre adjuster 408 may be selected based on psychoacoustic principles.

[0060] A playback device (e.g., a listener's mobile phone) may implement all of the components of the rebalancer 230. Alternatively, a capture device (e.g., a UGC content creator's mobile phone) may implement the direction estimator 402, and the playback device may implement the other components. In such an embodiment, the capture device may provide the weighting vector 420 to the playback device, for example, as metadata, along with the binaural object 266.

[0061] 5 is a block diagram illustrating additional details of mixer 240 (see FIG. 2). Mixer 240 includes a decorrelator 502, a mixing ratio calculator 504, and a mixer 506.

[0062] The decorrelator 502 receives the residual signal 268 and performs decorrelation on the residual signal 268 to generate a decorrelated residual signal 520 resulting from the decorrelation. The residual signal 268 has N+2 channels, and the decorrelated residual signal 520 has M channels, where M≧N+2. If M=N+2, the decorrelation operation may be skipped. However, increasing M generally provides a more distinctive perception to the listener, at the expense of increased processing time. The decorrelator can be implemented using delay lines. Another example implementation of the decorrelator 502 is given in Kendall, Gary S., “The decorrelation of audio signals and its impact on spatial imagery,” Computer Music Journal 19, no. 4 (1995): 71-87.

[0063] The mixing ratio calculator 504 receives the head orientation information 270 and generates a mixing matrix 522 based on the head orientation information 270. The mixing matrix 522 has a size M×2 and corresponds to the azimuth angle change Δθ and elevation angle change Δφ due to head rotation indicated by the head orientation information 270.

[0064] The mixer 506 receives the decorrelated residual signal 520 and the mixing matrix 522 and performs mixing to generate the residual signal 276. The mixer 506 may perform mixing as described by equation (5). JPEG2025529877000006.jpg17145

[0065] In equation (5), Y res (ω) is the frequency domain representation of the residual signal 276, and W res (Δθ, Δφ) is the mixing matrix 522, and X res(ω) is the frequency domain representation of the decorrelated residual signal 520. In other words, the residual signal 276 is proportional to the decorrelated residual signal 520 (which is based on the residual signal 268) and proportional to the mixing matrix 522 (which is based on the azimuth and elevation changes due to head rotation).

[0066] Device Architecture Example

[0067] 6 illustrates a device architecture 600 for implementing the features and processes described herein, according to an embodiment. The architecture 600 can be implemented in any electronic device, including, but not limited to, a desktop computer, consumer audio / visual (AV) equipment, wireless broadcasting equipment, mobile devices such as smartphones, tablet computers, laptop computers, wearable devices, etc. In the illustrated example embodiment, the architecture 600 is for a mobile phone. The architecture 600 includes a processor(s) 601, a peripherals interface 602, an audio subsystem 603, one or more speakers 604, one or more microphones 605, sensors 606 (e.g., accelerometers, gyros, barometers, magnetometers, cameras, etc.), a position processor 607 (e.g., a GNSS receiver, etc.), a wireless communication subsystem 608 (e.g., Wi-Fi, Bluetooth, cellular, etc.), and an I / O subsystem(s) 609 (including a touch controller 610 and other input controllers 611, a touch surface 612, and other input / control devices 613). Other architectures having more or fewer components may be used to implement the disclosed embodiments.

[0068] Memory interface 614 is coupled to processor 601, peripherals interface 602, and memory 615 (e.g., Flash, RAM, ROM, etc.). Memory 615 stores computer program instructions and data, including, but not limited to: operating system instructions 616, communications instructions 617, GUI instructions 618, sensor processing instructions 619, telephony instructions 620, electronic messaging instructions 621, web browsing instructions 622, voice processing instructions 623, GNSS / navigation instructions 624, and applications / data 625. Voice processing instructions 623 include instructions for performing the voice processing described herein.

[0069] According to an embodiment, architecture 600 may correspond to one or more playback devices, such as a mobile phone and earphones. In such an embodiment, device architecture 600 corresponds to a mobile phone, and audio subsystem 603 communicates wirelessly with speakers 604 implemented in the earphones. Sensors 606 generate head orientation information, for example, by tracking movement of the earphones. The earphones themselves include similar components to architecture 600 and can output binaural signal 278. Processor(s) 601 implement various functions of system 200, such as HRTF adjuster 220, rebalancer 230, mixer 240, and remixer 250.

[0070] Similarly, architecture 600 may correspond to one or more capture devices, such as a mobile phone and earphones. In such an embodiment, device architecture 600 corresponds to a mobile phone, and audio subsystem 603 communicates wirelessly with a microphone 605 implemented in the earphones. The earphones themselves may include components similar to architecture 600 and capture binaural signal 264. Processor(s) 601 implement various functions of system 200, such as object extractor 210.

[0071] Similarly, architecture 600 may correspond to a computer system implementing a cloud service. In such an embodiment, device architecture 600 corresponds to a computer system implementing object extractor 210. The computer system receives audio signals 262 and 264 from capture devices and transmits binaural object signal 266 and residual signal 268 to a playback device.

[0072] Figure 7 is a flowchart of a method 700 of audio processing. Method 700 may be performed by one or more devices (e.g., laptop computers, mobile phones, server computers, etc.) comprising components of architecture 600 of Figure 6, for example, by executing one or more computer programs to implement the functionality of system 200 (see Figure 2).

[0073] At 702, one or more playback devices receive user-generated content (UGC) captured by a group of interconnected capture devices. Each audio source of the UGC corresponds to a different characteristic in the audio scene. For example, a mobile phone 104 (see FIG. 1 ) and an earphone 106 may be wirelessly connected, with the mobile phone 104 capturing UGC video and UGC audio, and the earphone 106 capturing UGC binaural audio. The playback device (e.g., a listener's mobile phone or earphone) can receive the captured UGC.

[0074] At 704, the one or more playback devices receive information indicative of listener behavior of a user of the one or more playback devices from one or more sensors of the one or more playback devices. For example, the playback devices may be implemented by architecture 600 (see FIG. 6 ) in which sensor 606 includes a gyroscope that generates head orientation information corresponding to movements of the listener's head.

[0075] At 706, the UGC is adapted according to listener behavior, including compensating for characteristics of the audio source according to the listener behavior. For example, the playback device may implement system 200 (see FIG. 2) for adjusting captured UGC audio (audio signal 262 and binaural audio signal 264) according to head orientation information 270.

[0076] At 708, the adapted UGC (see 706) is rendered to provide an interactive experience for the listener with respect to the audio scene. For example, the playback device may implement remixer 250 (see FIG. 2) to render the results of adapting the captured UGC audio and generate modified binaural signal 278.

[0077] Method 700 may include additional steps corresponding to other functionality of the audio processing system described herein. One such functionality is object modification, for example, using HRTF adjuster 220 or rebalancer 230 (see FIGS. 2-4). HRTF adjustment may include, for example, extracting one or more objects from a given audio portion of UGC, as described herein with respect to object extractor 210 (see FIG. 2). HRTF adjustment may include calculating HRTF differences before and after a head rotation for a group of predefined positions, as described herein with respect to delta HRTF generator 304 (see FIG. 3). HRTF adjustment may include obtaining HRTF differences for a particular object by applying different weights to the HRTF differences for the group of predefined positions according to each orientation of the one or more objects, as described herein with respect to direction estimator 302 and delta HRTF calculator 306 (see FIG. 3). HRTF adjustment may include repositioning the particular object to a new position after head rotation, including applying the obtained HRTF difference to the particular object, for example, as described herein for object adjuster 308 (see FIG. 3).

[0078] Rebalancing may include extracting one or more objects from a given audio portion of the UGC, e.g., as described herein with respect to object extractor 210 (see FIG. 2). Rebalancing may include determining the orientation of each of the one or more objects, e.g., as described herein with respect to direction estimator 402 and rebalancing coefficient calculator 404 (see FIG. 4). Rebalancing may include rebalancing the objects according to head orientation information, e.g., by applying level and / or timbre adjustments, as described herein with respect to level adjuster 406 and timbre adjuster 408 (see FIG. 4).

[0079] Another such functionality is residual mixing, e.g., using mixer 240 (see FIGS. 2 and 5). Residual mixing may include obtaining a residual by removing an object from a given audio portion of the UGC, e.g., as described herein with respect to object extractor 210 for generating residual signal 268 (see FIG. 2). Residual mixing may include creating one or more additional channels by decorrelation, e.g., as described herein with respect to decorrelator 502 (see FIG. 5). Residual mixing may include mixing the residuals from different audio channels of different capture devices by applying a mixing ratio for each channel according to head orientation information, e.g., as described herein with respect to mixing ratio calculator 504 and mixer 506 (see FIG. 5).

[0080] Implementation details

[0081] The embodiments may be implemented in hardware, in executable modules stored on a computer-readable medium, or in a combination of both, such as a programmable logic array. Unless otherwise specified, steps performed by the embodiments need not inherently relate to any particular computer or other apparatus, although this may be the case in particular embodiments. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus, such as integrated circuits, to perform the required method steps. Thus, the embodiments may be implemented in one or more computer programs running on one or more programmable computer systems, each including at least one processor, at least one data storage system including volatile and nonvolatile memory and / or storage elements, at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.

[0082] Each such computer program is preferably stored on or downloaded to a general-purpose or dedicated programmable computer-readable storage medium or device (e.g., solid-state memory or media, magnetic media, optical media, etc.), and when the storage medium or device is read by a computer system, configures and operates the computer to perform the procedures described herein. The system of the present invention can also be considered to be embodied as a computer-readable storage medium configured with a computer program, the storage medium thus configured causing the computer system to operate in a specific, predefined manner, thereby performing the functions described herein. Software per se and intangible or ephemeral signals are excluded insofar as they are non-patentable subject matter.

[0083] Aspects of the systems described herein can be implemented in any suitable computer-based sound processing network environment for processing digital or digitized audio files. Portions of an adaptive audio system can include one or more networks made up of any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route data transmitted between computers. Such networks can be built based on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0084] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. Also, it should be noted that the various functions disclosed herein, in terms of their behavior, register transfers, logical components, and / or other characteristics, may be described using any number of combinations of hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media. The computer-readable media on which the data and / or instructions so formatted may be embodied include, but are not limited to, various forms of physical, non-transitory, non-volatile storage media, such as optical, magnetic, or semiconductor storage media.

[0085] The above description illustrates various embodiments of the present disclosure, along with examples of how aspects of the disclosure may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of the present disclosure, as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be adopted without departing from the spirit and scope of the present disclosure, as defined by the claims.

Claims

1. receiving, by one or more playback devices, user-generated content (UGC) captured by a group of interconnected capture devices, wherein each audio source in the UGC corresponds to a respective characteristic in an audio scene; receiving, by the one or more playback devices, information from one or more sensors of the one or more playback devices indicative of listener behavior of users of the one or more playback devices; Adapting the UGC according to the listener behavior, including compensating for characteristics of an audio source according to the listener behavior; providing a listener with an interactive experience regarding the audio scene by rendering the adapted UGC; and A computer-implemented method of speech processing, comprising:

2. The method of claim 1 , wherein the UGC includes a video stream and an immersive audio stream.

3. The method of claim 1 , wherein the one or more playback devices include a device for video playback and a connected device for audio playback.

4. 4. The method of claim 1, wherein the group of capture devices includes a first device that captures video and at least one channel of an audio stream, and a connected device that captures a binaural audio stream.

5. The method of claim 1 , wherein the listener behavior includes at least a head orientation relative to a screen of the one or more playback devices configured for video playback.

6. The method of claim 1 , wherein adapting the UGC comprises at least one of object modification and residual blending.

7. The method of claim 6 , wherein the object modification comprises at least one of head-related transfer function (HRTF) adjustment and object rebalancing.

8. The HRTF adjustment is extracting one or more objects from a given audio portion of the UGC; calculating HRTF differences before and after head rotation for a group of predefined positions; obtaining HRTF differences for a particular object by applying different weights to the HRTF differences for the group of predefined positions according to each orientation of the one or more objects; repositioning the particular object to a new position after the head rotation, including applying the obtained HRTF difference to the particular object; The method of claim 7 , comprising the action of:

9. The object rebalancing includes: extracting one or more objects from a given audio portion of the UGC; determining an orientation of each of the one or more objects; rebalancing the one or more objects according to head orientation information by applying level and / or timbre adjustments; The method of claim 7 , comprising the action of:

10. The residual blending is obtaining a residual by removing an object from a given audio portion of the UGC; creating one or more additional channels by decorrelation; blending the residuals from different audio channels of different capture devices by applying a blending ratio for each channel according to head orientation information; The method of claim 6 , comprising the action of:

11. Adapting the UGC includes: receiving one or more objects and a residual signal, the one or more objects and the residual signal being generated based on a given audio portion of the UGC, the given audio portion of the UGC including a first audio signal having at least one channel and a second audio signal being a binaural audio signal; performing object modification on the one or more objects based on head orientation information, wherein performing the object modification includes generating one or more modified objects; performing residual mixing on the residual signal based on the head orientation information, wherein performing residual mixing includes generating a mixed residual signal; remixing the one or more modified objects and the mixed residual signal, including generating a modified binaural signal based on the one or more modified objects and the mixed residual signal; 6. The method of claim 1, comprising:

12. The method of claim 11 , further comprising extracting the one or more objects and the residual signal from the given audio portion of the UGC.

13. performing the object modification, performing direction estimation to calculate a direction of arrival for each of the one or more objects; performing an HRTF adjustment for each of the one or more objects to adjust a given object for at least one of an azimuth angle change and an elevation angle change in the head orientation information based on a corresponding direction of arrival; performing object rebalancing for each of the one or more objects to adjust the given object for the elevation directional change in the head orientation information based on the corresponding direction of arrival; The method of claim 11 , comprising:

14. 14. The method of claim 13, wherein the HRTF adjustment is proportional to a function of at least one of an azimuth angle of the predefined position after the azimuth angle change in the head orientation information and an elevation angle of the predefined position after the elevation angle direction change in the head orientation information, and inversely proportional to a function of at least one of the azimuth angle and elevation angle of the predefined position.

15. 14. The method of claim 13, wherein the object rebalancing is proportional to a weighting vector and proportional to an activation function, the weighting vector being based on the orientation of the given object and the activation function being based on the elevation directional change in the head orientation information.

16. The performing of the residual mixing includes: decorrelating the residual signal, wherein the decorrelating comprises generating a decorrelated residual signal; and generating a mixing matrix based on the azimuth angle change and the elevation angle change in the head direction information; mixing the decorrelated residual signal with the mixing matrix in the frequency domain, including generating the mixed residual signal; The method of claim 11 , comprising:

17. The method of claim 16 , wherein the mixed residual signal is proportional to the mixing matrix and proportional to the decorrelated residual signal.

18. A non-transitory computer readable medium having stored thereon a computer program which, when executed by a processor, controls an apparatus to perform processes including the method of any one of claims 1 to 17.

19. 1. An apparatus for audio processing, comprising: A processor configured to control the device to perform processing including the method of any one of claims 1 to 17. An apparatus comprising:

20. the one or more playback devices a mobile phone having the processor; A pair of binaural earphones, 20. The apparatus of claim 19, comprising:

Citation Information

Patent Citations

  • Sound processing apparatus, sound source position control method and sound source position control program

    JP2015233252A

  • Video analysis-assisted generation of multichannel audio data

    JP2016513410A

  • Distance panning with near / far rendering

    JP2019523913A

  • Apparatus and method for rendering an audio signal for playback to a user

    JP2021522720A

  • Method and device for processing a binaural recording

    WO2022060891A1