Perceptual enhancement for binaural audio recording

The method addresses binaural audio capture challenges by using machine learning to reduce noise and artifacts, ensuring an optimal playback experience for user-generated content.

JP7834760B2Active Publication Date: 2026-03-24DOLBY LABORATORIES LICENSING CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing audiovisual capture systems face challenges in capturing binaural audio due to the use of mono or stereo microphones, and user-generated content often contains noise not present in professionally produced content, leading to discrepancies between audio and video streams.

Method used

A computer-implemented method using machine learning systems to calculate noise reduction gains for binaural audio capture, applying shared gains to reduce noise and artifacts, and simultaneously capturing video with binaural audio to enhance perceptual experience.

Benefits of technology

The method effectively reduces noise in binaural audio, maintaining spatial cues and providing an optimal playback experience by minimizing artifacts and noise discrepancies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007834760000010
    Figure 0007834760000010
  • Figure 0007834760000011
    Figure 0007834760000011
  • Figure 0007834760000012
    Figure 0007834760000012
Patent Text Reader

Abstract

A method of audio processing includes capturing a binaural audio signal, calculating a noise reduction gain using a machine learning model, and generating a modified binaural audio signal. The method may further include performing various corrections to the audio taking into account video captured by different cameras, such as a front camera and a rear camera. The method may further include performing a smooth transition of the binaural audio when switching between the front camera and the rear camera. The method may reduce noise in the binaural audio and improve a user's perception of the combined video and binaural audio.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-reference of related applications) This application claims priority to U.S. Provisional Patent Application No. 63 / 139,329 filed on 20 January 2020, U.S. Provisional Patent Application No. 63 / 287,730 filed on 9 December 2021, and PCT Application No. PCT / CN2020 / 138221 filed on 22 December 2020, all of which are incorporated herein by reference in their entirety.

[0002] This disclosure relates to audio processing, particularly noise suppression. [Background technology]

[0003] Unless otherwise stated herein, the approaches described in this section are not prior art to the claims of this application, nor will they be deemed prior art by being included in this section.

[0004] Devices for audiovisual capture are gaining popularity among consumers. Such devices include portable cameras such as Sony Action Cam® and GoPro® cameras, as well as mobile phones with integrated camera functionality. Generally, these devices capture audio simultaneously with video, for example, using a mono or stereo microphone. Audiovisual content sharing systems such as YouTube® and Twitch.tv® are also gaining popularity. Users then broadcast the captured audiovisual content simultaneously with the capture, or upload the captured audiovisual content to the content sharing system. Because this content is user-generated, it is called user-generated content (UGC), as opposed to professionally produced content (PGC), which is typically produced by professionals. UGC differs from PGC in that it is often created using consumer-grade equipment that is less expensive and has fewer features than professional equipment. Another difference between UGC and PGC is that UGC is often captured in uncontrolled environments, such as outdoors, while PGC is often captured in controlled environments, such as recording studios.

[0005] Binaural audio includes audio recorded using two microphones positioned at the user's ears. Captured binaural audio provides an immersive listening experience when played back through headphones. Compared to stereo audio, binaural audio also includes the shadows of the user's head and ears, resulting in a time difference and a level difference between the two ears when captured. [Overview of the Initiative]

[0006] Existing audiovisual capture systems have several problems. One problem is that many existing capture devices only include mono or stereo microphones, making binaural audio capture particularly difficult. Another problem is that because PGC is often captured in a controlled environment, UGC audio often contains stationary and transient noise that is not present in PGC audio. Yet another problem is that separate audio and video capture devices can result in audio and video streams that contradict human perception using sight and hearing.

[0007] The embodiment relates to capturing video simultaneously with binaural audio and performing perceptual enhancements such as noise reduction on the captured binaural audio. The resulting binaural audio, when subsequently consumed in combination with the captured video, is perceived in a different way than stereo or mono audio.

[0008] According to one embodiment, a computer-implemented method of audio processing includes capturing an audio signal having at least two channels, including a left channel and a right channel, using an audio capture device. The method further includes calculating a plurality of noise reduction gains for each of the at least two channels using a machine learning system. The method further includes calculating a plurality of shared noise reduction gains based on the plurality of noise reduction gains for each channel. The method further includes generating a modified audio signal by applying the plurality of shared noise reduction gains to each of the at least two channels.

[0009] As a result, noise can be reduced in the captured binaural audio.

[0010] Machine learning systems may use a monaural model, a binaural model, or both.

[0011] This method may further include capturing a video signal simultaneously with capturing an audio signal using a video capture device. This method may further include switching between a front camera and a rear camera, which includes smoothing left / right corrections to the audio signal using a first smoothing parameter and smoothing front / back corrections to the audio signal using a second smoothing parameter. Capturing a video signal simultaneously with capturing an audio signal may include performing corrections on the audio signal, which include at least one of left / right corrections, front / back corrections, and stereo image width control corrections. Stereo image width control corrections may include generating a center channel and side channels from the left and right channels of the audio signal, attenuating the side channels by a width adjustment coefficient, and generating a modified audio signal from the center channel and the attenuated side channels.

[0012] According to another embodiment, the device includes a processor. The processor is configured to control the device to implement one or more methods described herein. The device may further include details similar to one or more methods described herein.

[0013] According to another embodiment, a non-temporary computer-readable medium stores a computer program that, when executed by a processor, controls the device to perform processing including one or more methods described herein.

[0014] The following detailed explanation and attached diagrams will provide a further understanding of the properties and advantages of various implementations. [Brief explanation of the drawing]

[0015] [Figure 1]It is a standardized overhead view of the audio-visual capture system 100.

[0016] [Figure 2] It is a block diagram of the audio processing system 200.

[0017] [Figure 3] It is a block diagram of the audio processing system 300.

[0018] [Figure 4] It is a block diagram of the audio processing system 400.

[0019] [Figure 5] It is a block diagram of the audio processing system 500.

[0020] [Figure 6] It is a standardized overhead view showing binaural audio capture in the selfie mode using the video capture system 100 (see Figure 1).

[0021] [Figure 7] It is a graph showing an example of the magnitude response of a high shelf filter implemented using a bi-quad filter.

[0022] [Figure 8] It is a standardized overhead view showing various audio capture angles in the selfie mode.

[0023] [Figure 9] It is a graph of the attenuation coefficient α for different focal lengths f.

[0024] [Figure 10]This is a stylized overhead view showing binaural audio capture in normal mode using the video capture system 100 (see Figure 1).

[0025] [Figure 11] This is a device architecture 1100 for implementing the features and processing described herein, according to one embodiment.

[0026] [Figure 12] This is a flowchart of audio processing method 1200.

[0027] [Figure 13] This is a flowchart of audio processing method 1300. [Modes for carrying out the invention]

[0028] This section describes techniques relating to audio processing. The following description includes numerous examples and specific details to provide a complete understanding of the disclosure. However, it will be apparent to those skilled in the art that the disclosure as defined by the claims may include some or all of the features of these examples, individually or in combination with other features described below, and may also include modifications and equivalents of the features and concepts described herein.

[0029] The following description details various methods, processes, and procedures. Certain steps may be described in a specific order, primarily for convenience and clarity. A particular step may be repeated multiple times, may occur before or after other steps, or may occur concurrently with other steps, even if those steps are described in a different order. A second step must follow the first step only if the first step must be completed before the second step begins. Such situations will be specifically noted if they are not evident from the context.

[0030] This document uses the terms "and," "or," and "and / or." These terms are to be interpreted as having an inclusive meaning. For example, "A and B" may mean at least: "both A and B," or "at least both A and B." Another example is "A or B," which may mean at least: "at least A," "at least B," "both A and B," or "at least both A and B." Another example is "A and / or B," which may mean at least: "A and B," or "A or B." If an exclusive-or is intended, it will be specifically stated, for example, "either A or B," or "at most one of A or B."

[0031] This document describes various processing functions related to structures such as blocks, elements, components, and circuits. Generally, these structures can be implemented by a processor controlled by one or more computer programs.

[0032] Figure 1 is a stylized overhead view of the audiovisual capture system 100. Users generally use the audiovisual capture system 100 to capture audio and video in uncontrolled environments, for example, to capture user-generated content (UGC). The audiovisual capture system 100 includes a video capture device 102, a left earphone 104, and a right earphone 106.

[0033] The video capture device 102 generally includes a camera that captures video data. The video capture device 102 may include two cameras, called a front camera and a rear camera. The front camera, also called a selfie camera, is generally located on one side of the video capture device 102, for example, the side that includes a display screen or touchscreen. The rear camera is generally located on the opposite side from the front camera. The video capture device 102 may be a mobile phone, and as such, may have several additional components and functions such as a processor, volatile and non-volatile memory and storage, a radio, a microphone, and a speaker. For example, the video capture device 102 may be a mobile phone such as an Apple iPhone® mobile phone or a Samsung Galaxy® mobile phone. The video capture device 102 may generally be held by the user, mounted on the user's selfie stick or tripod, mounted on the user's shoulder mount, or attached to an aerial drone.

[0034] The left earphone 104 is positioned in the user's left ear, includes a microphone, and generally captures the left binaural signal. The left earphone 104 provides the left binaural signal to the video capture device 102 in order to capture audio data simultaneously with video data. The left earphone 104 may connect wirelessly to the video capture device 102 via an IEEE 802.15.1 standard protocol, such as the Bluetooth® protocol. Alternatively, the left earphone 104 may connect to another device (not shown) that receives both the audio data and the captured video data from the video capture device 102.

[0035] The right earphone 106 is positioned in the user's right ear, includes a microphone, and generally captures the right binaural signal. The right earphone 104 provides the right binaural signal to the video capture device 102 in a similar manner to that described above with respect to the left earphone 104. The right earphone 106 may otherwise be similar to the left earphone 104.

[0036] One use case for the audiovisual capture system 100 is for a user to walk down the street and capture binaural audio using earphones 104 and 106 while simultaneously capturing video using video capture device 102. The audiovisual capture system 100 then broadcasts the captured content or saves it for later editing or uploading. Another use case is recording podcasts, interviews, news reports, and speeches during meetings or events. In such situations, binaural recording can provide a desirable sense of spaciousness; however, the presence of ambient noise and the distance of other sources of interest from the person wearing earphones 104 and 106 often result in an overwhelming noise presence, leading to an unoptimal playback experience. While maintaining spatial cues in the recording, adequately reducing excessive noise is challenging but highly worthwhile.

[0037] The following sections describe in detail additional audio processing techniques implemented by the audiovisual capture system 100, for example, to perform noise reduction on captured binaural audio.

[0038] 1. Noise reduction of captured binaural audio

[0039] Figure 2 is a block diagram of the audio processing system 200. The audio processing system 200 can be implemented as a component of the audiovisual capture system 100 (see Figure 1), for example, as one or more computer programs executed by the processor of the video capture device 102. The audio processing system 200 includes a conversion system 202, a noise reduction system 204, a mixing system 206, and an inverse conversion system 208.

[0040] The conversion system 202 receives the left input signal 220 and the right input signal 222, performs signal conversion, and generates the converted left signal 224 and the converted right signal 226. The left input signal 220 generally corresponds to the signal captured by the left earphone 104, and the right input signal 222 generally corresponds to the signal captured by the right earphone 106. In other words, input signals 220 and 222 correspond to binaural signals, with the left input signal 220 corresponding to the left binaural signal and the right input signal 222 corresponding to the right binaural signal. The converted left input signal 224 corresponds to the converted left input signal 220, and the converted right input signal 226 corresponds to the converted right input signal 222.

[0041] Signal transformation generally converts an input signal from a first signal domain to a second signal domain. The first signal domain may be the time domain. The second signal domain may be the frequency domain. The signal transformation may be one or more of the following: Fourier transforms such as Fast Fourier Transform (FFT), Short-Time Fourier Transform (STFT), Discrete-Time Fourier Transform (DTFT), Discrete Fourier Transform (DFT), Discrete Sine Transform (DST), Discrete Cosine Transform (DCT); Quadrature Mirror Filter (QMF) transform; Complex Quadrature Mirror Filter (CQMF) transform; Hybrid Complex Quadrature Mirror Filter (HCQMF) transform; etc. The transformation system 202 may perform framing of the input signal before performing the transformation, and the transformation is performed in frames. The frame size may be between 5 milliseconds and 15 milliseconds, for example, 10 milliseconds. The transformation system 202 may output transformed signals 224 and 226 grouped into bands of the transformation domain. The number of bands may be between 15 and 25, for example, 20 bands.

[0042] The noise reduction system 204 receives the transformed left signal 224 and the transformed right signal 226, performs gain calculations, and generates left gain 230 and right gain 232. Generally, the noise reduction system 204 implements one or more machine learning systems to calculate the noise reduction gains 230 and 232. Specifically, the left gain 230 corresponds to the noise reduction gain applied to the transformed left signal 224, and the right gain 232 corresponds to the noise reduction gain applied to the transformed right signal 226. The noise reduction gains may be a shared noise reduction gain applied to both the left and right signals, or a single set of gains applied to both signals. Details of the machine learning systems and noise reduction gains are provided below, with particular reference to Figure 3-5.

[0043] The mixing system 206 receives the converted left signal 224, the converted right signal 226, the left gain 230, and the right gain 232, mixes them, and generates the mixed left signal 234 and the mixed right signal 236. Generally, the mixing system 206 mixes the converted left signal 224 and the left gain 230 to generate the mixed left signal 234, and mixes the converted right signal 226 and the right gain 232 to generate the mixed right signal 236. Further details of the mixing are provided below, with particular reference to Figure 3-5.

[0044] The inverse transform system 208 receives the mixed left signal 234 and the mixed right signal 236, performs an inverse transform, and generates a modified left signal 240 and a modified right signal 242. The inverse transform generally corresponds to the reverse of the signal transform performed by the transform system 202, transforming the signal back from the second signal domain to the first signal domain. For example, the inverse transform system 208 may transform the mixed signals 234 and 236 from the QMF domain to the time domain. As a result, the modified left signal 240 corresponds to a noise-reduced version of the left input signal 220, and the modified right signal 242 corresponds to a noise-reduced version of the right input signal 222.

[0045] Next, the audiovisual capture system 100 may output a modified left signal 240 and a modified right signal 242 along with the captured video signal as part of UGC generation. Additional details of the audio processing system 200 are provided below, with particular reference to Figure 3-5.

[0046] Figure 3 is a block diagram of the audio processing system 300. The audio processing system 300 is a more specific embodiment of the audio processing system 200 (see Figure 2). The audio processing system 300 may be implemented as a component of the audiovisual capture system 100 (see Figure 1), for example, as one or more computer programs executed by the processor of the video capture device 102. The audio processing system 300 includes conversion systems 302a and 302b, noise reduction systems 304a and 304b, gain calculation system 306, mixing systems 308a and 308b, and inverse conversion systems 310a and 310b.

[0047] The conversion systems 302a and 302b receive the left input signal 320 and the right input signal 322, perform signal conversion, and generate the converted left signal 324 and the converted right signal 326. In particular, conversion system 302a generates the converted left signal 324 based on the left input signal 320, and conversion system 302b generates the converted right signal 326 based on the right input signal 322. The input signals 320 and 322 correspond to the binaural signals captured by earphones 104 and 106 (see Figure 1). The signal conversion performed by conversion systems 302a and 302b generally corresponds to the signal conversion described above with respect to conversion system 202 (see Figure 2).

[0048] Noise reduction systems 304a and 304b receive the transformed left signal 324 and the transformed right signal 326, perform gain calculations, and generate the left gain 330 and the right gain 332. Specifically, noise reduction system 304a generates the left gain 330 based on the transformed left signal 324, and noise reduction system 304b generates the right gain 332 based on the transformed right signal 326. Noise reduction system 304a receives the transformed left signal 324, performs feature extraction on the transformed left signal 324 to extract a set of features, processes the set of features by inputting them into a trained model, and generates the left gain 330 as a result of processing the set of features. Processing features by inputting them into a trained model is sometimes called "classification". The noise reduction system 304b receives the transformed right signal 326, performs feature extraction on the transformed right signal 326 to extract a set of features, processes the set of features by inputting it into a trained model, and generates a right gain 332 as a result of processing the set of features.

[0049] A feature may include one or more of the following: temporal features, spectral features, time-frequency features, etc. Temporal features may include one or more of the following: automatic correction coefficient (ACC), linear predictive coding coefficient (LPCC), zero crossing coefficient (ZCR), etc. Spectral features may include one or more of the following: spectral centroid, spectral roll-off, spectral energy distribution, spectral flatness, spectral entropy, Mel-frequency cepstrum coefficient (MFCC), etc. Time-frequency features may include one or more of the following: spectral bundle, chroma, etc. A feature may also include statistical information on the other features mentioned above. These statistics may include the mean, standard deviation, and higher-order statistics, such as skewness and kurtosis. For example, a feature may include the mean and standard deviation of the spectral energy distribution.

[0050] The trained model can be implemented as part of a machine learning system. A machine learning system may include one or more neural networks, such as recurrent neural networks (RNNs) or convolutional neural networks (CNNs). The trained model receives extracted features as input, processes the extracted features, and outputs a gain as a result of processing the extracted features. It should be noted that both noise reduction systems 304a and 304b use the same trained model; for example, each noise reduction system implements a copy of the trained model. The trained model is trained offline using monaural training data, as will be further described below.

[0051] The gain calculation system 306 receives the left gain 330 and the right gain 332, combines gains 330 and 332 according to a mathematical function, and generates a shared gain 334. The mathematical function can be one or more, such as a maximum, average, range function, or comparison function. As an example, suppose the left gain 330, the right gain 332, and the shared gain 334 are each gain vectors, for example, a 20-band vector. For the maximum, the gain of band 1 of the shared gain 334 is the maximum of the gain of band 1 of the left gain 330 and the gain of band 1 of the right gain 332; the same applies to the other 19 bands. For the average, the gain of band 1 of the shared gain 334 is the average of the gain of band 1 of the left gain 330 and the gain of band 1 of the right gain 332; the same applies to the other 19 bands.

[0052] The range function applies a different function to each band based on the range of gains in each band, specifically gains 330 and 332. For example, if the gain in each band 1 of gains 330 and 332 is less than X1, the maximum is calculated. If the gain is between X1 and X2, the average is calculated. If the gain is greater than X2, the maximum is calculated.

[0053] The difference function applies a different function to each band based on a comparison of the gain difference between the gains in each band of gain 330 and 332. For example, if the gain difference in band 1 between gains 330 and 332 is less than X1, the average is calculated; if the gain difference is greater than or equal to X1, the maximum is calculated.

[0054] The audio processing system 300 uses a shared gain 334 instead of applying the left gain 330 to the converted left signal 324 and the right gain 332 to the converted right signal 326 in order to reduce any artifacts that may be present in quick-attack sounds. Binaurally captured quick-attack sounds may cross the frame boundaries of input signals 320 and 322 (as part of the operation of conversion systems 302a and 302b) due to the time difference between the two ears between the left and right microphones. In such cases, the gain of the quick-attack sound is processed at frame X of one channel and at frame X+1 of the other channel, which can result in artifacts. Calculating the maximum gain for a specific band of each channel using a shared gain results in a reduced perception of artifacts.

[0055] The noise reduction systems 304a and 304b, and the gain calculation system 306 may otherwise be similar to the noise reduction system 204 (see Figure 2).

[0056] Mixing systems 308a and 308b receive the converted left signal 324, the converted right signal 326, and the shared gain 334, apply the shared gain 334 to signals 324 and 326, and generate the mixed left signal 336 and the mixed right signal 338. Specifically, mixing system 308a generates the mixed left signal 336 by applying the shared gain 334 to the converted left signal 324, and mixing system 308b generates the mixed right signal 338 by applying the shared gain 334 to the converted right signal 326. For example, the transformed left signal 324 may have 20 bands, the shared gain 334 may be a gain vector with 20 bands, and the magnitude value of a given band of the mixed left signal 336 is obtained by multiplying the magnitude value of a given band of the transformed left signal 324 by the gain value of a given band of the shared gain 334. Mixing systems 308a and 308b may otherwise be similar to mixing system 206 (see Figure 2).

[0057] The inverse conversion systems 310a and 310b receive the mixed left signal 336 and the mixed right signal 338, perform inverse conversion, and generate the modified left signal 340 and the modified right signal 342. Specifically, the inverse conversion system 310a performs inverse conversion on the mixed left signal 336 to generate the modified left signal 340, and the inverse conversion system 310b performs inverse conversion on the mixed right signal 338 to generate the modified right signal 342. The inverse conversion performed by the inverse conversion systems 310a and 310b generally corresponds to the inverse conversion of the conversion performed by the conversion systems 302a and 302b, converting the signal back from the second signal domain to the first signal domain. The modified left signal 340 then corresponds to a noise-reduced version of the left input signal 320, and the modified right signal 342 corresponds to a noise-reduced version of the right input signal 322. The inverse conversion systems 310a and 310b may otherwise be similar to the inverse conversion system 208 (see Figure 2).

[0058] Monoaural Model Training

[0059] As described above, the noise reduction systems 304a and 304b use the trained model to generate left gain 330 and right gain 332 from the converted left signal 324 and the converted right signal 326. This trained model is trained offline using monaural training data. The offline training process may also be called the training phase and is contrasted with the operation phase, in which the trained model is used by the audio processing system 300 during normal operation. The training phase generally consists of four steps.

[0060] First, a training data set is generated. This training data set can be generated by mixing various monaural audio data source samples and various noise samples at different signal-to-noise ratios (SNRs). Monaural audio data source samples generally correspond to noise-free audio data, also known as clean audio data, including speech and music. Noise samples correspond to noisy audio data, including traffic noise, fan noise, airplane noise, construction noise, sirens, and baby crying. The training data can yield a corpus of approximately 100-200 hours, by mixing approximately 1-2 hours of source samples with 15-25 noise samples at an SNR of 5-10. Each source sample can be between 15 and 60 seconds long, and the SNR can range from -45 to 0 dB. For example, a given source sample of speech is 30 seconds long, and this source sample can be mixed with noise samples of traffic noise at 5 SNRs of -40, -30, -20, -10, and 0 dB, resulting in 600 seconds of training data in the training data corpus.

[0061] Secondly, features are extracted from the training data set. Generally, the feature extraction process is the same as that used during the operation of audio processing systems, such as 200 (see Figure 2) and 300 (see Figure 3), including the execution of transformations and the extraction of features in the second signal domain. The extracted features also correspond to the features used during the operation of the audio processing system.

[0062] Thirdly, the model is trained on a set of training data. Generally, training is done by adjusting the weights of the nodes in the model in relation to the model's output, which is compared to an ideal output. The ideal output corresponds to the gain required to adjust a noisy input to produce a noise-free output.

[0063] Finally, once the model is sufficiently trained, the resulting model is provided to an audio processing system, for example, Figure 200 or Figure 300, for use in the operation phase.

[0064] As mentioned above, the training data is monaural training data. This monaural training data yields a single model that the audio processing system 300 uses for each input channel. Specifically, noise reduction system 304a uses the trained model with the converted left signal 324 as input, and noise reduction system 304b uses the trained model with the converted right signal 326 as input; for example, systems 304a and 304b may each implement a copy of the trained model. The model can also be trained using binaural training data, as will be discussed later with respect to Figure 4-5.

[0065] Figure 4 is a block diagram of the audio processing system 400. The audio processing system 400 is a more specific embodiment of the audio processing system 200 (see Figure 2). The audio processing system 400 may be implemented as a component of the audiovisual capture system 100 (see Figure 1), for example, as one or more computer programs executed by the processor of the video capture device 102. The audio processing system 400 is similar to the audio processing system 300 (see Figure 3), but has differences related to the trained model, as detailed below. The audio processing system 400 includes conversion systems 402a and 402b, a noise reduction system 404, mixing systems 406a and 406b, and inverse conversion systems 408a and 408b.

[0066] The conversion systems 402a and 402b receive the left input signal 420 and the right input signal 422, perform signal conversion, and generate the converted left signal 424 and the converted right signal 426. The conversion systems 402a and 402b operate in the same manner as the conversion systems 302a and 302b (see Figure 3), and for brevity, that explanation will not be repeated.

[0067] The noise reduction system 404 receives the transformed left signal 424 and the transformed right signal 426, performs gain calculations, and generates a joint gain 430. The joint gain 430 is based on both the transformed left signal 424 and the transformed right signal 426. The noise reduction system 404 performs feature extraction on the transformed left signal 424 and the transformed right signal 426 to extract a joint set of features, processes the joint set of features by inputting it into a trained model, and generates the joint gain 430 as a result of processing the joint set of features. Therefore, the joint gain 430 corresponds to the shared gain and is sometimes called the shared gain 430. The noise reduction system 404 is otherwise similar to the noise reduction systems 304a and 304b (see Figure 3), and for brevity, its description will not be repeated. For example, the joint set of features may be the same features as those described above for the noise reduction systems 304a and 304b. The trained model is the same as the trained model described above with respect to noise reduction systems 304a and 304b, except that the trained model implemented by noise reduction system 404 is trained offline using binaural training data, as will be described later.

[0068] It should be noted that, unlike the audio processing system 300 (see Figure 3), the audio processing system 400 does not need a gain calculation system because it outputs a shared gain 430 as a result of the noise reduction system 404 being trained using binaural training data.

[0069] Mixing systems 406a and 406b receive the converted left signal 424, the converted right signal 426, and the shared gain 430, apply the shared gain 430 to signals 424 and 426, and generate the mixed left signal 434 and the mixed right signal 436. Specifically, mixing system 406a generates the mixed left signal 434 by applying the shared gain 430 to the converted left signal 424, and mixing system 406b generates the mixed right signal 436 by applying the shared gain 430 to the converted right signal 426. Mixing systems 406a and 406b are otherwise similar to mixing systems 308a and 308b (see Figure 3), and for brevity, their description will not be repeated.

[0070] The inverse conversion systems 408a and 408b receive the mixed left signal 434 and the mixed right signal 436, perform inverse conversion, and generate the modified left signal 440 and the modified right signal 442. In particular, the inverse conversion system 408a performs inverse conversion on the mixed left signal 434 to generate the modified left signal 440, and the inverse conversion system 408b performs inverse conversion on the mixed right signal 436 to generate the modified right signal 442. The inverse conversion performed by the inverse conversion systems 408a and 408b generally corresponds to the inverse conversion of the conversion performed by the conversion systems 402a and 402b, converting the signal back from the second signal domain to the first signal domain. As a result, the modified left signal 440 corresponds to a noise-reduced version of the left input signal 420, and the modified right signal 442 corresponds to a noise-reduced version of the right input signal 422. The inverse conversion systems 408a and 408b may otherwise be similar to the inverse conversion systems 310a and 310b (see Figure 3).

[0071] Binaural Model Training

[0072] As previously mentioned, the noise reduction system 404 uses the trained model to generate a shared gain 430 from the transformed left signal 424 and the transformed right signal 426. The trained model is trained offline using binaural training data. The use of binaural training data is in contrast to the use of monaural training data used when training the models of the noise reduction systems 304a and 304b (see Figure 3). Training the model using binaural training data is generally similar to training the model using monaural training data as described above with respect to Figure 3, and the training phase generally consists of four steps.

[0073] First, a set of training data is generated. The audio data source samples are binaural audio data source samples, instead of the monaural audio data source samples mentioned earlier in Figure 3. Mixing the binaural audio data source samples with noise samples at various SNRs yields a similar corpus of approximately 100-200 hours.

[0074] Secondly, features are extracted from the training data set. Features are extracted by combining binaural channels, for example, from the left and right channels. Extracting features by combining binaural channels is in contrast to extraction from a single channel, as is used when training models of noise reduction systems 304a and 304b (see Figure 3).

[0075] Thirdly, the model is trained on a set of training data. The training process is generally similar to the training process used when training the models for the noise reduction systems 304a and 304b (see Figure 3).

[0076] Finally, once the model is sufficiently trained, the resulting model is provided to an audio processing system, for example, 400 in Figure 4, for use in the operation phase.

[0077] Figure 5 is a block diagram of the audio processing system 500. The audio processing system 500 is a more specific embodiment of the audio processing system 200 (see Figure 2). The audio processing system 500 can be implemented as a component of the audiovisual capture system 100 (see Figure 1), for example, as one or more computer programs executed by the processor of the video capture device 102. The audio processing system 500 is similar to both the audio processing system 300 (see Figure 3) and the audio processing system 400 (see Figure 4), with differences related to the trained model, as detailed below. The audio processing system 500 includes conversion systems 502a and 502b, noise reduction systems 504a, 504b and 504c, gain calculation system 506, mixing systems 508a and 508b, and inverse conversion systems 510a and 510b.

[0078] The conversion systems 502a and 502b receive the left input signal 520 and the right input signal 522, perform signal conversion, and generate the converted left signal 524 and the converted right signal 526. The conversion systems 502a and 502b operate in the same manner as the conversion systems 302a and 302b (see Figure 3) or 402a and 402b (see Figure 4), and for brevity, that explanation will not be repeated.

[0079] Noise reduction systems 504a, 504b, and 504c receive the transformed left signal 524 and the transformed right signal 526, perform gain calculations, and generate a left gain 530, a right gain 532, and a combined gain 534. Specifically, noise reduction system 504a generates a left gain 530 based on the transformed left signal 524, noise reduction system 504b generates a right gain 532 based on the transformed right signal 326, and noise reduction system 504c generates a combined gain 534 based on both the transformed left signal 524 and the transformed right signal 526. Noise reduction system 504a receives the transformed left signal 524, performs feature extraction on the transformed left signal 524 to extract a set of features, processes the set of features by inputting it into a trained monaural model, and generates a left gain 530 as a result of processing the set of features. Noise reduction system 504b receives the transformed right signal 526, performs feature extraction on the transformed right signal 526 to extract a set of features, processes the set of features by inputting it into a trained monaural model, and generates a right gain 532 as a result of processing the set of features. Noise reduction system 504c receives the transformed left signal 524 and the transformed right signal 526, performs feature extraction on the transformed left signal 524 and the transformed right signal 526 to extract a set of features, processes the set of features by inputting it into a trained binaural model, and generates a combined gain 534 as a result of processing the set of features. Noise reduction systems 504a and 504b are otherwise similar to noise reduction systems 304a and 304b (see Figure 3), and noise reduction system 504c is otherwise similar to noise reduction system 404 (see Figure 4). For brevity, the explanation will not be repeated.

[0080] In summary, noise reduction systems 504a and 504b implement a machine learning system using a monaural model similar to that of audio processing system 300 (see Figure 3), while noise reduction system 504c implements a machine learning system using a binaural model similar to that of audio processing system 400 (see Figure 4). Therefore, audio processing system 500 can be viewed as a combination of audio processing systems 300 and 400.

[0081] The gain calculation system 506 receives the left gain 530, the right gain 532, and the combined gain 534, and combines the gains 530, 532, and 534 according to a mathematical function to generate a shared gain 536. The mathematical function can be one or more, such as a maximum, average, range function, or comparison function. The gains 530, 532, and 534 can be gain vectors of band gains, and the mathematical function is applied to each given band of the gains 530, 532, and 534. The gain calculation system 506 may otherwise be similar to the gain calculation system 306 (see Figure 3), and for brevity, its description will not be repeated.

[0082] Mixing systems 508a and 508b receive the converted left signal 524, the converted right signal 526, and the shared gain 536, apply the shared gain 536 to signals 524 and 526, and generate the mixed left signal 540 and the mixed right signal 542. In particular, mixing system 508a generates the mixed left signal 540 by applying the shared gain 536 to the converted left signal 524, and mixing system 508b generates the mixed right signal 542 by applying the shared gain 536 to the converted right signal 526. Mixing systems 508a and 508b may otherwise be the same as mixing systems 308a and 308b (see Figure 3), and for brevity, that explanation will not be repeated.

[0083] The inverse conversion systems 510a and 510b receive the mixed left signal 540 and the mixed right signal 542, perform inverse conversion, and generate the modified left signal 544 and the modified right signal 546. Specifically, the inverse conversion system 510a inversely converts the mixed left signal 540 to generate the modified left signal 544, and the inverse conversion system 510b inversely converts the mixed right signal 542 to generate the modified right signal 546. The inverse conversion performed by the inverse conversion systems 510a and 510b generally corresponds to the inverse conversion of the conversion performed by the conversion systems 502a and 502b, converting the signal back from the second signal domain to the first signal domain. As a result, the modified left signal 544 corresponds to a noise-reduced version of the left input signal 520, and the modified right signal 546 corresponds to a noise-reduced version of the right input signal 522. In other respects, the inverse conversion systems 510a and 510b may be the same as the inverse conversion systems 310a and 310b (see Figure 3) or 408a and 408b (see Figure 4).

[0084] Model training

[0085] As previously mentioned, the noise reduction systems 504a, 504b, and 504c use the trained monaural and binaural models to generate gains 530, 532, and 534 from the converted left signal 524 and the converted right signal 526, respectively. The training of the monaural models is generally similar to that of the models used by the noise reduction systems 304a and 304b (see Figure 3), and the training of the binaural models is generally similar to that of the models used by the noise reduction system 404 (see Figure 4), and for brevity, this explanation will not be repeated.

[0086] 2. Combined binaural audio and video capture

[0087] As mentioned earlier, UGC often includes combined audio and video capture. Simultaneous capture of video and binaural audio is particularly difficult. One such challenge is when binaural audio capture and video capture are performed by separate devices, for example, capturing video with a mobile phone and binaural audio with earphones. Mobile phones generally have two cameras: a front camera, also called a selfie camera, and a back camera, also called a main camera. When using the back (main) camera, this is sometimes called normal mode. When using the front (selfie) camera, this is sometimes called selfie mode. In normal mode, the user with the video capture device is behind the scene captured in the video. In selfie mode, the user with the video capture device is present in the scene captured in the video.

[0088] When binaural audio capture and video capture are performed on separate devices, there may be discrepancies between the captured video data and the captured binaural audio data compared to the human perception of the environment through sight and hearing. One example of such discrepancies is the perception of binaural audio captured simultaneously with video in normal mode versus the perception of binaural audio captured simultaneously with video in selfie mode. Another example of such discrepancies is the discontinuity that occurs when switching between normal mode and selfie mode. The following sections describe various processes for compensating for these discrepancies.

[0089] 3. Binaural audio capture in selfie mode

[0090] Figure 6 is a stylized overhead view showing binaural audio capture in selfie mode using video capture system 100 (see Figure 1). The video capture device 102 is in selfie mode and uses the front camera to capture video including the user in the scene. The user is wearing earphones 104 and 106 to capture binaural audio of the scene. The video capture device 102 is located between approximately 0.5m and 1.5m in front of the user, depending on whether the user is holding the video capture device 102 in their hand or using a selfie stick to hold the video capture device 102 in their hand. The video capture device 102 also captures other people near the user, for example, a person behind the user to the left, referred to as "Person Left," and a person behind the user to the right, referred to as "Person Right." Because the audio is captured binaurally, the listener perceives sounds emitted by Person Left as originating from behind and to the left, and sounds emitted by Person Right as originating from behind and to the right. This involves some corrections in selfie mode.

[0091] 3.1 Left / Right Correction

[0092] The opposite orientation of the user wearing earphones 104 and 106 and the front (selfie) camera of the video capture device 102 results in a left / right inversion of the captured binaural audio content. Consumers of the captured audiovisual content perceive the sound coming from the right earphone as coming from a source appearing on the left side of the video, and the sound coming from the left earphone as coming from a source appearing on the right side of the video, which contradicts our experience of seeing with our eyes and hearing with our ears.

[0093] Left / right correction involves taking the left channel from the input and transmitting it to the right channel of the output, or is expressed as the equation R'=L; or involves taking the right channel from the input and transmitting it to the left channel of the output, or is expressed as the equation L'=R.

[0094] 3.2 Front / Back Correction

[0095] When using the front (selfie) camera to record other speakers in the same scene, the user wearing earphones 104 and 106 and holding the video capture device 102 often stands slightly in front of the other speakers, i.e., closer to the camera. Therefore, in the case of captured binaural audio, the speech of the other speakers comes from behind the listener consuming the content. On the other hand, in captured video, all speakers appear in front.

[0096] Generally, to compensate for this and enhance perceptual consistency between audio and video, embodiments may implement front / back correction, which works to modify the spectral shape of sounds coming from behind the listener so that the sound is perceived in a similar way to sounds coming from the front.

[0097] Embodiments disclosed herein may implement spectral shape modification using a high-shelf filter. High-shelf filters can be constructed in various ways. For example, they may be implemented using an infinite impulse response (IIR) filter, such as a biquad filter.

[0098] Figure 7 is a graph showing an example of the magnitude response of a high-shelf filter implemented using a biquad filter. In Figure 7, the x-axis is frequency (kHz), and the y-axis is the magnitude of the loudness adjustment the filter applies to the signal. The high-shelf frequency in this example is approximately 3 kHz, which is a typical value considering the shading effect of the human head. Since the rear-captured audio is attenuated at high frequencies such as above 5 kHz, as shown in Figure 7, the filter implements a high-shelf that boosts these frequencies when the audio is corrected for forward input.

[0099] Embodiments disclosed herein may also perform spectral shape correction using an equalizer. The equalizer boosts or attenuates the input audio in one or more bands with different gains and may be implemented by an IIR filter or a finite impulse response (FIR) filter. The equalizer can shape the spectrum with greater precision, and in a typical configuration, the pre / post correction is an 8 to 12 dB boost in the frequency range of 3 to 8 kHz.

[0100] 3.3 Stereo Image Width Control

[0101] Figure 8 is a stylized overhead view showing various audio capture angles in selfie mode. Angle θ1 corresponds to the angle of the sound of the person on the right captured by the microphone of the video capture device 102 (see Figure 6), and angle θ2 corresponds to the angle of the sound of the person on the right captured by the right earphone 106. Compared to when the microphone is on the video capture device 102, earphones 104 and 106 are usually closer to the line in which the other speaker would normally be standing, so θ2 > θ1, which means that the speech of the other speaker is coming from a direction closer to the side, but based on the video scene, the viewer expects the speech to be coming from a direction closer to the center.

[0102] To address this problem, embodiments may implement stereo image width control to improve consistency between video and binaural audio recordings by compressing the perceived width of the binaural audio. In one implementation, compression is achieved by attenuating the side components of the binaural audio. First, the input binaural audio is converted to a middle-side representation according to equations (1.1) and (1.2).

number

[0103] In equations (1.1) and (1.2), L and R are the left and right channels of the input audio, for example, input signals 220 and 222 on the left and right in Figure 2, while M and S are the center and side components resulting from the conversion.

[0104] Next, the side channel S is attenuated by the attenuation coefficient α, and the processed output audio L' and R' are given by equations (2.1) and (2.2).

number

[0105] The attenuation coefficient α can be a function of the focal length f of the front (selfie) camera and is given by equation (3).

number

[0106] In equation (3), f c α=1 is the expected focal length = 1, i.e., the focal length to which the attenuation of the side component S does not apply, and is also called the baseline focal length; γ is the aggressiveness coefficient, which will be explained in more detail with reference to Figure 9.

[0107] Figure 9 is a graph of the attenuation coefficient α for different focal lengths f. In Figure 9, the x-axis represents focal lengths f in the range of 10 mm to 35 mm, and the y-axis represents the attenuation coefficient α and the baseline focal length f. c The aperture is 70mm, and the aggression factor γ can be selected from [1.2 1.5 2.0 2.5]. The aggression factor γ may be selectable by the device manufacturer to provide various options for the camera. For a typical front (selfie) camera on a smartphone with F=30mm, α is in the range of 0.5 to 0.7.

[0108] In summary, when video is captured at a smaller focal length, the video appears zoomed out, and the captured audio from both the left and right people appears to originate from the center of the video. Width control compensates for this by shrinking the audio scene to fit the video scene.

[0109] 4. Binaural audio capture in normal mode

[0110] Figure 10 is a stylized overhead view showing binaural audio capture in normal mode using the video capture system 100 (see Figure 1). The video capture device 102 is in normal mode and uses the rear camera to capture video that does not include the user in the scene. The user is wearing earphones 104 and 106 to capture binaural audio of the scene. In contrast to selfie mode (see Figures 6 and 8), in which the user is often captured within the video scene, in normal mode the user is rarely captured within the video scene. In normal mode, the user, wearing earphones 104 and 106 and holding the video capture device 102, is typically behind the video scene. Others are typically in front, as shown by the left and right people, to be captured in the video. Angle θ1 corresponds to the angle of the sound of the right person captured by the microphone of the video capture device 102, and angle θ2 corresponds to the angle of the sound of the right person captured by the right earphone 106.

[0111] In normal mode, the audio processing system does not need to perform either left / right correction or front / back correction, as may be done in selfie mode. Regarding stereo image width control, compared to when the microphone is on the video capture device 102, earphones 104 and 106 are usually farther away from the line in which other speakers would normally stand, so θ2 < θ1, and therefore the perceived width of the binaural audio can be slightly wider in this mode. However, the difference between θ1 and θ2 is not as large as in selfie mode, so for simplicity, a typical approach is to leave the binaural audio as is.

[0112] 5. Switching between normal mode and selfie mode

[0113] Different audio processing is often applied in normal mode compared to selfie mode. For example, left / right correction is performed in selfie mode but not in normal mode. When the user switches modes, it is beneficial for the audio processing system to make the transition smooth. Switching can occur during real-time operations, such as when capturing content for broadcast or streaming, and during non-real-time operations, such as when capturing content for later processing or uploading.

[0114] 5.1 Smoothing for left / right correction and stereo image width control

[0115] Recalling Section 3, we have L'=R and R'=L as formulas for left / right correction. These can be rewritten as formulas (4.1-4.4):

number

[0116] In equation (4.1-4.4), the attenuation coefficient α for left / right correction in selfie mode is -1.

[0117] Left / right correction is not required in normal mode, so in that mode, α = -1. Therefore, during the switch between normal mode and selfie mode, α switches between 1 and -1. To ensure a smooth transition, α should change its value gradually. An example of an equation for performing a smooth transition is given in equation (5):

number

[0118] In equation (5), t s t is the time it takes for the switch to be performed, and a 1-second transition time is sufficient for left / right correction switching. Therefore, the transition is t in the case of non-real-time. s Starting with -0.5, t s It ends at +0.5. In real time, equation (5) is t s It can be modified to start at a certain time and end at 1 second. The value of 1 second can be adjusted as needed, for example, within the range of 0.5 to 1.5 seconds.

[0119] Stereo image width control uses a similar set of equations, as shown in equations (6.1-6.4).

number

[0120] However, in equations (6.1-6.4), the attenuation coefficient α is in the range of 0.5 to 0.7 in selfie mode and 1.0 in normal mode.

[0121] In other words, stereo image width control involves generating a central channel M and side channels S, attenuating the side channels by a width adjustment coefficient α, and generating modified audio signals L' and R' from the attenuated central and side channels. The width adjustment coefficient is calculated based on the focal length of the video capture device, and the width adjustment coefficient can be updated in real time in response to changes in the focal length of the video capture device in real time.

[0122] Combining stereo image width control and smoothing of left / right correction results in an α value in the range of -0.5 to -0.7 for the self-timer mode and an α value of 1.0 in the normal mode. Assuming α = -0.5 as an example leads to Equation (7):

Equation

[0123] In Equation (7), t s is the time at which the switching occurs, and a transition time of 1 second functions well for the combination of left / right correction switching and stereo image width control switching. Therefore, in the non-real-time case, the transition starts at t s -0.5 and ends at t s +0.5. In the real-time case, Equation (7) can be modified to start at t<000,0010>and end in 1 second. The value of 1 second can be adjusted as needed, for example, within the range of 0.5 to 1.5 seconds.

[0124] 5.2 Smoothing of Front / Back Correction

[0125] As described in Sections 3 and 4, front / back correction is applied as spectral deformation in the self-timer mode, and front / back correction is not applied in the normal mode.

[0126] Let x org be the input of the front / back correction, and x fb be the output of the front / back correction. Then, the smoothed output of the front / back correction is given by Equation (8).

Equation

[0127] In Equation (8), α = 0 for the self-timer mode and α = 1 for the normal mode. An example of the equation for smooth transition is given by Equation (9).

Equation

[0128] In equation (9), t s This is the time it takes for the switch to occur, and a 6-second transition time is sufficient for pre / post correction. Therefore, the transition is not real-time. s Starting with -3, t s It ends with +3. In real time, equation (9) is t s It can be modified to start at 6 seconds and end at 6 seconds. The value of 6 seconds can be adjusted as needed, for example, within the range of 3 to 9 seconds.

[0129] For front / back smoothing, a longer transition time (e.g., 6 seconds) is used than for left / right and stereo image width smoothing (e.g., 1 second), because the front / back transition involves a change in timbre, which becomes less noticeable with a longer transition time.

[0130] 6. Exemplary Device Architecture

[0131] Figure 11 shows a device architecture 1100 for implementing the features and processes described herein, according to one embodiment. Architecture 1100 can be implemented in any electronic device, including but not limited to desktop computers, consumer audio / visual (AV) equipment, radio broadcasting equipment, and mobile devices such as smartphones, tablet computers, laptop computers, and wearable devices. In the exemplary embodiment shown, architecture 1100 is for a mobile phone. Architecture 1100 includes a processor(s) 1101, peripheral interface 1102, audio subsystem 1103, speaker 1104, microphone 1105, sensors 1106 (e.g., accelerometer, gyroscope, barometer, magnetometer, camera), location processor 1107 (e.g., GNSS receiver), wireless communication subsystem 1108 (e.g., Wi-Fi®, Bluetooth®, cellular), and I / O subsystem(s) 1109, a touch controller 1110 and other input controllers 1111, a touch surface 1112 and other input / control devices 1113. Other architectures with more or fewer components may also be used to implement the disclosed embodiments.

[0132] The memory interface 1114 is connected to the processor 1101, the peripheral interface 1102, and the memory 1115, such as flash, RAM, or ROM. The memory 1115 stores computer program instructions and data, including but not limited to operating system instructions 1116, communication instructions 1117, GUI instructions 1118, sensor processing instructions 1119, telephone instructions 1120, electronic message instructions 1121, web browsing instructions 1122, audio processing instructions 1123, GNSS / navigation instructions 1124, and application / data 1125. The audio processing instructions 1123 include instructions for performing the audio processing described herein.

[0133] According to one embodiment, the architecture 1100 may be compatible with a mobile phone connected to earphones that capture video data and binaural audio data (see Figure 1).

[0134] Figure 12 is a flowchart of the audio processing method 1200. Method 1200 can be implemented by a device, such as a laptop computer or mobile phone, having components of the architecture 1100 of Figure 11, to implement functions such as a video capture system 100 (see Figure 1) and an audio processing system 200 (see Figure 2) by running one or more computer programs.

[0135] In 1202, an audio signal is captured by an audio capture device. The audio signal has at least two channels, including a left channel and a right channel. For example, the left earphone 104 (see Figure 1) may capture the left channel (e.g., 220 in Figure 2), and the right earphone 106 may capture the right channel (e.g., 222 in Figure 2).

[0136] In step 1204, the noise reduction gain for each of at least two channels is calculated by a machine learning system. The machine learning system performs feature extraction, processes the extracted features by inputting them into a trained model, and may output a noise reduction gain as a result of processing the features. The trained model may be a monaural model, a binaural model, or both a monaural and a binaural model. In step 1206, a shared noise reduction gain is calculated based on the noise reduction gain of each channel.

[0137] Steps 1204 and 1206 can be performed as individual steps or as substeps of a combined operation. For example, noise reduction system 204 (see Figure 2) may calculate the left gain 230 and the right gain 232 as a shared noise reduction gain. As another example, noise reduction system 304a (see Figure 3) may generate the left gain 330, and noise reduction system 304b may generate the right gain 332; gain calculation system 306 may then generate a shared gain 334 by combining gains 330 and 332 according to a mathematical function. As yet another example, noise reduction system 404 (see Figure 4) may calculate a combined gain 430 as a shared noise reduction gain. As another example, noise reduction system 504a (see Figure 5) may generate a left gain 530, noise reduction system 504b may generate a right gain 532, and noise reduction system 504c may generate a combined gain 534; gain calculation system 506 may then generate a shared gain 536 by combining gains 530, 532, and 534 according to a mathematical function.

[0138] In 1208, the modified audio signal is generated by applying multiple shared noise reduction gains to each channel of at least two channels. For example, mixing system 206 (see Figure 2) can generate a mixed left signal 234 and a mixed right signal 236 by applying left gain 230 and right gain 232 to the converted left signal 224 and the converted right signal 226. As another example, mixing system 308a (see Figure 3) can generate a mixed left signal 336 by applying a shared gain 334 to the converted left signal 324, and mixing system 308b can generate a mixed right signal 338 by applying a shared gain 334 to the converted right signal 326. As another example, mixing system 406a (see Figure 4) may generate a mixed left signal 434 by applying a shared gain 430 to the converted left signal 424, and mixing system 406b may generate a mixed right signal 436 by applying a shared gain 430 to the converted right signal 426. As yet another example, mixing system 508a (see Figure 5) may generate a mixed left signal 540 by applying a shared gain 536 to the converted left signal 524, and mixing system 508b may generate a mixed right signal 542 by applying a shared gain 536 to the converted right signal 526.

[0139] Method 1200 may include additional steps corresponding to other functions of the audio processing system described herein. One such function is to convert an audio signal from a first signal domain to a second signal domain, perform audio processing in the second signal domain, and convert the processed audio signal back to the first signal domain, for example, using the conversion system 202 and inverse conversion system 208 in Figure 2. Another such function is simultaneous video and audio capture, including, for example, one or more of front / back correction, left / right correction, and stereo image width control correction, as described in Sections 3-4. Another such function is to smoothly switch between selfie mode and normal mode, including, for example, smoothing left / right correction using a first smoothing parameter and smoothing front / back correction using a second smoothing parameter, as described in Section 5.

[0140] 7. Alternative Embodiments

[0141] Many of the features are described above in combination, primarily due to the synergistic effects resulting from these combinations. Many features can be implemented independently of others, yet still offer advantages over existing systems.

[0142] 7.1 Single Camera System

[0143] While some of the features described here are in the context of video capture devices with two cameras, many of the features are also applicable to video capture devices with a single camera. For example, even a single-camera system can benefit from the binaural adjustments performed in normal mode, as described in Section 4.

[0144] 7.2 Smooth switching of video capture modes

[0145] Figure 13 is a flowchart of audio processing method 1300. Method 1200 (see Figure 12) performs noise reduction with smooth switching as an additional feature, as described in Section 5, although the smooth switching may be performed independently of the noise reduction. Method 1300 describes performing smooth switching independently of noise reduction. Method 1300 may be performed by a device, such as a laptop computer or mobile phone, having components of the architecture 1100 in Figure 11, for example, by running one or more computer programs to implement the functionality of the video capture system 100 (see Figure 1).

[0146] In 1302, an audio signal is captured by an audio capture device. The audio signal has at least two channels, including a left channel and a right channel. For example, the left earphone 104 (see Figure 1) may capture the left channel (e.g., 220 in Figure 2), and the right earphone 106 may capture the right channel (e.g., 222 in Figure 2).

[0147] In 1304, the video signal is captured by the video capture device at the same time as the audio signal is captured (see 1302). For example, the video capture device 102 (see Figure 1) can capture the video signal at the same time as the binaural audio signal is captured by the earphones 104 and 106.

[0148] In step 1306, the audio signal is corrected to produce a corrected audio signal. The correction may include one or more of the following: front / rear correction, left / right correction, and stereo image width correction.

[0149] In step 1308, the video signal is switched from the first camera mode to the second camera mode. For example, the video capture device 102 (see Figure 1) can switch from selfie mode (see Figures 6 and 8) to normal mode (see Figure 10), or from normal mode to selfie mode.

[0150] In 1310, smooth switching of the corrected audio signal occurs simultaneously with switching the video signal (see 1308). Smooth switching may be achieved using a first smoothing parameter to smooth one type of correction (e.g., left / right smoothing using equation (5), or combined left / right and stereo image width smoothing using equation (7)) and a second smoothing parameter to smooth another type of correction (e.g., front / back correction using equation (9)).

[0151] Implementation details

[0152] One embodiment may be implemented in hardware, an executable module stored in a computer-readable medium, or a combination of both, such as a programmable logic array. Unless otherwise specified, the steps performed by the embodiments do not necessarily have to be inherently associated with a specific computer or other device, even in a particular embodiment. In particular, various general-purpose machines may be used with programs written according to the teachings herein, or it may be more convenient to build a more specialized device, such as an integrated circuit, to perform the required method steps. Accordingly, the embodiments may be implemented in one or more computer programs that run on one or more programmable computer systems, each having at least one processor, at least one data storage system including volatile and non-volatile memory and / or memory elements, at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in known ways.

[0153] Each such computer program is preferably stored or downloaded to a storage medium or device readable by a general-purpose or special-purpose programmable computer, such as a solid-state memory or medium, magnetic or optical medium, in order to set up and operate the computer when the storage medium or device is read by the computer system in order to perform the procedures described herein. Furthermore, the system of the present invention may be implemented as a computer-readable storage medium configured with computer programs, and such a configured storage medium may be considered to cause the computer system to operate in a specific predetermined manner to perform the functions described herein. Software itself, and intangible or transient signals, are excluded to the extent that they are unpatentable subject matter.

[0154] The configurations of the systems described herein may be implemented in a suitable computer-based audio processing network environment for processing digital or digitized audio files. Parts of the adaptive audio system may include one or more networks having any number of individual machines, including one or more routers (not shown) that play a role in buffering and routing data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0155] One or more components, blocks, processes, or other functional components may be implemented through computer programs that control the execution of the system's processor-based computing devices. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware and firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media with respect to their operation, register transfers, logic components, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, various forms of physical, non-temporary, non-volatile storage media, such as optical, magnetic, or semiconductor storage media.

[0156] The above description illustrates various embodiments of the Disclosure, along with examples of how aspects of the Disclosure may be implemented. The above examples and embodiments should not be considered sole embodiments, but are presented to illustrate the flexibility and advantages of the Disclosure as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations and equivalents will be obvious to those skilled in the art and may be adopted without departing from the spirit and scope of the Disclosure as defined by the claims.

Claims

1. A computer-implemented method for audio processing, the method being: To capture an audio signal having at least two channels, including a left channel and a right channel, using an audio capture device; Converting the aforementioned audio signal from a first signal domain to a second signal domain, wherein the first signal domain is the time domain and the second signal domain is the frequency domain; A machine learning system calculates a plurality of noise reduction gains for each channel of the at least two channels, wherein the plurality of noise reduction gains are calculated based on the audio signal converted to the second signal domain; Calculating a plurality of shared noise reduction gains based on the plurality of noise reduction gains for each of the aforementioned channels; To generate a modified audio signal by applying the plurality of shared noise reduction gains to each channel of the at least two channels; and This includes converting the modified audio signal from the second signal domain to the first signal domain; method.

2. The calculation of the multiple noise reduction gains, the calculation of the multiple shared noise reduction gains, and the generation of the modified audio signal are performed simultaneously with the capture of the audio signal. The method according to claim 1.

3. The further includes storing the captured audio signal, The calculation of the plurality of noise reduction gains, the calculation of the shared noise reduction gain, and the generation of the modified audio signal are performed on the stored audio signal. The method according to claim 1.

4. The machine learning system can calculate the multiple noise reduction gains as follows: Perform feature extraction on each of the at least two channels so as to generate multiple features for each channel; Processing multiple features of each channel, wherein processing multiple features of each channel includes inputting multiple features of each channel into a machine learning model; and This includes, as a result of inputting the aforementioned features into the machine learning model, outputting the aforementioned plurality of noise reduction gains from the machine learning system; The method according to any one of claims 1 to 3.

5. The aforementioned machine learning model is a monaural model trained offline using monaural audio training data; The aforementioned features include a first set of features corresponding to the left channel and a second set of features corresponding to the right channel; The plurality of noise reduction gains include a first plurality of noise reduction gains corresponding to the first plurality of features, and a second plurality of noise reduction gains corresponding to the second plurality of features; The method according to claim 4.

6. The aforementioned machine learning model is a binaural model trained offline using binaural audio training data; The aforementioned features are features of the coupling that correspond to both the left channel and the right channel; The plurality of shared noise reduction gains arise from the plurality of features of the coupling corresponding to both the left channel and the right channel. The method according to claim 4.

7. The machine learning model includes a monaural model trained offline using monaural audio training data and a binaural model trained offline using binaural audio training data; The plurality of features include a first plurality of features corresponding to the left channel, a second plurality of features corresponding to the right channel, and a plurality of coupling features corresponding to both the left channel and the right channel; The plurality of noise reduction gains include a first plurality of noise reduction gains corresponding to the first plurality of features, a second plurality of noise reduction gains corresponding to the second plurality of features, and a plurality of coupling noise reduction gains corresponding to the coupling plurality of features. The method according to claim 4.

8. The audio capture device has a first earphone for capturing the left channel and a second earphone for capturing the right channel; The plurality of noise reduction gains include a first plurality of noise reduction gains and a second plurality of noise reduction gains; Calculating the plurality of shared noise reduction gains includes combining the first plurality of noise reduction gains and the second plurality of noise reduction gains according to a mathematical function. The method according to any one of claims 1 to 7.

9. The aforementioned mathematical function includes one or more of the following: mean, maximum, range function, and comparison function. The method according to claim 8.

10. The first plurality of noise reduction gains correspond to a first gain vector for the plurality of bands of the left channel, and the second plurality of noise reduction gains correspond to a second gain vector for the plurality of bands of the right channel; Calculating the plurality of shared noise reduction gains includes selecting the maximum gain for each of the plurality of bands from the first gain vector and the second gain vector. The method according to claim 8.

11. The aforementioned plurality of noise reduction gains further include a coupled plurality of noise reduction gains; Calculating the plurality of shared noise reduction gains includes combining the first plurality of noise reduction gains, the second plurality of noise reduction gains, and the combined plurality of noise reduction gains according to the mathematical function. The method according to claim 8.

12. The video capture device further includes capturing the audio signal and the video signal simultaneously, The video capture device includes a mobile phone, and the mobile phone includes a front camera and a rear camera. The method according to any one of claims 1 to 11.

13. The method further includes switching from a first mode using either the front camera or the rear camera to a second mode using the other of the front camera or the rear camera, wherein the switching includes smoothing the left / right correction of the audio signal using a first smoothing parameter and smoothing the front / back correction of the audio signal using a second smoothing parameter. The method according to claim 12.

14. Capturing the audio signal and simultaneously capturing the video signal includes performing corrections on the audio signal, the corrections including at least one of left / right correction, front / back correction, and stereo image width control correction. The method according to claim 12 or 13.

15. Performing the aforementioned stereo image width control correction means: To generate a center channel and side channels from the left channel and the right channel of the aforementioned audio signal; Attenuating the side channel by a width adjustment coefficient; and Including generating a modified audio signal from the central channel and the attenuated side channels; The method according to claim 14.

16. The width adjustment coefficient is calculated based on the focal length of the video capture device. The method according to claim 15.

17. The width adjustment coefficient is updated in real time in response to the video capture device changing the focal length in real time. The method according to claim 16.

18. A non-temporary computer-readable medium for storing a computer program that, when executed by a processor, controls the device to perform a process including the method according to any one of claims 1 to 17.

19. An apparatus for audio processing, wherein the apparatus is: The device has a processor, the processor is configured to control the device to perform a process including the method according to any one of claims 1 to 17, Device.

Citation Information

Patent Citations

  • Sound field spatial stabilizer with structured noise compensation

    EP2816816A1

  • Method and system for a multi-microphone noise reduction

    US20110305345A1

  • Sound field spatial stabilizer with echo spectral coherence compensation

    US20140376744A1

  • Wind noise reduction

    US20160155453A1

  • headset

    US20190387306A1