Method and apparatus for processing binaural recordings

CN116349252BActive Publication Date: 2026-08-18DOLBY LABORATORIES LICENSING CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180068152.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-21
Filing Date
2021-09-15
Publication Date
2026-08-18
Estimated Expiration
2041-09-15

AI Technical Summary

Technical Problem

[0003]双耳捕获设备的一个缺点是,双耳捕获设备对环境噪声非常敏感,这导致在渲染捕获的双耳信号时播放体验不佳

Benefits of technology

[0006] Based on the foregoing, the object of the present invention is to provide a method and apparatus for processing binaural audio signals more efficiently, and a method for rendering processed binaural audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116349252B_ABST
    Figure CN116349252B_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for processing a first audio signal and a second audio signal representing input binaural audio signals acquired by a binaural recording device. The invention further relates to a method for rendering binaural audio signals on a loudspeaker system. The method for processing binaural signals comprises extracting audio information from the first audio signal, calculating a band gain for reducing noise in the first audio signal, and applying the band gain to a corresponding band of the first audio signal according to a dynamic scaling factor to provide a first output audio signal. Wherein the value of the dynamic scaling factor is between zero and one and is selected to reduce quality degradation of the first audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and apparatus for processing binaural audio signals. Background Technology

[0002] In both user-generated content (UGC) and professionally generated content (PGC) fields, binaural capture devices are commonly used to capture audio. For example, binaural audio is recorded using a pair of microphones, each positioned in the earpiece of a pair of headphones worn by the user. Therefore, a binaural capture device captures sound at each corresponding ear of the user wearing the device. Consequently, binaural capture devices are generally adept at capturing a user's speech or audio perceived by the user. Binaural capture devices are therefore frequently used for recording podcasts, interviews, or meetings.

[0003] One drawback of binaural capture devices is that they are very sensitive to ambient noise, which results in a poor playback experience when rendering the captured binaural signals.

[0004] Another drawback of binaural capture devices is that, in addition to the user's voice, the captured audio sources of interest have very low signal strength, high noise, and high reverberation. Therefore, the intelligibility of other audio sources of interest within the captured binaural audio signal is reduced.

[0005] To circumvent these drawbacks, previous solutions involved complex audio processing algorithms that were computationally intensive to execute, making them particularly difficult to implement for low-latency communications or UGC where complex audio processing was challenging. Summary of the Invention

[0006] Based on the foregoing, the object of the present invention is to provide a method and apparatus for processing binaural audio signals more efficiently, and a method for rendering processed binaural audio signals.

[0007] According to a first aspect of the invention, a method is provided for processing a first audio signal and a second audio signal representing an input binaural audio signal. The binaural audio signal is acquired by a binaural recording device. The method includes: extracting audio information from the first audio signal, wherein the audio information includes at least a plurality of frequency bands representing the first audio signal; and calculating a frequency band gain for each frequency band to reduce noise in the first audio signal. Furthermore, the method includes applying the frequency band gain to the corresponding frequency band of the first audio signal according to a dynamic scaling factor to provide a first output audio signal. The value of the dynamic scaling factor is between zero and one, wherein a value of zero indicates that no frequency band gain is applied, and a value of one indicates that full-band gain is applied without modification. The dynamic scaling factor is selected to reduce quality degradation of the first audio signal, and the method further includes:

[0008] A second output audio signal is provided based on the second audio signal, and a binaural output audio signal is determined based on the first output audio signal and the second output audio signal.

[0009] The invention according to the first aspect is based at least in part on the understanding that the quality degradation of the output audio signal can be reduced by dynamically scaling the band gain. Regardless of the noise reduction method used to calculate the noise reduction band gain, the audio signal to which the band gain is applied will contain undesirable audio artifacts introduced by the noise reduction processing. To mitigate these audio artifacts, the band gain is applied dynamically according to a dynamic scaling factor. A static or predetermined scaling factor will not be able to reduce most possible audio signal quality degradation by implementing the band gain to a high degree that results in audio artifacts or to a low degree that suppresses noise reduction. The selection of the dynamic scaling factor can be based on the audio information and / or the band gain of the audio signal, so that a dynamic (non-static) scaling factor tailored after processing a particular audio signal can be used.

[0010] In some implementations, the dynamic scaling factor for each frequency band is based on the band gain associated with the corresponding frequency bands of the current time frame and previous time frames of the first audio signal.

[0011] A time frame refers to a portion of the time interval of the first audio signal. Accordingly, by analyzing the band gain of each frequency band in the current and previous time frames, the dynamic scaling factor is dynamically adjusted for the current first audio signal being processed. This optimizes the dynamic scaling factor to provide a first output audio signal with reduced quality degradation.

[0012] In some implementations, the method further includes processing additional audio signals from an additional recording device. This is achieved by synchronizing the additional audio signals with binaural audio signals and providing additional output audio signals based on the additional audio signals.

[0013] The additional recording device can be any device capable of recording at least a mono audio signal. For example, the additional recording device could be a user's smartphone. The additional audio signal can be used to enhance audio from a user wearing a binaural recording device or from a second source of interest. Because binaural recording devices are prone to picking up noise and reverberation from the surrounding environment, they are not suitable for recording audio from sources of interest other than the user wearing a binaural recording device (e.g., a respondent speaking to the user). Therefore, an additional recording device that records the additional audio signal can be used as a microphone to record audio from a second source of interest. Synchronizing the additional audio signal with the binaural signal, and combining the binaural signal with the synchronized additional audio signal, can facilitate, for example, clearer dialogue reproduction.

[0014] Some implementations also include processing bone vibration sensor signals acquired by the bone vibration sensors of the binaural recording device. By synchronizing the bone vibration sensor signals with the binaural audio signals and extracting the VAD probability of the additional audio signal, the source of the detected speech can be determined based on the VAD probability and the bone vibration sensor signals. If the source is a wearer of the binaural recording device with the bone vibration sensor, a first audio processing scheme is used to process the additional audio signal. If the source is not a wearer of the binaural recording device with the bone vibration sensor, a second audio processing scheme is used to process the additional audio signal. Using different processing schemes to process the additional audio signal allows for adaptive switching of gain levels and / or noise reduction processing depending on the source of the detected speech. This adaptive switching of audio processing schemes can be implemented in conjunction with the dynamic processing described above or using other general forms of audio processing and / or noise reduction methods.

[0015] For example, as a second aspect of the invention, a method is provided for processing a first audio signal, a second audio signal, and an additional audio signal, wherein the first and second audio signals represent input binaural audio signals acquired by a binaural recording device, and the additional audio signal is recorded by the additional recording device. The method includes: synchronizing the additional audio signal with the binaural audio signals; receiving a bone vibration sensor signal acquired by a bone vibration sensor of the binaural recording device; and synchronizing the bone vibration sensor signal with the binaural audio signals. Further, the method includes: extracting a VAD probability of the additional audio signal; and determining the source of the detected speech based on the VAD probability and the bone vibration sensor signal. If the source is a wearer of the binaural recording device with the bone vibration sensor, the additional audio signal is processed using a first audio processing scheme. If the source is not a wearer of the binaural recording device with the bone vibration sensor, the additional audio signal is processed using a second audio processing scheme. Additionally, an additional output audio signal is provided based on the processed additional audio signal, and a first output audio signal and a second output audio signal are provided based on the first and second audio signals from which the binaural output audio signal is determined.

[0016] Providing a first output audio signal and a second output audio signal may include performing audio processing and / or other forms of audio processing on the first audio signal and the second audio signal according to one aspect of the invention, such as noise cancellation and / or equalization.

[0017] According to a third aspect of the present invention, an audio processing apparatus is provided. The audio processing apparatus includes: a receiver configured to receive an input binaural audio signal comprising a first audio signal and a second audio signal; and an extraction unit configured to receive the first audio signal from the receiver and extract audio information from the first audio signal. The audio information includes at least a plurality of frequency bands representing a portion of the frequency composition of the first audio signal. The audio processing apparatus further includes: a processing unit configured to receive the audio information and calculate a band gain for each frequency band of the first audio signal, wherein the calculated band gain reduces noise in the first audio signal. An application unit of the audio processing apparatus is configured to apply the band gain to the corresponding frequency band of the first audio signal according to a dynamic scaling factor to provide a first output audio signal. The value of the dynamic scaling factor is between zero and one, wherein a value of zero indicates that no band gain is applied, and a value of one indicates that full-band gain is applied without modification. The dynamic scaling factor is selected to reduce the quality degradation of the first audio signal otherwise introduced by the noise reduction band gain. In this audio processing device, an additional processing module is configured to provide a second output audio signal based on the second audio signal, and an output stage is configured to determine a binaural output audio signal based on the first output audio signal and the second output audio signal.

[0018] The invention according to the second or third aspect has the same or equivalent embodiments and benefits as the invention according to the first aspect. Furthermore, any function described with respect to the processing method may have corresponding components present in a processing device or corresponding code for performing these functions in a computer program product. Attached Figure Description

[0019] The invention will be described in more detail with reference to the accompanying drawings, which illustrate embodiments of the invention according to the first or second aspect.

[0020] Figure 1 An exemplary binaural recording device and an additional recording device are described.

[0021] Figure 2 A binaural processing device according to some embodiments is described.

[0022] Figure 3 This is a flowchart illustrating a method for processing a first audio signal and a second audio signal according to an embodiment of the present invention.

[0023] Figure 4a This is a flowchart illustrating an alternative method for applying band gain based on a dynamic scaling factor.

[0024] Figure 4b This is a flowchart illustrating an alternative method for applying band gain based on a dynamic scaling factor.

[0025] Figure 5 The diagram illustrates the frequency bands representing a series of time frames of an audio signal.

[0026] Figure 6 This is a flowchart illustrating the estimation and processing of side and mid-signal signals according to some implementation methods.

[0027] Figure 7 This is a flowchart describing a rendering method according to one aspect of the present invention. Detailed Implementation

[0028] Figure 1 A user 4 wearing a binaural recording device 1 is depicted. The binaural recording device 1 may include two pairs of wired (not shown) or wireless microphones 2a, 2b, which may optionally be disposed in the respective earpieces of an earphone. The binaural recording device 1 records binaural audio signals comprising two audio signals, for example, a left audio signal and a right audio signal originating from the left microphone 2a and the right microphone 2b in each respective earpiece. In some embodiments, an additional recording device 31 records additional audio signals and / or a bone vibration sensor 11 records bone vibration signals. For example, the additional recording device 31 may be a microphone disposed in a user device 3 (e.g., a smartphone, tablet, or laptop), and the bone vibration sensor 11 may be disposed as an integrated part of the binaural recording device 1 (e.g., integrated in the earpiece as shown) or disposed externally (not shown). The additional recording device 31 may record a second source of interest (such as a second person speaking to the user 4). Alternatively, the additional recording device 31 may record the user 4's speech.

[0029] The bone vibration sensor signal from the bone vibration sensor 11 can indicate whether the user 4 wearing the binaural recording device 1 is speaking and / or the bone vibration sensor signal can be used to extract audio. Furthermore, the bone vibration sensor signal can be used in conjunction with the first and / or second audio signals to extract enhanced audio information.

[0030] The first and second audio signals recorded by the binaural recording device 1 can be time-synchronized by a binaural processing device 32 optionally configured in the user equipment 3, and additional audio signals and / or bone vibration sensor signals can be synchronized with the binaural audio signals by the binaural processing device 32. In some embodiments, the additional audio signals and / or bone vibration sensor signals are time-synchronized by the binaural processing device 32 using a software implementation. For example, synchronization between the binaural audio signals and the additional audio signals and / or bone vibration sensor signals is achieved by the processing device seeking a delay between the signals characterized by maximum correlation between the signals. Alternatively, each recorded data block or time frame representing a portion of the binaural audio signals and the additional audio signals and / or bone vibration sensor signals is associated with a timestamp, and these signals are synchronized by comparing the timestamps of each block.

[0031] Apart from signal time synchronization, any audio processing described below can be performed by the binaural processing device 32. The binaural processing device 32 can be wholly or partially integrated into the binaural recording device 1, and the user device 3 communicates with the binaural recording device 1 via wired or wireless (e.g., Bluetooth) communication. For example, the binaural processing device 32 of the user device 3 can receive, synchronize, and process all audio signals from the binaural recording device 1, any bone vibration sensor 11, and any additional recording device 31.

[0032] Further reference Figure 2 The image depicts a binaural processing device 32 according to some embodiments. The binaural processing device 32 is configured to receive binaural audio signals comprising two audio signals, for example, a left audio signal L and a right audio signal R recorded by a binaural recording device 1. In a synchronization module 321, the two audio signals L and R are synchronized. In some embodiments, the synchronization module 321 is integrated into the binaural recording device 1 and has further processing steps, such as synchronization with any bone vibration signals and / or additional audio signals performed by the user device 3.

[0033] Synchronization module 321 outputs a synchronized audio signal to optional transformation module 322. Optional transformation module 322 can extract audio information and / or alternative representations of the synchronized audio signals L and R. The alternative representations of the audio signals (referred to as A1 and B1) are provided to corresponding processing modules 323a and 323b. Each processing module 323a and 323b is configured to perform audio processing including noise reduction of the audio signal representations A1 and B1. In some embodiments, processing modules 323a and 323b perform processing equivalent to the first and second processing sequences described below.

[0034] The processed audio signals A2 and B2 output by signal processing modules 323a and 323b are provided to inverse transformation module 324, which performs an inverse transformation to regenerate processed audio signals PL and PR corresponding to the audio signals received at optional transformation module 322. In some embodiments, transformation module 322 and inverse transformation module 324 are not used, and the two audio signals L and R of the binaural recording device are processed in their original format.

[0035] The output stage 325 combines the first output audio signal PL and the second output audio signal PL into a binaural output audio signal representing the two output audio signals.

[0036] In some embodiments, the binaural processing device 32 incorporates the bone vibration sensor signal BV in the first and / or second processing modules 323a, 323b. Furthermore, the binaural processing device 32 may be further configured to receive an additional audio signal, synchronously and optionally transform the additional audio signal such that the additional audio signal is represented by at least one alternative representation of a first audio signal A1 and a second audio signal B1. Alternatively, in addition to the first and second processing modules 323a and 323b, a third processing module is added to process the additional audio signal and output the additional audio signal to an output stage 325, which generates a binaural output audio signal having auxiliary information representing the processed additional audio signal.

[0037] Figure 3 This is a flowchart illustrating a method according to some embodiments. At S1, input binaural audio signals, represented by a first audio signal A1 and a second audio signal B1, are received. The first and second audio signals can be synchronized left and right audio signals or alternative representations, such as side and mid audio signals. The first audio signal A1 is passed to a first processing sequence S2a, and the second audio signal B1 is passed to a second processing sequence S2b.

[0038] At step S21, audio information is extracted from the first audio signal A1. This audio information includes at least a representation of multiple frequency bands, each representing a portion of the frequency composition of the first audio signal A1. Furthermore, extracting audio information from the first audio signal A1 may include extracting acoustic parameters describing the first audio signal A1.

[0039] Extracting audio information at S21 may include first decomposing the first audio signal A1 into spectral information. The spectral information may be represented by a continuous or discrete spectrum, such as a Fourier spectrum or a filter bank (such as QMF). The spectral information may be represented by multiple binary numbers (bins), each binary number including a value, such that the multiple binary numbers represent discrete samples of the spectral information.

[0040] Secondly, the first audio signal A1 can be divided into multiple frequency bands, which may involve grouping binary numbers representing spectral information individually or in an overlapping manner to form the multiple frequency bands.

[0041] Spectral information can be used to extract frequency band features to be included in the audio information, such as Mel-frequency cepstral coefficients (MFCC) or Barker-frequency cepstral coefficients (BFCC). The frequency band harmonic features, fundamental frequency (F0), speech activity detection (VAD) probability, and signal-to-noise ratio (SNR) of the first audio signal A1 can be extracted by analyzing the spectral information of the first audio signal A1 and / or the first audio signal A1. Accordingly, the audio information may include one or more of the following for each frequency band of the first audio signal A1: frequency band harmonic features, fundamental frequency, VAD probability, and SNR.

[0042] At least based on the frequency bands representing the first audio signal A1 extracted at S21, a band gain BGain for each frequency band is calculated at S22. The calculated band gain BGain is used to reduce noise in the first audio signal A1. In some embodiments, calculating the band gain BGain includes predicting the band gain BGain based on the audio information using a trained neural network. This neural network may be a deep neural network and includes multiple neural network layers, each with multiple nodes. The neural network may be a fully connected neural network, a recurrent neural network, a convolutional neural network, or a combination thereof. A Wiener filter may be combined with the neural network to provide a final prediction of the band gain. Given at least a frequency band representing a portion of the first audio signal A1, the neural network is trained to predict the associated band gain BGain used for noise reduction. In some embodiments, the neural network (or a separate neural network) is further trained to also predict the VAD probability given at least a frequency band representing a portion of the frequency information of the first audio signal.

[0043] At S23, the band gain BGain from S22 is applied to the first audio signal A1 based on the dynamic scaling factor k obtained from S24 to form a first audio output signal A2 with reduced quality degradation. Specifically, the dynamic scaling factor k is selected at S24 based on the band gain BGain calculated at S22 to reduce quality degradation. By selecting the dynamic scaling factor k to reduce quality degradation, the calculated band gain BGain for each band can be adjusted according to the dynamic scaling factor k before being applied to the first audio signal A1 to provide a first output audio signal A2 with reduced quality degradation. The value of the dynamic scaling factor k is between zero and one and indicates the extent to which the calculated band gain is applied. In some embodiments, the dynamic scaling factor k for each band is based on at least one of the first audio signal A1, at least a portion of the audio information, and the calculated band gain BGain for each band.

[0044] Based on the second audio signal B1 of the binaural audio signal, a second output audio signal B2 is provided by processing the second audio signal B1 in a second processing sequence S2b. For example, the second processing sequence S2b may include performing separate processing (including, for example, noise reduction processing) on ​​the second audio signal B1 to form the second output audio signal B2. The separate processing of the second audio signal B1 may be equivalent to the processing of the first audio signal A1 in the first processing sequence S1a, and involves steps corresponding to steps S21, S22, S23, and S24.

[0045] In some implementations, the processing of the first audio signal A1 and the second audio signal B1 is coupled in the corresponding processing sequences S2a, S2b, for example, to apply a mono denoising model. A mono denoising model means that for each audio signal A1, B1, a corresponding set of denoising band gains BGain is calculated before reducing the band gain BGain to a single general set. This general set of band gains can be determined as the maximum, minimum, or average band gain for each band in all audio signals A1, B1. In other words, the calculated band gain BGain for each audio signal A1, B1 can initially be represented by a band gain matrix denoted as BGains(i,b), where i = 1: the number of audio signals, and b = 1: the number of bands. Accordingly, each row of BGains(i,b) includes all band gains of the signal, and each column includes the band gain for a given band of each audio signal. In a mono noise reduction matrix, the single-row band gain is extracted by merging each column into a single value (e.g., by finding the maximum value in each column). The same single-row band gain is then used for subsequent processing of all audio signals.

[0046] At S3, the first output audio signal A2 and the second output audio signal B2 are combined into a binaural output signal with reduced quality degradation.

[0047] Figure 3 The illustration also depicts a method according to some embodiments, wherein a bone vibration sensor signal BV is used in the processing of the first audio signal A1. The recorded signal from the bone vibration sensor is more resistant to ambient noise, and the bone vibration sensor signal can be used to extract additional audio information and / or enhanced audio information and / or enhanced bandwidth gain.

[0048] In some implementations, the bone vibration sensor signal BV is used to extract the VAD probability for each time frame or each frequency band of each time frame, or to provide an enhanced VAD probability extracted from the first audio signal A1 and the bone vibration sensor signal BV. At S21 and S22, the bone vibration sensor signal BV alone or in combination with the first audio signal A1 can be used to extract at least one of the following: spectral information, band gain, speech fundamental frequency, SNR, and VAD probability.

[0049] The bone vibration sensor signal BV can constitute a separate recording that supplements the first audio signal A1 and the second audio signal of the binaural audio signal. For example, the bone vibration sensor signal BV can be regarded as an additional audio signal and added to the binaural audio signal, or it can be provided as a separate output signal.

[0050] An enhanced first audio signal can be obtained from information derived from both the bone vibration sensor signal BV and the first audio signal A1. At S21, enhanced audio information (such as a more accurate representation of frequency composition) can be extracted from the enhanced first audio signal, and at S22, the enhanced band gain can be calculated based on this information. In some embodiments, in addition to the audio information, the bone vibration sensor signal BV is also provided to the neural network to predict the band gain and / or VAD probability at S22.

[0051] Similarly, in the processing of the second audio signal B2 in the second processing sequence S2b, the bone vibration sensor signal BV can be provided and taken into account.

[0052] Figure 4aThis is a flowchart illustrating how a band gain BGain is applied to the corresponding frequency band according to a dynamic scaling factor k at S23a. The band gain BGain calculated at S22 is provided together with the first audio signal A1, and at S231, the calculated band gain is applied to the first audio signal A1 to form a denoised first audio signal NA1. The denoised first audio signal NA1 may exhibit undesirable audio artifacts introduced by applying the band gain at S231. At S24, a dynamic scaling factor k is selected or calculated to reduce quality degradation, as will be described below. At S232, the denoised first audio signal NA1 is mixed with the (original) first audio signal A1 at a mixing ratio corresponding to the dynamic scaling factor k selected at S24 to apply the band gain according to the dynamic scaling factor k. Accordingly, the first output audio signal A2 is calculated as follows:

[0053] A2=k×A1+(1-k)NA1

[0054] This is based on a first audio signal A1, a noise-reduced first audio signal NA1, and a dynamic scaling factor k. Mixing can be performed for each frequency band of the first audio signal A1 with a corresponding dynamic scaling factor k. The dynamic scaling factor k can be the same for two or more frequency bands. After mixing the noise-reduced first audio signal NA1 with the first audio signal A1 at a mixing ratio equal to the dynamic scaling factor k, a first output audio signal A2 with reduced quality degradation is obtained.

[0055] Figure 4b The diagram illustrates an alternative method for applying the band gain BGain based on the dynamic scaling factor k. At S23b, the calculated band gain of the first audio signal A1 from S22, the selected dynamic scaling factor k from S24, and the first audio signal A1 are available. The dynamic scaling factor k indicates the extent to which the predicted band gain at S22 should be applied, and therefore, the first output audio signal is a weighted sum of the first audio signal A1 and the first audio signal A1 with the applied band gain BGain. That is, the first output audio signal A2 can be calculated using the following formula:

[0056] A2=kA1+(1-k)×BGain×A1=(k+(1-k)BGain)A1

[0057] in,

[0058] (k+(1-k)BGoin)

[0059] This is referred to as dynamic band gain. Accordingly, there is no need to calculate the first audio signal to be denoised and to perform a mixing of the denoised first audio signal with the first audio signal A1, because the dynamic band gain is calculated and applied to the first audio signal A1. The dynamic band gain for each band is extracted from the dynamic scaling factor k for each band and the calculated band gain BGain. When the dynamic band gain is applied to the first audio signal A1, a first output audio signal A2 with reduced quality degradation is formed.

[0060] Figure 5 The illustration shows a time-frame representation of an audio signal (e.g., a first audio signal). The audio signal is divided into multiple frames 101, 102, 103, and 104, represented by columns, and each time frame includes multiple frequency bands, represented by rows. For a specific frequency band 100, the calculated band gain (in linear units) is illustrated as 0.4, 0.6, and 0.7 for previous frames 101, 102, and 103, and as 0.8 for the current frame 104.

[0061] A method is provided for determining a dynamic scaling factor k based on calculated band gain. For example, the dynamic scaling factor k is based on the band gain calculated for the current (n+1) time frame 104 and previous (n, n-1, n-2) time frames 101, 102, 103 of the audio signal. In some embodiments, the dynamic scaling factor k for a specific band 100 of the current frame 104(n+1) is determined according to a weighted sum of gains G(n+1), wherein the weighted sum G(n+1) is calculated as follows:

[0062] G(n+1)=aG(n)+(1-a)BGain(n+1)

[0063] Here, 'a' is a constant that determines the extent to which the calculated band gain BGain(n+1) of the current frame 104 will modify the weighted sum of the gains G(n+1) of the current frame 104. The constant 'a' is between zero and one, preferably between 0.9 and 1, such as a = 0.99 or a = 0.9999. The constant 'a' can be 1–ε, where ε is in the range of 10. -1 Up to 10 -6 The initial value of G can be set to one. In other examples, the initial value of G is between 1 and 0.6, such as 0.8. It should be understood that the corresponding processing of previous frames 101, 102, and 103 may affect the value of G(n), and thus affect the final value of G(n+1) for the current frame 104. The dynamic scaling factor k can be linearly proportional to G(n+1), for example, the dynamic scaling factor k for the current frame 104 can be calculated as follows:

[0064] k = 1 - G(n+1).

[0065] In some implementations, the dynamic scaling factor k of the current frame 104 may be affected only by the gain T of the previous frames 101, 102, and 103 exceeding a predetermined threshold. Gain The influence of bandwidth gain. Predetermined threshold gain T Gain It can be between 0.3 and 0.7, and preferably around 0.5 (in linear units). This can be achieved by responding only to the calculated band gain BGain exceeding a predetermined threshold gain T. Gain This is achieved by updating the weighted sum of the gain G. Accordingly, the weighted sum of the gain G(n+1) for the current frame 104 is given by the following equation:

[0066]

[0067] Wherein, G(n) is subject to a gain T exceeding the threshold. Gain The influence of previous frames 101, 102, and 103.

[0068] As an example, when T Gain When the gain is 0.5, the calculated band gain of the frequency band 100 of the first frame 101 does not exceed the predetermined threshold gain T. Gain Because of 0.4 <T Gain Therefore, when the initial value of the weighted sum of gains G is one, the dynamic scaling factor k of the frequency band 100 of the first time frame 101 can be zero; for example, as described above, k = 1 - G. Thus, the frequency band 100 of the first processed frame 101 is equal to the calculated frequency band 100 of the first (denoised audio) audio signal. Since each of the subsequent time frames 102, 103, and 104 has a gain T exceeding a predetermined threshold... Gain However, the calculated band gain is less than one, so the processing of each subsequent frame 102, 103, 104 involves obtaining a lower G value, and in response, a larger dynamic scaling factor k. This means that the applied band gain will begin to deviate from the calculated band gain and approach the original audio signal of band 100 of frames 102, 103, 104.

[0069] It should be understood that, by Figure 5 Each frequency band represented in the row is associated with a corresponding weighted sum of frequency band gains G, which describe the frequency band gains of individual frequency bands in the current time frame 104 and previous time frames 101, 102, and 103.

[0070] Furthermore, the calculated band gain BGain(n+1) in response to the current frame 104 exceeds a predetermined threshold gain T. GainFurthermore, the calculated band gain BGain(n+1) also exceeds one (in linear units), and the calculated band gain BGain(n+1) can be set to a predetermined maximum value before updating the weighted sum of the band gains G(n+1). The predetermined maximum value can be one (in linear units), which means that the resulting dynamic mixing ratio k is guaranteed to remain in the range of zero to one.

[0071] For offline processing, a predetermined threshold gain T can be applied to each frequency band. Gain The average of all calculated band gain BGain is used to determine all time frames 101, 102, 103, and 104 (by...). Figure 5 The columns in the table represent the dynamic scaling factor k for each frequency band, to form a weighted sum of the frequency band gains G or to calculate the average frequency band gain from the dynamic scaling factor k.

[0072] In some implementations, the dynamic scaling factor can be further based on the VAD probability of each frequency band in each time frame 101, 102, 103, 104. In addition to the predetermined threshold gain T Gain In addition to the standard of weighted summation of the updated band gain G, the VAD probability can be further categorized. To this end, determining the dynamic scaling factor k may also include determining whether the VAD probability of band 100 in the current frame 104 exceeds a predetermined VAD probability threshold T. VAD Predetermined VAD probability threshold T VAD Between 0.4 (40%) and 0.6 (60%), preferably around 0.5 (50%). Accordingly, when determining the dynamic scaling factor k of the current frame 104, only the band gain BGain of the current frame 104 and the previous frames 101, 102, 103 are considered, where the audio signal is likely to represent speech.

[0073] By considering the band gain and optionally VAD probability of each frequency band in the current time frame 104 and previous time frames 101, 102, 103, the dynamic scaling factor k can be updated during online processing so that each frame (and each frequency band) of the audio signal obtains a suitable band gain BGain appropriate for reducing quality degradation given the available information. Accordingly, regardless of the audio signal being processed, the dynamic scaling factor can quickly approach a value suitable for reducing quality degradation in each additional processing time frame 101, 102, 103, 104.

[0074] For offline processing, the band gain and optionally audio information of frequency band 100 in all frames 101, 102, 103, and 104 of the audio signal can be analyzed to determine the dynamic scaling factor k for each frequency band, thereby determining the band gain to be applied to all frames of the audio signal. This can be achieved by applying a predetermined threshold gain T to each frequency band. gain and the predetermined probability threshold TVAD The average of all calculated band gains BGain is used to determine the dynamic scaling factor for each band across all time frames, forming a weighted sum of band gains G.

[0075] In a further example illustrated in Figure 4, the band gain of a specific frequency band 100 in the current frame 104 is calculated to be 0.8 (linear units), while the corresponding calculated band gains for the previous three frames 101, 102, and 103, in ascending order of time, are 0.4, 0.6, and 0.7 (linear units), respectively. At a predetermined threshold gain T... Gain When the gain is 0.5, the band gain of frames 102, 103, and 104 will affect the weighted sum of the band gain G of the current frame 104 and the resulting dynamic scaling factor k. When processing the previous frame 103, the band gain of frame 102 affected the weighted sum of the band gain G, while the band gain of frame 101 was lower than the threshold gain T. Gain And thus ignored. The selection of frames that influence the choice of the dynamic scaling factor k for the current frame 104 can differ, based on the VAD probability calculated for each frequency band of frames 101, 102, 103, and 104. For example, if the VAD probability of the previous frame 103 is lower than the probability threshold T... VAD Only frames 102 and 104 can influence the selection of the dynamic scaling factor k of the current frame 104. Frame 102 is ignored because its bandwidth gain is too low, and frame 103 is ignored because its VAD probability is too low.

[0076] Figure 6 The illustration shows a method for processing binaural audio signals received at S1, according to some embodiments. The audio signals of the binaural audio signals are a left audio signal L and a right audio signal R, or at least converted from alternative representations to left audio signal L and right audio signal R, and provided to S12 (optionally, provided to S12 via S11, as discussed below).

[0077] The left audio signal L and the right audio signal R are combined at S12 to form the middle audio signal M and the side audio signal S, which serve as alternative representations of the left audio signal L and the right audio signal R. The middle audio signal M is estimated by the sum of the left audio signal L and the right audio signal R. For example, the middle audio signal M can be estimated as:

[0078]

[0079] Similarly, the side audio signal S can be estimated by the difference between the left audio signal L and the right audio signal R. For example, the side audio signal S can be estimated as:

[0080]

[0081] Each or one of the estimated mid-audio signal M and side-audio signal S can constitute a first and / or a second audio signal and be processed according to the embodiments described in this disclosure. For example, it can be used... Figure 3 The processing sequences S2a and S2b shown process the side audio signal S and the mid-audio signal M separately. The audio processing of the side audio signal S may differ from that of the mid-audio signal M. In one embodiment, more aggressive noise reduction is employed in the processing of the side audio signal S at S2a compared to the processing of the mid-audio signal M at S2b. When the processed side audio signal PS and the processed mid-audio signal PM are recombined to form the processed binaural audio signal, more noise reduction is used in the side audio signal S to enhance signal quality, assuming a significant portion of the recording noise is present in the side audio signal S.

[0082] To recreate the processed versions of the original left audio signal L and right audio signal R—that is, the processed left audio signal PL and the processed right audio signal PR—the processed side audio signal PS and the processed mid audio signal PM can be recombined at S28 as a sum and difference to form the processed left audio signal PL and the processed right audio signal PR, respectively. For example, the processed left audio signal PL can be estimated as:

[0083]

[0084] The processed right audio signal PR can be estimated as follows:

[0085]

[0086] In some implementations, an additional audio signal from an additional recording device is received at S4. The additional audio signal is synchronized with the binaural audio signal and can be processed independently or in a manner coupled to the first and second audio signals (e.g., considered together with the first and second audio signals to provide a monophonic noise reduction model). For example, the processing of the additional audio signal can be equivalent to the processing of the first and second audio signals in the first processing sequence S2a and the second processing sequence S2b. A processed additional audio signal PA can be provided as supplementary information to the binaural output audio signal extracted at S28.

[0087] Alternatively, at S11, the additional audio signal is synchronized and mixed with the left audio signal L and the right audio signal R of the binaural audio signals. The additional audio signal A can be mixed with the left audio signal L and the right audio signal R separately with the same predetermined mixing ratio. For example, the mixing ratio of the additional audio signal A when mixed with the left audio signal L is 0.3, and the mixing ratio when mixed with the right audio signal R is 0.3. If it is determined that the additional audio signal A is likely to contain speech (e.g., by calculating the VAD probability), the predetermined mixing ratio can be increased by applying a mixing gain, such that, for example, the mixing ratio of the additional audio signal A when mixed with the left audio signal L is 0.7, and the mixing ratio when mixed with the right audio signal R is 0.7. The additional audio signal A can be preprocessed, for example, by noise reduction or VAD probability extraction before being mixed with the left audio signal L and the right audio signal R. The resulting binaural output audio signal obtained at S3 can help to more accurately recreate the audio of the second audio source of interest captured by the additional recording device.

[0088] In some implementations, the frequency responses of the binaural recording device and the additional recording device are obtained. The frequency response can be obtained by recording measurements representing the energy captured by each device for each frequency band. By comparing the frequency responses with equalization information associated with each device, which can be represented, for example, by an equalization curve, equalization information can be calculated and applied to at least one of the binaural audio signals (each of the first and second audio signals) and the additional audio signal. For example, the equalization information may include a gain for each frequency band, extracted by comparing the energy of each frequency band captured by the binaural recording device with the energy of each frequency band captured by the additional recording device.

[0089] Because binaural recording devices and auxiliary recording devices can have different frequency responses, the application of equalization information (e.g., equalization curves) allows for tonality matching between the binaural and auxiliary recording devices. Therefore, the mixing of audio sources captured by each recording is more uniform, which increases the intelligibility of the audio sources captured by the recording devices.

[0090] In some implementations, the mixing gain of the additional audio signal from S4 with the binaural audio signal and / or the binaural audio signal at S11 is adjusted based on the VAD probability. For example, the VAD probability of the additional audio signal can be extracted, and if the VAD probability indicates that the additional audio signal is likely to contain speech, a linear mixing gain greater than one can be applied to the additional audio signal when it is mixed with the binaural audio signals L and R at S11 to enhance, for example, the speech of a respondent near the additional recording device. Furthermore, if the VAD probability extracted for the mid-audio signal indicates that the mid-audio signal M is likely to contain speech, a linear gain greater than one can be applied to the mid-audio signal M at S28 to enhance, for example, the speech of a user wearing a binaural recording device.

[0091] Bone vibration sensor signal BV can be considered in the processing of binaural audio signals or in the processing of binaural audio signals and additional audio signals. Each processing sequence S2a, S2b can receive the bone vibration sensor signal BV according to the above description.

[0092] Alternatively or additionally, the bone vibration sensor signal BV can be used to determine the VAD probability or an enhanced VAD probability to manipulate the mixing of the binaural audio signal with the additional signal A at S11. For example, if the bone vibration sensor signal BV indicates that the user of the binaural recording device is unlikely to be speaking, a linear mixing gain greater than one can be applied at S11 to enhance the additional audio signal A. In some embodiments, the VAD estimated based on the bone vibration sensor signal BV is used to determine whether the speech originates from the user wearing the binaural recording device or from a second source of interest. For example, if the user of the binaural recording device wears a bone vibration sensor, and the VAD probability extracted from the bone vibration audio signal BV indicates that speech audio is likely to be present, then it is determined that the user wearing the binaural recording device is speaking. If the VAD probability extracted from the bone vibration audio signal BV indicates that speech audio is unlikely to be present, it can be determined that the user wearing the binaural recording device is not speaking. In response to determining that the user is not speaking, the additional audio signal and / or the lateral audio signal S is enhanced to highlight any audio from the surrounding environment, such as the interviewee speaking. In response to determining that the user is speaking, the mid-audio signal is enhanced to highlight the user's speech.

[0093] Alternatively, the additional audio signal is mixed with the same mixing ratio as the left audio signal L and the right audio signal R, the middle audio signal M can be extracted alone or mainly from the additional audio signal, and the side audio signal can be extracted alone or mainly from the left audio signal L and the right audio signal R.

[0094] In some implementations, the bone vibration sensor signal from the bone vibration sensor of the binaural recording device, together with the VAD probability of the extracted additional audio signal, is used to determine the source of the detected speech. For example, if the VAD of the additional audio signal is high, but the bone vibration sensor signal indicates little or no vibration, it can be determined that the source of the detected speech is not the wearer of the binaural recording device. Alternatively, if the VAD probability of the additional audio signal is high, and the bone vibration sensor signal indicates bone vibration associated with the speech, it can be determined that the source of the detected speech is the wearer of the binaural recording device.

[0095] Therefore, depending on the identified source of the detected speech, different noise reduction methods can be employed for the binaural audio signal and / or the additional audio signal. For example, when the speech source is the wearer of the additional recording device, a first noise reduction technique can be used, which is specifically designed to suppress the noise added to the channel between the wearer of the binaural recording device and the additional recording device. When the speech source is another source of interest, different noise reduction techniques are more suitable for reducing the noise in the channel between the other source of interest and the additional recording device.

[0096] Additionally or alternatively, depending on the detected source of the speech, the relative gain of the binaural audio signal and the additional audio signal can be modulated accordingly. For example, if the speech source is determined to be another source of interest, the gain of the additional audio signal relative to the binaural audio is increased. If the speech is determined to be from the wearer of the binaural audio signal, the gain of the additional audio signal relative to the binaural audio is decreased.

[0097] Figure 7 A flowchart illustrating a rendering method according to some embodiments is depicted. Besides playing binaural audio signals through headphones, using a speaker system (e.g., a HiFi system or surround sound system) or multiple speakers in a portable device is another common option. The portable device can be, for example, a tablet computer with four independent speakers, such as two top speakers and two bottom speakers, where each speaker is fed by a separate power amplifier. Therefore, a rendering method for rendering binaural audio signals to at least four speakers is provided.

[0098] In some implementations, the binaural audio signal comprises a pair of audio signals, such as a processed left audio signal PL and a right audio signal PR. The rendering of the binaural audio signal is based on two cascaded processes: applying translation information obtained at S205 and crosstalk cancellation information obtained at S210 to the binaural audio signal, and can generally be extended to render the binaural signal on an N-channel speaker system. Here, N is a natural number equal to or greater than four, and at least two speakers in the speaker system form a left and right speaker pair. The following N-channel rendering signal S can be obtained:

[0099]

[0100] Here, M is a translation matrix representing translation information of size N×2, and X is a crosstalk cancellation matrix of size N×N. The translation matrix indicates the amplitude ratio to be translated to the loudspeakers, and in some embodiments, the translation information indicates a centering translation of at least one left and right loudspeaker pair (equal row entries in the translation matrix M). Accordingly, binaural audio signals can be rendered on N-channel loudspeakers.

[0101] At S201, a binaural audio signal is obtained, and at S205, translation information (e.g., translation matrix M) is generated, indicating a centering translation of at least one left and right speaker pair of the speaker system.

[0102] In some implementations, in addition to the binaural audio signal with two audio signals (processed left audio signal PL and processed right audio signal PR) obtained at S201, a processed additional audio signal PA (derived from additional audio signal A recorded by an additional recording device) is also obtained at S202. The following N-channel rendering signal S can be obtained at S220:

[0103]

[0104] Wherein, M1 is a translation matrix (of size N×2) for the binaural audio signal, and M2 is a translation matrix (of size N×1) for the processed additional audio signal. The translation information represented by translation matrix M1 and translation information represented by translation matrix M2 can be set separately. For example, M1 can instruct a centering translation of at least one speaker pair, while M2 instructs a translation to all speakers. For instance, in a tablet computer with four speakers, M1 can instruct a translation to the top pair of speakers (to provide ambient audio), while M2 instructs a translation to all four speakers (to provide clear audio from a second source of interest). Accordingly, more intelligible speech from the binaural recording device and the additional recording device can be provided to the tablet computer user.

[0105] Parameters g1 and g2 indicate the corresponding mixing coefficients of the binaural audio signal and the additional audio signal, where the signal power level of the binaural audio signal relative to the additional audio signal is set.

[0106] The crosstalk cancellation matrix X1 represents the crosstalk cancellation information rendered to at least one pair of speakers by the binaural audio signal.

[0107] As described above, binaural audio signals accompanied by processed additional audio signals can be rendered onto an N-channel speaker system to more clearly recreate the speech of the user wearing binaural recording devices and a second source of interest (e.g., a respondent near the additional recording devices).

[0108] Therefore, the speaker system can render a binaural audio signal accompanied by an additional audio signal to highlight the audio from a second source of interest. By shifting the additional audio signal across all speakers, the additional audio signal can be clearly perceived, while simultaneously rendering the binaural signal on at least one speaker pair to provide ambient sound.

[0109] ******

[0110] In one embodiment, a system includes: one or more computer processors; and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operation of any one of the preceding method claims.

[0111] In one embodiment, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause one or more processors to perform the operations of any of the preceding method claims.

[0112] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via a communication unit, and / or installed from a removable medium.

[0113] Generally, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., a CPU combined with other components), thus enabling the control circuitry to perform the actions described in this disclosure. Some aspects can be implemented in hardware, while others can be implemented in firmware or software (e.g., control circuitry) that can be executed by a controller, microprocessor, or other computing device. Although various aspects of the exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, other computing devices, or some combination thereof, as non-limiting examples.

[0114] Furthermore, the various boxes shown in the flowchart can be viewed as method steps, and / or operations resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform associated functions. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code configured to perform the methods described above.

[0115] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0116] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that when executed by the processor of the computer or the processor of the other programmable data processing apparatus, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0117] While this document contains numerous details of specific implementation, these details should not be construed as limiting the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Specific features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features of a claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof. The logical flow depicted in the drawings does not require the specific order or ordered sequence shown to achieve the desired result. Additionally, other steps may be provided from the described flow, or steps may be removed, and other components may be added to or removed from the described system. Therefore, other implementations are within the scope of the appended claims.

Claims

1. A method for processing a first audio signal and a second audio signal representing input binaural audio signals acquired by a binaural recording device, the method comprising: Audio information is extracted from the first audio signal, wherein the audio information includes multiple frequency bands representing the first audio signal; Calculate the band gain for each band of the first audio signal to reduce noise in the first audio signal; For each frequency band of the first audio signal, calculate the corresponding Voice Activity Detection (VAD) probability; The frequency band gain is applied to the corresponding frequency band of the first audio signal according to the corresponding dynamic scaling factor to provide a first output audio signal, wherein: The value of the dynamic scaling factor is between zero and one, where a value of zero indicates that full-band gain is applied, and a value of one indicates that no band gain is applied. For each frequency band, the dynamic scaling factor is based on the frequency band gain associated with the corresponding frequency bands of the first audio signal whose VAD probability exceeds a predetermined VAD probability threshold in the current and previous time frames, and is not based on the frequency band gain associated with the corresponding frequency bands of the first audio signal whose VAD probability does not exceed the predetermined VAD probability threshold in the current and previous time frames. Providing the first output audio signal includes: The noise-reduced audio signal is calculated by applying the band gain to the corresponding band of the first audio signal, and Each frequency band of the first audio signal is mixed with the corresponding frequency band of the noise-reduced audio signal at a mixing ratio equal to the dynamic scaling factor to provide the first output audio signal; Perform noise reduction processing on the second audio signal to obtain a second output audio signal, and The binaural output audio signal is determined based on the first output audio signal and the second output audio signal.

2. The method according to claim 1, wherein, The noise reduction processing of the second audio signal includes separate processing steps corresponding to the processing steps of the first audio signal.

3. The method according to claim 1 or 2, wherein, Providing the first output audio signal includes: For each frequency band, the dynamic band gain is calculated as k + (1 - k)Bgain, where k is the dynamic scaling factor and Bgain is the calculated band gain. The dynamic band gain is applied to each frequency band of the first audio signal to provide the first output audio signal.

4. The method according to claim 1 or 2, wherein, The dynamic scaling factor for each frequency band is based on the frequency band gain exceeding a predetermined threshold gain for the corresponding frequency bands of the current time frame and the previous time frame.

5. The method according to claim 1 or 2, wherein, The dynamic scaling factor is based on a weighted sum of frequency band gains, the weighted sum including frequency band gains from previous time frames, and the method further includes: Determine that the band gain of a specific frequency band in the current time frame exceeds a predetermined threshold gain; If the frequency band gain associated with the specific frequency band of the current frame exceeds the predetermined threshold gain; The current weighted sum is calculated as a weighted sum of the frequency band gains of the current time frame, and the weighted sum includes the frequency band gains from previous time frames. If the frequency band gain associated with the specific frequency band of the current frame is lower than the predetermined threshold gain; The current weighted sum is calculated as a weighted sum that includes the band gain from previous time frames.

6. The method according to claim 1 or 2, wherein, The dynamic scaling factor is determined to be 1 - G, where G is a weighted sum of band gains, including at least the band gains from the bands of previous time frames.

7. The method according to claim 1 or 2, wherein, The determination of the dynamic scaling factor for each frequency band is performed offline, and each dynamic scaling factor is based on the frequency band gain associated with the corresponding frequency band for all time frames of the first audio signal.

8. The method of claim 7, further comprising: The dynamic scaling factor for each frequency band of the first audio signal is determined based on the average frequency band gain from all frames, where: The bandwidth gain exceeds a predetermined threshold gain, and the VAD probability exceeds a predetermined probability threshold.

9. The method according to claim 1 or 2, wherein, The two audio signals are a left channel audio signal and a right channel audio signal, and the method further includes: It is estimated that the first audio signal is a middle channel audio signal, and the middle signal is calculated from the sum of the left and right signals; The second audio signal is estimated to be a side channel audio signal, which is calculated from the difference between the left signal and the right signal; and The binaural output audio signal is determined by the following steps: The estimated left output audio signal is the sum of the middle output signal and the side output signal; and The right output audio signal is estimated to be the difference between the middle output signal and the side output signal.

10. The method of claim 1 or 2, further comprising processing additional audio signals from an additional recording device, wherein the first audio signal and the second audio signal are a left audio signal and a right audio signal, the method further comprising: Synchronize the additional audio signal with the binaural audio signal; as well as The additional audio signal is mixed with the left audio signal and the right audio signal.

11. The method of claim 10, further comprising processing the bone vibration sensor signal acquired by the bone vibration sensor, the method further comprising... Synchronize the bone vibration sensor signal with the binaural audio signal; and The gain of the additional audio signal is controlled based on the bone vibration sensor signal.

12. The method of claim 11, further comprising processing bone vibration sensor signals acquired by the bone vibration sensor of the binaural recording device, the method further comprising: Synchronize the bone vibration sensor signal with the binaural audio signal; Extract the VAD probability of the additional audio signal; The source of the detected speech is determined based on the VAD probability and the bone vibration sensor signal; If the source is a wearer of the binaural recording device with the bone vibration sensor, the additional audio signal is processed using a first audio processing scheme suitable for suppressing noise in the vocal tract between the wearer of the binaural recording device and the additional recording device. If the source is not the wearer of the binaural recording device with the bone vibration sensor, the additional audio signal is processed using a second audio processing scheme suitable for suppressing noise in the channel between the other source and the additional recording device.

13. The method according to claim 12, wherein, The first audio processing scheme and the second audio processing scheme apply different signal gains to the additional audio signal.

14. The method according to claim 1 or 2, wherein, The audio information also includes one or more of the following: The SNR of the first audio signal; The fundamental frequency of the first audio signal; The VAD probability of the first audio signal; Bone vibration sensor signal acquired by the bone vibration sensor; The fundamental frequency extracted from the bone vibration sensor signal acquired by the bone vibration sensor; and VAD probability extracted from bone vibration sensor signals acquired by bone vibration sensors.

15. The method of claim 14, further comprising: The gain of the first audio signal is controlled based on the VAD probability extracted from the bone vibration sensor signal.

16. The method according to claim 1 or 2, wherein, Calculating the band gain for each frequency band in the first audio signal involves predicting the band gain from the audio information using a trained neural network.

17. A method for processing a first audio signal and a second audio signal representing input binaural audio signals acquired by a binaural recording device and an additional audio signal from an additional recording device other than the binaural recording device, wherein the first and second input audio signals and the first and second output audio signals are left and right input audio signals and left and right output audio signals, respectively, the method comprising: The additional audio signal is synchronized with the binaural audio signal, wherein the binaural audio signal is processed by the method according to any one of claims 1 to 16; Receive bone vibration sensor signals acquired by the bone vibration sensor of the binaural recording device; Synchronize the bone vibration sensor signal with the binaural audio signal; Extract the VAD probability of the additional audio signal; The source of the detected speech is determined based on the VAD probability and the bone vibration sensor signal; If the source is a wearer of the binaural recording device with the bone vibration sensor, the gain of the additional audio signal from the additional recording device relative to the binaural audio signal from the binaural recording device is reduced to highlight the binaural audio signal from the binaural recording device. If the source is not the wearer of the binaural recording device with the bone vibration sensor, the gain of the additional audio signal from the additional recording device is increased relative to the binaural audio signal from the binaural recording device to highlight the additional audio signal from the additional recording device. Provides an additional output audio signal based on the processed additional audio signal; The additional output audio signal is mixed with the left and right audio signals to obtain the left and right output audio signals that form a binaural audio signal.

18. A computer program product comprising computer program code that, when executed on a computer, performs the method according to any one of the preceding claims.

19. An audio processing device, comprising: The receiver is configured to receive input binaural audio signals acquired by the binaural recording device, the input binaural audio signals including a first audio signal and a second audio signal; An extraction unit is configured to receive the first audio signal from the receiver and extract audio information from the first audio signal, the audio information including multiple frequency bands representing the first audio signal; The processing device is configured to receive the audio information and, for each frequency band of the first audio signal, calculate a frequency band gain for reducing noise in the first audio signal and a corresponding voice activity detection (VAD) probability. The application unit is configured to apply the band gain to a corresponding band of the first audio signal according to a dynamic scaling factor to provide a first output audio signal, wherein: The value of the dynamic scaling factor is between zero and one, where a value of zero indicates that full-band gain is applied, and a value of one indicates that no band gain is applied. For each frequency band, the dynamic scaling factor is based on the frequency band gain associated with the corresponding frequency bands of the first audio signal whose VAD probability exceeds a predetermined VAD probability threshold in the current and previous time frames, and is not based on the frequency band gain associated with the corresponding frequency bands of the first audio signal whose VAD probability does not exceed the predetermined VAD probability threshold in the current and previous time frames. Providing the first output audio signal includes: The noise-reduced audio signal is calculated by applying the band gain to the corresponding band of the first audio signal, and Each frequency band of the first audio signal is mixed with the corresponding frequency band of the noise-reduced audio signal at a mixing ratio equal to the dynamic scaling factor to provide the first output audio signal; An additional processing module is configured to perform noise reduction processing on the second audio signal to obtain a second output audio signal; and The output stage is configured to determine the binaural output audio signal based on the first output audio signal and the second output audio signal.

Citation Information

Patent Citations

  • Systems and methods of performing gain control

    CN104956437A

  • An audio signal processing apparatus for processing an input earpiece audio signal upon the basis of a microphone audio signal

    CN107533849A

  • Efficient noise reduction earphone with low power consumption and noise reduction system

    CN109195042A