A method and device for stereo audio signal delay estimation

By selecting different ITD estimation algorithms according to the noise type in the audio encoding device and using different weighting functions to process the frequency domain mutual power spectrum, the problem of low ITD estimation accuracy in the prior art is solved, and higher ITD estimation accuracy and more stable sound quality are achieved.

CN113948098BActive Publication Date: 2025-06-10HUAWEI TECH CO LTD

Patent Information

Application Number
CN202010700806.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-17
Publication Date
2025-06-10
Estimated Expiration
2040-07-17

AI Technical Summary

Technical Problem

The existing ITD estimation method has deteriorated performance in noisy environments, resulting in low ITD estimation accuracy of stereo audio signals, affecting the encoded sound quality.

Method used

By using an ITD estimation algorithm of different types of noise in the audio encoding device, specifically, if the noise signal of the current frame is correlation noise, a first algorithm is used; if it is diffuse noise, a second algorithm is used. The first algorithm and the second algorithm weight the frequency domain mutual power spectrum through different weighting functions respectively to improve the accuracy and stability of ITD estimation.

Benefits of technology

It greatly improves the ITD estimation accuracy and stability of stereo audio signals under diffuse noise and correlation noise conditions, reduces inter-frame discontinuity, maintains the phase of the stereo signal, improves the encoding sound image accuracy and stability, and enhances the sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113948098B_ABST
    Figure CN113948098B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for estimating the time delay of a stereophonic audio signal. The method may include: obtaining a current frame of the stereophonic audio signal, the current frame including a first-channel audio signal and a second-channel audio signal; if the signal type of the noise signal included in the current frame is a correlated noise signal type, estimating the inter-channel time difference of the current frame by using a first algorithm; if the signal type of the noise signal included in the current frame is a diffuse noise signal type, estimating the inter-channel time difference of the current frame by using a second algorithm; wherein, the first algorithm includes weighting the frequency-domain cross-power spectrum of the current frame by using a first weighting function, the second algorithm includes weighting the frequency-domain cross-power spectrum of the current frame by using a second weighting function, and the construction factors of the first weighting function and the second weighting function are different. In the present application, by using different ITD estimation algorithms for stereophonic audio signals containing different types of noise, the estimation accuracy of the ITD of the stereophonic audio signal is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio encoding and decoding, and particularly to a method and apparatus for estimating the time delay of a stereo audio signal. Background Art

[0002] In daily audio-visual communication systems, people not only pursue high-quality images but also high-quality audio. In voice and audio communication systems, single-channel audio can no longer meet people's needs. Stereo audio carries the position information of each sound source, improving the clarity, intelligibility, and realism of the audio, so it is increasingly favored by people.

[0003] In stereo audio encoding and decoding technology, parametric stereo encoding and decoding technology is a common audio encoding and decoding technology. Commonly used spatial parameters include inter-channel coherence (IC), inter-channel level difference (ILD), inter-channel time difference (ITD), inter-channel phase difference (IPD), etc. Among them, ILD and ITD contain the position information of the sound source. Accurately estimating the ILD and ITD information is crucial for the reconstruction of the encoded stereo sound image and sound field.

[0004] Currently, the most commonly used type of ITD estimation method is the generalized cross-correlation method because this type of algorithm has low complexity, good real-time performance, is easy to implement, and does not rely on other prior information of the stereo audio signal. However, in a noisy environment, the performance of several existing generalized cross-correlation algorithms degrades severely, resulting in low ITD estimation accuracy for stereo audio signals, causing problems such as inaccurate, unstable sound images, poor sense of space, and obvious head-in-the-middle effect in the decoded stereo audio signal in parametric encoding and decoding technology, seriously affecting the sound quality of the encoded stereo audio signal. Summary of the Invention

[0005] This application provides a method and apparatus for estimating the time delay of a stereo audio signal to improve the estimation accuracy of the inter-channel time difference of the stereo audio signal, thereby improving the accuracy and stability of the sound image of the decoded stereo audio signal and improving the sound quality.

[0006] In a first aspect, the present application provides a method for estimating the time delay of a stereo audio signal. This method can be applied to an audio coding device, which can be used in the audio coding part of an audio-video communication system involving stereo and multi-channel audio, or in the audio coding part of a virtual reality (VR) application program. The method may include: the audio coding device obtains the current frame of the stereo audio signal, and the current frame includes the first-channel audio signal and the second-channel audio signal; if the signal type of the noise signal included in the current frame is a correlated noise signal type, then a first algorithm is used to estimate the inter-channel time difference (ITD) between the first-channel audio signal and the second-channel audio signal; if the signal type of the noise signal included in the current frame is a diffuse noise signal type, then a second algorithm is used to estimate the ITD between the first-channel audio signal and the second-channel audio signal; wherein, the first algorithm includes weighting the cross-power spectrum in the frequency domain of the current frame using a first weighting function, and the second algorithm includes weighting the cross-power spectrum in the frequency domain of the current frame using a second weighting function, and the construction factors of the first weighting function and the second weighting function are different.

[0007] The above stereo audio signal can be an original stereo audio signal (including a left-channel audio signal and a right-channel audio signal), or a stereo audio signal composed of two audio signals in a multi-channel audio signal, or a stereo signal composed of two audio signals jointly generated by multiple audio signals in a multi-channel audio signal. Of course, there may be other forms of stereo audio signals, and the embodiments of the present application do not make specific limitations.

[0008] Optionally, the above audio coding device may specifically be a stereo coding device, which can constitute an independent stereo encoder; or it can be the core coding part of a multi-channel encoder, aiming to encode a stereo audio signal composed of two audio signals jointly generated by multiple signals in a multi-channel audio signal.

[0009] In some possible implementation manners, the current frame in the stereo signal obtained by the audio coding device can be a frequency-domain audio signal or a time-domain audio signal. If the current frame is a frequency-domain audio signal, the audio coding device can directly process the current frame in the frequency domain; if the current frame is a time-domain audio signal, the audio coding device can first perform a time-frequency transformation on the current frame in the time domain to obtain the current frame in the frequency domain, and then process the current frame in the frequency domain.

[0010] In this application, the audio encoding device adopts different ITD estimation algorithms for stereo audio signals containing different types of noise, greatly improving the accuracy and stability of ITD estimation for stereo audio signals under diffusive noise and correlated noise conditions, reducing the inter-frame discontinuity between downmixed stereo signals, while better maintaining the phase of the stereo signal. The sound image of the encoded stereo is more accurate and stable, with a stronger sense of realism, improving the auditory quality of the encoded stereo signal.

[0011] In some possible implementation manners, after obtaining the current frame of the stereo audio signal, the above method further includes: obtaining the noise coherence value of the current frame; if the noise coherence value is greater than or equal to a preset threshold, determining that the signal type of the noise signal included in the current frame is a correlated noise signal type; if the noise coherence value is less than the preset threshold, determining that the signal type of the noise signal included in the current frame is a diffusive noise signal type.

[0012] Optionally, the above preset threshold is an empirical value and can be set to, for example, 0.20, 0.25, 0.30.

[0013] In some possible implementation manners, obtaining the noise coherence value of the current frame may include: performing voice activity detection on the current frame; if the detection result indicates that the signal type of the current frame is a noise signal type, calculating the noise coherence value of the current frame; or, if the detection result indicates that the signal type of the current frame is a voice signal type, determining the noise coherence value of the previous frame of the current frame in the stereo audio signal as the noise coherence value of the current frame.

[0014] Optionally, the audio encoding device may calculate the value of voice activity detection in a time domain, a frequency domain, or a combination of time domain and frequency domain, and no specific limitation is made thereto.

[0015] In this application, after the audio encoding device calculates the noise coherence value of the current frame, it may further perform smoothing processing on it to reduce the error of noise coherence value estimation and improve the recognition accuracy of the noise type.

[0016] In some possible embodiments, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; estimating the inter-channel time difference between the first-channel audio signal and the second-channel audio signal by using the first algorithm includes: performing time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weighting the frequency-domain cross-power spectrum by using a first weighting function; obtaining an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, the amplitude weighting parameter, and the squared coherence value of the current frame.

[0017] In some possible embodiments, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; estimating the inter-channel time difference between the first-channel audio signal and the second-channel audio signal by using the first algorithm includes: calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weighting the frequency-domain cross-power spectrum by using a first weighting function; obtaining an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, the amplitude weighting parameter, and the squared coherence value of the current frame.

[0018] In some possible embodiments, the first weighting function Φ new_1 (k) satisfies the following formula:

[0019]

[0020] wherein, β is the amplitude weighting parameter, W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; Γ 2 (k) is the squared coherence value of the k-th frequency point of the current frame, X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), k is the frequency point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

[0021] In some possible embodiments, the first weighting function Φ new_1 (k) satisfies the following formula:

[0022]

[0023] where β is the amplitude weighting parameter, W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; Γ 2 (k) is the coherence squared value of the k-th frequency bin in the current frame, X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), k is the frequency bin index value, k = 0, 1, …, N DFT -1, and N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

[0024] Optionally, β ∈ [0, 1]. For example, β = 0.6, 0.7, 0.8, etc.

[0025] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal may be the first initial Wiener gain factor and / or the first improved Wiener gain factor of the first-channel frequency-domain signal; the Wiener gain factor corresponding to the second-channel frequency-domain signal may be the second initial Wiener gain factor and / or the second improved Wiener gain factor of the second-channel frequency-domain signal.

[0026] For example, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; then, after obtaining the current frame in the stereophonic audio signal, the above method further includes: obtaining an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determining the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtaining an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; determining the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

[0027] In this application, after being weighted by the Wiener gain factor, the weight of the correlated noise component in the cross-power spectrum of the stereo audio signal in the frequency domain is significantly reduced, and the correlation of the residual noise component is also greatly reduced. In most cases, the coherence squared value of the residual noise is much smaller than that of the target signal (such as a speech signal) in the stereo audio signal, so that the cross-correlation peak corresponding to the target signal will be more prominent, and the accuracy and stability of the ITD estimation of the stereo audio signal will be significantly improved.

[0028] In some possible implementation manners, the above first initial Wiener gain factor satisfies the following formula:

[0029]

[0030] The above second initial Wiener gain factor satisfies the following formula:

[0031]

[0032] Wherein, is the estimated value of the noise power spectrum of the first channel, is the estimated value of the noise power spectrum of the second channel; X 1 (k) is the frequency-domain signal of the first channel, X 2 (k) is the frequency-domain signal of the second channel, k is the frequency point index value, k = 0, 1,..., N DFT -1, N DFT is the total number of frequency points after the current frame is subjected to time-frequency transformation.

[0033] For another example, the Wiener gain factor corresponding to the frequency-domain signal of the first channel is the first improved Wiener gain factor of the frequency-domain signal of the first channel, and the Wiener gain factor corresponding to the frequency-domain signal of the second channel is the second improved Wiener gain factor of the frequency-domain signal of the second channel;

[0034] After obtaining the current frame in the stereo audio signal, the above method further includes: obtaining the above first initial Wiener gain factor and the above second Wiener gain factor; constructing a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; constructing a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

[0035] In this application, by constructing a binary masking function for the first initial Wiener gain factor corresponding to the frequency-domain signal of the first channel and the second initial Wiener gain factor corresponding to the frequency-domain signal of the second channel, the frequency points less affected by noise are screened out to improve the accuracy of the ITD estimation.

[0036] In some possible implementation manners, the above first improved Wiener gain factor Satisfy the following formula:

[0037]

[0038]

[0039] Wherein, μ 0 is the binary masking threshold of the Wiener gain factor, is the first initial Wiener gain factor; is the second initial Wiener gain factor.

[0040] Optionally, μ 0 ∈[0.5, 0.8]. For example, μ 0 = 0.5, 0.66, 0.75, 0.8, etc.

[0041] In some possible implementation manners, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; estimating the inter-channel time difference between the first-channel frequency-domain signal and the second-channel frequency-domain signal by using a second algorithm includes: performing time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain the first-channel frequency-domain signal and the second-channel frequency-domain signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weighting the frequency-domain cross-power spectrum by using a second weighting function to obtain an estimated value of the inter-channel time difference between the first-channel frequency-domain signal and the second-channel frequency-domain signal; wherein, the construction factors of the second weighting function include: an amplitude weighting parameter and the coherence square value of the current frame.

[0042] In some possible implementation manners, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; estimating the inter-channel time difference between the first-channel audio signal and the second-channel audio signal by using a second algorithm includes: calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weighting the frequency-domain cross-power spectrum by using a second weighting function; obtaining an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the second weighting function include: an amplitude weighting parameter and the coherence square value of the current frame.

[0043] In some possible implementation manners, the second weighting function Φ new_2 (k) satisfies the following formula:

[0044]

[0045] Wherein, β is the amplitude value weighting parameter, Γ 2 (k) is the coherence square value of the k-th frequency point of the current frame, X 1(k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), where k is the frequency-point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

[0046] Optionally, β ∈ [0, 1]. For example, β = 0.6, 0.7, 0.8, etc.

[0047] In a second aspect, the present application provides a method for estimating the time delay of a stereo audio signal. This method can be applied to an audio coding device, which can be used for the audio coding part in an audio-visual communication system involving stereo and multi-channel audio, or for the audio coding part in a VR application program. The method may include: the current frame includes a first-channel audio signal and a second-channel audio signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel audio signal and the second-channel audio signal; weighting the frequency-domain cross-power spectrum using a preset weighting function; obtaining an estimated value of the inter-channel time difference between the first-channel frequency-domain signal and the second-channel frequency-domain signal according to the weighted frequency-domain cross-power spectrum.

[0048] Among them, the preset weighting function includes a first weighting function or a second weighting function, and the construction factors of the first weighting function and the second weighting function are different.

[0049] Optionally, the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain corresponding to the second-channel frequency-domain signal, the amplitude weighting parameter, and the coherence square value of the current frame; the construction factors of the second weighting function include: the amplitude weighting parameter and the coherence square value of the current frame.

[0050] In some possible implementation manners, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel audio signal and the second-channel audio signal includes: performing time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal.

[0051] In some possible implementation manners, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal.

[0052] In some possible implementation manners, the first weighting function Φ new_1 (k) satisfies the following formula:

[0053]

[0054] Among them, β is the amplitude weighting parameter, and W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; Γ 2 (k) is the coherence square value of the k-th frequency point in the current frame, X 1 (k) is the first-channel frequency-domain signal, and X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), k is the frequency point index value, k = 0, 1,..., N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

[0055] In some possible implementation manners, the first weighting function Φ new_1 (k) satisfies the following formula:

[0056]

[0057] Among them, β is the amplitude weighting parameter, and W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; Γ 2 (k) is the coherence square value of the k-th frequency point in the current frame, X 1 (k) is the first-channel frequency-domain signal, and X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), k is the frequency point index value, k = 0, 1,..., N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

[0058] Optionally, β ∈ [0, 1]. For example, β = 0.6, 0.7, 0.8, etc.

[0059] In some possible implementation manners, the Wiener gain factor corresponding to the first-channel frequency-domain signal may be the first initial Wiener gain factor and / or the first improved Wiener gain factor of the first-channel frequency-domain signal; the Wiener gain factor corresponding to the second-channel frequency-domain signal may be the second initial Wiener gain factor and / or the second improved Wiener gain factor of the second-channel frequency-domain signal.

[0060] For example, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; after obtaining the current frame in the stereophonic audio signal, the above method further includes: obtaining an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determining the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtaining an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; determining the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

[0061] In some possible implementation manners, the first initial Wiener gain factor satisfies the following formula:

[0062]

[0063] The second initial Wiener gain factor satisfies the following formula:

[0064]

[0065] Wherein, is the estimated value of the first-channel noise power spectrum, is the estimated value of the second-channel noise power spectrum; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, k is the frequency point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after the current frame is subjected to time-frequency transformation.

[0066] For another example, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; after obtaining the current frame in the stereophonic audio signal, the above method further includes: obtaining the above first initial Wiener gain factor and the above second initial Wiener gain factor; constructing a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; constructing a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

[0067] In some possible implementation manners, the first improved Wiener gain factor satisfies the following formula:

[0068]

[0069]

[0070] Among them, μ 0 is the binary masking threshold of the Wiener gain factor, is the first Wiener gain factor; is the second Wiener gain factor.

[0071] Optionally, μ 0 ∈ [0.5, 0.8]. For example, μ 0 = 0.5, 0.66, 0.75, 0.8, etc.

[0072] In some possible implementation manners, the second weighting function Φ new_2 (k) satisfies the following formula:

[0073]

[0074] Among them, β is the amplitude weighting parameter, and Γ 2 (k) is the coherence square value of the k-th frequency point of the current frame, X 1 (k) is the frequency-domain signal of the first channel, and X 2 (k) is the frequency-domain signal of the second channel, is the conjugate function of X 2 (k), k is the frequency point index value, k = 0, 1,..., N DFT - 1, and N DFT is the total number of frequency points after the time-frequency transformation of the current frame.

[0075] Optionally, β ∈ [0, 1]. For example, β = 0.6, 0.7, 0.8, etc.

[0076] In a third aspect, the present application provides a stereo audio signal delay estimation device, which may be a chip or a system-on-chip in an audio coding device, or may also be a functional module in the audio coding device for implementing the method described in the first aspect or any possible implementation manner of the first aspect. For example, the stereo audio signal delay estimation device includes: a first acquisition module, configured to acquire a current frame of the stereo audio signal, where the current frame includes a first-channel audio signal and a second-channel audio signal; a first inter-channel time difference estimation module, configured to, if the signal type of the noise signal included in the current frame is a correlated noise signal type, estimate the inter-channel time difference between the first-channel audio signal and the second-channel audio signal by using a first algorithm; if the signal type of the noise signal included in the current frame is a diffuse noise signal type, estimate the inter-channel time difference between the first-channel audio signal and the second-channel audio signal by using a second algorithm; where the first algorithm includes weighting the cross-power spectrum in the frequency domain of the current frame by using a first weighting function, and the second algorithm includes weighting the cross-power spectrum in the frequency domain of the current frame by using a second weighting function, and the construction factors of the first weighting function and the second weighting function are different.

[0077] In some possible implementation manners, the above device further includes: a noise coherence value calculation module, configured to, after the first acquisition module acquires the current frame, acquire the noise coherence value of the current frame; if the noise coherence value is greater than or equal to a preset threshold, determine that the signal type of the noise signal included in the current frame is a correlated noise signal type; or, if the noise coherence value is less than the preset threshold, determine that the signal type of the noise signal included in the current frame is a diffuse noise signal type.

[0078] In some possible implementation manners, the above device further includes: a voice activity detection module, configured to perform voice activity detection on the current frame to obtain a detection result; the noise coherence value calculation module is specifically configured to, if the detection result indicates that the signal type of the current frame is a noise signal type, calculate the noise coherence value of the current frame; or, if the detection result indicates that the signal type of the current frame is a voice signal type, determine the noise coherence value of the previous frame of the current frame in the stereo audio signal as the noise coherence value of the current frame.

[0079] In the present application, the voice activity detection module may calculate the voice activity detection value in a time domain, a frequency domain, or a combination of the time domain and the frequency domain, and no specific limitation is made thereto.

[0080] In some possible embodiments, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the first inter-channel time difference estimation module is configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using a first weighting function; obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and the squared coherence value of the current frame.

[0081] In some possible embodiments, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; the first inter-channel time difference estimation module is configured to calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using a first weighting function; obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and the squared coherence value of the current frame.

[0082] In some possible embodiments, the first weighting function Φ new_1 (k) satisfies the following formula:

[0083]

[0084] wherein, β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the squared coherence value of the k-th frequency point of the current frame, k is the frequency point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

[0085] In some possible embodiments, the first weighting function Φ new_1 (k) satisfies the following formula:

[0086]

[0087] wherein, β is an amplitude weighting parameter, β ∈ [0, 1], and W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, and X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), and Γ 2 (k) is the squared coherence value of the k-th frequency bin in the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT -1, and N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

[0088] In some possible implementation manners, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; the first inter-channel time difference estimation module is specifically configured to, after the first obtaining module obtains the current frame, obtain an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determine the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtain an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; and determine the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

[0089] In some possible implementation manners, the first initial Wiener gain factor satisfies the following formula:

[0090]

[0091] The second initial Wiener gain factor satisfies the following formula:

[0092]

[0093] wherein, is the estimated value of the first-channel noise power spectrum, is the estimated value of the second-channel noise power spectrum; X 1 (k) is the first-channel frequency-domain signal, and X 2 (k) is the second-channel frequency-domain signal, k is the frequency bin index value, k = 0, 1, …, N DFT -1, and N DFTis the total number of frequency points after the current frame undergoes time-frequency transformation.

[0094] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; the first inter-channel time difference estimation module is specifically configured to, after the first acquisition module acquires the current frame, construct a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; and construct a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

[0095] In some possible embodiments, the first improved Wiener gain factor satisfies the following formula:

[0096]

[0097]

[0098] where μ 0 is the binary masking threshold of the Wiener gain factor, is the first initial Wiener gain factor; is the second initial Wiener gain factor.

[0099] In some possible embodiments, the first-channel audio signal is the first-channel time-domain signal, and the second-channel audio signal is the second-channel time-domain signal; the first inter-channel time difference estimation module is specifically configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain the first-channel frequency-domain signal and the second-channel frequency-domain signal; calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum using a second weighting function to obtain an estimated value of the inter-channel time difference; where the construction factors of the second weighting function include: an amplitude weighting parameter and the squared coherence value of the current frame.

[0100] In some possible embodiments, the first-channel audio signal is the first-channel frequency-domain signal, and the second-channel audio signal is the second-channel frequency-domain signal; the first inter-channel time difference estimation module is specifically configured to calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum using a second weighting function; and obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; where the construction factors of the second weighting function include: an amplitude weighting parameter and the squared coherence value of the current frame.

[0101] In some possible embodiments, the second weighting function Φ new_2 (k) satisfies the following formula:

[0102]

[0103] where β is the amplitude weighting parameter, β ∈ [0, 1], and X 1 (k) is the first-channel frequency-domain signal, and X 2 (k) is the second-channel frequency-domain signal. is the conjugate function of X 2 (k), and Γ 2 (k) is the coherence squared value of the k-th frequency bin in the current frame. k is the frequency bin index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

[0104] In a fourth aspect, the present application provides a stereo audio signal time delay estimation device, which can be a chip or a system-on-chip in an audio coding device, or can also be a functional module in the audio coding device for implementing the method described in the second aspect or any possible implementation manner of the second aspect. For example, the stereo audio signal time delay estimation device includes: a second acquisition module for acquiring the current frame in the stereo audio signal, where the current frame includes a first-channel audio signal and a second-channel audio signal; a second inter-channel time difference estimation module for calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel audio signal and the second-channel audio signal; weighting the frequency-domain cross-power spectrum using a preset weighting function; and obtaining an estimated value of the inter-channel time difference between the first-channel frequency-domain signal and the second-channel frequency-domain signal according to the weighted frequency-domain cross-power spectrum; where the preset weighting function is a first weighting function or a second weighting function, and the construction factors of the first weighting function and the second weighting function are different; the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain corresponding to the second-channel frequency-domain signal, the amplitude weighting parameter, and the coherence squared value of the current frame; the construction factors of the second weighting function include: the amplitude weighting parameter and the coherence squared value of the current frame.

[0105] In some possible implementation manners, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the second inter-channel time difference estimation module is configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; and calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal.

[0106] In some possible implementation manners, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal.

[0107] In some possible embodiments, the first weighting function Φ new_1 (k) satisfies the following formula:

[0108]

[0109] where β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the squared coherence value of the k-th frequency bin in the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT -1, and N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

[0110] In some possible embodiments, the first weighting function Φ new_1 (k) satisfies the following formula:

[0111]

[0112] where β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the squared coherence value of the k-th frequency bin in the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT -1, and N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

[0113] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; the second inter-channel time difference estimation module is specifically configured to, after the second acquisition module acquires the current frame, obtain an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determine the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtain an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; and determine the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

[0114] In some possible embodiments, the first initial Wiener gain factor satisfies the following formula:

[0115]

[0116] The second initial Wiener gain factor satisfies the following formula:

[0117]

[0118] where is the estimated value of the first-channel noise power spectrum, is the estimated value of the second-channel noise power spectrum; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, k is the frequency point index value, k = 0, 1,..., N DFT -1, N DFT is the total number of frequency points after the current frame is subjected to time-frequency transformation.

[0119] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; the second inter-channel time difference estimation module is specifically configured to, after the second acquisition module acquires the current frame, obtain the above-mentioned first initial Wiener gain factor and second initial Wiener gain factor; construct a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; and construct a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

[0120] In some possible embodiments, the first improved Wiener gain factor satisfies the following formula:

[0121]

[0122] The second improved Wiener gain factor Satisfies the following formula:

[0123]

[0124] Where μ 0 Is the binary masking threshold of the Wiener gain factor, Is the first initial Wiener gain factor; Is the second initial Wiener gain factor.

[0125] In some possible implementation manners, the second weighting function Φ new_2 (k) satisfies the following formula:

[0126]

[0127] Where β ∈ [0, 1], X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, Is the conjugate function of X 2 (k), Γ 2 (k) is the coherence square value of the k-th frequency point in the current frame, k is the frequency point index value, k = 0, 1, …, N DFT -1, N DFT Is the total number of frequency points after time-frequency transformation of the current frame.

[0128] In a fifth aspect, the present application provides an audio encoding apparatus, including: a non-volatile memory and a processor that are coupled to each other, and the processor calls program code stored in the memory to execute the stereophonic audio signal delay estimation method described in the first to second aspects and any one of them above.

[0129] In a sixth aspect, the present application provides a computer-readable storage medium, and the computer-readable storage medium stores instructions that, when run on a computer, are used to execute the stereophonic audio signal delay estimation method described in the first to second aspects and any one of them above.

[0130] In a seventh aspect, the present application provides a computer-readable storage medium, including an encoded bitstream, and the encoded bitstream includes the inter-channel time difference of the stereophonic audio signal obtained according to the stereophonic audio signal delay estimation method described in the first to second aspects and any one of their possible implementation manners above.

[0131] In an eighth aspect, the present application provides a computer program or a computer program product, and when the computer program or the computer program product is executed on a computer, the computer is enabled to implement the stereophonic audio signal delay estimation method described in the first to second aspects and any one of them above.

[0132] It should be understood that the technical solutions of the fourth to tenth aspects of this application are consistent with those of the first to second aspects of this application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar, so they will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0133] In order to more clearly illustrate the technical solutions in the embodiments of this application, the drawings required for use in the embodiments of this application or the background art will be described below.

[0134] Figure 1 It is a schematic flowchart of the parametric stereo encoding and decoding method in the frequency domain in the embodiments of this application;

[0135] Figure 2 It is a schematic flowchart of the generalized cross-correlation algorithm in the embodiments of this application;

[0136] Figure 3 It is a schematic flowchart of the method for estimating the delay of a stereophonic audio signal in the embodiments of this application Figure 1 ;

[0137] Figure 4 It is a schematic flowchart of the method for estimating the delay of a stereophonic audio signal in the embodiments of this application Figure 2 ;

[0138] Figure 5 It is a schematic flowchart of the method for estimating the delay of a stereophonic audio signal in the embodiments of this application Figure 3 ;

[0139] Figure 6 It is a schematic structural diagram of the device for estimating the delay of a stereophonic audio signal in the embodiments of this application;

[0140] Figure 7 It is a schematic structural diagram of the audio encoding device in the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0141] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. In the following description, reference is made to the accompanying drawings that form a part of the present application and show, by way of illustration, specific aspects of the embodiments of the present application or aspects in which the embodiments of the present application may be used. It should be understood that the embodiments of the present application may be used in other aspects and may include structural or logical changes not depicted in the drawings. For example, it should be understood that the disclosure of the described method may equally apply to the corresponding device or system for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units, such as functional units, to perform the one or more described method steps (e.g., one unit performs one or more steps, or multiple units, where each performs one or more of the multiple steps), even if such one or more units are not explicitly depicted or illustrated in the drawings. On the other hand, for example, if a specific device is described based on one or more units such as functional units, the corresponding method may include a step to perform the functionality of the one or more units (e.g., one step performs the functionality of the one or more units, or multiple steps, where each performs the functionality of one or more of the multiple units), even if such one or more steps are not explicitly depicted or illustrated in the drawings. Further, it should be understood that, unless otherwise explicitly stated, the features of the various exemplary embodiments and / or aspects described herein may be combined with each other.

[0142] In voice and audio communication systems, single-channel audio is increasingly unable to meet people's needs. Stereo audio, on the other hand, carries the location information of each sound source, improving the clarity, intelligibility, and realism of the audio. Therefore, it is increasingly favored by people.

[0143] In voice and audio communication systems, audio codec technology is a very crucial technology. Based on the auditory model, this technology uses the minimum energy perceptual distortion to express audio signals at as low a coding rate as possible for the transmission and storage of audio signals. Then, to meet the demand for high-quality audio, a series of stereo codec technologies have emerged.

[0144] Among them, one of the most commonly used stereo codec technologies is parametric stereo codec technology. The theoretical basis of this technology is the principle of spatial hearing. Specifically, during the audio encoding process, the original stereo audio signal is converted into a single-channel signal and some spatial parameters for representation, or the original stereo audio signal is converted into a single-channel signal, a residual signal, and some spatial parameters for representation; during the audio decoding process, the stereo audio signal is reconstructed through the decoded single-channel signal and spatial parameters, or through the decoded single-channel signal, residual signal, and spatial parameters.

[0145] Figure 1 This is a schematic flowchart of the parametric stereo encoding and decoding method in the frequency domain in the embodiments of the present application. Refer to Figure 1 As shown, this process may include:

[0146] S101: The encoding side performs time-frequency transformation (such as discrete Fourier transform (DFT)) on the first-channel audio signal and the second-channel audio signal of the current frame in the stereo audio signal to obtain the first-channel frequency-domain signal and the second-channel frequency-domain signal;

[0147] First of all, it should be noted that the input stereo audio signal obtained by the encoding side may include two audio signals, that is, the first-channel audio signal and the second-channel audio signal (such as the left-channel audio signal and the right-channel audio signal); the two audio signals included in the above stereo audio signal may also be two of the multi-channel audio signals or two audio signals jointly generated by multiple audio signals in the multi-channel audio signal, and no specific limitation is made thereto.

[0148] Here, when the encoding side encodes the stereo audio signal, it will perform frame division processing to obtain multiple audio frames and process them frame by frame.

[0149] S102: The encoding side extracts spatial parameters, downmix signals, and residual signals from the first-channel frequency-domain signal and the second-channel frequency-domain signal;

[0150] The above spatial parameters may include: inter-channel coherence (IC), inter-channel level difference (ILD), inter-channel time difference (ITD), inter-channel phase difference (IPD), etc.

[0151] S103: The encoding side encodes the spatial parameters, downmix signals, and residual signals respectively;

[0152] S104: The encoding side generates a frequency-domain parametric stereo bitstream according to the encoded spatial parameters, downmix signals, and residual signals;

[0153] S105: The encoding side sends the frequency-domain parametric stereo bitstream to the decoding side.

[0154] S106: The decoding side decodes the received frequency-domain parametric stereo bitstream to obtain the corresponding spatial parameters, downmix signals, and residual signals;

[0155] S107: The decoding side performs frequency-domain upmixing on the downmixed signal and the residual signal to obtain an upmixed signal;

[0156] S108: The decoding side synthesizes the upmixed signal and the spatial parameters to obtain a frequency-domain audio signal;

[0157] S109: The decoding side performs inverse time-frequency transformation (such as inverse discrete Fourier transform (IDFT)) on the frequency-domain audio signal in combination with the spatial parameters to obtain the first-channel audio signal and the second-channel audio signal of the current frame;

[0158] Further, the encoding side performs the above first to fifth steps on each audio frame in the stereo audio signal, and the decoding side performs the above sixth to ninth steps on each frame. In this way, the decoding side can obtain the first-channel audio signal and the second-channel audio signal of multiple audio frames, and further obtain the first-channel audio signal and the second-channel audio signal of the stereo audio signal.

[0159] In the above process of parametric stereo encoding and decoding, the ILD and ITD in the spatial parameters contain the position information of the sound source. Therefore, accurate estimation of ILD and ITD is crucial for the reconstruction of stereo sound image and sound field.

[0160] In the parametric stereo encoding technology, the most commonly used method for estimating ITD can be the generalized cross-correlation method, which has the advantages of low complexity, good real-time performance, easy implementation, and does not depend on other prior information of the stereo audio signal. Figure 2 For the flow diagram of the generalized cross-correlation algorithm in the embodiments of this application, see Figure 2 As shown, this method may include:

[0161] S201: The encoding side performs DFT on the stereo audio signal to obtain the first-channel frequency-domain signal and the second-channel frequency-domain signal;

[0162] S202: The encoding side calculates the frequency-domain cross-power spectrum and the frequency-domain weighting function of the two according to the first-channel frequency-domain signal and the second-channel frequency-domain signal;

[0163] S203: The encoding side weights the frequency-domain cross-power spectrum using the frequency-domain weighting function;

[0164] S204: The encoding side performs IDFT on the weighted frequency-domain cross-power spectrum to obtain the frequency-domain cross-correlation function;

[0165] S205: The encoding side performs peak detection on the frequency-domain cross-correlation function;

[0166] S206: The encoding side determines the estimated value of ITD according to the peak of the cross-correlation function.

[0167] In the above-mentioned generalized cross-correlation algorithm, the frequency-domain weighting function in the second step above can adopt the following several functions.

[0168] First, the frequency-domain weighting function in the second step above can be as shown in formula (1):

[0169]

[0170] where Φ PHAT (k) is the PHAT weighting function, X 1 (k) is the frequency-domain audio signal of the first-channel audio signal x 1 (n), that is, the first-channel frequency-domain signal; X 2 (k) is the frequency-domain audio signal of the second-channel audio signal x 2 (n), that is, the second-channel frequency-domain signal; is the cross-power spectrum of the first channel and the second channel; k is the frequency-point index value, k = 0, 1,..., N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

[0171] Correspondingly, the weighted generalized cross-correlation function can be as shown in formula (2):

[0172]

[0173] In practical applications, using the frequency-domain weighting function shown in formula (1) and the weighted generalized cross-correlation function shown in formula (2) for ITD estimation can be called the generalized cross-correlation with phase transformation (GCC-PHAT) algorithm. Since the energy of stereo audio signals varies greatly at different frequency points, the frequency points with low energy are greatly affected by noise, while the frequency points with high energy are less affected by noise. Then, in the GCC-PHAT algorithm, after the cross-power spectrum is weighted by the PHAT weighting function, the weighting values of each frequency point have exactly the same weight in the generalized cross-correlation function, resulting in the GCC-PHAT algorithm being very sensitive to noise signals. Even at medium and high signal-to-noise ratios, the performance of the GCC-PHAT algorithm will decline significantly. In addition, when there is one or several noise sources in the space, that is, when there are competing sound sources, there will be correlated noise signals in the stereo audio signals, weakening the peaks corresponding to the target signals (such as speech signals) in the current frame. Then, in some cases, for example, when the energy of the correlated noise signal is greater than the energy of the target signal or the noise source is closer to the microphone, the peak of the correlated noise signal will be greater than the peak corresponding to the target signal. At this time, the ITD estimation value of the stereo audio signal is the ITD estimation value of the noise signal. That is, in the presence of correlated noise, not only will the ITD estimation accuracy of the stereo audio signal decrease severely, but the ITD estimation value of the stereo audio signal will also continuously switch between the ITD value of the target signal and the ITD value of the noise signal, thus affecting the stability of the sound image of the encoded stereo audio signal.

[0174] Second, the frequency-domain weighting function in the second step above can also be as shown in formula (3):

[0175]

[0176] where β is the amplitude weighting parameter, and β ∈ [0, 1].

[0177] Correspondingly, the weighted generalized cross-correlation function can also be as shown in formula (4):

[0178]

[0179] In practical applications, using the frequency-domain weighting function shown in formula (3) and the weighted generalized cross-correlation function shown in formula (4) for ITD estimation can be called the GCC-PHAT-β algorithm. Since the optimal value of β is different for different types of noise signals and the differences between the optimal values are relatively large, the performance of the GCC-PHAT-β algorithm is different for different types of noise signals. Moreover, at medium and high signal-to-noise ratios, although the performance of the GCC-PHAT-β algorithm is improved to a certain extent, it cannot meet the requirements of parametric stereo codec technology for the ITD estimation accuracy. Further, in the presence of correlated noise, the performance of the GCC-PHAT-β algorithm will also seriously decline.

[0180] Thirdly, the frequency-domain weighting function in the second step above can also be as shown in formula (5):

[0181]

[0182] where Γ 2 (k) is the coherence square value of the k-th frequency bin in the current frame,

[0183] Correspondingly, the weighted generalized cross-correlation function can also be as shown in formula (6):

[0184]

[0185] In practical applications, using the frequency-domain weighting function shown in formula (5) and the weighted generalized cross-correlation function shown in formula (6) for ITD estimation can be called the GCC-PHAT-Coh algorithm. Under certain conditions, the coherence square values of most frequency bins in the correlated noise in the stereo audio signal will be greater than the coherence square value of the target signal in the current frame, which will cause the performance of the GCC-PHAT-Coh algorithm to seriously decline. Moreover, due to the extremely large energy differences at different frequency bins in the stereo audio signal, the GCC-PHAT-Coh algorithm does not consider the influence of the energy differences at different frequency bins on the algorithm performance, resulting in poor ITD estimation performance under some conditions.

[0186] As can be seen from the above, noise has a relatively serious impact on the performance of the generalized cross-correlation algorithm, resulting in a serious decline in the ITD estimation accuracy, and further causing problems such as inaccurate, unstable, poor sense of space, and obvious head-in-the-middle effect in the decoded stereo audio signal in parametric codec technology, seriously affecting the sound quality of the encoded stereo audio signal.

[0187] To solve the above problems, an embodiment of the present application provides a method for estimating the time delay of a stereophonic audio signal. This method can be applied to an audio coding device, which can be used for the audio coding part in an audio-visual communication system involving stereo and multi-channel, or for the audio coding part in a virtual reality (VR) application program.

[0188] In practical applications, the above audio coding device can be set in a terminal in an audio-visual communication system. For example, the terminal can be a device that provides voice or data connectivity to users. For example, it can also be called a user equipment (UE), a mobile station, a subscriber unit, a station (STA), or a terminal equipment (TE), etc. The terminal device can be a cellular phone, a personal digital assistant (PDA), a wireless modem, a handheld device, a laptop computer, a cordless phone, a wireless local loop (WLL) station, or a tablet computer, etc. With the development of wireless communication technology, any device that can access a wireless communication system, communicate with the network side of the wireless communication system, or communicate with other devices through the wireless communication system can be the terminal device in the embodiment of the present application. For example, the terminals and cars in intelligent transportation, the household devices in smart homes, the electricity meter reading instruments, voltage monitoring instruments, and environmental monitoring instruments in smart grids, the video monitoring instruments and cash registers in smart security networks, etc. The terminal device can be static and fixed or mobile.

[0189] Alternatively, the above audio encoder can also be set in a device with VR functions. For example, the device can be a smart phone, a tablet computer, a smart TV, a laptop computer, a personal computer, a wearable device (such as a VR glasses, a VR helmet, a VR hat), etc. that supports VR applications, or can also be set in a cloud server that communicates with the above device with VR functions. Of course, the above audio coding device can also be set on other devices with the function of storing and / or transmitting stereophonic audio signals. The embodiment of the present application does not make specific limitations.

[0190] In the embodiments of the present application, the stereo audio signal may be an original stereo audio signal (including a left-channel audio signal and a right-channel audio signal), or a stereo audio signal composed of two audio signals in a multi-channel audio signal, or a stereo signal composed of two audio signals jointly generated by multiple audio signals in a multi-channel audio signal. Of course, the stereo audio signal may also have other forms, which are not specifically limited in the embodiments of the present application. In the following embodiments, taking the stereo audio signal as the original stereo audio signal as an example for illustration, the stereo audio signal may include a left-channel time-domain signal and a right-channel time-domain signal in the time domain, and the stereo audio signal may include a left-channel frequency-domain signal and a right-channel frequency-domain signal in the frequency domain. Then, the first-channel audio signal in the following embodiments may be a left-channel audio signal (both in the time domain and in the frequency domain), the first-channel time-domain signal may be a left-channel time-domain signal, and the first-channel frequency-domain signal may be a left-channel frequency-domain signal; similarly, the second-channel audio signal may be a right-channel audio signal (both in the time domain and in the frequency domain), the second-channel time-domain signal may be a right-channel time-domain signal, and the second-channel frequency-domain signal may be a right-channel frequency-domain signal.

[0191] Optionally, the above audio encoding device may specifically be a stereo encoding device, which may constitute an independent stereo encoder; or it may be the core encoding part in a multi-channel encoder, aiming to encode a stereo audio signal composed of two audio signals jointly generated by multiple signals in a multi-channel audio signal.

[0192] The following describes the stereo audio signal delay estimation method provided by the embodiments of the present application.

[0193] First, the frequency-domain weighting function provided by the embodiments of the present application will be described.

[0194] In the embodiments of the present application, in order to improve the performance of the generalized cross-correlation algorithm, the frequency-domain weighting function (as shown in the above formulas (1), (3), and (5)) in the above several algorithms may be improved, and the improved frequency-domain weighting function may be, but not limited to, the following several functions.

[0195] First, the construction factors of the improved frequency-domain weighting function (i.e., the first weighting function) may include: a left-channel Wiener gain factor (i.e., the Wiener gain factor corresponding to the first-channel frequency-domain signal), a right-channel Wiener gain factor (i.e., the Wiener gain factor corresponding to the second-channel frequency-domain signal), and the coherence squared value of the current frame.

[0196] Here, the construction factor refers to a factor or factor used to construct the objective function. Then, when the objective function is the improved frequency-domain weighting function, its construction factor may be one or more functions used to construct the improved frequency-domain weighting function.

[0197] In practical applications, the first improved frequency-domain weighting function can be as shown in Formula (7):

[0198]

[0199] where Φ new_1 (k) is the first improved frequency-domain weighting function, β is the amplitude weighting parameter, β ∈ [0, 1]. For example, β = 0.6, 0.7, 0.8, etc., W x1 (k) is the left-channel Wiener gain factor; W x2 (k) is the right-channel Wiener gain factor; Γ 2 (k) is the coherence squared value of the k-th frequency point in the current frame,

[0200] In some possible embodiments, the first improved frequency-domain weighting function can also be as shown in Formula (8):

[0201]

[0202] Correspondingly, the generalized cross-correlation function weighted by the first improved frequency-domain weighting function can also be as shown in Formula (9):

[0203]

[0204] In some possible implementation manners, the above-mentioned left-channel Wiener gain factor may include a first initial Wiener gain factor and / or a first improved Wiener gain factor; the above-mentioned right-channel Wiener gain factor may include a second initial Wiener gain factor and / or a second improved Wiener gain factor.

[0205] In practical applications, the first initial Wiener gain factor can be determined by performing noise power spectrum estimation on X 1 (k). Specifically, when the left-channel Wiener gain factor includes the first initial Wiener gain factor, the above method may further include: First, the audio coding device may obtain an estimated value of the left-channel noise power spectrum according to the left-channel frequency-domain signal X 1 (k) of the current frame, and then determine the first initial Wiener gain factor according to the estimated value of the left-channel noise power spectrum; similarly, the second initial Wiener gain factor can also be determined by performing noise power spectrum estimation on X 2 (k). Specifically, when the right-channel Wiener gain factor includes the second initial Wiener gain factor, first, the audio coding device may obtain an estimated value of the right-channel noise power spectrum according to the right-channel frequency-domain signal X 2 (k) of the current frame, and determine the second initial Wiener gain factor according to the estimated value of the right-channel noise power spectrum.

[0206] In the above process of estimating the noise power spectrum of X(k) and X(k) for the current frame, algorithms such as the minimum statistics algorithm and the minimum tracking algorithm can be used to calculate. Of course, other algorithms can also be used to calculate the estimated values of the noise power spectra of X(k) and X(k), and the embodiments of the present application do not make specific limitations. 1 (k) and X 2 (k), and the estimated values of the noise power spectra of X 1 (k) and X 2 (k) are not specifically limited in the embodiments of the present application.

[0207] For example, the above first initial Wiener gain factor can be as shown in formula (10):

[0208]

[0209] The above second initial Wiener gain factor can be as shown in formula (11):

[0210]

[0211] Wherein, is the estimated value of the left-channel noise power spectrum, is the estimated value of the right-channel noise power spectrum.

[0212] In some possible implementation manners, in addition to directly using the first initial Wiener gain factor and the second initial Wiener gain factor to construct the first improved frequency-domain weighting function, the above left-channel Wiener gain factor and right-channel Wiener gain factor can also construct a corresponding binary masking function based on the first initial Wiener gain factor and the second initial Wiener gain factor to obtain the above first improved Wiener gain factor and second improved Wiener gain factor. The first improved frequency-domain weighting function constructed using the first improved Wiener gain factor and the second improved Wiener gain factor can filter out the frequency points less affected by noise, thereby improving the ITD estimation accuracy of the stereo audio signal.

[0213] Then, when the left-channel Wiener gain factor includes the first improved Wiener gain factor, the above method may further include: after obtaining the first initial Wiener gain factor, the audio coding device constructs a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; similarly, after obtaining the second initial Wiener gain factor, the audio coding device constructs a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

[0214] For example, the first improved Wiener gain factor can be as shown in formula (12):

[0215]

[0216] The second improved Wiener gain factor can be as shown in formula (13):

[0217]

[0218] where μ 0 is the binary masking threshold of the Wiener gain factor, μ 0 ∈[0.5, 0.8]. For example, μ 0 = 0.5, 0.66, 0.75, 0.8, etc.

[0219] Then, as can be seen from the above, the left-channel Wiener gain factor W x1 (k) can include and The right-channel Wiener gain factor W x2 (k) can include and Then, in the process of constructing the first improved frequency-domain weighting function as shown in formula (7) or (8), and can be substituted into formula (7) or (8), or and can be substituted into formula (7) or (8).

[0220] For example, and The first improved frequency-domain weighting function after substituting into formula (7) can be as shown in formula (14):

[0221]

[0222] Substituting and into formula (7), the first improved frequency-domain weighting function can be as shown in formula (15):

[0223]

[0224] In the embodiments of the present application, if the first improved frequency-domain weighting function is used to weight the frequency-domain cross-power spectrum of the current frame, after weighting by the Wiener gain factor, the weight of the correlation noise component in the frequency-domain cross-power spectrum of the stereophonic audio signal is greatly reduced, and the correlation of the residual noise component is also greatly reduced. In most cases, the coherent square value of the residual noise will be much smaller than the coherent square value of the target signal in the stereophonic audio signal, so that the cross-correlation peak corresponding to the target signal will be more prominent, and the accuracy and stability of the ITD estimation of the stereophonic audio signal will be greatly improved.

[0225] The construction factors of the second, improved frequency-domain weighting function (i.e., the second weighting function) may include: an amplitude weighting parameter β and the coherence squared value of the current frame.

[0226] In practical applications, the second improved frequency-domain weighting function may be as shown in Equation (16):

[0227]

[0228] where Φ new_2 is the second improved frequency-domain weighting function, β ∈ [0, 1]. For example, β = 0.6, 0.7, 0.8, etc.

[0229] Correspondingly, the generalized cross-correlation function weighted by the second improved frequency-domain weighting function may be as shown in Equation (17):

[0230]

[0231] In the embodiments of the present application, if the second improved frequency-domain weighting function is used to weight the frequency-domain cross-power spectrum of the current frame, it can ensure that the frequency points with large energy and high correlation have larger weights, and the frequency points with small energy or small correlation have smaller weights, thereby improving the accuracy of the ITD estimation of the stereophonic audio signal.

[0232] Secondly, a method for estimating the time delay of a stereophonic audio signal provided by the embodiments of the present application is introduced. This method is based on the above-mentioned improved frequency-domain weighting function to estimate the ITD value of the current frame.

[0233] Figure 3 For the flow schematic of the method for estimating the time delay of the stereophonic audio signal in the embodiments of the present application Figure 1 , as shown by the solid line in Figure 3 , this method may include:

[0234] S301: Obtain the current frame in the stereophonic audio signal;

[0235] where the current frame includes the left-channel audio signal and the right-channel audio signal.

[0236] The audio coding device obtains the input stereophonic audio signal. The stereophonic audio signal may include two audio signals, and these two audio signals may be time-domain audio signals or frequency-domain audio signals.

[0237] In one case, the two audio signals in the stereophonic audio signal are time-domain audio signals, that is, the left-channel time-domain signal and the right-channel time-domain signal (i.e., the first-channel time-domain signal and the second-channel time-domain signal). In this case, the above-mentioned stereophonic audio signal may be input through a sound sensor such as a microphone or a receiver. See Figure 3As shown by the dashed line in the figure, after S301, the method may further include: S302: Perform time-frequency transformation on the left-channel time-domain signal and the right-channel time-domain signal. Here, the audio encoding device performs frame division processing on the time-domain audio signal through S301 to obtain the current frame in the time domain. At this time, the current frame may include the left-channel time-domain signal and the right-channel time-domain signal. Then, the audio encoding device performs time-frequency transformation on the current frame in the time domain to obtain the current frame in the frequency domain. At this time, the current frame may include the left-channel frequency-domain signal and the right-channel frequency-domain signal (i.e., the first-channel frequency-domain signal and the second-channel frequency-domain signal).

[0238] In another case, the two audio signals in the stereo audio signal are frequency-domain audio signals, that is, the left-channel frequency-domain signal and the right-channel frequency-domain signal (i.e., the first-channel frequency-domain signal and the second-channel frequency-domain signal). In this case, the above stereo audio signal itself is two frequency-domain audio signals. Then, the audio encoding device can directly perform frame division processing on the stereo audio signal (i.e., the frequency-domain audio signal) in the frequency domain through S301 to obtain the current frame in the frequency domain. The current frame may include the left-channel frequency-domain signal and the right-channel frequency-domain signal (i.e., the first-channel frequency-domain signal and the second-channel frequency-domain signal).

[0239] It should be noted that in the subsequent description of the embodiments, if the stereo audio signal is a time-domain audio signal, the audio encoding device can perform time-frequency transformation on it to obtain the corresponding frequency-domain audio signal, and then process it in the frequency domain; if the stereo audio signal itself is a frequency-domain audio signal, the audio encoding device can directly process it in the frequency domain.

[0240] In practical applications, the left-channel time-domain signal after frame division processing in the current frame can be denoted as x 1 (n), and the right-channel time-domain signal after frame division processing in the current frame can be denoted as x 2 (n), where n is the sampling point.

[0241] In some possible implementation manners, after S301, the audio encoding device may further preprocess the current frame. For example, perform high-pass filtering on x 1 (n) and x 2 (n) respectively to obtain the preprocessed left-channel time-domain signal and the right-channel time-domain signal. The preprocessed left-channel time-domain signal is denoted as The preprocessed right-channel time-domain signal is denoted as Optionally, the high-pass filtering process can be an infinite impulse response (IIR) filter with a cut-off frequency of 20 Hz, or other types of filters. The embodiments of the present application do not make specific limitations.

[0242] Optionally, the audio encoding device may also perform time-frequency transformation on x 1 (n) and x 2 (n) to obtain X 1 (k) and X 2 (k); among them, the left-channel frequency-domain signal may be denoted as X 1 (k), and the right-channel frequency-domain signal may be denoted as X 2 (k).

[0243] Here, the audio encoding device may adopt time-frequency transformation algorithms such as DFT, fast Fourier transform (FFT), modified discrete cosine transform (MDCT), etc. to transform the time-domain signal into a frequency-domain signal. Of course, the audio encoding device may also adopt other time-frequency transformation algorithms, and the embodiments of the present application do not make specific limitations.

[0244] Assume that DFT is used to perform time-frequency transformation on the time-domain signals of the left and right channels. Specifically, the audio encoding device may perform DFT on x 1 (n) or to obtain X 1 (k); similarly, the audio encoding device may perform DFT on x 2 (n) or to obtain X 2 (k).

[0245] Furthermore, in order to overcome the problem of spectral aliasing, generally, the overlap-add method is used to process between the DFTs of adjacent two frames, and sometimes zero-padding is also performed on the input signal of the DFT.

[0246] S303: Calculate the frequency-domain cross-power spectrum of the current frame according to X 1 (k) and X 2 (k);

[0247] Here, the frequency-domain cross-power spectrum of the current frame may be as shown in formula (18):

[0248]

[0249] Among them, is the conjugate function of X 2 (k).

[0250] S304: Weight the frequency-domain cross-power spectrum by using a preset weighting function;

[0251] Here, the preset weighting function may refer to the above-mentioned improved frequency-domain weighting function, that is, the first improved frequency-domain weighting function Φ new_1Or the second improved frequency-domain weighting function Φ new_2 .

[0252] S304 can be understood as the audio coding device multiplying the improved weighting function by the frequency-domain power spectrum. Then, the weighted frequency-domain cross-power spectrum can be expressed as: Φ new_1 (k)C x1x2 (k) or Φ new_2 (k)C x1x2 (k).

[0253] In the embodiment of the present application, before executing S305, the audio coding device can also use X 1 (k) and X 2 (k) to calculate the improved frequency-domain weighting function (i.e., the preset weighting function).

[0254] S305: Perform an inverse time-frequency transform on the weighted frequency-domain cross-power spectrum to obtain the cross-correlation function;

[0255] The audio coding device can use the inverse time-frequency transform algorithm corresponding to the time-frequency transform algorithm adopted in S302 to transform the frequency-domain cross-power spectrum from the frequency domain to the time domain to obtain the cross-correlation function.

[0256] Here, the cross-correlation function corresponding to Φ new_1 (k)C x1x2 (k) can be as shown in formula (19):

[0257]

[0258] Or, the cross-correlation function corresponding to Φ new_2 (k)C x1x2 (k) can be as shown in formula (20):

[0259]

[0260] S306: Perform peak detection on the cross-correlation function;

[0261] After the audio coding device obtains the cross-correlation function through S306, it can determine the maximum value Δmax of the ITD (which can also be understood as the time range of ITD estimation) according to the preset sampling rate and the maximum distance between the sound sensor (i.e., microphone, receiver, etc.). For example, if the sampling points corresponding to Δmax are set to 5 ms, then, if the sampling rate of the stereo audio signal is 32 kHz, then Δmax = 160, that is, the maximum delay points of the left and right channels are 160 sampling points. Then, the audio coding device searches for the maximum peak of C x1x2 (n) within the range of n ∈ [-Δmax, Δmax], and the index value corresponding to this peak is the alternative value of the ITD for the current frame.

[0262] S307: Calculate the estimated value of the ITD of the current frame based on the peak value of the cross-correlation function.

[0263] The audio encoding device determines the alternative value of the ITD of the current frame according to the peak value of the cross-correlation function, and then combines the alternative value of the ITD of the current frame, the ITD value of the previous frame (i.e., historical information), the audio tail processing parameter, the correlation degree between the front and rear frames and other side information to determine the estimated value of the ITD of the current frame, so as to remove the outliers in the delay estimation.

[0264] Further, after the audio encoding device determines the estimated value of the ITD through S307, it can encode it and write it into the encoded bitstream of the stereophonic audio signal.

[0265] In the embodiments of the present application, if the first improved frequency-domain weighting function is used to weight the frequency-domain cross-power spectrum of the current frame, after being weighted by the Wiener gain factor, the weight of the correlated noise component in the frequency-domain cross-power spectrum of the stereophonic audio signal is greatly reduced, and the correlation of the residual noise component is also greatly reduced. In most cases, the coherence square value of the residual noise will be much smaller than the coherence square value of the target signal in the stereophonic audio signal, so that the cross-correlation peak corresponding to the target signal will be more prominent, and the accuracy and stability of the ITD estimation of the stereophonic audio signal will be greatly improved. If the second improved frequency-domain weighting function is used to weight the frequency-domain cross-power spectrum of the current frame, it can ensure that the frequency points with large energy and high correlation have larger weights, and the frequency points with small energy or small correlation have smaller weights, thereby improving the accuracy of the ITD estimation of the stereophonic audio signal.

[0266] Again, another method for estimating the delay of a stereophonic audio signal provided by the embodiments of the present application is introduced. This method uses different algorithms for ITD estimation for different types of noise signals in the stereophonic audio signal based on the above embodiments.

[0267] Figure 4 Flow schematic of the method for estimating the delay of the stereophonic audio signal in the embodiments of the present application Figure 2 , see Figure 4 As shown, this method may include

[0268] S401: Obtain the current frame of the stereophonic audio signal;

[0269] Here, for the implementation process of S401, refer to the description of S301, and no specific limitation is made here.

[0270] S402: Determine the signal type of the noise signal included in the current frame; if the signal type of the noise signal included in the current frame is the correlated noise signal type, then execute S403; if the signal type of the noise signal included in the current frame is the diffused noise signal type, then execute S404;

[0271] In a noisy environment, different noise signal types have different effects on the generalized cross-correlation algorithm. Therefore, in order to make full use of the performance of each generalized cross-correlation algorithm and improve the accuracy of ITD estimation, the audio coding device can determine the signal type of the noise signal included in the current frame, and then determine a suitable frequency-domain weighting function for the current frame from multiple frequency-domain weighting functions.

[0272] In practical applications, the above-mentioned correlated noise signal type refers to the noise signal type in which the correlation of the noise signals in the two audio signals of the stereo audio signal exceeds a certain degree. That is to say, the noise signal included in the current frame can be classified as a correlated noise signal; the above-mentioned diffused noise signal type refers to the noise signal type in which the correlation of the noise signals in the two audio signals of the stereo audio signal is lower than a certain degree. That is to say, the noise signal included in the current frame can be classified as a diffused noise signal.

[0273] In some possible implementation manners, the current frame may include both a correlated noise signal and a diffused noise signal. At this time, the audio coding device will determine the signal type of the main noise signal in the two noise signals as the signal type of the noise signal included in the current frame.

[0274] In some possible implementation manners, the audio coding device can determine the signal type of the noise signal included in the current frame by calculating the noise coherence value of the current frame. Then, S402 may include: obtaining the noise coherence value of the current frame; if the noise coherence value is greater than or equal to a preset threshold, it indicates that the noise signal included in the current frame has strong correlation. Then, the audio coding device can determine that the signal type of the noise signal included in the current frame is the correlated noise signal type; if the noise coherence value is less than the preset threshold, it indicates that the correlation of the noise signal included in the current frame is weak. Then, the audio coding device can determine that the signal type of the noise signal included in the current frame is the diffused noise signal type.

[0275] Here, the preset threshold of the noise coherence value is an empirical value, which can be set according to factors such as ITD estimation performance. For example, the preset threshold is set to 0.20, 0.25, 0.30, etc. Of course, it can also be set to other appropriate values, and the embodiments of the present application do not make specific limitations on this.

[0276] In practical applications, after the audio encoding device calculates the noise coherence value of the current frame, it can also perform smoothing processing on it to reduce the error in the estimation of the noise coherence value and improve the recognition accuracy of the noise type.

[0277] S403: Estimate the ITD value of the left-channel audio signal and the right-channel audio signal using the first algorithm;

[0278] Here, the first algorithm may include weighting the cross-power spectrum in the frequency domain of the current frame using the first weighting function; it may also include performing peak detection on the weighted cross-correlation function and estimating the ITD value of the current frame based on the peak of the weighted cross-correlation function.

[0279] After the audio encoding device determines through S402 that the signal type of the noise signal included in the current frame is a correlated noise signal type, it can use the first algorithm to estimate the ITD value of the current frame. For example, the audio encoding device selects to weight the cross-power spectrum in the frequency domain of the current frame using the first weighting function, and then performs peak detection on the weighted cross-correlation function, and estimates the ITD value of the current frame based on the peak of the weighted cross-correlation function.

[0280] In some possible embodiments, the first weighting function may be one or more of the frequency domain weighting functions and / or improved frequency domain weighting functions in the above one or more embodiments that perform better under the correlated noise condition, such as the frequency domain weighting function shown in formula (3), and the improved frequency domain weighting functions shown in formulas (7) and (8).

[0281] Preferably, the first weighting function may be the first improved frequency domain weighting function described in the above embodiments, such as the improved frequency domain weighting functions shown in formulas (7) and (8).

[0282] S404: Estimate the ITD value of the left-channel audio signal and the right-channel audio signal using the second algorithm.

[0283] Here, the second algorithm includes weighting the cross-power spectrum in the frequency domain of the current frame using the second weighting function, and may also include performing peak detection on the weighted cross-correlation function and estimating the ITD value of the current frame based on the peak of the weighted cross-correlation function.

[0284] Correspondingly, after the audio encoding device determines through S402 that the signal type of the noise signal included in the current frame is a diffuse noise signal type, it can use the second algorithm to estimate the ITD value of the current frame. For example, the audio encoding device can select to weight the cross-power spectrum in the frequency domain of the current frame using the second weighting function, and then perform peak detection on the weighted cross-correlation function, and estimate the ITD value of the current frame based on the peak of the weighted cross-correlation function.

[0285] In some possible embodiments, the second weighting function may be one or more weighting functions with better performance under diffusive noise conditions among the frequency-domain weighting functions and / or improved frequency-domain weighting functions in one or more of the above embodiments, such as the frequency-domain weighting function shown in Formula (5) and the improved frequency-domain weighting function shown in Formula (16).

[0286] Preferably, the second weighting function may be the second improved frequency-domain weighting function described in the above embodiments, that is, the improved frequency-domain weighting function shown in Formula (16).

[0287] In some possible embodiments, since the stereo audio signal includes both a speech signal and a noise signal, the signal type included in the current frame obtained by frame segmentation in S401 may be a speech signal or a noise signal. Then, in order to simplify the processing and further improve the accuracy of ITD estimation, before S402, the above method may further include: performing voice activity detection on the current frame to obtain a detection result; if the detection result indicates that the signal type of the current frame is a noise signal type, calculating the noise coherence value of the current frame; if the detection result indicates that the signal type of the current frame is a speech signal type, determining the noise coherence value of the previous frame of the current frame in the stereo audio signal as the noise coherence value of the current frame.

[0288] After obtaining the current frame, the audio coding device may perform voice activity detection (VAD) on the current frame to distinguish whether the main signal of the current frame is a speech signal or a noise signal. If it is detected that the current frame contains a noise signal, then in S402, the noise coherence value of the current frame can be directly calculated; if it is detected that the current frame contains a speech signal, then in S402, when calculating the noise coherence value, the noise coherence value combined with the historical frame, such as the noise coherence value of the previous frame of the current frame, can be determined as the noise coherence value of the current frame. Here, the previous frame of the current frame may contain a noise signal or a speech signal. If the previous frame still contains a speech signal, then the noise coherence value of the previous noise frame in the historical frame is determined as the noise coherence value of the current frame.

[0289] In the specific implementation process, the audio coding device may use various methods to perform VAD; when the value of VAD is 1, it indicates that the signal type of the current frame is a speech signal type; when the value of VAD is 0, it indicates that the signal type of the current frame is a noise signal type.

[0290] It should be noted that in the embodiments of the present application, the audio coding device may calculate the value of VAD in a time domain, a frequency domain, or a combination of time domain and frequency domain, and no specific limitation is made thereto.

[0291] The following uses specific examples to illustrate the above Figure 4 stereo audio signal delay estimation method shown.

[0292] Figure 5 The flow diagram of the stereo audio signal delay estimation method in the embodiment of the present application Figure 3 is as follows. The method may include:

[0293] S501: Frame the stereo audio signal to obtain x 1 (n) and x 2 (n) of the current frame;

[0294] S502: Perform DFT on x 1 (n) and x 2 (n) to obtain X 1 (k) and X 2 (k) of the current frame;

[0295] S503: Calculate the VAD value of the current frame according to x 1 (n) and x 2 (n) or X 1 (k) and X 2 (k) of the current frame; if VAD = 1, execute S504; if VAD = 0, execute S505;

[0296] Here, as shown by the dashed line in Figure 5 , S503 can be executed after S501 or after S502, and no specific limitation is made thereto.

[0297] S504: Calculate the noise coherence value Γ(k) of the current frame according to X 1 (k) and X 2 (k);

[0298] S505: Confirm the Γ m-1 (k) of the previous frame as the Γ(k) of the current frame;

[0299] Here, the Γ(k) of the current frame can also be expressed as Γ m (k), that is, the noise coherence value of the mth frame, where m is a positive integer.

[0300] S506: Compare the Γ(k) of the current frame with a preset threshold Γ thres ; if Γ(k) is greater than or equal to Γ thres , then execute S507; if Γ(k) is less than Γ thres , then execute S508;

[0301] S507: Use Φ new_1 (k) to process C x1x2(k) is weighted. At this time, the weighted cross-power spectrum in the frequency domain can be expressed as: Φ new_1 (k)C x1x2 (k);

[0302] S508: Use Φ PHAT-Coh (k) to weight C x1x2 (k) of the current frame. At this time, the weighted cross-power spectrum in the frequency domain can be expressed as: Φ PHAT-Coh (k)C x1x2 (k);

[0303] In practical applications, after S506, if it is determined to execute S507, X 1 (k) and X 2 (k) of the current frame can be used to calculate C x1x2 (k) and Φ new_1 (k) of the current frame; if it is determined to execute S508, X 1 (k) and X 2 (X) of the current frame can be used to calculate C x1x2 (k) and Φ PHAT-Coh (k).

[0304] S509: Perform IDFT on Φ new_1 (k)C x1x2 (k) or Φ PHAT-Coh (k)C x1x2 (k) to obtain the cross-correlation function G x1x2 (n);

[0305] Among them, G x1x2 (n) can be as shown in formula (6) or (9).

[0306] S510: Perform peak detection on G x1x2 (n);

[0307] S511: Calculate the estimated value of the ITD of the current frame according to the peak of G x1x2 (n).

[0308] So far, the ITD estimation process of the stereophonic audio signal is completed.

[0309] In some possible implementation manners, the above ITD estimation method can be applied not only to parametric stereo codec technologies, but also to technologies such as sound source localization, speech enhancement, and speech separation.

[0310] From the above, it can be seen that in the embodiment of the present application, the audio encoding device greatly improves the accuracy and stability of ITD estimation of stereo audio signals under diffuse noise and correlated noise conditions by adopting different ITD estimation algorithms for current frames containing different types of noise, reduces the inter-frame discontinuity between stereo downmix signals, and better maintains the phase of the stereo signal. The encoded stereo sound image is more accurate and stable, and the sense of reality is stronger, thereby improving the auditory quality of the encoded stereo signal.

[0311] Based on the same inventive concept, the embodiment of the present application provides a stereo audio signal delay estimation device, which can be a chip or system on chip in an audio encoding device, and can also be used in an audio encoding device to implement the above embodiments. Figure 4 The stereo audio signal delay estimation method shown and the functional modules of the method described in any possible implementation manner thereof. For example, Figure 6 This is a schematic diagram of the structure of the audio decoding device in the application embodiment, see Figure 6 As shown by the solid line, the stereo audio signal delay estimation device 600 includes: an acquisition module 601, which is used to obtain a current frame of the stereo audio signal, and the current frame includes a first channel audio signal and a second channel audio signal; an inter-channel time difference estimation module 602, which is used to use a first algorithm to estimate the inter-channel time difference between the first channel audio signal and the second channel audio signal if the signal type of the noise signal included in the current frame is a correlation noise signal type; if the signal type of the noise signal included in the current frame is a diffuse noise signal type, use a second algorithm to estimate the inter-channel time difference between the first channel audio signal and the second channel audio signal; wherein the first algorithm includes using a first weighting function to weight the frequency domain cross power spectrum of the current frame, and the second algorithm includes using a second weighting function to weight the frequency domain cross power spectrum of the current frame, and the construction factors of the first weighting function and the second weighting function are different.

[0312] In the embodiment of the present application, the current frame in the stereo signal obtained by the acquisition module 601 can be a frequency domain audio signal or a time domain audio signal. If the current frame is a frequency domain audio signal, the acquisition module 601 passes the current frame to the inter-channel time difference estimation module 602, and the inter-channel time difference estimation module 602 can directly process the current frame in the frequency domain; and if the current frame is a time domain audio signal, the acquisition module 601 can first perform a time-frequency transformation on the current frame in the time domain to obtain the current frame in the frequency domain, and then the acquisition module 601 passes the current frame in the frequency domain to the inter-channel time difference estimation module 602, and the inter-channel time difference estimation module 602 can process the current frame in the frequency domain.

[0313] In some possible implementations, see Figure 6As shown by the dashed line in the figure, the above device further includes: a noise coherence value calculation module 603, configured to obtain the noise coherence value of the current frame after the obtaining module 601 obtains the current frame; if the noise coherence value is greater than or equal to a preset threshold, determine that the signal type of the noise signal included in the current frame is a correlated noise signal type; or, if the noise coherence value is less than the preset threshold, determine that the signal type of the noise signal included in the current frame is a diffused noise signal type.

[0314] In some possible implementation manners, referring to Figure 6 As shown by the dashed line in the figure, the above device further includes: a voice activity detection module 604, configured to perform voice activity detection on the current frame to obtain a detection result; the noise coherence value calculation module 603 is specifically configured to calculate the noise coherence value of the current frame if the detection result indicates that the signal type of the current frame is a noise signal type; or, if the detection result indicates that the signal type of the current frame is a voice signal type, determine the noise coherence value of the previous frame of the current frame in the stereophonic audio signal as the noise coherence value of the current frame.

[0315] In the embodiments of the present application, the voice activity detection module 604 may calculate the VAD value in a time domain, a frequency domain, or a combination of the time domain and the frequency domain, and no specific limitation is made thereto. The obtaining module 601 may pass the current frame to the voice activity detection module 604 to perform VAD on the current frame.

[0316] In some possible implementation manners, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; an inter-channel time difference estimation module 602 is configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using a first weighting function; obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and the coherence square value of the current frame.

[0317] In some possible embodiments, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; the inter-channel time difference estimation module 602 is configured to calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using a first weighting function; and obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, the amplitude weighting parameter, and the coherence squared value of the current frame.

[0318] In some possible embodiments, the first weighting function Φ new_1 (k) satisfies the above formula (7).

[0319] In some other possible embodiments, the first weighting function Φ new_1 (k) satisfies the above formula (8).

[0320] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; the inter-channel time difference estimation module 602 is specifically configured to, after the obtaining module obtains the current frame, obtain an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determine the above first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtain an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; and determine the above second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

[0321] In some possible embodiments, the first initial Wiener gain factor satisfies the above formula (10), and the second initial Wiener gain factor satisfies the above formula (11).

[0322] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; the inter-channel time difference estimation module 602 is specifically configured to, after the obtaining module obtains the current frame, obtain the above first initial Wiener gain factor and second initial Wiener gain factor; construct a binary masking function for the above first initial Wiener gain factor to obtain the first improved Wiener gain factor; and construct a binary masking function for the above second initial Wiener gain factor to obtain the second improved Wiener gain factor.

[0323] In some possible embodiments, the first improved Wiener gain factor satisfies the above formula (12), and the second improved Wiener gain factor satisfies the above formula (13).

[0324] In some possible embodiments, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the inter-channel time difference estimation module 602 is specifically configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using a second weighting function to obtain an estimated value of the inter-channel time difference; wherein, the construction factors of the second weighting function include: an amplitude weighting parameter and the coherence square value of the current frame.

[0325] In some possible embodiments, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; the inter-channel time difference estimation module 602 is specifically configured to calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using a second weighting function; obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum; wherein, the construction factors of the second weighting function include: an amplitude weighting parameter and the coherence square value of the current frame.

[0326] In some possible embodiments, the second weighting function Φ new_2 (k) satisfies the above formula (16).

[0327] It should be noted that the specific implementation processes of the obtaining module 601, the inter-channel time difference estimation module 602, the noise coherence value calculation module 603, and the voice activity detection module 604 can refer to Figures 4 to 5 the detailed description of the embodiments. For the sake of simplicity of the specification, it will not be elaborated here.

[0328] In the embodiments of the present application, the obtaining module 601 mentioned may be a receiving interface, a receiving circuit, or a receiver, etc.; the inter-channel time difference estimation module 602, the noise coherence value calculation module 603, and the voice activity detection module 604 may be one or more processors.

[0329] Based on the same inventive concept, the embodiments of the present application provide a stereo audio signal delay estimation device, which may be a chip or a system-on-chip in an audio coding device, or may also be a functional module in an audio coding device for implementing the method and any possible embodiments thereof shown above Figure 3 of the stereo audio signal delay estimation method. For example, still referring toFigure 6 As shown in Figure 6 , the stereo audio signal time delay estimation device 600 includes: an obtaining module 601, configured to obtain a current frame in the stereo audio signal, where the current frame includes a first-channel audio signal and a second-channel audio signal; an inter-channel time difference estimation module 602, configured to calculate a frequency-domain cross-power spectrum of the current frame according to the first-channel audio signal and the second-channel audio signal; weight the frequency-domain cross-power spectrum by using a preset weighting function; and obtain an estimated value of the inter-channel time difference between the first-channel frequency-domain signal and the second-channel frequency-domain signal according to the weighted frequency-domain cross-power spectrum.

[0330] Wherein, the preset weighting function is a first weighting function or a second weighting function, and the construction factors of the first weighting function and the second weighting function are different; the construction factors of the first weighting function include: a Wiener gain factor corresponding to the first-channel frequency-domain signal, a Wiener gain corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and a coherence square value of the current frame; the construction factors of the second weighting function include: an amplitude weighting parameter and a coherence square value of the current frame.

[0331] In some possible implementation manners, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the inter-channel time difference estimation module 602 is configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; and calculate a frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal.

[0332] In some possible implementation manners, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal. At this time, the frequency-domain cross-power spectrum of the current frame can be directly calculated according to the first-channel audio signal and the second-channel audio signal.

[0333] In some possible implementation manners, the first weighting function Φ new_1 (k) satisfies the above formula (7).

[0334] In some other possible implementation manners, the first weighting function Φ new_1 (k) satisfies the above formula (8).

[0335] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; the inter-channel time difference estimation module 602 is specifically configured to, after the obtaining module 601 obtains the current frame, obtain an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determine the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtain an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; and determine the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

[0336] In some possible embodiments, the first initial Wiener gain factor satisfies the above formula (10), and the second initial Wiener gain factor satisfies the above formula (11).

[0337] In some possible embodiments, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; the inter-channel time difference estimation module 602 is specifically configured to, after the obtaining module 601 obtains the current frame, obtain the above first initial Wiener gain factor and second initial Wiener gain factor; construct a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; and construct a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

[0338] In some possible embodiments, the first improved Wiener gain factor satisfies the above formula (12), and the second improved Wiener gain factor satisfies the above formula (13).

[0339] In some possible embodiments, the second weighting function Φ new_2 (k) satisfies the above formula (16).

[0340] It should be noted that the specific implementation processes of the obtaining module 601 and the inter-channel time difference estimation module 602 can refer to Figure 3 the detailed description of the embodiments in, and for the sake of brevity of the specification, they will not be elaborated here.

[0341] In the embodiments of the present application, the obtaining module 601 mentioned may be a receiving interface, a receiving circuit, or a receiver, etc.; the inter-channel time difference estimation module 602 may be one or more processors.

[0342] Based on the same inventive concept, an embodiment of the present application provides an audio encoding device, which is consistent with the audio encoding device described in the above embodiment. Figure 7 It is a schematic structural diagram of the audio encoding device in the embodiment of the present application. Refer to Figure 7 As shown, the audio encoding device 700 includes: a non-volatile memory 701 and a processor 702 that are coupled to each other. The processor 702 calls the program code stored in the memory 701 to execute the operation steps of the method described in the above Figures 3 to 5 method for stereo audio signal delay estimation and any possible implementation manner thereof.

[0343] In some possible implementation manners, the audio encoding device may specifically be a stereo encoding device, and this device may constitute an independent stereo encoder; or it may be the core encoding part of a multi-channel encoder, aiming to encode a stereo audio signal composed of two audio signals jointly generated by multiple signals in a multi-channel frequency-domain signal.

[0344] In practical applications, the above audio encoding device may be implemented by using programmable devices such as application specific integrated circuit (ASIC), register transfer level (RTL), field programmable gate array (FPGA), etc. Of course, it may also be implemented by using other programmable devices, and the embodiments of the present application do not make specific limitations.

[0345] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores instructions, which are used to execute the operation steps of the method described in the above Figures 3 to 5 method for stereo audio signal delay estimation and any possible implementation manner thereof when the instructions run on a computer.

[0346] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, including an encoded bitstream, and the encoded bitstream includes the inter-channel time difference of the stereo audio signal obtained according to the method described in the above Figures 3 to 5 method for stereo audio signal delay estimation and any possible implementation manner thereof.

[0347] Based on the same inventive concept, an embodiment of the present application provides a computer program or a computer program product. When the computer program or the computer program product is executed on a computer, it enables the computer to implement the operation steps of the method described in the above Figures 3 to 5 method for stereo audio signal delay estimation and any possible implementation manner thereof.

[0348] Those skilled in the art will appreciate that the functions described in connection with the various illustrative logical blocks, modules, and algorithm steps disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions described by the various illustrative logical blocks, modules, and steps can be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium can include a computer-readable storage medium corresponding to a tangible medium, such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another (e.g., according to a communication protocol). In this way, the computer-readable medium generally can correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave. The data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this application. A computer program product can include a computer-readable medium.

[0349] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0350] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein may refer to any one of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described for the various illustrative logical blocks, modules, and steps described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Moreover, the techniques may be implemented entirely in one or more circuits or logic elements.

[0351] The techniques of this application may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or a group of ICs (e.g., a chipset). The various components, modules, or units described in this application are described to emphasize functional aspects of the devices for performing the disclosed techniques, but need not be implemented by distinct hardware units. In fact, as described above, the various units may be combined in a codec hardware unit with suitable software and / or firmware, or provided by interoperating hardware units including one or more processors as described above.

[0352] In the above embodiments, the descriptions of the various embodiments have their respective focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0353] As described above, the specific embodiments of this application are merely exemplary, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for estimating the time delay of a stereophonic audio signal, characterized in that, it includes: Obtain the current frame of the stereophonic audio signal, where the current frame includes a first-channel audio signal and a second-channel audio signal; If the signal type of the noise signal included in the current frame is a correlated noise signal type, then use a first algorithm to estimate the inter-channel time difference between the first-channel audio signal and the second-channel audio signal; If the signal type of the noise signal included in the current frame is a diffuse noise signal type, then use a second algorithm to estimate the inter-channel time difference between the first-channel audio signal and the second-channel audio signal; Wherein, the first algorithm includes weighting the cross-power spectrum in the frequency domain of the current frame using a first weighting function, and the second algorithm includes weighting the cross-power spectrum in the frequency domain of the current frame using a second weighting function, and the construction factors of the first weighting function and the second weighting function are different; The construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and the coherence square value of the current frame; The construction factors of the second weighting function include: an amplitude weighting parameter and the coherence square value of the current frame.

2. The method according to claim 1, characterized in that, after obtaining the current frame of the stereophonic audio signal, the method further includes: Obtain the noise coherence value of the current frame; If the noise coherence value is greater than or equal to a preset threshold, then determine that the signal type of the noise signal included in the current frame is the correlated noise signal type; If the noise coherence value is less than the preset threshold, then determine that the signal type of the noise signal included in the current frame is the diffuse noise signal type.

3. The method according to claim 2, characterized in that, obtaining the noise coherence value of the current frame includes: Perform voice activity detection on the current frame; If the detection result indicates that the signal type of the current frame is a noise signal type, then calculate the noise coherence value of the current frame; or, If the detection result indicates that the signal type of the current frame is a voice signal type, then determine the noise coherence value of the previous frame of the current frame in the stereophonic audio signal as the noise coherence value of the current frame.

4. The method according to claim 1, characterized in that, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; using the first algorithm to estimate the inter-channel time difference between the first-channel audio signal and the second-channel audio signal includes: Perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; Calculate the cross-power spectrum in the frequency domain of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; Weight the cross-power spectrum in the frequency domain using the first weighting function; Obtain an estimated value of the inter-channel time difference according to the weighted cross-power spectrum in the frequency domain.

5. The method according to claim 1, It is characterized in that the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; the step of estimating the inter-channel time difference between the first-channel audio signal and the second-channel audio signal by using the first algorithm includes: calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weighting the frequency-domain cross-power spectrum by using the first weighting function; obtaining an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum.

6. The method according to claim 4 or 5, It is characterized in that The first weighting function Φ new_1 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], and W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), and Γ 2 (k) is the squared coherence value of the k-th frequency bin of the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT - 1, and N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

7. The method according to claim 4 or 5, It is characterized in that The first weighting function Φ new_1 (k) satisfies the following formula: Among them, β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the squared coherence value of the k-th frequency point in the current frame, k is the frequency-point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

8. The method according to claim 4 or 5, It is characterized in that the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; after obtaining the current frame of the stereophonic audio signal, the method further includes: obtaining an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determining the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtaining an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; and determining the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

9. The method according to claim 8, It is characterized in that The first initial Wiener gain factor satisfies the following formula: The second initial Wiener gain factor satisfies the following formula: Among them, is the estimated value of the noise power spectrum of the first channel, is the estimated value of the noise power spectrum of the second channel; X 1 (k) is the frequency-domain signal of the first channel, X 2 (k) is the frequency-domain signal of the second channel, k is the frequency point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after the current frame is subjected to time-frequency transformation.

10. The method according to claim 4 or 5, It is characterized in that the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; after obtaining the current frame of the stereophonic audio signal, the method further includes: obtaining the first initial Wiener gain factor of the first-channel frequency-domain signal and the second initial Wiener gain factor of the second-channel frequency-domain signal; constructing a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; constructing a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

11. The method according to claim 10, It is characterized in that The first improved Wiener gain factor satisfies the following formula: The second improved Wiener gain factor satisfies the following formula: Among them, μ 0 is the binary masking threshold of the Wiener gain factor, is the first initial Wiener gain factor; is the second initial Wiener gain factor.

12. The method according to claim 1, It is characterized in that the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the step of estimating the inter-channel time difference between the first-channel audio signal and the second-channel audio signal by using the second algorithm includes: performing time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; The frequency-domain cross-power spectrum is weighted using the second weighting function to obtain an estimated value of the inter-channel time difference.

13. The method according to claim 1, wherein, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; the estimating the inter-channel time difference between the first-channel audio signal and the second-channel audio signal using the second algorithm includes: calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weighting the frequency-domain cross-power spectrum using the second weighting function; obtaining an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum.

14. The method according to claim 12 or 13, wherein, The second weighting function Φ new_2 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], X 1 (k) is the frequency-domain signal of the first channel, X 2 (k) is the frequency-domain signal of the second channel, is the conjugate function of X 2 (k), Γ 2 (k) is the coherence square value of the k-th frequency point in the current frame, k is the frequency point index value, k = 0, 1, …, N DFT - 1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

15. A method for estimating the time delay of a stereophonic audio signal, wherein, comprises: obtaining a current frame in the stereophonic audio signal, the current frame including a first-channel audio signal and a second-channel audio signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel audio signal and the second-channel audio signal; weighting the frequency-domain cross-power spectrum using a preset weighting function, the preset weighting function being a first weighting function or a second weighting function; obtaining an estimated value of the inter-channel time difference between the first-channel frequency-domain signal and the second-channel frequency-domain signal according to the weighted frequency-domain cross-power spectrum; wherein, if the signal type of the noise signal included in the current frame is a correlated noise signal type, the first weighting function is used to weight the frequency-domain cross-power spectrum; if the signal type of the noise signal included in the current frame is a diffuse noise signal type, the second weighting function is used to weight the frequency-domain cross-power spectrum; the construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and the coherence square value of the current frame; the construction factors of the second weighting function include: an amplitude weighting parameter and the coherence square value of the current frame; the construction factors of the first weighting function and the second weighting function are different.

16. The method according to claim 15, wherein, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel audio signal and the second-channel audio signal includes: performing time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculating the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal.

17. The method according to claim 15, wherein, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal.

18. The method according to any one of claims 15 to 16, wherein, The first weighting function Φ new_1 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the squared coherence value of the k-th frequency point in the current frame, k is the frequency-point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

19. The method according to any one of claims 15 to 16, characterized in that, the first weighting function Φ new_1 (k) satisfies the following formula: Among them, β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the coherence square value of the k-th frequency point in the current frame, k is the frequency-point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

20. The method according to any one of claims 15 to 17, characterized in that, the Wiener gain factor corresponding to the first channel frequency-domain signal is the first initial Wiener gain factor of the first channel frequency-domain signal, and the Wiener gain factor corresponding to the second channel frequency-domain signal is the second initial Wiener gain factor of the second channel frequency-domain signal; after obtaining the current frame in the stereo audio signal, the method further includes: obtaining an estimated value of the first channel noise power spectrum according to the first channel frequency-domain signal; determining the first initial Wiener gain factor according to the estimated value of the first channel noise power spectrum; obtaining an estimated value of the second channel noise power spectrum according to the second channel frequency-domain signal; and determining the second initial Wiener gain factor according to the estimated value of the second channel noise power spectrum.

21. The method according to claim 20, characterized in that, The first initial Wiener gain factor satisfies the following formula: The second initial Wiener gain factor satisfies the following formula: Among them, is the estimated value of the noise power spectrum of the first channel, is the estimated value of the noise power spectrum of the second channel; X 1 (k) is the frequency-domain signal of the first channel, X 2 (k) is the frequency-domain signal of the second channel, k is the frequency point index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency points after the current frame undergoes time-frequency transformation.

22. The method according to any one of claims 15 to 17, characterized in that, the Wiener gain factor corresponding to the first channel frequency-domain signal is the first improved Wiener gain factor of the first channel frequency-domain signal, and the Wiener gain factor corresponding to the second channel frequency-domain signal is the second improved Wiener gain factor of the second channel frequency-domain signal; after obtaining the current frame in the stereo audio signal, the method further includes: obtaining the first initial Wiener gain factor of the first channel frequency-domain signal and the second initial Wiener gain factor of the second channel frequency-domain signal; constructing a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; constructing a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

23. The method according to claim 22, characterized in that, The first improved Wiener gain factor satisfies the following formula: The second improved Wiener gain factor satisfies the following formula: where, μ 0 is the binary masking threshold of the Wiener gain factor, is the first initial Wiener gain factor; is the second initial Wiener gain factor.

24. The method according to any one of claims 15 to 17, 21, 23, characterized in that, The second weighting function Φ new_2 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor of the first channel; W x2 (k) is the Wiener gain factor of the second channel; X 1 (k) is the frequency-domain signal of the first channel, X 2 (k) is the frequency-domain signal of the second channel, is the conjugate function of X 2 (k), Γ 2 (k) is the coherence squared value of the k-th frequency point in the current frame, k is the frequency point index value, k = 0, 1, …, N DFT - 1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

25. A stereo audio signal time delay estimation device, characterized in that, comprising: a first obtaining module, configured to obtain a current frame of a stereo audio signal, where the current frame includes a first channel audio signal and a second channel audio signal; a first inter-channel time difference estimation module, configured to, if the signal type of the noise signal included in the current frame is a correlated noise signal type, estimate the inter-channel time difference between the first channel audio signal and the second channel audio signal by using a first algorithm; if the signal type of the noise signal included in the current frame is a diffuse noise signal type, estimate the inter-channel time difference between the first channel audio signal and the second channel audio signal by using a second algorithm; wherein, the first algorithm includes weighting the frequency-domain cross-power spectrum of the current frame by using a first weighting function, the second algorithm includes weighting the frequency-domain cross-power spectrum of the current frame by using a second weighting function, and the construction factors of the first weighting function and the second weighting function are different; The construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain factor corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and the squared coherence value of the current frame; The construction factors of the second weighting function include: an amplitude weighting parameter and the squared coherence value of the current frame.

26. The apparatus according to claim 25, wherein, the apparatus further includes: a noise coherence value calculation module, configured to obtain the noise coherence value of the current frame after the first obtaining module obtains the current frame; if the noise coherence value is greater than or equal to a preset threshold, determining that the signal type of the noise signal included in the current frame is the correlated noise signal type; or, if the noise coherence value is less than the preset threshold, determining that the signal type of the noise signal included in the current frame is the diffuse noise signal type.

27. The apparatus according to claim 26, wherein, the apparatus further includes: a voice endpoint detection module, configured to perform voice endpoint detection on the current frame; the noise coherence value calculation module is specifically configured to calculate the noise coherence value of the current frame if the detection result indicates that the signal type of the current frame is the noise signal type; or, if the detection result indicates that the signal type of the current frame is the voice signal type, determining the noise coherence value of the previous frame of the current frame in the stereophonic audio signal as the noise coherence value of the current frame.

28. The apparatus according to claim 25, wherein, the first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the first inter-channel time difference estimation module is configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using the first weighting function; and obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum.

29. The apparatus according to claim 25, wherein, the first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; the first inter-channel time difference estimation module is configured to calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using the first weighting function; and obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum.

30. The apparatus according to claim 28 or 29, wherein, The first weighting function Φ new_1 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], and W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), and Γ 2 (k) is the squared coherence value of the k-th frequency bin of the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT - 1, and N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

31. The apparatus according to claim 28 or 29, wherein, The first weighting function Φ new_1 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the coherence squared value of the k-th frequency bin in the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

32. The apparatus according to claim 28 or 29, wherein, The Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; The first inter-channel time difference estimation module is specifically configured to, after the first acquisition module acquires the current frame, obtain an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determine the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtain an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; and determine the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

33. The apparatus according to claim 32, wherein, The first initial Wiener gain factor satisfies the following formula: The second initial Wiener gain factor satisfies the following formula: wherein, is the estimated value of the noise power spectrum of the first channel, is the estimated value of the noise power spectrum of the second channel; X 1 (k) is the frequency-domain signal of the first channel, X 2 (k) is the frequency-domain signal of the second channel, k is the frequency point index value, k = 0, 1, …, N DFT -1, and N DFT is the total number of frequency points after the current frame undergoes time-frequency transformation.

34. The apparatus according to claim 28 or 29, wherein, The Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; The first inter-channel time difference estimation module is specifically configured to, after the first acquisition module acquires the current frame, obtain the first initial Wiener gain factor of the first-channel frequency-domain signal and the second initial Wiener gain factor of the second-channel frequency-domain signal; construct a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; and construct a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

35. The apparatus according to claim 34, wherein, The first improved Wiener gain factor satisfies the following formula: The second improved Wiener gain factor satisfies the following formula: Among them, μ 0 is the binary masking threshold of the Wiener gain factor, is the first initial Wiener gain factor; is the second initial Wiener gain factor.

36. The apparatus according to any one of claims 25 to 29, 33, and 35, wherein, The first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the first inter-channel time difference estimation module is specifically configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; and weight the frequency-domain cross-power spectrum by using the second weighting function to obtain an estimated value of the inter-channel time difference.

37. The apparatus according to any one of claims 25 to 29, 33, and 35, wherein, The first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal; the first inter-channel time difference estimation module is specifically configured to calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal; weight the frequency-domain cross-power spectrum by using the second weighting function; and obtain an estimated value of the inter-channel time difference according to the weighted frequency-domain cross-power spectrum.

38. The apparatus according to claim 37, wherein, The second weighting function Φ new_2 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the coherence square value of the k-th frequency point in the current frame, k is the frequency point index value, k = 0, 1, …, N DFT - 1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

39. A stereophonic audio signal time delay estimation device, characterized in that, it includes: A second acquisition module, configured to acquire a current frame in the stereophonic audio signal, where the current frame includes a first-channel audio signal and a second-channel audio signal; A second inter-channel time difference estimation module, configured to calculate a frequency-domain cross-power spectrum of the current frame according to the first-channel audio signal and the second-channel audio signal; weight the frequency-domain cross-power spectrum by using a preset weighting function, where the preset weighting function is a first weighting function or a second weighting function; obtain an estimated value of the inter-channel time difference between the first-channel frequency-domain signal and the second-channel frequency-domain signal according to the weighted frequency-domain cross-power spectrum; wherein, if the signal type of the noise signal included in the current frame is a correlated noise signal type, the first weighting function is used to weight the frequency-domain cross-power spectrum; if the signal type of the noise signal included in the current frame is a diffuse noise signal type, the second weighting function is used to weight the frequency-domain cross-power spectrum; The construction factors of the first weighting function include: the Wiener gain factor corresponding to the first-channel frequency-domain signal, the Wiener gain corresponding to the second-channel frequency-domain signal, an amplitude weighting parameter, and the coherence square value of the current frame; the construction factors of the second weighting function include: an amplitude weighting parameter and the coherence square value of the current frame; the construction factors of the first weighting function and the second weighting function are different.

40. The device according to claim 39, characterized in that, The first-channel audio signal is a first-channel time-domain signal, and the second-channel audio signal is a second-channel time-domain signal; the second inter-channel time difference estimation module is configured to perform time-frequency transformation on the first-channel time-domain signal and the second-channel time-domain signal to obtain a first-channel frequency-domain signal and a second-channel frequency-domain signal; calculate the frequency-domain cross-power spectrum of the current frame according to the first-channel frequency-domain signal and the second-channel frequency-domain signal.

41. The device according to claim 39, characterized in that, The first-channel audio signal is a first-channel frequency-domain signal, and the second-channel audio signal is a second-channel frequency-domain signal.

42. The device according to any one of claims 39 to 41, characterized in that, The first weighting function Φ new_1 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], and W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the squared coherence value of the k-th frequency bin in the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT - 1, and N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

43. The device according to any one of claims 39 to 41, characterized in that, the first weighting function Φ new_1 (k) satisfies the following formula: where β is the amplitude weighting parameter, β ∈ [0, 1], W x1 (k) is the Wiener gain factor corresponding to the first-channel frequency-domain signal; W x2 (k) is the Wiener gain factor corresponding to the second-channel frequency-domain signal; X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the squared coherence value of the k-th frequency bin in the current frame, k is the frequency bin index value, k = 0, 1, …, N DFT -1, N DFT is the total number of frequency bins after time-frequency transformation of the current frame.

44. The device according to any one of claims 39 to 41, characterized in that, The Wiener gain factor corresponding to the first-channel frequency-domain signal is the first initial Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second initial Wiener gain factor of the second-channel frequency-domain signal; The second inter-channel time difference estimation module is specifically configured to, after the second acquisition module acquires the current frame, obtain an estimated value of the first-channel noise power spectrum according to the first-channel frequency-domain signal; determine the first initial Wiener gain factor according to the estimated value of the first-channel noise power spectrum; obtain an estimated value of the second-channel noise power spectrum according to the second-channel frequency-domain signal; and determine the second initial Wiener gain factor according to the estimated value of the second-channel noise power spectrum.

45. The apparatus according to claim 44, wherein, The first initial Wiener gain factor satisfies the following formula: The second initial Wiener gain factor satisfies the following formula: Among them, is the estimated value of the noise power spectrum of the first channel, is the estimated value of the noise power spectrum of the second channel; X 1 (k) is the frequency-domain signal of the first channel, X 2 (k) is the frequency-domain signal of the second channel, k is the frequency point index value, k = 0, 1, …, N DFT -1, and N DFT is the total number of frequency points after the current frame undergoes time-frequency transformation.

46. The apparatus according to any one of claims 39 to 41, wherein, the Wiener gain factor corresponding to the first-channel frequency-domain signal is the first improved Wiener gain factor of the first-channel frequency-domain signal, and the Wiener gain factor corresponding to the second-channel frequency-domain signal is the second improved Wiener gain factor of the second-channel frequency-domain signal; The second inter-channel time difference estimation module is specifically configured to, after the second acquisition module acquires the current frame, obtain the first initial Wiener gain factor of the first-channel frequency-domain signal and the second initial Wiener gain factor of the second-channel frequency-domain signal; construct a binary masking function for the first initial Wiener gain factor to obtain the first improved Wiener gain factor; and construct a binary masking function for the second initial Wiener gain factor to obtain the second improved Wiener gain factor.

47. The apparatus according to claim 46, wherein, The first improved Wiener gain factor satisfies the following formula: The second improved Wiener gain factor satisfies the following formula: where μ 0 is the binary masking threshold of the Wiener gain factor, and is the first initial Wiener gain factor; is the second initial Wiener gain factor.

48. The apparatus according to any one of claims 39 to 41, 45, and 47, wherein, The second weighting function Φ new_2 (k) satisfies the following formula: where γ ∈ [0, 1], X 1 (k) is the first-channel frequency-domain signal, X 2 (k) is the second-channel frequency-domain signal, is the conjugate function of X 2 (k), Γ 2 (k) is the coherence square value of the k-th frequency point in the current frame, k is the frequency-point index value, k = 0, 1, …, N DFT - 1, N DFT is the total number of frequency points after time-frequency transformation of the current frame.

49. An audio coding apparatus, wherein, comprises: a non-volatile memory and a processor coupled to each other, and the processor calls program code stored in the memory to execute the stereo audio signal time delay estimation method described in any one of claims 1 to 24.

50. A computer storage medium, wherein, comprises a computer program, and when the computer program is executed on a computer, the computer is caused to execute the stereo audio signal time delay estimation method described in any one of claims 1 to 24.

51. A computer-readable storage medium, wherein, comprises an encoded bitstream, and the encoded bitstream includes the inter-channel time difference of the stereo audio signal obtained according to the stereo audio signal time delay estimation method described in any one of claims 1 to 24.

Citation Information

Patent Citations

  • Apparatus, method or computer program for estimating an inter-channel time difference

    WO2019193070A1

Cited By

  • Method and apparatus for estimating time delay of stereo audio signal

    WO2022012629A1