Adaptive window direct sound extraction method and device, electronic equipment and storage medium

By constructing an adaptive window function in direct sound extraction, adjusting the window length and coefficient according to frequency changes, the problem of window adjustment in the prior art is solved, and a more efficient direct sound extraction and analysis effect is achieved.

CN120220724APending Publication Date: 2025-06-27LINKPLAY TECHNOLOGY INC NANJING
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510459800.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art cannot realize adaptive adjustment of window functions according to the difference between high and low frequency signals in direct sound extraction, resulting in unsatisfactory processing effects, especially in reverb environments or multi-sound source scenarios.

Method used

By introducing frequency dependency factors, an adaptive window function is constructed so that its window length and window coefficient are adaptively adjusted with the frequency change. Use a longer window to ensure frequency resolution for low-frequency signals, and use a shorter window to improve time accuracy for high-frequency signals.

Benefits of technology

The optimal balance of time-frequency analysis is achieved, so that the extracted direct sound can retain more details, significantly improve signal analysis accuracy, and maintain stable performance in different acoustic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220724A_ABST
    Figure CN120220724A_ABST
Patent Text Reader

Abstract

The invention provides an adaptive window direct sound extraction method and device, electronic equipment and a storage medium, and the method comprises the steps: collecting an input audio signal, and carrying out the preprocessing of the audio signal; a self-adaptive window function is constructed based on the frequency difference of the audio signals, and the window length and the window coefficient of the self-adaptive window function can be automatically adjusted according to different frequency bands; the method comprises the following steps: converting an audio signal from a time domain to a frequency domain for representation through an adaptive window function, obtaining complete frequency domain data of the audio signal, using a longer window for a low-frequency signal by the adaptive window function so as to ensure the frequency resolution of a frequency domain signal, and using a shorter window for a high-frequency signal so as to improve the time precision of the frequency domain signal; and establishing a direct sound extraction tool according to the complete frequency domain data of the audio signal, and separating and extracting a direct sound signal from the audio signal. According to the method, the special window functions are designed for the signals with different frequencies, so that more details are reserved for the extracted direct sound, and the signal analysis precision is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio signal processing, and particularly relates to an adaptive window direct sound extraction method, device, equipment and storage medium. Background Art

[0002] In the field of audio signal processing, direct sound extraction is a key technology for separating the original sound source from secondary sounds such as environmental reflections and reverberations. Traditional methods mainly use the short-time Fourier transform (STFT) combined with a fixed window function for time-frequency analysis, and then use thresholds or statistical models for direct sound identification and extraction. However, the above methods face a dilemma in window length selection: a too long window will lead to a decrease in time resolution and make it difficult to accurately locate transient signals; a too short window will result in insufficient frequency resolution, especially for low-frequency signals. In addition, fixed window parameters cannot simultaneously meet the processing requirements of different frequency components. Low-frequency signals require a longer window to obtain sufficient frequency resolution, while high-frequency signals require a shorter window to retain time characteristics. Existing technologies usually adopt a compromise window length, resulting in unsatisfactory processing effects, especially in reverberant environments or multi-source scenarios. At the same time, traditional algorithms have high computational complexity and are difficult to achieve real-time processing on resource-constrained devices, limiting their applications in fields such as mobile devices and Internet of Things devices. Summary of the Invention

[0003] The present invention provides an adaptive window direct sound extraction method, device, equipment and storage medium to solve the technical problem that in the prior art, conventional direct sound extraction methods cannot adaptively adjust the window length and window coefficient of the window function according to the differences between high-frequency and low-frequency signals.

[0004] To solve the above problems, the technical solution of the present invention is: an adaptive window direct sound extraction method, including: Collecting an input audio signal and performing preprocessing on the audio signal; Based on the frequency difference of the audio signal, constructing an adaptive window function, where the adaptive window function is configured to automatically adjust its window length and window coefficient according to different frequency bands; Converting the audio signal from the time domain to the frequency domain representation through the adaptive window function to obtain the complete frequency domain data of the audio signal, where the adaptive window function uses a longer window for low-frequency signals to ensure the frequency resolution of the frequency domain signal, and uses a shorter window for high-frequency signals to improve the time accuracy of the frequency domain signal; Establishing a direct sound extraction tool according to the complete frequency domain data of the audio signal, and separating and extracting the direct sound signal from the audio signal through the direct sound extraction tool.

[0005] Preferably, performing preprocessing on the audio signal includes: Adjust the amplitude of the time-domain signal of the audio signal to the standard interval through maximum absolute value normalization processing; Perform frame segmentation on the normalized audio signal to obtain audio frame data.

[0006] Preferably, constructing the adaptive window function includes: Select a basic window function; Set a frequency-dependent factor, which is used to control the adaptive adjustment of the window length and window coefficient with the change of frequency. The frequency-dependent factor can respectively generate frequency-dependent values adapted to the signal frequency within the range from the high-frequency segment to the low-frequency segment of the signal; Based on the basic window function and the frequency-dependent factor, construct the adaptive window function adapted to different frequency segments, so that the window length and window coefficient of the adaptive window function change continuously with frequency.

[0007] Preferably, the basic window function includes a Hanning window, a Hamming window, a Blackman window or a Kaiser window.

[0008] Preferably, obtaining the complete frequency-domain data of the audio signal includes: Calculate the window length and window coefficient of the window function according to the target frequency segment in the audio frame, and construct the adaptive window function adapted to the target frequency segment; Intercept the signal segment with the corresponding window length from the audio frame, apply the adaptive window function, and perform short-time Fourier transform processing; Calculate the frequency-domain data of the audio frame at the target frequency segment; Integrate the frequency-domain data of all audio frames at different frequency segments in sequence to obtain the complete frequency-domain data of the audio signal.

[0009] Preferably, after obtaining the complete frequency-domain data of the audio signal, it includes: Perform smoothing processing on the frequency-domain data of the audio signal through the exponential moving average algorithm; Adjust the amplitude of the frequency-domain signal of the audio signal to the standard interval through maximum absolute value normalization processing.

[0010] Preferably, establishing the direct sound extraction tool according to the complete frequency-domain data of the audio signal includes: According to the characteristic pattern of the frequency-domain data of the audio signal, construct a direct sound judgment function, which is used to judge and distinguish the direct sound signal and the reverberation signal in the frequency-domain data of the audio signal; Based on the direct sound judgment function, establish a masking matrix, which is used to separate and extract the direct sound signal from the audio signal; Among them, the direct sound judgment function distinguishes the direct sound signal from the reverberation signal based on the transient peak characteristics of the frequency domain data of the audio signal.

[0011] Preferably, after obtaining the masking matrix, it includes: Select a basic smoothing window function; Calculate the window length of the window function according to the target frequency band, and construct a masking matrix smoothing processing function adapted to the target frequency band; Perform a convolution operation on the masking matrix in different frequency bands through the masking matrix smoothing processing function.

[0012] Preferably, separating and extracting the direct sound signal from the audio signal through the direct sound extraction tool includes: Apply the masking matrix to the frequency domain data of the audio signal; Perform an inverse short-time Fourier transform processing with a corresponding window length on the target frequency band in the frequency domain data; Calculate and extract the time domain data of each audio frame at the target frequency band and that is a direct sound signal; Overlap and add the time domain data of the direct sound signal in each audio frame in sequence to obtain the complete direct sound data of the audio signal.

[0013] Preferably, after obtaining the complete direct sound data of the audio signal, it includes: Eliminate the DC bias in the direct sound signal through the de-mean algorithm; Use a short-time high-pass filter to filter out the low-frequency noise in the direct sound signal; Adjust the signal amplitude of the direct sound signal to the standard range through peak normalization processing.

[0014] Preferably, the adaptive window direct sound extraction method further includes: Perform adaptive adjustment on the setting parameters of the direct sound judgment function and the masking matrix, including: Dynamically adjust the discrimination threshold of the direct sound judgment function according to the signal quality characteristics of the input audio signal; Dynamically adjust the masking characteristic parameters of the masking matrix according to the peak characteristics of the input audio signal; Perform scenario-based adaptive adjustment on the setting parameters of the direct sound judgment function and the masking matrix according to the environmental characteristics of the audio signal; Among them, the signal quality characteristics include the signal-to-noise ratio, and the environmental characteristics include the reverberation time, background noise, and signal steady-state characteristics.

[0015] Preferably, the adaptive window direct sound extraction method further includes: Evaluate the performance of the direct sound judgment function and the masking matrix and perform negative feedback adjustment, including: Define a set of performance indicators for the direct sound judgment function and the masking matrix, where the set of performance indicators includes clarity, distortion, degree of direct sound retention, and comprehensive performance score; Periodically monitor the performance of the direct sound judgment function and the masking matrix. When the performance of the direct sound judgment function and / or the masking matrix is lower than a preset threshold, adaptively correct the setting parameters of the direct sound judgment function and / or the masking matrix.

[0016] Based on the same concept, the present invention also provides an adaptive window direct sound extraction device that executes the adaptive window direct sound extraction method described in any one of the above, including: An acquisition module for acquiring an input audio signal and performing preprocessing on the audio signal; A construction module for constructing an adaptive window function based on the frequency difference of the audio signal, where the adaptive window function is configured to automatically adjust its window length and window coefficient according to different frequency bands; A conversion module for converting the audio signal from the time domain to the frequency domain representation through the adaptive window function to obtain the complete frequency domain data of the audio signal. Among them, the adaptive window function uses a longer window for low-frequency signals to ensure the frequency resolution of the frequency domain signal, and uses a shorter window for high-frequency signals to improve the time accuracy of the frequency domain signal; An extraction module for establishing a direct sound extraction tool based on the complete frequency domain data of the audio signal and separating and extracting the direct sound signal from the audio signal through the direct sound extraction tool.

[0017] Based on the same concept, the present invention also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the adaptive window direct sound extraction method described in any one of the above.

[0018] Based on the same concept, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the adaptive window direct sound extraction method described in any one of the above.

[0019] Due to the above technical solutions, the present invention has the following advantages and positive effects compared with the prior art: The present invention provides an adaptive window direct sound extraction method, apparatus, device, and storage medium. After collecting the input audio signal, by introducing a frequency-dependent factor, an adaptive window function is constructed, and the window length and window coefficient of the adaptive window function are configured to be adaptively adjusted according to the frequency change. For low-frequency signals, a longer window is used to ensure frequency resolution, and for high-frequency signals, a shorter window is used to improve time accuracy, effectively solving the technical problem of processing conflicts among different frequency components. By designing a dedicated window function for different frequency signals, the optimal balance of time-frequency analysis is achieved, enabling more details to be retained in the extracted direct sound and significantly improving the signal analysis accuracy. In addition, the direct sound judgment function and masking matrix provided by the present invention can automatically adjust their parameters according to the characteristics of the acoustic environment, ensuring stable performance of the direct sound extraction function in different application scenarios. Moreover, through memory management optimization, computational efficiency optimization, and instruction-level parallel computing, the present invention reduces the computational complexity by approximately 30% compared with traditional methods, supports real-time demand processing functions, and establishes a software-level application through the C language, ensuring compatibility and stability for cross-platform use. The above beneficial effects make the present invention have broad application prospects in the fields of speech recognition, sound source localization, stereo imaging, and audio post-production, especially suitable for professional audio processing scenarios that require high-quality direct sound, providing a new technical path for audio signal processing.

[0020] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0021] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 Schematic diagram of an embodiment of the adaptive window direct sound extraction method provided by an embodiment of the present disclosure; Figure 2 Schematic diagram of another embodiment of the adaptive window direct sound extraction method provided by an embodiment of the present disclosure; Figure 3Schematic diagram of the adaptive window direct sound extraction device provided by the embodiments of the present disclosure; Figure 4 Schematic diagram of an electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0024] To make the objectives, technical solutions, and advantages of the embodiments more clear, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present disclosure.

[0025] The embodiments provide an adaptive window direct sound extraction method, device, electronic device, and storage medium. It can be applied to any scenario of direct sound extraction.

[0026] In one embodiment of the present disclosure, the adaptive window direct sound extraction method can run on a terminal device or a server. Among them, the terminal device can be a local terminal device. When the adaptive window direct sound extraction method runs on the server, the method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device.

[0027] For ease of understanding, the specific process of this embodiment will be described below. Please refer to Figure 1 In an embodiment of the adaptive window direct sound extraction method in this embodiment, the following steps are included: Step 101: Collect the input audio signal and perform preprocessing on the audio signal; In one implementation, performing preprocessing on the audio signal includes normalizing the audio signal and performing frame segmentation. By normalizing the signal amplitude of the audio signal, it is ensured that there will be no numerical overflow problems in the subsequent processing, and all signals can be processed on the same scale. In addition, the continuous audio signal is segmented into multiple shorter time segments (frames) for subsequent time-frequency analysis.

[0028] Step 102: Based on the frequency difference of the audio signal, construct an adaptive window function, and the adaptive window function is configured to automatically adjust its window length and window coefficient according to different frequency bands; In one implementation, constructing the adaptive window function includes selecting a basic window function and setting a frequency-dependent factor, and generating an adaptive window function that can adaptively adjust function parameters for different frequency bands by combining the basic window function and the frequency-dependent factor.

[0029] Among them, a frequency bin is used to represent the discretized frequency components in the frequency domain. Each section represented by the frequency bin corresponds to a specific frequency value, that is, it reflects the energy magnitude of the frequency component.

[0030] If the sampling frequency of the signal is Fs (unit: Hz) and the number of sampling points is N, then the frequency resolution (i.e., the frequency interval between two adjacent frequency bins) is: Δf = Fs / N (Hz) The actual frequency corresponding to each frequency bin is: f k = k * Δf (k = 0, 1, 2, …, N / 2) Where k is the index of the frequency bin, which is used to identify the position of each frequency bin. The index k will be directly associated with the specific frequency value.

[0031] Step 103: Convert the audio signal from the time domain to the frequency domain representation through an adaptive window function to obtain the complete frequency domain data of the audio signal. Among them, the adaptive window function uses a longer window for low-frequency signals to ensure the frequency resolution of the frequency domain signal, and uses a shorter window for high-frequency signals to improve the time accuracy of the frequency domain signal; In one implementation, obtaining the complete frequency domain data of the audio signal includes applying an adaptive window function adapted to its frequency bin to each audio frame, performing short-time Fourier transform processing to convert the time domain representation of the audio signal into the frequency domain representation, and then integrating the frequency domain data of all audio frames at different frequency bins in sequence to obtain the complete frequency domain data of the audio signal, and performing smoothing processing and normalization processing on the frequency domain data of the audio signal.

[0032] Step 104: Establish a direct sound extraction tool based on the complete frequency domain data of the audio signal, and separate and extract the direct sound signal from the audio signal through the direct sound extraction tool.

[0033] In one implementation, establishing a direct sound extraction tool based on the complete frequency domain data of the audio signal includes constructing a direct sound judgment function and a masking matrix. After performing smoothing processing on the masking matrix, apply the masking matrix to the frequency domain data of the audio signal, perform inverse short-time Fourier transform processing with the corresponding window length on the target frequency bin in the frequency domain data, calculate and extract the time domain data of each audio frame at the target frequency bin and which is the direct sound signal, and finally perform overlapping addition in sequence on the time domain data of the direct sound signal in each audio frame to obtain the complete direct sound data of the audio signal, and perform post-processing on the direct sound signal data.

[0034] In this embodiment, by introducing a frequency-dependent factor, an adaptive window function is constructed, and the window length and window coefficient of the adaptive window function are configured to be adaptively adjusted according to the signal frequency. When facing low-frequency signals, a longer window is automatically used to ensure the frequency resolution, and when facing high-frequency signals, a shorter window is automatically used to improve the time accuracy. This effectively solves the technical problem of processing conflicts among different frequency components. By designing dedicated window functions for different frequency signals, the optimal balance of time-frequency analysis is achieved, so that in the direct sound extraction operation, more details of the extracted direct sound can be retained, and the accuracy of signal analysis is significantly improved.

[0035] Please refer to Figure 2 , another embodiment of the direct sound extraction method with an adaptive window in this embodiment includes: 201. Adjust the amplitude of the time-domain signal of the audio signal to the standard interval, and perform frame segmentation on the normalized audio signal to obtain audio frame data.

[0036] In one implementation, first receive the input audio signal x(n), where n represents the sampling point index.

[0037] Perform maximum absolute value normalization processing on the audio signal. The maximum absolute value normalization processing method is: x_norm(n) = x(n) / max(|x(n)|) where x_norm(n) is the audio signal after normalization processing, x(n) is the original input signal, max(|x(n)|) is the maximum absolute value in the entire audio signal x(n), and max(|x(n)|) can be obtained by pre-traversing the original input signal x(n).

[0038] For example, for an audio signal with a sampling rate of 44.1 kHz, first read the maximum absolute value in the entire audio signal, and then divide the value of each sampling point by this maximum absolute value. For example, if the maximum absolute value of the original signal is 0.8, then the value of each sampling point is divided by 0.8 for normalization processing to obtain the normalized audio signal.

[0039] By performing maximum absolute value normalization processing on the audio signal, it can be ensured that the signal amplitudes of all processed audio signals are within a suitable range, and the maximum absolute value of the signal is 1, avoiding numerical overflow problems in subsequent processing, and enabling all signals to be compared or processed on the same scale.

[0040] In one implementation, the normalized audio signal is segmented into frames according to a certain frame length and frame shift to generate a series of audio frames.

[0041] The implementation method of the frame segmentation process is: x_frame(m,n) = x_norm(n + mR) where m is the frame index; R is the frame shift; n is the sample index within the frame, n ∈ [0, N - 1]; N is the frame length; and x_frame(m,n) is the sample value with index n in the m-th frame.

[0042] For example, for an audio signal with a sampling rate of 44.1 kHz, the frame length can be selected as 2048 samples (about 46.4 ms), and the frame shift as 512 samples (about 11.6 ms). Therefore, there is an overlap of 1536 samples between every two frames, that is, there is an overlap of 75% (1536 / 2048) between adjacent frames, ensuring the continuity and smoothness of subsequent time-frequency analysis and being suitable for capturing the signal characteristics that change rapidly in the audio signal.

[0043] 202. Select a basic window function, set the frequency-dependent factor, and construct an adaptive window function adapted to different frequency bands.

[0044] In one implementation, first select a basic window function, and the basic window function includes any one of the Hanning window, Hamming window, Blackman window, or Kaiser window.

[0045] The general expression of the basic window function is: w(n, α) = f(n / N, α) where f represents the mathematical form of the window function, α is the shape parameter of the window function, N is the window length, and w(n,α) is the window function value at index n.

[0046] For example, if the Hanning window is selected as the basic window function, the specific expression of the basic window function is: w_hann(n) = 0.5 * (1 - cos(2πn / (N - 1))), n ∈ [0, N - 1] Subsequently, set the frequency-dependent factor β(k). The frequency-dependent factor β(k) can respectively generate frequency-dependent values adapted to the signal frequency in the range from the high-frequency band to the low-frequency band of the signal. The design criterion of the frequency-dependent factor β(k) is to consider the perception characteristics of the human ear for different frequencies and the time-frequency characteristic differences of the direct sound at different frequencies.

[0047] The expression of the frequency-dependent factor β(k) is: β(k) = β_min + (β_max - β_min) * g(k / K) where β_min is the minimum value of the frequency-dependent factor, β_max is the maximum value of the frequency-dependent factor; g(k / K) is a custom mapping function; K is the total number of frequency bands, and k is the index of the frequency band.

[0048] It should be noted that since the human ear's perception of frequency is non - linear, in the low - frequency region, the human ear is more sensitive to changes in the frequency of the audio signal, while in the high - frequency region, the human ear is relatively less sensitive to changes in the frequency of the audio signal. Therefore, in order to make the frequency - dependent factor better match the perception characteristics of the human ear, g(k / K) usually adopts non - linear functions such as logarithmic or power functions.

[0049] In one implementation, the mapping function g(k / K) is expressed as: g(k / K) = log10(1 + 9 * k / K) / log10(10) For example, for the case where the length of the Fourier transform is 2048, the total number of frequency segments K = 1024 (considering symmetry), and it is preset that β_min = 0.3 and β_max = 0.9.

[0050] When k = 0 (the lowest frequency of the signal), g(0)=0, so β(0)=β_min = 0.3; when k = K (the highest frequency of the signal), g(1)=1, so β(K)=β_max = 0.9; in the intermediate frequency region, due to the non - linear characteristics of the logarithmic function, the growth rate of the frequency - dependent factor β(k) is faster in the low - frequency region and slower in the high - frequency region, that is, the frequency - dependent value of the frequency - dependent factor β(k) can be adaptively adjusted according to the frequency segment difference.

[0051] Finally, based on the basic window function and the frequency - dependent factor, design the calculation formulas for the window length and window coefficient of the window function, and construct an adaptive window function adapted to different frequency segments.

[0052] In one implementation, the expression of the adaptive window function is: w(n, k) = w_base(n, α(β(k))) where w_base() is the basic window function, and α(β(k)) is the window parameter determined by the frequency - dependent factor β(k).

[0053] For example, if an adaptive window function is designed based on the Hanning window, first calculate the window length N(k) of the window function corresponding to the target frequency segment. The calculation method of the window length N(k) is: N(k) = N_max * (1 - β(k)) + N_min * β(k) where N_max is the maximum window length, preset N_max = 2048, and N_min is the minimum window function, preset N_min = 512. Therefore, when the frequency - dependent value of the frequency - dependent factor β(k) changes, the window length N(k) is adjusted synchronously.

[0054] Subsequently, calculate the window coefficient α(k) of the window function corresponding to the target frequency band. The calculation method of the window coefficient α(k) is as follows: α(k) = 0.5 + 0.3 * β(k) That is, in this embodiment, the frequency-dependent factor is used to control the adaptive adjustment of the window length and window coefficient with the change of frequency. The frequency-dependent factor can respectively generate frequency-dependent values adapted to the signal frequency in the range from the high-frequency band to the low-frequency band of the signal.

[0055] Finally, based on the basic window function and the frequency-dependent factor, construct an adaptive window function adapted to different frequency bands, so that the window length and window coefficient of the adaptive window function change continuously with frequency. The expression of the adaptive window function is: w(n,k) = (1-α(k))+α(k) * cos(2πn / (N(k)-1) - π), n∈[0, N(k)-1] In this embodiment, through the design of the frequency-dependent factor, the adaptive window function can automatically adjust the window length and coefficient according to the frequency characteristics of the audio signal, thereby optimizing the effect of time-frequency analysis. A longer window can be used in the low-frequency part to improve the frequency resolution of the signal, while a shorter window is used in the high-frequency part to keep the signal with better time accuracy. At the same time, by adjusting the window length and window coefficient, the best balance can be found between sidelobe suppression and main lobe width, improving the signal processing efficiency.

[0056] 203. Calculate the frequency-domain data of the audio frame at the target frequency band, and sequentially integrate the frequency-domain data of all audio frames at different frequency bands to obtain the complete frequency-domain data of the audio signal.

[0057] In one implementation, calculate the window length and window coefficient of the current window function according to the target frequency band in the audio frame to be processed, and construct an adaptive window function adapted to the target frequency band.

[0058] Intercept the signal segment with the corresponding window length from the audio frame, apply the adaptive window function matching the frequency, and perform the short-time Fourier transform processing (FR-STFT).

[0059] The expression of the short-time Fourier transform processing is: X(m, k) = Σ[n=0 to N(k)-1] x_frame(m, n) * w(n, k) * e^(-j2πkn / N(k)) Where X(m, k) is the complex spectrum at the target frequency band corresponding to index k in the m-th frame, and x_frame(m, n) is the value of the n-th sample in the m-th frame.

[0060] Subsequently, calculate the frequency-domain data of the audio frame at the target frequency band.

[0061] Repeat the above steps until the frequency-domain data at all frequency bands in all audio frames are obtained.

[0062] Integrate the frequency-domain data of all audio frames at different frequency bands in sequence to obtain the complete frequency-domain data of the audio signal, and obtain the complete spectral representation of the audio signal.

[0063] 204. Perform smoothing processing and normalization processing on the frequency-domain data of the audio signal.

[0064] In one implementation, the frequency-domain data of the audio signal is smoothed by the exponential moving average algorithm. The implementation of the exponential moving average algorithm is as follows: X_smooth(m, k) = λ * X_smooth(m - 1, k) + (1 - λ) * |X(m, k)| Where Xsmooth(m,k) is the smoothed spectral amplitude at the target frequency band corresponding to index k in the m-th frame; λ is the smoothing factor, and the preset value ranges from 0.7 to 0.9; |X(m, k)| is the original spectral amplitude.

[0065] During the process of smoothing the frequency-domain data of the audio signal, first perform the initialization step, so that X_smooth(0, k) = |X(0, k)|, that is, for the first frame (m = 0), the original spectral amplitude is used as the smoothed spectral amplitude. For subsequent frames (m > 0), use the smoothing factor λ = 0.8 for smoothing calculation, that is, X_smooth(m, k) = 0.8 * X_smooth(m - 1, k) + 0.2 * |X(m, k)|.

[0066] Smoothing the frequency-domain data of the audio signal by the exponential moving average algorithm can effectively reduce the fluctuations in the signal spectrum and improve the stability of subsequent direct sound extraction.

[0067] In one implementation, the amplitude of the frequency-domain signal of the audio signal is adjusted to the standard interval by the maximum absolute value normalization processing. The implementation of the maximum absolute value normalization processing is as follows: X_norm(m, k) = |X(m, k)| / (X_smooth(m, k) + ε) Where X_norm(m, k) is the normalized spectral amplitude at the target frequency band corresponding to index k in the m-th frame; ε is a very small constant (such as 1e-10) used to prevent division-by-zero errors in the calculation.

[0068] The amplitude of the frequency-domain signal of the audio signal is adjusted to the standard interval through maximum absolute value normalization, so that different frequency components have similar weights in subsequent processing, avoiding certain frequency components from dominating due to large amplitudes, and thus improving the stability and accuracy of signal processing.

[0069] 205. Construct a direct sound judgment function and a masking matrix.

[0070] In one implementation, since the direct sound is usually represented as a transient peak in the signal spectrum, the energy mutation characteristic of the direct sound can be used to quickly distinguish the direct sound from the reverberation. Therefore, a direct sound judgment function can be constructed according to the characteristic pattern of the frequency-domain data of the audio signal. The direct sound judgment function is used to judge and distinguish the direct sound signal and the reverberation signal in the frequency-domain data of the audio signal. Among them, the direct sound judgment function distinguishes the direct sound signal and the reverberation signal based on the transient peak characteristic of the frequency-domain data of the audio signal.

[0071] The expression of the direct sound judgment function is: P_direct(m, k) = sigmoid(α * (X_norm(m, k) - θ)) Where, P_direct(m, k) is the probability that it belongs to the direct sound at the target frequency segment corresponding to index k in the m-th frame; sigmoid() is the activation function. In this embodiment, the preset activation function is an S-shaped activation function of sigmoid(x) = 1 / (1+e^(-x)); α is a parameter that controls the steepness of the probability curve; θ is the discrimination threshold.

[0072] For example, preset α = 5 and θ = 1.5. For the normalized audio signal spectrum data X_norm(m, k), according to the direct sound judgment function, when the spectrum value of the time-frequency point in the audio signal is significantly higher than 1.5, it is marked as belonging to the direct sound, and the direct sound judgment function outputs a higher prediction probability value; when the spectrum value of the time-frequency point in the audio signal is significantly lower than 1.5, it is marked as not belonging to the direct sound, and the direct sound judgment function outputs a lower prediction probability value.

[0073] Subsequently, based on the direct sound judgment function, a masking matrix is further established. The masking matrix is used to separate and extract the direct sound signal from the audio signal.

[0074] In one implementation, the expression of the masking matrix M(m, k) is: M(m, k) = P_direct(m, k)^γ Where, γ is a parameter that controls the softness and hardness of the masking. The larger γ is, the closer the masking effect is to binary (hard masking); the smaller γ is, the smoother the masking effect is (soft masking).

[0075] Among them, hard masking is applicable to scenarios where it is necessary to clearly distinguish between direct sound and reverberation, such as speech enhancement or noise reduction tasks, which require removing as much background noise and reverberation as possible. Soft masking is applicable to scenarios where more details need to be retained, such as music processing or audio analysis, where it is desired to minimize information loss and maintain the natural characteristics of the signal.

[0076] For example, assuming γ = 2, for each signal time-frequency point (m, k), according to its direct sound probability Pdirect(m, k), the corresponding masking value M(m, k) is calculated as M(m, k) = P_direct(m, k)^2. For example, the masking value of a time-frequency point with a high probability (P = 0.9) is 0.81, and the masking value of a time-frequency point with a low probability (P = 0.1) is 0.01.

[0077] In one implementation, it also includes performing smoothing processing on the masking matrix adapted to the frequency band, which is used to reduce artifacts in the subsequent direct sound extraction process and improve the quality of the finally extracted direct sound component, especially suitable for application scenarios that require high-precision separation.

[0078] First, a basic smoothing window function is selected, such as using a triangular window as the smoothing window.

[0079] Subsequently, according to the target frequency band, the window length of the window function is calculated. The calculation method of the window length L(k) is as follows: L(k) = round(L_max * (1 - β(k)) + L_min * β(k)) Among them, it is preset that L_max = 5 and L_min = 1.

[0080] The expression of the triangular window is: h(i, k) = 1 - |i| / (L(k)+1), i ∈ [-L(k), L(k)] According to the target frequency band, the window length of the window function is calculated, and a masking matrix smoothing processing function adapted to the target frequency band is constructed. The expression of the masking matrix smoothing processing function is: M_smooth(m, k) = Σ[i=-L(k) to L(k)] h(i, k) * M(m - i, k) Among them, M_smooth(m, k) is the smoothed masking matrix at the target frequency band corresponding to index k in the m-th frame, and h(i, k) is the basic smoothing window function.

[0081] Finally, the convolution operation is performed on the masking matrix in different frequency bands through the masking matrix smoothing processing function to generate a smoothed masking matrix.

[0082] 206. Extract the time-domain data of each audio frame at the target frequency band and that is the direct sound signal to obtain the complete direct sound data of the audio signal.

[0083] In one implementation, apply a masking matrix to the frequency-domain data of the audio signal. The reconstruction expression of the frequency-domain data and the masking matrix is: Y(m, k) = M_smooth(m, k) * X(m, k) Subsequently, perform the inverse short-time Fourier transform processing (FR-ISTFT) with the corresponding window length on the target frequency band in the reconstructed frequency-domain data, calculate and extract the time-domain data of each audio frame at the target frequency band and that is the direct sound signal. The expression of the inverse short-time Fourier transform processing is: y_frame(m,n)=(1 / N(k))*Σ[k=0 to N(k) / 2]|Y(m,k)|*e^(j∠Y(m,k))*e^(j2πkn / N(k)) where y_frame(m, n) is the time-domain signal of the reconstructed direct sound, and ∠Y(m,k) is the phase of Y(m,k).

[0084] Finally, overlap and add the time-domain data of the direct sound signal in each audio frame in sequence to obtain the complete direct sound data of the audio signal.

[0085] In one implementation, during the process of overlapping and adding the reconstructed time-domain signals of each frame to obtain the final complete direct sound signal, since different frequency bands use windows of different lengths, a special method needs to be used to process the process of frame overlap and addition. The implementation method of the overlap and addition processing is: y(n) = Σ[m] y_frame(m, n-mR) * s(n-mR) where y(n) is the finally reconstructed time-domain signal; s(n) is a synthesis window function used to smooth the transition between frames; and R is the frame shift.

[0086] In the process of overlapping and adding the reconstructed time-domain signals of each frame to obtain the final complete direct sound signal, first, a suitable window function s(n) is selected to smooth the transition between frames. For example, a sine window is selected, s(n) = sin(πn / R)^2, where n ∈ [0, R - 1]. Initialize the output signal y(n) as a vector of all zeros, and its length should be sufficient to accommodate the superposition results of all frames. For each frame m, perform the following steps: Calculate the starting position of this frame in the output signal, start = m * R; Add this frame signal to the output signal, y(start + n) = y_frame(m, n) * s(n), where n ∈ [0, R - 1]; Repeat the above steps for all frames to obtain the final complete reconstructed signal.

[0087] Through the above overlapping and adding steps, the processed frame signals can be effectively recombined into a complete direct sound signal, while ensuring smooth transition between frames and improving the quality and continuity of the signal.

[0088] 207. Eliminate the DC bias and low-frequency noise in the direct sound signal, and adjust the signal amplitude of the direct sound signal to the standard range.

[0089] In one implementation, the DC bias in the direct sound signal is eliminated by the de-mean algorithm. The expression of the de-mean algorithm is: y_proc(n) = y(n) - mean(y) where mean(y) is the average value of the entire signal y(n). The calculation method of the signal average value mean(y) is: mean_y = (1 / L) * Σ[n = 0 to L - 1] y(n) where L is the signal length.

[0090] The DC bias is a constant offset existing in the signal, usually caused by system errors or cumulative effects during the processing. Therefore, in this embodiment, by using the de-mean algorithm, the average value of the signal is calculated and subtracted from each sample, so as to eliminate the signal offset, make the mean value of the signal return to zero, and thus avoid potential distortion.

[0091] In one implementation, a short-time high-pass filter is used to filter out the low-frequency noise in the direct sound signal. The expression for processing the signal by the short-time high-pass filter is: y_hp(n) = highpass(y_proc(n), f_cutoff) where f_cutoff is the cut-off frequency of the high-pass filter.

[0092] A high-pass filter can be used to remove components below a certain frequency (cutoff frequency), and these low-frequency components are usually considered as noise signals (such as environmental noise or mechanical vibration), thereby further improving the quality of the finally extracted direct sound signal.

[0093] In one implementation, the signal amplitude of the direct sound signal is adjusted to the standard range through peak normalization processing. The implementation method of peak normalization processing is as follows: y_out(n) = y_hp(n) / max(|y_hp(n)|) * target_level Where, max(|y_hp(n)|) is the maximum absolute value of the signal after high-pass filtering, and target_level is the target gain level.

[0094] Through peak normalization processing, it is ensured that the dynamic range of the signal is within a suitable standard range, and the maximum absolute value of all signal samples is scaled to the specified target level, which helps to avoid clipping distortion caused by too strong signals, and at the same time ensures that the signal has sufficient dynamic range.

[0095] 208. Adaptive adjustment is made to the setting parameters of the direct sound judgment function and the masking matrix.

[0096] In one implementation, according to the signal quality characteristics of the input audio signal, the discrimination threshold of the direct sound judgment function is dynamically adjusted, and the signal quality characteristics include the signal-to-noise ratio.

[0097] Specifically, according to the signal energy and noise floor of the input audio signal, the signal-to-noise ratio of the input audio signal is calculated, the signal-to-noise ratio is normalized, and when the signal-to-noise ratio exceeds the set threshold, the discrimination standard intensity of the direct sound judgment function for the direct sound signal and the reverberation signal is adaptively adjusted.

[0098] That is, first, the signal energy of the input audio signal is calculated. The calculation method of the signal energy E_signal is as follows: E_signal = Σ[n] x(n)^2 The noise floor of the input audio signal is calculated. The calculation method of the noise floor noise_floord is as follows: noise_floor = percentile(|x(n)|^2, 10) Subsequently, the signal-to-noise ratio of the input audio signal is calculated. The calculation method of the signal-to-noise ratio SNR is as follows: SNR = 10 * log10(E_signal / (noise_floor * signal_length) Among them, signal_length is the signal length.

[0099] Perform normalization processing on the signal-to-noise ratio SNR. The normalization processing method is as follows: SNR_norm = min(max(SNR / 30, 0), 1) Finally, adjust the discrimination threshold θ of the direct sound judgment function based on the signal-to-noise ratio SNR. The adjustment method is as follows: θ_adapt = θ_base + Δθ * (1 - SNR_norm) Among them, θ_base is the initial discrimination threshold, θ_adapt is the discrimination threshold after adaptive adjustment, and Δθ is the preset adjustment coefficient.

[0100] Therefore, in this embodiment, when the signal-to-noise ratio SNR is high, the discrimination threshold θ_adapt can be reduced so that the direct sound judgment function can more accurately identify the direct sound. Conversely, the discrimination threshold θ_adapt is increased to enhance the robustness of the direct sound judgment function.

[0101] In one implementation, according to the peak characteristics of the input audio signal, the masking characteristic parameters of the masking matrix are dynamically adjusted.

[0102] Specifically, calculate the peak factor of the audio signal according to the root mean square energy of the input audio signal, perform normalization processing on the peak factor of the audio signal, and adaptively adjust the masking softness and hardness parameters of the masking matrix according to the dynamic change range of the peak factor of the audio signal.

[0103] That is, first, calculate the root mean square energy of the input audio signal, which is used to represent the effective energy level of the signal. The calculation method of the root mean square energy RMS of the signal is as follows: E_rms = sqrt((1 / L) * Σ[n] x(n)^2) Subsequently, calculate the peak factor of the audio signal according to the root mean square energy of the input audio signal. The calculation method of the peak factor PF is as follows: PF = max(|x(n)|) / E_rms Perform normalization processing on the peak factor of the audio signal. The normalization processing method is as follows: DR_norm = min(max((PF - 3) / 10, 0), 1) Finally, according to the dynamic change range of the peak factor of the audio signal, adaptively adjust the masking softness and hardness parameter γ of the masking matrix. The adjustment method is as follows: γ_adapt = γ_base * (1 + Δγ * DR_norm) Among them, γ_base is the initial masking softness / hardness parameter, γ_adapt is the masking softness / hardness parameter after adaptive adjustment, and Δγ is the preset adjustment coefficient.

[0104] In this embodiment, the setting parameters of the masking matrix have an adaptive adjustment function. When dealing with signals with a large dynamic range, the anti-interference effect of the masking matrix can be significantly enhanced.

[0105] In one implementation, according to the environmental characteristics of the audio signal, scene-based adaptive adjustment is performed on the direct sound judgment function and the setting parameters of the masking matrix. The environmental characteristics include reverberation time, background noise, and signal steady-state characteristics.

[0106] Specifically, extract the acoustic environmental characteristics, evaluate the environment where the audio signal is located according to the reverberation time, background noise, and signal steady-state characteristic data. The environmental classification includes large-space indoor environment, small-space indoor environment, noisy environment, and general environment. According to the environmental classification, perform adaptive adjustment on the direct sound judgment function and the setting parameters of the masking matrix. The computer implementation method is as follows: Environmental type detection: env_type = detect_environment(x(n)); Select parameter set based on environmental type: params = select_params(env_type); Specifically, first, perform environmental feature extraction. Obtain the reverberation time in the input audio signal: RT60 = estimate_rt60(x(n)); obtain the background noise level: noise_level = estimate_noise(x(n)); obtain the signal steady-state characteristics: steadiness = estimate_steadiness(x(n)); Subsequently, evaluate the environment where the audio signal is located according to the reverberation time, background noise, and signal steady-state characteristic data. If RT60 > 0.8s and noise_level < -30dB, then determine that the audio environment is a "large-space indoor environment"; if RT60 < 0.3s and noise_level < -25dB, then determine that the audio environment is a "small-space indoor environment"; if noise_level > -20dB and steadiness > 0.7, then determine that the audio environment is a "noisy environment"; in other cases, determine that the audio environment is a "general environment".

[0107] Finally, adaptively adjust the setting parameters of the direct sound judgment function and the masking matrix according to the environmental classification. In the "large - space indoor environment", automatically configure the minimum value of the frequency - dependent factor β_min = 0.2, the maximum value of the frequency - dependent factor β_max = 0.85, the discrimination threshold θ of the direct sound judgment function = 1.7, and the masking soft - hard degree parameter γ of the masking matrix = 1.5; in the "small - space indoor environment", automatically configure the minimum value of the frequency - dependent factor β_min = 0.3, the maximum value of the frequency - dependent factor β_max = 0.9, the discrimination threshold θ of the direct sound judgment function = 1.3, and the masking soft - hard degree parameter γ of the masking matrix = 2.0; in the "noisy environment", automatically configure the minimum value of the frequency - dependent factor β_min = 0.4, the maximum value of the frequency - dependent factor β_max = 0.95, the discrimination threshold θ of the direct sound judgment function = 2.0, and the masking soft - hard degree parameter γ of the masking matrix = 2.5; in the "general environment", automatically configure the minimum value of the frequency - dependent factor β_min = 0.3, the maximum value of the frequency - dependent factor β_max = 0.9, the discrimination threshold θ of the direct sound judgment function = 1.5, and the masking soft - hard degree parameter γ of the masking matrix = 2.0.

[0108] In this embodiment, adaptively adjust the setting parameters of the direct sound judgment function and the masking matrix through environmental classification, and select the optimal parameter combination of the direct sound judgment function and the masking matrix according to the characteristics of different acoustic environments, so as to improve the performance of the direct sound judgment function and the masking matrix in different application environments.

[0109] 209. Evaluate and perform negative - feedback regulation on the performance of the direct sound judgment function and the masking matrix.

[0110] In one implementation, define a set of performance indicators for the direct sound judgment function and the masking matrix. The set of performance indicators includes clarity, distortion, direct - sound retention degree, and comprehensive performance score. The computer implementation of defining the set of performance indicators is: metrics = {clarity, distortion, directness,...} Periodically monitor the performance of the direct sound judgment function and the masking matrix. The computer implementation of periodic monitoring is: performance = evaluate_performance(y_out, metrics) When the performance of the direct sound judgment function and / or the masking matrix is lower than the preset threshold, adaptively correct the setting parameters of the direct sound judgment function and / or the masking matrix. The computer implementation of triggering self - correction when the performance is lower than the threshold is: if performance < threshold: params = auto_correct(params, performance) Specifically, during the calculation of performance metrics, the method for obtaining clarity is: clarity = measure_clarity(y_out); the method for obtaining distortion is: distortion = measure_distortion(y_out, x); the method for obtaining the degree of direct sound retention is: directness = measure_directness(y_out); the method for calculating the comprehensive performance score is: performance = 0.4 * clarity + 0.3 * (1 - distortion) + 0.3 * directness.

[0111] During the periodic monitoring process, the performance score of the direct sound judgment function and the masking matrix is calculated every 100 frames of signals processed. If the performance scores of the direct sound judgment function and / or the masking matrix are lower than the trigger coefficient, such as 0.6, for three consecutive times, self-correction is triggered.

[0112] The self-correction strategy is specifically as follows: if clarity < 0.5, the value of the discrimination threshold θ of the direct sound judgment function is reduced by 10%; if the degree of direct sound retention distortion > 0.3, the value of the masking softness / hardness parameter γ of the masking matrix is reduced by 15%; if the distortion directness < 0.5, the range of the frequency-dependent factor β is adjusted, and the value of the maximum value β_max of the frequency-dependent factor is increased by 5%; after any of the above adjustments is completed, the performance counter is reset, and the periodic monitoring task is repeated.

[0113] By evaluating the performance of the direct sound judgment function and the masking matrix and performing negative feedback adjustment, the dynamic adaptation of the direct sound judgment function and the masking matrix to environmental changes can be achieved, ensuring stability and effectiveness during long-term operation, and improving the quality of the finally obtained direct sound signal.

[0114] In addition, in this embodiment, the computer programming language is selected as C language, and memory management optimization is realized by using C language programming technology to reduce memory occupation and fragmentation and improve the algorithm operation efficiency.

[0115] The memory optimization strategy includes using a pre-allocated buffer to replace dynamic memory allocation, pre-allocating the maximum required memory, and implementing memory pool management. Using a circular buffer to reduce data copying, and using a circular index during frame processing to avoid unnecessary data movement. The computer implementation method is as follows: / / Use static allocation instead of malloc float buffer[MAX_BUFFER_SIZE]; float* circularBuffer = buffer; int bufferHead = 0; / / Access elements in the circular buffer float getValue(int index) { return circularBuffer[(bufferHead + index) % MAX_BUFFER_SIZE]; } / / Update the buffer head pointer void advanceHead(int steps) { bufferHead = (bufferHead + steps) % MAX_BUFFER_SIZE; } Improving the algorithm's running efficiency includes pre - calculating the adaptive window function. Pre - calculate the adaptive window function suitable for common frequency signals, and use interpolation to obtain the adaptive window function for intermediate frequencies.

[0116] In addition, it is also possible to improve the calculation efficiency, reduce the amount of computation and processing delay through algorithm optimization and instruction - level optimization.

[0117] Calculation optimization strategies include: using the FFT library to achieve efficient spectrum analysis, such as using optimized libraries like FFTW and KISS FFT, and using the symmetry of FFT to reduce the amount of computation. Implement vectorized calculation, such as using SIMD instructions for parallel calculation and optimizing the critical operation path.

[0118] Specifically, when reducing FFT calculation, use the KISS FFT library for transform calculation, group the frequency segments to avoid repeated calculation, and at the same time use the characteristics of real - signal FFT to reduce the amount of computation. During the process of performing vectorized calculation, use vectorized calculation for the application process of the adaptive window function and also for the application process of the masking matrix.

[0119] At the same time, this embodiment also realizes the design of cross - platform compatibility. The cross - platform design includes: using standard C language functions, avoiding using specific APIs of some platforms, such as using standard library functions and handling platform differences through conditional compilation. At the same time, implement a platform abstraction layer, such as encapsulating platform - specific functions and providing a unified interface.

[0120] Specifically, the computer implementation method of cross - platform design is: ifdef _WIN32 / / Windows-specific code elif defined(__APPLE__) / / macOS-specific code elif defined(__linux__) / / Linux-specific code else / / Default implementation endif The adaptive window direct sound extraction method in the above embodiment has been described. Next, the adaptive window direct sound extraction device in this embodiment will be described. Please refer to Figure 3 , an embodiment of the adaptive window direct sound extraction device in this embodiment includes: An acquisition module 301, configured to acquire an input audio signal and perform preprocessing on the audio signal; A construction module 302, configured to construct an adaptive window function based on the frequency difference of the audio signal. The adaptive window function is configured to automatically adjust its window length and window coefficient according to different frequency bands; A conversion module 303, configured to convert the audio signal from the time domain to a frequency domain representation through the adaptive window function, and obtain the complete frequency domain data of the audio signal. Among them, the adaptive window function uses a longer window for low-frequency signals to ensure the frequency resolution of the frequency domain signal, and uses a shorter window for high-frequency signals to improve the time accuracy of the frequency domain signal; An extraction module 304, configured to establish a direct sound extraction tool according to the complete frequency domain data of the audio signal, and separate and extract the direct sound signal from the audio signal through the direct sound extraction tool.

[0121] In this embodiment, after acquiring the input audio signal, by introducing a frequency-dependent factor, an adaptive window function is constructed, and the window length and window coefficient of the adaptive window function are configured to be adaptively adjusted according to the frequency change. A longer window is used for low-frequency signals to ensure the frequency resolution, and a shorter window is used for high-frequency signals to improve the time accuracy. This effectively solves the technical problem of processing conflicts among different frequency components. By designing a dedicated window function for different frequency signals, the optimal balance of time-frequency analysis is achieved, so that more details can be retained in the extracted direct sound, and the signal analysis accuracy is significantly improved.

[0122] Optionally, the acquisition module 301 can also be used for: Adjust the amplitude of the time domain signal of the audio signal to the standard interval through maximum absolute value normalization processing; perform frame division on the normalized audio signal to obtain audio frame data.

[0123] Optionally, the building block 302 can also be used for: Select a basic window function; set a frequency-dependent factor, which is used to control the adaptive adjustment of the window length and window coefficients with the change of frequency. The frequency-dependent factor can respectively generate frequency-dependent values adapted to the signal frequency in the range from the high-frequency band to the low-frequency band of the signal; based on the basic window function and the frequency-dependent factor, construct an adaptive window function adapted to different frequency bands, so that the window length and window coefficients of the adaptive window function change continuously with frequency.

[0124] Optionally, the conversion module 303 can also be used for: Calculate the window length and window coefficients of the window function according to the target frequency band in the audio frame, and construct an adaptive window function adapted to the target frequency band; intercept the signal segment corresponding to the window length from the audio frame, apply the adaptive window function, and perform short-time Fourier transform processing; calculate the frequency-domain data of the audio frame at the target frequency band; integrate the frequency-domain data of all audio frames at different frequency bands in sequence to obtain the complete frequency-domain data of the audio signal.

[0125] Moreover, perform smoothing processing on the frequency-domain data of the audio signal through the exponential moving average algorithm; adjust the amplitude of the frequency-domain signal of the audio signal to the standard interval through maximum absolute value normalization processing.

[0126] Optionally, the extraction module 304 can also be used for: According to the characteristic pattern of the frequency-domain data of the audio signal, construct a direct sound judgment function, which is used to judge and distinguish the direct sound signal and the reverberation signal in the frequency-domain data of the audio signal; based on the direct sound judgment function, establish a masking matrix, which is used to separate and extract the direct sound signal from the audio signal; among them, the direct sound judgment function distinguishes the direct sound signal and the reverberation signal based on the transient peak characteristics of the frequency-domain data of the audio signal.

[0127] In addition, select a basic smoothing window function; calculate the window length of the window function according to the target frequency band, and construct a masking matrix smoothing processing function adapted to the target frequency band; perform a convolution operation on the masking matrix in different frequency bands through the masking matrix smoothing processing function.

[0128] In addition, apply the masking matrix to the frequency-domain data of the audio signal; perform inverse short-time Fourier transform processing with the corresponding window length on the target frequency band in the frequency-domain data; calculate and extract the time-domain data of each audio frame at the target frequency band and being a direct sound signal; overlap and add the time-domain data of the direct sound signal in each audio frame in sequence to obtain the complete direct sound data of the audio signal.

[0129] Moreover, the DC bias in the direct sound signal is eliminated by the de-mean algorithm; the low-frequency noise in the direct sound signal is filtered out by using a short-time high-pass filter; and the signal amplitude of the direct sound signal is adjusted to the standard interval through peak normalization processing.

[0130] Optionally, the adaptive window direct sound extraction device further includes: An adjustment module 305, configured to adaptively adjust the setting parameters of the direct sound judgment function and the masking matrix.

[0131] Optionally, the adjustment module 305 can also be used for: Dynamically adjusting the discrimination parameters of the direct sound judgment function according to the signal quality characteristics of the input audio signal; dynamically adjusting the masking characteristic parameters of the masking matrix according to the peak characteristics of the input audio signal; performing scenario-based adaptive adjustment on the setting parameters of the direct sound judgment function and the masking matrix according to the environmental characteristics of the audio signal; wherein, the signal quality characteristics include the signal-to-noise ratio, and the environmental characteristics include the reverberation time, background noise, and signal steady-state characteristics.

[0132] Optionally, the adaptive window direct sound extraction device further includes: An evaluation module 306, configured to evaluate and perform negative feedback adjustment on the performance of the direct sound judgment function and the masking matrix.

[0133] Optionally, the evaluation module 306 can also be used for: Defining a set of performance indicators for the direct sound judgment function and the masking matrix, where the set of performance indicators includes clarity, distortion, direct sound retention degree, and comprehensive performance score; periodically monitoring the performance of the direct sound judgment function and the masking matrix, and when the performance of the direct sound judgment function and / or the masking matrix is lower than a preset threshold, performing adaptive correction on the setting parameters of the direct sound judgment function and / or the masking matrix.

[0134] The functions of each module and each unit in the above adaptive window direct sound extraction device correspond to the steps in the above embodiments of the adaptive window direct sound extraction method, and their functions and implementation processes will not be elaborated here one by one.

[0135] This embodiment further provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above adaptive window direct sound extraction method. This electronic device can be a server or a terminal device.

[0136] See Figure 4As shown, the electronic device includes a processor 400 and a memory 401. The memory 401 stores machine-executable instructions that can be executed by the processor 400. The processor 400 executes the machine-executable instructions to implement the above adaptive window direct sound extraction method.

[0137] Furthermore, Figure 4 The electronic device shown further includes a bus 402 and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected through the bus 402.

[0138] Among them, the memory 401 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 403 (which can be wired or wireless). The Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 402 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a single bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0139] The processor 400 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 400 or the instructions in the form of software. The above-mentioned processor 400 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in this embodiment. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with this embodiment can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401 and combines its hardware to complete the steps of the adaptive window direct sound extraction method.

[0140] This embodiment also provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions cause the processor to implement the steps of the above-mentioned adaptive window direct sound extraction method.

[0141] The computer program product of the adaptive window direct sound extraction method, device, electronic device and storage medium provided in this embodiment includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiment. For the specific implementation, reference can be made to the method embodiment, which will not be elaborated here.

[0142] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.

[0143] In addition, in the description of this embodiment, unless otherwise clearly specified and defined, the terms "installation" and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in this disclosure can be understood according to specific circumstances.

[0144] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0145] In the description of this disclosure, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this disclosure. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0146] Finally, it should be noted that the above embodiments are only specific implementation manners of this disclosure to illustrate the technical solutions of this disclosure, rather than limiting it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art within the technical scope disclosed by this disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of this embodiment, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. An adaptive window direct sound extraction method, characterized in that: include: Collecting input audio signals and performing preprocessing on the audio signals; Based on the frequency difference of the audio signal, an adaptive window function is constructed, wherein the adaptive window function is configured to automatically adjust its window length and window coefficient according to different frequency bands; The audio signal is converted from the time domain to the frequency domain through the adaptive window function to obtain the complete frequency domain data of the audio signal, wherein the adaptive window function uses a longer window for the low-frequency signal to ensure the frequency resolution of the frequency domain signal, and uses a shorter window for the high-frequency signal to improve the time accuracy of the frequency domain signal; A direct sound extraction tool is established according to the complete frequency domain data of the audio signal, and the direct sound signal is separated and extracted from the audio signal by the direct sound extraction tool.

2. The adaptive window direct sound extraction method according to claim 1, characterized in that: Perform preprocessing on the audio signal, including: The time domain signal amplitude of the audio signal is adjusted to a standard range through maximum absolute value normalization processing; The normalized audio signal is framed to obtain audio frame data.

3. The adaptive window direct sound extraction method according to claim 1, characterized in that: Constructing the adaptive window function includes: Select the basic window function; Setting a frequency dependency factor, wherein the frequency dependency factor is used to control the window length and the window coefficient to be adaptively adjusted with frequency changes, and the frequency dependency factor can respectively respond to generate frequency dependency values ​​adapted to the signal frequency in the range from the high frequency band of the signal to the low frequency band of the signal; Based on the basic window function and the frequency dependence factor, the adaptive window function adapted to different frequency segments is constructed so that the window length and window coefficient of the adaptive window function change continuously with the frequency.

4. The adaptive window direct sound extraction method according to claim 3, characterized in that: The basic window function includes a Hanning window, a Hamming window, a Blackman window or a Kaiser window.

5. The adaptive window direct sound extraction method according to claim 1, characterized in that: Get complete frequency domain data of audio signals, including: Calculate the window length and window coefficient of the window function according to the target frequency segment in the audio frame, and construct the adaptive window function adapted to the target frequency segment; Cutting a signal segment corresponding to the window length from the audio frame, applying the adaptive window function, and performing short-time Fourier transform processing; Calculate the frequency domain data of the audio frame at the target frequency band; The frequency domain data of all audio frames at different frequency bands are sequentially integrated to obtain the complete frequency domain data of the audio signal.

6. The adaptive window direct sound extraction method according to claim 5, characterized in that: After obtaining the complete frequency domain data of the audio signal, including: Performing smoothing on the frequency domain data of the audio signal by using an exponential moving average algorithm; The frequency domain signal amplitude of the audio signal is adjusted to within the standard range through maximum absolute value normalization processing.

7. The adaptive window direct sound extraction method according to claim 1, characterized in that: The direct sound extraction tool is established according to the complete frequency domain data of the audio signal, including: Constructing a direct sound judgment function according to a characteristic pattern of the audio signal frequency domain data, wherein the direct sound judgment function is used to judge and distinguish between a direct sound signal and a reverberation signal in the audio signal frequency domain data; Based on the direct sound judgment function, a masking matrix is ​​established, wherein the masking matrix is ​​used to separate and extract the direct sound signal from the audio signal; The direct sound judgment function distinguishes the direct sound signal from the reverberation signal based on the transient peak characteristics of the audio signal frequency domain data.

8. The adaptive window direct sound extraction method according to claim 7, characterized in that: After obtaining the masking matrix, the following steps are performed: Select the basic smoothing window function; Calculate the window length of the window function according to the target frequency band, and construct a masking matrix smoothing function adapted to the target frequency band; A convolution operation is performed on the masking matrices in different frequency bands by the masking matrix smoothing function.

9. The adaptive window direct sound extraction method according to claim 8, characterized in that: Separating and extracting a direct sound signal from an audio signal by using the direct sound extraction tool includes: applying the masking matrix to frequency domain data of an audio signal; Performing an inverse short-time Fourier transform process of a corresponding window length on a target frequency segment in the frequency domain data; Calculate and extract the time domain data of each audio frame at the target frequency band, which is the direct sound signal; The time domain data of the direct sound signal in each audio frame are overlapped and added in sequence to obtain the complete direct sound data of the audio signal.

10. The adaptive window direct sound extraction method according to claim 9, characterized in that: After obtaining the complete direct sound data of the audio signal, including: The DC offset in the direct sound signal is eliminated by the de-averaging algorithm; Use a short-time high-pass filter to filter out low-frequency noise in the direct sound signal; The signal amplitude of the direct sound signal is adjusted to within the standard range through peak normalization processing.

11. The adaptive window direct sound extraction method according to claim 7, characterized in that: Also includes: Adaptively adjusting the setting parameters of the direct sound judgment function and the masking matrix includes: Dynamically adjusting the discrimination threshold of the direct sound judgment function according to the signal quality characteristics of the input audio signal; Dynamically adjusting the masking characteristic parameters of the masking matrix according to the peak characteristics of the input audio signal; According to the environmental characteristics of the audio signal, the setting parameters of the direct sound judgment function and the masking matrix are adjusted in a scene-based adaptive manner; The signal quality characteristics include a signal-to-noise ratio, and the environmental characteristics include a reverberation time, background noise, and a signal steady-state characteristic.

12. The adaptive window direct sound extraction method according to claim 7, characterized in that: Also includes: The performance of the direct sound judgment function and the masking matrix is ​​evaluated and negatively feedback adjusted, including: Defining a set of performance indicators of the direct sound judgment function and the masking matrix, wherein the set of performance indicators includes clarity, distortion, degree of direct sound retention, and a comprehensive performance score; The performance of the direct sound judgment function and the masking matrix is ​​periodically monitored, and when the performance of the direct sound judgment function and / or the masking matrix is ​​lower than a preset threshold, the setting parameters of the direct sound judgment function and / or the masking matrix are adaptively corrected.

13. An adaptive window direct sound extraction device, characterized in that: Executing the adaptive window direct sound extraction method according to any one of claims 1 to 12, comprising: An acquisition module, used for acquiring input audio signals and performing preprocessing on the audio signals; A construction module, used to construct an adaptive window function based on the frequency difference of the audio signal, wherein the adaptive window function is configured to automatically adjust its window length and window coefficient according to different frequency bands; A conversion module, used to convert the audio signal from the time domain to the frequency domain through the adaptive window function to obtain the complete frequency domain data of the audio signal, wherein the adaptive window function uses a longer window for the low-frequency signal to ensure the frequency resolution of the frequency domain signal, and uses a shorter window for the high-frequency signal to improve the time accuracy of the frequency domain signal; The extraction module is used to establish a direct sound extraction tool based on the complete frequency domain data of the audio signal, and separate and extract the direct sound signal from the audio signal through the direct sound extraction tool.

14. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the adaptive window direct sound extraction method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the adaptive window direct sound extraction method according to any one of claims 1 to 12.

Citation Information

Cited By

  • Voice signal feature extraction method and system

    CN121922150A

  • A feature extraction method and system for speech signals

    CN121922150B

  • Room sound calibration method based on intermediate frequency reverberation time and related equipment

    CN122027974A