Multi-Microphone Array Beamforming Signal Enhancement Method and Device

Through the time-frequency conversion and spatial distribution compensation of multi-microphone arrays, combined with complex spatiotemporal attention beamforming network and adversarial environment simulator, the beam direction is optimized, which solves the problem of poor signal enhancement effect in complex acoustic environments, and achieves high-quality voice signal output and system robustness.

CN119811408BActive Publication Date: 2025-08-01DONGGUAN HUAZE ELECTRONIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510294713.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-08-01
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing multi-microphone array beamforming signal enhancement method has poor effect in complex acoustic environments, lacks adaptability to dynamic acoustic environments, and cannot effectively optimize beam direction and gain parameters, resulting in poor signal enhancement effect.

Method used

By acquiring the multi-channel time domain signals of the multi-microphone array for time-frequency conversion, combining spatial distribution vectors for deformation compensation, using complex spatiotemporal attention beamforming network and adversarial environment simulator to optimize beam direction and gain parameters, dynamic band enhancement, and low-latency optimization.

Benefits of technology

It improves the quality of the sound field consistency signal, enhances the extraction accuracy of the target voice signal and the accuracy of speech recognition, reduces noise interference and echo effects, and improves the robustness and user experience of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811408B_ABST
    Figure CN119811408B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-microphone array beamforming signal enhancement method and apparatus. The method includes obtaining multi-channel time-domain signals collected by a multi-microphone array, performing time-frequency conversion to obtain complex time-frequency domain signals; obtaining the spatial distribution vector of the multi-microphone array, performing deformation compensation processing on the complex time-frequency domain signals to obtain a sound field consistency signal; extracting the sound source direction characteristics of the sound field consistency signal and performing complex spatio-temporal attention beamforming network processing to obtain a beamforming signal; optimizing the beam direction and gain parameters of the beamforming signal through a preset adversarial environment simulator to obtain a dynamic frequency band enhanced signal; performing inverse time-frequency conversion on the dynamic frequency band enhanced signal and performing low-latency optimization to obtain a target quality voice output signal. The present invention can perform more accurate deformation compensation of signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of signal enhancement, and particularly relates to a multi-microphone array beamforming signal enhancement method and device. Background Art

[0002] As a technology for improving the quality of speech signals and enhancing speech recognition ability, the multi-microphone array beamforming signal enhancement method has been widely applied in the fields of modern communication and speech processing. With the popularization of intelligent devices and speech interaction systems, how to accurately extract and enhance the target speech signal in a complex acoustic environment has become one of the research focuses. Existing signal enhancement methods usually rely on a single processing technology, such as simple beamforming or noise suppression algorithms, while ignoring the spatial information of the multi-microphone array and its consistent influence on the sound field. The application of this single technology may lead to unsatisfactory signal enhancement effects, thus affecting the performance of the speech recognition system. In addition, current methods often lack the ability to adapt to dynamic acoustic environments and cannot effectively optimize the beam direction and gain parameters, so ideal signal enhancement effects cannot be achieved in complex environments. The existence of these problems urgently requires a signal enhancement method that comprehensively considers the spatial distribution of the multi-microphone array, the characteristics of the dynamic sound field, and the adaptability to complex environments to improve the quality of the target speech signal and the robustness of the system. Summary of the Invention

[0003] The main object of the present invention is to provide a multi-microphone array beamforming signal enhancement method and device, which can more accurately perform signal deformation compensation and improve the quality of the sound field consistent signal.

[0004] To achieve the above object, the present invention provides a multi-microphone array beamforming signal enhancement method, including:

[0005] Obtain multi-channel time-domain signals collected by a multi-microphone array, perform time-frequency conversion to obtain complex time-frequency domain signals;

[0006] Obtain the spatial distribution vector of the multi-microphone array, and perform deformation compensation processing on the complex time-frequency domain signals to obtain sound field consistent signals;

[0007] Extract the source direction characteristics of the sound field consistent signals, and perform complex spatio-temporal attention beamforming network processing to obtain beamforming signals;

[0008] Optimize the beam direction and gain parameters of the beamforming signals through a preset adversarial environment simulator to obtain dynamically band-enhanced signals;

[0009] Perform inverse time-frequency conversion on the dynamically band-enhanced signals and perform low-latency optimization to obtain target quality speech output signals.

[0010] Further, obtaining multi-channel time-domain signals collected by a multi-microphone array and performing time-frequency conversion to obtain complex time-frequency domain signals includes:

[0011] Performing endpoint detection and noise power estimation on the multi-channel time-domain signals to obtain preprocessed signals;

[0012] Framing and windowing the preprocessed signals according to preset frame length and frame shift parameters to obtain initial complex frequency domain signals;

[0013] Performing spectrum smoothing on the initial complex frequency domain signals to obtain smoothed frequency domain signals;

[0014] Performing non-linear sub-band division on the smoothed frequency domain signals to obtain sub-band signals;

[0015] Performing geometric structure phase compensation on the sub-band signals to obtain sub-band corrected signals;

[0016] Performing spectral subtraction denoising on the sub-band corrected signals based on a preset noise power spectrum and performing complex time-frequency domain feature representation conversion to obtain the complex time-frequency domain signals.

[0017] Further, obtaining the spatial distribution vector of the multi-microphone array and performing deformation compensation processing on the complex time-frequency domain signals to obtain sound field consistency signals includes:

[0018] Obtaining the physical coordinates and orientation parameters of the multi-microphone array and performing spatial construction to obtain the spatial distribution vector;

[0019] Performing spatial transfer function mapping calculation on the complex time-frequency domain signals according to the spatial distribution vector to obtain array acoustic characteristic mapping information;

[0020] Performing singular value decomposition on the array acoustic characteristic mapping information and performing phase and amplitude non-linear correction to obtain preliminary compensation signals;

[0021] Calculating the cross-power spectral density matrix based on the preliminary compensation signals and performing sound field geometric consistency constraint construction to obtain sound field topology characterization;

[0022] Performing complex domain tensor decomposition on the sound field topology characterization to obtain deformation compensation coefficients;

[0023] Performing phase and amplitude alignment on the complex time-frequency domain signals according to the deformation compensation coefficients to obtain the sound field consistency signals.

[0024] Further, extracting sound source direction features from the sound field consistency signals and performing complex spatio-temporal attention beamforming network processing to obtain beamforming signals includes:

[0025] Perform adaptive wavelet packet transform decomposition on the sound field consistency signal to obtain a multi-level time-frequency representation of the sound field;

[0026] Perform complex domain spatial covariance calculation on the multi-level time-frequency representation of the sound field to obtain a frequency band phase difference spectrum;

[0027] Perform high-order singular value decomposition on the frequency band phase difference spectrum to obtain an orthogonal projection structure;

[0028] Construct a hyperbolic direction estimation function based on the orthogonal projection structure and perform gradient polarization search to obtain a sound source direction vector;

[0029] Align the phase of the sound source direction vector and the sound field consistency signal to obtain an initial beam gain coefficient;

[0030] Construct a complex domain attention mask based on the initial beam gain coefficient and perform channel weighting on the sound field consistency signal to obtain a beamforming signal.

[0031] Further, the optimization of the beam direction and gain parameters of the beamforming signal by a preset adversarial environment simulator to obtain a dynamic frequency band enhancement signal includes:

[0032] Simulate various noise interferences and reverberation conditions on the beamforming signal based on the adversarial environment simulator to obtain diverse environmental interference samples;

[0033] Perform gradient descent optimization calculation on the diverse environmental interference samples and adjust the beam direction parameters and frequency band gain coefficients according to a preset signal-to-noise ratio objective function to obtain a parameter optimization matrix;

[0034] Reconstruct the directionality of the beamforming signal based on the parameter optimization matrix to obtain a direction enhancement signal;

[0035] Match the acoustic features of the direction enhancement signal and preset speech data to obtain corresponding target actual short-time speech data;

[0036] Extract features from the target actual short-time speech data based on a preset meta-learning online calibration algorithm to obtain a target feature adaptation matrix;

[0037] Perform convolution operation on the target feature adaptation matrix and the direction enhancement signal to obtain a spectrum correction signal;

[0038] Perform cross-frequency band consistency processing and phase frequency band gain on the spectrum correction signal to obtain the dynamic frequency band enhancement signal.

[0039] Further, simulating various noise interferences and reverberation conditions on the beamforming signal based on the adversarial environment simulator to obtain diverse environmental interference samples, including:

[0040] Performing time-frequency transformation on the beamforming signal through the adversarial environment simulator to obtain a time-frequency domain representation signal matrix;

[0041] Constructing a three-dimensional acoustic propagation structure according to the time-frequency domain representation signal matrix and setting random reflection coefficients to obtain a set of spatial reverberation parameters;

[0042] Randomly extracting various types of noise sources from a preset environmental noise library, and performing time-domain stretching and frequency-domain offset to obtain a set of deformed noise sources;

[0043] Performing position mapping on the set of deformed noise sources according to a preset spatial directivity distribution function to obtain a multi-directional noise distribution matrix;

[0044] Calculating the acoustic propagation time delay and energy attenuation according to the multi-directional noise distribution matrix and the set of spatial reverberation parameters to obtain a noise propagation characteristic vector;

[0045] Performing reverberation processing on the time-frequency domain representation signal matrix based on the set of spatial reverberation parameters, and performing interference superposition with the noise propagation characteristic vector to obtain the diverse environmental interference samples.

[0046] Further, extracting features from the target actual short-time speech data based on a preset meta-learning online calibration algorithm to obtain a target feature adaptation matrix, including:

[0047] Performing short-time window framing and windowing on the target actual short-time speech data to obtain an overlapping segmented signal sequence;

[0048] Calculating Mel-frequency cepstral coefficients for the overlapping segmented signal sequence to obtain a segmented feature description vector;

[0049] Constructing a meta-learning task sampling set for the segmented feature description vector to obtain a feature support set and a feature query set;

[0050] Constructing a meta-learning inner-loop optimization function based on the feature support set to obtain environmental adaptation parameters;

[0051] Performing MAML gradient update on the feature query set according to the environmental adaptation parameters to obtain task model parameters;

[0052] Performing meta-learning outer-loop optimization and online learning rate adjustment on the task model parameters to obtain a dynamic learning rate matrix;

[0053] Perform prototype network mapping on the segmented feature description vectors according to the dynamic learning rate matrix to obtain an environmental perception feature representation;

[0054] Construct a metric learning embedding space for the environmental perception feature representation to obtain a speech feature embedding matrix;

[0055] Perform Bayesian credible interval constraint on the speech feature embedding matrix to obtain a target feature adaptation matrix.

[0056] Further, the inverse time-frequency conversion of the dynamic band-enhanced signal and low-latency optimization to obtain a target quality speech output signal includes:

[0057] Perform an overlap-and-add inverse transform process on the dynamic band-enhanced signal to obtain an initial time-domain signal;

[0058] Perform inter-frame phase consistency compensation on the initial time-domain signal to obtain a phase-continuous intermediate signal;

[0059] Perform buffer dynamic adjustment according to the phase-continuous intermediate signal to obtain a variable signal buffer;

[0060] Perform non-linear amplitude normalization on the variable signal buffer to obtain a dynamic range compression signal;

[0061] Calculate a time-domain transient feature vector according to the dynamic range compression signal and perform time-domain jitter elimination to obtain a jitter-corrected signal;

[0062] Perform multi-stage cascaded band reconstruction on the jitter-corrected signal to obtain a frequency response equalization signal;

[0063] Perform device pre-compensation processing on the frequency response equalization signal to obtain a device matching signal;

[0064] Eliminate noise and artifacts from the device matching signal according to a preset frame selective discard algorithm to obtain the target quality speech output signal.

[0065] The present invention also provides a multi-microphone array beamforming signal enhancement device applied to the multi-microphone array beamforming signal enhancement method described in any one of the above, including:

[0066] An acquisition module, which is used to acquire multi-channel time-domain signals collected by a multi-microphone array and perform time-frequency conversion to obtain complex time-frequency domain signals;

[0067] An analysis module, which is used to obtain the spatial distribution vector of the multi-microphone array and perform deformation compensation processing on the complex time-frequency domain signals to obtain a sound field consistency signal;

[0068] An association module, which is used to extract the sound source direction features of the sound field consistency signal and perform complex spatio-temporal attention beamforming network processing to obtain a beamforming signal;

[0069] A processing module, which is used to optimize the beam direction and gain parameters of the beamforming signal through a preset adversarial environment simulator to obtain a dynamic frequency band enhanced signal;

[0070] A control module, which is used to perform inverse time-frequency conversion on the dynamic frequency band enhanced signal and perform low-latency optimization to obtain a target quality voice output signal.

[0071] A multi-microphone array beamforming signal enhancement method and device provided by the present invention have the following beneficial effects:

[0072] By performing time-frequency conversion on the time-domain signals collected by the multi-microphone array and combining the spatial distribution vector information of the array, signal deformation compensation can be performed more accurately, improving the quality of the sound field consistency signal. This method can effectively improve the extraction accuracy of the target voice signal by extracting the direction features of the sound source and applying the complex spatio-temporal attention beamforming network, avoiding signal attenuation or distortion in traditional signal enhancement methods. By introducing a preset adversarial environment simulator to optimize the beam direction and gain parameters, it can adapt to dynamic and complex acoustic environments, realizing the adaptive adjustment of the beam direction and gain parameters, and further improving the clarity of the signal and the accuracy of speech recognition. Based on this technology, noise interference and echo effects can be effectively reduced, ensuring high-quality voice output signals in different environments and reducing the problems of decreased sound quality and speech distortion in traditional methods. In addition, the low-latency optimization design of the dynamic frequency band enhanced signal helps the smooth operation of the real-time voice interaction system, enhancing the robustness and user experience of the system. By comprehensively considering spatio-temporal information and dynamic sound field features, this method can flexibly adapt to complex and changing acoustic environments, improving the overall performance and application effects of the multi-microphone array system. Description of the Drawings

[0073] Figure 1 is a flowchart of a multi-microphone array beamforming signal enhancement method provided by the present invention;

[0074] Figure 2 is a structural diagram of a multi-microphone array beamforming signal enhancement device provided by the present invention.

[0075] The realization, functional characteristics and advantages of the purpose of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0076] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the present invention.

[0077] Next, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners.

[0078] Referring to Figure 1 As shown, 1. A multi-microphone array beamforming signal enhancement method includes:

[0079] Step S1: Obtain multi-channel time-domain signals collected by a multi-microphone array, perform time-frequency conversion, and obtain complex time-frequency domain signals;

[0080] Step S2: Obtain the spatial distribution vector of the multi-microphone array, perform deformation compensation processing on the complex time-frequency domain signals, and obtain sound field consistency signals;

[0081] Step S3: Extract the source direction characteristics of the sound field consistency signals, and perform complex spatio-temporal attention beamforming network processing to obtain beamforming signals;

[0082] Step S4: Optimize the beam direction and gain parameters of the beamforming signals through a preset adversarial environment simulator to obtain dynamically band-enhanced signals;

[0083] Step S5: Perform inverse time-frequency conversion on the dynamically band-enhanced signals and perform low-latency optimization to obtain target-quality speech output signals.

[0084] Based on the above steps, the detailed step process is as follows:

[0085] Step S1: The first step of multi-microphone array beamforming is to obtain the original acoustic signal and convert it into a representation suitable for processing. The microphone array consists of multiple microphones arranged in a specific geometry, and each microphone captures the time-domain representation of the sound wave. These time-domain signals contain source information as well as background noise and reverberation. To effectively process these signals, it is necessary to convert them into the time-frequency domain. Time-frequency conversion usually uses the Short-Time Fourier Transform (STFT), which divides the time-domain signal into short-time frames and applies the Fourier transform to each frame. Specifically, a window function (such as a Hamming window or a Hanning window) is applied to the time-domain signal of each microphone channel to reduce spectral leakage. After the window function is applied, the Fast Fourier Transform (FFT) is performed to generate a complex time-frequency domain representation. This complex representation contains amplitude and phase information, which is crucial for subsequent beamforming. The result of time-frequency conversion is a three-dimensional tensor with dimensions [number of microphones × number of time frames × number of frequency points], and each element is a complex number. The selection of the frame length and the frame shift step affects the trade-off between time resolution and frequency resolution and needs to be adjusted according to the application scenario. Typical speech processing may use a frame length of 20 - 30 ms and a frame shift of 10 - 15 ms.

[0086] Step S2: The spatial distribution vector of the multi-microphone array describes the relative position relationship of each microphone in three-dimensional space. This spatial distribution information is crucial for understanding how sound waves arrive at different parts of the array. In practical applications, there may be problems such as shape deformation, installation errors, or inconsistent microphone characteristics in the microphone array, which will affect the consistent capture of the sound field. The deformation compensation process estimates the calibration coefficient of each microphone through the spatial distribution vector to compensate for these errors. The compensation process is based on the acoustic propagation model and applies phase and amplitude correction to the complex time-frequency domain signal. The correction factor usually includes a complex coefficient matrix related to frequency, which is used to adjust the phase and amplitude relationship between each microphone channel. Deformation compensation may also involve calibration of microphone sensitivity differences and frequency response equalization. The resulting sound field consistent signal should ideally show that if the sound source comes from a specific direction, the signals of all microphones should be completely coherent with the corresponding phase delay. This consistency provides a key basis for sound source direction estimation and beamforming in the next step.

[0087] Step S3: Use the calibrated sound field consistency signal for sound source localization and enhancement. The sound source direction feature extraction process is based on the phase information analysis of the sound field consistency signal, and calculates the possible direction of the sound source through the phase difference between the microphone signals. Commonly used methods include Generalized Cross-Correlation (GCC), Multi-Channel Cross-Power Spectral Density (CPSD) analysis, or a deep learning-based direction feature extraction network. The extracted direction features are usually represented as the probability distribution or feature vector of azimuth and elevation angles. The complex spatio-temporal attention beamforming network is an innovative signal processing architecture that combines the advantages of traditional beamforming techniques and deep learning. This network applies an attention mechanism to both the time dimension and the space dimension, and can dynamically adjust the attention to different time-frequency points and space directions. The network processing includes the calculation of a complex masking matrix, which assigns different complex weights to each time-frequency point according to the sound source direction and time-frequency characteristics. By applying this masking matrix to the sound field consistency signal, the signal from the target direction can be enhanced, and the interference and noise from other directions can be suppressed. The quality of the beamformed signal depends on the accuracy of the direction features and the optimization degree of the network parameters.

[0088] Step S4: The adversarial environment simulator is a specially designed module for simulating various acoustic interference situations that may occur in real-world environments, such as reverberation, non-stationary noise, competing speakers, etc. The simulator adopts an adversarial training framework, in which the generator is responsible for creating various complex acoustic scenarios, and the discriminator evaluates the performance of the beamforming algorithm in these scenarios. Based on this adversarial mechanism, the system can discover potential weaknesses in the beamforming process. The beam direction optimization process maximizes the suppression of interference sources while maintaining the enhancement of the target signal by fine-tuning the width and direction of the beam main lobe. The gain parameter optimization adjusts the gain distribution of different frequency bands to balance the relationship between signal amplification and noise suppression. The optimization process uses gradient descent or evolutionary algorithms and iteratively optimizes based on objective metrics such as signal-to-noise ratio and speech intelligibility. The dynamically band-enhanced signal is the result of optimization, with adaptive gain adjustment characteristics for different frequency components, which can improve the intelligibility of speech while maintaining its naturalness. This step enables the beamforming technology to shift from static optimization to dynamic adaptation, significantly enhancing the robustness of the system in changing environments.

[0089] Step S5: The inverse time-frequency conversion uses the inverse short-time Fourier transform (ISTFT) to convert the complex time-frequency domain signal back to the time domain. The conversion process includes performing the inverse Fourier transform on each time-frequency frame and then using the overlap-add method to synthesize a continuous time-domain signal. To ensure signal quality, spectral smoothing techniques are usually employed to reduce artifacts caused by phase discontinuities. Low-latency optimization is a key requirement for real-time voice interaction systems, and the traditional STFT-ISTFT framework may introduce a relatively large processing delay. Low-latency optimization strategies include: using shorter analysis windows and frame shift lengths, implementing a streaming processing architecture to reduce buffer delays, and using frequency-domain prediction techniques to estimate the content of subsequent frames in advance. In addition, computational resource optimization is also an important aspect of reducing system latency, including techniques such as algorithm parallelization and GPU acceleration. The finally output target-quality voice signal meets the low-latency requirements while maintaining good speech clarity and naturalness. Quality assessment is usually comprehensively judged by combining subjective listening tests (such as MOS scores) and objective metrics (such as PESQ, STOI). This step enables the entire system to meet the requirements of application scenarios with high real-time requirements, such as conference calls and smart speakers.

[0090] Step S6: As the final step of the multi-microphone array beamforming signal enhancement method, the output signal is finally polished and quality-verified. Post-processing includes noise suppression residual artifact processing, dynamic range compression, and dereverberation enhancement. Artifact processing eliminates the musical noise and spectral discontinuities that may be introduced during the beamforming process through spectral smoothing and transient detection techniques. Dynamic range compression adjusts the dynamic characteristics of the signal according to the application scenario to ensure consistent audibility on different playback devices. The dereverberation enhancement technology compensates for the deficiency of beamforming in reverberation suppression and further improves speech clarity through blind room impulse response estimation and inverse filtering. The quality assessment link establishes a multi-dimensional evaluation system, including two major categories: objective evaluation and subjective evaluation. Objective evaluation uses standardized metrics such as PESQ, STOI, SNR, etc., and subjective evaluation uses methods such as MUSHRA tests or ABX comparisons, inviting professional listeners and ordinary users to participate in the evaluation. The evaluation results are fed back to the previous steps to form a closed-loop optimization mechanism. This step ensures that the finally output voice signal can meet the quality standards expected by users in different application scenarios, and at the same time provides a scientific basis for the continuous optimization of the system.

[0091] The present invention provides a multi-microphone array beamforming signal enhancement method. By performing time-to-frequency conversion on the time-domain signals collected by the multi-microphone array and combining the spatial distribution vector information of the array, it can more accurately compensate for signal deformation and improve the quality of the sound field consistency signal. By extracting the directional characteristics of the sound source and applying a complex spatiotemporal attention beamforming network, this method can effectively improve the extraction accuracy of the target speech signal and avoid the signal attenuation or distortion in traditional signal enhancement methods. By introducing a preset adversarial environment simulator, the beam direction and gain parameters are optimized, which can adapt to dynamic and complex acoustic environments, achieving adaptive adjustment of the beam direction and gain parameters, further improving signal clarity and speech recognition accuracy. Based on this technology, it can effectively reduce noise interference and echo effects, thereby ensuring high-quality speech output signals in different environments, reducing the sound quality degradation and speech distortion problems encountered in traditional methods. In addition, the low-latency optimization design of the dynamic frequency band enhancement signal contributes to the smooth operation of the real-time voice interaction system, enhancing the system robustness and user experience. By comprehensively considering spatiotemporal information and dynamic sound field characteristics, this method can flexibly adapt to complex and changing acoustic environments, improving the overall performance and application effect of the multi-microphone array system.

[0092] In one embodiment, a multi-channel time domain signal collected by a multi-microphone array is obtained, and time-frequency conversion is performed to obtain a complex time-frequency domain signal, including:

[0093] After acquiring the multi-channel time-domain signal from the multi-microphone array, endpoint detection and noise power estimation are performed to produce a preprocessed signal. Endpoint detection identifies the start and end points of valid speech segments in the speech signal, distinguishing valid speech from background noise. Endpoint detection is typically based on features such as short-term energy and zero-crossing rate, and determines whether the current frame contains valid speech by setting an appropriate threshold. Noise power estimation estimates the power spectral density of background noise during periods of inactivity, providing a reference for subsequent signal enhancement. Noise power estimation uses a minimum statistics method or a method based on voice activity detection to update the noise power during periods of inactivity. The preprocessed signal is the signal after endpoint detection and noise power estimation. This signal retains the valid speech portion and includes the noise power estimation result.

[0094] Frame the preprocessed signal according to the preset frame length and frame shift parameters and apply windowing to obtain the initial complex frequency-domain signal. Framing is to divide a long speech signal into short-time frames, and the signal within each frame has relative stationarity. The preset frame length is usually 20 - 30 milliseconds (for a signal with a sampling rate of 16 kHz, the frame length is 320 - 480 sampling points), and the frame shift is half of the frame length or less (such as 10 - 15 milliseconds, corresponding to 160 - 240 sampling points) to ensure a certain overlap between adjacent frames. Windowing is to multiply each frame of the signal by a window function (such as Hamming window, Hanning window, etc.) to reduce spectral leakage. After framing and windowing, perform a fast Fourier transform (FFT) on each frame of the signal to convert it to the frequency domain and obtain the initial complex frequency-domain signal. The initial complex frequency-domain signal contains amplitude information and phase information, providing a basis for subsequent spectral processing.

[0095] Perform spectral smoothing on the initial complex frequency-domain signal to obtain the smoothed frequency-domain signal. Spectral smoothing is to perform weighted averaging on the energy spectra of adjacent frequency points in the frequency domain to reduce the volatility of the spectrum and improve the stability of spectral estimation. Spectral smoothing uses a sliding window smoothing method in the time-frequency domain. For each frequency point, consider several adjacent frequency points and the corresponding frequency points in the previous few frames for weighted averaging. The smoothed frequency-domain signal has better smoothness and continuity than the initial complex frequency-domain signal, reduces the influence of random noise, and is beneficial for subsequent signal processing.

[0096] Perform non-linear sub-band partitioning on the smoothed frequency-domain signal to obtain sub-band signals. Non-linear sub-band partitioning means dividing the frequency-domain signal into multiple sub-bands according to the auditory characteristics of the human ear. The sub-band partitioning in the low-frequency part is finer, and the sub-band partitioning in the high-frequency part is coarser. This partitioning method conforms to the perception characteristics of the human ear for sounds of different frequencies. The low-frequency part contains more speech information and requires more refined processing. Non-linear sub-band partitioning uses the Mel frequency scale or the Bark frequency scale to convert the linear frequency scale to the perceptual frequency scale. For a signal with a sampling rate of 16 kHz, it is usually divided into 20 - 30 sub-bands. The sub-band signal is the signal after non-linear sub-band partitioning, and each sub-band contains the signal components within a specific frequency range.

[0097] Perform geometric structure phase compensation on the subband signals to obtain subband corrected signals. Geometric structure phase compensation refers to adjusting the phases of the signals received by each microphone according to the geometric layout of the microphone array and the sound source direction, so that the sound source signals in the target direction are phase-aligned, thereby enhancing the target sound source signals. Geometric structure phase compensation calculates the corresponding phase differences based on the time delay differences from the sound source to each microphone and corrects the phases of the microphone signals. For a linear microphone array, phase compensation is usually based on the plane wave assumption; for a circular or other shaped microphone array, phase compensation takes into account more complex geometric relationships. The subband corrected signals are the signals after geometric structure phase compensation, and the sound source signals in the target direction received by each microphone are phase-aligned.

[0098] Perform spectral subtraction denoising on the subband corrected signals based on a preset noise power spectrum and perform complex time-frequency domain feature representation conversion to obtain complex time-frequency domain signals. Spectral subtraction denoising is a commonly used speech enhancement technique that obtains the enhanced speech signal power spectrum by subtracting the estimated noise power spectrum from the observed signal power spectrum. The preset noise power spectrum is the background noise power spectrum estimated during the period without speech activity. Spectral subtraction denoising adopts over-subtraction or under-subtraction strategies and sets a spectral floor to avoid generating musical noise. Complex time-frequency domain feature representation conversion converts the enhanced signal from the power spectrum domain back to the complex time-frequency domain, facilitating subsequent beamforming and signal reconstruction. The complex time-frequency domain signals retain the amplitude and phase information of the enhanced signal, providing the necessary information for beamforming and inverse transformation.

[0099] In this embodiment, through a systematic signal processing flow, efficient target sound source extraction is achieved in a complex acoustic environment. By combining endpoint detection and noise power estimation in the preprocessing method, the effective speech and background noise are accurately distinguished, laying a foundation for subsequent processing. The nonlinear subband division technique is adopted to finely divide the frequency domain signal according to the human ear auditory characteristics. The low-frequency part is more refined and the high-frequency part is relatively rough, effectively improving the speech perception quality. Geometric structure phase compensation makes full use of the spatial information of the microphone array to achieve phase alignment of the sound source signals in the target direction, significantly enhancing the spatial selectivity of the target sound source. The spectral smoothing technique reduces the spectral fluctuations and improves the spectral estimation stability, while the spectral subtraction denoising based on the preset noise power spectrum further suppresses the background noise interference. The organic combination of these techniques forms an integrated signal enhancement solution, which improves the signal clarity and intelligibility while maintaining the speech naturalness, and is particularly suitable for speech interaction scenarios in complex acoustic environments such as conference systems and smart speakers.

[0100] In one embodiment, obtain the spatial distribution vector of the multi-microphone array, and perform deformation compensation processing on the complex time-frequency domain signals to obtain sound field consistency signals, including:

[0101] By measuring or setting the position coordinates of each microphone element in three-dimensional space and its pickup directivity angle parameters, a complete spatial geometric structure model is constructed. These physical coordinates and orientation parameters are integrated into a unified spatial coordinate system to generate a spatial distribution vector. The spatial distribution vector contains the precise spatial positions and directivity information of each pickup element in the microphone array, providing the basic data for subsequent acoustic characteristic mapping.

[0102] After obtaining the spatial distribution vector, the spatial transfer function mapping calculation is performed on the transformed complex time-frequency domain signal using this vector. This mapping process associates the complex time-frequency domain signal received by each microphone with its spatial position and direction characteristics, generating array acoustic characteristic mapping information. The array acoustic characteristic mapping information characterizes the characteristics of the interaction between sound waves and the microphone array during spatial propagation, including information such as the time difference, amplitude attenuation, and phase change of sound waves arriving at each microphone.

[0103] Perform a singular value decomposition (SVD) operation on the array acoustic characteristic mapping information to decompose the complex mapping information into several orthogonal eigencomponents. Through this decomposition, the system identifies the main acoustic characteristic patterns and minor interference components. Subsequently, the system performs non-linear correction of the phase and amplitude on the decomposition result to eliminate or reduce the distortion caused by uneven microphone spacing, sensitivity differences, and hardware errors, thereby obtaining a preliminary compensation signal. The preliminary compensation signal has higher spatial consistency than the original signal but is not yet fully aligned.

[0104] Based on the preliminary compensation signal, the system calculates the cross-power spectral density matrix (CPSD), which describes the correlation and energy distribution between different microphone signals. Combining the physical characteristics and propagation laws of the sound field, the system constructs geometric consistency constraint conditions for the sound field, and these constraint conditions reflect the geometric relationships that the sound field should present at the positions of each microphone in an ideal state. By applying these constraints, the system generates a sound field topology representation, which depicts the geometric relationships between the sound source, the sound field spatial structure, and the microphone array.

[0105] The sound field topology representation undergoes complex domain tensor decomposition processing to decompose the sound field characteristics into amplitude components and phase components in the complex domain, thereby obtaining deformation compensation coefficients. The deformation compensation coefficients contain fine correction parameters for each frequency point and each microphone, which are used to offset the non-ideal deformation factors in the sound field.

[0106] Use the deformation compensation coefficients to perform precise alignment operations on the phase and amplitude of the complex time-frequency domain signal. Phase alignment ensures that the phase differences between signals from the same sound source among different microphones are appropriately compensated, while amplitude alignment makes the signal intensities received by each microphone reach a balanced consistency. After this processing, the system obtains a sound field consistency signal, which has high spatial consistency and provides ideal basic data for subsequent beamforming processing.

[0107] Compared with the original signal, the spatial coherence of the sound field consistency signal is significantly improved, reducing the influence of environmental noise and reverberation, enhancing the quality of the target sound source signal, and thus achieving more accurate sound source localization and signal enhancement effects in complex acoustic environments.

[0108] In this embodiment, by obtaining the spatial distribution vector of the multi-microphone array and performing deformation compensation processing, this method realizes the high-precision spatial consistency alignment of the microphone array signals, significantly improving the accuracy and stability of beamforming. Using the spatial transfer function mapping calculation combined with the singular value decomposition technique, it effectively solves the problem of acoustic characteristic distortion caused by element position errors, sensitivity differences, and hardware inconsistencies in the actual deployment of the microphone array. By constructing the sound field geometric consistency constraint and performing complex domain tensor decomposition, this method can accurately identify and compensate for the non-linear deformation in the sound field, overcoming the limitations of traditional methods in dealing with complex acoustic environments. The precise alignment processing of phase and amplitude makes the signals from the same sound source highly consistent among different microphones, thus significantly suppressing environmental noise and reverberation interference while maintaining the integrity of the target sound source signal.

[0109] In one embodiment, extracting the sound source direction characteristics of the sound field consistency signal and performing complex spatio-temporal attention beamforming network processing to obtain the beamforming signal, including:

[0110] The sound field consistency signal collected by the multi-microphone array is decomposed through adaptive wavelet packet transform to generate a multi-level sound field time-frequency representation. Adaptive wavelet packet transform is a time-frequency analysis method that decomposes the signal at different frequencies and time scales and is suitable for the processing of non-stationary signals. This transform automatically adjusts the decomposition level and wavelet basis function according to the local characteristics of the signal to obtain the best time-frequency resolution. During the transformation process, each microphone signal is mapped to multiple frequency sub-bands to form a time-frequency atomic structure. These time-frequency atoms retain the time and frequency characteristics of the original signal, effectively capturing the energy distribution and phase information of the sound field in each frequency band. The multi-level sound field time-frequency representation contains the complete time-frequency characteristics of the sound field in different frequency intervals, providing a basis for subsequent spatial feature analysis.

[0111] The multi-level sound field time-frequency representation is sent to the complex-domain spatial covariance calculation unit to calculate the phase difference spectrum of each frequency band. The complex-domain spatial covariance is a statistical description of the phase relationship between different microphone signals, reflecting the characteristics of sound wave propagation in space. During the calculation process, for each frequency band, a complex covariance matrix is formed among the multi-microphone signals, and this matrix contains the phase difference information between microphone pairs. The frequency band phase difference spectrum is a three-dimensional tensor structure with dimensions of frequency, space, and time. It records the spatial phase change patterns in the sound field, and these patterns are closely related to the direction and distance of the sound source. The frequency band phase difference spectrum has a high signal-to-noise ratio and can maintain relatively stable direction information even in a noisy environment.

[0112] The frequency band phase difference spectrum undergoes high-order singular value decomposition to obtain an orthogonal projection structure. High-order singular value decomposition is a multi-dimensional data analysis technique applicable to processing tensor data with three or more dimensions. This decomposition decomposes the frequency band phase difference spectrum tensor into a series of orthogonal subspace projections, and each projection represents an independent spatial mode in the sound field. The orthogonal projection structure consists of the main singular matrix and the corresponding singular values. The main singular matrix reflects the main spatial characteristics in the frequency band phase difference spectrum, while the singular values represent the significance of each characteristic. The orthogonal projection structure eliminates redundant information, highlights the main direction characteristics in the sound field, and effectively improves the accuracy and robustness of subsequent direction estimation.

[0113] Based on the orthogonal projection structure, a hyperbolic direction estimation function is constructed and gradient polarization search is carried out to obtain the sound source direction vector. The hyperbolic direction estimation function is a non-linear mapping that maps spatial coordinates to direction similarity scores, forming a hyperbolic surface. This function utilizes the orthogonal projection structure of the frequency band phase difference spectrum to form a direction response surface in three-dimensional space. Gradient polarization search is an algorithm for finding the optimal point on this response surface. By iteratively calculating the gradient direction and moving along the gradient direction, it finally converges to the extreme point of the direction response surface. The sound source direction vector is a three-dimensional unit vector pointing to the spatial position of the sound source, containing azimuth and elevation angle information. The calculation of the direction vector takes into account the fusion of multi-frequency band information to ensure the stability of direction estimation in various acoustic environments.

[0114] The sound source direction vector is phase-aligned with the sound field consistency signal to obtain the initial beam gain coefficients. Phase alignment is a core step in beamforming, which makes the phase of each microphone signal consistent in a specific direction, thereby enhancing the sound source signal in that direction. During the alignment process, according to the sound source direction vector and the geometric layout of the microphone array, the time delay difference from the sound source to each microphone is calculated, and the phase of each microphone signal is adjusted accordingly. The initial beam gain coefficients are a set of complex weights that contain amplitude and phase information and are used to weighted synthesize each microphone signal. These coefficients make the signals in the target direction phase-align and superimpose, while the signals in non-target directions cancel each other out due to inconsistent phases, thus forming a spatial selectivity gain.

[0115] Based on the initial beam gain coefficients, a complex domain attention mask is constructed and the sound field consistency signal is channel-weighted to obtain the beamforming signal. The complex domain attention mask is a spatial filter operating in the complex plane, which assigns different weights to signals in different directions and frequencies. The mask construction process utilizes the initial beam gain coefficients and further optimizes the spatial filtering characteristics through an adaptive learning mechanism. This mask not only considers the amplitude of the signal but also retains the phase information, ensuring phase consistency during the beamforming process. Channel weighting is the process of applying the complex domain attention mask to the sound field consistency signal. After each channel signal is weighted, complex domain superposition is performed to form the final beamforming output. The beamforming signal has a high signal-to-noise ratio and clarity, effectively suppressing environmental noise and reverberation interference and enhancing the sound source signal in the target direction. The beamforming signal is used in subsequent applications such as speech recognition, sound source localization, and acoustic monitoring, significantly improving the performance of these applications in complex acoustic environments.

[0116] In this embodiment, through the adaptive wavelet packet transform decomposition of the sound field consistency signal, a multi-level sound field time-frequency representation is obtained, effectively capturing the energy distribution and phase information of the sound field in each frequency band, providing a solid foundation for accurate sound source localization. The frequency band phase difference spectrum calculated using complex domain spatial covariance can maintain stable direction information in a high-noise environment, significantly improving the anti-interference ability of the system. The orthogonal projection structure generated by high-order singular value decomposition eliminates redundant information, highlights the main direction features in the sound field, and enhances the accuracy of direction estimation. The method based on the hyperbolic direction estimation function and gradient polarization search realizes high-precision calculation of the sound source direction vector, adapting to the changes in complex acoustic environments. The initial beam gain coefficients obtained through phase alignment make the signals in the target direction phase-align and superimpose, while the signals in non-target directions cancel each other out. Finally, the construction and application of the complex domain attention mask optimize the spatial filtering characteristics while retaining the phase information, generating a beamforming signal with a high signal-to-noise ratio and clarity, effectively suppressing environmental noise and reverberation interference.

[0117] In one embodiment, the beam direction and gain parameters of the beamforming signal are optimized by a preset adversarial environment simulator to obtain a dynamic frequency band enhanced signal, including:

[0118] The simulator includes a noise generation module, a reverberation simulation module, and an environmental condition parameter library for creating diverse environmental interference samples.

[0119] The adversarial environment simulator simulates various noise interferences and reverberation conditions on the beamforming signal to generate diverse environmental interference samples. These samples include various noise types such as white noise, pink noise, mechanical noise, and human voices, as well as reverberation conditions under different room sizes, materials, and shapes. The environmental interference samples form an interference matrix through interference source localization and interference intensity mapping, which contains interference information in both the spatial and frequency dimensions.

[0120] When performing gradient descent optimization calculations on the diverse environmental interference samples, the stochastic gradient descent algorithm is used to adjust the beam direction parameters and frequency band gain coefficients according to a preset signal-to-noise ratio objective function. The signal-to-noise ratio objective function is defined as the logarithm of the ratio of the desired signal power to the noise power, and the threshold of this function is set to 15 dB to achieve adaptive adjustment of the beam direction. During the optimization process, the adjustment range of the beam direction parameters is ±30 degrees, and the adjustment range of the frequency band gain coefficients is from 0.5 to 2.0. A parameter optimization matrix is obtained through 300 iterations of calculation. The parameter optimization matrix contains the beam direction angle correction values and the adjustment values of each frequency band gain coefficient for subsequent signal reconstruction.

[0121] Based on the parameter optimization matrix, the beamforming signal is directionally reconstructed to obtain a directionally enhanced signal. During the directional reconstruction process, the subband decomposition technique is used to decompose the signal into 32 subbands, and the corresponding direction parameter corrections are independently applied to each subband. The directionally enhanced signal has more accurate sound source localization ability and stronger noise suppression effect compared to the original beamforming signal, especially prominent in a low signal-to-noise ratio environment.

[0122] Based on the directionally enhanced signal and the preset speech data, acoustic feature matching is performed to obtain the corresponding target actual short-time speech data. The acoustic feature matching process uses Mel Frequency Cepstral Coefficients (MFCC) feature extraction and Dynamic Time Warping (DTW) distance calculation to retrieve the most similar speech segment from the preset speech database. The preset speech database contains 8000 standard speech samples covering speakers of different genders, ages, and languages. The target actual short-time speech data represents the ideal speech signal characteristics for guiding subsequent signal enhancement processing.

[0123] Extract features from the target actual short-time speech data based on a preset meta-learning online calibration algorithm to obtain a target feature adaptation matrix. The meta-learning online calibration algorithm adopts a model-agnostic adaptive learning mechanism, including three core components: a feature extraction layer, a feature mapping layer, and an adaptation layer. The feature extraction layer uses the short-time Fourier transform to extract time-frequency features. The feature mapping layer identifies key feature points through an attention mechanism. The adaptation layer generates a 16×16-dimensional adaptation matrix. The target feature adaptation matrix contains the gain and phase adjustment parameters required for spectrum correction and has a high-dimensional feature expression ability.

[0124] Perform a convolution operation on the target feature adaptation matrix and the direction enhancement signal to obtain a spectrum correction signal. The convolution operation adopts a two-dimensional convolution method, with a convolution kernel size of 5×5, a stride of 1, and a padding mode of "same". During the convolution process, each element of the adaptation matrix is multiplied and accumulated with the corresponding frequency-time point of the signal to achieve fine adjustment of the spectrum. The spectrum correction signal has a clearer performance in harmonic structure and formants, effectively eliminating the spectrum distortion caused by environmental interference.

[0125] Perform cross-band consistency processing and phase-band gain on the spectrum correction signal to obtain a dynamic band enhancement signal. The cross-band consistency processing adopts an adjacent band smoothing algorithm, using a 5-point smoothing window to eliminate the discontinuity between bands. The phase-band gain adopts a group delay correction technique to maintain the coherence of the signal phase, with a phase adjustment range of ±0.2π. The dynamic band enhancement signal shows a low-frequency enhancement of 3 - 6 dB, a mid-frequency enhancement of 1 - 3 dB, and a high-frequency enhancement of 4 - 8 dB in the spectrum, and the overall enhancement characteristics are dynamically adjusted according to the environment.

[0126] In this embodiment, the beamforming signal is optimized for beam direction and gain parameters through an adversarial environment simulator, which can significantly improve the quality and clarity of the signal in various complex noise and reverberation environments. This method effectively adjusts the beam direction and band gain by simulating a variety of environmental interference samples and combining the gradient descent optimization algorithm, enabling the beamforming signal to adaptively adjust in the face of different interferences, thereby improving the robustness and adaptability of the system. Acoustic feature matching based on the target actual short-time speech data enables the system to more accurately restore the time-frequency features of the speech signal, further enhancing the intelligibility of the speech. Extracting and generating an adaptation matrix for the target features through the meta-learning online calibration algorithm effectively improves the spectrum quality of the signal, making the effect of the spectrum correction signal more significant. The final dynamic band enhancement signal can provide stronger enhancement effects in different frequency bands, especially outstanding in low signal-to-noise ratio environments.

[0127] In one embodiment, various noise interferences and reverberation conditions are simulated for the beamforming signal based on an adversarial environment simulator to obtain diverse environmental interference samples, including:

[0128] The adversarial environment simulator performs time-frequency transformation operations on the beamforming signal. Time-frequency transformation is a process of converting a time-domain signal into a time-frequency domain representation. The short-time Fourier transform (STFT) is used to decompose the beamforming signal into representations in the time and frequency dimensions. The result of the time-frequency transformation is a time-frequency domain representation signal matrix, which contains information about the energy distribution of the signal at different time points and frequencies.

[0129] A three-dimensional acoustic propagation structure is constructed based on the time-frequency domain representation signal matrix. The three-dimensional acoustic propagation structure is a mathematical model that simulates the propagation characteristics of sound waves in three-dimensional space. This structure takes into account factors such as spatial geometry, obstacle positions, and surface materials. During the construction process, the adversarial environment simulator sets random reflection coefficients to simulate the reflection characteristics of sound waves on surfaces of different materials. The randomization of the reflection coefficients ensures the diversity and authenticity of the simulated environment. Through this step, the system generates a set of spatial reverberation parameters, which includes key parameters such as room impulse response (RIR), early reflections, and late reverberation.

[0130] The adversarial environment simulator randomly selects various types of noise sources from a preset environmental noise library. The environmental noise library is a dataset containing recordings of various real environmental noises, such as traffic noise, crowd noise, machine noise, etc. Time-domain stretching is performed on these noise sources to change the duration and time characteristics of the noise; at the same time, frequency-domain offset processing is performed to adjust the frequency characteristics of the noise. Phase vocoder technology is used for time-domain stretching to keep the timbre characteristics of the noise unchanged; frequency-domain offset is achieved through spectral shifting. After these processes, the system obtains a set of deformed noise sources, which has higher diversity and unpredictability.

[0131] The adversarial environment simulator performs position mapping on the set of deformed noise sources according to a preset spatial directivity distribution function. The spatial directivity distribution function describes the distribution law of noise sources in space and defines the probability of a noise source appearing in a specific direction based on a probability density function. The position mapping process assigns the deformed noise sources to various directions in three-dimensional space, generating a multi-directional noise distribution matrix. This matrix represents the spatial distribution of noise sources from different directions, increasing the spatial diversity of noise interference.

[0132] The specific process of performing position mapping on the set of deformed noise sources according to a preset spatial directivity distribution function is as follows:

[0133] The spatial directivity distribution function is a mathematical description of the distribution law of noise sources in three-dimensional space. The position mapping process first assigns spatial position coordinates to each deformed noise source , where A represents the horizontal angle (azimuth angle), C represents the vertical angle (elevation angle), and r represents the distance.

[0134] For each noise source ni, the system samples the azimuth and elevation angles from the spatial directivity distribution function . This distribution function can be a uniform distribution, a Gaussian distribution, or a custom distribution based on actual environmental measurements. The distance parameter r is sampled based on a preset distance distribution function D(r), which defines the possible distance range between the noise source and the microphone array.

[0135] The mathematical expression for position mapping is:

[0136] ;

[0137] where M represents the mapping function, represents the i-th deformed noise source, represents the angle value sampled from the directivity distribution function, represents the distance value sampled from the distance distribution function.

[0138] After completing the position mapping for all deformed noise sources, the system constructs a multi-directional noise distribution matrix N. Each row of this matrix contains the position information and noise characteristic parameters of a noise source, forming a complete spatial distribution representation. This matrix will be used for subsequent calculations of acoustic propagation delay and energy attenuation to ensure that the simulated noise interference has real spatial characteristics. <able>

[0139] The adversarial environment simulator calculates the acoustic propagation delay and energy attenuation based on the multi-directional noise distribution matrix and the set of spatial reverberation parameters. Acoustic propagation delay refers to the time required for sound waves to propagate from the noise source to the receiving point, which is related to the source distance and the medium propagation speed; energy attenuation describes the degree of energy loss of sound waves during propagation, which is related to the propagation distance and the environmental absorption coefficient. By calculating the propagation path characteristics of each noise source to the microphone array, the system generates a noise propagation characteristic vector, which contains the delay, attenuation, and direction information of each noise source propagating to the receiving point.

[0140] The adversarial environment simulator performs reverberation processing on the time-frequency domain representation signal matrix based on the set of spatial reverberation parameters. Reverberation processing simulates the acoustic effect generated by multiple reflections of sound waves in an enclosed space and is achieved through convolution operation with the room impulse response. The processed signal is superimposed with the noise propagation characteristic vector to simulate the situation where the target signal is simultaneously interfered by multiple noise sources and reverberation in a real environment. The superimposition process takes into account the spatial position, propagation characteristics, and energy magnitude of the noise sources to ensure that the generated interference pattern is more realistic. Finally, diverse environmental interference samples are obtained, which reflect the signal characteristics in various complex acoustic environments.

[0141] The generation of diverse environmental interference samples greatly enriches the diversity of training data, enabling the beamforming algorithm to learn the ability to adapt to various complex environmental conditions, thereby improving the robustness and generalization ability of the system in practical applications. Through these samples generated by the adversarial environment simulator, the beamforming system can still maintain good signal enhancement effects under unseen noise and reverberation conditions.

[0142] In this embodiment, by introducing an adversarial environment simulator into the multi-microphone array beamforming signal enhancement method, various noise interferences and reverberation conditions can be effectively simulated, and diverse environmental interference samples are generated. This method realizes the accurate simulation of complex acoustic environments through time-frequency transformation, constructing a three-dimensional acoustic propagation structure, randomly extracting deformed noise sources, position mapping, and calculating acoustic propagation time delay and energy attenuation, thereby greatly enhancing the robustness and adaptability of beamforming signal enhancement. By performing time-domain stretching and frequency-domain offset on real environmental noise, the generated noise sources have higher diversity, which can effectively improve the system's ability to handle unexpected noise interferences. The introduction of the spatial reverberation parameter set and the multi-directional noise distribution matrix makes the generated interference samples more realistic, reflecting the reverberation and noise characteristics in the real environment. Finally, through reverberation processing and interference superposition on the time-frequency domain representation signal matrix, the generated diverse environmental interference samples greatly enrich the training data, contribute to the performance of the beamforming algorithm in complex environments, and thus enhance the signal enhancement effect and generalization ability of the system.

[0143] In one embodiment, based on a preset meta-learning online calibration algorithm, feature extraction is performed on target actual short-time speech data to obtain a target feature adaptation matrix, including:

[0144] Perform short-time window framing and windowing on the target actual short-time speech data to obtain an overlapping segmented signal sequence. Specifically, the original speech signal is framed using a 25-millisecond Hamming window, and the overlap between adjacent frames is 10 milliseconds. This processing is based on the assumption that the speech signal has quasi-steady-state characteristics in a short time, and the continuity and smooth transition of the signal are ensured through overlapping framing. After each speech frame is windowed, the edge decays, reducing the phenomenon of spectral leakage. The overlapping segmented signal sequence is a time-series set composed of multiple overlapping short-time speech segments, retaining the temporal characteristics and local spectral characteristics of the original speech.

[0145] Calculate the Mel-frequency cepstral coefficients (MFCCs) for the overlapping segmented signal sequences to obtain the segmented feature description vectors. The MFCC calculation process includes performing a fast Fourier transform (FFT) on each segmented signal to obtain the power spectrum, converting the linear frequency scale to the Mel frequency scale, and applying the discrete cosine transform (DCT) to obtain the cepstral coefficients. Usually, 13 - 26 dimensional MFCCs are extracted, and then combined with first-order and second-order difference coefficients to form a 39 - 78 dimensional feature vector. The segmented feature description vectors are a set of multi-dimensional vectors that can effectively characterize the acoustic characteristics of speech and capture the key information of speech in the spectral domain.

[0146] Construct a meta-learning task sampling set from the segmented feature description vectors to obtain a feature support set and a feature query set. The construction of the meta-learning task sampling set adopts the N-way K-shot strategy. Randomly select N classes from the segmented feature description vectors, with each class containing K samples to form the feature support set, and select some other samples to form the feature query set. The feature support set is used for model parameter initialization and inner-loop optimization, and the feature query set is used to evaluate the model's generalization ability and outer-loop optimization. Here, "class" refers to different acoustic environments or speaker feature patterns.

[0147] Construct a meta-learning inner-loop optimization function based on the feature support set to obtain the environment adaptation parameters. The meta-learning inner-loop optimization function uses the gradient descent algorithm to perform local optimization on the feature support set. The loss function is set as the prototype network metric loss, and the optimization goal is to minimize the Euclidean distance between the samples in the feature support set and the corresponding class prototypes. The environment adaptation parameters include the embedding function parameter matrix and the metric space transformation matrix under a specific environment, reflecting the adaptive adjustment results of the model to the current acoustic environment.

[0148] Perform MAML (Model-Agnostic Meta-Learning) gradient update on the feature query set according to the environment adaptation parameters to obtain the task model parameters. The MAML gradient update calculates the gradient of the loss function on the feature query set with respect to the environment adaptation parameters and uses the second-order derivative optimization method to adjust the parameters. The update process includes three steps: forward propagation to calculate the loss, backward propagation to calculate the gradient, and parameter update. The task model parameters are the set of model weights after environment adaptation and have the ability to quickly adapt to new environments.

[0149] Meta-learning outer-loop optimization is to integrate and optimize the meta-learning results of multiple tasks to improve the model's generalization ability and adaptability. The specific process is as follows:

[0150] Sample the tasks in the task pool to form multiple tasks, each task containing a feature support set and a feature query set. For each task, calculate the loss function on the feature query set based on the current task model parameters. The loss function usually uses cross-entropy loss or mean squared error, reflecting the performance of the model on this task.

[0151] Calculate the gradient of the loss function for each task to obtain the gradient of the loss function with respect to the model parameters. Average the gradients of all tasks to obtain the global gradient. Based on the global gradient, update the global parameters of the model through the gradient descent algorithm. The updated model parameters can better adapt to different tasks and improve the generalization ability of the model.

[0152] The above process of task sampling, loss calculation, gradient calculation, and parameter update is iterated repeatedly in the outer loop until the loss function converges or reaches the preset number of iterations. Through multiple iterations, the model parameters are gradually optimized and can better adapt to different acoustic environments and signal characteristics.

[0153] The purpose of online learning rate adjustment is to dynamically adjust the learning rate of the model so that it has the optimal learning rate in different environments and tasks. The specific process is as follows:

[0154] Online learning rate adjustment adopts the Bayesian optimization framework. Bayesian optimization is a global optimization method based on a probabilistic model. By constructing a probabilistic distribution model of the objective function, the optimal learning rate parameter is selected. The Bayesian optimization framework includes three main parts: prior distribution, likelihood function, and posterior distribution.

[0155] Based on the historical optimization trajectory and the current performance metrics, construct the prior distribution of the learning rate parameter. The prior distribution reflects the initial estimate of the learning rate parameter and can be described using a Gaussian process or other probabilistic distribution models.

[0156] Based on the loss function and gradient information of the current task, update the posterior distribution of the learning rate parameter. The posterior distribution combines the prior information and the observed data of the current task and more accurately reflects the optimal value of the learning rate parameter.

[0157] By maximizing the likelihood function of the posterior distribution, select the optimal learning rate parameter. The optimal learning rate parameter is the learning rate value that minimizes the loss of the model on the current task.

[0158] According to the optimal learning rate parameter, generate a dynamic learning rate matrix. The dynamic learning rate matrix contains adaptive learning rate values for different parameter layers and different tasks. The learning rate of each parameter layer is adjusted according to its importance and variability on different tasks, enabling the model to learn at the optimal rate in different environments and tasks.

[0159] Perform prototype network mapping on the segmented feature description vectors according to the dynamic learning rate matrix to obtain the environmental perception feature representation. The prototype network mapping process includes two main steps: feature embedding and prototype calculation. Feature embedding uses a deep neural network to map the segmented feature description vectors into a high-dimensional embedding space; prototype calculation obtains the class prototype vectors by averaging the embedding vectors of samples in the same class. The environmental perception feature representation is a set of high-dimensional feature vectors with environmental adaptability generated under the prototype network framework, retaining the key features of the speech signal while filtering out the influence of environmental noise.

[0160] Construct a metric learning embedding space for the environmental perception feature representation to obtain the speech feature embedding matrix. The construction of the metric learning embedding space uses the Mahalanobis distance metric framework, and by learning the optimal distance metric matrix, the distance between samples of the same class is minimized, and the distance between samples of different classes is maximized.

[0161] The process of constructing the metric learning embedding space includes three main links: initializing the metric matrix, iteratively optimizing the metric parameters, and verifying the metric performance. In the initialization stage, calculate the covariance matrix based on the environmental perception feature representation, and use the eigenvalue decomposition method to obtain the initial distance metric matrix. The covariance matrix reflects the correlation between different feature dimensions, and its inverse matrix is used to construct the Mahalanobis distance metric. The initial distance metric matrix serves as the initial parameter for metric learning, providing an initial metric benchmark for subsequent optimization.

[0162] In the iterative optimization stage, use the contrastive loss function to update the distance metric matrix. The optimization goal is to minimize the Mahalanobis distance between samples of the same class and maximize the Mahalanobis distance between samples of different classes. During the optimization process, construct positive and negative sample pairs, where the positive sample pairs come from the same class and the negative sample pairs come from different classes. Construct multiple positive and negative sample pairs through random sampling, calculate the gradient of the loss function, and use the gradient descent method to iteratively adjust the distance metric matrix to make the optimization goal gradually converge.

[0163] In the stage of verifying the metric performance, use classification accuracy and clustering metric indicators to evaluate the quality of the optimized metric space. The classification accuracy is evaluated by the performance of the nearest neighbor classifier in the metric space. The clustering metric indicators include within-class distance, between-class distance, and silhouette coefficient. The within-class distance measures the tightness of samples in the same class, the between-class distance measures the separation of samples in different classes, and the silhouette coefficient comprehensively evaluates the clustering effect. Evaluate the optimized metric learning embedding space through the above indicators, and adjust the training strategy according to the evaluation results.

[0164] The speech feature embedding matrix is a data structure containing multiple speech feature vectors. These feature vectors have good within-class aggregation and between-class separability in the learned metric space, ensuring the discrimination ability of speech features in different environments in the metric space.

[0165] Perform Bayesian credible interval constraint on the speech feature embedding matrix to obtain the target feature adaptation matrix. Bayesian credible interval constraint estimates the uncertainty interval of parameters by establishing a probability distribution model of feature parameters, and filters out unreliable feature representations. The constraint process includes four key steps: constructing the prior distribution of feature parameters, calculating the posterior distribution, determining the credible interval, and screening feature vectors. The target feature adaptation matrix is a set of feature representations screened by credibility, with higher robustness and environmental adaptability, providing a reliable feature basis for subsequent beamforming processing.

[0166] In this embodiment, by adopting a meta-learning online calibration algorithm for feature extraction in the multi-microphone array beamforming signal enhancement method, it can effectively adapt to different acoustic environments and improve the clarity and intelligibility of speech signals. Through short-time window framing and windowing processing, as well as Mel-frequency cepstral coefficient calculation, the time and spectral features of the original speech are carefully retained. By constructing a meta-learning task sampling set and an inner-loop optimization function, the model parameters can be quickly adjusted to adapt to the new acoustic environment, enhancing the flexibility and adaptability of the model. MAML gradient update and meta-learning outer-loop optimization ensure the efficient learning ability of the model on different tasks, and through online learning rate adjustment, the learning rate adaptively changes, further improving the training efficiency and stability of the model. Through the construction of a metric learning embedding space, speech features have good intra-class aggregation and inter-class separability in the high-dimensional metric space, ensuring the discrimination ability of speech features in different environments. The entire method not only improves the speech recognition accuracy in noisy environments but also enhances the robustness and practicality of the system under complex acoustic conditions.

[0167] In one embodiment, perform inverse time-frequency conversion on the dynamic band-enhanced signal and perform low-latency optimization to obtain the target quality speech output signal, including:

[0168] The dynamic band-enhanced signal is converted into an initial time-domain signal through overlap-and-add inverse transformation. Overlap-and-add inverse transformation is a method of converting a frequency-domain signal back to the time domain. This processing process uses the inverse short-time Fourier transform (ISTFT). In this process, the frequency-domain data is divided into several frames, the inverse Fourier transform is performed on each frame of data, and overlap-and-add is performed in the time domain. The overlap factor is usually set to 50% or 75% to ensure smooth transition. The initial time-domain signal retains the basic time-domain characteristics of the original sound but may have problems with phase discontinuity.

[0169] The initial time-domain signal is processed through inter-frame phase consistency compensation to obtain a phase-continuous intermediate signal. This step aims to address the possible phase discontinuity problem between adjacent frames, avoiding the generation of artificial synthesized sound effects and distortion. Phase consistency compensation is achieved by calculating the phase difference between adjacent frames and applying a phase correction factor to achieve smooth phase transitions. For each frequency component, the system calculates its phase continuity metric, and when a phase jump exceeding a preset threshold (usually π / 2) is detected, phase correction is performed. The phase-continuous intermediate signal has more natural sound characteristics, reducing the traces of artificial synthesis.

[0170] The phase-continuous intermediate signal undergoes buffer dynamic adjustment processing to obtain a variable signal buffer. Buffer dynamic adjustment is based on signal characteristics and processing requirements, adaptively adjusting the buffer size. Based on voice activity detection (VAD), a smaller buffer is used in the silent segment to reduce latency, while a larger buffer is used in the voice active segment to ensure processing quality. The buffer size adjustment range is between 10 ms and 50 ms, and the adjustment step size is 5 ms. The variable signal buffer effectively balances the relationship between processing latency and signal quality.

[0171] The variable signal buffer undergoes non-linear amplitude normalization processing to obtain a dynamic range compression signal. Non-linear amplitude normalization uses logarithmic compression or hyperbolic tangent functions to perform non-linear mapping on the signal amplitude, compressing high-amplitude signals while enhancing low-amplitude signals. The normalization parameters are dynamically adjusted according to the statistical characteristics of the signal, including the compression ratio (1.5:1 to 3:1) and threshold (-20 dB to -30 dB). The dynamic range compression signal has a more balanced sound energy distribution, enhancing the audibility and intelligibility of speech.

[0172] The dynamic range compression signal undergoes time-domain transient feature vector calculation and time-domain jitter elimination to obtain a jitter-corrected signal. The time-domain transient feature vector is calculated by analyzing parameters such as the short-time energy change, zero-crossing rate, and spectral entropy of the signal. Jitter detection is based on the degree of mutation of the feature vector. When jitter is detected (the feature vector change exceeds a preset threshold of 25%), an adaptive smoothing filter is applied for jitter suppression. The jitter elimination process retains the transient characteristics of the speech signal while removing instability. The jitter-corrected signal has a more stable speech quality.

[0173] The jitter correction signal undergoes multi - stage cascaded frequency band reconstruction processing to obtain the frequency response equalization signal. The multi - stage cascaded frequency band reconstruction uses a segmented multi - resolution filter bank to decompose the signal into multiple sub - frequency bands (usually 4 to 8 sub - frequency bands), which are enhanced separately and then synthesized. The low - frequency band (0 - 1 kHz) focuses on retaining the fundamental frequency and timbre information, the middle - frequency band (1 - 4 kHz) focuses on improving speech clarity, and the high - frequency band (4 - 8 kHz) focuses on restoring speech details. The frequency response equalization signal has a more comprehensive and balanced spectrum distribution, and the clarity and naturalness of the speech are significantly improved.

[0174] The frequency response equalization signal undergoes device pre - compensation processing to obtain the device - matching signal. The device pre - compensation is based on the frequency response characteristics of the target playback device and applies inverse filtering technology for pre - correction to compensate for the distortion that the device may introduce. The compensation processing includes frequency response correction, phase correction, and dynamic range adjustment, etc. The correction parameters are obtained from the device characteristic database, and different compensation templates are applied for different types of devices (such as headphones, speakers, or mobile devices). The device - matching signal is optimized for a specific playback device to ensure the best auditory experience in the actual usage environment.

[0175] The device - matching signal undergoes the frame - selective discarding algorithm processing to achieve noise and artifact elimination, and finally obtains the target - quality speech output signal. The frame - selective discarding algorithm scores the quality of each frame of the signal based on signal quality evaluation metrics. These evaluation metrics include signal - to - noise ratio (SNR), voice activity detection (VAD), and spectral distortion ratio (SDR). In the SNR evaluation, the system calculates the ratio between the energy of each frame of the signal and the background noise. When the SNR is lower than a preset threshold (such as 10 dB), the quality of that frame is considered poor. The voice presence probability is determined by the voice activity detection (VAD) algorithm. When the voice presence probability is lower than a preset threshold (such as 0.4), that frame is considered a non - speech frame. The spectral distortion ratio is evaluated by comparing the spectral differences between the current frame and its adjacent frames before and after. When the spectral distortion ratio exceeds a preset threshold (such as 0.2), artifacts are considered to exist in that frame.

[0176] When the signal quality score is lower than the preset threshold, that frame will be discarded or replaced with an interpolated synthesized frame. The interpolated synthesized frame is generated based on the data of the adjacent frames before and after through linear interpolation or spline interpolation techniques. This processing method ensures the continuity and natural transition of the signal.

[0177] The noise artifact cancellation process uses wavelet decomposition and reconstruction techniques. First, the device-matched signal is decomposed by wavelet decomposition into detail and approximation components at several levels. The detail components contain high-frequency noise and artifact information, while the approximation components retain the main features of the signal. In each level of detail components, the system applies a thresholding method (such as soft threshold or hard threshold) for noise suppression, and the threshold is dynamically set according to the noise standard deviation. The processed detail components are recombined with the approximation components, and the denoised signal is obtained through wavelet reconstruction.

[0178] During the noise artifact cancellation process, the system also combines directional filtering techniques to perform directional processing on the detected artifact regions. The directional filtering technique is based on the transient characteristics and local statistical properties of the signal, and adaptively adjusts the filter parameters to maximize the suppression of artifacts while retaining the natural characteristics of the speech. The processed signal is iteratively optimized multiple times to obtain the target quality speech output signal.

[0179] In this embodiment, through a series of signal processing steps, the speech signal quality of the multi-microphone array beamforming system is effectively improved. Through the frame selective discard algorithm, the quality of each frame of the signal can be accurately evaluated, and dynamic adjustment is performed according to the signal quality score. Low-quality frames are discarded or replaced to ensure the coherence and naturalness of the speech signal. This process effectively eliminates noise interference and artifacts, avoids artificial synthesis traces, and makes the speech signal clearer and more natural. Through wavelet decomposition and reconstruction techniques, noise suppression can be refined during the denoising process, improving the signal fidelity. In addition, by combining directional filtering techniques, the system adaptively suppresses artifacts and effectively retains the original features of the speech, further improving the intelligibility of the speech.

[0180] Refer to Figure 2 As shown, the present invention also provides a multi-microphone array beamforming signal enhancement device, which is applied to the multi-microphone array beamforming signal enhancement method of any one of the above, and includes:

[0181] An acquisition module, which is used to acquire multi-channel time-domain signals collected by the multi-microphone array, perform time-frequency conversion, and obtain complex time-frequency domain signals;

[0182] An analysis module, which is used to obtain the spatial distribution vector of the multi-microphone array, perform deformation compensation processing on the complex time-frequency domain signals, and obtain a sound field consistency signal;

[0183] An association module, which is used to extract the sound source direction characteristics of the sound field consistency signal and perform complex spatio-temporal attention beamforming network processing to obtain a beamforming signal;

[0184] A processing module, which is used to optimize the beam direction and gain parameters of the beamforming signal through a preset adversarial environment simulator to obtain a dynamically band-enhanced signal;

[0185] A control module, which is used to perform inverse time-frequency conversion on the dynamically band-enhanced signal and perform low-latency optimization to obtain a target-quality voice output signal.

[0186] A multi-microphone array beamforming signal enhancement device provided by the present invention can perform more accurate signal deformation compensation by performing time-frequency conversion on the time-domain signal collected by the multi-microphone array and combining the spatial distribution vector information of the array, improving the quality of the sound field consistency signal. This method can effectively improve the extraction accuracy of the target voice signal by extracting the direction characteristics of the sound source and applying a complex spatio-temporal attention beamforming network, avoiding the occurrence of signal attenuation or distortion in traditional signal enhancement methods. By introducing a preset adversarial environment simulator to optimize the beam direction and gain parameters, it can adapt to a dynamic and complex acoustic environment, realizing the adaptive adjustment of the beam direction and gain parameters, and further improving the clarity of the signal and the accuracy of speech recognition. Based on this technology, it can effectively reduce noise interference and echo effects, thus ensuring high-quality voice output signals in different environments and reducing the problems of sound quality degradation and speech distortion in traditional methods. In addition, the low-latency optimization design of the dynamically band-enhanced signal helps the smooth operation of the real-time voice interaction system, enhancing the robustness and user experience of the system. By comprehensively considering spatio-temporal information and dynamic sound field characteristics, this method can flexibly adapt to complex and changing acoustic environments, improving the overall performance and application effect of the multi-microphone array system.

[0187] In this embodiment, the processor and the memory can be connected through a bus or other means. The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive. The processor may be a general-purpose processor, such as a central processing unit, a digital signal processor, an application-specific integrated circuit, or one or more integrated circuits configured to implement the embodiments of the present invention. <==

[0188] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system and each module can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0189] The above are only the preferred embodiments of the present invention, and do not thereby limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present invention.

Claims

1. A method for enhancing multi-microphone array beamforming signals, characterized in that Including: Obtain multi-channel time-domain signals collected by a multi-microphone array, perform time-frequency conversion to obtain complex time-frequency domain signals; Obtain the spatial distribution vector of the multi-microphone array, perform deformation compensation processing on the complex time-frequency domain signals to obtain sound field consistency signals; Extract source direction features from the sound field consistency signals, and perform complex spatio-temporal attention beamforming network processing to obtain beamforming signals; Optimize the beam direction and gain parameters of the beamforming signals through a preset adversarial environment simulator to obtain dynamic band enhancement signals; Perform inverse time-frequency conversion on the dynamic band enhancement signals and perform low-latency optimization to obtain target quality speech output signals; The extracting source direction features from the sound field consistency signals and performing complex spatio-temporal attention beamforming network processing to obtain beamforming signals includes: Perform adaptive wavelet packet transform decomposition on the sound field consistency signals to obtain multi-level sound field time-frequency representations; Perform complex-domain spatial covariance calculation on the multi-level sound field time-frequency representations to obtain band phase difference spectra; Perform high-order singular value decomposition on the band phase difference spectra to obtain an orthogonal projection structure; Construct a hyperbolic direction estimation function based on the orthogonal projection structure and perform gradient polarization search to obtain a source direction vector; Align the phases of the source direction vector and the sound field consistency signals to obtain initial beam gain coefficients; Construct a complex-domain attention mask based on the initial beam gain coefficients and perform channel weighting on the sound field consistency signals to obtain beamforming signals.

2. The multi-microphone array beamforming signal enhancement method according to claim 1, wherein The obtaining multi-channel time-domain signals collected by a multi-microphone array, performing time-frequency conversion to obtain complex time-frequency domain signals includes: Perform endpoint detection and noise power estimation on the multi-channel time-domain signals to obtain preprocessed signals; Frame and window the preprocessed signals according to preset frame length and frame shift parameters to obtain initial complex frequency domain signals; Perform spectral smoothing on the initial complex frequency domain signals to obtain smoothed frequency domain signals; Perform non-linear sub-band division on the smoothed frequency domain signals to obtain sub-band signals; Perform geometric structure phase compensation on the sub-band signals to obtain sub-band corrected signals; Perform spectral subtraction denoising on the sub-band corrected signals based on a preset noise power spectrum and perform complex time-frequency domain feature representation conversion to obtain the complex time-frequency domain signals.

3. The multi-microphone array beamforming signal enhancement method according to claim 1, wherein The obtaining the spatial distribution vector of the multi-microphone array, performing deformation compensation processing on the complex time-frequency domain signals to obtain sound field consistency signals includes: Obtain the physical coordinates and orientation parameters of the multi-microphone array, perform spatial construction to obtain the spatial distribution vector; Perform spatial transfer function mapping calculation on the complex time-frequency domain signals according to the spatial distribution vector to obtain array acoustic characteristic mapping information; Perform singular value decomposition on the array acoustic characteristic mapping information and perform phase and amplitude non-linear correction to obtain preliminary compensation signals; Calculate the cross-power spectral density matrix based on the preliminary compensation signals and perform sound field geometric consistency constraint construction to obtain sound field topology representations; Perform complex domain tensor decomposition on the acoustic field topological characterization to obtain a deformation compensation coefficient; Align the phase and amplitude of the complex time-frequency domain signal according to the deformation compensation coefficient to obtain the acoustic field consistency signal.

4. The multi-microphone array beamforming signal enhancement method according to claim 1, wherein The optimization of the beam direction and gain parameters of the beamforming signal through a preset adversarial environment simulator to obtain a dynamic frequency band enhancement signal includes: Simulate various noise interferences and reverberation conditions on the beamforming signal based on the adversarial environment simulator to obtain diversified environmental interference samples; Perform gradient descent optimization calculation on the diversified environmental interference samples, and adjust the beam direction parameters and frequency band gain coefficients according to a preset signal-to-noise ratio objective function to obtain a parameter optimization matrix; Reconstruct the directivity of the beamforming signal according to the parameter optimization matrix to obtain a direction enhancement signal; Match the acoustic features of the direction enhancement signal with preset speech data to obtain corresponding target actual short-time speech data; Extract features from the target actual short-time speech data based on a preset meta-learning online calibration algorithm to obtain a target feature adaptation matrix; Perform convolution operation on the target feature adaptation matrix and the direction enhancement signal to obtain a spectrum correction signal; Perform cross-frequency band consistency processing and phase-frequency band gain on the spectrum correction signal to obtain the dynamic frequency band enhancement signal.

5. The multi-microphone array beamforming signal enhancement method according to claim 4, wherein The simulation of various noise interferences and reverberation conditions on the beamforming signal based on the adversarial environment simulator to obtain diversified environmental interference samples includes: Perform time-frequency transformation on the beamforming signal through the adversarial environment simulator to obtain a time-frequency domain representation signal matrix; Construct a three-dimensional acoustic propagation structure according to the time-frequency domain representation signal matrix and set random reflection coefficients to obtain a set of spatial reverberation parameters; Randomly extract various types of noise sources from a preset environmental noise library, and perform time-domain stretching and frequency-domain offset to obtain a set of deformed noise sources; Perform position mapping on the set of deformed noise sources according to a preset spatial directivity distribution function to obtain a multi-directional noise distribution matrix; Calculate the acoustic propagation time delay and energy attenuation according to the multi-directional noise distribution matrix and the set of spatial reverberation parameters to obtain a noise propagation characteristic vector; Perform reverberation processing on the time-frequency domain representation signal matrix based on the set of spatial reverberation parameters, and perform interference superposition with the noise propagation characteristic vector to obtain the diversified environmental interference samples.

6. The multi-microphone array beamforming signal enhancement method according to claim 4, wherein The extraction of features from the target actual short-time speech data based on a preset meta-learning online calibration algorithm to obtain a target feature adaptation matrix includes: Perform short-time window framing and windowing processing on the target actual short-time speech data to obtain an overlapping segmented signal sequence; Calculate the Mel-frequency cepstral coefficients of the overlapping segmented signal sequence to obtain a segmented feature description vector; Construct a meta-learning task sampling set for the segmented feature description vector to obtain a feature support set and a feature query set; Construct a meta-learning inner loop optimization function based on the feature support set to obtain environment adaptation parameters; Perform MAML gradient update on the feature query set according to the environment adaptation parameters to obtain task model parameters; Perform meta - learning outer - loop optimization and online learning rate adjustment on the task model parameters to obtain a dynamic learning rate matrix; Perform prototype network mapping on the segmented feature description vectors according to the dynamic learning rate matrix to obtain an environment - aware feature representation; Construct a metric - learning embedding space for the environment - aware feature representation to obtain a speech feature embedding matrix; Perform Bayesian credible interval constraint on the speech feature embedding matrix to obtain a target feature adaptation matrix.

7. The multi-microphone array beamforming signal enhancement method according to claim 1, wherein The inverse time - frequency conversion and low - latency optimization of the dynamic band - enhanced signal to obtain a target - quality speech output signal include: Perform an overlap - add inverse transform process on the dynamic band - enhanced signal to obtain an initial time - domain signal; Perform inter - frame phase consistency compensation on the initial time - domain signal to obtain a phase - continuous intermediate signal; Perform buffer dynamic adjustment according to the phase - continuous intermediate signal to obtain a variable signal buffer; Perform non - linear amplitude normalization on the variable signal buffer to obtain a dynamic range compression signal; Calculate the time - domain transient feature vector according to the dynamic range compression signal and perform time - domain jitter cancellation to obtain a jitter - corrected signal; Perform multi - stage cascaded band reconstruction on the jitter - corrected signal to obtain a frequency - response equalized signal; Perform device pre - compensation processing on the frequency - response equalized signal to obtain a device - matched signal; Eliminate noise and artifacts from the device - matched signal according to a preset frame - selective discard algorithm to obtain the target - quality speech output signal.

8. A multi-microphone array beamforming signal enhancement device, characterized in that, Applied to the multi - microphone array beamforming signal enhancement method described in any one of claims 1 - 7 above, it includes: An acquisition module, which is used to acquire multi - channel time - domain signals collected by a multi - microphone array, perform time - frequency conversion to obtain complex time - frequency domain signals; An analysis module, which is used to obtain the spatial distribution vector of the multi - microphone array, perform deformation compensation processing on the complex time - frequency domain signals to obtain a sound - field consistency signal; An association module, which is used to extract source - direction features from the sound - field consistency signal and perform complex spatio - temporal attention beamforming network processing to obtain a beamforming signal; A processing module, which is used to optimize the beam direction and gain parameters of the beamforming signal through a preset adversarial environment simulator to obtain a dynamic band - enhanced signal; A control module, which is used to perform inverse time - frequency conversion on the dynamic band - enhanced signal and perform low - latency optimization to obtain a target - quality speech output signal; The extraction of source - direction features from the sound - field consistency signal and the performance of complex spatio - temporal attention beamforming network processing to obtain a beamforming signal include: Perform adaptive wavelet packet transform decomposition on the sound - field consistency signal to obtain a multi - level sound - field time - frequency representation; Perform complex - domain spatial covariance calculation on the multi - level sound - field time - frequency representation to obtain a band - phase difference spectrum; Perform high - order singular value decomposition on the band - phase difference spectrum to obtain an orthogonal projection structure; Construct a hyperbolic direction estimation function according to the orthogonal projection structure and perform gradient polarization search to obtain a source - direction vector; Perform phase alignment on the sound source direction vector and the sound field consistency signal to obtain an initial beam gain coefficient; Construct a complex domain attention mask according to the initial beam gain coefficient, and perform channel weighting on the sound field consistency signal to obtain a beamforming signal.

Citation Information

Patent Citations

  • Speech enhancement method, device and equipment

    CN112712818A

  • Microphone array calibration method

    CN116567515A

  • Adaptive beam forming anti-interference method based on large-scale array system

    CN118249869A