Intelligent microphone pickup and speech enhancement method, system and device
Audio signals are collected through dual silicon microphone arrays, vocal signal correlation feature extraction and sound source positioning, combined with frequency characteristic analysis and dynamic gain processing, the problem of multiple acoustic interference coupling in complex acoustic environments is solved, high-quality speech enhancement and noise suppression are achieved, and system performance is significantly improved.
Patent Information
- Application Number
- CN202510663002.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speech signal processing methods are difficult to effectively solve the coupling problem of multiple acoustic interference in complex acoustic environments, resulting in speech signal distortion and system performance degradation.
The audio signal is collected by a dual-silicon microphone array, and the correlation characteristics of the vocal signal are extracted through cross-correlation analysis, and the sound source positioning and target vocal signal extraction are performed. Then, frequency characteristic analysis and segmented dynamic gain processing are performed, combined with the adaptive filtering of noise components, noise suppression parameters are generated, and the speech enhancement signal and noise suppression parameters are input to the echo cancellation model for joint optimization.
It effectively improves the quality of voice acquisition in complex acoustic environments, ensures the clarity and nature of voice signals, reduces background noise interference, significantly improves the overall performance of the system, and can better adapt to the needs of practical application scenarios.
Smart Images

Figure CN120186515A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microphones, and particularly to an intelligent microphone sound pickup and voice enhancement method, system and device. Background Art
[0002] As an important way of human-computer interaction, intelligent voice interaction technology plays a key role in many application scenarios such as smart home and remote conferencing. With the continuous improvement of people's requirements for voice interaction quality, how to achieve clear and reliable voice collection and enhancement in complex acoustic environments has become one of the key research areas. Existing voice signal processing methods often only focus on a single technical link, such as simple noise cancellation or echo suppression, while ignoring the dynamic change characteristics of the acoustic environment and the coupling effect of multiple interference sources. This isolated processing method is likely to cause distortion of the voice signal, reduce the overall performance of the system, and cannot meet the requirements for voice quality in practical applications. Summary of the Invention
[0003] The main object of the present invention is to provide an intelligent microphone sound pickup and voice enhancement method, system and device, which can effectively solve the coupling problem of multiple acoustic interferences.
[0004] To achieve the above object, the present invention provides an intelligent microphone sound pickup and voice enhancement method, including: Obtaining an audio signal collected by a dual-silicon microphone array, performing cross-correlation analysis on the audio signal to obtain a corresponding correlation feature of the human voice signal; Performing sound source localization analysis on the audio signal based on the correlation feature of the human voice signal to obtain a corresponding target human voice signal; Performing frequency characteristic analysis on the target human voice signal, and performing segmented dynamic gain processing on the signal according to a preset frequency band range to obtain a corresponding voice enhancement signal; Performing adaptive filtering processing on the noise component of the target human voice signal to obtain a corresponding noise suppression parameter; Inputting the voice enhancement signal and the noise suppression parameter into a preset echo cancellation model for joint optimization to obtain a corresponding output human voice signal.
[0005] Further, the obtaining an audio signal collected by a dual-silicon microphone array, performing cross-correlation analysis on the audio signal to obtain a corresponding correlation feature of the human voice signal includes: Performing cross-correlation calculation between microphones of the dual-silicon microphone array to obtain corresponding human voice correlation data; Performing sub-band decomposition on the audio signal to obtain a plurality of sub-band signals; Extract the phase difference feature of the frequency band sub-signal based on the voice correlation data to obtain the corresponding phase difference feature spectrum; Perform non-linear transformation on the phase difference feature spectrum to obtain a frequency band correlation matrix; Calculate the energy concentration degree of the frequency band correlation matrix to obtain the frequency band energy feature; Perform joint analysis on the frequency band energy feature and the phase difference feature spectrum to obtain a joint feature vector; Perform signal feature fusion processing on the joint feature vector, the voice signal correlation feature.
[0006] Further, perform sound source localization analysis on the audio signal based on the voice signal correlation feature to obtain the corresponding target voice signal, including: Perform delay analysis on the voice signal correlation feature to obtain the corresponding signal delay difference; Perform response analysis on the audio signal according to the signal delay difference to obtain the corresponding sound source direction response value; Perform phase difference operation on the sound source direction response value to obtain a direction weight coefficient; Perform signal recognition on the audio signal based on the direction weight coefficient to obtain the corresponding initial target signal; Perform delay compensation and loudness compensation on the audio signal according to the direction weight coefficient to obtain a compensation signal; Perform convolution operation on the initial target signal and the compensation signal to obtain the corresponding target voice signal.
[0007] Further, perform frequency characteristic analysis on the target voice signal and perform segmented dynamic gain processing on the signal according to the preset frequency band range to obtain the corresponding voice enhancement signal, including: Perform frequency band division processing on the target voice signal according to the preset frequency band range to obtain the corresponding voice frequency band signal; Calculate the signal amplitude of the voice frequency band signal to obtain the corresponding signal amplitude range; Perform segmented processing on the voice frequency band signal according to the signal amplitude range to obtain a first signal amplitude interval segment, a second signal amplitude interval segment, and a third signal amplitude interval segment; Perform a 50 dB gain process on the first signal amplitude interval segment to obtain the corresponding first gain signal; Perform a 30 dB gain process on the second signal amplitude interval segment to obtain the corresponding second gain signal; Perform a 15 dB gain process on the third signal amplitude interval segment to obtain the corresponding third gain signal; Perform a synthesis process on the first gain signal, the second gain signal, and the third gain signal to obtain the speech enhancement signal; Among them, the amplitude range of the first signal amplitude interval segment is [-70dB, -50dB], the amplitude range of the second signal amplitude interval segment is [-50dB, -30dB], and the amplitude range of the third signal amplitude interval segment is [-30dB, -15dB].
[0008] Further, the adaptive filtering process for the noise component of the target human voice signal to obtain the corresponding noise suppression parameter includes: Perform a spectrum analysis on the target human voice signal to obtain the corresponding frequency-domain feature sequence; Perform a sub-band decomposition process on the target human voice signal according to the frequency-domain feature sequence to obtain a plurality of frequency sub-band signals; Perform a short-time feature distribution analysis on the frequency sub-band signals to obtain the corresponding sub-band distribution features; Perform a sliding window calculation on the frequency sub-band signals according to the sub-band distribution features to obtain the corresponding noise statistical parameters; Perform an adaptive threshold update on the noise statistical parameters to obtain the corresponding noise suppression threshold; Perform an adaptive filtering process on the frequency sub-band signals according to the noise suppression threshold to obtain the noise suppression parameter.
[0009] Further, the input of the speech enhancement signal and the noise suppression parameter into a preset echo cancellation model for joint optimization to obtain the corresponding output human voice signal includes: Input the speech enhancement signal into the echo cancellation model, and perform frame division and time-frequency transformation processing on the speech enhancement signal through the signal preprocessing layer of the echo cancellation model to obtain the corresponding frequency-domain feature information; Perform an acoustic feature analysis on the speech enhancement signal through the feature analysis layer of the echo cancellation model to obtain the corresponding acoustic feature parameters; Input the noise suppression parameter into the echo cancellation model, and perform a voice activity detection on the acoustic feature parameters through the double-talk detection layer of the echo cancellation model to obtain the corresponding voice state information; Perform a non-linear echo path analysis and feature extraction on the frequency-domain feature information through the echo path layer of the echo cancellation model to obtain the corresponding echo feature information; Perform a joint enhancement on the acoustic feature parameters, the voice state information, the echo feature information, and the noise suppression parameter through the joint optimization layer of the echo cancellation model to obtain the corresponding enhanced speech signal; The inverse time-frequency transformation and frame overlap and add processing are performed on the enhanced speech signal through the signal reconstruction layer of the echo cancellation model to obtain and output the output human voice signal.
[0010] Further, the joint optimization layer of the echo cancellation model jointly enhances the acoustic feature parameters, the speech state information, the echo feature information, and the noise suppression parameters to obtain corresponding enhanced speech signals, including: Perform multi-dimensional feature fusion on the acoustic feature parameters to obtain corresponding acoustic feature vectors; Adjust the weights of the acoustic feature vectors according to the speech state information to obtain corresponding weighted feature matrices; Perform non-linear mapping on the echo feature information to obtain corresponding echo estimation parameters; Perform joint filtering processing on the weighted feature matrix according to the echo estimation parameters and the noise suppression parameters to obtain corresponding filtering coefficients; Correct the noise suppression parameters according to the filtering coefficients to obtain corresponding corrected suppression parameters; Perform residual echo compensation according to the corrected suppression parameters and the echo estimation parameters to obtain corresponding corrected echo signals; Calculate the spectral gain of the corrected echo signal to obtain corresponding spectral gain coefficients; Perform spectral subtraction processing on the corrected echo signal according to the spectral gain coefficients to obtain the enhanced speech signal.
[0011] The present invention also provides an intelligent microphone sound pickup and speech enhancement system, which is applied to the intelligent microphone sound pickup and speech enhancement method described in any one of the above, including: An acquisition module, which is used to acquire the audio signal collected by the dual silicon microphone array, perform cross-correlation analysis on the audio signal, and obtain corresponding human voice signal correlation characteristics; An analysis module, which is used to perform sound source localization analysis on the audio signal based on the human voice signal correlation characteristics to obtain corresponding target human voice signals; An association module, which is used to perform frequency characteristic analysis on the target human voice signal and perform segmented dynamic gain processing on the signal according to a preset frequency band range to obtain corresponding speech enhancement signals; A processing module, which is used to perform adaptive filtering processing on the noise components of the target human voice signal to obtain corresponding noise suppression parameters; A control module, which is used to input the speech enhancement signal and the noise suppression parameters into a preset echo cancellation model for joint optimization to obtain corresponding output human voice signals.
[0012] The present invention also provides an intelligent microphone sound pickup and voice enhancement device, which is characterized by comprising: A memory for storing programs; A processor for executing the programs to implement each step of an intelligent microphone sound pickup and voice enhancement method as described in any one of the above.
[0013] An intelligent microphone sound pickup and voice enhancement method, system and device provided by the present invention have the following beneficial effects: Through cross-correlation analysis and sound source localization of the microphone audio signal, precise capture of the target human voice signal is achieved, effectively improving the voice collection quality in a complex acoustic environment. Through frequency band analysis and dynamic gain processing of the target human voice signal, refined enhancement of voice signals in different frequency bands is realized, ensuring the clarity and naturalness of the voice signal. Based on the adaptive noise filtering processing of the target human voice signal, the interference of background noise is effectively reduced, and the anti-noise ability of the system in a complex environment is improved. By jointly optimizing the voice enhancement signal and the noise suppression parameter, dynamic adjustment of the echo cancellation model is achieved, effectively solving the coupling problem of multiple acoustic interferences and significantly improving the overall performance of the system. Through the organic combination of multi-dimensional signal processing and optimization strategies, the limitations of traditional single processing methods are overcome, and the system can better meet the requirements of various actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a flowchart of an intelligent microphone sound pickup and voice enhancement method provided by the present invention; Figure 2 is a structural diagram of an intelligent microphone sound pickup and voice enhancement system provided by the present invention; Figure 3 is a structural diagram of an intelligent microphone sound pickup and voice enhancement device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] Next, with reference to the accompanying drawings and specific embodiments, the present invention will be further described.
[0017] Referring to Figure 1 as shown, an intelligent microphone sound pickup and voice enhancement method provided by the present invention includes: Step S1: Obtain the audio signal collected by the dual-silicon microphone array, perform cross-correlation analysis on the audio signal, and obtain the corresponding correlation characteristics of the human voice signal; Step S2: Based on the correlation characteristics of the human voice signal, perform sound source localization analysis on the audio signal to obtain the corresponding target human voice signal; Step S3: Perform frequency characteristic analysis on the target human voice signal, and perform segmented dynamic gain processing on the signal according to a preset frequency band range to obtain the corresponding voice enhancement signal; Step S4: Perform adaptive filtering processing on the noise components of the target human voice signal to obtain the corresponding noise suppression parameters; Step S5: Input the voice enhancement signal and the noise suppression parameters into a preset echo cancellation model for joint optimization to obtain the corresponding output human voice signal.
[0018] Based on the above steps, the detailed step process is as follows: Step S1: The dual-silicon microphones are installed at a fixed spacing (usually within 4 cm), and the two microphones simultaneously collect the sound field signal in the environment. The collected audio signal will be converted into a digital signal and preprocessed, including signal normalization and preliminary noise reduction. Perform cross-correlation analysis on the signals collected by the two microphones. This process uses the cross-correlation function to calculate the similarity between the two signals. Specifically, by calculating the signal correlation coefficients at different time delays, a correlation coefficient sequence is formed. This sequence reflects the time difference characteristics of the two microphone signals, and these characteristics are used to distinguish the human voice and environmental noise. Since the human voice signal has strong correlation between the two microphones, while environmental noise usually exhibits random characteristics, the correlation characteristics of the human voice signal are extracted through cross-correlation analysis.
[0019] Step S2: Based on the correlation characteristics of the human voice signal, calculate the time delay between the two microphone signals. Since the different source positions will cause the time difference of the sound wave arriving at the two microphones, this time difference is determined by the position of the cross-correlation peak. Combining the known microphone spacing and the sound propagation speed, use geometric operation methods to calculate the direction angle of the sound source. After obtaining the sound source direction, establish a spatial filter, which can enhance the sound signal from the target direction, while suppressing the interference signals from other directions, and extract the target human voice signal. This process involves the spatial characteristics of the sound field and the influence of room reverberation to ensure the accuracy of sound source localization. The extracted target human voice signal will be used as the input for subsequent voice enhancement processing.
[0020] Step S3: After obtaining the target human voice signal, perform frequency characteristic analysis and dynamic gain processing. Conduct spectrum analysis on the target human voice signal, convert the time-domain signal to the frequency domain through fast Fourier transform (FFT), and analyze the energy distribution of the signal in different frequency bands. The main frequency range of the human voice signal is usually between 400 Hz and 4 kHz. Based on the preset frequency band range, adopt a segmented dynamic gain strategy for processing: for the weak signal segment from -70 dB to -50 dB, perform a 50 dB gain boost; for the signal segment from -50 dB to -30 dB, perform a 30 dB gain boost; for the signal segment from -30 dB to -15 dB, perform a 15 dB gain boost. During the processing, monitor the peak value of the signal in real time to ensure that the gain-adjusted signal does not clip. The finally obtained voice enhancement signal has better clarity and audibility while maintaining the original voice characteristics. This signal will serve as the basis for subsequent noise suppression processing.
[0021] Step S4: Based on the obtained enhanced signal, analyze its noise characteristics. This process uses an adaptive filtering algorithm, usually the least mean square (LMS) or normalized least mean square (NLMS) algorithm. Estimate the spectral characteristics of the background noise in the non-speech segment and establish a noise statistical model. In the speech segment, update the filter coefficients adaptively to track the changes in the noise characteristics in real time. This adaptive processing method can effectively cope with noise changes in different environments. During the processing, dynamically adjust the parameters of the filter according to the estimated result of the signal-to-noise ratio. Increase the filtering intensity when the signal-to-noise ratio is low, and decrease the filtering intensity when the signal-to-noise ratio is high, so as to effectively suppress noise while retaining the speech signal to the greatest extent. The finally obtained noise suppression parameters include the gain factor and filter coefficients for each frequency band, and these parameters will be used for subsequent echo cancellation processing.
[0022] Step S5: Input the voice enhancement signal and the noise suppression parameters into the echo cancellation model. The model will establish the transfer function of the acoustic environment and predict the possible echo components. During the processing, evaluate the clarity and echo level of the signal in real time and dynamically adjust the processing parameters. When obvious echo is detected, increase the intensity of echo cancellation; when the signal is relatively clear, appropriately reduce the processing intensity to avoid speech distortion caused by overprocessing. The finally output human voice signal has a low noise level and echo interference while maintaining good voice quality. This optimization process fully considers the balance among voice enhancement, noise suppression, and echo cancellation to ensure that the finally output voice signal has the best comprehensive performance.
[0023] An intelligent microphone sound pickup and speech enhancement method provided by the present invention realizes the precise capture of the target human voice signal through cross-correlation analysis and sound source localization of the microphone audio signal, effectively improving the speech acquisition quality in complex acoustic environments. Through frequency band analysis and dynamic gain processing of the target human voice signal, refined enhancement of speech signals in different frequency bands is realized, ensuring the clarity and naturalness of the speech signal. Based on the adaptive noise filtering processing of the target human voice signal, the interference of background noise is effectively reduced, and the anti-noise ability of the system in complex environments is improved. By jointly optimizing the speech enhancement signal and the noise suppression parameter, dynamic adjustment of the echo cancellation model is realized, effectively solving the coupling problem of multiple acoustic interferences, and significantly improving the overall performance of the system. Through the organic combination of multi-dimensional signal processing and optimization strategies, the limitations of traditional single processing methods are overcome, and the system can better meet the requirements of various actual application scenarios.
[0024] In one embodiment, the audio signal collected by the dual-silicon microphone array is obtained, and cross-correlation analysis is performed on the audio signal to obtain the corresponding human voice signal correlation characteristics, including: The dual-silicon microphone array is installed with a fixed spacing of 4 cm, the sampling frequency is set to 16 kHz, and the sampling accuracy is 16 bit. The audio signal collected by the microphone array enters the digital signal processing unit after analog-to-digital conversion.
[0025] The cross-correlation calculation of the dual-silicon microphone array adopts the frequency domain analysis method based on the short-time Fourier transform. Assume that the sampling signals of microphone A and microphone B are x(n) and y(n) respectively. Frame processing with a frame length of 512 points and a frame shift of 256 points is performed on these two signals, and a Hamming window function is used for windowing. A 512-point FFT transform is performed on each frame of the signal to obtain the corresponding frequency domain signals X(k) and Y(k).
[0026] By calculating the cross-correlation function of X(k) and Y(k) , where represents the expectation operation, is the conjugate complex number of Y(k), is the time delay. The cross-correlation calculation result forms a human voice correlation data matrix.
[0027] The sub-band decomposition of the audio signal adopts the wavelet packet decomposition method. The original audio signal is decomposed by 3-layer wavelet packet decomposition using the db4 wavelet basis function, and the signal is decomposed into 8 frequency bands. The bandwidth of each frequency band is 1 kHz, covering the frequency range of 0 - 8 kHz. The sub-band signals obtained after sub-band decomposition retain the time-frequency characteristics of the original signal in different frequency bands.
[0028] Phase difference feature extraction is based on vocal correlation data and frequency band sub-signals. The phase difference between two microphone signals is calculated for each frequency band sub-signal using the generalized cross-spectrum method. The phase difference calculation formula is , where represents the phase angle calculation. Statistical analysis is performed on the calculated phase differences to generate a 128-dimensional phase difference feature spectrum.
[0029] The frequency band correlation matrix is obtained by performing a non-linear mapping on the phase difference feature spectrum. The hyperbolic tangent function is used to normalize the phase difference features, mapping the value range to the [-1, 1] interval. The mapped features form an 8×16 frequency band correlation matrix, and each element in the matrix represents the correlation degree of the corresponding frequency band.
[0030] The calculation of frequency band energy features uses an entropy-based energy concentration analysis method. The entropy is calculated for each row of the frequency band correlation matrix to obtain an 8-dimensional frequency band energy feature vector. The smaller the entropy value, the higher the energy concentration, and the corresponding frequency band is more likely to contain vocal components.
[0031] The joint feature vector is formed by feature concatenation. The 8-dimensional frequency band energy feature vector is concatenated with the 128-dimensional phase difference feature spectrum to obtain a 136-dimensional joint feature vector. This feature vector contains the complete information of the frequency band energy distribution and the time-varying characteristics of the phase difference.
[0032] Principal component analysis is used for signal feature fusion. Principal component analysis is performed on the 136-dimensional joint feature vector, and the principal components with a cumulative contribution rate reaching 95% are selected as the final vocal signal correlation features. The dimension of the feature after dimensionality reduction is usually between 20 and 30, effectively reducing the computational complexity of subsequent processing.
[0033] In this embodiment, by adopting a dual-silicon microphone array configuration with a fixed spacing of 4 cm, combined with a sampling frequency of 16 kHz and a sampling precision of 16 bits, the spatial characteristics of vocal signals can be accurately captured, effectively improving the accuracy of sound source localization. The cross-correlation calculation is performed using a frequency domain analysis method based on the short-time Fourier transform, combined with a frame length of 512 points and a frame shift of 256 points, which not only ensures real-time performance but also provides sufficient frequency resolution. The signal is divided into 8 frequency bands through 3-layer wavelet packet decomposition, each with a bandwidth of 1 kHz, achieving a fine coverage of the entire 0-8 kHz frequency band, making the extraction of vocal features more complete. The entropy-based energy concentration analysis method can accurately identify the frequency band where the main components of the voice are located. Through feature concatenation and principal component analysis, the 136-dimensional features are compressed to 20-30 dimensions, significantly reducing the computational complexity while retaining 95% of the information volume and improving the real-time processing ability of the system. This solution effectively improves the accuracy of vocal recognition and the processing efficiency of the system through multi-level signal analysis and feature extraction.
[0034] In one embodiment, source localization analysis is performed on an audio signal based on the correlation characteristics of a vocal signal to obtain a corresponding target vocal signal, including: Through localization analysis of the audio signal, accurate localization and enhancement of the vocal signal are achieved. During the localization analysis process, when performing delay analysis on the correlation characteristics of the vocal signal, the audio signal collected by the microphone array passes through the source localization module, and the cross-correlation function between signals of different channels is calculated. This cross-correlation function reflects the time delay relationship between the signals. Based on the peak position of the cross-correlation function, the signal time delay difference is determined. The time delay difference characterizes the propagation time difference from the sound source to different microphones. The calculation of the time delay difference uses the generalized cross-correlation method, which introduces a phase transformation weighting function to improve the accuracy of time delay estimation.
[0035] After obtaining the signal time delay difference, the source localization module performs response analysis on the audio signal. The response analysis is based on the beamforming algorithm, and the delay-and-sum beamforming method is used to construct a directional response function. This response function describes the energy distribution of the sound source in different directions, and the sound source direction response value is determined by searching for the extreme point of the response function. The calculation of the direction response value uses the minimum variance distortionless response criterion, which minimizes the output noise variance on the premise of keeping the desired signal undistorted.
[0036] The source localization module performs phase difference operations on the sound source direction response value to calculate the phase difference characteristics in each direction. The phase difference characteristics have a corresponding relationship with the sound source direction, and the direction weight coefficient is calculated based on this characteristic. The calculation of the direction weight coefficient uses a spatial filtering method based on the phase difference, which combines the sound source direction and spatial phase information to improve the direction performance. The value range of the weight coefficient is from 0 to 1, and the larger the weight value, the more likely the direction is the target sound source direction.
[0037] The signal recognition module processes the audio signal based on the direction weight coefficient to extract the initial target signal. The signal recognition uses a method based on time-frequency masking to weight the time-frequency representation of the audio signal and suppress the interference signals in non-target directions. The extraction of the initial target signal uses the Wiener filtering criterion, which estimates the target signal in the sense of minimum mean square error.
[0038] The compensation processing module performs time delay compensation and loudness compensation on the audio signal according to the direction weight coefficient. The time delay compensation aligns the signals of each channel based on the sound source direction to eliminate the phase distortion caused by the propagation time delay. The loudness compensation normalizes the signal amplitude based on the sound source distance to compensate for the energy loss caused by the distance attenuation. The generation of the compensation signal uses an adaptive compensation algorithm, which can automatically adjust the compensation parameters according to the actual environment.
[0039] The signal synthesis module performs a convolution operation on the initial target signal and the compensation signal to obtain the target human voice signal. The convolution operation uses the fast convolution algorithm, where signal multiplication is performed in the frequency domain and then transformed back to the time domain. This operation process reconstructs the time-frequency characteristics of the signal, improving the quality of the target human voice signal. The output of the target human voice signal uses adaptive gain control to ensure that the amplitude of the output signal is within an appropriate range.
[0040] In this embodiment, through the sound source localization analysis based on the correlation characteristics of the human voice signal, accurate recognition and enhancement of the target human voice signal are achieved. The generalized cross-correlation method is used to calculate the signal delay difference, and a phase transformation weighting function is introduced to improve the delay estimation accuracy. In the response analysis, the minimum variance distortionless response criterion is adopted to minimize the output noise variance while keeping the desired signal undistorted, improving the sound source localization effect. The spatial filtering method based on the phase difference calculates the direction weight coefficient, and combines the sound source direction and spatial phase information to enhance the direction performance. The signal recognition method based on time-frequency masking and the Wiener filtering criterion are used to effectively suppress the interference signals in non-target directions. Through the adaptive compensation algorithm for delay and loudness compensation, it automatically adapts to the actual environmental changes, improving the system robustness. The application of the fast convolution algorithm and adaptive gain control optimizes the reconstruction effect of the time-frequency characteristics of the target human voice signal, ensuring the quality of the output signal. The overall scheme significantly improves the recognition accuracy and enhancement effect of the human voice signal, and has good practical value.
[0041] In one embodiment, the frequency characteristics of the target human voice signal are analyzed, and the signal is subjected to segmented dynamic gain processing according to a preset frequency band range to obtain the corresponding speech enhancement signal, including: Perform a frequency band division operation on the input target human voice signal. Based on the preset frequency band range parameters, the target human voice signal is decomposed into speech frequency band signals of multiple different frequency bands. This frequency band division process is implemented using a digital filter bank to ensure the accurate separation of signals in each frequency band.
[0042] After obtaining the speech frequency band signals, the implementation method performs amplitude calculation and analysis on the signals in each frequency band. By calculating the effective value and peak value of the signal, the amplitude change range of the signal in the time domain is determined. This calculation process uses a sliding time window method to track the amplitude characteristics of the signal in real time.
[0043] According to the calculated signal amplitude range, the speech frequency band signals are divided into three interval segments. The range of the signal components in the first signal amplitude interval segment is [-70dB, -50dB].
[0044] The range of the signal components in the second signal amplitude interval segment is [-50dB, -30dB], and the range of the signal components in the third signal amplitude interval segment is [-30dB, -15dB]. This segmentation process is performed based on the signal amplitude threshold decision method.
[0045] For the signal components in each interval, different degrees of gain processing are applied respectively. A 50 dB gain is applied to the first signal amplitude interval to enhance the weaker signal components; a 30 dB gain is applied to the second signal amplitude interval to enhance the medium-strength signals; a 15 dB gain is applied to the third signal amplitude interval to moderately amplify the stronger signal components. The gain processing is implemented using an adjustable gain amplifier circuit.
[0046] After completing the segmented gain processing, the three interval signals after gain are synthesized. Using the digital signal superposition method, the first gain signal, the second gain signal, and the third gain signal are integrated into a complete voice enhancement signal for output. This synthesis process ensures smooth signal transition and avoids distortion.
[0047] In this embodiment, by implementing frequency band division and segmented dynamic gain processing on the target human voice signal, the quality and recognizability of the voice signal are significantly improved. The signal amplitude range is divided into three intervals, and a differential gain strategy is implemented for voice signals of different intensities, effectively solving the problem of limited signal dynamic range in traditional voice enhancement methods. A 50 dB high gain processing is adopted for weak signals, a 30 dB gain is applied to medium-strength signals, and a 15 dB gain is used for stronger signals. This hierarchical gain control scheme enables voice signals of various intensities to be reasonably enhanced, avoiding signal distortion or over-amplification. The use of a digital filter bank for frequency band division and an adjustable gain amplifier to achieve dynamic gain ensures the accuracy and reliability of signal processing. In the signal synthesis link, the digital superposition method is adopted to ensure the coherence and naturalness of the enhanced voice signal. Overall, the voice enhancement method provided by the present invention significantly improves the sound pickup effect of the intelligent microphone, and has practical value and promotional significance.
[0048] In one embodiment, an adaptive filtering process for noise components is performed on the target human voice signal to obtain corresponding noise suppression parameters, including: The target human voice signal undergoes spectral analysis processing to obtain a frequency domain feature sequence. The spectral analysis uses the short-time Fourier transform (STFT) method, dividing the target human voice signal into several frames, with each frame signal length of 25 ms and an adjacent frame overlap rate of 50%. After applying a Hamming window to each frame of signal, a fast Fourier transform (FFT) process is performed, with the number of transform points being 512 points, thereby obtaining the frequency domain feature sequence of the target human voice signal.
[0049] Based on the obtained frequency domain feature sequence, the Mel filter bank is used to perform sub-band decomposition on the target human voice signal. The Mel filter bank contains 24 triangular filters, covering a frequency range of 0 - 8 kHz. By performing a convolution operation on the frequency domain feature sequence and the frequency responses of each filter, 24 frequency sub-band signals are obtained.
[0050] Perform short-time feature distribution analysis on each frequency sub-band signal. Calculate statistical features such as energy distribution, zero-crossing rate, and spectral flatness of each sub-band signal within a 32-ms time window to form a sub-band distribution feature vector. The sub-band distribution features reflect the time-frequency characteristics of the signal within that frequency sub-band.
[0051] Based on the sub-band distribution features, perform sliding window calculations on the frequency sub-band signals. Use a 64-ms rectangular window with a window sliding step of 16 ms. Within each sliding window, statistically calculate parameters such as the mean, variance, and peak factor of the sub-band signal as the noise statistical parameters for that sub-band.
[0052] Adaptively update the noise statistical parameters according to the speech enhanced signal. Set the initial noise suppression threshold to -40 dB and dynamically adjust the threshold as the signal characteristics change. When speech activity is detected, the threshold remains unchanged; when background noise is detected, the threshold gradually increases at a rate of 0.1 dB / frame until it reaches the maximum value of -20 dB.
[0053] Perform Wiener filtering on the frequency sub-band signals according to the updated noise suppression threshold. The gain function of the Wiener filter uses a soft thresholding method. For frequency components with a signal-to-noise ratio lower than the threshold, its gain decreases as the signal-to-noise ratio decreases. After filtering, obtain the noise suppression parameters for each sub-band for subsequent speech enhancement processing.
[0054] In one embodiment, by performing spectral analysis on the target human voice signal and using a Mel filter bank for sub-band decomposition, a fine division of different frequency components is achieved, improving the capture accuracy of the frequency-domain characteristics of the speech signal. Based on the dual feature extraction mechanism of short-time feature distribution analysis and sliding window calculation, the dynamic change characteristics of the signal within each frequency sub-band can be accurately characterized, providing a reliable basis for the calculation of noise suppression parameters. By adopting an adaptive threshold update strategy, the system can dynamically adjust the noise suppression threshold according to the speech activity state, effectively suppressing background noise while maintaining the integrity of the speech signal. Through the soft threshold gain control of the Wiener filter, the detailed features of the speech signal are maximally retained while reducing noise, avoiding the speech distortion problem caused by traditional hard threshold methods. This method significantly improves the speech enhancement effect of intelligent microphones in complex noise environments and has strong environmental adaptability and processing stability.
[0055] In one embodiment, input the speech enhanced signal and the noise suppression parameters into a preset echo cancellation model for joint optimization to obtain the corresponding output human voice signal, including: The echo cancellation model receives the speech enhanced signal and the noise suppression parameters as inputs and outputs a human voice signal after multiple layers of processing. The echo cancellation model includes a signal preprocessing layer, a feature analysis layer, a double-talk detection layer, an echo path layer, a joint optimization layer, and a signal reconstruction layer.
[0056] The signal preprocessing layer performs frame segmentation on the input voice enhancement signal, segmenting the signal with a frame length of 20 ms and a frame shift of 10 ms. A Hamming window function is applied to the segmented voice signal for windowing, and then the fast Fourier transform is used to convert the time-domain signal to the frequency domain to obtain frequency-domain feature information. The frequency-domain feature information includes amplitude spectrum features and phase spectrum features.
[0057] The feature analysis layer extracts acoustic feature parameters based on the voice enhancement signal. This layer calculates time-domain features such as the short-time energy, zero-crossing rate, and fundamental period of the voice signal, and simultaneously extracts frequency-domain features such as Mel cepstral coefficients and linear prediction coefficients. These features together constitute a set of acoustic feature parameters for subsequent voice state judgment.
[0058] The double-talk detection layer analyzes the acoustic feature parameters in combination with noise suppression parameters. This layer judges the voice state of the current frame signal based on preset voice activity detection rules. The detection rules include multiple rules such as energy threshold judgment, spectral entropy judgment, and zero-crossing rate judgment. When the voice judgment conditions are met, the current frame is marked as the voice state; otherwise, it is marked as the non-voice state.
[0059] The echo path layer performs non-linear echo path modeling on the frequency-domain feature information. This layer uses an adaptive filter to model the echo propagation path and extracts echo feature information. The echo feature information includes feature parameters such as echo attenuation coefficient and reverberation time. During the modeling process, the characteristics of the room impulse response are combined to accurately estimate the echo signal.
[0060] The joint optimization layer jointly processes the acoustic feature parameters, voice state information, echo feature information, and noise suppression parameters. This layer optimally fuses the various features based on the Wiener filtering criterion to suppress echo and background noise components and enhance the target voice components. The optimization process is carried out iteratively until the preset convergence conditions are met.
[0061] The signal reconstruction layer performs an inverse transformation on the enhanced voice signal. This layer uses the inverse fast Fourier transform to convert the frequency-domain signal to the time domain and performs frame overlap-and-add processing. A triangular window is used for weighting during the overlap-and-add process to ensure the smoothness of signal reconstruction. Finally, the output human voice signal after joint optimization is output.
[0062] In this embodiment, by setting a multi-level joint optimization processing structure in the echo cancellation model, the deep fusion of the speech enhancement signal and the noise suppression parameter is achieved, and the quality of the output human voice signal is improved. The signal preprocessing layer adopts a reasonable frame division and transformation strategy to ensure the accuracy of the frequency-domain feature extraction. The feature analysis layer extracts multi-dimensional acoustic feature parameters, providing a reliable basis for speech state judgment. The double-talk detection layer judges the speech activity based on multiple detection rules, improving the accuracy of speech state recognition. The echo path layer models the echo propagation characteristics by using the adaptive filtering method, enhancing the echo suppression effect. The joint optimization layer fuses and optimizes the multi-dimensional features based on the Wiener filtering criterion, effectively suppressing the echo and background noise interference. The signal reconstruction layer reconstructs the time-domain signal by using the weighted overlap and add method, ensuring the continuity and smoothness of the output signal. The overall scheme improves the system robustness through the multi-level joint optimization structure, ensuring the speech enhancement effect in a complex acoustic environment.
[0063] In one embodiment, the acoustic feature parameters, speech state information, echo feature information, and noise suppression parameter are jointly enhanced by the joint optimization layer of the echo cancellation model to obtain the corresponding enhanced speech signal, including: When the joint optimization layer of the echo cancellation model processes the acoustic features, the feature fusion module is used to perform a non-linear combination of the extracted multi-dimensional acoustic features (including time-domain features, frequency-domain features, cepstrum features, etc.) to generate a high-dimensional acoustic feature vector. This feature vector contains the complete characterization information of the speech signal in different feature domains. Based on the speech state recognition result, the joint optimization layer sets the weight coefficients of different feature components to construct a weighted feature matrix. The weight coefficients are dynamically adjusted according to the speech state (such as voiced, unvoiced, silent, etc.), highlighting the important feature components.
[0064] In the echo feature processing section, the joint optimization layer uses a deep neural network to perform a non-linear transformation on the echo features extracted from the far-end reference signal to estimate the feature parameters of the echo path. This parameter reflects the attenuation, delay, and other characteristics of the echo signal during transmission. The joint optimization layer uses the echo estimation parameter and the noise suppression parameter as the calculation basis for the filter coefficients to perform adaptive filtering on the weighted feature matrix. The filtering process is based on the least mean square error criterion to achieve the optimization and enhancement of the feature matrix.
[0065] During the noise suppression parameter correction process, the joint optimization layer corrects the initial noise suppression parameter according to the filtered coefficients. The correction process comprehensively considers the time-varying characteristics of the environmental noise to ensure the noise suppression effect. The corrected suppression parameter and the echo estimation parameter jointly act on the residual echo cancellation module to compensate for the residual echo component caused by non-linear distortion. The compensated signal passes through the spectral gain calculation module, which calculates the optimal spectral gain coefficient based on the Wiener filtering theory.
[0066] In the final stage of speech enhancement, the joint optimization layer uses spectral subtraction to process the corrected echo signal. Spectral subtraction modulates the signal spectrum based on the spectral gain coefficient to suppress the noise and echo components while retaining the effective speech components. After spectral subtraction processing, an enhanced speech signal is output, which has a higher signal-to-noise ratio and speech clarity.
[0067] In this embodiment, by adopting the joint optimization layer of the echo cancellation model in the intelligent microphone, the joint enhancement processing of acoustic feature parameters, speech state information, echo feature information, and noise suppression parameters is realized. The joint optimization layer adopts the multi-dimensional feature fusion technology to effectively integrate multi-dimensional acoustic feature information such as time-domain features, frequency-domain features, and cepstrum features, and improves the feature representation ability of the speech signal. Based on the dynamic weight adjustment mechanism of the speech state, the system can adaptively optimize the feature parameters for different speech states, enhancing the adaptability of the system to complex speech scenarios. Through the non-linear mapping processing of the echo features by the deep neural network, the accuracy of the echo path estimation is improved. Combining with the adaptive filtering technology based on the least mean square error criterion, the echo cancellation effect is significantly improved. By adopting the joint optimization method of spectral subtraction and spectral gain, while ensuring the integrity of the speech signal, the noise and residual echo components are effectively suppressed, and the output enhanced speech signal has a higher signal-to-noise ratio and speech clarity.
[0068] Refer to Figure 2 As shown in the figure, the present invention also provides an intelligent microphone sound pickup and speech enhancement system, which is applied to the intelligent microphone sound pickup and speech enhancement method of any one of the above, and includes: A collection module, which is used to obtain the audio signal collected by the dual-silicon microphone array, perform cross-correlation analysis on the audio signal, and obtain the corresponding human voice signal correlation characteristics; An analysis module, which is used to perform sound source localization analysis on the audio signal based on the human voice signal correlation characteristics to obtain the corresponding target human voice signal; An association module, which is used to perform frequency characteristic analysis on the target human voice signal and perform segmented dynamic gain processing on the signal according to the preset frequency band range to obtain the corresponding speech enhancement signal; A processing module, which is used to perform adaptive filtering processing on the noise components of the target human voice signal to obtain the corresponding noise suppression parameters; A control module, which is used to input the speech enhancement signal and the noise suppression parameters into the preset echo cancellation model for joint optimization to obtain the corresponding output human voice signal.
[0069] An intelligent microphone sound pickup and speech enhancement system provided by the present invention realizes the precise capture of target human voice signals through cross-correlation analysis and sound source localization of microphone audio signals, effectively improving the quality of speech collection in complex acoustic environments. Through frequency band analysis and dynamic gain processing of the target human voice signals, the refined enhancement of speech signals in different frequency bands is realized, ensuring the clarity and naturalness of the speech signals. Based on the adaptive noise filtering processing of the target human voice signals, the interference of background noise is effectively reduced, and the anti-noise ability of the system in complex environments is improved. By jointly optimizing the speech enhancement signal and the noise suppression parameter, the dynamic adjustment of the echo cancellation model is realized, effectively solving the coupling problem of multiple acoustic interferences and significantly improving the overall performance of the system. Through the organic combination of multi-dimensional signal processing and optimization strategies, the limitations of traditional single processing methods are overcome, and the system can better meet the requirements of various actual application scenarios.
[0070] Referring to Figure 3 As shown, the present invention also provides an intelligent microphone sound pickup and speech enhancement device, including: A memory for storing programs; A processor for executing the programs to implement each step of an intelligent microphone sound pickup and speech enhancement method as described in any one of the above.
[0071] In this embodiment, the processor and the memory can be connected through a bus or other means. The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid-state drive. The processor may be a general-purpose processor, such as a central processing unit, a digital signal processor, an application-specific integrated circuit, or one or more integrated circuits configured to implement the embodiments of the present invention.
[0072] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system and each module can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0073] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specifications and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An intelligent microphone sound pickup and voice enhancement method, characterized in that, Including: Obtain the audio signal collected by the dual silicon microphone array, perform cross-correlation analysis on the audio signal to obtain the corresponding correlation characteristics of the human voice signal; Based on the correlation characteristics of the human voice signal, perform sound source localization analysis on the audio signal to obtain the corresponding target human voice signal; Perform frequency characteristic analysis on the target human voice signal, and perform segmented dynamic gain processing on the signal according to a preset frequency band range to obtain the corresponding speech enhancement signal; Perform adaptive filtering processing on the noise components of the target human voice signal to obtain the corresponding noise suppression parameters; Input the speech enhancement signal and the noise suppression parameters into a preset echo cancellation model for joint optimization to obtain the corresponding output human voice signal.
2. The intelligent microphone sound pickup and voice enhancement method according to claim 1, characterized in that, The obtaining the audio signal collected by the dual silicon microphone array, performing cross-correlation analysis on the audio signal to obtain the corresponding correlation characteristics of the human voice signal includes: Perform cross-correlation calculation between microphones on the dual silicon microphone array to obtain the corresponding human voice correlation data; Perform sub-band decomposition on the audio signal to obtain multiple sub-band signals; Extract phase difference features from the sub-band signals according to the human voice correlation data to obtain the corresponding phase difference feature spectrum; Perform non-linear transformation processing on the phase difference feature spectrum to obtain a frequency band correlation matrix; Calculate the energy aggregation degree of the frequency band correlation matrix to obtain the frequency band energy characteristics; Perform joint analysis on the frequency band energy characteristics and the phase difference feature spectrum to obtain a joint feature vector; Perform signal feature fusion processing on the joint feature vector, the correlation characteristics of the human voice signal.
3. The intelligent microphone sound pickup and voice enhancement method according to claim 1, characterized in that, The performing sound source localization analysis on the audio signal based on the correlation characteristics of the human voice signal to obtain the corresponding target human voice signal includes: Perform delay analysis on the correlation characteristics of the human voice signal to obtain the corresponding signal delay difference; Perform response analysis on the audio signal according to the signal delay difference to obtain the corresponding sound source direction response value; Perform phase difference operation on the sound source direction response value to obtain a direction weight coefficient; Based on the direction weight coefficient, perform signal recognition on the audio signal to obtain the corresponding initial target signal; Perform delay compensation and loudness compensation on the audio signal according to the direction weight coefficient to obtain a compensation signal; Perform convolution operation on the initial target signal and the compensation signal to obtain the corresponding target human voice signal.
4. The intelligent microphone sound pickup and voice enhancement method according to claim 1, characterized in that, The performing frequency characteristic analysis on the target human voice signal, and performing segmented dynamic gain processing on the signal according to a preset frequency band range to obtain the corresponding speech enhancement signal includes: Perform frequency band division processing on the target human voice signal according to a preset frequency band range to obtain the corresponding speech frequency band signal; Calculate the signal amplitude of the speech frequency band signal to obtain the corresponding signal amplitude range; Perform segmented processing on the speech frequency band signal according to the signal amplitude range to obtain a first signal amplitude interval segment, a second signal amplitude interval segment, and a third signal amplitude interval segment; Perform a 50 dB gain process on the first signal amplitude interval segment to obtain the corresponding first gain signal; Perform a 30 dB gain processing on the second signal amplitude interval segment to obtain a corresponding second gain signal; Perform a 15 dB gain processing on the third signal amplitude interval segment to obtain a corresponding third gain signal; Perform a synthesis processing on the first gain signal, the second gain signal, and the third gain signal to obtain the voice enhancement signal; Wherein, the amplitude range of the first signal amplitude interval segment is [-70 dB, -50 dB], the amplitude range of the second signal amplitude interval segment is [-50 dB, -30 dB], and the amplitude range of the third signal amplitude interval segment is [-30 dB, -15 dB].
5. The intelligent microphone sound pickup and voice enhancement method according to claim 1, characterized in that, Performing an adaptive filtering process on the noise component of the target human voice signal to obtain a corresponding noise suppression parameter, including: Performing a spectrum analysis on the target human voice signal to obtain a corresponding frequency domain feature sequence; Performing a sub-band decomposition process on the target human voice signal according to the frequency domain feature sequence to obtain a plurality of frequency sub-band signals; Performing a short-time feature distribution analysis on the frequency sub-band signals to obtain corresponding sub-band distribution features; Performing a sliding window calculation on the frequency sub-band signals according to the sub-band distribution features to obtain corresponding noise statistical parameters; Performing an adaptive threshold update on the noise statistical parameters to obtain a corresponding noise suppression threshold; Performing an adaptive filtering process on the frequency sub-band signals according to the noise suppression threshold to obtain the noise suppression parameter.
6. The intelligent microphone sound pickup and voice enhancement method according to claim 1, wherein, Inputting the voice enhancement signal and the noise suppression parameter into a preset echo cancellation model for joint optimization to obtain a corresponding output human voice signal, including: Inputting the voice enhancement signal into the echo cancellation model, and performing a frame splitting and time-frequency transformation process on the voice enhancement signal through the signal preprocessing layer of the echo cancellation model to obtain corresponding frequency domain feature information; Performing an acoustic feature analysis on the voice enhancement signal through the feature analysis layer of the echo cancellation model to obtain corresponding acoustic feature parameters; Inputting the noise suppression parameter into the echo cancellation model, and performing a voice activity detection on the acoustic feature parameters through the double-talk detection layer of the echo cancellation model to obtain corresponding voice state information; Performing a non-linear echo path analysis and feature extraction on the frequency domain feature information through the echo path layer of the echo cancellation model to obtain corresponding echo feature information; Jointly enhancing the acoustic feature parameters, the voice state information, the echo feature information, and the noise suppression parameter through the joint optimization layer of the echo cancellation model to obtain a corresponding enhanced voice signal; Performing an inverse time-frequency transformation and a frame overlap and add process on the enhanced voice signal through the signal reconstruction layer of the echo cancellation model to obtain and output the output human voice signal.
7. The intelligent microphone sound pickup and voice enhancement method according to claim 6, wherein, Jointly enhancing the acoustic feature parameters, the voice state information, the echo feature information, and the noise suppression parameter through the joint optimization layer of the echo cancellation model to obtain a corresponding enhanced voice signal, including: Performing a multi-dimensional feature fusion on the acoustic feature parameters to obtain corresponding acoustic feature vectors; Adjust the weights of the acoustic feature vectors according to the voice state information to obtain a corresponding weighted feature matrix; Perform a non-linear mapping on the echo feature information to obtain corresponding echo estimation parameters; Perform a joint filtering process on the weighted feature matrix according to the echo estimation parameters and the noise suppression parameters to obtain corresponding filtering coefficients; Modify the noise suppression parameters according to the filtering coefficients to obtain corresponding modified suppression parameters; Perform residual echo compensation according to the modified suppression parameters and the echo estimation parameters to obtain corresponding modified echo signals; Calculate the spectral gain of the modified echo signal to obtain corresponding spectral gain coefficients; Perform spectral subtraction on the modified echo signal according to the spectral gain coefficients to obtain the enhanced speech signal.
8. An intelligent microphone sound pickup and voice enhancement system, wherein, Applied to the intelligent microphone sound pickup and speech enhancement method according to any one of claims 1-7 above, including: An acquisition module, which is used to acquire the audio signal collected by the dual-silicon microphone array, perform cross-correlation analysis on the audio signal, and obtain the corresponding voice signal correlation characteristics; An analysis module, which is used to perform sound source localization analysis on the audio signal based on the voice signal correlation characteristics to obtain the corresponding target voice signal; An association module, which is used to perform frequency characteristic analysis on the target voice signal, and perform segmented dynamic gain processing on the signal according to a preset frequency band range to obtain the corresponding speech enhancement signal; A processing module, which is used to perform adaptive filtering processing on the noise components of the target voice signal to obtain corresponding noise suppression parameters; A control module, which is used to input the speech enhancement signal and the noise suppression parameters into a preset echo cancellation model for joint optimization to obtain the corresponding output voice signal.
9. An intelligent microphone sound pickup and voice enhancement device, wherein, Including: A memory for storing programs; A processor for executing the program to implement each step of an intelligent microphone sound pickup and speech enhancement method according to any one of claims 1-7.
Citation Information
Cited By
Audio noise suppression method based on deep learning and intelligent sound equipment
CN120412619A
Audio processing system of photographic camera
CN120600041A
Wireless honeycomb cluster microphone cooperative control method and system
CN120692490A
Wind noise suppression method and device, equipment and storage medium
CN120748428A
Audio signal processing method for multi-person conference scene of distributed multi-microphone sound amplification system
CN120980430A