Multi-mode sensing intelligent microphone array signal processing method and system

By employing a multimodal sensing-based intelligent microphone array signal processing method, combined with audiovisual information fusion and environmental adaptation mechanisms, the problem of speech separation and enhancement in complex acoustic environments was solved, achieving efficient speech separation and enhancement effects and improving the system's adaptability and computational efficiency in dynamic environments.

CN120808810AActive Publication Date: 2025-10-17GUANGZHOU OPSMEN TECH CO LTD

Patent Information

Application Number
CN202511337012.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-10-17
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing microphone array signal processing technologies struggle to effectively fuse visual and acoustic information in complex acoustic environments, leading to decreased speech recognition and interaction performance. Furthermore, they lack environmental adaptability, have high computational complexity, and are difficult to deploy efficiently on resource-constrained devices.

Method used

A multimodal sensing intelligent microphone array signal processing method is adopted. Through deep fusion and collaborative processing of audiovisual multimodal information, combined with topology-enhanced adversarial generative network architecture and environment adaptation mechanism, an audiovisual topology feature space is constructed. Speech separation is performed using multidimensional discriminative adversarial generative network, and the processing parameters are dynamically adjusted in real time based on the environmental state.

Benefits of technology

High-quality speech separation and enhancement are achieved in complex acoustic environments, which improves the accuracy and naturalness of speech separation, reduces computational complexity, and improves the system's adaptability and computational efficiency in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808810A_ABST
    Figure CN120808810A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode sensing intelligent microphone array signal processing method and system, and belongs to the technical field of signal processing, and the method comprises the steps: obtaining a visual signal through employing a multi-directional visual sensor, and obtaining a sound signal through employing a microphone array; extracting visual features and acoustic features; constructing an audio-visual topological feature space, mapping visual features and acoustic features to the space, and establishing a sound source probability distribution model; processing the sound signal by adopting a multi-dimensional discriminant adversarial generative network, and separating out a target voice signal; the acoustic environment state is evaluated in real time, and processing parameters are dynamically adjusted; quality evaluation is carried out on the separated multiple paths of target voice signals, the voice signal with the highest quality is selected as output, audio-visual multi-mode information is deeply fused and cooperatively processed, and a topology enhanced adversarial generative network architecture and an environment self-adaptive mechanism are combined, so that the voice separation effect in a complex environment is remarkably improved; and more than 85% of speech intelligibility can still be kept in a scene that six persons speak simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of signal processing, in particular to a multi-modal perception intelligent microphone array signal processing method and system, which is applied to acoustic signal processing scenes such as speech separation, enhancement and noise reduction. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, intelligent voice interaction has become an important way of human-computer interaction. However, in the actual application environment, due to the existence of multiple speaker interference, environmental noise and reverberation, etc., the speech signal is often seriously polluted, which leads to a significant decline in speech recognition and interaction performance, which is known as the cocktail party effect in the field of signal processing.

[0003] Traditional microphone array signal processing technology mainly relies on acoustic information, such as spatial filtering technology based on beamforming, blind source separation technology based on independent component analysis, etc. These technologies perform well under ideal conditions, but their performance is limited in complex acoustic environments. On the one hand, traditional beamforming technology needs to accurately estimate the direction of the sound source, which is difficult to accurately locate in a multi-source, strong reverberation environment; on the other hand, blind source separation technology has high requirements for the independence and sparsity of signals, and it is difficult to deal with strongly correlated interference and non-sparse noise.

[0004] In recent years, deep learning technology has made significant progress in the field of speech processing, such as speech separation and enhancement methods based on deep neural networks. However, these methods are mainly based on a single acoustic modality and lack effective use of visual and other modal information, while the human perception system is precisely through the fusion of multi-modal information to achieve efficient sound source separation. In addition, existing deep learning methods are mostly static models, lack of adaptability to dynamically changing environments, and have high computational complexity, making it difficult to efficiently deploy on resource-constrained devices.

[0005] Therefore, it is urgent to develop an intelligent microphone array signal processing method and system that can effectively fuse audio-visual multi-modal information and have environmental adaptability to solve the key technical problems of speech separation and enhancement in complex environments. SUMMARY

[0006] The purpose of the present application is to provide a multi-modal perception intelligent microphone array signal processing method and system, which realizes high-quality speech separation and enhancement in complex acoustic environments through deep fusion and collaborative processing of audio-visual multi-modal information, combined with a topologically enhanced generative adversarial network architecture and an environmental adaptation mechanism.

[0007] The present application proposes a multi-modal perception intelligent microphone array signal processing method, which comprises: using a multi-directional visual sensor to obtain visual signals and using a microphone array to obtain sound signals; extracting visual features of the visual signals and acoustic features of the sound signals, wherein the visual features include visual orientation features, face features and lip shape features, and the acoustic features include time-frequency features and voiceprint features; constructing an audio-visual topological feature space, mapping the visual features and the acoustic features to the audio-visual topological feature space, and establishing a sound source probability distribution model based on the audio-visual topological feature space; processing the sound signals by using a multi-dimensional discriminative generative adversarial network to separate target speech signals, wherein the multi-dimensional discriminative generative adversarial network includes a multi-level discriminator structure and a generator based on audio-visual feature guidance; real-time evaluating acoustic environment states and dynamically adjusting processing parameters of the multi-dimensional discriminative generative adversarial network; evaluating qualities of the separated multi-path target speech signals and selecting a speech signal with the highest quality as an output.

[0008] Preferably, the acquiring the visual signals by using multi-directional visual sensors and the acquiring the sound signals by using a microphone array specifically include: acquiring 360° panoramic visual information by using visual sensors arranged in six directions of front, back, left, right, up and down; acquiring sound field information by using a microphone array configurable in linear, ring or spherical geometric arrangement; time-synchronizing the visual signals and the sound signals to ensure a synchronization accuracy better than 1 millisecond.

[0009] Preferably, the extracting the visual features of the visual signals and the acoustic features of the sound signals specifically include: performing orientation recognition, face detection and lip shape tracking on the visual signals to extract visual orientation features, face features and lip shape features; performing time-frequency analysis, voiceprint extraction and acoustic fingerprint generation on the sound signals to extract time-frequency features and voiceprint features; performing noise reduction, enhancement and normalization processing on the extracted visual features and acoustic features.

[0010] Preferably, the constructing the audio-visual topological feature space specifically includes: defining a 128- to 256-dimensional joint feature space in which a Riemann distance metric is defined; projecting the visual features and the acoustic features to the joint feature space by nonlinear dimensionality reduction; optimizing feature distance metrics by a contrastive learning method to enhance intra-class cohesion and inter-class separability; establishing an accurate correspondence between visual features and acoustic features to maximize mutual information between audio-visual features.

[0011] As preferred, the multi-dimensional discriminative generative adversarial network comprises: a phoneme-level discriminator for evaluating the authenticity of short-time speech segments; a prosody-level discriminator for evaluating the naturalness of speech rhythm and intonation; a speaker-level discriminator for verifying the consistency of separated speech and speaker characteristics; an environment-level discriminator for ensuring the coordination of speech and the current acoustic environment; The generator based on audio-visual feature guidance adopts a U-shaped network structure, including multi-level jump connection and multi-path processing flow, and guides the sound separation process through visual condition embedding and voiceprint condition control.

[0012] As preferred, the processing process of the multi-dimensional discriminative generative adversarial network comprises: extract discriminative features on multiple time scales, dynamically focus on the most discriminative feature area through attention mechanism; adaptively adjust the discriminative threshold, quantify the reliability of the discriminative result, and weight and fuse the multi-layer discriminative results according to the reliability; decompose the sound signal into multiple time-frequency components, and distribute each time-frequency component to different sound sources based on the audio-visual features; construct a soft masking matrix for each sound source, and eliminate residual cross interference through post-processing.

[0013] As preferred, the real-time evaluation of the acoustic environment state dynamically adjusts the processing parameters of the multi-dimensional discriminative generative adversarial network, specifically comprising: extracting acoustic environment parameters such as reverberation time, early reflection density, and background noise level; identifying sound source distribution characteristics such as sound source number, spatial distribution, and motion state; matching the current environment with predefined environment categories to generate an environment feature vector; analyzing environmental parameters on a continuous time window to identify significant change points in environmental parameters; adjusting filter parameters, gain control parameters, and network weight parameters according to the environment state to ensure the continuity of the parameter adjustment process.

[0014] As preferred, the quality evaluation of the separated multi-path target speech signal selects the highest quality speech signal as the output, specifically comprising: calculating objective indicators such as signal-to-noise ratio, speech articulation index, and cepstrum distance; evaluating perceptual indicators such as subjective mean opinion score and speech naturalness; comprehensive multi-index calculation of weighted quality score; compare the quality of the separation results of different microphone channels, and select the channel with the highest quality score. The selected voice signal is subjected to customized enhancement processing such as spectral smoothing, harmonic structure enhancement and sub-band gain adjustment.

[0015] As preferred, it further comprises: According to the application scene, a preset configuration is selected, including a conference scene configuration, a mobile scene configuration and a noise environment configuration; Dynamically allocate computing resources, and allocate more resources to key modules in the processing chain; Real-time monitoring of the quality indicators of the output voice, triggering corrective measures when the quality drops below a predetermined threshold; Continuously learn the environmental characteristics and optimize the parameter configuration for the specific environment.

[0016] The multi-modal perception intelligent microphone array signal processing system comprises: A multi-modal perception layer comprising a multi-directional visual sensor and a microphone array for acquiring visual signals and sound signals; A feature processing layer comprising a visual feature processing module and an acoustic feature processing module for extracting visual features and acoustic features; A multi-modal fusion layer comprising a topological feature mapping engine and a probabilistic manifold modeling unit for constructing an audio-visual topological feature space and establishing a sound source probability distribution model; An intelligent separation layer comprising a multi-dimensional adversarial generative network and a dynamic balance controller for separating target voice signals; An environment adaptation layer comprising an environment state evaluator and a parameter self-adaptive adjuster for evaluating the environment state and dynamically adjusting the processing parameters; An output processing layer comprising a quality evaluation unit and a sound synthesizer for evaluating the voice quality and generating the final output signal; The system realizes multi-modal perception intelligent microphone array signal processing through the cooperative work of the multi-modal perception layer, the feature processing layer, the multi-modal fusion layer, the intelligent separation layer, the environment adaptation layer and the output processing layer.

[0017] The present application has the following advantages: 1. By constructing an audio-visual topological feature space, the visual features and acoustic features are mapped to a common high-dimensional feature space, and a sound source probability distribution model is established, effectively solving the problem of performance limitation of traditional voice separation technology in multi-speaker scenarios, so that the system can still maintain a voice intelligibility of more than 85% in a 6-person simultaneous speaking scenario.

[0018] 2. Adopting multi-dimensional discriminative generative adversarial network, discriminative features in multiple dimensions such as phonemes, prosody, speakers and environments, and combining with audio-visual feature guided generator, the precision and naturalness of speech separation are significantly improved, and compared with traditional methods, the signal-to-noise ratio can be improved by 9-12dB in high noise environment (signal-to-noise ratio <-5dB).

[0019] 3. Real-time evaluation and parameter adaptive adjustment of environmental state are realized, so that the system can quickly adapt to the dynamic changes of the acoustic environment, and in the fast changing environment scene (such as mobile scene), the adaptive speed is improved by 3 times, and the steady-state performance is improved by 40%.

[0020] 4. Multi-scale fusion optimization strategy and dynamic computing resource allocation mechanism are adopted, which significantly reduces the computational complexity while ensuring processing performance, and compared with traditional deep learning methods, the computational complexity is reduced by 35% and the power consumption is reduced by 42% under the same performance. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 The method flowchart of the application is shown in the figure. Figure 2 The overall architecture schematic diagram of the multi-modal perception intelligent microphone array signal processing system of the application is shown in the figure. DETAILED DESCRIPTION

[0022] Please refer to Figure 1 - Figure 2 The application will be further described in detail below in combination with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only for illustration and explanation of the application, and are not intended to limit the application.

[0023] Referring to Figure 2 The multi-modal perception intelligent microphone array signal processing system 100 provided by the application includes a multi-modal perception layer 110, a feature processing layer 120, a multi-modal fusion layer 130, an intelligent separation layer 140, an environment adaptation layer 150 and an output processing layer 160.

[0024] The multi-modal perception layer 110 includes a multi-directional vision sensor 111 and a microphone array 112 for acquiring visual and acoustic signals. The feature processing layer 120 includes a visual feature processing module 121 and an acoustic feature processing module 122 for extracting visual and acoustic features. The multi-modal fusion layer 130 includes a topological feature mapping engine 131 and a probabilistic manifold modeling unit 132 for constructing an audio-visual topological feature space and establishing a sound source probability distribution model. The intelligent separation layer 140 includes a multi-dimensional adversarial generative network 141 and a dynamic balancing controller 142 for separating target speech signals. The environment adaptation layer 150 includes an environment state evaluator 151 and a parameter self-adaptive adjuster 152 for evaluating environment states and dynamically adjusting processing parameters. The output processing layer 160 includes a quality evaluation unit 161 and a sound synthesizer 162 for evaluating speech quality and generating final output signals.

[0025] Preferably, the multi-modal perception layer 110 is responsible for collecting raw audio-visual data, the feature processing layer 120 extracts feature representations of each modality, the multi-modal fusion layer 130 implements deep fusion of audio-visual information, the intelligent separation layer 140 performs speech separation based on fused features, the environment adaptation layer 150 implements adaptive adjustment of the system to environmental changes, and the output processing layer 160 is responsible for quality evaluation and final output generation. The six levels form a complete processing pipeline, and multi-modal intelligent speech processing is achieved through collaborative work.

[0026] Please refer to Figure 1 According to one embodiment of the present application, a multi-modal perception intelligent microphone array signal processing method is provided, and the detailed process is as follows: The method of the present application first acquires visual signals using a multi-directional vision sensor and acquires acoustic signals using a microphone array.

[0027] In a preferred embodiment of the present application, the multi-directional vision sensor is arranged in front, back, left, right, up and down six directions to realize 360° panoramic visual information acquisition. The sampling frequency of the vision sensor is 30-60 frames per second, and the resolution can be dynamically adjusted, with a maximum support of 1080p. After visual data acquisition, preprocessing is performed, including operations such as distortion removal, illumination equalization and background segmentation, to improve the accuracy of subsequent processing.

[0028] The microphone array can be configured in linear, ring or spherical geometry, and can be flexibly selected according to specific application scenarios. For example, in a conference room scenario, a ring-shaped microphone array is preferably used to realize omnidirectional sound field acquisition; while in a mobile device, a linear microphone array is preferably used to balance performance and integration. The microphone sampling frequency can be set to 16kHz or 48kHz, and the quantization accuracy is 24 bits, to ensure the capture of high-quality acoustic signals. Acoustic data preprocessing includes operations such as pre-filtering, DC offset elimination and gain balancing.

[0029] The present application adopts a hardware timestamp mechanism based on the Precision Time Protocol (PTP) to realize time synchronization of audio-visual data, ensuring synchronization accuracy better than 1 millisecond. The system maintains a sliding time window, with a window length preferably set to 100-500 milliseconds to balance processing delay and system responsiveness. For the sampling rate difference between audio and visual modalities, asynchronous processing and resampling techniques are used to ensure time alignment of the data.

[0030] The method extracts visual features of visual signals and acoustic features of sound signals, wherein the visual features include visual orientation features, facial features, and lip shape features, and the acoustic features include time-frequency features and voiceprint features.

[0031] The visual feature extraction adopts a hierarchical processing strategy. First, the orientation feature extraction calculates the precise spatial position of the speaker through multi-camera triangulation technology, generating a three-dimensional orientation vector describing the speaker relative to the microphone array, which is dynamically updated at a frequency of 30 Hz to realize real-time tracking of the speaker's movement. Second, the facial feature extraction uses a multi-scale feature pyramid technique to extract facial features at different resolutions, calculates the pitch, yaw, and roll angles of the head in three degrees of freedom, and generates a visual identity embedding vector of the speaker. Finally, the lip shape feature extraction accurately segments the lip contour through an adaptive region growing algorithm, captures the time sequence features of lip shape changes, and maps the lip shape changes to the viseme classification in phonetics.

[0032] The acoustic feature extraction adopts a multi-dimensional processing method. The time-frequency feature extraction simultaneously performs short-time Fourier transform and wavelet packet transform to extract the modulation characteristics and spectral envelope of the sound signal. The parameters of the short-time Fourier transform are set as follows: frame length 25 milliseconds, frame shift 10 milliseconds, and Hanning window function. The voiceprint feature extraction mainly analyzes the speaker's vocal tract characteristics, prosodic features (pitch, intensity, speech rate, etc.), and suprasegmental features, and constructs a compact speaker identity representation. In addition, environmental acoustic features are extracted, including room impulse response estimation, background noise statistical feature analysis, and sound source direction estimation.

[0033] The extracted raw features are denoised, enhanced, and normalized to improve feature quality and consistency. Specifically, for visual features, spatial filtering and temporal smoothing techniques are used to reduce jitter; for acoustic features, spectral subtraction and adaptive normalization techniques are used to enhance signal quality. The final extracted visual features and acoustic features have dimensions of 64-128 and 128-256, respectively.

[0034] The method constructs an audio-visual topological feature space, maps the visual features and acoustic features to the audio-visual topological feature space, and establishes a sound source probability distribution model based on the audio-visual topological feature space.

[0035] In an embodiment of the present application, the audio-visual topological feature space is a 128-256 dimensional joint feature space. The feature space adopts Riemannian distance metric, which is defined as follows: , wherein: is the Riemannian distance between two points and in the feature space; and are two points in the feature space, respectively dimensional vectors; and are the th dimensional components of the point and the point , respectively; is the dimension of the feature space, ranging from 128 to 256; is the weight factor of the th dimensional feature, reflecting the importance of the dimensional feature, with an initial value of 1, and then dynamically adjusted according to the feature discriminability, usually ranging from 0.5 to 2.0.

[0036] The process of mapping visual features and acoustic features to the joint feature space adopts a nonlinear dimensionality reduction technique. First, the original feature vectors are preprocessed to ensure consistent feature scales across different modalities. Then, the features are projected to the common space through mapping functions: , wherein: is the original visual feature vector, with a dimension of 64-128; is the original acoustic feature vector, with a dimension of 128-256; and are the mapping functions for visual features and acoustic features, respectively; and are the mapped visual feature and acoustic feature representations, both with a dimension of 128-256. The mapping functions are implemented using a multilayer perceptron, with a typical structure containing 3-5 hidden layers, each with 256-512 neurons, and a ReLU (Rectified Linear Unit) activation function.

[0037] The feature distance metric is optimized through a contrastive learning method to enhance intra-class cohesion and inter-class separability. The contrastive loss function is defined as: , wherein: is the contrastive loss function; is the cosine similarity function, which calculates the cosine value of the angle between two vectors, ranging from -1 to 1; is a temperature parameter, controlling the smoothness of the distribution, usually set to 0.07; represents the acoustic feature of the th negative sample; is the number of negative samples, usually set to 64-128; is a natural exponential function.

[0038] Finally, an accurate correspondence between the visual feature and the acoustic feature is established, and the mutual information between the audio-visual features is maximized. The mutual information maximization objective function is: , wherein: is the mutual information maximization loss function; represents the mutual information between the visual feature and the acoustic feature ; represents the entropy of the visual feature , measuring its uncertainty; represents the conditional entropy of the visual feature given the acoustic feature . In actual implementation, the mutual information is approximated by a mutual information neural estimator.

[0039] Through the above process, a high-dimensional feature space that preserves the topological relationship of audio-visual features is constructed, and on this basis, a sound source probability distribution model is established, providing strong prior information for subsequent speech separation.

[0040] The method of the present application adopts a multi-dimensional discriminative generative adversarial network to process sound signals and separate target speech signals.

[0041] The multi-dimensional discriminative generative adversarial network of the present application includes a multi-level discriminator structure and a generator based on audio-visual feature guidance. The multi-level discriminator structure includes a phoneme-level discriminator, a prosody-level discriminator, a speaker-level discriminator, and an environment-level discriminator, which are respectively responsible for evaluating different dimensions of speech characteristics.

[0042] The phoneme-level discriminator focuses on the authenticity discrimination of short-time speech segments (10-30 milliseconds), mainly evaluating the local acoustic characteristics of speech. The prosody-level discriminator evaluates the naturalness of speech rhythm and intonation in a longer time period (100-500 milliseconds), focusing on suprasegmental features. The speaker-level discriminator verifies the consistency of the separated speech and the speaker features, ensuring the preservation of speech identity characteristics. The environment-level discriminator ensures the coordination of the speech and the current acoustic environment, avoiding the unnatural feeling caused by the mismatch of environmental characteristics.

[0043] The discriminator extracts discriminative features at multiple time scales, and uses an attention mechanism to dynamically focus on the most discriminative feature regions. The calculation formula of the attention mechanism is: , wherein: is the output matrix of the attention mechanism; is the query matrix, dimension ; is the key matrix, dimension ; is the value matrix, dimension ; is the query sequence length; is the key / value sequence length; , and are the dimensions of the query, key and value vectors, usually set to ; denotes the matrix is multiplied by the transpose of the matrix , resulting in a dimension ; is a scaling factor to prevent the softmax gradient from vanishing due to large inner product values; is a normalization function that converts each row vector into a probability distribution. The discriminator also employs an adaptive threshold adjustment technique to dynamically adjust the discrimination threshold according to the complexity of the environment, quantifying the reliability of the discrimination result, and weighting and fusing the multi-layer discrimination results according to the reliability.

[0044] The generator based on audio-visual feature guidance adopts a U-shaped network structure, containing multi-level skip connections and multi-path processing streams. The network depth is 8-12 layers, with the number of channels starting from 64 and reaching 512 in the deepest layer. The skip connection adopts a residual connection method to alleviate the gradient vanishing problem and preserve detailed information. The generator guides the sound separation process through visual condition embedding and voiceprint condition control. The condition embedding is realized through a FiLM (Feature Intelligent Linear Modulation) layer: , , wherein: is the scaling parameter vector, with the same dimension as the feature vector ; is the offset parameter vector, with the same dimension as the feature vector ; is the condition vector (visual feature or voiceprint feature), with a dimension of 128-256; and are mapping functions, usually single-layer or double-layer fully connected networks; is the input feature vector; is the modulated feature vector; denotes element-wise multiplication (Hadamard product). and The scaling and offset of features are controlled separately to achieve the modulation of features by conditional information.

[0045] The processing of the multi-dimensional discriminative adversarial generative network includes: decomposing the sound signal into multiple time-frequency components, assigning each time-frequency component to a different sound source based on audio-visual features, constructing a soft masking matrix for each sound source, and eliminating residual cross-interference through post-processing. The masking matrix generation formula is: , in: For the The sound source at time and frequency The masking value at , range is [0,1]; For the estimated The sound source at time and frequency The complex spectrum at ; Indicates plural The square of the modulus, i.e. the power spectrum; is the total number of sound sources; Indicates that all sum of the sound sources; is a small positive number, usually set to , which is used to prevent numerical instability caused by the denominator being zero. Residual cross-interference elimination is achieved by combining spectral subtraction and adaptive filtering.

[0046] Through the processing of multi-dimensional discriminative adversarial generative networks, the system can effectively separate the target speech from the mixed speech signal while maintaining the naturalness and identity characteristics of the speech.

[0047] The method of the present invention evaluates the state of the acoustic environment in real time and dynamically adjusts the processing parameters of the multi-dimensional discriminant adversarial generation network.

[0048] In an embodiment of the present invention, environmental status assessment includes extracting acoustic environmental parameters such as reverberation time (RT60), early reflection density, and background noise level, as well as identifying sound source distribution characteristics such as the number of sound sources, spatial distribution, and motion state. Reverberation time estimation uses the energy decay method: , in: Reverberation time, in seconds (s), represents the time required for the sound energy to decay by 60dB; is the slope of the room impulse response energy decay curve, in dB / s (decibel / second), estimated from the energy decay curve by linear regression method; Decibel number representing the need for attenuation. Early reflection density is estimated by the number of reflection peaks within the first 50-80 milliseconds of the impulse response. Background noise level is estimated by the minimum statistic method: , where: is the noise power spectrum estimate at frequency ; is the short-time Fourier transform coefficient of the observed signal at time and frequency ; denotes the modulus square of the complex , i.e. the power spectrum; denotes taking the minimum value within the time window ; is the time window length, typically set to 1-2 seconds, corresponding to 100-200 frames (assuming a frame shift of 10 milliseconds).

[0049] The system matches the current environment with predefined environment categories, generating an environment feature vector. Predefined environment categories include quiet indoor (background noise < 30 dB, RT60 < 0.3 s), general indoor (background noise 30-50 dB, RT60 0.3-0.8 s), noisy indoor (background noise > 50 dB, RT60 0.3-0.8 s), outdoor (background noise varies greatly, RT60 approaches 0), and moving scene (parameters change rapidly).

[0050] Environment dynamic change detection uses a sliding window analysis technique to analyze environmental parameters over a continuous time window (typical value 0.5-2 seconds) to identify significant change points in environmental parameters. The change point detection algorithm is based on the CUSUM (Cumulative Sum) method: , where: is the cumulative sum statistic at the current time ; is the cumulative sum statistic at the previous time , with an initial value ; is the observation value (such as background noise level or RT60) at time ; is the target mean value, i.e. the expected parameter steady-state value; is the sensitivity parameter, controlling the detection sensitivity, typically set to 0.5; takes the maximum value between 0 and the parameter, ensuring that the cumulative sum is non-negative. When exceeds a pre-set threshold (typical value 5), it is determined as a change point.

[0051] According to the environmental state, the system dynamically adjusts filter parameters, gain control parameters and network weight parameters.Parameter adjustment adopts a continuous parameter adjustment mechanism to ensure the continuity of the parameter adjustment process and avoid sound artifacts caused by mutations.The mathematical expression of the smooth transition strategy is: , Wherein: is the parameter value at the current time ; is the parameter value at the previous time ; is the target parameter value, which is determined according to the current environmental state; is a smoothing factor that controls the parameter update speed, with a value range of [0, 1], usually set to 0.05-0.2, dynamically adjusted according to the environmental change speed, taking a larger value when the environment changes fast and a smaller value when the environment is stable.

[0052] The system also implements environmental feedback closed-loop control, which adjusts the parameter direction through signal-to-noise ratio improvement, speech intelligibility, artifact level and other effect evaluation indicators, and realizes self-optimization of the parameters.The objective function of the closed-loop control is: , Wherein: is the comprehensive performance score of parameter , the higher the better the performance; is the signal-to-noise ratio evaluation function when using parameter , with a value range of [0, 1], obtained by signal-to-noise ratio gain normalization; is the intelligibility evaluation function when using parameter , with a value range of [0, 1], obtained by CSII or PESQ score normalization; is the artifact level evaluation function when using parameter , with a value range of [0, 1], the larger the value, the more serious the artifact; , and are weight coefficients that control the importance of each indicator, satisfying , typical values are 0.4, 0.4 and 0.2 respectively.

[0053] Through the above environmental self-adaptive mechanism, the system can real-time perceive environmental changes and make corresponding adjustments to maintain stable performance in a dynamic environment.

[0054] The method of the present application evaluates the quality of the separated multi-channel target speech signals and selects the highest quality speech signal as the output.

[0055] The present invention establishes a comprehensive speech quality assessment index system, including objective indicators, perceptual indicators and comprehensive quality score. Objective indicators include signal-to-noise ratio, speech clarity index and cepstral distance, etc. The calculation formula is as follows: Signal-to-noise ratio (SNR) calculation: , in: is the signal-to-noise ratio, in decibels (dB); The first Amplitude of sampling points; is the estimated speech signal Amplitude of sampling points; For all sampling points sum; For signal The square of represents the signal energy; is the square of the estimation error, which represents the noise energy; is the logarithmic function with base 10.

[0056] Speech Intelligibility Index (CSII) calculation: , in: is the speech clarity index, with a value range of [0,1], and a larger value indicates a higher clarity; is the original signal in time frame and frequency The short-time Fourier transform coefficients at ; To estimate the signal in the time frame and frequency The short-time Fourier transform coefficients at ; for The complex conjugate of and Respectively and The square of the modulus; For all timeframes and frequency points Summation.

[0057] Cepstral distance (CD) calculation: , in: is the cepstrum distance, in dB; The original voice cepstral coefficients; To estimate the speech cepstral coefficients; is the dimension of the cepstral coefficient, usually 13; denotes the summation of all cepstral coefficients; is the natural logarithm The result of acting on 10 is approximately 2.303; is the coefficient for converting the natural logarithm scale to the decibel scale, which is approximately 4.343.

[0058] The perceptual indicators are obtained by subjective mean opinion score (MOS) and speech naturalness score, etc. In the actual system, these subjective scores are predicted by training models.

[0059] The comprehensive quality score is calculated by weighting: , wherein: is the comprehensive quality score, the value range is [0, 1], and the larger the value is, the higher the quality is; 、 、 and respectively represent the signal-to-noise ratio, speech articulation index, cepstral distance and subjective average score normalized to the interval [0, 1]; The cepstral distance is converted into a positive indicator, and the larger the value is, the higher the quality is; 、 、 and are weight coefficients, satisfying , and the empirical values are 0.25, 0.3, 0.2 and 0.25, respectively.

[0060] The system compares the quality of the separation results of different microphone channels, and selects the channel with the highest quality score as the output. In order to further improve the output quality, the system performs customized enhancement processing on the selected speech signal, including spectral smoothing, harmonic structure enhancement and sub-band gain adjustment, etc.

[0061] The spectral smoothing adopts a time-frequency filter: , wherein: is the smoothed spectrum, the value at time and frequency ; is the original spectrum, the value at time and frequency ; is the filter weight, representing the weight coefficient at time offset and frequency offset ; denotes the summation within the time window and the frequency window ; and are the window half-widths in time and frequency direction respectively, typical values are 1 and 2, corresponding to a 3x5 filter window; the denominator term is used for normalization, ensuring the sum of filter weights is 1.

[0062] Harmonic structure enhancement is achieved by strengthening the harmonic peaks of the speech signal: , where: is the enhanced spectrum, value at time t and frequency f; is the original spectrum, value at time t and frequency f; is the harmonic saliency map, value range is [0, 1], the larger the value, the more significant the harmonic characteristics of the time-frequency point; is the enhancement factor, controlling the enhancement strength, usually set to 0.2-0.5; is the enhancement coefficient, range is [1, 1+α]. The sub-band gain adjustment adjusts the gain of each frequency band according to the importance of the speech content: , where: is the final output spectrum, value at time t and frequency f; is the original spectrum, value at time t and frequency f; is the gain factor for frequency f, set according to the importance of the frequency band. Typical gain settings are: low frequency band (<300Hz) gain 0.8, speech main frequency band (300-3400Hz) gain 1.2, high frequency band (>3400Hz) gain 0.9.

[0063] Through the above speech quality evaluation and enhancement processing, the system outputs high-quality separated speech signals, meeting the needs of various application scenarios.

[0064] The method of the present application also includes measures such as selecting preset configurations according to application scenarios, dynamically allocating computing resources, monitoring output quality in real time, and continuously learning environmental characteristics, to further improve the practicality of the system.

[0065] The present application selects preset configurations according to application scenarios, including conference scenario configuration, mobile scenario configuration, and noise environment configuration. The conference scenario configuration optimizes the multi-person alternating speaking scenario, especially strengthening the speaker switching detection and identity maintenance; the mobile scenario configuration adapts to environmental changes in motion, enhancing the speed of environmental adaptation; the noise environment configuration enhances the performance in extreme noise environments, improving the anti-noise ability.

[0066] The system realizes dynamic computing resource allocation, and allocates more resources to key modules in the processing chain. The resource allocation strategy is based on task priority and real-time requirements, and an adaptive algorithm is used to adjust the computing resource quota of each module. For example, in a high-noise environment, the resource quota of the noise suppression module is increased; in a multi-speaker scene, the resource quota of the speech separation module is increased.

[0067] The system continuously monitors the quality indicators of the output speech, and triggers corrective measures when the quality drops below a predetermined threshold (usually set to 20%). The corrective measures include parameter readjustment, backup processing chain activation, and user feedback prompts. In addition, the system also continuously learns the environmental characteristics, optimizes the parameter configuration for specific environments, and realizes long-term performance improvement.

[0068] Through the above optimization measures, the system realizes good practicability and user experience while maintaining high performance.

[0069] The multi-modal perception intelligent microphone array signal processing system provided by the application comprises a multi-modal perception layer, a feature processing layer, a multi-modal fusion layer, an intelligent separation layer, an environment adaptation layer and an output processing layer.

[0070] The multi-modal perception layer comprises a multi-directional vision sensor and a microphone array, which are used to acquire visual signals and sound signals. The feature processing layer comprises a visual feature processing module and an acoustic feature processing module, which are used to extract visual features and acoustic features. The multi-modal fusion layer comprises a topological feature mapping engine and a probabilistic manifold modeling unit, which are used to construct an audio-visual topological feature space and establish a sound source probability distribution model. The intelligent separation layer comprises a multi-dimensional generative adversarial network and a dynamic balancing controller, which are used to separate target speech signals. The environment adaptation layer comprises an environment state evaluator and a parameter self-adaptive adjuster, which are used to evaluate the environment state and dynamically adjust the processing parameters. The output processing layer comprises a quality evaluation unit and a sound synthesizer, which are used to evaluate the speech quality and generate the final output signal.

[0071] The system realizes multi-modal perception intelligent microphone array signal processing through the cooperative work of the multi-modal perception layer, the feature processing layer, the multi-modal fusion layer, the intelligent separation layer, the environment adaptation layer and the output processing layer. The modular design of the system architecture makes it have good expansibility and maintainability, and can adapt to the needs of different application scenarios.

[0072] The hardware implementation of the system can be based on a heterogeneous computing platform of CPU+GPU / DSP, supporting resource-aware scheduling and power consumption control. The software architecture adopts a pipeline processing mode, supporting real-time processing and low-latency response. The system provides standardized API interfaces, facilitating integration with other systems and secondary development.

[0073] The multi-modal perception intelligent microphone array signal processing method and system of the application is suitable for various application scenarios, including intelligent conference systems, intelligent vehicle systems, intelligent home devices, and mobile intelligent terminals.

[0074] In the intelligent conference system, the application can realize functions such as multi-person conference voice accurate separation, remote participant voice enhancement, and automatic transcription of conference records. Tests show that in a complex conference scene with 6 people speaking at the same time, the system can still maintain a speech intelligibility of more than 85%, which is more than 40% higher than traditional methods.

[0075] In the intelligent vehicle system, the application can realize functions such as driver voice command recognition, multi-passenger interaction management, and road noise suppression. Tests show that under high-speed driving conditions (vehicle noise about 75dB), the speech recognition accuracy of the system can still maintain more than 90%, which is more than 30% higher than traditional methods.

[0076] In the intelligent home device, the application can realize functions such as far-field voice interaction, multi-user voice differentiation, and voice recognition under home appliance noise. Tests show that under the condition of a distance of 5 meters and a TV volume of 60dB, the voice wake-up success rate of the system can still maintain more than 95%, which is more than 25% higher than traditional methods.

[0077] In the mobile intelligent terminal, the application can realize functions such as handheld device call quality improvement, video conference voice enhancement, and noisy environment voice memo. Tests show that in a street environment (environmental noise about 70dB), the speech call quality score (MOS) of the system reaches 3.8, which is 0.8 points higher than traditional methods.

[0078] Overall, the application has achieved significant improvement in voice quality indicators: the signal-to-noise ratio is improved by an average of 15-20dB, the speech intelligibility (PESQ score) is improved by 1.5-2.0, and the speech intelligibility is improved by 40%-60% in noisy environments. In terms of system performance indicators, the application realizes a processing delay of less than 50 milliseconds in a streaming processing mode, can run in real time on a mid-end processor, and can adapt within 1-2 seconds after environmental changes.

[0079] Compared with the prior art, the application improves the separation accuracy by 65% in a multi-speaker scenario, improves the performance by 80% in a low signal-to-noise ratio environment, improves the computing efficiency by 40%, and improves the adaptability by 200%. These advantages are mainly due to the multi-modal coordination effect, environmental adaptation mechanism, and efficient algorithm implementation of the application.

[0080] The application provides a multi-modal perception intelligent microphone array signal processing method and system, through deep fusion and collaborative processing of audio-visual multi-modal information, combining a topology enhanced generative adversarial network architecture and an environment adaptive mechanism, high-quality speech separation and enhancement in a complex acoustic environment are realized. The application has significant advantages in multi-speaker scenes, low signal-to-noise ratio environments and rapidly changing environments, and brings a qualitative leap to intelligent speech interaction technology.

[0081] The main innovations of the application include: audio-visual topology feature space construction and mapping technology, multi-dimensional discriminative generative adversarial network architecture, environment adaptive mechanism and multi-scale fusion optimization strategy. These innovations synergize with each other to solve the key technical problems of speech separation and enhancement in complex environments.

[0082] The application has good practicability and popularization value, and can be widely applied to intelligent conference, intelligent vehicle, intelligent home and mobile intelligent terminal and various scenes, and brings significant improvement to speech interaction experience.

[0083] The above only describes the preferred embodiments of the application and is not used to limit the application. The application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.

Claims

1. A multimodal sensing intelligent microphone array signal processing method, characterized in that: include: Using multi-directional visual sensors to obtain visual signals and microphone arrays to obtain sound signals; Extracting visual features of the visual signal and acoustic features of the sound signal, wherein the visual features include visual orientation features, facial features, and lip shape features, and the acoustic features include time-frequency features and voiceprint features; Constructing an audio-visual topological feature space, mapping the visual features and the acoustic features to the audio-visual topological feature space, and establishing a sound source probability distribution model based on the audio-visual topological feature space; Processing the sound signal using a multi-dimensional discriminant adversarial generative network to separate the target speech signal, wherein the multi-dimensional discriminant adversarial generative network includes a multi-layer discriminator structure and a generator guided by audio-visual features; Real-time evaluation of the acoustic environment state and dynamic adjustment of processing parameters of the multi-dimensional discriminant adversarial generative network; The quality of the separated multi-channel target speech signals is evaluated and the speech signal with the highest quality is selected as the output.

2. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that: The method of acquiring visual signals by using a multi-directional visual sensor and acquiring sound signals by using a microphone array specifically includes: Use visual sensors arranged in six directions: front, back, left, right, up, and down to collect 360° panoramic visual information; Acquire sound field information using microphone arrays that can be configured in linear, annular, or spherical geometries; The visual signal and the sound signal are time synchronized to ensure a synchronization accuracy better than 1 millisecond.

3. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that: The extracting of the visual features of the visual signal and the acoustic features of the sound signal specifically includes: Performing orientation recognition, facial detection, and lip tracking on the visual signal to extract visual orientation features, facial features, and lip features; Performing time-frequency analysis, voiceprint extraction, and acoustic fingerprint generation on the sound signal to extract time-frequency features and voiceprint features; The extracted visual features and acoustic features are subjected to noise reduction, enhancement and normalization processing.

4. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that: The construction of the audio-visual topological feature space specifically includes: Define a joint feature space of 128 to 256 dimensions, and define a Riemannian distance metric in the feature space; Projecting the visual features and the acoustic features into the joint feature space by nonlinear dimensionality reduction; Optimize feature distance metrics through contrastive learning methods to enhance intra-class aggregation and inter-class separation; Establish accurate correspondence between visual features and acoustic features, and maximize the mutual information between visual and audio features.

5. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that: The multi-dimensional discriminative adversarial generation network includes: Phoneme-level discriminator, used to assess the authenticity of short speech segments; a prosody-level discriminator to assess the naturalness of speech rhythm and intonation; Speaker-level discriminator, used to verify the consistency of separated speech and speaker characteristics; An environment-level discriminator, used to ensure the coordination of speech with the current acoustic environment; The generator based on audio-visual feature guidance adopts a U-shaped network structure, including multi-stage jump connections and multi-path processing flows, and guides the sound separation process through visual condition embedding and voiceprint condition control.

6. The multimodal sensing intelligent microphone array signal processing method according to claim 5, characterized in that: The processing process of the multi-dimensional discriminant adversarial generative network includes: Extract discriminative features at multiple time scales and dynamically focus on the most discriminative feature areas through the attention mechanism; Adaptively adjust the discrimination threshold, quantify the credibility of the discrimination results, and perform weighted fusion of multi-layer discrimination results based on the credibility; Decomposing the sound signal into multiple time-frequency components, and assigning each time-frequency component to a different sound source based on the audiovisual features; A soft masking matrix is ​​constructed for each sound source, and residual cross-interference is eliminated through post-processing.

7. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that: The real-time evaluation of the acoustic environment state and the dynamic adjustment of the processing parameters of the multi-dimensional discriminant adversarial generative network specifically include: Extract acoustic environment parameters such as reverberation time, early reflection density, and background noise level; Identify sound source distribution characteristics such as the number of sound sources, spatial distribution, and motion status; Match the current environment with predefined environment categories to generate an environment feature vector; Analyze environmental parameters over continuous time windows and identify significant change points of environmental parameters; Adjust the filter parameters, gain control parameters and network weight parameters according to the environmental status to ensure the continuity of the parameter adjustment process.

8. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that: The performing quality evaluation on the separated multi-channel target speech signals and selecting the speech signal with the highest quality as output specifically includes: Calculate objective indicators such as signal-to-noise ratio, speech clarity index, and cepstral distance; Evaluate perceptual metrics such as subjective mean opinion score and speech naturalness; Calculate weighted quality scores by integrating multiple indicators; Compare the quality of the separation results of different microphone channels and select the channel with the highest quality score; Perform customized enhancement processing such as spectrum smoothing, harmonic structure enhancement, and sub-band gain adjustment on the selected speech signal.

9. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that: Also includes: Select preset configurations based on application scenarios, including conference scenario configurations, mobile scenario configurations, and noise environment configurations; Dynamically allocate computing resources to allocate more resources to key modules in the processing chain; Monitor the quality indicators of output voice in real time and trigger corrective measures when the quality drops beyond a predetermined threshold; Continuously learn environmental characteristics and optimize parameter configurations for specific environments.

10. A multimodal sensing intelligent microphone array signal processing system, which implements the multimodal sensing intelligent microphone array signal processing method according to any one of claims 1 to 9, characterized in that: include: Multimodal perception layer, including multi-directional visual sensors and microphone arrays, used to acquire visual and sound signals; Feature processing layer, including visual feature processing module and acoustic feature processing module, used to extract visual features and acoustic features; The multimodal fusion layer includes a topological feature mapping engine and a probabilistic manifold modeling unit, which is used to construct the audio-visual topological feature space and establish a sound source probability distribution model; Intelligent separation layer, including a multi-dimensional adversarial generative network and a dynamic equalization controller, for separating the target speech signal; The environmental adaptation layer includes an environmental state evaluator and a parameter adaptive adjuster, which are used to evaluate the environmental state and dynamically adjust the processing parameters; Output processing layer, including quality assessment unit and voice synthesizer, used to assess speech quality and generate the final output signal; The system realizes multimodal perception intelligent microphone array signal processing through the collaborative work of the multimodal perception layer, feature processing layer, multimodal fusion layer, intelligent separation layer, environment adaptation layer and output processing layer.

Citation Information

Patent Citations

  • Multi-sound-source separation system and method based on generative adversarial network

    CN116312609A

  • Multi-microphone array beamforming signal enhancement method and device

    CN119811408A

  • Method and system of environment sensitive automatic speech recognition

    US20160284349A1

  • Multi-microphone audio signal unifier and methods therefor

    US20240412750A1

Cited By

  • Dynamic feature enhancement method for noise robustness speech recognition

    CN121096327A

  • Mandarin pronunciation real-time correction method and system based on multi-modal streaming learning

    CN121393456A

  • Mandarin pronunciation real-time correction method and system based on multi-modal streaming learning

    CN121393456B

  • Self-adaptive voice noise reduction method and system in multi-noise scene

    CN121565189A

  • Self-adaptive speech enhancement method and device guided by acoustic environment perception

    CN121789701A