Multimodal perception based smart microphone array signal processing method and system
By employing a multimodal sensing-based intelligent microphone array signal processing method, and utilizing audiovisual information fusion and environmental adaptation technologies, the problem of speech separation and enhancement in complex acoustic environments using traditional microphone arrays is solved, achieving high-quality speech signal processing.
Patent Information
- Application Number
- CN202511337012.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Traditional microphone array signal processing techniques struggle to effectively separate and enhance speech in complex acoustic environments, especially in multi-speaker, strong reverberation, and high-noise environments where performance is limited. Furthermore, existing deep learning methods lack multimodal information fusion and environmental adaptation capabilities.
Visual and acoustic signals are acquired using multi-directional visual sensors and microphone arrays. Speech separation is performed through audiovisual topological feature space mapping and multi-dimensional discriminative adversarial generative networks. Combined with an environment adaptive mechanism, processing parameters are dynamically adjusted to achieve deep fusion and collaborative processing of multimodal information.
It significantly improves speech separation accuracy and naturalness in complex acoustic environments, reduces computational complexity, enhances the system's adaptability and speech quality in dynamic environments, and improves signal-to-noise ratio and speech intelligibility.
Smart Images

Figure CN120808810B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of signal processing, in particular to a multi-modal perception intelligent microphone array signal processing method and system, which is applied to acoustic signal processing scenes such as speech separation, enhancement and noise reduction. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, intelligent voice interaction has become an important way of human-computer interaction. However, in the actual application environment, due to the existence of multiple speaker interference, environmental noise and reverberation, etc., the speech signal is often seriously polluted, which leads to a significant decline in speech recognition and interaction performance, which is known as the cocktail party effect in the field of signal processing.
[0003] Traditional microphone array signal processing technology mainly relies on acoustic information, such as spatial filtering technology based on beamforming, blind source separation technology based on independent component analysis, etc. These technologies perform well under ideal conditions, but their performance is limited in complex acoustic environments. On the one hand, traditional beamforming technology needs to accurately estimate the direction of the sound source, which is difficult to accurately locate in a multi-source, strong reverberation environment; on the other hand, blind source separation technology has high requirements for the independence and sparsity of signals, and it is difficult to deal with strongly correlated interference and non-sparse noise.
[0004] In recent years, deep learning technology has made significant progress in the field of speech processing, such as speech separation and enhancement methods based on deep neural networks. However, these methods are mainly based on a single acoustic modality and lack effective use of visual and other modal information, while the human perception system precisely achieves efficient sound source separation through the fusion of multi-modal information. In addition, existing deep learning methods are mostly static models, lacking the ability to adapt to dynamically changing environments, and have high computational complexity, making it difficult to efficiently deploy on resource-constrained devices.
[0005] Therefore, it is urgent to develop an intelligent microphone array signal processing method and system that can effectively fuse audio-visual multi-modal information and have environmental adaptation capability to solve the key technical problems of speech separation and enhancement in complex environments. SUMMARY
[0006] The purpose of the present application is to provide a multi-modal perception intelligent microphone array signal processing method and system, which realizes high-quality speech separation and enhancement in complex acoustic environments through deep fusion and collaborative processing of audio-visual multi-modal information, combined with a topologically enhanced generative adversarial network architecture and an environmental adaptation mechanism.
[0007] The present application proposes a multi-modal perception intelligent microphone array signal processing method, which comprises:
[0008] A multi-directional visual sensor is used to obtain visual signals and a microphone array is used to obtain sound signals;
[0009] extracting visual features of the visual signals and acoustic features of the sound signals, wherein the visual features include visual orientation features, face features and lip shape features, and the acoustic features include time-frequency features and voiceprint features;
[0010] constructing an audio-visual topological feature space, mapping the visual features and the acoustic features to the audio-visual topological feature space, and establishing a sound source probability distribution model based on the audio-visual topological feature space;
[0011] processing the sound signals by using a multi-dimensional discriminative generative adversarial network to separate target speech signals, wherein the multi-dimensional discriminative generative adversarial network includes a multi-level discriminator structure and a generator based on audio-visual feature guidance;
[0012] real-time evaluating an acoustic environment state and dynamically adjusting processing parameters of the multi-dimensional discriminative generative adversarial network;
[0013] performing quality evaluation on the separated multi-path target speech signals and selecting a speech signal with the highest quality as an output.
[0014] As a preferred embodiment, the acquiring of the visual signals by using multi-orientation visual sensors and the acquiring of the sound signals by using a microphone array specifically include:
[0015] acquiring 360° panoramic visual information by using visual sensors arranged in front, back, left, right, up and down directions;
[0016] acquiring sound field information by using a microphone array that can be configured in linear, annular or spherical geometric arrangement;
[0017] performing time synchronization on the visual signals and the sound signals to ensure that the synchronization accuracy is better than 1 millisecond.
[0018] As a preferred embodiment, the extracting of the visual features of the visual signals and the acoustic features of the sound signals specifically include:
[0019] performing orientation recognition, face detection and lip shape tracking on the visual signals to extract visual orientation features, face features and lip shape features;
[0020] performing time-frequency analysis, voiceprint extraction and acoustic fingerprint generation on the sound signals to extract time-frequency features and voiceprint features;
[0021] performing noise reduction, enhancement and normalization processing on the extracted visual features and acoustic features.
[0022] As a preferred embodiment, the constructing of the audio-visual topological feature space specifically includes:
[0023] define a joint feature space of 128 to 256 dimensions in which a Riemannian distance metric is defined;
[0024] project the visual features and the acoustic features into the joint feature space by nonlinear dimension reduction;
[0025] optimize the feature distance metric by a contrastive learning method to enhance intra-class cohesion and inter-class separation;
[0026] establish an accurate correspondence between visual features and acoustic features, and maximize the mutual information between audio-visual features.
[0027] Preferably, the multi-dimensional discriminative generative adversarial network comprises:
[0028] a phoneme-level discriminator for evaluating the authenticity of short-time speech segments;
[0029] a prosody-level discriminator for evaluating the naturalness of speech rhythm and intonation;
[0030] a speaker-level discriminator for verifying the consistency of separated speech and speaker features;
[0031] an environment-level discriminator for ensuring the coordination of speech and the current acoustic environment;
[0032] The audio-visual feature guided generator adopts a U-shaped network structure, including multi-level skip connections and multi-path processing streams, and guides the sound separation process through visual condition embedding and voiceprint condition control.
[0033] Preferably, the processing process of the multi-dimensional discriminative generative adversarial network comprises:
[0034] extract discriminative features on multiple time scales and dynamically focus on the most discriminative feature regions through attention mechanisms;
[0035] adaptively adjust the discriminative threshold, quantify the reliability of the discriminative results, and weight and fuse the multi-layer discriminative results according to the reliability;
[0036] decompose the sound signal into multiple time-frequency components, and assign each time-frequency component to a different sound source based on the audio-visual features;
[0037] construct a soft masking matrix for each sound source, and eliminate residual cross-interference through post-processing.
[0038] Preferably, the real-time evaluation of the acoustic environment state dynamically adjusts the processing parameters of the multi-dimensional discriminative generative adversarial network, specifically comprising:
[0039] extracting acoustic environment parameters such as reverberation time, early reflection density, and background noise level;
[0040] Identify the number of sound sources, spatial distribution, and motion state of the sound source distribution characteristics;
[0041] Match the current environment with the predefined environment category to generate an environment feature vector;
[0042] Analyze the environment parameters over a continuous time window to identify significant changes in the environment parameters;
[0043] Adjust the filter parameters, gain control parameters, and network weight parameters according to the environment state to ensure the continuity of the parameter adjustment process.
[0044] As a preferred, the quality evaluation of the separated multi-path target speech signal, the highest quality speech signal is selected as the output, specifically includes:
[0045] Calculate objective indicators such as signal-to-noise ratio, speech articulation index, and cepstrum distance;
[0046] Evaluate perceptual indicators such as subjective mean opinion score and speech naturalness;
[0047] Calculate the weighted quality score by integrating multiple indicators;
[0048] Compare the quality of the separation results of different microphone channels and select the channel with the highest quality score;
[0049] Perform customized enhancement processing such as spectral smoothing, harmonic structure enhancement, and sub-band gain adjustment on the selected speech signal.
[0050] As a preferred, it also includes:
[0051] Select a preset configuration according to the application scenario, including conference scenario configuration, mobile scenario configuration, and noise environment configuration;
[0052] Dynamically allocate computing resources, and allocate more resources to key modules in the processing chain;
[0053] Real-time monitoring of the quality indicators of the output speech, and triggering corrective measures when the quality drops below a predetermined threshold;
[0054] Continuously learn the environmental characteristics and optimize the parameter configuration for specific environments.
[0055] A multi-modal perception intelligent microphone array signal processing system, comprising:
[0056] A multi-modal perception layer including a multi-directional vision sensor and a microphone array for acquiring visual and acoustic signals;
[0057] A feature processing layer including a visual feature processing module and an acoustic feature processing module for extracting visual and acoustic features;
[0058] A multi-modal fusion layer including a topological feature mapping engine and a probabilistic manifold modeling unit is configured to construct an audio-visual topological feature space and establish a sound source probability distribution model.
[0059] An intelligent separation layer including a multi-dimensional generative adversarial network and a dynamic balancing controller is configured to separate a target speech signal.
[0060] An environment adaptation layer including an environment state evaluator and a parameter adaptive adjuster is configured to evaluate an environment state and dynamically adjust processing parameters.
[0061] An output processing layer including a quality evaluation unit and a sound synthesizer is configured to evaluate speech quality and generate a final output signal.
[0062] The system realizes multi-modal perception intelligent microphone array signal processing through the cooperative work of the multi-modal perception layer, the feature processing layer, the multi-modal fusion layer, the intelligent separation layer, the environment adaptation layer, and the output processing layer.
[0063] The present application has the following beneficial effects:
[0064] 1. By constructing an audio-visual topological feature space, the visual features and acoustic features are mapped to a common high-dimensional feature space, and a sound source probability distribution model is established, effectively solving the problem of performance limitation of traditional speech separation technology in a multi-speaker scene, so that the system can still maintain a speech intelligibility of more than 85% in a 6-person simultaneous speaking scene.
[0065] 2. A multi-dimensional discriminative generative adversarial network is used to perform feature discrimination in multiple dimensions such as phonemes, prosody, speakers, and environment, and a generator guided by audio-visual features is used to significantly improve the precision and naturalness of speech separation. In a high-noise environment (signal-to-noise ratio <-5dB), the signal-to-noise ratio can be improved by 9-12dB compared with traditional methods.
[0066] 3. Real-time evaluation of the environment state and adaptive adjustment of the parameters are realized, so that the system can quickly adapt to the dynamic changes of the acoustic environment, and in a rapidly changing environment (such as a moving scene), the adaptation speed is improved by 3 times and the steady-state performance is improved by 40%.
[0067] 4. A multi-scale fusion optimization strategy and a dynamic computing resource allocation mechanism are used to significantly reduce the computational complexity while ensuring processing performance. Compared with traditional deep learning methods, the computational complexity is reduced by 35% and the power consumption is reduced by 42% under the same performance. BRIEF DESCRIPTION OF DRAWINGS
[0068] Figure 1 The method flowchart of the present application;
[0069] Figure 2 The overall architecture schematic diagram of the multi-modal perception intelligent microphone array signal processing system of the present application. Detailed Implementation
[0070] Please refer to Figure 1 - Figure 2 The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the scope of the invention.
[0071] Reference Figure 2 The multimodal sensing intelligent microphone array signal processing system 100 provided by the present invention includes a multimodal sensing layer 110, a feature processing layer 120, a multimodal fusion layer 130, an intelligent separation layer 140, an environmental adaptation layer 150, and an output processing layer 160.
[0072] The multimodal perception layer 110 includes a multi-directional visual sensor 111 and a microphone array 112 for acquiring visual and acoustic signals. The feature processing layer 120 includes a visual feature processing module 121 and an acoustic feature processing module 122 for extracting visual and acoustic features. The multimodal fusion layer 130 includes a topology feature mapping engine 131 and a probabilistic manifold modeling unit 132 for constructing an audiovisual topology feature space and establishing a sound source probability distribution model. The intelligent separation layer 140 includes a multidimensional generative adversarial network 141 and a dynamic equalization controller 142 for separating the target speech signal. The environment adaptation layer 150 includes an environment state evaluator 151 and a parameter adaptive adjuster 152 for evaluating the environment state and dynamically adjusting processing parameters. The output processing layer 160 includes a quality evaluation unit 161 and a voice synthesizer 162 for evaluating speech quality and generating the final output signal.
[0073] Preferably, the multimodal perception layer 110 is responsible for collecting raw audiovisual data, the feature processing layer 120 extracts feature representations of each modality, the multimodal fusion layer 130 achieves deep fusion of audiovisual information, the intelligent separation layer 140 performs speech separation based on the fused features, the environmental adaptation layer 150 enables the system to adaptively adjust to environmental changes, and the output processing layer 160 is responsible for quality assessment and final output generation. These six layers form a complete processing pipeline, achieving multimodal intelligent speech processing through collaborative work.
[0074] Please refer to Figure 1 According to an embodiment of the present invention, a signal processing method for a multimodal sensing smart microphone array is provided, the detailed process of which is as follows:
[0075] The method of the present invention first uses a multi-directional visual sensor to acquire visual signals and a microphone array to acquire sound signals.
[0076] In a preferred embodiment of the present application, the multi-directional visual sensor is arranged in front, back, left, right, up and down six directions to realize 360° panoramic visual information collection. The sampling frequency of the visual sensor is 30-60 frames per second, and the resolution can be dynamically adjusted, with a maximum support of 1080p. After visual data collection, preprocessing is performed, including distortion removal, light equalization and background segmentation, etc. to improve the accuracy of subsequent processing.
[0077] The microphone array can be configured in linear, ring or spherical geometry, and can be flexibly selected according to specific application scenarios. For example, in a conference room scenario, a ring-shaped microphone array is preferred to realize omnidirectional sound field collection; while in a mobile device, a linear microphone array is preferred to balance performance and integration. The microphone sampling frequency can be set to 16kHz or 48kHz, and the quantization accuracy is 24 bits to ensure the capture of high-quality sound signals. The acoustic data preprocessing includes pre-filtering, DC offset elimination and gain balancing operations.
[0078] The present application uses a hardware timestamp mechanism based on the Precision Time Protocol (PTP) to realize time synchronization of audio-visual data, ensuring synchronization accuracy better than 1 millisecond. The system maintains a sliding time window, and the window length is preferably set to 100-500 milliseconds to balance processing delay and system responsiveness. For the sampling rate difference between audio and visual modalities, asynchronous processing and resampling techniques are used to handle the time alignment of the data.
[0079] The method of the present application extracts visual features of visual signals and acoustic features of sound signals, wherein the visual features include visual orientation features, face features and lip shape features, and the acoustic features include time-frequency features and voiceprint features.
[0080] The visual feature extraction adopts a hierarchical processing strategy. First, the orientation feature extraction calculates the accurate spatial position of the speaker through multi-camera triangulation technology to generate a three-dimensional orientation vector describing the speaker relative to the microphone array, which is dynamically updated at a frequency of 30Hz to realize real-time tracking of the speaker's movement. Second, the face feature extraction uses a multi-scale feature pyramid technique to extract face features at different resolutions, calculates the pitch, yaw and roll angles of the head, and generates a visual identity embedding vector of the speaker. Finally, the lip shape feature extraction accurately segments the lip contour through an adaptive region growing algorithm, captures the time sequence features of the lip shape changes, and maps the lip shape changes to the visual place classification in phonetics.
[0081] Acoustic feature extraction employs a multi-dimensional processing method. Time-frequency feature extraction simultaneously performs short-time Fourier transform and wavelet packet transform to extract the modulation characteristics and spectral envelope of the sound signal. The parameters for the short-time Fourier transform are set as follows: frame length 25 milliseconds, frame shift 10 milliseconds, and a Hanning window function. Voiceprint feature extraction primarily analyzes the speaker's vocal tract characteristics, prosodic features (pitch, intensity, speech rate, etc.), and suprasegmental features to construct a compact speaker identity representation. Furthermore, environmental acoustic features are extracted, including room impulse response estimation, background noise statistical characteristics analysis, and sound source direction estimation.
[0082] The extracted raw features are subjected to noise reduction, enhancement, and normalization to improve feature quality and consistency. Specifically, for visual features, spatial filtering and temporal smoothing techniques are used to reduce jitter; for acoustic features, spectral subtraction and adaptive normalization techniques are used to enhance signal quality. The final extracted visual and acoustic features have dimensions of 64-128 and 128-256, respectively.
[0083] The method of this invention constructs an audiovisual topological feature space, maps visual and acoustic features to the audiovisual topological feature space, and establishes a sound source probability distribution model based on the audiovisual topological feature space.
[0084] In one embodiment of the present invention, the audiovisual topology feature space is a joint feature space of 128 to 256 dimensions. This feature space uses a Riemann distance metric, defined as follows:
[0085] ,
[0086] in: For two points in the feature space and The Riemann distance between them; and Let be two points in the feature space, namely... dimensional vector; and Points and points The Dimensional components; is the dimension of the feature space, with values ranging from 128 to 256; For the first The weight factor of a feature dimension reflects the importance of that feature dimension. The initial value is set to 1, and then dynamically adjusted according to the feature discrimination. The range is usually from 0.5 to 2.0.
[0087] The process of mapping visual and acoustic features to a joint feature space employs a nonlinear dimensionality reduction technique. First, the original feature vectors are preprocessed to ensure consistent scale across different modalities. Then, the features are projected onto the common space using a mapping function.
[0088] ,
[0089] wherein: is the original visual feature vector with dimension 64-128; is the original acoustic feature vector with dimension 128-256; and are the mapping functions for visual and acoustic features, respectively; and are the mapped visual and acoustic feature representations with dimension 128-256. The mapping functions are implemented by multi-layer perceptron with typical structure containing 3-5 hidden layers, each layer with 256-512 neurons and ReLU (Rectified Linear Unit) as activation function.
[0090] The feature distance metric is optimized by contrastive learning method to enhance the intra-class cohesion and inter-class separability. The contrastive loss function is defined as:
[0091] ,
[0092] wherein: is the contrastive loss function; is the cosine similarity function, which calculates the cosine value of the included angle between two vectors, ranging from -1 to 1; is the temperature parameter, which controls the smoothness of the distribution, usually set to 0.07; represents the acoustic feature of the th negative sample; is the number of negative samples, usually set to 64-128; is the natural exponential function.
[0093] Finally, the precise correspondence between visual and acoustic features is established, and the mutual information between audio-visual features is maximized. The mutual information maximization objective function is:
[0094] ,
[0095] wherein: is the mutual information maximization loss function; represents the mutual information between the visual feature and the acoustic feature ; represents the entropy of the visual feature , which measures its uncertainty; represents the conditional entropy of the visual feature given the acoustic feature . In actual implementation, the mutual information is approximated by a mutual information neural estimator.
[0096] Through the above process, a high-dimensional feature space maintaining the topological relationship of audio-visual features is constructed, and on this basis, a sound source probability distribution model is established, which provides strong prior information for subsequent speech separation.
[0097] The method of the application adopts a multi-dimensional discriminative generative adversarial network to process a sound signal and separate a target speech signal.
[0098] The multi-dimensional discriminative generative adversarial network of the application comprises a multi-level discriminator structure and a generator based on audio-visual feature guidance. The multi-level discriminator structure comprises a phoneme-level discriminator, a prosody-level discriminator, a speaker-level discriminator and an environment-level discriminator, which are respectively responsible for evaluating different dimensions of speech characteristics.
[0099] The phoneme-level discriminator focuses on the authenticity discrimination of a short-time speech segment (10-30 milliseconds) and mainly evaluates the local acoustic characteristics of the speech. The prosody-level discriminator evaluates the naturalness of speech rhythm and intonation in a relatively long time period (100-500 milliseconds) and focuses on suprasegmental features. The speaker-level discriminator verifies the consistency of the separated speech and the speaker features, ensuring the preservation of the speech identity characteristics. The environment-level discriminator ensures the coordination of the speech and the current acoustic environment, avoiding the unnatural feeling caused by the mismatch of environmental characteristics.
[0100] The discriminator extracts discriminative features on multiple time scales and adopts an attention mechanism to dynamically focus on the most discriminative feature area. The calculation formula of the attention mechanism is as follows:
[0101]
[0102] Among them: is the output matrix of the attention mechanism; is the query matrix, the dimension is ; is the key matrix, the dimension is ; is the value matrix, the dimension is ; is the length of the query sequence; is the length of the key / value sequence; , and are the dimensions of the query, key and value vectors, usually set to ; represents the multiplication of matrix and the transpose of matrix , and the result dimension is ; is a scaling factor used to prevent the softmax gradient from disappearing due to the large inner product value; is a normalization function that converts each row vector into a probability distribution. The discriminator also adopts an adaptive threshold adjustment technique to dynamically adjust the discrimination threshold according to the environmental complexity, quantify the reliability of the discrimination result, and weight and fuse the multi-layer discrimination results according to the reliability.
[0103] The generator based on audio-visual feature guidance adopts a U-shaped network structure, including multi-stage jump connections and multi-path processing flows. The network depth is 8-12 layers, and the number of channels starts from 64, reaching 512 in the deepest layer. The jump connection adopts a residual connection mode to alleviate the gradient disappearance problem and retain detailed information. The generator guides the sound separation process through visual condition embedding and voiceprint condition control. The condition embedding is realized through a FiLM (Feature Intelligent Linear Modulation) layer:
[0104] ,
[0105] ,
[0106] wherein: is a scaling parameter vector, with the same dimension as the feature vector ; is an offset parameter vector, with the same dimension as the feature vector ; is a condition vector (visual feature or voiceprint feature), with a dimension of 128-256; and are mapping functions, usually single-layer or double-layer fully connected networks; is an input feature vector; is a modulated feature vector; represents element-wise multiplication (Hadamard product). and control the scaling and offset of the features, respectively, to realize the modulation of the condition information on the features.
[0107] The processing process of the multi-dimensional discriminative adversarial generative network includes: decomposing the sound signal into multiple time-frequency components, assigning each time-frequency component to different sound sources based on audio-visual features, constructing a soft masking matrix for each sound source, and eliminating residual cross interference through post-processing. The masking matrix generation formula is:
[0108] ,
[0109] wherein: is the masking value of the th sound source at time and frequency , ranging from 0 to 1; is the estimated th sound source at time and frequency complex spectrum; representing the modulus square of complex number , i.e. power spectrum; total number of sound sources; representing the summation of all sound sources; a small positive number, usually set as , to prevent numerical instability caused by zero denominator. Residual cross-interference cancellation adopts the method combining spectral subtraction and adaptive filtering.
[0110] Through the processing of the multi-dimensional discriminative generative adversarial network, the system can effectively separate the target speech in the mixed speech signal while maintaining the naturalness and identity characteristics of the speech.
[0111] The method of the present application evaluates the acoustic environment state in real time and dynamically adjusts the processing parameters of the multi-dimensional discriminative generative adversarial network.
[0112] In the embodiments of the present application, the environment state evaluation includes extracting acoustic environment parameters such as reverberation time (RT60), early reflection density, background noise level, and identifying sound source distribution characteristics such as sound source number, spatial distribution and motion state. The reverberation time estimation adopts the energy decay method:
[0113] ,
[0114] wherein: is the reverberation time, with the unit of seconds (s), representing the time required for sound energy to decay by 60 dB; is the slope of the room impulse response energy decay curve, with the unit of dB / s (decibel / second), which is estimated from the energy decay curve by linear regression method; represents the number of decibels that need to be decayed. The early reflection density is estimated by the number of reflection peaks within 50-80 milliseconds before the impulse response. The background noise level is estimated by the minimum statistical method:
[0115] ,
[0116] wherein: is the noise power spectrum estimate at frequency ; is the short-time Fourier transform coefficient of the observation signal at time and frequency ; represents the modulus square of complex number , i.e. power spectrum; represents the minimum value within the time window ; is the length of the time window, usually set as 1-2 seconds, corresponding to 100-200 frames (assuming a frame shift of 10 milliseconds).
[0117] The system matches the current environment with predefined environment categories, generating an environment feature vector. The predefined environment categories include quiet indoor (background noise < 30 dB, RT60 < 0.3 s), general indoor (background noise 30-50 dB, RT60 0.3-0.8 s), noisy indoor (background noise > 50 dB, RT60 0.3-0.8 s), outdoor (background noise varies greatly, RT60 approaches 0), and moving scene (parameters change rapidly).
[0118] The environment dynamic change detection uses a sliding window analysis technique to analyze the environment parameters over a continuous time window (typical value: 0.5-2 seconds) and identify significant change points in the environment parameters. The change point detection algorithm is based on the CUSUM (Cumulative Sum) method:
[0119] ,
[0120] where: is the cumulative sum statistic at the current time ; is the cumulative sum statistic at the previous time , with an initial value ; is the observation value (such as background noise level or RT60) at time ; is the target mean value, i.e., the expected parameter steady-state value; is the sensitivity parameter, which controls the detection sensitivity, usually set to 0.5; takes the maximum value between 0 and the parameter, ensuring that the cumulative sum is non-negative. When exceeds the preset threshold (typical value: 5), it is determined as a change point.
[0121] According to the environment state, the system dynamically adjusts the filter parameters, gain control parameters, and network weight parameters. The parameter adjustment uses a continuous parameter adjustment mechanism to ensure the continuity of the parameter adjustment process and avoid sound artifacts caused by sudden changes. The mathematical expression of the smooth transition strategy is:
[0122] ,
[0123] where: is the parameter value at the current time ; is the parameter value at the previous time ; is the target parameter value, determined according to the current environment state; This is a smoothing factor that controls the update speed of parameters. Its value range is [0,1], and it is usually set to 0.05-0.2. It is dynamically adjusted according to the rate of environmental change, taking a larger value when the environment changes rapidly and a smaller value when the environment is stable.
[0124] The system also implements environmental feedback closed-loop control. Through performance evaluation indicators such as signal-to-noise ratio improvement, speech clarity, and artifact reduction, it adjusts parameters accordingly to achieve self-optimization. The objective function of the closed-loop control is:
[0125] ,
[0126] in: For parameters The overall performance score is as follows: the higher the score, the better the performance. For using parameters The signal-to-noise ratio (SNR) evaluation function is obtained by normalizing the SNR gain, with a range of [0,1]. For using parameters The sharpness evaluation function at that time has a value range of [0,1] and is obtained by normalizing the CSII or PESQ scores. For using parameters The artifact severity evaluation function is defined with a range of [0,1]. A larger value indicates a more severe artifact. , and The weighting coefficients control the importance of each indicator and satisfy the following requirements: Typical values are 0.4, 0.4 and 0.2, respectively.
[0127] Through the above environmental adaptation mechanism, the system can perceive environmental changes in real time and make corresponding adjustments to maintain stable performance in dynamic environments.
[0128] The method of this invention performs quality assessment on the separated multi-channel target speech signals and selects the highest quality speech signal as the output.
[0129] This invention establishes a comprehensive speech quality evaluation index system, including objective indicators, perceptual indicators, and a comprehensive quality score. Objective indicators include signal-to-noise ratio, speech clarity index, and cepstral distance, calculated using the following formulas:
[0130] Signal-to-noise ratio (SNR) calculation:
[0131] ,
[0132] in: Signal-to-noise ratio, in decibels (dB). The first clean speech signal Amplitude of each sampling point; The estimated speech signal's amplitude at the th sample point; The sum of all sample points ; The square of the signal , representing the signal energy; The square of the estimation error, representing the noise energy; The logarithm function with base 10.
[0133] The calculation of the speech intelligibility index (CSII):
[0134] ,
[0135] Where: The speech intelligibility index, with a value range of [0, 1], the larger the value, the higher the intelligibility; The short-time Fourier transform coefficient of the original signal at time frame and frequency ; The short-time Fourier transform coefficient of the estimated signal at time frame and frequency ; The complex conjugate of ; And represent the modulus square of and , respectively; The sum of all time frames and frequency points .
[0136] The calculation of the cepstral distance (CD):
[0137] ,
[0138] Where: The cepstral distance, with units of dB; The original speech's th cepstral coefficient; The estimated speech's th cepstral coefficient; The dimension of the cepstral coefficient, usually 13; The sum of all cepstral coefficients; The natural logarithm acting on 10, approximately 2.303; The coefficient for converting the natural logarithm scale to the decibel scale, approximately 4.343.
[0139] The perceptual metrics are obtained by subjective evaluation methods such as Mean Opinion Score (MOS) and Perceptual Speech Quality Measure (PSQM). In the actual system, these subjective scores are predicted by training models.
[0140] The comprehensive quality score is calculated by weighting:
[0141]
[0142] is the comprehensive quality score, with a value range of [0, 1], and a larger value indicates higher quality; and represent the signal-to-noise ratio, speech articulation index, cepstral distance, and subjective average score normalized to the [0, 1] interval, respectively; The cepstral distance is converted to a positive indicator, with a larger value indicating higher quality; are weight coefficients, satisfying The empirical values are 0.25, 0.3, 0.2, and 0.25, respectively.
[0143] The system compares the quality of the separation results of different microphone channels and selects the channel with the highest quality score as the output. To further improve the output quality, the system performs customized enhancement processing on the selected speech signal, including spectral smoothing, harmonic structure enhancement, and sub-band gain adjustment.
[0144] Spectral smoothing uses a time-frequency filter:
[0145]
[0146] is the smoothed spectrum, with a value at time and frequency ; is the original spectrum, with a value at time and frequency ; is the filter weight, representing the weight coefficient at time offset and frequency offset ; represents summation within the time window and frequency window ; and are the half-widths of the time and frequency windows, respectively, with typical values of 1 and 2, corresponding to a 3x5 filter window; the denominator term is used for normalization to ensure that the filter weight sum is 1.
[0147] Harmonic structure enhancement is achieved by strengthening the harmonic peaks of the speech signal:
[0148] ,
[0149] wherein: is the enhanced spectrum, the value at time t and frequency f; is the original spectrum, the value at time t and frequency f; is the harmonic saliency map, the value domain is [0, 1], the larger the value, the more significant the harmonic characteristics of the time-frequency point; is the enhancement factor, which controls the enhancement intensity, usually set to 0.2-0.5; is the enhancement coefficient, the range is [1, 1+α]. The sub-band gain adjustment adjusts the gain of each frequency band according to the importance of the speech content:
[0150] ,
[0151] wherein: is the final output spectrum, the value at time t and frequency f; is the original spectrum, the value at time t and frequency f; is the gain factor of frequency f, which is set according to the importance of the frequency band. The typical gain setting is: low frequency band (<300Hz) gain 0.8, speech main frequency band (300-3400Hz) gain 1.2, high frequency band (>3400Hz) gain 0.9.
[0152] Through the above speech quality evaluation and enhancement processing, the system outputs high-quality separated speech signals, meeting the needs of various application scenarios.
[0153] The method of the present application also includes measures such as selecting a preset configuration according to the application scenario, dynamically allocating computing resources, monitoring the output quality in real time, and continuously learning the environmental characteristics, to further improve the practicality of the system.
[0154] The present application selects a preset configuration according to the application scenario, including conference scenario configuration, mobile scenario configuration and noise environment configuration, etc. The conference scenario configuration optimizes the multi-person alternating speaking scenario, especially strengthens the speaker switching detection and identity maintenance; the mobile scenario configuration adapts to the environmental changes in motion, enhances the environmental adaptation speed; the noise environment configuration enhances the performance in extreme noise environment, improves the anti-noise ability.
[0155] The system realizes dynamic computing resource allocation, allocating more resources to the key modules in the processing chain. The resource allocation strategy is based on task priority and real-time requirements, and an adaptive algorithm is used to adjust the computing resource quota of each module. For example, in a high-noise environment, increase the resource quota of the noise suppression module; in a multi-speaker scenario, increase the resource quota of the speech separation module.
[0156] The system continuously monitors the quality indicators of the output speech, and triggers corrective measures when the quality drops below a predetermined threshold (usually set to 20%). The corrective measures include parameter readjustment, backup processing chain activation, and user feedback prompts. In addition, the system also continuously learns the environmental characteristics, optimizes the parameter configuration for specific environments, and improves the long-term performance.
[0157] Through the above optimization measures, the system maintains high performance while achieving good practicality and user experience.
[0158] The multi-modal perception intelligent microphone array signal processing system provided by the application includes a multi-modal perception layer, a feature processing layer, a multi-modal fusion layer, an intelligent separation layer, an environment adaptation layer, and an output processing layer.
[0159] The multi-modal perception layer includes a multi-directional vision sensor and a microphone array for acquiring visual and acoustic signals. The feature processing layer includes a visual feature processing module and an acoustic feature processing module for extracting visual and acoustic features. The multi-modal fusion layer includes a topological feature mapping engine and a probabilistic manifold modeling unit for constructing an audio-visual topological feature space and establishing a sound source probability distribution model. The intelligent separation layer includes a multi-dimensional generative adversarial network and a dynamic balancing controller for separating target speech signals. The environment adaptation layer includes an environment state evaluator and a parameter self-adaptive adjuster for evaluating the environment state and dynamically adjusting the processing parameters. The output processing layer includes a quality evaluation unit and a sound synthesizer for evaluating the speech quality and generating the final output signal.
[0160] The system realizes multi-modal perception intelligent microphone array signal processing through the collaborative work of the multi-modal perception layer, the feature processing layer, the multi-modal fusion layer, the intelligent separation layer, the environment adaptation layer, and the output processing layer. The modular design of the system architecture makes it have good scalability and maintainability, and can adapt to the needs of different application scenarios.
[0161] The hardware implementation of the system can be based on a heterogeneous computing platform of CPU+GPU / DSP, supporting resource-aware scheduling and power consumption control. The software architecture adopts a pipeline processing mode, supporting real-time processing and low-latency response. The system provides standardized API interfaces, facilitating integration with other systems and secondary development.
[0162] The multi-modal perception intelligent microphone array signal processing method and system of the application are applicable to various application scenarios, including intelligent conference systems, intelligent vehicle systems, smart home devices, and mobile intelligent terminals.
[0163] In the intelligent conference system, the application can realize the functions of multi-person conference voice accurate separation, remote participant voice enhancement and conference record automatic transcription. Tests show that in the complex conference scene of 6 people speaking at the same time, the system can still maintain more than 85% of the speech intelligibility, which is more than 40% higher than the traditional method.
[0164] In the intelligent vehicle system, the application can realize the functions of driver voice command recognition, multi-passenger interaction management and road noise suppression. Tests show that under the condition of high-speed driving (vehicle noise about 75dB), the voice recognition accuracy of the system can still maintain more than 90%, which is more than 30% higher than the traditional method.
[0165] In the intelligent home device, the application can realize the functions of far-field voice interaction, multi-user voice distinction and voice recognition under the noise of household appliances. Tests show that under the condition of 5 meters away and TV volume of 60dB, the voice wake-up success rate of the system can still maintain more than 95%, which is more than 25% higher than the traditional method.
[0166] In the mobile intelligent terminal, the application can realize the functions of handheld device call quality improvement, video conference voice enhancement and noisy environment voice memo. Tests show that in the street environment (environmental noise about 70dB), the voice call quality score (MOS) of the system reaches 3.8, which is 0.8 points higher than the traditional method.
[0167] Overall, the application has achieved significant improvement in voice quality indicators: the signal-to-noise ratio is improved by an average of 15-20dB, the speech intelligibility (PESQ score) is improved by 1.5-2.0, and the speech intelligibility is improved by 40%-60% in noisy environments. In terms of system performance indicators, the application realizes a processing delay of less than 50 milliseconds in a streaming processing mode, can run in real time on a mid-end processor, and can adapt within 1-2 seconds after environmental changes.
[0168] Compared with the prior art, the application improves the separation accuracy by 65% in a multi-speaker scene, improves the performance by 80% in a low signal-to-noise ratio environment, improves the computing efficiency by 40%, and improves the adaptability by 200%. These advantages are mainly due to the multi-modal collaborative effect, environmental adaptation mechanism and efficient algorithm implementation of the application.
[0169] The application provides a multi-modal perception intelligent microphone array signal processing method and system, which realizes high-quality voice separation and enhancement in complex acoustic environments through deep fusion and collaborative processing of audio-visual multi-modal information, combined with a topology enhanced generative adversarial network architecture and an environmental adaptation mechanism. The application has significant advantages in multi-speaker scenes, low signal-to-noise ratio environments and rapidly changing environments, and brings a qualitative leap to intelligent voice interaction technology.
[0170] The main innovations of the application include: audio-visual topology feature space construction and mapping technology, multi-dimensional discriminative generative network architecture, environment adaptive mechanism and multi-scale fusion optimization strategy. These innovations synergize together to solve the key technical problems of speech separation and enhancement in complex environments.
[0171] The application has good practicability and popularization value, and can be widely applied to intelligent conference, intelligent vehicle, smart home and mobile intelligent terminal and various scenes, thereby bringing significant improvement for voice interaction experience.
[0172] The above only describes the preferred embodiments of the application and is not used to limit the application. The application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the application shall be included in the protection scope of the application.
Claims
1. A multimodal sensing intelligent microphone array signal processing method, characterized in that, include: A multi-directional visual sensor is used to acquire visual signals, and a microphone array is used to acquire sound signals; The visual features of the visual signal and the acoustic features of the sound signal are extracted, wherein the visual features include visual orientation features, facial features and lip shape features, and the acoustic features include time-frequency features and voiceprint features; Construct an audiovisual topology feature space, map the visual features and the acoustic features to the audiovisual topology feature space, and establish a sound source probability distribution model based on the audiovisual topology feature space; The sound signal is processed by a multidimensional discriminant adversarial generative network to separate the target speech signal. The multidimensional discriminant adversarial generative network includes a multi-level discriminator structure and a generator guided by audiovisual features. The acoustic environment is assessed in real time, and the processing parameters of the multidimensional discriminant adversarial generative network are dynamically adjusted. The quality of the separated multi-channel target speech signals is evaluated, and the highest quality speech signal is selected as the output. The construction of the audiovisual topological feature space specifically includes: Define a joint feature space of 128 to 256 dimensions, and define the Riemann distance metric in this feature space; The visual features and the acoustic features are projected into the joint feature space through nonlinear dimensionality reduction; By optimizing the feature distance metric through contrastive learning methods, intra-class clustering and inter-class separation are enhanced. Establish a precise correspondence between visual and acoustic features to maximize the mutual information between audiovisual features.
2. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that, The specific methods for acquiring visual signals using a multi-directional visual sensor and acquiring sound signals using a microphone array include: Visual sensors positioned in six directions—front, back, left, right, top, and bottom—are used to collect 360° panoramic visual information. Sound field information is acquired using a microphone array that can be configured to be linear, circular, or spherical in geometry. The visual signal and the sound signal are synchronized in time to ensure a synchronization accuracy better than 1 millisecond.
3. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that, The extraction of visual features from the visual signal and acoustic features from the sound signal specifically includes: The visual signals are subjected to orientation recognition, face detection and lip tracking to extract visual orientation features, facial features and lip features; The sound signal is subjected to time-frequency analysis, voiceprint extraction, and acoustic fingerprint generation to extract time-frequency features and voiceprint features; The extracted visual and acoustic features are subjected to noise reduction, enhancement, and normalization processing.
4. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that, The multidimensional discriminative adversarial generative network includes: Phoneme-level discriminator is used to evaluate the authenticity of short speech segments; A prosodic discriminator is used to assess the naturalness of speech rhythm and intonation; Speaker-level discriminator is used to verify the consistency between the separated speech and speaker features; An environmental discriminator is used to ensure the consistency of speech with the current acoustic environment; The generator based on audiovisual features adopts a U-shaped network structure, which includes multi-level hop connections and multi-path processing flow, and guides the sound separation process through visual conditional embedding and voiceprint conditional control.
5. The multimodal sensing intelligent microphone array signal processing method according to claim 4, characterized in that, The processing steps of the multidimensional discriminative adversarial generative network include: Discriminative features are extracted at multiple time scales, and the most discriminative feature regions are dynamically focused on through an attention mechanism. The discrimination threshold is adaptively adjusted to quantify the credibility of the discrimination results, and the multi-level discrimination results are weighted and fused based on the credibility. The sound signal is decomposed into multiple time-frequency components, and each time-frequency component is assigned to a different sound source based on the audiovisual characteristics; A soft masking matrix is constructed for each sound source, and residual cross-interference is eliminated through post-processing.
6. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that, The real-time assessment of the acoustic environment and the dynamic adjustment of the processing parameters of the multidimensional discriminative adversarial generative network specifically include: Extract acoustic environmental parameters such as reverberation time, early reflection density, and background noise level; Identify the number of sound sources, their spatial distribution, and the characteristics of sound source distribution in motion. The current environment is matched with predefined environment categories to generate an environment feature vector; Analyze environmental parameters over a continuous time window to identify points of significant change in environmental parameters; Adjust filter parameters, gain control parameters, and network weight parameters according to environmental conditions to ensure the continuity of the parameter adjustment process.
7. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that, The process of evaluating the quality of the separated multi-channel target speech signals and selecting the highest quality speech signal as the output specifically includes: Calculate objective metrics such as signal-to-noise ratio, speech intelligibility index, and cepstral distance; The subjective average opinion score and the perceived naturalness of speech were evaluated. A weighted quality score is calculated by combining multiple indicators. Compare the quality of the separation results of different microphone channels and select the channel with the highest quality score; The selected speech signal is subjected to customized enhancement processing, including spectral smoothing, harmonic structure enhancement, and subband gain adjustment.
8. The multimodal sensing intelligent microphone array signal processing method according to claim 1, characterized in that, Also includes: Select the preset configuration based on the application scenario, including meeting scenario configuration, mobile scenario configuration, and noisy environment configuration; Dynamically allocate computing resources, distributing more resources to key modules in the processing chain; Real-time monitoring of output speech quality indicators; triggering corrective measures when the quality deteriorates beyond a predetermined threshold. Continuously learn about environment characteristics and optimize parameter configurations for specific environments.
9. A multimodal sensing intelligent microphone array signal processing system, implementing the multimodal sensing intelligent microphone array signal processing method according to any one of claims 1-8, characterized in that, include: The multimodal perception layer includes a multi-directional visual sensor and a microphone array for acquiring visual and sound signals; The feature processing layer includes a visual feature processing module and an acoustic feature processing module, which are used to extract visual and acoustic features; The multimodal fusion layer, including a topology feature mapping engine and a probabilistic manifold modeling unit, is used to construct an audiovisual topology feature space and establish a sound source probability distribution model. The intelligent separation layer, comprising a multidimensional generative adversarial network and a dynamic equalization controller, is used to separate the target speech signal; The environmental adaptation layer includes an environmental state evaluator and a parameter adaptive adjuster, which are used to evaluate the environmental state and dynamically adjust the processing parameters; The output processing layer includes a quality assessment unit and a speech synthesizer, which are used to assess speech quality and generate the final output signal; The system achieves intelligent microphone array signal processing with multimodal perception through the coordinated operation of the multimodal perception layer, feature processing layer, multimodal fusion layer, intelligent separation layer, environmental adaptation layer, and output processing layer.
Citation Information
Patent Citations
Multi-sound-source separation system and method based on generative adversarial network
CN116312609A
Multi-microphone array beamforming signal enhancement method and device
CN119811408A