Speech enhancement and high-precision recognition method and system in complex environment

By combining microphone arrays and deep neural network models, the problems of poor speech signal quality and low recognition accuracy in complex environments are solved, achieving high-precision speech enhancement and recognition, and improving the clarity and accuracy of voice interaction.

CN121641016APending Publication Date: 2026-03-10INTELLIGENT INTER CONNECTION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Noise interference in complex environments severely affects the clarity and recognition accuracy of speech signals, resulting in poor speech signal quality and low recognition accuracy.

Method used

A microphone array is used for signal acquisition and preprocessing. A deep neural network model is used for noise power spectral density estimation and reverberation parameter analysis. A deep neural network model is constructed for speech masking, and high-precision recognition is achieved through acoustic and language models.

Benefits of technology

It improves the quality of voice signals and recognition accuracy, enabling clear, accurate, and real-time voice interaction, thereby enhancing work efficiency and user experience in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641016A_ABST
    Figure CN121641016A_ABST
Patent Text Reader

Abstract

The invention provides a voice enhancement and high-precision recognition method and system in a complex environment, and relates to the technical field of voice processing, and the method comprises the steps: collecting a time domain signal in an off-road parking sentry box environment for preprocessing, detecting a mute segment signal in a standard time domain signal for noise power spectral density estimation, and obtaining a noise power spectral density value; a reverberation parameter is obtained by combining voice onset information and noise spatial correlation estimation, prediction is performed by using a deep neural network model, voice masking is applied to microphone array signals to perform enhancement processing, adaptive feature extraction is performed on time domain enhanced voice signals, and a voice signal is obtained. And performing high-precision recognition on the voice adaptive feature sequence based on an acoustic model and a language model, and outputting a target recognition text. The technical problems of poor voice signal quality and low recognition accuracy in a complex noise environment in the prior art are solved. The technical effects of improving the voice signal quality and the recognition accuracy and realizing clear, accurate and real-time voice interaction are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, specifically to a method and system for speech enhancement and high-precision recognition in complex environments. Background Technology

[0002] In modern society, voice interaction technology has been widely applied in smart devices, communication systems, public safety, and other fields. However, noise interference in complex environments severely affects the clarity and recognition accuracy of voice signals. In off-street parking booth environments, this is typically accompanied by continuous and high-intensity traffic noise, such as car engine sounds, horns, and tire friction noise, as well as environmental background noise, such as wind noise, conversations, and potential equipment operating noise. This noise can severely mask voice signals, making it difficult for both parties to hear each other clearly, resulting in poor voice signal quality and low recognition accuracy.

[0003] Existing technologies suffer from poor speech signal quality and low recognition accuracy in complex noise environments. Summary of the Invention

[0004] The purpose of this application is to provide a method and system for speech enhancement and high-precision recognition in complex environments, in order to solve the technical problems of poor speech signal quality and low recognition accuracy in complex noise environments in existing technologies.

[0005] In view of the above problems, this application provides a method and system for speech enhancement and high-precision recognition in complex environments.

[0006] The first aspect of this application provides a method for speech enhancement and high-precision recognition in complex environments. The method includes: real-time detection of silence segments in a standard time-domain signal; estimation of the noise power spectral density of the silence segments to obtain a non-stationary noise power spectral density; and simultaneous estimation of reverberation parameters by combining speech onset information and noise spatial correlation. A deep neural network model is constructed, using the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameters as input; the deep neural network model is trained and predicted using noisy speech data from an off-street parking environment to obtain a speech mask. The speech mask is applied to a microphone array signal for enhancement processing to obtain a time-domain enhanced speech signal; adaptive feature extraction is performed on the time-domain enhanced speech signal to obtain a speech adaptive feature sequence. High-precision recognition of the speech adaptive feature sequence is performed based on an acoustic model and a language model to output target recognition text.

[0007] A second aspect of this application provides a speech enhancement and high-precision recognition system for complex environments. The system includes: a signal acquisition module for synchronously acquiring time-domain signals in an off-street parking booth environment using a microphone array, wherein the microphone array includes M microphones; preprocessing the time-domain signals to obtain a standard time-domain signal; a parameter acquisition module for real-time detection of silence segments in the standard time-domain signal; estimating the noise power spectral density of the silence segments to obtain a non-stationary noise power spectral density; and simultaneously obtaining reverberation parameters by combining speech onset information and noise spatial correlation estimation; and a data prediction module. The system is used to construct a deep neural network model, which takes the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameters as inputs. The deep neural network model is trained and predicted using noisy speech data from an off-street parking environment to obtain a speech mask. A feature extraction module applies the speech mask to a microphone array signal for enhancement, obtaining a time-domain enhanced speech signal. Adaptive feature extraction is performed on the time-domain enhanced speech signal to obtain a speech adaptive feature sequence. A text output module performs high-precision recognition of the speech adaptive feature sequence based on an acoustic model and a language model, outputting the target recognition text.

[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0009] The method provided in this application synchronously acquires time-domain signals in an off-street parking booth environment using a microphone array, wherein the microphone array includes M microphones. The time-domain signals are preprocessed to obtain a standard time-domain signal; silence segments in the standard time-domain signal are detected in real time, and noise power spectral density is estimated for the silence segments to obtain a non-stationary noise power spectral density. Simultaneously, reverberation parameters are obtained by combining speech onset information and noise spatial correlation estimation; a deep neural network model is constructed, using the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameters as input representations. The deep neural network model is trained and predicted using noisy speech data from the off-street parking environment to obtain a speech mask; the speech mask is applied to the microphone array signal for enhancement processing to obtain a time-domain enhanced speech signal; adaptive feature extraction is performed on the time-domain enhanced speech signal to obtain a speech adaptive feature sequence; high-precision recognition is performed on the speech adaptive feature sequence based on an acoustic model and a language model to output target recognition text. This achieves the technical effect of improving speech signal quality and recognition accuracy, realizing clear, accurate, and real-time voice interaction, and improving work efficiency and user experience in complex scenarios.

[0010] The above description is merely an overview of the technical solution of this application. To enable a clearer understanding of the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0012] Figure 1 A flowchart illustrating the speech enhancement and high-precision recognition method for complex environments provided in this application.

[0013] Figure 2 A schematic diagram of the structure of the speech enhancement and high-precision recognition system in complex environments provided in this application.

[0014] Figure labeling: Signal acquisition module 11, parameter acquisition module 12, data prediction module 13, feature extraction module 14, text output module 15. Detailed Implementation

[0015] This application provides a method and system for speech enhancement and high-precision recognition in complex environments, addressing the technical problems of poor speech signal quality and low recognition accuracy in existing technologies under complex noise conditions. It achieves the technical effect of improving speech signal quality and recognition accuracy, enabling clear, accurate, and real-time voice interaction, and enhancing work efficiency and user experience in complex scenarios.

[0016] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. It should also be noted that, for ease of description, only the parts related to the present invention are shown in the accompanying drawings, not all of them.

[0017] Example 1, as Figure 1As shown, this application provides a method for speech enhancement and high-precision recognition in complex environments, which includes:

[0018] A microphone array is used to synchronously acquire time-domain signals in an off-street parking booth environment. The microphone array includes M microphones. The time-domain signals are preprocessed to obtain standard time-domain signals.

[0019] Furthermore, obtaining the standard time-domain signal includes: performing low-pass anti-aliasing filtering on the time-domain signal to obtain a usable time-domain signal; performing ADC conversion on the usable time-domain signal to obtain a time-domain digital signal; and performing gain control compensation on the time-domain digital signal to obtain the standard time-domain signal.

[0020] Specifically, a microphone array is used to synchronously acquire time-domain signals in a roadside parking booth environment. The microphone array consists of M reasonably distributed microphones. The M microphones achieve multi-angle acquisition of the sound field through spatial distribution, obtaining multiple time-domain signals containing spatial information. The acquired time-domain signals are denoted as x1(t), x2(t), ..., x M (t), where t is time. A time-domain signal is a signal form that uses time as the independent variable to reflect the changes in the value of a speech signal at different times. It records the original information of sound fluctuations over time, including the entire content of the speech and interference components such as environmental noise. In the off-street parking booth environment, conversations, vehicle noise, and environmental noise are all collected in the form of time-domain signals.

[0021] The acquired time-domain signals undergo preprocessing, which includes low-pass anti-aliasing filtering on multiple acquired time-domain signals. This low-pass anti-aliasing filter limits the signal bandwidth and suppresses spectral components above the Nyquist frequency, thereby preventing aliasing distortion during analog-to-digital conversion. The low-pass anti-aliasing filter employs a finite impulse response or an infinite impulse response, with a cutoff frequency set at half the sampling rate. For example, the cutoff frequency is set to 8kHz or 16kHz, depending on the subsequent sampling rate. The low-pass anti-aliasing filter effectively preserves the main frequency components of the speech, obtaining a usable time-domain signal.

[0022] The usable time-domain signal after low-pass anti-aliasing filtering undergoes ADC conversion with a unified sampling rate, such as 16kHz or 24kHz. The analog voltage signal is converted into a discrete digital signal sequence to obtain a time-domain digital signal. Due to differences in sensitivity among the microphones in the microphone array, and the influence of factors such as temperature, humidity, and installation angle in practice, amplitude inconsistencies occur between channels. Therefore, gain control is applied to the obtained time-domain digital signal to compensate for signal strength differences between different microphones or in different directions, ensuring that the signal amplitude acquired by each microphone is within a suitable range. This avoids signal saturation distortion due to excessive amplitude or signal masking due to insufficient amplitude. Automatic gain control (AGC) can be used for gain control, performing amplitude matching and dynamic range adjustment for each channel signal. AGC detects the amplitude of the time-domain digital signal and compares it with preset thresholds, such as a lower threshold of 0.2V and an upper threshold of 1.5V. It outputs a control signal indicating whether the gain needs to be increased, decreased, or maintained. Based on the comparison result, an adjustment command is generated to achieve gain control compensation and obtain a standard time-domain signal. This standard time-domain signal features noise control, amplitude equalization, and sampling consistency.

[0023] Through efficient noise suppression and reverberation cancellation, the signal-to-noise ratio of the speech signal acquired from the microphone is significantly improved, reducing the interference of noise on the main speech information, making the speech clearer and easier to understand.

[0024] The silent segment signal in the standard time domain signal is detected in real time, and the noise power spectral density of the silent segment signal is estimated to obtain the non-stationary noise power spectral density. At the same time, the reverberation parameters are obtained by combining the speech onset information and noise spatial correlation estimation.

[0025] Furthermore, the non-stationary noise power spectral density is obtained by: converting the silent segment signal to the frequency domain using short-time Fourier transform and performing spectrum calculation to obtain multiple microphone spectra; and using a dynamic weighting strategy to perform PSD estimation on the multiple microphone spectra to obtain the non-stationary noise power spectral density.

[0026] Specifically, methods such as energy thresholding, zero-crossing rate, or short-time autocorrelation function are used to detect silent segments in standard time-domain signals in real time. A silent segment refers to a time interval in the speech information where invalid speech is emitted and the microphone signal contains only ambient noise. For example, when the energy of several consecutive frames is below a set threshold and the spectrum changes stably, it is determined to be a silent segment. This detection enables real-time differentiation between speech and noise intervals, ensuring that noise is based solely on pure noise signals and avoiding estimation bias caused by speech leakage.

[0027] After detecting the silence segment signal, the noise power spectral density of the silence segment signal is estimated: First, the time-frequency signal of the silence segment signal is converted into the frequency domain using a short-time Fourier transform. For each microphone i in the microphone array, its spectrum X is calculated in time frame n. i( (k, n), where k is the frequency point. Multiple microphone spectra are obtained through spectrum calculation. Then, a power spectral density (PSD) estimation method using a dynamic weighting strategy is employed to estimate the PSD of the multiple microphone spectra. For example, a correlation algorithm is used to perform correlation analysis on the spectra of multiple microphones, calculating the correlation coefficients of different microphones at different time and frequency points. The correlation coefficient reflects the similarity between the spectra of two microphones; the larger the correlation coefficient, the more similar the noise signals received by the two microphones at that time and frequency point. Based on the correlation analysis results and combined with the statistical characteristics of the noise, the weight of each microphone at different time and frequency points is calculated. The weight calculation is, for example, based on the minimum mean square error criterion, using an optimization algorithm to minimize the error between the weighted spectrum and the true noise power spectral density. Using the calculated weights, the spectral contributions of different channels are weighted, the spectrum of each microphone is multiplied by its corresponding weight, and all weighted spectra are summed to obtain the non-stationary noise power spectral density.

[0028] Simultaneously, reverberation parameter estimation is performed, including spatial correlation analysis using speech onset information and noise. Onset information refers to the correlation characteristics at the start of the speech signal, used to determine the time boundary of speech and noise alternation. Near the speech onset, the relative proportion of direct sound and early reflections is relatively high. Short-time energy and short-time zero-crossing analysis are used to extract onset information. For example, by setting thresholds, when the short-time energy and short-time zero-crossing rate exceed the threshold, it is determined as the start of speech, and the onset information is extracted as zero. Spatial correlation reflects the similarity of noise between different microphones. The spatial correlation of noise between microphones is used to estimate reverberation parameters. The cross-correlation function reflects the similarity of two signals at different time delays. By analyzing the peak position and amplitude of the auto-cross-correlation function, the similarity between noise signals from different microphones is obtained, and the spatial correlation of noise is estimated. For example, when two microphones are far apart, the received noise signals have a strong correlation. By integrating the speech onset information and the spatial correlation estimation results of noise, the reverberation parameters are obtained by using the maximum likelihood estimation or minimum mean square error estimation method. The reverberation parameters include the reverberation time, such as a rough estimate of the reverberation time T60, which reflects the time required for the sound energy to decay to 1 / 6000 of the initial energy.

[0029] By analyzing and processing the silent segment signals in the standard time domain signal, and combining the speech onset information and noise spatial correlation estimation to obtain reverberation parameters, the extracted speech features are corrected and compensated, thereby improving the accuracy and reliability of speech recognition and ensuring accurate recognition of speech commands in complex environments such as roadside parking booths.

[0030] A deep neural network model is constructed, and the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameters are used as input representations. The deep neural network model is trained and predicted using noisy speech data from off-street parking environments to obtain speech masking.

[0031] Furthermore, the deep neural network model adopts a deep complex value neural network or a U-Net variant structure for time-frequency domain masking prediction, and includes multiple convolutional layers, pooling layers, residual connections, and gated linear units.

[0032] Specifically, a deep neural network model is constructed for the accurate separation and masking prediction of speech and noise components. This deep neural network model employs a deep complex value neural network or a variant of U-Net for time-frequency domain masking prediction. The deep complex value neural network is a neural network structure capable of simultaneously analyzing both amplitude and phase information of the speech signal. Traditional real-valued networks, by processing only the amplitude spectrum and ignoring phase features, limit the enhancement of speech naturalness and spatial consistency. Deep complex neural networks introduce complex number operations, maintaining complex form in operations such as convolution, normalization, and activation. Structurally, the deep complex value neural network contains multiple convolutional and pooling layers. The input and output of each layer consist of real and imaginary parts, and the complex convolution operation follows the rules of complex multiplication. The real part contains similarity information between the speech and noise and the cosine reference signal at different frequencies, while the imaginary part provides supplementary information about the signal's phase. Deep complex value neural networks also introduce residual connections to alleviate the gradient decay problem in deep networks. At the same time, gated linear units are embedded in each convolutional block to control the information flow path and realize dynamic filtering of feature channels.

[0033] The U-Net variant for time-frequency domain masking prediction is based on an encoder-decoder symmetric framework. Structurally, the encoder consists of multiple layers of two-dimensional convolutional and pooling layers. The convolutional layers extract local correlation features of speech in the time and frequency dimensions, while the pooling layers progressively reduce the time-frequency resolution and capture global noise. Features from the encoder's end are fed into the decoder via residual connections. The decoder progressively recovers the time-frequency resolution using deconvolutional or upsampling layers. During decoding, skip connections are used to directly concatenate features from each encoder layer with the decoder layer, achieving a fusion of global structural information and local details, avoiding over-smoothing or loss of speech details in the enhancement structure. Gated linear units or residual blocks are introduced into each convolutional module to achieve dynamic weighting and information preservation between channels.

[0034] The standard time-domain signal, non-stationary noise power density, and reverberation parameters are used as joint inputs, and the deep neural network model is trained end-to-end using noisy speech data from off-street parking environments. In the training structure, the input is noisy speech data from off-street parking environments, and the target is the corresponding noiseless speech. The network weights are continuously optimized through backpropagation, enabling the deep neural network model to learn speech masking patterns under different noise types and reverberation conditions. The trained deep neural network analyzes the joint input, and the output layer predicts the speech masking for each frequency point. The speech masking refers to a matrix used to distinguish between speech signals and noise / reverberation signals. The matrix elements represent the energy ratio of the speech signal and noise signal at a specific time and frequency point, ranging from 0 to 1, and are used to suppress noise components and preserve speech components. A speech masking value close to 1 indicates that the time and frequency point is mainly composed of speech components and should be preserved; a value close to 0 indicates that the time and frequency point is mainly composed of noise or reverberation and should be suppressed.

[0035] By inputting multi-dimensional features and analyzing them through deep neural networks, the accuracy of speech analysis in complex environments is improved, ensuring accurate and efficient speech interaction and management in complex environments.

[0036] Furthermore, obtaining speech masking includes: performing a short-time Fourier transform on the standard time-domain signal to obtain a complex spectrum; stacking the spectra of the multiple microphones on time frames and frequency points to form a multi-channel frequency domain input feature map; selecting a model composite loss function, using the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameter as input representations, and combining the model composite loss function with noisy speech data from off-street parking environments to train and predict the deep neural network model to obtain speech masking.

[0037] Specifically, a short-time Fourier transform (SFT) is performed on the standard time-domain signal. The SFT decomposes the continuous time-domain signal into a series of time frames and frequency points, while preserving amplitude and phase information, converting it into a complex spectrum. This complex spectrum contains the amplitude and phase information of the signal at different times and frequencies. The complex spectra of M microphones are stacked at time frames and frequency points. For each time frame, the amplitude and phase spectral values ​​of each frequency point of different microphones within that time frame are arranged sequentially according to the microphone channel order to form a time frame vector. Then, the time frame vectors corresponding to all time frames are concatenated in chronological order to form a multi-channel frequency domain input feature map. This multi-channel frequency domain input feature map contains the amplitude and phase information of the different channel speech signals at different times, as well as spatial feature information.

[0038] The loss function employs a composite loss function combining amplitude and phase errors, such as generalized likelihood ratio (GLR) loss or improved short-time signal-to-noise ratio loss, or incorporates constraints on speech naturalness, such as Mel-spectrum-based loss, into the loss function. Using this composite loss function, the deep neural network is trained end-to-end with a large amount of noisy speech data recorded in a similar off-street parking environment. The composite loss function measures the difference between the deep neural network's predictions and the actual results, improving the performance of the deep neural network model. Furthermore, using the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameters as inputs, the deep neural network continuously adjusts its parameters based on the learned relationship between the input representation and the speech mask, minimizing the composite loss function to accurately predict and output the speech mask.

[0039] By leveraging the powerful learning capabilities of deep neural network models, speech masking can be accurately predicted, providing precise metrics for speech enhancement processing. This enables effective analysis of speech and noise, significantly improving speech quality and achieving clear, accurate, and real-time speech interaction, thereby enhancing work efficiency and user experience in complex scenarios.

[0040] The speech masking is applied to the microphone array signal for enhancement processing to obtain a time-domain enhanced speech signal. Adaptive feature extraction is then performed on the time-domain enhanced speech signal to obtain a speech adaptive feature sequence.

[0041] Furthermore, the speech adaptive feature sequence is obtained by: calculating the Mel-frequency cepstral coefficients of the time-domain enhanced speech signal to obtain the basic speech MFCC features; introducing adaptive feature mapping technology in combination with the non-stationary noise power spectral density to perform adaptive feature extraction on the basic speech MFCC features to obtain the speech adaptive feature sequence.

[0042] Specifically, after obtaining the speech mask M(k,n), it is applied to the microphone array signal for enhancement. For example, to apply the speech mask to the microphone array signal, firstly, beamforming is used to merge the multi-channel signals into a single enhanced signal Y(k,n) pointing towards the sound source. Then, the enhanced spectrum Ẽ(k,n) is obtained by calculating Ẽ(k,n) = M(k,n) × Y(k,n). An inverse short-time Fourier transform is performed on the enhanced spectrum, and the time-domain enhanced speech signal is reconstructed from the frequency domain signal using an overlap-addition method. Specifically, the time-domain enhanced speech signal undergoes pre-emphasis processing to enhance the high-frequency components of the signal to compensate for the attenuation of high-frequency components during transmission. Then, the pre-emphasized signal is divided into short-time frames, and a window function, such as a Hamming window, is added to both ends of each frame to reduce spectral leakage. A short-time Fourier transform is performed on each frame to obtain the enhanced spectrum. The parameter settings include: the frame length of the inverse short-time Fourier transform is usually 256 or 512 points, the frame shift is half the frame length, and there is 50% overlap. The resulting time-domain enhanced speech signal is closer to real speech in amplitude and phase, while significantly reducing environmental noise interference.

[0043] An adaptive feature extraction method is used to process the time-domain enhanced speech signal. First, Mel-frequency cepstral coefficients (MFCCs) are calculated. These MFCCs simulate the frequency perception characteristics of the human ear by filtering, logarithmically compressing, and performing discrete cosine transforms on the Mel-frequency scale to obtain the fundamental MFCC features reflecting the timbre and resonance characteristics of the speech. The MFCC calculation process is as follows: The time-domain enhanced speech signal is filtered through a Mel-filter bank. The Mel-filters are uniformly distributed on the Mel-frequency scale, simulating the sensitivity of the human ear to different frequencies. The logarithmic energy of the output of each filter is taken to obtain the logarithmic Mel spectrum. Finally, a discrete cosine transform is performed on the logarithmic Mel spectrum to obtain the MFCC coefficients. The Mel-filter bank typically has 26 components. The 12th-order cepstral coefficients are retained, and their energy / logarithmic energy is added. The first and second-order differences are then calculated, resulting in a total of 39 features.

[0044] An adaptive feature mapping technique is introduced, combining non-stationary noise power spectral density as an environmental prior with the fundamental speech MFCC features to dynamically adjust the fundamental speech MFCC features. For example, in frequency regions with strong noise, the responding speech MFCC features are weighted and suppressed or normalized to further suppress residual interference, while in relatively stable noise or speech-dominant regions, the original feature information is preserved. Through adaptive processing, an adaptive speech feature sequence is obtained.

[0045] By enhancing the signal and performing adaptive feature extraction, the resulting adaptive speech features can more accurately reflect the essential characteristics of speech, while effectively suppressing noise interference, achieving high-precision speech recognition, and ensuring the accuracy of functions such as instruction transmission and information confirmation.

[0046] The speech adaptive feature sequence is recognized with high precision based on acoustic and language models, and the target recognition text is output.

[0047] Furthermore, the acoustic model adopts an end-to-end connected temporal classification model or an encoder-decoder architecture based on an attention mechanism, wherein the encoder part uses a deep RNN or Transformer to process the speech adaptive feature sequence, the decoder part is used to predict text labels, and the language model integrates a statistical language model or a neural network language model.

[0048] Specifically, the acoustic model employs either an end-to-end connected temporal classification model or an encoder-decoder architecture based on an attention mechanism. The end-to-end connected temporal classification model does not require precise correspondence between speech frames and label sequences; instead, it uses an end-to-end connected temporal classification loss function to allow the network to automatically learn alignment during training, making it suitable for speech recognition tasks with inconsistent input lengths. The attention-based encoder-decoder architecture utilizes a deep RNN or Transformer in the encoder part to process adaptive speech feature sequences and extract context-dependent high-dimensional representations. Deep RNNs, through multi-layered recurrent structures, can capture temporal dependencies in speech feature sequences. Each RNN unit performs a non-linear transformation on the input features and combines the information from the current time step with the hidden state from the previous time step, progressively extracting an abstract representation with temporal features. For example, when processing consecutive speech frames, a deep RNN can remember information from previous frames, providing contextual reference for feature extraction in the current frame, which helps to more accurately understand the speech content.

[0049] The Transformer architecture, based on a self-attention mechanism, enables parallel processing of speech adaptive feature sequences, improving computational efficiency. The self-attention mechanism allows the acoustic model to automatically consider other relevant features in the sequence when processing each feature, thereby better capturing long-distance dependencies between features. For example, in a speech segment, the feature of a certain syllable may be related to the feature of another syllable located further away. The Transformer can discover these dependencies through its self-attention mechanism, improving the accuracy of feature representation.

[0050] The decoder predicts the corresponding text label based on the feature representation output by the encoder. In the encoder-decoder architecture based on an attention mechanism, the decoder selects relevant feature information from the feature sequence output by the encoder for reference at each step when predicting the text label. This attention mechanism enables the decoder to generate the corresponding text more accurately based on speech features, improving recognition accuracy in complex environments.

[0051] The language model integrates statistical language models or neural network language models, such as N-gram, LSTM-LM, or Transformer-LM, to constrain recognition results and improve grammatical accuracy and speech fluency. Statistical language models use N-grams to count the probability of word sequences, providing grammatical and word order constraints. Neural network language models learn language context features based on deep networks, capturing long-distance dependencies and improving the recognition of rare words or semantically complex sentences. In processing adaptive feature sequences of speech, the acoustic model obtains the probability distribution of each phoneme / character / word for each time frame, outputting the label probability. This is then decoded using the conditional probability from the language model. A search algorithm, such as BeamSearch, is used to obtain the final target text for recognition, which is then output. Specifically, by setting the beam width to represent the number of optimal candidate sequences to retain simultaneously, the joint probability of possible labels for the next time step is calculated for each current candidate sequence. All extended sequences are sorted by probability, and the beam width bar with the highest probability is retained as the next candidate. This process is repeated until the sequence length reaches a termination marker or a specific condition. Finally, the sequence with the highest probability is selected from the bundle for recognition and converted into target recognition text. The output format is phoneme sequence, word or complete sentence.

[0052] By analyzing and recognizing adaptive feature sequences through acoustic and language models, the enhanced speech signals are made more robust. Even in noisy and reverberant environments, high-precision and high-reliability speech recognition can be achieved, enabling clear, accurate, and real-time speech interaction. This ensures the accuracy of functions such as command transmission and information confirmation, improving work efficiency and user experience in complex environments.

[0053] Furthermore, the method also includes: monitoring input signal index parameters, including signal-to-noise ratio estimation, recognition error rate, and processing delay, and dynamically adjusting system parameters based on the input signal index parameters.

[0054] Specifically, the system continuously monitors input signal parameters, including signal-to-noise ratio (SNR) estimation, recognition error rate, and processing latency. SNR estimation reflects the ratio of speech energy to noise energy in the speech signal, reflecting the intensity of environmental noise and the effectiveness of enhancement processing. It can be achieved by detecting pure noise frames using silence segments, calculating noise power, calculating the total power for segments containing speech, and obtaining the SNR estimation result using the SNR estimation formula. Recognition error rate reflects the deviation between the output text and the true label, used to evaluate recognition accuracy, and can be calculated using word error rate or phoneme error rate. Processing latency reflects the time required for the speech signal to go from input to final recognition output, ensuring the system meets real-time requirements.

[0055] Based on monitoring results, system parameters are dynamically adjusted according to rules. For example, when a sharp increase in noise level is detected, the gain of speech enhancement can be temporarily increased or a more aggressive noise suppression strategy can be switched; when the recognition error rate remains high, the parameters of feature extraction or the weights of the language model can be adjusted; when the processing delay approaches the threshold, the short-time Fourier transform frame length or frame rate can be adjusted. Through continuous monitoring combined with adaptive mechanisms, the system can adapt to complex environments with different noise levels and reverberation conditions, ensuring stable and reliable performance under various operating conditions and achieving clear, accurate, and real-time high-precision speech recognition.

[0056] Example 2, based on the same inventive concept as the speech enhancement and high-precision recognition method in complex environments described in the foregoing examples, such as... Figure 2 As shown, this application provides a speech enhancement and high-precision recognition system for complex environments, wherein the speech enhancement and high-precision recognition system for complex environments includes:

[0057] The signal acquisition module 11 is used to synchronously acquire time-domain signals in the off-street parking booth environment using a microphone array, wherein the microphone array includes M microphones. The time-domain signals are preprocessed to obtain a standard time-domain signal. The parameter acquisition module 12 is used to detect silent segments in the standard time-domain signal in real time, estimate the noise power spectral density of the silent segments to obtain a non-stationary noise power spectral density, and simultaneously obtain reverberation parameters by combining speech onset information and noise spatial correlation estimation. The data prediction module 13 is used to construct a deep neural network model, using the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameters as input representations. The deep neural network model is trained and predicted using noisy speech data from the off-street parking environment to obtain a speech mask. The feature extraction module 14 is used to apply the speech mask to the microphone array signal for enhancement processing to obtain a time-domain enhanced speech signal. Adaptive feature extraction is performed on the time-domain enhanced speech signal to obtain a speech adaptive feature sequence. The text output module 15 is used to perform high-precision recognition of the speech adaptive feature sequence based on an acoustic model and a language model, and output target recognition text.

[0058] Furthermore, the signal acquisition module 11 is also used to: perform low-pass anti-aliasing filtering on the time-domain signal to obtain a usable time-domain signal; perform ADC conversion on the usable time-domain signal to obtain a time-domain digital signal; and perform gain control compensation on the time-domain digital signal to obtain a standard time-domain signal.

[0059] Furthermore, the parameter acquisition module 12 is also used to: convert the silent segment signal to the frequency domain using short-time Fourier transform to perform spectrum calculation, thereby obtaining multiple microphone spectra; and to perform PSD estimation on the multiple microphone spectra using a dynamic weighting strategy to obtain the non-stationary noise power spectral density.

[0060] Furthermore, the data prediction module 13 is also used to: the deep neural network model adopts a deep complex value neural network or a U-Net variant structure for time-frequency domain masking prediction, which includes multiple convolutional layers, pooling layers, residual connections and gated linear units.

[0061] Furthermore, the data prediction module 13 is also used to: perform a short-time Fourier transform on the standard time-domain signal to obtain a complex spectrum; stack the multiple microphone spectra on time frames and frequency points to form a multi-channel frequency domain input feature map; select a model composite loss function, use the standard time-domain signal, the non-stationary noise power spectral density, and the reverberation parameter as input representations, and combine the model composite loss function with noisy speech data from off-street parking environments to train and predict the deep neural network model to obtain speech masking.

[0062] Furthermore, the feature extraction module 14 is also used to: calculate the Mel frequency cepstral coefficients of the time-domain enhanced speech signal to obtain the basic features of speech MFCC; introduce adaptive feature mapping technology to combine the non-stationary noise power spectral density to perform adaptive feature extraction on the basic features of speech MFCC, and obtain an adaptive feature sequence of speech.

[0063] Furthermore, the text output module 15 is also used for: the acoustic model adopts an end-to-end connected temporal classification model or an encoder-decoder architecture based on an attention mechanism, wherein the encoder part uses a deep RNN or Transformer to process the speech adaptive feature sequence, the decoder part is used to predict text labels, and the language model integrates a statistical language model or a neural network language model.

[0064] Furthermore, the system is also used to: monitor input signal index parameters, including signal-to-noise ratio estimation, recognition error rate, and processing delay, and dynamically adjust system parameters based on the input signal index parameters.

[0065] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The speech enhancement and high-precision recognition method and specific examples in complex environments described in the foregoing Embodiment 1 are also applicable to the speech enhancement and high-precision recognition system in complex environments in this embodiment. Through the foregoing detailed description of the speech enhancement and high-precision recognition method in complex environments, those skilled in the art can clearly understand the speech enhancement and high-precision recognition system in complex environments in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.

[0066] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0067] Obviously, those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A voice enhancement and high-precision recognition method in a complex environment, characterized in that, The method comprises: Synchronously collecting time domain signals in an off-road parking post environment by using a microphone array, wherein the microphone array comprises M microphones, preprocessing the time domain signals to obtain standard time domain signals; Real-time detecting mute section signals in the standard time domain signals, performing noise power spectrum density estimation on the mute section signals to obtain non-stationary noise power spectrum density, and simultaneously combining voice onset information and noise spatial correlation estimation to obtain reverberation parameters; Constructing a deep neural network model, taking the standard time domain signals, the non-stationary noise power spectrum density and the reverberation parameters as input representations, training and predicting the deep neural network model by using off-road parking environment noisy voice data to obtain voice masking; Applying the voice masking to microphone array signal for enhancement processing to obtain time domain enhanced voice signals, performing adaptive feature extraction on the time domain enhanced voice signals to obtain voice adaptive feature sequences; Performing high-precision recognition on the voice adaptive feature sequences based on an acoustic model and a language model to output target recognition text.

2. The voice enhancement and high-precision recognition method in a complex environment according to claim 1, characterized in that, The standard time domain signals are obtained, comprising: Performing low-pass anti-aliasing filtering on the time domain signals to obtain available time domain signals; Performing ADC conversion on the available time domain signals to obtain time domain digital signals; Performing gain control compensation on the time domain digital signals to obtain standard time domain signals.

3. The voice enhancement and high-precision recognition method in a complex environment according to claim 1, characterized in that, The non-stationary noise power spectrum density is obtained, comprising: Converting the mute section signals to the frequency domain by using short-time Fourier transform for spectrum calculation to obtain multiple microphone spectrums; Performing PSD estimation on the multiple microphone spectrums by using a dynamic weighting strategy to obtain non-stationary noise power spectrum density.

4. The voice enhancement and high-precision recognition method in a complex environment according to claim 1, characterized in that, The deep neural network model adopts a deep complex value neural network or a U-Net variant structure of time-frequency domain masking prediction, and comprises multiple convolution layers, pooling layers, residual connections and gated linear units.

5. The voice enhancement and high-precision recognition method in a complex environment according to claim 3, characterized in that, The voice masking is obtained, comprising: Performing short-time Fourier transform on the standard time domain signals to obtain complex spectrums; Stacking the multiple microphone spectrums in time frames and frequency points to form a multi-channel frequency domain input feature map; Selecting a model composite loss function, taking the standard time domain signals, the non-stationary noise power spectrum density and the reverberation parameters as input representations, combining the model composite loss function, and training and predicting the deep neural network model by using off-road parking environment noisy voice data to obtain voice masking.

6. The voice enhancement and high-precision recognition method in a complex environment according to claim 1, characterized in that, The voice adaptive feature sequences are obtained, comprising: Performing mel-frequency cepstrum coefficient calculation on the time domain enhanced voice signals to obtain voice MFCC basic features; Introducing adaptive feature mapping technology to perform adaptive feature extraction on the voice MFCC basic features in combination with the non-stationary noise power spectrum density to obtain voice adaptive feature sequences.

7. The voice enhancement and high-precision recognition method in a complex environment according to claim 1, characterized in that, The acoustic model adopts an end-to-end connection time sequence classification model or an encoder-decoder architecture based on an attention mechanism, wherein a deep RNN or a Transformer is used in the encoder part to process the voice adaptive feature sequences, and the decoder part is used for predicting text labels, and the language model integrates a statistical language model or a neural network language model.

8. The voice enhancement and high-precision recognition method in a complex environment according to claim 1, characterized in that, The method further comprises: monitoring input signal index parameters, including signal-to-noise ratio estimation, recognition error rate, and processing delay, and dynamically adjusting system parameters based on the input signal index parameters.

9. A speech enhancement and high-precision recognition system in a complex environment, characterized by, Steps for implementing the speech enhancement and high-precision recognition method in a complex environment according to any one of claims 1-8, comprising: a signal acquisition module for synchronously acquiring time-domain signals in an off-road parking booth environment using a microphone array, wherein the microphone array includes M microphones, and the time-domain signals are preprocessed to obtain standard time-domain signals; a parameter acquisition module for real-time detection of silence segment signals in the standard time-domain signals, noise power spectral density estimation of the silence segment signals to obtain non-stationary noise power spectral density, and combination of speech onset information and noise spatial correlation estimation to obtain reverberation parameters; a data prediction module for constructing a deep neural network model, using the standard time-domain signals, the non-stationary noise power spectral density, and the reverberation parameters as input representations, training and predicting the deep neural network model using off-road parking environment noisy speech data, and obtaining speech masking; a feature extraction module for applying the speech masking to microphone array signal enhancement processing to obtain time-domain enhanced speech signals, and performing adaptive feature extraction on the time-domain enhanced speech signals to obtain a sequence of speech adaptive features; a text output module for high-precision recognition of the sequence of speech adaptive features based on an acoustic model and a language model, and outputting target recognition text.

Citation Information

Cited By

  • Decision boundary constraint guided voice confrontation sample generation method

    CN121862078A

  • Dynamic noise adaptive speech recognition method and device based on large language model, equipment and medium

    CN121983034A