Real-time speech recognition across languages

By employing multimodal sound pickup and temporal alignment techniques, combined with deep domain adaptive networks and lightweight CNN-Transformer networks, the acoustic differences and noise robustness issues in cross-language speech recognition are addressed, enabling real-time processing and recognition of cross-language speech and improving recognition accuracy and system stability.

CN120833782BActive Publication Date: 2025-11-21GUANGZHOU SIZHENG ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511317719.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-11-21
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Real-time speech recognition in cross-language scenarios faces challenges such as acoustic differences, temporal loss of multimodal data, insufficient noise robustness, contradiction between real-time performance and resource constraints, and lack of system stability and health monitoring. Existing technologies struggle to dynamically adapt to language-specific features and environmental noise, resulting in low recognition accuracy and high latency.

Method used

We employ multimodal sound pickup and data preprocessing, multimodal temporal alignment, cross-lingual feature adaptation, noise-robust sound pickup, real-time speech recognition, and health monitoring and anomaly handling. Through multimodal sound pickup arrays, deep domain adaptive networks, lightweight CNN-Transformer networks, and health score calculation, we achieve cross-lingual feature adaptation and noise robustness enhancement, ensuring system stability.

Benefits of technology

It achieves high-precision synchronization and recognition of cross-language voice data, improves noise resistance and recognition accuracy in complex environments, meets the requirements of low-latency interaction, and ensures the stability and reliability of the system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833782B_ABST
    Figure CN120833782B_ABST
Patent Text Reader

Abstract

The application discloses a real-time speech recognition and pickup method and system across languages, and relates to the technical field of speech recognition and signal processing. The method comprises the following steps: collecting speech and VAD data, and combining environmental parameters to perform pre-emphasis, frame division and noise removal processing. Linear interpolation and DTW algorithm are adopted to compensate for transmission delay. Acoustic features and a lightweight CNN network are used to detect languages, and a DAN network is used to map general features and language-specific features to a unified space. Based on classified noise types, a beamforming algorithm is dynamically selected, and an adaptive filter is combined to improve the signal-to-noise ratio. Frame-level pipeline control delay is controlled, and pickup parameters are optimized according to WER and SNR feedback. The system health degree is evaluated by weighting device status and processing quality indicators, and an abnormal processing strategy is triggered. The system improves the cross-language recognition accuracy and noise resistance, and is suitable for real-time speech interaction in a multilingual and high-noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech recognition and signal processing technology, specifically relating to a cross-language real-time speech recognition pickup method and system. Background Technology

[0002] Speech recognition technology, as a core component of human-computer interaction, has made significant progress in recent years driven by deep learning and multimodal fusion. However, real-time speech recognition in cross-language scenarios still faces many challenges:

[0003] Cross-linguistic acoustic differences: The phonemes, prosody, and spectral characteristics of different languages ​​vary significantly, making it difficult for universal acoustic models to achieve multilingual recognition accuracy. Existing technologies rely on fixed feature extraction methods but fail to dynamically adapt to language-specific features, limiting the recognition capability of low-resource languages.

[0004] Multimodal data timing discrepancy: In traditional audio pickup systems, the difference in sampling frequencies between the microphone array and the VAD sensor can easily lead to timing errors, affecting the accuracy of voice start and stop detection. Existing alignment methods do not consider the physical delay of the transmission path and require correction through complex algorithms such as DTW.

[0005] Insufficient robustness to noisy environments: interference from multiple speakers, impulse noise, and steady-state noise can significantly reduce the signal-to-noise ratio, while existing beamforming algorithms lack dynamic adaptation to noise types.

[0006] Adaptive filtering techniques are prone to getting stuck in local optima in complex scenarios and have high computational overhead, making it difficult to meet real-time requirements.

[0007] The conflict between real-time performance and resource constraints: Cloud-based speech recognition relies on high computing power, but privacy risks and network latency limit its application on edge devices. While edge models can be deployed on embedded devices, the balance between model complexity and pruning rate is not yet mature, and the accuracy of low-resource speech recognition remains low.

[0008] Lack of system stability and health monitoring: The existing system lacks a joint assessment mechanism for equipment status and processing quality, making it difficult to automatically degrade in the event of hardware failure or sudden environmental changes. The setting of health thresholds lacks quantitative basis, resulting in a lag in the response of anomaly handling strategies.

[0009] Limitations of existing technologies: They rely on predefined language models, making online fine-tuning impossible to handle mixed or minority language scenarios; they do not dynamically adjust pickup parameters based on environmental awareness data, resulting in limited anti-interference capabilities. Frame-level pipeline design and computing power allocation strategies are not yet mature, and latency can easily exceed the 300ms threshold in high-noise scenarios.

[0010] To address the aforementioned problems, this invention proposes a cross-language real-time speech recognition pickup method and system. Summary of the Invention

[0011] In order to overcome the shortcomings and deficiencies of the existing technology, the first objective of this invention is to provide a cross-language real-time speech recognition pickup method; the second objective of this invention is to provide a cross-language real-time speech recognition pickup system.

[0012] The first objective of this invention is achieved through the following technical solution:

[0013] A cross-language real-time speech recognition method includes the following steps:

[0014] Multimodal sound pickup and data preprocessing, multimodal temporal alignment, cross-linguistic feature adaptation, noise-robust sound pickup, real-time speech recognition, and health monitoring and anomaly handling;

[0015] Preferably, multimodal sound pickup and data preprocessing involves acquiring microphone speech data and VAD sensor data through a multimodal sound pickup array, obtaining noise level and speaker count through an environmental perception interface, performing DC removal, pre-emphasis, and frame segmentation on the microphone data, and smoothing filtering on the VAD data. Preferably, multimodal timing alignment corrects timing errors through linear interpolation, transmission delay compensation, and the DTW algorithm. Preferably, cross-language feature adaptation achieves cross-language feature unification through language detection, feature extraction, and DAN network mapping. Preferably, noise-robust sound pickup suppresses residual noise through noise classification, beamforming algorithm selection, and adaptive filtering. Preferably, real-time speech recognition achieves real-time recognition and feedback optimization parameters through a lightweight CNN-Transformer network. Preferably, health monitoring and anomaly handling ensure stable system operation through health score calculation and a tiered strategy.

[0016] Preferably, in the pre-emphasis processing, a pre-emphasis coefficient is set for Chinese to highlight high-frequency unvoiced sounds, and a pre-emphasis coefficient is set for English to balance high-frequency and low-frequency energy; preferably, the frame processing adopts a 20ms frame length and a 10ms frame shift, and uses the Hanning window function for weighting to reduce spectral leakage; preferably, the smoothing filtering of VAD data adopts a 5th-order Butterworth low-pass filter with a cutoff frequency of 100Hz to remove high-frequency noise.

[0017] Preferably, in multimodal timing alignment, linear interpolation adapts the VAD sensor data to the sampling rate; transmission delay compensation calculates the delay difference by measuring the physical length of the data transmission path between the microphone and the VAD sensor and the signal propagation speed, and adjusts the timestamp so that the timing error between the microphone voice frame and the VAD start / stop signal does not exceed a set value; the DTW algorithm constructs a time series matrix of the microphone voice frame and the VAD signal, calculates the Euclidean distance and dynamically plans the optimal path, and adjusts the VAD signal time series to correct the timing error.

[0018] Preferably, language detection constructs a feature vector by extracting the fundamental frequency F0 and formant F1 features of speech, and inputs it into a lightweight language classification network. This network is trained on a multilingual dataset containing Chinese, English, Japanese and other minority languages, and uses the cross-entropy loss function and the Adam optimizer.

[0019] Preferably, in cross-language feature adaptation, MFCC calculation is achieved through short-time Fourier transform, Mel filter bank, and discrete cosine transform; for Chinese focused tone features, the center frequency shift of filters 5-15 is adjusted; for English stressed accent features, the center frequency concentration of filters 12-25 is adjusted; for minority languages, the center frequency of the first 10 filters is optimized through a genetic algorithm; when a minority language is detected, the K halving strategy is triggered, K=512, and K=1024 is restored in multi-speaker interference scenarios.

[0020] Preferably, the number of neurons in the input layer of the DAN network matches the dimension of the source language-specific features, the hidden layer 1 contains 256 ReLU activated neurons, the hidden layer 2 contains 128 ReLU activated neurons, and the output layer contains 64 linear activated neurons; the mapping matrix W is optimized through cross-language pre-training and online fine-tuning, and is jointly trained by combining MMD loss and classification loss.

[0021] Preferably, in the noise-robust sound pickup, noise classification uses a CNN-LSTM network, with the input being a log-Mel spectrum. The CNN part contains two layers of convolution and pooling, and the LSTM part is a bidirectional LSTM layer. The output is the probability of four types of noise: steady-state, broadband, impulse, and multi-speaker. The beamforming algorithm is selected according to the noise type: steady-state noise uses delay-summation, impulse noise uses MVDR, and multi-speaker noise uses sparse beamforming. In the adaptive filtering, steady-state noise uses ANF, broadband noise uses improved spectral subtraction, and multi-speaker interference uses a combination of VAD and speaker separation.

[0022] Preferably, real-time speech recognition adopts frame-level pipeline processing. When the microphone acquires frame t, it preprocesses frame t−1 and recognizes frame t−2, with a delay not exceeding a set value. Preferably, the health score is a weighted sum of device status indicators and processing quality indicators, with the weights determined by AHP. When the health score is less than the set value, a grading strategy is triggered: for mild abnormalities, the microphone gain is increased; for moderate abnormalities, the backup VAD sensor is switched; and for severe abnormalities, the single microphone + basic noise suppression mode is switched.

[0023] The second objective of this invention is achieved through the following technical solution:

[0024] A cross-language real-time speech recognition and pickup system includes a main control module, a pickup and perception module, a cross-language speech processing engine, and a real-time recognition and feedback module.

[0025] Preferably, the main control module, serving as the core scheduling hub, is deployed on an edge computing device and includes a synchronization scheduling unit, a resource allocation unit, and an instruction interaction unit. Preferably, the synchronization scheduling unit achieves hardware clock synchronization through the PTP protocol, combined with software dynamic compensation, to control the synchronization error to not exceed a set value. Preferably, the resource allocation unit dynamically allocates CPU / GPU computing power according to the language type and environmental noise level, reducing the CPU computing power for recognition tasks when the language has low resources, and reserving a set proportion of redundant computing power for noise suppression when the noise level exceeds a set value. Preferably, the instruction interaction unit parses user instructions through natural language processing, issues language adaptation instructions, and provides feedback on the sound pickup quality.

[0026] Preferably, the sound pickup and sensing module includes a 6-8 channel ring microphone array and 2 channels VAD sensors, in conjunction with a data preprocessing unit and an environmental sensing interface; the data preprocessing unit is used for DC removal, pre-emphasis, and frame segmentation processing, high-frequency compensation for voiceless consonants in Japanese to improve frequency band gain, and adding sliding window smoothing filtering in low signal-to-noise ratio scenarios; the environmental sensing interface is used to read noise level and number of speakers.

[0027] Preferably, the cross-language speech processing engine includes a language feature adaptation engine, a noise-robust sound pickup engine, and a multimodal temporal alignment engine;

[0028] Preferably, the real-time recognition and feedback module includes a lightweight CNN-Transformer network, which controls the end-to-end latency through a three-stage parallel pipeline and achieves closed-loop optimization in conjunction with the recognition quality evaluation unit. The three-stage parallel pipeline includes acquiring t frames, preprocessing t-1 frames, and recognizing t-2 frames.

[0029] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0030] 1. This invention effectively solves the timing synchronization problem caused by the difference in sampling frequency between the microphone and the VAD sensor through linear interpolation, transmission delay compensation and dynamic programming algorithm, realizes high-precision synchronization of voice data and control signals, provides a stable and reliable input foundation for subsequent cross-language processing, accurately synchronizes multimodal timing and improves the reliability of sound pickup.

[0031] 2. This invention combines Deep Domain Adaptive Network (DAN) and noise classification technology to automatically adapt to the acoustic feature differences of different languages ​​and dynamically adjust the processing algorithm according to the type of environmental noise, thereby improving the noise resistance and multilingual recognition accuracy in complex environments, and enhancing the dynamic adaptation of cross-language features and noise robustness.

[0032] 3. This invention adopts a pruned and compressed CNN-Transformer network and a frame-level pipelined parallel processing architecture, combined with dynamic computing power scheduling, to reduce end-to-end processing latency while maintaining high recognition accuracy. This meets the dual requirements of low latency and high reliability in cross-language real-time voice interaction scenarios. The lightweight model and real-time processing optimization ensure low-latency interaction. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 A flowchart of the cross-language real-time speech recognition pickup method of the present invention is shown;

[0035] Figure 2 A block diagram of the cross-language real-time speech recognition pickup system of the present invention is shown;

[0036] Figure 3 A flowchart of the cross-language feature adaptation of the present invention is shown. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more exemplary embodiments. Numerous specific details are provided in the following description to give a full understanding of exemplary embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, steps, etc., can be employed. In other instances, well-known structures, methods, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0039] Example 1:

[0040] Reference Figure 1As shown, this invention provides a cross-language real-time speech recognition pickup method. This method includes the following steps, achieving real-time optimization of the entire process from pickup and processing to recognition. The specific steps are as follows:

[0041] Step S1: Multimodal sound pickup and data preprocessing.

[0042] A multimodal pickup array acquires microphone speech data and VAD sensor data, while an environmental perception interface reads noise levels and the number of speakers. The data preprocessing unit performs DC removal, pre-emphasis (adapted to the current language), and framing (20ms frame length, 10ms frame shift) on the microphone data. In the pre-emphasis stage, for Chinese, based on its speech energy distribution characteristics, the pre-emphasis coefficient is typically set to 0.95-0.98 to highlight the high-frequency components and improve unvoiced speech recognition; for English, the pre-emphasis coefficient is generally between 0.93-0.96 to balance high-frequency and low-frequency energy. During framing, a Hanning window function is used to weight each frame to reduce spectral leakage. The VAD data is smoothed using a 5th-order Butterworth low-pass filter with a cutoff frequency of 100Hz to remove high-frequency noise interference, and the preprocessed data is output.

[0043] Step S2: Multimodal timing alignment.

[0044] The multimodal timing alignment engine receives preprocessed data and adapts the VAD data to 48kHz using linear interpolation. In practice, it assumes the VAD sensor data is a discrete sequence. The corresponding time point is To interpolate it to the time point corresponding to a 48kHz sampling rate For any Find satisfaction If i, then the interpolated value .

[0045] To compensate for transmission delay, the transmission delay is calculated by measuring the physical length of the data transmission path between the microphone and the VAD sensor, and the signal propagation speed (the speed of an electrical signal in a wire is approximately the speed of light, c). Assume the microphone data transmission path length is... The data transmission path length of the VAD sensor is Then the transmission delay difference The timestamp is adjusted accordingly to ensure that the timing error between the microphone voice frame and the VAD start / stop signal is ≤1ms.

[0046] If the error exceeds the standard, the DTW algorithm is used for correction. In the DTW algorithm, a time series matrix of microphone speech frames and VAD signals is constructed, the distance of each element in the matrix is ​​calculated (using Euclidean distance), the optimal path is found through dynamic programming algorithm, the VAD signal time series is adjusted according to the optimal path, and the aligned pickup data and speech start and stop markers (T0, T1) are output.

[0047] Step S3: Cross-language feature adaptation.

[0048] See Figure 3 As shown, the process of cross-language feature adaptation is as follows:

[0049] The language feature adaptation engine performs language detection (acoustic features + model matching) on ​​the aligned data to determine the target language. During acoustic feature detection, features such as the fundamental frequency (F0) and formants (F1) of the speech are extracted to construct a feature vector. For Chinese, the fundamental frequency (F0) ranges from 100-500Hz, and the formants (F1) range from approximately 300-700Hz; for English, the fundamental frequency (F0) ranges from 80-300Hz, and the formants (F1) ranges from approximately 200-600Hz.

[0050] Acoustic feature vectors were input into a lightweight language classification network (3-layer CNN + 1-layer fully connected layer). This network was trained on a multilingual dataset containing 1 million speech samples, with common languages ​​such as Chinese, English, and Japanese each accounting for 20%, and other less common languages ​​accounting for the remaining 20%. During training, the cross-entropy loss function and Adam optimizer were used, with a learning rate of 0.001. After 100 epochs of training, the model converged, achieving a language detection accuracy of ≥98%.

[0051] Extract general acoustic features (MFCC, PLP) and language-specific features (such as Chinese tones). MFCC calculation is performed using Short-Time Fourier Transform (STFT) → Mel-Filter Bank → Discrete Cosine Transform (DCT), with the specific formula as follows: ;in, For the nth dimension MFCC coefficients (n=1,2,...,13); The STFT spectrum of the speech signal; Let m be the response function of the m-th Mel filter (covering a frequency range that matches the speech acoustic band). The number of Mel filters; This represents the number of STFT spectral points. It is a dynamic value. ,in, The highest acoustic characteristic frequency of the target language (8000Hz for Chinese and 7000Hz for English). For the sound pickup sampling rate, The frame length corresponds to 1024 sampling points, which is rounded up to 1024 points because 48kHz×0.02s=960. When a less common language (such as Korean or Japanese) is detected, the K halving strategy (K=512) is triggered because the acoustic frequency band is more concentrated (Korean 70-6000Hz) to reduce the amount of computation. If a multi-speaker interference scene is identified, K=1024 is restored to preserve high-frequency details.

[0052] Mel filter Acoustic characteristics of different languages ​​are matched layer by layer using center frequencies: Number of filters M=40 (general configuration, covering 20-8000Hz), center frequency of the m-th filter. satisfy: ,in, The center frequency of the Mel scale is determined through linguistic feature statistics:

[0053] Chinese: Focus on tone features (300-700Hz), shift the center frequency of filters 5-15 (40 in total) to this range (e.g., adjust the center frequency of filter 10 from 500Hz to 550Hz) to improve tone recognition;

[0054] English: Enhance the accent features (1000-3000Hz), and concentrate the center frequencies of filters 12-25 into this range (e.g., adjust the center frequency of filter 20 from 2000Hz to 2200Hz).

[0055] For less commonly spoken languages ​​(such as Vietnamese): Based on the acoustic distribution of the pre-trained corpus, the center frequencies of the first 10 filters are automatically optimized using a genetic algorithm to adapt to their unique consonant bands.

[0056] PLP is based on an auditory model and is obtained through critical band analysis → equal loudness pre-emphasis → intensity-loudness conversion → linear prediction analysis. The key formula is: ;in, The linear prediction coefficients (derived from the autocorrelation function) Solved using the Levinson-Durbin algorithm); The speech signal after equal loudness pre-emphasis (pre-emphasis coefficient α=0.97); This is the prediction order (usually 12).

[0057] The DAN network is mapped to a unified feature space. The number of neurons in the input layer of the DAN network is the same as the dimension of the source language-specific features. For example, if the dimension of Chinese tone features is 5, then the number of neurons in the input layer is 5. Hidden layer 1 has 256 neurons and uses the ReLU activation function; hidden layer 2 has 128 neurons and uses the ReLU activation function; the number of neurons in the output layer is 64 and uses the linear activation function.

[0058] Training and Optimization: This DAN network is jointly optimized using a domain-adaptive loss function (MAD) and a classification loss function. It employs a Domain Adaptive Network (DAN) to map language-specific features to a unified feature space. The DAN learns a mapping matrix W by minimizing the feature distribution difference between the source language (e.g., Chinese) and the target language (e.g., English), thus integrating the language-specific features... Mapping to a unified feature space The formula is: ;in, For the mapping matrix ( To unify feature dimensions, (Source language feature dimension) ; This is the bias term. The mapping process is optimized through cross-linguistic corpus pre-training (fixed W) and online fine-tuning (dynamically adjusting W) to ensure... Preserve common features across languages .

[0059] The training rules for the DAN network mapping matrix W and bias b are as follows:

[0060] (1) Initialization method

[0061] Mapping matrix W: Orthogonally initialized to avoid gradient vanishing in the early stages of training. The formula is: ;in, For source language feature dimensions (such as Chinese tone features) ); To unify feature dimensions; Sampling is performed for uniform distribution.

[0062] Bias b: Initialized as a vector of all zeros ( (This can be dynamically adjusted through training.)

[0063] (2) Training process update rules.

[0064] Cross-language pre-training:

[0065] Corpus size: 100,000 cross-language parallel corpora were used (50,000 Chinese-English, 30,000 Chinese-Japanese, and 20,000 English-Korean). Each corpus contains source language-specific features (such as Chinese tones) and target language text annotations.

[0066] Fixed W training: Freeze W, only update the classification layer parameters, use MMD loss + cross-entropy loss for joint optimization, learning rate η=0.01, batch size=128, train for 500 epochs to initially align the feature distribution.

[0067] Online fine-tuning of trigger conditions and cycles:

[0068] Triggering conditions: When the confidence score Conf < 0.8 or the word error rate WER > 15% for 10 consecutive frames of recognition results, it is determined to be a cross-linguistic feature mismatch, triggering W fine-tuning;

[0069] Fine-tuning cycle: Each fine-tuning lasts for 5 epochs, the learning rate decays to η=0.001, batchsize=64, and only the weights of the last layer of W are updated (to avoid overfitting since the front-end features have been pre-trained).

[0070] Step S4: Noise-robust pickup.

[0071] A noise-robust sound pickup engine classifies noise based on the adapted features (CNN-LSTM network). A cross-scene noise dataset is constructed, covering four typical noise types: steady-state noise (e.g., air conditioner sound, frequency 200-800Hz), broadband noise (e.g., crowd noise, frequency 500-5000Hz), impulse noise (e.g., door closing sound, duration ≤50ms), and multi-speaker interference (2-3 people speaking in a mixed manner). The dataset contains 10,000 noise samples in a 1:1:1:1 ratio, with 5,000 used for training, 3,000 for validation, and 2,000 for testing.

[0072] The input to the CNN-LSTM network is a log-Mel spectrogram obtained through a short-time Fourier transform, with dimensions of [number of time frames] × [number of Mel filter banks]. For example, for a 20ms audio frame using 128 Mel filters, the input feature dimension of a single sample is (T, 128), where T is the number of consecutive time frames. CNN part: The first convolutional layer uses 64 3×3 kernels with a stride of (1,1), followed by a ReLU activation function and a max-pooling layer (pooling size 2×2); the second convolutional layer uses 128 3×3 kernels with a stride of (1,1), followed by a ReLU activation function and a max-pooling layer (pooling size 2×2). LSTM part: The feature sequence output by the CNN is expanded along the time dimension and input into a bidirectional LSTM (Bi-LSTM) layer with 64 neurons. Finally, a fully connected layer is connected, and the Softmax activation function is used to output the probability distribution of four types of noise (steady-state, wideband, impulse, and multi-speaker).

[0073] Select a beamforming algorithm based on the noise type (e.g., use sparse beamforming for multi-speaker scenarios) and combine it with adaptive filtering to suppress residual noise.

[0074] Delay-sum beamforming is used to calculate the target speaker's position (horizontal angle) based on the time delay difference (TDOA) of the microphone array. +pitch angle φ), the formula is: ;in, Let TDOA be the difference between the i-th microphone and the reference microphone; The speed of sound (340 m / s); The TDOA is the microphone coordinate difference. The delay time of each microphone signal is determined by the calculated TDOA, and then weighted and summed to achieve beamforming gain.

[0075] Circular array microphone coordinate difference The calculation is as follows:

[0076] (1) Array physical layout

[0077] A 6-channel circular microphone array is used, deployed on the top of the pickup terminal, with a radius r = 5cm (to balance the pickup range and hardware size). The microphones are numbered i = 1, 2, ..., 6, corresponding to polar angles. (Uniformly distributed).

[0078] (2) Coordinate difference Derivation

[0079] Establish a polar coordinate system with the reference microphone (i=1) as the origin, and the coordinates of the i-th microphone are: Coordinate difference with reference microphone for: Substitution have to:

[0080] i=1 (reference microphone): (It has no inherent difference);

[0081] i=2,6: ,

[0082] i=3,5: ,

[0083] i=4: ,

[0084] Beamforming Gain Microphone weight For each microphone signal The weighted summation is achieved using the following formula: Where N is the number of microphones (6-8 channels); To determine the position of the target speaker (horizontal angle) Pitch angle The weights for dynamically adjusted noise type are:

[0085] Adaptive beamforming: Dynamically adjusts the beammask based on noise type: Delay-sum beamforming is used for steady-state noise scenarios, Minimum Variance Distortionless Response (MVDR) beamforming is used for impulse noise scenarios, and sparse beamforming (focusing on the target speaker) is used for multi-speaker scenarios. Beamforming gain... Microphone weight For each microphone signal The weighted summation is achieved using the following formula: Where N is the number of microphones (6-8 channels); To determine the position of the target speaker (horizontal angle) Pitch angle The weights for dynamically adjusted noise type are:

[0086] Steady-state noise scenario: For fixed delay - summation weights ( f is the speech frequency. (for TDOA)

[0087] Impulse noise scenario: Minimum variance weights optimized for the MVDR algorithm ( , The noise covariance matrix is... (as the guiding vector)

[0088] Multi-speaker scenarios: Weighting of sparse beamforming (only retaining the microphone signal in the direction of the target speaker).

[0089] Adaptive filtering suppresses residual noise:

[0090] Steady-state noise is controlled by an adaptive notch filter (ANF) that tracks the fundamental frequency of the noise and covers the 1st to 3rd harmonics.

[0091] Broadband noise is addressed using an improved spectral subtraction method, incorporating cross-lingual speech spectral characteristics (such as tone band protection in Mandarin). The noise spectrum estimate is subtracted from the noisy speech spectrum, and a language-adaptive protection coefficient is introduced to obtain the denoised spectrum. ;in, This is the denoised spectrum; The spectrum of noisy speech; For noise spectrum estimation; Protection coefficient (Chinese = 1.2, English = 1.1, Japanese = 1.0). This is the lower limit of the spectrum (to prevent negative spectrum, usually set to 0.1). Output target speech data with a signal-to-noise ratio ≥ 25dB.

[0092] Multi-speaker interference: Speech activity detection (VAD) is combined with speaker separation to retain target language speech frames and remove interfering speech frames.

[0093] Step S5: Real-time speech recognition.

[0094] The real-time recognition and feedback module inputs the target speech data into a lightweight CNN-Transformer network, achieving real-time recognition (latency ≤300ms) through frame-level pipeline processing. In the frame-level pipeline, when the microphone captures frame t, frame t−1 is preprocessed and frame t−2 is recognized, with parallel processing reducing latency by 15%. The lightweight CNN-Transformer network adopts an improved architecture and is trained on a large-scale multilingual dataset (containing 1 million speech samples covering more than 10 languages), using the cross-entropy loss function, the Adagrad optimizer, and a learning rate of 0.005. The model converges after 500 epochs of training.

[0095] The recognition quality assessment unit calculates WER and SNR. If WER > 15%, it provides feedback to optimize the pickup parameters (e.g., adjust microphone gain). When calculating WER, the recognition result is compared with the real text, and the minimum edit distance is calculated using a dynamic programming algorithm. When calculating the SNR, a power spectrum estimation-based method is used to first estimate the power of the speech signal. and noise power , .

[0096] S6: Health monitoring and abnormality handling.

[0097] The health monitoring module calculates the system's health score in real time. The health score H is determined by the device status indicators. Device status indicators (microphone operating status, VAD sensor response speed, network transmission rate) and processing quality indicators The weighted sum of (language detection accuracy, sound pickup signal-to-noise ratio, and recognition delay) is given by the following formula: ;in, For equipment status indicators; To process quality indicators; The normalized value (0-100) of the status index of the m-th device. Let M be the normalized value (0-100) of the k-th processing quality index, and M and K be the equipment status and the number of processing quality indices, respectively.

[0098] Different strategies are triggered based on health score:

[0099] Mild abnormality ( ): If the sensitivity of a single microphone decreases (signal-to-noise ratio drops to 20-25dB), it triggers a fine-tuning of device parameters (increasing the gain of that microphone by 2dB).

[0100] Moderate abnormality ( If the VAD sensor is offline, trigger the switchover to the backup device (activate the backup VAD sensor, with a switching delay of ≤50ms).

[0101] Severe abnormality ( If half of the microphone array fails or the recognition delay is greater than 500ms, the system will be downgraded (switched to "single microphone + basic noise suppression" mode) and a fault warning will be sent to the user terminal.

[0102] The weights of each indicator were determined using the Analytic Hierarchy Process (AHP). An exemplary weight allocation determined by AHP is as follows: microphone operating status (0.15), VAD response speed (0.15), network speed (0.10), speech detection accuracy (0.20), sound pickup signal-to-noise ratio (0.20), and recognition latency (0.20). After analyzing the system's operation in 100 different scenarios, when the health score was <80, the probability of system failure or performance degradation exceeded 80%. Therefore, 80 was selected as the threshold.

[0103] If the system recovers from an anomaly, it retrieves historical health parameters to ensure stable operation of the voice pickup and recognition process. Historical parameters are saved every 30 seconds, retaining the most recent 100 sets of historical parameters. When the system recovers from an anomaly, it prioritizes retrieving parameters from the most recent healthy state, with a recovery time of ≤10ms. Finally, the recognition result is output to the user terminal, completing the entire cross-language real-time speech recognition and voice pickup process.

[0104] The beneficial effects of this embodiment are as follows: by using multimodal sound pickup and timing alignment technology, timing errors are reduced, enabling real-time processing and recognition of cross-language speech; by combining noise classification and adaptive noise reduction algorithms, noise resistance is significantly enhanced; dynamic language feature adaptation mechanism supports high-accuracy recognition in multilingual scenarios; lightweight model optimization and frame-level pipeline design ensure low-latency response; health monitoring and anomaly handling ensure system stability, making it suitable for cross-language real-time interaction needs in complex environments, and comprehensively improving the reliability and applicability of the sound pickup-recognition process.

[0105] Example 2:

[0106] See Figure 2 As shown, the cross-language real-time speech recognition and pickup system of this embodiment aims to solve the problems of large differences in acoustic features in cross-language scenarios (such as Chinese, English, Japanese, etc.), insufficient robustness to environmental noise (such as background noise in public places, interference from multiple speakers), temporal loss of multimodal pickup data, and difficulty in balancing speech recognition and real-time pickup. The system includes:

[0107] 2.1 Main control module.

[0108] As the core scheduling hub of the system, deployed on edge computing devices (such as embedded microphone terminals and lightweight cloud nodes), it possesses low-latency coordination and dynamic resource allocation capabilities, and is responsible for coordinating the entire process of microphone perception, cross-language processing, real-time recognition, and status monitoring. Internally, it includes:

[0109] Synchronization Scheduling Unit: This unit synchronizes the timestamps of the multimodal audio pickup devices and the speech recognition module through hardware clock synchronization (such as the PTP protocol) and software dynamic compensation. During hardware clock synchronization, the PTP protocol is used to transmit high-precision time information between devices by accurately measuring network transmission latency, achieving sub-microsecond synchronization accuracy. Software dynamic compensation periodically adjusts the local clock based on the clock drift rate between devices, ensuring the timing consistency between audio pickup data and recognition commands, with synchronization errors controlled within 5ms.

[0110] Resource allocation unit: Dynamically allocates CPU / GPU computing power based on language type (e.g., low-resource language / high-resource language) and environmental noise level (e.g., quiet / noisy). Through complexity analysis of different language recognition models, the number of parameters in a low-resource language model is approximately 50% of that in a high-resource language model, resulting in a corresponding reduction in computational load. When a low-resource language is detected, the CPU computing power allocated to the recognition task is reduced by 30%; when the noise level is >70dB, 20% redundant computing power is reserved for noise suppression. Priority is given to cross-language feature adaptation and real-time recognition (overall sound pickup-recognition latency ≤300ms).

[0111] Command Interaction Unit: Receives user speech recognition requests (e.g., "Switch to Japanese voice pickup") and performs semantic analysis and understanding of the user's commands using natural language processing technology. First, the user's command is segmented into words, then matched against a predefined command template library to determine the command type and parameters. The language adaptation command is then sent to the voice pickup perception module, with simultaneous feedback on voice pickup quality (e.g., "Current voice pickup signal-to-noise ratio ≥ 25dB, recognition ready").

[0112] 2.2 Sound pickup and sensing module.

[0113] Responsible for collecting cross-language speech signals and environmental data to provide high-quality input for cross-language processing, including:

[0114] Multimodal microphone array: Composed of a 6-8 channel circular microphone array (deployed on the top / edge of the terminal) and 2 channels of speech activity detection (VAD) sensors (attached to the side of the microphone terminal). The microphones have a sampling rate of 48kHz and a bit depth of 24bit (covering multiple language acoustic bands: Chinese 80-8000Hz, English 100-7000Hz), while the VAD sensors have a sampling rate of 1kHz (detecting the start and stop of speech). In the circular microphone array, the microphone spacing is determined based on the Nyquist sampling theorem and the actual application scenario, typically 5-10cm, to ensure effective acquisition of speech signals from different directions.

[0115] The data preprocessing unit performs DC removal, pre-emphasis (for different language high-frequency features, such as high-frequency compensation for voiceless consonants in Japanese), and frame segmentation (20ms frame length, 10ms frame shift) on the raw audio data to filter out microphone hardware noise. For high-frequency compensation of voiceless consonants in Japanese, a boost filter is used to increase the gain by 5-10dB in the 8000-10000Hz frequency band. For low signal-to-noise ratio scenarios (<15dB), an additional 5ms sliding window smoothing filter is added, using a mean filtering algorithm to average the data within the window.

[0116] Environmental perception interface: Reads environmental sensor data (noise level, number of speakers, temperature, humidity) as dynamic environmental parameters for cross-language voice pickup adaptation. Noise level is measured by a sound pressure level sensor, and the number of speakers is determined by a combination of microphone array beamforming technology and speech activity detection algorithm. When multiple speakers (≥2) are detected, a multi-target voice pickup mode is triggered, which simultaneously focuses on the directions of multiple speakers by adjusting the parameters of the beamforming algorithm.

[0117] 2.3 Cross-language speech processing engine.

[0118] The core innovative module of this invention runs on the computing core of the main control module and achieves robust sound pickup through a cross-language adaptation algorithm, including:

[0119] Language Feature Adaptation Engine: Real-time detection of the language of speech, extraction of cross-language universal acoustic features and language-specific features, and elimination of the impact of language acoustic differences on sound pickup through an adaptive model.

[0120] Noise-robust pickup engine: Based on noise classification and beamforming technology, it suppresses environmental noise (such as steady-state noise, impulse noise, and multi-speaker interference) and focuses on the target language speech.

[0121] Multimodal timing alignment engine: Aligns microphone array voice data with VAD sensor signals, solves timing discrepancies caused by differences in sampling frequencies of multiple devices, and ensures accurate matching between voice start and stop times and pickup data.

[0122] 2.4 Real-time recognition and feedback module.

[0123] Real-time cross-language recognition is achieved based on adapted voice pickup data. Low latency and high accuracy are ensured through a "lightweight model + dynamic optimization + quality closed loop," specifically including:

[0124] (1) The lightweight recognition network adopts an improved CNN-Transformer network (pruning rate of 30%), which supports real-time recognition of 10+ languages ​​(Chinese, English, Japanese, Korean, Spanish, French, etc.). The core design revolves around "cross-language adaptation + simplified computing power", as follows:

[0125] Network structure details:

[0126] The CNN encoder consists of three layers responsible for extracting local acoustic features. Layer 1 has 64 convolutional kernels (3×3, stride 1), and layers 2 and 3 have 128 convolutional kernels (3×3, stride 1). Each layer is followed by a Batch Normalization (BN) layer and a ReLU activation function. Finally, global average pooling compresses the feature dimension to 256 dimensions. To address cross-lingual acoustic differences, a language-adaptive convolutional kernel is added after the third convolutional layer—an additional tone-sensitive kernel for Chinese (covering the 80-500Hz frequency band) and an accent-sensitive kernel for English (covering the 1000-3000Hz frequency band). A dynamic activation mechanism is used to match the current target language.

[0127] Transformer Decoder: It adopts a 4-layer encoder-decoder structure, and the number of multi-head attention heads is pruned from the original 8 to 5 (removing 3 redundant low-frequency attention heads). The hidden layer dimension is 512 (768 dimensions before pruning). It introduces cross-language attention masking, which applies a mask (weight attenuation of 80%) to non-target language acoustic feature regions based on the unified feature space output by the language feature adaptation engine, focusing on the target language speech information.

[0128] Output layer: It adopts a design that combines a shared vocabulary and a language-specific vocabulary. The shared vocabulary contains 10+ language-common symbols (such as numbers and punctuation marks), while the language-specific vocabulary is customized for the characteristics of each language (such as Chinese Pinyin, English letters, and Japanese kana). The corresponding vocabulary is dynamically activated based on the language detection results to reduce invalid calculations.

[0129] Pruning and optimization details: Pruning covers CNN convolutional kernels (removing 30% of low-frequency redundant convolutional kernels), Transformer attention heads (removing 3), and fully connected layer weights; after pruning, the model size is compressed from 280MB to 196MB, improving inference speed. The pruning uses an iterative pruning algorithm based on weight magnitude. First, the 30% of weights with the smallest absolute value are removed from the pre-trained model, and then the pruned model is retrained to restore accuracy.

[0130] Pre-training and fine-tuning: Pre-training is performed on a corpus containing 500,000 cross-lingual speech samples (150,000 each for Chinese and English, 80,000 each for Japanese and Korean, and 40,000 for other languages). The cross-entropy loss and CTC (connection-temporal classification) loss are used for joint optimization. For low-resource languages ​​(such as minority languages), the pre-trained parameters of high-resource languages ​​(Chinese and English) are reused through transfer learning. Only the language-specific word surface is fine-tuned, which can reduce the amount of training data by 60%.

[0131] (2) Delay control unit.

[0132] By employing a "three-level parallelism + dynamic computing power scheduling" approach, the end-to-end latency of sound pickup and recognition is controlled within 200-300ms. Specific strategies include:

[0133] Parallel frame-level pipeline: A three-stage pipeline of "acquiring frame t → preprocessing frame t-1 → identifying frame t-2" is adopted, with each stage taking ≤80ms, and the overall latency is reduced after stacking.

[0134] Dynamic computing power allocation: The CPU (quad-core ARM Cortex-A53) and GPU (Mali-G52) on the edge terminal (such as embedded device) work together; in simple scenarios (noise ≤50dB + high resource language), only CPU inference is enabled (latency ≤200ms); in complex scenarios (noise >50dB + low resource language), GPU acceleration is triggered (two computing units are enabled, latency ≤300ms).

[0135] Data compression and transmission: After adaptation, the features adopt inter-frame differential coding, and only the feature difference values ​​of adjacent frames are transmitted (compression rate of 60%), which reduces the data transmission time in edge-cloud collaborative scenarios (from 50ms to 20ms).

[0136] (3) A multi-index + closed-loop feedback mechanism is constructed to identify the quality assessment unit, and the sound pickup and recognition parameters are optimized in real time. Specifically:

[0137] Core evaluation metric: Word Error Rate (WER): Calculated using dynamic programming to determine the minimum edit distance (number of insertions / deletions / replacements) between the identified text and the labeled text. The formula is as follows: The requirements are: WER ≤ 10% for high-resource languages ​​and WER ≤ 18% for low-resource languages.

[0138] Recognition confidence (Conf): Calculated based on the probability distribution output by the Transformer decoder, using the following formula: ;in, Let be the probability distribution of the i-th token, and N be the total number of tokens. Conf ≥ 0.85 is required.

[0139] Signal-to-noise ratio (SNR): The SNR value output by the noise-robust pickup engine, requiring SNR ≥ 20dB.

[0140] Feedback optimization strategy:

[0141] If WER > 15% and SNR < 20dB: Send a gain boost command to the noise robust pickup engine (increase microphone array gain by 3dB) and switch to the MVDR beamforming algorithm.

[0142] If WER > 15% and Conf < 0.8: Send a model fine-tuning instruction to the language feature adaptation engine (update the DAN network mapping matrix, learning rate 0.001, fine-tune for 5 epochs).

[0143] If WER ≤ 10% for 5 consecutive frames: save the current parameters (microphone gain, beamforming weight, recognition network pruning ratio) as a scene template, and call it directly when similar scenes (same language + noise level ± 5dB) are detected later, reducing startup delay.

[0144] The beneficial effects of this embodiment are as follows: reducing timing errors through multimodal timing alignment, improving multilingual recognition accuracy by combining cross-language feature adaptation, enhancing noise resistance by utilizing noise classification and adaptive noise reduction, ensuring low latency through lightweight models and pipeline design, ensuring system stability through health monitoring, making it suitable for cross-language real-time interaction in complex environments, and comprehensively improving the reliability and applicability of the sound pickup-recognition process.

[0145] All formulas in this invention are dimensionless and calculated numerically. The preset parameters in the formulas can be set by those skilled in the art according to the actual situation.

[0146] The weighting coefficients of this invention are used to measure the degree of influence of different factors or variables on a certain outcome or decision. The weighting coefficient is defined as the numerical value assigned to each factor when comparing and evaluating multiple factors, reflecting their importance or priority. These weighting coefficients can be determined according to specific circumstances and needs, and are usually jointly formulated and confirmed by professionals or relevant stakeholders. By reasonably setting the weighting coefficients, programs or systems can be helped to make decisions or predictions more accurately.

[0147] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

[0148] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A cross-language real-time speech recognition and pickup method, characterized in that, Includes the following steps: Multimodal sound pickup and data preprocessing, multimodal temporal alignment, cross-linguistic feature adaptation, noise-robust sound pickup, real-time speech recognition, and health monitoring and anomaly handling; The multimodal sound pickup and data preprocessing process involves acquiring microphone speech data and VAD sensor data through a multimodal sound pickup array, combining this with an environmental perception interface to obtain noise levels and the number of speakers, and then performing DC removal, pre-emphasis, and frame segmentation on the microphone data, while smoothing and filtering the VAD data. Multimodal timing alignment corrects timing errors through linear interpolation, transmission delay compensation, and the DTW algorithm. Cross-language feature adaptation achieves cross-language feature unification through language detection, feature extraction, and DAN network mapping. Noise-robust sound pickup suppresses residual noise through noise classification, beamforming algorithm selection, and adaptive filtering. Real-time speech recognition achieves real-time recognition and feeds back optimized parameters through a lightweight CNN-Transformer network. Health monitoring and anomaly handling ensure stable system operation through health score calculation and a tiered strategy. In the noise robust pickup process, noise classification uses a CNN-LSTM network. The input is a log-Mel spectrum. The CNN part contains two layers of convolution and pooling, and the LSTM part is a bidirectional LSTM layer. The output is the probability of four types of noise: steady-state, broadband, impulse, and multi-speaker. The beamforming algorithm is selected according to the noise type: delay-summation is used for steady-state noise, MVDR is used for impulse noise, and sparse beamforming is used for multi-speaker. In the adaptive filtering, ANF is used for steady-state noise, modified spectral subtraction is used for broadband noise, and VAD combined with speaker separation is used for multi-speaker interference.

2. The cross-language real-time speech recognition pickup method according to claim 1, characterized in that, In the pre-emphasis processing, a pre-emphasis coefficient is set for Chinese to highlight high-frequency unvoiced sounds, and a pre-emphasis coefficient is set for English to balance high-frequency and low-frequency energy; the frame segmentation processing uses a 20ms frame length and a 10ms frame shift, and uses a Hanning window function for weighting to reduce spectral leakage; the smoothing filtering of the VAD data uses a 5th-order Butterworth low-pass filter with a cutoff frequency of 100Hz to remove high-frequency noise.

3. The cross-language real-time speech recognition pickup method according to claim 1, characterized in that, In the multimodal timing alignment, linear interpolation adapts the VAD sensor data to the sampling rate; transmission delay compensation calculates the delay difference by measuring the physical length of the data transmission path between the microphone and the VAD sensor and the signal propagation speed, and adjusts the timestamp to ensure that the timing error between the microphone voice frame and the VAD start / stop signal does not exceed a set value; the DTW algorithm constructs a time series matrix of the microphone voice frame and the VAD signal, calculates the Euclidean distance and dynamically plans the optimal path, and adjusts the VAD signal time series to correct the timing error.

4. The cross-language real-time speech recognition pickup method according to claim 1, characterized in that, The language detection constructs a feature vector by extracting the fundamental frequency F0 and formant F1 features of the speech, and inputs it into a lightweight language classification network. This network is trained on a multilingual dataset containing Chinese, English, Japanese and other minority languages, and uses the cross-entropy loss function and the Adam optimizer.

5. The cross-language real-time speech recognition pickup method according to claim 1, characterized in that, In the cross-language feature adaptation, MFCC calculation is achieved through short-time Fourier transform, Mel filter bank, and discrete cosine transform; for Chinese focused tone features, the center frequency shift of filters 5-15 is adjusted; for English emphasized accent features, the center frequency of filters 12-25 is concentrated; for minority languages, the center frequency of the first 10 filters is optimized through a genetic algorithm; when a minority language is detected, the K halving strategy is triggered, K=512, and K=1024 is restored in the case of multi-speaker interference.

6. The cross-language real-time speech recognition pickup method according to claim 1, characterized in that, The number of neurons in the input layer of the DAN network matches the dimension of the source language-specific features. Hidden layer 1 contains 256 ReLU activated neurons, hidden layer 2 contains 128 ReLU activated neurons, and the output layer contains 64 linear activated neurons. The mapping matrix W is optimized through cross-language pre-training and online fine-tuning, and is jointly trained by combining MMD loss and classification loss.

7. The cross-language real-time speech recognition pickup method according to claim 1, characterized in that, The real-time speech recognition adopts frame-level pipeline processing. When the microphone acquires frame t, it preprocesses frame t-1 and recognizes frame t-2, with a delay not exceeding a set value. The health score is a weighted sum of device status indicators and processing quality indicators, with the weights determined by AHP. When the health score is less than the set value, a grading strategy is triggered: for mild abnormalities, the microphone gain is increased; for moderate abnormalities, the backup VAD sensor is switched; and for severe abnormalities, the single microphone + basic noise suppression mode is switched.

8. A cross-language real-time speech recognition pickup system, used to implement the cross-language real-time speech recognition pickup method as described in claim 1, characterized in that, The system includes a main control module, a sound pickup and perception module, a cross-language speech processing engine, and a real-time recognition and feedback module; The main control module, serving as the core scheduling hub, is deployed on an edge computing device and includes a synchronization scheduling unit, a resource allocation unit, and an instruction interaction unit. The synchronization scheduling unit achieves hardware clock synchronization through the PTP protocol, combined with software dynamic compensation, to control the synchronization error to not exceed a set value. The resource allocation unit dynamically allocates CPU / GPU computing power based on language type and environmental noise level, reducing CPU computing power for recognition tasks when the language has low resources, and reserving a set proportion of redundant computing power for noise suppression when the noise level exceeds a set value. The instruction interaction unit parses user instructions through natural language processing, issues language adaptation instructions, and provides feedback on sound pickup quality. The sound pickup and sensing module includes a 6-8 channel ring microphone array and 2 channels VAD sensors, along with a data preprocessing unit and an environmental sensing interface. The data preprocessing unit is used for DC removal, pre-emphasis, and frame processing. High-frequency compensation for voiceless consonants in Japanese increases the frequency band gain, and sliding window smoothing filtering is added for low signal-to-noise ratio scenarios. The environmental sensing interface is used to read noise levels and the number of speakers; The cross-language speech processing engine includes a language feature adaptation engine, a noise-robust pickup engine, and a multimodal temporal alignment engine. The real-time recognition and feedback module includes a lightweight CNN-Transformer network, which controls the end-to-end latency through a three-stage parallel pipeline and achieves closed-loop optimization in conjunction with a recognition quality evaluation unit. The three-stage parallel pipeline includes acquiring t frames, preprocessing t-1 frames, and recognizing t-2 frames.

Citation Information

Patent Citations

  • Target voice real-time separation method in multi-person voice environment

    CN116259331A

  • Voice recognition method and system of intelligent voice robot

    CN119517012A