A cloud low-delay speech recognition system and method fusing RNNoise denoising and SileroVAD
By employing RNNoise noise reduction and SileroVAD's cloud-based low-latency speech recognition method, the problems of low speech recognition accuracy and invalid data transmission in complex environments were solved, achieving high-precision, low-latency speech recognition and improving the quality and usability of the recognition results.
CN122392514APending Publication Date: 2026-07-14
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-14
Smart Images

Figure CN122392514A_ABST
Abstract
The present application relates to the technical field of audio processing, and in particular to a cloud low-delay speech recognition system and method fusing RNNoise noise reduction and Silero VAD, constructing an audio processing pipeline in the cloud, sequentially executing audio stream receiving and framing, real-time noise reduction based on RNNoise, saving historical audio by using a pre-buffer, performing robust voice activity detection by using Silero VAD with a delay confirmation mechanism, judging and selecting performance energy attenuation based on multiple features, and setting a silence merging window to realize intelligent merging of speech segments, and finally submitting complete semantic segments to an ASR engine for recognition. The speech quality and signal-to-noise ratio in a mixed environment of noise and far and near fields are significantly improved; invalid data transmission and computing resource consumption are effectively reduced, and the context coherence of the recognized text is ensured; at the same time, the strong dependence on network real-time performance is reduced, thereby realizing high-accuracy, low-delay and resource-efficient cloud speech recognition as a whole.
Need to check novelty before this filing date? Find Prior Art