A cloud low-delay speech recognition system and method fusing RNNoise denoising and SileroVAD

By employing RNNoise noise reduction and SileroVAD's cloud-based low-latency speech recognition method, the problems of low speech recognition accuracy and invalid data transmission in complex environments were solved, achieving high-precision, low-latency speech recognition and improving the quality and usability of the recognition results.

CN122392514APending Publication Date: 2026-07-14

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-04-15
Publication Date
2026-07-14

Smart Images

  • Figure CN122392514A_ABST
    Figure CN122392514A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audio processing, and in particular to a cloud low-delay speech recognition system and method fusing RNNoise noise reduction and Silero VAD, constructing an audio processing pipeline in the cloud, sequentially executing audio stream receiving and framing, real-time noise reduction based on RNNoise, saving historical audio by using a pre-buffer, performing robust voice activity detection by using Silero VAD with a delay confirmation mechanism, judging and selecting performance energy attenuation based on multiple features, and setting a silence merging window to realize intelligent merging of speech segments, and finally submitting complete semantic segments to an ASR engine for recognition. The speech quality and signal-to-noise ratio in a mixed environment of noise and far and near fields are significantly improved; invalid data transmission and computing resource consumption are effectively reduced, and the context coherence of the recognized text is ensured; at the same time, the strong dependence on network real-time performance is reduced, thereby realizing high-accuracy, low-delay and resource-efficient cloud speech recognition as a whole.
Need to check novelty before this filing date? Find Prior Art