WebRTC voice enhancement system and method based on multi-modal large model
Patent Information
- Application Number
- CN202511134961.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-08-14
AI Technical Summary
[0010]本发明的技术任务是针对以上不足之处,提供基于多模态大模型的WebRTC语音增强系统及方法,在保障语音语义完整性的同时,实现高精度噪声抑制与毫秒级延迟,满足工业巡检、车载通信等场景对高保真、低延迟、强鲁棒性的需求
[0039]1、多模态信息融合提高在低信噪比环境下的语音清晰度:本发明通过融合视觉唇动特征、音频信号及语义文本信息,在信号微弱或受噪声干扰的极端环境中,仍可重构清晰语音内容。尤其是在传统音频增强方法易丢失语义或误判语音边界的情况下,视觉与文本模态有效弥补音频模态不足,实现语义完整性更强、噪声误杀更少的语音增强效果,显著提升了低信噪比通信质量。
Smart Images

Figure CN121011196B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and real-time communication, specifically to a WebRTC speech enhancement system and method based on a multimodal large model. Background Technology
[0002] WebRTC (Web Real-Time Communication), as a core open standard supporting real-time audio and video interaction on browsers, enables peer-to-peer audio and video streaming through a JavaScript API, achieving low-latency communication in browsers and mobile devices without the need for plugins. Its core value lies in end-to-end encrypted transmission, cross-platform compatibility, and resilience to weak network conditions, and it has been widely applied in scenarios such as online conferencing and telemedicine. However, in complex environments, traditional audio signal processing-based speech enhancement technologies suffer from significant performance bottlenecks when faced with sudden noise and low signal-to-noise ratios.
[0003] Multimodal Large Language Model (MLLM) integrates visual, speech, and textual modal information, enabling enhanced environmental understanding and semantic reconstruction capabilities. Its core capabilities include cross-modal representation unification (mapping images, audio, and text to a shared latent space) and dynamic knowledge modeling (combining retrieval-enhanced generation mechanisms to achieve real-time adaptation to unknown noise patterns in the environment).
[0004] Although some research has attempted to incorporate visual or semantic information into the speech enhancement process, the following problems still exist in practical applications:
[0005] (1) Shallow use of visual information: Most solutions only use visual information for speech activity detection (VAD) and lack a mechanism for joint modeling of lip movement and audio, resulting in unstable performance under conditions of limited or occluded camera view.
[0006] (2) Lack of semantic restoration mechanism: Existing noise reduction methods are unable to fully utilize the semantic information of the text, and often suffer from semantic false killing (such as loss of conference terms or emotional information), which is particularly evident in low signal-to-noise ratio environments.
[0007] (3) The problem of static noise library is prominent: Most systems rely on preset templates or static modeling, which makes it difficult to dynamically adapt to new environmental noises, especially in dynamic sound fields such as industrial inspection and vehicle environment.
[0008] (4) Lack of multimodal decision-making mechanism: Current technology lacks a unified decision-making mechanism that integrates audio, visual and semantic modalities. It cannot adaptively switch strategies when faced with modal loss (such as camera failure), and the robustness of the system is limited.
[0009] Therefore, the industry urgently needs an intelligent voice enhancement system that integrates multimodal perception and dynamic noise knowledge management, taking into account accuracy, timeliness and privacy protection, in order to achieve high-quality real-time communication in the WebRTC environment. Summary of the Invention
[0010] The technical objective of this invention is to address the above-mentioned shortcomings by providing a WebRTC speech enhancement system and method based on a multimodal large model. This system ensures the integrity of speech semantics while achieving high-precision noise suppression and millisecond-level latency, meeting the requirements of high fidelity, low latency, and strong robustness in scenarios such as industrial inspection and vehicle communication.
[0011] The technical solution adopted by this invention to solve its technical problem is:
[0012] A WebRTC speech enhancement system based on a multimodal large model, comprising:
[0013] The audio and video acquisition module is used to synchronously acquire the user's original voice signal and corresponding video image data through the WebRTC protocol stack, and achieve high-precision alignment through timestamp marking and caching mechanism;
[0014] The multimodal feature extraction module is used to extract audio features from the speech signal, visual lip movement features from the video image, and generate text semantic features through the speech recognition engine.
[0015] The noise matching and update module is used to extract the embedding vector based on the currently collected audio data and match it with the preset noise fingerprint database. When the matching confidence is lower than the threshold, the federated learning mechanism is triggered to update the noise template.
[0016] The multimodal semantic perception enhancement module includes a multimodal modeling network with a Transformer structure, which uses an attention mechanism to fuse the above three types of modal information and enhance speech.
[0017] The audio reconstruction module is used to reconstruct waveform audio from the enhanced spectrogram through inverse transformation;
[0018] The WebRTC integration module is used to insert enhanced audio into the WebRTC audio processing chain and transmit it back to the remote end in real time via the RTP protocol.
[0019] This system integrates multimodal information such as visual lip movement features, audio streams, and semantic text. Combined with a dynamic noise knowledge base and lightweight edge-based model deployment, it achieves higher precision noise suppression and preservation of speech semantic integrity, making it particularly suitable for real-time audio and video communication in low signal-to-noise ratio (SNR) scenarios. The system mainly includes multimodal collaborative input, a dynamic noise retrieval and suppression mechanism, and a deployment strategy adapted to edge computing environments. Compared with traditional methods, this system has significant advantages in handling non-steady-state noise, semantic corruption under low SNR, and cloud communication latency. By improving noise reduction through multimodal fusion and generative modeling, and enhancing environmental adaptability by combining dynamic noise updates and reinforcement learning mechanisms, it ultimately achieves a low-latency, privacy-friendly integrated enhancement system on edge devices.
[0020] Furthermore, the audio and video acquisition module uses the getUserMedia API interface to acquire audio and video streams, with an audio sampling rate of 48kHz, a video resolution of no less than 720p, a frame rate of no less than 30fps, and achieves audio and video frame-level time alignment through frame interpolation and resampling algorithms.
[0021] Furthermore, the audio features include Mel frequency cepstral coefficients (MFCC) and log-Mel spectrogram, obtained through STFT and Mel filter, with 40-dimensional features extracted per frame; the visual features are obtained by motion encoding of the lip ROI region based on a 3D convolutional network; and the text features are generated by a lightweight ASR engine.
[0022] Furthermore, the noise matching and updating module includes:
[0023] A noise fingerprint database stores the spectral embedding vectors of multiple known environmental noises;
[0024] The similarity matching unit uses cosine similarity to perform nearest neighbor retrieval and determine the noise category;
[0025] The federated learning update unit is used to upload noise features to the cloud and perform unsupervised clustering to generate new noise templates when the matching confidence is lower than a set threshold.
[0026] Furthermore, the Transformer network of the multimodal semantic perception enhancement module includes audio, visual and text branches, and achieves feature alignment and fusion through a gated cross-attention mechanism to enhance speech spectrum quality.
[0027] Furthermore, the audio reconstruction module includes inverse Mel transform and inverse short-time Fourier transform (iSTFT), and combines sliding window and overlapping windowing strategies to reconstruct the speech waveform, so as to maintain the continuity and naturalness of the speech signal.
[0028] Furthermore, the WebRTC integration module embeds the voice enhancement module as a custom AudioProcessor plugin into the WebRTC audio processing link, and uses shared memory and asynchronous queue mechanisms to achieve concurrent communication with the main thread.
[0029] Furthermore, the system supports deployment based on the ONNX model format and combines GPU or edge TPU for inference acceleration, so that the overall latency of voice enhancement does not exceed 200ms.
[0030] Furthermore, the system includes a multimodal fault-tolerant mechanism that automatically switches to audio-text dual-modal enhancement mode when the camera is disconnected or video frames are lost, and dynamically adjusts the enhancement strategy based on reinforcement learning algorithms.
[0031] This invention also claims a WebRTC speech enhancement method based on a multimodal large model, which is implemented based on the above system and includes the following steps:
[0032] Step S1: Acquire synchronized audio and video streams and perform time alignment;
[0033] Step S2: Extract features from three modalities: audio, visual, and text.
[0034] Step S3: Match noise categories based on the fingerprint database, and update the noise template through federated learning when a match fails;
[0035] Step S4: Use a multimodal Transformer model for semantic awareness modeling and speech enhancement;
[0036] Step S5: Reconstruct the enhanced spectrogram into an audio waveform;
[0037] Step S6: Embed the enhanced voice into the WebRTC processing chain and transmit it back to the remote receiver in real time via the RTP channel.
[0038] Compared with existing technologies, the WebRTC speech enhancement system and method based on a multimodal large model of the present invention have the following advantages:
[0039] 1. Multimodal information fusion improves speech clarity in low signal-to-noise ratio environments: This invention, by fusing visual lip movement features, audio signals, and semantic text information, can still reconstruct clear speech content in extreme environments with weak signals or noise interference. Especially when traditional audio enhancement methods are prone to losing semantics or misjudging speech boundaries, the visual and text modalities effectively compensate for the deficiencies of the audio modalities, achieving a speech enhancement effect with stronger semantic integrity and fewer false positives from noise, significantly improving the quality of low signal-to-noise ratio communication.
[0040] 2. Dynamic noise knowledge base enables environmental adaptation: To address the problem that static noise templates cannot adapt to new noise types, the system identifies unknown noise through real-time spectrum embedding vector matching (cosine similarity threshold 0.75) and triggers a federated learning mechanism: After the noise fingerprint is anonymously uploaded at the edge, unsupervised clustering is performed in the cloud to generate a new template.
[0041] 3. Multimodal fault tolerance mechanism enhances system robustness: Existing systems experience a sharp performance drop when visual signals are interrupted. This solution designs a multimodal decision engine: when camera failure is detected, it automatically switches to audio-semantic dual-modal mode and dynamically adjusts the noise reduction strategy weights through reinforcement learning; if visual recovery occurs, it seamlessly switches back to trimodal fusion, which can significantly improve the reliability of equipment failure-sensitive scenarios such as industrial inspection.
[0042] 4. Supports edge inference deployment, reducing communication latency and cloud dependence: This invention optimizes the model architecture and inference chain for edge devices, supporting lightweight deployment in ONNX format and GPU / TPU acceleration, with typical inference latency controlled within 200ms. Compared to traditional methods that require uploading audio to the cloud for processing, the system can directly complete enhancement tasks on the terminal, significantly reducing communication latency and cloud resource consumption, while meeting privacy protection and local processing requirements, making it more suitable for scenarios with strict requirements for timeliness and data security.
[0043] 5. Deep integration with WebRTC, supporting end-to-end real-time voice enhancement transmission: This invention embeds a custom AudioProcessor plugin into the WebRTC audio processing chain, achieving seamless integration between the enhancement module and the WebRTC framework. The system ensures real-time interaction of the audio stream through asynchronous queues and shared memory technology. Enhanced voice can directly enter the RTP transmission link, avoiding additional latency and data copying overhead. Compared to external enhancement modules, the overall architecture is more compact, with better latency control and system-level compatibility, making it suitable for various real-time communication scenarios. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the overall structure of the WebRTC speech enhancement system based on a multimodal large model provided in an embodiment of the present invention;
[0045] Figure 2 This is a diagram illustrating the multimodal feature extraction structure provided in an embodiment of the present invention. Detailed Implementation
[0046] This invention provides a WebRTC speech enhancement system based on a multimodal large model. The system includes:
[0047] The audio and video acquisition module is used to synchronously acquire the user's original voice signal and corresponding video image data through the WebRTC protocol stack, and achieve high-precision alignment through timestamp marking and caching mechanism;
[0048] The multimodal feature extraction module is used to extract audio features from the speech signal, visual lip movement features from the video image, and generate text semantic features through the speech recognition engine.
[0049] The noise matching and update module is used to extract the embedding vector based on the currently collected audio data and match it with the preset noise fingerprint database. When the matching confidence is lower than the threshold, the federated learning mechanism is triggered to update the noise template.
[0050] The multimodal semantic perception enhancement module includes a multimodal modeling network with a Transformer structure, which uses an attention mechanism to fuse the above three types of modal information and enhance speech.
[0051] The audio reconstruction module is used to reconstruct waveform audio from the enhanced spectrogram through inverse transformation;
[0052] The WebRTC integration module is used to insert enhanced audio into the WebRTC audio processing chain and transmit it back to the remote end in real time via the RTP protocol.
[0053] The audio and video acquisition module uses the getUserMedia API interface to acquire audio and video streams. The audio sampling rate is 48kHz, the video resolution is no less than 720p, the frame rate is no less than 30fps, and the audio and video frame-level time alignment is achieved through frame interpolation and resampling algorithms.
[0054] The audio features include Mel frequency cepstral coefficients (MFCC) and log-Mel spectrogram, obtained through STFT and Mel filter, with 40-dimensional features extracted per frame; the visual features are obtained by motion coding of the lip ROI region based on a 3D convolutional network; and the text features are generated by a lightweight ASR engine.
[0055] The noise matching and updating module includes:
[0056] A noise fingerprint database stores the spectral embedding vectors of multiple known environmental noises;
[0057] The similarity matching unit uses cosine similarity to perform nearest neighbor retrieval and determine the noise category;
[0058] The federated learning update unit is used to upload noise features to the cloud and perform unsupervised clustering to generate new noise templates when the matching confidence is lower than a set threshold.
[0059] The Transformer network of the multimodal semantic perception enhancement module includes audio, visual and text branches, and achieves feature alignment and fusion through a gated cross-attention mechanism to enhance speech spectrum quality.
[0060] The audio reconstruction module includes inverse Mel transform and inverse short-time Fourier transform (iSTFT), and combines sliding window and overlapping windowing strategies to reconstruct the speech waveform in order to maintain the continuity and naturalness of the speech signal.
[0061] The WebRTC integration module embeds the voice enhancement module as a custom AudioProcessor plugin into the WebRTC audio processing chain, and uses shared memory and asynchronous queue mechanisms to achieve concurrent communication with the main thread.
[0062] The system supports deployment based on the ONNX model format and combines GPU or edge TPU for inference acceleration, so that the overall latency of voice enhancement does not exceed 200ms.
[0063] The system includes a multimodal fault-tolerant mechanism. When the camera is disconnected or video frames are lost, it automatically switches to an audio-text dual-modal enhancement mode and dynamically adjusts the enhancement strategy based on a reinforcement learning algorithm.
[0064] This system addresses the issues of instability and poor generalization in traditional speech enhancement algorithms under low signal-to-noise ratio and complex scenarios. By fusing visual lip movement features, audio signals, and semantic text information, and introducing dynamic noise matching and incremental update mechanisms, combined with a multimodal Transformer semantic-aware modeling network, it significantly improves audio enhancement quality while ensuring real-time performance. Deeply integrated with WebRTC, the system can be widely applied to high real-time audio and video communication scenarios such as online conferencing, distance education, and customer call centers.
[0065] Combination Figure 1 and Figure 2 As shown, the specific implementation of this system is as follows.
[0066] Step S1: Synchronous acquisition of audio and video data.
[0067] In this step, the system utilizes the WebRTC protocol stack to synchronously acquire the user's raw audio and corresponding video images, ensuring time alignment accuracy to facilitate subsequent joint modeling of multimodal features.
[0068] (1) Standardized data collection method:
[0069] The system uses the getUserMedia API in the WebRTC standard acquisition interface to obtain the microphone audio stream (PCM encoded) and the video stream (30fps or higher, 720p or 1080p) captured by the camera from the terminal device. Internally, the system employs a synchronous timestamp marking and audio / video frame buffering mechanism to ensure the alignment of lip movement images with corresponding speech segments in the time dimension.
[0070] (2) Time alignment:
[0071] To further ensure the accuracy of multimodal modeling, a frame rate matching module and an audio-video synchronization optimization algorithm are integrated. By adjusting the rhythm and step size of the two types of data through video frame interpolation and audio resampling mechanisms, contextual modeling based on synchronized input can be performed during the semantic modeling stage. Assume the audio signal sampling is A(t), and the video frame timestamp is t. v The linear alignment interpolation form is as follows:
[0072]
[0073] Step S2: Multimodal feature extraction.
[0074] This step extracts three types of features from the collected data: audio, video, and text semantic features, which are then used for subsequent fusion in a large model.
[0075] (1) Audio feature extraction:
[0076] In the audio channel, the system employs a stacked CNN network or CRNN architecture based on residual connections to extract the time-spectral features of the audio. Features such as Mel-frequency cepstral coefficients (MFCC) and log-Mel spectrograms are extracted from the speech stream, where the Mel frequency transformation formula is:
[0077]
[0078] Each audio frame (25ms window, 10ms frame shift) will have 13-dimensional MFCC and 40-dimensional Mel features extracted and concatenated to form a vector sequence. The original audio signal will be processed by pre-emphasis, short-time Fourier transform (STFT), and Mel filter to form a Mel spectrogram.
[0079] (2) Visual lip movement feature extraction:
[0080] In the video channel, the lip movement image sequence first undergoes face detection and mouth region ROI cropping, and then is input into a 3D convolutional-based visual coding network to extract lip movement features. This feature encoding considers temporal continuity and spatial dynamic changes, preserving clear pronunciation contour features.
[0081] (3) Textual semantic features:
[0082] By integrating a lightweight ASR engine, real-time speech is initially transcribed to extract semantic text sequences. The lightweight ASR engine can quickly convert speech into text and can adapt to different accents and speech rates to a certain extent, providing textual basis for subsequent semantic perception enhancement modeling.
[0083] Step S3: Dynamic noise library matching and incremental update.
[0084] This step senses the current type of ambient noise in real time and dynamically updates the noise template library through a federated learning mechanism, thereby improving the system's ability to adapt to non-static noise.
[0085] (1) Matching with the noisy fingerprint database:
[0086] The system constructs a pre-built noise fingerprint library containing 200 types of typical environmental background noise (such as traffic noise, fan noise, keyboard noise, equipment operation noise, etc.). Each type of noise is processed by a residual spectrum encoder to extract its spectrum embedding vector and stored in the vector database.
[0087] After the system acquires a real-time audio segment, it extracts its corresponding embedding vector and performs nearest neighbor matching with all vectors in the template library. The matching similarity is calculated using cosine similarity. If the maximum matching confidence is lower than a preset threshold of 0.75, it is considered as unknown noise.
[0088] (2) Incremental Federated Learning:
[0089] To address the issue of delayed template library updates, the system initiates an incremental federated learning process: edge terminals anonymously encrypt unknown noise fingerprints and upload them to the cloud. The cloud then performs unsupervised clustering on similar unknown fingerprints, automatically generating new noise templates, which are then asynchronously pushed to each device after being updated.
[0090] Step S4: Semantic awareness enhancement modeling.
[0091] The speech enhancement context modeling and noise removal are based on a multimodal large model architecture. The overall network adopts a Transformer structure and is divided into two parts: encoder and decoder.
[0092] The encoder first models the temporal contextual relationships of audio, video, and semantic features using a self-attention mechanism. For example, for audio features, the self-attention mechanism can capture the correlation between different time points in the speech, thus better understanding the semantic information of the speech. Subsequently, cross-modal interaction is achieved through a multimodal attention fusion layer, namely, visual-guided audio alignment and semantic-guided voiceprint enhancement. The specific implementation includes a multimodal cross-attention and feature alignment layer based on a gating mechanism. For example, when visual lip movement features indicate that the user is speaking a specific syllable, the system can align the audio features with it through a multimodal cross-attention mechanism, thereby more accurately enhancing the speech signal of that syllable.
[0093] The decoder section reconstructs the spectrum of the modeling output using a deconvolutional network or a Transformer decoder, ultimately generating a clear speech spectrum to restore the denoised audio. For example, the feature vector processed by the encoder can be deconvolved by the decoder to reconstruct a clearer and cleaner speech spectrum, thereby improving speech quality.
[0094] Step S5: Large model inference and speech enhancement output.
[0095] (1) Audio restoration and completion:
[0096] The system reconstructs the enhanced spectrogram output by the model into an enhanced waveform signal through inverse Mel-Transform and inverse STFT (iSTFT). During the reconstruction process, the system strictly follows the conversion process from Mel-Transform to time-domain signal to ensure that the reconstructed audio signal is as close as possible to the original speech signal in waveform. Simultaneously, the system employs a sliding window and overlapping windowing strategy to ensure output continuity and smoothness, avoiding auditory discomfort caused by inter-frame jumps. For example, during continuous playback of the speech signal, the sliding window and overlapping windowing strategy ensure a natural transition between each audio frame, without noticeable stutters or abrupt switches.
[0097] (2) Inference acceleration and delay control:
[0098] To reduce latency, the system supports model deployment based on the ONNX inference framework and utilizes GPUs or edge TPUs to accelerate the inference process, keeping the overall inference latency below 200ms to meet the real-time interactive requirements of WebRTC. The final output audio stream will be used as enhanced audio for the WebRTC module to transmit back or play.
[0099] Step S6: WebRTC Real-time Integration and Backhaul Control.
[0100] (1) Enhanced module integration:
[0101] The enhancement module is embedded as a custom AudioProcessor plugin in the WebRTC audio processing chain, replacing the original audio stream or inserting it after the original signal. The system uses shared memory and asynchronous queue mechanisms to achieve bidirectional flow of audio data, ensuring that the enhancement module does not block the WebRTC main thread.
[0102] (2) Sound quality adaptive strategy:
[0103] To adapt to network bandwidth fluctuations, the system incorporates an adaptive audio quality strategy. When bandwidth decreases or the system detects increased latency, it automatically switches to a lightweight model or a simplified modeling channel. For example, in poor network conditions, the system can switch to a lightweight speech enhancement model; although the enhancement effect may be slightly reduced, it ensures smooth voice communication. Once network conditions return to normal, the system can automatically switch back to the full model, providing high-quality speech enhancement.
[0104] (3) Real-time transmission:
[0105] The system ultimately transmits the enhanced audio to the remote receiver in real time via WebRTC's RTP channel, achieving end-to-end low-latency, robust voice communication. The RTP channel provides an efficient transmission mechanism for audio data, ensuring that audio data reaches the remote receiver quickly and accurately, thus achieving high-quality real-time voice communication. For example, in online meetings, the enhanced audio is transmitted to each participant's terminal device via the RTP channel, allowing each participant to clearly hear the speaker's voice.
[0106] This system addresses the core bottlenecks faced by traditional WebRTC speech enhancement technology in complex acoustic environments: failure to suppress burst noise, low signal-to-noise ratio and semantic impairment, and excessively high cloud processing latency. By fusing multimodal information from visual lip movement features, audio streams, and semantic text, and combining a dynamic noise knowledge base with lightweight edge model deployment, a real-time speech enhancement system with environmental adaptability is constructed. This system achieves high-precision noise suppression and millisecond-level latency while ensuring the integrity of speech semantics, meeting the stringent requirements of high fidelity, low latency, and strong robustness in scenarios such as industrial inspection and vehicle communication.
[0107] This invention also provides a WebRTC speech enhancement method based on a multimodal large model. This method is implemented based on the WebRTC speech enhancement system based on a multimodal large model described in the above embodiments, and the implementation of this method includes the following steps:
[0108] Step S1: Acquire synchronized audio and video streams and perform time alignment;
[0109] Step S2: Extract features from three modalities: audio, visual, and text.
[0110] Step S3: Match noise categories based on the fingerprint database, and update the noise template through federated learning when a match fails;
[0111] Step S4: Use a multimodal Transformer model for semantic awareness modeling and speech enhancement;
[0112] Step S5: Reconstruct the enhanced spectrogram into an audio waveform;
[0113] Step S6: Embed the enhanced voice into the WebRTC processing chain and transmit it back to the remote receiver in real time via the RTP channel.
[0114] This method integrates multimodal information such as visual lip movement features, audio streams, and semantic text. Combined with a dynamic noise knowledge base and lightweight edge model deployment, it achieves higher precision noise suppression and preservation of speech semantic integrity, making it particularly suitable for real-time audio and video communication in low signal-to-noise ratio (SNR) scenarios. Compared to traditional methods, the system exhibits significant advantages in handling non-stationary noise, semantic corruption under low SNR, and cloud communication latency. By enhancing noise reduction through multimodal fusion and generative modeling, and improving environmental adaptability through dynamic noise updates and reinforcement learning mechanisms, it ultimately achieves a low-latency, privacy-friendly integrated enhancement system on edge devices.
[0115] The use of multimodal fusion and generative models effectively improves noise suppression accuracy. It also ensures robustness in complex environments through dynamic noise databases and reinforcement learning mechanisms. Meanwhile, WebRTC's real-time integration design and adaptive strategies can achieve a balance between low latency and privacy security.
[0116] Through the specific embodiments described above, those skilled in the art can easily implement the present invention. However, it should be understood that the present invention is not limited to the specific embodiments described above. Based on the disclosed embodiments, those skilled in the art can arbitrarily combine different technical features to achieve different technical solutions.
[0117] Except for the technical features described in the specification, all other technologies are known to those skilled in the art.
Claims
1. A WebRTC speech enhancement system based on a multimodal large model, characterized in that, The system includes: The audio and video acquisition module is used to synchronously acquire the user's original voice signal and corresponding video image data through the WebRTC protocol stack, and achieve high-precision alignment through timestamp marking and caching mechanism; The multimodal feature extraction module is used to extract audio features from the speech signal, visual lip movement features from the video image, and generate text semantic features through the speech recognition engine. The noise matching and updating module includes: a noise fingerprint database for storing spectral embedding vectors obtained by processing multiple known environmental noises using a residual spectrum encoder; a similarity matching unit for extracting the embedding vector of the currently acquired audio using the residual spectrum encoder, calculating the cosine similarity between the embedding vector and each spectral embedding vector in the noise fingerprint database, and determining the matching confidence and noise category based on the maximum cosine similarity; and a noise template collaborative updating unit for anonymously encrypting and uploading the unknown noise fingerprint corresponding to the currently acquired audio to the cloud when the matching confidence is lower than a set threshold, so that the cloud can perform unsupervised clustering on similar unknown noise fingerprints to generate new noise templates, and asynchronously push the new noise templates to each device to update the noise fingerprint database of each device. The updated noise fingerprint database is used for nearest neighbor retrieval of subsequent acquired audio. The multimodal semantic perception enhancement module includes a multimodal modeling network with a Transformer structure, comprising an audio branch, a visual branch, and a text branch. The audio, visual, and text branches respectively model the temporal context of their respective modalities using a self-attention mechanism, resulting in audio context features, visual context features, and text context features. A gated cross-attention mechanism takes the audio context features, visual context features, and text context features as input, performs cross-modal feature alignment and fusion, and outputs enhanced features for reconstructing the target speech spectrum to improve speech spectrum quality. The audio reconstruction module is used to reconstruct waveform audio from the enhanced spectrogram through inverse transformation; The WebRTC integration module is used to insert the waveform audio reconstructed by the audio reconstruction module into the WebRTC audio processing link through a custom AudioProcessor plugin. It uses shared memory to pass the audio to be enhanced and the enhanced audio, and uses an asynchronous queue to enable the voice enhancement processing to be executed concurrently with the WebRTC main thread. The enhanced audio is then transmitted back to the remote end in real time via the RTP protocol.
2. The WebRTC speech enhancement system based on a multimodal large model according to claim 1, characterized in that, The audio and video acquisition module uses the getUserMedia API interface to acquire audio and video streams. The audio sampling rate is 48kHz, the video resolution is no less than 720p, the frame rate is no less than 30fps, and the audio and video frame-level time alignment is achieved through frame interpolation and resampling algorithms.
3. The WebRTC speech enhancement system based on a multimodal large model according to claim 1, characterized in that, The audio features include Mel frequency cepstral coefficients and logarithmic power spectrum, obtained through STFT and Mel filter. For each frame of audio, 13-dimensional Mel frequency cepstral coefficient features and 40-dimensional Mel features are extracted and concatenated to form an audio feature vector sequence. The visual lip movement features are obtained by extracting motion encoding of the lip ROI region based on a 3D convolutional network. The text semantic features are generated by a lightweight ASR engine.
4. The WebRTC speech enhancement system based on a multimodal large model according to claim 1, characterized in that, The audio reconstruction module includes inverse Mel transform and inverse short-time Fourier transform, and combines sliding window and overlapping windowing strategies to reconstruct the speech waveform in order to maintain the continuity and naturalness of the speech signal.
5. The WebRTC speech enhancement system based on a multimodal large model according to claim 1, characterized in that, The system supports deployment based on the ONNX model format and combines GPUs or edge TPUs for inference acceleration, ensuring that the model inference latency of the multimodal modeling network does not exceed 200ms.
6. The WebRTC speech enhancement system based on a multimodal large model according to claim 1, characterized in that, The system includes a multimodal fault tolerance mechanism. When a camera disconnection or video frame loss is detected, the visual branch is blocked and the system switches to an audio-text dual-modal enhancement mode. The noise reduction strategy weights are dynamically adjusted based on a reinforcement learning algorithm. When visual input is detected to be restored, the audio, visual, and text trimodal enhancement modes are restored.
7. A WebRTC speech enhancement method based on a multimodal large model, characterized in that, The method includes the following steps: Step S1: Synchronously collect the user's original voice signal and corresponding video image data through the WebRTC protocol stack, and achieve high-precision alignment through timestamp marking and caching mechanism; Step S2: Extract audio features from the speech signal, extract visual lip movement features from the video image, and generate text semantic features through the speech recognition engine; Step S3: Extract the embedding vector of the currently acquired audio using the residual spectrum encoder, calculate the cosine similarity between the embedding vector and each spectrum embedding vector in the noise fingerprint database, and determine the noise category based on the maximum cosine similarity; when the matching confidence corresponding to the maximum cosine similarity is lower than a set threshold, anonymously encrypt and upload the unknown noise fingerprint to the cloud, and the cloud performs unsupervised clustering on similar unknown noise fingerprints to generate new noise templates, and asynchronously pushes the new noise templates to each device to update the noise fingerprint database; Step S4: Use a multimodal Transformer model for semantic perception modeling and speech enhancement, specifically including: modeling the temporal context of the corresponding modality through self-attention mechanisms in the audio branch, visual branch, and text branch respectively, to obtain audio context features, visual context features, and text context features; and performing cross-modal alignment and fusion of the three context features through a gated cross-attention mechanism to output enhanced features for reconstructing the spectrum of the target speech. Step S5: Reconstruct the audio waveform from the enhanced spectrogram using inverse transform; Step S6: Insert the reconstructed waveform audio into the WebRTC audio processing link using a custom AudioProcessor plugin. Use shared memory to pass the audio to be enhanced and the enhanced audio, and use an asynchronous queue to enable the voice enhancement processing to be executed concurrently with the WebRTC main thread. Send the enhanced audio back to the remote receiving end in real time through the RTP channel.
Citation Information
Patent Citations
Speech enhancement method and device, electronic equipment and computer readable storage medium
CN114333863A