A conference large-screen audio and video self-optimization method and system based on deep learning
By using domain-specific parameter acquisition and multimodal fusion networks, combined with edge-cloud collaborative computing, the problems of insufficient real-time performance and strong network dependence in existing audio and video optimization solutions are solved, achieving highly robust audio and video self-optimization effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN QICHANG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-04-21
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies suffer from problems such as insufficient real-time performance, weak cross-modal collaboration capabilities, and over-reliance on preset strategies in complex acoustic environments, dynamic lighting conditions, and network fluctuation scenarios. They are difficult to adapt to dynamic environmental changes, have strong network dependence, and result in significant fluctuations in synchronization and quality.
It employs domain-specific parameter acquisition, hierarchical cross-modal analysis, and dynamic strategy generation. By combining the spatiotemporal features of audio and video with a multimodal fusion network (MFN) to mine semantic correlations, it achieves edge-cloud collaborative computing, reduces reliance on preset rules, and supports personalized optimization.
It achieves highly robust audio and video self-optimization, reduces latency and resource consumption, improves voice signal-to-noise ratio, image stability and user satisfaction, and adapts to real-time optimization needs in complex environments.
Smart Images

Figure CN120358387B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent deep optimization of audio and video, and specifically discloses a method and system for self-optimization of audio and video on a large conference screen based on deep learning. Background Technology
[0002] With the increasing prevalence of remote collaboration and hybrid on-site work models, the audio and video quality of conferencing systems has become a core factor affecting communication efficiency. Existing technologies for audio and video optimization mostly focus on single-modal enhancement or fixed-scene adaptation. However, in complex acoustic environments, dynamic lighting conditions, and network fluctuation scenarios, problems remain, including insufficient real-time performance, weak cross-modal collaboration capabilities, and over-reliance on preset strategies. The following analysis, based on relevant patent technologies, examines the limitations of existing solutions:
[0003] CN111866439A (Conference device, system, and operation method for optimizing audio and video experience)
[0004] This solution suppresses echoes through microphone array beamforming technology and optimizes the physical layout of audio devices to reduce interference, significantly improving the naturalness of full-duplex calls. However, its core reliance on hardware topology design (such as symmetrically distributed microphones and speakers) results in high deployment costs in non-fixed scenarios (such as mobile conferencing and open office environments), and the algorithm lacks robustness to sudden noise (such as keystrokes and temporary device access). Furthermore, its video module only reduces distortion by installing the camera at a low position and does not address adaptive processing for dynamic lighting or image jitter.
[0005] CN118075418A (Method, Apparatus, Equipment and Storage Medium for Optimizing Video Conferencing Content Output)
[0006] This invention proposes an AI-based method for optimizing the entire audio and video workflow. Through preprocessing, detection of items to be optimized, and secondary verification, it ensures output quality, making it particularly suitable for complex scenarios combining outdoor monitoring and conferencing. However, its drawbacks are:
[0007] Relying on preset optimization strategies and fixed thresholds makes it difficult to cope with dynamic changes in environmental parameters (such as sudden changes in lighting or multiple people speaking at the same time).
[0008] Deep learning models (such as CNN+LSTM) have a high computational load, which limits their real-time performance and makes them unable to meet the requirements of low-latency meetings.
[0009] Lacking cross-modal collaborative analysis capabilities, audio and video streams are only aligned by timestamps, without deeply exploring the semantic correlation between sound and image.
[0010] CN118214825A (Integrated audio and video equipment and system for conferencing)
[0011] This solution integrates online and offline meeting functions through a modular design, supports one-click switching and multi-device interconnection, and utilizes automatic tracking cameras and omnidirectional microphones for adaptive optimization. However, its limitations are as follows:
[0012] It is highly dependent on network stability, and audio and video synchronization and quality fluctuate significantly in weak network environments;
[0013] Overly high system integration leads to complex troubleshooting, upgrades, and maintenance.
[0014] The optimization strategy is rule-driven (such as fixed noise reduction intensity) and lacks the ability to dynamically learn from users' personalized needs.
[0015] In summary, existing technologies share the following common shortcomings:
[0016] Insufficient scene adaptability: Fixed hardware layouts or preset strategies are unable to cope with dynamic environmental changes;
[0017] Lack of cross-modal collaboration: Audio and video optimization is carried out in isolation, failing to fully utilize the correlation between audio and video to improve the overall experience;
[0018] The conflict between real-time performance and resource efficiency: complex model calculations lead to latency, while simplification strategies sacrifice optimization results;
[0019] Strong network dependence: Cloud processing and multi-route architecture have poor stability in weak network scenarios. Summary of the Invention
[0020] To address the aforementioned issues, this invention discloses a deep learning-based method and system for self-optimizing audio and video on large-screen conference displays. Through domain-specific parameter acquisition, hierarchical cross-modal analysis, and dynamic strategy generation, it achieves the following breakthroughs: a time-sharing and partitioned acquisition mechanism: differentiated processing of acoustic and video parameters to improve data effectiveness; a multimodal fusion network (MFN): combining audio-visual spatiotemporal features to mine semantic correlations and enhance environmental adaptability; edge-cloud collaborative computing: balancing latency and accuracy through local real-time processing and cloud-based model iteration; and reinforcement learning-driven dynamic control: reducing reliance on preset rules and achieving personalized optimization.
[0021] This solution systematically addresses the shortcomings of existing technologies, such as insufficient real-time performance, rigid scenarios, and strong network dependence, providing a highly robust audio and video self-optimization solution for hybrid office scenarios.
[0022] This invention includes the following technical solutions:
[0023] A deep learning-based method for self-optimization of audio and video on large-screen conferencing displays includes the following steps:
[0024] Step S1: Collect multimodal parameters by domain, including:
[0025] The first set of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time (RT60), and sound source direction angle;
[0026] The second set of parameters: video dynamic parameters, including dynamic range of illumination intensity, image motion vector, and facial key point displacement;
[0027] The first set of parameters and the second set of parameters are collected independently through a time-division sampling mechanism, and the sampling frequency of the first set of parameters is 1.5-2 times that of the second set of parameters;
[0028] Step S2: Perform hierarchical progressive analysis on the multimodal parameters, including:
[0029] First-level analysis: Generate a noise suppression weight matrix based on acoustic parameters and calculate the reverberation cancellation coefficient;
[0030] Second-level analysis: Based on video dynamic parameters, the image jitter features are extracted by optical flow method, and the image stability score is generated by combining the displacement of facial key points.
[0031] Step S3: Integrate acoustic and video analysis results, and perform cross-modal correlation assessment using a deep learning model, including:
[0032] Input the noise suppression weight matrix, reverberation cancellation coefficient, and image stability score to the multimodal fusion network (MFN), and output the audio-visual synchronization quality index and scene adaptation level;
[0033] Step S4: Based on the output of step S3, dynamically generate a combination of optimization strategies, including:
[0034] Select the audio enhancement mode (directional beamforming / omnidirectional noise reduction) and video optimization mode (dynamic HDR / motion compensation frame interpolation) according to the scene adaptation level;
[0035] If the audio-visual synchronization quality index is lower than the threshold, the timing calibration module is triggered to align the audio and video streams.
[0036] Step S5: Real-time feedback and adjustments, including:
[0037] The audio noise reduction intensity and video frame rate are adjusted by an adaptive PID controller to respond to sudden changes in environmental parameters;
[0038] Step S6: Output the optimized audio and video streams to the conference screen and record the optimization strategy parameters to the cloud knowledge base;
[0039] Step S7: Periodically verify the optimization effect, including:
[0040] The weights of the deep learning model are updated based on user interaction data (such as speech clarity scores and video smoothness feedback).
[0041] Furthermore, in the aforementioned deep learning-based self-optimization method for audio and video on large conference screens, step S1 includes:
[0042] The sound source direction angle is calculated using the time delay difference (TDOA) algorithm of the distributed microphone array, and the microphone array is deployed at the edge of the large screen in an asymmetric topology.
[0043] The dynamic range of light intensity is achieved by capturing highlight and shadow areas in a time-division manner using dual exposure sensors and fusing them into HDR metadata.
[0044] Furthermore, in the aforementioned deep learning-based self-optimization method for audio and video on large-screen conference displays, step S2 includes:
[0045] The noise suppression weight matrix is generated through a frequency domain mask, specifically by dividing the background noise spectrum into sub-bands and dynamically allocating the suppression coefficients of each sub-band based on the signal-to-noise ratio (SNR).
[0046] The image stability score is evaluated using a convolutional neural network (CNN) classifier, with the input being the spatiotemporal correlation matrix of optical flow features and facial key points.
[0047] Furthermore, in the aforementioned deep learning-based self-optimization method for audio and video on large-screen conference displays, step S3 includes:
[0048] The multimodal fusion network (MFN) is a cascaded structure, including:
[0049] Level 1: Bidirectional LSTM network extracts temporal features of acoustic parameters;
[0050] Level 2: 3D convolutional networks extract spatiotemporal features of video dynamic parameters;
[0051] Level 3: The attention mechanism module integrates cross-modal features and outputs correlation evaluation results.
[0052] Furthermore, in the aforementioned deep learning-based self-optimization method for audio and video on large-screen conference displays, step S4 includes:
[0053] The timing calibration module uses the Dynamic Time Warping (DTW) algorithm to align the audio and video streams and compensates for inter-frame differences caused by network latency through interpolation.
[0054] Furthermore, this invention discloses a deep learning-based self-optimizing audio and video system for large-screen conferences, comprising the following modules:
[0055] Domain-based acquisition module: used for time-division acquisition of environmental acoustic parameters and video dynamic parameters, including a distributed microphone array, dual exposure sensor and motion vector detection unit;
[0056] Hierarchical analysis module: Connects to the domain acquisition module, performs noise suppression weight matrix calculation, reverberation elimination, and image stability score generation;
[0057] Multimodal Fusion Network (MFN) module: Connects to the hierarchical analysis module and evaluates the audio-visual correlation through a deep learning model;
[0058] Dynamic strategy generation module: Connects to the MFN module and generates audio and video optimization strategies based on the scene adaptation level;
[0059] Real-time adjustment module: Connects to the dynamic strategy generation module, including an adaptive PID controller and a timing calibration unit;
[0060] Effect verification module: Connects the real-time adjustment module to the user terminal, collects feedback data and updates model weights;
[0061] Cloud-based knowledge base module: Connects bidirectionally to all modules, storing optimization strategy parameters and user interaction data.
[0062] The present invention also discloses a computer-readable storage medium storing a computer program, characterized in that, when the program is executed by a processor, it implements the steps of the method according to any one of claims 1-6, and the storage medium is a distributed architecture, comprising:
[0063] Local cache unit: stores raw parameters and intermediate analysis results acquired in real time;
[0064] Edge computing unit: Deploys MFN model and adaptive PID controller for low-latency data processing;
[0065] Cloud storage unit: Periodically synchronizes and optimizes strategy parameters and user feedback data.
[0066] The present invention also discloses a computing device, characterized in that it comprises:
[0067] Multi-core processor: used for parallel execution of acoustic parameter processing and video dynamic analysis;
[0068] GPU acceleration unit: Dedicated to running multimodal fusion networks (MFN) and reinforcement learning agents;
[0069] Hardware encoder: Real-time compression and optimization of audio and video streams, supporting H.265 / AV1 encoding protocols.
[0070] Compared with existing technologies, the present invention has the following advantages and beneficial effects:
[0071] This invention discloses a deep learning-based method and system for self-optimizing audio and video on a large-screen conference platform. Through a domain-based acquisition mechanism and an edge computing architecture (with acoustic parameter sampling frequency 1.5-2 times that of video), it achieves a breakthrough in real-time performance, reducing video processing latency to 45ms (a 62.5% reduction) and audio response time to 85ms. Combined with cross-modal collaborative optimization using a multimodal fusion network (MFN), the audio-visual synchronization error is reduced to only 18ms in weak network environments, the number of frames for dynamic lighting overexposure recovery is reduced by 62.5%, and the speech signal-to-noise ratio is improved to 22.5dB. Through a reinforcement learning-driven adaptive PID controller, the dynamic range of illumination reaches 120dB, reverberation time is optimized to 0.5s, motion blur is reduced by 75%, and smooth 60fps output is supported. The edge-cloud collaborative architecture reduces resource consumption by 30%, controls weak network synchronization error within 25ms, and reduces cloud dependency by 60%. User feedback loop and modular design resulted in a satisfaction score of 4.7 / 5, improved troubleshooting efficiency by 70%, shortened model iteration cycle to days, and achieved highly robust and low-latency audio and video self-optimization. Attached Figure Description
[0072] Figure 1 The present invention discloses a flowchart of a deep learning-based self-optimization method for audio and video on a large conference screen;
[0073] Figure 2 A schematic diagram of the self-optimizing audio and video system for large-screen conferencing displays based on deep learning, according to this invention. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below. However, it should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concept of the invention. All raw materials used in the embodiments of this invention are commercially available.
[0075] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the embodiments.
[0076] Example 1
[0077] A deep learning-based method for self-optimization of audio and video on large-screen conferencing displays, such as... Figure 1 As shown, it includes the following steps:
[0078] Step S1: Collect multimodal parameters by domain, including:
[0079] The first set of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time (RT60), and sound source direction angle;
[0080] The second set of parameters: video dynamic parameters, including dynamic range of illumination intensity, image motion vector, and facial key point displacement;
[0081] The first set of parameters and the second set of parameters are collected independently through a time-division sampling mechanism, and the sampling frequency of the first set of parameters is 1.5-2 times that of the second set of parameters;
[0082] Step S2: Perform hierarchical progressive analysis on the multimodal parameters, including:
[0083] First-level analysis: Generate a noise suppression weight matrix based on acoustic parameters and calculate the reverberation cancellation coefficient;
[0084] Second-level analysis: Based on video dynamic parameters, the image jitter features are extracted by optical flow method, and the image stability score is generated by combining the displacement of facial key points.
[0085] Step S3: Integrate acoustic and video analysis results, and perform cross-modal correlation assessment using a deep learning model, including:
[0086] Input the noise suppression weight matrix, reverberation cancellation coefficient, and image stability score to the multimodal fusion network (MFN), and output the audio-visual synchronization quality index and scene adaptation level;
[0087] Step S4: Based on the output of step S3, dynamically generate a combination of optimization strategies, including:
[0088] Select the audio enhancement mode (directional beamforming / omnidirectional noise reduction) and video optimization mode (dynamic HDR / motion compensation frame interpolation) according to the scene adaptation level;
[0089] If the audio-visual synchronization quality index is lower than the threshold, the timing calibration module is triggered to align the audio and video streams.
[0090] Step S5: Real-time feedback and adjustments, including:
[0091] The audio noise reduction intensity and video frame rate are adjusted by an adaptive PID controller to respond to sudden changes in environmental parameters;
[0092] Step S6: Output the optimized audio and video streams to the conference screen and record the optimization strategy parameters to the cloud knowledge base;
[0093] Step S7: Periodically verify the optimization effect, including:
[0094] The weights of the deep learning model are updated based on user interaction data (such as speech clarity scores and video smoothness feedback).
[0095] Example 2
[0096] A deep learning-based method for self-optimization of audio and video on large-screen conferencing displays includes the following steps:
[0097] Step S1: Collect multimodal parameters by domain, including:
[0098] The first set of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time (RT60), and sound source direction angle;
[0099] The second set of parameters: video dynamic parameters, including dynamic range of illumination intensity, image motion vector, and facial key point displacement;
[0100] The first set of parameters and the second set of parameters are collected independently through a time-division sampling mechanism, and the sampling frequency of the first set of parameters is 1.5-2 times that of the second set of parameters;
[0101] Parameter analysis in step S1:
[0102] 1. Sound parameters
[0103] 1) Background noise spectrum
[0104] Technical principle: The frequency domain distribution of environmental noise is decomposed by FFT (Fast Fourier Transform) to identify low-frequency steady-state noise (such as air conditioner noise) and high-frequency transient noise (such as keyboard typing).
[0105] Implementation method: A distributed microphone array is used to acquire raw audio signals, and energy distribution features are extracted by frequency band.
[0106] Optimization effect: Provides a data foundation for subsequent frequency domain mask generation, achieving accurate sub-band noise suppression.
[0107] 2) Reverberation time (RT60)
[0108] Technical principle: Characterizing the sound field attenuation rate, the calculation formula is RT60 = 0.161V / A (V is the room volume, A is the sound-absorbing area).
[0109] Implementation method: The attenuation curve is calculated by using the impulse response measurement method combined with the time-domain signal of the microphone array.
[0110] Optimization effect: Dynamically adjust the reverberation cancellation coefficient to avoid a decrease in speech clarity.
[0111] 3) Sound source direction angle
[0112] Technical principle: Based on the TDOA (Time Delay Occurrence Algorithm), the location of the sound source is calculated by the time difference of arrival of the signals from the microphone pair.
[0113] Implementation method: Asymmetric topology deployment of microphone array (main microphone spacing 15cm, auxiliary microphone spacing 10cm).
[0114] Optimization effect: Supports directional beamforming technology to improve the signal-to-noise ratio of target speech.
[0115] 2 Video Parameters
[0116] 1) Dynamic range of light intensity
[0117] Technical principle: Dual exposure sensors capture highlight (short exposure) and shadow (long exposure) areas in a time-division manner, and fuse them into HDR metadata.
[0118] Implementation method: Time-division sampling of sensors combined with image fusion algorithm
[0119] Optimization effect: Dynamic HDR mode adapts to sudden changes in brightness in meeting room scenes.
[0120] 2) Image motion vector
[0121] Technical principle: Calculates the motion trajectory of pixels between adjacent frames using optical flow.
[0122] Implementation method: Based on the Lucas-Kanade algorithm for GPU parallel computing
[0123] Optimization effect: Provides a basis for motion trajectory prediction for motion compensation frame interpolation.
[0124] 3) Facial key point displacement
[0125] Technical principle: CNN detects 68 key points on a face and analyzes micro-expression changes.
[0126] Implementation method: Real-time tracking using a lightweight MobileNet model
[0127] Optimization effect: Combine facial features to assess image stability and avoid facial expression distortion.
[0128] Step S2: Perform hierarchical progressive analysis on the multimodal parameters, including:
[0129] First-level analysis: Generate a noise suppression weight matrix based on acoustic parameters and calculate the reverberation cancellation coefficient;
[0130] Second-level analysis: Based on video dynamic parameters, the image jitter features are extracted by optical flow method, and the image stability score is generated by combining the displacement of facial key points.
[0131] Key Algorithm Analysis
[0132] 1. Generation of noise suppression weight matrix
[0133] Algorithm flow:
[0134] The audio signal is framed and converted to the frequency domain (STFT).
[0135] Dynamically assigned inhibition coefficient: α_k=1 / (1+e^(-β(SNR_k-θ)))
[0136] SNR_k: Signal-to-noise ratio of the k-th subband (unit: dB), calculated as follows:
[0137] SNR_k = 10 × log 10 (S_k / N_k), where S_k is the speech energy and N_k is the noise energy.
[0138] β: Slope factor (typical value 3-5), controls the steepness of the transition zone.
[0139] θ: Signal-to-noise ratio threshold (typical value 5-10dB), below which strong suppression is activated.
[0140] effect:
[0141] When SNR_k < θ, α_k approaches 0 (noise is completely suppressed);
[0142] When SNR_k > θ, α_k approaches 1 (preserving the speech signal).
[0143] Advantages: The Sigmoid function achieves smooth subband transitions, avoiding musical noise.
[0144] Step S3: Integrate acoustic and video analysis results, and perform cross-modal correlation assessment using a deep learning model, including:
[0145] Input the noise suppression weight matrix, reverberation cancellation coefficient, and image stability score to the multimodal fusion network (MFN), and output the audio-visual synchronization quality index and scene adaptation level;
[0146] Multimodal Fusion Network (MFN)
[0147] Network structure:
[0148] Level 1 (Bidirectional LSTM): Extracting acoustic temporal features
[0149] Level 2 (3D-CNN): Capturing video spatiotemporal features
[0150] Level 3 (Attention Mechanism): Calculate cross-modal feature weights
[0151] Function: To quantify the audio-visual synchronization quality index (0-1 scale)
[0152] Step S4: Based on the output of step S3, dynamically generate a combination of optimization strategies, including:
[0153] Select the audio enhancement mode (directional beamforming / omnidirectional noise reduction) and video optimization mode (dynamic HDR / motion compensation frame interpolation) according to the scene adaptation level;
[0154] If the audio-visual synchronization quality index is lower than the threshold, the timing calibration module is triggered to align the audio and video streams.
[0155] Step S5: Real-time feedback and adjustments, including:
[0156] The audio noise reduction intensity and video frame rate are adjusted by an adaptive PID controller to respond to sudden changes in environmental parameters;
[0157] Step S6: Output the optimized audio and video streams to the conference screen and record the optimization strategy parameters to the cloud knowledge base;
[0158] Step S7: Periodically verify the optimization effect, including:
[0159] The weights of the deep learning model are updated based on user interaction data (such as speech clarity scores and video smoothness feedback).
[0160] In step S1:
[0161] The sound source direction angle is calculated using the time delay difference (TDOA) algorithm of the distributed microphone array, and the microphone array is deployed at the edge of the large screen in an asymmetric topology.
[0162] The dynamic range of light intensity is achieved by capturing highlight and shadow areas in a time-division manner using dual exposure sensors and fusing them into HDR metadata.
[0163] In step S2:
[0164] The noise suppression weight matrix is generated through a frequency domain mask, specifically by dividing the background noise spectrum into sub-bands and dynamically allocating the suppression coefficients of each sub-band based on the signal-to-noise ratio (SNR).
[0165] The image stability score is evaluated using a convolutional neural network (CNN) classifier, with the input being the spatiotemporal correlation matrix of optical flow features and facial key points.
[0166] In step S3:
[0167] The multimodal fusion network (MFN) is a cascaded structure, including:
[0168] Level 1: Bidirectional LSTM network extracts temporal features of acoustic parameters;
[0169] Level 2: 3D convolutional networks extract spatiotemporal features of video dynamic parameters;
[0170] Level 3: The attention mechanism module integrates cross-modal features and outputs correlation evaluation results.
[0171] In step S4:
[0172] The timing calibration module uses the Dynamic Time Warping (DTW) algorithm to align the audio and video streams and compensates for inter-frame differences caused by network latency through interpolation.
[0173] Dynamic Time Warping (DTW)
[0174] Principle: Dynamic programming aligns the time axis and minimizes the path cost function.
[0175] D(i,j)=cost(i,j)+min(D(i-1,j),D(i,j-1),D(i-1,j-1))
[0176] Parameter description:
[0177] cost(i,j): The synchronization error between audio frame i and video frame j (unit: ms), calculated as follows:
[0178] cost(i,j)=|t_audio(i)-t_video(j)|.
[0179] D(i,j): Cumulative path cost, used to find the optimal alignment path.
[0180] Constraints:
[0181] The sliding window constraint |ij|≤5 controls the computational complexity to O(N).
[0182] Optimization: Sliding window limits computational complexity
[0183] In step S5:
[0184] The parameters of the adaptive PID controller are dynamically adjusted by the reinforcement learning agent, and the reward function is generated based on the real-time audio-visual synchronization quality index and user feedback data.
[0185] The parameter adjustment mechanism of the adaptive PID controller is as follows:
[0186] Reinforcement learning agent: using the synchronization index as the state and PID parameters as the actions.
[0187] Reward function: R = w1·SyncScore + w2·UserFeedback - w3·EnergyCost
[0188] Parameter description:
[0189] SyncScore: Audio-visual synchronization quality index (0-1 scale), output by MFN.
[0190] UserFeedback: User ratings (0-5 points), collected in real time through the interactive interface.
[0191] EnergyCost: System power consumption (unit: watts), monitored by hardware sensors.
[0192] w1, w2, w3: Weight coefficients (default values 0.5, 0.3, 0.2), dynamically adjusting the priority of the optimization target.
[0193] Online learning: PPO algorithm update policy network.
[0194] Example 3
[0195] A deep learning-based self-optimizing audio and video system for large-screen conferencing systems, such as... Figure 2 As shown, it includes the following modules:
[0196] Domain-based acquisition module: used for time-division acquisition of environmental acoustic parameters and video dynamic parameters, including a distributed microphone array, dual exposure sensor and motion vector detection unit;
[0197] Hierarchical analysis module: Connects to the domain acquisition module, performs noise suppression weight matrix calculation, reverberation elimination, and image stability score generation;
[0198] Multimodal Fusion Network (MFN) module: Connects to the hierarchical analysis module and evaluates the audio-visual correlation through a deep learning model;
[0199] Dynamic strategy generation module: Connects to the MFN module and generates audio and video optimization strategies based on the scene adaptation level;
[0200] Real-time adjustment module: Connects to the dynamic strategy generation module, including an adaptive PID controller and a timing calibration unit;
[0201] Effect verification module: Connects the real-time adjustment module to the user terminal, collects feedback data and updates model weights;
[0202] Cloud-based knowledge base module: Connects bidirectionally to all modules, storing optimization strategy parameters and user interaction data.
[0203] Example 4
[0204] Application examples
[0205] Deployment of Intelligent Conference Room Audio and Video Self-Optimization System
[0206] Scene Description
[0207] A smart conference room at the headquarters of a multinational corporation, with an area of 60 square meters and a capacity of 20 people, faces the following typical problems:
[0208] 1) Complex acoustic environment: continuous noise from central air conditioning (45dB), speech overlap caused by multiple people speaking at the same time;
[0209] 2) Dynamic changes in lighting: The floor-to-ceiling windows on the east side result in strong morning light (>1000 lux), while the projection area is not well lit (<100 lux);
[0210] 3) Frequent network fluctuations: Cross-border video conferences often encounter network latency of more than 200ms.
[0211] Hardware deployment
[0212] a. Audio acquisition system:
[0213] Distributed microphone array: 4 ReSpeaker Mic Array v2.0 arranged in a ring (main spacing 15cm, auxiliary spacing 10cm);
[0214] Acoustic sensor: Equipped with MiniDSP UMIK-1 for measuring RT60, sampling rate 48kHz.
[0215] b. Video capture system:
[0216] Dual exposure camera: Sony IMX586 sensor, supporting 120dB dynamic range;
[0217] Motion detection unit: NVIDIA Jetson Xavier NX runs optical flow algorithms.
[0218] c. Edge computing devices:
[0219] AI-accelerated server: Intel i9-12900K + 2×NVIDIA RTX A6000;
[0220] Real-time controller: Arduino Due + Xilinx Zynq UltraScale + MPSoC.
[0221] Implementation process
[0222] Step S1: Domain Parameter Acquisition
[0223] Acoustic parameter acquisition (100Hz):
[0224] Background noise spectrum: The air conditioner noise was decomposed into low frequency band (50-500Hz, energy percentage 70%) by FFT;
[0225] RT60 measurement: The impulse response shows that it takes 0.8 seconds to decay to -60dB, and the trigger reverberation cancellation coefficient α = 0.7;
[0226] Sound source localization: The TDOA algorithm locates the speaker's position (angle error <3°).
[0227] Video parameter acquisition (60Hz):
[0228] HDR fusion: Short exposure (1 / 1000s) captures the scene outside the window, long exposure (1 / 30s) enhances the details of the projected area;
[0229] Motion vector analysis: Optical flow method detected projection screen jitter (maximum displacement 15 pixels / frame);
[0230] Facial landmark tracking: MobileNet v3 detected 3 people frowning (corner of the mouth drooping >5 pixels).
[0231] Step S2: Hierarchical Progressive Analysis
[0232] Acoustic treatment:
[0233] In the noise suppression weight matrix, the air conditioner noise sub-band (200-400Hz) is assigned α = 0.3;
[0234] The reverberation cancellation factor is dynamically adjusted to 0.65 (to avoid a hollow feeling in the voice).
[0235] Video processing:
[0236] Optical flow features are input into ResNet-18, and the output jitter score is 7.2 (triggering motion compensation);
[0237] The spatiotemporal matrix of facial key points shows a frowning frequency of 2Hz, which is judged as a risk of image distortion.
[0238] Step S3: Cross-modal fusion
[0239] MFN Network Inference:
[0240] Bidirectional LSTM is used to extract acoustic temporal features (hidden state dimension 256);
[0241] 3D-CNN captures video motion trajectories (3×3×3 convolutional kernels, stride 2);
[0242] After the attention mechanism is weighted, the output synchronization index is 0.82, and the scene adaptation level is "medium".
[0243] Step S4: Dynamic Policy Generation
[0244] Optimization strategy:
[0245] Audio mode: Directional beamforming (main lobe width 40°) + omnidirectional noise reduction (suppression of non-speech audio segments <-25dB);
[0246] Video mode: Dynamic HDR (luminance mapping curve γ = 2.2) + motion compensation interpolation (output frame rate 60fps);
[0247] Timing calibration: DTW alignment of the offset frame (interpolation compensation 32ms).
[0248] Step S5: Real-time feedback and adjustment
[0249] Sudden noise response (door suddenly closes, noise peak 65dB):
[0250] The PID controller increases the noise reduction intensity within 80ms (Kp=1.2, Ti=1.5s);
[0251] The video frame rate was reduced to 45fps to prioritize audio quality.
[0252] User feedback intervention:
[0253] Attendees submitted a clarity rating of 4.5 / 5 via iPad Pro;
[0254] The reinforcement learning agent updates the reward function weights (w1 = 0.6, w2 = 0.4).
[0255] Step S6 / S7: Output and Verification
[0256] Optimize output:
[0257] The speech signal-to-noise ratio has been improved to 28dB (originally 12dB);
[0258] The dynamic range of the image has been expanded to 18 stops, and motion blur has been reduced by 62%.
[0259] Periodic verification:
[0260] Simulated cross-border network latency (250ms): synchronization error controlled within 25ms;
[0261] After model iteration, the MFN synchronization index increased to 0.89.
[0262] Example 5
[0263] System Implementation
[0264] Module configuration and connection:
[0265] Domain-specific data acquisition module:
[0266] Distributed microphone array (4 channels), dual exposure sensor (Sony IMX586), FPGA motion detection unit;
[0267] The output is connected to the hierarchical analysis module via a PCIe interface.
[0268] Hierarchical Analysis Module:
[0269] The acoustic DSP chip (TI TMS320C6748) performs frequency domain mask calculations;
[0270] The GPU (NVIDIA Jetson AGX) runs the CNN classifier.
[0271] MFN module:
[0272] Deployed on an edge server (Intel Xeon + 4×RTX 6000), it supports TensorRT acceleration.
[0273] Dynamic strategy generation module:
[0274] The strategy library pre-stores 10 optimization modes, which can be switched in real time via API calls.
[0275] Real-time adjustment module:
[0276] Adaptive PID controller (Arduino Due), timing calibration unit (Xilinx Zynq FPGA).
[0277] Effect verification module:
[0278] The user terminal (iPad Pro) collects feedback, and the data is transmitted back via Wi-Fi 6.
[0279] Cloud-based knowledge base module:
[0280] AWS S3 buckets + Lambda functions support distributed model training.
[0281] Comparative Example 1
[0282] (Refer to CN111866439A)
[0283] Solution: Use a symmetrical microphone array + fixed beamforming algorithm, and install the camera at a low position.
[0284] defect:
[0285] Insufficient suppression of burst noise: Under 60dB keyboard noise, the speech signal-to-noise ratio is only improved by 8dB (15dB in this invention);
[0286] Failure in mobile scenarios: When the direction angle deviation of the sound source is greater than 20°, the main lobe of the beamforming shifts, resulting in sound pickup failure.
[0287] Comparative Example 2
[0288] (Refer to CN118075418A)
[0289] Solution: Based on CNN+LSTM preprocessing + secondary detection process, with a fixed optimization threshold.
[0290] defect:
[0291] Poor real-time performance: 1080p video processing latency reaches 120ms (45ms in this invention);
[0292] Dynamic lighting processing failed: The image was overexposed under 1000 lux strong light (this invention maintains details through dynamic adjustment via HDR).
[0293] Comparative Example 3
[0294] (Refer to CN118214825A)
[0295] Solution: An integrated conferencing system that relies on a cloud-based policy library and a multi-router architecture.
[0296] defect:
[0297] Network dependency: Audio and video synchronization error >80ms under 100ms network jitter (this invention reduces it to 20ms through edge computing);
[0298] High maintenance costs: Module failure requires complete replacement (this invention supports hot-swapping of independent modules).
[0299] Test Example 1
[0300] Speech intelligibility in noisy environments
[0301] Conditions: Simulated conference room (background noise 55dB, reverberation RT60 = 1.2s), 5 people speaking simultaneously.
[0302] The results are shown in Table 1.
[0303] Table 1 Comparison of speech intelligibility in noisy environments.
[0304] index Example 4 Comparative Example 1 Comparative Example 2 Speech signal-to-noise ratio (dB) 22.5 14.3 18.7 Echo suppression rate (%) 92 85 88 Response time (ms) 85 120 150
[0305] Test Example 2
[0306] Dynamic lighting image stability
[0307] Conditions: The light intensity suddenly increases from 100 lux to 1000 lux, and the camera moves rapidly.
[0308] The results are shown in Table 2.
[0309] Table 2 Comparison of stability of dynamic lighting images
[0310] index Example 4 Comparative Example 2 Comparative Example 3 Image overexposure recovery frame rate 3 8 6 Motion blur reduction rate (%) 75 40 55 Face tracking accuracy (%) 98 82 90
[0311] Test Example 3
[0312] Synchronization in weak network environments
[0313] Conditions: Simulate network jitter (latency fluctuation 50-200ms, packet loss rate 5%).
[0314] index Example 4 Comparative Example 3 Audio / video synchronization error (ms) 18 75 Frame rate stability (fps) 58±2 45±10 User satisfaction rating ( / 5) 4.7 3.2
[0315] This invention, through its innovative architecture of domain-specific acquisition, multimodal fusion, and edge-cloud collaboration, significantly outperforms existing technologies in the following aspects:
[0316] Improved real-time performance: Video processing latency reduced by 62.5%, and voice response time shortened by 29%;
[0317] Enhanced scene adaptability: Key indicators (signal-to-noise ratio, image stability) are improved by over 40% in dynamic lighting and noise environments;
[0318] Reduced network dependence: Synchronization errors are reduced by 76% in weak network scenarios, and offline optimization mode is supported;
[0319] User experience optimization: The subjective user rating reached 4.7 / 5, far exceeding the comparison ratio (the highest being 3.2 / 5).
[0320] This solution provides a robust, low-latency audio and video self-optimization solution for conferencing systems in complex environments.
[0321] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, any changes and modifications made to the embodiments described herein based on the innovative concept of the present invention, or equivalent structural or procedural transformations made using the content of the present invention specification, directly or indirectly applying the above technical solutions to other related technical fields, are all included within the scope of protection of the present invention patent.
Claims
1. A deep learning-based method for self-optimizing audio and video on a large conference screen, characterized in that, Includes the following steps: Step S1: Collect multimodal parameters by domain, including: The first set of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time, and sound source direction angle; The second set of parameters: video dynamic parameters, including dynamic range of illumination intensity, image motion vector, and facial key point displacement; The first set of parameters and the second set of parameters are collected independently through a time-division sampling mechanism, and the sampling frequency of the first set of parameters is 1.5-2 times that of the second set of parameters; Step S2: Perform hierarchical progressive analysis on the multimodal parameters, including: First-level analysis: Generate a noise suppression weight matrix based on acoustic parameters and calculate the reverberation cancellation coefficient; Second-level analysis: Based on video dynamic parameters, the image jitter features are extracted by optical flow method, and the image stability score is generated by combining the displacement of facial key points. The noise suppression weight matrix is generated through a frequency domain mask, specifically by dividing the background noise spectrum into sub-bands and dynamically allocating the suppression coefficients of each sub-band based on the signal-to-noise ratio. The image stability score is evaluated by a convolutional neural network (CNN) classifier, with the input being the spatiotemporal correlation matrix between optical flow features and facial key points. Step S3: Integrate acoustic and video analysis results, and perform cross-modal correlation assessment using a deep learning model, including: The noise suppression weight matrix, reverberation cancellation coefficient, and image stability score are input to the multimodal fusion network (MFN), and the audio-visual synchronization quality index and scene adaptation level are output. The MFN is a cascaded structure, including: Level 1: Bidirectional LSTM network extracts temporal features of acoustic parameters; Level 2: 3D convolutional networks extract spatiotemporal features of video dynamic parameters; Level 3: The attention mechanism module integrates cross-modal features and outputs correlation evaluation results; Step S4: Based on the output of step S3, dynamically generate a combination of optimization strategies, including: Select the audio enhancement mode and video optimization mode according to the scene adaptation level; If the audio-visual synchronization quality index is lower than the threshold, the timing calibration module is triggered to align the audio and video streams. The timing calibration module uses the Dynamic Time Warping (DTW) algorithm to align the audio and video streams and compensates for inter-frame differences caused by network latency through interpolation. Step S5: Real-time feedback and adjustments, including: The audio noise reduction intensity and video frame rate are adjusted by an adaptive PID controller to respond to sudden changes in environmental parameters; the parameters of the adaptive PID controller are dynamically adjusted by a reinforcement learning agent, and the reward function is generated based on the real-time audio-visual synchronization quality index and user feedback data. Step S6: Output the optimized audio and video streams to the conference screen and record the optimization strategy parameters to the cloud knowledge base; Step S7: Periodically verify the optimization effect, including: The weights of the deep learning model are updated based on user interaction data.
2. The method according to claim 1, characterized in that, In step S1: The sound source direction angle is calculated using the time delay difference (TDOA) algorithm of the distributed microphone array, and the microphone array is deployed at the edge of the large screen in an asymmetric topology. The dynamic range of light intensity is achieved by capturing highlight and shadow areas in a time-division manner using dual exposure sensors and fusing them into HDR metadata.
3. A self-optimizing audio and video system for large-screen conferencing displays based on deep learning, characterized in that: The method applied to claim 1 or 2 includes the following modules: Domain-based acquisition module: used for time-division acquisition of environmental acoustic parameters and video dynamic parameters, including a distributed microphone array, dual exposure sensor and motion vector detection unit; Hierarchical analysis module: Connects to the domain acquisition module, performs noise suppression weight matrix calculation, reverberation elimination, and image stability score generation; Multimodal Fusion Network (MFN) module: Connects to the hierarchical analysis module, and evaluates the audio-visual correlation through a deep learning model; Dynamic strategy generation module: Connects to the MFN module and generates audio and video optimization strategies based on the scene adaptation level; Real-time adjustment module: Connects to the dynamic strategy generation module, including an adaptive PID controller and a timing calibration unit; Effect verification module: Connects the real-time adjustment module to the user terminal, collects feedback data and updates model weights; Cloud-based knowledge base module: Connects bidirectionally to all modules, storing optimization strategy parameters and user interaction data.
4. A computer-readable storage medium storing a computer program, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-2, and the storage medium is a distributed architecture, including: Local cache unit: stores raw parameters and intermediate analysis results acquired in real time; Edge computing unit: Deploys MFN model and adaptive PID controller for low-latency data processing; Cloud storage unit: Periodically synchronizes and optimizes strategy parameters and user feedback data.
5. A computing device, characterized in that, The steps of implementing the method according to any one of claims 1-2 include: Multi-core processor: used for parallel execution of acoustic parameter processing and video dynamic analysis; GPU acceleration unit: Dedicated to running Multimodal Fusion Network (MFN) and reinforcement learning agents; Hardware encoder: Real-time compression and optimization of audio and video streams, supporting H.265 / AV1 encoding protocols.
Citation Information
Patent Citations
Conference device and system for optimizing audio and video experience and operation method thereof
CN111866439A
Audio and video integrated device and system for conference
CN118214825A
Monitoring audio and video joint optimization method based on cross-modal attention mechanism
CN119172500A
Noise reduction and audio-visual speech activity detection
US20060224382A1
Visual tracking system for active object
US20240062580A1