Conference large-screen audio and video self-optimization method and system based on deep learning

Through the method of domain-based multimodal parameter acquisition and hierarchical integration, combined with edge-cloud collaborative computing and reinforcement learning, the problems of insufficient real-time and scene adaptability of the existing audio and video optimization solutions are solved, and a highly robust audio and video self-optimization effect is achieved.

CN120358387AActive Publication Date: 2025-07-22SHENZHEN QICHANG INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510497430.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-22
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the existing technology, in complex acoustic environments, dynamic lighting conditions and network fluctuations, the audio and video optimization solutions of the conference system have problems such as insufficient real-time, weak cross-modal collaboration capabilities, and excessive dependence on preset strategies, making it difficult to adapt to changes in the dynamic environment.

Method used

The method of domain-based multimodal parameter acquisition and hierarchical fusion is adopted, and acoustic and video parameters are differentiated through the time-sharing sampling mechanism, and cross-modal correlation evaluation is used to combine edge-cloud collaborative computing and reinforcement learning-driven dynamic control to generate a personalized optimization strategy.

Benefits of technology

It realizes highly robust audio-visual self-optimization in complex environments, reducing latency, improving voice signal-to-noise ratio, reducing dynamic lighting overexposure recovery frames, enhancing audio-visual synchronization in network environments, and improving user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358387A_ABST
    Figure CN120358387A_ABST
Patent Text Reader

Abstract

The invention discloses a conference large-screen audio and video self-optimization method and system based on deep learning, and belongs to the technical field of intelligent audio and video processing. Aiming at the problems of insufficient real-time performance, weak cross-modal collaboration, poor dynamic scene adaptability and the like in the prior art, an innovative architecture of domain multi-modal parameter acquisition and hierarchical fusion is provided. The method comprises the following steps: differentially acquiring acoustic parameters (background noise spectrum and sound source direction angle) and video parameters (illumination dynamic range and face key point displacement) through a time-sharing sampling mechanism; and generating a noise suppression weight matrix by using a frequency domain mask and extracting image stability characteristics by using an optical flow method. Experiments show that according to the scheme, the voice signal-to-noise ratio is increased to 22.5 dB in a 55dB noise environment, the audio and video synchronization error in a weak network scene is reduced to 18ms, the dynamic illumination overexposure recovery frame number is reduced by 62.5%, the scheme is obviously superior to a traditional scheme, and a high-robustness and low-delay audio and video self-optimization solution is provided for a mixed office scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent deep optimization of audio and video, and specifically discloses a method and system for self-optimization of conference large-screen audio and video based on deep learning. Background Art

[0002] With the popularization of remote collaboration and hybrid in-office work models, the audio and video quality of conference systems has become a core factor affecting communication efficiency. In the prior art, solutions for audio and video optimization mainly focus on single-modal enhancement or fixed-scenario adaptation. However, in complex acoustic environments, dynamic lighting conditions, and network fluctuation scenarios, there are still problems such as insufficient real-time performance, weak cross-modal collaboration ability, and over-reliance on preset strategies. The limitations of existing solutions are analyzed below in combination with relevant patented technologies:

[0003] CN111866439A (Conference device, system, and operating method for optimizing audio and video experience)

[0004] This solution suppresses echoes through microphone array beamforming technology and optimizes the physical layout of audio devices to reduce interference, significantly improving the naturalness of full-duplex calls. However, its core relies on hardware topology design (such as symmetrically distributed microphones and speakers), which has a high deployment cost in non-fixed scenarios (such as mobile conferences and open office environments), and the algorithm has insufficient robustness to sudden noises (such as keyboard typing and temporary device access). In addition, its video module only reduces distortion by installing the camera at a low position and does not involve adaptive processing of dynamic lighting or picture jitter.

[0005] CN118075418A (Video conference content output optimization method, device, equipment, and storage medium thereof)

[0006] This invention proposes a full-process optimization method for audio and video based on an AI model, ensuring output quality through preprocessing, detection of items to be optimized, and secondary verification, especially suitable for complex scenarios combining outdoor monitoring and conferences. However, its defects are as follows:

[0007] It relies on preset optimization strategies and fixed thresholds and is difficult to cope with dynamic changes in environmental parameters (such as sudden changes in lighting and multiple people speaking simultaneously);

[0008] The computational load of deep learning models (such as CNN + LSTM) is relatively high, resulting in limited real-time performance and unable to meet the requirements of low-latency conferences;

[0009] It lacks the ability of cross-modal collaborative analysis, and the audio and video streams are only aligned by timestamps without deeply exploring the semantic correlation between sound and picture.

[0010] CN118214825A (Integrated audio and video device and system for conferences)

[0011] This solution integrates online and offline meeting functions through modular design, supports one-key switching and multi-device interconnection, and uses automatic tracking cameras, omnidirectional microphones, etc. to achieve adaptive optimization. However, its limitations are as follows:

[0012] Highly dependent on network stability, with significant fluctuations in audio and video synchronization and quality in weak network environments;

[0013] The system integration degree is too high, resulting in complex troubleshooting and upgrade maintenance;

[0014] The optimization strategy is rule-driven (such as fixed noise reduction intensity), lacking the ability to dynamically learn the personalized needs of users.

[0015] In summary, the existing technologies have the following common defects:

[0016] Insufficient scene adaptability: Fixed hardware layouts or preset strategies are difficult to cope with dynamic environmental changes;

[0017] Lack of cross-modal collaboration: Audio and video optimization is carried out in isolation, without making full use of the sound and picture correlation to improve the overall experience;

[0018] Contradiction between real-time performance and resource efficiency: Complex model calculations lead to delays, while simplified strategies sacrifice optimization effects;

[0019] Strong network dependence: Cloud processing and multi-routing architectures have poor stability in weak network scenarios. Summary of the Invention

[0020] To address the above problems, the present invention discloses a method and system for self-optimizing audio and video on a conference large screen based on deep learning. Through domain-specific parameter collection, hierarchical cross-modal analysis, and dynamic policy generation, the following breakthroughs are achieved: Time-sharing and zoning acquisition mechanism: Differentiated processing of acoustic and video parameters to improve data effectiveness; Multi-modal fusion network (MFN): Combining sound and picture spatio-temporal features to mine semantic relevance and enhance environmental adaptability; Edge-cloud collaborative computing: Balancing latency and accuracy through local real-time processing and cloud model iteration; Reinforcement learning-driven dynamic control: Reducing dependence on preset rules and achieving personalized optimization.

[0021] This solution systematically solves the deficiencies of insufficient real-time performance, rigid scenarios, and strong network dependence in the existing technologies, and provides a highly robust audio and video self-optimization solution for hybrid work scenarios.

[0022] The present invention includes the following technical solutions:

[0023] A method for self-optimizing audio and video on a conference large screen based on deep learning, comprising the following steps:

[0024] Step S1: Collect multi-modal parameters in different domains, including:

[0025] The first set of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time (RT60), and sound source direction angle;

[0026] The second set of parameters: video dynamic parameters, including dynamic range of illumination intensity, picture motion vector, and displacement of face key points;

[0027] The first set of parameters and the second set of parameters are independently collected through a time-division sampling mechanism, and the sampling frequency of the first set of parameters is 1.5 - 2 times that of the second set of parameters;

[0028] Step S2: Perform hierarchical progressive analysis on the multi-modal parameters, including:

[0029] First-layer analysis: Generate a noise suppression weight matrix based on the acoustic parameters, and calculate the reverberation cancellation coefficient;

[0030] Second-layer analysis: Based on the video dynamic parameters, extract the picture jitter characteristics through the optical flow method, and generate a picture stability score in combination with the displacement of face key points;

[0031] Step S3: Integrate the acoustic and video analysis results, and perform cross-modal correlation evaluation through a deep learning model, including:

[0032] Input the noise suppression weight matrix, reverberation cancellation coefficient, and picture stability score into the multi-modal fusion network (MFN), and output the audio-visual synchronization quality index and the scene adaptation level;

[0033] Step S4: Dynamically generate an optimized strategy combination based on the output result of Step S3, including:

[0034] Select the audio enhancement mode (directional beamforming / omnidirectional noise reduction) and video optimization mode (dynamic HDR / motion compensation frame interpolation) according to the scene adaptation level;

[0035] If the audio-visual synchronization quality index is lower than the threshold, trigger the timing calibration module to align the audio and video streams;

[0036] Step S5: Provide real-time feedback adjustment, including:

[0037] Adjust the audio noise reduction intensity and video frame rate through an adaptive PID controller to respond to sudden changes in environmental parameters;

[0038] Step S6: Output the optimized audio and video streams to the conference large screen, and record the optimized strategy parameters in the cloud knowledge base;

[0039] Step S7: Periodically verify the optimization effect, including:

[0040] Update the weights of the deep learning model based on user interaction data (such as speech clarity score, picture smoothness feedback).

[0041] Further, in the above method for self-optimizing conference large-screen audio and video based on deep learning, in step S1:

[0042] The sound source direction angle is calculated by the time delay difference (TDOA) algorithm of the distributed microphone array, and the microphone array is deployed at the edge of the large screen in an asymmetric topology;

[0043] The dynamic range of the light intensity is captured by a dual-exposure sensor for the highlight and dark regions at different times and fused into HDR metadata.

[0044] Further, in the above method for self-optimizing conference large-screen audio and video based on deep learning, in step S2:

[0045] The noise suppression weight matrix is generated by a frequency domain mask, specifically: the background noise spectrum is divided into sub-bands, and sub-band suppression coefficients are dynamically allocated based on the signal-to-noise ratio (SNR);

[0046] The picture stability score is evaluated by a convolutional neural network (CNN) classifier, and the input is the spatio-temporal correlation matrix of the optical flow features and the facial key points.

[0047] Further, in the above method for self-optimizing conference large-screen audio and video based on deep learning, in step S3:

[0048] The multi-modal fusion network (MFN) is a cascade structure, including:

[0049] The first stage: A bidirectional LSTM network extracts the temporal features of the acoustic parameters;

[0050] The second stage: A 3D convolutional network extracts the spatio-temporal features of the video dynamic parameters;

[0051] The third stage: An attention mechanism module fuses the cross-modal features and outputs the correlation evaluation result.

[0052] Further, in the above method for self-optimizing conference large-screen audio and video based on deep learning, in step S4:

[0053] The temporal calibration module uses the dynamic time warping (DTW) algorithm to align the audio and video streams and compensates for the inter-frame differences caused by network latency through interpolation.

[0054] Further, the present invention discloses a system for self-optimizing conference large-screen audio and video based on deep learning, including the following modules:

[0055] Domain-separated acquisition module: used to acquire environmental acoustic parameters and video dynamic parameters at different times, including a distributed microphone array, a dual-exposure sensor, and a motion vector detection unit;

[0056] Hierarchical Analysis Module: Connects to the sub-domain acquisition module, performs noise suppression weight matrix calculation, reverberation elimination, and generates picture stability scores;

[0057] Multi-modal Fusion Network (MFN) Module: Connects to the hierarchical analysis module, and evaluates the audio-visual correlation through a deep learning model;

[0058] Dynamic Policy Generation Module: Connects to the MFN module, and generates audio-visual optimization strategies according to the scene adaptation level;

[0059] Real-time Adjustment Module: Connects to the dynamic policy generation module, including an adaptive PID controller and a timing calibration unit;

[0060] Effect Verification Module: Connects to the real-time adjustment module and the user terminal, collects feedback data and updates the model weights;

[0061] Cloud Knowledge Base Module: Bidirectionally connects to all modules, and stores optimization strategy parameters and user interaction data.

[0062] The present invention also discloses a computer-readable storage medium storing a computer program, characterized in that when the program is executed by a processor, it implements the steps of the method according to any one of claims 1-6, and the storage medium has a distributed architecture, including:

[0063] Local Cache Unit: Stores the originally collected parameters and intermediate analysis results in real time;

[0064] Edge Computing Unit: Deploys the MFN model and the adaptive PID controller for low-latency data processing;

[0065] Cloud Storage Unit: Periodically synchronizes optimization strategy parameters and user feedback data.

[0066] The present invention also discloses a computing device, characterized in that it includes:

[0067] Multi-core Processor: Used for parallel execution of acoustic parameter processing and video dynamic analysis;

[0068] GPU Acceleration Unit: Specifically used for running the multi-modal fusion network (MFN) and the reinforcement learning agent;

[0069] Hardware Encoder: Real-time compresses the optimized audio-visual stream, supporting the H.265 / AV1 encoding protocol.

[0070] Compared with the existing technologies, the present invention has the following advantages and beneficial effects:

[0071] The present invention discloses a method and system for self-optimizing audio and video on a conference large screen based on deep learning. Through a domain-based acquisition mechanism and an edge computing architecture (the sampling frequency of acoustic parameters is 1.5 - 2 times that of the video), real-time breakthrough is achieved, the video processing delay is reduced to 45 ms (a 62.5% reduction), and the audio response time is shortened to 85 ms. Combining cross-modal collaborative optimization of a multi-modal fusion network (MFN), the sound and picture synchronization error is only 18 ms in a weak network environment, the number of frames for dynamic overexposed light restoration is reduced by 62.5%, and the voice signal-to-noise ratio is increased to 22.5 dB. Through an adaptive PID controller driven by reinforcement learning, the dynamic range of light reaches 120 dB, the reverberation time is optimized to 0.5 s, the motion blur is reduced by 75%, and smooth output at 60 fps is supported. The edge-cloud collaborative architecture reduces resource consumption by 30%, the weak network synchronization error is controlled within 25 ms, and the cloud dependence is reduced by 60%. The user feedback closed-loop and modular design enable the satisfaction score to reach 4.7 / 5, the fault troubleshooting efficiency is increased by 70%, and the model iteration cycle is shortened to a daily level, achieving high-robustness and low-latency audio and video self-optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 Flowchart of a method for self-optimizing audio and video on a conference large screen based on deep learning disclosed by the present invention;

[0073] Figure 2 Schematic diagram of a system for self-optimizing audio and video on a conference large screen based on deep learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention. All raw materials in the embodiments of the present invention can be obtained through commercial channels.

[0075] It should be noted that, without conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below in conjunction with the embodiments.

[0076] Embodiment 1

[0077] A method for self-optimizing audio and video on a conference large screen based on deep learning, as Figure 1 shown, includes the following steps:

[0078] Step S1: Collect multi-modal parameters in domains, including:

[0079] The first group of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time (RT60), and sound source direction angle;

[0080] The second set of parameters: video dynamic parameters, including the dynamic range of light intensity, the motion vector of the picture, and the displacement of facial key points;

[0081] The first set of parameters and the second set of parameters are independently collected through a time-division sampling mechanism, and the sampling frequency of the first set of parameters is 1.5 - 2 times that of the second set of parameters;

[0082] Step S2: Perform hierarchical progressive analysis on the multi-modal parameters, including:

[0083] The first layer of analysis: Generate a noise suppression weight matrix based on the acoustic parameters and calculate the reverberation cancellation coefficient;

[0084] The second layer of analysis: Based on the video dynamic parameters, extract the picture jitter characteristics through the optical flow method, and generate a picture stability score in combination with the displacement of facial key points;

[0085] Step S3: Integrate the acoustic and video analysis results and perform cross-modal correlation evaluation through a deep learning model, including:

[0086] Input the noise suppression weight matrix, the reverberation cancellation coefficient, and the picture stability score into the multi-modal fusion network (MFN), and output the audio-visual synchronization quality index and the scene adaptation level;

[0087] Step S4: Dynamically generate an optimization strategy combination based on the output result of Step S3, including:

[0088] Select the audio enhancement mode (directional beamforming / omnidirectional noise reduction) and the video optimization mode (dynamic HDR / motion compensation frame interpolation) according to the scene adaptation level;

[0089] If the audio-visual synchronization quality index is lower than the threshold, trigger the timing calibration module to align the audio and video streams;

[0090] Step S5: Real-time feedback adjustment, including:

[0091] Adjust the audio noise reduction intensity and the video frame rate through an adaptive PID controller to respond to sudden changes in environmental parameters;

[0092] Step S6: Output the optimized audio and video streams to the conference large screen and record the optimization strategy parameters in the cloud knowledge base;

[0093] Step S7: Periodically verify the optimization effect, including:

[0094] Update the weights of the deep learning model based on user interaction data (such as speech clarity score, picture smoothness feedback).

[0095] Embodiment 2

[0096] A method for self-optimizing audio and video on a conference large screen based on deep learning, comprising the following steps:

[0097] Step S1: Collect multi-modal parameters by domain, including:

[0098] The first group of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time (RT60), and sound source direction angle;

[0099] The second group of parameters: video dynamic parameters, including light intensity dynamic range, picture motion vector, and displacement of facial key points;

[0100] The first group of parameters and the second group of parameters are independently collected through a time-sharing sampling mechanism, and the sampling frequency of the first group of parameters is 1.5 - 2 times that of the second group of parameters;

[0101] Parameter analysis in Step S1:

[0102] 1 Sound parameters

[0103] 1) Background noise spectrum

[0104] Technical principle: Decompose the frequency domain distribution of environmental noise through FFT (Fast Fourier Transform) to identify low-frequency steady-state noise (such as air conditioner noise) and high-frequency transient noise (such as keyboard keystrokes)

[0105] Implementation method: A distributed microphone array collects the original audio signal and extracts the energy distribution characteristics in frequency bands

[0106] Optimization effect: Provide a data basis for subsequent frequency domain mask generation and achieve precise sub-band noise suppression

[0107] 2) Reverberation time (RT60)

[0108] Technical principle: Characterize the sound field decay rate, and the calculation formula is RT60 = 0.161V / A (V is the room volume, A is the sound absorption area)

[0109] Implementation method: Through the impulse response measurement method, combine the time-domain signal of the microphone array to calculate the decay curve

[0110] Optimization effect: Dynamically adjust the reverberation cancellation coefficient to avoid the decline of speech clarity

[0111] 3) Sound source direction angle

[0112] Technical principle: Based on the TDOA (Time Difference of Arrival) algorithm, calculate the sound source azimuth through the time difference of arrival of signals from microphone pairs

[0113] Implementation method: Deploy the microphone array in an asymmetric topology (the main microphone spacing is 15 cm, and the auxiliary microphone spacing is 10 cm)

[0114] Optimization effect: Support for directional beamforming technology to improve the signal-to-noise ratio of the target voice.

[0115] 2 Video parameters

[0116] 1) Dynamic range of light intensity

[0117] Technical principle: The dual-exposure sensor captures the highlight (short exposure) and dark area (long exposure) regions at different times and fuses them into HDR metadata

[0118] Implementation method: The sensor samples at different times and cooperates with the image fusion algorithm

[0119] Optimization effect: The dynamic HDR mode adapts to the sudden changes in light and darkness in the meeting room

[0120] 2) Picture motion vector

[0121] Technical principle: Calculate the pixel motion trajectory between adjacent frames through the optical flow method

[0122] Implementation method: Based on the Lucas-Kanade algorithm for parallel computing on the GPU

[0123] Optimization effect: Provide a basis for predicting the motion trajectory for motion compensation and frame interpolation

[0124] 3) Displacement of facial key points

[0125] Technical principle: The CNN detects 68 key points of the face and analyzes the micro-expression changes

[0126] Implementation method: The lightweight MobileNet model for real-time tracking

[0127] Optimization effect: Combine facial features to evaluate the picture stability and avoid expression distortion

[0128] Step S2: Perform a hierarchical progressive analysis on the multi-modal parameters, including:

[0129] First-layer analysis: Generate a noise suppression weight matrix based on the acoustic parameters and calculate the reverberation cancellation coefficient;

[0130] Second-layer analysis: Based on the video dynamic parameters, extract the picture jitter features through the optical flow method, and generate a picture stability score by combining the displacement of facial key points;

[0131] Key algorithm analysis

[0132] 1. Generation of noise suppression weight matrix

[0133] Algorithm process:

[0134] The audio signal is framed and transformed into the frequency domain (STFT)

[0135] Dynamic allocation suppression coefficient: α_k = 1 / (1 + e^(-β(SNR_k - θ)))

[0136] SNR_k: Signal-to-noise ratio of the k-th sub-band (unit: dB), calculated as:

[0137] SNR_k = 10×log 10 (S_k / N_k), where S_k is the speech energy and N_k is the noise energy.

[0138] β: Slope factor (typical value 3 - 5), controlling the steepness of the transition band.

[0139] θ: Signal-to-noise ratio threshold (typical value 5 - 10 dB), below which strong suppression is activated.

[0140] Function:

[0141] When SNR_k < θ, α_k approaches 0 (completely suppressing noise);

[0142] When SNR_k > θ, α_k approaches 1 (retaining the speech signal).

[0143] Advantage: The Sigmoid function realizes smooth transition of sub-bands, avoiding musical noise

[0144] Step S3: Integrate the acoustic and video analysis results, and conduct cross-modal correlation evaluation through a deep learning model, including:

[0145] Input the noise suppression weight matrix, reverberation cancellation coefficient, and picture stability score into the multi-modal fusion network (MFN), and output the audio-visual synchronization quality index and scene adaptation level;

[0146] Multi-modal fusion network (MFN)

[0147] Network structure:

[0148] The first level (bidirectional LSTM): Extract acoustic temporal features

[0149] The second level (3D-CNN): Capture video spatio-temporal features

[0150] The third level (attention mechanism): Calculate cross-modal feature weights

[0151] Function: Quantify the audio-visual synchronization quality index (0 - 1 scale)

[0152] Step S4: Based on the output result of Step S3, dynamically generate an optimized strategy combination, including:

[0153] Select the audio enhancement mode (directional beamforming / omnidirectional noise reduction) and video optimization mode (dynamic HDR / motion compensation frame interpolation) according to the scene adaptation level;

[0154] If the audio-visual synchronization quality index is lower than the threshold, trigger the timing calibration module to align the audio-visual stream;

[0155] Step S5: Real-time feedback adjustment, including:

[0156] Adjust the audio noise reduction intensity and video frame rate through an adaptive PID controller to respond to sudden changes in environmental parameters;

[0157] Step S6: Output the optimized audio-visual stream to the conference large screen and record the optimization strategy parameters in the cloud knowledge base;

[0158] Step S7: Periodically verify the optimization effect, including:

[0159] Update the weights of the deep learning model based on user interaction data (such as speech clarity score, picture smoothness feedback).

[0160] In the said step S1:

[0161] The sound source direction angle is calculated by the time delay difference (TDOA) algorithm of the distributed microphone array, and the microphone array is deployed at the edge of the large screen in an asymmetric topology;

[0162] The dynamic range of the light intensity is captured by a dual-exposure sensor for the highlight and dark regions at different times and fused into HDR metadata.

[0163] In the said step S2:

[0164] The noise suppression weight matrix is generated by a frequency domain mask, specifically: the background noise spectrum is divided into sub-bands, and sub-band suppression coefficients are dynamically allocated based on the signal-to-noise ratio (SNR);

[0165] The picture stability score is evaluated by a convolutional neural network (CNN) classifier, and the input is the spatio-temporal correlation matrix of the optical flow features and the key points of the human face.

[0166] In the said step S3:

[0167] The multi-modal fusion network (MFN) is a cascade structure, including:

[0168] The first level: A bidirectional LSTM network extracts the temporal features of the acoustic parameters;

[0169] The second level: A 3D convolutional network extracts the spatio-temporal features of the video dynamic parameters;

[0170] The third level: An attention mechanism module fuses the cross-modal features and outputs the correlation evaluation result.

[0171] In the said step S4:

[0172] The timing calibration module uses the Dynamic Time Warping (DTW) algorithm to align the audio-visual stream and compensates for the inter-frame differences caused by network latency through interpolation.

[0173] Dynamic Time Warping (DTW)

[0174] Principle: Align the time axis using dynamic programming to minimize the path cost function:

[0175] D(i,j) = cost(i,j) + min(D(i - 1,j), D(i,j - 1), D(i - 1,j - 1))

[0176] Parameter description:

[0177] cost(i,j): The synchronization error between audio frame i and video frame j (unit: ms), calculated as:

[0178] cost(i,j) = |t_audio(i) - t_video(j)|.

[0179] D(i,j): The cumulative path cost, used to find the optimal alignment path.

[0180] Constraint conditions:

[0181] The sliding window is restricted by |i - j| ≤ 5 to control the computational complexity to O(N).

[0182] Optimization: The sliding window restricts the computational complexity

[0183] In step S5:

[0184] The parameters of the adaptive PID controller are dynamically adjusted by the reinforcement learning agent, and the reward function is generated based on the real-time audio-visual synchronization quality index and user feedback data.

[0185] The parameter adjustment mechanism of the adaptive PID controller is as follows:

[0186] Reinforcement learning agent: Uses the synchronization index as the state and the PID parameters as the action

[0187] Reward function: R = w1·SyncScore + w2·UserFeedback - w3·EnergyCost

[0188] Parameter description:

[0189] SyncScore: The audio-visual synchronization quality index (scaled from 0 to 1), output by the MFN.

[0190] UserFeedback: The user rating (scaled from 0 to 5 points), collected in real time through the interaction interface.

[0191] EnergyCost: System power consumption (unit: watt), monitored by hardware sensors.

[0192] w1, w2, w3: Weight coefficients (default values 0.5, 0.3, 0.2), dynamically adjust the optimization target priority.

[0193] Online learning: Update the policy network using the PPO algorithm.

[0194] Example 3

[0195] A conference large-screen audio-visual self-optimization system based on deep learning, as Figure 2 shown, includes the following modules:

[0196] Domain-based acquisition module: Used to collect environmental acoustic parameters and video dynamic parameters at different times, including a distributed microphone array, a dual-exposure sensor, and a motion vector detection unit;

[0197] Hierarchical analysis module: Connected to the domain-based acquisition module, performs noise suppression weight matrix calculation, reverberation elimination, and generation of picture stability scores;

[0198] Multi-modal fusion network (MFN) module: Connected to the hierarchical analysis module, evaluates the audio-visual correlation through a deep learning model;

[0199] Dynamic policy generation module: Connected to the MFN module, generates audio-visual optimization policies according to the scene adaptation level;

[0200] Real-time adjustment module: Connected to the dynamic policy generation module, includes an adaptive PID controller and a timing calibration unit;

[0201] Effect verification module: Connected to the real-time adjustment module and the user terminal, collects feedback data and updates the model weights;

[0202] Cloud knowledge base module: Bidirectionally connected to all modules, stores optimization policy parameters and user interaction data.

[0203] Example 4

[0204] Application example

[0205] Deployment of the intelligent conference room audio-visual self-optimization system

[0206] Scene description

[0207] The intelligent conference room of the headquarters of a multinational enterprise, with an area of 60㎡ and a capacity of 20 people, faces the following typical problems:

[0208] 1) Complex acoustic environment: Continuous noise from the central air conditioner (45dB), overlapping voices caused by multiple people speaking simultaneously;

[0209] 2) Dynamic changes in lighting: The floor-to-ceiling windows on the east side cause strong morning light to shine in (>1000 lux), resulting in insufficient lighting in the projection area (<100 lux).

[0210] 3) Frequent network fluctuations: Video conferences across countries often encounter network delays of over 200 ms.

[0211] Hardware Deployment

[0212] a. Audio acquisition system:

[0213] Distributed microphone array: 4 ReSpeaker Mic Array v2.0 arranged in a ring (main spacing 15 cm, auxiliary spacing 10 cm);

[0214] Acoustic sensor: Equipped with a MiniDSP UMIK-1 to measure RT60, sampling rate 48 kHz.

[0215] b. Video acquisition system:

[0216] Dual-exposure camera: Sony IMX586 sensor, supporting a dynamic range of 120 dB;

[0217] Motion detection unit: NVIDIA Jetson Xavier NX running an optical flow algorithm.

[0218] c. Edge computing device:

[0219] AI acceleration server: Intel i9-12900K + 2×NVIDIA RTX A6000;

[0220] Real-time controller: Arduino Due + Xilinx Zynq UltraScale+MPSoC.

[0221] Implementation process

[0222] Step S1: Domain-specific parameter acquisition

[0223] Acoustic parameter acquisition (100 Hz):

[0224] Background noise spectrum: Decompose the air conditioner noise into the low-frequency band (50 - 500 Hz, energy ratio 70%) through FFT;

[0225] RT60 measurement: The impulse response shows that it takes 0.8 seconds to decay to -60 dB, and the trigger reverberation cancellation coefficient α = 0.7;

[0226] Sound source localization: Locate the speaker's azimuth using the TDOA algorithm (angle error <3°).

[0227] Video parameter acquisition (60 Hz):

[0228] HDR Fusion: The short exposure (1 / 1000s) captures the scenery outside the window, and the long exposure (1 / 30s) enhances the details in the projection area;

[0229] Motion Vector Analysis: The optical flow method detects the jitter of the projection screen (maximum displacement 15 pixels / frame);

[0230] Facial Keypoint Tracking: MobileNet v3 detects that 3 people are frowning (the corners of the mouth droop > 5 pixels).

[0231] Step S2: Hierarchical Progressive Analysis

[0232] Acoustic Processing:

[0233] In the noise suppression weight matrix, the air conditioner noise sub-band (200 - 400Hz) is assigned α = 0.3;

[0234] The reverberation cancellation coefficient is dynamically adjusted to 0.65 (to avoid the sense of voice hollowness).

[0235] Video Processing:

[0236] The optical flow features are input into ResNet-18, and the output jitter score is 7.2 (triggering motion compensation);

[0237] The spatio-temporal matrix of facial keypoints shows a frowning frequency of 2Hz, which is determined as a risk of picture distortion.

[0238] Step S3: Cross-modal Fusion

[0239] MFN Network Inference:

[0240] Bidirectional LSTM extracts the acoustic temporal features (hidden state dimension 256);

[0241] 3D-CNN captures the video motion trajectory (convolution kernel 3×3×3, stride 2);

[0242] After weighted by the attention mechanism, the output synchronization index is 0.82, and the scene adaptation level is "medium".

[0243] Step S4: Dynamic Policy Generation

[0244] Optimization Strategy:

[0245] Audio Mode: Directional beamforming (main lobe width 40°) + omnidirectional noise reduction (suppressing non-speech frequency bands < -25dB);

[0246] Video Mode: Dynamic HDR (brightness mapping curve γ = 2.2) + motion compensation interpolation (output frame rate 60fps);

[0247] Temporal Calibration: DTW aligns the deviation frames (interpolation compensation 32ms).

[0248] Step S5: Real-time feedback adjustment

[0249] Sudden noise response (door closes suddenly, noise peak 65dB):

[0250] The PID controller increases the noise reduction intensity within 80ms (Kp = 1.2, Ti = 1.5s);

[0251] The video frame rate drops to 45fps to prioritize audio quality.

[0252] User feedback intervention:

[0253] The participants submit a clarity score of 4.5 / 5 through iPad Pro;

[0254] The reinforcement learning agent updates the reward function weights (w1 = 0.6, w2 = 0.4).

[0255] Steps S6 / S7: Output and verification

[0256] Optimized output:

[0257] The voice signal-to-noise ratio is increased to 28dB (original 12dB);

[0258] The dynamic range of the picture is extended to 18 stops, and the motion blur is reduced by 62%.

[0259] Periodic verification:

[0260] Simulate cross-border network latency (250ms): The synchronization error is controlled within 25ms;

[0261] After model iteration, the MFN synchronization index is increased to 0.89.

[0262] Example 5

[0263] System implementation

[0264] Module configuration and connection:

[0265] Domain-based acquisition module:

[0266] Distributed microphone array (4 channels), dual-exposure sensor (Sony IMX586), FPGA motion detection unit;

[0267] The output end is connected to the hierarchical analysis module through the PCIe interface.

[0268] Hierarchical analysis module:

[0269] The acoustic DSP chip (TI TMS320C6748) performs frequency-domain mask calculation;

[0270] The GPU (NVIDIA Jetson AGX) runs the CNN classifier.

[0271] MFN module:

[0272] Deployed on the edge server (Intel Xeon + 4×RTX 6000), supporting TensorRT acceleration.

[0273] Dynamic policy generation module:

[0274] The policy library pre-stores 10 optimization modes and switches them in real time through API calls.

[0275] Real-time adjustment module:

[0276] Adaptive PID controller (Arduino Due), timing calibration unit (Xilinx Zynq FPGA).

[0277] Effect verification module:

[0278] The user terminal (iPad Pro) collects feedback and the data is transmitted back via Wi-Fi 6.

[0279] Cloud knowledge base module:

[0280] AWS S3 bucket + Lambda function, supporting distributed model training.

[0281] Comparative example 1

[0282] (Refer to CN111866439A)

[0283] Solution: Adopt a symmetric microphone array + fixed beamforming algorithm, and install the camera at a low position.

[0284] Defects:

[0285] Insufficient burst noise suppression: Under 60 dB keyboard sound, the voice signal-to-noise ratio only increases by 8 dB (15 dB in the present invention);

[0286] Failure in mobile scenarios: When the deviation of the sound source direction angle > 20°, the main lobe of beamforming shifts, resulting in failed sound pickup.

[0287] Comparative example 2

[0288] (Refer to CN118075418A)

[0289] Solution: Based on the preprocessing + secondary detection process of CNN + LSTM, with a fixed optimization threshold.

[0290] Defects:

[0291] Poor real-time performance: The processing delay of 1080p video reaches 120 ms (45 ms in the present invention);

[0292] Failed dynamic light processing: The picture is overexposed under strong light of 1000 lux (the present invention maintains details through HDR dynamic adjustment).

[0293] Comparative Example 3

[0294] (Refer to CN118214825A)

[0295] Solution: An integrated conference system, relying on a cloud policy library + multi-router architecture.

[0296] Defects:

[0297] Network dependence: The audio-visual synchronization error is > 80 ms under a network jitter of 100 ms (reduced to 20 ms by edge computing in the present invention);

[0298] High maintenance cost: The module needs to be replaced as a whole in case of failure (the present invention supports hot plugging of independent modules).

[0299] Test Case 1

[0300] Speech clarity in a noisy environment

[0301] Condition: Simulate a meeting room (background noise 55 dB, reverberation RT60 = 1.2 s), and 5 people speak simultaneously.

[0302] The results are shown in Table 1

[0303] Table 1 Comparison of speech clarity in a noisy environment

[0304] Index Example 4 Comparative Example 1 Comparative Example 2 Speech Signal-to-Noise Ratio (dB) 22.5 14.3 18.7 Echo Suppression Rate (%) 92 85 88 Response Time (ms) 85 120 150

[0305] Test Case 2

[0306] Stability of the dynamic light picture

[0307] Condition: The light suddenly increases from 100 lux to 1000 lux, and the camera pans quickly.

[0308] The results are shown in Table 2

[0309] Table 2 Comparison of the stability of the dynamic light picture

[0310] Index Example 4 Comparative Example 2 Comparative Example 3 Number of Frames for Overexposed Picture Recovery 3 8 6 Motion Blur Reduction Rate (%) 75 40 55 Face Tracking Accuracy Rate (%) 98 82 90

[0311] Test Case 3

[0312] Synchronization in a weak network environment

[0313] Condition: Simulate network jitter (delay fluctuation 50 - 200 ms, packet loss rate 5%).

[0314] Index Example 4 Comparative Example 3 Audio-Video Synchronization Error (ms) 18 75 Frame Rate Stability (fps) 58±2 45±10 User Satisfaction Score ( / 5) 4.7 3.2

[0315] Through the innovative architecture of domain-based acquisition + multi-modal fusion + edge-cloud collaboration, the present invention is significantly superior to the prior art in the following aspects:

[0316] Improved real-time performance: The video processing delay is reduced by 62.5%, and the voice response time is shortened by 29%.

[0317] Enhanced scene adaptability: In dynamic lighting and noisy environments, the key indicators (signal-to-noise ratio, image stability) are improved by more than 40%.

[0318] Reduced network dependence: The synchronization error in weak network scenarios is reduced by 76%, and an offline optimization mode is supported.

[0319] Optimized user experience: The subjective user score reaches 4.7 / 5, far exceeding the comparative example (the highest is 3.2 / 5).

[0320] This solution provides a highly robust and low-latency audio-video self-optimization solution for conference systems in complex environments.

[0321] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Therefore, based on the innovative concept of the present invention, any changes and modifications made to the embodiments described herein, or equivalent structural or equivalent process transformations made using the content of the specification of the present invention, and directly or indirectly applying the above technical solutions to other related technical fields, are all included in the protection scope of the present invention patent.

Claims

1. A method for self-optimization of conference large-screen audio and video based on deep learning, characterized in that, It includes the following steps: Step S1: Collect multi-modal parameters in domains, including: The first group of parameters: environmental acoustic parameters, including background noise spectrum, reverberation time, and sound source direction angle; The second group of parameters: video dynamic parameters, including dynamic range of light intensity, picture motion vector, and displacement of facial key points; The first group of parameters and the second group of parameters are independently collected through a time-sharing sampling mechanism, and the sampling frequency of the first group of parameters is 1.5 - 2 times that of the second group of parameters; Step S2: Conduct hierarchical progressive analysis on the multi-modal parameters, including: The first layer of analysis: Generate a noise suppression weight matrix based on the acoustic parameters and calculate the reverberation cancellation coefficient; The second layer of analysis: Based on the video dynamic parameters, extract the picture jitter features by the optical flow method, and generate a picture stability score in combination with the displacement of facial key points; Step S3: Integrate the acoustic and video analysis results, and conduct cross-modal correlation evaluation through a deep learning model, including: Input the noise suppression weight matrix, reverberation cancellation coefficient, and picture stability score into the multi-modal fusion network MFN, and output the audio-visual synchronization quality index and the scene adaptation level; Step S4: Dynamically generate an optimized strategy combination based on the output result of Step S3, including: Select the audio enhancement mode and video optimization mode according to the scene adaptation level; If the audio-visual synchronization quality index is lower than the threshold, trigger the timing calibration module to align the audio and video streams; Step S5: Conduct real-time feedback adjustment, including: Adjust the audio noise reduction intensity and video frame rate through an adaptive PID controller to respond to sudden changes in environmental parameters; Step S6: Output the optimized audio and video streams to the conference large screen, and record the optimized strategy parameters in the cloud knowledge base; Step S7: Periodically verify the optimization effect, including: Update the weights of the deep learning model based on user interaction data.

2. The method according to claim 1, wherein In the said Step S1: The sound source direction angle is calculated by the time delay difference TDOA algorithm of the distributed microphone array, and the microphone array is deployed at the edge of the large screen in an asymmetric topology; The dynamic range of light intensity is captured by a dual-exposure sensor at different times for the highlight and dark regions, and fused into HDR metadata.

3. The method according to claim 1, characterized in that In the said Step S2: The noise suppression weight matrix is generated by a frequency domain mask, specifically: divide the background noise spectrum into sub-bands, and dynamically allocate sub-band suppression coefficients based on the signal-to-noise ratio; The picture stability score is evaluated by a convolutional neural network CNN classifier, and the input is the spatio-temporal correlation matrix of the optical flow features and facial key points.

4. The method according to claim 1, characterized in that In the said Step S3: The multi-modal fusion network MFN is a cascaded structure, including: The first stage: A bidirectional LSTM network extracts the temporal features of the acoustic parameters; The second stage: A 3D convolutional network extracts the spatio-temporal features of the video dynamic parameters; The third stage: An attention mechanism module fuses the cross-modal features and outputs the correlation evaluation result.

5. The method according to claim 1, wherein In the said Step S4: The timing calibration module uses the dynamic time warping DTW algorithm to align the audio and video streams, and compensates for the inter-frame differences caused by network latency through interpolation.

6. The method according to claim 1, characterized in that In the said Step S5: The parameters of the adaptive PID controller are dynamically adjusted by a reinforcement learning agent, and the reward function is generated based on the real-time audio-visual synchronization quality index and user feedback data.

7. A conference large-screen audio and video self-optimization system based on deep learning, characterized in that, It includes the following modules: Domain Acquisition Module: Used to collect environmental acoustic parameters and video dynamic parameters at different times, including a distributed microphone array, a dual-exposure sensor, and a motion vector detection unit; Hierarchical Analysis Module: Connected to the Domain Acquisition Module, performing noise suppression weight matrix calculation, reverberation elimination, and generation of picture stability scores; Multi-modal Fusion Network MFN Module: Connected to the Hierarchical Analysis Module, evaluating the audio-visual correlation through a deep learning model; Dynamic Policy Generation Module: Connected to the MFN Module, generating audio-visual optimization policies according to the scene adaptation level; Real-time Adjustment Module: Connected to the Dynamic Policy Generation Module, including an adaptive PID controller and a timing calibration unit; Effect Verification Module: Connected to the Real-time Adjustment Module and the user terminal, collecting feedback data and updating the model weights; Cloud Knowledge Base Module: Bidirectionally connected to all modules, storing optimization policy parameters and user interaction data.

8. A computer-readable storage medium storing a computer program, characterized in that, When the program is executed by a processor, it implements the steps of the method described in any one of claims 1-6, and the storage medium has a distributed architecture, including: Local Cache Unit: Storing the originally collected parameters and intermediate analysis results in real time; Edge Computing Unit: Deploying the MFN model and the adaptive PID controller for low-latency data processing; Cloud Storage Unit: Periodically synchronizing the optimization policy parameters and user feedback data.

9. A computing device, characterized in that, Including: Multi-core Processor: Used to execute acoustic parameter processing and video dynamic analysis in parallel; GPU Acceleration Unit: Specifically used to run the multi-modal fusion network MFN and the reinforcement learning agent; Hardware Encoder: Real-time compressing the optimized audio-visual stream, supporting the H.265 / AV1 encoding protocol.

Citation Information

Patent Citations

  • Conference device and system for optimizing audio and video experience and operation method thereof

    CN111866439A

  • Audio and video integrated device and system for conference

    CN118214825A

  • Spokesman positioning and tracking method and device based on audio and video features

    CN113740803A

  • Audio-visual speech enhancement method based on cross attention and model building method thereof

    CN117854535A

  • Sound source positioning method and system for remote video conference

    CN118264906A

Cited By

  • Student psychological health dynamic assessment method and system based on deep learning

    CN120753654A