Cascade conference treasure based on improved audio and video algorithm and application thereof
By improving the cascading conference treasure system of audio and video algorithms, the synchronization and efficient resource scheduling of multi-terminal audio and video are realized, and the synchronization and scalability problems of traditional systems in multi-level node deployment are solved, improving the conference experience and efficiency.
Patent Information
- Application Number
- CN202510839238.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Traditional audio and video conferencing systems are difficult to achieve audio and video synchronization and efficient resource scheduling in multi-terminal conference scenarios, and are insufficient in scalability and cannot support the cascading deployment of multi-level nodes, which affects the conference experience and efficiency.
A cascading conference treasure system based on improved audio and video algorithms is adopted, including a first acquisition module for unified encoding and division of audio and video data, a first screening module filters key frames, a first judgment module calculates the delay time between nodes, a synchronization module corrects, and integrates it into a unified stream output through the output module.
Improve audio and video synchronization accuracy, reduce system delay error, enhance stability and adaptability, realize hierarchical output management, and support flexible control of multi-level conference scenarios.
Smart Images

Figure CN120358323A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio - video conferencing, and particularly relates to a cascading conference device based on improved audio - video algorithms and its applications. Background Art
[0002] With the popularization of remote work, online education, and distributed collaborative work, audio - video communication devices are increasingly widely used in modern information interaction. Especially in the scenario of multi - party remote conferences, higher requirements are put forward for the audio - video quality, stability, and compatibility of the devices. Traditional conference systems mostly adopt the method of directly connecting a single terminal to the server. When the number of participating nodes increases or the network environment is complex, problems such as audio - video asynchrony, video freezing, and echo interference often occur, affecting the conference experience and communication efficiency.
[0003] To solve the above problems, some improvement measures have been proposed in the prior art. For example, the video transmission quality is optimized through a bandwidth adaptation mechanism, or the voice clarity is improved through an echo cancellation algorithm. However, these methods mostly optimize for a single node or a local network, and it is difficult to achieve audio - video synchronization and efficient resource scheduling in a large - scale distributed multi - terminal conference system. At the same time, existing conference devices have deficiencies in scalability and cannot flexibly support the cascaded deployment of multi - level nodes, making it difficult to meet the complex conference requirements of enterprises across regions and networks.
[0004] In addition, in the multi - terminal conference scenario, the encoding and decoding methods, timing synchronization processing mechanisms, network packet loss recovery capabilities, etc. of the audio - video data of each terminal also directly affect the overall system operation efficiency and user experience. Due to the lack of unified, efficient, and scalable algorithm support, traditional conference systems are difficult to achieve the coordinated operation and cascaded management of multi - level devices while ensuring low - latency and high - quality communication.
[0005] Therefore, there is an urgent need to propose a cascading conference device system based on improved audio - video algorithms. By optimizing the synchronization strategy, enhancing the network adaptation ability, and the module - level cascaded communication strategy, multi - terminal high - efficiency coordination, high - quality audio - video transmission, and intelligent deployment of a wide - area conference system are realized to solve the deficiencies of the prior art. Summary of the Invention
[0006] In view of the above problems, the object of the present invention is to provide: A cascading conference device based on improved audio - video algorithms, comprising: A first acquisition module, configured to acquire the audio - video data of each level of conference nodes, uniformly encode the audio data, and divide the video data into key frames and non - key frames; The first screening module is used to set a reference time interval, screen key frames within the reference time interval as merged data, and arrange the merged data in order of priority; the priority is determined in advance according to the affiliation, participation role or speaking order of the participants; The first determination module is used to calculate the offset in the merged data and, taking this offset as a reference, determine the delay duration between nodes; the offset is the difference in acquisition times of adjacent nodes in the merged data; The synchronization module is used to output a delay instruction based on the determined delay duration and perform synchronization correction on the video and audio; The output module is used to integrate the data after synchronization correction into a unified audio-visual stream and output it to a specified terminal for display.
[0007] In a preferred technical solution, the first acquisition module further includes: a frame determination unit, which is used to divide the video data into key frames and non-key frames according to the frame coding type after acquiring the video data, and the frame determination unit determines the key frames and non-key frames in the following ways: T1: By parsing the frame type field in the video coding format, identify frames with complete image information that do not depend on other frames for image reconstruction, and label them as key frames; label frames that need to rely on forward or bidirectional reference frames for image reconstruction as non-key frames; T2: Set a key frame extraction interval parameter, extract image frames within a preset time period, and make a dynamic judgment based on the degree of image change between adjacent frames. When the image change amplitude exceeds the threshold, determine this frame as a key frame, otherwise determine it as a non-key frame; T3: When it is detected that the meeting scene changes, including speaker switching, shared content insertion or sudden change in the picture structure, force the generation of key frames to ensure subsequent data synchronization and playback continuity.
[0008] In a preferred technical solution, the first screening module specifically includes: The frame classification unit is used to divide the key frames by node and arrange them in the order of node numbers; The interval screening unit is used to set a reference time interval and screen the key frames within this interval to determine the merged data.
[0009] In a preferred technical solution, the first determination module includes: The offset calculation unit is used to count the difference in acquisition times of adjacent nodes in the merged data as the offset; The delay analysis unit is used to calculate the difference between the maximum and minimum offsets and judge the delay duration between nodes based on this.
[0010] In a preferred technical solution, the delay analysis unit includes: A tolerance zone construction module, which is used to expand outward to form an offset tolerance zone based on the start time and end time of the offset interval; An offset correction module, which is used to remove or correct the original offset within the tolerance zone to generate an updated offset; A duration recalculation module, which is used to recalculate the delay duration according to the updated offset.
[0011] In a preferred technical solution, the output module includes: An arrangement unit, which is used to uniformly arrange the synchronized and corrected audio and video data and generate a merged audio-visual stream; An output control unit, which is used to select a terminal as an output node and display the audio-visual stream as an output picture in sequence according to the node priority level.
[0012] In a preferred technical solution, the output module further includes: A performance monitoring unit, which is used to continue receiving the video stream from the upper-level node after the output node outputs the audio-visual stream, and calculate the receiving rate and real-time load usage of this node; A performance evaluation unit, which is used to perform synchronous operations based on the data to obtain a performance score and judge whether the node is in a synchronous output state; A feedback correction unit, which is used to correct its receiving rate and load usage when it is judged that the node is not in a synchronous output state, and regenerate the audio-visual stream to be output; A downstream correction module, which is used to use the audio-visual stream as a new reference to correct the receiving rate and load conditions of the lower-level nodes in sequence, and return the updated result as feedback information to the upper-level node until the first node.
[0013] The present invention also provides a cascaded conference method based on an improved audio-visual algorithm, which is implemented based on the above cascaded conference device, and includes the following steps: S1. Collect audio-visual data: Collect the audio-visual data of each level of conference nodes, uniformly encode the audio data, and divide the video data into key frames and non-key frames; S2. Screen and merge data: Set a reference time interval, screen the key frames within the reference time interval as merged data, and arrange the merged data in sequence according to the priority level; the priority level is determined in advance according to the subordination relationship, participating role or speaking order of the participants; S3. Determine the delay duration: Calculate the offset in the merged data, and use this offset as a reference to determine the delay duration between nodes; S4. Synchronous correction: Output a delay instruction based on the determined delay duration, and perform synchronous correction on the video and audio; S5. Uniformly output the audio-video stream: Integrate the data after synchronous correction into a unified audio-video stream according to the timestamps, and output it to the specified terminal for display.
[0014] In a preferred technical solution, the step S2 includes the following sub-steps: S21. Divide the key frames by nodes and arrange them in the order of node numbers; S22. Set a reference time interval, filter the key frames within this interval, and determine them as merged data.
[0015] In a preferred technical solution, the step S3 includes the following sub-steps: S31. Statistically calculate the acquisition time difference between adjacent nodes in the merged data as the offset; S32. Calculate the maximum offset and the minimum offset, and judge the delay duration between the output node and the upper-level node based on the difference between the two; S33. Construct an offset tolerance zone, remove or correct the original offset within this zone, and obtain the updated offset; S34. Recalculate the delay duration with the updated offset.
[0016] In a preferred technical solution, the step S4 includes the following sub-steps: S41. Obtain a delay instruction, which contains the delay duration between each node and its corresponding synchronization identifier; S42. Adjust the arrangement order of the video frames of each node according to the delay duration, so that the key frames and the audio segments are aligned on the time axis; S43. Perform delay compensation or clipping processing on the audio signals of each node to make the audio outputs of each node synchronized; S44. Integrate the processed video frames and audio data to form a synchronized video stream and audio stream.
[0017] Beneficial effects The present invention provides a cascaded conference system and method based on an improved audio-video algorithm, having the following beneficial effects: 1. Improve the audio-video synchronization accuracy: Through the unified encoding of audio data and the key frame division processing of video data, precise alignment between different nodes is achieved, avoiding the phenomena of picture misalignment and voice lag, and improving the experience coherence of multi-node cascaded conferences.
[0018] 2. Effectively reduce system latency error: By constructing a determination mechanism for offset and latency duration and introducing an offset tolerance zone correction strategy, the present invention can dynamically correct the time deviation between nodes, thereby significantly reducing the latency error caused by network fluctuations or device differences during the cascading process.
[0019] 3. Enhance system stability and adaptability: During the audio-video stream output process, by using a performance scoring mechanism to monitor the reception rate and load usage of each node in real time, it can intelligently judge and correct the state of asynchronous nodes, ensuring the stable operation and smooth output of the overall system in a dynamic environment.
[0020] 4. Achieve hierarchical output management: The present invention supports the orderly sorting and output of audio-video data according to the priority level of nodes, improving the control flexibility and display efficiency in multi-level meeting scenarios, and is applicable to complex application environments such as large groups, remote teaching, and collaborative office.
[0021] 5. The feedback closed-loop mechanism improves the system's self-adaptive ability: Through the feedback mechanism of downward correction and the upper-level nodes, the present invention can form a closed-loop dynamic adjustment path, ensuring rapid response and completion of synchronization correction when the node state changes, and improving the overall fault tolerance and self-healing ability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the system structure of the present invention; Figure 2 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To deepen the understanding of the present invention, the following will further elaborate on the present invention in combination with embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.
[0024] Embodiment 1 According to Figure 1 As shown, this embodiment provides a cascading conference device based on an improved audio-video algorithm, which is used to achieve efficient acquisition, synchronous processing, and unified output of audio-video data in a multi-level conference system, improving the audio-video coordination consistency and stability between conference terminals.
[0025] The cascading conference device includes the following modules: The first acquisition module is used to acquire the audio-video data of each level of conference nodes. After the audio data acquisition is completed, it is uniformly encoded to ensure the audio format consistency between nodes, facilitating subsequent synchronous processing. After the video data is acquired, it is divided into key frames and non-key frames by a frame determination unit.
[0026] The frame determination unit is used to determine key frames and non-key frames based on the following three methods: T1. Analyze the frame type field in the video coding format (such as H.264 or H.265), identify the frames with complete image information and do not rely on other frames for image reconstruction, and label them as key frames; the frames that need to rely on other reference frames for reconstruction are labeled as non-key frames; T2. Set the key frame extraction interval parameter, extract image frames within a preset time period, and judge according to the degree of image difference between the current frame and the previous frame. If the image change amplitude exceeds the set threshold, then determine this frame as a key frame, otherwise it is a non-key frame; T3. When the system detects that the meeting scene changes (such as speaker switching, shared content insertion, sudden change in picture structure, etc.), immediately force the generation of key frames to ensure data integrity and synchronization during the change of meeting content.
[0027] The first screening module is used to screen and sort the key frames in the above-mentioned collected video data. Preferably, the first screening module includes: The frame classification unit is used to classify the key frames according to the nodes they belong to and arrange them in the order of node numbers to ensure that the data of each node has a unified time series structure before merging; The interval screening unit is used to set a reference time interval and screen the key frames within this time interval to determine the data set (i.e., merged data) that can participate in the merging process.
[0028] The first determination module is used to calculate the time difference between nodes based on the merged data, and thus determine the audio-visual synchronization delay between each node. The first determination module includes: The offset calculation unit is used to count the acquisition time difference corresponding to adjacent nodes in the merged data as the offset between each node; the offset is the acquisition time difference between adjacent nodes in the merged data; The delay analysis unit is used to obtain the maximum offset and the minimum offset, and use the difference between the two as the basis for judging the node delay duration.
[0029] The delay analysis unit further includes: The tolerance zone construction module is used to expand outward based on the start time and end time of the above offset interval to form an offset tolerance zone with an allowable error range; The offset correction module is used to eliminate or correct the original offset within this tolerance interval, filter abnormal or distorted time difference data, and obtain the corrected offset; The duration recalculation module is used to recalculate the delay duration between nodes based on the updated offset to ensure the accuracy and robustness of the delay determination result.
[0030] The synchronization module is used to output a synchronization control instruction based on the delay determination result, and uniformly adjust the corresponding video frame sequence and audio track, so that each node outputs synchronously on the unified time axis.
[0031] The output module is used to integrate the data after synchronization correction into a unified audio-visual stream and output it to a specified terminal for display. Preferably, the output module includes: The sorting unit is used to integrate the synchronized audio data and video data according to a unified timing structure and generate a merged standard audio-visual stream; The output control unit is used to select the target terminal as the output node, and sort the merged audio-visual stream according to the priority information of each node and then output it as a display screen; the priority level is determined in advance according to the subordination relationship, participation role or speaking order of the participants.
[0032] The output module further includes: The performance monitoring unit is used to continue receiving the audio-visual stream from the upper-level node after the output node completes data output, and calculate the receiving rate and real-time load situation of the current node; The performance evaluation unit is used to perform synchronization evaluation based on the above two parameters, obtain the current performance score of the node, and use it to judge whether it is in an effective synchronization state; The feedback correction unit is used to perform adaptive correction according to its receiving rate and load situation when it is found that the node is not in a synchronized state, and regenerate the standard audio-visual stream to be output; The downlink correction module is used to use this audio-visual stream as a new benchmark to sequentially correct the lower-level nodes, and send the corrected result back to the upper-level node through the feedback link until the first node, forming a closed-loop feedback mechanism to ensure the continuous effectiveness of the synchronization state of the entire conference link.
[0033] Through the above implementation manner, the present invention can realize intelligent identification, screening, delay determination, synchronization correction and unified output of audio-visual streams of different nodes in a multi-level conference system, significantly improve the audio-visual consistency and interaction stability of the system, and have good real-time performance, scalability and fault tolerance.
[0034] Embodiment 2 This embodiment provides a cascaded conference method based on an improved audio-visual algorithm, which is implemented relying on the above cascaded conference device, aiming to solve technical problems such as audio-visual asynchronization, picture delay, inconsistent output, etc. between multi-level conference nodes, and realize audio-visual consistency, link stability and multi-terminal collaboration during the conference process.
[0035] This method includes the following steps: S1. Collect audio-visual data: In the conference initialization stage, audio and video capture modules are sequentially started by conference nodes at all levels to capture local microphone audio input and camera video streams respectively. The system uniformly performs standardized encoding processing on the captured audio data, such as using AAC or Opus format for compression encoding, to ensure that the audio output by different devices is consistent in format and rate.
[0036] In terms of video data, the system introduces a frame determination mechanism to divide the original video stream into key frames (I-frames) and non-key frames (P-frames / B-frames). Key frames have image self-containment and serve as reference frames for subsequent screening and synchronization. The frame determination methods include: identifying the frame type by parsing the video coding protocol fields; setting the key frame extraction interval (such as 1 frame per second); and dynamically adjusting the key frame generation frequency in combination with the inter-frame image change degree and conference scene changes (such as speaker switching, picture insertion).
[0037] S2. Screen and merge data: To achieve the time-series merging of multi-node videos, the system sets a unified reference time interval (such as 5 seconds). Within this time interval, the following sub-steps are executed: S21. Classify and sort the captured key frames according to the node numbers to ensure a clear and definite data structure on the time axis; S22. Screen out the key frames within this time interval from each node as the merged data set to form a merged frame set with a unified time base and multi-source structure for subsequent delay judgment and synchronization processing.
[0038] S3. Determine the delay duration: After obtaining the merged key frame data of multiple nodes, the system performs time difference analysis based on the frame capture time to determine the audio and video delay durations between different nodes. The function expression is: Any two nodes , The time difference between key frames ; The maximum delay duration ; The maximum capture time among all nodes:
[0039] The minimum capture time among all nodes:
[0040] Among them, represents the total number of conference nodes participating in the merge; represents the th node's capture timestamp of the uploaded key frame This step specifically includes: S31. Statistically calculate the difference in the acquisition time of key frames between any two adjacent nodes within the same time interval, and preliminarily calculate the offset; S32. Extract the maximum and minimum values from all node offsets, calculate their difference, and use it as the maximum delay between nodes; S33. To eliminate errors caused by network jitter or abnormal acquisition, construct an offset tolerance zone (such as ±200 ms), and eliminate abnormal offset data; S34. Re-evaluate the remaining offsets, and use the updated data as the final delay duration for generating subsequent synchronization instructions.
[0041] S4. Synchronization correction: Based on the delay data obtained in the previous step, the system generates synchronization instructions and distributes them to all conference nodes to unify the audio-visual output timeline of each node. The correction process includes: S41. Parse the delay instructions to obtain the delay duration between each node and the reference benchmark and the corresponding synchronization identifier; S42. Adjust the arrangement order of the video frames of each node. Through frame buffering and delay frame interpolation means, align the key frames with the audio segments at the same time stamp; S43. Perform duration compensation on the audio tracks of the nodes, and crop the redundant front segments or fill in the blank segments when necessary to achieve audio track synchronization; S44. After completing the timeline reconstruction, integrate the video frames and audio data together to generate synchronized video and audio streams, ensuring that the output of each node is consistent, continuous, and uninterrupted.
[0042] S5. Unified output of audio-visual streams: Integrate the synchronized multi-node audio-visual data and output it to the target terminal. The system determines the display node (such as the speaker or the main venue) according to the set hierarchical priority strategy, and arranges and outputs the audio-visual streams in the order of node priority levels. The output method supports functions such as multi-view display, primary-secondary screen switching, and automatic update of voice-activated views, ensuring that end users can obtain a clear, low-latency, and well-synchronized conference experience; the hierarchical priority strategy satisfies: participants with a higher subordination relationship have a higher priority level, the conference host or speaker has a higher priority level, and the first speaker according to the conference agenda has a higher priority level.
[0043] In the further implementation process, the system can also dynamically detect the reception rate and CPU / GPU load of each output node, and judge whether it is in a synchronized state based on real-time scoring; if there is a trend of desynchronization, the system can start a feedback mechanism, perform adaptive correction at the node level, and complete the closed-loop repair of the link through upstream and downstream conduction feedback.
[0044] In summary, the method of this embodiment can achieve high-precision audio and video alignment and unified output in complex multi-node conference scenarios, improve the efficiency of remote collaboration, and has good real-time performance, adaptability, and system robustness.
[0045] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A cascaded conference device based on an improved audio-video algorithm, characterized in that Including: A first acquisition module, configured to acquire audio and video data of each level of conference nodes, uniformly encode the audio data, and divide the video data into key frames and non-key frames; A first screening module, configured to set a reference time interval, screen key frames within the reference time interval as merged data, and arrange the merged data in order of priority level; the priority level is determined in advance according to the subordination relationship, participation role or speaking order of the participants; A first determination module, configured to calculate the offset in the merged data, and use this offset as a reference to determine the delay duration between nodes; the offset is the acquisition time difference between adjacent nodes in the merged data; A synchronization module, configured to output a delay instruction based on the determined delay duration, and perform synchronous correction on the video and audio; An output module, configured to integrate the data after synchronous correction into a unified audio and video stream, and output it to a specified terminal for display.
2. The cascade conference device based on an improved audio - video algorithm according to claim 1, wherein: The first screening module specifically includes: A frame classification unit, configured to divide the key frames by node and arrange them in the order of node numbers; An interval screening unit, configured to set a reference time interval, and screen the key frames within this interval to determine them as merged data.
3. The cascade conference device based on an improved audio-video algorithm according to claim 1, wherein: The first determination module includes: An offset calculation unit, configured to count the acquisition time difference between adjacent nodes in the merged data as the offset; A delay analysis unit, configured to calculate the difference between the maximum and minimum offsets, and use this to judge the delay duration between nodes.
4. The cascade conference device based on the improved audio and video algorithm according to claim 3, characterized in that: The delay analysis unit includes: A tolerance zone construction module, configured to expand outward based on the start time and end time of the offset interval to form an offset tolerance zone; An offset correction module, configured to remove or correct the original offset within this tolerance zone to generate an updated offset; A duration recalculation module, configured to recalculate the delay duration according to the updated offset.
5. The cascade conference device based on the improved audio and video algorithm according to claim 1, characterized in that: The output module includes: An arrangement unit, configured to uniformly arrange the audio and video data after synchronous correction, and generate a merged audio and video stream; An output control unit, configured to select a terminal as the output node, and display this audio and video stream as an output picture in order of the node priority level.
6. The cascade conference device based on an improved audio - video algorithm according to claim 5, wherein: The output module further includes: A performance monitoring unit, configured to continue to receive the video stream from the upper-level node after the output node outputs the audio and video stream, and calculate the reception rate and real-time load usage of this node; A performance evaluation unit, configured to perform synchronous operations based on the data, obtain a performance score, and judge whether the node is in a synchronous output state; A feedback correction unit, configured to correct its reception rate and load usage when it is judged that the node is not in a synchronous output state, and regenerate the audio and video stream to be output; A downstream correction module, configured to use the audio and video stream as a new reference, sequentially correct the reception rate and load conditions of the lower-level nodes, and return the updated result as feedback information to the upper-level node until the first node.
7. A cascaded conference method based on an improved audio-visual algorithm, implemented based on the cascaded conference device described in any one of claims 1-6, characterized in that, Including the following steps: S1. Acquire audio and video data: Acquire audio and video data of each level of conference nodes, uniformly encode the audio data, and divide the video data into key frames and non-key frames; S2. Screen and merge data: Set a reference time interval, filter out the key frames within the reference time interval as merged data, and arrange the merged data in order of priority level; S3. Determine the delay duration: Calculate the offset in the merged data, and use this offset as a reference to determine the delay duration between nodes; S4. Synchronization correction: Based on the determined delay duration, output a delay instruction, and perform synchronization correction on the video and audio; S5. Uniformly output the audio-video stream: Integrate the data after synchronization correction into a unified audio-video stream according to the time stamp, and output it to a specified terminal for display.
8. The cascade conference method based on an improved audio-video algorithm according to claim 7, characterized in that: The step S2 includes the following sub-steps: S21. Divide the key frames by node and arrange them in the order of node numbers; S22. Set a reference time interval, filter out the key frames within this interval, and determine them as merged data.
9. The cascade meeting method based on an improved audio and video algorithm according to claim 8, characterized in that: The step S3 includes the following sub-steps: S31. Statistically calculate the acquisition time difference between adjacent nodes in the merged data as the offset; S32. Calculate the maximum offset and the minimum offset, and judge the delay duration between the output node and the upper-level node based on the difference between the two; S33. Construct an offset tolerance zone, remove or correct the original offset within this zone to obtain an updated offset; S34. Recalculate the delay duration with the updated offset.
10. The cascade conference method based on the improved audio and video algorithm according to claim 9, wherein: The step S4 includes the following sub-steps: S41. Obtain a delay instruction, which contains the delay duration between each node and its corresponding synchronization flag; S42. Adjust the arrangement order of the video frames of each node according to the delay duration to align the key frames with the audio segments on the time axis; S43. Perform delay compensation or clipping processing on the audio signals of each node to make the audio outputs of each node synchronized; S44. Integrate the processed video frames and audio data to form a synchronized video stream and audio stream.
Citation Information
Patent Citations
Image processing method and device
CN108255299A
Message display method and system, wearable device and storage medium
CN109814723A
Audio and video synchronization method and device, electronic equipment and storage medium
CN117097936A
Synchronous acquisition method and system for sound signals and video signals
CN119052545A
Method and apparatus for controlling a multimedia conference by an application server
US20110185021A1
Cited By
Video-based encoder delay automatic measurement method
CN121442088A