A cascade conference system based on improved audio and video algorithm and its application

By improving the cascading conference treasure of audio and video algorithms, unified encoding, synchronous correction and priority output of multi-terminal audio and video data is achieved, which solves the synchronization and scalability problems of traditional conference systems under multi-level nodes, and improves conference experience and efficiency.

CN120358323BActive Publication Date: 2025-08-22HANGZHOU ALMIGHTY DIGIT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510839238.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Traditional conference systems are difficult to achieve audio and video synchronization and efficient resource scheduling in multi-terminal conference scenarios, and are insufficient in scalability and cannot flexibly support the cascading deployment of multi-level nodes, which affects the conference experience and communication efficiency.

Method used

The cascading conference treasure based on improved audio and video algorithm is adopted. Through the acquisition module, the audio and video data is uniformly encoded, the keyframes are filtered, the delay time is determined, the synchronization module performs correction, and the audio and video stream is output according to priority levels through the output module, supporting the coordinated operation of multi-level nodes.

Benefits of technology

Improve audio and video synchronization accuracy, reduce system delay error, enhance stability and adaptability, realize hierarchical output management, and support efficient collaboration in complex application environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358323B_ABST
    Figure CN120358323B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of audio and video conferencing, and specifically relates to a cascade conference system based on an improved audio and video algorithm and its application, which is suitable for audio and video synchronization and output control in multi-level conference scenarios. The system includes a first acquisition module, a first screening module, a first determination module, a synchronization module and an output module, which are respectively used to collect node audio and video data, screen key frames, calculate delay duration, perform synchronization correction and uniformly output audio and video streams. Among them, the video data is divided into key frames and non-key frames through a frame determination mechanism, and the delay between nodes is determined based on the time offset, and then the audio and video time axis alignment is achieved through synchronization instructions, and finally integrated into a unified data stream for terminal display. This method improves the synchronization accuracy, stability and scalability of the multi-node conference system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio and video conferencing, and specifically relates to a cascade conference device based on an improved audio and video algorithm and its application. Background Art

[0002] With the rise of remote work, online education, and distributed collaborative office environments, audio and video communication equipment is increasingly being used in modern information exchange. This is especially true in multi-party remote conferencing scenarios, which place higher demands on the equipment's audio and video quality, stability, and compatibility. Traditional conferencing systems often rely on a single terminal connected directly to the server. As the number of participating nodes increases or the network environment becomes complex, issues such as audio and video asynchrony, video freezes, and echo interference can easily occur, impacting the meeting experience and communication efficiency.

[0003] To address these issues, existing technologies have proposed several improvements, such as optimizing video transmission quality through bandwidth adaptation mechanisms or improving speech clarity through echo cancellation algorithms. However, these approaches primarily target single nodes or local networks, making it difficult to achieve audio and video synchronization and efficient resource scheduling in large, distributed, multi-terminal conferencing systems. Furthermore, existing conferencing equipment lacks scalability and cannot flexibly support cascaded deployment of multiple nodes, making it difficult to meet the complex, cross-regional, and cross-network conferencing requirements of enterprises.

[0004] Furthermore, in multi-terminal conferencing scenarios, the encoding and decoding methods for audio and video data, timing synchronization mechanisms, and network packet loss recovery capabilities of each terminal directly impact the overall system's operational efficiency and user experience. Due to the lack of unified, efficient, and scalable algorithm support, traditional conferencing systems struggle to achieve coordinated operation and cascade management of multiple layers of devices while ensuring low latency and high-quality communication.

[0005] Therefore, it is urgent to propose a cascade conference system based on improved audio and video algorithms. By optimizing synchronization strategies, enhancing network adaptability and modular cascade communication strategies, efficient multi-terminal collaboration, high-quality audio and video transmission, and intelligent deployment of wide-area conference systems can be achieved to address the shortcomings of existing technologies. Summary of the Invention

[0006] In view of the above problems, the present invention aims to propose a cascade conference system based on an improved audio and video algorithm, comprising:

[0007] The first acquisition module is used to collect audio and video data of conference nodes at all levels, uniformly encode the audio data, and divide the video data into key frames and non-key frames;

[0008] A first screening module is configured to set a reference time interval, screen key frames within the reference time interval as merged data, and arrange the merged data in order of priority; the priority levels are predetermined based on the affiliation of the participants, the roles of the participants, or the order of speaking;

[0009] A first determination module is configured to calculate an offset in the merged data and use the offset as a reference to determine the delay duration between nodes; the offset is the difference in acquisition time between adjacent nodes in the merged data;

[0010] A synchronization module, configured to output a delay instruction based on the determined delay duration and perform synchronization correction on the video and audio;

[0011] The output module is used to integrate the synchronized and corrected data into a unified audio and video stream and output it to the designated terminal for display.

[0012] In a preferred technical solution, the first acquisition module further includes: a frame determination unit, configured to, after acquiring the video data, divide the video data into key frames and non-key frames according to the frame coding type, wherein the frame determination unit determines the key frames and non-key frames in the following manner:

[0013] T1: By parsing the frame type field in the video encoding format, frames with complete image information and that do not rely on other frames for image reconstruction are identified as key frames. Frames that require forward or bidirectional reference frames for image reconstruction are identified as non-key frames.

[0014] T2: Set the key frame extraction interval parameter, extract image frames within the preset time period, and make dynamic judgments based on the degree of image change between adjacent frames. When the image change amplitude exceeds the threshold, the frame is determined as a key frame, otherwise it is determined as a non-key frame;

[0015] T3: When a change in the conference scene is detected, including speaker switching, shared content insertion, or sudden change in the screen structure, key frames are forcibly generated to ensure subsequent data synchronization and playback continuity.

[0016] In a preferred technical solution, the first screening module specifically includes:

[0017] A frame classification unit is used to divide the key frames into nodes and arrange them in order according to the node numbers;

[0018] The interval screening unit is used to set a reference time interval and screen the key frames within the interval to determine them as merged data.

[0019] In a preferred technical solution, the first determination module includes:

[0020] The offset calculation unit is used to calculate the difference in collection time between adjacent nodes in the merged data as the offset;

[0021] The delay analysis unit is used to calculate the difference between the maximum and minimum offsets and use it to determine the delay duration between nodes.

[0022] In a preferred technical solution, the delay analysis unit includes:

[0023] A tolerance zone construction module is used to expand outward to form an offset tolerance zone based on the start time and end time of the offset interval;

[0024] An offset correction module, configured to remove or correct the original offset within the tolerance zone to generate an updated offset;

[0025] The duration recalculation module is used to recalculate the delay duration according to the updated offset.

[0026] In a preferred technical solution, the output module includes:

[0027] The collating unit is used to uniformly collate the synchronized and corrected audio and video data and generate a combined audio and video stream;

[0028] The output control unit is used to select a terminal as an output node and display the audio and video stream as an output screen in sequence according to the node priority.

[0029] In a preferred technical solution, the output module further includes:

[0030] The performance monitoring unit is used to continue receiving the video stream from the upper node after the output node outputs the audio and video stream, and calculate the receiving rate and real-time load usage of the node;

[0031] a performance evaluation unit, configured to perform synchronization operations based on the data, obtain a performance score, and determine whether the node is in a synchronous output state;

[0032] A feedback correction unit, configured to correct the receiving rate and load usage of a node when it is determined that the node is not in a synchronous output state, and regenerate the audio and video stream to be output;

[0033] The downlink correction module is used to use the audio and video stream as a new benchmark to correct the receiving rate and load conditions of the lower-level nodes in turn, and return the updated results as feedback information to the upper-level node until the first node.

[0034] The present invention also provides a cascade conference method based on an improved audio and video algorithm, which is implemented based on the cascade conference treasure and includes the following steps:

[0035] S1. Collect audio and video data:

[0036] Collect audio and video data from conference nodes at all levels, uniformly encode the audio data, and divide the video data into key frames and non-key frames;

[0037] S2. Filter and merge data:

[0038] Setting a reference time interval, selecting key frames within the reference time interval as merged data, and arranging the merged data in order of priority; the priority levels are predetermined based on the affiliation of participants, their roles, or the order of speaking;

[0039] S3. Determine the delay time:

[0040] Calculate the offset in the merged data and use it as a reference to determine the delay between nodes;

[0041] S4, synchronous correction:

[0042] Based on the determined delay duration, a delay instruction is output, and synchronization correction is performed on the video and audio;

[0043] S5. Unified output of audio and video streams:

[0044] The synchronized and corrected data is integrated into a unified audio and video stream according to the timestamp, and output to the designated terminal for display.

[0045] In a preferred technical solution, step S2 includes the following sub-steps:

[0046] S21, dividing the key frames into nodes and arranging them in order according to the node numbers;

[0047] S22: Set a reference time interval, and filter key frames within the interval to determine them as merged data.

[0048] In a preferred technical solution, step S3 includes the following sub-steps:

[0049] S31, counting the acquisition time differences of adjacent nodes in the merged data as an offset;

[0050] S32. Calculate the maximum offset and the minimum offset, and determine the delay between the output node and the upper node based on the difference between the two.

[0051] S33, constructing an offset tolerance zone, removing or correcting the original offset within the zone to obtain an updated offset;

[0052] S34. Recalculate the delay duration using the updated offset.

[0053] In a preferred technical solution, step S4 includes the following sub-steps:

[0054] S41, obtaining a delay instruction, which includes the delay duration between each node and its corresponding synchronization identifier;

[0055] S42, adjusting the order of the video frames of each node according to the delay duration so that the key frames and the audio clips are aligned on the time axis;

[0056] S43, performing delay compensation or clipping processing on the audio signal of each node to make the audio output of each node synchronized;

[0057] S44: Integrate the processed video frames and audio data to form synchronized video streams and audio streams.

[0058] Beneficial effects

[0059] The present invention provides a cascade conference system and method based on an improved audio and video algorithm, which has the following beneficial effects:

[0060] 1. Improved audio and video synchronization accuracy: Through unified encoding of audio data and key frame division processing of video data, precise alignment between different nodes is achieved, avoiding image dislocation and voice lag, and improving the experience consistency of multi-node cascade meetings.

[0061] 2. Effectively reduce system delay error: This invention constructs a determination mechanism for offset and delay duration and introduces an offset tolerance zone correction strategy to dynamically correct the time deviation between nodes, thereby significantly reducing the delay error caused by network fluctuations or device differences during the cascade process.

[0062] 3. Enhanced system stability and adaptability: During the audio and video stream output process, the receiving rate and load usage of each node are monitored in real time through a performance scoring mechanism. It can intelligently judge and correct the status of asynchronous nodes, ensuring the stable operation and smooth output of the entire system in a dynamic environment.

[0063] 4. Realize hierarchical output management: This invention supports the orderly organization and output of audio and video data according to node priority, improving the control flexibility and presentation efficiency in multi-level conference scenarios. It is suitable for complex application environments such as large groups, distance learning, and collaborative offices.

[0064] 5. Feedback closed-loop mechanism improves system adaptability: Through the feedback mechanism of downstream correction and upper-level nodes, the present invention can form a closed-loop dynamic adjustment path to ensure rapid response and completion of synchronous correction when the node status changes, thereby improving the overall fault tolerance and self-healing capabilities of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 Schematic diagram of the system structure of the present invention;

[0066] Figure 2 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0067] In order to deepen the understanding of the present invention, the present invention will be further described in detail below with reference to the examples. The examples are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention.

[0068] Example 1

[0069] according to Figure 1 As shown, this embodiment provides a cascade conference treasure based on an improved audio and video algorithm, which is used to achieve efficient collection, synchronous processing and unified output of audio and video data in a multi-level conference system, and improve the coordination consistency and stability of audio and video between conference terminals.

[0070] The cascade conference treasure includes the following modules:

[0071] The first acquisition module is responsible for collecting audio and video data from all levels of conference nodes. After audio data is collected, it is uniformly encoded to ensure audio format consistency across all nodes, facilitating subsequent synchronous processing. After video data is collected, the frame determination unit classifies the video data into key frames and non-key frames.

[0072] The frame determination unit is used to determine key frames and non-key frames based on the following three methods:

[0073] T1. Parse the frame type field in the video encoding format (such as H.264 or H.265) to identify frames that have complete image information and do not rely on other frames for image reconstruction. These frames are marked as key frames. Frames that require other reference frames for reconstruction are marked as non-key frames.

[0074] T2. Set the key frame extraction interval parameter, extract image frames within the preset time period, and judge based on the image difference between the current frame and the previous frame. If the image change exceeds the set threshold, the frame is determined to be a key frame, otherwise it is a non-key frame;

[0075] T3. When the system detects a change in the conference scene (such as speaker switching, shared content insertion, sudden change in screen structure, etc.), it immediately forces the generation of key frames to ensure data integrity and synchronization during the change of conference content.

[0076] The first screening module is used to screen and sort the key frames in the video data collected above. Preferably, the first screening module includes:

[0077] The frame classification unit is used to classify key frames according to the nodes they belong to and arrange them in the order of node numbers to ensure that the data of each node has a unified time series structure before merging;

[0078] The interval screening unit is used to set a reference time interval and screen the key frames within the time interval to determine the data set that can participate in the merging process (ie, the merged data).

[0079] The first determination module is used to calculate the time difference between nodes based on the merged data, so as to determine the audio and video synchronization delay between each node. The first determination module includes:

[0080] An offset calculation unit, configured to calculate the acquisition time differences corresponding to adjacent nodes in the merged data as offsets between the nodes; the offsets are the acquisition time differences between adjacent nodes in the merged data;

[0081] The delay analysis unit is used to obtain the maximum offset and the minimum offset, and use the difference between the two as the basis for determining the node delay duration.

[0082] The delay analysis unit also includes:

[0083] A tolerance zone construction module is used to expand outward based on the start time and end time of the above offset interval to form an offset tolerance zone with an allowable error range;

[0084] An offset correction module is used to remove or correct the original offset within the tolerance interval, filter out abnormal or distorted time difference data, and obtain a corrected offset;

[0085] The duration recalculation module is used to recalculate the delay duration between nodes based on the updated offset to ensure the accuracy and robustness of the delay determination results.

[0086] The synchronization module is used to output synchronization control instructions based on the delay determination result, and uniformly adjust the corresponding video frame sequence and audio track so that each node can be output synchronously on a unified timeline.

[0087] The output module is used to integrate the synchronized and corrected data into a unified audio and video stream and output it to a designated terminal for display. Preferably, the output module includes:

[0088] The collating unit is used to integrate the synchronized and corrected audio data and video data according to a unified timing structure and generate a combined standard audio and video stream;

[0089] The output control unit is used to select the target terminal as the output node, and sort the merged audio and video streams according to the priority information of each node and output them as a display screen; the priority is predetermined according to the affiliation of the participants, the role of the participants or the order of speaking.

[0090] The output module also includes:

[0091] The performance monitoring unit is used to continue receiving audio and video streams from the upper node after the output node completes data output, and calculate the receiving rate and real-time load of the current node;

[0092] A performance evaluation unit is used to perform synchronization evaluation based on the above two parameters to obtain the current performance score of the node to determine whether it is in a valid synchronization state;

[0093] A feedback correction unit is used to perform adaptive correction based on the receiving rate and load conditions when it finds that the node is not in a synchronized state, and regenerate the standard audio and video stream to be output;

[0094] The downlink correction module is used to use the audio and video stream as a new benchmark to correct the lower-level nodes in sequence, and transmit the corrected results back to the upper-level node through the feedback link until the first node, forming a closed-loop feedback mechanism to ensure that the synchronization status of the entire conference link remains valid.

[0095] Through the above-mentioned implementation methods, the present invention can realize intelligent identification, screening, delay determination, synchronization correction and unified output of audio and video streams of different nodes in a multi-level conferencing system, significantly improving the audio and video consistency and interactive stability of the system, and having good real-time, scalability and fault tolerance.

[0096] Example 2

[0097] This embodiment provides a cascade conference method based on an improved audio and video algorithm, which is implemented based on the above-mentioned cascade conference device. It aims to solve technical problems such as audio and video asynchrony, picture delay, and inconsistent output between multi-level conference nodes, and achieve consistent audio and video, stable links and multi-terminal collaboration during the conference process.

[0098] The method comprises the following steps:

[0099] S1. Collect audio and video data:

[0100] During the conference initialization phase, each node in the conference system sequentially activates the audio and video capture modules to collect local microphone audio input and camera video streams. The system then applies standardized encoding to the collected audio data, such as using AAC or Opus compression, to ensure consistent audio output format and rate across different devices.

[0101] Regarding video data, the system incorporates a frame determination mechanism, dividing the raw video stream into keyframes (I-frames) and non-keyframes (P-frames / B-frames). Keyframes are self-contained and serve as reference frames for subsequent filtering and synchronization. This frame determination method includes: identifying frame types by parsing the video encoding protocol fields; setting a keyframe extraction interval (e.g., 1 frame per second); and dynamically adjusting the keyframe generation frequency based on inter-frame image changes and conference scene changes (e.g., speaker switching, screen insertion).

[0102] S2. Filter and merge data:

[0103] To achieve the time series merging of multi-node videos, the system sets a unified reference time interval (for example, 5 seconds). Within this time interval, the following sub-steps are performed:

[0104] S21, classifying and sorting the collected key frames according to the node numbers to ensure that the data structure on the time axis is clear and unambiguous;

[0105] S22. Filter out key frames within the time interval from each node as a merged data set to form a merged frame set with a unified time base and multi-source structure for subsequent delay judgment and synchronization processing.

[0106] S3. Determine the delay time:

[0107] After obtaining the merged keyframe data of multiple nodes, the system performs time difference analysis based on the frame acquisition time to determine the audio and video delay duration between different nodes. The function is expressed as:

[0108] Any two nodes 、 The time difference between key frames ;

[0109] Maximum delay time ;

[0110] Maximum acquisition time among all nodes:

[0111] Minimum collection time among all nodes:

[0112] in, Indicates the total number of conference nodes participating in the merger; Indicates the The acquisition timestamp of the key frame uploaded by each node

[0113] This step specifically includes:

[0114] S31, counting the acquisition time difference of key frames of any two adjacent nodes in the same interval, and preliminarily calculating the offset;

[0115] S32. Extract the maximum and minimum values ​​from all node offsets and calculate their difference as the maximum delay between nodes.

[0116] S33. To eliminate errors caused by network jitter or acquisition anomalies, an offset tolerance zone (e.g., ±200ms) is established to eliminate abnormal offset data.

[0117] S34. Re-evaluate the remaining offset and use the updated data as the final delay duration for generating subsequent synchronization instructions.

[0118] S4, synchronous correction:

[0119] Based on the delay data obtained in the previous step, the system generates synchronization instructions and distributes them to all conference nodes to unify the audio and video output timelines of each node. The correction process includes:

[0120] S41, parsing the delay instruction to obtain the delay duration and corresponding synchronization identifier between each node and the reference benchmark;

[0121] S42: Adjust the arrangement order of the video frames of each node, and align the key frames and the audio clips at the same timestamp by means of frame buffering and delayed frame insertion;

[0122] S43. Compensate the duration of the audio track of the node, trim the leading redundant segments or fill the blank segments if necessary, to achieve audio track synchronization;

[0123] S44. After completing the time axis reconstruction, the video frames and audio data are integrated together to generate synchronized video streams and audio streams to ensure that the output of each node is consistent, continuous, and uninterrupted.

[0124] S5. Unified output of audio and video streams:

[0125] The system integrates synchronized multi-node audio and video data and outputs it to the target terminal. The system determines the presentation node (such as the speaker or main venue) based on the set hierarchical priority strategy and outputs the audio and video streams in an orderly arranged order according to the node priority. The output supports multi-view presentation, primary and secondary screen switching, and voice-activated automatic view updates, ensuring that end users receive a clear, low-latency, and well-synchronized conference experience. The hierarchical priority strategy ensures that participants with higher affiliations have higher priority, the conference host or speaker has higher priority, and the first speaker according to the conference agenda has higher priority.

[0126] During further implementation, the system can also dynamically detect the receiving rate and CPU / GPU load of each output node, and determine whether it is in a synchronous state based on real-time scoring; if an asynchronous trend occurs, the system can activate the feedback mechanism, perform node-level adaptive correction, and complete the link closed-loop repair through upstream and downstream feedback.

[0127] In summary, the method of this embodiment can achieve high-precision audio and video alignment and unified output in complex multi-node conference scenarios, improve remote collaboration efficiency, and has good real-time performance, adaptability, and system robustness.

[0128] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A cascade conference treasure based on an improved audio and video algorithm, characterized in that: include: The first acquisition module is used to collect audio and video data of conference nodes at all levels, uniformly encode the audio data, and divide the video data into key frames and non-key frames; A first screening module is configured to set a reference time interval, screen key frames within the reference time interval as merged data, and arrange the merged data in order of priority; the priority levels are predetermined based on the affiliation of the participants, the roles of the participants, or the order of speaking; A first determination module is configured to calculate an offset in the merged data and use the offset as a reference to determine the delay duration between nodes; the offset is the difference in acquisition time between adjacent nodes in the merged data; A synchronization module, configured to output a delay instruction based on the determined delay duration and perform synchronization correction on the video and audio; The output module is used to integrate the synchronized and corrected data into a unified audio and video stream and output it to the designated terminal for display.

2. The cascade conference treasure based on the improved audio and video algorithm according to claim 1 is characterized by: The first screening module specifically includes: A frame classification unit is used to divide the key frames into nodes and arrange them in order according to the node numbers; The interval screening unit is used to set a reference time interval and screen the key frames within the interval to determine them as merged data.

3. The cascade conference treasure based on the improved audio and video algorithm according to claim 1 is characterized by: The first determination module includes: The offset calculation unit is used to calculate the difference in collection time between adjacent nodes in the merged data as the offset; The delay analysis unit is used to calculate the difference between the maximum and minimum offsets and use it to determine the delay duration between nodes.

4. The cascade conference system based on the improved audio and video algorithm according to claim 3 is characterized by: The delay analysis unit includes: A tolerance zone construction module is used to expand outward to form an offset tolerance zone based on the start time and end time of the offset interval; An offset correction module, configured to remove or correct the original offset within the tolerance zone to generate an updated offset; The duration recalculation module is used to recalculate the delay duration according to the updated offset.

5. The cascade conference treasure based on the improved audio and video algorithm according to claim 1 is characterized by: The output module includes: The collating unit is used to uniformly collate the synchronized and corrected audio and video data and generate a combined audio and video stream; The output control unit is used to select a terminal as an output node and display the audio and video stream as an output screen in sequence according to the node priority.

6. The cascade conference treasure based on the improved audio and video algorithm according to claim 5 is characterized by: The output module also includes: The performance monitoring unit is used to continue receiving the video stream from the upper node after the output node outputs the audio and video stream, and calculate the receiving rate and real-time load usage of the node; a performance evaluation unit, configured to perform synchronization operations based on the data, obtain a performance score, and determine whether the node is in a synchronous output state; A feedback correction unit, configured to correct the receiving rate and load usage of a node when it is determined that the node is not in a synchronous output state, and regenerate the audio and video stream to be output; The downlink correction module is used to use the audio and video stream as a new benchmark to correct the receiving rate and load conditions of the lower-level nodes in turn, and return the updated results as feedback information to the upper-level node until the first node.

7. A cascade conference method based on an improved audio and video algorithm, implemented based on the cascade conference treasure according to any one of claims 1 to 6, characterized in that: The following steps are involved: S1. Collect audio and video data: Collect audio and video data from conference nodes at all levels, uniformly encode the audio data, and divide the video data into key frames and non-key frames; S2. Filter and merge data: Set the reference time interval, filter the key frames within the reference time interval as the merged data, and sort the merged data in order of priority; S3. Determine the delay time: Calculate the offset in the merged data and use it as a reference to determine the delay between nodes; S4, synchronous correction: Based on the determined delay duration, a delay instruction is output, and synchronization correction is performed on the video and audio; S5. Unified output of audio and video streams: The synchronized and corrected data is integrated into a unified audio and video stream according to the timestamp, and output to the designated terminal for display.

8. The cascade conferencing method based on the improved audio and video algorithm according to claim 7, characterized in that: The step S2 includes the following sub-steps: S21, dividing the key frames into nodes and arranging them in order according to the node numbers; S22: Set a reference time interval, and filter key frames within the interval to determine them as merged data.

9. The cascade conferencing method based on the improved audio and video algorithm according to claim 8, characterized in that: The step S3 includes the following sub-steps: S31, counting the acquisition time differences of adjacent nodes in the merged data as an offset; S32. Calculate the maximum offset and the minimum offset, and determine the delay between the output node and the upper node based on the difference between the two. S33, constructing an offset tolerance zone, removing or correcting the original offset within the zone to obtain an updated offset; S34. Recalculate the delay duration using the updated offset.

10. The cascade conferencing method based on the improved audio and video algorithm according to claim 9, characterized in that: The step S4 includes the following sub-steps: S41, obtaining a delay instruction, which includes the delay duration between each node and its corresponding synchronization identifier; S42, adjusting the order of the video frames of each node according to the delay duration so that the key frames and the audio clips are aligned on the time axis; S43, performing delay compensation or clipping processing on the audio signal of each node to make the audio output of each node synchronized; S44: Integrate the processed video frames and audio data to form synchronized video streams and audio streams.

Citation Information

Patent Citations

  • Image processing method and device

    CN108255299A

  • Message display method and system, wearable device and storage medium

    CN109814723A