Conference content intelligent synchronization method and system

Through multimodal data stream processing with global coarse alignment and local fine alignment, combined with dynamic resource allocation, the synchronization instability problem of traditional conference content synchronization systems in complex network environments is solved, and efficient resource optimization configuration and timely transmission of key information are achieved.

CN120583254BActive Publication Date: 2025-10-03GUANGZHOU DAZZLE VIEW INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511090741.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-03
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Traditional conference content synchronization systems have difficulty accurately identifying sudden network degradation in complex network environments, resulting in problems such as audio and video asynchrony, delayed document operations, and loss of key information. Existing technologies lack multi-dimensional network assessment and dynamic resource allocation strategies.

Method used

Global coarse alignment of multimodal data streams is achieved through the server timeline, local timing deviations are corrected by combining dynamic interpolation compensation algorithms, multi-dimensional network indicators are calculated to construct a comprehensive load score, and a three-level dynamic resource allocation strategy is implemented to prioritize the transmission of core data streams.

Benefits of technology

It significantly improves the accuracy and stability of conference content synchronization in complex network environments, achieves optimized resource allocation in high-load or weak network scenarios, and ensures the smoothness and real-time nature of meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583254B_ABST
    Figure CN120583254B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for intelligent synchronization of conference content, which relate to the technical field of intelligent synchronization of conference content. The present invention realizes global coarse alignment of multimodal data streams through a server timeline, and then dynamically compensates for local timing deviations of audio streams and document streams based on the video stream. Then, the audio stability score, video complexity score and text importance score in the conference synchronization content are calculated, and each score is corrected by the calculated global time synchronization error and local time synchronization error. Then, a dynamic network jitter coefficient calculation model is constructed. Finally, the obtained network jitter coefficient is weightedly fused with the audio stability score, video complexity score and text importance score to obtain a comprehensive load score. Through a three-level dynamic resource allocation mechanism that links the comprehensive load score with the network status, the technical defects of the traditional solution, such as the single network evaluation dimension, multimodal timing misalignment and rigid resource allocation, are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of conference content intelligent synchronization, and in particular to a conference content intelligent synchronization method and system thereof. Background Art

[0002] In traditional conference content synchronization systems, the stability of real-time audio, video, and document collaboration is highly dependent on the network environment. However, due to the inherent complexity of network transmission, such as inter-regional node latency variations, sudden traffic congestion, and wireless signal fluctuations, the system often faces multiple challenges such as latency jitter, packet loss, and dynamic bandwidth fluctuations. These challenges frequently lead to audio and video desynchronization, delayed document operations, and loss of critical information. Existing technologies typically use a single network metric to assess network status, lacking quantitative analysis of the multi-dimensional characteristics of network jitter, making it difficult to accurately identify scenarios of sudden network degradation. Due to differences in transmission protocols and clock skew, audio, video, and document-related operations are prone to cross-modal timing misalignment. For example, if the timestamps of a voice commentary and a PowerPoint presentation page-turning instruction are not precisely aligned, the speaker's speech may advance before the slide transition, or participant annotations may be delayed, severely impacting meeting continuity. Existing synchronization technologies often rely on simple timestamp matching, lacking a unified global clock reference calibration and dynamic interpolation compensation mechanisms, making it difficult to eliminate the cumulative errors caused by cross-device and cross-protocol errors. In addition, traditional resource allocation strategies are static and fixed, and do not take into account the impact of deviations in the timeline alignment of multimodal data. They are also unable to dynamically adjust data transmission priorities based on real-time network conditions, resulting in the loss of critical information in high-load or weak network environments.

[0003] Therefore, it is necessary to provide a conference content intelligent synchronization method and system to solve the above technical problems. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides a method and system for intelligent synchronization of conference content, which achieves the beneficial effect of dynamically allocating resources according to conference content.

[0005] The present invention provides a method for intelligent synchronization of conference content, the synchronization method comprising the following steps:

[0006] S1: Use the server timeline to perform global timeline coarse alignment on the audio, video, and document data streams during the conference, obtain a multimodal time-aligned dataset, and calculate the global time synchronization error at preset time intervals.

[0007] S2: Using the time axis of the video stream after the global time axis coarse alignment calibration as the reference time axis, perform local time axis fine alignment calibration, and calculate the local time synchronization error at a preset time interval;

[0008] S3: Based on the global time synchronization error and the local time synchronization error, the audio stability score, video complexity score, and text importance score of the multimodal time alignment dataset are extracted at preset time intervals;

[0009] S4: Calculate the network jitter coefficient at the same time within the preset time interval, and calculate the comprehensive load score by combining the audio stability score, video complexity score, and text importance score;

[0010] S5: Divide the processing channels into at least three levels and dynamically allocate resources during the conference content synchronization process based on the comprehensive load score.

[0011] Preferably, the local time axis fine alignment calibration performs time axis alignment on the audio stream and document data stream after the global time axis coarse alignment calibration by using a dynamic interpolation compensation algorithm.

[0012] Preferably, step S2 further includes:

[0013] Based on the global time synchronization error, the local time synchronization error and the preset synchronization error threshold, the data streams that do not meet the preset conditions among the audio stream, the video stream and the document data stream are frozen.

[0014] Preferably, the step of obtaining the audio stability score includes the following steps:

[0015] Obtaining a volume level score, a frequency change rate score, and a voice activity detection value score of a voice stream within a preset time interval;

[0016] The volume level score, the frequency change rate score and the voice activity detection value score are weighted and fused according to the preset weights to obtain the audio basic score;

[0017] An audio stream correction coefficient is calculated based on the local time synchronization error, and the audio basic score is multiplied by the audio stream correction coefficient to obtain an audio stability score.

[0018] Preferably, the step of obtaining the video complexity score includes the following steps:

[0019] Obtain the frame rate stability score, resolution level score, and color change rate score of the video stream within a preset time interval;

[0020] The frame rate stability score, resolution level score and color change rate score are weighted and fused according to the preset weights to obtain the basic video score;

[0021] A video stream correction coefficient is calculated based on the global time synchronization error, and the video basic score is multiplied by the video stream correction coefficient to obtain a video complexity score.

[0022] Preferably, the step of obtaining the text importance score includes the following steps:

[0023] Obtaining a text update frequency score and a character length score after update of a document data stream within a preset time interval;

[0024] The text update frequency score and the updated character length score are weighted and fused according to the preset weights to obtain the basic score of the document data;

[0025] A document data flow correction coefficient is calculated based on the local time synchronization error, and the document data basic score is multiplied by the document data flow correction coefficient to obtain a text importance score.

[0026] Preferably, the network jitter coefficient is calculated based on the delay standard deviation, bandwidth utilization, average packet loss rate and a preset network jitter coefficient calculation formula within a preset time interval. The preset network jitter coefficient calculation formula is:

[0027]

[0028] in, is the network jitter coefficient, is the delay standard deviation, is the average packet loss rate, is the bandwidth utilization, 、 、 are weight coefficients respectively.

[0029] Preferably, step S4 includes a fuse degradation strategy, when the network jitter coefficient exceeds a preset network jitter threshold, the video stream is downsampled to 15fps and the audio stream is switched to Opus 16kbps encoding.

[0030] Preferably, step S5 includes dynamically adjusting the maximum concurrent number of processing channels based on the network jitter coefficient and a preset configuration strategy.

[0031] The present invention provides a conference content intelligent synchronization system, which is applied to a conference content intelligent synchronization method. The synchronization system includes:

[0032] The multimodal timing coarse alignment module is used to perform global time axis coarse alignment calibration on the audio stream, video stream and document data stream during the conference using the server time axis, obtain a multimodal time alignment dataset, and calculate the global time synchronization error at preset time intervals;

[0033] The local timing fine alignment module is used to perform local time axis fine alignment calibration based on the time axis of the video stream after the global time axis coarse alignment calibration, and calculate the local time synchronization error at a preset time interval;

[0034] A multimodal quality assessment module is used to extract the audio stability score, video complexity score, and text importance score of the multimodal time-aligned dataset at preset time intervals based on the global time synchronization error and the local time synchronization error;

[0035] A comprehensive load scoring module is used to simultaneously calculate the network jitter coefficient within a preset time interval and calculate a comprehensive load score based on the audio stability score, video complexity score, and text importance score;

[0036] The resource dynamic scheduling module is used to divide at least three levels of processing channels based on the comprehensive load score to dynamically allocate resources during the conference content synchronization process.

[0037] Compared with related technologies, the method and system for intelligent synchronization of conference content provided by the present invention have the following beneficial effects:

[0038] The present invention first realizes the global coarse alignment of multimodal data streams through the server timeline, and then dynamically compensates the local timing deviation of audio and document streams based on the video stream, effectively eliminating problems such as audio and video asynchrony and document operation lag. Then, the audio stability score, video complexity score and text importance score in the multimodal data set in the conference synchronization content are calculated, and the scores are corrected by the calculated global time synchronization error and local time synchronization error. Then, a dynamic network jitter coefficient calculation model is constructed by integrating multi-dimensional network indicators, including delay standard deviation, average packet loss rate and bandwidth utilization, which significantly improves the synchronization of conference content in complex network environments. The accuracy and stability of the network are finally evaluated. The obtained network jitter coefficient is weightedly integrated with the audio stability score, video complexity score and text importance score to obtain a comprehensive load score. Through the three-level dynamic resource allocation mechanism linked to the error-aware multimodal quality score and the network status, the circuit breaker degradation strategy is automatically triggered when the network degrades, giving priority to the transmission of core data streams while taking into account the efficient compression processing of non-critical data, realizing the optimal allocation of resources in high-load or weak network scenarios, taking into account the smoothness, real-time performance and resource utilization of the meeting, and solving the technical defects of the traditional solution such as the single network evaluation dimension, multimodal timing misalignment and rigid resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flow chart of a method for intelligent synchronization of conference content according to the present invention;

[0040] Figure 2 This is a module structure diagram of a conference content intelligent synchronization system of the present invention. DETAILED DESCRIPTION

[0041] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all of the structures. Furthermore, the embodiments of the present invention and the features of the embodiments may be combined with one another unless there is a conflict.

[0042] It should also be noted that, for ease of description, only portions relevant to the present invention are shown in the accompanying drawings, rather than all of the contents. Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the various operations (or steps) as being processed sequentially, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the various operations can be rearranged. The process can be terminated when its operations are completed, but may also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0043] Example 1

[0044] The present invention provides a method for intelligent synchronization of conference content. In the specific implementation process, Figure 1 As shown, it shows a flow chart of a method for intelligent synchronization of conference content, including:

[0045] Step S1: Use the server time axis to perform global time axis coarse alignment calibration on the audio stream, video stream and document data stream during the meeting to obtain a multimodal time alignment dataset, and calculate the global time synchronization error at a preset time interval.

[0046] During the specific implementation process, the specific method of global time axis coarse alignment calibration is that when the server receives the audio stream, video stream and document data stream sent by the terminal, it will discard the original timestamp of each data stream collection terminal, and mark it as the server timestamp of the time when the data arrives at the server, and calculate the sum of the deviations between the original timestamps of the voice stream, document data stream and video stream and the server time axis within the preset time interval to obtain the global time synchronization error. The global time synchronization error reflects the transmission delay difference of multimodal data. Due to the different network delays of different data streams, the time when they arrive at the server is not consistent.

[0047] Step S2: Using the time axis of the video stream after the global time axis coarse alignment calibration as the reference time axis, perform local time axis fine alignment calibration, and calculate the local time synchronization error at a preset time interval.

[0048] Specifically, the local time axis fine alignment calibration uses a dynamic interpolation compensation algorithm to align the time axes of the audio stream and document data stream after the global time axis coarse alignment calibration.

[0049] Specifically, step S2 further includes:

[0050] Based on the global time synchronization error, the local time synchronization error and the preset synchronization error threshold, the data streams that do not meet the preset conditions among the audio stream, the video stream and the document data stream are frozen.

[0051] During the specific implementation process, after the global time axis coarse alignment calibration, it is also necessary to perform fine time axis alignment calibration between the audio stream, video stream and document data stream. The time axis of the video stream after the global time axis coarse alignment calibration is used as the reference time axis, and the time axis of the audio stream and the document data stream are aligned with the reference time axis through a dynamic interpolation compensation algorithm. At the same time, the synchronization error between the audio stream and the document data stream after interpolation compensation and the reference time axis is calculated within a preset time interval, and the local time synchronization error is obtained. In addition, a synchronization error threshold is preset, and the synchronization error threshold includes a global synchronization error threshold and a local synchronization error threshold. For example, when the global time synchronization error exceeds the global synchronization error threshold and lasts for two time intervals or the local time synchronization error exceeds the local synchronization error threshold, it indicates that the time axis synchronization is seriously out of sync, the video stream is switched to a static frame, and only the last frame per second is retained. The real-time synchronization of the document data stream is stopped, and only the final version snapshot is transmitted, and the audio stream is switched to a silent state.

[0052] Step S3: Based on the global time synchronization error and the local time synchronization error, the audio stability score, the video complexity score, and the text importance score of the multimodal time alignment dataset are extracted at preset time intervals.

[0053] During the specific implementation process, the audio stability score, video complexity score and text importance score are calculated based on the characteristics of the video stream, audio stream and document data stream in the multimodal time alignment dataset, and each score is corrected by the global time synchronization error and the local time synchronization error for the calculation of subsequent steps.

[0054] Specifically, the step of obtaining the audio stability score includes the following steps:

[0055] Obtaining a volume level score, a frequency change rate score, and a voice activity detection value score of a voice stream within a preset time interval;

[0056] The volume level score, the frequency change rate score and the voice activity detection value score are weighted and fused according to the preset weights to obtain the audio basic score;

[0057] An audio stream correction coefficient is calculated based on the local time synchronization error, and the audio basic score is multiplied by the audio stream correction coefficient to obtain the audio stability score.

[0058] During the specific implementation process, the audio stream is collected within a preset time interval, the continuous audio stream data is covered with a fixed sliding time window, the RMS energy value of the audio frame within the preset time interval is calculated, and the audio level score is obtained after normalization. The Euclidean distance of the MFCC vectors of adjacent frames within the preset time interval is calculated, and the frequency change rate score is obtained after normalization. The proportion of the number of valid audio frames within the preset time interval is statistically calculated and normalized to obtain a voice activity detection value score. The volume level score, the frequency change rate score, and the voice activity detection value score are weighted and fused according to preset weights to obtain an audio basic score. For example, the preset weighted weight of the volume level is 0.4, the weighted weight of the frequency change rate is 0.3, and the weighted weight of the voice activity detection value is 0.3. Then, the audio stream correction coefficient is calculated based on the local time synchronization error. The calculation formula of the audio stream correction coefficient is:

[0059]

[0060] in, is the audio stream correction coefficient, is the local time synchronization error, 200 is the audio stream correction threshold, when When the delay is greater than 200ms, the audio stability score is forced to decay to half of its original value. When the duration is less than or equal to 200ms, the audio stability score is obtained by multiplying the audio base score by the audio stream correction factor.

[0061] Specifically, the step of obtaining the video complexity score includes the following steps:

[0062] Obtain the frame rate stability score, resolution level score, and color change rate score of the video stream within a preset time interval;

[0063] The frame rate stability score, resolution level score and color change rate score are weighted and fused according to the preset weights to obtain the basic video score;

[0064] The video stream correction coefficient is calculated based on the global time synchronization error, and the video complexity score is obtained by multiplying the video basic score with the video stream correction coefficient.

[0065] During the specific implementation process, the actual transmission frame rate and the nominal frame rate in the window are counted within a preset time interval, and the absolute difference between the actual transmission frame rate and the nominal frame rate is subtracted from 1 to obtain the ratio of the nominal frame rate. The frame rate stability score is obtained based on the preset resolution score mapping table within the preset time interval. The HSV histogram of adjacent frames is extracted within the preset time interval, and the Bhattacharyya distance is calculated and normalized to obtain the color change rate score. The frame rate stability score, resolution level score and color change rate score are weighted and fused according to the preset weighting weight to obtain the video basic score. Then, the video stream correction coefficient is calculated based on the global time synchronization error. The calculation formula of the video stream correction coefficient is:

[0066]

[0067] in, is the video stream correction coefficient, is the global time synchronization error, 500 is the video stream correction threshold, when When it is greater than 500ms, it is forced to set is 0.3, when When the duration is less than or equal to 500ms, the video complexity score is obtained by multiplying the video basic score by the video stream correction coefficient.

[0068] Specifically, the steps for obtaining the text importance score include the following steps:

[0069] Obtaining a text update frequency score and a character length score after update of a document data stream within a preset time interval;

[0070] The text update frequency score and the updated character length score are weighted and fused according to the preset weights to obtain the basic score of the document data;

[0071] The document data flow correction coefficient is calculated based on the local time synchronization error, and the document data basic score is multiplied by the document data flow correction coefficient to obtain the text importance score.

[0072] During the specific implementation process, the number of operations on the document data stream is counted within a preset time interval, and the ratio of the number of operations to the preset maximum number of operations is calculated to obtain a text update frequency score. The total number of added and deleted characters is counted within a preset time interval, and the ratio of the total number of deleted and added characters to the total number of characters in the document is calculated to obtain an updated character length score. The text update frequency score and the updated character length score are weighted and fused according to a preset weighted weight to obtain a document data basic score. Then, the document data stream correction coefficient is calculated based on the local time synchronization error. The calculation formula of the document data stream correction coefficient is:

[0073]

[0074] in, is the document data flow correction coefficient, is the local time synchronization error, 150 is the error tolerance threshold, and 150ms is the human perception critical value of document operation synchronization delay. When the document operation delay exceeds 150ms, 83% of users will notice a lag. 100 is used to control the decay rate of the document data flow correction coefficient between 150ms and 250ms. The document data basic score is multiplied by the document data flow correction coefficient to obtain the text importance score.

[0075] Step S4: Calculate the network jitter coefficient simultaneously within a preset time interval, and calculate the comprehensive load score by combining the audio stability score, the video complexity score, and the text importance score.

[0076] Specifically, the network jitter coefficient is calculated based on the delay standard deviation, bandwidth utilization, average packet loss rate and a preset network jitter coefficient calculation formula within a preset time interval. The preset network jitter coefficient calculation formula is:

[0077]

[0078] in, is the network jitter coefficient, is the delay standard deviation, is the average packet loss rate, is the bandwidth utilization, 、 、 are weight coefficients respectively.

[0079] Specifically, step S4 includes a fuse degradation strategy. When the network jitter coefficient exceeds a preset network jitter threshold, the video stream is downsampled to 15fps and the audio stream is switched to Opus 16kbps encoding.

[0080] During the specific implementation process, the network conditions where the conference content is synchronized must also be considered. The network jitter coefficient is calculated using the preset network jitter coefficient calculation formula by combining the delay standard deviation, average packet loss rate, and bandwidth utilization within the preset time interval. The comprehensive load score is calculated by combining the audio stability score, video complexity score, and text importance score. The calculation formula for the comprehensive load score is:

[0081]

[0082] in, For the comprehensive load score, Rate audio stability, Score the video complexity, Score the importance of the text, is the network jitter coefficient, 、 、 They are the weight coefficients of the audio stability score, video complexity score, and text importance score, respectively. The comprehensive load score is calculated by the calculation formula of the comprehensive load score and used as the basis for resource allocation in conference synchronization in subsequent steps. In addition, step S4 includes a fuse degradation strategy with a preset network jitter threshold, which is calibrated based on historical data and corresponds to high-latency fluctuation scenarios. When the network jitter coefficient of multiple consecutive time intervals is greater than the network jitter threshold, the fuse degradation strategy is triggered, the frame rate of the video stream is downsampled to 15fps, and the encoding of the audio stream is switched to Opus encoding, and the bit rate is reduced to 16kbps.

[0083] Step S5: Divide the processing channels into at least three levels and dynamically allocate resources during the conference content synchronization process based on the comprehensive load score.

[0084] Specifically, step S5 includes dynamically adjusting the maximum concurrent number of processing channels based on the network jitter coefficient and a preset configuration strategy.

[0085] In specific implementations, three processing channels are exemplified: core, basic, and enhanced. The core channel maintains a minimum 10% bandwidth allocation, used to transmit 720p 15fps video streams, 16kHz 24kbps audio streams, document data streams, and document page turning instructions. The basic channel is used to transmit historical logs during conference synchronization, the last frame of each second of video, and an 8kHz 16kbps audio stream. The enhanced channel is used to transmit 1080p 30fps video streams, 48kHz audio streams, and real-time operations of document data streams. The overall load score is greater than 0. 8, the allocated bandwidth is 72% for the core channel, 25% for the enhanced channel, and 3% for the basic channel. When the comprehensive load score is greater than or equal to 0.5 and less than 0.8, the allocated bandwidth is 36% for the core channel, 48% for the enhanced channel, and 16% for the basic channel. When the comprehensive load score is less than 0.5, the allocated bandwidth is: 10% for the core channel, 42% for the enhanced channel, and 48% for the basic channel. In addition, based on the network jitter coefficient and the preset dynamic threshold trigger mechanism, for example, when the network jitter coefficient is greater than 15, the number of concurrency is reduced to achieve network quality-driven elastic scaling, prioritize core channel resources, and optimize the global balance between bandwidth and computing load.

[0086] The working principle of the intelligent synchronization method of conference content provided by the present invention is as follows:

[0087] The present invention realizes intelligent synchronization of multimodal conference content through layered time alignment and dynamic resource allocation: first, the audio, video and document streams are globally coarsely aligned with the server time axis to calculate the global time synchronization error; then, based on the video stream, the audio and document streams are locally finely aligned through a dynamic interpolation algorithm to generate a local time synchronization error; based on the two-level error data, an audio stability score, a video complexity score and a text importance score are constructed respectively, and the network jitter coefficient is simultaneously introduced. Combined with the multimodal score, a dynamic comprehensive load score is generated, and three-level resource regulation is implemented accordingly. Under high load, fuse degradation is implemented, video is downsampled to 15fps, audio is switched to 16kbps encoding, and elastic expansion is implemented under low load. While ensuring the real-time performance of core data, the purpose of reducing the cross-modal synchronization error rate is achieved.

[0088] Example 2

[0089] The present invention provides a conference content intelligent synchronization system. In the specific implementation process, Figure 2 As shown, it shows a module structure diagram of a conference content intelligent synchronization system, including:

[0090] The multimodal timing coarse alignment module 100 is used to perform global time axis coarse alignment calibration on the audio stream, video stream and document data stream during the conference using the server time axis, obtain a multimodal time alignment dataset, and calculate the global time synchronization error at preset time intervals;

[0091] The local timing fine alignment module 200 is used to perform local time axis fine alignment calibration based on the time axis of the video stream after the global time axis coarse alignment calibration, and calculate the local time synchronization error at a preset time interval;

[0092] A multimodal quality assessment module 300 is configured to extract an audio stability score, a video complexity score, and a text importance score of the multimodal time-aligned dataset at preset time intervals based on a global time synchronization error and a local time synchronization error;

[0093] A comprehensive load scoring module 400 is configured to simultaneously calculate the network jitter coefficient within a preset time interval and calculate a comprehensive load score by combining the audio stability score, the video complexity score, and the text importance score;

[0094] The resource dynamic scheduling module 500 is used to divide at least three levels of processing channels based on the comprehensive load score to perform dynamic resource allocation during the conference content synchronization process.

[0095] The working principle of the intelligent synchronization system for conference content provided by the present invention is as follows:

[0096] The present invention achieves efficient conference synchronization through multi-level collaborative timing alignment and dynamic resource scheduling. The multimodal timing coarse alignment module 100 performs global coarse-grained alignment of audio, video and document streams based on server time, generates a unified timestamp data set and calculates the global synchronization error; the local timing fine alignment module 200 uses the video stream as the reference axis, compensates for the local timing deviation of the audio and document streams through a dynamic interpolation algorithm, and synchronously detects local errors; the multimodal quality assessment module 300 combines the global time synchronization error and the local time synchronization error to construct an error correction scoring model for audio stability frequency division, video complexity frequency division and text importance scoring; the comprehensive load scoring module 400 integrates the network jitter coefficient and the multimodal score to generate a dynamic comprehensive load score; the resource dynamic scheduling module 500 dynamically adjusts the bandwidth share and concurrency number of the divided three-level processing channels according to the comprehensive load score, so as to achieve efficient synchronization of multimodal data and optimal resource allocation in a complex network environment.

[0097] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0098] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0099] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

Claims

1. A method for intelligent synchronization of conference content, characterized in that: The synchronization method comprises the following steps: S1: Use the server timeline to perform global timeline coarse alignment on the audio, video, and document data streams during the conference, obtain a multimodal time-aligned dataset, and calculate the global time synchronization error at preset time intervals. S2: Using the time axis of the video stream after the global time axis coarse alignment calibration as the reference time axis, perform local time axis fine alignment calibration, and calculate the local time synchronization error at a preset time interval; S3: Based on the global time synchronization error and the local time synchronization error, the audio stability score, video complexity score, and text importance score of the multimodal time alignment dataset are extracted at preset time intervals; S4: Calculate the network jitter coefficient at the same time within the preset time interval, and calculate the comprehensive load score by combining the audio stability score, video complexity score, and text importance score; S5: Divide the processing channels into at least three levels and dynamically allocate resources during the conference content synchronization process based on the comprehensive load score.

2. A method for intelligent synchronization of conference content according to claim 1, characterized in that: The local time axis fine alignment calibration performs time axis alignment on the audio stream and document data stream after the global time axis coarse alignment calibration through a dynamic interpolation compensation algorithm.

3. A method for intelligent synchronization of conference content according to claim 2, characterized in that: Step S2 further includes: Based on the global time synchronization error, the local time synchronization error and the preset synchronization error threshold, the data streams that do not meet the preset conditions among the audio stream, the video stream and the document data stream are frozen.

4. A method for intelligent synchronization of conference content according to claim 3, characterized in that: The step of obtaining the audio stability score includes the following steps: Obtaining a volume level score, a frequency change rate score, and a voice activity detection value score of a voice stream within a preset time interval; The volume level score, the frequency change rate score and the voice activity detection value score are weighted and fused according to the preset weights to obtain the audio basic score; An audio stream correction coefficient is calculated based on the local time synchronization error, and the audio basic score is multiplied by the audio stream correction coefficient to obtain an audio stability score.

5. A method for intelligent synchronization of conference content according to claim 4, characterized in that: The step of obtaining the video complexity score includes the following steps: Obtain the frame rate stability score, resolution level score, and color change rate score of the video stream within a preset time interval; The frame rate stability score, resolution level score and color change rate score are weighted and fused according to the preset weights to obtain the basic video score; A video stream correction coefficient is calculated based on the global time synchronization error, and the video basic score is multiplied by the video stream correction coefficient to obtain a video complexity score.

6. A method for intelligent synchronization of conference content according to claim 5, characterized in that: The step of obtaining the text importance score comprises the following steps: Obtaining a text update frequency score and a character length score after update of a document data stream within a preset time interval; The text update frequency score and the updated character length score are weighted and fused according to the preset weights to obtain the basic score of the document data; A document data flow correction coefficient is calculated based on the local time synchronization error, and the document data basic score is multiplied by the document data flow correction coefficient to obtain a text importance score.

7. A method for intelligent synchronization of conference content according to claim 6, characterized in that: The network jitter coefficient is calculated based on the delay standard deviation, bandwidth utilization, average packet loss rate and a preset network jitter coefficient calculation formula within a preset time interval. The preset network jitter coefficient calculation formula is: Among them, C net is the network jitter coefficient, σ d is the delay standard deviation, L r is the average packet loss rate, B u is the bandwidth utilization, ω1, ω2, and ω3 are weight coefficients respectively.

8. A method for intelligent synchronization of conference content according to claim 7, characterized in that: Step S4 includes a fuse degradation strategy. When the network jitter coefficient exceeds a preset network jitter threshold, the video stream is downsampled to 15fps, the encoding of the audio stream is switched to Opus encoding, and the bit rate is reduced to 16kbps.

9. A method for intelligent synchronization of conference content according to claim 8, characterized in that: Step S5 includes dynamically adjusting the maximum concurrent number of processing channels based on the network jitter coefficient and a preset configuration strategy.

10. A conference content intelligent synchronization system, characterized in that: The method for intelligent synchronization of conference content according to any one of claims 1 to 9 is applied, wherein the synchronization system comprises: The multimodal timing coarse alignment module is used to perform global time axis coarse alignment calibration on the audio stream, video stream and document data stream during the conference using the server time axis, obtain a multimodal time alignment dataset, and calculate the global time synchronization error at preset time intervals; The local timing fine alignment module is used to perform local time axis fine alignment calibration based on the time axis of the video stream after the global time axis coarse alignment calibration, and calculate the local time synchronization error at a preset time interval; A multimodal quality assessment module is used to extract the audio stability score, video complexity score, and text importance score of the multimodal time-aligned dataset at preset time intervals based on the global time synchronization error and the local time synchronization error; The comprehensive load scoring module is used to simultaneously calculate the network jitter coefficient within a preset time interval, and calculate the comprehensive load score in combination with the audio stability score, video complexity score and text importance score; the resource dynamic scheduling module is used to divide at least three levels of processing channels based on the comprehensive load score to dynamically allocate resources during the conference content synchronization process.

Citation Information

Patent Citations

  • Conference room audio and video synchronous visualization control method and system

    CN119728912A

  • Data processing method and system for near-field wireless transmission

    CN119835627A