Real-time audio and video synchronization monitoring system for video call

CN121193983BActive Publication Date: 2026-09-15DONGGUAN CLOUD FOX INTELLIGENCE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510835632.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2026-09-15
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

[0004]传统系统主要依赖静态的同步机制,如时间戳标记和缓冲区管理,在实际应用中,尤其是在网络波动大、带宽不稳定的环境下,静态同步机制难以应对突发的延迟变化,导致音视频数据的不同步现象频发,影响用户的交互体验,例如,在远程医疗和直播系统中,由于音视频不同步,会导致误解或信息传递错误,从而影响决策和用户满意度,传统系统在处理动态变化方面存在局限,尤其是在没有实时监控和调整同步差异的功能时,其同步控制的精确性和可靠性受到限制

Benefits of technology

[0042] In this invention, by extracting and analyzing the timestamp difference of audio and video frames in real time and establishing a transmission delay interval comparison sequence, the time difference in audio and video transmission can be effectively identified, enhancing the monitoring accuracy of audio and video asynchrony. By adjusting the timestamp offset in real time, the synchronization accuracy of audio and video during playback is ensured. In dynamic network environments, the offset rate is monitored and adjusted in real time, especially under highly volatile network conditions, ensuring the continuity and smoothness of the interactive experience. By analyzing the concentrated distribution segments of the offset difference and marking the synchronization fluctuation segments, the synchronization status of audio and video frames can be more precisely controlled, avoiding a large-scale decline in user experience due to the accumulation of synchronization errors. This improves the system's response speed to audio and video synchronization control, enhances the processing capability of audio and video data under unstable network conditions, and ensures the integrity of information transmission and the smoothness of communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121193983B_ABST
    Figure CN121193983B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio-video synchronization, in particular to a real-time audio-video synchronization monitoring system for video calls, which comprises a frame time extraction module, a main time sequence locking module, a delay structure analysis module, a synchronization node generation module and a synchronization control instruction module.In the application, the time stamp difference calculation of real-time extraction and analysis of audio-video frames is used to establish a transmission delay interval comparison sequence, which can effectively identify the time difference in audio-video transmission, enhance the monitoring accuracy of the audio-video asynchronization phenomenon, adjust the time stamp offset in real time, ensure the synchronization accuracy of the audio-video in the playing process, monitor and adjust the offset rate in real time under a dynamic network environment, ensure the continuity and smoothness of the interactive experience, analyze the offset difference value centralized distribution section, mark the synchronization fluctuation section, more accurately control the synchronization state of the audio-video frames, and avoid the large-scale user experience decline caused by the synchronization error accumulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio and video synchronization technology, and in particular to a real-time audio and video synchronization monitoring system for video calls. Background Technology

[0002] The field of audio-video synchronization technology involves controlling the timing consistency of audio and video signals across multiple processing stages, including acquisition, encoding, transmission, decoding, and presentation. This technology primarily addresses audio-video asynchrony caused by variations in processing paths, bandwidth, network jitter, or transmission delays, ensuring that audio and video are synchronized and coordinated during playback at the receiving end. Synchronization technologies typically include timestamp marking, synchronization buffer management, clock calibration mechanisms, and latency compensation algorithms, and are widely used in real-time or near-real-time interactive audio-video applications such as video conferencing, live streaming systems, streaming media playback, distance education, and telemedicine. The core objective of this field is to achieve accurate matching and synchronized presentation of audio frames and corresponding video frames, thereby improving the consistency of the user's audiovisual experience and the efficiency of information comprehension in interactive scenarios.

[0003] The real-time audio and video synchronization monitoring system for video calls is a system used to synchronize and monitor the quality of audio and video signals during video communication. While enabling real-time communication between two or more parties, the system provides monitoring capabilities for audio and video synchronization, promptly identifying and adjusting audio-video desynchronization issues to ensure the integrity of information transmission and smooth communication during the call. Its applications include improving the quality of video call experiences and supporting the automatic identification and control of synchronization anomalies in remote communication. It is suitable for applications requiring stable synchronized audio and video calls, such as remote conferencing, online customer service, and remote collaboration.

[0004] Traditional systems primarily rely on static synchronization mechanisms, such as timestamps and buffer management. In practical applications, especially in environments with large network fluctuations and unstable bandwidth, these static mechanisms struggle to cope with sudden changes in latency, leading to frequent audio and video data asynchrony. This negatively impacts the user experience. For example, in telemedicine and live streaming systems, audio-video asynchrony can cause misunderstandings or information transmission errors, affecting decision-making and user satisfaction. Traditional systems also have limitations in handling dynamic changes, particularly lacking real-time monitoring and adjustment capabilities for synchronization discrepancies. This restricts the accuracy and reliability of their synchronization control. Furthermore, it increases the system's adaptability to different network conditions, potentially affecting the accurate transmission and processing efficiency of information at critical moments. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a real-time audio and video synchronization monitoring system for video calls.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a real-time audio and video synchronization monitoring system for video calls, the system comprising:

[0007] The frame time extraction module obtains the parameter information of each audio frame and image frame during the video call, extracts the sending timestamp and receiving timestamp, calculates the difference, constructs a comparison sequence of the transmission delay interval of audio frames and image frames, and obtains synchronization difference sequence data.

[0008] The main timing locking module calls the synchronization difference sequence data, calculates the frame time interval offset and offset rate between two adjacent frames, compares them with the synchronization stability reference threshold, sets the reference frame time reference group, and generates the image frame main time axis reference adjustment parameters.

[0009] The delay structure analysis module adjusts the parameters based on the main time axis reference of the image frame, compares the time axis reference difference between the audio frame and the image frame, filters out the concentrated distribution area of ​​the offset difference, and marks the corresponding time point interval as the synchronization fluctuation segment to obtain the multi-frame audio and video delay fluctuation interval labeling information.

[0010] The synchronization node generation module calls the multi-frame audio and video delay fluctuation interval annotation information, judges the receiving time difference of adjacent image frames and audio frames after sorting, extracts the frame pair with the smallest time difference, records the timestamp and marks it as a candidate synchronization control node, and generates a candidate synchronization control node list.

[0011] As a further aspect of the present invention, the synchronization difference sequence data includes a list of audio frame-image frame time difference comparisons, channel sequence number mapping information, and an image frame transmission delay distribution table. The image frame main time axis reference adjustment parameters specifically include a reference frame time offset index, a reference offset replacement parameter, and a frame time drift identifier label value. The multi-frame audio and video delay fluctuation interval labeling information includes a synchronization critical difference distribution record, an inter-frame fluctuation amplitude index table within a time period, and a fluctuation segment time node index list. The candidate synchronization control node list specifically includes a frame pair minimum reception time difference identifier, a node timestamp comparison table, and a synchronization control candidate index set.

[0012] As a further aspect of the present invention, the frame time extraction module includes:

[0013] The frame parameter acquisition submodule acquires the parameter information of each audio frame and image frame during the video call, including the encoding type identifier, sending timestamp, receiving timestamp and receiving channel sequence number. It performs time difference processing on the sending timestamp and receiving timestamp of each frame, filters complete frame information groups and removes incomplete frame records, establishes a set of valid frame indexes based on the data integrity judgment flag, and generates a set of valid frame index numbers.

[0014] The channel matching and recognition submodule extracts the receiving channel sequence number corresponding to each voice frame and image frame based on the set of valid frame index numbers, performs frame group pairing according to the channel sequence number, records the encoding type identifier for each group of matching frames, and marks the frame groups with the same type identifier as data blocks to be compared, thereby generating a set of channel synchronization matching identifiers.

[0015] The delay difference generation submodule calls the channel synchronization matching identifier set, extracts the transmission timestamp and reception timestamp of the voice frame and image frame in each frame pair, normalizes the timestamp difference and establishes the time series corresponding to the frame pair, calculates the voice and image transmission delay difference through the time series position index matching relationship and generates interval records to obtain synchronization difference sequence data.

[0016] As a further aspect of the present invention, the main timing locking module includes:

[0017] The transmission interval calculation submodule calls the synchronization difference sequence data, obtains the reception time difference between any two adjacent image frames based on the reception timestamp of the image frame, performs the calculation of the difference ratio between consecutive frames for the time difference between consecutive frames, compares the ratio with the image frame interval offset reference threshold, marks the frame group with the ratio greater than the image frame interval offset reference threshold, and generates the image frame interval offset rate sequence.

[0018] The frame drift recognition submodule divides the time period according to the continuity of the frame group index based on the image frame interval offset rate sequence, extracts the average offset rate of each continuous segment and judges it with the image frame synchronization stability threshold, filters the time period with the average offset rate exceeding the image frame synchronization stability threshold, marks the time period as drift segment, and generates the image frame drift time period interval value.

[0019] The baseline offset construction submodule calls the image frame drift time interval value, calculates the set of differences between the receiving time and the sending time in each segment according to the receiving time and sending time of the corresponding frame group, locates the minimum value frame index position in the difference set, sets the receiving time and sending time value of the corresponding frame as the reference frame time reference group, calculates the image frame time difference offset, establishes the image frame time axis baseline offset structure, and obtains the image frame main time axis reference adjustment parameters.

[0020] As a further aspect of the present invention, the formula for calculating the image frame time difference offset is:

[0021] ;

[0022] in, Representing the In the time period of the th time period The time difference offset between the reception and transmission of a group of image frames. Indicates the first Within the time period, the first In the frame The reception time of each image frame Indicates the transmission time of the corresponding frame. Indicates the first The first time period The average of the time differences between receiving and sending all frames in the group. Indicates the sequence number of the frame in the group. This indicates the number of image frames within the group. Indicates the drift time period number, Indicates the frame group number within the segment.

[0023] As a further aspect of the present invention, the delay structure analysis module includes:

[0024] The reference difference extraction submodule obtains the reception time of the audio frame and the reference offset of the image frame time axis within the time period of each image frame according to the reference adjustment parameters of the main time axis of the image frame. It subtracts the corresponding reference offset of the image frame from the reception time of each audio frame to construct the time axis difference sequence of the audio frame and the image frame, and generates the reference difference sequence of audio and video frames.

[0025] The offset determination and filtering submodule calls the audio and video frame reference difference value sequence, obtains the concentration of difference values ​​in each time period and calculates the maximum local offset difference value, determines whether the maximum value exceeds the synchronization critical difference threshold, marks the interval where the offset difference value exceeds the threshold, and generates the segment value of the offset difference concentration interval.

[0026] The fluctuation segment labeling and classification submodule extracts the start and end time points and audio frame sequence index of each interval segment based on the segment value of the interval in the offset difference set, and labels the segments that meet the continuity and threshold conditions as synchronous fluctuation segments, establishes a corresponding time period index mapping table, and obtains multi-frame audio and video delay fluctuation interval labeling information.

[0027] As a further aspect of the present invention, the synchronization node generation module includes:

[0028] The frame timing sorting submodule calls the multi-frame audio and video delay fluctuation interval annotation information, extracts the reception timestamp and frame number of each image frame and audio frame in each segment, sorts the reception timestamps of the image frames and audio frames in each segment in ascending order according to the frame number, establishes a frame pair sequence index mapping table, and generates a sorted frame timing index set.

[0029] The frame pair difference extraction submodule obtains the received timestamps between adjacent image frames and audio frames based on the sorted frame time sequence index set, performs time difference calculation for each pair of image frames and audio frames, filters the frame pairs with the smallest time difference, records the received timestamps of the corresponding frame pairs and marks them as candidate node frames, and generates the minimum difference frame pair timestamp group.

[0030] The synchronization node aggregation submodule calls the minimum difference frame pair timestamp group, extracts the timestamp and frame pair index number of the candidate node frame corresponding to each segment, and aggregates the information of each candidate node frame in time axis order to construct a candidate synchronization control node sequence and obtain a candidate synchronization control node list.

[0031] As a further aspect of the present invention, the system further includes:

[0032] The synchronization control instruction module obtains the candidate synchronization control node list, and calls the playback timestamp of the image frame and audio frame to be output in the current playback queue according to the timestamp of each pair of image frames and audio frames. It compares the offset value of the playback timestamp with the corresponding timestamp in the node list. If the offset value exceeds the acceptable synchronization error range, it adjusts the playback timestamp by adding or subtracting the offset value to generate a frame-level audio and video synchronization control instruction set.

[0033] The frame-level audio and video synchronization control instruction set includes image frame playback offset adjustment parameters, voice frame playback time correction parameters, and synchronization tolerance range correction identifiers.

[0034] As a further aspect of the present invention, the synchronization control command module includes:

[0035] The playback offset extraction submodule calls the candidate synchronization control node list to obtain the playback timestamps of image frames and audio frames in the current playback queue. It calculates the offset value between the playback timestamp of each pair of frames and the timestamp of the corresponding candidate node, calculates the audio and video frame synchronization offset value, establishes an index table of playback offset values ​​for image frames and audio frames, and generates a group of audio and video frame playback offset values.

[0036] The error interval judgment submodule extracts each pair of offset values ​​based on the audio and video frame playback offset value group and compares them with the set synchronization error tolerance range. It marks the frame pairs whose offset values ​​exceed the synchronization error tolerance range, records the corresponding frame pair index and offset direction, and generates a frame synchronization error identifier set.

[0037] The control instruction generation submodule calls the frame synchronization error identifier set, adds or subtracts offset values ​​according to the frame type based on the playback time marker and offset direction information of the frame pair, and establishes an adjusted playback time marker sequence. The adjustment results are summarized to construct a frame-level synchronization control record table, and a frame-level audio and video synchronization control instruction set is obtained.

[0038] As a further aspect of the present invention, the formula for calculating the audio and video frame synchronization offset value is as follows:

[0039] ;

[0040] in, Representing the The first in the section The playback offset values ​​between image frames and audio frames. Representing the Section 1 The playback time marker for the image frame in the current playback queue, This represents the timestamp of the corresponding image frame in the list of candidate synchronization control nodes. Representing the Section 1 The playback time of each audio frame in the current playback queue is marked. This represents the timestamp of the corresponding voice frame in the list of candidate synchronization control nodes. For synchronization section numbering, The frame pair number within the segment. Indicates the image frame type marker, This indicates the voice frame type marker.

[0041] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0042] In this invention, by extracting and analyzing the timestamp difference of audio and video frames in real time and establishing a transmission delay interval comparison sequence, the time difference in audio and video transmission can be effectively identified, enhancing the monitoring accuracy of audio and video asynchrony. By adjusting the timestamp offset in real time, the synchronization accuracy of audio and video during playback is ensured. In dynamic network environments, the offset rate is monitored and adjusted in real time, especially under highly volatile network conditions, ensuring the continuity and smoothness of the interactive experience. By analyzing the concentrated distribution segments of the offset difference and marking the synchronization fluctuation segments, the synchronization status of audio and video frames can be more precisely controlled, avoiding a large-scale decline in user experience due to the accumulation of synchronization errors. This improves the system's response speed to audio and video synchronization control, enhances the processing capability of audio and video data under unstable network conditions, and ensures the integrity of information transmission and the smoothness of communication. Attached Figure Description

[0043] Figure 1 This is a system flowchart of the present invention;

[0044] Figure 2 This is a schematic diagram of the system framework of the present invention;

[0045] Figure 3This is a flowchart of the frame time extraction module of the present invention;

[0046] Figure 4 This is a flowchart of the main timing locking module of the present invention;

[0047] Figure 5 This is a flowchart of the delayed structure analysis module of the present invention;

[0048] Figure 6 This is a flowchart of the synchronization node generation module of the present invention;

[0049] Figure 7 This is a flowchart of the synchronization control instruction module of the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0051] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0052] Please see Figure 1 A real-time audio and video synchronization monitoring system for video calls, comprising:

[0053] The frame time extraction module obtains the parameter information of each audio frame and image frame during the video call, including the encoding type identifier, sending timestamp, receiving timestamp and receiving channel sequence number. It performs synchronization index matching based on the receiving channel sequence number of the image frame and audio frame, calculates the difference between the sending timestamp and receiving timestamp extracted from the matched frame group, constructs the transmission delay interval comparison sequence of audio frame and image frame, and obtains the synchronization difference sequence data.

[0054] The main timing locking module calls the synchronization difference sequence data, calculates the frame time interval offset and offset rate between two adjacent frames based on the transmission interval difference of consecutive image frames, and compares the offset rate with the synchronization stability reference threshold. If the continuous offset rate exceeds the threshold, the time period of the corresponding frame is marked as the time axis drift segment, and the reception time and transmission time value of the corresponding frame are set as the reference frame time reference group to generate the main time axis reference adjustment parameters of the image frame.

[0055] The delay structure analysis module adjusts parameters based on the main time axis reference of the image frame, extracts the offset between the reception time of the audio frame and the reference time within the same time period, compares the time axis reference difference between the audio frame and the image frame, determines whether the offset difference crosses the synchronization critical difference threshold, filters out the concentrated distribution area of ​​the offset difference, and marks the corresponding time point interval as the synchronization fluctuation segment, thus obtaining the multi-frame audio and video delay fluctuation interval labeling information.

[0056] The synchronization node generation module calls the multi-frame audio and video delay fluctuation interval labeling information, sorts the received timestamps of image frames and audio frames in the segment in ascending order of frame number, judges the received time difference of adjacent image frames and audio frames after sorting, extracts the frame pair with the smallest time difference, records the timestamp and marks it as a candidate synchronization control node, and aggregates each segment's candidate frame pair in turn to generate a candidate synchronization control node list.

[0057] The synchronization control instruction module obtains a list of candidate synchronization control nodes. Based on the timestamps of each pair of image frames and audio frames, it calls the playback timestamps of the image frames and audio frames that are about to be output in the current playback queue. It compares the offset value of the playback timestamp with the corresponding timestamp in the node list. If the offset value exceeds the acceptable synchronization error range, it adjusts the playback timestamp by adding or subtracting the offset value, and generates a frame-level audio and video synchronization control instruction set.

[0058] The synchronization difference sequence data includes a list of time differences between audio and image frames, channel sequence number mapping information, and an image frame transmission delay distribution table. The image frame main time axis reference adjustment parameters are specifically the reference frame time offset index, reference offset replacement parameters, and frame time drift label values. The multi-frame audio and video delay fluctuation interval labeling information includes a synchronization critical difference distribution record, an inter-frame fluctuation amplitude index table within a time period, and a fluctuation segment time node index list. The candidate synchronization control node list is specifically the frame pair minimum reception time difference identifier, a node timestamp comparison table, and a synchronization control candidate index set. The frame-level audio and video synchronization control instruction set includes image frame playback offset adjustment parameters, audio frame playback time correction parameters, and synchronization tolerance range correction identifiers.

[0059] Please see Figure 2 and Figure 3The frame time extraction module includes a frame parameter acquisition submodule, a channel matching and recognition submodule, and a delay difference generation submodule;

[0060] The frame parameter acquisition submodule acquires the parameter information of each audio frame and image frame during the video call, including the encoding type identifier, sending timestamp, receiving timestamp and receiving channel sequence number. It performs time difference processing on the sending timestamp and receiving timestamp of each frame, filters complete frame information groups and removes incomplete frame records, establishes a set of valid frame indexes based on the data integrity judgment flag, and generates a set of valid frame index numbers.

[0061] The frame parameter acquisition submodule obtains parameter information for each audio and image frame during a video call. Based on the system's independent recognition mechanism for audio and image streams, it connects to the audio acquisition module and image acquisition module respectively. It acquires the metadata of each frame in real time through the frame buffer. This metadata includes the encoding type identifier, transmission timestamp, reception timestamp, and reception channel sequence number. For the encoding type identifier, the corresponding compression method and format number can be directly parsed from a hard-coded configuration table. For example, if the audio frame's encoding type is "G.729", its type identifier encoding value is "0011"; if the image frame's encoding type is "H.264", its type identifier encoding value is "1010". This encoding value is written to the frame record data block. The transmission and reception timestamps are synchronously written by the system clocks of the sending and receiving ends during frame encapsulation and decapsulation, respectively. For example, a certain audio frame's transmission time is set to 10.002 seconds, and its reception time to 10.00 seconds. If the time difference is 0.004 seconds, it is stored in the buffer dataset. Then, the decoder assigns the receiving channel sequence number according to the channel identifier. For example, channels A, B, and C correspond to numbers 1, 2, and 3 respectively, so the image frame received by channel B is recorded as 2. After this recording, the system proceeds to the time difference processing stage. The system calculates the difference between the sending and receiving timestamps of each frame, obtains the time difference, and stores it in a queue. Next, the system filters complete frame information groups based on whether they simultaneously possess the encoding type identifier, sending timestamp, receiving timestamp, and receiving channel sequence number. If any field is missing, the frame is marked as incomplete and removed from the record set. For example, if an image frame lacks the receiving timestamp field, the record is directly removed. The remaining frame information constitutes a complete frame set. The system assigns a valid frame index number to each complete frame record, starting from 0 and incrementing. Finally, a valid frame index set is constructed. This set is the set of valid frame index numbers.

[0062] The channel matching and recognition submodule extracts the receiving channel sequence number corresponding to each voice frame and image frame based on the set of valid frame index numbers, performs frame group pairing according to the channel sequence number, records the encoding type identifier of each matching frame group, and marks the frame group with the same type identifier as the data block to be compared, thereby generating a set of channel synchronization matching identifiers.

[0063] The channel matching and recognition submodule, based on the set of valid frame index numbers, first locates and accesses each audio frame and image frame from the filtered complete frame records by index. It then extracts the recorded receiving channel sequence number for each frame. Frame 0 is designated as an audio frame with receiving channel number 1, and frame 1 as an image frame with receiving channel number 1. If frame 0 and frame 1 are considered to have a channel match, their frame group pairing relationship is recorded and stored in the frame pairing buffer. The frame group pairing rule is that the receiving channel numbers are the same and the time difference is less than or equal to the set synchronization window. The synchronization window As a key judgment threshold, it is set with reference to the maximum allowable frame synchronization deviation of a video call system. Under the WebRTC architecture, the maximum allowable synchronization deviation of audio and video frames is usually no more than 10 milliseconds. A reasonable setting is 0.01 seconds. Frames outside this range will not match. For example, if frame 2 is an audio frame with a transmission time of 20.000 seconds, and image frame 3 is a frame with the same receiving channel number with a receiving time of 20.008 seconds, then their time difference is 0.008 seconds, which satisfies the condition. The requirement is to form a pair of matching frames. For each pair of matching frames, the system further records its encoding type identifier. For example, frame 2 is identified as "0011" and frame 3 as "1010". If the identifiers match (both are "1010"), then the frame group is marked as a data block to be compared. The marking method is to record the corresponding index and its status field "match successful" in the matching identifier set; otherwise, it is marked as "match failed". The set is constructed using frame indices, such as... Furthermore, a set of channel synchronization matching identifiers is generated for subsequent delay difference calculation.

[0064] The delay difference generation submodule calls the channel synchronization matching identifier set, extracts the transmission timestamp and reception timestamp of the voice frame and image frame in each frame pair, normalizes the timestamp difference and establishes the time series corresponding to the frame pair, calculates the voice and image transmission delay difference through the time series position index matching relationship and generates interval records to obtain synchronization difference sequence data.

[0065] The delay difference generation submodule calls the channel synchronization matching identifier set. First, it iterates through each pair of frame groups in the set, extracting the transmission and reception timestamps of the audio and image frames in each pair, and denoting them as follows: , , , Calculate their time difference, for example, for frame pair (4,5), where Second, Second, Second, If the time is seconds, then the audio frame latency is... seconds, image frame delay is The time difference is the difference in delay between the two frames, and the result is 0 seconds. Normalization is performed on the time difference of all frame groups using the following normalization formula:

[0066] ;

[0067] in This is the delay difference for the current frame group. , These represent the minimum and maximum values ​​in the set of delay differences, respectively. The set of frame group delay differences is defined as follows: The normalized value of the current frame group delay difference of 0.004 is then... The normalized result is stored in the frame pair time series. The system constructs a time series mapping table according to the frame pair index. For example, frame pair index 0 corresponds to a time difference of 0.5, index 1 corresponds to 0.25, etc. By calculating the difference between adjacent frame pair indices in the time series, the delay change value between consecutive frame pairs is obtained and an interval record is generated. For example, frame pair The latency value is 0.5, and the frame pair is... If the value is 0.25, then the interval difference is 0.25. This interval value is recorded in the synchronization difference sequence to form a complete synchronization difference sequence data.

[0068] Please see Figure 2 and Figure 4 The main timing locking module includes a transmission interval calculation submodule, a frame drift identification submodule, and a baseline offset construction submodule;

[0069] The transmission interval calculation submodule calls the synchronization difference sequence data, obtains the reception time difference between any two adjacent image frames based on the reception timestamp of the image frame, performs the calculation of the difference ratio between consecutive frames for the time difference between consecutive frames, compares the ratio with the image frame interval offset reference threshold, marks the frame group with the ratio greater than the image frame interval offset reference threshold, and generates the image frame interval offset rate sequence.

[0070] The transmission interval calculation submodule calls the synchronization difference sequence data. Based on the received timestamps of the image frames, it first extracts the received time values ​​of all image frames from the frame sequence in ascending order of time, and constructs a received time array. For any two adjacent image frames and Calculate its receiving time difference For example, when the received timestamp of the image frame is Second, If the time difference is 0.033 seconds, then the time difference ratio between the preceding and following frames is calculated. Difference from the previous value To perform ratio operations, the ratio is defined as follows: ,like seconds, then When network jitter exists between consecutive frames, the difference ratio will shift. For example, if the next difference is 0.066 seconds, the corresponding ratio will be 2.000. Image frame interval offset reference threshold Compare them. The setting is based on the stable variation range of the normal frame interval. Referring to the current system's image frame rate of 30 frames / second, the corresponding average frame interval is approximately 0.033 seconds. The system's normal jitter ratio is controlled within ±20%, therefore, it can be set... That is, if the frame interval is increased by 20%, an offset is considered to exist, and vice versa. When the ratio is 1.30, it is marked as an offset frame group because it is greater than 1.20. The system stores the index of the frame group in the image frame interval offset rate sequence, and finally constructs a complete image frame interval offset rate sequence.

[0071] The frame drift recognition submodule is based on the image frame interval offset rate sequence. It divides the time period according to the continuity of the frame group index, extracts the average offset rate of each continuous segment and judges it with the image frame synchronization stability threshold. It filters the time period with the average offset rate exceeding the image frame synchronization stability threshold and marks the time period as the drift segment, generating the image frame drift time period interval value.

[0072] The frame drift recognition submodule, based on the image frame interval offset rate sequence, divides time periods according to the continuity of frame group indices, performs continuous frame group index difference calculations, and determines whether adjacent frame group indices are continuous. If a frame group index exists... This forms a continuous segment, and the offset value sequence of each frame within this segment is extracted. And calculate the average offset rate of that segment. The calculation method is as follows: ,in This refers to the number of frame groups within that segment. For example, if... , , ,but averaging the offset Image frame synchronization stability threshold Make a judgment. The settings are based on the minimum allowable offset strength for image frame reception stability. Statistical analysis of image frame frequency fluctuations shows that fluctuations exceeding 25% affect frame synchronization; therefore, proper settings are necessary. ,like If 1.30 > 1.25, it is determined to be a drift time period. The start and end time boundaries of the frame index of this segment are recorded as the start and end interval values ​​according to the received timestamp of the image frame, and recorded in the image frame drift time period interval value.

[0073] The baseline offset construction submodule calls the image frame drift time interval value, calculates the set of differences between the receiving time and the sending time in each segment based on the receiving time and sending time of the corresponding frame group, locates the minimum value frame index position in the difference set, sets the receiving time and sending time value of the corresponding frame as the reference frame time reference group, calculates the image frame time difference offset, establishes the image frame time axis baseline offset structure, and obtains the image frame main time axis reference adjustment parameters;

[0074] The formula for calculating the temporal difference offset of image frames is:

[0075] ;

[0076] in, Representing the In the time period of the th time period The time difference offset between the reception and transmission of a group of image frames. Indicates the first Within the time period, the first In the frame The reception time of each image frame Indicates the transmission time of the corresponding frame. Indicates the first The first time period The average of the time differences between receiving and sending all frames in the group. Indicates the sequence number of the frame in the group. This indicates the number of image frames within the group. Indicates the drift time period number, Indicates the frame group number within the segment;

[0077] The baseline offset construction submodule calls the image frame drift time interval value and constructs the reception time series based on the reception time and transmission time of each corresponding frame group. With transmission time series Calculate the time difference between each corresponding frame, that is, the time difference of each frame. This forms a set of frame time difference values, and then the frame index position corresponding to the smallest difference value is located in this set. Set the existence of frame group number Group and time period numbered as Number of frames within a group The reception time of the frame group is:

[0078] ;

[0079] The corresponding sending time is:

[0080] ;

[0081] The time differences between receiving and transmitting each frame are as follows:

[0082] ;

[0083] ;

[0084] ;

[0085] ;

[0086] The difference corresponding to frame 1 is the minimum value, which is the reference frame. As a baseline frame time reference group, used for time axis alignment, the image frame time difference offset is then calculated according to the following formula:

[0087] ;

[0088] in:

[0089] , where is the number of the current frame group;

[0090] , which is the average time difference between receiving and sending.

[0091] Substitute the square of the difference between each frame and the average value:

[0092] Frame 1: ;

[0093] Frame 2: ;

[0094] Frame 3: Same as above, 0.00000625;

[0095] Frame 4: Same as above, 0.00000625;

[0096] Sum the above squared values ​​and divide by the number of frames:

[0097] ;

[0098] Therefore, the time difference offset between receiving and transmitting the second frame in time period 1 is: This value will be written into the image frame time axis baseline offset structure as a quantitative index value of the time axis offset of the image frame segment, forming the image frame main time axis reference adjustment parameter.

[0099] formula:

[0100] ;

[0101] The formula is a standard deviation structure, and its function is to measure a given frame group (the first frame). Group) in a certain time period (the first group) The degree to which the difference between the receiving and sending time of each image frame within a segment deviates from the mean, i.e., the degree of dispersion of time synchronization.

[0102] Meaning: The The difference between the actual reception time and the transmission time of a frame image, i.e., the time consumed by the frame's transmission over the network. Purpose: It is a fundamental element constituting a time difference set, providing raw data for drift measurement.

[0103] Meaning: The Time period The mean of the receive and transmit time differences of all frames in a frame group. Purpose: To serve as a reference standard, measuring the degree of deviation of each frame's time difference from the group's average, and quantifying consistency.

[0104] Meaning: The squared error of the time difference between each frame and the mean. Purpose: To emphasize that the larger the deviation, the more significant the impact on the overall fluctuation, and to help highlight abnormal drift frames.

[0105] Meaning: Average the squared deviations of all frames within a group. Purpose: To eliminate the influence of frame number differences and make the offsets of different frame groups comparable.

[0106] Meaning: Calculates the square root of the deviation, bringing the unit back to the original dimension of "time difference" (e.g., milliseconds). Purpose: Facilitates comparison with a set "image frame synchronization stability threshold" to assess whether a group of frames has experienced significant drift.

[0107] Formula output value The meaning is to indicate a certain period of time (the first... A certain frame group (segment) The degree of dispersion in the receive-transmit time difference offset of all image frames within a group. A larger value indicates more severe network fluctuations or clock asynchrony, reflecting instability in the synchronization state. It can be used to identify drift segments, establish corrective baselines, and assist in the stability control of the time axis.

[0108] The formula's position within the system:

[0109] This formula is used in a core calculation step of the "baseline offset construction submodule" in the patented solution, supporting the following functions: identifying the frame group with the smallest time difference as the reference frame within each image frame drift time period; and quantifying the drift degree... Select a time reference group with minimal drift; then establish an "image frame time axis baseline offset structure" and output parameters for time synchronization adjustment.

[0110] Please see Figure 2 and Figure 5 The delayed structure analysis module includes a benchmark difference extraction submodule, an offset determination and filtering submodule, and a fluctuation segment labeling and classification submodule.

[0111] The reference difference extraction submodule adjusts the parameters of the main time axis of the image frame according to the reference parameters, obtains the reception time of the audio frame and the time axis reference offset of the image frame within the time period of each image frame, subtracts the corresponding image frame reference offset from the reception time of each audio frame, constructs the time axis difference sequence of the audio frame and the image frame, and generates the audio and video frame reference difference sequence.

[0112] The baseline difference extraction submodule, based on the image frame's main time axis baseline adjustment parameters, first obtains the time period number to which each image frame belongs and retrieves the corresponding baseline offset value sequence. This offset is provided by the previously constructed time axis baseline offset structure. Then, for all audio frames within the time period, it sorts them by received timestamp and sequentially matches the corresponding time period of the image frame, extracting the baseline offset value for that segment. Subtraction is then performed to obtain the time axis adjustment result of the audio frame, constructing a time axis difference sequence between the audio frame and the image frame. Specifically, let the image frame's... The time axis reference offset value corresponding to the segment is The voice frame reception time is The time axis difference of the corresponding image frames is then... For example, let the reference offset value of the second image frame be... If the reception time of the voice frame is , , The corresponding time axis differences are respectively , , The above time differences form a time axis difference sequence in sequence. The data are then written into the audio and video frame baseline difference value sequence in the order of the voice frame index.

[0113] The offset determination and filtering submodule calls the audio and video frame baseline difference value sequence, obtains the concentration of difference values ​​in each time period, calculates the maximum local offset difference value, determines whether the maximum value exceeds the synchronization critical difference threshold, marks the interval where the offset difference value exceeds the threshold, and generates the segment value of the offset difference concentration interval.

[0114] The offset determination and filtering submodule calls the audio and video frame baseline difference value sequence to obtain the difference value set for each time period. It then extracts the difference subsequence for each time period by sorting by frame index, performs an extreme value retrieval operation to extract the maximum value element, and sets it as... To determine the baseline value for offset, the statistical central tendency of all differences within that time period is then calculated. The degree of variation is calculated using the standard deviation method to further assess the consistency of delay within a local time period. The calculation formula is as follows: ,in The average difference value is given by the difference subsequence. ,but The corresponding standard deviation is: ;

[0115] Determine the maximum difference within this segment. Does it exceed the synchronization critical difference threshold? This threshold The value should be set according to the audio / video synchronization error tolerance range, usually controlled within 50 milliseconds. ,Depend on If the threshold is not exceeded, the segment is not marked. If another segment has a maximum difference of 0.210 seconds and an average of 0.149 seconds, then the difference is 0.061 seconds, which exceeds the threshold. It is determined to be a delayed offset segment, and the segment value of the offset difference range is generated.

[0116] The fluctuation segment labeling and classification submodule extracts the start and end time points and audio frame sequence index of each interval segment based on the segment value of the interval segment in the offset difference set, and labels the segments that meet the continuity and threshold conditions as synchronous fluctuation segments, establishes a corresponding time period index mapping table, and obtains multi-frame audio and video delay fluctuation interval labeling information.

[0117] The fluctuation segment labeling and classification submodule traverses the start and end frame indices of each segment based on the segment values ​​within the concentrated interval of offset differences. It extracts the receiving timestamps corresponding to the speech frame indices as the start and end time points of the segment and judges the sequential continuity of all frame indices within the segment. If the index difference is continuously not greater than 1, it is marked as a continuous sequence segment. Then, it checks whether the inter-frame difference continuously satisfies the delay difference being greater than a set threshold. The system verifies the conditions. If at least three consecutive frames within a segment meet the conditions, it is defined as a fluctuation segment that satisfies the continuity and delay offset conditions. For example, if the frame indices {11, 12, 13} within the segment correspond to differences of {0.189, 0.193, 0.201}, all of which are greater than 0.050 seconds, the segment meets the conditions and is marked as a synchronous fluctuation segment. The system writes the segment's start and end time values ​​and the audio frame index into the annotation record table and constructs a mapping table structure from the time period number to the audio frame index, ultimately forming multi-frame audio and video delay fluctuation interval annotation information.

[0118] Please see Figure 2 and Figure 6 The synchronization node generation module includes a frame timing sorting submodule, a frame pair difference extraction submodule, and a synchronization node aggregation submodule.

[0119] The frame timing sorting submodule calls the multi-frame audio and video delay fluctuation interval annotation information, extracts the reception timestamp and frame number of each image frame and audio frame in each segment, sorts the reception timestamps of the image frames and audio frames in each segment in ascending order according to the frame number, establishes a frame pair sequence index mapping table, and generates a sorted frame timing index set.

[0120] The frame timing sorting submodule calls the multi-frame audio and video delay fluctuation interval annotation information, extracts the start and end time points of each segment, filters the corresponding image and audio frame data according to the time range of the annotated segment, retrieves the reception timestamps and frame numbers of the image and audio frames in each segment, and constructs image frame sets respectively. With voice frame set ,in , For image and audio frame numbers. , To assign a timestamp to the received data, the system sorts the frame data within each set in ascending order by frame number. After sorting, a frame time sequence index structure is established. Specifically, the sorted image frame sequence and audio frame sequence are mapped to a frame pair sequence index table using a double index structure, i.e., sorted by image frame... With the nearest speech frame Pairing, for example, if the image frame numbers are 3, 5, and 8, and the audio frame numbers are 4, 6, and 7, then the matching pair is: Finally, the mapping structure is recorded in the sorted frame time sequence index set for subsequent frame difference processing.

[0121] The frame pair difference extraction submodule obtains the received timestamps between adjacent image frames and audio frames based on the sorted frame time sequence index set, performs time difference calculation for each pair of image frames and audio frames, filters the frame pairs with the smallest time difference, records the received timestamps of the corresponding frame pairs and marks them as candidate node frames, and generates the minimum difference frame pair timestamp group.

[0122] The frame pair difference extraction submodule, based on the sorted frame time sequence index set, obtains the received timestamp difference between each pair of image frames and audio frames, and defines the difference between each pair of frames as _____. Iterate through all frame pairs in the index set, perform difference calculations on each pair, and filter out the pairs with the smallest difference. Assume there exists a sorted set of pairs. The corresponding receiving timestamps are respectively , , Calculate the difference. , , The system selects the frame pair with the smallest difference, i.e., (3,4), and assigns the received timestamp values ​​of this frame pair to the appropriate values. and Record the candidate node frame timestamp and the frame sequence number. Simultaneously, the timestamp group of the frame pair with the smallest difference is written, and in this way, the frame pair difference extraction and the minimum difference pair filtering are completed in each segment.

[0123] The synchronization node aggregation submodule calls the minimum difference frame pair timestamp group, extracts the timestamp and frame pair index number of the candidate node frame corresponding to each segment, and aggregates the information of each candidate node frame in time axis order to construct a candidate synchronization control node sequence and obtain a candidate synchronization control node list.

[0124] The synchronization node aggregation submodule calls the minimum difference frame pair timestamp group to extract the received timestamps of candidate frame pairs segment by segment. With frame number All candidate frames are reordered in ascending order of image frame timestamps. The system then constructs a candidate synchronization control node structure, with the following structure: ,in As a sorting criterion, for example, if three candidate pairs of segments are Arranged in ascending order of time as described above, and finally aggregated in array form to construct a candidate synchronization control node sequence. The system writes this sequence into the candidate synchronization control node list for use in the construction of reference synchronization anchors in subsequent synchronization control strategies.

[0125] Please see Figure 2 and Figure 7 The synchronization control instruction module includes a playback offset extraction submodule, an error interval judgment submodule, and a control instruction generation submodule;

[0126] The playback offset extraction submodule calls the candidate synchronization control node list, obtains the playback timestamps of image frames and audio frames in the current playback queue, calculates the offset value between the playback timestamp of each pair of frames and the timestamp of the corresponding candidate node, calculates the audio and video frame synchronization offset value, establishes an index table of playback offset values ​​for image frames and audio frames, and generates a group of audio and video frame playback offset values.

[0127] The formula for calculating the audio and video frame synchronization offset is:

[0128] ;

[0129] in, Representing the The first in the section The playback offset values ​​between image frames and audio frames. Representing the Section 1 The playback time marker for the image frame in the current playback queue, This represents the timestamp of the corresponding image frame in the list of candidate synchronization control nodes. Representing the Section 1 The playback time of each audio frame in the current playback queue is marked. This represents the timestamp of the corresponding voice frame in the list of candidate synchronization control nodes. For synchronization section numbering, The frame pair number within the segment. Indicates the image frame type marker, Indicates the voice frame type marker;

[0130] The playback offset extraction submodule calls the candidate synchronization control node list, first based on the image frame timestamps recorded in the synchronization control nodes. With audio frame timestamps Extract the corresponding frame pair index and retrieve the corresponding playback time marker for the frame pair in the current playback queue. , Construct a structure comparing the original playback records with node timestamps, and calculate the playback offset value for each pair of frames. Using formula For example: Suppose the second pair of frames in the first segment has the following playback timestamps: , The timestamp corresponding to the candidate synchronization node is , Then the playback offset value is:

[0131] ;

[0132] The system will match the offset value with the frame pair index. Write the audio / video frame playback offset value index table, and... Include the audio and video frame playback offset value group.

[0133] This formula is used for synchronization offset detection during audio and video playback. By comparing the deviation between the actual playback time of image frames and audio frames in the playback queue and the marked time in the candidate synchronization control nodes, the audio and video playback offset value of the current frame pair is calculated, providing a quantitative basis for the generation of subsequent synchronization control commands.

[0134] Formula expression and meaning:

[0135] ;

[0136] The formula calculates the first In the first synchronization segment The absolute value of the average playback offset between image frames and audio frames.

[0137] The function of each part in the formula:

[0138] Meaning: Image frame playback offset value. Calculation method: Playback timestamp of image frames in the current playback queue (…). Subtract its timestamp in the candidate synchronization control node list. Purpose: To reflect whether the image frame lags or advances relative to the preset time reference (synchronization point) during playback.

[0139] Meaning: Audio frame playback offset value. Calculation method: Audio frame playback time stamp ( Subtract its corresponding synchronization node timestamp () Purpose: To reflect the degree of deviation between the synchronization of audio frame playback and the reference time point.

[0140] Meaning: Calculate the arithmetic mean of the offsets between image frames and audio frames. Purpose: Construct a comprehensive offset between image and audio frames to describe the degree of audio-visual synchronization and avoid misjudgment caused by a single frame deviation.

[0141] Meaning: Take the absolute value. Purpose: To eliminate the influence of offset direction and retain only the magnitude of the offset, used for comparison with the allowable range of synchronization error.

[0142] Output The practical significance:

[0143] This represents the offset for audio-video synchronization, specifically the first... In the first synchronization segment The average value of the playback time reference error between frames. The larger the value, the more obvious the time difference between the audio and video frames, which may cause synchronization anomalies such as lip-sync misalignment, image delay, or premature playback.

[0144] The value will then be compared with the system's set "synchronization error tolerance range" to indicate whether synchronization adjustment is needed.

[0145] The error interval judgment submodule extracts each pair of offset values ​​based on the audio and video frame playback offset value group and compares them with the set synchronization error tolerance range. Frame pairs with offset values ​​exceeding the synchronization error tolerance range are marked, and the corresponding frame pair index and offset direction are recorded to generate a frame synchronization error identifier set.

[0146] The error interval determination submodule is based on the audio and video frame playback offset value group, and for each offset value... The system's set synchronization error tolerance threshold Perform a difference assessment; this threshold... The reference standard sets the maximum voice synchronization error tolerance to 150 milliseconds. Considering that video frames have better error tolerance than voice frames, the offset tolerance in synchronization control is reasonably reduced to [missing value]. The system executes for each item and The size comparison operation, if the offset value exceeds Then mark the frame pair as a synchronization error frame, as in the example above. If the time limit is less than 0.050 seconds, no flag is set. If another frame pair is as follows: If it exceeds the threshold, the index should be... With offset direction information, i.e. , The error identifiers are recorded to indicate that the image playback is premature and the voice playback is delayed. The directional information is used for the generation of subsequent control commands.

[0147] The control instruction generation submodule calls the frame synchronization error identifier set, adds or subtracts offset values ​​according to the frame type based on the playback time stamp and offset direction information of the frame pair, and establishes an adjusted playback time stamp sequence. The adjustment results are summarized to construct a frame-level synchronization control record table, and a frame-level audio and video synchronization control instruction set is obtained.

[0148] The control command generation submodule calls the frame synchronization error flag set, reads the index and offset direction information of frame pairs one by one, and determines whether the playback direction needs to be moved forward or delayed. If the image frame playback time is earlier than the node timestamp, then... Then the playback time of the image frames needs to be adjusted backward. If the playback time of the audio frame is later than the node timestamp, then... Then the playback time of the audio frame needs to be adjusted forward. This method is used to construct the adjusted playback time markers. For example, for the aforementioned frame pair (1,3), if Therefore, the image frame needs to be delayed by 0.061 seconds, and the audio frame needs to be advanced by 0.061 seconds. The adjusted markers are as follows:

[0149] ;

[0150] ;

[0151] The system categorizes and summarizes all adjusted playback marker pairs into image frame adjustment sequences and audio frame adjustment sequences according to frame type, records them in the frame-level synchronization control record table, and outputs them as a frame-level audio and video synchronization control instruction set.

[0152] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A real-time audio and video synchronous monitoring system for video calls, characterized in that, The system includes: The frame time extraction module obtains the parameter information of each audio frame and image frame during the video call, extracts the sending timestamp and receiving timestamp, calculates the difference, constructs a comparison sequence of the transmission delay interval of audio frames and image frames, and obtains synchronization difference sequence data. The main timing locking module calls the synchronization difference sequence data, calculates the frame time interval offset and offset rate between two adjacent frames, compares them with the synchronization stability reference threshold, sets the reference frame time reference group, and generates the image frame main time axis reference adjustment parameters. The delay structure analysis module adjusts the parameters based on the main time axis reference of the image frame, compares the time axis reference difference between the audio frame and the image frame, filters out the concentrated distribution area of ​​the offset difference, and marks the corresponding time point interval as the synchronization fluctuation segment to obtain the multi-frame audio and video delay fluctuation interval labeling information. The synchronization node generation module calls the multi-frame audio and video delay fluctuation interval labeling information, judges the receiving time difference of adjacent image frames and audio frames after sorting, extracts the frame pair with the smallest time difference, records the timestamp and marks it as a candidate synchronization control node, and generates a candidate synchronization control node list. The synchronization control instruction module obtains the candidate synchronization control node list, and calls the playback time marker of the image frame and audio frame to be output in the current playback queue according to the timestamp of each pair of image frames and audio frames. It compares the offset value of the playback time marker with the corresponding timestamp in the node list. If the offset value exceeds the acceptable synchronization error range, it adjusts the playback time marker by adding or subtracting the offset value to generate a frame-level audio and video synchronization control instruction set. The frame-level audio and video synchronization control instruction set includes image frame playback offset adjustment parameters, voice frame playback time correction parameters, and synchronization tolerance range correction identifiers.

2. The real-time audio and video synchronization monitoring system for video calls according to claim 1, characterized in that, The synchronization difference sequence data includes a list of audio frame and image frame time difference comparisons, channel sequence number mapping information, and an image frame transmission delay distribution table. The image frame main time axis reference adjustment parameters specifically include a reference frame time offset index, a reference offset replacement parameter, and a frame time drift identifier label value. The multi-frame audio and video delay fluctuation interval labeling information includes a synchronization critical difference distribution record, an inter-frame fluctuation amplitude index table within a time period, and a fluctuation segment time node index list. The candidate synchronization control node list specifically includes a frame pair minimum reception time difference identifier, a node timestamp comparison table, and a synchronization control candidate index set.

3. The real-time audio and video synchronization monitoring system for video calls according to claim 2, characterized in that, The frame time extraction module includes: The frame parameter acquisition submodule acquires the parameter information of each audio frame and image frame during the video call, including the encoding type identifier, sending timestamp, receiving timestamp and receiving channel sequence number. It performs time difference processing on the sending timestamp and receiving timestamp of each frame, filters complete frame information groups and removes incomplete frame records, establishes a set of valid frame indexes based on the data integrity judgment flag, and generates a set of valid frame index numbers. The channel matching and recognition submodule extracts the receiving channel sequence number corresponding to each voice frame and image frame based on the set of valid frame index numbers, performs frame group pairing according to the channel sequence number, records the encoding type identifier for each group of matching frames, and marks the frame groups with the same type identifier as data blocks to be compared, thereby generating a set of channel synchronization matching identifiers. The delay difference generation submodule calls the channel synchronization matching identifier set, extracts the transmission timestamp and reception timestamp of the voice frame and image frame in each frame pair, normalizes the timestamp difference and establishes the time series corresponding to the frame pair, calculates the voice and image transmission delay difference through the time series position index matching relationship and generates interval records to obtain synchronization difference sequence data.

4. The real-time audio and video synchronization monitoring system for video calls according to claim 3, characterized in that, The master timing locking module includes: The transmission interval calculation submodule calls the synchronization difference sequence data, obtains the reception time difference between any two adjacent image frames based on the reception timestamp of the image frame, performs the calculation of the difference ratio between consecutive frames for the time difference between consecutive frames, compares the ratio with the image frame interval offset reference threshold, marks the frame group with the ratio greater than the image frame interval offset reference threshold, and generates the image frame interval offset rate sequence. The frame drift recognition submodule divides the time period according to the continuity of the frame group index based on the image frame interval offset rate sequence, extracts the average offset rate of each continuous segment and judges it with the image frame synchronization stability threshold, filters the time period with the average offset rate exceeding the image frame synchronization stability threshold, marks the time period as drift segment, and generates the image frame drift time period interval value. The baseline offset construction submodule calls the image frame drift time interval value, calculates the set of differences between the receiving time and the sending time in each segment according to the receiving time and sending time of the corresponding frame group, locates the minimum value frame index position in the difference set, sets the receiving time and sending time value of the corresponding frame as the reference frame time reference group, calculates the image frame time difference offset, establishes the image frame time axis baseline offset structure, and obtains the image frame main time axis reference adjustment parameters.

5. The real-time audio and video synchronization monitoring system for video calls according to claim 4, characterized in that, The formula for calculating the temporal difference offset of image frames is: ; in, Representing the In the time period of the th time period The time difference offset between the reception and transmission of a group of image frames. Indicates the first Within the time period, the first In the frame The reception time of each image frame Indicates the transmission time of the corresponding frame. Indicates the first The first time period The average of the time differences between receiving and sending all frames in the group. Indicates the sequence number of the frame in the group. Indicates the number of image frames within a group. Indicates the drift time period number, Indicates the frame group number within the segment.

6. The real-time audio and video synchronization monitoring system for video calls according to claim 5, characterized in that, The delay structure analysis module includes: The reference difference extraction submodule obtains the reception time of the audio frame and the reference offset of the image frame time axis within the time period of each image frame according to the reference adjustment parameters of the main time axis of the image frame. It subtracts the corresponding reference offset of the image frame from the reception time of each audio frame to construct the time axis difference sequence of the audio frame and the image frame, and generates the reference difference sequence of audio and video frames. The offset determination and filtering submodule calls the audio and video frame reference difference value sequence, obtains the concentration of difference values ​​in each time period and calculates the maximum local offset difference value, determines whether the maximum value exceeds the synchronization critical difference threshold, marks the interval where the offset difference value exceeds the threshold, and generates the segment value of the offset difference concentration interval. The fluctuation segment labeling and classification submodule extracts the start and end time points and audio frame sequence index of each interval segment based on the segment value of the interval in the offset difference set, and labels the segments that meet the continuity and threshold conditions as synchronous fluctuation segments, establishes a corresponding time period index mapping table, and obtains multi-frame audio and video delay fluctuation interval labeling information.

7. The real-time audio and video synchronization monitoring system for video calls according to claim 6, characterized in that, The synchronization node generation module includes: The frame timing sorting submodule calls the multi-frame audio and video delay fluctuation interval annotation information, extracts the reception timestamp and frame number of each image frame and audio frame in each segment, sorts the reception timestamps of the image frames and audio frames in each segment in ascending order according to the frame number, establishes a frame pair sequence index mapping table, and generates a sorted frame timing index set. The frame pair difference extraction submodule obtains the received timestamps between adjacent image frames and audio frames based on the sorted frame time sequence index set, performs time difference calculation for each pair of image frames and audio frames, filters the frame pairs with the smallest time difference, records the received timestamps of the corresponding frame pairs and marks them as candidate node frames, and generates the minimum difference frame pair timestamp group. The synchronization node aggregation submodule calls the minimum difference frame pair timestamp group, extracts the timestamp and frame pair index number of the candidate node frame corresponding to each segment, and aggregates the information of each candidate node frame in time axis order to construct a candidate synchronization control node sequence and obtain a candidate synchronization control node list.

8. The real-time audio and video synchronization monitoring system for video calls according to claim 7, characterized in that, The synchronization control command module includes: The playback offset extraction submodule calls the candidate synchronization control node list to obtain the playback timestamps of image frames and audio frames in the current playback queue. It calculates the offset value between the playback timestamp of each pair of frames and the timestamp of the corresponding candidate node, calculates the audio and video frame synchronization offset value, establishes an index table of playback offset values ​​for image frames and audio frames, and generates a group of audio and video frame playback offset values. The error interval judgment submodule extracts each pair of offset values ​​based on the audio and video frame playback offset value group and compares them with the set synchronization error tolerance range. It marks the frame pairs whose offset values ​​exceed the synchronization error tolerance range, records the corresponding frame pair index and offset direction, and generates a frame synchronization error identifier set. The control instruction generation submodule calls the frame synchronization error identifier set, adds or subtracts offset values ​​according to the frame type based on the playback time marker and offset direction information of the frame pair, and establishes an adjusted playback time marker sequence. The adjustment results are summarized to construct a frame-level synchronization control record table, and a frame-level audio and video synchronization control instruction set is obtained.

9. The real-time audio and video synchronization monitoring system for video calls according to claim 8, characterized in that, The formula for calculating the audio / video frame synchronization offset is: ; in, Representing the The first in the section The playback offset values ​​between image frames and audio frames. Representing the Section 1 The playback time marker for the image frame in the current playback queue, This represents the timestamp of the corresponding image frame in the list of candidate synchronization control nodes. Representing the Section 1 The playback time of each audio frame in the current playback queue is marked. This represents the timestamp of the corresponding voice frame in the list of candidate synchronization control nodes. For synchronization section numbering, The frame pair number within the segment. Indicates the image frame type marker, This indicates the voice frame type marker.

Citation Information

Patent Citations

  • H.324M protocol-based 3G video telephone audio and video synchronization device and method thereof

    CN101742548A

  • System and method for extracting caption time axis from multi-audio-track video file

    CN111212319A