A deep learning-based audio-video synchronicity detection method and system

By employing a deep learning-based audio-video synchronization detection method, which utilizes multimodal feature extraction and the SyncNet deep dual-stream network, the method addresses the insufficient accuracy of existing audio-video synchronization detection techniques, achieving efficient synchronization determination and generating highly reliable detection reports.

CN121691794BActive Publication Date: 2026-07-21SUYUAN TECHNOLOGY (HUNAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUYUAN TECHNOLOGY (HUNAN) CO LTD
Filing Date
2025-12-16
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing audio-video synchronization detection methods struggle to achieve high accuracy and reliability in complex scenarios, particularly in situations with background noise and unclear speech, failing to effectively address the issues of high accuracy and reliability in audio-video synchronization.

Method used

A deep learning-based audio-video synchronization detection method is adopted. By separating the audio and video streams, multimodal audio and video features are generated using face detection and feature extraction. Non-target speech interference segments are removed, and the SyncNet deep dual-stream network is used for synchronization discrimination. The anomaly type is determined by combining lip movement features and audio energy, a structured anomaly log is generated, and synchronization is verified by a temporal alignment algorithm. Finally, a detection report is generated.

Benefits of technology

It achieves efficient extraction and determination of audio and video synchronization, improves the accuracy of synchronization detection, and generates high-quality detection reports, ensuring the high credibility and judicial acceptance of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121691794B_ABST
    Figure CN121691794B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep learning's audio-video synchronism detection method and system, it is related to digital audio-video processing technical field, including the audio-video file to be detected of acquisition, video stream and audio stream separation is carried out, and utilize face detection and feature extraction, generate multimodal audio-video feature;According to pure data packet, using SyncNet depth double-flow network carries out synchronism discrimination, respectively extracts lip movement feature and speech feature, generates synchronous discrimination result packet;According to detection report, using time sequence alignment algorithm carries out verification to audio-video synchronism, and completes result solidification by time stamp anchoring, obtains synchronism determination result.The application is discriminated by using SyncNet depth double-flow network to carry out synchronism, and the efficient extraction and synchronism determination of audio-video lip movement feature and speech feature are realized in combination with deep learning, it is helpful to accurately identify the matching degree between target speech and mouth shape, improves the accuracy of synchronism detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital audio and video processing technology, and in particular to a method and system for detecting audio and video synchronization based on deep learning. Background Technology

[0002] With the development of deep learning and computer vision, audio-video synchronization detection has gradually become a research hotspot in multimedia technology. In the fields of speech recognition, video editing, virtual reality, and the judicial system, higher demands are being placed on the accurate detection of audio-video synchronization. Existing audio-video synchronization detection methods are constantly being optimized and gradually incorporating deep learning to improve detection accuracy and adaptability. Deep learning-based audio-video synchronization detection methods automatically extract video and audio features and utilize neural networks to determine synchronization. These methods are widely used in video review, virtual anchors, and voice matching scenarios, providing significant performance improvements.

[0003] Existing audio-video synchronization detection methods still have some shortcomings, especially in accurately determining audio-video synchronization. Traditional synchronization detection methods mainly rely on simple feature matching and rule-based algorithms, which are often unable to cope with complex scene changes, such as background noise, unclear speech, and video frame loss; their judgment accuracy in variable environments is insufficient, and they cannot effectively solve the problems of high accuracy and high reliability in audio-video synchronization determination. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a deep learning-based audio-video synchronization detection method to address the problem of insufficient accuracy and precision in synchronization detection.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a deep learning-based audio-video synchronization detection method, comprising: separating the video stream and audio stream of the acquired audio-video file to be detected, and generating multimodal audio-video features using face detection and feature extraction; performing voiceprint recognition screening on the audio window based on the multimodal audio-video features to remove non-target speech interference segments and obtain a clean data packet; using the clean data packet, employing a SyncNet deep dual-stream network for synchronization discrimination, extracting lip movement features and speech features respectively, and generating a synchronization discrimination result packet; identifying abnormal frames in the audio-video based on the synchronization discrimination result packet, and determining the abnormality type by combining lip movement amplitude and audio energy to obtain a structured abnormality log; encapsulating and integrating the abnormality log and the synchronization discrimination result packet to generate a detection report; and using a time-series alignment algorithm to verify the audio-video synchronization based on the detection report, and completing the result solidification through timestamp anchoring to obtain a synchronization determination result.

[0007] As a preferred embodiment of the deep learning-based audio-video synchronization detection method of the present invention, the audio-video file to be detected refers to the original audio-video data for audio-video synchronization detection.

[0008] As a preferred embodiment of the deep learning-based audio-video synchronization detection method described in this invention, the steps include: separating the video stream and audio stream of the acquired audio-video file to be detected, and generating multimodal audio-video features using face detection and feature extraction. The specific steps are as follows: For the audio and video files to be tested, the video stream and audio stream are separated to obtain video frames and audio signals; Based on video frames, face detection is performed on each frame of the image and the lip region is cropped to obtain lip movement features; By analyzing the audio signal, a Mel spectrogram is generated and spectral features are extracted. Combining lip movement features and spectral features, multimodal audio and video features are generated.

[0009] As a preferred embodiment of the deep learning-based audio-video synchronization detection method of the present invention, the step of performing voiceprint recognition screening on the audio window based on multimodal audio-video features to remove non-target speech interference segments and obtain clean data packets includes the following specific steps. Based on multimodal audio and video features, the audio window is filtered through voiceprint recognition to identify the target speech segment, remove background noise and irrelevant audio, and obtain audio data packets; Based on audio data packets, a speech feature analysis method is used to filter and process audio segments. By extracting speech features and removing background noise and non-target speech interference segments, speech data is obtained. The audio data packet and voice data are synchronized and verified to obtain a clean data packet.

[0010] As a preferred embodiment of the deep learning-based audio-video synchronization detection method described in this invention, the SyncNet deep dual-stream network refers to the method of detecting the synchronization between audio signals and video signals through feature encoding methods and temporal alignment algorithms.

[0011] As a preferred embodiment of the deep learning-based audio-video synchronization detection method described in this invention, the following steps are taken: Synchronization discrimination is performed using a SyncNet deep dual-stream network based on clean data packets, extracting lip movement features and speech features respectively, and generating a synchronization discrimination result packet. A feature encoding method is used to jointly encode the lip region and audio segments in the video stream to obtain a multimodal feature sequence; Based on multimodal feature sequences, a temporal alignment algorithm is used to calculate the correlation and synchronicity of lip movement features and speech features within each time window, and obtain a frame-level synchronicity score sequence. The synchronization status of the frame-level synchronization scoring sequence is determined to obtain the synchronization discrimination result packet.

[0012] As a preferred embodiment of the deep learning-based audio-video synchronization detection method described in this invention, the steps of identifying abnormal frames in the audio-video based on the synchronization discrimination result packet, and determining the anomaly type by combining lip movement amplitude and audio energy to obtain a structured anomaly log are as follows. A timestamp anchoring method is used to align audio energy with lip movement amplitude to obtain a multimodal frame-level feature set; Based on the abnormal inter-frame differences in lip movement amplitude and audio energy, abnormal frames are screened in the multimodal frame-level feature set to obtain a multimodal abnormal frame set. The multimodal abnormal frame set is restructured to obtain a structured abnormal log for audio and video location and synchronization discrimination.

[0013] As a preferred embodiment of the deep learning-based audio / video synchronization detection method of the present invention, the specific steps for encapsulating and integrating the abnormal logs and synchronization discrimination result packets to generate a detection report are as follows. The structured anomaly logs and the synchronization discrimination result packets are correlated and integrated to obtain the anomaly event synchronization dataset; Based on the statistical induction of the abnormal event synchronization dataset, the lip movement amplitude, audio energy and abnormality type label of each abnormal event are aggregated and organized to obtain the statistical summary data of abnormal events. The statistical summary data of abnormal events is analyzed and classified, and the data is mined by combining audio and video features. The data is then organized by paragraphs, tables, and timelines to generate detection reports.

[0014] As a preferred embodiment of the deep learning-based audio-video synchronization detection method described in this invention, the following steps are taken: Based on the detection report, a time-series alignment algorithm is used to verify the audio-video synchronization, and the result is solidified through timestamp anchoring to obtain the synchronization determination result. Based on the detection report, a time-series alignment algorithm is used to perform deviation analysis and alignment correction on the time series of the audio track and the video track, and an alignment result set is obtained. By aligning the result set, the positional relationship between audio segments and video frames is fixed by time stamping and offset correction using the timestamp anchoring method, and the anchoring result package with time coordinates and synchronization status marks is obtained. The synchronization of the anchoring result packet is checked to evaluate the consistency of audio and video status in different time periods, identify abnormal situations, and obtain the synchronization judgment result.

[0015] Secondly, this invention provides a deep learning-based audio-video synchronization detection system, comprising: a separation and extraction template, which separates the video stream and audio stream of the acquired audio-video file to be detected, and generates multimodal audio-video features using face detection and feature extraction; a screening and removal template, which performs voiceprint recognition screening on the audio window based on the multimodal audio-video features, removes non-target speech interference segments, and obtains a clean data packet; a discrimination and extraction template, which uses a SyncNet deep dual-stream network to perform synchronization discrimination based on the clean data packet, extracts lip movement features and speech features respectively, and generates a synchronization discrimination result packet; an identification and judgment template, which identifies abnormal frames in the audio-video based on the synchronization discrimination result packet, and judges the abnormality type by combining lip movement amplitude and audio energy, and obtains a structured abnormality log; a data integration template, which encapsulates and integrates the abnormal log and the synchronization discrimination result packet to generate a detection report; and a verification and solidification template, which uses a time-series alignment algorithm to verify the audio-video synchronization based on the detection report, and solidifies the result by timestamp anchoring to obtain a synchronization judgment result.

[0016] The beneficial effects of this invention are as follows: By employing the SyncNet deep dual-stream network for synchronization discrimination, combined with deep learning, efficient extraction and synchronization determination of audio and video lip movement features and speech features are achieved, which helps to accurately identify the matching degree between target speech and lip shape, and improves the accuracy of synchronization detection; by encapsulating and integrating the synchronization discrimination result package, comprehensive recording and structured management of the synchronization discrimination results are achieved, which helps to generate high-quality detection reports, provides data support for verification and evidence preservation, and ensures the high credibility and judicial acceptance of the results. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a deep learning-based audio-video synchronization detection method.

[0019] Figure 2 This is a schematic diagram of a deep learning-based audio-video synchronization detection system.

[0020] Figure 3 A flowchart for generating multimodal audio and video features.

[0021] Figure 4 This is a flowchart for anomaly detection and report generation. Detailed Implementation

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a deep learning-based audio-video synchronization detection method, comprising the following steps: S1: Separate the video stream and audio stream from the collected audio and video files to be detected, and generate multimodal audio and video features using face detection and feature extraction.

[0026] S1.1: The audio and video file to be detected refers to the original audio and video data for audio and video synchronization detection.

[0027] S1.2: Separate the video stream from the audio stream of the audio / video file to be tested to obtain video frames and audio signals.

[0028] Furthermore, the audio and video files to be detected are subjected to audio and video demultiplexing processing. The mixed encapsulated audio and video data are split into independent video tracks and audio tracks. The video tracks are sampled and decoded to generate a continuous sequence of video frames, and the audio tracks are sampled and decoded to generate a corresponding audio signal sequence. This completes the separation of the video stream and the audio stream, and obtains video frames and audio signals.

[0029] It should be noted that audio and video demultiplexing refers to separating the video stream and audio stream in the audio and video file to be detected by decapsulation (decomposing the encapsulation format in the audio and video file and extracting the video stream and audio stream), and then performing feature extraction and analysis on the video and audio separately.

[0030] Sampling decoding refers to extracting audio and video data from audio and video files and performing decoding processing to convert audio and video data into video frame and audio signal formats.

[0031] Video streams and audio streams refer to independent data streams separated from the audio and video files to be detected. Video streams contain frame-by-frame image data, while audio streams contain speech signal data.

[0032] Video tracks and audio tracks refer to the timelines in the audio and video files to be tested, which respectively carry the video frame sequence and the audio signal sequence.

[0033] S1.3: Based on video frames, perform face detection on each frame and crop the lip region to obtain lip movement features.

[0034] Furthermore, face detection is performed frame by frame on the separated video frames to locate the face regions in the images. Within the detected face regions, the coordinates of the mouth contour are identified. Based on the coordinates of the mouth contour, the lip region is cropped to obtain a local image sequence of lip movements, which is then used as lip movement features.

[0035] Specifically, face detection refers to analyzing each frame of a video image using feature encoding methods to identify faces in the image and extract the coordinates of the lip region.

[0036] S1.4: By analyzing the audio signal, a Mel spectrogram is generated and spectral features are extracted. Combined with lip movement features and spectral features, multimodal audio and video features are generated.

[0037] Furthermore, the separated audio signals are processed into frames according to time windows. Each frame of audio signal is analyzed using a time alignment algorithm to generate a Mel spectrogram and extract spectral features. The extracted spectral features are then time-aligned and spliced ​​with the lip movement features to generate multimodal audio and video features.

[0038] It should be noted that a time window refers to the time period during which audio and video are processed frame by frame, and each time window contains audio frames and video frames.

[0039] Frame segmentation refers to dividing the audio signal into consecutive audio frames according to a set frame length (e.g., 20ms) and a frame shift length (e.g., 10ms), and analyzing the Mel spectrogram characteristics of each frame.

[0040] S2: Based on the multimodal audio and video characteristics, perform voiceprint recognition and filtering on the audio window to remove non-target speech interference segments and obtain a clean data packet.

[0041] S2.1: Based on the multimodal audio and video features, the audio window is filtered through voiceprint recognition to identify the target speech segment, remove background noise and irrelevant audio, and obtain the audio data packet.

[0042] Furthermore, based on the multimodal audio and video features, voiceprint recognition filtering is performed on the audio window. Audio time segments that match the voiceprint of the target speaker are marked as target speech segments, while unmatched audio segments are regarded as non-target speech interference segments. Interference segments and background noise are removed from the audio stream, and the audio content corresponding to the target speech segment is retained to obtain audio data packets.

[0043] It should be noted that multimodal audio and video features refer to the combination of extracting Mel spectrogram features from audio signals and lip movement features from video frames to form a feature package containing both audio and video information.

[0044] Matching and non-matching are determined by calculating the similarity between the voiceprint features extracted from the audio window and the voiceprint template of the target speaker. When the similarity (using cosine similarity) between the voiceprint features in the audio window and the voiceprint template of the target speaker exceeds a preset threshold (e.g., 0.8), the audio window matches the target speaker and is marked as the target speech segment. When the similarity is below the threshold, the audio window does not match the target speaker and is marked as a non-target speech interference segment.

[0045] Voiceprint recognition refers to the process of analyzing audio signals, extracting voiceprint features (laryngeal structure, accent, and intonation), comparing them with the voiceprint template of the target speaker, filtering out audio segments that match the target speaker, eliminating interference segments that are not part of the target speech, and obtaining clear speech data.

[0046] The target speaker refers to the speaker who is pre-registered during the audio-visual synchronization detection process.

[0047] S2.2: Based on audio data packets, the audio segments are filtered and processed using speech feature analysis methods. By extracting speech features, background noise and non-target speech interference segments are removed to obtain speech data.

[0048] Furthermore, based on the audio data packet, the audio segments in the audio data packet are extracted by frame segmentation using the speech feature analysis method. The time-frequency distribution characteristics are distinguished from speech activities and non-speech components by frequency domain analysis. The audio parts that do not have the target speech features are divided into background noise and non-target speech interference segments and removed. The speech content that conforms to the speech feature analysis method is retained to obtain speech data.

[0049] Specifically, speech feature analysis methods refer to the process of performing frame-by-frame and frequency domain analysis on audio signals, generating Mel spectrograms, and extracting spectral features and speaker features to distinguish target speech from background noise and non-target speech interference segments.

[0050] Audio segments lacking target speech features refer to segments that, after comparison using voiceprint recognition and speech feature analysis methods, do not match the pre-registered target speaker's voiceprint template in terms of spectral morphology, voiceprint features, and temporal variations.

[0051] S2.3: Perform synchronization verification on the audio data packet and the voice data to obtain a clean data packet.

[0052] Furthermore, by aligning the audio data packets and speech data through synchronous verification, the consistency of the target speech segment in the audio data packet with the speech features extracted from the speech data on the time axis is verified. Segments with time mismatch and inconsistency are eliminated, and the aligned audio content is retained, thereby obtaining a clean data packet.

[0053] It should be noted that synchronization verification refers to the deviation analysis and alignment correction of the time series of audio data packets and voice data, and the evaluation and confirmation of the consistency of audio and video status in each time period.

[0054] A time mismatch or inconsistency segment refers to a segment where the absolute value of the audio's time offset relative to the video exceeds an acceptable delay range. For example, a delay of no more than approximately 30 milliseconds is acceptable, while a time interval exceeding this range is considered a mismatch or inconsistency segment.

[0055] S3: Based on the clean data packet, the SyncNet deep dual-stream network is used to determine synchronization, extracting lip movement features and speech features respectively, and generating synchronization determination result packets.

[0056] S3.1: SyncNet deep dual-stream network refers to a network that detects the synchronization between audio and video signals through feature encoding methods and timing alignment algorithms.

[0057] It should be noted that the SyncNet deep dual-stream network processes both audio and video streams simultaneously to determine the synchronization between them. In SyncNet, "dual-stream" refers to the simultaneous processing of both audio and video information streams, extracting features from each through independent neural networks, and then combining these extracted features to evaluate synchronization.

[0058] Specifically, the video stream extracts lip movement features from video frames and processes them using feature encoding methods. The input to the video stream is each frame of video image. The feature encoding method extracts lip movement features from the image, focuses on the lip region in the image, and captures the movement of the lips.

[0059] Audio streams extract speech features from audio signals. The audio signals are processed by feature coding methods for deep learning. The audio stream uses the spectral features of each audio segment to perform deep learning on the speech in the audio signal, including the pitch, volume, and rate of speech.

[0060] It should be noted that training the SyncNet deep two-stream network includes data preparation, feature extraction, temporal alignment, synchronization determination, loss function, backpropagation, and optimization.

[0061] Data preparation requires a large amount of audio and video data for training the SyncNet deep two-stream network. Each audio and video data includes a video segment and the corresponding audio. In each video frame, the time point of the pronunciation and the corresponding speech information are marked. Training the SyncNet deep two-stream network with audio and video data can learn the temporal relationship between audio and video.

[0062] Feature extraction is performed during the training of the SyncNet deep dual-stream network. Video and audio are input into the video stream and audio stream, respectively. The video stream uses feature encoding to extract lip movement features, and the audio stream uses feature encoding to extract audio features. For the video stream, lip movement features are extracted from each frame of the image. These lip movement features are used to determine clear articulation actions in the video. For the audio stream, the audio signal is input, and the spectral features of the audio are extracted using feature encoding. The SyncNet deep dual-stream network is then used to extract the time-series features of the speech.

[0063] Temporal alignment is a component of the SyncNet deep dual-stream network. It evaluates synchronization by aligning the features of the audio and video streams. The temporal alignment algorithm finds the temporal relationship between matching audio and video frames by calculating the similarity (cosine similarity) between audio and video features.

[0064] Synchronization determination: The SyncNet deep dual-stream network calculates the matching degree between the features of the audio stream and the video stream to obtain the synchronization determination result; the SyncNet deep dual-stream network outputs a synchronization score.

[0065] The loss function is used to guide network optimization during the training of the SyncNet deep two-stream network. The loss function is based on the difference between the synchronization score and the true label (synchronized and desynchronized). By optimizing the loss function, the gap between predicted synchronization and actual synchronization is quantified, enabling the SyncNet deep two-stream network to evaluate the synchronization of audio and video.

[0066] During backpropagation and optimization training, the SyncNet deep two-stream network gradually improves its accuracy in judging audio and video synchronization through multiple iterations. In each forward propagation, the SyncNet deep two-stream network calculates synchronization scores based on the features of the audio and video streams and compares them with the true labels (synchronized and desynchronized). By calculating the loss function, the SyncNet deep two-stream network can quantify the difference between the predicted synchronization scores and the true labels. In the backpropagation stage, the SyncNet deep two-stream network uses the gradient information of the loss function to adjust the parameters of each layer and gradually optimize the feature extraction method and temporal alignment algorithm. During training, the SyncNet deep two-stream network continuously improves its feature extraction capability and synchronization judgment capability, thereby enhancing the detection of synchronization between audio and video signals.

[0067] S3.2: The lip region and audio segments in the video stream are jointly encoded using a feature encoding method to obtain a multimodal feature sequence.

[0068] Furthermore, a feature encoding method is used to encode the lip region and the corresponding audio segment in the video stream. The continuous video frames and audio segments of the lip region are encoded separately. The lip movement features obtained by face detection and feature extraction are used as input for the lip region. The audio segments are generated by generating Mel spectrograms and extracting spectral features as input. After aligning the lip movement features and spectral features in the time dimension, the feature encoding method is used for joint encoding to form a multimodal feature sequence containing lip movement features and speech features.

[0069] It should be noted that the feature encoding method refers to the process of using the SyncNet deep dual-stream network to extract spatiotemporal features from lip movement video sequences, extracting spectral features from the Mel spectrogram, and mapping lip movement features and speech features to a shared embedding space for synchronization discrimination.

[0070] The time dimension refers to the chronological order and positional relationship of audio signals and video frame sequences on the time axis, describing the arrangement and alignment of lip movement features and speech features over time.

[0071] S3.3: Based on the multimodal feature sequence, a temporal alignment algorithm is used to calculate the correlation and synchronicity of lip movement features and speech features in each time window to obtain a frame-level synchronicity score sequence.

[0072] Furthermore, the lip movement features and speech features corresponding to the same time window in the multimodal feature sequence are input into the temporal alignment algorithm. By calculating the alignment degree between the lip movement features and speech features in the time dimension, the correlation degree within the time window is obtained. Based on the correlation degree, the synchronization of each time window is evaluated, and a value representing the degree of synchronization matching is generated. The matching degree values ​​of all time windows are arranged in chronological order to form a frame-level synchronization score sequence.

[0073] Synchronization score, expressed as: ; in, Indicates the first Synchronicity score for a time period, indicating the synchronicity score within the time period. The degree of audio-visual synchronization within the device; the higher the score, the better the audio-visual synchronization within a given time period. It is a time period The number of frames in the time interval indicates the number of frames contained within each time interval; It is the first The synchronization difference of a frame indicates the degree of synchronization difference in a frame. The degree of matching between audio and video; the smaller the value, the better the frame synchronization. It refers to a time period after the time window of audio and video data is divided; It refers to the index of each frame in the audio and video sequence within a time window, the first... The frame is an abnormal synchronization frame.

[0074] Specifically, a time window refers to a continuous interval extracted from a multimodal feature sequence according to a fixed duration, and the same continuous interval simultaneously contains lip movement features and speech features.

[0075] Timing alignment algorithms refer to the deviation analysis and alignment correction of the time series of audio tracks and video tracks to obtain the degree of alignment between audio segments and video frames on the time axis.

[0076] Alignment level refers to the cosine similarity and corresponding offset between lip movement features and spectral features, reflecting the temporal synchronization between the audio track and the video track. Alignment level is quantified by cosine similarity and temporal offset. The cosine similarity value ranges from 0 to 1. When the cosine similarity is not less than 0.75 and the absolute value of the audio and video temporal offset does not exceed 30 milliseconds, the alignment level is considered to meet the synchronization requirements.

[0077] The degree of association refers to the measurement of lip movement features and speech features in a shared embedding space within the same time window.

[0078] Synchronization assessment refers to evaluating the consistency between lip movement features and speech features within each time window based on multimodal feature sequences, using feature encoding methods and temporal alignment algorithms, to determine the synchronization status of audio and video over a time period. For example, in a video, when a speaker pronounces "hello," the speech features in the audio signal should correspond precisely to that pronunciation time point. Feature encoding methods are used to extract lip movement features from the video and compare them with the speech features in the audio signal; temporal alignment algorithms are used to align the time windows of the audio and video, checking for lip movement and speech synchronization. When the timing of lip movement and speech features is consistent, audio and video are considered synchronized; when discrepancies exist (speech precedence and lag), it is considered asynchronous.

[0079] The matching degree value refers to the cosine similarity calculated from the lip movement feature and the speech feature. It quantifies the synchronization of audio and video within a time window. The value ranges from -1 to 1. The larger the value, the higher the matching degree. When the matching degree value is greater than or equal to 0.75, it can be used as a reference value for audio and video synchronization.

[0080] S3.4: Determine the synchronization status of the frame-level synchronization scoring sequence and obtain the synchronization discrimination result packet.

[0081] Furthermore, based on the frame-level synchronization score sequence, and according to the distribution characteristics of the synchronization score in the time dimension, the synchronization status of each frame is determined. Frames whose synchronization scores meet the synchronization conditions are marked as synchronized, and frames whose synchronization scores do not meet the synchronization conditions are marked as asynchronous. The synchronization status marks of all frames are encapsulated with the corresponding timestamps to obtain a synchronization discrimination result package.

[0082] It should be noted that the synchronicity score is obtained by calculating the audio window cosine similarity between lip movement features and spectral features in the shared embedding space using the SyncNet deep two-stream network.

[0083] S4: Based on the synchronous discrimination result packet, identify abnormal frames in audio and video, and combine lip movement amplitude and audio energy to determine the abnormal type and obtain structured abnormal logs.

[0084] S4.1: The audio energy and lip movement amplitude are aligned using the timestamp anchoring method to obtain a multimodal frame-level feature set.

[0085] Furthermore, a timestamp anchoring method is used to align the audio energy calculation results with the lip movement amplitude extracted from the corresponding video frame frame by frame according to the timestamp. Abnormal frames in the audio and video are identified, so that the audio energy and lip movement amplitude at each time point are correlated, and a multimodal frame-level feature set of time-aligned audio energy and lip movement amplitude is obtained.

[0086] The abnormal frame screening calculation is expressed as follows: ; in, It is the first Abnormal frame scoring of frames It is the first The amplitude of lip movements in a frame. Is with the first The amplitude of lip movements in a frame, for example, the first frame. A frame is a video lip-syncing action (the clear lip movements of the mouth). It will be higher because it contributes significantly to synchronization. It is the first The audio energy of a frame The number of all frames It is the index of the synchronization score frame.

[0087] Specifically, timestamp information refers to the time stamps used to mark the corresponding positions of each audio segment and video frame in the time series of the audio and video tracks.

[0088] The correspondence between audio energy and lip movement amplitude refers to pairing the audio energy at the same time point with the lip movement amplitude in the corresponding video frame using timestamp information on the timeline.

[0089] S4.2: Based on the abnormal inter-frame differences in lip movement amplitude and audio energy, perform abnormal frame screening on the multimodal frame-level feature set to obtain a multimodal abnormal frame set.

[0090] Furthermore, based on the abnormal inter-frame differences in lip movement amplitude and audio energy, the time-aligned multimodal frame-level feature set of audio energy and lip movement amplitude is used to screen for abnormal frames. The screening obtains the variation differences between lip movement amplitude and audio energy, and frames that do not match lip movement amplitude and audio energy are identified based on the variation differences. These mismatched frames are marked as abnormal frames, and all abnormal frames are collected to form a multimodal abnormal frame set.

[0091] It should be noted that the difference in variation refers to the comparison of the magnitude and trend of lip movement amplitude and audio energy on the same time axis. For example, there are asynchronous states where the lip movement amplitude increases or decreases significantly while the audio energy remains unchanged or changes in the opposite direction.

[0092] S4.3: Perform structured reorganization on the multimodal abnormal frame set to obtain a structured abnormal log for audio and video position and synchronization discrimination.

[0093] Furthermore, based on the lip movement amplitude, audio energy, and timing information corresponding to each abnormal frame in the multimodal abnormal frame set, and according to the positional relationship between the audio segment and the video frame obtained by the timestamp anchoring method, the abnormal frames are organized in chronological order into record items containing frame position, lip movement state, audio state, and synchronization discrimination results. The record items are uniformly arranged through a structured format to obtain a structured abnormal log of audio and video position and synchronization discrimination.

[0094] Specifically, frame position refers to the time point and sequence number of a video frame based on its timestamp information.

[0095] Lip movement status refers to whether the lips are in the process of opening and closing, and the magnitude of the opening and closing motion in an abnormal frame.

[0096] Audio state refers to the energy strength of the audio segment corresponding to the abnormal frame in speech features.

[0097] Synchronization discrimination result refers to the judgment information that the audio and lip movement are synchronized and marked as abnormal at the time point of the abnormal frame.

[0098] Record items refer to independent fields in structured exception logs that record frame position, lip movement status, audio status, and synchronization judgment results, providing a complete description of the basic information of audio and video synchronization status at a given time point.

[0099] S5: Encapsulate and integrate the abnormal logs and synchronization judgment result packets to generate a detection report.

[0100] S5.1: The structured anomaly logs and the synchronization discrimination result package are correlated and integrated to obtain the anomaly event synchronization dataset.

[0101] Furthermore, based on the audio and video locations and synchronization discrimination results recorded in the structured anomaly log, the frame-level synchronization score sequences contained in the synchronization discrimination result package are timestamped, and the anomaly events at the same time location are associated and integrated with the synchronization status data to form an anomaly event synchronization dataset containing the time of anomaly occurrence, lip movement amplitude, audio energy, anomaly type label, and synchronization discrimination results.

[0102] It should be noted that timestamp alignment refers to using a timestamp anchoring method to time-stamp and offset the audio segments in the synchronization discrimination result packet with the corresponding video frames.

[0103] S5.2: Based on the statistical summary of the abnormal event synchronization dataset, the lip movement amplitude, audio energy and abnormality type label of each abnormal event are aggregated and organized to obtain the statistical summary data of abnormal events.

[0104] Furthermore, based on the abnormal event synchronization dataset, the lip movement amplitude, audio energy, and abnormality type label corresponding to each abnormal event are extracted item by item. The abnormal events are then categorized and summarized according to their temporal order (the order in which audio and video frames and audio segments are arranged chronologically) and category attributes (the synchronization status markers and classification information of abnormal frames and audio segments and abnormality type labels). The lip movement amplitude and audio energy of abnormal events of the same type and in adjacent time periods are centrally arranged and statistically analyzed to form a summary record of abnormal event location, lip movement amplitude range, audio energy range, and abnormality type label, thus obtaining the statistical summary data of abnormal events.

[0105] It should be noted that item-by-item extraction refers to performing face detection, lip cropping, and Mel-spectral feature extraction on each frame of video image and corresponding audio segment in chronological order, with each time segment obtaining corresponding lip movement features and speech features.

[0106] S5.3: Analyze and classify the statistical summary data of abnormal events, mine the data by combining audio and video features, organize the data by paragraphs, tables and timelines, and generate a detection report.

[0107] Furthermore, based on the lip movement amplitude, audio energy, and anomaly type label of each anomaly event included in the statistical summary data, the anomaly events are classified and analyzed. At the same time, the corresponding audio and video feature information in the multimodal audio and video features is used for correlation mining. The analysis results are divided into paragraphs according to the time sequence, and the statistical summary data of anomaly events is presented in a structured table format. The time of anomaly occurrence is aligned with the audio and video timeline to form a visual organizational structure with time coordinates, and a detection report is generated.

[0108] Specifically, the analysis results refer to the classification and analysis of the lip movement amplitude, audio energy, and abnormality type label of each abnormal event in the statistical summary data of abnormal events, and the correlation mining obtained by combining the corresponding audio and video feature information in the multimodal audio and video features. The results are conclusive data reflecting the relationship between abnormal event features and audio and video time (structured data that directly presents audio and video synchronization anomalies in the detection report).

[0109] Association mining refers to analyzing audio and video features based on statistical summary data of abnormal events, combined with lip movement amplitude, audio energy, and abnormal type labels, to uncover the inherent correspondence between audio and video states in different time periods.

[0110] Correspondence refers to the one-to-one correspondence and matching of lip movement amplitude, audio energy, and abnormality type in the same time period.

[0111] A detection report is a document that organizes the abnormal event synchronization dataset and statistical summary data in the form of paragraphs, tables and timelines after the data is encapsulated and integrated based on the abnormal logs and synchronization discrimination results. It centrally presents the lip movement amplitude, audio energy, abnormal type labels and audio-visual synchronization judgment results for each time period.

[0112] S6: Based on the test report, the timing alignment algorithm is used to verify the audio and video synchronization, and the result is solidified by timestamp anchoring to obtain the synchronization judgment result.

[0113] S6.1: Based on the detection report, a time-series alignment algorithm is used to perform deviation analysis and alignment correction on the time series of the audio track and the video track, and an alignment result set is obtained.

[0114] Furthermore, based on the synchronization discrimination result packet and structured anomaly log contained in the detection report, the time series information corresponding to the audio track and the video track is extracted. The temporal alignment algorithm is used to align the temporal relationship between the synchronization discrimination result packet and the structured anomaly log frame by frame to obtain the time offset between the audio segment and the video frame. Based on the time offset, the audio track or video track is adjusted and aligned to obtain the alignment result set.

[0115] It should be noted that sliding adjustment refers to shifting the audio and video tracks forward and backward on the time axis based on the obtained time offset. By continuously shifting the positional relationship between audio segments and video frames, the corresponding audio segments and video frames are realigned in the time sequence.

[0116] S6.2: By aligning the result set, the positional relationship between the audio segment and the video frame is time-stamped and offset-corrected using the timestamp anchoring method, and the anchoring result package with time coordinates and synchronization status marks is obtained.

[0117] Furthermore, based on the time-series alignment information of the audio and video tracks contained in the alignment result set, the time coordinates of each audio segment and the corresponding video frame on the time axis are determined using the timestamp anchoring method. The time offset is corrected for the positional relationship with the synchronization status, and the corrected time coordinates are fixedly bound to the synchronization status marker to form an anchoring result package of time coordinates and synchronization status marker.

[0118] It should be noted that, fixed binding refers to, after completing the time offset correction of audio segments and video frames, fixing and recording the corrected time coordinates and synchronization status markers on the time axis, and organizing them into an anchoring result package for verification and output.

[0119] S6.3: Perform synchronization verification on the anchoring result packet, evaluate the consistency of audio and video status in each time period, confirm abnormal situations, and obtain synchronization judgment results.

[0120] Furthermore, based on the time coordinates and synchronization status markers contained in the anchoring result package, the synchronization of the correspondence between the audio track and the video track in each time period is checked. The positional relationship between the audio segment and the video frame is aligned by the time coordinates. In the case of audio and video asynchrony, the consistency of audio and video status in each time period is evaluated according to the synchronization status markers. Time periods with inconsistencies are marked as abnormal situations. The verification results of all time periods are integrated to obtain the synchronization judgment result.

[0121] It should be noted that synchronization verification refers to using a time-series alignment algorithm to analyze and compare the time relationship between the audio track and the video track in each time period within the anchored result packet, to determine the consistency of the audio and video states on the entire time axis and to identify any abnormalities.

[0122] Synchronization determination results refer to the final determination information output by the deep learning-based audio-video synchronization detection method. The synchronization determination results include a qualitative conclusion on the overall synchronization status of audio and video, detailed identification of whether audio and video are synchronized or not in each time period, specific time coordinates of abnormal situations, corresponding abnormal types, and the matching status of lip movement amplitude and audio energy. The synchronization determination results solidify the positional relationship between audio segments and video frames through a timestamp anchoring method, clearly indicating the synchronization consistency of audio and video streams in each time segment in the form of time coordinates and synchronization status markers, and performing structured annotation on identified abnormal frames and abnormal time periods to form a complete synchronization assessment conclusion that is traceable, verifiable, and localizable.

[0123] This embodiment also provides a deep learning-based audio-video synchronization detection system, including: The template is extracted by separating the video stream from the audio stream of the collected audio and video files to be detected, and multimodal audio and video features are generated by face detection and feature extraction. The template is filtered and removed. Based on the multimodal audio and video features, the audio window is filtered by voiceprint recognition to remove non-target speech interference segments and obtain a clean data packet. The template is extracted and discriminated. Based on the clean data packet, the SyncNet deep dual-stream network is used to perform synchronization discrimination, extracting lip movement features and speech features respectively, and generating synchronization discrimination result packets. The identification and judgment template is used to identify abnormal frames in audio and video based on the synchronous discrimination result package. The abnormality type is determined by combining the lip movement amplitude and audio energy to obtain a structured abnormality log. The data integration template encapsulates and integrates anomaly logs and synchronization judgment result packages to generate a detection report. The template is validated and solidified. Based on the test report, the timing alignment algorithm is used to verify the audio and video synchronization. The result is solidified by timestamp anchoring to obtain the synchronization judgment result.

[0124] In summary, this invention achieves efficient extraction and synchronization determination of audio / video lip movement features and speech features by employing the SyncNet deep dual-stream network for synchronization discrimination. This helps to accurately identify the matching degree between target speech and lip movements, improving the accuracy of synchronization detection. Furthermore, by encapsulating and integrating the synchronization discrimination result package, it achieves comprehensive recording and structured management of the synchronization discrimination results. This helps to generate high-quality detection reports, providing data support for verification and evidence preservation, and ensuring the high credibility and judicial acceptance of the results.

[0125] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A deep learning-based audio-video synchronization detection method, characterized in that: include, The collected audio and video files to be detected are separated into video and audio streams, and multimodal audio and video features are generated using face detection and feature extraction. Based on the multimodal audio and video features, the audio window is filtered using voiceprint recognition to remove non-target speech interference segments and obtain a clean data packet. The specific steps are as follows: Based on multimodal audio and video features, the audio window is filtered through voiceprint recognition to identify the target speech segment, remove background noise and irrelevant audio, and obtain audio data packets; Based on audio data packets, a speech feature analysis method is used to filter and process audio segments. By extracting speech features and removing background noise and non-target speech interference segments, speech data is obtained. The audio data packet and voice data are synchronized and verified to obtain a clean data packet; Based on the clean data packet, the SyncNet deep dual-stream network is used for synchronization discrimination. Lip movement features and speech features are extracted separately to generate a synchronization discrimination result packet. The specific steps are as follows: A feature encoding method is used to jointly encode the lip region and audio segments in the video stream to obtain a multimodal feature sequence; Based on multimodal feature sequences, a temporal alignment algorithm is used to calculate the correlation and synchronicity of lip movement features and speech features within each time window, and obtain a frame-level synchronicity score sequence. The synchronization status of the frame-level synchronization scoring sequence is determined to obtain the synchronization discrimination result packet. Based on the synchronous discrimination result packet, abnormal frames in audio and video are identified, and the abnormality type is determined by combining lip movement amplitude and audio energy to obtain a structured abnormality log. The abnormal logs and synchronization judgment result packets are encapsulated and integrated to generate a detection report; According to the test report, the timing alignment algorithm was used to verify the synchronization of audio and video, and the result was solidified by timestamp anchoring to obtain the synchronization judgment result.

2. The deep learning-based audio-video synchronization detection method as described in claim 1, characterized in that: The audio and video files to be detected refer to the original audio and video data for audio and video synchronization detection.

3. The deep learning-based audio-video synchronization detection method as described in claim 2, characterized in that: The process involves separating the video and audio streams of the collected audio and video files to be detected, and generating multimodal audio and video features using face detection and feature extraction. The specific steps are as follows: For the audio and video files to be tested, the video stream and audio stream are separated to obtain video frames and audio signals; Based on video frames, face detection is performed on each frame of the image and the lip region is cropped to obtain lip movement features; By analyzing the audio signal, a Mel spectrogram is generated and spectral features are extracted. Combining lip movement features and spectral features, multimodal audio and video features are generated.

4. The deep learning-based audio-video synchronization detection method as described in claim 1, characterized in that: The SyncNet deep dual-stream network refers to the method of detecting the synchronization between audio and video signals through feature encoding and timing alignment algorithms.

5. The deep learning-based audio-video synchronization detection method as described in claim 1, characterized in that: The process involves identifying abnormal frames in the audio and video based on the synchronization discrimination result packet, and determining the anomaly type by combining lip movement amplitude and audio energy to obtain a structured anomaly log. The specific steps are as follows. A timestamp anchoring method is used to align audio energy with lip movement amplitude to obtain a multimodal frame-level feature set; Based on the abnormal inter-frame differences in lip movement amplitude and audio energy, abnormal frames are screened in the multimodal frame-level feature set to obtain a multimodal abnormal frame set. The multimodal abnormal frame set is restructured to obtain a structured abnormal log for audio and video location and synchronization discrimination.

6. The deep learning-based audio-video synchronization detection method as described in claim 5, characterized in that: The specific steps for encapsulating and integrating the anomaly logs and synchronization judgment result packets to generate a detection report are as follows. The structured anomaly logs and the synchronization discrimination result packets are correlated and integrated to obtain the anomaly event synchronization dataset; Based on the statistical induction of the abnormal event synchronization dataset, the lip movement amplitude, audio energy and abnormality type label of each abnormal event are aggregated and organized to obtain the statistical summary data of abnormal events. The statistical summary data of abnormal events is analyzed and classified, and the data is mined by combining audio and video features. The data is then organized by paragraphs, tables, and timelines to generate detection reports.

7. The deep learning-based audio-video synchronization detection method as described in claim 6, characterized in that: Based on the detection report, a time-series alignment algorithm is used to verify the audio-video synchronization, and the result is solidified by timestamp anchoring to obtain the synchronization judgment result. The specific steps are as follows. Based on the detection report, a time-series alignment algorithm is used to perform deviation analysis and alignment correction on the time series of the audio track and the video track, and an alignment result set is obtained. By aligning the result set, the positional relationship between audio segments and video frames is fixed by time stamping and offset correction using the timestamp anchoring method, and the anchoring result package with time coordinates and synchronization status marks is obtained. The synchronization of the anchoring result packet is checked to evaluate the consistency of audio and video status in different time periods, identify abnormal situations, and obtain the synchronization judgment result.

8. A deep learning-based audio-video synchronization detection system, based on the deep learning-based audio-video synchronization detection method according to any one of claims 1 to 7, characterized in that: include, The template is extracted by separating the video stream from the audio stream of the collected audio and video files to be detected, and multimodal audio and video features are generated by face detection and feature extraction. The template is filtered and removed. Based on the multimodal audio and video features, the audio window is filtered by voiceprint recognition to remove non-target speech interference segments and obtain a clean data packet. The template is extracted and discriminated. Based on the clean data packet, the SyncNet deep dual-stream network is used to perform synchronization discrimination, extracting lip movement features and speech features respectively, and generating synchronization discrimination result packets. The identification and judgment template is used to identify abnormal frames in audio and video based on the synchronous discrimination result package. The abnormality type is determined by combining the lip movement amplitude and audio energy to obtain a structured abnormality log. The data integration template encapsulates and integrates anomaly logs and synchronization judgment result packages to generate a detection report. The template is validated and solidified. Based on the test report, the timing alignment algorithm is used to verify the audio and video synchronization. The result is solidified by timestamp anchoring to obtain the synchronization judgment result.