Video detection method and video detection system

By detecting the wrong frequency frames in the video and counting the wrong frequency frame rate, the problem of lack of virtual character video detection technology in the prior art is solved, and the effective detection of virtual character videos and the effect of preventing abuse is achieved.

CN120014519APending Publication Date: 2025-05-16SHANGHAI TONGSHI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111531.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology lacks the detection technology of virtual character videos, which leads to the malicious use of virtual human generation technology, causing ethical and security issues.

Method used

By detecting the wrong frequency frames in the video, counting the number of wrong frequency frames, and determining whether the video is a virtual character video based on the wrong frequency frame rate. The method includes establishing a fault frequency detection model based on audio features and facial features of the character, and performing fault frequency detection and statistics on the video through the model.

Benefits of technology

The detection of virtual character videos is realized, the abuse of virtual character videos is avoided, and the ethics and security of online audio and video interaction is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014519A_ABST
    Figure CN120014519A_ABST
Patent Text Reader

Abstract

The invention provides a video detection method and system, and the method comprises the steps: detecting wrong frequency frames in a video, carrying out the statistics of the wrong frequency frames, judging whether the video is a virtual character video or not according to the statistical result of the wrong frequency frames, and achieving the detection of the virtual character video. And abuse of the virtual character video is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video detection technology, and in particular to a video detection method and a video detection system. Background Art

[0002] With the continuous advancement of virtual human generation technology, digital humans and simulated humans have been widely used in entertainment, education, medical care and other fields.

[0003] The publicly available virtual human generation technologies, such as StyleGAN face generation technology, DeepFake video generation technology, and Wav2Lip audio and video synchronization generation technology, mostly focus on the optimization of voice-driven facial expression generation algorithms, aiming to improve naturalness and synchronization, thereby improving the simulation of virtual characters.

[0004] However, there is currently a lack of detection technology for virtual character videos, which has led to the malicious use of virtual human generation technology, causing ethical and safety issues when conducting online audio and video interactions through virtual simulated characters.

[0005] Therefore, it is necessary to provide a new video detection method and a video detection system to solve the above problems existing in the prior art. Summary of the invention

[0006] The purpose of the present invention is to provide a video detection method and a video detection system to realize the detection of virtual character videos.

[0007] To achieve the above object, the video detection method of the present invention comprises the following steps:

[0008] S1: Detecting a frequency error frame in a video, wherein the video includes a plurality of video frames and a plurality of audio frames, the video frames correspond to the audio frames one by one, and when the video frame does not correspond to the audio frame, it is determined to be the frequency error frame;

[0009] S2: Counting the error frames, and then judging whether the video is a virtual character video according to the statistical result of the error frames.

[0010] The beneficial effect of the video detection method is to detect the error frequency frames in the video, count the error frequency frames, and then judge whether the video is a virtual character video based on the statistical results of the error frequency frames, thereby realizing the detection of virtual character videos and avoiding the abuse of virtual character videos.

[0011] Optionally, in step S1, detecting a frequency error frame in a video includes:

[0012] The video is subjected to frequency error detection by using a frequency error detection model to obtain frequency error frames in the video.

[0013] Optionally, performing frequency error detection on the video by using a frequency error detection model to obtain frequency error frames in the video includes:

[0014] When the error frame detection model determines that the meaning expressed by the character's lip shape in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, which is the error frame.

[0015] Optionally, when it is determined through the error frame detection model that the meaning expressed by the mouth shape of the character in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, that is, the error frame, including:

[0016] An audio frame stage is defined as a period from a start audio frame to an end audio frame of a certain text. In an audio frame stage, a frequency error frame detection model is used to determine whether the start audio frame corresponds to the mouth shape of a character in the corresponding video frame.

[0017] If it is determined by the error frame detection model that the starting audio frame does not correspond to the mouth shape of the person in the corresponding video frame, then the meaning expressed by the mouth shape of the person in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame;

[0018] Determining whether the end audio frame corresponds to the mouth shape of the character in the corresponding video frame through a frequency error frame detection model;

[0019] If the error frame detection model determines that the end audio frame does not correspond to the lip shape of the character in the corresponding video frame, then the meaning expressed by the lip shape of the character in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame.

[0020] Optionally, when it is determined through the error frame detection model that the meaning expressed by the mouth shape of the character in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, that is, the error frame, including:

[0021] Obtain a video frame when the person's mouth is opened to the maximum when pronouncing a word, and then determine whether the person's mouth shape in the video frame corresponds to the corresponding audio frame through a frequency error frame detection model;

[0022] If the error frame detection model determines that the lip shape of the character in the video frame does not correspond to the corresponding audio frame, then the meaning expressed by the lip shape of the character in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame.

[0023] Optionally, in step S2, counting the frequency error frames, and then judging whether the video is a virtual character video according to the statistical result of the frequency error frames, includes:

[0024] Counting the frequency-error frames to obtain the number of frequency-error frames, and counting the number of all detected video frames to obtain the total number of video frames;

[0025] Calculating the error frequency frame rate according to the number of the error frequency frames and the total number of the video frames;

[0026] It is determined whether the video is a virtual character video according to the error frequency.

[0027] Optionally, the performing statistics on the frequency error frames includes:

[0028] The frequency error frames are counted by using an offset statistics method.

[0029] Optionally, in step S2, counting the frequency error frames, and then judging whether the video is a virtual character video according to the statistical result of the frequency error frames, includes:

[0030] Counting the frequency error frames to obtain the number of offset frames and the number of frequency error frames of the frequency error frames;

[0031] Whether the video is a virtual character video is determined according to the offset frame number and the error frame number of all the error frames.

[0032] Optionally, the video detection method further includes a frequency error detection model training step, and the frequency error detection model training step includes:

[0033] S21: Establish a training error frequency detection model based on audio features and facial features of characters;

[0034] S22: Training the to-be-trained frequency error detection model through database training data to obtain a frequency error detection model, wherein the database training data includes a plurality of audio features and a plurality of facial features of a person matching the audio features.

[0035] The present invention also provides a video detection system, comprising a detection unit and a statistical judgment unit, wherein the detection unit is used to detect error frequency frames in a video, wherein the video comprises a plurality of video frames and a plurality of audio frames, and the video frames correspond to the audio frames one-to-one. When the video frames do not correspond to the audio frames, they are judged to be the error frequency frames, and the statistical judgment unit is used to perform statistics on the error frequency frames, and then judge whether the video is a virtual character video based on the statistical results of the error frequency frames.

[0036] The beneficial effect of the video detection system is that: the detection unit detects the error frequency frames in the video, the statistical judgment unit counts the error frequency frames, and then judges whether the video is a virtual character video based on the statistical results of the error frequency frames, thereby realizing the detection of virtual character videos and avoiding the abuse of virtual character videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A flowchart of a video detection method in some embodiments of the present invention;

[0038] Figure 2 A schematic diagram of the timing of coordination of video frames and audio frames in a real video in the prior art;

[0039] Figure 3 Schematic diagram of the timing of coordination of audio frames and video frames in a virtual character video in some embodiments of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with general skills in the field to which the present invention belongs. "Including" and similar words used in this article mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0041] In view of the problems existing in the prior art, an embodiment of the present invention provides a video detection method. Figure 1 , the video detection method comprises the following steps:

[0042] S1: Detecting frequency error frames in a video, wherein the video includes a plurality of video frames and a plurality of audio frames, the video frames correspond to the audio frames one by one, and when the video frames do not correspond to the audio frames, they are determined to be frequency error frames, and both the video frames and the audio frames can be frequency error frames;

[0043] S2: Counting the error frames, and then judging whether the video is a virtual character video according to the statistical result of the error frames.

[0044] Figure 2 FIG. 1 is a schematic diagram of the timing of the coordination of video frames and audio frames in a real video in the prior art. Figure 2, the real video includes video frame A, video frame B, video frame C, video frame D, video frame E, video frame F, video frame G, video frame H, video frame I, video frame J, video frame K, video frame L, video frame M, video frame N, video frame O, video frame P, video frame Q, video frame R, audio frame a, audio frame b, audio frame c, audio frame d, audio frame e, audio frame f, audio frame g, audio frame h, audio frame i, audio frame j, audio frame k, audio frame l, audio frame m, audio frame n, audio frame o, audio frame p, audio frame q and audio frame r.

[0045] Reference Figure 2 , video frame A corresponds to audio frame a, video frame B corresponds to audio frame b, video frame C corresponds to audio frame c, video frame D corresponds to audio frame d, video frame E corresponds to audio frame e, video frame F corresponds to audio frame f, video frame G corresponds to audio frame g, video frame H corresponds to audio frame h, video frame I corresponds to audio frame i, video frame J corresponds to audio frame j, video frame K corresponds to audio frame k, video frame L corresponds to audio frame l, video frame M corresponds to audio frame m, video frame N corresponds to audio frame n, video frame O corresponds to audio frame o, video frame P corresponds to audio frame p, video frame Q corresponds to audio frame q, and video frame R corresponds to audio frame r.

[0046] Figure 3 FIG. 1 is a schematic diagram of the timing of the coordination of audio frames and video frames in a virtual character video in some embodiments of the present invention. Figure 3 , video frame A and video frame B have no corresponding audio frames, video frame C corresponds to audio frame a, video frame D corresponds to audio frame b, video frame E corresponds to audio frame c, video frame F corresponds to audio frame d, video frame G corresponds to audio frame e, video frame H corresponds to audio frame f, video frame I corresponds to audio frame g, video frame J corresponds to audio frame h, video frame K corresponds to audio frame i, video frame L corresponds to audio frame j, video frame M corresponds to audio frame k, video frame N corresponds to audio frame l, video frame O corresponds to audio frame m, video frame P corresponds to audio frame n, video frame Q corresponds to audio frame o, video frame R corresponds to audio frame p, audio frame q and audio frame r have no corresponding video frames.

[0047] Reference Figure 2 It can be clearly seen that when the network is normal, the meaning expressed by the video frame is the same as the meaning expressed by the audio frame, that is, the video frame corresponds to the audio frame one by one.

[0048] Reference Figure 3It can be clearly seen that the meaning expressed by the video frame is not completely different from the meaning expressed by the audio frame, that is, the video frame and the audio frame do not correspond one to one, and there is a frequency error between the video frame and the audio frame. The frequency error frames are counted, and then the video is judged whether it is a virtual character video based on the statistical results of the frequency error frames, so as to detect the virtual character video and avoid the abuse of the virtual character video.

[0049] In some embodiments, in step S1, detecting frequency error frames in a video includes: performing frequency error detection on the video using a frequency error detection model to obtain frequency error frames in the video.

[0050] In some embodiments, a video is subjected to frequency error detection through a frequency error detection model to obtain a frequency error frame in the video, including: when the frequency error frame detection model determines that the meaning expressed by the character's lip shape in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, that is, the frequency error frame.

[0051] In some embodiments, when it is determined through the error frame detection model that the meaning expressed by the mouth shape of the character in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, that is, the error frame, including:

[0052] An audio frame stage is defined as a period from a start audio frame to an end audio frame of a certain text. In an audio frame stage, a frequency error frame detection model is used to determine whether the start audio frame corresponds to the mouth shape of a character in the corresponding video frame.

[0053] If it is determined by the error frame detection model that the starting audio frame does not correspond to the mouth shape of the person in the corresponding video frame, then the meaning expressed by the mouth shape of the person in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame;

[0054] Determining whether the end audio frame corresponds to the mouth shape of the character in the corresponding video frame through a frequency error frame detection model;

[0055] If the error frame detection model determines that the end audio frame does not correspond to the lip shape of the character in the corresponding video frame, then the meaning expressed by the lip shape of the character in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame.

[0056] In some embodiments, when it is determined through the error frame detection model that the meaning expressed by the mouth shape of the character in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, that is, the error frame, including:

[0057] Obtain a video frame when the person's mouth is opened to the maximum when pronouncing a word, and then determine whether the person's mouth shape in the video frame corresponds to the corresponding audio frame through a frequency error frame detection model;

[0058] If the error frame detection model determines that the lip shape of the character in the video frame does not correspond to the corresponding audio frame, then the meaning expressed by the lip shape of the character in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame.

[0059] In some embodiments, in step S2, counting the error frames, and then judging whether the video is a virtual character video according to the statistical results of the error frames, includes:

[0060] Counting the frequency-error frames to obtain the number of frequency-error frames, and counting the number of all detected video frames to obtain the total number of video frames;

[0061] Calculating the error frequency frame rate according to the number of the error frequency frames and the total number of the video frames;

[0062] It is determined whether the video is a virtual character video according to the error frequency.

[0063] In some embodiments, the counting of the frequency error frames includes:

[0064] The frequency error frames are counted by using an offset statistics method.

[0065] In some embodiments, in step S2, counting the error frames, and then judging whether the video is a virtual character video according to the statistical results of the error frames, includes:

[0066] The error frequency frame is counted to obtain the offset frame number and the error frequency frame number of the error frequency frame, wherein the offset frame number is the number of frames that the frame currently corresponding to the current error frequency frame differs from the frame that should have been corresponding to it;

[0067] Whether the video is a virtual character video is determined according to the offset frame number and the error frame number of all the error frames.

[0068] Specifically, the error frequency frames are counted in time periods to obtain the offset frame number and the error frequency frame number of the error frequency frames, and a normal distribution calculation is performed based on the offset frame number, the error frequency frame number and the duration of the time period. The calculation result is compared with a threshold to determine whether the video is a virtual character video.

[0069] In some embodiments, the video detection method further includes a frequency error detection model training step, and the frequency error detection model training step includes:

[0070] S21: Establishing a training error frequency detection model based on audio features and facial features of a person, wherein the audio features include but are not limited to Chinese and English, and the facial features of a person include but are not limited to mouth shape, tongue position, facial muscles, etc.;

[0071] S22: Training the to-be-trained frequency error detection model through database training data to obtain a frequency error detection model, wherein the database training data includes a plurality of audio features and a plurality of facial features of a person matching the audio features.

[0072] The present invention also provides a video detection system for implementing the video detection method, comprising a detection unit and a statistical judgment unit, wherein the detection unit is used to detect error frequency frames in the video, wherein the video comprises a plurality of video frames and a plurality of audio frames, and the video frames correspond to the audio frames one-to-one. When the video frame does not correspond to the audio frame, it is judged to be the error frequency frame, and the statistical judgment unit is used to perform statistics on the error frequency frames, and then judge whether the video is a virtual character video based on the statistical results of the error frequency frames.

[0073] Although the embodiments of the present invention are described in detail above, it is obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and may be implemented or realized in a variety of ways.

Claims

1. A video detection method, characterized in that: The following steps are involved: S1: Detecting a frequency error frame in a video, wherein the video includes a plurality of video frames and a plurality of audio frames, the video frames correspond to the audio frames one by one, and when the video frame does not correspond to the audio frame, it is determined to be the frequency error frame; S2: Counting the error frames, and then judging whether the video is a virtual character video according to the statistical results of the error frames.

2. The video detection method according to claim 1, characterized in that: In step S1, detecting the frequency error frame in the video includes: The video is subjected to frequency error detection by using a frequency error detection model to obtain frequency error frames in the video.

3. The video detection method according to claim 2, characterized in that: Performing frequency error detection on the video by using a frequency error detection model to obtain frequency error frames in the video, including: When the error frame detection model determines that the meaning expressed by the character's lip shape in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, which is the error frame.

4. The video detection method according to claim 3, characterized in that: When it is determined through the error frame detection model that the meaning expressed by the mouth shape of the character in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, that is, the error frame, including: An audio frame stage is defined as a period from a start audio frame to an end audio frame of a certain text. In an audio frame stage, a frequency error frame detection model is used to determine whether the start audio frame corresponds to the mouth shape of a character in the corresponding video frame. If it is determined by the error frame detection model that the starting audio frame does not correspond to the mouth shape of the person in the corresponding video frame, then the meaning expressed by the mouth shape of the person in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame; Determining whether the end audio frame corresponds to the mouth shape of the character in the corresponding video frame through a frequency error frame detection model; If the error frame detection model determines that the end audio frame does not correspond to the lip shape of the character in the corresponding video frame, then the meaning expressed by the lip shape of the character in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame.

5. The video detection method according to claim 3, characterized in that: When it is determined through the error frame detection model that the meaning expressed by the mouth shape of the character in the video frame is different from the meaning expressed by the audio frame, it is determined that the video frame does not correspond to the audio frame, that is, the error frame, including: Obtain a video frame when the person's mouth is opened to the maximum when pronouncing a word, and then determine whether the person's mouth shape in the video frame corresponds to the corresponding audio frame through a frequency error frame detection model; If the error frame detection model determines that the lip shape of the character in the video frame does not correspond to the corresponding audio frame, then the meaning expressed by the lip shape of the character in the video frame is different from the meaning expressed by the audio frame, and it is determined that the video frame does not correspond to the audio frame.

6. The video detection method according to claim 1, characterized in that: In step S2, statistics are collected on the error frames, and then it is determined whether the video is a virtual character video according to the statistical results of the error frames, including: Counting the frequency-error frames to obtain the number of frequency-error frames, and counting the number of all detected video frames to obtain the total number of video frames; Calculating the error frequency frame rate according to the number of the error frequency frames and the total number of the video frames; It is determined whether the video is a virtual character video according to the error frequency.

7. The video detection method according to claim 6, characterized in that: The counting of the frequency error frames includes: The frequency error frames are counted by using an offset statistics method.

8. The video detection method according to claim 1, characterized in that: In step S2, statistics are collected on the error frames, and then it is determined whether the video is a virtual character video according to the statistical results of the error frames, including: Counting the frequency error frames to obtain the number of offset frames and the number of frequency error frames of the frequency error frames; Whether the video is a virtual character video is determined according to the offset frame number and the error frame number of all the error frames.

9. The video detection method according to claim 1, characterized in that: The method further includes a frequency error detection model training step, wherein the frequency error detection model training step includes: S21: Establish a training error frequency detection model based on audio features and facial features of characters; S22: Training the to-be-trained frequency error detection model through database training data to obtain a frequency error detection model, wherein the database training data includes a plurality of audio features and a plurality of facial features of a person matching the audio features.

10. A video detection system, characterized in that: It includes a detection unit and a statistical judgment unit, the detection unit is used to detect error frames in a video, wherein the video includes a plurality of video frames and a plurality of audio frames, the video frames correspond to the audio frames one by one, when the video frames do not correspond to the audio frames, they are judged to be the error frames, the statistical judgment unit is used to perform statistics on the error frames, and then judge whether the video is a virtual character video based on the statistical results of the error frames.