Deepfake video detection method and device, and storage medium

By deploying multiple face recognition models and fake video detection methods in the face recognition system, and combining facial action recognition and coarse and fine detection strategies, the problem of the single method of network video fake detection is solved, and efficient and accurate fake video detection is achieved.

CN116645710BActive Publication Date: 2025-12-05BEIJING REALAI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211206017.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-12-05
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

In existing technologies, the methods for detecting forgery in online videos are limited and cannot flexibly cope with various forgery methods, resulting in low detection efficiency and easy missed detections. It is also impossible to efficiently and accurately identify forgery traces in diverse online videos.

Method used

By deploying multiple face recognition models and fake video detection methods in the face recognition system, the appropriate detection method can be dynamically switched by recognizing the facial movements of the target face. Combining coarse and fine detection strategies can improve the flexibility and accuracy of detection and avoid missed detections.

Benefits of technology

It enables efficient and accurate detection of forged videos when faced with diverse forgery methods, improving detection efficiency and coverage, avoiding missed detections, and is applicable to diverse video scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645710B_ABST
    Figure CN116645710B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the technical field of image processing, and provides a deep fake video detection method and device and a storage medium, the method comprises the following steps: obtaining a first to-be-detected video from at least one video source party, identifying a first face action of a target face in at least one first face image in the first to-be-detected video; determining a first detection mode from a plurality of preset detection modes based on the first face action; calling the first detection mode to identify the true or false of the target face in the to-be-detected video, and outputting a first detection result. When facing different fake videos for detection, the scheme can be switched to the adaptive fake video detection mode at will, especially when facing network videos with various fake modes, the scheme can be dynamically switched to the true or false detection mode adaptive to the current to-be-detected video, at least one network video adopting at least one fake mode can be detected. The detection flexibility is high, the scheme is suitable for diversified videos, the test coverage is wide, and the scheme can effectively avoid missed detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of image processing, and in particular to a deep fake video detection method and device and a storage medium. BACKGROUND

[0002] Based on black production tools, it is easy to make a face prosthesis and bypass the face recognition system by wearing the face prosthesis. At present, there are many N times of modification of network videos. Such N times of modified network videos are usually synthesized by using PS, deep fake, etc. For example, AI face changing. At present, the following problems exist:

[0003] 1. At present, for different network videos or the same network video, from the first frame to the last frame, only a fixed detection method can be used to detect the video segment with a fake trace in the network video. The use scene is single, and it is not flexible to cope with network videos of various sources or detection business requirements. For the case where there are a large number of detection requirements and various video sources, it is not convenient to batch mark, audit and filter the video segments or network videos with fake traces.

[0004] 2. In addition, each detection tool for detecting the authenticity of a network video can only detect one fixed fake trace. For example, a PS detection tool can only detect the fake trace of a network video synthesized by using a PS technology, and a deep fake detection tool can only detect the fake trace of a network video synthesized by using a deep fake related technology, and it is easy to miss detection.

[0005] Therefore, due to the single detection method, when facing network videos with various fake methods, it is not possible to detect all network videos with various fake methods by using the same detection tool. SUMMARY

[0006] Embodiments of the present application provide a deep fake video detection method, device and storage medium, which can switch to an adaptive fake video detection method when detecting fake videos with different fake methods. In particular, when facing network videos with various fake methods, the authenticity detection of the current video to be detected can be dynamically switched to detect at least one network video with at least one fake method. The detection is flexible, suitable for diversified videos, and has a wide test coverage, which can effectively avoid missing detection of videos made by using a certain fake method.

[0007] In a first aspect, the embodiments of the present application provide a deep fake video detection method, comprising:

[0008] obtaining a first video to be detected from at least one video source, wherein the first video to be detected comprises a plurality of first face images of at least one target face;

[0009] identify a first facial action of a target face in at least one of the first face images in the first to-be-detected video;

[0010] determine a first detection manner from a plurality of preset detection manners based on the first facial action;

[0011] invoke the first detection manner to identify whether the target face in the to-be-detected video is fake or not, and output a first detection result.

[0012] In a possible implementation, before the identifying the first facial action of the target face in at least one of the first face images in the first to-be-detected video, the method further includes:

[0013] detecting at least one video frame in the to-be-detected video;

[0014] if it is detected that the target face in a target video frame in the to-be-detected video meets a preset fake trace condition, ending the detection of the remaining video frames in the first to-be-detected video, and generating a first detection result based on the fact that there is a video frame meeting the preset fake trace condition in the first to-be-detected video; wherein the target video frame is any video frame in the to-be-detected video, and the first detection result indicates that the to-be-detected video is a fake video.

[0015] In a possible implementation, before the determining the target detection manner from the plurality of preset detection manners based on the facial action, the method further includes:

[0016] analyzing the historical detection results according to fake trace types to obtain a fake manner corresponding to the historical processed video;

[0017] performing proportion analysis on fake manners of target video frames meeting a preset fake trace condition according to the historical detection results, and setting a preset mark on a video source side of the historical processed video according to the proportion of fake manners.

[0018] In a possible implementation, after the setting the preset mark on the video source side of the historical processed video according to the proportion of fake manners, the method includes:

[0019] obtaining a first to-be-detected video from a first video source side;

[0020] determining a second detection manner according to a preset correspondence relationship and a channel identifier of the first video source side, wherein the preset correspondence relationship includes a correspondence relationship among a preset mark, a default detection manner, and a channel identifier;

[0021] performing fake trace detection on the first to-be-detected video according to the second detection manner to obtain a second detection result.

[0022] In a possible implementation manner, after obtaining the second detection result, the method further includes:

[0023] obtaining first analysis data of the first to-be-detected video and second analysis data of the second detection result;

[0024] if it is determined according to the first analysis data and the second analysis data that the second analysis data has at least one of an exception, a false positive, or a missed detection, identifying a second facial action of a target face in at least one of the first face images in the first to-be-detected video;

[0025] determining a third detection manner from a plurality of preset detection manners based on the second facial action;

[0026] calling the third detection manner to identify whether the target face in the to-be-detected video is true or false, and outputting a third detection result.

[0027] In a possible implementation manner, after the third detection manner is determined from the plurality of preset detection manners based on the second facial action, the method further includes:

[0028] updating the third detection manner to the preset correspondence.

[0029] In a possible implementation manner, the method further includes:

[0030] determining a target segment in the first to-be-detected video that meets a preset fake condition, the target segment including at least one video frame that is continuous or spaced in playback time;

[0031] analyzing playback content corresponding to the target segment;

[0032] if a matching degree of the playback content and video description information is lower than a first threshold, or the matching degree belongs to the video description information but a proportion of the target segment is lower than a preset proportion, determining that the target segment is a non-key segment, setting a first mark for the first to-be-detected video, and the first mark is used to indicate that the first to-be-detected video is classified as a normal video.

[0033] In a possible implementation manner, the method further includes:

[0034] determining a target segment in the first to-be-detected video that meets a preset fake condition, the target segment including at least one video frame that is continuous or spaced in playback time;

[0035] If the proportion of the target slice in the first video to be detected is less than a second threshold, and the content weight of at least one video frame in the target slice is lower than a preset weight, and the forged object does not belong to the target object, the first video to be detected is classified as a normal video.

[0036] In a possible implementation, the first face motion of a target face in at least one of the first face images in the first video to be detected is identified, and a first detection manner is determined from a plurality of preset detection manners based on the first face motion, including:

[0037] A first set is filtered out from the first video to be detected, and the first set includes a plurality of face images that meet a preset face motion;

[0038] Forgery trace detection is performed on each face image in the first set.

[0039] In a possible implementation, after the preset mark is set for the video source side according to the proportion of the forgery manner, the method further includes:

[0040] A third video to be detected is obtained from a first video source side;

[0041] A fourth detection manner is determined according to historical detection data of the first video source side in a historical period, the fourth detection manner including a default detection manner or a priority detection manner;

[0042] Forgery trace detection is performed on the third video to be detected according to the fourth detection manner, to obtain a fourth detection result.

[0043] In a second aspect, an embodiment of the present application provides a video detection device having a function of implementing the deep fake video detection method provided in the first aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, which can be software and / or hardware.

[0044] In some embodiments, the video detection device includes:

[0045] An input and output module is configured to obtain a first video to be detected from at least one video source side, the first video to be detected including a plurality of first face images of at least one target face;

[0046] The processing module is configured to identify a first facial action of a target face in at least one first face image in the first to-be-detected video acquired by the input / output module, determine a first detection manner from a plurality of preset detection manners based on the first facial action, call the first detection manner to identify whether the target face in the to-be-detected video is true or false, and output a first detection result through the input / output module.

[0047] In a third aspect, an embodiment of the present application provides a computer device, comprising at least one connected processor, a memory and a transceiver, wherein the memory is configured to store a computer program, and the processor is configured to call the computer program in the memory to execute the method provided in the first aspect and various possible designs in the first aspect.

[0048] In yet another aspect, an embodiment of the present application provides a computer readable storage medium comprising instructions which, when executed on a computer, cause the computer to perform the method provided in the first aspect and various possible designs in the first aspect.

[0049] In yet another aspect, an embodiment of the present application provides a computer program product or computer program comprising computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the method provided in the first aspect and various possible designs in the first aspect.

[0050] Compared with the prior art, in the scheme provided by the embodiment of the application, multiple face recognition models and at least two fake video detection methods are deployed in the face recognition system. In one aspect, since each fake video detection method can well and specifically detect the fake video made by the corresponding fake means, the embodiment of the application can switch to the adaptive fake video detection method when facing fake videos of different fake methods for detection, especially when facing network videos of various fake methods, the adaptive fake detection method for the current to-be-detected video can be dynamically switched to, and at least one network video made by at least one fake method can be detected. In addition, the detection is flexible and suitable for diversified videos, and the test coverage is wide, which can effectively avoid missing some videos made by a certain fake method; in another aspect, before the target detection method is determined, the face action of the target face in the to-be-detected video is recognized first, and the state of the face action of the target face can represent the action state of the target face in at least two time windows. Generally, the dynamic face state is used for fake detection, and the dynamic face action is usually made by a deep fake method, which cannot be well detected by a general detection method. Therefore, the embodiment of the application determines the adaptive target detection method according to the state of the face action, and then detects the target face according to the target detection method. Under this detection strategy, the presence of the fake fragment in the first to-be-detected video can be accurately judged, and the detection efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0051] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of exemplary embodiments of the present application taken in conjunction with the accompanying drawings, in which like reference characters designate the same parts throughout the drawings.

[0052] Figure 1a is a schematic diagram of an application environment in which the deep fake video detection method is implemented in the embodiment of the application;

[0053] Figure 1b is a schematic diagram of an application environment in which the deep fake video detection method is implemented in the embodiment of the application;

[0054] Figure 2 is a schematic diagram of an application environment in which the deep fake video detection method is implemented in the embodiment of the application;

[0055] Figure 3 is a schematic diagram of an application environment in which the deep fake video detection method is implemented in the embodiment of the application;

[0056] Figure 4 is a schematic diagram of an application environment in which the deep fake video detection method is implemented in the embodiment of the application;

[0057] Figure 5 is a schematic diagram for visualizing the detection result in an embodiment of the present application;

[0058] Figure 6 is a schematic diagram for visualizing the detection result in an embodiment of the present application;

[0059] Figure 7 is a schematic diagram for visualizing the detection result in an embodiment of the present application;

[0060] Figure 8 is a structural schematic diagram of a video detection device for implementing the deep fake video detection method in an embodiment of the present application;

[0061] Figure 9 is a structural schematic diagram of a computer device for implementing the deep fake video detection method in an embodiment of the present application;

[0062] Figure 10 is a structural schematic diagram of a mobile phone for implementing the deep fake video detection method in an embodiment of the present application;

[0063] Figure 11 is a structural schematic diagram of a server for implementing the deep fake video detection method in an embodiment of the present application. DETAILED DESCRIPTION

[0064] The terms "first", "second", etc. in the description and claims of the present application and in the above drawings are used to distinguish similar objects (for example, the first face image and the second face image in the embodiments of the present application represent face images corresponding to different identities, respectively) and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or modules does not have to be limited to only those steps or modules clearly listed, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The division of modules in the embodiments of the present application is only a logical division, and in actual application, there can be another division manner, for example, a plurality of modules can be combined or integrated into another system, or some features can be ignored or not executed, in addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, the indirect coupling or communication connection between the modules can be electrical or other similar forms, which are not limited in the embodiments of the present application. In addition, the modules or sub-modules described as separate components can or can not be physically separated, can or can not be physical modules, or can be distributed into a plurality of circuit modules, and some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0065] The embodiments of the present application provide a deep fake video detection method and device and a storage medium, which can be applied to a video detection scene. The scheme can be executed by a server or a terminal. In some embodiments, the scheme is applied to an application scenario as shown in Figure 1a When the deep fake video detection method is implemented based on the application environment as shown in Figure 1a , the specific process can refer to Figure 1b . The server can detect the video from the data source for fake traces. Specifically,

[0066] The data source is a party that saves the video, for example, a short video platform, a social platform, a network database, etc. The data source in the embodiments of the present application can be a server or a terminal, which is not limited in the embodiments of the present application.

[0067] After obtaining the video to be detected from at least one data source, the server runs a pre-deployed, pre-trained image processing model to identify facial movements of the target face in at least one face image within the video. Based on the identified facial movement states, it predicts the appropriate detection method and then performs forgery detection on the video. The server can deploy at least two types of detection tools, such as a Photoshop detection tool and a deep fake detection tool, where the deep fake detection tool includes n deep fake detection algorithms (…). Figure 1b (All are labeled as models, and no distinction is made between them). PS detection tools and deep fakery detection tools can be deployed separately or integrated; this application's embodiments do not limit this. Figure 1b As shown, the server can select a target model from the deep fake detection tool as the detection method for the video to be detected by predicting the fake. The target model outputs the detection result of fake traces. For example, the detection result may include: the playback time point with fake traces and the video frame with fake traces.

[0068] It should be specifically noted that the server involved in this application embodiment (such as the video detection device mentioned above) can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The video detection device involved in this application embodiment can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, personal digital assistant, etc., but is not limited to these. The video detection device and the server can be directly or indirectly connected via wired or wireless communication, and this application embodiment does not impose any restrictions.

[0069] The solutions in this application can be implemented based on technologies such as Artificial Intelligence (AI), Natural Language Processing (NLP), and Machine Learning (ML), as illustrated in the following embodiments:

[0070] AI, or Artificial Intelligence, refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, Artificial Intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine capable of reacting in a manner similar to human intelligence. Artificial Intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0071] AI technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0072] NLP is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. Natural Language Processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close connection with linguistic research. Natural Language Processing techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0073] To address the aforementioned video detection problem in the field of artificial intelligence, the embodiments of this application mainly adopt the following technical solution: deploying multiple face recognition models and PS detection strategies in the face recognition system, recognizing the facial movements of the target face in the video to be detected, determining the appropriate target detection method based on the state of the facial movements, and performing real or fake detection on the target face according to the target detection method.

[0074] The following is in conjunction with the appendix Figures 2 to 7 The technical solutions of the embodiments of this application are described in detail.

[0075] See Figure 2 This is a flowchart illustrating a deepfake video detection method according to an embodiment of this application. The method can be executed by a service server, which can be various video detection platforms. The embodiments of this application mainly include... Figure 2 Steps 101 to 104 of the example are explained as follows:

[0076] 101. Obtain the first video to be detected from at least one video source.

[0077] The first to-be-detected video includes at least one target face image of a target face. The target face refers to a face with facial action in the first to-be-detected video. For example, the upper body head portrait of a girl "Mei" appears in consecutive n frames, and "Mei" successively changes from an expressionless face to a pupil-enlarged face, a mouth-opened face, and a laughing face in the consecutive n frames. It can be understood that the target face can correspond to the same user identity or at least two user identities, and the embodiments of the present application do not limit this.

[0078] 102. Identify a first facial action of a target face in at least one first face image in the first to-be-detected video.

[0079] The facial action can include smiling, opening the mouth, blinking, turning the head, lowering the head, raising the head, frowning, etc. Figure 3 The interface change schematic diagram of one of the live body verification tasks is shown. Figure 3 After clicking start, the user is prompted to look straight at the mobile phone screen and keep still, and then the user is prompted to turn the head, and then the live body verification of the user is started.

[0080] 103. Determine a first detection mode from a plurality of preset detection modes based on the first facial action.

[0081] It can be understood that the video detection device in the embodiments of the present application predeploys a plurality of preset detection modes to meet the detection needs of videos of various forgery modes.

[0082] Specifically, the target detection mode corresponding to at least one target model can be determined according to the state (for example, static or dynamic) of the facial action. The target video frame corresponding to any playback moment can be randomly selected, and the facial action in the selected target video frame is detected for forgery traces. Specifically, the following two cases can be included:

[0083] (1) If it is static, such as closed mouth, frontal face, and normal open eyes, a general detection mode (which can be a PS detection mode or a machine learning model) is used to detect whether it is a PS forgery.

[0084] For example, a photo ID, the face in the photo ID is PSed to a complex background in a video to obtain a forged video.

[0085] Reason: black production can usually obtain an ID photo, and the photo ID is a regular format photo, so the face PSed from the photo ID to a complex background will be more strange or unnatural. For example, a living background, which is more complex than a photo ID background.

[0086] Since the deep fake trace of the PS is not obvious, the face recognition system cannot identify the fake trace from machine vision. Therefore, a detection tool specifically for detecting PS traces is generally used to detect the authenticity of the target object in the static video.

[0087] (2) If there is such a dynamic as action, it is determined that there is a deep fake trace, and a deep fake detection model is called to identify the authenticity.

[0088] The action can include nodding, opening the mouth, shaking the head, closing the eyes, opening the eyes, etc., wherein the nodding and shaking of the head can be represented by a face posture angle, for example, the nodding uses the pitch angle; the shaking of the head uses the yaw angle, which involves the side face angle.

[0089] Specifically, whether it is dynamic can be determined according to the head turning angle, the mouth opening amplitude, the eye opening amplitude, the nodding amplitude, etc.

[0090] 104. calling the first detection method to identify the authenticity of the target face in the to-be-detected video, and outputting a first detection result.

[0091] Meanwhile, in order to facilitate subsequent analysis, a first mark can also be set for the to-be-detected video according to the detection result, and the second mark can include at least one of the authenticity identification, the fake type, and the fake frame amount.

[0092] Compared with the prior art, in the embodiments of the present application, a plurality of face recognition models and at least two fake video detection methods are deployed in the face recognition system. On the one hand, since each fake video detection method can well and specifically detect the fake video made by the corresponding fake means, the embodiments of the present application can switch to the adaptive fake video detection method when facing fake videos of different fake methods, especially when facing network videos of various fake methods, the authenticity detection can be dynamically switched to adapt to the current to-be-detected video, and at least one network video made by at least one fake method can be detected. In addition, the detection is flexible, suitable for diversified videos, and has wide test coverage, which can effectively avoid missing some videos made by a certain fake method; on the other hand, since the face action of the target face in the to-be-detected video is identified before the target detection method is determined, the state of the face action of the target face can represent the action state of the target face in at least two time windows. Since, generally speaking, the authenticity detection based on the dynamic face state is more accurate, and the dynamic face action is usually forged by a deep fake method, it cannot be well detected by a general detection method, so the embodiments of the present application determine the adaptive target detection method according to the state of the face action, and then detect the authenticity of the target face according to the target detection method. Under this detection strategy, it can not only ensure an accurate judgment of whether there is a fake fragment in the above-mentioned first to-be-detected video, but also improve the detection efficiency.

[0093] In some embodiments of the present application, considering that the depth of the deep fake technology adopted by different video source parties is different, or there is a certain rule, or the picture tone of the forged video is different, in order to better adapt to each video to be detected, the target detection method suitable for the current video to be detected can also be selected first. Specifically, at least one deep fake detection algorithm can be deployed in the detection tool, and each deep fake detection algorithm corresponds to a deep fake scene.

[0094] In some embodiments of the present application, in order to improve the detection efficiency of the forged video, the embodiments of the present application mainly start from the following first type of detection strategy and second type of detection strategy:

[0095] The first type of detection strategy is to detect the slices (including at least one video frame) in the video to be detected, which mainly includes strategy 1- strategy 3:

[0096] Strategy 1: rough detection

[0097] The video to be detected can be subjected to rough detection. Specifically, before recognizing the face action of the target face in the first face image in the first video to be detected, the method further comprises:

[0098] Detecting at least one video frame in the video to be detected;

[0099] If it is detected that the target face in the target video frame in the video to be detected meets the preset fake trace condition, the detection of the remaining video frames in the first video to be detected is ended, and a first detection result is generated based on the fact that there is a video frame in the first video to be detected that meets the preset fake trace condition; wherein the target video frame is any video frame in the video to be detected, and the first detection result indicates that the video to be detected is a fake video.

[0100] As can be seen, as long as it is detected that the face in at least one frame meets the preset fake trace condition, the detection of the remaining video frames in the video to be detected can be stopped, and it is preliminarily determined or directly determined that the video to be detected has a video frame that meets the preset fake trace condition. The detection result can be directly output, and the detection result is used to indicate that the video to be detected is a fake video. This detection mechanism can effectively improve the detection efficiency.

[0101] Strategy 2: hierarchical detection

[0102] In some embodiments, considering the detection accuracy and targeted analysis of the video to be detected, a hierarchical detection strategy can be used, that is, the forged video determined by rough detection can be subjected to fine detection (i.e., detection method based on face action) on the basis of the above-mentioned rough detection. Specifically, after obtaining the second detection result, the method further comprises:

[0103] obtaining first analysis data of the first to-be-detected video and second analysis data of the second detection result;

[0104] If it is determined according to the first analysis data and the second analysis data that the second analysis data has at least one of the following: an anomaly, a false judgment or a missed detection, a second facial action of a target face in at least one of the first face images in the first to-be-detected video is identified;

[0105] A third detection manner is determined from a plurality of preset detection manners based on the second facial action;

[0106] The third detection manner is called to identify whether the target face in the to-be-detected video is true or false, and a third detection result is output.

[0107] As can be seen, on the basis of the above coarse detection, the forged video determined after the coarse detection is further detected. For example, the action determination and the true or false trace analysis are performed on the remaining video frames in the to-be-detected video, and then the true or false detection result corresponding to each video frame is obtained, so that the hierarchical detection strategy can effectively improve the detection efficiency and accuracy.

[0108] Strategy 3: Decision detection manner based on preset correspondence

[0109] In some other embodiments, the proportion of the forgery means in the to-be-detected video can also be analyzed, and the video source party can be marked. After a new to-be-detected video is obtained from the corresponding video source party, the default detection manner can be directly called to detect the new to-be-detected video from the video source party according to the preset correspondence between the first mark, the default detection manner and the video source party ID (as shown in Table 1). Specifically, before the target detection manner is determined from a plurality of preset detection manners based on the facial action, the method further comprises:

[0110] The historical detection results are analyzed according to the types of the forgery traces, and the forgery manner corresponding to the historical processed video is obtained;

[0111] The proportion of the forgery manner of the target video frame meeting the preset forgery trace condition is analyzed according to the historical detection results, and a preset mark is set for the video source party of the historical processed video according to the proportion of the forgery manner.

[0112] Correspondingly, after the preset mark is set for the video source party of the historical processed video according to the proportion of the forgery manner, the method comprises:

[0113] A first to-be-detected video is obtained from a first video source party, and the first to-be-detected video is multiple.

[0114] According to the preset correspondence relationship and the channel identifier of the first video source party, a second detection manner is determined, and the preset correspondence relationship includes a preset mark, a default detection manner, and a corresponding relationship of a channel identifier.

[0115] According to the second detection manner, the first to-be-detected video is detected for a forgery trace to obtain a second detection result.

[0116] Pre-set first mark Default detection mode Video source side ID Correspondence 1 a0 Deep pseudo detection mode 10223 Correspondence 2 b0 PS 20384 … … Others …

[0117] Table 1

[0118] It can be seen that by pre-setting the above-mentioned preset correspondence relationship, on the one hand, when facing the first to-be-detected video, the decision-making time for selecting the second detection manner for detecting the first to-be-detected video can be saved; on the other hand, even if a new to-be-detected video with partial forgery traces (other forgery means) is not detected, the overall detection efficiency is improved to some extent, especially when a large number of videos are detected.

[0119] In some other embodiments, if it is found through analysis of a large number of video detection results that most videos cannot be detected to have forgery traces after the default detection manner is selected according to the preset correspondence relationship, then the detection manner determined based on the face motion in the above-mentioned embodiments needs to be re-used, and then the default detection manner in the preset correspondence relationship shown in Table 1 is updated. Alternatively, the final detection manner is determined according to historical detection data, which will be introduced as follows:

[0120] (1) Determining the detection manner based on the face motion, and then updating the preset correspondence relationship

[0121] As shown in Figure 4 After obtaining the second detection result, the method further includes:

[0122] Obtaining first analysis data of the first to-be-detected video and second analysis data of the second detection result;

[0123] If it is determined according to the first analysis data and the second analysis data that the second analysis data has at least one of the following: abnormality, misjudgment, or missed detection, a second face motion of a target face in at least one of the first face images in the first to-be-detected video is recognized;

[0124] Based on the second face motion, a third detection manner is determined from a plurality of preset detection manners;

[0125] The third detection manner is called to identify the true or false of the target face in the to-be-detected video, and a third detection result is output.

[0126] Correspondingly, after determining the third detection manner from the plurality of preset detection manners based on the second facial action, the third detection manner can also be updated to the preset corresponding relationship, for example, updating Table 1 above.

[0127] It can be seen that, on the basis of determining the default detection manner based on the preset corresponding relationship shown in Table 1 above and performing batch and rapid preliminary detection based on the default detection manner, the embodiment further adopts the detection manner determined based on the facial action in the above embodiment to perform fine detection, so as to further detect some or all of the missed fragments with fake traces on the basis of the second detection result. On the one hand, the detection accuracy can be improved through the two detections; on the other hand, after obtaining the third detection result, the preset corresponding relationship can be updated based on the third detection result, so that the subsequent target detection manner based on the continuously updated preset corresponding relationship can provide a higher matching degree and better detection effect in rapid detection, and form a positive feedback to continuously optimize the preset corresponding relationship.

[0128] (2) determining the final detection manner according to historical detection data:

[0129] For example, after setting a preset mark on the video source side of the historical processing video according to the proportion of the forgery manner, the method further comprises:

[0130] obtaining a third to-be-detected video from a first video source side;

[0131] obtaining historical detection data of the first video source side in a historical period according to the channel identifier of the first video source side to determine a fourth detection manner, the fourth detection manner including a default detection manner or a priority detection manner;

[0132] performing forgery trace detection on the third to-be-detected video according to the fourth detection manner to obtain a fourth detection result.

[0133] It can be seen that, before determining the detection manner, the fourth detection manner is determined based on the historical detection data of the first video source side. Since the historical detection data has a certain representativeness, it can represent the tendency and adaptability of the image monitoring device to the detection manner of the video from the first video source side in the historical period, so the fourth detection manner determined based on the historical detection data can better continue to detect the subsequent video from the first video source side, on the one hand, reducing the decision time of the detection manner, on the other hand, ensuring the detection accuracy and detection efficiency.

[0134] In some embodiments, the proportion of the target video segment in which the fake trace exists in the to-be-detected video can be further analyzed, so as to analyze each video source party. For example, the source party of the to-be-detected video with a proportion greater than a first threshold can be set with a second mark, and the amount of video of the source party set with the second mark is counted, and each source party is graded according to the statistical data. The second mark can include a fake type and a fake frame proportion. The second mark can be used for video credit level, public opinion analysis, etc., and the present application does not limit the marking method and content.

[0135] In some embodiments, even if there are a small number of target segments (i.e., at least one video frame in which a fake trace exists), if the target segment does not affect the user's evaluation or content output after watching the entire to-be-detected video, the to-be-detected video can be classified as a normal video. Specifically, the present application also provides the following a and b detection strategies:

[0136] Detection strategy a: judging based on content relevance of video segments

[0137] (1) Determine the target segment in the first to-be-detected video that meets the preset fake condition, the target segment including at least one video frame with continuous or interval playing time;

[0138] (2) If the proportion of the target segment in the first to-be-detected video is less than a second threshold, and the content weight of at least one video frame in the target segment is lower than a preset weight, and the object that is fake does not belong to a target object, the first to-be-detected video is classified as a normal video.

[0139] For example, the playing time of the to-be-detected video is 1 hour, and there is a fake trace in the target video frame corresponding to the playing time of 36:23:56, but the target video frame does not affect the user's correct understanding of the entire to-be-detected video. For example, the to-be-detected video is a 1-hour speech video of a specific person a, and the face of a person in the target video frame corresponding to the playing time of 36:23:56 is replaced (which can be the specific person a who is speaking, other specific person b, audience / staff, etc.), but the target video frame does not affect the user's correct understanding of the entire speech video and the identification of the binding relationship between the speech behavior of the specific person a and the speech video.

[0140] If the to-be-detected video is directly marked as described above, it will be too harsh and unnecessary. Therefore, it can be determined whether the proportion of the target segment in the to-be-detected video is less than a second threshold, and the content weight of at least one video frame in the target segment is lower than a preset weight, and the object that is fake does not belong to a target object, and the to-be-detected video is classified as a normal video and is not filtered out.

[0141] Therefore, through the detection strategy a, the forged video which really affects the expression of the video content can be more accurately distinguished, so as to reduce the case that the entire video to be detected is determined as a forged video as long as the fragment with deep forgery traces is detected from the video to be detected, that is, to reduce the influence of some ambiguous forgery detection fragments on the overall judgment of the video to be detected.

[0142] Detection strategy b: judging based on the importance of the video fragment in the video to be detected

[0143] Specifically, the following steps are included:

[0144] (1) determining a target fragment in the first video to be detected that meets a preset forgery condition, the target fragment including at least one video frame that is continuous or intermittent in playing time;

[0145] (2) analyzing the playing content corresponding to the target fragment;

[0146] (3) if the matching degree of the playing content with the video description information is lower than a first threshold, or the matching degree belongs to the video description information but the proportion of the target fragment is lower than a preset proportion, determining that the target fragment is a non-key fragment, and setting a first mark for the first video to be detected, the first mark being used to indicate that the first video to be detected is classified as a normal video.

[0147] It can be seen that, through the detection strategy b, the forged video which really affects the expression of the video content can be more accurately distinguished, so as to reduce the influence of some ambiguous forgery detection fragments on the overall judgment of the video to be detected.

[0148] Second type of detection strategy: focusing on detecting the video frame with dynamic face action

[0149] Since the accuracy of the method of determining whether the video to be detected is a forged video by judging whether the video frame with face action has forgery traces is higher than that of the method of determining whether the video to be detected is a forged video by judging whether the video frame in a silent state is a forged video, in order to ensure the detection accuracy while improving the detection efficiency, when it is determined that there is a face feature fragment in the video to be detected, a deep forgery detection method can be preferentially used for detection.

[0150] For example, a first set is first screened from the first video to be detected, and then forgery trace detection is performed on each face image in the first set. The first set includes a plurality of face images that meet a preset face action, and frame skipping or frame dropping detection can be performed, and the specific diagram and method are not limited.

[0151] Any technical feature mentioned in the embodiments of any one of the Figures 7 of the drawings also applies to the Figures 8 to 11The subsequent similar places will not be repeated.

[0152] The above describes a deep fake video detection method in an embodiment of the application. The following introduces a video detection device for executing the deep fake video detection method.

[0153] Referring to Figure 8 As Figure 8 Fig. 4 shows a structural schematic diagram of a video detection device 40 according to an embodiment of the application. The video detection device 40 can be applied to detect the traces of forgery of network videos from different data source parties to detect at least one network video that has been modified, so as to ultimately purify the network and avoid unnecessary public opinion fermentation and ensure the authenticity of network videos. The video detection device 40 in the embodiment of the application can implement the steps of the deep fake video detection method performed by the video detection device 40 in any of the corresponding embodiments. The functions implemented by the video detection device 40 can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. The video detection device 40 can include an input / output module 401 and a processing module 402. The functions of the input / output module 401 and the processing module 402 can be referred to the functions of the input / output module 101 and the processing module 102 in the video detection device 10 in the embodiment of the application. Figure 6 The functions implemented by the video detection device 40 can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. The video detection device 40 can include an input / output module 401 and a processing module 402. The functions of the input / output module 401 and the processing module 402 can be referred to the functions of the input / output module 101 and the processing module 102 in the video detection device 10 in the embodiment of the application. Figure 7 The operations performed in any of the corresponding embodiments will not be repeated here.

[0154] In some embodiments, the input / output module 401 can be configured to obtain a first to-be-detected video from at least one video source party, the first to-be-detected video including a plurality of first face images of at least one target face;

[0155] The processing module 402 can be configured to identify a first facial action of the target face in at least one of the first face images in the first to-be-detected video obtained by the input / output module 401; determine a first detection manner from a plurality of preset detection manners based on the first facial action; call the first detection manner to identify the authenticity of the target face in the to-be-detected video, and output a first detection result through the input / output module.

[0156] In some embodiments, before the processing module 402 identifies the facial action of the target face in at least one of the first face images in the first to-be-detected video, the processing module 402 is further configured to:

[0157] detect at least one video frame in the to-be-detected video;

[0158] If it is detected that a target face in a target video frame in the to-be-detected video meets a preset forgery trace condition, the detection on the remaining video frames in the first to-be-detected video is ended, and a first detection result is generated based on the first to-be-detected video meeting the preset forgery trace condition; wherein the target video frame is any video frame in the to-be-detected video, and the first detection result indicates that the to-be-detected video is a forged video.

[0159] In some embodiments, before the processing module 402 determines the target detection manner from the plurality of preset detection manners based on the facial action, the processing module 402 is further configured to:

[0160] analyze the historical detection results according to the forgery trace types, to obtain a forgery manner corresponding to the historical processed video;

[0161] According to the historical detection results, the forgery manner of the target video frame meeting the preset forgery trace condition is analyzed, and a preset mark is set for the video source side of the historical processed video according to the forgery manner proportion.

[0162] In some embodiments, after the processing module 402 sets the preset mark for the video source side of the historical processed video according to the forgery manner proportion, the processing module 402 is further configured to:

[0163] obtain a first to-be-detected video from a first video source side through the input and output module 401;

[0164] determine a second detection manner according to a preset correspondence relationship and a channel identifier of the first video source side, wherein the preset correspondence relationship includes a correspondence relationship among a preset mark, a default detection manner, and a channel identifier;

[0165] perform forgery trace detection on the first to-be-detected video according to the second detection manner, to obtain a second detection result.

[0166] In some embodiments, after the processing module 402 obtains the second detection result, the processing module 402 is further configured to:

[0167] obtain first analysis data of the first to-be-detected video and second analysis data of the second detection result;

[0168] If it is determined that the second analysis data has at least one of the following: abnormality, misjudgment, or missed detection, according to the first analysis data and the second analysis data, a second facial action of a target face in at least one of the first face images in the first to-be-detected video is identified;

[0169] determine a third detection manner from a plurality of preset detection manners based on the second facial action;

[0170] The third detection manner is called to identify authenticity of a target face in the to-be-detected video, and a third detection result is output.

[0171] In some embodiments, after determining the third detection manner from the plurality of preset detection manners based on the second facial action, the processing module 402 is further configured to:

[0172] update the third detection manner into the preset correspondence relationship.

[0173] In some embodiments, the processing module 402 is further configured to:

[0174] determine a target segment in the first to-be-detected video that meets a preset forgery condition, the target segment including at least one video frame that is continuous or intermittent in playback time;

[0175] analyze playback content corresponding to the target segment;

[0176] if a matching degree of the playback content with video description information is lower than a first threshold value, or the matching degree belongs to the video description information but a proportion of the target segment is lower than a preset proportion, determine that the target segment is a non-key segment, and set a first mark for the first to-be-detected video, the first mark being used to indicate that the first to-be-detected video is classified as a normal video.

[0177] In some embodiments, the processing module 402 is further configured to:

[0178] determine a target segment in the first to-be-detected video that meets a preset forgery condition, the target segment including at least one video frame that is continuous or intermittent in playback time;

[0179] if a proportion of the target segment in the first to-be-detected video is less than a second threshold value, and a content weight of at least one video frame in the target segment is lower than a preset weight, and a forged object does not belong to a target object, classify the first to-be-detected video as a normal video.

[0180] In some embodiments, the processing module 402 is specifically configured to:

[0181] screen a first set from the first to-be-detected video, the first set including a plurality of face images that meet a preset facial action;

[0182] perform forgery trace detection on each face image in the first set.

[0183] In some embodiments, after setting a preset mark for a video source side of the historical processing video according to a forgery manner proportion, the processing module 402 is further configured to:

[0184] obtaining a third to-be-detected video from the first video source party through the input and output module 401;

[0185] obtaining historical detection data of the first video source party in a historical period according to a channel identifier of the first video source party to determine a fourth detection manner, the fourth detection manner including a default detection manner or a priority detection manner;

[0186] performing a fake trace detection on the third to-be-detected video according to the fourth detection manner to obtain a fourth detection result.

[0187] As to the video detection device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the method, and will not be described in detail here.

[0188] From the above Figure 8 As can be known from the example video detection device 40, a plurality of face recognition models and at least two fake video detection manners are deployed in the face recognition system. In one aspect, since each fake video detection manner can well and pertinently detect fake videos made by corresponding fake means, the embodiments of the present application can arbitrarily switch to an adaptive fake video detection manner when facing fake videos made by different fake manners, especially when facing network videos with various fake manners, can dynamically switch to an adaptive authenticity detection manner for the current to-be-detected video, and can detect at least one network video made by at least one fake manner. In addition, the detection is flexible and suitable for diversified videos, and has a wide test coverage, which can effectively avoid missing some videos made by a certain fake manner; in another aspect, since the face action of the target face in the to-be-detected video is recognized before the target detection manner is determined, the state of the face action of the target face can represent the action state of the target face in at least two time windows. Since, generally speaking, authenticity detection based on dynamic face state is more accurate, and dynamic face action is usually made by deep fake manner, it cannot be well detected by general detection manner, so the embodiments of the present application determine an adaptive target detection manner according to the state of the face action, and then perform authenticity detection on the target face according to the target detection manner. Under this detection strategy, it can not only accurately judge whether there is a fake segment in the first to-be-detected video, but also improve the detection efficiency.

[0189] The video detection device 40 for performing the deep fake video detection method in the embodiments of the present application is described from the perspective of hardware processing. It should be noted that, in the embodiments of the present application Figure 7The input / output module 401 in the embodiment shown corresponds to an entity device such as an input / output unit, a transceiver, a radio frequency circuit, a communication module, and an output interface. The processing module 402 corresponds to an entity device such as a processor. Figure 7 The video detection apparatus 40 shown can have a structure as shown in Figure 9 The video detection apparatus 40 shown can have a structure as shown in Figure 7 The video detection apparatus 40 shown can have a structure as shown in Figure 9 The video detection apparatus 40 shown can have a structure as shown in Figure 9 The processor and the transceiver in the video detection apparatus 40 can realize the same or similar functions of the input / output module 401 and the processing module 402 provided by the apparatus embodiment corresponding to the video detection apparatus 40, Figure 9 The memory in the video detection apparatus 40 stores a computer program required by the processor to execute the deep fake video detection method.

[0190] The embodiment of the present application also provides another video detection apparatus. As shown in Figure 10 For ease of illustration, only parts related to the embodiment of the present application are shown, and specific technical details not disclosed are referred to the method part of the embodiment of the present application. The video detection apparatus can be any video detection apparatus such as a mobile phone, a tablet computer, a personal digital assistant (English full name: Personal Digital Assistant, English abbreviation: PDA), a point of sales (English full name: Point of Sales, English abbreviation: POS), and a vehicle-mounted computer. Taking the mobile phone as an example of the video detection apparatus:

[0191] Figure 10 The diagram shown is a block diagram of part of the structure of the mobile phone related to the video detection apparatus provided by the embodiment of the present application. Referring to Figure 10 , the mobile phone includes a radio frequency (English full name: Radio Frequency, English abbreviation: RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 780, an audio circuit 760, a wireless fidelity (English full name: wireless-fidelity, English abbreviation: Wi-Fi) module 7100, a processor 780, and a power supply 790, and the like. Those skilled in the art can understand that Figure 7 The structure of the mobile phone shown in the embodiment of the present application does not constitute a limitation on the mobile phone, and can include more or fewer components than those shown, or combine certain components, or different arrangement of components.

[0192] The components of the mobile phone will be specifically introduced as follows: Figure 10

[0193] ​The RF circuit 710 can be used for receiving and sending signals in the process of information or communication, in particular, receiving the downlink information from the base station and processing by the processor 780; in addition, sending the uplink data to the base station. Generally, the RF circuit 710 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 710 can also communicate with the network and other devices through wireless communication. The above-mentioned wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0194] The memory 720 can be used to store software programs and modules, and the processor 780 can execute various function applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 720 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0195] The input unit 730 can be used to receive input digital or character information, and to generate key signal input with respect to user setting of the mobile phone and control of the function. Specifically, the input unit 730 can include a touch panel 731 and other input device 732. The touch panel 731, also called a touch screen, can collect a touch operation (such as an operation of a user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 731) of the user on or near it, and drive the corresponding connection device according to the pre-set program. Optionally, the touch panel 731 can include two parts of a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, and converts it into touch coordinates and sends it to the processor 780, and can receive the command from the processor 780 and execute it. In addition, the touch panel 731 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 731, the input unit 730 can also include other input device 732. Specifically, the other input device 732 can include one or more of a physical keyboard, a function key (such as a volume control button, an on-off button, etc.), a trackball, a mouse, a joystick, etc.

[0196] The display unit 740 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 740 can include a display panel 741, which can be configured in the form of a liquid crystal display (English full name: Liquid Crystal Display, English abbreviation: LCD), an organic light-emitting diode (English full name: Organic Light-Emitting Diode, English abbreviation: OLED), etc. Further, the touch panel 731 can cover the display panel 741, and when the touch panel 731 detects a touch operation on or near it, it is transmitted to the processor 780 to determine the type of touch event, and then the processor 780 provides corresponding visual output on the display panel 741 according to the type of touch event. Although in the Figure 7 , the touch panel 731 and the display panel 741 are realized as two independent components to realize the input and output functions of the mobile phone, but in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.

[0197] The phone can also include at least one sensor 780, such as an optical sensor, a motion sensor, and other sensors. Specifically, the optical sensor can include an ambient light sensor to adjust the brightness of the display panel 741 according to the brightness of ambient light, and a proximity sensor to turn off the display panel 741 and / or the backlight when the phone is moved to the ear. As one of the motion sensors, the accelerometer sensor can detect the magnitude of acceleration in each direction (usually three axes), and when at rest, the magnitude and direction of gravity, which can be used for applications that identify the phone posture (such as switching between landscape and portrait screens, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometers, taps), and the like. As for other sensors that the phone can also be configured, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, and the like, will not be described here.

[0198] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the phone. The audio circuit 760 can convert the received audio data into an electrical signal, transmit it to the speaker 761, and convert it into a sound signal output by the speaker 761; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and converted into audio data, which is then output to the processor 780 for processing, and then transmitted to another phone via the RF circuit 710, or output to the memory 720 for further processing.

[0199] Wi-Fi belongs to a short-range wireless transmission technology. The phone can help users send and receive emails, browse web pages, and access streaming media through the Wi-Fi module 7100, which provides users with wireless broadband Internet access. Although Figure 9 Although the Wi-Fi module 7100 is shown, it is understood that it does not belong to the necessary components of the phone, and can be omitted as needed without changing the essence of the application.

[0200] The processor 780 is the control center of the phone, which connects all parts of the phone through various interfaces and lines, executes various functions of the phone and processes data by running or executing software programs and / or modules stored in the memory 720, and calling data stored in the memory 720, thereby monitoring the phone as a whole. Optionally, the processor 780 can include one or more processing units; preferably, the processor 780 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application program, and the modem processor mainly processes wireless communication. It is understood that the above-mentioned modem processor can also not be integrated into the processor 780.

[0201] The mobile phone further includes a power supply 790 (such as a battery) for supplying power to each component, which can be logically connected to the processor 780 through a power management system, so as to realize functions such as power management, discharge management, and power consumption management through the power management system.

[0202] Although not shown, the mobile phone can further include a camera, a Bluetooth module, and the like, which will not be described herein.

[0203] In the embodiments of the present application, the processor 780 included in the mobile phone further has a function of controlling the method flow executed by the video detection device 40 shown in the above. Figure 10 The steps executed by the video detection device in the above embodiments can be based on the structure of the mobile phone shown in the above. Figure 10 For example, the processor 722 executes the following operations by calling instructions in the memory 732:

[0204] obtains a first to-be-detected video from at least one video source through the input unit 730, the first to-be-detected video including a plurality of first face images of at least one target face;

[0205] recognizes a first facial action of a target face in at least one first face image in the first to-be-detected video obtained by the input / output module 401, determines a first detection manner from a plurality of preset detection manners based on the first facial action, calls the first detection manner to recognize a true or false target face in the to-be-detected video, and outputs a first detection result through the input unit 730.

[0206] The embodiments of the present application further provide another video detection device for implementing the above deep fake video detection method, as shown in the above. Figure 11 Figure 11 is a schematic diagram of a server structure provided by the embodiments of the present application. The server 1020 can have a large difference due to different configurations or performances, and can include one or more central processing units (English: central processing units, English: CPU) 1022 (for example, one or more processors) and a memory 1032, one or more storage media 1030 (for example, one or more mass storage devices) for storing application programs 1042 or data 1044. Among them, the memory 1032 and the storage medium 1030 can be temporary storage or persistent storage. The programs stored in the storage medium 1030 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the server. Further, the central processing unit 1022 can be configured to communicate with the storage medium 1030 and execute a series of instruction operations in the storage medium 1030 on the server 1020.

[0207] ​The server 1020 can also include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0208] The steps performed by the service server (e.g. Figure 7 the video detection apparatus 40) in the above embodiments can be based on the structure of the server 1020 as shown in FIG. 10. For example, the steps performed by the video detection apparatus 40 in the above embodiments can be based on the server structure as shown in FIG. 10. For example, the processor 1022 performs the following operations by invoking instructions in the memory 1032: Figure 11 Figure 7 The steps performed by the service server (e.g. Figure 11 the video detection apparatus 40) in the above embodiments can be based on the structure of the server 1020 as shown in FIG. 10. For example, the steps performed by the video detection apparatus 40 in the above embodiments can be based on the server structure as shown in FIG. 10. For example, the processor 1022 performs the following operations by invoking instructions in the memory 1032:

[0209] obtaining, by an input / output interface 1058, a first to-be-detected video from at least one video source, the first to-be-detected video including a plurality of first face images of at least one target face;

[0210] identifying a first facial action of a target face in at least one of the first face images in the first to-be-detected video obtained by the input / output module 401; determining a first detection manner from a plurality of preset detection manners based on the first facial action; calling the first detection manner to identify whether the target face in the to-be-detected video is true or false, and outputting a first detection result through the input / output interface 1058.

[0211] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0212] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, apparatus and module described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0213] ​In several embodiments provided in the embodiments of the present application, it should be understood that the disclosed system, device and method can be implemented by other manners. For example, the above-described device embodiments are merely illustrative, for example, the division of the modules is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the shown or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.

[0214] The modules described as separate components can or can not be physically separated, and the components shown as modules can or can not be physical modules, that is, can be located in one place, or can be distributed to a plurality of network modules. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0215] In addition, each functional module in each embodiment of the embodiments of the present application can be integrated in one processing module, or each module can exist physically, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can be stored in a computer readable storage medium.

[0216] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, all or part can be realized in the form of a computer program product.

[0217] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on the computer, the flow or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can store or be integrated into a data storage device such as a server, data center, etc. containing one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0218] The above describes the technical solutions provided by the embodiments of the present application in detail. The principles and implementation manners of the embodiments of the present application are described by applying specific examples. The above examples are only used to help understand the method and core idea of the embodiments of the present application; at the same time, for those skilled in the art, according to the idea of the embodiments of the present application, the specific implementation manner and application range will be changed, and the above description should not be understood as a limitation of the embodiments of the present application.

Claims

1. A method for detecting deepfake videos, characterized in that, The method includes: A first video to be detected is obtained from at least one video source, the first video to be detected including multiple first face images of at least one target face; Detect at least one video frame in the video to be detected; If a target face in a target video frame of the video to be detected matches a preset forgery condition, the detection of the remaining video frames in the first video to be detected is terminated, and a first detection result is generated based on the existence of video frames in the first video to be detected that match the preset forgery condition; wherein, the target video frame is any video frame in the video to be detected, and the first detection result indicates that the video to be detected is a forged video. Identify the first facial action of the target face in at least one of the first face images in the first video to be detected; By analyzing the historical detection results according to the type of forgery traces, the forgery methods corresponding to the historical processed videos can be obtained; Based on the historical detection results, the proportion of forgery methods for target video frames that meet the preset forgery trace conditions is analyzed, and a preset mark is set for the video source of the historically processed video based on the proportion of forgery methods. Obtain the first video to be tested from the first video source; A second detection method is determined based on a preset correspondence and the channel identifier of the first video source. The preset correspondence includes a correspondence between a preset marker, a default detection method, and a channel identifier. The first video to be detected is subjected to forgery detection according to the second detection method to obtain the second detection result; The first detection method is determined from a variety of preset detection methods based on the first facial movement; The first detection method is invoked to identify whether the target face in the video to be detected is real or fake, and the first detection result is output.

2. The deepfake video detection method according to claim 1, characterized in that, After obtaining the second detection result, the method further includes: Acquire the first analysis data of the first video to be detected and the second analysis data of the second detection result; If, based on the first analysis data and the second analysis data, it is determined that the second analysis data contains at least one of the following: anomaly, misjudgment, or missed detection, then the second facial action of the target face in at least one of the first face images in the first video to be detected is identified. A third detection method is determined from multiple preset detection methods based on the second facial movement; The third detection method is invoked to identify whether the target face in the video to be detected is real or fake, and the third detection result is output.

3. The deepfake video detection method according to claim 1 or 2, characterized in that, The method further includes: Identify a target segment in the first video to be detected that meets the preset forgery conditions, wherein the target segment includes at least one video frame with continuous or intermittent playback time; Analyze the playback content corresponding to the target segment; If the matching degree between the playback content and the video description information is lower than a first threshold, or if the matching degree belongs to the video description information but the proportion of the target segment is lower than a preset proportion, then the target segment is determined to be a non-critical segment, and the first video to be detected is set with a first mark. The first mark is used to indicate that the first video to be detected is classified as a normal video.

4. The deepfake video detection method according to claim 1 or 2, characterized in that, The method further includes: Identify a target segment in the first video to be detected that meets the preset forgery conditions, wherein the target segment includes at least one video frame with continuous or intermittent playback time; If the proportion of the target segment in the first video to be detected is less than the second threshold, and the content weight of at least one video frame in the target segment is lower than the preset weight, and the forged object does not belong to the target object, then the first video to be detected is classified as a normal video.

5. The deepfake video detection method according to claim 1 or 2, characterized in that, The first facial action of the target face in at least one of the first face images in the first video to be detected is identified; Based on the first facial movement, a first detection method is determined from multiple preset detection methods, including: A first set is selected from the first video to be detected, the first set including multiple face images that conform to preset face actions; Forgery detection is performed on each face image in the first set.

6. The deepfake video detection method according to claim 1, characterized in that, After setting a preset marker for the video source of the historically processed video based on the proportion of forgery methods, the method further includes: Obtain the third video to be tested from the first video source; The fourth detection method is determined by obtaining the historical detection data of the first video source within a historical time period based on the channel identifier of the first video source. The fourth detection method includes a default detection method or a priority detection method. The third video to be detected is subjected to forgery detection using the fourth detection method to obtain the fourth detection result.

7. A video detection device, comprising: An input / output module is used to acquire a first video to be detected from at least one video source, wherein the first video to be detected includes multiple first face images of at least one target face; The processing module is used to detect at least one video frame in the video to be detected; if a target face in a target video frame in the video to be detected is found to meet a preset forgery trace condition, the detection of the remaining video frames in the first video to be detected is stopped, and a first detection result is generated based on the existence of video frames in the first video to be detected that meet the preset forgery trace condition; wherein, the target video frame is any video frame in the video to be detected, and the first detection result indicates that the video to be detected is a forged video; the module identifies the first facial action of the target face in at least one first face image in the first video to be detected obtained by the input / output module; the module analyzes the historical detection results according to the forgery trace type to obtain the forgery method corresponding to the historical processed video; the module analyzes the proportion of forgery methods of target video frames that meet the preset forgery trace condition according to the historical detection results, and sets a preset mark for the video source of the historical processed video according to the proportion of forgery methods; the module obtains the first video to be detected from the first video source through the input / output module; and determines the second detection method according to the preset correspondence relationship and the channel identifier of the first video source, wherein the preset correspondence relationship includes the correspondence relationship between the preset mark, the default detection method, and the channel identifier. The first video to be detected is subjected to forgery detection according to the second detection method to obtain a second detection result; a first detection method is determined from a variety of preset detection methods based on the first facial movements; the first detection method is called to identify the authenticity of the target face in the video to be detected, and the first detection result is output through the input / output module.

8. The apparatus according to claim 7, characterized in that, After obtaining the second detection result, the processing module is further used for: Acquire the first analysis data of the first video to be detected and the second analysis data of the second detection result; If, based on the first analysis data and the second analysis data, it is determined that the second analysis data contains at least one of the following: anomaly, misjudgment, or missed detection, then the second facial action of the target face in at least one of the first face images in the first video to be detected is identified. A third detection method is determined from multiple preset detection methods based on the second facial movement; The third detection method is invoked to identify whether the target face in the video to be detected is real or fake, and the third detection result is output.

9. The apparatus according to claim 8, characterized in that, After determining the third detection method from multiple preset detection methods based on the second facial movement, the processing module is further used for: Update the third detection method to the preset correspondence.

10. The apparatus according to any one of claims 7-9, characterized in that, The processing module is also used for: Identify a target segment in the first video to be detected that meets the preset forgery conditions, wherein the target segment includes at least one video frame with continuous or intermittent playback time; Analyze the playback content corresponding to the target segment; If the matching degree between the playback content and the video description information is lower than a first threshold, or if the matching degree belongs to the video description information but the proportion of the target segment is lower than a preset proportion, then the target segment is determined to be a non-critical segment, and the first video to be detected is set with a first mark. The first mark is used to indicate that the first video to be detected is classified as a normal video.

11. The apparatus according to any one of claims 7-9, characterized in that, The processing module is also used for: Identify a target segment in the first video to be detected that meets the preset forgery conditions, wherein the target segment includes at least one video frame with continuous or intermittent playback time; If the proportion of the target segment in the first video to be detected is less than the second threshold, and the content weight of at least one video frame in the target segment is lower than the preset weight, and the forged object does not belong to the target object, then the first video to be detected is classified as a normal video.

12. The apparatus according to any one of claims 7-9, characterized in that, The processing module is specifically used for: A first set is selected from the first video to be detected, the first set including multiple face images that conform to preset face actions; Forgery detection is performed on each face image in the first set.

13. The apparatus according to claim 7, characterized in that, After setting a preset marker for the video source of the historical processed video based on the proportion of forgery methods, the processing module is further used for: The third video to be detected is obtained from the first video source through the input / output module; The fourth detection method is determined by obtaining the historical detection data of the first video source within a historical time period based on the channel identifier of the first video source. The fourth detection method includes a default detection method or a priority detection method. The third video to be detected is subjected to forgery detection using the fourth detection method to obtain the fourth detection result.

14. A computing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the method as described in any one of claims 1-6.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for detecting authenticity of figure in video, electronic equipment and storage medium

    CN111444873A