Video forgery detection method and device, equipment and medium

By extracting multimodal feature information of video and performing feature fusion, the problem of low accuracy of video forgery detection in the prior art is solved, and more efficient video forgery detection is achieved.

CN120032287APending Publication Date: 2025-05-23ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411074127.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art usually relies on single modal features in video forgery detection, resulting in low detection accuracy and difficulty in effectively distinguishing authenticity.

Method used

By extracting the multimodal feature information of the video, including airspace feature information, time domain feature information and frequency domain feature information, and performing feature fusion, input to the forgery detection network for video forgery detection.

Benefits of technology

It improves the accuracy of video forgery detection and can more effectively distinguish between real videos and fake videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032287A_ABST
    Figure CN120032287A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video forgery detection method and device, equipment and a medium. The method comprises the following steps: acquiring a to-be-detected video; extracting target feature information of the to-be-detected video; wherein the target feature information comprises spatial domain feature information, time domain feature information and frequency domain feature information; and according to the target feature information, performing forgery detection on the to-be-detected video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of forgery detection, and more specifically, to a video forgery detection method, device, equipment and medium. Background Art

[0002] In the related art, deep learning image forgery algorithms can generate forged images that are close to real images, so deep learning image forgery algorithms can be applied to different fields. However, due to the advancement of deep learning image forgery algorithms, realistic images that are increasingly difficult to distinguish between true and false have appeared, which may lead to social suspicion of authenticity and credibility. Therefore, forgery detection of images can be performed based on forgery detection models.

[0003] In the related art, when performing video forgery detection, it is usually performed by extracting a single modal feature of the video, and performing video forgery detection based on the extracted single modal feature through a forgery detection model, resulting in low accuracy of video forgery detection. Summary of the invention

[0004] The embodiments of the present disclosure provide a video forgery detection method, apparatus, device and medium.

[0005] According to a first aspect of the present disclosure, a video forgery detection method is provided, the method comprising:

[0006] Get the video to be detected;

[0007] Extracting target feature information of the video to be detected; wherein the target feature information includes spatial domain feature information, temporal domain feature information and frequency domain feature information;

[0008] The video to be detected is subjected to forgery detection according to the target feature information.

[0009] Optionally, extracting target feature information of the video to be detected includes:

[0010] Respectively extracting a single frame image, a video clip and an audio clip of the video to be detected;

[0011] Obtaining spatial feature information of the video to be detected according to the extracted single-frame image;

[0012] Obtaining time domain feature information of the video to be detected according to the extracted video clip;

[0013] The frequency domain feature information of the video to be detected is obtained according to the extracted audio segment.

[0014] Optionally, obtaining the spatial feature information of the video to be detected according to the extracted single-frame image includes:

[0015] Inputting the extracted single-frame image into a preset first feature extraction network to obtain spatial domain feature information corresponding to the single-frame image;

[0016] The spatial feature information corresponding to the single-frame image is subjected to feature splicing and feature fusion to obtain the spatial feature information of the video to be detected.

[0017] Optionally, obtaining the time domain feature information of the video to be detected according to the extracted video clip includes:

[0018] The extracted video clip is input into a preset second feature extraction network to obtain time domain feature information corresponding to the video clip as the time domain feature information of the video to be detected.

[0019] Optionally, obtaining frequency domain feature information of the video to be detected according to the extracted audio segment includes:

[0020] Inputting the extracted audio segment into a preset third feature extraction network to obtain frequency domain feature information corresponding to the audio segment;

[0021] The frequency domain feature information corresponding to the audio clip is subjected to feature splicing and feature fusion to obtain the frequency domain feature information of the video to be detected.

[0022] Optionally, performing forgery detection on the video to be detected according to the target feature information includes:

[0023] Performing feature fusion on the spatial domain feature information, the time domain feature information and the frequency domain feature information to obtain fused feature information;

[0024] The fused feature information is input into a preset forgery detection network for forgery detection.

[0025] According to a second aspect of the present disclosure, a video forgery detection device is provided, the device comprising:

[0026] An acquisition module is used to acquire the video to be detected;

[0027] An extraction module, used to extract target feature information of the video to be detected; wherein the target feature information includes spatial domain feature information, temporal domain feature information and frequency domain feature information;

[0028] The video to be detected is subjected to forgery detection according to the target feature information.

[0029] According to a third aspect of the present disclosure, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store an executable computer program; and the computer program is used to control the processor to execute the method according to the first aspect of the present disclosure.

[0030] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect of the present disclosure is implemented.

[0031] According to the video forgery detection method of the embodiment of the present disclosure, the electronic device first obtains the video to be detected, and extracts the target feature information of the video to be detected, the target feature information includes spatial feature information, temporal feature information and frequency feature information, and performs forgery detection on the video to be detected according to the target feature information. Through this embodiment, it combines the multimodal features of the video to be detected, namely the three dimensional features of spatial feature information, temporal feature information and frequency feature information to perform video forgery detection, which can improve the accuracy of video forgery detection.

[0032] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0034] Figure 1 is a schematic diagram of the hardware configuration of an electronic device according to an embodiment of the present disclosure;

[0035] Figure 2 is a flow chart of a video forgery detection method according to an embodiment of the present disclosure;

[0036] Figure 3 is a flowchart of a video forgery detection method according to an example of the present disclosure;

[0037] Figure 4 is a principle block diagram of a video forgery detection device according to an embodiment of the present disclosure;

[0038] Figure 5 is a schematic diagram of the hardware configuration of an electronic device according to another embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.

[0040] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0041] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.

[0042] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0043] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0044] <Hardware Configuration>

[0045] Figure 1 is a block diagram of a hardware configuration of an electronic device 1000 according to an embodiment of the present disclosure.

[0046] The electronic device 1000 may be a terminal device, a portable computer, a desktop computer, a server, a server cluster, etc. Figure 1 As shown, the electronic device 1000 may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, etc. Among them, the processor 1100 may be a processor CPU, a microprocessor MCU, etc. The memory 1200 includes, for example, a ROM (read-only memory), a RAM (random access memory), a non-volatile memory such as a hard disk, etc. The interface device 1300 includes, for example, a USB interface, a headphone interface, etc. The communication device 1400 is capable of wired or wireless communication, and specifically may include Wifi communication, Bluetooth communication, 2G / 3G / 4G / 5G communication, etc. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, etc. The input device 1600 may include, for example, a touch screen, a keyboard, etc. The user may input / output voice information through the speaker 1700 and the microphone 1800.

[0047] Figure 1 The electronic device shown is merely illustrative and does not in any way imply any limitation on the present disclosure, its application or use. In the embodiments of the present disclosure, the memory 1200 of the electronic device 1000 is used to store instructions, and the instructions are used to control the processor 1100 to operate to perform any one of the video forgery detection methods provided in the embodiments of the present disclosure. It should be understood by those skilled in the art that althoughFigure 1 In the electronic device 1000, multiple devices are shown, but the present disclosure may only involve some of the devices, for example, the electronic device 1000 only involves the processor 1100 and the storage device 1200. A technician can design instructions according to the scheme disclosed in the present disclosure. How instructions control the processor to operate is well known in the art, so it will not be described in detail here.

[0048] <Method Example>

[0049] In this embodiment, a video forgery detection method is provided. The video forgery detection method can be implemented by an electronic device. The electronic device can be as follows: Figure 1 The electronic device 1000 shown may be a terminal device, a portable computer, a desktop computer, a server, a server cluster, etc.

[0050] according to Figure 2 As shown, the video forgery detection method of the embodiment of the present disclosure may include the following steps S2100 to S2300.

[0051] Step S2100, obtaining the video to be detected.

[0052] The video to be detected is a video that needs to be detected for forgery. For example, the video to be detected can be a video that needs to be detected for face forgery. The video to be detected can be a real video or a forged video generated based on a deep learning forgery algorithm. The video forgery detection method of the embodiment of the present disclosure can detect forgery on any video to be detected.

[0053] After executing the above step S2100 to obtain the video to be detected, proceed to:

[0054] Step S2200: extracting target feature information of the video to be detected.

[0055] The target feature information includes spatial domain feature information, time domain feature information and frequency domain feature information.

[0056] The above-mentioned spatial feature information can be understood as intra-frame spatial information, which is used to perceive image texture abnormality information, image quality difference information, image facial sensory unnatural information, etc.

[0057] The above-mentioned time domain feature information can be understood as inter-frame time domain information, which is used to perceive visually unnatural features between image frames, such as facial position jitter, edge contour discontinuity and sudden changes, etc.

[0058] The above frequency domain feature information can be used to perceive the differences between the tampered audio and the real audio frequency domain in terms of high-frequency noise, naturalness, and timbre consistency.

[0059] In this embodiment, the electronic device can perform video forgery identification by combining three dimensional feature information, namely, spatial feature information, time feature information, and frequency feature information.

[0060] In an optional embodiment, the step S2200 of extracting the target feature information of the video to be detected may further include the following steps S2210 to S2240:

[0061] Step S2210: extracting a single frame image, a video clip and an audio clip of the video to be detected respectively.

[0062] Reference Figure 3 , the electronic device can extract single-frame images, video clips and audio clips from the video to be detected.

[0063] Step S2220: Obtain spatial feature information of the video to be detected based on the extracted single-frame image.

[0064] In this step S2220, when the electronic device obtains the spatial feature information of the video to be detected based on the extracted single-frame image, it can first input the extracted single-frame image into a preset first feature extraction network to obtain the spatial feature information corresponding to the single-frame image; and perform feature splicing and feature fusion on the spatial feature information corresponding to the single-frame image to obtain the spatial feature information of the video to be detected.

[0065] The preset first feature extraction network is used to extract spatial feature information of a single-frame image, that is, the input of the preset first feature extraction network is a single-frame image, and the output is the spatial feature information of the single-frame image. The preset first feature extraction network is a two-dimensional feature extraction network, and the preset first feature extraction network is a feature extraction layer based on a convolutional neural network.

[0066] Reference Figure 3 The electronic device inputs the extracted single-frame image into a preset first feature extraction network, outputs the spatial feature information of the single-frame image through the preset first feature extraction network, then splices the spatial feature information of multiple single-frame images, and inputs the spatial feature information into a long short-term memory (LSTM) network for feature fusion to obtain the spatial feature information of the video to be detected.

[0067] It should be noted that the purpose of feature fusion of spatial feature information based on LSTM network is to represent space in time.

[0068] Step S2230: obtaining time domain feature information of the video to be detected according to the extracted video clip.

[0069] Specifically, when the electronic device obtains the time domain feature information of the video to be detected based on the extracted video clip, it can input the extracted video clip into a preset second feature extraction network to obtain the time domain feature information corresponding to the video clip as the time domain feature information of the video to be detected.

[0070] The preset second feature extraction network is used to extract the time domain feature information of the video clip, that is, the input of the preset second feature extraction network is the video clip, and the output is the time domain feature information of the video clip. The preset second feature extraction network is a three-dimensional feature extraction network, and the preset second feature extraction network is a feature extraction layer based on a convolutional neural network.

[0071] Reference Figure 3 The electronic device can input the extracted video clip into a preset second feature extraction network, output the time domain feature information of the video clip through the preset second feature extraction network, and obtain the time domain feature information of the video to be detected.

[0072] It should be noted that since the time domain feature information itself is represented by the time dimension, there is no need to perform feature fusion based on the LSTM network.

[0073] Step S2240: Obtain frequency domain feature information of the video to be detected according to the extracted audio segment.

[0074] Specifically, when the electronic device obtains the frequency domain feature information of the video to be detected based on the extracted audio clip, it can first input the extracted audio clip into the third feature extraction network to obtain the frequency domain feature information corresponding to the audio clip; and perform feature splicing and feature fusion on the frequency domain feature information corresponding to the audio clip to obtain the frequency domain feature information of the video to be detected.

[0075] The preset third feature extraction network is used to extract spatial feature information of the audio clip, that is, the input of the preset third feature extraction network is the audio clip, and the output is the frequency domain feature information of the audio clip. The preset third feature extraction network is a two-dimensional feature extraction network, and the preset third feature extraction network is a feature extraction layer based on a convolutional neural network.

[0076] Reference Figure 3 The electronic device can input the extracted audio clip into a preset third feature extraction network, output the frequency domain feature information of the audio clip through the preset third feature extraction network, then splice the frequency domain feature information of multiple audio clips, and input them into the long short-term memory LSTM network for feature fusion to obtain the frequency domain feature information of the video to be detected.

[0077] It should be noted that the purpose of feature fusion of frequency domain feature information based on LSTM network is to represent the frequency domain in time.

[0078] After executing the above step S2200 to extract the target feature information of the video to be detected, proceed to:

[0079] Step S2300: performing forgery detection on the video to be detected according to the target feature information.

[0080] In this embodiment, when the electronic device performs forgery detection on the video to be detected based on the target feature information, it can fuse the spatial feature information, the time domain feature information and the frequency domain feature information to obtain fused feature information; and input the fused feature information into a preset forgery detection network for forgery detection.

[0081] Among them, the preset forgery detection network is used to perform forgery detection on the video to be detected. The input of the preset forgery detection network is the fusion feature information, and the output is the detection result of whether the video to be detected is a forged video.

[0082] It should be noted that the output detection result can not only indicate whether the video to be detected is a forged video, but also indicate the classification of the forged video.

[0083] For example, if the output detection results are: the image forgery detection result is "0", and the audio forgery detection result is "1", it indicates that the video to be detected is a forged video, and there is a forged image in the forged video.

[0084] For another example, if the output detection results are: the image forgery detection result is "1", and the audio forgery detection result is "0", it indicates that the video to be detected is a forged video, and there is forged audio in the forged video.

[0085] For another example, if the output detection results are: the image forgery detection result is "0", and the audio forgery detection result is "0", it indicates that the video to be detected is a forged video, and there are forged images and forged audio in the forged video.

[0086] For another example, if the output detection results are: the image forgery detection result is "1", and the audio forgery detection result is "1", it indicates that the video to be detected is a real video.

[0087] Reference Figure 3 The electronic device can fuse the features of the three dimensions to obtain fused feature information, and perform forgery detection through a preset video forgery detection network to output the detection result of the video to be detected.

[0088] According to the method of the embodiment of the present disclosure, the electronic device first obtains the video to be detected, and extracts the target feature information of the video to be detected, the target feature information includes spatial feature information, temporal feature information and frequency feature information, and performs forgery detection on the video to be detected according to the target feature information. Through this embodiment, it combines the multimodal features of the video to be detected, namely the three dimensional features of spatial feature information, temporal feature information and frequency feature information to perform video forgery detection, which can improve the accuracy of video forgery detection.

[0089] <Example>

[0090] Next, refer to Figure 3 , an example of a video forgery detection method is shown, in which example, the video forgery detection method may further include:

[0091] Step 1: Get the video to be detected.

[0092] Step 2: extract the single-frame image, video clip and audio clip of the video to be detected respectively.

[0093] Step 3: Input the extracted single-frame image into a preset first feature extraction network, output the spatial feature information of the single-frame image through the preset first feature extraction network, then perform feature splicing on the spatial feature information of multiple single-frame images, and input them into the LSTM network for feature fusion.

[0094] Step 4: input the extracted video clip into a preset second feature extraction network, and output the time domain feature information of the video clip through the preset second feature extraction network.

[0095] Step 5: Input the extracted audio clip into a preset third feature extraction network, output the frequency domain feature information of the audio clip through the preset third feature extraction network, then perform feature splicing on the frequency domain feature information of multiple audio clips, and input it into the LSTM network for feature fusion.

[0096] Step 6: Perform cross-modal feature fusion on the spatial domain feature information, the time domain feature information and the frequency domain feature information to obtain fused feature information, and input the fused feature information into a preset forgery detection network for forgery detection.

[0097] From this example, we can see that a heterogeneous network based on time-space-frequency domain feature learning is proposed. On the one hand, this heterogeneous network can not only suppress the interference of static background information, but also efficiently mine multi-domain inconsistent features. On the other hand, this heterogeneous network can fully consider the forgery characteristics of the three dimensions of intra-frame, inter-frame and frequency domain, effectively improving the generalization ability and robustness of the model across compression rates, forgery methods and modalities.

[0098] <Device Example>

[0099] In this embodiment, a video forgery detection device 400 is provided. Figure 4 As shown, the video forgery detection device 400 may include an acquisition module 410 , an extraction module 420 and a detection module 430 .

[0100] An acquisition module 410 is used to acquire a video to be detected;

[0101] The extraction module 420 is used to extract the target feature information of the video to be detected; wherein the target feature information includes spatial domain feature information, temporal domain feature information and frequency domain feature information;

[0102] The detection module 430 is used to perform forgery detection on the video to be detected according to the target feature information.

[0103] In one embodiment, the extraction module 420 is specifically used to extract the single-frame image, video clip and audio clip of the video to be detected respectively; obtain the spatial domain feature information of the video to be detected based on the extracted single-frame image; obtain the time domain feature information of the video to be detected based on the extracted video clip; obtain the frequency domain feature information of the video to be detected based on the extracted audio clip.

[0104] In one embodiment, the extraction module 420 is specifically used to input the extracted single-frame image into a preset first feature extraction network to obtain spatial feature information corresponding to the single-frame image; perform feature stitching and feature fusion on the spatial feature information corresponding to the single-frame image to obtain the spatial feature information of the video to be detected.

[0105] In one embodiment, the extraction module 420 is specifically configured to input the extracted video segment into a preset second feature extraction network to obtain time domain feature information corresponding to the video segment as the time domain feature information of the video to be detected.

[0106] In one embodiment, the extraction module 420 is specifically used to input the extracted audio segment into a preset third feature extraction network to obtain frequency domain feature information corresponding to the audio segment; perform feature splicing and feature fusion on the frequency domain feature information corresponding to the audio segment to obtain the frequency domain feature information of the video to be detected.

[0107] In one embodiment, the detection module 430 is specifically used to perform feature fusion on the spatial domain feature information, the time domain feature information and the frequency domain feature information to obtain fused feature information; and input the fused feature information into a preset forgery detection network for forgery detection.

[0108] According to this embodiment, the electronic device first obtains the video to be detected, and extracts the target feature information of the video to be detected, the target feature information includes spatial feature information, temporal feature information and frequency feature information, and performs forgery detection on the video to be detected according to the target feature information. Through this embodiment, it combines the multimodal features of the video to be detected, namely the three dimensional features of spatial feature information, temporal feature information and frequency feature information to perform video forgery detection, which can improve the accuracy of video forgery detection.

[0109] <Equipment Embodiment>

[0110] Figure 5 FIG. 1 is a schematic diagram of the hardware structure of an electronic device according to an embodiment. Figure 5 As shown, the electronic device 500 includes a processor 510 and a memory 520 .

[0111] The memory 520 may be used to store executable computer instructions.

[0112] The processor 510 may be configured to execute the video forgery detection method according to the method embodiment of the present disclosure under the control of the executable computer instructions.

[0113] The electronic device 500 may be Figure 1 The electronic device 1000 shown may also be a device with other hardware structures, which is not limited here.

[0114] In another embodiment, the electronic device 500 may include the above video forgery detection device 400 .

[0115] In one embodiment, each module of the above video forgery detection apparatus 400 may be implemented by the processor 510 executing computer instructions stored in the memory 520 .

[0116] <Computer Readable Storage Medium>

[0117] The embodiment of the present disclosure further provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the video forgery detection method provided by the embodiment of the present disclosure is executed.

[0118] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0119] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0120] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0121] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0122] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0123] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0124] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0125] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of an instruction, and the module, a program segment or a part of an instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of the boxes in the block diagram and / or the flowchart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by combining software and hardware.

[0126] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the marketplace, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present disclosure is defined by the appended claims.

Claims

1. A video forgery detection method, characterized in that: The method comprises: Get the video to be detected; Extracting target feature information of the video to be detected; wherein the target feature information includes spatial domain feature information, temporal domain feature information and frequency domain feature information; The video to be detected is subjected to forgery detection according to the target feature information.

2. The method according to claim 1, characterized in that The step of extracting target feature information of the video to be detected includes: Respectively extracting a single frame image, a video clip and an audio clip of the video to be detected; Obtaining spatial feature information of the video to be detected according to the extracted single-frame image; Obtaining time domain feature information of the video to be detected according to the extracted video clip; The frequency domain feature information of the video to be detected is obtained according to the extracted audio segment.

3. The method according to claim 2, characterized in that The step of obtaining spatial feature information of the video to be detected based on the extracted single-frame image includes: Inputting the extracted single-frame image into a preset first feature extraction network to obtain spatial domain feature information corresponding to the single-frame image; The spatial feature information corresponding to the single-frame image is subjected to feature splicing and feature fusion to obtain the spatial feature information of the video to be detected.

4. The method according to claim 2, characterized in that: The step of obtaining the time domain feature information of the video to be detected according to the extracted video clip includes: The extracted video clip is input into a preset second feature extraction network to obtain time domain feature information corresponding to the video clip as the time domain feature information of the video to be detected.

5. The method according to claim 2, characterized in that: The step of obtaining frequency domain feature information of the video to be detected according to the extracted audio segment includes: Inputting the extracted audio segment into a preset third feature extraction network to obtain frequency domain feature information corresponding to the audio segment; The frequency domain feature information corresponding to the audio clip is subjected to feature splicing and feature fusion to obtain the frequency domain feature information of the video to be detected.

6. The method according to claim 1, characterized in that The step of performing forgery detection on the video to be detected according to the target feature information includes: Performing feature fusion on the spatial domain feature information, the time domain feature information and the frequency domain feature information to obtain fused feature information; The fused feature information is input into a preset forgery detection network for forgery detection.

7. A video forgery detection device, characterized in that: The device comprises: An acquisition module is used to acquire the video to be detected; An extraction module, used to extract target feature information of the video to be detected; wherein the target feature information includes spatial domain feature information, temporal domain feature information and frequency domain feature information; The video to be detected is subjected to forgery detection according to the target feature information.

8. The device according to claim 7, characterized in that The extraction module is specifically used for: Respectively extracting a single frame image, a video clip and an audio clip of the video to be detected; Obtaining spatial feature information of the video to be detected according to the extracted single-frame image; Obtaining time domain feature information of the video to be detected according to the extracted video clip; The frequency domain feature information of the video to be detected is obtained according to the extracted audio segment.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store an executable computer program; and the computer program is used to control the processor to execute the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are executed by a processor to perform the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Micro-expression intelligent recognition system based on face image

    CN121074991A

  • A micro-expression intelligent recognition system based on facial images

    CN121074991B