Forgery video detection method and system

By dividing the video into video segments, the global and local features of the space-time plane layout are extracted, and the problem that the existing technology cannot fully cover the forged traces is solved, and effective capture and forgery detection of subtle time-time inconsistencies of videos is achieved.

CN120220030AInactive Publication Date: 2025-06-27INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510322798.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120220030A_ABST
    Figure CN120220030A_ABST
Patent Text Reader

Abstract

The invention discloses a forged video detection method and system. The method comprises the following steps: dividing a video to be detected into a plurality of video segments with equal duration; for each video segment in the plurality of video segments, selecting a predetermined number of image frames with continuous time from the current video segment; for each image frame in the preset number of image frames, intercepting an area of the target object from the current image frame to obtain an intercepted frame corresponding to the current image frame; splicing the preset number of intercepted frames clockwise according to a time sequence to obtain a space-time plane layout of the current video segment; scanning the space-time plane layout in multiple directions to obtain global space-time features of the current video segment; for each intercepted frame in the space-time plane layout, obtaining a local feature of the current intercepted frame based on the space-time features of other intercepted frames except the current intercepted frame in the space-time plane layout; and obtaining a forgery detection result based on the global spatial-temporal features of all the video segments and the local features of the intercepted frame of each video segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of computer vision technology, and more specifically, to a method and system for detecting forged videos. Background Art

[0002] With the development of artificial intelligence technology, especially the widespread application of the Internet and digital products, the acquisition, editing, and transmission of visual media materials such as videos and images have become more convenient. However, security issues such as identity forgery have become increasingly serious. In particular, the simulation degree of forged target objects (such as face images, eye images, etc.) and forged videos generated by artificial intelligence technologies such as deep learning is getting higher and higher, making it difficult for ordinary people to distinguish their authenticity, thus posing a huge threat to social trust and security. Therefore, it is particularly important to develop effective target object forgery detection technologies to address these risks.

[0003] The generated videos usually have both spatial and temporal artifacts simultaneously. Modeling only a single dimension (spatial or temporal) may not be able to comprehensively capture all types of artifacts. Some researchers have proposed separately learning the spatial information and temporal information of videos and then fusing the two. Such methods can effectively capture the multi-dimensional features of videos, but the computational complexity is relatively high. In addition, existing spatio-temporal modeling methods often focus on the extraction of global features rather than specifically targeting local details, which means that even if there are obvious inconsistencies in local regions, these methods may not be able to effectively capture this information. Some other researchers perform block processing on single-frame images to extract multiple non-overlapping local features. This method only extracts local features for each frame, focusing on learning spatial features and not considering the temporal information in the video, which may limit its forgery detection ability when dealing with dynamic videos. Recently, with the rise of the Mamba model, corresponding forgery detection methods have emerged. This model partitions within a single-frame image and scans each partition using a specific structured state space sequence model to extract local features of the image. However, the Mamba model does not explicitly consider the temporal relationship during the local feature extraction process but relies on subsequent global temporal modeling to process the information in the time dimension. This approach ignores the temporal dependence between local features and may lead to insufficient capture of the local detail inconsistencies in forged videos. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method and system for detecting forged videos, which can effectively solve the problem that the prior art cannot comprehensively cover all forgery traces and still has deficiencies in capturing the subtle spatio-temporal inconsistencies in videos.

[0005] In one general aspect, there is provided a forged video detection method, including: dividing a video to be detected into a plurality of video segments with equal time lengths; for each of the plurality of video segments, performing the following processing: selecting a predetermined number of consecutively-timed image frames from the current video segment; for each of the predetermined number of image frames, intercepting the region of the target object from the current image frame to obtain an intercepted frame corresponding to the current image frame; splicing the predetermined number of intercepted frames in chronological order in a clockwise manner to obtain the spatio-temporal plane layout of the current video segment; scanning the spatio-temporal plane layout in multiple directions to obtain the global spatio-temporal features of the current video segment; for each intercepted frame in the spatio-temporal plane layout, obtaining the local features of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame; and obtaining the forged detection result of the target object in the video to be detected based on the global spatio-temporal features of all video segments and the local features of the intercepted frames of each video segment.

[0006] Optionally, before obtaining the local features of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame, it further includes: scanning each intercepted frame in the spatio-temporal plane layout in multiple directions to obtain the total spatio-temporal features of the predetermined number of intercepted frames, where the total spatio-temporal features include the spatio-temporal features of each of the predetermined number of intercepted frames; and obtaining the local features of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame, including: obtaining the local features of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the total spatio-temporal features.

[0007] Optionally, obtaining the local features of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the total spatio-temporal features includes: removing the spatio-temporal features of the current intercepted frame from the total spatio-temporal features through mask processing to obtain the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame; and splicing the spatio-temporal features of other intercepted frames in chronological order to obtain the local features of the current intercepted frame.

[0008] Optionally, obtaining the forged detection result of the target object in the video to be detected based on the global spatio-temporal features of all video segments and the local features of the intercepted frames of each video segment includes: for each of all video segments, fusing the global spatio-temporal features of the video segment and the local features of the intercepted frames corresponding to the video segment to obtain the fused features of the video segment; and inputting the fused features of all video segments into a preset classifier to obtain the forged detection result of the target object in the video to be detected.

[0009] Optionally, the preset classifier can be trained in the following manner: for each video segment sample in a video sample to be detected, perform the following processing: based on the global spatio-temporal features of the current video segment sample, determine the cross-entropy loss function and the global information loss; based on the local features of each intercepted frame of the current video segment sample, determine the local information loss; perform a weighted sum of the cross-entropy loss, the global information loss, and the local information loss to obtain a multi-task loss; based on the multi-task losses of all video segment samples, adjust the parameters of the preset classifier.

[0010] Optionally, the weights of the cross-entropy loss, the global information loss, and the local information loss are determined in the following manner:

[0011] where represents the weight of any one of the cross-entropy loss, the global information loss, and the local information loss, and is proportional to the importance of the features included in any one of the losses.

[0012] In another general aspect, a forged video detection system is provided, including: a division unit configured to divide a video to be detected into a plurality of video segments of equal duration; a feature acquisition unit configured to, for each of the plurality of video segments, perform the following processing: select a predetermined number of image frames with continuous time from the current video segment; for each of the predetermined number of image frames, intercept the region of the target object from the current image frame to obtain an intercepted frame corresponding to the current image frame; splice the predetermined number of intercepted frames in chronological order in a clockwise manner to obtain the spatio-temporal plane layout of the current video segment; scan the spatio-temporal plane layout in multiple directions to obtain the global spatio-temporal features of the current video segment; for each intercepted frame in the spatio-temporal plane layout, based on the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame, obtain the local features of the current intercepted frame; a detection unit configured to obtain a forged detection result of the target object in the video to be detected based on the global spatio-temporal features of all video segments and the local features of the intercepted frames of each video segment.

[0013] Optionally, the feature acquisition unit is further configured to scan each intercepted frame in the spatio-temporal plane layout in multiple directions respectively to obtain the total spatio-temporal features of the predetermined number of intercepted frames, where the total spatio-temporal features include the spatio-temporal features of each of the predetermined number of intercepted frames; based on the spatio-temporal features of other intercepted frames in the total spatio-temporal features, obtain the local features of the current intercepted frame.

[0014] Optionally, the feature acquisition unit is further configured to remove the spatio-temporal features of the current intercepted frame from the total spatio-temporal features through mask processing, so as to obtain the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame; splice the spatio-temporal features of other intercepted frames in chronological order to obtain the local features of the current intercepted frame.

[0015] Optionally, the detection unit is further configured to, for each video segment in all video segments, fuse the global spatio-temporal features of the video segment and the local features of the intercepted frame corresponding to the video segment to obtain the fused features of the video segment; input the fused features of all video segments into a preset classifier to obtain the forgery detection result of the target object in the video to be detected.

[0016] Optionally, the detection unit is further configured to determine the cross-entropy loss function and the global information loss based on the global spatio-temporal features of the video segment; determine the local information loss based on the local features of each intercepted frame of the video segment; determine the weights of the cross-entropy loss, the global information loss, and the local information loss respectively based on the importance degrees of the global spatio-temporal features of the video segment and the local features of each intercepted frame of the video segment; perform weighted summation on the cross-entropy loss, the global information loss, and the local information loss based on the determined weights to obtain the multi-task loss; fuse the global spatio-temporal features of the video segment and the local features of each intercepted frame corresponding to the video segment based on the multi-task loss to obtain the fused features of the video segment.

[0017] Optionally, the weights of the cross-entropy loss, the global information loss, and the local information loss are determined in the following manner:

[0018] wherein, represents the weight of any one of the cross-entropy loss, the global information loss, and the local information loss, is proportional to the importance degree of the features included in any one of the losses.

[0019] In another general aspect, there is provided a computer-readable storage medium storing instructions, wherein when the instructions are run by at least one computing device, at least one computing device is prompted to execute any one of the forgery video detection methods as described above.

[0020] In another general aspect, there is provided a system including at least one computing device and at least one storage device storing instructions, wherein when the instructions are run by at least one computing device, at least one computing device is prompted to execute any one of the forgery video detection methods as described above.

[0021] In another general aspect, there is provided a computer program product including computer instructions, and when the computer instructions are executed by a processor, the forgery video detection method as described above is implemented.

[0022] According to an embodiment of the present disclosure, a forged video detection method and system splice the intercepted frames of a predetermined number of image frames of each video segment in a clockwise manner in chronological order to obtain a corresponding spatio-temporal plane layout, so that the original video time series information is transformed into a static spatial structure, thereby significantly reducing the computational amount without frame-by-frame processing; and, the present disclosure scans the spatio-temporal plane layout in multiple aspects to obtain the global spatio-temporal features of the current video segment, so that the spatio-temporal plane layout is converted into a multi-directional feature sequence, thereby realizing the extraction of global spatio-temporal features while maintaining a low computational complexity; furthermore, the present disclosure also obtains the local features of the current intercepted frame based on the spatio-temporal features of other image frames, so that not only the static information of a single frame can be captured, but also the dynamic association between frames can be established; finally, the global spatio-temporal features of the current video segment and the local features of the corresponding intercepted frame are fused for final forged video detection, so that all forged traces can be comprehensively covered and the capture ability of subtle inconsistent features in forged videos can be improved. Therefore, through the present disclosure, the problems that the prior art cannot comprehensively cover all forged traces and still has deficiencies in capturing subtle spatio-temporal inconsistencies in videos can be effectively solved.

[0023] Some other aspects and / or advantages of the general concept of the present disclosure will be described in part in the following description, and some will be clear from the description, or can be learned through the implementation of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Through the following description with reference to the drawings showing embodiments, the above and other objects and features of the embodiments of the present disclosure will become clearer, where: Figure 1 is a flowchart showing the forged video detection method according to an embodiment of the present disclosure; Figure 2 is a schematic diagram showing the sequential splicing of consecutive frames according to an embodiment of the present disclosure; Figure 3 is a schematic diagram showing the multi-selection state space module according to an embodiment of the present disclosure; Figure 4 is a schematic diagram showing the sequential three-frame local module according to an embodiment of the present disclosure; Figure 5 is a schematic diagram showing the dual-branch spatio-temporal plane network model according to an embodiment of the present disclosure; Figure 6 is a block diagram showing the forged video detection system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The following specific embodiments are provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, after understanding the disclosure of this application, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but rather may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a specific order. Additionally, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0026] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many viable ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of this application.

[0027] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more thereof.

[0028] Although terms such as "first," "second," and "third" may be used herein to describe various components, elements, regions, layers, or parts, these components, elements, regions, layers, or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, element, region, layer, or part from another. Thus, a first component, first element, first region, first layer, or first part referred to in the examples described herein may also be referred to as a second component, second element, second region, second layer, or second part without departing from the teachings of the examples.

[0029] In the specification, when an element (such as a layer, region, or substrate) is described as "on," "connected to," or "coupled to" another element, the element may be directly "on," directly "connected to," or "coupled to" the other element, or there may be one or more other elements therebetween. In contrast, when an element is described as "directly on," "directly connected to," or "directly coupled to" another element, there may be no other elements therebetween.

[0030] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising," "including," and "having" specify the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0031] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding this disclosure. Unless explicitly defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and shall not be interpreted in an idealized or overly formal manner.

[0032] In addition, in the description of the examples, when a detailed description of a related structure or function that is considered well-known would cause an ambiguous interpretation of this disclosure, such a detailed description will be omitted.

[0033] In recent years, with the development of generative artificial intelligence technology, the realism of video forgery has been continuously improved, posing a huge challenge to social trust and security. Some researchers have proposed multi-dimensional feature extraction methods that combine spatial and temporal information to capture forgery traces in videos; other researchers have adopted a local feature extraction approach, dividing the video into multiple non-overlapping regions to learn local features. However, there are still problems that cannot be ignored: separately modeling spatial or temporal features cannot comprehensively cover all forgery traces, and existing methods still have deficiencies in capturing subtle spatio-temporal inconsistencies in videos.

[0034] This disclosure proposes a dual-branch spatio-temporal feature extraction method, that is, using a global branch to extract global spatio-temporal features and using a local branch to extract local features. Specifically, the intercepted frames of a predetermined number of image frames of each video segment are spliced in a clockwise manner in chronological order to obtain the corresponding spatio-temporal plane layout; then, the global spatio-temporal features of the video segment are obtained by scanning the spatio-temporal plane layout in multiple directions, which can effectively extract the global spatio-temporal features of the video while maintaining linear complexity; moreover, the local features of the current intercepted frame are obtained through the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout, which can improve the ability to capture subtle inconsistent features in forged videos; finally, the global spatio-temporal features and local features are integrated for final detection to obtain the corresponding forgery detection results, improving the generalization performance on different datasets.

[0035] The forgery video detection method and system of this disclosure will be described in detail below with reference to the accompanying drawings.

[0036] This disclosure proposes a forgery video detection method, Figure 1 is a flowchart showing the forgery video detection method of the embodiments of this disclosure. Referring to Figure 1 , the forgery video detection method includes the following steps: In step S101, the video to be detected is divided into multiple video segments of equal duration.

[0037] As an example, after obtaining the video to be detected, the video to be detected can be divided into several equally long video segments in chronological order, so as to obtain multiple video segments of equal length in time. For example, given the video to be detected , where T is the number of frames of the video to be detected, C is the number of channels of the video to be detected, H and W are the height and width of the image frame in the video to be detected respectively. The given video to be detected can be divided into N equal segments, and the length of each segment is , which is convenient for subsequent processing. It should be noted that N is a positive integer.

[0038] In step S102, for each of the multiple video segments, the following processing is performed: select a predetermined number of consecutive image frames from the current video segment; for each of the predetermined number of image frames, intercept the region of the target object from the current image frame to obtain the intercepted frame corresponding to the current image frame; splice the predetermined number of intercepted frames in chronological order in a clockwise manner to obtain the spatio-temporal plane layout of the current video segment; scan the spatio-temporal plane layout in multiple directions to obtain the global spatio-temporal feature of the current video segment; for each intercepted frame in the spatio-temporal plane layout, based on the spatio-temporal features of the other intercepted frames except the current intercepted frame in the spatio-temporal plane layout, obtain the local feature of the current intercepted frame.

[0039] As an example, the above target object can be a human face, can be eyes, or can be other objects, and the present disclosure does not limit this.

[0040] As an example, taking the target object as a human face, for each video segment, several consecutive frames (such as 4 frames) can be randomly selected from it to capture the time series information of the video; for each image frame in the selected consecutive frames, the human face region can be cropped respectively, that is, the human face region is intercepted. For example, in the process of cropping the human face region, the MTCNN (Multi-task Cascaded Convolutional Neural Networks) face detection algorithm can be used, or other algorithms can also be adopted, as long as the human face region can be cropped, and the present disclosure does not limit this.

[0041] Specifically, taking the example of using the MTCNN face detection algorithm to crop faces, for the faces in each image frame of the above-mentioned consecutive frames, a large-area face bounding box can be extracted, that is, the length and width of the cropped face region are increased by 30% in size, so as to eliminate background interference and focus on facial features, thereby improving the detection accuracy. Specifically, the size from the center point to the border input to the MTCNN face detection algorithm is increased by 15%, so that the length and width of the cropped face region are both increased by 30% in size.

[0042] As an example, still taking the target object as a face, the intercepted face frames (i.e., the intercepted frames in the above-mentioned embodiments) are spliced in a clockwise manner in chronological order into a specific spatio-temporal plane layout. For example, when the consecutive frames contain t image frames, after the face frames corresponding to the t image frames are spliced, they are respectively the 1st frame, the 2nd frame, the 3rd frame... the tth frame when viewed in clockwise order; assuming t = 4, for the construction of the spatio-temporal plane layout in this embodiment, the face frames are spliced in a clockwise direction in chronological order to form a 2×2 image plane, as Figure 2 shown, converting the original video time series information into a static spatial structure, which not only avoids the computational overhead of frame-by-frame processing and significantly reduces the computational amount, but also enables the model to effectively capture the spatio-temporal dynamic features in the video.

[0043] Specifically, scale and splice the t face frames into a thumbnail layout, that is, the spatio-temporal plane layout, to ensure that the spatio-temporal information of consecutive frames is retained in a single layout. The size of the scaled face frames can be , then the size of the spliced spatio-temporal plane layout is the same as the size of the intercepted face frames, so that the size of the spatio-temporal plane layout meets the processing requirements of the model and avoids inaccurate subsequent detection.

[0044] It should be noted that the present disclosure can also splice the consecutive frames in a counterclockwise direction in chronological order, and the present disclosure does not limit this.

[0045] The spatio-temporal plane layout of the present disclosure can fuse the time information and spatial information of the video in a single image, thereby reducing the computational complexity and retaining the spatio-temporal consistency of the video frames. This layout enables the model to capture the features at multiple time points in one processing process by splicing consecutive frames into a static layout, avoiding the high computational overhead of frame-by-frame processing. At the same time, through face detection and background cropping, the spatio-temporal plane layout focuses on the facial area and eliminates the background interference information, ensuring that the model focuses on the key area of forgery detection. Generally speaking, the spatio-temporal plane layout not only improves the efficiency of the model, but also enhances the feature extraction effect, providing efficient and accurate data input for subsequent modules.

[0046] As an example, after obtaining the spatio-temporal plane layout, the global branch can be entered, that is, a multi-directional global scan of the spatio-temporal plane layout is performed to capture the spatio-temporal features of the spliced image frames from multiple angles, thereby generating rich global spatio-temporal features. For example, for the above multi-directional global scan, four-way scanning can be used. For example, the spatio-temporal plane layout can be scanned in the diagonal directions such as from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right. Each pixel can capture rich context information from different angles, enabling each intercepted frame to capture its own spatio-temporal relationship and share the feature information of other intercepted frames, thereby achieving effective extraction of global spatio-temporal features while maintaining a low computational complexity. Specifically, as Figure 3 shown, the global spatio-temporal features of the spatio-temporal plane layout can be extracted through a multi-selection state space module, that is, the spatio-temporal plane layout is input into the multi-selection state space module, and the multi-selection state space module uses a four-way scanning strategy (from top left to bottom right, from bottom right to top left, from top right to bottom left, from bottom left to top right) to perform a multi-directional scan on the spatio-temporal plane layout, and then the global spatio-temporal feature F corresponding to the spatio-temporal plane layout can be output global .

[0047] The above multi-selection state space model can include a multi-stage Mamba feature extractor, a visual state space module, a four-way scanning sub-module, and a multi-directional feature fusion sub-module. Among them, the multi-stage Mamba feature extractor combines with the visual state space module to perform layer-by-layer extraction and preservation of the spatio-temporal plane layout. Different from traditional methods that only retain the features of the last layer, the present disclosure can obtain richer spatial and temporal information from each level, which helps to identify subtle inconsistencies in forged videos. During the multi-stage feature extraction process, the visual state space module is used to integrate the multi-directional features of the intermediate layer of the model generated by the four-way scanning sub-module; the four-way scanning sub-module performs a four-way scan on each pixel in the spatio-temporal plane layout through a 2D selective scanning mechanism to capture its context information from multiple directions, so as to improve the multi-angle coverage of global features and the comprehensiveness of detection; the multi-directional feature fusion sub-module integrates the multi-directional features of the last layer of the model generated by the four-way scanning sub-module, that is, through weighted fusion of different direction features, it can reasonably balance the information in each direction during the integration process, making the extracted features more complete and expressive, thereby obtaining unified global spatio-temporal features.

[0048] The above 2D selective scanning mechanism gradually scans the spatio-temporal plane from four directions (such as from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right), enabling each pixel to integrate spatio-temporal information in different directions, thereby enhancing the diversity and integrity of feature expression while maintaining a low computational complexity.

[0049] The above multi-choice state space module can be implemented based on Mamba's structured state space model. The structured state space sequence model based on Mamba maps a one-dimensional function or sequence through a continuous system to the output , and introduces an intermediate hidden state . This process represents the input data through ordinary differential equations, where the change of the hidden state is described by the following formula:

[0050] where is the evolution matrix of the continuous system, and and are the projection matrices of the continuous system. To convert Mamba's structured state space sequence model from a continuous model to a discretized model, the Mamba framework utilizes the time scale parameter ∆, and converts the matrices A and B to their discrete equivalent forms A and B , obtaining the discrete update equation:

[0051] and finally obtaining:

[0052] It should be noted that traditional structured state space models are usually limited to traversing in a single direction during image processing, resulting in limited ability to capture global spatio-temporal features. The multi-choice state space module can perform multi-directional scanning on the input spatio-temporal plane layout in four directions (from top left to bottom right, from bottom right to top left, from top right to bottom left, from bottom left to top right), converting the image information into multiple sequences. This multi-directional scanning strategy ensures that each pixel can integrate its own spatio-temporal relationship from different angles and share the feature information of other frames, thereby generating a feature representation with global perception.

[0053] As an example, for the intercepted frames in the spatio-temporal plane layout, the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout can be utilized to generate the local features with temporal dependence of the current intercepted frame. For example, the local features of each intercepted frame can be obtained through masking and splicing operations, and the present disclosure does not limit this.

[0054] According to an embodiment of the present disclosure, before obtaining the local feature of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame, each intercepted frame in the spatio-temporal plane layout is scanned in multiple directions respectively to obtain the total spatio-temporal feature of a predetermined number of intercepted frames, where the total spatio-temporal feature includes the spatio-temporal features of each intercepted frame among the predetermined number of intercepted frames. In this case, obtaining the local feature of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame may include: obtaining the local feature of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the total spatio-temporal feature. Through this embodiment, each intercepted frame is also scanned in multiple directions to ensure that each frame can integrate spatio-temporal features in different directions at the local level.

[0055] As an example, each frame in the spatio-temporal plane layout is individually scanned in multiple directions through a multi-selection state space module to obtain local spatio-temporal information of each frame from multiple directions. It should be noted that the multi-selection state space module can separately scan each frame in multiple directions through the position of each frame in the spatio-temporal plane layout, so as to obtain the total spatio-temporal feature of the spatio-temporal plane layout, that is, the total spatio-temporal feature in which the spatio-temporal features of all frames in the spatio-temporal plane layout are arranged according to the position of the spatio-temporal plane layout, and the total spatio-temporal feature includes the independent spatio-temporal features of each frame.

[0056] As an example, the above local branch operation can be implemented by a sequence of three-frame local module, as Figure 4 shown, the sequence of three-frame local module inputs the spatio-temporal plane layout into the multi-selection state space module, and the module processes each frame by four-way scanning, enabling each pixel to obtain its own context information from multiple directions, ensuring that each frame can integrate spatio-temporal features from different directions at the local level, and then obtaining the local feature of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the total spatio-temporal feature.

[0057] According to an embodiment of the present disclosure, obtaining the local feature of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the total spatio-temporal feature may include: removing the spatio-temporal feature of the current intercepted frame from the total spatio-temporal feature through mask processing to obtain the spatio-temporal features of other intercepted frames in the spatio-temporal plane layout except the current intercepted frame; splicing the spatio-temporal features of other intercepted frames in chronological order to obtain the local feature of the current intercepted frame. Through this embodiment, when generating the local feature for each frame, only the information of the remaining three frames is retained, which can enhance the dependence on temporal information, eliminate the influence of the current intercepted frame on the temporal features, and then focus on the temporal features between the remaining frames, strengthening the modeling of the dynamic relationship between frames; moreover, splicing the spatio-temporal features of other intercepted frames except the current intercepted frame in chronological order as the local feature of the current intercepted frame makes the local feature contain both the static features of each remaining frame and effectively integrates the dynamic associations between frames.

[0058] As an example, such as Figure 4 shown, after obtaining the total spatio-temporal features, for each frame, the three-frame local module of the sequence excludes the current frame from the spatio-temporal plane layout through a masking operation, and only retains the temporal information of the remaining three frames to remove the interference of the current frame on the temporal features, so as to focus on the temporal information between frames. Then, the spatio-temporal features of the remaining three frames are concatenated in chronological order to generate local features with temporal dependence, thereby ensuring that the local branch can not only capture the static information of a single frame, but also capture the dynamic relationship between frames. Finally, the local features F corresponding to each frame are generated in strict chronological order local,i , significantly enhancing the model's detection ability for local detail inconsistencies in video forgeries and enhancing the overall accuracy and robustness of forgery detection.

[0059] It should be noted that the three-frame local module of the sequence successfully enhances the model's ability to capture the temporal correlation and dynamic changes of local features without significantly increasing the computational complexity. This module uses frame-by-frame masking and temporal concatenation operations to ensure that each local feature contains both the spatial information of a single frame and effectively reflects the dynamic correlation between frames, thereby enhancing the temporal sensitivity of the model in forgery detection. This module design also effectively strengthens the model's detection ability for subtle forgery traces in videos, providing a solid technical support for improving the overall detection accuracy and robustness.

[0060] In step S103, based on the global spatio-temporal features of all video segments and the local features of the intercepted frames of each video segment, the forgery detection result of the target object in the video to be detected is obtained.

[0061] According to the embodiments of the present disclosure, based on the global spatio-temporal features of all video segments and the local features of the intercepted frames of each video segment, obtaining the forgery detection result of the target object in the video to be detected may include: for each video segment among all video segments, fusing the global spatio-temporal features of the video segment and the local features of the intercepted frame corresponding to the video segment to obtain the fused features of the video segment; inputting the fused features of all video segments into a preset classifier to obtain the forgery detection result of the target object in the video to be detected. Through this embodiment, fusing the global spatio-temporal features and local features can improve the classification performance and generalization performance of video forgery detection.

[0062] As an example, splicing technology can be used to fuse the global spatio-temporal features and local features together to form the final discriminant feature. Then, the final discriminant feature is input into a preset classifier for classification tasks to obtain the forgery detection result of the target object in the video to be detected.

[0063] As an example, the preset classifier can be Swin Transformer, CNN, Xception, etc., and the present disclosure does not limit this.

[0064] It should be noted that the above preset classifier needs to be trained in advance. The specific training process can rely on a multi-task loss function, and the multi-task loss function can include multiple loss functions, such as a cross-entropy loss function , a global information loss function , and a local information loss function , etc., and the present disclosure does not limit this.

[0065] The training process of the preset classifier will be introduced separately below: According to an embodiment of the present disclosure, the preset classifier can be trained in the following manner: for each video segment sample in a video sample to be detected, the following processing is performed: based on the global spatio-temporal features of the current video segment sample, determine the cross-entropy loss function and the global information loss; based on the local features of each intercepted frame of the current video segment sample, determine the local information loss; perform weighted summation on the cross-entropy loss, the global information loss, and the local information loss to obtain a multi-task loss; based on the multi-task losses of all video segment samples, adjust the parameters of the preset classifier. Through the design of the multi-task loss function in this embodiment, different requirements for global spatio-temporal features and local features can be balancedly met, thereby improving the overall performance of the trained preset classifier for video forgery detection.

[0066] As an example, after obtaining the global spatio-temporal features and local features of each video segment in a video to be detected, for any video segment, first, input the global spatio-temporal features of the video segment into the preset classifier for processing to generate a predicted forgery detection result corresponding to the global spatio-temporal features, and additionally input the local features of the video segment into the preset classifier for processing to generate a predicted forgery detection result corresponding to the local features; then, according to these two predicted forgery detection results, combined with the multi-task loss function obtain the multi-task loss. Specifically, the multi-task loss function can include the following three main loss terms: the global information loss function, the local information loss function, and the cross-entropy loss function. The three loss terms will be introduced separately below.

[0067] Global information loss function : It is used to maintain the information integrity of the global spatio-temporal features during the prediction process, that is, calculate the information loss between the original global spatio-temporal features and the compressed global spatio-temporal features (i.e., the global spatio-temporal features after dimensionality reduction processing). This loss term measures the distribution difference between the original global spatio-temporal features and the compressed global spatio-temporal features through KL divergence, so as to ensure that important global information is not lost after dimensionality reduction processing. Global information loss function It can be expressed as follows:

[0068] Wherein, represents the predicted value obtained by predicting the original global spatio-temporal features, represents the predicted value obtained by predicting the compressed global spatio-temporal features, KL represents the Kullback-Leibler divergence, and T is the temperature parameter that controls the smoothness of the softmax output.

[0069] Local information loss function : It is used to maintain the independence of local features and the integrity of information, that is, to calculate the similarity between pairwise local features among multiple local features. The local features generated for each frame in the local branch should be unique to reduce the impact of redundant information on the overall classification. This loss term calculates the mutual information of the local features for each frame to ensure the temporal consistency and feature independence of the local features. Local information loss function It can be expressed as follows:

[0070] Wherein, , are the predicted forgery detection results corresponding to pairwise local features.

[0071] Cross-entropy loss function : It is used to measure the classification accuracy of the finally fused features. In the model, after fusing the compressed global features and local features, the fused features are obtained, and these fused features pass through a linear layer (or other classifier) to obtain the predicted class probability distribution. By calculating the difference between the class probability distribution predicted by the model and the true label y, it is ensured that the model can correctly distinguish between real videos and forged videos after fusing global and local features. Cross-entropy loss function It can be expressed as follows:

[0072] Wherein, represents the probability of the th class in the true label is the probability of the th class output by the preset classifier; N is the total number of classes. It should be noted that in this disclosure, a binary classification task of videos is performed. When detecting a video to be detected, the th class refers to the class in the classification task. For example, the class can be "real" or "forged", N is the total number of classes, and if the class is "real" or "forged", then N = 2; Refers to the one-hot encoded probability of the "real" category (such as [1, 0] or [0, 1]).

[0073] However, the multi-task loss function can be calculated as follows:

[0074] Wherein, 、 、 are the weight coefficients of each loss term.

[0075] After obtaining the multi-task loss corresponding to each video segment, the parameters of the preset classifier can be adjusted according to the multi-task loss to achieve the training of the preset classifier.

[0076] It should be noted that the weights of each loss function in the above multi-task loss function can be set as needed, or an adaptive weighted fusion mechanism can be used to determine them. The present disclosure does not limit this.

[0077] As an example, using an adaptive weighted fusion mechanism to determine the weight coefficients of each loss term is equivalent to providing a fusion mechanism for loss functions. This mechanism will ultimately learn a weight for the loss function, which will affect the learning process of the classifier for global spatio-temporal features and local features, and then optimize the ability of the classifier to extract features. Loss fusion has balanced the learning weights of global spatio-temporal features and local features, so that the classifier will not overly rely on a certain feature (global spatio-temporal feature or local feature) during the training process. In this way, in the final classification stage, the global spatio-temporal features and local features can be directly concatenated, and the concatenated features can be directly input into the classifier for classification.

[0078] The following introduces the process of using an adaptive weighted fusion mechanism to determine the weight of each loss function in the multi-task loss function: In the adaptive fusion mechanism, the contributions of each loss term to the overall loss are different under different video segments, that is, under different video segments, the weights of each loss function in the multi-task loss function can be dynamically adjusted, so as to dynamically adjust the fusion weights of global spatio-temporal features and local features, balance their influences in the decision-making process, and finally determine the optimal fusion weights of each loss term in the multi-task loss function. The preset classifier is trained through the multi-task loss function corresponding to the optimal fusion weights, so that the trained preset classifier can evenly meet the different requirements for global spatio-temporal features and local features, thereby improving the detection accuracy and generalization ability of the preset classifier for different forgery methods.

[0079] Specifically, the adaptive weighted fusion mechanism may include a multi-task loss function and a dynamic weight adjustment module for optimizing the detection effect in different datasets and scenarios; and the dynamic weight adjustment module may use the AutomaticWeightedLoss function to dynamically adjust the weights of each loss term, and this function has an adaptive weight mechanism.

[0080] According to an embodiment of the present disclosure, the weights of the cross-entropy loss, the global information loss, and the local information loss can be determined in the following manner:

[0081] Wherein, represents the weight of any one of the cross-entropy loss, the global information loss, and the local information loss, and is proportional to the importance of the features included in any one loss. Through this formula, the weights of the individual loss functions can be dynamically adjusted, so that the global spatio-temporal features and local features can be better fused.

[0082] As an example, initially, the weights of all loss terms are equal and can be initialized to 1. As the adaptive mechanism progresses, during backpropagation, the weights of each loss term will be automatically adjusted according to the gradients of the global spatio-temporal features and local features in the preset classifier respectively, as well as the changes of each loss term. That is, the importance of the features included in a loss is determined based on the magnitude of the gradient of the features and the change of the loss. For example, if a loss term (such as the global information loss) gradually decreases during training, and the gradient of the global spatio-temporal features in the preset classifier is large, it indicates that the global spatio-temporal features contribute more to the task. In other words, this loss term has "converged well". Therefore, the weight of this loss term in the multi-task loss function may be reduced, that is, the classifier believes that the relevant information of this loss term has been effectively learned, and this loss term has played a great role. Therefore, this loss term does not need to occupy too much model resources (such as a larger weight). At this time, by reducing the weight of the global information loss function in the multi-task loss function, other features (such as local features) can be given more attention, further improving the performance. On the contrary, if the loss of a certain loss term is large, the weight of this loss term in the multi-task loss function will be increased to strengthen the influence of this loss term.

[0083] Specifically, this embodiment introduces a learnable parameter to dynamically adjust the weight coefficient , and the weight coefficient is calculated by the following formula:

[0084] This formula ensures that the weight coefficient It can be adaptively adjusted according to the requirements of the loss terms during training. For important loss terms, by adjusting the value to increase its weight, it is ensured that the proportion of the corresponding information in the decision-making is enhanced. For example, when local detailed information is more important, the weight of the local information loss is correspondingly increased, so as to flexibly meet the requirements of different forgery features.

[0085] Finally, the classifier can find the optimal fusion weights of each loss term. Through these optimal fusion weights, the training of the above-mentioned preset classifier is carried out, making the training process more efficient and obtaining a preset classifier with better results.

[0086] To facilitate the understanding of the above embodiments, the following is combined with Figure 5 for a systematic description.

[0087] Figure 5 Figure 14 shows a dual-branch spatio-temporal plane network model. The model internally includes two branches, namely a global branch and a uniform branch. The model includes a splicing spatio-temporal plane layout module, a multi-choice state space module based on Mamba, and a sequential three-frame local module.

[0088] The splicing spatio-temporal plane layout module divides the video to be detected into multiple video segments of equal duration. For each video segment, several consecutive frames can be randomly selected from it (for example, 4 frames can be selected) to capture the time series information of the video; for each image frame in the selected consecutive frames, the face region can be respectively cropped to obtain the corresponding face frame (i.e., the intercepted frame in the above embodiments); the intercepted face frames are spliced into a specific spatio-temporal plane layout in a clockwise manner according to the time sequence. For example, when the consecutive frames contain 4 image frames, the face frames corresponding to the 4 image frames are spliced to form a 2×2 image plane, which are the 1st frame, the 2nd frame, the 3rd frame, and the 4th frame in clockwise order.

[0089] The multi-choice state space module is used to extract the global features of each video segment. For each video segment, the video segment is input into the multi-choice state space module to obtain the global spatio-temporal features of the video segment. Specifically, the multi-choice state space module adopts a four-way scanning strategy (from top left to bottom right, from bottom right to top left, from top right to bottom left, from bottom left to top right) to perform multi-directional scanning on the corresponding spatio-temporal plane layout to obtain the global spatio-temporal feature F global .

[0090] The three-frame local module of the sequence includes parts such as frame-by-frame input processing, masking operation, and temporal stitching of local features. First, in the frame-by-frame input processing part, that is, the three-frame local module of the sequence performs multi-directional scanning on each frame in the spatio-temporal plane layout through a multi-selection state space module, and obtains the local spatio-temporal features of each frame from multiple directions, that is, the total spatio-temporal features, including the spatio-temporal features of each of the 4 frames; secondly, in the masking operation part, that is, the three-frame local module of the sequence applies masking processing to each frame in the spatio-temporal plane layout, excluding the current frame and only retaining the spatio-temporal features of the other three frames to remove the interference of the current frame on the temporal features, as Figure 5 described within the dashed box, the masking result of each frame is the spatio-temporal features of the other three frames; finally, in the part of temporal stitching of local features, that is, the spatio-temporal features of the three remaining frames are stitched in chronological order to generate the local feature F local,i , to capture the dynamic correlation between frames, thereby generating local features with temporal dependence, improving the model's ability to capture the dynamic relationship between video frames, and thus helping to detect subtle forgery traces in the video.

[0091] In summary, the present disclosure proposes a dual-branch face forgery detection method guided by a multi-selection state space, which combines a global branch and a local branch to effectively extract the global spatio-temporal features and local features in the video to improve the accuracy of forgery video detection; moreover, a multi-selection state space module is also used to scan the spatio-temporal plane layout from multiple directions to capture rich global spatio-temporal features; furthermore, the three-frame local module of the sequence is also combined to enhance the dynamic relationship between frames through masking and temporal stitching operations, and fully capture the inconsistencies in local details. Finally, the global spatio-temporal features and local features are dynamically fused through an adaptive weighted fusion mechanism to balance the roles of the two in forgery video detection, effectively improving the classification performance and generalization ability of the model, and achieving high-precision classification detection of forgery videos.

[0092] Figure 6 is a block diagram showing a forgery video detection system according to an embodiment of the present disclosure, as Figure 6 shown, the system includes a partitioning unit 60, a feature acquisition unit 62, and a detection unit 64.

[0093] The division unit 60 is configured to divide the video to be detected into multiple video segments with equal durations; the feature acquisition unit 62 is configured to perform the following processing for each video segment among the multiple video segments: select a predetermined number of temporally consecutive image frames from the current video segment; for each image frame among the predetermined number of image frames, intercept the region of the target object from the current image frame to obtain an intercepted frame corresponding to the current image frame; splice the predetermined number of intercepted frames in a clockwise manner according to the time order to obtain the spatio-temporal plane layout of the current video segment; scan the spatio-temporal plane layout in multiple directions to obtain the global spatio-temporal feature of the current video segment; for each intercepted frame in the spatio-temporal plane layout, obtain the local feature of the current intercepted frame based on the spatio-temporal features of other intercepted frames except the current intercepted frame in the spatio-temporal plane layout; the detection unit 64 is configured to obtain the forgery detection result of the target object in the video to be detected based on the global spatio-temporal features of all video segments and the local features of the intercepted frames of each video segment.

[0094] According to an embodiment of the present disclosure, the feature acquisition unit 62 is further configured to scan each intercepted frame in the spatio-temporal plane layout in multiple directions respectively to obtain the total spatio-temporal feature of the predetermined number of intercepted frames, where the total spatio-temporal feature includes the spatio-temporal features of each intercepted frame among the predetermined number of intercepted frames; and obtain the local feature of the current intercepted frame based on the spatio-temporal features of other intercepted frames in the total spatio-temporal feature.

[0095] According to an embodiment of the present disclosure, the feature acquisition unit 62 is further configured to remove the spatio-temporal feature of the current intercepted frame from the total spatio-temporal feature through mask processing to obtain the spatio-temporal features of other intercepted frames except the current intercepted frame in the spatio-temporal plane layout; and splice the spatio-temporal features of other intercepted frames in time order to obtain the local feature of the current intercepted frame.

[0096] According to an embodiment of the present disclosure, the detection unit 64 is further configured to, for each video segment among all video segments, fuse the global spatio-temporal feature of the video segment and the local feature of the intercepted frame corresponding to the video segment to obtain the fusion feature of the video segment; and input the fusion features of all video segments into a preset classifier to obtain the forgery detection result of the target object in the video to be detected.

[0097] According to an embodiment of the present disclosure, the above device further includes a training unit configured to train the preset classifier in the following manner: for each video segment sample in a video sample to be detected, perform the following processing: determine the cross-entropy loss function and the global information loss based on the global spatio-temporal feature of the current video segment sample; determine the local information loss based on the local feature of each intercepted frame of the current video segment sample; perform a weighted sum of the cross-entropy loss, the global information loss, and the local information loss to obtain a multi-task loss; and adjust the parameters of the preset classifier based on the multi-task loss of all video segment samples.

[0098] According to an embodiment of the present disclosure, the weights of the cross-entropy loss, the global information loss, and the local information loss are determined as follows:

[0099] wherein, represents the weight of any one of the cross-entropy loss, the global information loss, and the local information loss, and is proportional to the importance of the features included in any one of the losses.

[0100] According to an embodiment of the present disclosure, there is provided a computer-readable storage medium storing instructions, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the forged video detection method according to any one of the above embodiments.

[0101] According to an embodiment of the present disclosure, there is provided a system including at least one computing device and at least one storage device storing instructions, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the forged video detection method according to any one of the above embodiments. Such a system can be used for real-time image processing in an industrial production environment.

[0102] According to an embodiment of the present disclosure, there is provided a computer program product including computer instructions, which implement the forged video detection method as described above when executed by a processor.

[0103] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure as defined by the claims and their equivalents.

Claims

1. A method for detecting forged videos, characterized in that: include: Divide the video to be detected into multiple video segments of equal length; For each video segment in the plurality of video segments, the following processing is performed: Selecting a predetermined number of image frames that are continuous in time from the current video segment; For each image frame in the predetermined number of image frames, intercepting a region of the target object from a current image frame to obtain an intercepted frame corresponding to the current image frame; splicing the predetermined number of intercepted frames in a clockwise manner in chronological order to obtain a spatiotemporal plane layout of the current video segment; Scanning the spatiotemporal plane layout in multiple directions to obtain global spatiotemporal features of the current video segment; For each intercepted frame in the space-time plane layout, based on the space-time features of other intercepted frames in the space-time plane layout except the current intercepted frame, obtain the local features of the current intercepted frame; Based on the global spatiotemporal features of all video segments and the local features of the captured frames of each video segment, a forgery detection result of the target object in the video to be detected is obtained.

2. The forged video detection method according to claim 1, characterized in that: Before obtaining the local features of the current intercepted frame based on the spatiotemporal features of other intercepted frames except the current intercepted frame in the spatiotemporal plane layout, the method further includes: Scanning each intercepted frame in the space-time plane layout in multiple directions to obtain the total space-time features of the predetermined number of intercepted frames, wherein the total space-time features include the space-time features of each intercepted frame in the predetermined number of intercepted frames; Wherein, the obtaining of the local features of the current intercepted frame based on the spatiotemporal features of other intercepted frames except the current intercepted frame in the spatiotemporal plane layout includes: Based on the spatiotemporal features of other captured frames in the total spatiotemporal features, local features of the current captured frame are obtained.

3. The forged video detection method according to claim 2, characterized in that: The obtaining of the local features of the current intercepted frame based on the spatiotemporal features of other intercepted frames in the total spatiotemporal features includes: Removing the spatiotemporal features of the current intercepted frame from the total spatiotemporal features by mask processing to obtain the spatiotemporal features of other intercepted frames in the spatiotemporal plane layout except the current intercepted frame; The spatiotemporal features of the other captured frames are spliced ​​in time sequence to obtain the local features of the current captured frame.

4. The forged video detection method according to claim 1, characterized in that: The method of obtaining a forgery detection result of a target object in the video to be detected based on the global spatiotemporal features of all video segments and the local features of the captured frames of each video segment includes: For each video segment in all video segments, fusing the global spatiotemporal features of the video segment with the local features of the captured frames corresponding to the video segment to obtain a fused feature of the video segment; The fused features of all video segments are input into a preset classifier to obtain a forgery detection result of the target object in the video to be detected.

5. The forged video detection method according to claim 4, characterized in that: The preset classifier can be trained in the following way: For each video segment sample in a video sample to be detected, perform the following processing: Determining a cross entropy loss function and a global information loss based on the global spatiotemporal characteristics of the current video segment sample; Determining local information loss based on local features of each captured frame of the current video segment sample; Performing a weighted summation on the cross entropy loss, the global information loss, and the local information loss to obtain a multi-task loss; Based on the multi-task loss of all video segment samples, the parameters of the preset classifier are adjusted.

6. The forged video detection method according to claim 5, characterized in that: The weights of the cross entropy loss, the global information loss and the local information loss are determined in the following manner: in, represents the weight of any one of the cross entropy loss, the global information loss and the local information loss, It is proportional to the importance of the features included in any of the losses.

7. A forged video detection system, characterized in that: include: A division unit, configured to divide the video to be detected into a plurality of video segments of equal duration; The feature acquisition unit is configured to perform the following processing for each of the multiple video segments: selecting a predetermined number of image frames that are continuous in time from the current video segment; for each of the predetermined number of image frames, intercepting a region of the target object from the current image frame to obtain an intercepted frame corresponding to the current image frame; splicing the predetermined number of intercepted frames in a clockwise manner in chronological order to obtain a spatiotemporal plane layout of the current video segment; Scanning the spatiotemporal plane layout in multiple directions to obtain global spatiotemporal features of the current video segment; For each intercepted frame in the space-time plane layout, based on the space-time features of other intercepted frames in the space-time plane layout except the current intercepted frame, obtain the local features of the current intercepted frame; The detection unit is configured to obtain a forgery detection result of the target object in the video to be detected based on the global spatiotemporal features of all video segments and the local features of the captured frames of each video segment.

8. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one computing device, the at least one computing device is prompted to perform the forged video detection method according to any one of claims 1 to 6.

9. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that: When the instructions are executed by the at least one computing device, the at least one computing device is prompted to perform the forged video detection method according to any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the forged video detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video detection method, device, equipment, medium and product

    CN116958846A

  • Space-time combination detection method, device and equipment for deeply-forged video

    CN116994175A

  • Adaptive scene monitoring video target detection method and system based on feature fusion

    CN117237867A

  • Deep counterfeit video detection method and device based on time sequence difference

    CN117496392A