A method and system for generating an intraframe tampered video

By acquiring the background layer and mask sequence of video clips and using Adobe Effect to adjust brightness and contrast, high-quality tampered videos are generated, solving the problems of high cost and low quality in existing technologies and realizing efficient pre-training of video tampering localization models.

CN119603434BActive Publication Date: 2025-12-05SHENZHEN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411574227.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-12-05
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

Existing technologies for generating intra-frame tampered videos are costly and of low quality, requiring significant manpower and time for the production and post-processing of tampered video samples.

Method used

By acquiring the background layer and mask sequence of the original video clips, and using Adobe Effect software for preprocessing and brightness and contrast adjustments, a high-quality tampered video is generated.

Benefits of technology

It significantly reduces labor and time costs, provides abundant high-quality data for pre-training video tampering localization models, and improves the efficiency and accuracy of tampered video detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119603434B_ABST
    Figure CN119603434B_ABST
Patent Text Reader

Abstract

The application discloses an intra-frame tampered video generation method and system, comprising the following steps: cutting a given original video into original video segments, obtaining background layers of each segment, and performing preprocessing; obtaining a mask sequence of a tampered target from each segment; obtaining a corresponding foreground layer of each segment according to the mask sequence, and calculating a brightness value and a contrast difference according to the foreground layer and the background layer to obtain the brightness value and the contrast difference value of each frame of video in each segment; using an automatic tracking tool to obtain mask region attributes corresponding to the mask sequence, and completely copying the mask region attributes from an adjustment layer to an original video layer; performing brightness and contrast adjustment processing on the mask region of the original video layer to obtain a post-processed tampered video. The application can provide rich high-quality data for pre-training of a video tampering positioning model, greatly reduce artificial and time costs, and has important application value in video tampering positioning evidence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video tampering, in particular to an intra-frame tampered video generation method and system. BACKGROUND

[0002] Video plays an important role in public safety, social media, data transmission and other fields. With the development of video editing technology, video tampering problems have become increasingly prominent. With the help of video editing software, ordinary people can also tamper with videos at a low learning cost. Adobe Effect is the most widely used video special effect making software. And under the premise of sufficient time, tamperers modify the content of each frame of the video and post-process it, so that the tampered video will not leave obvious visual tampering traces. In the era of self-media, such tampered videos can easily spread to the public through the Internet, and the uninformed will inevitably be misled, and even cause serious social impact.

[0003] Intra-frame tampering refers to modifying the content of a frame in a video, adding or removing objects, changing the color or texture of objects, etc. in a frame of image. However, in the actual intra-frame tampering detection task of the video, a large number of tampered videos are required as reference samples, which not only need to be modified frame by frame, but also need to be post-processed, and the tampering quality needs to be focused on, which requires a lot of manpower and time cost.

[0004] Therefore, the prior art still needs to be improved. SUMMARY

[0005] The technical problem to be solved by the present application is that, in view of the defects of the prior art, the present application provides an intra-frame tampered video generation method and system to solve the problems of high production cost and low quality of existing tampered video reference samples.

[0006] The technical solution adopted by the present application to solve the technical problem is as follows:

[0007] In a first aspect, the present application provides an intra-frame tampered video generation method, comprising:

[0008] cutting a given original video into original video segments, obtaining the background layer of each segment, and pre-processing;

[0009] obtaining a mask sequence of the tampered target from each segment;

[0010] obtaining the foreground layer corresponding to each segment according to the mask sequence, and calculating the brightness value and contrast difference according to the foreground layer and the background layer to obtain the brightness value and contrast difference value of each frame of video in each segment;

[0011] The automatic tracking tool is used to obtain mask region attributes corresponding to the mask sequence, and the mask region attributes are completely copied from the adjustment layer to the original video layer.

[0012] The original video layer mask region is subjected to brightness and contrast adjustment processing to obtain the post-processed tampered video.

[0013] In an implementation manner, the given original video slice is divided into original video segments, the background layers of the segments are obtained, and preprocessing is performed, including:

[0014] The given original video slice is divided into original video segments; wherein the original video segments include a plurality of target of divisible regions; the target is a moving or stationary target, and each frame in which the target continuously appears in the original video segment is a frame to be tampered with.

[0015] The background layers of the segments are obtained, and preprocessing is performed, including adjusting the video frame rate, video resolution, and video duration to the same value.

[0016] In an implementation manner, the background layers of the segments are obtained, including:

[0017] A video frame of the same length as the original video segment is selected as a first background layer; or another video of the same frame number, the same resolution, and similar content as the original video is selected as the first background layer.

[0018] A long-time span video with the same scene as the original video is selected, and after segmentation, random sampling is performed on all frames of each segment, and after random recombination, N background frame sequences consistent with the frame number of the original video are obtained as a second background layer.

[0019] The first background layer and the second background layer are packaged into the background layers corresponding to the segments.

[0020] In an implementation manner, the mask sequence of the tampered target obtained from the segments includes:

[0021] An area segmentation model is used to extract n target mask regions from a frame in which the target appears completely.

[0022] The n target mask regions of the obtained video are sequentially input into the original video as a target segmentation model to obtain a mask sequence corresponding to the n target mask regions.

[0023] In an implementation manner, the mask sequence is a sequence constituted by masks of binary images; wherein the mask sequence corresponds to the same frame number and image resolution as the original video, the value of the mask sequence is 0 in a background area, and the value of the mask sequence is 1 in a foreground area.

[0024] In an implementation manner, the foreground layer corresponding to each segment is obtained according to the mask sequence, and the brightness value and the contrast difference are calculated according to the foreground layer and the background layer, including:

[0025] The original video, the obtained background layer and the mask sequence are loaded, the mask sequence is dilated one by one, and the boundary is extracted to obtain the inner and outer boundaries of the mask area;

[0026] The first pixel set and the second pixel set are selected from the inner and outer boundaries respectively;

[0027] The brightness value and the contrast difference value of the nearest point in the first pixel set and the second pixel set are calculated to generate a difference value set;

[0028] The difference value set is analyzed by using a clustering algorithm to determine the brightness value and the contrast difference value of each frame;

[0029] The brightness value and the contrast difference value are adjusted by smoothing filtering to ensure the visual coherence of the tampered video, and the result is saved as a brightness-contrast list for adjusting the background layer in a video processing software.

[0030] In an implementation manner, the brightness and contrast adjustment processing is performed on the mask area of the original video layer to obtain a post-processed tampered video, including:

[0031] The original video layer is selected, and the video target segmentation mask is converted into a mask attribute by calling the automatic tracking tool;

[0032] The mask area in the original video layer is tampered, and the mask feathering and mask expansion values are adjusted;

[0033] According to the calculated brightness value and contrast difference, the brightness and contrast adjustment tool is used to adjust the brightness and contrast of the background layer, and the background layer is placed at the bottom end of the original video layer to generate the tampered video.

[0034] In a second aspect, the application provides an intra-frame tampered video generation system, including:

[0035] An original video segment module is configured to segment a given original video into original video segments, obtain the background layer of each segment, and perform preprocessing.

[0036] The mask sequence acquisition module is configured to acquire a mask sequence of the tampered target from each segment.

[0037] The brightness and contrast difference calculation module is configured to acquire a foreground layer corresponding to each segment according to the mask sequence, and calculate a brightness value and a contrast difference according to the foreground layer and a background layer, so as to obtain the brightness value and the contrast difference value of each frame of video in each segment.

[0038] The mask region attribute acquisition module is configured to acquire a mask region attribute corresponding to the mask sequence by using an automatic tracking tool, and completely copy the mask region attribute from the adjustment layer to the original video layer.

[0039] The brightness and contrast adjustment module is configured to perform brightness and contrast adjustment processing on the mask region of the original video layer, so as to obtain the post-processed tampered video.

[0040] In a third aspect, the present application provides a terminal, comprising a processor and a memory, wherein the memory stores an intra-frame tampered video generation program, and the intra-frame tampered video generation program is used to implement the operations of the intra-frame tampered video generation method according to the first aspect when executed by the processor.

[0041] In a fourth aspect, the present application further provides a medium, which is a computer readable storage medium, and the medium stores an intra-frame tampered video generation program, and the intra-frame tampered video generation program is used to implement the operations of the intra-frame tampered video generation method according to the first aspect when executed by a processor.

[0042] The technical scheme of the present application has the following effects:

[0043] The present application acquires the background layer of each original video segment and performs preprocessing, acquires the mask sequence of the tampered target from each segment, acquires the foreground layer corresponding to each segment according to the mask sequence, calculates the brightness value and the contrast difference for adjusting the background layer in Adobe Effect, acquires the mask region attribute corresponding to the mask sequence by using an automatic tracking tool, completely copies the mask region attribute from the adjustment layer to the original video layer, performs brightness and contrast adjustment processing on the mask region of the original video layer, and obtains the post-processed tampered video. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained according to the structures shown in the drawings without creative labor.

[0045] Figure 1 is a flow chart of the method for generating an intra-frame tampered video in the present application.

[0046] Figure 2 is a flow chart of the specific implementation of the method for generating an Adobe Effect tampered video in the present application.

[0047] Figure 3 is a flow chart of the step of obtaining a video target segmentation mask in the present application.

[0048] Figure 4 is a comparison chart of the automatically generated tampered video frame and the original video frame in the method for generating an Adobe Effect tampered video in the present application.

[0049] Figure 5 is a comparison chart of the automatically generated tampered video frame and the tampered region mask in the method for generating an Adobe Effect tampered video in the present application.

[0050] Figure 6 is a functional principle diagram of a terminal in an implementation manner of the present application.

[0051] The purposes, functional features and advantages of the present application will be further described with reference to the accompanying drawings. DETAILED DESCRIPTION

[0052] In order to make the purposes, technical solutions and advantages of the present application clearer and more explicit, the present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.

[0053] Exemplary method

[0054] The existing video intra-frame tampering detection model needs a large amount of intra-frame tampering sample data for training, so as to quickly and accurately detect the tampered part in the video in the actual video intra-frame tampering detection task. Intra-frame tampering refers to modifying the content of a frame in a video, adding or removing objects, changing the color or texture of objects, etc. in an image. However, in the actual video intra-frame tampering detection task, a large number of tampered videos are needed as reference samples, not only the content of each frame needs to be modified, but also post-processing is needed, and the tampering quality needs to be focused on, which requires a large amount of manpower and time cost.

[0055] To solve the above technical problems, the embodiment of the present application provides an intra-frame tampered video generation method, which obtains the background layer of each original video segment and performs preprocessing; the mask sequence of the tampered target can be obtained from each segment; and the foreground layer corresponding to each segment is obtained according to the mask sequence, and the brightness value and the contrast difference used for adjusting the background layer in Adobe Effect (a graphic video processing software) are calculated; and the mask region attribute corresponding to the mask sequence is obtained by using an automatic tracking tool, and the mask region attribute is completely copied from the adjustment layer to the original video layer; the brightness and contrast of the original video layer mask region are adjusted and processed to obtain the post-processed tampered video. Therefore, the method provided in the embodiment of the present application can provide rich high-quality data for pre-training of a video tampering positioning model, greatly reduce the labor and time cost, and has important application value in video tampering positioning evidence.

[0056] As shown in Figure 1 The embodiment of the present application provides an intra-frame tampered video generation method, which includes the following steps:

[0057] Step S100, the given original video is cut into original video segments, the background layer of each segment is obtained, and preprocessing is performed.

[0058] In this embodiment, the intra-frame tampered video generation method obtains a tampered target mask sequence through automatic processing of the original video, then imports the original video, the mask sequence and the background layer into the Adobe Effect software to generate a post-processed tampered video. The entire process is realized by using an Adobe Effect script program. The tampered video dataset generated in this embodiment can be used for pre-training of a video intra-frame tampering detection deep model.

[0059] As shown in Figure 2 As an example, the method in this embodiment needs to prepare an original video, select a tampered target, and select a mask sequence acquisition method in an actual application scenario; then, a background layer is selected according to the need, and the brightness and contrast difference is calculated according to the original video and the background layer; finally, the AE mask attribute of the original video layer is converted from the mask sequence by using the video target segmentation mask, a tampered video is constructed, and the tampered video is post-processed and exported.

[0060] Specifically, in one implementation manner of this embodiment, step S100 includes the following steps:

[0061] Step S101, the given original video is cut into original video segments; wherein the original video segments include a target of a plurality of divisible regions;

[0062] In the embodiment, the original video clip needs to include N (N>2) divisible regions of interest, which can be moving or static objects such as people, wall decorations, or doors, and each object appearing in the original video clip can be a frame to be tampered with. The tampered objects should appear continuously in time to provide inter-frame characteristics of the tampering behavior.

[0063] In step S102, the background layers of each clip are obtained and preprocessed to adjust the video frame rate, video resolution, and video duration to the same value.

[0064] Specifically, in one implementation of the embodiment, the background layers of each clip obtained in step S102 include the following steps:

[0065] In step S102a, a video frame of the same length as the original video clip is selected as the first background layer; or another video with the same number of frames, the same resolution, and similar content as the original video is selected as the first background layer.

[0066] In step S102b, a longer time span video with the same scene as the original video is selected, and random sampling is performed on all frames of each segment after segmentation. After random recombination, N segments of background frame sequences (i.e., background PNG sequences) consistent with the number of frames of the original video are obtained as the second background layer.

[0067] In step S102c, the first background layer and the second background layer are packaged as the background layer corresponding to each clip.

[0068] As an example, in the embodiment, the tampered video generation method based on the video target segmentation mask is as follows:

[0069] Step 1: Prepare a tampered video containing tampered objects in advance; for example, as shown in FIG. 1, a video monitoring scene is shown in FIG. 2, and it is assumed that a person is selected as the tampered object, and the following two background layers need to be generated: Figure 5 Figure 5 As shown in FIG. 3, the background layer one (i.e., the first background layer) is selected as a frame in which the person does not appear in the scene; a video clip with the same number of frames as the original video in which the person does not appear is selected as the background layer one.

[0070] As shown in FIG. 4, the background layer two (i.e., the second background layer) is selected as a longer time span video with the same scene as the original video; after segmentation, random sampling is performed on all frames of each segment; after random recombination, N segments of background frame sequences (i.e., background PNG sequences) consistent with the number of frames of the original video are obtained as the second background layer.

[0071] ​Background layer two (i.e., the second background layer): Since the background layer one is relatively single, it is taken from the original video scene, and after adding post-processing operation to one background picture, it can be applied to each frame, so the pixel similarity of the tampered area of each frame is high. Especially when different target masks have intersecting areas, after repeated tampering of the same original video, most of the areas present the same statistical characteristics, which is not conducive to the pre-training of the model. In order to increase the generalization ability of the pre-training data set, the background layer is constructed in the following way: a video with a long time span that is the same as or similar to the original video scene is selected, and after segmentation, random sampling is performed on all frames of each segment of the video, and after random recombination, N background PNG sequences consistent with the number of frames of the original video are obtained as the background layer two;

[0072] It is worth mentioning that in the generation of the above two background layers, if the background layer is a video, the video frame rate, video resolution and video duration of the background layer need to be adjusted to be consistent with the original video; if the background layer is an image, the image resolution is adjusted to be consistent with the original video.

[0073] The traditional manual tampering video method cannot generate tampered videos on a large scale. Under the premise of focusing on tampering quality, it is often necessary to select the tampered area frame by frame and perform post-processing adjustment, which is low in efficiency and consumes a lot of time and energy. In the present embodiment, an automatic process is realized by combining the Adobe Effect tool, the mask sequence of the tampered target can be obtained from each segment by generating the above two background layers, and the subsequent tampering is performed by using the generated background layer and the obtained mask sequence, so that a large number of tampered videos can be generated in a relatively short time, greatly saving time and labor cost.

[0074] As shown in Figure 1 , the present embodiment provides a frame-in tampered video generation method, which comprises the following steps:

[0075] Step S200, obtaining a mask sequence of a tampered target from each segment.

[0076] In the present embodiment, the SAM large model and the XMem model are used to obtain the mask sequence of the tampered target from each segment; by using the SAM large model and the XMem model to assist in the segmentation of the video target, it is possible to tamper with a specific target in the video, compared with the random area tampering of the traditional method, the mask sequence acquisition method in the present embodiment is more semantic information and high-quality visual effect on one hand, and more similar to the steps and ideas of actual video tampering on the other hand, which can provide more features of actual tampered videos.

[0077] Specifically, in one implementation manner of the present embodiment, step S200 comprises the following steps:

[0078] Step S201, using a region segmentation model (i.e., a SAM model) to extract n target mask regions from a frame in which the target completely appears;

[0079] Step S202, taking the n target mask regions of the video as inputs of the original video as a target segmentation model (i.e., an XMem model, which is a video object segmentation architecture for long videos), to obtain a mask sequence corresponding to the n target mask regions.

[0080] It can be understood that the mask sequence in the embodiment is a sequence composed of binary image masks; the mask sequence has the same frame number and image resolution as the original video, the value of the mask sequence is 0 in the background region, and the value of the mask sequence is 1 in the foreground region.

[0081] As an example, the tampered video generation method based on the video target segmentation mask further includes the following:

[0082] The second step is shown in the flowchart of Figure 3 As shown, first, initialize the SAM model and the XMem model, select the video file to be processed through the GUI interface, and extract all frames from the video and initialize the video state using the set code. Then, according to the selected initial frame of the video, perform target segmentation to generate initial object masks mask_1 to mask_k. Finally, after preliminary screening, get the simplified mask_1 to mask_n.

[0083] The initial screening takes the center of mass of the mask and the area size of the mask as the evaluation criteria, that is, the mask meets the range, the scores are arranged in descending order, and the top n scores are selected.

[0084]

[0085]

[0086] According to the above n masks, the target object to be tracked is determined, and the object in the video is tracked to generate mask tracking results masks_1 to masks_n.

[0087] Finally, the tracking results are saved, including the mask binary image of each frame, and a YAML file is generated to record video information and mask duration frame interval metadata;

[0088] It is worth mentioning that the mask generated by the above scheme is a binary image with the same size as the original video, and each frame of the video corresponds to a mask, and finally stored in the form of a PNG sequence. Assuming that an initial video selects a target for tracking, and after screening, there are targets, and the number of video frames is , then the original video can generate a total of tampered mask sequences.

[0089] As shown in Figure 1 , the embodiment of the application provides a method for generating an intra-frame tampered video, comprising the following steps:

[0090] Step S300, obtaining the foreground layer corresponding to each segment according to the mask sequence, and calculating the brightness value and contrast difference according to the foreground layer and the background layer to obtain the brightness value and contrast difference value of each frame of video in each segment.

[0091] Specifically, in one implementation manner of the embodiment, step S300 comprises the following steps:

[0092] Step S301, loading the original video, the obtained background layer and the mask sequence, and performing dilation processing on the mask sequence one by one, and performing boundary extraction to obtain the inner and outer boundaries of the mask region; The default value is 3, and the dilation kernel size can be customized;

[0093] Step S302, respectively selecting a first pixel point set of N 3*3 pixel points and a second pixel point set of N 3*3 pixel points from the inner and outer boundaries at equal distances; wherein N is the number of internal boundary pixel points divided by , The value can be defined by yourself according to the need; the first pixel point set and the second pixel point set are used for brightness contrast calculation;

[0094] Step S303, calculating the brightness value and contrast difference value of the nearest points in the first pixel point set and the second pixel point set to generate a difference value set;

[0095] Step S304, using a clustering algorithm (i.e. K-means clustering algorithm) to analyze the difference value set to determine the brightness value and contrast difference value of each frame;

[0096] Step S305, smoothing filter adjusting the brightness value and contrast difference value to ensure the visual coherence of the tampered video, and saving the result as a brightness contrast list; the brightness contrast list is used for adjusting the background layer in the video processing software (i.e. Adobe Effect).

[0097] ​As an example, the tampered video generation method based on the video target segmentation mask further includes the following:

[0098] Step 3: Calculate the brightness and contrast adjustment values of the background layer and the original video frame by frame, and the steps are as follows:

[0099] 1) Load the original video, the background layer, and cyclically load k mask sequences with a length of For each mask sequence, the following processing is performed.

[0100] 2) Perform dilation processing on the mask sequence one by one, and then perform boundary extraction to obtain the inner and outer boundaries of the mask region. The default value is 3, and the dilation kernel size can be customized.

[0101] 3) Respectively select N 3*3 pixel sets A and N 3*3 pixel sets B (N is the number of internal boundary pixels divided by , The value of N can be defined by the user as needed) from the equidistant inner and outer boundaries, which are used for brightness and contrast calculation.

[0102] 4) Calculate the brightness and contrast difference of the nearest points in A and B to generate a difference set Delta.

[0103] 5) Use K-Means clustering analysis on Delta to determine the brightness and contrast difference of each frame.

[0104] 6) Smooth the filter adjustment brightness and contrast difference to ensure the visual coherence of the tampered video, and save the results as a brightness and contrast list for adjusting the background layer in Adobe Effect and save to i.txt n txt files, respectively recording the brightness and contrast adjustment values (d_brightness, d_contrast) of each video frame in which the person appears.

[0105] As shown in Figure 1 , the embodiment of the application provides an intra-frame tampered video generation method, including the following steps:

[0106] Step S400: using an automatic tracking tool to obtain mask region attributes corresponding to the mask sequence, and copying the mask region attributes from the adjustment layer to the original video layer.

[0107] Step S500: adjusting the brightness and contrast of the mask region of the original video layer to obtain a post-processed tampered video.

[0108] ​In the embodiment, the Auto Trace tool of Adobe Effect is used as the automatic tracking tool to analyze the mask sequence, obtain the mask region attribute of Adobe Effect, and then completely copy the mask region attribute of Adobe Effect from the adjustment layer to the original video layer; after the tampered video is constructed, the mask region of the original video layer is post-processed, that is, the brightness and contrast of the background layer are processed to obtain the tampered video after Adobe Effect post-processing.

[0109] Specifically, in one implementation manner of the embodiment, step S500 includes the following steps:

[0110] Step S501, selecting the original video layer, calling the automatic tracking tool (i.e., the Auto Trace tool) to convert the video target segmentation mask into the mask attribute of Adobe Effect;

[0111] Step S502, tampering the mask region of Adobe Effect in the original video layer and adjusting the mask feathering and mask expansion values;

[0112] Step S503, according to the calculated brightness value and contrast difference, using the brightness and contrast adjustment tool to adjust the brightness and contrast of the background layer, and placing the background layer at the bottom end of the original video layer to generate the tampered video.

[0113] As an example, the tampered video generation manner based on the video target segmentation mask further includes the following:

[0114] Fourth step: opening Adobe Effect, importing the original video, first importing the background layer one to perform video tampering, second importing the background layer two to perform video tampering, and the subsequent steps are as follows:

[0115] 1) According to the read txt file, n compositions are created in the original video layer, if it is the background layer one, the background layer is directly imported, if it is the background layer two, the ith background layer is imported into the ith composition .

[0116] The ith mask sequence uses the Auto trace tool to create mask attributes frame by frame, and the mask is copied to the original layer frame by frame, so as to realize tampering of the mask region of Adobe Effect in the original video layer and adjusting the mask feathering and mask expansion values. The brightness and contrast adjustment tool is called to adjust the background layer, that is, according to i.txt Each pair of (d_brightness, d_contrast) adjustment values obtained by the file are used to adjust the video, and the background layer is placed at the bottom of the original video layer, and the tampered video is exported and generated.

[0117] As Figure 4 shown, Figure 4 In the figure, the visual effect comparison between the tampered video frame (left) and the original video frame (right) is generated. After post-processing such as brightness contrast adjustment, mask feathering and mask expansion, it is difficult to accurately find the tampered pixel area visually. As Figure 5 shown, Figure 5 In the figure, the left column is the binary image of the tampered area, and the right column is the corresponding tampered video frame.

[0118] In order to illustrate that the automatically generated tampered video in this embodiment can play an important role in video forensics method. This embodiment builds a deep model MVSS-Net applied in the field of forensics, then uses the tampered video in the form of PNG sequence as training image sample, trains on the MVSS-Net model, and then tests the performance of the trained model on the artificial tampering database.

[0119] Table 1

[0120]

[0121] As can be seen from Table 1, using the data set generated by the method as the pre-training data set of the image tampering detection model can enable the MVSS-Net model to learn to some extent the actual manual tampering features, thereby improving the video tampering detection capability for Adobe Effect software.

[0122] The existing automatic generation method usually relies on ordinary video frame processing or simple AI technology, and cannot simulate video frame characteristics similar to the Adobe Effect tampering process, resulting in significant differences in statistical properties between the generated video and the real Adobe Effect tampered video. In this embodiment, Adobe Effect is still used for tampering fundamentally, and other tools only provide auxiliary role, so the generated tampered video is closer to the real tampered video in statistical properties. In addition, this embodiment also provides a UI for adjusting the tampered area in an interactive form. After adjustment, a high-quality video target segmentation mask can be provided, and then Adobe Effect is used for automatic video tampering to generate a video similar to manual high-quality tampering. Therefore, this embodiment can provide rich high-quality data for pre-training of video tampering positioning model, greatly reducing the labor and time cost, and has important application value in video tampering positioning forensics.

[0123] The embodiment achieves the following technical effects through the above technical solutions:

[0124] The embodiment obtains the background layer of each original video segment and performs preprocessing, obtains the mask sequence of the tampered target from each segment, obtains the corresponding foreground layer of each segment according to the mask sequence, calculates the brightness value and the contrast difference for adjusting the background layer in the Adobe Effect, obtains the mask region attribute corresponding to the mask sequence by using an automatic tracking tool, and completely copies the mask region attribute from the adjustment layer to the original video layer, adjusts the brightness and the contrast of the mask region of the original video layer, and obtains the post-processed tampered video. The embodiment can provide rich high-quality data for pre-training of a video tampering positioning model, greatly reduces the labor and time cost, and has important application value in video tampering positioning evidence.

[0125] Exemplary device

[0126] Based on the above embodiment, the application further provides an intra-frame tampered video generation system, comprising:

[0127] An original video segment module is configured to slice a given original video into original video segments, obtain the background layer of each segment, and perform preprocessing;

[0128] A mask sequence acquisition module is configured to obtain the mask sequence of the tampered target from each segment;

[0129] A brightness and contrast difference calculation module is configured to obtain the corresponding foreground layer of each segment according to the mask sequence, and calculate the brightness value and the contrast difference according to the foreground layer and the background layer, to obtain the brightness value and the contrast difference value of each frame of video in each segment;

[0130] A mask region attribute acquisition module is configured to obtain the mask region attribute corresponding to the mask sequence by using an automatic tracking tool, and completely copy the mask region attribute from the adjustment layer to the original video layer;

[0131] A brightness and contrast adjustment module is configured to adjust the brightness and the contrast of the mask region of the original video layer, and obtain the post-processed tampered video.

[0132] The embodiment achieves the following technical effects through the above technical solutions:

[0133] The embodiment obtains the background layer of each original video segment and performs preprocessing; the mask sequence of the tampered target can be obtained from each segment; and the corresponding foreground layer of each segment is obtained according to the mask sequence, the brightness value and the contrast difference used for adjusting the background layer in the Adobe Effect are calculated; and the mask region attribute corresponding to the mask sequence is obtained by using the automatic tracking tool, and the mask region attribute is completely copied from the adjustment layer to the original video layer; the brightness and contrast of the original video layer mask region are adjusted and processed to obtain the post-processed tampered video. The embodiment can provide rich high-quality data for pre-training of a video tampering positioning model, greatly reduce the labor and time cost, and has important application value in video tampering positioning evidence.

[0134] Based on the above embodiment, a terminal is also provided, and a principle block diagram thereof can be as shown in Figure 6

[0135] The terminal includes a processor, a memory, an interface, a display screen and a communication module connected through a system bus; the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operating system and the computer program in the storage medium to run; the interface is used to connect external devices; the display screen is used to display corresponding information; and the communication module is used to communicate with a cloud server or other devices.

[0136] The computer program is executed by the processor to implement the operations of the intra-frame tampered video generation method.

[0137] Those skilled in the art can understand that, Figure 6 The principle block diagram shown in the above embodiment is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied; specifically, the terminal can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0138] In one embodiment, a terminal is provided, which includes a processor and a memory, and the memory stores an intra-frame tampered video generation program, and the intra-frame tampered video generation program is executed by the processor to implement the operations of the intra-frame tampered video generation method as above.

[0139] In one embodiment, a storage medium is provided, and the storage medium stores an intra-frame tampered video generation program, and the intra-frame tampered video generation program is executed by the processor to implement the operations of the intra-frame tampered video generation method as above.

[0140] ​Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments of the present application can include non-volatile and volatile memory.

[0141] In summary, the present application provides an intra-frame tampered video generation method and system, comprising: cutting a given original video into original video segments, obtaining the background layer of each segment, and performing preprocessing; obtaining the mask sequence of the tampered target from each segment; obtaining the foreground layer corresponding to each segment according to the mask sequence, and calculating the brightness value and contrast difference according to the foreground layer and the background layer to obtain the brightness value and contrast difference value of each frame of video in each segment; using an automatic tracking tool to obtain the mask region attribute corresponding to the mask sequence, and completely copying the mask region attribute from the adjustment layer to the original video layer; adjusting the brightness and contrast of the original video layer mask region to obtain the post-processed tampered video. The present application can provide rich high-quality data for the pre-training of the video tampering positioning model, greatly reduce the labor and time cost, and has important application value in video tampering positioning evidence.

[0142] It should be understood that the application of the present application is not limited to the above examples, and those skilled in the art can make improvements or changes according to the above description, and all these improvements and changes shall belong to the protection scope of the claims of the present application.

Claims

1. An intra-frame tampered video generation method, characterized by, The method comprises the following steps: slicing a given original video into original video segments, obtaining the background layers of each segment, and performing preprocessing; obtaining a mask sequence of the tampered target from each segment; obtaining the foreground layer corresponding to each segment according to the mask sequence, and calculating the brightness value and contrast difference according to the foreground layer and the background layer to obtain the brightness value and contrast difference value of each frame of video in each segment; using an automatic tracking tool to obtain the mask region attribute corresponding to the mask sequence, and completely copying the mask region attribute from the adjustment layer to the original video layer; based on the brightness value and contrast difference value, adjusting the brightness and contrast of the mask region of the original video layer to obtain the processed tampered video.

2. The intra-forged video generation method of claim 1, wherein, The method comprises the following steps: slicing a given original video into original video segments, obtaining the background layers of each segment, and performing preprocessing; slicing a given original video into original video segments; wherein the original video segment includes a plurality of target subareas; the target is a moving or stationary target, and each target appears continuously in the original video segment. The frame is a video frame to be tampered with; 3. The intra-forged video generation method of claim 2, wherein, obtaining the background layer of each segment and performing preprocessing, adjusting the video frame rate, video resolution and video duration to the same value. The method comprises the following steps: selecting a segment of the original video segment, and selecting the same length of video frames from the remaining segments of the original video segment as the first background layer; or selecting another segment of video with the same number of frames, the same resolution and similar content as the original video as the first background layer; selecting a longer time span video with the same scene as the original video, and randomly sampling all frames of each segment after segmentation to obtain N background frame sequences consistent with the number of frames of the original video as the second background layer; 4. The intra-forged video generation method of claim 1, wherein, packaging the first background layer and the second background layer into the background layer corresponding to each segment. The method comprises the following steps: using a region segmentation model to extract n target mask regions from a frame in which the target appears completely; 5. The intra-tampered video generation method of claim 1 or 4, wherein, inputting the n target mask regions of the obtained video into the original video as a target segmentation model one by one to obtain a mask sequence corresponding to the n target mask regions.

6. The intra-forged video generation method of claim 1, wherein, The mask sequence is a sequence composed of binary image masks; wherein the number of frames and the image resolution of the mask sequence are the same as the number of frames and the image resolution of the original video, the region with a value of 0 in the mask sequence is a background region, and the region with a value of 1 in the mask sequence is a foreground region. The method comprises the following steps: loading the original video, the obtained background layer and the mask sequence, performing dilation processing on the mask sequence one by one, and performing boundary extraction to obtain the inner and outer boundaries of the mask region; respectively selecting a first pixel set and a second pixel set from the inner and outer boundaries at equal distances; Calculate the luminance value and contrast difference of the nearest points in the first pixel set and the second pixel set, and generate a difference value set; Analyze the difference value set using a clustering algorithm to determine the luminance value and contrast difference of each frame; Smooth filter to adjust the luminance value and contrast difference, ensure the visual coherence of the tampered video, and save the result as a luminance contrast list for adjusting the background layer in the video processing software.

7. The intra-fraudulent video generation method of claim 1, wherein, The luminance and contrast adjustment processing of the original video layer mask region based on the luminance value and contrast difference value, to obtain the processed tampered video, comprising: Select the original video layer, call the automatic tracking tool to convert the video target segmentation mask to a mask attribute; Tamper with the mask region in the original video layer and adjust the mask feathering and mask expansion value; According to the calculated luminance value and contrast difference, use the luminance and contrast adjustment tool to adjust the luminance and contrast of the background layer, and place the background layer at the bottom of the original video layer, to generate the tampered video.

8. An intra-frame tampered video generation system characterized by, Comprising: The original video segment module is used to cut the given original video into original video segments, obtain the background layer of each segment, and pre-process; The mask sequence acquisition module is used to obtain the mask sequence of the tampered target from each segment; The luminance and contrast difference calculation module is used to obtain the foreground layer corresponding to each segment according to the mask sequence, and calculate the luminance value and contrast difference according to the foreground layer and the background layer, to obtain the luminance value and contrast difference value of each frame of video in each segment; The mask region attribute acquisition module is used to use the automatic tracking tool to obtain the mask region attribute corresponding to the mask sequence, and completely copy the mask region attribute from the adjustment layer to the original video layer; The luminance and contrast adjustment module is used to adjust the luminance and contrast of the original video layer mask region based on the luminance value and contrast difference value, to obtain the processed tampered video.

9. A terminal, characterized by comprising: Comprising: A processor and a memory, the memory stores a frame-in tampered video generation program, the frame-in tampered video generation program is executed by the processor to realize the operation of the frame-in tampered video generation method in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a frame-in tampered video generation program, the frame-in tampered video generation program is executed by the processor to realize the operation of the frame-in tampered video generation method in any one of claims 1-7.

Citation Information

Patent Citations

  • Video coding and decoding, searching method and device

    CN109495749A

  • Tampered video generation method and system based on semantic guidance

    CN117201839A