A generative edge trace reduction method based on diffusion model
By adopting a generative edge trace reduction method based on diffusion model in image repair technology, using DDPM and manifold constraints, combined with edge dynamic hierarchy and image fusion technology, the problems of inconsistent edge effects and blurred effects in image repair process are solved, and a more effective falsified image trace reduction effect is achieved.
Patent Information
- Application Number
- CN202411668920.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The prior art has problems of inconsistent edge effects, inconsistent textures and blur effects during image repair, and it has failed to effectively resist trace reduction of forged image detection.
Using a generative edge trace reduction method based on diffusion model, the generative edge trace reduction network is constructed and trained, and the denoising diffusion probability model (DDPM) and manifold constraints are used to combine edge dynamic hierarchy and image fusion technology to optimize the generation sampling range of edge pixel points and the mutual influence between the generated pixel points.
It effectively improves the trace reduction effect of forged images, and is suitable for different types of forgery methods such as copy-paste, splicing and image modification, showing wide applicability and effectiveness, and can perform superiorly on both the single forgery type detection method network and the mixed forgery type detection method network.
Smart Images

Figure CN119169038B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to information security and image processing technology, and specifically relates to a generative edge trace reduction method based on a diffusion model. Background Art
[0002] With the rapid development of digital image processing technology and image editing tools, digital images can be easily tampered with without leaving obvious traces of modification. Various image modification software has emerged, and tampered images generated for various purposes are everywhere on the Internet. In order to deal with these common image tampering strategies, many corresponding detection methods have been derived. For forged images of types such as splicing, copy-paste, copy-move, and traditional diffusion, these tampering detection methods have achieved high detection accuracy.
[0003] Among the three common image tampering methods (splicing, copy-paste, and image restoration), we are more concerned about the image restoration type of tampering. Image restoration technology can effectively repair damaged areas in images and remove stains and scratches in images. If the damaged area in the image is regarded as a target that needs to be removed, then image restoration technology is also a means of image tampering. Copy-paste can directly copy a larger area to cover the content that needs to be tampered with, which is more convenient for deleting objects, but it will leave obvious duplicates or areas similar to the tampered image, while tampering based on image restoration is not easy to be visually detected.
[0004] The emergence of tampering technology based on image restoration has provided a new idea for image forgery technology. In recent years, the growing strength of generative models has also led to the rapid development of image restoration technology. Image restoration technology can be divided into two main categories: one is sample-based method and the other is partial differential equation-based method. The former often selects a certain area of the background image to fill the unknown area. The background image sample that fills the unknown area is often considered to have similar structure and texture similarity to the edge of the unknown area. The latter simulates the heat diffusion in physics through partial differential equations (PDE) to smoothly propagate the local image structure from the outside to the inside of the gap. However, this method will bring about a blurring effect during the generation process, so it is suitable for image restoration of small areas. However, the target object in many image tampering samples is a small area, so diffusion-based image restoration technology can be used as a powerful tampering tool.
[0005] Diffusion-based methods can make the image target repair area more reasonable to a certain extent. At present, for image tampering schemes based on traditional diffusion methods, by classifying the different features of the image repair area and the unmanipulated area in the residual domain, it is possible to effectively detect whether the image has been tampered with by image repair. However, since its target area is usually a large rectangular filling repair area, the use of traditional diffusion filling will inevitably produce a smooth blurred area, and the blur effect will change the texture of the repair area.
[0006] Since the advent of the denoising diffusion probability model (DDPM), corresponding image restoration schemes have also been proposed. The denoising diffusion probability model with pre-trained weights can not only predict the possible pixel values of the target area in the image by training the network, but also combine the image pixels around the target area to jointly generate the area. Therefore, the new image restoration scheme based on the denoising diffusion probability model can not only generate a restored image without blurring effect, but also obtain a more reasonable and consistent structure with the surrounding images.
[0007] However, the above-mentioned new image restoration method still has some defects, such as inconsistent texture on the edge effect, large color span at both ends of the restoration area, etc., it does not take into account the trace reduction work to resist forged image detection, and the edge restoration of the target area produces inconsistent edge structure, etc. Summary of the invention
[0008] Purpose of the invention: The purpose of the present invention is to solve the deficiencies in the prior art and to provide a generative edge trace reduction method based on a diffusion model.
[0009] Technical solution: A generative edge trace reduction method based on a diffusion model of the present invention constructs and trains a generative edge trace reduction network, and inputs a tampered image X and its corresponding target forged region mask mask(X) into the generative edge trace reduction network. The specific steps are as follows:
[0010] Step 1: In the first round of processing, the target forged area mask (X) is first converted into a gray image, and then input into the target edge acquisition and stratification module. After processing by the target edge acquisition and stratification module, the edge point coordinates E of the forged area are obtained. eg ;
[0011] Step 2: Tamper with the image X and edge point coordinates E eg An edge generation module is input, wherein the edge generation module is based on a denoising diffusion probability model DDPM and adopts a manifold constraint, updates the tampered image X through the edge generation module, and outputs an image G;
[0012] The output image G obtained in the first round is input to the edge detection module, and the image loss of the detected edge and the actual edge, the undetected edge pixel value and the corresponding position K are obtained by the edge detection module, and the image loss, the undetected edge pixel value and the corresponding position K are returned to the target edge acquisition and stratification module and the edge image fusion module;
[0013] Step 3: From the second round, the target edge acquisition and stratification module receives the image loss, the undetected edge pixel value and the corresponding position K, and the edge point coordinates E obtained in step 1 eg Subtract K to get the detected edge pixel position P, add one pixel distance to P to get the edge neighboring layer L, edge point coordinates E eg Combined with the edge layer L to form the edge adjacent layer combination E zh Then the updated image G is combined with the edge neighbor layer E zh Input to edge generation module;
[0014] At the same time, the edge image fusion module receives the image G, the undetected edge pixel values and the corresponding positions K, and outputs the fusion image GR;
[0015] Step 4: Starting from the second round, the edge detection module receives the fused image GR, calculates the loss image, the undetected edge pixel values and the corresponding position K again, and returns them to the target edge acquisition and stratification module and the edge image fusion module.
[0016] Furthermore, the first round of processing of the grayscale image by the target edge acquisition and stratification module is as follows:
[0017] Use the Get_mask_edge method to fix the x-axis and traverse the y-axis image of the grayscale image to obtain the pixel value change point, then fix the y-axis and traverse the pixel value change point obtained by the x-axis, and combine the two obtained pixel value change points to obtain the complete edge point coordinates E eg ; Set the edge point coordinates E eg The target edge layering network is passed in, and the remaining pixels within a distance of one pixel around each point coordinate are recorded as the edge neighboring layer. Since it is necessary to distinguish the pixels in the neighboring pixel layer as belonging to the inner and outer sides of the edge, especially to deal with the inner and outer distinction of irregular edges, the y-axis is fixed and the edge point coordinates on the x-axis are combined into point pairs. The edge neighboring layer pixels located in the edge point pair are marked as the inner edge neighboring layer, and other pixels are marked as the outer edge neighboring layer.
[0018] From the second round onwards, the target edge acquisition and stratification module will use the edge point coordinates E egSubtract the undetected edge pixel position K returned by the edge detection module to obtain the detected edge pixel position P, and output the edge neighboring layer L with a pixel distance added to P. The distance θ of the edge neighboring layer L is cumulative, that is, the first round E eg The value of θ for each pixel is 0. If the edge neighboring layer θ of a pixel point q on P in the previous round expands by one pixel, and q is still on P in this round, the θ of point q will expand by another pixel.
[0019] The present invention sets an adjustable distance selection parameter θ for each edge pixel point. The distance selection parameter θ can dynamically select multiple groups of edge neighboring layers by detecting the edge pixel loss function returned by the network, and dynamically update the target area sampling range selection in the reverse generation process of the diffusion model in the next module according to task requirements.
[0020] The target edge acquisition and stratification module in the present invention has the following functions: first, by tampering with the grayscale image of the target area mask of the image to accurately determine the edge position to be generated without being disturbed by noise points, the precise edge point position can improve the texture consistency of the generated edge and background image. Secondly, the adjustable distance selection parameter θ for each pixel setting point in the edge stratification greatly improves the generation freedom of each edge pixel point, so that the generation of each edge pixel point will not be limited by the influence of the generation of the edge pixels on both sides, that is, the generation sampling range of each edge pixel point selects the positive promotion edge neighboring layer that is most conducive to the generation of the point.
[0021] Furthermore, the diffusion model of the edge generation module includes a forward diffusion process and a reverse diffusion process, the details of which are as follows;
[0022] , Each step of the forward diffusion process is a Gaussian translation;
[0023] (1)
[0024] in is a fixed variance schedule, which refers to a linearly increasing noise schedule ; Indicates how the initial value of the data x changes from time step t-1 to time t during the forward diffusion process. is the added Gaussian perturbation noise;
[0025] Formula (1) adds Gaussian noise to the latent variable to find , through (1) it can be deduced that when given clean data hour, The sampling of will be expressed in closed form:
[0026] (2)
[0027] in and ; s, t represent the time step of the diffusion model;
[0028] Expressed as and A linear combination of: (3)
[0029] is the Gaussian disturbance noise that conforms to the normal distribution;
[0030] The back diffusion process is modeled by a neural network, which predicts the parameters of the Gaussian distribution and ; The reverse diffusion process has the same functional form as the forward process and is expressed as a Gaussian transform with a learned mean and fixed variance, expressed as:
[0031] (4)
[0032] In addition, by Decompose into Sum Noise Approximator The linear combination of , the generation process is expressed as:
[0033] (5)
[0034] In the above formula, Each back-diffusion step is random, represents a pair of neural networks with the same input and output dimensions The predicted value at time step t, and the noise predicted by the neural network at each step is used for the denoising process in formula (5);
[0035] , in the backward diffusion process, an additional correction term inspired by the manifold constraint is used to make the backward iteration close to the manifold, that is, the initial data manifold of the original tampered image is recorded as M, and the forward diffusion process deviates from the data manifold M and gradually approaches the noise data manifold .
[0036] Step 2 uses the concept of data manifold to represent the forward diffusion trajectory of edge pixels in DDPM. The data manifold is used to describe the true distribution of the original image data. As the noise increases, the image gradually moves away from its true data manifold and enters a state containing more noise. The image will gradually change from a clear state to a completely random state. During the forward diffusion process, each pixel will experience a process from clarity to blur. This process is regarded as moving along a certain trajectory on the data manifold. This trajectory will eventually lead the image into a state composed entirely of noise. In short, the data manifold composed of all pixels of the image is forced to be locally assumed to be a linear structure, so as to obtain a better reconstruction effect.
[0037] Furthermore, starting from the second round, the edge image fusion module retains the appropriate edge pixel positions and corresponding pixel values of this round, and uses the retained pixels to replace the pixels at the corresponding positions on the original tampered image X. The remaining parts are sent to the target edge acquisition and layering module to update the edge range, and then passed to the edge generation module to generate the new edge and corresponding pixel value images, and the fusion image GR is obtained by fusion; the expression is as follows:
[0038] ;
[0039] ;
[0040] ;
[0041] In the above formula, It refers to the original tampered image X after removing the image edge. It means that for time step t, the pixel at position p on the forged edge is the result inferred by the reverse diffusion process and has the same noise level as the forward diffusion at time t.
[0042] The details are as follows: After obtaining the edge points of the target area, the target edge acquisition and stratification module sets the coordinates of the edge points to E eg , and set an adjustable distance selection parameter θ for each pixel point, directly converting E eg The edge layer of the state is sent to the second part edge generation module, and the output result of the edge generation module is recorded as , the edge image fusion module does not process this output, It will be directly sent to the edge detection module, and the error function result of the detected edge and the actual edge will be obtained through the edge detection module, and the result will be returned to the target edge acquisition and layering module and the edge image fusion module; at this time, the decision diagram of the edge image fusion module will save the edge pixel positions that were not correctly identified by the detection network in the first round And the pixel value V0 at that position, replace and update the corresponding pixel value on the original input image X The target edge acquisition and stratification module converts the pixel range E eg Adjust to E zh , E zh The output is passed to the edge generation module to obtain a new round of generated graph G1. The edge image fusion module fuses the output and the saved appropriate pixel value V0 according to the fusion decision graph provided by the edge image fusion module; the operation of each round is the same, and each round of image fusion module updates the fusion decision graph through the loss function value and guides the fusion of the generated edge and the appropriate edge image in that round.
[0043] Furthermore, the image GR output by the edge image fusion module is input into the edge detection module. The edge detection module processes the undetected edge pixel values and the corresponding positions K and the image loss. The image loss is used to update the edge sampling range of the next round. Through multiple iterations, a suitable generative forged region edge is finally obtained. The specific method is as follows: the edge detection module is based on an MVSS-Net network with pre-trained weights, including multiple ResNet blocks. A Sobel layer is cascaded between two adjacent ResNet blocks. An edge residual block ERB is arranged after each Sobel layer. The features output by each cascaded edge residual block ERB are summed and combined with the features output by the next cascaded edge residual block ERB, and then the shallow features and the deep features are cascaded to realize the edge detection output of the generative forged region edge.
[0044] The above-mentioned progressive combination of features from different ResNet blocks for edge detection can cascade shallow and deep networks to achieve comprehensive detection. The introduction of the Sobel layer can enhance edge-related patterns. To prevent the impact of accumulation, the combined features must pass through another ERB before the next round of feature combination.
[0045] In order to improve the sensitivity of edge pixel-level operation detection and better return to the layered module to update the edge range, the edge pixel-level loss Dice loss is used to calculate the image loss , the formula is as follows:
[0046] ;
[0047] in It indicates the edge i A binary label of whether a pixel is manipulated; manipulated pixels are marked as 0 and unmarked pixels are marked as 1. The Dice loss function effectively learns edge loss from extremely unbalanced data.
[0048] Beneficial effects: The present invention effectively improves the effect of reducing forged image traces, and shows wide applicability and effectiveness in different types of forgery methods such as copy-paste, splicing, and image modification. It performs well against both single forgery type detection method networks and mixed forgery type detection method networks. The present invention applies a strategy combining a diffusion model with manifold constraints with edge dynamic stratification and image fusion methods to the task of reducing forged image traces, and fully considers the generation sampling range of edge points and the mutual influence between generated pixels, thereby effectively improving the structural consistency between edge pixels and background areas. At the same time, the present invention uses different detection and positioning networks as edge detection modules, which not only obtains the most robust generated edges through a combination of multiple detection networks, but also can specifically improve the forged image trace reduction capability of a certain network. The edge image fusion module further enhances the performance of the model in the task of reducing forged image traces by replacing appropriate pixels at corresponding positions and adjusting the decision graph reasonably. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is the overall network structure diagram of the present invention;
[0050] Figure 2 This is a schematic diagram of the target edge acquisition and layering module processing flow in an embodiment of the present invention;
[0051] Figure 3 This is a schematic diagram of the processing flow of the edge generation module in an embodiment of the present invention;
[0052] Figure 4 This is a network architecture diagram of an edge image fusion module in an embodiment of the present invention;
[0053] Figure 5 This is a network architecture diagram of an edge detection module in an embodiment of the present invention;
[0054] Figure 6 This is a comparison chart of detection accuracy under different image fusion rounds in the embodiment;
[0055] Figure 7 It is a diagram showing the detection effects of different schemes in the embodiments. DETAILED DESCRIPTION
[0056] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.
[0057] like Figure 1 As shown, the generative edge trace reduction method based on the diffusion model of the present invention constructs and trains a generative edge trace reduction network, and inputs the tampered image X and its corresponding target forged area mask mask(X) into the generative edge trace reduction network. The specific steps are:
[0058] Step 1: In the first round of processing, the target forged area mask (X) is first converted into a gray image, and then input into the target edge acquisition and stratification module. After processing by the target edge acquisition and stratification module, the edge point coordinates E of the forged area are obtained. eg ;
[0059] Step 2: Tamper with the image X and edge point coordinates E eg An edge generation module is input, wherein the edge generation module is based on a denoising diffusion probability model DDPM and adopts a manifold constraint, updates the tampered image X through the edge generation module, and outputs an image G;
[0060] The output image G obtained in the first round is input to the edge detection module, and the image loss of the detected edge and the actual edge, the undetected edge pixel value and the corresponding position K are obtained by the edge detection module, and the image loss, the undetected edge pixel value and the corresponding position K are returned to the target edge acquisition and stratification module and the edge image fusion module;
[0061] Step 3: From the second round, the target edge acquisition and stratification module receives the image loss, the undetected edge pixel value and the corresponding position K, and the edge point coordinates E obtained in step 1 eg Subtract K to get the detected edge pixel position P, add one pixel distance to P to get the edge neighboring layer L, and the edge position Eeg and edge layer L are combined to form the edge neighboring layer combination E zh Then the updated image G is combined with the edge neighbor layer E zh Input to edge generation module;
[0062] At the same time, the edge image fusion module receives the image G, the undetected edge pixel values and the corresponding positions K, and outputs the fusion image GR;
[0063] Step 4: Starting from the second round, the edge detection module receives the fused image GR, calculates the loss image, the undetected edge pixel values and the corresponding position K again, and returns them to the target edge acquisition and stratification module and the edge image fusion module.
[0064] In the embodiment, when the forged region edge is generated in the first round, the input of the target edge acquisition and layering module is the target forged region mask (X) of the known tampered image X, and the obtained forged region edge position E eg The output is sent to the edge generation module. After that, the input of each round of target edge acquisition and stratification module is the image loss, the undetected edge pixel position K and the pixel value returned by the previous round of edge detection module, and then E eg Subtract K to get the detected edge pixel position P, add an edge neighboring layer L with a pixel distance to P, and connect the edge neighboring layer LL and edge E egAfter combining, they form E zh The output of each round of target edge acquisition and stratification module is E zh .
[0065] When generating the forged region edge in the first round, the input of the edge generation module is the tampered image X and the edge and adjacent layer combination E output by the target edge acquisition and layering module zh After that, from the second round onwards, the input of the edge generation module is the edge and adjacent layer combination E zh , after each round of updating, the tampered graph X is outputted by the edge generation module as the newly generated image G.
[0066] The edge image fusion module does not participate in the operation of generating the forged area edge in the first round; but starting from the second round, the input of each round is the generated image G output by the edge generation module, the image loss returned by the edge detection module in the previous round, the position K and pixel value of the undetected edge pixel, and the output of each round is the fusion image GR.
[0067] The input of the edge detection module is GR, and the output is the loss function and the undetected edge pixel value and position K.
[0068] like Figure 2 As shown in FIG. 1 , the first round of processing of the grayscale target forged area mask (X) by the target edge acquisition and stratification module is as follows:
[0069] Use the Get_mask_edge method to fix the x-axis and traverse the y-axis image of the grayscale image to obtain the pixel value change point, then fix the y-axis and traverse the pixel value change point obtained by the x-axis, and combine the two obtained pixel value change points to obtain the complete edge point coordinates E eg ; Set the edge point coordinates E eg The target edge layering network is passed in, and the remaining pixels within a distance of one pixel around each point coordinate are recorded as the edge neighboring layer. Since it is necessary to distinguish the pixels in the neighboring pixel layer as belonging to the inner and outer sides of the edge, especially to deal with the inner and outer distinction of irregular edges, the y-axis is fixed and the edge point coordinates on the x-axis are combined into point pairs. The edge neighboring layer pixels located in the edge point pair are marked as the inner edge neighboring layer, and other pixels are marked as the outer edge neighboring layer.
[0070] From the second round onwards, the target edge acquisition and stratification module will use the edge point coordinates E eg Subtract the undetected edge pixel position K returned by the edge detection module to obtain the detected edge pixel position P, and output the edge neighboring layer L with a pixel distance added to P. The distance θ of the edge neighboring layer L is cumulative, that is, the first round E egThe value of θ for each pixel is 0. If the edge neighboring layer θ of a pixel point q on P in the previous round expands by one pixel, and q is still on P in this round, the θ of point q will expand by another pixel.
[0071] Figure 2 The left side of the first row is the irregular grayscale target forged area mask (X); Figure 2 The right side of the first row shows the complete edge image obtained after the mask edge acquisition block is processed; Figure 2 The left side of the second row is the image details. The dotted line in the figure indicates that on the fixed y-axis, every two edge points passed by the straight line form a point pair. The edge neighboring layer pixels inside the point pair are marked as the inner edge neighboring layer (Mask_edge_layer(x)), and the remaining points are marked as the outer edge neighboring layer (Mask_edge_layer(x); Figure 2 The second row on the right shows the resulting edge stratification.
[0072] Furthermore, the diffusion model of the edge generation module includes a forward diffusion process and a reverse diffusion process, the details of which are as follows;
[0073] , Each step of the forward diffusion process is a Gaussian translation;
[0074] (1)
[0075] in is a fixed variance schedule, which refers to a linearly increasing noise schedule ;
[0076] Formula (1) adds Gaussian noise to the latent variable to find , through (1) it can be deduced that when given clean data hour, The sampling of will be expressed in closed form:
[0077] (2)
[0078] in and ;
[0079] Expressed as and A linear combination of: (3)
[0080] The back diffusion process is modeled by a neural network, which predicts the parameters of the Gaussian distribution and ; The reverse diffusion process has the same functional form as the forward process and is expressed as a Gaussian transform with a learned mean and fixed variance, expressed as:
[0081] (4)
[0082] In addition, by Decompose into Sum Noise Approximator The linear combination of , the generation process is expressed as:
[0083] (5)
[0084] In the above formula, Each back-diffusion step is random, Represents a neural network with the same input and output dimensions;
[0085] , in the backward diffusion process, an additional correction term inspired by the manifold constraint is used to make the backward iteration close to the manifold, that is, the initial data manifold of the original tampered image is recorded as M, and the forward diffusion process deviates from the data manifold M and gradually approaches the noise data manifold ,like Figure 3 shown.
[0086] Figure 3 (a) in the figure explains the forward diffusion and reverse generation process of the diffusion model from the data manifold. The forward process means that the data gradually diffuses from the clear state data manifold M to the noisy state data manifold Mi, and the reverse process means that the data gradually diffuses from the noisy state data manifold M to the noisy state data manifold Mi. ∞ Gradually restore to the clear state data manifold M. Figure 3 (b) shows that compared with traditional diffusion generation, the addition of manifold constraints effectively reduces the back-diffusion steps required to leave the manifold.
[0087] In this embodiment, starting from the second round, the edge image fusion module retains the appropriate edge pixel positions and corresponding pixel values of this round, and uses the retained pixels to replace the pixels at the corresponding positions on the original tampered image X. The remaining parts are sent to the target edge acquisition and layering module to update the edge range, and then passed to the edge generation module to generate new edges and corresponding pixel value images, and fused to obtain the fusion image GR; expressed as follows:
[0088] ;
[0089] ;
[0090] ;
[0091] In the above formula, It refers to the original tampered image X after removing the image edge.
[0092] The specific details are as follows: After the first round of target edge acquisition and stratification module obtains the edge point of the target area, the coordinates of the edge point are set to E eg , and set an adjustable distance selection parameter θ for each pixel point, directly converting E eg The edge layer of the state is sent to the second part edge generation module, and the output result of the edge generation module is recorded as , the edge image fusion module does not process this output, It will be directly sent to the edge detection module, and the error function result of the detected edge and the actual edge will be obtained through the edge detection module, and the result will be returned to the target edge acquisition and layering module and the edge image fusion module; at this time, the decision diagram of the edge image fusion module will save the edge pixel positions that were not correctly identified by the detection network in the first round And the pixel value V0 at that position, replace and update the corresponding pixel value on the original input image X The target edge acquisition and stratification module converts the pixel range E eg Adjust to E zh , E zh The output is passed to the edge generation module to obtain a new round of generated graph G1. The edge image fusion module fuses the output and the saved appropriate pixel value V0 according to the fusion decision graph provided by the edge image fusion module; the operation of each round is the same, and each round of image fusion module updates the fusion decision graph through the loss function value and guides the fusion of the generated edge and the appropriate edge image in that round.
[0093] like Figure 4 As shown, the pixel on the left of the edge detection module output The rest is sent to the target edge acquisition and layering module to update the range, and then to the edge generation module. The newly generated edges and The pixel value images are fused to obtain a new edge image.
[0094] like Figure 5As shown, the image GR output by the edge image fusion module is input into the edge detection module. The edge detection module processes the undetected edge pixel values and the corresponding positions K and the image loss. The image loss is used to update the edge sampling range of the next round. After multiple iterations, the appropriate generative forged region edge is finally obtained. The specific method is as follows: the edge detection module is based on the MVSS-Net network with pre-trained weights, including multiple ResNet blocks. Sobel layers are cascaded between two adjacent ResNet blocks. An edge residual block ERB is arranged after each Sobel layer. The features output by each cascaded edge residual block ERB are summed and combined with the features output by the next cascaded edge residual block ERB, and then the shallow features and deep features are cascaded to achieve edge detection output of generative forged region edges.
[0095] Figure 5 (a) is the overall structure diagram of the edge image fusion module. Figure 5 (b) is the Sobel layer structure diagram; Figure 5 (c) in the figure is the structure diagram of the edge residual block ERB.
[0096] Use Dice loss function to calculate image loss , the formula is as follows:
[0097] ;
[0098] in It indicates the edge i A binary label indicating whether a pixel is manipulated; manipulated pixels are marked as 0, and unmarked pixels are marked as 1.
[0099] In order to verify the technical effect of the technical solution of the present invention, this embodiment uses CASIA v1.0, CASIA v2.0, and 200 image modification forged data based on the diffusion method to eliminate image forgery traces. CASIA mainly focuses on splicing and copy-paste images. The tampered areas selected are small and fine, and some forged images are post-processed by filtering and blurring. It is divided into two versions: CASIAv2.0 (5123 samples) for training and CASIA v1.0 (921 samples) for testing; both provide binary groundtruth for evaluation. In order to distinguish the impact of different types of tampering methods on the present invention, the data of the data set is subdivided as follows.
[0100] Table 1 Three types of forged images and their numbers in the forged dataset
[0101]
[0102] The experimental configuration of this embodiment is: implemented by the open source PyTorch deep learning framework and trained using a single NVIDIA GeForce RTX 3060. Considering the configuration of the server, the image size is adjusted to 256×256 to facilitate the calculation of the diffusion model.
[0103] The experimental results and analysis of this embodiment are as follows. The first part is the ablation experiment, which includes the detail ablation experiment of edge acquisition and layering method, the ablation experiment of image fusion method and the ablation experiment of replacement detection network; the second part is the detection experiment of different forged image types, which aims to measure the effectiveness of the method proposed in this paper for different types of forged images.
[0104] The main goal of this embodiment is to reduce the traces of forgery. and the original fake image The images are sent to multiple forgery detection networks together. Since the technical solution of the present invention pays more attention to the image detection effect on the edge of the target area, a forged image identification network with forged area positioning is specially selected. Considering the time cost and experimental efficiency, the number of rounds of edge generation in the edge generation module is fixed to 3, and the generated image of the last round is taken as the final output of the proposed technical solution.
[0105] For pixel-level manipulation detection, pixel-level precision and recall are calculated, and their F1 is calculated. For image-level manipulation detection, in order to measure the missed detection rate and false alarm rate, sensitivity, specificity and their F1 are calculated. Due to the randomness of the experiment, this embodiment conducts 10 independent trainings on each detection network model, and takes the average of the detection accuracy of these 10 experimental results as the detection result of the model for the forged image obtained by the proposed trace reduction method, thereby reducing the impact of experimental errors.
[0106] This embodiment shows a total of three ablation experiments: edge acquisition and layering method detail ablation experiment, image fusion method ablation experiment and replacement detection network ablation experiment.
[0107] 1) Detail ablation experiment of edge acquisition and layering method. The preprocessing method of the present invention is to obtain the precise position of the edge of the target area through the mask grayscale image of the target area of the forged image. Since it is a trace reduction work for the forged image, the mask of the target area can usually be obtained before the experiment. The setting of the target edge acquisition and layering module can make each edge pixel value obtain the most appropriate generation sampling range. If the target edge acquisition and layering modules are cancelled, it is equivalent to the ordinary generative image restoration based on the diffusion model.
[0108] Table 2 shows the impact of removing the edge layering module on the accuracy of trace reduction of the proposed method, where the F1 value represents the original detection accuracy, and F1(XF ) value represents the detection accuracy of the forged image XF after removing the layered module, and the forged image XF after the complete experiment G The detection accuracy is F1 (X G ).
[0109] Table 2 Effect of removing layered modules on detection accuracy in different forgery detection and positioning networks
[0110]
[0111] 2) Explore the influence of the number of image fusion rounds on the trace reduction effect through the ablation experiment of the edge image fusion method. Since the technical solution of the present invention is to continuously update the edge pixel generation sampling range through multiple rounds of loss value feedback, the selection of the number of rounds has a high impact on the efficiency and time complexity of the entire experiment. Too few rounds may not allow each edge pixel to obtain the best generation sampling range, and too many rounds may make the experiment take too long. In this embodiment, the number of experimental rounds of the data set is set to 8 rounds, and the forged images generated in each round are sent to different detection networks. The detection accuracy is shown in Table 3:
[0112] Table 3 Effect of different image fusion rounds on detection accuracy
[0113]
[0114] From the analysis of the experimental results in Table 3, it can be found that when the number of image fusion rounds exceeds 3, the image detection accuracy tends to stabilize, and the detection accuracy decreases slightly with the increase of subsequent rounds. Considering the time cost of the experiment, this paper takes 3 rounds of image fusion as the optimal number of image fusion rounds.
[0115] 3) Ablation experiment of replacing detection network. In order to explore the influence of different types of detection networks acting as edge detection modules on the generative forged edges, this embodiment sets an ablation experiment of replacing the detection network.
[0116] The detection network of the edge detection module in the technical solution of the present invention is composed of edge detection branches on a single detection network (MVSS-Net). The loss function returned by the detection branch guides the update of the edge pixel range of the edge layering module, thereby affecting different edge generation effects.
[0117] To verify that the effectiveness is independent of the selection of the detection network, this embodiment replaces multiple forgery detection and positioning networks as the detection network of the edge detection module for testing. After the number of image fusion rounds is fixed at 3, the forgery detection network in the edge detection module is replaced, and the other modules remain unchanged. The obtained forged image is sent to the detection network to calculate the F1 value. Table 4 shows the experimental results on the data set when different detection and positioning networks serve as edge detection modules.
Claims
1. A generative edge trace reduction method based on a diffusion model, characterized in that: Construct and train a generative edge trace reduction network, input the tampered image X and its corresponding target forged area mask mask(X) into the generative edge trace reduction network, the specific steps are: Step 1: In the first round of processing, the target forged area mask (X) is first converted into a gray image, and then input into the target edge acquisition and stratification module. After processing by the target edge acquisition and stratification module, the edge point coordinates E of the forged area are obtained. eg ; Step 2: Tamper with the image X and edge point coordinates E eg An edge generation module is input, wherein the edge generation module is based on a denoising diffusion probability model DDPM and adopts a manifold constraint, updates the tampered image X through the edge generation module, and outputs an image G; The output image G obtained in the first round is input to the edge detection module, and the image loss of the detected edge and the actual edge, the undetected edge pixel value and the corresponding position K are obtained by the edge detection module, and the image loss, the undetected edge pixel value and the corresponding position K are returned to the target edge acquisition and stratification module and the edge image fusion module; Step 3: From the second round, the target edge acquisition and stratification module receives the image loss, the undetected edge pixel value and the corresponding position K, and the edge point coordinates E obtained in step 1 eg Subtract K to get the detected edge pixel position P, add one pixel distance to P to get the edge neighboring layer L, and the edge position Eeg and edge layer L are combined to form the edge neighboring layer combination E zh Then the updated image G is combined with the edge neighbor layer E zh Input to edge generation module; At the same time, the edge image fusion module receives the image G, the undetected edge pixel values and the corresponding positions K, and outputs the fusion image GR; Step 4: Starting from the second round, the edge detection module receives the fused image GR, calculates the loss image, the undetected edge pixel values and the corresponding position K again, and returns them to the target edge acquisition and stratification module and the edge image fusion module.
2. The generative edge trace reduction method based on the diffusion model according to claim 1, characterized in that: The first round of processing of the grayscale target forged area mask (X) by the target edge acquisition and stratification module is as follows: Use the Get_mask_edge method to fix the x-axis and traverse the y-axis image of the grayscale image to obtain the pixel value change point, then fix the y-axis and traverse the pixel value change point obtained by the x-axis, and combine the two obtained pixel value change points to obtain the complete edge point coordinates E eg ; Set the edge point coordinates E eg The target edge layering network is passed in, and the remaining pixels within a distance of one pixel around each point coordinate are recorded as the edge neighboring layer. Since it is necessary to distinguish the pixels in the neighboring pixel layer as belonging to the inner and outer sides of the edge, especially to deal with the inner and outer distinction of irregular edges, the y-axis is fixed and the edge point coordinates on the x-axis are combined into point pairs. The edge neighboring layer pixels located in the edge point pair are marked as the inner edge neighboring layer, and other pixels are marked as the outer edge neighboring layer. From the second round onwards, the target edge acquisition and stratification module will use the edge point coordinates E eg Subtract the undetected edge pixel position K returned by the edge detection module to obtain the detected edge pixel position P, and output the edge neighboring layer L with a pixel distance added to P. The distance θ of the edge neighboring layer L is cumulative, that is, the first round E eg The value of θ for each pixel is 0. If the edge neighboring layer θ of a pixel point q on P in the previous round expands by one pixel, and q is still on P in this round, the θ of point q will expand by another pixel.
3. The generative edge trace reduction method based on diffusion model according to claim 1, characterized in that: The diffusion model of the edge generation module includes the forward diffusion process and the reverse diffusion process, the details of which are as follows; , Each step of the forward diffusion process is a Gaussian translation; (1) in is a fixed variance schedule, which refers to a linearly increasing noise schedule , Indicates how the initial value of the data x changes from time step t-1 to time t during the forward diffusion process. is the added Gaussian perturbation noise; Formula (1) adds Gaussian noise to the latent variable to find , through (1) it can be deduced that when given clean data When , the sampling of will be expressed in closed form: (2) in and ; Here s and t represent the time step of the diffusion model; Expressed as and A linear combination of: (3) is the Gaussian disturbance noise that conforms to the normal distribution; The back diffusion process is modeled by a neural network, which predicts the parameters of the Gaussian distribution and ; The reverse diffusion process has the same functional form as the forward process and is expressed as a Gaussian transform with a learned mean and fixed variance, expressed as: (4) In addition, by Decompose into Sum Noise Approximator The linear combination of , the generation process is expressed as: (5) In the above formula, Each back-diffusion step is random, represents a pair of neural networks with the same input and output dimensions The predicted value at time step t; , in the backward diffusion process, an additional correction term inspired by the manifold constraint is used to make the backward iteration close to the manifold, that is, the initial data manifold of the original tampered image is recorded as M, and the forward diffusion process deviates from the data manifold M and gradually approaches the noise data manifold .
4. The generative edge trace reduction method based on diffusion model according to claim 1, characterized in that: From the second round onwards, the edge image fusion module retains the appropriate edge pixel positions and corresponding pixel values of this round, and uses the retained pixels to replace the pixels at the corresponding positions on the original tampered image X. The rest is sent to the target edge acquisition and layering module to update the edge range, and then passed to the edge generation module to generate the new edge and corresponding pixel value images, and the fusion image GR is obtained by fusion; it is expressed as follows: ; ; ; In the above formula, It refers to the original tampered image X after removing the image edge. It means that for time step t, the pixel at position p on the forged edge is the result inferred by the reverse diffusion process and has the same noise level as the forward diffusion at time t.
5. The generative edge trace reduction method based on diffusion model according to claim 1, characterized in that: The image GR output by the edge image fusion module is input into the edge detection module. The edge detection module processes the undetected edge pixel values and the corresponding positions K and image losses. The image losses are used to update the edge sampling range of the next round. The appropriate generative forged region edge is finally obtained through multiple iterations. The specific method is as follows: the edge detection module is based on the MVSS-Net network with pre-trained weights, including multiple ResNet blocks. Sobel layers are cascaded between two adjacent ResNet blocks. An edge residual block ERB is arranged after each Sobel layer. The features output by each cascaded edge residual block ERB are summed and combined with the features output by the next cascaded edge residual block ERB, and then the shallow features and deep features are cascaded to achieve edge detection output of generative forged region edge. Use Dice loss function to calculate image loss , the formula is as follows: ; in It indicates the edge i A binary label indicating whether a pixel is manipulated; manipulated pixels are marked as 0, and unmarked pixels are marked as 1.
Citation Information
Patent Citations
Infrared blind pixel compensation method based on generative adversarial network
CN111369449A
Detection method and system for certificate document image tampering
CN114419633A