An image detection method and device

The proposed detection model improves image tampering detection accuracy by employing region and boundary detection layers to analyze pixel value differences, enhancing precision in identifying tampered regions and boundaries.

CN113763405BActive Publication Date: 2025-07-15BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110142944.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-02
Publication Date
2025-07-15
Estimated Expiration
2041-02-02

AI Technical Summary

Technical Problem

When detecting image tampering, the prior art focuses only on the local features of the image, resulting in low detection accuracy.

Method used

An image detection method is adopted to obtain training samples, including training images, area labels and boundary labels, and use the detection model to perform area detection and boundary detection. Combined with feature extraction layer, area detection layer and boundary detection layer, tampered areas and tampered boundaries are identified to improve detection accuracy.

Benefits of technology

By paying attention to the overall and local feature differences of the image, the tampering area is accurately determined, which improves the accuracy of image detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113763405B_ABST
    Figure CN113763405B_ABST
Patent Text Reader

Abstract

The present invention discloses an image detection method and apparatus, relating to the field of computer technology. A specific embodiment of the method includes: obtaining training samples; wherein, the training samples include: training images, region labels, and boundary labels; inputting the training images into a detection model to obtain region detection results and boundary detection results; training the detection model according to the region labels, the boundary labels, the region detection results, and the boundary detection results; and determining whether a detection image is tampered with based on the trained detection model. This embodiment can improve the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to an image detection method and apparatus. Background Art

[0002] In actual application scenarios, criminals synthesize the content of multiple images into one image, changing the original meaning of the image and misleading users. For example, in an e-commerce platform, merchants tamper with the original image to attract consumers. Therefore, how to detect whether an image has been tampered with has become an urgent problem to be solved.

[0003] The prior art identifies whether an image has been tampered with through edge detection.

[0004] However, this method only focuses on the local features of the image, and its detection accuracy is low. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide an image detection method and apparatus, which can improve the detection accuracy.

[0006] In a first aspect, embodiments of the present invention provide an image detection method, including:

[0007] Obtaining a training sample; wherein, the training sample includes: a training image, a region label, and a boundary label;

[0008] Inputting the training image into a detection model to obtain a region detection result and a boundary detection result;

[0009] Training the detection model according to the region label, the boundary label, the region detection result, and the boundary detection result;

[0010] Based on the trained detection model, determining whether a detection image has been tampered with.

[0011] Optionally,

[0012] The detection model includes: a feature extraction layer, a region detection layer, and a boundary detection layer;

[0013] The inputting the training image into the detection model to obtain a region detection result and a boundary detection result includes:

[0014] Inputting the training image into the feature extraction layer to extract a high-order feature map and a low-order feature map from the training image;

[0015] Inputting the high-order feature map and the low-order feature map into the region detection layer to obtain the region detection result;

[0016] Input the high-order feature map and the low-order feature map into the boundary detection layer to obtain the boundary detection result.

[0017] Optionally,

[0018] The inputting the training image into the feature extraction layer to extract a high-order feature map and a low-order feature map from the training image includes:

[0019] Input the training image into the backbone network to obtain the low-order feature map and the first feature map;

[0020] Extract multi-scale features from the first feature map based on the multi-scale network to obtain a plurality of second feature maps;

[0021] Input the plurality of second feature maps after splicing into the first convolutional layer to obtain the high-order feature map;

[0022] Wherein, the backbone network includes: a first multi-channel convolutional layer and a depthwise separable convolutional layer; the first convolutional layer is a 1×1 convolutional layer.

[0023] Optionally,

[0024] The multi-scale network includes: an atrous convolutional layer, a second convolutional layer, and a pooling layer;

[0025] Wherein, the second convolutional layer is a 1×1 convolutional layer.

[0026] Optionally,

[0027] The region detection layer includes: a first feature fusion layer, a region anomaly analysis layer, and a first result output layer;

[0028] The inputting the high-order feature map and the low-order feature map into the region detection layer to obtain the region detection result includes:

[0029] Input the high-order feature map and the low-order feature map into the first feature fusion layer to obtain a third feature map;

[0030] Determine a fourth feature map according to the third feature map and the region anomaly analysis layer; wherein, the fourth feature map is used to characterize the pixel value difference between the tampered region and the background region in the third feature map;

[0031] Input the fourth feature map into the first result output layer to obtain the region detection result.

[0032] Optionally,

[0033] The inputting the high-order feature map and the low-order feature map into the first feature fusion layer to obtain a third feature map includes:

[0034] Input the low-order feature map into the third convolutional layer to obtain the fifth feature map;

[0035] Upsample the high-order feature map to obtain the sixth feature map;

[0036] Concatenate the fifth feature map and the sixth feature map and input them into the second multi-channel convolutional layer to obtain the third feature map;

[0037] Among them, the third convolutional layer is a 1×1 convolutional layer.

[0038] Optionally,

[0039] Determining the fourth feature map according to the third feature map and the regional anomaly analysis layer includes:

[0040] Calculate the average pixel value of the third feature map according to the pixel values of each pixel coordinate in the third feature map;

[0041] Determine the difference between the pixel value of each pixel coordinate and the average pixel value;

[0042] Calculate the standard deviation of the pixel values of the third feature map according to the difference between the pixel value of each pixel coordinate and the average pixel value;

[0043] Calculate the normalized pixel value of each pixel coordinate according to the standard deviation of the pixel values and the difference between the pixel value of each pixel coordinate and the average pixel value;

[0044] Determine the fourth feature map according to the normalized pixel values of each pixel coordinate.

[0045] Optionally,

[0046] The step of inputting the fourth feature map into the first result output layer to obtain the regional detection result includes:

[0047] Input the fourth feature map into the fourth convolutional layer to obtain the seventh feature map;

[0048] Upsample the seventh feature map to obtain the eighth feature map;

[0049] Input the eighth feature map into the activation function to obtain the regional detection result;

[0050] Among them, the fourth convolutional layer is a 1×1 convolutional layer.

[0051] Optionally,

[0052] The boundary detection layer includes: a second feature fusion layer, a boundary anomaly analysis layer, and a second result output layer;

[0053] Inputting the high-order feature map and the low-order feature map into the boundary detection layer to obtain the boundary detection result includes:

[0054] Inputting the high-order feature map and the low-order feature map into the second feature fusion layer to obtain a ninth feature map;

[0055] Determining a tenth feature map according to the ninth feature map and the boundary anomaly analysis layer; wherein, the tenth feature map is used to represent the pixel value difference between the tampered area and the background area within the detection window;

[0056] Inputting the tenth feature map into the second result output layer to obtain the boundary detection result.

[0057] Optionally,

[0058] The step of inputting the high-order feature map and the low-order feature map into the second feature fusion layer to obtain a ninth feature map includes:

[0059] Inputting the low-order feature map into a fifth convolutional layer to obtain an eleventh feature map;

[0060] Upsampling the high-order feature map to obtain a twelfth feature map;

[0061] Concatenating the eleventh feature map and the twelfth feature map and inputting the concatenated result into a third multi-channel convolutional layer to obtain the ninth feature map;

[0062] Wherein, the fifth convolutional layer is a 1×1 convolutional layer.

[0063] Optionally,

[0064] The step of determining a tenth feature map according to the ninth feature map and the boundary anomaly analysis layer includes:

[0065] Calculating the average pixel value of the detection window according to the pixel values of each pixel coordinate in the ninth feature map within the detection window;

[0066] Determining the difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located;

[0067] Calculating the standard deviation of the pixel values of the ninth feature map;

[0068] Calculating the normalized pixel value of the pixel coordinate within the detection window according to the standard deviation of the pixel values and the difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located;

[0069] Determining the tenth feature map according to the normalized pixel values of the pixel coordinates within the detection window.

[0070] Optionally,

[0071] Inputting the tenth feature map into the second result output layer to obtain the boundary detection result includes:

[0072] Inputting the tenth feature map into a sixth convolutional layer to obtain a thirteenth feature map;

[0073] Performing upsampling on the thirteenth feature map to obtain a fourteenth feature map;

[0074] Inputting the fourteenth feature map into an activation function to obtain the region detection result;

[0075] Wherein, the sixth convolutional layer is a 1×1 convolutional layer.

[0076] Optionally,

[0077] Obtaining the training samples includes:

[0078] Obtaining the training image and the region label;

[0079] Performing a dilation operation on the region label to obtain a dilated image;

[0080] Performing an erosion operation on the region label to obtain an eroded image;

[0081] Determining the boundary label according to the dilated image and the eroded image.

[0082] Optionally,

[0083] Further includes:

[0084] Obtaining pre-training samples;

[0085] Pre-training the detection model based on the pre-training samples;

[0086] Inputting the training image into the detection model to obtain a region detection result and a boundary detection result includes:

[0087] Inputting the training image into the pre-trained detection model to obtain the region detection result and the boundary detection result.

[0088] In a second aspect, an embodiment of the present invention provides an image detection device, including:

[0089] An acquisition module configured to acquire training samples; wherein, the training samples include: a training image, a region label, and a boundary label;

[0090] A training module, configured to input the training image into a detection model to obtain a region detection result and a boundary detection result; and train the detection model according to the region label, the boundary label, the region detection result, and the boundary detection result.

[0091] A detection module, configured to determine whether a detection image is tampered with based on the trained detection model.

[0092] In a third aspect, an embodiment of the present invention provides an electronic device, including:

[0093] One or more processors;

[0094] A storage device for storing one or more programs,

[0095] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any of the above embodiments.

[0096] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method as described in any of the above embodiments is implemented.

[0097] One of the above embodiments of the invention has the following advantages or beneficial effects: By performing boundary detection and region detection on an image based on a detection model, the region detection identifies a tampered region based on the feature differences between the tampered region and the background region of the entire image, focusing on the overall features of the image; the boundary detection identifies a tampered boundary based on the feature differences on both sides of the tampered boundary, focusing on local features. The boundary detection can assist the region detection to more accurately determine the tampered region and improve the accuracy of image detection.

[0098] The further effects of the above non-conventional optional manners will be described in combination with specific embodiments below. Description of the Drawings

[0099] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0100] Figure 1 is a flowchart of an image detection method provided by an embodiment of the present invention;

[0101] Figure 2 is a flowchart of an image detection method provided by an embodiment of the present invention;

[0102] Figure 3 is an architecture diagram of a detection model provided by an embodiment of the present invention;

[0103] Figure 4(a) is a schematic diagram of a regional label provided by an embodiment of the present invention;

[0104] Figure 4 (b) is a schematic diagram of a dilated image provided by an embodiment of the present invention;

[0105] Figure 4 (c) is a schematic diagram of an eroded image provided by an embodiment of the present invention;

[0106] Figure 4 (d) is a schematic diagram of a boundary label provided by an embodiment of the present invention;

[0107] Figure 5 is a schematic structural diagram of a backbone network provided by an embodiment of the present invention;

[0108] Figure 6 is a schematic structural diagram of a multi-scale network provided by an embodiment of the present invention;

[0109] Figure 7 is a schematic structural diagram of an image detection device provided by an embodiment of the present invention;

[0110] Figure 8 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0111] Figure 9 is a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0112] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0113] Edge detection focuses on the local features of the image within the detection frame and does not consider the feature differences between the tampered area and the background area of the entire image. Therefore, the accuracy of its detection results needs to be further improved.

[0114] In view of this, as Figure 1 shown, the embodiments of the present invention provide an image detection method, including:

[0115] Step 101: Obtain training samples; wherein, the training samples include: training images, regional labels, and boundary labels.

[0116] To improve the training effect, the training samples in the embodiments of the present invention are from the tampered image dataset CASIA 2.0. CASIA 2.0 can provide more than five thousand tampered images and involves various tampering methods and image formats, which can meet the training requirements of the embodiments of the present invention. In actual application scenarios, the CASIA 1.0 dataset with a relatively small number of samples can also be selected according to the actual situation, etc.

[0117] The dataset includes training images and region labels, and the boundary labels are determined according to the region labels.

[0118] Step 102: Input the training images into the detection model to obtain region detection results and boundary detection results.

[0119] The region detection results are the predicted tampered regions, and the boundary detection results are the predicted tampered boundaries.

[0120] Step 103: Train the detection model according to the region labels, boundary labels, region detection results, and boundary detection results.

[0121] Determine the loss value according to the region labels, boundary labels, region detection results, boundary detection results, and a preset loss function; adjust the parameters of the detection model according to the loss value.

[0122] In actual application scenarios, to ensure the prediction quality of the detection model, during the training process, the detection model is tested using test samples to determine the prediction effect of the detection model. Specifically, the quantity ratio of training samples to prediction samples can be 9:1. The test samples can be from CASIA 1.0 and the Columbia dataset.

[0123] Step 104: Based on the trained detection model, determine whether the detected image is tampered.

[0124] In the embodiments of the present invention, boundary detection and region detection are performed on the image based on the detection model. The region detection is based on the feature differences between the tampered regions and the background regions of the entire image to identify the tampered regions, which focuses on the overall features of the image; the boundary detection is based on the feature differences on both sides of the tampered boundary to identify the tampered boundary, which focuses on the local features. The boundary detection can assist the region detection to more accurately determine the tampered regions and improve the accuracy of image detection.

[0125] In an embodiment of the present invention, the detection model includes: a feature extraction layer, a region detection layer, and a boundary detection layer;

[0126] Inputting the training images into the detection model to obtain region detection results and boundary detection results includes:

[0127] Input the training image into the feature extraction layer to extract high-order feature maps and low-order feature maps from the training image;

[0128] Input the high-order feature maps and low-order feature maps into the region detection layer to obtain the region detection result;

[0129] Input the high-order feature maps and low-order feature maps into the boundary detection layer to obtain the boundary detection result.

[0130] In the embodiment of the present invention, the region detection result and the boundary detection result are determined based on the detection model. Based on different functions, the detection model can be divided into a feature extraction layer, a region detection layer, and a boundary detection layer. Among them, the feature extraction layer is used to extract high-order features and low-order features from the training image. The extracted high-order features form high-order feature maps, and the extracted low-order features form low-order feature maps. The low-order features have higher resolution and contain position information, detail information, etc.; the high-order features have more semantic information but lower resolution. The region detection layer is used to detect the tampered region, and the boundary detection layer is used to detect the tampered boundary.

[0131] In the embodiment of the present invention, high-order features and low-order features are used for boundary detection and region detection, considering the multi-dimensional features in the training image, improving the accuracy of region detection and boundary detection, and further improving the accuracy of the tampering recognition result.

[0132] In an embodiment of the present invention, inputting the training image into the feature extraction layer to extract high-order feature maps and low-order feature maps from the training image includes:

[0133] Input the training image into the backbone network to obtain a low-order feature map and a first feature map;

[0134] Extract multi-scale features from the first feature map based on the multi-scale network to obtain multiple second feature maps;

[0135] Input the concatenated multiple second feature maps into the first convolutional layer to obtain high-order feature maps;

[0136] Among them, the backbone network includes: a first multi-channel convolutional layer and a depthwise separable convolutional layer; the first convolutional layer is a 1×1 convolutional layer.

[0137] In the embodiment of the present invention, the feature extraction layer includes a backbone network, a multi-scale network, and a first convolutional layer. The backbone network is used to extract features from the training image, which can be implemented by a multi-channel convolutional layer and a depthwise separable convolution. The combination of the first multi-channel convolutional layer and the depthwise separable convolutional layer can improve the feature extraction efficiency and quality. In the embodiment of the present invention, the backbone network may include multiple first multi-channel convolutional layers and multiple depthwise separable convolutional layers, and the depthwise separable convolutional layer may also be replaced with a first multi-channel convolutional layer or other types of convolutional layers.

[0138] In the embodiment of the present invention, the accuracy of region detection is improved by extracting multi-scale features, thereby improving the accuracy and reliability of the image detection result.

[0139] The first convolutional layer is used to fuse the second feature map to obtain a high-order feature map.

[0140] In an embodiment of the present invention, the multi-scale network includes: an atrous convolutional layer, a second convolutional layer, and a pooling layer;

[0141] Among them, the second convolutional layer is a 1×1 convolutional layer.

[0142] In the embodiment of the present invention, the atrous convolutional layer enlarges the receptive field, enabling the multi-scale network to output more abundant information, thereby improving the model training effect. The embodiment of the present invention may include multiple atrous convolutional layers, such as three or four.

[0143] In an embodiment of the present invention, the region detection layer includes: a first feature fusion layer, a region anomaly analysis layer, and a first result output layer;

[0144] Inputting the high-order feature map and the low-order feature map into the region detection layer to obtain a region detection result, including:

[0145] Inputting the high-order feature map and the low-order feature map into the first feature fusion layer to obtain a third feature map;

[0146] Determining a fourth feature map according to the third feature map and the region anomaly analysis layer; wherein, the fourth feature map is used to characterize the pixel value difference between the tampered region and the background region in the third feature map;

[0147] Inputting the fourth feature map into the first result output layer to obtain a region detection result.

[0148] In the embodiment of the present invention, the range of the tampered region is determined based on the pixel value difference between the tampered region and the background region. In the process of calculating the pixel value difference, the pixel values of each pixel coordinate in the third feature map are considered, and the tampered region can be recognized from a global perspective.

[0149] In an embodiment of the present invention, inputting the high-order feature map and the low-order feature map into the first feature fusion layer to obtain a third feature map, including:

[0150] Inputting the low-order feature map into the third convolutional layer to obtain a fifth feature map;

[0151] Upsampling the high-order feature map to obtain a sixth feature map;

[0152] Concatenating the fifth feature map and the sixth feature map and inputting them into the second multi-channel convolutional layer to obtain a third feature map;

[0153] Among them, the third convolutional layer is a 1×1 convolutional layer.

[0154] In the embodiment of the present invention, the 1×1 convolutional layer is used to perform feature fusion and compression on the low-order feature map, so as to remove redundant features and improve the model training effect. In addition, bilinear interpolation or transposed convolution can be used to upsample the high-order feature map to enlarge the size of the high-order feature map. The fifth feature map and the sixth feature map can be concatenated along the Z axis and then feature fusion is performed through the 1×1 convolutional layer.

[0155] In an embodiment of the present invention, determining the fourth feature map according to the third feature map and the regional anomaly analysis layer includes:

[0156] Calculating the average pixel value of the third feature map according to the pixel values of each pixel coordinate in the third feature map;

[0157] Determining the difference between the pixel value of each pixel coordinate and the average pixel value;

[0158] Calculating the standard deviation of the pixel values of the third feature map according to the difference between the pixel value of each pixel coordinate and the average pixel value;

[0159] Calculating the normalized pixel value of each pixel coordinate according to the standard deviation of the pixel values and the difference between the pixel value of each pixel coordinate and the average pixel value;

[0160] Determining the fourth feature map according to the normalized pixel values of each pixel coordinate.

[0161] In the embodiment of the present invention, the difference degree between the pixel value of the pixel coordinate and the average pixel value of the third feature map is characterized by the normalized pixel value. The greater the difference degree, the greater the possibility that the pixel coordinate is located in the tampered area. In the embodiment of the present invention, whether the pixel coordinate is in the tampered area is identified through the difference of the pixel values, which can improve the accuracy of area recognition.

[0162] In an embodiment of the present invention, inputting the fourth feature map into the first result output layer to obtain the region detection result includes:

[0163] Inputting the fourth feature map into the fourth convolutional layer to obtain the seventh feature map;

[0164] Upsampling the seventh feature map to obtain the eighth feature map;

[0165] Inputting the eighth feature map into the activation function to obtain the region detection result;

[0166] Among them, the fourth convolutional layer is a 1×1 convolutional layer.

[0167] In the embodiment of the present invention, the image is enlarged by upsampling, and the pixel values are mapped to the range between 0 and 1 through an activation function. The mapping results of the pixel coordinates of each pixel in the eighth feature map constitute the region detection result. In the embodiment of the present invention, dimensionality reduction and amplification are performed after calculating the normalized pixel values, which can ensure the accuracy of the pixel values used in the calculation process and improve the accuracy and reliability of the region detection result. The activation function used can be the sigmoid function, the softmax function, etc.

[0168] In an embodiment of the present invention, the boundary detection layer includes: a second feature fusion layer, a boundary anomaly analysis layer, and a second result output layer;

[0169] Inputting the high-order feature map and the low-order feature map into the boundary detection layer to obtain the boundary detection result, including:

[0170] Inputting the high-order feature map and the low-order feature map into the second feature fusion layer to obtain the ninth feature map;

[0171] Determining the tenth feature map according to the ninth feature map and the boundary anomaly analysis layer; wherein, the tenth feature map is used to characterize the pixel value difference between the tampered area and the background area within the detection window;

[0172] Inputting the tenth feature map into the second result output layer to obtain the boundary detection result.

[0173] In the embodiment of the present invention, the range of the tampered area is determined based on the pixel value difference between the tampered area and the background area within the detection window. Different from region detection, the embodiment of the present invention considers the pixel values of each pixel coordinate in the detection window and can assist the region detection process to determine the tampered area from a local perspective.

[0174] In an embodiment of the present invention, inputting the high-order feature map and the low-order feature map into the second feature fusion layer to obtain the ninth feature map includes:

[0175] Inputting the low-order feature map into the fifth convolutional layer to obtain the eleventh feature map;

[0176] Performing upsampling on the high-order feature map to obtain the twelfth feature map;

[0177] Concatenating the eleventh feature map and the twelfth feature map and inputting them into the third multi-channel convolutional layer to obtain the ninth feature map;

[0178] Wherein, the fifth convolutional layer is a 1×1 convolutional layer.

[0179] Similar to the region detection part, in the embodiments of the present invention, a 1×1 convolutional layer is used to perform feature fusion and compression on the low-order feature map, so as to remove redundant features and improve the model training effect. In addition, bilinear interpolation or transposed convolution can be used to upsample the high-order feature map to enlarge the size of the high-order feature map. The eleventh feature map and the twelfth feature map can be concatenated along the Z-axis and then feature fusion is performed through a 1×1 convolutional layer.

[0180] In one embodiment of the present invention, determining the tenth feature map according to the ninth feature map and the boundary anomaly analysis layer includes:

[0181] Calculating the average pixel value of the detection window according to the pixel values of each pixel coordinate in the ninth feature map within the detection window;

[0182] Determining the difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located;

[0183] Calculating the standard deviation of the pixel values of the ninth feature map;

[0184] Calculating the normalized pixel value of the pixel coordinate within the detection window according to the standard deviation of the pixel values and the difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located;

[0185] Determining the tenth feature map according to the normalized pixel values of the pixel coordinates within the detection window.

[0186] The embodiments of the present invention focus on the pixel value differences within the local area of the detection window. On both sides of the tampering boundary, there are pixel value differences, and these differences can be determined by calculating the normalized pixel values of the pixel coordinates within the detection window. In actual application scenarios, the size of the detection window can be adjusted as needed.

[0187] In one embodiment of the present invention, inputting the tenth feature map into the second result output layer to obtain the boundary detection result includes:

[0188] Inputting the tenth feature map into the sixth convolutional layer to obtain the thirteenth feature map;

[0189] Upsampling the thirteenth feature map to obtain the fourteenth feature map;

[0190] Inputting the fourteenth feature map into an activation function to obtain the region detection result;

[0191] Among them, the sixth convolutional layer is a 1×1 convolutional layer.

[0192] The second result output layer is similar to the first result output layer. In the embodiments of the present invention, the image is magnified by upsampling, and the pixel values are mapped to the range between 0 and 1 through an activation function. The mapping results of the pixel coordinates of the fourteenth feature map constitute the region detection result. The activation function used can be the sigmoid function, the softmax function, etc.

[0193] In one embodiment of the present invention, obtaining training samples includes:

[0194] Obtaining training images and region labels;

[0195] Performing a dilation operation on the region label to obtain a dilated image;

[0196] Performing an erosion operation on the region label to obtain an eroded image;

[0197] Determining a boundary label according to the dilated image and the eroded image.

[0198] In the embodiments of the present invention, since there is no boundary label in CASIA 2.0, therefore, the embodiments of the present invention generate boundary labels based on region labels. The difference between the dilated image and the eroded image is the boundary label. The dilation operation and the erosion operation can be implemented using a 7x7 window. Through the embodiments of the present invention, boundary labels can be obtained more conveniently, improving the model training efficiency.

[0199] In one embodiment of the present invention, the method further includes: obtaining pre-training samples; pre-training the detection model based on the pre-training samples;

[0200] Inputting the training image into the detection model to obtain a region detection result and a boundary detection result, including:

[0201] Inputting the training image into the pre-trained detection model to obtain a region detection result and a boundary detection result.

[0202] In the embodiments of the present invention, the pre-training samples can be constructed from the original images in the COCO dataset. For example, select an image as the original image in the COCO dataset, and then cut out an object from another image and paste it onto the original image after operations such as rotation and magnification. The embodiments of the present invention improve the training effect of the detection model through pre-training, thereby improving the accuracy of tampered image detection.

[0203] As Figure 2 shown, the embodiments of the present invention provide an image detection method, including:

[0204] Step 201: Obtain pre-training samples.

[0205] Select the original images from the COCO dataset, crop the object images from another image, and paste the object images into the original images after rotation and other operations to obtain the pre-training samples.

[0206] Step 202: Pre-train the detection model based on the pre-training samples.

[0207] The architecture of the detection model is as Figure 3 shown, and the following embodiments will elaborate on its architecture in detail.

[0208] Step 203: Obtain the training images and region labels.

[0209] Obtain the training images and region labels from CASIA 2.0.

[0210] Step 204: Perform dilation operation on the region labels to obtain the dilated image.

[0211] Step 205: Perform erosion operation on the region labels to obtain the eroded image.

[0212] The window size used for the dilation operation and the erosion operation is 7×7.

[0213] As Figure 4 shown, from left to right are the region labels, the dilated image, the eroded image, and the boundary labels.

[0214] Step 206: Determine the boundary labels according to the dilated image and the eroded image.

[0215] The training images, region labels, and boundary labels constitute the training samples.

[0216] Step 207: Input the training images into the feature extraction layer to extract the high-order feature map and the low-order feature map from the training images.

[0217] Specifically, input the training images into the backbone network. The structure of the backbone network is as Figure 5 shown. It can be seen from the figure that the backbone network includes an input layer, an intermediate layer, and an output layer. The input layer includes five multi-channel convolutional layers and nine depthwise separable convolutional layers. Taking "Conv 32, 3x3, stride2" as an example, Conv 32 indicates that the output channels of the multi-channel convolutional layer are 32, the convolutional kernel is 3x3, and the stride is 2. The intermediate layer includes 16 identical depthwise separable convolutional layers. The output layer includes one multi-channel convolutional layer and six depthwise separable convolutional layers. The low-order feature map is output by the third depthwise separable convolutional layer in the intermediate layer, and the first feature map is output by the output layer.

[0218] Reference Figure 3, the first feature map is successively input into three dilated convolutional layers with a convolution kernel of 3x3 and dilation rates of 6, 12, and 18, a 1x1 convolutional layer, and a pooling layer to obtain multiple second feature maps. The second feature maps are concatenated along the Z-axis and input into a 1x1 convolutional layer to obtain a high-order feature map that fuses features of different scales. In the embodiment of the present invention, the size of the low-order feature map is 1 / 4 of the training image, and the size of the high-order feature map is 1 / 16 of the training image. The multi-scale network can also be Figure 6 the structure shown, which includes a 1x1 convolutional layer and dilated convolutions with dilation rates of 1, 2, and 5. In Figure 6 , the convolutional layers in each horizontal row share the convolution kernel parameters so that the same object has the same feature expression ability at different scales.

[0219] Step 208: Input the high-order feature map and the low-order feature map into the region detection layer to obtain a region detection result.

[0220] Specifically, input the low-order feature map into a 1x1 convolutional layer to obtain a fifth feature map. Perform bilinear interpolation on the high-order feature map to magnify it by four times to obtain a sixth feature map. Concatenate the fifth feature map and the sixth feature map along the Z-axis, and then input them into a 3x3 convolutional layer to fuse the features to obtain a third feature map.

[0221] According to the pixel values of each pixel coordinate in the third feature map, calculate the average pixel value of the third feature map as shown in Equation (1).

[0222]

[0223] Among them, F[i, j] is used to represent the pixel value of the pixel coordinate (i, j) in the third feature map, H is used to represent the height of the third feature map, W is used to represent the width of the third feature map, and μ f is used to represent the average pixel value of the third feature map.

[0224] Determine the difference between the pixel value of each pixel coordinate and the average pixel value as shown in Equation (2).

[0225] D f [i, j] = F[i, j] - μ f (2)

[0226] Among them, D f [i, j] is used to represent the difference between the pixel value of the pixel coordinate (i, j) and the average pixel value.

[0227] According to the difference between the pixel value of each pixel coordinate and the average pixel value, calculate the standard deviation of the pixel values of the third feature map.

[0228] Calculate the normalized pixel values of each pixel coordinate according to the standard deviation of pixel values and the differences between the pixel values of each pixel coordinate and the average pixel value, as shown in Equation (3).

[0229] Z f [i,j] = D f [i,j] / max(σ f , ε + ω σ1 ) (3)

[0230] Among them, σ f is used to represent the standard deviation of pixel values of the third feature map, ε is 10 -5 , ω σ1 is the first vector that can be continuously adjusted through the training process.

[0231] Determine the fourth feature map according to the normalized pixel values of each pixel coordinate.

[0232] Input the fourth feature map into a 1×1 convolutional layer to obtain the seventh feature map. Enlarge the seventh feature map by four times through bilinear interpolation to obtain the eighth feature map. Input the eighth feature map into the sigmoid function to obtain the region detection result.

[0233] Step 209: Input the high-order feature map and the low-order feature map into the boundary detection layer to obtain the boundary detection result.

[0234] Specifically, input the low-order feature map into a 1x1 convolutional layer to obtain the eleventh feature map. Perform bilinear interpolation on the high-order feature map to enlarge it by four times to obtain the twelfth feature map. Concatenate the eleventh feature map and the twelfth feature map along the Z-axis, and then input them into a 3x3 convolutional layer to fuse the features to obtain the ninth feature map.

[0235] Calculate the average pixel value of the detection window according to the pixel values of each pixel coordinate in the ninth feature map within the detection window, as shown in Equation (4).

[0236]

[0237] Among them, is used to represent the average pixel value of the detection window with a height of 7 and a width of 7.

[0238] Determine the difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located, as shown in Equation (5).

[0239]

[0240] Among them, is used to represent the difference between the pixel value of the pixel coordinate (i, j) and the average pixel value of the detection window where the pixel coordinate is located.

[0241] Calculate the standard deviation of the pixel values of the ninth feature map.

[0242] Calculate the normalized pixel values of the pixel coordinates within the detection window based on the standard deviation of the pixel values and the difference between the pixel values of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located, as shown in Equation (6).

[0243]

[0244] In the embodiment of the present invention, the ninth feature map is the same as the third feature map, and the standard deviations of their pixel values are the same. ω σ2 is the second vector that can be continuously adjusted through the training process.

[0245] Determine the tenth feature map based on the normalized pixel values of the pixel coordinates within the detection window.

[0246] Input the tenth feature map into a 1×1 convolutional layer to obtain the thirteenth feature map. Enlarge the thirteenth feature map by four times through bilinear interpolation to obtain the fourteenth feature map. Input the fourteenth feature map into the sigmoid function to obtain the boundary detection result.

[0247] Step 210: Train the detection model according to the region label, boundary label, region detection result, and boundary detection result.

[0248] Based on the difference between the region label and the region detection result, the difference between the predicted tampered region and the actual tampered region can be determined; based on the difference between the boundary label and the boundary detection result, the difference between the predicted tampered boundary and the actual tampered boundary can be determined. The embodiment of the present invention uses a cross-entropy loss function, including two parts: region detection and boundary detection, as shown in Equations (7)-(9).

[0249]

[0250]

[0251]

[0252] Among them, m is used to represent the number of training samples, is used to represent the region detection result of training sample k, is used to represent the boundary detection result of training sample k, is used to represent the value of the region label corresponding to the pixel coordinate (i, j) in training sample k, is used to represent the region detection result of the pixel coordinate (i, j) in training sample k, is used to represent the value of the boundary label corresponding to the pixel coordinate (i, j) in training sample k, is used to represent the boundary detection result of the pixel coordinate (i, j) in training sample k.

[0253] The loss value can be calculated through formulas (7)-(9), and the parameters of the detection model are adjusted according to the loss value.

[0254] Step 211: Based on the trained detection model, determine whether the detected image is tampered with.

[0255] The detection model can map the input pixel values between 0 and 1. If the value output by the Sigmoid function is greater than the set value (0.5 in the embodiment of the present invention), it is determined that the pixel coordinates are in the tampered area; otherwise, they are in the background area.

[0256] In the embodiment of the present invention, Columbia and CASIA 1.0 are used as the test sample sets, and the performance of the trained detection model is evaluated through the F1 score. The test results are shown in Table 1. It can be seen from Table 1 that compared with other models, the detection model trained in the embodiment of the present invention has the highest F1 score, indicating that its performance is better than other models. Among them, RGB-N is a method for detecting tampered images based on two-stream Faster R-CNN, NOI1 is a method for detecting tampered images based on noise inconsistency, which uses high-pass wavelet coefficients to simulate local noise, CFA is a CFA mode estimation method, which uses nearby pixels to approximate the camera filter array mode and then generates the tampering probability of each pixel. DCT is a method for detecting JPEG image tampering based on the difference in DCT coefficient histograms.

[0257] Table 1 F1 scores of different models

[0258] Columbia CASIA 1.0 Detection model 0.747 0.435 RGB-N 0.697 0.408 NOI1 0.574 0.263 DCT 0.520 0.301 CFA 0.503 0.212

[0259] As Figure 7 shown, the embodiment of the present invention provides an image detection device, including:

[0260] An acquisition module 701, configured to acquire training samples; wherein, the training samples include: training images, region labels, and boundary labels;

[0261] A training module 702, configured to input the training images into the detection model to obtain region detection results and boundary detection results; and train the detection model according to the region labels, boundary labels, region detection results, and boundary detection results;

[0262] A detection module 703, configured to determine whether the detected image is tampered with based on the trained detection model.

[0263] In an embodiment of the present invention, the detection model includes: a feature extraction layer, a region detection layer, and a boundary detection layer;

[0264] The training module 702 is configured to input a training image into a feature extraction layer to extract a high-order feature map and a low-order feature map from the training image; input the high-order feature map and the low-order feature map into a region detection layer to obtain a region detection result; and input the high-order feature map and the low-order feature map into a boundary detection layer to obtain a boundary detection result.

[0265] In an embodiment of the present invention, the training module 702 is configured to input a training image into a backbone network to obtain a low-order feature map and a first feature map; extract multi-scale features from the first feature map based on a multi-scale network to obtain a plurality of second feature maps; splice the plurality of second feature maps and input them into a first convolutional layer to obtain a high-order feature map; wherein, the backbone network includes: a first multi-channel convolutional layer and a depthwise separable convolutional layer; and the first convolutional layer is a 1×1 convolutional layer.

[0266] In an embodiment of the present invention, the multi-scale network includes: an atrous convolutional layer, a second convolutional layer, and a pooling layer; wherein, the second convolutional layer is a 1×1 convolutional layer.

[0267] In an embodiment of the present invention, the region detection layer includes: a first feature fusion layer, a region anomaly analysis layer, and a first result output layer; the training module 702 is configured to input the high-order feature map and the low-order feature map into the first feature fusion layer to obtain a third feature map; determine a fourth feature map according to the third feature map and the region anomaly analysis layer; wherein, the fourth feature map is used to characterize the pixel value difference between the tampered region and the background region in the third feature map; and input the fourth feature map into the first result output layer to obtain a region detection result.

[0268] In an embodiment of the present invention, the training module 702 is configured to input the low-order feature map into a third convolutional layer to obtain a fifth feature map; perform upsampling on the high-order feature map to obtain a sixth feature map; splice the fifth feature map and the sixth feature map and input them into a second multi-channel convolutional layer to obtain a third feature map; wherein, the third convolutional layer is a 1×1 convolutional layer.

[0269] In an embodiment of the present invention, the training module 702 is configured to calculate the average pixel value of the third feature map according to the pixel values of each pixel coordinate in the third feature map; determine the difference between the pixel value of each pixel coordinate and the average pixel value; calculate the standard deviation of the pixel values of the third feature map according to the difference between the pixel value of each pixel coordinate and the average pixel value; calculate the normalized pixel value of each pixel coordinate according to the standard deviation of the pixel values and the difference between the pixel value of each pixel coordinate and the average pixel value; and determine the fourth feature map according to the normalized pixel values of each pixel coordinate.

[0270] In one embodiment of the present invention, the training module 702 is configured to input the fourth feature map into the fourth convolutional layer to obtain the seventh feature map; perform upsampling on the seventh feature map to obtain the eighth feature map; input the eighth feature map into the activation function to obtain the region detection result; wherein, the fourth convolutional layer is a 1×1 convolutional layer.

[0271] In one embodiment of the present invention, the boundary detection layer includes: a second feature fusion layer, a boundary anomaly analysis layer, and a second result output layer; the training module 702 is configured to input the high-order feature map and the low-order feature map into the second feature fusion layer to obtain the ninth feature map; determine the tenth feature map according to the ninth feature map and the boundary anomaly analysis layer; wherein, the tenth feature map is used to represent the pixel value difference between the tampered region and the background region within the detection window; input the tenth feature map into the second result output layer to obtain the boundary detection result.

[0272] In one embodiment of the present invention, the training module 702 is configured to input the low-order feature map into the fifth convolutional layer to obtain the eleventh feature map; perform upsampling on the high-order feature map to obtain the twelfth feature map; splice the eleventh feature map and the twelfth feature map and then input them into the third multi-channel convolutional layer to obtain the ninth feature map; wherein, the fifth convolutional layer is a 1×1 convolutional layer.

[0273] In one embodiment of the present invention, the training module 702 is configured to calculate the average pixel value of the detection window according to the pixel values of each pixel coordinate in the ninth feature map within the detection window; determine the difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located; calculate the standard deviation of the pixel values of the ninth feature map; calculate the normalized pixel value of the pixel coordinate within the detection window according to the standard deviation of the pixel values and the difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located; determine the tenth feature map according to the normalized pixel value of the pixel coordinate within the detection window.

[0274] In one embodiment of the present invention, the training module 702 is configured to input the tenth feature map into the sixth convolutional layer to obtain the thirteenth feature map; perform upsampling on the thirteenth feature map to obtain the fourteenth feature map; input the fourteenth feature map into the activation function to obtain the region detection result; wherein, the sixth convolutional layer is a 1×1 convolutional layer.

[0275] In one embodiment of the present invention, the acquisition module 701 is configured to acquire a training image and a region label; perform a dilation operation on the region label to obtain a dilated image; perform an erosion operation on the region label to obtain an eroded image; determine a boundary label according to the dilated image and the eroded image.

[0276] In one embodiment of the present invention, an acquisition module 701 is configured to acquire pre-training samples; pre-train a detection model based on the pre-training samples; and input a training image into the pre-trained detection model to obtain a region detection result and a boundary detection result.

[0277] An embodiment of the present invention provides an electronic device, including:

[0278] One or more processors;

[0279] A storage device for storing one or more programs,

[0280] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any of the above embodiments.

[0281] An embodiment of the present invention provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method as described in any of the above embodiments is implemented.

[0282] Figure 8 An exemplary system architecture 800 to which the image detection method or image detection device according to the embodiments of the present invention can be applied is shown.

[0283] As Figure 8 shown, the system architecture 800 may include terminal devices 801, 802, 803, a network 804, and a server 805. The network 804 is used to provide a medium for a communication link between the terminal devices 801, 802, 803 and the server 805. The network 804 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0284] Users may use the terminal devices 801, 802, 803 to interact with the server 805 through the network 804 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 801, 802, 803, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0285] The terminal devices 801, 802, 803 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0286] The server 805 may be a server that provides various services. For example, it may be a background management server (merely an example) that supports shopping websites browsed by users using the terminal devices 801, 802, and 803. The background management server may analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - merely examples) to the terminal devices.

[0287] It should be noted that the image detection method provided by the embodiments of the present invention is generally executed by the server 805. Correspondingly, the image detection device is generally disposed in the server 805.

[0288] It should be understood that Figure 8 the numbers of the terminal devices, networks, and servers in

[0289] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 9 Figure 9

[0290] Figure 9 As

[0291] shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage section 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0292] The following components are connected to the I / O interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as required. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as required, so that the computer program read from it can be installed into the storage portion 908 as required.

[0292] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above functions defined in the system of the present invention are executed.

[0293] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0294] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0295] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a sending module, an obtaining module, a determining module, and a first processing module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the sending module can also be described as "a module that sends a picture acquisition request to the connected server".

[0296] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes:

[0297] Obtain training samples; wherein, the training samples include: training images, region labels, and boundary labels;

[0298] Input the training images into a detection model to obtain a region detection result and a boundary detection result;

[0299] Train the detection model according to the region labels, the boundary labels, the region detection result, and the boundary detection result;

[0300] Based on the trained detection model, determine whether a detection image is tampered with.

[0301] According to the technical solution of the embodiment of the present invention, boundary detection and region detection are performed on an image based on a detection model. The region detection identifies a tampered region based on the feature difference between the tampered region and the background region of the entire image, and it focuses on the overall features of the image. The boundary detection identifies a tampered boundary based on the feature difference between both sides of the tampered boundary. The boundary detection can assist the region detection to more accurately determine the tampered region and improve the accuracy of image detection.

[0302] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image detection method, characterized in that Including: Obtaining training samples; wherein, the training samples include: training images, region labels, and boundary labels; Inputting the training images into a detection model to obtain region detection results and boundary detection results, including: inputting the training images into a feature extraction layer to extract high-order feature maps and low-order feature maps from the training images; inputting the high-order feature maps and the low-order feature maps into a region detection layer to obtain the region detection results; inputting the high-order feature maps and the low-order feature maps into a boundary detection layer to obtain the boundary detection results; the detection model includes: a feature extraction layer, a region detection layer, and a boundary detection layer; Training the detection model according to the region labels, the boundary labels, the region detection results, and the boundary detection results; Based on the trained detection model, determining whether a detection image is tampered with; Wherein, the obtaining of the training samples includes: obtaining the training images and the region labels; performing a dilation operation on the region labels to obtain a dilated image; performing an erosion operation on the region labels to obtain an eroded image; determining the boundary labels according to the difference between the dilated image and the eroded image; The inputting of the training images into the feature extraction layer to extract high-order feature maps and low-order feature maps from the training images includes: inputting the training images into a backbone network to obtain the low-order feature maps and a first feature map; extracting multi-scale features from the first feature map based on a multi-scale network to obtain a plurality of second feature maps; splicing the plurality of second feature maps and inputting them into a first convolutional layer to obtain the high-order feature maps; wherein, the backbone network includes: a first multi-channel convolutional layer and a depthwise separable convolutional layer; the first convolutional layer is a 1×1 convolutional layer; the multi-scale network includes: an atrous convolutional layer, a second convolutional layer, and a pooling layer; wherein, the second convolutional layer is a 1×1 convolutional layer; The region detection layer includes: a first feature fusion layer, a region anomaly analysis layer, and a first result output layer; The inputting of the high-order feature maps and the low-order feature maps into the region detection layer to obtain the region detection results includes: inputting the high-order feature maps and the low-order feature maps into the first feature fusion layer to obtain a third feature map; determining a fourth feature map according to the third feature map and the region anomaly analysis layer; wherein, the fourth feature map is used to characterize the pixel value difference between the tampered region and the background region in the third feature map; inputting the fourth feature map into the first result output layer to obtain the region detection results.

2. The method according to claim 1, wherein The inputting of the high-order feature maps and the low-order feature maps into the first feature fusion layer to obtain a third feature map includes: Inputting the low-order feature maps into a third convolutional layer to obtain a fifth feature map; Performing upsampling on the high-order feature maps to obtain a sixth feature map; Splicing the fifth feature map and the sixth feature map and inputting them into a second multi-channel convolutional layer to obtain the third feature map; Wherein, the third convolutional layer is a 1×1 convolutional layer.

3. The method according to claim 1, wherein determining a fourth feature map based on the third feature map and the region anomaly analysis layer includes: calculating an average pixel value of the third feature map according to pixel values of each pixel coordinate in the third feature map; determining a difference between the pixel value of each pixel coordinate and the average pixel value; calculating a standard deviation of pixel values of the third feature map according to the difference between the pixel value of each pixel coordinate and the average pixel value; calculating a normalized pixel value of each pixel coordinate according to the standard deviation of pixel values and the difference between the pixel value of each pixel coordinate and the average pixel value; determining the fourth feature map according to the normalized pixel value of each pixel coordinate.

4. The method according to any one of claims 1-3, wherein inputting the fourth feature map into the first result output layer to obtain the region detection result includes: inputting the fourth feature map into a fourth convolutional layer to obtain a seventh feature map; performing upsampling on the seventh feature map to obtain an eighth feature map; inputting the eighth feature map into an activation function to obtain the region detection result; wherein the fourth convolutional layer is a 1×1 convolutional layer.

5. The method according to claim 1, wherein the boundary detection layer includes: a second feature fusion layer, a boundary anomaly analysis layer, and a second result output layer; inputting the high-order feature map and the low-order feature map into the boundary detection layer to obtain the boundary detection result includes: inputting the high-order feature map and the low-order feature map into the second feature fusion layer to obtain a ninth feature map; determining a tenth feature map according to the ninth feature map and the boundary anomaly analysis layer; wherein the tenth feature map is used to characterize the pixel value difference between the tampered region and the background region within the detection window; inputting the tenth feature map into the second result output layer to obtain the boundary detection result.

6. The method according to claim 5, wherein inputting the high-order feature map and the low-order feature map into the second feature fusion layer to obtain a ninth feature map includes: inputting the low-order feature map into a fifth convolutional layer to obtain an eleventh feature map; performing upsampling on the high-order feature map to obtain a twelfth feature map; concatenating the eleventh feature map and the twelfth feature map and inputting the concatenated result into a third multi-channel convolutional layer to obtain the ninth feature map; wherein the fifth convolutional layer is a 1×1 convolutional layer.

7. The method according to claim 5, wherein determining a tenth feature map according to the ninth feature map and the boundary anomaly analysis layer includes: calculating an average pixel value of the detection window according to pixel values of each pixel coordinate in the ninth feature map within the detection window; determining a difference between the pixel value of each pixel coordinate and the average pixel value of the detection window where the pixel coordinate is located; calculating a standard deviation of pixel values of the ninth feature map; Calculate the normalized pixel values of the pixel coordinates within the detection window according to the standard deviation of the pixel values and the differences between the pixel values of each of the pixel coordinates and the average pixel value of the detection window where the pixel coordinates are located. Determine the tenth feature map according to the normalized pixel values of the pixel coordinates within the detection window.

8. The method according to any one of claims 5-7, wherein The inputting the tenth feature map into the second result output layer to obtain the boundary detection result includes: Inputting the tenth feature map into a sixth convolutional layer to obtain a thirteenth feature map; Performing upsampling on the thirteenth feature map to obtain a fourteenth feature map; Inputting the fourteenth feature map into an activation function to obtain the region detection result; wherein the sixth convolutional layer is a 1×1 convolutional layer.

9. The method according to claim 1, wherein Further comprising: Obtaining pre-training samples; Pre-training the detection model based on the pre-training samples; The inputting the training image into the detection model to obtain the region detection result and the boundary detection result includes: Inputting the training image into the pre-trained detection model to obtain the region detection result and the boundary detection result.

10. An image detection device, characterized in that, Comprising: An acquisition module configured to acquire training samples; wherein the training samples include: training images, region labels, and boundary labels; A training module configured to input the training image into a detection model to obtain a region detection result and a boundary detection result, including: inputting the training image into a feature extraction layer to extract a high-order feature map and a low-order feature map from the training image; inputting the high-order feature map and the low-order feature map into a region detection layer to obtain the region detection result; inputting the high-order feature map and the low-order feature map into a boundary detection layer to obtain the boundary detection result; the detection model includes: a feature extraction layer, a region detection layer, and a boundary detection layer; training the detection model according to the region label, the boundary label, the region detection result, and the boundary detection result; wherein the acquisition of the training samples includes: acquiring the training image and the region label; performing a dilation operation on the region label to obtain a dilated image; performing an erosion operation on the region label to obtain an eroded image; determining the boundary label according to the difference between the dilated image and the eroded image; The inputting the training image into the feature extraction layer to extract a high-order feature map and a low-order feature map from the training image includes: inputting the training image into a backbone network to obtain the low-order feature map and a first feature map; extracting multi-scale features from the first feature map based on a multi-scale network to obtain a plurality of second feature maps; splicing the plurality of second feature maps and inputting them into a first convolutional layer to obtain the high-order feature map; wherein the backbone network includes: a first multi-channel convolutional layer and a depthwise separable convolutional layer; the first convolutional layer is a 1×1 convolutional layer; the multi-scale network includes: an atrous convolutional layer, a second convolutional layer, and a pooling layer; wherein the second convolutional layer is a 1×1 convolutional layer; The region detection layer includes: a first feature fusion layer, a region anomaly analysis layer, and a first result output layer; the step of inputting the high-order feature map and the low-order feature map into the region detection layer to obtain the region detection result includes: inputting the high-order feature map and the low-order feature map into the first feature fusion layer to obtain a third feature map; determining a fourth feature map according to the third feature map and the region anomaly analysis layer; wherein, the fourth feature map is used to characterize the pixel value difference between the tampered region and the background region in the third feature map; inputting the fourth feature map into the first result output layer to obtain the region detection result; The detection module is configured to determine whether a detection image is tampered with based on the trained detection model.

11. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.

12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Image detection method and device, computer equipment and storage medium

    CN111738244A