A method, device, and storage medium for locating an image inpainting area
By combining the multi-scale cross-fusion technology of the noise domain and local binary mode features with fast Fourier convolution, efficient and accurate positioning of the image repair area is achieved, and the problem of unclear positioning in the prior art is solved.
Patent Information
- Application Number
- CN202210268461.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-18
AI Technical Summary
The image repair method based on deep learning in the prior art is not effective in positioning sample block image repair, while the traditional method requires a lot of time to calculate and may require manual operation, resulting in unclear positioning of the image repair area.
By extracting the noise characteristics and local binary mode features of the image from the noise domain, combining fast Fourier convolution and multi-scale cross-fusion technology, efficient and accurate positioning of the image repair area is achieved.
It can capture different repair traces of sample block image repair and deep learning image repair at the same time, and efficiently and accurately locate the repaired areas in the image, solving the problem of unclear area edges.
Smart Images

Figure CN114693913B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image restoration, and particularly relates to a method, device and storage medium for locating an image restoration area. Background Art
[0002] Image restoration refers to the process of reconstructing the lost or damaged parts in an image. For object removal, generally, sample block-based and deep learning-based image restoration methods can be used, and they both have good restoration effects.
[0003] The deep learning-based image restoration method has poor effect in locating sample block image restoration, while the traditional method for locating sample block restoration requires a large amount of time for calculation and even manual operation. Summary of the Invention
[0004] The present invention aims to at least solve the technical problems existing in the prior art. For this purpose, the present invention provides a method, device and storage medium for locating an image restoration area, which can efficiently and accurately locate the restored area in an image and can also solve the problem that the edge of the currently located image restoration area is not clear.
[0005] In a first aspect of the present invention, a method for locating an image restoration area is provided, including the steps of:
[0006] Obtaining an original image to be located;
[0007] Converting the original image into a noise domain to obtain a first image, and extracting a first feature map from the first image; wherein, the first feature map includes the inconsistency features between the source area and the noise restoration area in the first image;
[0008] Extracting local binary pattern features from the original image, merging the local binary pattern features with the original image to obtain a second image, and extracting a second feature map from the second image; wherein, the second feature map includes the inconsistency features between the source area and the image restoration area in the second image;
[0009] Fusing the first feature map and the second feature map to obtain a fused feature map;
[0010] Locating the restoration area in the fused feature map.
[0011] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects:
[0012] This method first extracts the noise features of the image from the noise domain and uses local binary patterns to extract image features to enhance the repair traces. Since both the noise features and the image features are the differential features between the repaired area and the source area, fusing the noise features and the image features can obtain richer features. Finally, based on the fused features, positioning is performed to achieve the positioning of image repair. This method can simultaneously capture the different repair traces of sample block image repair and deep learning image repair, can efficiently and accurately locate the repaired area in the image, and can also solve the problem of unclear regional edges in current image repair positioning.
[0013] According to some embodiments of the present invention, the converting the original image into a noise domain to obtain a first image and extracting a first feature map from the first image includes:
[0014] Extracting local noise descriptors in the original image according to a model filter rich in steganalysis to obtain a first image;
[0015] Extracting a first feature map from the first image according to a plurality of residual blocks, wherein the plurality of residual blocks are connected in sequence, and the first image is input into the first residual block to obtain the corresponding first feature map output by each residual block.
[0016] According to some embodiments of the present invention, the plurality of residual blocks are three residual blocks, each residual block includes two bottleneck units, each bottleneck unit includes three consecutive convolutional layers and an identity shortcut connection, normalization and ReLu activation are performed before each convolutional operation, and the sum sizes of the three convolutional layers are 1×1, 3×3, and 1×1 respectively; the convolutional layer strides of the first bottleneck unit are all 1, and the last convolutional stride of the second bottleneck unit is 2.
[0017] According to some embodiments of the present invention, the extracting a second feature map from the second image includes:
[0018] After performing channel expansion on the second image, performing continuous downsampling multiple times;
[0019] Extracting local and global features from the downsampled second image according to fast Fourier convolution;
[0020] Performing continuous upsampling multiple times on the extracted local and global features to obtain the corresponding second feature map output by each upsampling.
[0021] According to some embodiments of the present invention, local and global features are extracted from the downsampled second image through a plurality of fast Fourier convolution blocks with the same configuration, wherein each fast Fourier convolution block extracts features in the following manner:
[0022] Perform two ordinary convolutions on the downsampled second image to extract the first local feature and the second local feature of the second image; perform one ordinary convolution and one frequency-domain convolution on the downsampled second image to obtain the third local feature and the global feature;
[0023] Add the first local feature and the third local feature to obtain the local feature of the second image; add the second local feature and the global feature to obtain the global feature of the second image.
[0024] According to some embodiments of the present invention, perform three consecutive upsamplings on the extracted local and global features.
[0025] According to some embodiments of the present invention, the fusing the first feature map and the second feature map to obtain a fused feature map includes:
[0026] Perform the first multi-scale cross-fusion on the first feature map: perform upsampling, downsampling, and no change on each of the first feature maps respectively to obtain intermediate feature maps; connect the feature maps with the same size in the intermediate feature maps to obtain a noise cross-fusion feature map;
[0027] Perform the second multi-scale cross-fusion on the second feature map: perform upsampling, downsampling, and no change on each of the second feature maps respectively to obtain intermediate feature maps; connect the feature maps with the same size in the intermediate feature maps to obtain an image cross-fusion feature map;
[0028] Connect the noise cross-fusion feature map and the image cross-fusion feature map, and perform the third multi-scale cross-fusion after connection to obtain a fused feature map.
[0029] According to some embodiments of the present invention, the positioning of the repaired area in the fused feature map includes:
[0030] Perform one convolution on the fused feature map;
[0031] Perform per-pixel prediction and positioning on the repaired area in the convolved fused feature map according to logistic regression.
[0032] In a second aspect of the present invention, there is provided a positioning device for an image repair area, including at least one control processor and a memory for communicatingly connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the positioning method of the image repair area as described above.
[0033] In a third aspect of the present invention, there is provided a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the method for locating an image repair area as described above.
[0034] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0036] Figure 1 are schematic diagrams of an original image, an object-removed repair map, and a noise domain image provided by the first embodiment of the present invention;
[0037] Figure 2 is Figure 1 an enlarged schematic diagram of (c);
[0038] Figure 3 are comparative schematic diagrams of a detection area provided by the first embodiment of the present invention;
[0039] Figure 4 is a flowchart of a method for locating an image repair area provided by the first embodiment of the present invention;
[0040] Figure 5 is Figure 4 a specific flowchart of step S102 in;
[0041] Figure 6 is Figure 4 a specific flowchart of step S103 in;
[0042] Figure 7 is a structural diagram of a fast Fourier convolution block provided by the second embodiment of the present invention;
[0043] Figure 8 is a structural diagram of a spectrum converter provided by the second embodiment of the present invention;
[0044] Figure 9 is Figure 4 a specific flowchart of step S104 in;
[0045] Figure 10 is Figure 4 a specific flowchart of step S105 in;
[0046] Figure 11 is a structural block diagram of a method for locating an image repair area provided by the third embodiment of the present invention;
[0047] Figure 12 It is a schematic flowchart of a method for locating an image repair area provided by the third embodiment of the present invention;
[0048] Figure 13 It is a schematic flowchart of feature extraction for the first feature extraction module provided by the third embodiment of the present invention to extract the inconsistency features between the noise repair area and the source area;
[0049] Figure 14 It is a schematic flowchart of feature extraction for the second feature extraction module provided by the third embodiment of the present invention to extract the inconsistency features between the image repair area and the source area. Detailed implementation manners
[0050] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0051] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "axial", "radial", "circumferential", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the features defined as "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0052] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0053] Before introducing the embodiments of the present invention, a brief introduction to the existing solutions of the present invention will be given first:
[0054] Image inpainting refers to the process of reconstructing the lost or damaged parts in an image. For object removal, generally, sample block-based and deep learning-based image inpainting methods can be used, and both of them can have good inpainting effects.
[0055] First, image inpainting;
[0056] 1. Sample block-based image inpainting;
[0057] Taking A. Criminisi, P. Perez and K. Toyama, “Region filling and object removal by exemplar-based image inpainting,” IEEE Transaction on Image Process, vol. 13, no. 9, pp. 1200–1212, Sep. 2004 as an example:
[0058] The missing region in the image is the target region, and the boundary of the target region is called the filling front. The non-missing region is the source region (the whole image minus the target region). Each pixel in the image has a color value (the unfilled pixels are empty) and a confidence value. The steps of sample block-based image inpainting are repeated in three steps: a. Calculate the priority of the block centered on the pixel on the filling front to determine which block is filled first; b. Propagate the texture and structure information, find the block in the source region that is most similar to the block to be filled first, and copy the pixel values; c. Update the confidence in the filled block.
[0059] First, define the block size, which is defaulted to 9*9.
[0060] a. Calculate the priority;
[0061] The priority P(p) of the block Ψp centered at point p:
[0062] Confidence C(p): Initialize the confidence in the source region to 1 and the target region to 0. C(p) is the sum of all confidences in the block divided by the block size. Preferentially fill the blocks with more known points around them.
[0063] Data D(p): Multiply the delivery direction perpendicular to this point by the normal direction and then normalize. The points with the gradient direction perpendicular to the contour and relatively large gradients have higher priorities;
[0064] b. Find the block in the source region with the smallest sum of squared differences from the filled pixels of the block with the highest priority, and fill it from the source region to the target region;
[0065] c. Update the confidence value of each pixel in the filled region to the confidence of the block calculated in (1).
[0066] 2. Deep learning image inpainting;
[0067] Refer to Figure 1 , taking Liu H, Jiang B, Song Y, et al. Rethinking Image Inpainting via a Mutual Encoder-Decoder with Feature Equalizations[J]. 2020. as an example:
[0068] Use a mutual encoder-decoder to jointly generate an end-to-end image to fill in the missing regions of the image. The deep encoder represents the structural features of the input image respectively, and the shallow encoder represents the texture features of the input image. Connect the features of the two branches for channel and spatial equalization, supplement the equalized features into the features in the decoder, and generate the output image through skip connections.
[0069] Second, sample block image inpainting forensics (localization);
[0070] 1. Traditional methods:
[0071] (1) Taking Wu Q, Sun S J, Zhu W, et al. Detection of digital doctoring in exemplar-based inpainted images[C] / / 2008 International Conference on Machine Learning and Cybernetics. IEEE, 2008, 3: 1222-1226. as an example:
[0072] First, manually extract the suspicious region and repeat the following three steps:
[0073] a. Calculate the block matching degree. First, subtract the two blocks and take the absolute value to obtain the difference array; mark the elements in the difference array equal to 0 as 0, otherwise as 1; take the maximum zero connection length (in the 8-neighborhood) as the matching degree; b. Calculate the fuzzy membership degree: Select the ascending trapezoidal membership function to map the set of similar tampered regions to the interval [0, 1]; c. Divide the cut set: The forged region is likely to be connected rather than a set of very small blocks or single pixels. Use morphological operations to remove the isolated blocks that may be introduced by noise or compression.
[0074] (2) Take Bacchuwar K S, Ramakrishnan K R. A jump patch-block match algorithm for multiple forgery detection [C] / / 2013 International Mutli-Conference on Automation, Computing, Communication, Control and Compressed Sensing (iMac4s). IEEE, 2013:723-728. as an example:
[0075] First, manually select the suspicious area, convert the suspicious area to the Y-Cb-Cr format, and only use the Y component. The steps in (1) are only carried out when the median difference of the block pair is less than the median threshold. This method can skip a large number of blocks, thus reducing the algorithm time.
[0076] (3) Take Chang I C, Yu J C, Chang C C. A forgery detection algorithm for exemplar-based inpainting images using multi-region relation [J]. Image and Vision Computing, 2013, 31(1):57-71. as an example:
[0077] a. It is not necessary to manually extract the suspicious area first. First, perform suspicious area detection: calculate the difference array of the block pair, use the maximum zero connection length as the block similarity, and normalize it to the suspicious degree; b. Vector filtering: The vector pointing from the target block to the corresponding similar block is the similarity vector. Since regions with consistent textures such as the sky and grass will lead to incorrect results, the similarity vector regions with short lengths and in the same direction are removed. The repaired region is the similarity vector with a long length and divergent directions. Then this similar block is the suspicious area; c. Divide the multi-region relationship in the suspicious area. The one-to-many concentrated relationship group is the repaired region.
[0078] In order to speed up the search speed, a two-stage search strategy is proposed: first, convert the color to the key value through weight transformation, first limit the search range through the similar key value, and then perform block similarity calculation.
[0079] (4) Take Liang Z, Yang G, Ding
[0080] a. Apply zero connectivity features to search for suspicious blocks. First, use center pixel mapping to collect blocks with equal center pixel values, and then perform block similarity calculation to speed up the search. b. Use the maximum connected component label instead of the previous overall block label to improve accuracy. c. Vector filtering: same as (3). d. Fragment splicing detection distinguishes the tampered area from the reference area.
[0081] 2. Deep learning methods;
[0082] (1) Take Zhu X, Qian Y, Zhao X, et al. A deep learning approach to patch-based image inpainting forensics[J]. Signal Processing: Image Communication, 2018, 67: 90-99. as an example:
[0083] It is proposed to use convolutional neural networks to detect sample block image restoration. The convolutional neural network is built according to the encoder-decoder network structure to predict the restoration probability of each pixel in the image. There is no need to consider feature extraction and classifier design, and no post-processing is required like traditional methods.
[0084] (2)Lu M, Niu SA Detection Approach Using LSTM-CNN for Object Removal Caused by Exemplar-Based Image Inpainting[J].Electronics, 2020,9(5):858.[1]Lu M,Niu SA Detection Approach Using LSTM-CNN for Object RemovalCaused by Exemplar-Based Image Inpainting[J].Electronics,2020,9(5):858. For example:
[0085] A method for forensic identification of sample block image inpainting based on Long Short-Term Memory (LSTM)-Convolutional Neural Network (CNN) is proposed.
[0086] First, a Convolutional Neural Network is used to identify suspicious similar blocks in the tampered image, and LSTM is utilized to distinguish texture-consistent regions and suspicious regions, thus greatly reducing the false positive rate.
[0087] Since the detection of the deep learning method (1) depends on the similarity of blocks, there are abnormally similar blocks in normal images such as blue sky, grassland and other uniform regions, which are easily confused during feature classification, resulting in difficulty in correctly identifying the classification of pixels. Therefore, a module composed of a convolutional layer and three stacked LSTMs is proposed to eliminate the misclassification caused by real texture-consistent regions. The convolutional layer extracts different low-level features from the image, and the LSTM network is used to capture the dependencies between sequential pixels, and this is used to detect the boundaries between different blocks in the image, accurately distinguishing normal regions and false report regions.
[0088] In summary, the image inpainting method based on deep learning has poor effect in locating sample block image inpainting, while the traditional method for locating sample block inpainting requires a large amount of time for calculation and even manual operation. There is currently a lack of a method that can detect both sample block image inpainting and deep learning image inpainting.
[0089] To solve the above technical deficiencies, the present invention enhances the inpainting traces by extracting the noise features of the image and extracting the image features from the noise domain; then fuses the noise features and the image features; finally, based on the fused features, the inpainting location is realized. By simultaneously capturing different inpainting traces of sample block image inpainting and deep learning image inpainting, the inpainted area in the image can be efficiently and accurately located, and the problem of unclear region edges in current image inpainting location can also be solved. The following is a detailed introduction of the embodiments:
[0090] The first embodiment;
[0091] Referring to Figures 1 to 3 , Figure 1 (a) is the original image, Figure 1 (b) is the inpainting image of the original image with an object removed, Figure 1 (c) is the image in the noise domain. Figure 2 It is an enlarged schematic diagram of the image in the noise domain. Figure 3 It is a schematic diagram of the detection result of the original image, where Figure 3 (a) is the calibration ground truth (the area that should actually be detected), Figure 3 (b) the detected area is the detection result of the noise flow, Figure 3(c) is the detection area of the original image. It can be clearly seen that the noise stream captures repair features that cannot be captured by the usage image, while the edges of the usage image are clearer and not blurred.
[0092] Therefore, referring to Figure 4 , an embodiment of the present invention provides a method for locating (forensics) an image repair area, including the steps of:
[0093] Step S101, obtain the original image to be located.
[0094] Step S102, convert the original image into the noise domain to obtain a first image, and extract a first feature map from the first image; wherein, the first feature map contains the inconsistency features between the source area and the noise repair area in the first image. The first image refers to the image obtained after converting the original image into the noise domain. The noise repair area inherits the noise pattern of the encoder-synthesized content, while the source area retains the noise pattern from the real content. Extracting the inconsistent features (features with differences) can distinguish the noise repair area and the source area, thereby providing richer information for subsequent location of the image repair area.
[0095] Step S103, extract local binary pattern features from the original image, merge the local binary pattern features with the original image to obtain a second image, and extract a second feature map from the second image; wherein, the second feature map contains the inconsistency features between the source area and the image repair area in the second image. Similarly to step S102, extracting the inconsistent features can distinguish the image repair area and the source area, facilitating providing richer information for subsequent location of the image repair area
[0096] Step S104, fuse the first feature map and the second feature map to obtain a fused feature map.
[0097] Step S105, locate the repair area in the fused feature map.
[0098] The noise domain can capture repair features that cannot be captured by the usage image. This method first extracts the noise features of the image from the noise domain and uses local binary patterns to extract image features to enhance the repair traces. Then, since both the noise features and the image features are the inconsistent features between the repair area and the source area, fusing the noise features and the image features can obtain richer features. Finally, based on the fused features for location, the location of image repair is achieved. This method can simultaneously capture different repair traces of sample block image repair and deep learning image repair, can efficiently and accurately locate the repaired area in the image, and can also solve the problem of unclear edges of the repaired area in current image repair location.
[0099] Second embodiment;
[0100] Reference Figure 5 , based on the above first embodiment, converting the original image into the noise domain to obtain a first image and extracting a first feature map from the first image in step S102 specifically includes:
[0101] Step S1021: Extract local noise descriptors in the original image according to a model filter rich in steganalysis to obtain a first image. Here, converting into the noise domain through a model (SRM) filter rich in steganalysis can show the repair traces that may be invisible in the RGB channels of the image.
[0102] Step S1022: Extract a first feature map from the first image according to multiple residual blocks, where the multiple residual blocks are connected in sequence, the first image is input into the first residual block, and a first feature map corresponding to the output of each residual block is obtained. Since with the increase of convolutional layers, although the feature map can obtain more semantic information, some image detail information and resolution will be lost. Step S1022 uses multiple residual blocks to output their corresponding features in sequence, and feature maps of different scales can be obtained, and multi-scale fusion of feature maps of different scales can be realized in subsequent steps. Multi-scale feature fusion can avoid the loss of detail information and provide richer information for locating the repaired area of the image.
[0103] In step S1022, three residual blocks are connected in sequence. Each residual block includes two bottleneck units. Each bottleneck unit includes three consecutive convolutional layers and an identity shortcut connection. Normalization and ReLu activation are performed before each convolution operation. The sizes of the sums of the three convolutional layers are 1×1, 2×2, and 1×1 respectively; the stride of the first two convolutional layers is 1, and the stride of the last convolutional layer is 2 to merge and reduce the spatial resolution.
[0104] Reference Figure 6 , the extraction of the second feature map from the second image in step S103 specifically includes the steps:
[0105] Step S1031: After channel expansion of the second image, perform continuous multiple downsamplings to reduce the image.
[0106] Step S1032: Extract local and global features from the downsampled second image according to fast Fourier convolution.
[0107] In step S1032, local and global features are extracted from the downsampled second image through multiple fast Fourier convolution blocks with the same configuration. Among them, each fast Fourier convolution block extracts features in the following way:
[0108] Step S1032a: Perform two ordinary convolutions on the downsampled second image to extract the first local feature and the second local feature of the second image; perform one ordinary convolution and one frequency-domain convolution on the downsampled second image to obtain the third local feature and the global feature.
[0109] Step S1032b: Add the first local feature and the third local feature to obtain the local feature of the second image; add the second local feature and the global feature to obtain the global feature of the second image.
[0110] As Figure 7 (Structure diagram of the fast Fourier convolution block) shown, the fast Fourier convolution (FFC) block first divides the feature map into two parts in the channel dimension. One part performs two ordinary convolutions to extract the local features F1 and F2 of the image, and the other part is further divided into an ordinary convolution to extract the local feature F3 of the image and a convolution in the frequency domain to extract the global feature F4 of the image. Adding F1 and F3 gives the local feature of the extracted image, and adding F2 and F4 gives the global feature of the extracted image. This enables the fast Fourier convolution block to capture complementary information with different receptive fields.
[0111] As Figure 8 (Structure diagram of the spectrum converter) shown, due to the presence of the spectrum converter, the fast Fourier convolution block can extract the global feature; the spectrum converter transforms the image from the spatial domain to the frequency domain. A point in the frequency domain corresponds to all the information in the entire spatial domain corresponding to that frequency. Convolution in the frequency domain extracts the global feature. Convolution in the spatial domain extracts the local feature of the image due to the limitation of the convolution kernel size. In this embodiment, the specific implementation manner of the spectrum converter includes: first using a 1×1 convolution to halve the number of channels, then transforming the original spatial feature to the frequency domain using the Fourier transform (which will convert the real function to a complex function), connecting the real part and the imaginary part in the frequency-domain feature, then passing through convolution, batch normalization, and the rectified linear unit function, separating the real part and the imaginary part again, applying the inverse Fourier transform to convert to the spatial domain and adding it to the source spatial-domain feature map. Finally, using a 1×1 convolution to restore the number of channels.
[0112] The Fourier transform is shown in Equation (1):
[0113]
[0114] where (x, y) represents the position in the spatial domain, (u, v) is the position in the frequency domain, s() is a sign function, u, v represent the positions in the frequency domain, and the inverse Fourier transform is calculated as:
[0115]
[0116] Step S1033: Continuously perform upsampling on the extracted local and global features multiple times to obtain a second feature map corresponding to the output of each upsampling.
[0117] It should be noted that the number of upsampling times in step S1033 corresponds to the number of residual blocks in the above step S1022. For example, in the above embodiment, three residual blocks output three first feature maps of different scales, and here three upsamplings are performed to obtain three second feature maps of different scales, so as to achieve multi-scale cross-fusion in subsequent steps.
[0118] Refer to Figure 9 , in some embodiments, since the above steps S1022 and S1033 can obtain feature maps of different scales, step S104 performs multi-scale cross-fusion, which specifically includes the following steps:
[0119] Step S1041: Perform the first multi-scale cross-fusion on the first feature map: Upsample, downsample, and keep unchanged each first feature map respectively to obtain intermediate feature maps; Connect the feature maps with the same size in the intermediate feature maps to obtain a noise cross-fusion feature map. Since the above step S1022 obtains first feature maps of different scales, and the first feature maps include features with differences between the noise repair region and the source region, therefore, step S1041 performs multi-scale cross-fusion on the first feature maps of different scales, avoiding the loss of detail information caused by the increase in the number of convolutional layers (although the feature maps can obtain more semantic information with the increase in the number of convolutional layers, some image detail information will be lost and the resolution will be reduced), and being able to provide richer information for later locating the repair region of the image.
[0120] Step S1042: Perform the second multi-scale cross-fusion on the second feature map: Upsample, downsample, and keep unchanged each second feature map respectively to obtain intermediate feature maps; Connect the feature maps with the same size in the intermediate feature maps to obtain an image cross-fusion feature map. Since the above step S1033 obtains second feature maps of different scales, and the second feature maps include features with differences between the image repair region and the source region, therefore, step S1042 performs multi-scale cross-fusion on the second feature maps of different scales, avoiding the loss of detail information caused by the increase in the number of convolutional layers, and being able to provide richer information for later locating the repair region of the image.
[0121] Step S1043: Connect the noise cross-fusion feature map and the image cross-fusion feature map, and perform a third multi-scale cross-fusion after connection to obtain a fused feature map. The noise features and the image features are different information. Performing multi-scale cross-fusion separately is to fuse the feature information of different scales in the noise and the image respectively. Since the multi-scale cross-fusion in Step S1041 and Step S1042 will output feature maps of different scales, Step S1043 is to finally and effectively fuse the fusion features between the noise and the image.
[0122] Referring to Figure 10 , Step S105 specifically includes the following steps:
[0123] Step S1051: Perform a convolution on the fused feature map. This is used to remove the artifacts that may be introduced by upsampling.
[0124] Step S1052: Perform per-pixel prediction and localization on the repaired area in the convolved fused feature map according to logistic regression.
[0125] Third Embodiment;
[0126] The image inpainting method based on deep learning has poor effect on inpainting the location sample block image, while the traditional method for inpainting the location sample block requires a lot of time for calculation and even requires manual operation. Currently, there is a lack of a method that can detect both the inpainting of the sample block image and the inpainting of the deep learning image.
[0127] Referring to Figures 11 to 14 , for the convenience of understanding by those skilled in the art, an embodiment of the present invention provides a positioning system for an image inpainting area. This system executes a positioning method for an image inpainting area. The system includes an enhancement block, an extraction block, and a decision block, where:
[0128] 1. The enhancement block includes a noise domain conversion module and an image feature addition module, which are respectively used to extract noise features using a model (SRM) filter rich in implicit writing analysis and extract image features using local binary pattern (LBP) to enhance the repair traces. 2. The extraction block includes a first feature extraction module and a second feature extraction module, which are respectively used to extract the repair traces in the noise features and the repair traces in the image features. 3. The multi-scale cross-fusion module is used in the decision block to effectively fuse the noise features and the image features, and efficiently and accurately realizes the positioning of image inpainting.
[0129] Referring to Figure 11 and Figure 12 , the method steps executed by this system are as follows:
[0130] Step S201: The noise domain conversion module converts the image to the noise domain using the SRM filter. The noise domain conversion module uses a model filter rich in steganalysis to extract the local noise descriptors of the original image, which can reveal the repair traces that may be invisible in the RGB channels of the image.
[0131] The kernel size of the model (SRM) filter rich in steganalysis is defined as 5×5×3, as follows:
[0132]
[0133] Feed the local noise descriptors into the first feature extraction module.
[0134] Step S202: The image feature addition module extracts the local binary pattern features of the original image and merges them with the original image. For a given original image (RGB image), calculate the local binary pattern features that describe the local texture features of the image. The calculation formula is as shown in Equation (3):
[0135]
[0136] where (x c , y c ) represents the central pixel, i c is the gray value of the central pixel, i p is the gray value of the adjacent pixel of the central pixel, and s(g) is a sign function, calculated as:
[0137]
[0138] Add the feature map to the original image, and after addition, convert it to a 4-channel image and feed it into the second feature extraction module.
[0139] Step S203: The first feature extraction module extracts the inconsistency features between the noise repair area and the source area. The noise repair area inherits the noise pattern of the encoder-synthesized content, while the source area retains the noise pattern from the real content. The flowchart for extracting features is as Figure 13 shown.
[0140] The first feature extraction module uses three residual blocks to extract features. Each block consists of two bottleneck units. In each bottleneck unit, there are three consecutive convolutional layers and an identity shortcut connection. Batch normalization and ReLu activation are performed before each convolutional operation. The sum sizes of the three convolutional layers are 1×1, 3×3, and 1×1 respectively. Most of the layer strides are 1, and the stride of the last convolutional layer in each block is 2 to merge and reduce the spatial resolution. Finally, 512 feature maps are learned, and their spatial resolution is 1 / 8 of the input image, as shown in Table 1 below.
[0141]
[0142] Table 1
[0143] Step S204: The second feature extraction module extracts the inconsistency features between the image repair area and the source area. The flowchart of its feature extraction is as Figure 14 shown.
[0144] In the module, a fast Fourier convolutional block is mainly used to extract features. First, the feature map of the RGB image merged with the local binary pattern image is channel-expanded. The 4 channels are expanded to 64 channels using a 1×1 convolution, and then the feature map is downsampled 3 times by a factor of 2, reducing the size of the feature map to 1 / 8 of the original, and the number of channels is increased to 512. Then, 9 fast Fourier convolutional blocks are used to extract global and local features, and finally the generated feature map is downsampled 3 times by a factor of 2. The network structure parameters are shown in Table 2.
[0145]
[0146]
[0147] Table 2
[0148] In "Fast Fourier Convolution #1 - #9" in Table 2, fast Fourier convolutional blocks with the same configuration are stacked to extract features, and its structure is as Figure 7 shown. It has been described in detail in the above embodiments and will not be elaborated here.
[0149] The convolution in the frequency domain in the fast Fourier convolutional block is a spectrum converter, and its structure is as Figure 8 shown. It has been described in detail in the above embodiments and will not be elaborated here.
[0150] Step S205: For the three output feature maps of the noise in step S203 and the three output feature maps of the image in step S204, multi-scale cross-fusion is performed on the noise features and the image features.
[0151] Step S2051: Taking the three output feature maps Fn1, Fn2, and Fn3 of the noise as examples, where Fn1 is 1 / 8 of the original image, Fn2 is 1 / 4 of the original image, and Fn3 is 1 / 2 of the original image. All upsampling and downsampling in the description are by a factor of 2. Since the output feature maps are of three scales, corresponding transformations (upsampling, downsampling, or no change) are performed on the three feature maps respectively. Fn1 is upsampled twice to obtain Fn1_1, upsampled once to obtain Fn1_2, and remains unchanged as Fn1_3; Fn2 is upsampled once to obtain Fn2_1, remains unchanged as Fn2_2, and is downsampled once to obtain Fn2_3; Fn3 remains unchanged as Fn3_1, is downsampled once to be Fn3_2, and is downsampled twice to be Fn3_3. In this way, nine feature maps are obtained, and the feature maps of the same size are concatenated to obtain the noise cross-fusion feature maps: (Fn1_1 + Fn2_1 + Fn3_1), (Fn1_2 + Fn2_2 + Fn3_2), (Fn1_3 + Fn2_3 + Fn3_3).
[0152] Step S2052: The same operations are performed on the three output feature maps of the image to obtain the image cross-fusion feature maps.
[0153] Step S2053: The noise cross-fusion feature maps and the image cross-fusion feature maps are concatenated, and then multi-scale feature fusion is performed once more.
[0154] Step S2054: The smallest feature map is upsampled twice, the medium-sized feature map is upsampled once, and then concatenated with the largest feature map. In this way, features of different scales are fused.
[0155] Step S2056: The transposed convolution is performed on the concatenated feature map once more for upsampling by a factor of 2. The resolution is enlarged.
[0156] Step S206: The features of different scales are convolved once to remove the artifacts that may be brought by upsampling, and finally the location of the repaired area is achieved through logistic regression to obtain accurate results.
[0157] The system and method provided by this embodiment have the following beneficial effects:
[0158] This system is based on deep learning and can greatly reduce the computing time consumed by traditional methods for location. Compared with the less used deep learning solutions currently, the network structure and corresponding methods of this system can capture different repair traces of sample block image repair and deep learning image repair, and can efficiently and accurately locate the repaired area in the image. Since the network structure of this system can extract rich features, it can also solve the problem of unclear edges in the area of image repair location currently.
[0159] In this system and method, the extracted image features and noise features are both inconsistent features between the repaired area and the source area. The differences in features are relatively large. By fusing the noise features and image features, richer features can be obtained.
[0160] When this system and method fuse features, they respectively extract the multi-scale feature maps of noise and the multi-scale feature maps of images, which can retain more detailed information. Finally, the multi-scale cross-fusion of the multi-scale feature maps of the two is performed, and the obtained fused feature map has richer information. Moreover, when extracting the multi-scale feature map of the image, a fast Fourier convolution block is used to efficiently and quickly extract the repair features of the image and texture from the global and local aspects, and complementary information with different receptive fields can be captured.
[0161] Fourth Embodiment;
[0162] This application also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it realizes: the method for locating the image repair area as described above.
[0163] The processor and the memory can be connected through a bus or other means.
[0164] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0165] The non-transitory software programs and instructions required to implement the method for locating the image repair area in the above embodiments are stored in the memory. When executed by the processor, they execute the method for locating the image repair area in the above embodiments. For example, they execute the method steps S101 to step S105 described above Figure 4 in, Figure 12 and the method steps S201 to step S206 in.
[0166] Fifth Embodiment;
[0167] This application also provides a computer-readable storage medium, storing computer-executable instructions for executing: the method for locating the image repair area as described above.
[0168] The computer-readable storage medium stores computer-executable instructions that are executed by a processor or a controller, for example, executed by a processor in the above-described electronic device embodiments, and can cause the above processor to execute the method for locating an image repair area in the above embodiments. For example, for example, execute the Figure 4 method steps S101 to S105 in Figure 12 and method steps S201 to S206 in
[0169] Those of ordinary skill in the art will appreciate that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing data, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired data and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any data delivery medium.
[0170] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0171] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for locating an image inpainting area, characterized in that, Including the steps: Obtain the original image to be located; Convert the original image into a noise domain to obtain a first image, and extract a first feature map from the first image; wherein, the first feature map contains the inconsistent features with differences between the source region and the noise repair region in the first image; the source region is the non-missing region; the noise repair region is the repaired region in the first image; Extract local binary pattern features from the original image, merge the local binary pattern features with the original image to obtain a second image, and extract a second feature map from the second image; wherein, the second feature map contains the inconsistent features with differences between the source region and the image repair region in the second image; the image repair region is the repaired region in the second image; the extracting the second feature map from the second image includes: After performing channel expansion on the second image, perform continuous multiple downsamplings; extract local and global features from the downsampled second image according to fast Fourier convolution; Perform continuous multiple upsamplings on the extracted local and global features to obtain a second feature map corresponding to the output of each upsampling; extracting local and global features from the downsampled second image through multiple fast Fourier convolution blocks with the same configuration, wherein each of the fast Fourier convolution blocks extracts features in the following manner: Perform two ordinary convolutions on the downsampled second image to extract the first local feature and the second local feature of the second image; perform one ordinary convolution and one frequency domain convolution on the downsampled second image to obtain the third local feature and the global feature; Add the first local feature and the third local feature to obtain the local feature of the second image; add the second local feature and the global feature to obtain the global feature of the second image; Fuse the first feature map and the second feature map to obtain a fused feature map; Locate the repaired region in the fused feature map.
2. The method for locating an image restoration area according to claim 1, wherein The converting the original image into a noise domain to obtain a first image, and extracting a first feature map from the first image includes: Extract the local noise descriptor in the original image according to the model filter rich in steganalysis to obtain a first image; Extract a first feature map from the first image according to multiple residual blocks, wherein the multiple residual blocks are connected in sequence, and the first image is input into the first residual block to obtain the first feature map corresponding to the output of each residual block.
3. The method for locating an image restoration area according to claim 2, wherein The multiple residual blocks are three residual blocks, each of the residual blocks includes two bottleneck units, each of the bottleneck units includes three consecutive convolutional layers and an identity shortcut connection, normalization and ReLu activation are performed before each convolutional operation, and the sum sizes of the three convolutional layers are , , ; the convolutional layer strides of the first bottleneck unit are all 1, and the last convolutional stride of the second bottleneck unit is 2.
4. The method for locating an image restoration area according to claim 1, wherein Perform continuous three upsamplings on the extracted local and global features.
5. The method for locating an image repair area according to claim 4, wherein The fusing the first feature map and the second feature map to obtain a fused feature map includes: Perform the first multi-scale cross fusion on the first feature map: perform upsampling, downsampling, and no change on each of the first feature maps respectively to obtain intermediate feature maps; connect the feature maps with the same size in the intermediate feature maps to obtain a noise cross fusion feature map; Perform a second multi-scale cross-fusion on the second feature map: perform upsampling, downsampling, and no change on each of the second feature maps respectively to obtain intermediate feature maps; connect the feature maps with the same size in the intermediate feature maps to obtain an image cross-fusion feature map. Connect the noise cross-fusion feature map and the image cross-fusion feature map, and perform a third multi-scale cross-fusion after connection to obtain a fusion feature map.
6. The method for locating an image restoration area according to any one of claims 1 to 5, characterized in that, The positioning of the repaired area in the fusion feature map includes: Perform one convolution on the fusion feature map. Perform pixel-by-pixel prediction and positioning on the repaired area in the convolved fusion feature map according to logistic regression.
7. A positioning device for an image restoration area, characterized in that, Comprising at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the positioning method of the image repair area according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the positioning method of the image repair area according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image modification area positioning method and system based on deep learning, and storage medium
CN113850307A