Small target difference detection method for view angle change image pair

By using the key point detection model and the descriptive sub-computation model, combined with homography matrix transformation and block amplification technology, the problem of difficult to detect small and medium-sized target differences in the existing technology is solved, and more efficient and accurate difference detection is achieved.

CN120219705APending Publication Date: 2025-06-27NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510227238.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing difference detection methods are difficult to accurately detect small target differences in image pairs with changes in view angles, and they are demanding on the corresponding image.

Method used

The key point detection model and descriptive sub-computation model are used to find the representative set of points in the image through the key point detection model, and the feature vectors are calculated for these points; then the homography matrix transformation and block amplification technology are used to improve the resolution and accuracy of the detection.

Benefits of technology

Accurate detection of small and medium-sized object differences by viewing angle changing images is achieved, and different objects that are smaller than traditional methods can be detected, and the detection efficiency and accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219705A_ABST
    Figure CN120219705A_ABST
Patent Text Reader

Abstract

The invention discloses a small target difference detection method for a view angle change image pair. The method comprises the following steps: inputting a pair of images with view angle change characteristics into a key point detection model and a descriptor calculation model; obtaining a rough matching point set, and calculating a homography matrix; respectively cutting the two images according to the coordinate ranges of the two rough matching point sets, transforming the images corresponding to the rough matching point sets kpts1 by using a homography matrix to obtain detection areas of the two images, partitioning the detection areas of the two images, and amplifying the resolution; pixel points are sampled on the corresponding image blocks; inputting the image block after the resolution is amplified into a descriptor calculation model to obtain a descriptor; calculating the similarity of each pair of corresponding sampling point descriptors; constructing a difference point list according to a similarity threshold value; removing noise points in the fuzzy region; converting the coordinates of all the difference points into coordinate positions in the original image; and clustering the coordinate positions of all the difference points, and representing the position coordinates of the difference region by using a clustering result. According to the invention, smaller difference objects in the view angle change image pairs can be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a method for detecting small target differences in a pair of images with perspective changes. Background Art

[0002] The detection of small target differences in a pair of images with perspective changes is an important research direction in the field of computer vision, and has important applications in fields including transportation, industry, medicine, photography, etc. Existing difference detection methods include human eye observation, pixel-by-pixel comparison method based on pixels, comparison method based on histograms, and detection method based on deep learning.

[0003] Currently, except for the inefficient human eye observation difference method, the difference detection algorithm cannot well detect the small target differences in a pair of images with perspective changes. The change in the target position caused by the perspective change brings great difficulties to the detection, and the pixel area of the small target difference is relatively small, which greatly reduces the accuracy of the difference detection. In addition, these methods do not make special contributions to the small target differences in a pair of images with perspective changes that frequently appear in various tasks. They require precise correspondence between images and have strict requirements for images.

[0004] In summary, there is a need for an accurate and fast method for detecting small target differences in a pair of images with perspective changes to solve the above difficulties. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for detecting small target differences in a pair of images with perspective changes to solve the problem that it is difficult to detect small target differences in a pair of images with perspective changes.

[0006] The technical solution for achieving the purpose of the present invention is as follows:

[0007] A method for detecting small target differences in a pair of images with perspective changes, comprising:

[0008] Step 1: Input a pair of images with perspective change features into a key point detection model and a descriptor calculation model; the key point detection model is used to find a representative point set in the image; the descriptor calculation model is used to calculate a feature vector for the points detected by the key point detection model to represent the key points;

[0009] After inputting the pair of images into the key point detection model and the descriptor calculation model, a rough matching point set kpts1 and kpts2 are obtained, and the homography matrix H corresponding to the pair of images is calculated using the findHomography function and the ransac algorithm of opencv;

[0010] Step 2: Crop the two images according to the coordinate ranges of the two sets of roughly matched points respectively. Transform the image corresponding to the set of roughly matched points kpts1 using the homography matrix H calculated in Step 1. The image corresponding to the set of roughly matched points kpts2 does not need to be transformed, and the detection regions of the two images are obtained. The detection regions are the first image after cropping and transformation and the second image after cropping. Divide the detection regions of the two images into blocks and increase the resolution of the image blocks.

[0011] Step 3: After increasing the resolution, sample pixel points on the corresponding image blocks from the two images. Input the image blocks with increased resolution into the descriptor calculation model to obtain the descriptors of the sampled points. Calculate the similarity of each pair of corresponding sampled point descriptors.

[0012] Step 4: Construct a list of different points according to the similarity threshold.

[0013] Step 5: Calculate the pixel standard deviation for the pixel region around each determined different point to remove the noise in the blurred region.

[0014] Step 6: Convert the coordinates of all different points to the coordinate positions in the original image.

[0015] Step 7: Cluster the coordinate positions of all different points, and use the clustering results to represent the coordinate positions of this different region.

[0016] Compared with the prior art, the significant advantages of the present invention are:

[0017] The key point detection model and descriptor calculation model trained by the present invention can detect more key points compared with the traditional manually made method. The key point detection model uses the denoising diffusion model to extract features. This feature extraction method can capture more detailed features related to key points, providing a feature representation including local details and global information for key point detection. The feature extraction method with multiple time steps can provide feature representations of different attention regions for key point detection. The descriptor model designed by the present invention is simple and lightweight, and the storage space occupied by the trained weights is smaller than that of other models. The present invention divides the image into blocks and enlarges them, and can detect smaller different objects than other difference detection algorithms. Description of the Drawings

[0018] Figure 1 It is a training process diagram of the denoising diffusion probability model, the key point detection model, and the descriptor calculation model.

[0019] Figure 2 It is a schematic diagram of feature extraction of the denoising diffusion probability model.

[0020] Figure 3 It is the feature aggregation module and feature decoding module of the key point detection model.

[0021] Figure 4 It is a structural schematic diagram of the descriptor calculation model.

[0022] Figure 5 It is a schematic diagram for determining the detection area in difference detection.

[0023] Figure 6 It is a schematic diagram for block division, resolution magnification, and calculation of the similarity of sampling points in difference detection.

[0024] Figure 7 It is a schematic diagram for determining the list of difference points and the difference area in difference detection.

[0025] Figure 8 It is a schematic diagram of the detection result of this embodiment. Specific implementation manner

[0026] The present invention will be further introduced below in conjunction with the accompanying drawings and specific embodiments.

[0027] Combined with Figures 1-8 , a small target difference detection method for a pair of images with perspective change of the present invention includes the following steps:

[0028] Step 1: Input a pair of images with perspective change features into the key point detection model and the descriptor calculation model to detect the key points and the descriptors corresponding to the key points;

[0029] 1. The purpose of the key point detection model is to find a representative set of points in the image. These points are unique, distinct, and distinguishable, usually being corner points, intersection points, and rounded corners in the image. The key point detection model is divided into two parts: a pre-trained denoising diffusion probabilistic model and a key point detection model. The pre-trained denoising diffusion probabilistic model is trained by the DDPM using a virtual synthetic dataset containing a large number of geometric shapes. The Denoising Diffusion Probabilistic Model (DDPM) is a new type of deep learning model based on probabilistic generation. Its core idea is to gradually recover data from noise by simulating the physical diffusion process. This model mainly includes two processes: the forward diffusion process and the reverse denoising process. In the forward diffusion process, the original data is transformed into pure noise by gradually adding Gaussian noise; in the reverse denoising process, a UNET network is trained to gradually reconstruct the original data from the noise. The virtual synthetic dataset is a synthetic dataset containing polygons, line segments, intersection lines, astroid lines, and checkerboards. These basic shapes cover most cases where corner points, intersection points, and rounded corner features appear. The virtual synthetic dataset can help the DDPM capture the detailed features of these geometric shapes, thereby helping the key point detection model detect key points. After the denoising diffusion probabilistic model is trained, its weights do not need to be changed. The feature representation composed of each feature map of the upsampling part of the UNET extracted from the denoising process of the pre-trained denoising diffusion probabilistic model is used as the pre-input for the next key point detection model. The extracted feature representation is divided into 5 layers in total, and the size of the feature map of each layer doubles compared to the previous layer. Three time steps are required to extract the feature map from the pre-trained denoising diffusion probabilistic model. The three time steps make the extracted feature map have features with multiple attention regions and multiple sizes. where f 0:4 represents the 5-layer feature maps from f0 to f4 at each of the three time steps from t1 to t3. t1 to t3 represent the three time steps. Extracts the feature representations at the three time steps from t1 to t3 for the pre-trained denoising diffusion probabilistic model. It is divided into three parts at the three time steps from t1 to t3. Each part represents the feature representation at its respective time step. Each part has 5-layer feature maps from f0 to f4. img represents the noisy input image received by the pre-trained denoising model during the denoising process. Represents the process of extracting features at the three time steps from t1 to t3 in the denoising process of the pre-trained denoising diffusion probabilistic model. The key point detection model first aggregates the feature representations extracted from the pre-trained denoising diffusion probabilistic model in the order of time step first and then size. The feature maps of the same layer at the three time steps from t1 to t3 Concatenate them using the torch.cat function. Here, n represents the n-th layer among the 5 layers of feature maps from f0 to f4: f n represents the concatenation result. Then, pass the concatenation result f n through a BLOCK module consisting of two convolutional layers, a SE attention mechanism module, and an upsampling module to obtain the calculation result of this layer of feature map: where represents the final calculation result of the n-th layer, F BLOCK represents the BLOCK module, F SE represents the SE attention mechanism module, F INTERPOLATE represents the upsampling module. The calculation result of each layer will be added to the concatenation result f n+1 of the next layer using the torch.cat function, and then operations are performed. After the aggregation of the five-layer feature representations is completed, the calculation result of the last 5th layer is input into a decoding module composed of two convolutional layers and a RELU activation function to obtain a two-channel result: out represents the final output of the key point detection module, F CONV1 represents the first convolution, F CONV2 represents the second convolution, F RELU represents the RELU activation function.

[0030] The training dataset of the key point detection model is a virtual synthetic dataset and a combined dataset containing pure noise images and elliptical images. Before training the key point detection model, different degrees of Gaussian noise are added to each image. These newly added images and noise addition operations are to increase the robustness of the key point detection model. In this combined dataset, for the virtual synthetic dataset part, the corner points, intersection points, and rounded corners of the geometric figures in the images are considered as the key points to be detected. A pure black image with the same size as the images in the dataset is created as the label for the training process, and the positions of the corner points, intersection points, and rounded corners of the geometric figures in the image are set to white.

[0031] The loss function of the key point detection model is L = CELoss(y pt - y label ), where CELoss is the cross-entropy loss function, y pt is the output result of the key point detection network, y label is the true label, and L is the calculated loss;

[0032] 2. The purpose of the descriptor calculation model is to calculate a feature vector for the detected key points to represent these key points. The present invention uses a 128-dimensional vector to represent the key points. The feature vector needs to satisfy that the corresponding key point descriptors of multiple images from different perspectives are similar. The structure of the descriptor calculation model is as follows. Input normalization layer: Use the "InstanceNorm2d" layer to normalize the input image, and the number of normalized channels is 1 (before the input image, average the three channels of the input RGB image into one channel); Feature extraction layer 1: The feature_extraction1 layer, which consists of 4 sub-modules submodel, gradually increases the number of channels (1 -> 6 -> 12 -> 24 -> 48), and uses a stride of 2 to downsample the feature map in the second and fourth sub-module layers; Feature extraction layer 2: The feature_extraction2 layer, which consists of 2 sub-modules submodel, increases the number of channels (48 -> 48 -> 96); Feature extraction layer 3: The feature_extraction3 layer, which consists of 3 sub-modules submodel, further increases the number of channels (96 -> 128 -> 128 -> 128), and does not use padding in the last sub-module submodel layer; Feature extraction layer 4: The feature_extraction4 layer, which consists of 3 sub-modules submodel, keeps the number of channels unchanged (128 -> 128 -> 128), and uses a stride of 2 to downsample the feature map in the first sub-module submodel layer; Feature extraction layer 5: The feature_extraction5 layer, which consists of 4 sub-modules submodel, increases the number of channels (128 -> 256 -> 256 -> 256 -> 128), and does not use padding in the last sub-module submodel layer; Feature fusion and output layer: Upsample the outputs of feature extraction layer 4 and feature extraction layer 5 to make their sizes consistent with the output of feature extraction layer 3, add the outputs of feature extraction layer 3, feature extraction layer 4 and feature extraction layer 5, and then process them through a des module composed of two sub-module submodel layers and a convolutional layer, and finally output a feature map with 128 channels; In the structure of the entire descriptor calculation model, each sub-module submodel is composed of a convolutional layer, a batch normalization layer and a ReLU activation function. The convolutional layer extracts local features on the feature map of each intermediate process through a sliding window. The batch normalization layer is used to calculate the mean and variance of the batch data of each channel of the intermediate process feature map, and then normalize the data of each channel into a distribution with a mean of 0 and a variance of 1. The ReLU activation function is used to increase the non-linear expression ability of the model.The size of the descriptor calculation result generated by the descriptor calculation model is [b, 128, h / 8, w / 8], where b is the image batch, h is the width of the image, and w is the length of the image.

[0033] The training dataset of the descriptor calculation model is composed of four datasets: COCO2017, SUN2012, MPII HUMAN POSE, and PASCAL VOC. Among them, COCO2017 accounts for 15%, SUN2012 accounts for 15%, MPII HUMAN POSE accounts for 40%, and PASCAL VOC accounts for 30%. The descriptor calculation model uses self-supervised training. First, a random image imga is sampled from the combined dataset, and a homography matrix H is randomly generated. The image imga is transformed by H to obtain the image imgb. Both imga and imgb are input into the descriptor calculation model to obtain the descriptor calculation results descriptor1 and descriptor2 corresponding to the two images. The pixel coordinates are transformed by H to obtain the transformed coordinates of imga and imgb and their corresponding original coordinates. Then, through these two coordinate sets, the one-to-one corresponding descriptor sets des1 and des2 of the two images are obtained.

[0034] The loss function of the descriptor calculation network is:

[0035]

[0036] L = NLLLoss(log(softmax(similarity)), y des ) + NLLLoss(log(softmax(similarity T )), y des )

[0037] where L is the calculated loss, F.normalize is the normalization function in torch, des1 is the descriptor set of imga, des2 is the descriptor set of img2, dim is the dimension of des1 and des2, softmax is the softmax activation function, similarity is the cosine similarity matrix of des1 and des2, NLLLoss is the negative log-likelihood loss function, and y des is the label matrix with the same shape as the cosine similarity matrix, y des has a diagonal of 1, and y des has 0s in all other positions, and T is the transpose.

[0038] 3. Input the perspective change image pair img1 and img2 for which differences are to be detected into the trained keypoint detection model and descriptor calculation model, obtaining two keypoint detection outputs keypoint_map1 and keypoint_map2 of the keypoint detection module, as well as two descriptor outputs descriptor1 and descriptor2 of the descriptor calculation model. Both keypoint_map1 and keypoint_map2 are two-channel, where the second channel represents the heat value of the pixel point being a keypoint. The higher the heat value of a certain pixel point, the higher the probability that the pixel point is a keypoint. Select the top 10,000 pixel points with high heat values in keypoint_map1 and keypoint_map2 as candidate pixel points. Perform a sparse two-dimensional interpolation operation on descriptor1 and descriptor2 (the size of the descriptor model result is one-eighth of the input image) to obtain the descriptors of the candidate pixel points at the input image size. Sparse two-dimensional interpolation: First, normalize the candidate pixel points of img1 and img2, and then use the grid_sample function of torch to interpolate to obtain the descriptors des1 and des2 of the candidate pixel points of img1 and img2. Use the mutual nearest neighbor search algorithm of cosine similarity to match the candidate pixel points to obtain two rough matching point sets kpts1 and kpts2. Use the findHomography function and ransac algorithm of opencv to calculate the homography matrix H corresponding to img1 and img2.

[0039] Step 2: Crop the images img1 and img2 respectively according to the coordinate ranges of the rough matching point sets kpts1 and kpts2 to obtain the overlapping regions of the two images. Transform the first cropped image with the homography matrix H calculated in step 1 of 3 to obtain the detection regions of the two images. The detection regions are the image img1_p obtained by cropping and transforming img1 and the image img2_p obtained by cropping img2. Divide img1_p and img2_p into blocks respectively, and magnify the resolution of each image block to enhance the detailed features of small target differences in the images. The size of the image block and the magnification factor of the resolution are determined by the complexity of the image. The more complex the image, the more blocks are divided. The image blocks of img1_p and img2_p correspond one by one.

[0040] Step 3: Sample pixel points on an image patch after upsampling the resolution of img1_p. The sampling method is to uniformly select pixel points from the upper left corner of the image patch from left to right and from top to bottom according to the sampling granularity determined by the image complexity and the estimated difference region range to obtain the sampled pixel point set pts1. Among them, the more complex the image, the smaller the sampling granularity, and the more similar the images, the smaller the sampling granularity. Input the image patch after upsampling the resolution into the descriptor calculation model and obtain the descriptor set des1_pt1 of the sampled pixel point set pts1 according to the sparse two-dimensional interpolation method in Step 1. Perform the same operation on the corresponding image patch of image img2_p to obtain the sampled pixel points pts2 and the descriptor set des2_pt2 of the sampled pixel points. Calculate the similarity of each corresponding sampled point descriptor. This similarity is the L2 distance, and the form is as follows: where difference is the L2 distance similarity, i is the dimension number of the descriptor, des1_pt and des2_pt are the descriptors of the corresponding sampled points from des1_pt1 and des1_pt2 respectively, des1_pt[i] is the value of the i-th dimension in des1_pt, and des2_pt[i] is the value of the i-th dimension in des2_pt.

[0041] Step 4: Construct a list of difference points according to the similarity threshold. Compare the similarity of the corresponding sampled points with the threshold. The point pairs with the similarity of the corresponding points greater than or equal to the threshold are considered as difference points and added to the list of difference points, and the point pairs with similarity less than the threshold are considered as similar points and not processed. When no difference point in a pair of image patches is added to the list of difference points, a smaller similarity threshold will be selected to reconstruct the list of difference points. If no difference point is added to the list of difference points after selecting a smaller similarity threshold once, it is determined that there is no difference region in this pair of image patches.

[0042] Step 5: Calculate the pixel standard deviation of the pixel region around each determined difference point to remove the noise in the blurred region; extract the region R with a size of N×N centered on the difference point (x, y); calculate the average value μ of all pixel values in region R, and the form is as follows: where (x, y) represents the pixel point in region R, and I(x, y) represents the pixel value of the pixel point (x, y); calculate the average value of the squares of the differences between all pixel values I(x, y) in region R and the average value μ, that is, the variance σ 2 , and the form is as follows: Compare the pixel standard deviation σ of region R with the color change threshold. Those less than the threshold are considered as blurred regions, and the difference points at the center of this region are removed from the list of difference points.

[0043] Step 6: Convert the coordinates of all difference points from the divided and upsampled image patches to the coordinate positions of the original image img2. where x patch represents the horizontal coordinate of the difference point in the image block, and y patch represents the vertical coordinate of the difference point in the image block, w patch represents the length of the segmented image, and h patch represents the width of the segmented image, and w resized represents the length of the magnified image block, and h resized represents the width of the magnified image block.

[0044] x orginal = xM + grid x × w patch + x patch , y orginal = yM + grid y × h patch + y patch , where x orginal represents

[0045] the horizontal coordinate of the difference point in the original image, and y orginal represents the vertical coordinate of the difference point in the original image, xM represents the horizontal coordinate of the upper left corner point of the cropped img2 in img2, yM represents the vertical coordinate of the upper left corner point of the cropped img2 in img2, and grid x represents the horizontal position of the image block where the difference point is located among all image blocks, and grid y represents the vertical position of the image block where the difference point is located among all image blocks.

[0046] Step Seven: Cluster the coordinate positions of all difference points in the original image. The clustering result is to calculate the central position coordinates of each difference point cluster, and the central position coordinates of each difference point cluster are used to represent the position coordinates of this difference region.

[0047] Example:

[0048] Combined with Figures 1-8 , a small target difference detection method for a pair of images with perspective changes according to the present invention includes the following steps:

[0049] Step One: Input a pair of images with perspective change features into a key point detection model and a descriptor calculation model to detect key points and descriptors corresponding to the key points.

[0050] First, describe the training process of the key point detection model and the descriptor calculation model in this embodiment. As Figure 1As shown in the figure, in step 101, a denoising diffusion probabilistic model is trained with a virtual synthetic dataset containing polygons, polyhedrons, line segments, intersection lines, astroids, broken lines, and checkerboard geometric figures. In this embodiment, the number of iterations for training the denoising diffusion probabilistic model is 170,000 times, the GPU used is NVIDIA 3090, and the learning rate is 1e-05. In step 102, as Figure 2 shown, the feature representation formed by combining each feature map of the upsampling part of UNET is extracted from the denoising process of the denoising diffusion probabilistic model as the pre-input of the next key point detection model. During the denoising process, the upsampling part of UNET performs upsampling four times in total. Before and after each upsampling layer, there is a sub-module composed of three-layer residual structure and attention mechanism, and the results of each sub-module are output as the feature representation of the key points to be detected.

[0051] In step 103, the virtual synthetic dataset used to train DDPM before is added with pure noise images and elliptical images, and different degrees of Gaussian noise are added to each image to form a new combined dataset. In this part of the images in the virtual synthetic dataset in this combined dataset, the corner points, intersection points, and rounded corners of geometric figures are considered as key points to be detected. A pure black image with the same size as the images in the dataset is created as the label for the training process, and the positions of the vertices of geometric figures in the image are set to white. First, an image is selected from the dataset and input into the diffusion model, and a noise image img t is created, and the noise variance is determined by the time step. Among them, the noise ε~N(0,I), and the noise variance γ t about the time step t is given by , where α i is a scheduler that controls the noise variance added at each step, and t is each time step. Select the time step t∈{50, 100, 400}, and obtain the feature representation of the input image in the manner of step 102: to represent the five-layer features at time step t, and founction DDPM represents the feature extraction method described in step 102. The structure of the key point detection model is as Figure 3 shown. The feature representation generated by the pre-trained denoising diffusion probabilistic model is input into the feature aggregation module of the key point detection model as shown in the figure. First, the feature maps of the nth layer at different time steps are concatenated using the torch.cat function, where t1 = 50, t2 = 100, t3 = 400: Then, the concatenated result f n passes through a BOLCK module composed of two layers of convolution, a SE attention mechanism module, and an upsampling module to obtain the calculation result of this layer of feature map: Among them represents the final calculation result of each layer, F BLOCK represents the BLOCK module, F SE represents the SE attention mechanism module, F INTERPOLATE represents the upsampling module. The calculation result of each layer will be added to the result f concatenated by the torch.cat function with the next layer, and then n+1 perform the operation. After the five-layer feature representations are all aggregated, the calculation result of the fifth layer is input into the feature decoding module composed of two convolutional layers and the RELU activation function to obtain a two-channel result: Among them, out represents the final output of the key point detection module, F CONV1 represents the first convolution, F CONV2 represents the second convolution, F RELU represents the RELU activation function. The key point detection result of each iteration result is obtained for each iteration, and this result is compared with the label image. Minimize the loss function L = CELoss(y - y label ), where CELoss is the cross-entropy loss function, y is the output result of the key point detection model, and y label is the label image, and L is the calculated loss;

[0052] The structure of the descriptor sub-model in step 104 is as follows: Input normalization layer: The "InstanceNorm2d" layer is used to normalize the input image, and the number of normalized channels is 1 (before the input image, the three channels of the input RGB image are averaged into one channel); Feature extraction layer 1: The feature_extraction1 layer, which consists of 4 sub-modules submodel, gradually increases the number of channels (1->6->12->24->48), and uses a stride of 2 to downsample the feature map in the second and fourth sub-module layers; Feature extraction layer 2: The feature_extraction2 layer, which consists of 2 sub-modules submodel, increases the number of channels (48->48->96); Feature extraction layer 3: The feature_extraction3 layer, which consists of 3 sub-modules submodel, further increases the number of channels (96->128->128->128), and no padding is used in the last sub-module submodel layer. Feature extraction layer 4: The feature_extraction4 layer, which consists of 3 sub-modules submodel, keeps the number of channels unchanged (128->128->128), and uses a stride of 2 to downsample the feature map in the first sub-module submodel layer. Feature extraction layer 5: The feature_extraction5 layer, which consists of 4 sub-modules submodel, increases the number of channels (128->256->256->256->128), and no padding is used in the last sub-module submodel layer; Feature fusion and output layer: Upsample the outputs of feature extraction layer 4 and feature extraction layer 5 to make their sizes consistent with the output of feature extraction layer 3, add the outputs of feature extraction layer 3, feature extraction layer 4 and feature extraction layer 5, and then process them through a des module consisting of two sub-module submodel layers and a convolutional layer, and finally output a feature map with 128 channels; In the structure of the entire descriptor calculation model, each sub-module submodel is composed of a convolutional layer, a batch normalization layer and a ReLU activation function. The size of the descriptor calculation result generated by the descriptor calculation model is [b, 128, h / 8, w / 8], where b is the image batch, h is the width of the image, and w is the length of the image.

[0053] Combine the images from four datasets, namely COCO2017, SUN2012, MPII HUMAN POSE, and PASCAL VOC, into a new dataset. Among them, COCO2017 accounts for 15%, SUN2012 accounts for 15%, MPII HUMAN POSE accounts for 40%, and PASCAL VOC accounts for 30%. Training method of the descriptor calculation model: First, randomly sample an image imga from the new dataset and randomly generate a homography matrix H. Transform the image imga with H to obtain the image imgb. Input both imga and imgb into the descriptor calculation model to obtain the descriptor calculation results descriptor1 and descriptor2 corresponding to the two images. The strategy for randomly generating a homography matrix H is to randomly sample and generate a homography matrix composed of the combination of perspective, scaling, rotation, and translation. Transform the pixel coordinates with H to obtain the transformed coordinates of imga and imgb and their corresponding original coordinates. First, obtain the inverse homography matrix H_inv of the homography matrix H. Map the coordinates of imgb back to the coordinates of the input image imga. Use torch.meshgrid to create the grid coordinates grid of the input image imga. Transform grid through the inverse homography matrix H_inv. The transformed coordinates grid_new are the coordinates in the input image imga corresponding to each point in the target image imgb. Filter out the coordinates that exceed the boundary of the target image after transformation to ensure that only valid coordinate pairs are retained. Obtain the descriptor sets des1 and des2 corresponding to the two images through these two coordinate sets. Minimize the cosine similarity between des1 and des2 to train the descriptor calculation model. The loss function is:

[0054]

[0055]

[0056] L = NLLLoss(log(softmax(similarity)), y des ) + NLLLoss(log(softmax(similarity T )), y des )

[0057] where L is the calculated loss, F.normalize is the normalization function in torch, des1 is the descriptor set of imga, des2 is the descriptor set of img2, dim is the dimension of des1 and des2, softmax is the softmax activation function, similarity is the cosine similarity matrix of des1 and des2, NLLLoss is the negative log-likelihood loss function, y desis a label matrix with the same shape as the cosine similarity matrix, y des The diagonal is 1, y des and the rest of the positions are 0.

[0058] As Figure 5 shown, there are obvious perspective changes between Image A and Image B. There are 3 obvious differences in the images, and there are supplementary and regional parts in the images.

[0059] Step 2: In Step 201, input Image A and Image B into the pre-trained denoising diffusion probabilistic model to obtain feature representations. Input the feature representations into the keypoint detection model to obtain 2-channel keypoint heatmaps keypoint_map1 and keypoint_map2. Select the pixels with the top 10,000 high heat values in keypoint_map1 and keypoint_map2 as candidate pixels. Input Image A and Image B into the descriptor calculation model to obtain outputs descriptor1 and descriptor2. Perform sparse 2D interpolation on descriptor1 and descriptor2 to obtain the descriptors of the candidate pixels at the input image size. Sparse 2D interpolation: First, normalize the candidate pixels of Image A and Image B, and then use the grid_sample function of torch to interpolate to obtain the descriptors des1 and des2 of the candidate pixels of Image A and Image B. Use the mutual nearest neighbor search algorithm of cosine similarity to match the 10,000 candidate pixels to obtain the rough matching point sets kpts1 and kpts2. Use the findHomography function and ransac algorithm of opencv to calculate the homography matrix H corresponding to Image A and Image B and the fine matching point sets P1 and P2. Step 202: Crop Image A and B according to the coordinate ranges of the rough matching point sets kpts1 and kpts2 to obtain the overlapping regions of the two images. Transform the first cropped image with the previously calculated homography matrix H to obtain the detection regions of the two images. The detection regions are the image img1_p obtained by cropping and transforming Image A and the image img2_p obtained by cropping img2.

[0060] Step 3: As Figure 6 shown, in Step 301, divide the detection regions img1_p and img2_p into blocks. In this embodiment, the number of blocks is 2 image blocks in the horizontal direction and 2 image blocks in the vertical direction, for a total of 4 image blocks. Both parts of the image blocks are as Figure 6One-to-one correspondence as shown on the right. Step 302, use the resize function of opencv to enlarge the resolution of the image patch. In this embodiment, the magnification factor is 5 times. Uniformly sample pixel points in the enlarged image patch. The sampling method is to start from the upper left corner of the image patch, from left to right, from top to bottom, and select a pixel point every sampling interval. In this embodiment, the sampling interval is 0.01 times the resolution of the enlarged image. Input the image patch with the enlarged resolution into the descriptor calculation model and obtain the descriptor des1_pt1 of the sampled pixel point set pts1 according to the sparse two-dimensional interpolation method. Perform the same operation on the corresponding image patch of the image img2_p to obtain the sampled pixel points pts2 and the descriptor des2_pt2 of the sampled pixel points. Calculate the similarity of the corresponding sampled points:

[0061] where difference is the L2 distance similarity, i is the number of dimensions of the descriptor, des1_pt and des2_pt are the descriptors of the corresponding sampled points from des1_pt1 and des1_pt2 respectively, des1_pt[i] is the value of the i-th dimension in des1_pt, and des2_pt[i] is the value of the i-th dimension in des2_pt.

[0062] Step Four. As Figure 7 shown, in step 401, construct a list of difference points according to the similarity threshold. Compare the similarity of the corresponding sampled points with the threshold. The corresponding points with similarity greater than or equal to the threshold are considered as difference points and added to the list of difference points, and those less than the threshold are considered as similar points and not processed; when no difference point in an image patch is added to the list of difference points, a smaller similarity threshold will be selected to reconstruct the list of difference points. In this embodiment, the similarity threshold is 1.3, and the smaller similarity threshold is 1.

[0063] Step Five. In step 402, calculate the pixel standard deviation of the pixel region around each determined difference point to remove the noise in the blurred region; extract the region R with a size of N×N centered on the difference point (x, y); calculate the average value μ of all pixel values in the region R, in the following form: where (x, y) represents the pixel point in the region R, and I(x, y) represents the pixel value of the pixel point (x, y); calculate the average value of the squares of the differences between all pixel values I(x, y) in the region R and the average value μ, that is, the variance σ 2 , in the following form: Compare the pixel standard deviation σ of the region R with the color change threshold. Those less than the threshold are considered as blurred regions, and the difference points at the centers of these regions are removed from the list of difference points.

[0064] Step Six. In step 403, convert the coordinates of all difference points from the divided and enlarged-resolution image patches to the coordinate positions of the original images A and B where x patch represents the horizontal coordinate of the difference point in the image block, and y patch represents the vertical coordinate of the difference point in the image block, w patch represents the length of the segmented image, h patch represents the width of the segmented image, w resized represents the length of the magnified image block, h resized represents the width of the magnified image block.

[0065] x orginal = xM + grid x × w patch + x patch and y orginal = yM + grid y × h patch + y patch where x orginal represents

[0066] the horizontal coordinate of the difference point in the original image, y orginal represents the vertical coordinate of the difference point in the original image, xM represents the horizontal coordinate of the upper left corner point of the cropped img1 in img1, yM represents the vertical coordinate of the upper left corner point of the cropped Image A in Image B, and grid x represents the horizontal position of the image block where the difference point is located among all image blocks, and grid y grid x represents the vertical position of the image block where the difference point is located among all image blocks.

[0067] Step 7: Step 404 clusters the coordinate positions of all the difference points in the original image. The clustering method is as follows: First, initialize the list of difference points, and add a classification label (initially 0, indicating unclassified) to each point; randomly select a starting point, randomly select a point from the list of difference points as the initial clustering center, and record the coordinates of this point as the current center center; perform mean shift iteration. In each iteration, calculate the average shift direction of all points within the neighborhood of the current center center, move the current center in this shift direction, update the center position, and repeat the above process until the center position no longer changes significantly (the shift amount is less than the threshold) or reaches the maximum number of iterations (10 times); judge the new clustering center. If the distance between the newly found center and the existing clustering centers is greater than a certain threshold, then take it as a new clustering center, otherwise, ignore this center; classify all sample points, traverse all sample points, calculate the distance from each point to all clustering centers, assign each point to the clustering center with the closest distance, and update the number of samples of the clustering center; output the results, return the classification label of each sample point, as well as the position and number of samples of each clustering center. Calculating the center position coordinates of each difference point cluster represents the position coordinates of this difference. Use the circle function in opencv to circle the clustering centers to display the difference detection results.

[0068] As Figure 8 shown, the difference detection results of Image A and Image B are circled by white circular frames in the figure. This invention patent can make good adaptability to the perspective changes, illumination changes, and the difference regions of the three small targets in the embodiments, and detect the difference objects in the embodiments.

[0069] The above series of detailed descriptions are only specific descriptions of the feasible implementation manners of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit the protection scope of the present invention. The scope of the present invention is defined by the claims. Those skilled in the art can make modifications, substitutions, and improvements without departing from the spirit and principle of the present invention, and these improvements, modifications, and substitutions are also regarded as within the protection scope of the present invention. The content not detailedly described in this specification belongs to the prior art well-known to those skilled in the art.

Claims

1. A small target difference detection method for a pair of images with changing viewing angles, characterized in that: include: Step 1: Input a pair of images with perspective change features into a key point detection model and a descriptor calculation model; The key point detection model is used to find a representative point set in the image; the descriptor calculation model is used to calculate a feature vector representing the key point for the point detected by the key point detection model; After inputting the image pair into the key point detection model and the descriptor calculation model, the rough matching point sets kpts1 and kpts2 are obtained, and the homography matrix H corresponding to the image pair is calculated using the opencv findHomography function and the ransac algorithm; Step 2: Crop two images respectively according to the coordinate range of the two coarse matching point sets, transform the image corresponding to the coarse matching point set kpts1 with the homography matrix H calculated in step 1, and the image corresponding to the coarse matching point set kpts2 does not need to be transformed, and obtain the detection areas of the two images, which are the cropped and transformed first image and the cropped second image; divide the detection areas of the two images into blocks, and enlarge the resolution of the image blocks; Step 3, after enlarging the resolution, sample pixel points on the corresponding image blocks from the two images; input the image blocks after enlarging the resolution into the descriptor calculation model to obtain the descriptors of the sampling points; calculate the similarity of each pair of corresponding sampling point descriptors; Step 4: Build a list of differences based on the similarity threshold; Step 5: Calculate the pixel standard deviation for the pixel area around each determined difference point to remove noise in the blurred area; Step 6: Convert the coordinates of all difference points into the coordinate positions in the original image; Step 7: Cluster the coordinates of all the difference points and use the clustering results to represent the position coordinates of the difference area.

2. The small target difference detection method for a pair of images with changing viewing angles according to claim 1, characterized in that: The key point detection model includes a denoising diffusion probability model and a key point detection model; The denoising diffusion probability model is used to provide multi-time step and multi-scale feature representation for the key point detection model; The key point detection model is used to perform feature aggregation and feature decoding on the multi-time step and multi-scale feature representation provided by the denoising diffusion probability model to obtain a key point detection result.

3. The small target difference detection method for perspective change image pairs according to claim 2, characterized in that: Provide multi-time-step and multi-scale feature representations for the key point detection model from the denoising process; extract each feature map of the UNET upsampling part from the denoising process of the pre-trained denoising diffusion probability model, combine them into a feature representation, and then use the feature representation as the pre-input of the next key point detection model; the extracted feature representation is divided into 5 layers, and the size of the feature map of each layer is doubled compared to the previous layer: Among them, f 0:4 Represents the 5-layer feature map of each time step f0 to f4 in the three time steps from t1 to t3; t1 to t3 represent three time steps; Extract feature representations of three time steps from t1 to t3 for the pre-trained denoising diffusion probability model. It is divided into three parts of three time steps from t1 to t3. Each part represents the feature representation of its own time step. Each part has 5 layers of feature maps from f0 to f4. img represents the noisy input image received by the pre-trained denoising model during the denoising process. It represents the process of extracting features by the pre-trained denoising diffusion probability model in the three time steps t1 to t3 of the denoising process.

4. The small target difference detection method for perspective-changing image pairs according to claim 2, characterized in that: The key point detection model first aggregates the feature representations extracted by the pre-trained denoising diffusion probability model in the order of time step first and size second; the feature maps of the same layer at three time steps from t1 to t3 are aggregated. Use torch.cat function to concatenate them, where n represents the nth layer in the 5-layer feature map from f0 to f4: f n Represents the splicing result; then, the splicing result f n After a BOLCK module consisting of two layers of convolution, an SE attention mechanism module and an upsampling module, the calculation result of this layer of feature map is obtained: in Indicates the final calculation result of the nth layer, F BLOCK Indicates BLOCK module, F SE represents the SE attention mechanism module, F INTERPOLATE Represents the upsampling module; the calculation results of each layer Will be concatenated with the result f of the next layer using the torch.cat function n+1 Add, then proceed Operation: When all five layers of feature representation are aggregated, the calculation results of the last 5th layer are Input to the decoding module consisting of two convolutional layers and RELU activation function, and get a two-channel result: out represents the final output of the key point detection module, F CONV1 represents the first convolution, F CONV2 represents the second convolution, F RELU Represents the RELU activation function.

5. The small target difference detection method for perspective change image pairs according to claim 1, characterized in that: The descriptor calculation model includes: Input normalization layer, which is used to normalize each channel of the input image pair independently; Feature extraction layer 1 is used to increase the number of channels of the intermediate process feature map from 1 to 48 and halve the spatial resolution of the image; Feature extraction layer 2, used to increase the number of channels of the intermediate process feature map from 48 to 96; Feature extraction layer 3, used to increase the number of channels of the intermediate process feature map from 96 to 128; Feature extraction layer 4, used to halve the spatial resolution of the image; Feature extraction layer 5 is used to increase the number of channels of the intermediate process feature map from 128 to 256 and then reduce it back to 128; The feature fusion and output layer is used to fuse the features of feature extraction layer three, feature extraction layer four and feature extraction layer five, and finally output a 128-channel feature map.

6. The small target difference detection method for perspective-changing image pairs according to claim 5, characterized in that: Feature extraction layer 1 to feature extraction layer 5 include 4, 2, 3, 3, and 4 submodules, respectively. Each submodule contains a convolution layer, a batch normalization layer, and a ReLU activation function. The convolution layer extracts local features on the feature map of each intermediate process through a sliding window; The batch normalization layer is used to calculate the mean and variance of each channel batch data of the intermediate process feature map, and then normalize the data of each channel to a distribution with a mean of 0 and a variance of 1; The ReLU activation function is used to increase the nonlinear expression ability of the model.

7. The small target difference detection method for perspective-changing image pairs according to claim 1, characterized in that: Pixel points are sampled on the image block from the first image. The sampling method is to uniformly select pixels from the upper left corner of the image block from left to right and from top to bottom according to the sampling granularity determined by the image complexity and the estimated difference area range, so as to obtain the sampling pixel point set pts1; the image block after the resolution is enlarged is input into the descriptor calculation model and the descriptor set des1_pt1 of the sampling pixel point set pts1 is obtained; the same operation is performed on the corresponding image block from the second image to obtain the sampling pixel point pts2 and the descriptor set des2_pt2 of the sampling pixel point.

8. The small target difference detection method for a pair of images with changing viewing angles according to claim 7, characterized in that: Calculate the similarity of the descriptor of each corresponding sampling point in the following form: Where difference is the L2 distance similarity, i is the number of dimensions of the descriptor, des1_pt and des2_pt are the descriptors from the corresponding sampling points in des1_pt1 and des1_pt2 respectively, des1_pt[i] is the value of the i-th dimension in des1_pt, and des2_pt[i] is the value of the i-th dimension in des2_pt.

9. The small target difference detection method for perspective-changing image pairs according to claim 1, characterized in that: Calculate the pixel standard deviation for the pixel area around each determined difference point to remove the noise in the blurred area: extract the N×N area R centered on the difference pixel point (x, y); calculate the average value μ of all pixel values ​​in area R: Where I(x,y) represents the pixel value of the difference pixel (x,y); calculate the average of the squares of the differences between all pixel values ​​I(x,y) in the region R and the average value μ, that is, the variance σ 2 : Compare the pixel standard deviation σ in region R with the color change threshold. The area with a value less than the threshold is considered to be a blurred area, and the difference point at the center of this area is removed from the difference point list.

10. The small target difference detection method of the perspective change image pair according to claim 1, characterized in that: Convert all difference point coordinates into the coordinate positions in the original image: Convert all difference point coordinates from the divided and enlarged resolution image blocks into the coordinate positions of the original image img2. where x patch Indicates the horizontal coordinate of the difference point in the image block, y patch Represents the vertical coordinate of the difference point in the image block, w patch Indicates the length of the block image, h patch Indicates the width of the block image, w resized Indicates the length of the enlarged image block, h resized Indicates the width of the image block after enlargement; x orginal =xM+grid x × patch +x patch ,y orginal =yM+grid y ×h patch +y patch , where x orginal express The horizontal coordinate of the difference point in the original image, y orginal Indicates the vertical coordinates of the difference point in the original image, xM indicates the horizontal coordinates of the upper left corner of img2 after cropping, yM indicates the vertical coordinates of the upper left corner of img2 after cropping, grid x Indicates the horizontal position of the image block where the difference point is located among all image blocks, grid y Indicates the vertical position of the image block where the difference point is located among all image blocks.