Image matching method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202210631040.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-06-06
AI Technical Summary
[0003]然而,拍摄视角、光照、拍摄场景和拍摄器材等外部环境的变化,往往会对同一物体在不同时间的成像造成影响,再加上噪声、旋转等对图像的影响,都会导致搜索图像和目标图像之间存在一定程度的差异,进而增加匹配的难度
[0065]本申请的方案中,首先分别获取与搜索图像对应的搜索图像特征图,以及与目标图像对应的目标图像特征图。再对搜索图像特征图与目标图像特征图进行互相关操作,得到互相关得分图,并对互相关得分图进行尺寸变换,得到第一得分图,也即得到了搜索图像特征图与目标图像特征图之间的相似度得分图,得分越高,相似度就越高,又由于第一得分图的尺寸与搜索图像的尺寸相同,则第一得分图中每个得分在第一得分图中的位置,相当于该得分在搜索图像中的位置。相应的,搜索图像中与目标图像的相关度最高的第一目标区域,也即第一得分图中得分最高的区域。如此,根据第一得分图,就可以从搜索图像中直接确定出与目标图像的相关度最高的第一图像区域,将其作为与目标图像相匹配的匹配区域,实现了对搜索图像中与目标图像相匹配的图像区域的准确定位,提高了图像匹配效率。
Smart Images

Figure CN115187797B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to an image matching method, apparatus, electronic device, and storage medium. Background Technology
[0002] Image matching is a popular application in image processing. It involves using specific algorithms to find image regions in a search image that match a target image. The search image is usually large, while the target image is relatively small.
[0003] However, changes in the external environment, such as shooting angle, lighting, shooting scene, and shooting equipment, often affect the imaging of the same object at different times. Combined with the effects of noise and rotation, these factors can lead to differences between the search image and the target image, increasing the difficulty of matching. Therefore, accurately identifying the image region that matches the target image from the search image is a crucial problem that urgently needs to be solved. Summary of the Invention
[0004] In view of this, this application provides an image matching method, apparatus, electronic device and storage medium that can accurately find the image region that matches the target image from the search image.
[0005] To achieve the above objectives, this application adopts the following technical solution:
[0006] The first aspect of this application provides an image matching method, comprising:
[0007] Obtain the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, respectively;
[0008] A cross-correlation operation is performed on the feature map of the search image and the feature map of the target image to obtain a cross-correlation score map. The cross-correlation score map is then resized to obtain a first score map, the size of which is the same as the size of the search image.
[0009] Based on the first score map, the first image region with the highest relevance to the target image is determined from the search image, and is used as the image region that matches the target image.
[0010] Optionally, the target image includes multiple images of different sizes obtained by cropping the original target image;
[0011] Perform a cross-correlation operation on the search image feature map and the target image feature map to obtain a cross-correlation score map, including:
[0012] The search image feature map is cross-correlated with the target image feature map of different sizes to obtain the cross-correlation sub-score map.
[0013] The cross-correlation sub-score maps are fused to obtain the cross-correlation score map.
[0014] Optionally, the search image feature map corresponding to the search image and the target image feature map corresponding to the target image are obtained respectively, including:
[0015] The search image and the target image are respectively input into a pre-trained image feature extraction network to obtain the feature map of the search image and the feature map of the target image.
[0016] The image feature extraction network extracts features from the input image through multiple hidden layers and fuses the image features extracted from different hidden layers to obtain a feature map of the input image.
[0017] Optionally, a cross-correlation operation is performed on the search image feature map and the target image feature map, including:
[0018] The feature map of the search image is filled with feature points.
[0019] A cross-correlation operation is performed between the filled search image feature map and the target image feature map.
[0020] Optionally, feature point filling processing is performed on the search image feature map, including:
[0021] Using preset feature values, feature points are filled around the edges of the search image feature map to obtain the filled search image feature map.
[0022] Optionally, feature points are filled around the edges of the search image feature map using preset feature values to obtain a filled search image feature map, including:
[0023] Using preset feature values, feature points of a first thickness are filled on both sides of the length direction of the search image feature map, and feature points of a second thickness are filled on both sides of the width direction of the search image feature map, to obtain the filled search image feature map.
[0024] Wherein, the first thickness is half the length of the target image feature map, and the second thickness is half the width of the target image feature map.
[0025] Optionally, the preset feature value is the average feature value of each feature point in the search image feature map.
[0026] Optional, also includes:
[0027] The target image is cropped to obtain a target sub-image;
[0028] Obtain the target sub-image feature map corresponding to the target sub-image;
[0029] A cross-correlation operation is performed on the feature map of the search image and the feature map of the target sub-image, and the cross-correlation score map obtained by the cross-correlation operation is resized to obtain a second score map, the size of which is the same as the size of the search image;
[0030] Based on the second score map, the second image region with the highest correlation to the target sub-image is determined from the first image region, and is used as the image region that matches the target image.
[0031] Optionally, the cross-correlation score map is resized to obtain a first score map, including:
[0032] The first score map is obtained by performing bicubic interpolation on the cross-correlation score map.
[0033] A second aspect of this application provides an image matching apparatus, comprising:
[0034] The acquisition module is used to acquire the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, respectively.
[0035] The cross-correlation module is used to perform cross-correlation operation on the feature map of the search image and the feature map of the target image to obtain a cross-correlation score map, and to perform size transformation on the cross-correlation score map to obtain a first score map, wherein the size of the first score map is the same as the size of the search image;
[0036] The determination module is used to determine, based on the first score map, a first image region in the search image that has the highest relevance to the target image, as a matching region that matches the target image.
[0037] Optionally, the target image includes multiple images of different sizes obtained by cropping the original target image; correspondingly, when performing a cross-correlation operation on the search image feature map and the target image feature map to obtain a cross-correlation score map, the cross-correlation module is specifically used for:
[0038] The search image feature map is cross-correlated with the target image feature map of different sizes to obtain the cross-correlation sub-score map.
[0039] The cross-correlation sub-score maps are fused to obtain the cross-correlation score map.
[0040] Optionally, when acquiring the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, the acquisition module is specifically used for:
[0041] The search image and the target image are respectively input into a pre-trained image feature extraction network to obtain the feature map of the search image and the feature map of the target image.
[0042] The image feature extraction network extracts features from the input image through multiple hidden layers and fuses the image features extracted from different hidden layers to obtain a feature map of the input image.
[0043] Optionally, when performing a cross-correlation operation on the search image feature map and the target image feature map, the cross-correlation module is specifically used for:
[0044] The feature map of the search image is filled with feature points.
[0045] A cross-correlation operation is performed between the filled search image feature map and the target image feature map.
[0046] Optionally, when performing feature point filling processing on the search image feature map, the cross-correlation module is specifically used for:
[0047] Using preset feature values, feature points are filled around the edges of the search image feature map to obtain the filled search image feature map.
[0048] Optionally, when filling feature points around the edges of the search image feature map using preset feature values to obtain the filled search image feature map, the cross-correlation module is specifically used for:
[0049] Using preset feature values, feature points of a first thickness are filled on both sides of the length direction of the search image feature map, and feature points of a second thickness are filled on both sides of the width direction of the search image feature map, to obtain the filled search image feature map.
[0050] Wherein, the first thickness is half the length of the target image feature map, and the second thickness is half the width of the target image feature map.
[0051] Optionally, the preset feature value is the average feature value of each feature point in the search image feature map.
[0052] Optionally, the device further includes an iteration module, which is specifically used for:
[0053] The target image is cropped to obtain a target sub-image;
[0054] Obtain the target sub-image feature map corresponding to the target sub-image;
[0055] A cross-correlation operation is performed on the feature map of the search image and the feature map of the target sub-image, and the cross-correlation score map obtained by the cross-correlation operation is resized to obtain a second score map, the size of which is the same as the size of the search image;
[0056] Based on the second score map, the second image region with the highest correlation to the target sub-image is determined from the first image region, and is used as the image region that matches the target image.
[0057] Optionally, when performing a size transformation on the cross-correlation score map to obtain the first score map, the cross-correlation module is specifically used for:
[0058] The first score map is obtained by performing bicubic interpolation on the cross-correlation score map.
[0059] A third aspect of this application provides an electronic device, comprising:
[0060] A processor, and a memory connected to the processor;
[0061] The memory is used to store computer programs;
[0062] The processor is used to invoke and execute the computer program in the memory to perform the image matching method as described in the first aspect of this application.
[0063] A fourth aspect of this application provides a storage medium storing a computer program that, when executed by a processor, implements the steps of the image matching method as described in the first aspect of this application.
[0064] The technical solution provided in this application may include the following beneficial effects:
[0065] In this application, the scheme first obtains the search image feature map corresponding to the search image and the target image feature map corresponding to the target image. Then, a cross-correlation operation is performed on the search image feature map and the target image feature map to obtain a cross-correlation score map. The cross-correlation score map is then resized to obtain a first score map, which is the similarity score map between the search image feature map and the target image feature map. A higher score indicates a higher similarity. Since the size of the first score map is the same as the size of the search image, the position of each score in the first score map corresponds to its position in the search image. Correspondingly, the first target region in the search image with the highest correlation to the target image is also the region with the highest score in the first score map. Thus, based on the first score map, the first image region with the highest correlation to the target image can be directly determined from the search image and used as the matching region to match the target image. This achieves accurate localization of the image region in the search image that matches the target image, improving image matching efficiency. Attached Figure Description
[0066] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0067] Figure 1 This is a flowchart of an image matching method provided in one embodiment of this application.
[0068] Figure 2 This is a schematic diagram of a Siamese network structure provided in one embodiment of this application.
[0069] Figure 3 This is a schematic diagram of edge filling of a search image feature map x according to an embodiment of this application.
[0070] Figure 4 This is a schematic diagram of edge filling of a search image feature map x provided in another embodiment of this application.
[0071] Figure 5 This is a schematic diagram of the structure of an image matching device provided in one embodiment of this application.
[0072] Figure 6 This is a structural block diagram of an electronic device provided in one embodiment of this application. Detailed Implementation
[0073] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other implementation methods obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0074] Current mainstream image matching algorithms mainly fall into three categories: template matching, feature matching, and deep learning-based target finding. Traditional template matching is pixel-based, similar to 2D convolution operations. It slides the target image across the search image and compares each position, finally returning a grayscale image where each pixel value represents the degree of matching between that region and the template. Feature matching uses traditional algorithms such as Scale-Invariant Feature Transform (SIFT) and Functional Link Neural Networks (FLANN) to describe the region surrounding the feature, thereby finding the same feature in other images. Deep learning-based target finding utilizes deep learning object detection to locate the target, training a model through a network to identify the desired target.
[0075] Extensive research has led to the development of various algorithms in image matching, each with its own advantages and disadvantages. Traditional template matching, while simple, has limitations; changes in the rotation or size of the target object can affect the matching results, compromising accuracy. Feature matching, using traditional algorithms like SIFT and FLANN, typically involves high computational costs and struggles to achieve usable matching accuracy. Deep learning-based approaches have become mainstream in recent years. While deep learning-based object detection can achieve good matching results in certain specific situations, it places high demands on the construction of training data, and transferability is difficult to guarantee; trained models often only apply to specific scenarios.
[0076] Based on this, embodiments of this application provide an image matching method, such as... Figure 1 As shown, the image matching method includes at least the following implementation steps:
[0077] S101. Obtain the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, respectively.
[0078] In practice, when performing image matching between the search image and the target image, feature extraction can be performed on both images first. Specifically, feature extraction of the search image yields a search image feature map, and feature extraction of the target image yields a target image feature map.
[0079] To ensure matching accuracy, the same feature extraction method can be used to extract features from both the search image and the target image. For example, the same ResNet network can be used as the feature extraction network. The search image is input into this network to obtain the search image feature map; the target image is input into the network to obtain the target image feature map.
[0080] S102. Perform a cross-correlation operation on the feature map of the search image and the feature map of the target image to obtain a cross-correlation score map, and perform a size transformation on the cross-correlation score map to obtain a first score map. The size of the first score map is the same as the size of the search image.
[0081] Specifically, image matching methods can be implemented based on Siamese networks. Its structure is as follows: Figure 2 As shown, the target image (query) and the search image (gallery) are respectively input into the same feature extraction network. In the process, the query passes through a feature extraction network. The target image feature map z is then obtained, and the gallery image feature map x is obtained after passing through a feature extraction network. A cross-correlation operation is then performed on the target image feature map z and the search image feature map x to obtain the cross-correlation score map f.
[0082] The cross-correlation score map f is a score map representing the feature similarity between corresponding feature points of the target image feature map z and the search image feature map x. In the cross-correlation score map f, the higher the score of a feature point, the higher the similarity between the target image and the search image at the image location corresponding to that feature point, that is, the higher the matching degree of the feature region corresponding to that location.
[0083] After obtaining the cross-correlation score map, the cross-correlation score map can be resized so that the resized score map is the same size as the search image. This allows for a better correspondence between the score of each position in the cross-correlation score map and the position in the search image, thus quickly determining the matching region of the target image.
[0084] In practice, the size of the cross-correlation score map is generally smaller than the size of the search image. To make the size of the cross-correlation score map the same as the size of the search image, it can be upsampled to enlarge it to the same size as the search image. There are various upsampling methods, such as interpolation and deconvolution.
[0085] Preferably, when performing a size transformation on the cross-correlation score map to obtain the first score map, bicubic interpolation can be performed on the cross-correlation score map to obtain the first score map. Using the bicubic interpolation algorithm for upsampling can enlarge the image without adding image information, ensuring that the first score map and the search image can be aligned, so that the score at each position in the first score map is the matching degree score between the target image and the corresponding image region in the search image.
[0086] S103. Based on the first score map, determine the first image region with the highest relevance to the target image from the search images, and use it as the image region that matches the target image.
[0087] Since the size of the first score image is the same as the size of the search image, the score at each position in the first score image can represent the matching degree score between the target image and the corresponding image region in the search image. Therefore, after aligning the first score image with the search image, the positions in the first score image correspond one-to-one with the positions in the search image; that is, the position of a point in the first score image is also the position of that point in the search image. Thus, the point with the highest value in the first score image is the point with the highest matching degree, and the region in the search image corresponding to this point is the region with the highest relevance to the target image, i.e., the first image region. This first image region can be used as the image region in the search image that matches the target image.
[0088] In this embodiment, firstly, the search image feature map corresponding to the search image and the target image feature map corresponding to the target image are obtained. Then, a cross-correlation operation is performed on the search image feature map and the target image feature map to obtain a cross-correlation score map. The cross-correlation score map is then resized to obtain a first score map, which is the similarity score map between the search image feature map and the target image feature map. A higher score indicates a higher similarity. Since the size of the first score map is the same as the size of the search image, the position of each score in the first score map corresponds to its position in the search image. Correspondingly, the first target region in the search image with the highest correlation to the target image is also the region with the highest score in the first score map. Thus, based on the first score map, the first image region with the highest correlation to the target image can be directly determined from the search image and used as the matching region to match the target image, achieving accurate localization of the image region in the search image that matches the target image.
[0089] To improve matching performance, in some embodiments, target images of multiple sizes can be selected for matching. That is, the target images mentioned above may include multiple images of different sizes obtained by cropping the original target image.
[0090] Accordingly, when performing cross-correlation operations on the search image feature map and the target image feature map to obtain a cross-correlation score map, the search image feature map can be cross-correlated with target image feature maps of different sizes to obtain each cross-correlation sub-score map; then, the cross-correlation sub-score maps can be fused to obtain the cross-correlation score map.
[0091] When cropping the original target image to obtain multiple target images of different sizes, the cropping size of the target image mainly considers the following points: First, it is necessary to ensure that the size of the target image after feature extraction is odd, so that the size of the target image feature map obtained after feature extraction is also odd, that is, the target image feature map has a center point. This ensures that after the target image feature map and the search image feature map are cross-correlated, the size of the cross-correlation score map is consistent with the size of the search image feature map, which facilitates the subsequent alignment of the first score map with the search image. Second, it is necessary to ensure that the size of the feature maps of multiple target images of different sizes are adjacent odd sizes to improve the matching effect. For example, if there are three target images with sizes of 141*141, 161*161 and 199*199, the size of the target image feature maps obtained after feature extraction are 9*9, 11*11 and 13*13, where 9, 11 and 13 are adjacent odd numbers, and 9*9, 11*11 and 13*13 are adjacent odd sizes.
[0092] It's important to understand that convolution kernels are generally set to an odd number of pixels. This serves two purposes: first, to ensure that the anchor point is exactly in the center, facilitating sliding convolution with the center pixel as the standard and preventing positional information from shifting; and second, to ensure padding by adding extra zero layers between images, making the two sides of the image symmetrical so that the output image is the same size as the input.
[0093] Of course, the embodiments of this application only take the example of the target image having an odd size and multiple target images having consecutive adjacent odd sizes. In some other embodiments, the target image may not have an odd size, and similarly, the multiple target images may not have consecutive odd sizes.
[0094] In implementation, a reference point can be first determined from the original target image. Then, based on this reference point, the original target image can be cropped at multiple different sizes, resulting in multiple target images of different sizes. The reference point can be used as the center point of the target image. Cropping the original target image at different sizes using the reference point as the center point yields multiple target images of different odd-numbered sizes. The number of target images can be set according to actual needs and is not limited here.
[0095] After performing cross-correlation operations between the search image feature map and the target image feature maps of different sizes, we can obtain the cross-correlation sub-score maps. To ensure the accuracy of the final cross-correlation score map, the various cross-correlation sub-score maps can be fused to obtain the final cross-correlation score map.
[0096] In implementation, the score maps of each cross-correlation sub-map are merged, which can be achieved by superimposing the scores from each cross-correlation sub-map. Specifically, based on each score position, the scores from different cross-correlation sub-maps are added together to obtain the total score for each position, which is also the score for each position in the cross-correlation score map.
[0097] In some embodiments, to further improve matching accuracy, when acquiring the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, the search image and the target image can be input into a pre-trained image feature extraction network to obtain the search image feature map and the target image feature map, respectively. The image feature extraction network extracts features from the input image through multiple hidden layers and fuses the image features extracted from different hidden layers to obtain the feature map of the input image.
[0098] In implementation, the image feature extraction network can employ the ResNet50 network. Fusing the features extracted from different hidden layers of the ResNet50 network can improve feature extraction capabilities to some extent, thereby increasing matching accuracy. The feature fusion method can be concat (stack). Experiments have shown that concatting the features from the last two layers (conv4_x and conv5_x) of the ResNet50 network yields the best-performing feature map. Therefore, when using the ResNet50 network for feature extraction, concatting the conv4_x and conv5_x layers can extract better-performing target image and search image feature maps.
[0099] In some embodiments, when performing cross-correlation operations on the search image feature map and the target image feature map, feature point filling processing can be performed on the search image feature map first; then, cross-correlation operations can be performed on the filled search image feature map and the target image feature map.
[0100] When performing cross-correlation operations on the search image feature map and the target image feature map, in order to ensure that the cross-correlation scores of feature points at all locations in both the search image feature map and the target image feature map are obtained, and that the size of the cross-correlation score map is consistent with the size of the search image feature map, the center point of the target image feature map needs to be able to slide across all points in the search image feature map. Accordingly, feature point filling processing needs to be performed on the search image feature map so that the size of the cross-correlation score map obtained after the feature point filling operation with the target image feature map is consistent with the size of the search image feature map before feature point filling.
[0101] During implementation, in order to avoid the cross-correlation effect being affected after feature points are filled, preset feature values can be used to fill feature points around the edge of the search image feature map to obtain the filled search image feature map.
[0102] For example, Figure 3 As shown, feature points with a value of 0 can be filled around the edges of the search image feature map x to obtain the filled search image feature map x'.
[0103] By padding the feature maps of the search image, the size of the search image feature map can be increased, ensuring that during cross-correlation operations, the center point of the target image feature map can slide across all points of the search image feature map before padding. Furthermore, using preset feature values for padding can prevent the scores of feature points at the edges of the search image feature map from being underestimated, thus avoiding negative impacts on the cross-correlation performance.
[0104] The preset feature value can be the average feature value of each feature point in the search image feature map, that is, the mean value of the feature points in the search image feature map.
[0105] For example, Figure 4 As shown, in this embodiment, feature points with values equal to the mean of feature points in the search image feature map x are filled around the edge of the search image feature map x to obtain the filled search image feature map x'.
[0106] Padding the search image feature map with the mean of feature points can effectively improve the cross-correlation operation, and the larger the size of the search image feature map, the more significant the improvement. Of course, the preset feature value is not limited to the mean of feature points in the search image feature map; in some other embodiments, the preset feature value can also be other feature values.
[0107] In specific implementation, feature points are filled around the edge of the search image feature map using preset feature values to obtain the filled search image feature map. This can be achieved by filling feature points of a first thickness on both sides of the length direction of the search image feature map and filling feature points of a second thickness on both sides of the width direction of the search image feature map using preset feature values to obtain the filled search image feature map. The first thickness is half the size of the target image feature map in the length direction, and the second thickness is half the size of the target image feature map in the width direction.
[0108] like Figure 4 As shown, filling the search image feature map with feature points of a first thickness on both sides of its length direction and with feature points of a second thickness on both sides of its width direction ensures that when the filled search image feature map is cross-correlated with the target image feature map, the center point of the target image feature map can slide from the edge of the filled search image feature map, thus avoiding the loss of feature information and improving the accuracy of matching.
[0109] On the other hand, when acquiring the feature map of the target image corresponding to the target image, the target image is often smaller than the search image. Experiments have shown that using a cropped, smaller target image for matching results in higher accuracy. However, when the target image is too small, there is less feature information, which may lead to multiple image regions in the search image that are highly similar to the target image, making it impossible to determine the correct image region.
[0110] To further improve matching accuracy, a preliminary matching can be performed using a large target image to determine the first image region. Based on this, the target image is then cropped to obtain a target sub-image; subsequently, the feature map corresponding to the target sub-image is obtained; then, a cross-correlation operation is performed between the search image feature map and the target sub-image feature map, and the size of the cross-correlation score map obtained from the cross-correlation operation is transformed to obtain a second score map, the size of which is the same as the size of the search image; based on the second score map, the second image region with the highest correlation to the target sub-image is determined from the first image region, and this region is taken as the image region that matches the target image.
[0111] By first performing preliminary matching using a large target image, the first image region is determined, which is the range of the image region that matches the target image. Then, based on this range, the target image is cropped to obtain a smaller target sub-image. The smaller target sub-image is then used for matching again to determine the image region that matches the target sub-image from the first image region determined in the previous step. This sub-image is then used as the image region that matches the target image after further filtering.
[0112] Following the above process, the target image is iteratively cropped, and the image region matching the target image is iteratively obtained from the obtained image region. In this way, the situation where a small target image is easily trapped in a region with high local similarity in the search image is avoided to a certain extent, and the matching accuracy is effectively improved.
[0113] Additionally, it's important to understand that in some embodiments, the search image can be the original search image or an image obtained by scaling the original search image proportionally. When the search image is obtained by scaling the original search image proportionally, the size of the final first score image or second score image needs to be consistent with the size of the original search image. That is, after performing the size transformation on the cross-correlation score image to obtain the first score image, and the size of the first score image is the same as the size of the search image, the first score image also needs to undergo a proportional scaling process that is the opposite of scaling the original search image (if the search image is obtained by reducing the original search image by a factor of K, then the first score image is enlarged by a factor of K; if the search image is obtained by enlarging the original search image by a factor of K, then the first score image is reduced by a factor of K) to make the size of the first score image consistent with the size of the original search image, i.e., to achieve alignment between the first score image and the original search image. Thus, based on the first score image, the first image region with the highest relevance to the target image can be determined from the original search image as the image region that matches the target image.
[0114] Similarly, after obtaining the second score map and the size of the second score map is the same as the size of the search image, the second score map needs to be scaled in the opposite way to make the size of the second score map consistent with the size of the original search image, that is, to achieve alignment between the second score map and the original search image. In this way, based on the second score map, the second image region with the highest relevance to the target sub-image can be determined from the first image region as the image region that matches the target image.
[0115] Corresponding to the image matching method described above, embodiments of this application also disclose an image matching device, such as... Figure 5 As shown, the device may include: an acquisition module 501, used to acquire a search image feature map corresponding to a search image and a target image feature map corresponding to a target image; a cross-correlation module 502, used to perform a cross-correlation operation on the search image feature map and the target image feature map to obtain a cross-correlation score map, and to perform a size transformation on the cross-correlation score map to obtain a first score map, the size of the first score map being the same as the size of the search image; and a determination module 503, used to determine, based on the first score map, a first image region in the search image that has the highest correlation with the target image, as a matching region that matches the target image.
[0116] The image matching apparatus proposed in this application can obtain a search image feature map corresponding to the search image and a target image feature map corresponding to the target image through the acquisition module 501. Then, the cross-correlation module 502 performs a cross-correlation operation on the search image feature map and the target image feature map, and performs a size transformation on the obtained cross-correlation score map to obtain a first score map, making the size of the first score map consistent with that of the search image. Finally, the determination module 503 determines the matching region that matches the target image, realizing the accurate positioning of the image region in the search image that matches the target image.
[0117] Optionally, the target image includes multiple images of different sizes obtained by cropping the original target image; correspondingly, when performing a cross-correlation operation on the search image feature map and the target image feature map to obtain a cross-correlation score map, the cross-correlation module 502 is specifically used to: perform a cross-correlation operation on the search image feature map with each target image feature map of different sizes to obtain each cross-correlation sub-score map; and fuse each cross-correlation sub-score map to obtain a cross-correlation score map.
[0118] Optionally, when acquiring the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, the acquisition module 501 is specifically used to: input the search image and the target image into a pre-trained image feature extraction network to obtain the search image feature map and the target image feature map; wherein, the image feature extraction network extracts features of the input image through multiple hidden layers and fuses the image features extracted by different hidden layers to obtain the feature map of the input image.
[0119] Optionally, when performing cross-correlation operations on the search image feature map and the target image feature map, the cross-correlation module 502 is specifically used for: performing feature point filling processing on the search image feature map; and performing cross-correlation operations on the filled search image feature map and the target image feature map.
[0120] Optionally, when performing feature point filling processing on the search image feature map, the cross-correlation module 502 is specifically used to: fill feature points around the edge of the search image feature map using preset feature values to obtain the filled search image feature map.
[0121] Optionally, when filling feature points around the edge of the search image feature map using preset feature values to obtain the filled search image feature map, the cross-correlation module 502 is specifically used to: fill feature points of a first thickness on both sides of the length direction of the search image feature map using preset feature values, and fill feature points of a second thickness on both sides of the width direction of the search image feature map to obtain the filled search image feature map; wherein, the first thickness is half the size of the target image feature map in the length direction, and the second thickness is half the size of the target image feature map in the width direction.
[0122] The preset feature value is the average feature value of each feature point in the search image feature map.
[0123] Optionally, the image matching device may further include an iteration module, which is specifically used for: cropping the target image to obtain a target sub-image; obtaining a target sub-image feature map corresponding to the target sub-image; performing a cross-correlation operation on the search image feature map and the target sub-image feature map, and performing a size transformation on the cross-correlation score map obtained by the cross-correlation operation to obtain a second score map, the size of the second score map being the same as the size of the search image; and determining, based on the second score map, a second image region with the highest correlation to the target sub-image from the first image region as the image region that matches the target image.
[0124] Optionally, when performing a size transformation on the cross-correlation score map to obtain the first score map, the cross-correlation module 502 is specifically used to: perform bicubic interpolation on the cross-correlation score map to obtain the first score map.
[0125] It should be understood that the specific implementation of the image matching device provided in the embodiments of this application can refer to the specific implementation of the image matching method described in the corresponding embodiments above, and will not be repeated here.
[0126] Figure 6 The diagram shown is a block diagram of an electronic device 600 for performing an image matching method according to an exemplary embodiment of this application.
[0127] Reference Figure 6 The electronic device 600 includes a processing component 601, which further includes one or more processors, and memory resources represented by memory 602 for storing instructions, such as application programs, that can be executed by the processing component 601. The application programs stored in memory 602 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 601 is configured to execute instructions to perform the image matching method described in any of the above embodiments.
[0128] Electronic device 600 may also include a power supply component configured to perform power management of electronic device 600, a wired or wireless network interface configured to connect electronic device 600 to a network, and an input / output (I / O) interface. Electronic device 600 may operate based on an operating system stored in memory 602, such as Windows Server™, Mac OSX™, Unix™, Linux™, FreeBSD™, or similar.
[0129] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device 600, enables the electronic device 600 to perform any of the image matching methods described in the above embodiments. The image matching method includes: acquiring a search image feature map corresponding to a search image and a target image feature map corresponding to a target image; performing a cross-correlation operation on the search image feature map and the target image feature map to obtain a cross-correlation score map, and performing a size transformation on the cross-correlation score map to obtain a first score map, the size of the first score map being the same as the size of the search image; and determining, based on the first score map, a first image region from the search image with the highest correlation to the target image, as the image region matching the target image.
[0130] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0131] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0132] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0133] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0135] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0136] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program verification codes, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0137] It should be noted that in the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0138] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications or equivalent substitutions made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An image matching method, characterized in that, include: Obtain the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, respectively; The target image includes multiple images of different sizes obtained by cropping the original target image; The cropping process includes: determining reference points in the original target image; cropping the original target image to multiple different sizes based on the reference points to obtain multiple target images of different sizes; A cross-correlation operation is performed on the feature map of the search image and the feature map of the target image to obtain a cross-correlation score map. The cross-correlation score map is then resized to obtain a first score map, the size of which is the same as the size of the search image. Specifically, a cross-correlation operation is performed between the search image feature map and the target image feature map to obtain a cross-correlation score map, including: The search image feature map is cross-correlated with the target image feature map of different sizes to obtain the cross-correlation sub-score map. The cross-correlation sub-score maps are fused to obtain the cross-correlation score map; Based on the first score map, the first image region with the highest relevance to the target image is determined from the search image, and is used as the image region that matches the target image.
2. The method according to claim 1, characterized in that, Obtain the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, including: The search image and the target image are respectively input into a pre-trained image feature extraction network to obtain the feature map of the search image and the feature map of the target image. The image feature extraction network extracts features from the input image through multiple hidden layers and fuses the image features extracted from different hidden layers to obtain a feature map of the input image.
3. The method according to claim 1, characterized in that, Performing a cross-correlation operation between the search image feature map and the target image feature map includes: The feature map of the search image is filled with feature points. A cross-correlation operation is performed between the filled search image feature map and the target image feature map.
4. The method according to claim 3, characterized in that, The feature map of the search image is filled with feature points, including: Using preset feature values, feature points are filled around the edges of the search image feature map to obtain the filled search image feature map.
5. The method according to claim 4, characterized in that, Using preset feature values, feature points are filled around the edges of the search image feature map to obtain the filled search image feature map, including: Using preset feature values, feature points of a first thickness are filled on both sides of the length direction of the search image feature map, and feature points of a second thickness are filled on both sides of the width direction of the search image feature map, to obtain the filled search image feature map. Wherein, the first thickness is half the length of the target image feature map, and the second thickness is half the width of the target image feature map.
6. The method according to claim 4 or 5, characterized in that, The preset feature value is the average feature value of each feature point in the search image feature map.
7. The method according to claim 1, characterized in that, Also includes: The target image is cropped to obtain a target sub-image; Obtain the target sub-image feature map corresponding to the target sub-image; A cross-correlation operation is performed on the feature map of the search image and the feature map of the target sub-image, and the cross-correlation score map obtained by the cross-correlation operation is resized to obtain a second score map, the size of which is the same as the size of the search image; Based on the second score map, the second image region with the highest correlation to the target sub-image is determined from the first image region, and is used as the image region that matches the target image.
8. The method according to claim 1, characterized in that, The cross-correlation score map is resized to obtain a first score map, including: The first score map is obtained by performing bicubic interpolation on the cross-correlation score map.
9. An image matching device, characterized in that, include: The acquisition module is used to acquire the search image feature map corresponding to the search image and the target image feature map corresponding to the target image, respectively. The target image includes multiple images of different sizes obtained by cropping the original target image; The cropping process includes: determining reference points in the original target image; cropping the original target image to multiple different sizes based on the reference points to obtain multiple target images of different sizes; The cross-correlation module is used to perform cross-correlation operation on the feature map of the search image and the feature map of the target image to obtain a cross-correlation score map, and to perform size transformation on the cross-correlation score map to obtain a first score map, wherein the size of the first score map is the same as the size of the search image; Specifically, a cross-correlation operation is performed between the search image feature map and the target image feature map to obtain a cross-correlation score map, including: The search image feature map is cross-correlated with the target image feature map of different sizes to obtain the cross-correlation sub-score map. The cross-correlation sub-score maps are fused to obtain the cross-correlation score map; The determining module is used to determine, based on the first score map, a first image region with the highest relevance to the target image from the search image, as the image region that matches the target image.
10. An electronic device, characterized in that, include: A processor, and a memory connected to the processor; The memory is used to store computer programs; The processor is used to call and execute the computer program in the memory to perform the image matching method as described in any one of claims 1-8.
11. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the image matching method as described in any one of claims 1-8.
Citation Information
Patent Citations
Tracking method and related product
CN110503662A
Twin double-path target tracking method
CN111260688A
Image matching method and device, computer equipment and storage medium
CN111666974A
Target tracking method based on SIAM-FC network
CN112767440A
Image matching method and device, storage medium and electronic device
CN114548218A