Image matching method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202111667861.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-12-31
AI Technical Summary
[0003]相关技术中的图像匹配方法主要包括基于特征的匹配方法和基于模板匹配的方法,其中,基于特征的匹配方法常提取关键点周围一定邻域内的局部特征信息描述符,通过比较描述符来确定匹配点,常使用尺度不变特征转换(scale invariant featuretransformation,SIFT)描述符,但由于SIFT描述符是基于影像局部邻域的梯度分布描述关键点,导致对异源影像的匹配精度较低;基于模板匹配的方法一般基于图像的边缘线提取或图像的互相关信息,如果边缘提取不够准确则影响图像的匹配精度,而基于图像互相关的方法由于SAR图像和可见光图像的成像机理差距大,导致有时图像匹配精度偏低
[0040]根据本公开的图像匹配方法、装置、电子设备及存储介质,一方面,通过第一、第二边缘线强度图预测模型分别获取第一边缘线强度图和第二边缘线强度图,可提高图像的边缘线提取的准确性,从而提高基于边缘线强度图进行图像匹配的匹配精度;另一方面,通过图像匹配模型对模板图像和待匹配图像进行匹配,可改善基于边缘线提取的方法进行模板匹配时由于信息丢失造成的匹配精度偏低的问题,从而提高图像匹配的匹配精度。更进一步,通过将这两方面的匹配结果结合而得到最终的匹配位置,可弥补二者之间可能存在的不足,从而在极大程度上提高图像匹配的匹配精度。
Smart Images

Figure CN116416263B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image matching, and more specifically, to an image matching method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the rapid development of remote sensing technology, Earth observation images from various types of sensors, such as visible light, infrared, and synthetic aperture radar (SAR), are becoming increasingly abundant. Images from different platforms and sensors have a certain degree of complementarity, providing a massive source of data for in-depth mining of remote sensing information and big data analysis. Among these, image matching is the core issue for further processing of heterogeneous images.
[0003] Image matching methods in related technologies mainly include feature-based matching methods and template-based matching methods. Feature-based matching methods often extract local feature information descriptors within a certain neighborhood around key points and determine matching points by comparing descriptors. Scale-invariant feature transformation (SIFT) descriptors are commonly used. However, since SIFT descriptors describe key points based on the gradient distribution of the local neighborhood of the image, the matching accuracy for images from different sources is relatively low. Template-based matching methods generally rely on edge line extraction or cross-correlation of images. If the edge extraction is not accurate enough, it will affect the image matching accuracy. Furthermore, image cross-correlation-based methods sometimes have low image matching accuracy due to the large difference in imaging mechanisms between SAR images and visible light images. Summary of the Invention
[0004] This disclosure provides an image matching method, apparatus, electronic device, and storage medium to at least solve the problems in the aforementioned related technologies.
[0005] According to a first aspect of the present disclosure, an image matching method is provided, comprising: acquiring an image to be matched and a template image; obtaining a first edge line intensity map based on the image to be matched using a first edge line intensity map prediction model; obtaining a second edge line intensity map based on the template image using a second edge line intensity map prediction model; performing template matching on the first edge line intensity map using the second edge line intensity map as a template to obtain a first matching result; acquiring a target image multiple times from the image to be matched, wherein the target image and the template image have the same size; matching the template image and the image to be matched using an image matching model based on the acquired target image and the template image to obtain a second matching result; and obtaining the matching position between the image to be matched and the template image based on the first matching result and the second matching result.
[0006] Optionally, obtaining a first edge line intensity map based on the image to be matched using a first edge line intensity map prediction model includes: obtaining at least one first edge line intensity map based on the image to be matched using at least one first edge line intensity map prediction model; obtaining a second edge line intensity map based on the template image using a second edge line intensity map prediction model includes: obtaining at least one second edge line intensity map based on the template image using at least one second edge line intensity map prediction model.
[0007] Optionally, the step of using the second edge line intensity map as a template to perform template matching on the first edge line intensity map to obtain a first matching result includes: using the at least one second edge line intensity map as a template to perform template matching on the at least one first edge line intensity map to obtain at least one matching result; and obtaining the first matching result based on the at least one matching result.
[0008] Optionally, obtaining a first edge line intensity map based on the image to be matched using a first edge line intensity map prediction model includes: performing a multi-scale transformation on the image to be matched; obtaining multiple first transformed edge line intensity maps based on the image to be matched at each scale using the first edge line intensity map prediction model; and obtaining the first edge line intensity map based on the multiple first transformed edge line intensity maps. The obtaining a second edge line intensity map based on the template image using a second edge line intensity map prediction model includes: performing a multi-scale transformation on the template image, wherein the multi-scale transformation performed on the template image is of the same transformation type as the multi-scale transformation performed on the image to be matched; obtaining multiple second transformed edge line intensity maps based on the template image at each scale using the second edge line intensity map prediction model; and obtaining the second edge line intensity map based on the multiple second transformed edge line intensity maps.
[0009] Optionally, the training process of the first edge line intensity map prediction model includes: acquiring a first training dataset, wherein the first training dataset includes a first image and a real edge line intensity map corresponding to the first image, and the type of the first image is the same as the type of the image to be matched; obtaining a first estimated map based on the first image using the first edge line intensity map prediction model; calculating a first loss based on the first estimated map and the real edge line intensity map; and training the first edge line intensity map prediction model by adjusting the model parameters of the first edge line intensity map prediction model according to the first loss.
[0010] Optionally, the first loss includes cross-entropy loss, cross-correlation loss, and correlation loss, and the first loss is expressed as:
[0011] L=w1FL+w2CCORRLOSS+w3CCOEFFLOSS
[0012] Wherein, FL represents the cross-entropy loss; CCORRLOSS represents the cross-correlation loss; CCOEFFLOSS represents the correlation loss; and w1, W2, and W3 represent different weights.
[0013] Optionally, after training the first edge intensity map prediction model, the method further includes: acquiring a second training dataset, wherein the second training dataset includes a second image, a true edge intensity map corresponding to the second image, and an erroneous edge intensity map, the type of the second image being the same as the type of the image to be matched, and the erroneous edge intensity map being an edge intensity map that deviates from the true matching position when performing template matching on the first edge intensity map output by the trained first edge intensity map prediction model; obtaining a second estimated map based on the second image using the first edge intensity map prediction model; calculating a second loss based on the second estimated map, the true edge intensity map corresponding to the second image, and the erroneous edge intensity map; and training the first edge intensity map prediction model by adjusting the model parameters of the first edge intensity map prediction model according to the second loss.
[0014] Optionally, the second loss is expressed as:
[0015] HEMLOSS=2+CCOEFF NORMED neg -CCOEFF NORMED pos
[0016] Among them, CCOEFF NORMED neg This represents the correlation coefficient between the second estimated map and the erroneous edge line intensity map; CCOEFF NORMED pos This represents the correlation coefficient between the second estimated map and the actual edge line intensity map corresponding to the second image.
[0017] Optionally, the training process of the image matching model includes: acquiring a third training dataset, wherein the third training dataset includes multiple third images, fourth images, positive sample pairs, and negative sample pairs, the type of the third images is the same as the type of the image to be matched, the type of the fourth images is the same as the type of the template image, each third image and each fourth image has the same size, the positive sample pairs are mutually matched third and fourth images, and the negative sample pairs are mutually unmatched third and fourth images; obtaining an estimated matching probability based on any third and fourth images through the image matching model; calculating a third loss based on the estimated matching probability, the label values of the positive sample pairs, and the label values of the negative sample pairs; and training the image matching model by adjusting the model parameters of the image matching model according to the third loss.
[0018] Optionally, the third loss is expressed as:
[0019] FL(p)=-|yp| β ((1-y)log(1-p)+ylog(p))
[0020] Where y represents the label value of the positive sample pair or the label value of the negative sample pair; p represents the estimated matching probability; -|yp| β This indicates dynamic weights.
[0021] According to a second aspect of the present disclosure, an image matching apparatus is provided, comprising: an image acquisition unit configured to acquire an image to be matched and a template image; a first edge line intensity map acquisition unit configured to obtain a first edge line intensity map based on the image to be matched using a first edge line intensity map prediction model; a second edge line intensity map acquisition unit configured to obtain a second edge line intensity map based on the template image using a second edge line intensity map prediction model; a first matching result determination unit configured to perform template matching on the first edge line intensity map using the second edge line intensity map as a template to obtain a first matching result; a target image acquisition unit configured to acquire a target image multiple times from the image to be matched, wherein the target image and the template image have the same size; a second matching result determination unit configured to match the template image and the image to be matched based on the acquired target image and the template image using an image matching model to obtain a second matching result; and a matching position determination unit configured to determine the matching position between the image to be matched and the template image based on the first matching result and the second matching result.
[0022] Optionally, the first edge line intensity map acquisition unit is configured to: obtain at least one first edge line intensity map based on the image to be matched using at least one first edge line intensity map prediction model; the second edge line intensity map acquisition unit is configured to: obtain at least one second edge line intensity map based on the template image using at least one second edge line intensity map prediction model.
[0023] Optionally, the first matching result determination unit is configured to: use the at least one second edge line intensity map as a template to perform template matching on the at least one first edge line intensity map to obtain at least one matching result; and obtain the first matching result based on the at least one matching result.
[0024] Optionally, the first edge line intensity map acquisition unit is configured to: perform multi-scale transformation on the image to be matched; based on the image to be matched at each scale, obtain multiple first transformed edge line intensity maps through the first edge line intensity map prediction model; and obtain the first edge line intensity map based on the multiple first transformed edge line intensity maps. The second edge line intensity map acquisition unit is configured to: perform multi-scale transformation on the template image, wherein the multi-scale transformation performed on the template image is of the same transformation type as the multi-scale transformation performed on the image to be matched; based on the template image at each scale, obtain multiple second transformed edge line intensity maps through the second edge line intensity map prediction model; and obtain the second edge line intensity map based on the multiple second transformed edge line intensity maps.
[0025] Optionally, the training process of the first edge line intensity map prediction model includes: acquiring a first training dataset, wherein the first training dataset includes a first image and a real edge line intensity map corresponding to the first image, and the type of the first image is the same as the type of the image to be matched; obtaining a first estimated map based on the first image using the first edge line intensity map prediction model; calculating a first loss based on the first estimated map and the real edge line intensity map; and training the first edge line intensity map prediction model by adjusting the model parameters of the first edge line intensity map prediction model according to the first loss.
[0026] Optionally, the first loss includes cross-entropy loss, cross-correlation loss, and correlation loss, and the first loss is expressed as:
[0027] L=w1FL+w2CCORRLOSS+w3CCOEFFLOSS
[0028] Wherein, FL represents the cross-entropy loss; CCORRLOSS represents the cross-correlation loss; CCOEFFLOSS represents the correlation loss; and w1, W2, and W3 represent different weights.
[0029] Optionally, after training the first edge intensity map prediction model, the method further includes: acquiring a second training dataset, wherein the second training dataset includes a second image, a true edge intensity map corresponding to the second image, and an erroneous edge intensity map, the type of the second image being the same as the type of the image to be matched, and the erroneous edge intensity map being an edge intensity map that deviates from the true matching position when performing template matching on the first edge intensity map output by the trained first edge intensity map prediction model; obtaining a second estimated map based on the second image using the first edge intensity map prediction model; calculating a second loss based on the second estimated map, the true edge intensity map corresponding to the second image, and the erroneous edge intensity map; and training the first edge intensity map prediction model by adjusting the model parameters of the first edge intensity map prediction model according to the second loss.
[0030] Optionally, the second loss is expressed as:
[0031] HEMLOSS=2+CCOEFF NORMED neg -CCOEFF NORMED pos
[0032] Among them, CCOEFF NORMED neg This represents the correlation coefficient between the second estimated map and the erroneous edge line intensity map; CCOEFF NORMED pos This represents the correlation coefficient between the second estimated map and the actual edge line intensity map corresponding to the second image.
[0033] Optionally, the training process of the image matching model includes: acquiring a third training dataset, wherein the third training dataset includes multiple third images, fourth images, positive sample pairs, and negative sample pairs, the type of the third images is the same as the type of the image to be matched, the type of the fourth images is the same as the type of the template image, each third image and each fourth image has the same size, the positive sample pairs are mutually matched third and fourth images, and the negative sample pairs are mutually unmatched third and fourth images; obtaining an estimated matching probability based on any third and fourth images through the image matching model; calculating a third loss based on the estimated matching probability, the label values of the positive sample pairs, and the label values of the negative sample pairs; and training the image matching model by adjusting the model parameters of the image matching model according to the third loss.
[0034] Optionally, the third loss is expressed as:
[0035] FL(p)=-|yp| β ((1-y)log(1-p)+ylog(p))
[0036] Where y represents the label value of the positive sample pair or the label value of the negative sample pair; p represents the estimated matching probability; -|yp| β This indicates dynamic weights.
[0037] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform an image matching method according to the present disclosure.
[0038] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores instructions which, when executed by at least one processor, cause the at least one processor to perform an image matching method according to the present disclosure.
[0039] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects:
[0040] According to the image matching method, apparatus, electronic device, and storage medium disclosed herein, on the one hand, by obtaining the first edge line intensity map and the second edge line intensity map respectively through the first and second edge line intensity map prediction models, the accuracy of edge line extraction of the image can be improved, thereby improving the matching accuracy of image matching based on edge line intensity maps. On the other hand, by matching the template image and the image to be matched through the image matching model, the problem of low matching accuracy caused by information loss when performing template matching based on edge line extraction methods can be improved, thereby improving the matching accuracy of image matching. Furthermore, by combining the matching results of these two aspects to obtain the final matching position, the potential deficiencies between the two can be compensated for, thereby greatly improving the matching accuracy of image matching.
[0041] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0043] Figure 1 This is a flowchart illustrating an image matching method according to an exemplary embodiment of the present disclosure.
[0044] Figure 2 This is a schematic diagram illustrating a true edge line intensity map and an erroneous edge line intensity map of a visible light image according to an exemplary embodiment of the present disclosure.
[0045] Figure 3 This is a schematic diagram illustrating the construction process of training data for first and second edge line intensity map prediction models according to exemplary embodiments of the present disclosure.
[0046] Figure 4 This is a schematic diagram illustrating the process of matching SAR images and visible light images based on first and second edge line intensity map prediction models according to an exemplary embodiment of the present disclosure.
[0047] Figure 5 This is a schematic diagram illustrating the process of matching SAR images and visible light images based on an image allocation model according to an exemplary embodiment of the present disclosure.
[0048] Figure 6 This is a block diagram illustrating an image matching apparatus according to an exemplary embodiment of the present disclosure.
[0049] Figure 7 This is a block diagram of an electronic device 700 according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0050] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0051] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0052] It should be noted that the phrase "at least one of several items" in this disclosure refers to three parallel cases: "any one of the several items", "a combination of any number of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. As another example, "performing at least one of step one and step two" indicates the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing both step one and step two.
[0053] To address the issue of low matching accuracy in heterogeneous images, this disclosure proposes an image matching method, apparatus, electronic device, and storage medium. Specifically, on one hand, by obtaining the first and second edge line intensity maps respectively through first and second edge line intensity map prediction models, the accuracy of edge line extraction can be improved, thereby increasing the matching accuracy of image matching based on edge line intensity maps. On the other hand, by matching the template image and the image to be matched through an image matching model, the problem of low matching accuracy caused by information loss when using edge line extraction-based template matching can be mitigated, thus improving the matching accuracy of image matching. Furthermore, by combining the matching results from these two aspects to obtain the final matching position, any potential deficiencies between the two methods can be compensated for, thereby significantly improving the matching accuracy of image matching. The following will refer to... Figures 1 to 7 The present disclosure provides a detailed description of an image matching method, apparatus, electronic device, and storage medium according to exemplary embodiments thereof.
[0054] Figure 1 This is a flowchart illustrating an image matching method according to an exemplary embodiment of the present disclosure.
[0055] Reference Figure 1 In step 101, the image to be matched and the template image can be obtained. Here, the image to be matched and the template image are heterogeneous images (i.e., images acquired by sensors or platforms using different imaging methods), such as visible light images, SAR images, or infrared images, etc., without limitation. The template image refers to the reference template used in image matching, and the image to be matched refers to the object being matched in the image matching process. In specific implementations, the size of the template image is smaller than the size of the image to be matched. The final result of the matching is to find the position and region of the template image in the image to be matched.
[0056] In step 102, a first edge line intensity map can be obtained based on the image to be matched using a first edge line intensity map prediction model.
[0057] In step 103, a second edge line intensity map can be obtained based on the template image using the second edge line intensity map prediction model. Here, an edge line refers to a location in the image where the pixel value changes abruptly, and the edge line intensity refers to the pixel value corresponding to that edge line. By finding all the locations in an image where the pixel value changes abruptly, the edge line intensity map corresponding to that image is formed.
[0058] According to an exemplary embodiment of this disclosure, the first edge line intensity map prediction model may adopt the U-net network structure framework, and its backbone network may adopt Resnest50. Of course, any convolutional neural network framework, such as common image segmentation networks, may also be used, without limitation. The first edge line intensity map prediction model is pre-trained. Here, the training process of the first edge line intensity map prediction model may be as follows: First, a first training dataset is obtained, wherein the first training dataset includes a first image and the real edge line intensity map corresponding to the first image, and the type of the first image is the same as the type of the image to be matched. In the specific implementation process, to obtain better training results, multiple first images are used. Traditional unsupervised edge extraction methods (such as phase consistency algorithms) can be employed to extract the edge lines of the first images. After obtaining the edge lines, to avoid inaccurate edge lines due to thresholding, the edge lines are not divided into binary values of 0 or 1 based on the threshold. Furthermore, to reduce the training difficulty of the first edge line intensity map prediction model, the pixel values of the extracted edge lines can be normalized to between 0 and 1, thereby obtaining the true edge line intensity map, which is used as the label during the training process. Then, based on the first image, the first edge line intensity map prediction model can be used to obtain the first estimated map (i.e., the estimated edge line intensity map corresponding to the first image). Here, to prevent overfitting, random data augmentation (such as random scale changes, random translations, random rotations, and random flips) can be applied to the first image and its corresponding label. The first image after data augmentation is then input into the first edge line intensity map prediction model to obtain the first estimated map. Of course, the first image can also be directly input into the first edge line intensity map prediction model; there are no restrictions on this. Subsequently, a first loss can be calculated based on the first estimated map and the true edge line intensity map. Here, to improve the prediction accuracy of the first edge line intensity map prediction model, the first loss may include cross-entropy loss, cross-correlation loss, and correlation loss. In some embodiments, the first loss can be represented by a first loss function, which, for example, but not limited to, can be expressed as:
[0059] L=w1FL+w2CCORRLOSS+w3CCOEFFLOSS
[0060] Wherein, FL represents the cross-entropy loss function (i.e., represents the cross-entropy loss); CCORRLOSS represents the cross-correlation loss function (i.e., represents the cross-correlation loss); CCOEFFLOSS represents the correlation loss function (i.e., represents the correlation loss); w1, w2, and w3 represent different weights. Here, the cross-entropy loss function is used to calculate the cross-entropy loss between the first estimated image and its corresponding label pixel by pixel. Since the label contains a large number of non-edge pixels (i.e., edge intensity is 0), the cross-entropy loss function can be dynamically weighted. In some embodiments, the cross-entropy loss function, for example, but not limited to, can be expressed as:
[0061] FL(x)=-|x′-x| β ((1-x′)log(1-x)+x′log(x))
[0062] Where x represents the first estimated map, x′ represents the label corresponding to the first estimated map, and -|x′-x| β This indicates dynamic weights.
[0063] The cross-correlation loss function measures the normalized cross-correlation between the first estimated map and its corresponding label. In some embodiments, the cross-correlation loss function, for example, but not limited to, may be expressed as:
[0064] CCORRLOSS = 1 - CCORR NORMED
[0065] Among them, CCOEFF NORMED The normalized cross-correlation coefficient between the first estimated plot and its corresponding label is expressed in the relevant techniques and will not be repeated here.
[0066] The correlation loss function measures the normalized correlation between the first estimated map and its corresponding label. In some embodiments, the correlation loss function, for example, but not limited to, may be expressed as:
[0067] CCOEFFLOSS = 1 - CCOEFF NORMED
[0068] Among them, CCOEFF NORMED The normalized cross-correlation coefficient between the first estimated map and its corresponding label is expressed in related technologies and will not be repeated here. Furthermore, this disclosure does not limit the first loss function; any possible loss function (e.g., determining the squared error, absolute error, etc. between the first estimated map and its corresponding label pixel-by-pixel) can be used to train the first edge line intensity map prediction model of this disclosure.
[0069] Finally, the edge line intensity map prediction model corresponding to the image to be matched can be trained by adjusting the model parameters of the model model based on the first loss. In some embodiments, the model parameters of the first edge line intensity map prediction model can be adjusted by backpropagation of the first loss calculated by the first loss function. Furthermore, during model training, a batch of first images can be used to adjust (or update) the model parameters of the first edge line intensity map prediction model, and the model parameters of the first edge line intensity map prediction model are iteratively adjusted (or updated) with the goal of minimizing the value of the first loss function until the first edge line intensity map prediction model converges.
[0070] According to an exemplary embodiment of this disclosure, after training the first edge line intensity map prediction model, to further improve the prediction accuracy of the model, it can be fine-tuned, that is, the model can be trained a second time. Specifically, the second training process is as follows: First, a second training dataset is obtained, wherein the second training dataset includes a second image, the true edge line intensity map corresponding to the second image, and erroneous edge line intensity maps. Here, the type of the second image is the same as the type of the image to be matched, and the erroneous edge line intensity map is the edge line intensity map obtained when performing template matching on the edge line intensity map output by the trained first edge line intensity map prediction model, which deviates from the true matching position. Specifically, after the first edge line intensity map prediction model is trained, the first edge line intensity map obtained by the trained first edge line intensity map prediction model is used for image matching tests. During the test, there may be a situation where the matched edge line intensity map is not the edge line intensity map corresponding to the true template image in the image to be matched, for example... Figure 2 This is a schematic diagram illustrating a true edge intensity map and an erroneous edge intensity map of a visible light image according to an exemplary embodiment of the present disclosure, with reference to... Figure 2 , Figure 2 (a) is the edge line intensity map of the visible light image (the image to be matched). Figure 2 (b) is the edge line intensity map of the SAR image (template image). Figure 2 (a) The area defined by the middle frame line A is Figure 2 (b) in Figure 2 In (a), the corresponding actual matching region is defined by box B, which is the region obtained during the actual image matching test. Figure 2 (b) in Figure 2 The edge line intensity map corresponding to (a) is incorrect, and there is a discrepancy between the two.
[0071] Then, based on the second image, a second estimated image can be obtained using the trained first edge intensity map prediction model. Similar to the aforementioned training process, to prevent overfitting, random data augmentation (e.g., random scaling, random translation, random rotation, random flipping, color space transformation, image sharpening, or image blurring) can be applied to the second image, the corresponding true edge intensity map, and the erroneous edge intensity map. The data-augmented first image is then input into the first edge intensity map prediction model to obtain the second estimated image. Afterward, a second loss can be calculated based on the second estimated image, the corresponding true edge intensity map, and the erroneous edge intensity map. The first edge intensity map prediction model is then trained by adjusting its parameters according to the second loss. Here, the second loss can be represented by a second loss function, which, for example, but not limited to, can be expressed as:
[0072] HEMLOSS=2+CCOEFF NORMED neg -CCOEFF NORMED pos
[0073] Among them, CCOEFF NORMED neg This represents the correlation coefficient between the second estimated plot and the erroneous edge line intensity plot; CCOEFF NoRMED pos This represents the correlation coefficient between the second estimated map and the corresponding true edge line intensity map of the second image. Similarly, CCOEFF... NORMED neg and CCOEFF NORMED pos The expression for this has already been demonstrated in relevant technologies and will not be elaborated upon here.
[0074] Since the structure and training process of the second edge line intensity map prediction model are similar to those of the first edge line intensity map prediction model, with the only difference being the training data and training labels, the structure and training process of the second edge line intensity map prediction model will not be described in detail here for the sake of brevity. It should be noted that after the second edge line intensity map prediction model is trained, it is also fine-tuned (the fine-tuning process is similar to that of the first edge line intensity map prediction model) to improve the prediction accuracy of the second edge line intensity map prediction model.
[0075] According to exemplary embodiments of this disclosure, at least one first edge line intensity map can be obtained based on the image to be matched using at least one first edge line intensity map prediction model, and at least one second edge line intensity map can be obtained based on the template image using at least one second edge line intensity map prediction model. Specifically, when using a neural network model to predict the edge line intensity map of an image, some detailed information of the image may be lost due to the construction of the training data or the characteristics of the neural network model itself, resulting in room for improvement in the accuracy of the edge line intensity map predicted by the neural network model. In this case, to improve the matching accuracy between the template image and the image to be matched, multiple (e.g., two) first edge line intensity map prediction models and second edge line intensity map prediction models can be trained for the image to be matched and the template image, respectively, using images of the same image type as the image to be matched and images of the same image type as the template image. Thus, multiple first edge line intensity maps or second edge line intensity maps can be obtained for the same image to be matched or template image. Using multiple first edge line intensity maps and second edge line intensity maps for image matching can yield multiple matching results, and the matching position determined based on these multiple matching results is closer to the actual situation.
[0076] In some embodiments, the number of first and second edge line intensity map prediction models can be determined based on the constructed training data, and can be combined with Figure 3 Describe it. Figure 3 This is a schematic diagram illustrating the construction process of training data for first and second edge line intensity map prediction models according to exemplary embodiments of the present disclosure, wherein the training data consists of two types of images that match each other, such as visible light images and SAR images. (Refer to...) Figure 3 For SAR images, the edge intensity map of the visible light image region matched by the SAR image in the labeled training samples can be used as the edge intensity map of the SAR image. Figure 3 (1) is a visible light image. Figure 3 (2) is a SAR image, where Figure 3 (1) The part defined by the middle frame and Figure 3 SAR image matching in (2), Figure 3 (3) The phase consistency algorithm is used to... Figure 3 (1) Edge line intensity map obtained by edge line extraction from visible light image. Figure 3 (4) is from Figure 3The image cropped from the portion defined by the frame in (3) represents the edge intensity map of the SAR image. Thus, by performing this method on all visible light image and SAR image matching pairs in the training samples, a dataset of edge intensity maps of visible light images (which can be denoted as dataset DATA_OPTICAL_A) and a dataset of edge intensity maps of SAR images (which can be denoted as dataset DATA_SAR_A) can be constructed.
[0077] Similarly, for visible light images, unsupervised edge line extraction methods (e.g., phase consistency algorithms) can be used to extract the edge line intensity map of the SAR image as the label of the SAR image. The edge line intensity map of the visible light image is the edge line intensity map of the matching region in the matched SAR image. Two edge line intensity map datasets can be obtained, which can be denoted as DATA_SAR_B and DATA_OPTICAL_B, respectively. Using this method, four training datasets can be constructed: two for visible light images (DATA_OPTICAL_A and DATA_OPTICAL_B) and two for SAR images (DATA_SAR_A and DATA_SAR_B). For each training dataset, an edge line intensity map prediction model can be trained. Specifically, the edge line intensity map prediction model for visible light images trained using datasets DATA_OPTICAL_A and DATA_SAR_A (denoted as MODEL_OPTICAL_A) corresponds to the edge line intensity map prediction model for SAR images (MODEL_SAR_A), and the edge line intensity map prediction model for visible light images trained using datasets DATA_OPTICAL_B and DATA_SAR_B (denoted as MODEL_OPTICAL_B) corresponds to the edge line intensity map prediction model for SAR images (denoted as MODEL_SAR_B).
[0078] In other embodiments, instead of determining the number of first and second edge line intensity map prediction models based on the constructed training data, multiple first and second edge line intensity map prediction models can be trained directly using different neural network structures and loss functions, without any limitation.
[0079] According to exemplary embodiments of this disclosure, to improve the accuracy of the predicted first edge line intensity map and the second edge line intensity map, a multi-scale prediction method can be used for prediction. Specifically, the image to be matched can first be transformed at multiple scales (e.g., but not limited to, enlarging or reducing the image, sub-pixel sampling the image to obtain a thumbnail, etc.), and then, based on the image to be matched at each scale, multiple first transformed edge line intensity maps are obtained through the first edge line intensity map prediction model. Finally, the first edge line intensity map is obtained based on the multiple first transformed edge line intensity maps (e.g., by taking the average of the multiple first transformed edge line intensity maps). Similarly, the template image can first undergo multi-scale transformation (e.g., but not limited to, enlarging or reducing the image, or sub-pixel sampling to obtain a thumbnail). Here, the multi-scale transformation performed on the template image is of the same type as the multi-scale transformation performed on the image to be matched. Then, based on the template image at each scale, multiple second-transformed edge line intensity maps are obtained through a second edge line intensity map prediction model. Finally, based on the multiple second-transformed edge line intensity maps (e.g., averaging the multiple second-transformed edge line intensity maps), a second edge line intensity map is obtained. Here, since the saliency of multiple features of the image differs at different scales, the features among the multiple edge line intensity maps obtained also differ. Therefore, the accuracy of the first and second edge line intensity maps obtained based on multiple first-transformed or second-transformed edge line intensity maps can be significantly improved compared to a single scale.
[0080] In other embodiments, different neural network structures and loss functions may be used to train multiple first and second edge line intensity map prediction models, and the weighted average of the edge line intensity maps output by the multiple first and second edge line intensity map prediction models may be used as the final first and second edge line intensity maps, without limitation.
[0081] It should be noted that, in the specific implementation process, there is no restriction on the execution order of steps 102 and 103. That is to say, step 102 can be executed before step 103, or after step 103, or simultaneously with step 103.
[0082] Return to reference Figure 1 In step 104, the second edge line intensity map can be used as a template to perform template matching on the first edge line intensity map to obtain the first matching result.
[0083] According to an exemplary embodiment of this disclosure, to improve the matching accuracy between a template image and an image to be matched, at least one second edge line intensity map can be used as a template to perform template matching on at least one first edge line intensity map, obtaining at least one matching result (e.g., a matching score map). Based on the at least one matching result (e.g., adding the at least one matching result together, or averaging the sums, etc.), a first matching result is obtained. Here, a template matching method based on correlation coefficients can be used for template matching, or other methods for measuring the similarity between two images can be used, such as, but not limited to, variance matching, cross-correlation matching, etc., without limitation. In some embodiments, the template image and the image to be matched are, for example, a SAR image and a visible light image, respectively, then combined with the... Figure 3 The relevant description states that the prediction models for the first edge line intensity map of the image to be matched are MODEL_OPTICAL_A and MODEL_OPTICAL_B, and the prediction models for the second edge line intensity map of the template image are MODEL_SAR_A and MODEL_SAR_B. Figure 4 This is a schematic diagram illustrating the process of matching SAR images and visible light images based on first and second edge line intensity map prediction models according to an exemplary embodiment of the present disclosure. (Refer to...) Figure 4 The width and height of a visible light image can be denoted as [W]. opt H opt The width and height of the SAR image are denoted as [W]. sar H sar ], will [W opt H opt Input the corresponding MODEL_OPTICAL_A and MODEL_OPTICAL_B respectively to obtain the edge intensity maps optical_edge_a and optical_edge_b of the predicted visible light image; then input [W sar H sar Input the corresponding MODEL_SAR_A and MODEL_SAR_B respectively to obtain the edge intensity maps sar_edge_a and sar_edge_b of the predicted SAR image. A template matching method based on correlation coefficients can be used, with optical_edge_a as the object to be matched and sar_edge_a as the template, to perform template matching, resulting in a template of size [W]. opt -W sar +1,H opt -H sar The matching score map (which can be denoted as score_map_a) of size [+1] is obtained, and template matching is performed using optical_edge_b as the object to be matched and sar_edge_b as the template, resulting in a template matching of size [W]. opt -Wsar +1,H opt -H sar The matching score map (which can be denoted as score_map_b) with a value of +1], when weighted and averaged over score_map_a and score_map_b, yields a result of size [W]. opt -W sar +1,H opt -H sar The final matching score map of +1] (i.e., the first matching result, which can be denoted as score_map).
[0084] Return to reference Figure 1 In step 105, the target image can be obtained multiple times from the image to be matched. Here, the target image and the template image have the same size.
[0085] In step 106, based on the acquired target image and template image, the template image and the image to be matched can be matched using an image matching model to obtain a second matching result.
[0086] According to exemplary embodiments of this disclosure, the image matching model may employ a Resnest50-based classification framework. Of course, any convolutional neural network framework, such as common image classification networks, may also be used, without limitation. In some embodiments, the training process of the image matching model is as follows: First, a third training dataset is obtained, wherein the third training dataset includes multiple third images, fourth images, positive sample pairs, and negative sample pairs. Here, the type of the third image is the same as the type of the image to be matched, the type of the fourth image is the same as the type of the template image, each third image and each fourth image have the same size, positive sample pairs are mutually matched third and fourth images, and negative sample pairs are mutually unmatched third and fourth images. Specifically, the third training dataset can be constructed as follows: For two types of images that have been matched, a region is randomly cropped from one type of image according to a certain range of scale and aspect ratio. Then, regions that match the cropped region are found from the matched other type of image, and the two form a positive sample pair (the label value can be recorded as 1). Regions in the other type of image that do not match the cropped region form a negative sample pair with the cropped region (the label value can be recorded as 0). In this way, a large number of positive and negative sample pairs can be constructed. The region cropped from one type of image can be used as the third image, and the region cropped from the other type of image can be used as the fourth image.
[0087] Then, based on any third and fourth images, the estimated matching probability can be obtained through the image matching model. Here, the estimated matching probability is the matching probability between the third and fourth images. To prevent overfitting, the third and fourth images can be subjected to random data augmentation processing with the same parameters (e.g., random scaling, random translation, random rotation, random flipping, color space transformation, image sharpening, and image blurring, etc.). The processed third and fourth images are then superimposed on each other in channels to form a multi-channel image, which is then input into the image matching model to obtain the matching probability between them. Alternatively, the third and fourth images can be directly superimposed on each other in channels and then input into the image matching model. Alternatively, the third and fourth images can be input into two separate feature extraction backbone networks, and the extracted feature maps can be concatenated together, passed through several convolutional layers, and then the estimated matching probability can be output. There are no restrictions on this approach. To reduce the training difficulty of the image matching model, Gaussian distribution normalization can be applied to the negative sample pairs.
[0088] Then, a third loss can be calculated based on the estimated matching probability, the label values of the positive sample pairs corresponding to the third and fourth images, and the label values of the negative sample pairs. Here, the third loss can be represented by a third loss function, which, for example, but not limited to, can be expressed as:
[0089] FL(p)=-|yp| β ((1-y)log(1-p)+ylog(p))
[0090] Where y represents the label value of a positive sample pair or the label value of a negative sample pair; p represents the estimated matching probability; -|yp| β The dynamic weights are indicated. Furthermore, this disclosure does not limit the third loss function; any possible loss function (e.g., judging the squared error, absolute error, etc. between the estimated matching probability and the label value of the corresponding positive sample pair or the label value of the negative sample pair pixel by pixel) can be used to train the image matching model of this disclosure.
[0091] Finally, the image matching model can be trained by adjusting its parameters based on the third loss. That is, the model parameters can be adjusted via backpropagation using the third loss (e.g., a value calculated using the third loss function). Furthermore, during model training, batches of second and third images can be used to adjust (or update) the model parameters, iteratively adjusting (or updating) them with the goal of minimizing the third loss until the image matching model converges.
[0092] After obtaining a trained image matching model, image matching can be performed using this model. First, target images of the same size as the template image can be obtained multiple times from the image to be matched. Here, a sliding window approach can be used to obtain images of the same size as the template image. Then, based on the obtained target images and template images, the image matching model can be used to match the template image and the image to be matched to obtain a second matching result (e.g., a matching score image).
[0093] In some embodiments, to improve the accuracy of image matching, multiple image matching models can be trained using different neural network structures and loss functions. When matching the template image and the image to be matched, multiple image matching models are used to perform the matching, resulting in multiple matching results. The weighted average of the multiple matching results is then used as the second matching result.
[0094] Figure 5 This is a schematic diagram illustrating the process of matching SAR images and visible light images based on an image allocation model according to an exemplary embodiment of the present disclosure.
[0095] Reference Figure 5 The SAR image is the template image, and the visible light image is the image to be matched. The width and height of the visible light image are denoted as [W]. opt H opt The width and height of a SAR image are denoted as [W]. sar H sar [opt] Cropping an image of the same size as the SAR image from a visible light image using a sliding window. (x,y) Then put opt (x,y) Here, (x, y) represents the starting position of the sliding window clipping, and the traversal range of x is [0, W]. opt -W sar The traversal range of y is [0, H]. opt -H sar The sample pairs (x, y) of the SAR image and the SAR image are input into the image matching model (MatchModel) to predict the probability value of a window at the starting position (x, y) matching the SAR image. By sliding all the windows, a window of size [W] can be obtained. opt -H sar +1,H opt -H sar The matching probability map of +1] (i.e., the second matching result) can be denoted as prob_map.
[0096] It should be noted that, in the specific implementation process, there is no restriction on the execution order of "steps 102, 103, and 104" and "steps 105 and 106". That is to say, "steps 102, 103, and 104" can be executed before "steps 105 and 106", or after "steps 105 and 106", or simultaneously with "steps 105 and 106".
[0097] Return to reference Figure 1 In step 107, the matching position between the image to be matched and the template image can be obtained based on the first matching result and the second matching result. According to an exemplary embodiment of this disclosure, the first matching result and the second matching result can be added together to obtain the final result, and the position with the largest value in the final result can be determined as the matching position between the image to be matched and the template image. Of course, the average value can also be taken after adding them together as the final result for the current position, and there is no limitation on this. Specifically, each result in the first matching result and the second matching result represents the degree of matching between the current position in the image to be matched and the upper left corner position of the template image. The higher the value of the result, the higher the degree of matching. Here, since the first matching result is obtained based on the first and second edge line intensity map prediction models, and the second matching result is obtained based on the image matching model, the matching position between the image to be matched and the template image can be obtained by combining the first matching result and the second matching result, which can greatly improve the matching accuracy of image matching.
[0098] Figure 6 This is a block diagram illustrating an image matching apparatus according to an exemplary embodiment of the present disclosure.
[0099] Reference Figure 6 An image matching apparatus 600 according to an exemplary embodiment of the present disclosure may include an image acquisition unit 601, a first edge line intensity map acquisition unit 602, a second edge line intensity map acquisition unit 603, a first matching result determination unit 604, a target image acquisition unit 605, a second matching result determination unit 606, and a matching position determination unit 607.
[0100] The image acquisition unit 601 can acquire the image to be matched and the template image. Here, the image to be matched and the template image are heterogeneous images (i.e., images acquired by sensors or platforms using different imaging methods), such as visible light images, SAR images, or infrared images, etc., without limitation. The template image refers to the reference template used in image matching, and the image to be matched refers to the object being matched in the image matching process. In specific implementations, the size of the template image is smaller than the size of the image to be matched, and the final result of the matching is to find the position and region of the template image corresponding to the image to be matched.
[0101] The first edge line intensity map acquisition unit 602 can obtain a first edge line intensity map based on the image to be matched using a first edge line intensity map prediction model. The second edge line intensity map acquisition unit 603 can obtain a second edge line intensity map based on a template image using a second edge line intensity map prediction model. Here, an edge line refers to a location in an image where pixel values change abruptly, and the edge line intensity refers to the pixel value corresponding to that edge line. By finding all locations in an image where pixel values change abruptly, the edge line intensity map corresponding to that image is formed.
[0102] According to an exemplary embodiment of this disclosure, the first edge line intensity map prediction model may adopt the U-net network structure framework, and its backbone network may adopt Resnest50. Of course, any convolutional neural network framework, such as common image segmentation networks, may also be used, without limitation. The first edge line intensity map prediction model is pre-trained. Here, the training process of the first edge line intensity map prediction model may be as follows: First, a first training dataset is obtained, wherein the first training dataset includes a first image and the real edge line intensity map corresponding to the first image, and the type of the first image is the same as the type of the image to be matched. In the specific implementation process, to obtain better training results, multiple first images are used. Traditional unsupervised edge extraction methods (such as phase consistency algorithms) can be employed to extract the edge lines of the first images. After obtaining the edge lines, to avoid inaccurate edge lines due to thresholding, the edge lines are not divided into binary values of 0 or 1 based on the threshold. Furthermore, to reduce the training difficulty of the first edge line intensity map prediction model, the pixel values of the extracted edge lines can be normalized to between 0 and 1, thereby obtaining the true edge line intensity map, which is used as the label during the training process. Then, based on the first image, the first edge line intensity map prediction model can be used to obtain the first estimated map (i.e., the estimated edge line intensity map corresponding to the first image). Here, to prevent overfitting, random data augmentation (such as random scale changes, random translations, random rotations, and random flips) can be applied to the first image and its corresponding label. The first image after data augmentation is then input into the first edge line intensity map prediction model to obtain the first estimated map. Of course, the first image can also be directly input into the first edge line intensity map prediction model; there are no restrictions on this. Subsequently, a first loss can be calculated based on the first estimated map and the true edge line intensity map. Here, to improve the prediction accuracy of the first edge line intensity map prediction model, the first loss may include cross-entropy loss, cross-correlation loss, and correlation loss. In some embodiments, the first loss can be represented by a first loss function, which, for example, but not limited to, can be expressed as:
[0103] L=w1FL+w2CCORRLOSS+w3CCOEFFLOSS
[0104] Wherein, FL represents the cross-entropy loss function (i.e., represents the cross-entropy loss); CCORRLOSS represents the cross-correlation loss function (i.e., represents the cross-correlation loss); CCOEFFLOSS represents the correlation loss function (i.e., represents the correlation loss); w1, w2, and w3 represent different weights. Here, the cross-entropy loss function is used to calculate the cross-entropy loss between the first estimated image and its corresponding label pixel by pixel. Since the label contains a large number of non-edge pixels (i.e., edge intensity is 0), the cross-entropy loss function can be dynamically weighted. In some embodiments, the cross-entropy loss function, for example, but not limited to, can be expressed as:
[0105] FL(x)=-|x′-x| β ((1-x′)log(1-x)+x′log(x))
[0106] Where x represents the first estimated map, x′ represents the label corresponding to the first estimated map, and -|x′-x| β This indicates dynamic weights.
[0107] The cross-correlation loss function measures the normalized cross-correlation between the first estimated map and its corresponding label. In some embodiments, the cross-correlation loss function, for example, but not limited to, may be expressed as:
[0108] CCORRLOSS = 1 - CCORR NORMED
[0109] Among them, CCOEFF NORMED The normalized cross-correlation coefficient between the first estimated plot and its corresponding label is expressed in the relevant techniques and will not be repeated here.
[0110] The correlation loss function measures the normalized correlation between the first estimated map and its corresponding label. In some embodiments, the correlation loss function, for example, but not limited to, may be expressed as:
[0111] CCOEFFLOSS = 1 - CCOEFF NORMED
[0112] Among them, CCOEFF NORMED The normalized cross-correlation coefficient between the first estimated map and its corresponding label is expressed in related technologies and will not be repeated here. Furthermore, this disclosure does not limit the first loss function; any possible loss function (e.g., determining the squared error, absolute error, etc. between the first estimated map and its corresponding label pixel-by-pixel) can be used to train the first edge line intensity map prediction model of this disclosure.
[0113] Finally, the edge line intensity map prediction model corresponding to the image to be matched can be trained by adjusting the model parameters of the model model based on the first loss. In some embodiments, the model parameters of the first edge line intensity map prediction model can be adjusted by backpropagation of the first loss calculated by the first loss function. Furthermore, during model training, a batch of first images can be used to adjust (or update) the model parameters of the first edge line intensity map prediction model, and the model parameters of the first edge line intensity map prediction model are iteratively adjusted (or updated) with the goal of minimizing the value of the first loss function until the first edge line intensity map prediction model converges.
[0114] According to an exemplary embodiment of this disclosure, after training the first edge line intensity map prediction model, to further improve the prediction accuracy of the model, it can be fine-tuned, that is, the model can be trained a second time. Specifically, the second training process is as follows: First, a second training dataset is obtained, wherein the second training dataset includes a second image, the real edge line intensity map corresponding to the second image, and an incorrect edge line intensity map. Here, the type of the second image is the same as the type of the image to be matched, and the incorrect edge line intensity map is the edge line intensity map obtained when performing template matching on the edge line intensity map output by the trained first edge line intensity map prediction model, which deviates from the real matching position. Specifically, after the first edge line intensity map prediction model is trained, the first edge line intensity map obtained by the trained first edge line intensity map prediction model is used to perform image matching tests. During the test, there may be a situation where the matched edge line intensity map is not the edge line intensity map corresponding to the real template image in the image to be matched (for example, see reference). Figure 2 Then, based on the second image, a second estimated image can be obtained using the trained first edge intensity map prediction model. Similar to the aforementioned training process, to prevent overfitting, random data augmentation (e.g., random scale changes, random translations, random rotations, random flips, color space transformations, image sharpening, or image blurring) can be applied to the second image, the corresponding true edge intensity map, and the erroneous edge intensity map. The data-augmented first image is then input into the first edge intensity map prediction model to obtain the second estimated image. Afterward, a second loss can be calculated based on the second estimated image, the corresponding true edge intensity map, and the erroneous edge intensity map. The first edge intensity map prediction model is then trained by adjusting its parameters according to the second loss. Here, the second loss can be represented by a second loss function, which, for example but not limited to, can be expressed as:
[0115] HEMLOSS=2+CCOEFF NORMED neg -CCOEFF NORMED pos
[0116] Among them, CCOEFF NORMED neg This represents the correlation coefficient between the second estimated plot and the erroneous edge line intensity plot; CCOEFF NORMED pos This represents the correlation coefficient between the second estimated map and the corresponding true edge line intensity map of the second image. Similarly, CCOEFF... NORMED neg and CCOEFF NoRMED pos The expression for this has already been demonstrated in relevant technologies and will not be elaborated upon here.
[0117] Since the structure and training process of the second edge line intensity map prediction model are similar to those of the first edge line intensity map prediction model, with the only difference being the training data and training labels, the structure and training process of the second edge line intensity map prediction model will not be described in detail here for the sake of brevity. It should be noted that after the second edge line intensity map prediction model is trained, it is also fine-tuned (the fine-tuning process is similar to that of the first edge line intensity map prediction model) to improve the prediction accuracy of the second edge line intensity map prediction model.
[0118] According to an exemplary embodiment of the present disclosure, the first edge line intensity map acquisition unit 602 may obtain at least one first edge line intensity map based on the image to be matched by at least one first edge line intensity map prediction model, and the second edge line intensity map acquisition unit 603 may obtain at least one second edge line intensity map based on the template image by at least one second edge line intensity map prediction model. Specifically, when using a neural network model to predict the edge intensity map of an image, some detailed information of the image may be lost due to the construction of the training data or the characteristics of the neural network model itself, resulting in room for improvement in the accuracy of the edge intensity map predicted by the neural network model. In this case, in order to improve the matching accuracy between the template image and the image to be matched, multiple (e.g., two) first edge intensity map prediction models and second edge intensity map prediction models can be trained for the image to be matched and the template image, respectively, using images of the same image type as the image to be matched and images of the same image type as the template image. Thus, for the same image to be matched, the first edge intensity map acquisition unit 602 can obtain multiple first edge intensity maps, and for the same template image, the second edge intensity map acquisition unit 603 can obtain a second edge intensity map. By using multiple first edge intensity maps and second edge intensity maps for image matching, multiple matching results can be obtained. The matching position determined based on these multiple matching results is closer to the actual situation.
[0119] In some embodiments, the number of first and second edge line intensity map prediction models can be determined based on the constructed training data (e.g., refer to the method embodiments regarding...). Figure 3 The description of the first and second edge line intensity map prediction models will not be repeated here. In other embodiments, the number of first and second edge line intensity map prediction models may not be determined based on the constructed training data, but rather multiple first and second edge line intensity map prediction models may be trained directly using different neural network structures and loss functions, without any limitation.
[0120] According to an exemplary embodiment of this disclosure, the first edge line intensity map acquisition unit 602 can perform multi-scale transformations on the image to be matched (e.g., but not limited to, enlarging or reducing the image, or sub-pixel sampling of the image to obtain a thumbnail, etc.), and then, based on the image to be matched at each scale, obtain multiple first transformed edge line intensity maps through a first edge line intensity map prediction model. Finally, based on the multiple first transformed edge line intensity maps (e.g., by taking the average of the multiple first transformed edge line intensity maps), a first edge line intensity map is obtained. The second edge line intensity map acquisition unit 603 can perform multi-scale transformations on the template image (e.g., but not limited to, enlarging or reducing the image, or sub-pixel sampling of the image to obtain a thumbnail, etc.). Here, the multi-scale transformation performed on the template image is of the same transformation type as the multi-scale transformation performed on the image to be matched. Then, based on the template image at each scale, obtain multiple second transformed edge line intensity maps through a second edge line intensity map prediction model. Finally, based on the multiple second transformed edge line intensity maps (e.g., by taking the average of the multiple second transformed edge line intensity maps), a second edge line intensity map is obtained. Here, since the saliency of multiple features of images at different scales is different, the features of the multiple edge line intensity maps obtained are also different. Therefore, the accuracy of the first edge line intensity map and the second edge line intensity map obtained based on multiple first transformed edge line intensity maps or second transformed edge line intensity maps can be significantly improved compared with single scale.
[0121] In other embodiments, different neural network structures and loss functions may be used to train multiple first and second edge line intensity map prediction models. The first edge line intensity map acquisition unit 602 may use the weighted average of the edge line intensity maps output by multiple first edge line intensity map prediction models as the final first edge line intensity map, and the second edge line intensity map acquisition unit 603 may use the weighted average of the edge line intensity maps output by multiple second edge line intensity map prediction models as the final second edge line intensity map. There are no restrictions on this.
[0122] It should be noted that, in the specific implementation process, the execution order of the first edge line intensity map acquisition unit 602 and the second edge line intensity map acquisition unit 603 is not limited. That is to say, the first edge line intensity map acquisition unit 602 may be executed before the second edge line intensity map acquisition unit 603, or may be executed after the second edge line intensity map acquisition unit 603, or may be executed simultaneously with the second edge line intensity map acquisition unit 603.
[0123] The first matching result determination unit 604 can use the second edge line intensity map as a template to perform template matching on the first edge line intensity map to obtain the first matching result.
[0124] According to an exemplary embodiment of this disclosure, to improve the matching accuracy between the template image and the image to be matched, the first matching result determination unit 604 can use at least one second edge line intensity map as a template, perform template matching on at least one first edge line intensity map, and obtain at least one matching result (e.g., a matching score map). Based on the at least one matching result (e.g., adding the at least one matching result together, or averaging the sums), a first matching result is obtained. Here, a template matching method based on correlation coefficients can be used for template matching, or other methods for measuring the similarity between two images can be used, such as, but not limited to, variance matching, cross-correlation matching, etc., without limitation.
[0125] The target image acquisition unit 605 can acquire the target image multiple times from the image to be matched. Here, the target image and the template image have the same size.
[0126] The second matching result determination unit 606 can match the template image and the image to be matched based on the acquired target image and template image through an image matching model to obtain the second matching result.
[0127] According to exemplary embodiments of this disclosure, the image matching model may employ a Resnest50-based classification framework. Of course, any convolutional neural network framework, such as common image classification networks, may also be used, without limitation. In some embodiments, the training process of the image matching model is as follows: First, a third training dataset is obtained, wherein the third training dataset includes multiple third images, fourth images, positive sample pairs, and negative sample pairs. Here, the type of the third image is the same as the type of the image to be matched, the type of the fourth image is the same as the type of the template image, each third image and each fourth image have the same size, positive sample pairs are mutually matched third and fourth images, and negative sample pairs are mutually unmatched third and fourth images. Specifically, the third training dataset can be constructed as follows: For two types of images that have been matched, a region is randomly cropped from one type of image according to a certain range of scale and aspect ratio. Then, regions that match the cropped region are found from the matched other type of image, and the two form a positive sample pair (the label value can be recorded as 1). Regions in the other type of image that do not match the cropped region form a negative sample pair with the cropped region (the label value can be recorded as 0). In this way, a large number of positive and negative sample pairs can be constructed. The region cropped from one type of image can be used as the third image, and the region cropped from the other type of image can be used as the fourth image.
[0128] Then, based on any third and fourth images, the estimated matching probability can be obtained through the image matching model. Here, the estimated matching probability is the matching probability between the third and fourth images. To prevent overfitting, the third and fourth images can be subjected to random data augmentation processing with the same parameters (e.g., random scaling, random translation, random rotation, random flipping, color space transformation, image sharpening, and image blurring, etc.). The processed third and fourth images are then superimposed on each other in channels to form a multi-channel image, which is then input into the image matching model to obtain the matching probability between them. Alternatively, the third and fourth images can be directly superimposed on each other in channels and then input into the image matching model. Alternatively, the third and fourth images can be input into two separate feature extraction backbone networks, and the extracted feature maps can be concatenated together, passed through several convolutional layers, and then the estimated matching probability can be output. There are no restrictions on this approach. To reduce the training difficulty of the image matching model, Gaussian distribution normalization can be applied to the negative sample pairs.
[0129] Then, a third loss can be calculated based on the estimated matching probability, the label values of the positive sample pairs corresponding to the third and fourth images, and the label values of the negative sample pairs. Here, the third loss can be represented by a third loss function, which, for example, but not limited to, can be expressed as:
[0130] FL(p)=-|yp| β ((1-y)log(1-p)+ylog(p))
[0131] Where y represents the label value of a positive sample pair or the label value of a negative sample pair; p represents the estimated matching probability; -|yp| β The dynamic weights are indicated. Furthermore, this disclosure does not limit the third loss function; any possible loss function (e.g., judging the squared error, absolute error, etc. between the estimated matching probability and the label value of the corresponding positive sample pair or the label value of the negative sample pair pixel by pixel) can be used to train the image matching model of this disclosure.
[0132] Finally, the image matching model can be trained by adjusting its parameters based on the third loss. That is, the model parameters can be adjusted via backpropagation using the third loss (e.g., a value calculated using the third loss function). Furthermore, during model training, batches of second and third images can be used to adjust (or update) the model parameters, iteratively adjusting (or updating) them with the goal of minimizing the third loss until the image matching model converges.
[0133] After obtaining a trained image matching model, image matching can be performed using this model. First, target images of the same size as the template image can be obtained multiple times from the image to be matched. Here, a sliding window approach can be used to obtain images of the same size as the template image. Then, based on the obtained target images and template images, the image matching model can be used to match the template image and the image to be matched to obtain a second matching result (e.g., a matching score image).
[0134] In some embodiments, to improve the accuracy of image matching, multiple image matching models can be trained using different neural network structures and loss functions. When matching the template image and the image to be matched, the second matching result determination unit 606 can use multiple image matching models to perform matching, obtain multiple matching results, and take the weighted average of the multiple matching results as the second matching result.
[0135] It should be noted that, in the specific implementation process, the execution order of the "first edge line intensity map acquisition unit 602, second edge line intensity map acquisition unit 603, first matching result determination unit 604" and the "target image acquisition unit 605, second matching result determination unit 606" is not limited. That is to say, the "first edge line intensity map acquisition unit 602, second edge line intensity map acquisition unit 603, first matching result determination unit 604" can be executed before the "target image acquisition unit 605, second matching result determination unit 606", or after the "target image acquisition unit 605, second matching result determination unit 606", or can be executed simultaneously with the "target image acquisition unit 605, second matching result determination unit 606".
[0136] The matching position determination unit 607 can obtain the matching position between the image to be matched and the template image based on the first matching result and the second matching result. According to an exemplary embodiment of this disclosure, the matching position determination unit 607 can add the first matching result and the second matching result to obtain the final result, and determine the position with the largest value in the final result as the matching position between the image to be matched and the template image. Of course, it is also possible to add them and take the average value as the final result of the current position, and there is no limitation on this.
[0137] Figure 7 This is a block diagram of an electronic device 700 according to an exemplary embodiment of the present disclosure.
[0138] Reference Figure 7The electronic device 700 includes at least one memory 701 and at least one processor 702. The at least one memory 701 stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor 702, an image matching method according to an exemplary embodiment of the present disclosure is performed.
[0139] As an example, electronic device 700 may be a PC, tablet, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, electronic device 700 is not necessarily a single electronic device, but may be a collection of any devices or circuits capable of executing the aforementioned instructions (or instruction sets) individually or in combination. Electronic device 700 may also be part of an integrated control system or system manager, or may be configured to interconnect with a portable electronic device locally or remotely (e.g., via wireless transmission) through an interface.
[0140] In electronic device 700, processor 702 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, processor may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, etc.
[0141] The processor 702 can execute instructions or code stored in the memory 701, which can also store data. Instructions and data can also be sent and received via a network through a network interface device, which can employ any known transmission protocol.
[0142] The memory 701 can be integrated with the processor 702, for example, by placing RAM or flash memory within an integrated circuit microprocessor. Alternatively, the memory 701 can include a separate device, such as an external disk drive, a storage array, or other storage device usable by any database system. The memory 701 and the processor 702 can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., enabling the processor 702 to read files stored in the memory.
[0143] In addition, the electronic device 700 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device 700 can be interconnected via a bus and / or network.
[0144] According to exemplary embodiments of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, they cause at least one processor to perform an image matching method according to the present disclosure. Examples of computer-readable storage media herein include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, servers, etc. Furthermore, in one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0145] According to the image matching method, apparatus, electronic device, and storage medium disclosed herein, on the one hand, by obtaining the first edge line intensity map and the second edge line intensity map respectively through the first and second edge line intensity map prediction models, the accuracy of edge line extraction of the image can be improved, thereby improving the matching accuracy of image matching based on edge line intensity maps. On the other hand, by matching the template image and the image to be matched through the image matching model, the problem of low matching accuracy caused by information loss when performing template matching based on edge line extraction methods can be improved, thereby improving the matching accuracy of image matching. Furthermore, by combining the matching results of these two aspects to obtain the final matching position, the potential deficiencies between the two can be compensated for, thereby greatly improving the matching accuracy of image matching.
[0146] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0147] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image matching method, characterized in that, include: Obtain the image to be matched and the template image, wherein the size of the template image is smaller than the size of the image to be matched; Based on the image to be matched, a first edge line intensity map is obtained through a first edge line intensity map prediction model. The first edge line intensity map prediction model is trained using an image of the same type as the image to be matched. Based on the template image, a second edge line intensity map is obtained through a second edge line intensity map prediction model, which is trained using an image of the same type as the template image. Using the second edge line intensity map as a template, template matching is performed on the first edge line intensity map to obtain a first matching result; The target image is obtained multiple times from the image to be matched, and the target image has the same size as the template image; Based on the acquired target image and template image, the template image and the image to be matched are matched using an image matching model to obtain a second matching result; Based on the first matching result and the second matching result, the matching position between the image to be matched and the template image is obtained.
2. The image matching method as described in claim 1, characterized in that, The step of obtaining the first edge line intensity map based on the image to be matched using the first edge line intensity map prediction model includes: Based on the image to be matched, at least one first edge line intensity map is obtained through at least one first edge line intensity map prediction model; The step of obtaining the second edge line intensity map based on the template image using the second edge line intensity map prediction model includes: Based on the template image, at least one second edge line intensity map is obtained through at least one second edge line intensity map prediction model.
3. The image matching method as described in claim 2, characterized in that, The step of using the second edge line intensity map as a template to perform template matching on the first edge line intensity map to obtain a first matching result includes: Using the at least one second edge line intensity map as a template, template matching is performed on the at least one first edge line intensity map to obtain at least one matching result; The first matching result is obtained based on the at least one matching result.
4. The image matching method according to any one of claims 1 to 3, characterized in that, The step of obtaining the first edge line intensity map based on the image to be matched using the first edge line intensity map prediction model includes: Perform multi-scale transformation on the image to be matched; Based on the image to be matched at each scale, multiple first transformed edge line intensity maps are obtained through the first edge line intensity map prediction model; The first edge line intensity map is obtained based on the plurality of first transformed edge line intensity maps; The step of obtaining the second edge line intensity map based on the template image using the second edge line intensity map prediction model includes: The template image is subjected to a multi-scale transformation, wherein the multi-scale transformation performed on the template image is of the same type as the multi-scale transformation performed on the image to be matched; Based on the template image at each scale, multiple second-transformed edge line intensity maps are obtained through the second edge line intensity map prediction model; The second edge line intensity map is obtained based on the plurality of second transformed edge line intensity maps.
5. The image matching method as described in claim 1, characterized in that, The training process of the first edge line intensity map prediction model includes: Obtain a first training dataset, wherein the first training dataset includes a first image and a real edge intensity map corresponding to the first image, and the type of the first image is the same as the type of the image to be matched; Based on the first image, a first estimated image is obtained using the first edge line intensity map prediction model; The first loss is calculated based on the first estimated map and the actual edge line intensity map; The first edge line intensity map prediction model is trained by adjusting the model parameters of the first edge line intensity map prediction model according to the first loss.
6. The image matching method as described in claim 5, characterized in that, The first loss includes cross-entropy loss, cross-correlation loss, and correlation loss, and is expressed as: in, This represents the cross-entropy loss; This represents the cross-correlation loss; This represents the correlation loss; , and This indicates different weights.
7. The image matching method as described in claim 5, characterized in that, After training the first edge line intensity map prediction model, the method further includes: Obtain a second training dataset, wherein the second training dataset includes a second image, a true edge intensity map corresponding to the second image, and an incorrect edge intensity map. The type of the second image is the same as the type of the image to be matched. The incorrect edge intensity map is an edge intensity map that deviates from the true matching position when performing template matching on the first edge intensity map output by the trained first edge intensity map prediction model. Based on the second image, a second estimated image is obtained using the first edge line intensity map prediction model; The second loss is calculated based on the second estimated map, the true edge intensity map corresponding to the second image, and the erroneous edge intensity map. The first edge line intensity map prediction model is trained by adjusting the model parameters of the first edge line intensity map prediction model according to the second loss.
8. The image matching method as described in claim 7, characterized in that, The second loss is expressed as: in, This represents the correlation coefficient between the second estimated map and the erroneous edge line intensity map; This represents the correlation coefficient between the second estimated map and the actual edge line intensity map corresponding to the second image.
9. The image matching method as described in claim 1, characterized in that, The training process of the image matching model includes: Obtain a third training dataset, wherein the third training dataset includes multiple third images, fourth images, positive sample pairs, and negative sample pairs. The type of the third images is the same as the type of the images to be matched, the type of the fourth images is the same as the type of the template images, each third image and each fourth image has the same size, the positive sample pairs are mutually matched third and fourth images, and the negative sample pairs are mutually unmatched third and fourth images. Based on any of the third and fourth images, the estimated matching probability is obtained through the image matching model; The third loss is calculated based on the estimated matching probability, the label value of the positive sample pair, and the label value of the negative sample pair; The image matching model is trained by adjusting its parameters according to the third loss.
10. The image matching method as described in claim 9, characterized in that, The third loss is represented as: in, This represents the label value of the positive sample pair or the label value of the negative sample pair; This represents the estimated matching probability; This indicates dynamic weights.
11. An image matching device, characterized in that, include: The image acquisition unit is configured to acquire an image to be matched and a template image, wherein the size of the template image is smaller than the size of the image to be matched. The first edge line intensity map acquisition unit is configured to: obtain a first edge line intensity map based on the image to be matched by a first edge line intensity map prediction model, wherein the first edge line intensity map prediction model is trained using an image of the same type as the image to be matched; The second edge line intensity map acquisition unit is configured to: obtain a second edge line intensity map based on the template image using a second edge line intensity map prediction model, wherein the second edge line intensity map prediction model is trained using an image of the same type as the template image; The first matching result determination unit is configured to: use the second edge line intensity map as a template to perform template matching on the first edge line intensity map to obtain the first matching result; The target image acquisition unit is configured to acquire a target image multiple times from the image to be matched, wherein the target image and the template image have the same size; The second matching result determination unit is configured to: based on the acquired target image and the template image, match the template image and the image to be matched using an image matching model to obtain a second matching result; The matching position determination unit is configured to: obtain the matching position between the image to be matched and the template image based on the first matching result and the second matching result.
12. The image matching apparatus as claimed in claim 11, characterized in that, The first edge line intensity map acquisition unit is configured as follows: Based on the image to be matched, at least one first edge line intensity map is obtained through at least one first edge line intensity map prediction model; The second edge line intensity map acquisition unit is configured to: obtain at least one second edge line intensity map based on the template image and through at least one second edge line intensity map prediction model.
13. The image matching apparatus as claimed in claim 12, characterized in that, The first matching result determination unit is configured as follows: Using the at least one second edge line intensity map as a template, template matching is performed on the at least one first edge line intensity map to obtain at least one matching result; The first matching result is obtained based on the at least one matching result.
14. The image matching apparatus according to any one of claims 11 to 13, characterized in that, The first edge line intensity map acquisition unit is configured as follows: Perform multi-scale transformation on the image to be matched; Based on the image to be matched at each scale, multiple first transformed edge line intensity maps are obtained through the first edge line intensity map prediction model; The first edge line intensity map is obtained based on the plurality of first transformed edge line intensity maps; The second edge line intensity map acquisition unit is configured as follows: The template image is subjected to a multi-scale transformation, wherein the multi-scale transformation performed on the template image is of the same type as the multi-scale transformation performed on the image to be matched; Based on the template image at each scale, multiple second-transformed edge line intensity maps are obtained through the second edge line intensity map prediction model; The second edge line intensity map is obtained based on the plurality of second transformed edge line intensity maps.
15. The image matching apparatus as claimed in claim 11, characterized in that, The training process of the first edge line intensity map prediction model includes: Obtain a first training dataset, wherein the first training dataset includes a first image and a real edge intensity map corresponding to the first image, and the type of the first image is the same as the type of the image to be matched; Based on the first image, a first estimated image is obtained using the first edge line intensity map prediction model; The first loss is calculated based on the first estimated map and the actual edge line intensity map; The first edge line intensity map prediction model is trained by adjusting the model parameters of the first edge line intensity map prediction model according to the first loss.
16. The image matching apparatus as claimed in claim 15, characterized in that, The first loss includes cross-entropy loss, cross-correlation loss, and correlation loss, and is expressed as: in, This represents the cross-entropy loss; This represents the cross-correlation loss; This represents the correlation loss; , and This indicates different weights.
17. The image matching apparatus as claimed in claim 15, characterized in that, After training the first edge line intensity map prediction model is completed, the following steps are also included: Obtain a second training dataset, wherein the second training dataset includes a second image, a true edge intensity map corresponding to the second image, and an incorrect edge intensity map. The type of the second image is the same as the type of the image to be matched. The incorrect edge intensity map is an edge intensity map that deviates from the true matching position when performing template matching on the first edge intensity map output by the trained first edge intensity map prediction model. Based on the second image, a second estimated image is obtained using the first edge line intensity map prediction model; The second loss is calculated based on the second estimated map, the true edge intensity map corresponding to the second image, and the erroneous edge intensity map. The first edge line intensity map prediction model is trained by adjusting the model parameters of the first edge line intensity map prediction model according to the second loss.
18. The image matching apparatus as claimed in claim 17, characterized in that, The second loss is expressed as: in, This represents the correlation coefficient between the second estimated map and the erroneous edge line intensity map; This represents the correlation coefficient between the second estimated map and the actual edge line intensity map corresponding to the second image.
19. The image matching apparatus as claimed in claim 11, characterized in that, The training process of the image matching model includes: Obtain a third training dataset, wherein the third training dataset includes multiple third images, fourth images, positive sample pairs, and negative sample pairs. The type of the third images is the same as the type of the images to be matched, the type of the fourth images is the same as the type of the template images, each third image and each fourth image has the same size, the positive sample pairs are mutually matched third and fourth images, and the negative sample pairs are mutually unmatched third and fourth images. Based on any of the third and fourth images, the estimated matching probability is obtained through the image matching model; The third loss is calculated based on the estimated matching probability, the label value of the positive sample pair, and the label value of the negative sample pair; The image matching model is trained by adjusting its parameters according to the third loss.
20. The image matching apparatus as claimed in claim 19, characterized in that, The third loss is represented as: in, This represents the label value of the positive sample pair or the label value of the negative sample pair; This represents the estimated matching probability; This indicates dynamic weights.
21. An electronic device, characterized in that, include: At least one processor; At least one memory that stores computer-executable instructions. The computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform the image matching method as described in any one of claims 1 to 10.
22. A computer-readable storage medium for storing instructions, characterized in that, When the instruction is executed by at least one processor, it causes the at least one processor to perform the image matching method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Different-source image matching method based on template matching and twin neural network optimization
CN112801141A