An image localization method including any target and background
By interpolation scaling and depth feature extraction of images, combined with correlation calculation, the problems of low image positioning accuracy and poor generalization under illumination changes and rotation transformation in the prior art are solved, and efficient and accurate image positioning is achieved.
Patent Information
- Application Number
- CN202411287160.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-09-13
AI Technical Summary
The existing image positioning methods have low accuracy in lighting changes, angle rotation and affine transformation, and poor generalization of training models, resulting in large matching errors and high time costs.
The template and search images are scaled by interpolation method, and deep feature extraction is performed using open source or specific data-trained feature backbone extraction networks. The correlation is calculated based on Euclidean distance, dot product or cosine similarity, and the best matching result is obtained through non-maximum suppression.
It improves the light insensitivity and rotation transformation adaptability of image positioning, improves matching accuracy, reduces training time costs, and enhances model generalization.
Smart Images

Figure CN119090963B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image localization, and specifically relates to an image localization method including any target and background. Background Art
[0002] In existing image localization methods, there are the following methods:
[0003] Method 1: The Chinese patent document with the application number 202111255618.9 provides a method and system for heterologous image matching and localization based on a Siamese network and supervised training. The template image T to be matched is traversed in the search image I, pixel features are compared, and then correlation calculation is performed. However, in this method, if the illumination of the template image to be matched and the search image changes, the method fails.
[0004] Method 2: The Chinese patent with the application number 202410202593.3 provides an image retrieval method and related device. Multiple-resolution template images T are subjected to feature extraction to obtain their respective multi-size features FT, and then the search image is subjected to feature extraction to obtain the feature to be searched FI. The feature to be searched FI is combined with the multi-scale features FT for image matching to obtain a search result, thereby locating the position of the template image in the search image. However, in this method, the search image only outputs a single search feature, and the obtained search feature is single, resulting in a large matching error.
[0005] Method 3: The Chinese patent with the application number 202310359338.5 provides a heterologous image matching method based on a Siamese neural network, including (1) performing feature extraction on the template image T to be matched and the search image I to obtain corresponding feature maps FT and FI; (2) traversing the obtained template feature FT to be matched on the search feature map FI, and performing a correlation operation at each traversed position to obtain a matching cross-correlation map Mgen; (3) building a regression network Netreg to perform position regression on the matching correlation map Mgen and output the relative position estimation of the template image T to be matched in the search image I; (4) constructing an ideal cross-correlation map Mide according to the actual position information as the supervision information of the matching cross-correlation map; (5) building a discriminant network Netdis, constructing a discriminant loss and an overall generation loss function, and training the network. However, in this method, a regression Siamese network is used for training, so the trained model has poor generalization ability for unknown images, resulting in overfitting. The template feature to be matched is only matched with a single search feature, and the single search feature contains insufficient semantic information, making it difficult for the network model to fit, and the accuracy of the model is low. Moreover, training the model takes a long time. Summary of the Invention
[0006] The object of the present invention is to overcome the deficiencies of the prior art and provide an image positioning method including any target and background.
[0007] The object of the present invention is achieved by the following technical solutions:
[0008] The present invention provides an image positioning method including any target and background, comprising:
[0009] S1. Scale the template image by interpolation to obtain a first template image. The size of the original template image is , and the size of the first template image is ;
[0010] S2. Use a first model to infer the first template image to obtain the depth features of the first template image. The size of the depth features of the first template image is ;
[0011] S3. Scale the search image by interpolation to obtain a first search image. The size of the original search image is , and the size of the first search image is ;
[0012] S4. Use the first model to infer the first search image to obtain the depth features of the first search image. The size of the depth features of the first search image is ;
[0013] S5. Calculate the correlation result between the depth features of the first template image and the depth features of the first search image, and select the result with the largest correlation as the matching result. The calculation includes:
[0014] Move the depth features of the first template image left to right or top to bottom in the depth features of the first search image. During the movement, calculate the correlation result between the depth features of the first template image and the corresponding depth features of the first search image. According to the width and height of the depth features of the first template image , select the depth features of the first template image with the result of the largest correlation as the first matching result. Map the coordinates of the depth features of the first search image corresponding to the first matching result to the image coordinates of the first search image;
[0015] S6. Repeat steps S1 - S5 for N times, and sequentially replace the first model with the Nth model to obtain the Nth matching result;
[0016] S7. Use non - maximum suppression to obtain the best matching result.
[0017] Further, the interpolation method includes nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation.
[0018] Further, the first model to the Nth model are feature backbone extraction networks trained with open-source data or feature extraction backbone networks for object detection and object segmentation trained with specific data.
[0019] Further, the method for inferring the first search image or the first template image using the first model includes any one or a combination of a convolution operator, a residual operator, a multi-layer perceptron operator, a fully connected network operator, or a Transformer operator.
[0020] Further, the method for calculating the correlation result between the depth feature of the first template image and the depth feature of the corresponding first search image includes the Euclidean distance: in an n-dimensional real vector space, A and B are two points in the n-dimensional real vector space, the coordinates of point A are , and the coordinates of point B are , and the Euclidean distance between point A and point B is defined as: ;
[0021] where d(A, B) represents the Euclidean distance between point A and point B, and the calculated Euclidean distance is used as the correlation result between the depth feature of the first template image and the depth feature of the corresponding first search image.
[0022] Further, the method for calculating the correlation result between the depth feature of the first template image and the depth feature of the corresponding first search image further includes: dot product, cosine similarity, or absolute value error.
[0023] The beneficial effects of the present invention are as follows:
[0024] 1) The present invention solves the problem of insensitivity to illumination, and can still be successfully recognized when the image undergoes angular rotation and affine transformation;
[0025] 2) The present invention uses search image features of multiple scales, making the matching accuracy higher;
[0026] 3) The present invention does not require training a network model, has better generalization, and saves time costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of an image localization method including any object and background according to an embodiment of the present invention;
[0028] Figure 2 is the first template image of an image localization method including any object and background according to an embodiment of the present invention;
[0029] Figure 3 Search Image 1 for an image localization method including any target and background according to an embodiment of the present invention;
[0030] Figure 4 Search Image 2 for an image localization method including any target and background according to an embodiment of the present invention;
[0031] Figure 5 Template Image 2 for an image localization method including any target and background according to an embodiment of the present invention. Detailed implementation manners
[0032] Next, in combination with the embodiments, the technical solutions of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0033] Exemplarily, the flowchart of the present invention is as Figure 1 shown. The present invention includes the following steps: the step of obtaining the depth feature of the first template image, the step of obtaining the depth feature of the first search image, and the step of obtaining the best matching result. In the present invention, there is no sequence between the step of obtaining the depth feature of the first template image and the step of obtaining the depth feature of the first search image.
[0034] Specifically, it is possible to first process Template Image 1 as Figure 2 shown to obtain the depth feature of the first template image of the template image. After that, process Search Image 1 as Figure 3 shown to obtain the depth feature of the first search image of the search image. Finally, perform the step of obtaining the best matching result; it is also possible to first perform the step of obtaining the depth feature of the first search image, process Search Image 2 as Figure 4 shown to obtain the depth feature of the first search image of the search image. After that, perform the step of obtaining the depth feature of the first template image, process Template Image 2 as Figure 5 shown to obtain the depth feature of the first template image of the template image. Finally, perform the step of obtaining the best matching result.
[0035] Specifically, the step of obtaining the depth feature of the first template image includes the following sub-steps: Scale the template image by an interpolation method, where the interpolation method includes methods such as nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation, and any one of them can be selected. The size of the original template image is ; corresponding to the width, height, and number of layers of the original template image, obtaining the scaled first template image, and the size of the first template image is , corresponding to the width, height, and number of layers of the first template image, where the width is greater than 12, and the height is also greater than 12. Then, the first template image is inferred through the first model. The first model can adopt a feature backbone extraction network trained with open-source data (such as: ResNet, Vit, DeiT, and Unet), or a feature extraction backbone network for object detection and object segmentation trained with specific data. For example, for specific data training: infrared and visible light, in time and coordinates, infrared and visible light correspond to data such as vehicles, airplanes, or houses. In this way, when training the model, both infrared and visible light are involved in training at the same time; the inference method can be a combination of one or more of the following operators: convolution operator, residual operator, multi-layer perceptron operator, fully connected network operator, or Transformer operator. Each open-source model contains different operators, so the included calculation processes are also different, thereby obtaining the depth features of the first template image. The size of the depth features of the first template image is , corresponding to the width, height, and number of layers of the depth features of the first template image.
[0036] Specifically, the steps for obtaining the depth features of the first search image include the following sub-steps: scaling the search image through an interpolation method, where the interpolation method includes methods such as nearest-neighbor interpolation, bilinear interpolation, or bicubic interpolation, and any one of them can be selected. The size of the original search image is , corresponding to the width, height, and number of layers of the original search image, obtaining the scaled first search image, and the size of the first search image is , corresponding to the width, height, and number of layers of the first search image, where the width is greater than 12, and the height is also greater than 12. The size difference between the first template image and the first search image is X times, where 1 < X < the resolution of the original search image ; Then, the first search image is inferred through the first model. The first model can adopt a feature backbone extraction network trained with open-source data (such as: ResNet, Vit, DeiT, and Unet), or a feature extraction backbone network for object detection and object segmentation trained with specific data. For example, for infrared and visible light, in time and coordinates, there are data such as vehicles, airplanes, or houses corresponding to infrared and visible light. Thus, when training the model, both infrared and visible light are involved in the training simultaneously. The inference method can be a combination of one or more of the following operators: convolution operator, residual operator, multi-layer perceptron operator, fully connected network operator, or Transformer operator. Each open-source model contains different operators, so the calculation processes are also different, thereby obtaining the depth features of the first search image. The size of the depth features of the first search image is , corresponding to the width, height, and number of layers of the depth features of the first search image.
[0037] Specifically, the steps to obtain the best matching result include the following sub-steps: Move the depth features of the first template image in the depth features of the first search image. The moving method can be from left to right or from top to bottom. During the moving process, calculate the correlation result between the depth features of the first template image and the depth features of the first search image at the corresponding position. The calculation methods include methods such as Euclidean distance, dot product, cosine similarity, or absolute value error. According to the width and height of the size of the depth features of the first template image , select the depth features of the first template image with the largest correlation result as the first matching result. The coordinates of the depth features of the first search image corresponding to the first matching result are mapped to the image coordinates of the first search image; then repeat the above steps, and change the first model to the Nth model, and finally obtain the Nth matching result; finally, through non-maximum suppression, obtain the best matching result.
[0038] The method of non-maximum suppression is as follows:
[0039] 1. Sorting: First, sort in descending order according to the confidence of each bounding box (usually a comprehensive indicator of classification probability and localization accuracy). The bounding box with the highest confidence is considered the most likely to correctly detect the target.
[0040] 2. Selection: Select the bounding box with the highest confidence from the sorted list, mark it as selected, and add it to the final detection result list.
[0041] 3. Calculate IOU: For each remaining bounding box, calculate its IOU with the selected bounding box.
[0042] 4. Comparison and elimination: If the IOU of a certain bounding box and the selected box exceeds the preset threshold (e.g., 0.5 or 0.7), it is considered that the two boxes represent the same object. Then, according to the principle of lower confidence, this low-confidence bounding box is eliminated.
[0043] Repeat steps 2-4: Continue to select the one with the highest confidence among the remaining bounding boxes, and repeat the IOU calculation and elimination process until all bounding boxes have been checked.
[0044] Exemplarily, after obtaining the first search image by interpolating and scaling an arbitrarily selected interpolation method for the search image, the size of the first search image is: 1920*1080*3, corresponding to the width * height * number of layers of the first search image. After the inference (calculation of the first model), the depth feature of the first search image is obtained, and the size of the depth feature of the first search image is: 960*540*50; after obtaining the first template image by interpolating and scaling an arbitrarily selected interpolation method for the template image, the size of the first template image is: 50*50*3, corresponding to the width * height * number of layers of the first template image. After the first template image passes through the inference (calculation of the first model), the depth feature of the first template image is obtained, and the size of the depth feature of the first template image is: 25*25*50.
[0045] Slide the depth feature of the first template image on the depth feature of the first search image from left to right and from top to bottom in turn. The size of the sliding window is the width and height 25*25 of the depth feature size of the first template image, and calculate it with the sliding window feature value of the corresponding depth feature of the first search image in turn; for example, calculate it through the Euclidean distance, which is to subtract the same-dimensional numerical values corresponding to the same point. In the n-dimensional space (n-dimensional real vector space), point A and point B are two points in the n-dimensional space. The coordinates of point A are , and the coordinates of point B are , and the Euclidean distance between point A and point B is defined as: ; wherein, d(A, B) represents the Euclidean distance between point A and point B, and the calculated Euclidean distance is used as the correlation result between the depth feature of the first template image and the depth feature of the corresponding first search image. The size of the sliding window is 25*25, which is the width and height of the depth feature size of the first template image. The depth feature of the first template image with the largest correlation result is selected as the first matching result. The coordinates of the first matching result are: (w0, w1, h0, h1) = (100, 125, 100, 125), where (w0, w1, h0, h1) are the coordinate values relative to the depth feature of the first search image, and the image coordinates mapped to the first search image (1920*1080*3) are (W0, W1, H0, H1). The calculation method of the image coordinates mapped to the first search image (1920*1080*3) is: w0:W0 = 960:1920, so W0 = w0 * 1920 / 960. Calculate W1 in the same way; h0:H0 = 540:1080, so H0 = h0 * 1080 / 540. Calculate H1 in the same way.
[0046] Repeat the above steps, change the first model to the Nth model in sequence, and finally obtain the Nth matching result. Finally, through non-maximum suppression, obtain the best matching result
[0047] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. An image positioning method including any target and background, characterized in that, Including: S1. Scale the template image through interpolation to obtain the first template image. The size of the original template image is , where represents the width of the original template image, represents the height of the original template image. The size of the first template image is , represents the width of the first template image, represents the height of the first template image, represents the number of layers of the template image. The template image includes the original template image and the first template image; S2. Use the first model to perform inference on the first template image to obtain the depth features of the first template image. The size of the depth features of the first template image is ; represents the width of the depth features of the first template image, represents the height of the depth features of the first template image; S3. Scale the search image through interpolation to obtain the first search image. The size of the original search image is ; represents the width of the original search image, represents the height of the original search image, represents the number of layers of the original search image. The size of the first search image is ; represents the width of the first search image, represents the height of the first search image, represents the number of layers of the first search image; S4. Use the first model to perform inference on the first search image to obtain the depth features of the first search image. The size of the depth features of the first search image is ; represents the width of the depth features of the first search image, represents the height of the depth features of the first search image, represents the number of layers of the depth features. The depth features include the depth features of the first search image and the depth features of the first template image; S5. Calculate the correlation result between the depth feature of the first template image and the depth feature of the first search image, and select the result with the largest correlation as the matching result. The calculation includes: Move the depth feature of the first template image left to right or top to bottom in the depth feature of the first search image. During the movement, calculate the correlation result between the depth feature of the first template image and the corresponding depth feature of the first search image. According to the width and height of the size of the depth feature of the first template image , select the depth feature of the first template image with the largest correlation result as the first matching result, and map the coordinates of the depth feature of the first search image corresponding to the first matching result to the image coordinates of the first search image; S6. Repeat steps S1 - S5 for N times, successively replace the first model with the Nth model, and obtain the Nth matching result; S7. Use non - maximum suppression to obtain the best matching result.
2. The image localization method including any target and background according to claim 1, wherein: The interpolation method includes nearest - neighbor interpolation, bilinear interpolation, or bicubic interpolation.
3. A method for image localization including any target and background according to claim 1, characterized in that: The first model to the Nth model are feature backbone extraction networks trained with open - source data or feature extraction backbone networks for object detection and object segmentation trained with specific data.
4. A method for image localization including any target and background according to claim 1, characterized in that: The method of using the first model to perform inference on the first search image or the first template image includes any one or a combination of multiple of the convolution operator, residual operator, multi - layer perceptron operator, fully - connected network operator, or Transformer operator.
5. A method for image localization including any target and background according to claim 1, characterized in that, The method for calculating the correlation result between the depth feature of the first template image and the corresponding depth feature of the first search image includes the Euclidean distance: in an n-dimensional real vector space, A and B are two points in the n-dimensional real vector space, and the coordinates of point A are , and the coordinates of point B are . The Euclidean distance between point A and point B is defined as: ; where d(A, B) represents the Euclidean distance between point A and point B, and the calculated Euclidean distance is used as the correlation result between the depth feature of the first template image and the corresponding depth feature of the first search image.
6. A method for image localization including any target and background according to claim 5, characterized in that, The method of calculating the correlation result between the depth feature of the first template image and the corresponding depth feature of the first search image also includes: dot product, cosine similarity, or absolute error.
Citation Information
Patent Citations
Heterogeneous image matching and positioning method and system based on twin network and supervised training
CN114022729B
Heterogenous image matching method based on twin neural network
CN116385747A
Image retrieval method and related device
CN117788842A
Fast and dense image matching method and system
CN113033708A
Different-source image target positioning method based on gradient direction
CN118314336A