An image matching method, device and electronic device

By adjusting the threshold T of feature point extraction in image matching, and using the self-attention mechanism and cross-attention mechanism of graph neural network, the problem of difficulty in extracting feature points in areas with weak texture is solved, and the accuracy and wide application of image matching are improved.

CN114972820BActive Publication Date: 2025-06-24HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210622207.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-06-24
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively extract feature points in areas with weak textures in image matching, resulting in uneven distribution of feature points, which in turn affects the accuracy of image matching, especially in applications such as SLAM systems.

Method used

By building an image pyramid, the image is divided into multiple image blocks, and the threshold T of feature point extraction is adjusted within the image blocks whose texture does not meet the setting requirements to ensure that feature points are extracted in these areas. At the same time, the graph neural network is used to determine the matching relationship between images.

Benefits of technology

Improve the accuracy of image matching and ensure that areas with weak textures can also extract feature points, making the image matching method more suitable for other tasks such as SLAM systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972820B_ABST
    Figure CN114972820B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an image matching method, apparatus and electronic device. In the embodiments of the present application, a corresponding first image pyramid is constructed for a first image to be matched, and for each component image in each layer of the first image pyramid, the component image is divided into image blocks of L*M. When feature points cannot be extracted in an image block whose texture does not meet the set requirements, the set threshold T originally used for feature point extraction is adjusted to enable feature points to be extracted in the image block whose texture does not meet the set requirements, so as to ensure that feature points can also be extracted in image blocks with weak texture, ensure that the finally extracted feature points are evenly distributed, and further improve the accuracy of image matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and particularly to an image matching method, apparatus, and electronic device. Background Art

[0002] Image matching refers to finding similar targets between two or more images through a certain matching algorithm. For example, in two-dimensional image matching, the correlation coefficient of windows of the same size in the target area and the search area is compared, and the center point of the window corresponding to the largest correlation coefficient in the search area is taken as the target point.

[0003] Currently, the commonly used image matching algorithm is a feature point-based matching method. However, when the existing feature point-based matching method is used, it is extracted from the image area with stronger texture. Considering that the characteristics of the image area with stronger texture are different from those of the image area with weaker texture (also referred to as the image area where the texture does not meet the set requirements), when the original method for extracting feature points in the image area with stronger texture is applied to the image area with weaker texture (also referred to as the image area where the texture does not meet the set requirements), few feature points will be extracted or no feature points can be extracted. This results in extremely uneven distribution of the extracted feature points. And the unevenness of the feature points will lead to inaccurate image matching and even limit the application of the image matching method, such as not being able to be used in other tasks such as the SLAM system. Summary of the Invention

[0004] Embodiments of this application disclose an image matching method, apparatus, and electronic device to improve the accuracy of image matching.

[0005] Embodiments of this application provide an image matching method, which includes:

[0006] For each component image in each layer of the first image pyramid corresponding to the first image to be matched, divide the component image into image blocks of L*M, extract feature points from at least one image block of the component image, and obtain feature point information corresponding to the feature points; wherein, the number of image blocks in the component image is greater than the number of feature points L required to be extracted from the component image, and when no feature points can be extracted in the image block where the texture does not meet the set requirements, adjust the originally set threshold T for extracting feature points to achieve extracting feature points in the image block where the texture does not meet the set requirements; the feature point information at least includes the position of the pixel point determined as a feature point in the component image and the confidence that the pixel point is determined as a feature point; each layer in the first image pyramid has a corresponding component image, and different layers correspond to component images of different sizes, and the component images of different sizes are obtained by performing size conversion on the first image;

[0007] For each feature point in the composed image on each layer of the first image pyramid, determine the feature point descriptor corresponding to the feature point;

[0008] Input the feature point information and feature point descriptors corresponding to the feature points extracted from the composed images on each layer of the first image pyramid into a trained graph neural network, so that the graph neural network uses the self-attention mechanism and the cross-attention mechanism to determine and output the matching relationship between the first image and the second image as the target.

[0009] An embodiment of the present application provides an image matching device, and the device includes:

[0010] An image processing unit, for each composed image on each layer of the first image pyramid corresponding to the first image to be matched, divide the composed image into image blocks of L*M, extract feature points from at least one image block of the composed image, and obtain the feature point information corresponding to the feature points; wherein, the number of image blocks in the composed image is greater than the number of feature points L required to be extracted from the composed image, and when no feature points can be extracted from an image block where the texture does not meet the set requirements, adjust the set threshold T originally used to extract feature points, so as to extract feature points in the image block where the texture does not meet the set requirements; the feature point information at least includes the position of the pixel point determined as a feature point in the composed image and the confidence that the pixel point is determined as a feature point; each layer in the first image pyramid has a corresponding composed image, and the composed images corresponding to different layers have different sizes, and the composed images of different sizes are obtained by performing size conversion on the first image;

[0011] A determination unit, for each feature point in the composed image on each layer of the first image pyramid, determine the feature point descriptor corresponding to the feature point;

[0012] A matching unit, for inputting the feature point information and feature point descriptors corresponding to the feature points extracted from the composed images on each layer of the first image pyramid into a trained graph neural network, so that the graph neural network uses the self-attention mechanism and the cross-attention mechanism to determine and output the matching relationship between the first image and the second image as the target.

[0013] An embodiment of the present application provides an electronic device, and the electronic device includes: a processor and a memory.

[0014] Wherein, the memory is used to store machine-executable instructions;

[0015] The processor is used to read and execute the machine-executable instructions stored in the memory to implement the above method.

[0016] As can be seen from the above description, in this embodiment, for the first image to be matched, a corresponding first image pyramid is constructed. For each component image in the first image pyramid, the component image is divided into L*M image patches. When no feature points can be extracted from the image patches whose texture does not meet the set requirements, the set threshold T originally used for feature point extraction is adjusted to enable feature points to be extracted from the image patches whose texture does not meet the set requirements, so as to ensure that feature points can also be extracted from the image patches with weak texture, ensure that the finally extracted feature points are evenly distributed, and thus improve the accuracy of image matching.

[0017] Further, in this embodiment, since the finally extracted feature points are evenly distributed during image matching, this further ensures that this image matching method can be better used in other tasks such as SLAM systems, improving the application universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the specification and used together with the specification to explain the principles of the specification.

[0019] Figure 1 It is a flowchart of the method provided by the embodiment of the present application;

[0020] Figure 2 It is a flowchart of feature point extraction provided by the embodiment of the present application;

[0021] Figure 3 It is another flowchart of feature point extraction provided by the embodiment of the present application;

[0022] Figure 4 It is a flowchart for determining the feature point descriptor provided by the embodiment of the present application;

[0023] Figure 5 It is a schematic structural diagram of the MLP example provided by the embodiment of the present application;

[0024] Figure 6 It is a schematic structural diagram of the device provided by the embodiment of the present application;

[0025] Figure 7 It is a schematic diagram of the hardware structure of the device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of the devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0027] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0028] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0029] To enable those skilled in the art to better understand the technical solutions provided by the embodiments of this application and to make the above-mentioned objects, features, and advantages of the embodiments of this application more obvious and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0030] See Figure 1 , Figure 1 which is a flowchart of the method provided by the embodiment of this application. This method is applied to an electronic device. As an example, the electronic device may be a device such as a PC, and this embodiment does not specifically limit it.

[0031] As Figure 1 shown, this process may include the following steps:

[0032] Step 101, for each component image in each layer of the first image pyramid corresponding to the first image to be matched, divide the component image into multiple image blocks of size M1*M2, extract feature points from at least one image block of the component image, and obtain the feature point information corresponding to the feature points; wherein, when no feature points can be extracted from an image block where the texture does not meet the set requirements, adjust the set threshold T originally used for extracting feature points to enable feature points to be extracted from the image block where the texture does not meet the set requirements.

[0033] Optionally, in this embodiment, the first image pyramid may be constructed for the first image first. The first image pyramid has H layers. H is greater than 1. In this embodiment, each layer in the first image pyramid has a corresponding component image. As an example, different layers correspond to different component images.

[0034] Optionally, in this embodiment, the constituent images corresponding to at least H - 1 layers are obtained by performing image conversion on the first image, and the constituent image corresponding to the remaining one layer may be the first image; alternatively, the constituent images corresponding to H layers are all obtained by performing image conversion on the first image to achieve different constituent images corresponding to different layers.

[0035] Taking image conversion as size conversion as an example, the H layers correspond to constituent images of different sizes. The constituent images of different sizes corresponding to at least H - 1 layers are all obtained by performing size conversion on the first image.

[0036] After constructing the above first image pyramid, in this embodiment, as described in step 101, for the constituent image of each layer in the first image pyramid, the constituent image is divided into multiple image blocks of size M1 * M2. Here, M1 and M2 can be equal or unequal, and can be specifically set according to actual needs. For example, M1 and M2 are set to 30, and the constituent image is divided into multiple image blocks of size 30 * 30. In this embodiment, the number of image blocks in the constituent image is greater than the number L of feature points required to be extracted from the constituent image to retain the feature point with the maximum confidence in each image block. For the specific description of feature point extraction, see the following text and will not be elaborated here.

[0037] After dividing the image blocks, feature points can be extracted from at least one image block of the constituent image. Optionally, in this embodiment, the feature points are, for example, points where the image gray value changes drastically in the above - mentioned constituent image or points with a large curvature on the image edge, such as corner points, edges, blocks, etc. This embodiment does not specifically limit. Figure 2 An example is used to describe how to extract feature points and will not be elaborated here.

[0038] It should be noted that when feature points cannot be extracted in an image block with weak texture (an image block whose texture does not meet the set requirements), to ensure that the finally extracted feature points are evenly distributed, the originally set threshold T for extracting feature points can be adjusted here to achieve extracting feature points in an image block whose texture does not meet the set requirements, such as an image block with weak texture. For the specific process, see Figure 2 the shown process and will not be elaborated here.

[0039] Step 102, for each feature point in the constituent image of each layer in the first image pyramid, determine the feature point descriptor corresponding to the feature point.

[0040] Optionally, in this embodiment, there are many methods to determine the feature point descriptor corresponding to the feature point. For example, the Brief descriptor is used to determine the feature point descriptor corresponding to the feature point. For easy understanding, Figure 4An example is described to illustrate how to use the Brief descriptor to determine the feature point descriptors corresponding to the feature points, which will not be elaborated here for the time being.

[0041] Step 103: Input the feature point information and feature point descriptors corresponding to the feature points in each layer of the first image pyramid, and the feature point information and feature point descriptors corresponding to the feature points in each layer of the second image pyramid of the obtained second image as the target into the trained graph neural network, so that the graph neural network uses the self-attention mechanism and the cross-attention mechanism to determine the matching relationship between the first image and the second image and output it.

[0042] In this embodiment, the second image pyramid corresponding to the second image and the feature point information corresponding to the feature points in each layer of the second image pyramid are obtained according to the description of a similar step 101, and the feature point descriptors corresponding to the feature points in each layer of the second image pyramid are obtained according to the description of a similar step 102, which will not be elaborated here.

[0043] In this embodiment, the graph neural network aggregates the information of intra-image feature points and inter-image feature points based on the self-attention mechanism and the cross-attention mechanism to output the matching relationship between the first image and the second image. Here, the intra-image feature points (Intra-image edges, ε self ) refer to the feature point i on each layer of the first image pyramid connecting to other feature points in the same image (the feature point satisfies the matching condition for matching with the feature point i). The inter-image feature points (Inter-image deges, ε cross ) refer to the feature point i on each layer of the first image pyramid to some feature points on the second image (the feature point satisfies the matching condition for matching with the feature point i). This will be described when training the graph neural network based on supervised learning, which will not be elaborated here for the time being.

[0044] As an embodiment, in this embodiment, the above matching relationship may at least include: the matching confidence between the feature points in the first image and the feature points in the second image. In this embodiment, the above matching relationship does not directly perform Softmax classification on the network output. The above matching relationship may include the matching confidence between the feature points in the first image and the feature points in the second image, which is convenient for backend visualization and selection.

[0045] So far, the Figure 1 shown process is completed.

[0046] Through Figure 1As can be seen from the process shown, in this embodiment, a corresponding first image pyramid is constructed for the first image to be matched. For each component image in each layer of the first image pyramid, the component image is divided into image blocks of L*M. When no feature points can be extracted in the image blocks where the texture does not meet the set requirements, the originally set threshold T for extracting feature points is adjusted to enable feature points to be extracted in the image blocks where the texture does not meet the set requirements, so as to ensure that feature points can also be extracted in the image blocks with weak texture, ensure that the finally extracted feature points are evenly distributed, and further improve the accuracy of image matching.

[0047] Furthermore, in this embodiment, since the finally extracted feature points are evenly distributed during image matching, this further ensures that this image matching method can be better used in other tasks such as SLAM systems, improving the application universality.

[0048] Next, the Figure 2 process shown in this application embodiment will be described:

[0049] Refer to Figure 2 , Figure 2 which is the flowchart of feature point extraction provided by this application embodiment. As an embodiment, Figure 2 the process shown extracts feature points by improving the existing feature point detection algorithm FAST.

[0050] As Figure 2 shown, the process may include the following steps:

[0051] Step 201, for each image block in each component image on each layer of the first image pyramid, perform the following step 202.

[0052] That is to say, in this embodiment, feature points are extracted for each image block separately. The specific extraction criteria are shown in the following steps 202 to 204.

[0053] Step 202, for each pixel point in the image block, construct a circle with the pixel point as the center and a set value r1 as the radius in the image block, and determine whether the boundary of the circle meets the set conditions. If so, perform step 203; if not, perform step 204.

[0054] In this embodiment, the set radius r1 can be set according to actual needs, such as set to 3 pixel points, etc., and this embodiment does not specifically limit it.

[0055] The set conditions are related to the gray value of the pixel point. Suppose the gray value of the pixel point is I P , and the set condition is that the gray values of n consecutive pixel points on the circle boundary are greater than I P +T or less than I P -T.

[0056] For example, draw a circle with the pixel point P as the center and a radius of 3 pixel points. If the gray value of each pixel point among n consecutive pixel points on the boundary of the circle, such as 12 consecutive pixel points, is greater than I P +T or less than I P -T, then step 203 is executed; otherwise, step 204 is executed.

[0057] Step 203: Determine the pixel point as the target pixel point and determine the confidence level of the pixel point as a feature point. Then step 205 is executed.

[0058] This step 203 is carried out under the condition that the boundary of the circle meets the set conditions. However, for some image blocks with textures that do not meet the set requirements (such as image blocks with weak textures), according to the above set threshold T, it is impossible to determine the target pixel point. At this time, the following step 204 is executed, that is, adjust the above set threshold T.

[0059] Step 204: If there is no record of the number of threshold changes currently, or there is a record of the number of threshold changes but the recorded number of threshold changes is less than the preset maximum number of threshold changes, then adjust the above set threshold T and return to judge whether the boundary of the circle meets the set conditions. Then step 205 is executed.

[0060] In this embodiment, the adjustment of the above set threshold T is also limited. For example, it is achieved by limiting the number of threshold changes, etc. Among them, if there is no record of the number of threshold changes for the pixel point currently (indicating that the above set threshold T has not been adjusted yet), or there is a record of the number of threshold changes but the recorded number of threshold changes is less than the preset maximum number of threshold changes, then adjust the above set threshold T and return to judge whether the boundary of the circle meets the set conditions. Optionally, in this embodiment, adjusting the above set threshold T can be: reducing the above set threshold T.

[0061] Based on the above description, in this embodiment, after adjusting the above set threshold T, the method further includes: when there is no record of the number of threshold changes for the pixel point currently, record the number of threshold changes and set the number of threshold changes to a set value, such as 1; when there is a record of the number of threshold changes currently, increase the recorded number of threshold changes by a set value, such as 1.

[0062] It should be noted that in this embodiment, when the recorded number of threshold changes for the pixel point is equal to the preset maximum number of threshold changes, and at this time, if the pixel point still cannot be determined as the target pixel point, then it can be directly defaulted that the pixel point is not the target pixel point, and the step of judging whether the pixel point is the target pixel point is directly ended.

[0063] In this embodiment, after determining that a pixel is a target pixel in step 203 above, if there is a record of the number of threshold changes for this pixel currently, the record of the number of threshold changes is deleted.

[0064] It should also be noted that, as an embodiment, in this embodiment, to save resources, for each pixel, the adjustment of the above-mentioned set threshold T can be restricted to only once. As long as it is determined after adjusting the above-mentioned set threshold T that the boundary of the circle with this pixel as the center and the set value r1 as the radius does not meet the set conditions, then it is defaulted that this pixel is not a target pixel.

[0065] Step 205: Extract feature points from at least one image block that makes up the image according to the target pixels in each image block.

[0066] Optionally, as an embodiment, the target pixels in each image block can be directly determined as feature points.

[0067] Optionally, as another embodiment, after determining the target pixels above, there may be many target pixels crowded together, which affects the uniformity of the feature points. In response to this situation, the target pixels can also be made uniform through a uniform strategy such as the quadtree strategy. This will be described by way of example below and will not be elaborated here for the time being. Figure 3 For example, it will not be elaborated here for the time being.

[0068] So far, the Figure 2 shown process is completed.

[0069] Through the Figure 2 shown process, it is realized how to extract feature points from at least one image block that makes up the image.

[0070] Next, the Figure 3 shown process will be described:

[0071] Refer to Figure 3 , Figure 3 which is another flowchart of feature point extraction provided by the embodiment of the present application. As Figure 3 shown, this process may include the following steps:

[0072] Step 301: For each composed image corresponding to each layer in the first image pyramid, according to the size of the composed image, at least one image region is divided from the composed image, the size of the image region is a specified size, and the following step 302 is executed for each image region.

[0073] Here, the specified size can be set according to actual needs, such as 640*480, etc., and this embodiment does not specifically limit it.

[0074] Step 302: Take the image region as the current region; divide the current region into K sub-regions, delete the sub-regions where no target pixel points exist, and determine whether the total number of sub-regions is greater than the above-mentioned L. If not, execute Step 303; if so, execute Step 304.

[0075] If the above-mentioned uniform strategy is a quadtree strategy, K here can be 4.

[0076] Step 303: For each sub-region, when the number of target pixel points in the sub-region is greater than 1, take the sub-region as the current region and return to the step of dividing the current region into K sub-regions in the above Step 302. When there is only one target pixel point in the sub-region, determine the target pixel point as a feature point.

[0077] Step 304: For each sub-region, when the number of target pixel points in the sub-region is greater than 1, select a target pixel point with the highest confidence as a feature point from the sub-region. When there is only one target pixel point in the sub-region, determine the target pixel point as a feature point.

[0078] Thus far, the Figure 3 shown process is completed.

[0079] Through the Figure 3 shown process, the extraction of feature points from at least one image block of the composed image is realized according to the uniform strategy based on the target pixel points in each image block.

[0080] Next, the Figure 4 shown process will be described:

[0081] Refer to Figure 4 , Figure 4 which is the flowchart of the feature point descriptor provided by the embodiment of the present application. As Figure 4 shown, this process may include the following steps:

[0082] Step 401: For each feature point, determine the centroid of the neighborhood of the feature point.

[0083] In this embodiment, the gray centroid method can be used to calculate the direction of each feature point to solve the problem of rotation invariance. In order to calculate the direction of each feature point, first determine the centroid of the neighborhood of each feature point. Here, the neighborhood refers to the image region on the composed image that is at a set distance r2 from the feature point.

[0084] Taking the gray centroid method as an example, here, the centroid of the neighborhood can be represented by the following formula 1:

[0085]

[0086] where m pq = ∑x,y x p y q I(x, y)p, q = {0, 1}, m pq represents the moment of the neighborhood. I(x, y) is the gray value of the feature point at the position (x, y) within the neighborhood.

[0087] Step 402, determine the direction between the feature point and the centroid as the direction of the feature point. Then perform Step 403.

[0088] For example, the direction of the feature point is represented by the following formula 2:

[0089]

[0090] where θ represents the direction of the feature point.

[0091] Step 403, rotate the image window centered on the feature point on the composed image by the above θ, select N pairs of pixel points within the image window, and determine the descriptor of the feature point based on the feature values of the two different pixel points in each pair of pixel points under the same feature attribute.

[0092] Taking the BRIEF descriptor as an example, in this embodiment, first perform Gaussian blur on the above - composed image to reduce the interference of noise; then take an image window of size S*S centered on the feature point (S is generally taken as 31), rotate the image window by the above θ along the direction of the feature point, take N pairs of pixel points (such as N is 256) within the rotated image window, for each pair of pixel points, compare the magnitudes of their gray values, if they are close or the same, return 1, otherwise, return 0, and finally form a binary code of length N, such as 256, which is the feature point descriptor.

[0093] So far, the Figure 4 shown process is completed.

[0094] Through Figure 4 the shown process, how to determine the feature point descriptor is realized.

[0095] Next, the graph neural network provided by the embodiments of the present application will be described:

[0096] As an embodiment, the graph neural network can be trained through the following steps:

[0097] 1), Feature point encoding: This step mainly combines the position information corresponding to each feature point on the training image and the feature point descriptor through a multi - layer perceptron structure. Figure 5 Illustrate the multi - layer perceptron structure by way of example. The training feature points on the training image include the feature points on each layer of the composed image in the training pyramid corresponding to the training image.

[0098] For example, for training image A, let the set of feature points in image A be K and the set of feature descriptors be D. Then the feature point encoding corresponding to the feature points can be:

[0099] x i = d i + MLP(P i )

[0100] where d i ∈ R D , represents the feature descriptor (i.e., visual information) of the i-th feature point. MLP is the network structure of the multi-layer perceptron ( Figure 5 as shown in the example). P i represents the feature point information of the i-th feature point in image A (including the position (x, y) of the pixel point as the i-th feature point in the two-dimensional coordinate system and the confidence c of the pixel point as the feature point). P i = (x, y, c) i . For the target image matching the training image, the processing method is similar.

[0101] 2), Graph structure: The nodes of the graph structure are the feature points on the training image. The graph structure involves the intra-image edges (ε self ) feature points and inter-image edges (ε cross ) feature points described above to ensure that a multi-graph neural network will eventually be generated. The multi-graph neural network starts from the high-dimensional state of each node and aggregates the messages of all given edges of all nodes simultaneously to update the feature points on each layer of the image pyramid corresponding to the training image. The update of the feature points on each layer of the composed image is represented by the following formula:

[0102]

[0103] where is the intermediate representation of the i-th feature point on the l-th layer of the composed image in the image pyramid corresponding to image A (training image). m ε→i is the result of aggregating all feature points {j: (i, j) ∈ ε} (ε ∈ {ε self , ε cross}). [·||·] represents the concatenation operation. The processing of the target image is similar.

[0104] 3), Attention mechanism: Determine the matching descriptor for training feature points to match with the second image as the target based on the self-attention mechanism and cross-attention mechanism.

[0105] Similar to data retrieval, for an element i, the query queue is q i , according to the attributes of the element (kj ) Retrieve the values v of certain elements (which elements) j . The information is calculated as the weighted average of these values:

[0106]

[0107] where α ij is the similarity of softmax in terms of key query, represented by:

[0108]

[0109] Applied to this embodiment, k j , q i and v j are represented as follows:

[0110]

[0111]

[0112] where, (Q, S) ∈ (A, B) 2 , A and B represent the training image and the target image respectively. W1, W2, W3 represent the network weights, and b i represents the network bias;

[0113] After L times of self / cross attention mechanisms, the matching descriptor corresponding to image A (training image) is obtained. The matching descriptor is represented by:

[0114]

[0115] 4), the loss function:

[0116] After calculating the matching descriptor, the corresponding assignment matrix P can be constructed.

[0117] Optionally, in this embodiment, by maximizing the score ∑ i,j S i,j P i,j the assignment matrix can be obtained, where S i,j is realized by the inner product of f i A and f i B :

[0118] Using the supervised learning method, according to the specification, the optimal model parameters of the graph neural network are finally learned. As an example, for instance, taking the matching of image A and image B as an example, the loss function takes into account the matching feature points and also takes into account the non-matching feature points and The loss function is expressed as follows:

[0119]

[0120] Among them, P i,j represents parameters related to the feature points matched on images A and B, such as the matching degree, etc. P i,N+1 and P M+1,j respectively represent parameters related to the feature points on image A that do not match (do not match the feature points on image B), such as the matching degree, etc., and parameters related to the feature points on image B that do not match (do not match the feature points on image A), such as the matching degree, etc.

[0121] Through the above-mentioned supervised learning, the accuracy and the matching recall rate are maximized. Finally, the optimal model parameters of the graph neural network can be trained to obtain the optimal graph neural network. After that, the trained graph neural network can be used for image matching.

[0122] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described:

[0123] Refer to Figure 6 , Figure 6 which is the structural diagram of the device provided by the embodiments of the present application. The device may include:

[0124] An image processing unit, which is used to divide each constituent image corresponding to each layer of the first image pyramid corresponding to the first image to be matched into multiple image blocks of size M1*M2, extract feature points from at least one image block of the constituent image, and obtain the feature point information corresponding to the feature points; among them, the number of image blocks in the constituent image is greater than the number L of feature points required to be extracted from the constituent image, and when no feature points can be extracted from an image block where the texture does not meet the set requirements, the set threshold T originally used for extracting feature points is adjusted to achieve extracting feature points from an image block where the texture does not meet the set requirements; the feature point information at least includes the confidence that the pixel point in the constituent image is determined as a feature point and the position of the pixel point; each layer of the first image pyramid has a corresponding constituent image, and at least one layer of the corresponding constituent image is obtained by performing image conversion on the first image;

[0125] A determination unit, which is used to determine the feature point descriptor corresponding to each feature point in the constituent image on each layer of the first image pyramid;

[0126] A matching unit, configured to input the feature point information and feature point descriptors corresponding to the feature points in each layer of the composed images in the first image pyramid, and the feature point information and feature point descriptors corresponding to the feature points in each layer of the composed images in the second image pyramid obtained as the target into a trained graph neural network, so that the graph neural network determines the matching relationship between the first image and the second image by using the self-attention mechanism and the cross-attention mechanism and outputs it.

[0127] Optionally, extracting feature points from at least one image patch of the composed image includes:

[0128] Performing the following steps for each image patch in the composed image on each layer in the first image pyramid:

[0129] For each pixel point in the image patch, the gray value of the pixel point is I P , a circle is constructed in the image patch with the pixel point as the center and a set value r1 as the radius, and it is judged whether the boundary of the circle meets the set conditions. The set conditions are that the gray value of each of the n consecutive pixel points on the circle boundary is greater than I P +T or less than I P -T, where T is the set threshold initially set;

[0130] If so, mark the pixel point and determine the confidence of the pixel point as a feature point. Among them, the marked pixel point is recorded as the target pixel point;

[0131] If not, there is no corresponding record of the threshold change times for the pixel point currently, or there is a record of the threshold change times but the recorded threshold change times are less than the preset maximum number of threshold change times, then adjust the set threshold T and return to judge whether the boundary of the circle meets the set conditions;

[0132] Extract feature points from at least one image patch of the composed image according to the target pixel points in each image patch.

[0133] Optionally, extracting feature points from at least one image patch of the composed image according to the target pixel points in each image patch includes:

[0134] For each composed image corresponding to each layer in the first image pyramid, at least one image region is divided from the composed image according to the size of the composed image. The size of the image region is a specified size, and the following steps are performed for each image region:

[0135] Take the image region as the current region; divide the current region into K sub-regions, delete the sub-regions without target pixel points, and judge whether the total number of sub-regions is greater than L;

[0136] If not, for each sub-region, when the number of target pixel points in the sub-region is greater than 1, then use the sub-region as the current region, and return to the step of dividing the current region into K sub-regions. When there is only one target pixel point in the sub-region, determine the target pixel point as the feature point;

[0137] If so, for each sub-region, when the number of target pixel points in the sub-region is greater than 1, select a target pixel point with the highest confidence from the sub-region as the feature point. When there is only one target pixel point in the sub-region, determine the target pixel point as the feature point.

[0138] Optionally, determining the feature point descriptor corresponding to each feature point in each component image of the first image pyramid includes:

[0139] For each feature point, determine the centroid of the neighborhood of the feature point. The neighborhood refers to the image region on the component image that is at a set distance r2 from the feature point, and determine the direction from the feature point to the centroid as the direction of the feature point;

[0140] Rotate the image window centered on the feature point on the component image by θ along the direction of the feature point, select N pairs of pixel points in the rotated image window, and determine the descriptor of the feature point according to the feature values of the two different pixel points in each pair of pixel points under the same feature attribute; where θ is the angle formed by the feature point to the centroid.

[0141] Optionally, determining the centroid of the neighborhood of the feature point includes:

[0142] The centroid of the neighborhood is represented by the following formula:

[0143]

[0144] where m pq =∑ x,y x p y q I(x,y)p,q={0,1}, and I(x,y) is the gray value of the feature point at the position (x,y) in the neighborhood.

[0145] Optionally, the graph neural network is trained in the following manner:

[0146] Based on the network structure of the multi-layer perceptron, combine the feature point information and the feature point descriptor corresponding to the training feature points on each training image to obtain the training feature point coding information; the training feature point coding information is represented by the following formula: x i =d i +MLP(P i ), where di denotes the feature point descriptor corresponding to the i-th training feature point, P i denotes the feature point information corresponding to the i-th training feature point, and MLP represents the network structure of a multi-layer perceptron; the training feature points on the training image include the feature points on each component image in the training pyramid corresponding to the training image;

[0147] Determine the matching descriptors for each training feature point on the training image for matching the training image and the target image based on the self-attention mechanism and the cross-attention mechanism; the matching descriptors are represented by the following formula: where W represents the network weight and b represents the network bias; represents the intermediate representation of the i-th training feature point on the l-th component image in the training pyramid, and A represents the training image;

[0148] Determine the matching matrix based on the matching descriptors of each training feature point; optimize the matching matrix according to the specified loss function to learn the optimal model parameters of the graph neural network; the optimal model parameters at least include the optimal network weight and the optimal network bias.

[0149] Thus, the structure description of the Figure 6 shown device is completed.

[0150] Correspondingly, an embodiment of the present application also provides Figure 6 the hardware structure diagram of the shown device, specifically as Figure 7 shown, and the electronic device can be the device of the above implementation method. As Figure 7 shown, the hardware structure includes: a processor and a memory.

[0151] Among them, the memory is used to store machine-executable instructions;

[0152] The processor is used to read and execute the machine-executable instructions stored in the memory to implement the method embodiment of the corresponding network congestion control as shown above.

[0153] As an embodiment, the memory can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the memory can be: volatile memory, non-volatile memory or similar storage media. Specifically, the memory can be RAM (Radom Access Memory), flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disc, DVD, etc.), or similar storage media, or a combination thereof.

[0154] Thus, the description of the Figure 7 shown electronic device is completed.

[0155] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.

Claims

1. An image matching method, characterized in that, The method includes: For each component image corresponding to each layer in the first image pyramid corresponding to the first image to be matched, divide the component image into multiple image patches of size M1*M2, extract feature points from at least one image patch of the component image, and obtain the feature point information corresponding to the feature points; wherein, the number of image patches in the component image is greater than the number of feature points L required to be extracted from the component image, and when no feature points can be extracted from an image patch where the texture does not meet the set requirements, adjust the set threshold T originally used for extracting feature points to enable feature points to be extracted from the image patch where the texture does not meet the set requirements; the feature point information at least includes the position of the pixel point determined as a feature point in the component image and the confidence level of the pixel point being determined as a feature point; each layer in the first image pyramid has a corresponding component image, and at least one layer of the corresponding component image is obtained by performing image transformation on the first image; For each feature point in the component image on each layer in the first image pyramid, determine the feature point descriptor corresponding to the feature point; Input the feature point information and feature point descriptors corresponding to the feature points in the component images on each layer in the first image pyramid, and the feature point information and feature point descriptors corresponding to the feature points in the component images on each layer in the second image pyramid corresponding to the obtained second image as the target into the trained graph neural network, so that the graph neural network uses the self-attention mechanism and cross-attention mechanism to determine the matching relationship between the first image and the second image and output it.

2. The method according to claim 1, wherein The extracting feature points from at least one image patch of the component image includes: For each image patch in the component image on each layer in the first image pyramid, perform the following steps: For each pixel point in the image block, the gray level of the pixel point is I P , a circle with the pixel point as the center and a set value r1 as the radius is constructed in the image block, and it is judged whether the boundary of the circle meets the set conditions. The set conditions are that the gray level value of each pixel point among n consecutive pixel points on the circle boundary is greater than I P +T or less than I P -T; if so, the pixel point is determined as the target pixel point, and the confidence of the pixel point as a feature point is determined; if not, if there is no corresponding record of the number of threshold changes for the pixel point currently, or there is a record of the number of threshold changes but the recorded number of threshold changes is less than the preset maximum number of threshold changes, the set threshold T is adjusted, and it is returned to judge whether the boundary of the circle meets the set conditions; Extract feature points from at least one image patch of the component image according to the target pixel points in each image patch.

3. The method according to claim 2, characterized in that, The extracting feature points from at least one image patch of the component image according to the target pixel points in each image patch includes: For each component image corresponding to each layer in the first image pyramid, divide at least one image region from the component image according to the size of the component image, the size of the image region is the specified size, and for each image region, perform the following steps: Take the image region as the current region; divide the current region into K sub-regions, delete the sub-regions without target pixel points, and determine whether the total number of sub-regions is greater than the L; If not, for each sub-region, when the number of target pixel points in the sub-region is greater than 1, take the sub-region as the current region and return to the step of dividing the current region into K sub-regions, when there is only one target pixel point in the sub-region, determine the target pixel point as a feature point; If so, for each sub-region, when the number of target pixel points in the sub-region is greater than 1, select the target pixel point with the highest confidence level as a feature point from the sub-region, when there is only one target pixel point in the sub-region, determine the target pixel point as a feature point.

4. The method according to claim 1, wherein For each feature point in the composed image on each layer of the first image pyramid, determining the feature point descriptor corresponding to the feature point includes: For each feature point, determining the centroid of the neighborhood of the feature point, where the neighborhood refers to the image area on the composed image that is at a set distance r2 from the feature point, and determining the direction from the feature point to the centroid as the direction of the feature point; Rotating the image window centered on the feature point on the composed image by θ along the direction of the feature point, selecting N pixel point pairs within the rotated image window, and determining the descriptor of the feature point according to the feature values of two different pixel points in each pixel point pair under the same feature attribute; where θ is the angle formed by the feature point and the centroid.

5. The method according to claim 4, characterized in that, The determining the centroid of the neighborhood of the feature point includes: The centroid of the neighborhood is represented by the following formula: where m pq = ∑ x,y x p y q I(x, y)p, q = {0, 1}, and I(x, y) is the gray value of the feature point at the position (x, y) in the neighborhood.

6. The method according to claim 1, characterized in that, The graph neural network is trained in the following manner: Combining the feature point information and the feature point descriptor corresponding to the training feature points on each training image based on the network structure of the multi-layer perceptron to obtain the training feature point encoding information; The training feature point encoding information is represented by the following formula: x i = d i + MLP(P i ), where d i represents the feature point descriptor corresponding to the i-th training feature point, P i represents the feature point information corresponding to the i-th training feature point, and MLP represents the network structure of a multi-layer perceptron; the training feature points on the training image include the feature points on each component image in the training pyramid corresponding to the training image; Determine the matching descriptors corresponding to each training feature point on the training image for matching the training image and the target image based on the self-attention mechanism and the cross-attention mechanism; the matching descriptors are represented by the following formula: where W represents the network weight and b represents the network bias, represents the intermediate representation of the i-th training feature point on the l-th component image in the training pyramid, and A represents the training image; Determining the matching matrix according to the matching descriptors of each training feature point; optimizing the matching matrix according to the specified loss function to learn the optimal model parameters of the graph neural network; the optimal model parameters at least include the optimal network weights and the optimal network bias.

7. An image matching device, characterized in that, The device includes: An image processing unit, configured to divide the composed image on each layer of the first image pyramid corresponding to the first image to be matched into image blocks of L*M, extract feature points from at least one image block of the composed image, and obtain the feature point information corresponding to the feature points; where the number of image blocks in the composed image is greater than the number of feature points L required to be extracted from the composed image, and when no feature points can be extracted from an image block where the texture does not meet the set requirements, adjusting the set threshold T originally used to extract feature points to enable feature points to be extracted from the image block where the texture does not meet the set requirements; the feature point information at least includes the position of the pixel point determined as a feature point in the composed image and the confidence that the pixel point is determined as a feature point; each layer of the first image pyramid has a corresponding composed image, and the composed images corresponding to different layers have different sizes, and the composed images of different sizes are obtained by performing size conversion on the first image; A determination unit, configured to determine the feature point descriptor corresponding to each feature point in the composed image on each layer of the first image pyramid; A matching unit, configured to input the feature point information and the feature point descriptor corresponding to the feature points in the composed images of each layer of the first image pyramid, and the feature point information and the feature point descriptor corresponding to the feature points in the composed images of each layer of the second first image pyramid corresponding to the obtained second image as the target into the trained graph neural network, so that the graph neural network determines the matching relationship between the first image and the second image as the target by using the self-attention mechanism and the cross-attention mechanism and outputs it.

8. The device according to claim 7, characterized in that, The extracting feature points from at least one image block of the composed image includes: Perform the following steps for each image patch in the composed image on each layer of the first image pyramid: For each pixel in the image block, the gray level of the pixel is I P , a circle with the pixel as the center and a set value r1 as the radius is constructed in the image block, and it is determined whether the boundary of the circle meets the set conditions. The set conditions are that the gray level value of each of the n consecutive pixels on the circle boundary is greater than I P +T or less than I P -T, where T is the set threshold set initially; If so, identify the pixel point and determine the confidence of the pixel point as a feature point, where the identified pixel point is denoted as the target pixel point; If not, there is currently no record of the number of threshold changes corresponding to the pixel point, or there is a record of the number of threshold changes but the recorded number of threshold changes is less than the preset maximum number of threshold changes, then adjust the set threshold T and return to determine whether the boundary of the circle meets the set conditions; Extract feature points from at least one image patch of the composed image based on the target pixel points in each image patch.

9. The device according to claim 8, characterized in that, The extracting feature points from at least one image patch of the composed image based on the target pixel points in each image patch includes: For the composed image corresponding to each layer of the first image pyramid, divide at least one image region from the composed image according to the size of the composed image, where the size of the image region is the specified size, and perform the following steps for each image region: Take the image region as the current region; divide the current region into K sub-regions, delete the sub-regions without target pixel points, and determine whether the total number of sub-regions is greater than L; If not, for each sub-region, when the number of target pixel points in the sub-region is greater than 1, take the sub-region as the current region and return to the step of dividing the current region into K sub-regions, and when there is only one target pixel point in the sub-region, determine the target pixel point as a feature point; If so, for each sub-region, when the number of target pixel points in the sub-region is greater than 1, select a target pixel point with the highest confidence as a feature point from the sub-region, and when there is only one target pixel point in the sub-region, determine the target pixel point as a feature point.

10. An electronic device, characterized in that, The electronic device includes: a processor and a memory; Wherein, the memory is used to store machine-executable instructions; The processor is used to read and execute the machine-executable instructions stored in the memory to implement any method as claimed in claims 1 to 6.

Citation Information

Patent Citations

  • Image feature matching method and device, computer equipment and storage medium

    CN113139490A

  • Feature point extraction method and device, image reconstruction method and device

    CN113837202A