Image matching method and device based on graph neural network fusion model
By adopting a method based on graph neural network fusion model in image matching, using entity detection and graph structure to extract spatial relationship features, the problem of low image matching accuracy in complex scenarios is solved, and higher matching accuracy and stability are achieved.
Patent Information
- Application Number
- CN202010622911.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-06-30
AI Technical Summary
In complex scenes, the image matching accuracy is not high, and the variable information interference is large, resulting in a decrease in image matching accuracy.
The image matching method based on the graph neural network fusion model is adopted to construct the graph structure to extract spatial relationship features through entity detection and visual feature extraction, and the visual feature is fused with spatial relationship features, and the similarity of image pairs is measured using Euclidean distance.
Effectively eliminate interference from irrelevant content in the image, improve the accuracy and stability of image matching, and is suitable for image matching tasks in different scenarios.
Smart Images

Figure CN113869338B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image matching, and in particular to an image matching method and device based on a graph neural network fusion model. Background Art
[0002] Image matching refers to a method that determines the similarity and consistency of image pairs by the correspondence between image content, features, structure, etc. Image matching has a wide range of applications in the field of computer vision, such as image-based navigation, location recognition, and closed-loop detection in SLAM. Image matching methods can be mainly divided into two categories: the first category is based on artificially designed features, using image descriptors to represent local features, such as SIFT, SURF, and ORB; and further using aggregation models to aggregate all descriptor features in the image, such as Bag-of-Words (BoW), Vector of Locally Aggregated Descriptors (VLAD), and Fisher Vector (FV). The similarity of image pairs is obtained by calculating the distance between image aggregated descriptors. The second category is based on deep learning methods, which usually use convolutional neural networks (CNNs) to extract image features. Convolutional neural networks can mine the intrinsic features of images layer by layer, and have achieved better performance in image matching than methods based on artificially designed features. The deep learning-based method mainly includes two steps: feature extraction and similarity measurement. The input of the network is an image pair, and the output is the similarity score of the image pair. Commonly used structures include Siamese network structure, matchnet, etc.
[0003] Most current image matching algorithms focus on the entire image, but some images may contain a large amount of variable information, which limits the accuracy of image matching. Figure 1 The two images (a) and (b) contain a large number of pedestrians, vehicles and other volatile contents. These contents vary greatly at different times, so they may cause great interference to image matching. Figure 1 In image pairs, only a small number of geographical entities are relatively fixed, such as buildings, traffic signs, etc. These geographical entities play a positive role in image matching. Therefore, how to ignore the changeable pedestrians, vehicles and other contents and focus on static geographical entities has a good promotion to improve the image matching accuracy. Summary of the invention
[0004] The purpose of this application is to provide an image matching method based on a graph neural network fusion model to solve the problem of low image matching accuracy in complex scenes.
[0005] To achieve the above object, the solution of the present invention includes:
[0006] An image matching method based on a graph neural network fusion model, the steps are as follows:
[0007] 1) Using an entity detection algorithm to detect entity targets in an image, and obtaining image blocks containing entities; and using a visual feature extraction network to extract visual features of the entity;
[0008] 2) All entities in the image are constructed into a graph, and the graph neural network is used to extract spatial relationship features;
[0009] 3) Aggregating the visual features of the entity to obtain a visual feature aggregation vector; using a feature fusion network, fusing the visual feature aggregation vector with the spatial relationship feature to obtain a feature vector of the entire image, and calculating the Euclidean distance of the feature vector to measure the similarity between the image pairs. If the Euclidean distance of the feature vector of the image pair is less than a threshold, the image pair is predicted to be a match, otherwise it is a mismatch.
[0010] Furthermore, the entity detection algorithm is an Edge Boxers algorithm.
[0011] Furthermore, the visual feature extraction network is a deep convolutional network based on a twin network structure.
[0012] Furthermore, constructing all entities in the image into a graph includes: taking all entities in an image as nodes of the graph; connecting all entities in pairs as edges of the graph; and taking the ratio of the Euclidean distance between the center points of the entities to the length of the diagonal of the image as the weight of the edge.
[0013] Furthermore, the bag-of-words model is used to aggregate the visual features of the entity image blocks.
[0014] Furthermore, the feature fusion network is a fully connected network based on a twin network structure.
[0015] In addition, the present invention also provides an image matching device based on a graph neural network fusion model, comprising a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the above method.
[0016] In order to solve the problem that there are many objects in complex scenes and it is difficult to extract discriminative features between image pairs, the present invention proposes an image matching method based on a graph neural network fusion model. The present invention does not start from the perspective of directly matching the entire image, but converts the image matching task into entity matching, which effectively avoids the interference of irrelevant content in the image on the matching. Experimental results show that the image matching method based on the graph neural network fusion model achieves better matching performance on most subsets of the data set, and has good stability and versatility for image matching tasks in different scenarios. Compared with the matching of the entire image, the use of entities for matching can effectively eliminate the interference of non-corresponding content and achieve higher matching accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the matched image pairs and the corresponding entities between the images;
[0018] Figure 2 It is a diagram of the visual feature extraction network structure;
[0019] Figure 3 It is a schematic diagram of the method of constructing the graph;
[0020] Figure 4 It is a graph neural network structure diagram;
[0021] Figure 5 It is the BoW-GNN fusion network structure;
[0022] Figure 6 It is the overall flow chart of the present invention. DETAILED DESCRIPTION
[0023] To solve the problems raised by the background technology, the entities in the matching image pair can be used to complete the matching task of the entire image. At the same time, the development of graph neural network (GNN) in recent years has provided a good tool for the representation learning of graph structured data, and also provided a new idea for the extraction of spatial relationship features of entities in images.
[0024] The method of the present invention uses entities to perform matching tasks, which can effectively eliminate the interference of irrelevant content in the image. The present invention uses graph structure data based on entity construction to train the graph neural network, and the obtained network model can effectively extract the spatial relationship between entities in the image. After fusing the visual features and spatial relationship features of the entity, the similarity between the image pairs is measured to determine whether the image pairs match. Experimental results show that compared with existing methods, the image matching method proposed in the present invention has better performance. The following is a detailed description:
[0025] The image matching method based on the graph neural network fusion model mainly includes three parts:
[0026] 1) Visual feature extraction
[0027] The Edge Boxes algorithm is used to detect entity targets in images and obtain image blocks containing entities: After using the Edge Boxes algorithm for entity detection in each image, multiple image blocks containing entities will be obtained. Then, the visual feature extraction network is used to extract the visual features of the entity.
[0028] Entities are detected based on the Edge Boxes algorithm, which groups the edges of the same objects based on the edge clustering method and detects entities based on the bounding boxes of the edge groups.
[0029] In order to make the entity targets detected by the algorithm more prominent, the minimum score of the detection box is set to 0.11, the minimum size of the detected image block is 64x64, and the maximum size is 256x256; in order to keep the number of detected entities within an appropriate range, the maximum number of entity detections is set to 50.
[0030] The visual feature extraction network structure is as follows Figure 2 As shown, it is built based on the twin network structure, including two branches with the same structure and shared weights.
[0031] In order to fully extract the visual features of the entity, each branch of the constructed twin network consists of 5 convolutional layers and 2 fully connected layers. At the same time, a hole convolution layer with the same convolution kernel size is set in parallel with the first three convolutional layers to expand the receptive field of the convolution kernel. The network structure is shown in Table 1:
[0032] Table 1 shows the parameters of each layer of each branch of the twin network. The output size is height × width × channel. Specifically, C(3) represents a convolutional layer with a filter size of 3×3 (step size of 1), MP(2) represents a maximum pooling layer with a size of 2×2 (step size of 2), R represents a nonlinear activation layer, and the rectified linear unit (ReLU) is selected to complete the nonlinear activation operation. L2_norm represents the normalization operation of the output features, and F represents the fully connected layer. The Dilated_Conv layer is a dilated convolutional layer with a dilated interval of 2.
[0033] Table 1
[0034]
[0035] The network gradually extracts visual features at different levels by performing convolution operations on the input image blocks layer by layer, so that the output feature vector has a strong degree of discrimination, which is conducive to the calculation of loss values using loss functions during training.
[0036] In order to reduce the distance between matching entity features and increase the distance between mismatching entity features during training, we use the contrastive loss function during training, using the Euclidean distance between feature vectors to evaluate the similarity between image pairs. If the distance is small, the similarity between image pairs is high; if the distance is large, the similarity between image pairs is low.
[0037] For all image sets T = {X i}, is a pair of image feature vectors, Y is the training label of the image pair, if and If it is matched, then Y=1; if and If it does not match, then Y=0. express and The Euclidean distance of
[0038]
[0039] Can be abbreviated as D W , we define the loss function as:
[0040]
[0041] Where m (m>0) is the boundary value. The loss function L only works when the distance between non-corresponding samples is less than the boundary value. If the boundary value is not set, all D W It tends to be a constant 0, which makes it impossible to continue the training. In the present invention, m is set to 2.0 during the training process.
[0042] In order to have a network model with good usability when performing entity visual feature extraction operations, we use the Multi-View Stereo dataset (MVS) to pre-train the network. The dataset contains about 1.5 million image blocks and 500,000 3D points. All image blocks are grayscale images of size 64×64, and each image block is associated with a specific 3D point. If two image blocks (a sample image pair) are associated with the same 3D point, the two image blocks are corresponding (positive samples), otherwise they are not corresponding (negative samples). The training label indicates the matching of the image pairs. The matching image pair label is 1, and the mismatched label is 0.
[0043] The actual training uses the sample list provided with the data set, where the number of positive samples and negative samples is equal. The training process of this embodiment uses 50,000 samples from the Liberty subset of the data set, divided into 1,000 batches, 50 samples per batch, and trained for 200 epochs. During the training process, the network is optimized using the stochastic gradient descent method, and the network model with the minimum loss value of the verification data is saved. Because the two branches of the network have the same structure and share weights, only the model of one branch needs to be saved.
[0044] 2) Spatial relationship feature extraction
[0045] Before extracting features of the spatial relationship of objects in an image, we first need to construct all entity representations in the image in the form of a graph. Considering that there are a small number of entities in each image (less than 50), to facilitate the training of the graph network, we use all entities in an image as nodes of the graph; connect all entities in pairs as edges of the graph; and use the ratio of the Euclidean distance between the center points of the entities to the length of the diagonal of the image as the weight of the edge (so that the weight remains in the range of 0-1). The construction of the graph is as follows: Figure 3 shown.
[0046] In the figure, the lines are the edges of the graph, and the ratio of the Euclidean distance between the center points of the entities to the diagonal length of the entire image is defined as the weight of the edge. The figure only shows the connections of some entities.
[0047] Based on the spatial relationship of entities in the image, they are constructed in the form of a graph, and the graph neural network is used to extract the spatial relationship features. The graph neural network structure used in the present invention is as follows: Figure 4 shown.
[0048] For the tth network layer in k-GNN, the calculation process of each node feature is as follows:
[0049]
[0050] Among them, W 1 (t) is, W 2 (t) is the parameter matrix, σ represents the nonlinear activation function (ReLU is used in this invention); s represents a subgraph consisting of k nodes, u is the neighbor subgraph of this subgraph s, N L (s) is the local neighborhood of s; for the node set V(G) of graph G, the definition of the neighbor subgraph is as follows:
[0051] N(s)={t∈[V(G)] k ||s∩t|=k-1}
[0052] For a subgraph consisting of k nodes, its neighbor subgraph must have only k-1 common nodes with it. For a k-layer GNN, the features learned by the (k-1)-layer GNN are used as the input features of this layer. After each layer of GNN, the output is pooled and the outputs of each network layer are aggregated through the fully connected network layer.
[0053] In graph-level classification tasks, using a multi-layer structure can achieve higher-level information aggregation. Compared with a single-layer GNN algorithm, using a multi-layer k-GNN structure can hierarchically combine graph feature representations learned at different granularities.
[0054] 3) Fusion of visual features and spatial relationship features
[0055] In order to fully describe the image information, we fuse the visual features of the entity with the spatial features between them to further enhance the algorithm's ability to represent image features. After using the Edge Boxes algorithm for object detection in each image, multiple image blocks containing entities will be obtained. In order to perform further feature fusion, a single vector representation of the visual features of multiple entities in each image is required. The present invention uses the Bag of Words (BoW) model to aggregate entity visual features.
[0056] In order to fuse the visual feature aggregation vector obtained by the bag-of-words model with the spatial relationship feature vector obtained by the graph neural network, and perform feature fitting on the fused vector, the present invention uses a 4-layer fully connected network to construct a feature fusion network. In order to enable the feature fusion network to obtain the ability to distinguish positive and negative samples, the graph neural network is connected to the feature fusion network and trained using the twin network model. The overall network structure is as follows: Figure 5 As shown in the figure. Based on the twin network structure, the spatial relationship feature network and the feature fusion network are jointly trained to simplify the training process and reduce the computational overhead. During the training process, the number of training samples for each subset is 16,000, each batch has 100 samples, and the training is repeated for 1,500 epochs. During the training process, the stochastic gradient descent method is used to optimize the network.
[0057] After fusing the visual features and spatial relationship features of the entity, the Euclidean distance of the feature vector is calculated to measure the similarity between the image pairs, and a threshold is set for the distance. If the Euclidean distance of the image pair feature vector is less than the threshold, the image pair is predicted to be a match.
[0058] In summary, in order to solve the problem that there are many objects in complex scenes and it is difficult to extract discriminative features between image pairs, the present invention proposes an image matching method based on a graph neural network fusion model. The present invention does not start from the perspective of directly matching the entire image, but converts the image matching task into entity matching, which effectively avoids the interference of irrelevant content in the image on the matching. Figure 6 As shown, the method of the present invention mainly includes three parts: 1) visual feature extraction (using the Edge Boxes algorithm to detect entity targets in the image, and using a convolutional network trained based on a twin network to extract the visual features of the entity); 2) spatial relationship feature extraction (based on the spatial relationship of the entity in the image, it is constructed in the form of a graph, and the graph neural network is used to extract its spatial relationship features). 3) Fusion of visual features and spatial relationship features (using a feature fusion network to fuse the visual features and spatial relationship features of the entity, and using Euclidean distance to complete the discrimination of image matching relationships). The present invention first uses the Edge Boxes algorithm to detect entities that may be contained in the image, and uses a pre-trained feature extraction network to extract the visual features of each entity; then, based on the defined spatial relationship model, each entity is represented as a graph structure and input into the twin graph neural network; finally, the obtained image vector representation is fused with the image BoW (Bag of Words) bag of words model feature vector, and the matching relationship between image pairs is discriminated based on the distance between the fused vectors. The feature vector obtained by the convolutional neural network describes the visual features of the entity itself; the feature vector obtained by the graph neural network describes the spatial relationship features between entities in the image. These two features are complementary in describing the relationship between image pairs, so the fusion model has a stronger ability to solve image matching problems. Experimental results show that the image matching method based on the graph neural network fusion model has achieved better matching performance on most subsets of the data set, and has good stability and versatility for image matching tasks in different scenarios. Compared with matching the entire image, using entities for matching can effectively eliminate the interference of irrelevant content and achieve higher matching accuracy.
[0059] The specific entity detection algorithm, convolutional neural network structure, aggregation model, feature fusion network structure, etc. adopted in this embodiment are all preferred methods. As other implementation methods, they can also be implemented using similar technologies in the prior art.
[0060] The present invention also provides an image matching device based on a graph neural network fusion model, including a memory and a processor, and a computer program stored in the memory and running on the processor. The processor is coupled to the memory, and the processor is used to run the program instructions stored in the memory to implement the above method. Since the method is described clearly and completely in the method example, it will not be repeated in this embodiment.
[0061] That is, the method in the above method embodiment can be implemented by computer program instructions. These computer program instructions can be provided to a processor (such as a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device, etc.), so that the processor executes these instructions to generate functions specified by the above method.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An image matching method based on graph neural network fusion model, It is characterized in that Here are the steps: 1) Using an entity detection algorithm to detect entity targets in an image to obtain entity image blocks; using a visual feature extraction network to extract visual features of the entity image blocks; 2) All entity image blocks in the image are formed into a graph, and the graph neural network is further used to extract spatial relationship features; The method of forming a graph from all entity image blocks in an image includes: taking all entities in an image as nodes of the graph; connecting all entities in pairs as edges of the graph; taking the ratio of the Euclidean distance between the center points of the entities to the length of the diagonal line of the image as the weight of the edge; 3) Aggregating the visual features of the entity image block to obtain a visual feature aggregation vector; using a feature fusion network, fusing the visual feature aggregation vector with the spatial relationship feature to obtain a feature vector of the image pair after fusion, and calculating the Euclidean distance of the feature vector to measure the similarity between the image pairs. If the Euclidean distance of the feature vector of the image pair is less than a threshold, the image pair is predicted to be a match, otherwise it is a mismatch.
2. According to claim 1, the image matching method based on the graph neural network fusion model, It is characterized in that The entity detection algorithm is the Edge Boxes algorithm.
3. According to claim 2, the image matching method based on the graph neural network fusion model, It is characterized in that The visual feature extraction network is a deep convolutional network based on a twin network structure.
4. The image matching method based on the graph neural network fusion model according to claim 1, It is characterized in that The bag-of-words model is used to aggregate the visual features of entity image patches.
5. According to claim 4, the image matching method based on the graph neural network fusion model, It is characterized in that The feature fusion network is a deep convolutional network based on a twin network structure.
6. An image matching device based on a graph neural network fusion model, comprising a processor and a memory, It is characterized in that The processor executes the computer program stored in the memory to implement the method according to any one of claims 1 to 5.