Lightweight graph matching model training method, graph matching method and system based on state space model
Patent Information
- Application Number
- CN202511103272.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-08-07
AI Technical Summary
这些限制使得现有GNN方法难以在智能摄像头、穿戴设备等典型边缘场景中落地
[0031]图神经网络:图神经网络是一类专门用于处理图结构数据的深度学习模型,能够捕捉节点之间的关系(边)和全局拓扑信息,广泛应用于社交网络、分子结构、推荐系统等场景。
Smart Images

Figure CN121170543B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, specifically to a lightweight graph matching model training method, graph matching method, and system based on a state-space model. Background Technology
[0002] Graph matching, as a key technology in graph data processing, has driven the development of graph representation learning, provided new approaches to solving problems with non-Euclidean spatial data, and contributed to the improvement of the theoretical framework of geometric deep learning. In practical applications, graph matching technology has been widely used in various fields such as computer vision, bioinformatics, and social network analysis, as well as important scenarios such as medical image analysis, protein network alignment, and cross-platform user identification. Graph matching not only effectively handles the irregular features of graph-structured data but also demonstrates significant value in numerous industrial applications.
[0003] Traditional graph matching methods are mainly divided into two categories: (1) Combinatorial optimization methods: The matching problem is modeled as a quadratic assignment problem and solved using integer programming, spectral methods, etc. These methods have strong theoretical guarantees, but the computational complexity is as high as O(n³) or higher, making it difficult to extend to large-scale graphs; (2) Deep learning-based methods: Deep learning is used to automatically learn graph features and matching strategies using graph neural networks. Although the performance is superior, the graph neural networks of existing methods have the problem of over-parameterization, with the number of parameters often reaching millions, which cannot meet the real-time requirements of mobile / edge devices.
[0004] Current lightweighting techniques mainly revolve around deep neural network model compression, including pruning, quantization, and knowledge distillation. However, directly applying these techniques to graph matching can lead to a significant decrease in topology awareness. For example, pruning may disrupt the crucial adjacency matrix structure, and quantization may lose subtle differences in node features.
[0005] While deep learning-based graph neural networks (GNNs) perform well in modeling graph-structured data, their practical deployment still faces significant challenges, especially in resource-constrained edge computing scenarios. Typical GNN architectures (such as GAT and GIN) can have millions of parameters per layer, and the model size grows exponentially after multiple layers are stacked. For example, the standard GraphSAGE in a fully connected configuration has more than 4.7M parameters when processing the Cora dataset, while the more complex GATv2 model has as many as 21M parameters on the OGB-arxiv dataset. The large number of parameters directly leads to an increase in model file size, exceeding 500MB, far exceeding the capacity of the 1-4GB of available memory on edge devices such as Raspberry Pi 4B. These limitations make it difficult for existing GNN methods to be implemented in typical edge scenarios such as smart cameras and wearable devices. (2) Directly applying network lightweighting techniques (such as pruning and quantization) to the model can lead to the loss of key topological feature information, which seriously affects the performance of graph matching. Specifically, while these lightweight methods can effectively compress model size and improve inference speed, they can damage the topological relationships and higher-order semantic features in graph-structured data. Pruning operations may remove neural connections that are sensitive to topological features; the quantization process introduces numerical errors, causing deviations in the similarity calculation between nodes, ultimately leading to a decrease in the accuracy of graph matching. Summary of the Invention
[0006] The purpose of this invention is to provide a lightweight graph matching method and system based on a state-space model to solve at least one of the technical problems existing in the background art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides a lightweight graph matching model training method based on a state-space model, comprising:
[0009] We obtained open-source datasets for multiple graph matching tasks as training sets and preprocessed the training sets: we selected photo pairs A and B with the same number of point pairs in the same category, recorded the true matching relationship of the key point sequences in the photo pairs, and then shuffled the feature sequences of the two key points.
[0010] The preprocessed point-pair feature sequence is converted into a tensor data type that can be learned by the neural network. The initial key point feature sequence and position feature sequence of Figure A and Figure B are randomly set to zero respectively, and high-dimensional mapping is performed. The key point features and position features are summed element by element to form an initial feature sequence that can be updated and learned.
[0011] The initial feature sequences of graphs A and B are expanded by rows and columns respectively to form four different expanded sequences. The adjacency matrix structure of the graphs is learned, and the initial feature sequences of the two graphs are updated using the adjacency matrix.
[0012] The updated feature sequences of Figure A and Figure B are concatenated to form a fused feature sequence that learns from each other. The prediction path of the sequence is reconstructed, and the matching matrix is learned. Different matching matrices are output for each path. All output matching matrices are summed element-wise, averaged, and the final soft correspondence matching matrix is calculated. The feature sequences of Figure A and Figure B are updated again using the obtained soft correspondence matching matrix.
[0013] The predicted matching matrix is calculated using the Hungarian algorithm to obtain a discrete Boolean matrix that represents the one-to-one correspondence between the key points of Figure A and Figure B.
[0014] Using the obtained Boolean matrix, the loss during the model learning process is calculated through the cross-entropy loss function, and gradient backpropagation is performed to train and update the model weights, and the process is iterated and optimized.
[0015] As a further limitation of the first aspect of the present invention, the VGG16 network is used to extract key point visual features, and the key point visual features and position information are encoded and randomly zeroed in the feature dimension; the key point coordinates are converted into high-dimensional position embeddings, the position is encoded as a 2D mapping to 128 dimensions, the two input key point feature sequences and position feature sequences are converted into tensor types, and the key point feature sequences and position feature sequences are summed element-wise to obtain the enhanced feature sequence.
[0016] As a further limitation of the first aspect of the present invention, the initial n×d dimensional initial feature sequence is expanded by extending it in both the row and column directions by a factor of n times the length of the point sequence, resulting in n 2 For a given keypoint sequence of an image, generate two distinct n-dimensional feature sequences. 2 ×d-dimensional feature sequences, each image pair contains two images, for a total of four n... 2 ×d-dimensional feature sequences; the expanded four feature sequences are then used to learn graph structure features, i.e., adjacency matrices, to obtain the row-predicted adjacency matrices and column-predicted adjacency matrices for graph A, and the row-predicted adjacency matrices and column-predicted adjacency matrices for graph B, for a total of four n×n adjacency matrices predicted by rows and columns respectively; for each graph with different row and column expansions, there are two different corresponding adjacency matrices. These two adjacency matrices are summed element-wise and averaged to form a symmetric matrix, i.e., the final adjacency matrix; this operation is performed simultaneously on graph A and graph B, and the corresponding initial feature sequences are updated using their respective adjacency matrices, and then summed element-wise with the positional feature sequences to obtain the updated feature sequences of graph A and graph B.
[0017] As a further limitation of the first aspect of the present invention, the updated two feature sequences are expanded in terms of row and column dimensions, and the expanded feature sequences are concatenated along the feature dimensions to obtain n that can be used to learn matching relationships. 2 The fusion feature sequence is 2d×2D. Two zigzag continuous prediction paths are manually set, and the fusion feature sequence is rearranged according to the prediction paths and fed into the multi-path Mamba module for matching matrix prediction. Each path outputs a prediction result of a matching matrix. The average of each output matching matrix is calculated and passed through the Sinkhorn algorithm to obtain the final n×n soft correspondence matrix, which is the model output. The matching matrix is used to update the two feature sequences of Figure A and Figure B respectively. Finally, the updated feature sequence is summed element-wise with the feature sequence updated by the adjacency matrix and fed into the ReLU function for activation, which serves as the initial feature sequence for the next round.
[0018] As a further limitation of the first aspect of the present invention, the predicted soft matching matrix and the true matching matrix are vectorized, and the model loss is calculated using the cross-entropy loss function. A weight w=5 is used to balance the positive and negative samples. This loss is used to quantify the difference between the model's prediction result and the true result. By minimizing the loss function, the optimal value of the model parameters is found, thereby improving the model's prediction accuracy. During the training process, the weighted cross-entropy is calculated as a supervision signal, and the model parameters are optimized through gradient backpropagation. The model parameters are continuously adjusted to reduce the value of the loss function, and the model gradually learns the inherent rules and features of the graph feature data.
[0019] Secondly, the present invention provides a lightweight graph matching method, comprising:
[0020] Get the image pairs to be matched;
[0021] The acquired image pairs are processed using a pre-trained matching model to obtain graph matching results; wherein the matching model is obtained using the model training method described in the first aspect.
[0022] Thirdly, the present invention provides a lightweight graph matching system, comprising:
[0023] The acquisition module is used to acquire image pairs to be matched;
[0024] The processing module is used to process the acquired image pairs using a pre-trained matching model to obtain image matching results; wherein the matching model is obtained using the model training method described in the first aspect.
[0025] Fourthly, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the lightweight graph matching method as described in the second aspect.
[0026] Fifthly, the present invention provides a computer device including a memory and a processor, the processor and the memory communicating with each other, the memory storing program instructions executable by the processor, and the processor calling the program instructions to execute the lightweight graph matching method as described in the second aspect.
[0027] In a sixth aspect, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the lightweight graph matching method as described in the second aspect.
[0028] Graph matching: Graph matching is an important task in computer vision, pattern recognition and machine learning, which aims to find the node correspondence between two graph structures so that their topology and node / edge attributes are as consistent as possible.
[0029] State-space model (SSM): SSM is a mathematical model used to describe dynamic systems and is widely used in control theory, signal processing, time series analysis, machine learning, and other fields. It links the system's state (a set of hidden variables) with observation data, and describes the system's evolution process through state equations and observation equations.
[0030] Lightweighting: Lightweighting refers to achieving efficient, low-power, and low-latency computing tasks in resource-constrained environments such as mobile devices, embedded systems, and edge computing through algorithm optimization, model compression, or hardware adaptation. Its core objective is to balance performance and resource consumption, making it suitable for scenarios with high real-time requirements or demanding deployment conditions.
[0031] Graph Neural Networks: Graph neural networks are a class of deep learning models specifically designed to process graph-structured data. They can capture the relationships (edges) between nodes and global topological information, and are widely used in scenarios such as social networks, molecular structures, and recommender systems.
[0032] The beneficial effects of this invention are: it reduces the memory usage of model inference, requiring only a single read / write operation; the model has a very low number of parameters, resulting in less hardware resource consumption and making it more suitable for deployment and inference in environments with limited hardware resources; by reducing the graph task to a sequence prediction task and performing feature learning, the computational load required for model training and inference is effectively reduced, as is the memory consumption required for inference, while ensuring matching accuracy.
[0033] The advantages of additional aspects of the invention will be set forth more clearly in the following description or will be learned by practice of the invention. Attached Figure Description
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a flowchart of the lightweight graph matching method based on the state-space model according to an embodiment of the present invention.
[0036] Figure 2 This is a schematic diagram of the directional Mamba module network structure according to an embodiment of the present invention. Detailed Implementation
[0037] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0038] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0039] It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as here.
[0040] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.
[0041] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0042] To facilitate understanding of the present invention, the present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments. However, the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0043] Those skilled in the art should understand that the accompanying drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0044] This invention addresses the application bottlenecks of existing graph matching methods in resource-constrained scenarios by proposing a lightweight graph matching network design method based on a State Space Model (SSM). Traditional graph matching networks typically rely on complex network architectures, making them difficult to run in real-time on lightweight devices. Existing deep learning models have a large number of parameters and high inference memory consumption. Furthermore, some lightweight methods, such as pruning and quantization, significantly reduce accuracy when applied to graph matching, especially for topology-sensitive tasks, where the lack of an efficient feature transfer mechanism leads to severe loss of structural information. To solve these problems, this invention proposes a graph neural network design method that creatively introduces the temporal modeling capabilities of SSM into the optimization of graph matching tasks through a sequential modeling approach. This reduces the dimensionality of the two-dimensional graph matching task to a one-dimensional sequence prediction task, utilizing the temporal modeling capabilities of SSM to process graph structural information. Simultaneously, benefiting from the hardware-aware capabilities of the SSM architecture, this invention significantly reduces the number of parameters and memory consumption required for inference while maintaining matching accuracy compared to existing methods. It also reduces the resource consumption of graph matching tasks and expands their application scope, enabling them to better meet the needs of real-world scenarios. This invention uses a Mamba-based SSM to design the network structure. Thanks to the SSM architecture, this invention greatly reduces the number of model parameters; avoids O(n2) computational overhead; maintains inference speed and accuracy; and Mamba's hardware-friendly awareness significantly reduces memory consumption during inference, enabling it to run in resource-constrained environments.
[0045] Example 1
[0046] In this embodiment 1, a lightweight graph matching system is first provided, including: an acquisition module for acquiring image pairs to be matched; and a processing module for processing the acquired image pairs using a pre-trained matching model to obtain graph matching results.
[0047] In this embodiment, a lightweight graph matching method is implemented using the system described above, including: acquiring image pairs to be matched using an acquisition module; and processing the acquired image pairs using a pre-trained matching model using a processing module to obtain graph matching results.
[0048] like Figure 1 As shown, in this embodiment, the matching model is a lightweight graph matching model based on a state-space model, and its training method includes the following steps:
[0049] S1: Three open-source datasets for graph matching tasks—Willow, Pascal VOC, and Spair-71K—were used as training and testing samples. Preprocessing of the model data samples mainly included: selecting photo pairs with the same number of point pairs within the same category: Image A and Image B; recording the true matching relationships of the keypoint sequences in the photo pairs; then shuffling the feature sequences of the two keypoints before feeding them into the model for training. In testing, a certain number of samples from each category were selected from the test samples for testing, and the model's prediction accuracy was calculated.
[0050] S2: The model receives the initial point-pair feature sequence input, converts it into a tensor data type that can be learned by the neural network, randomly sets the initial key point feature sequence and position feature sequence of Figure A and Figure B to zero respectively, performs high-dimensional mapping, and sums the key point features and position features element by element to form an initial feature sequence that can be updated and learned.
[0051] S3: Expand the initial feature sequences of graphs A and B by row and column respectively to form four different expanded sequences. These expanded sequences are then fed into four Mamba modules (deep learning modules based on state-space model (SSM)) to learn the adjacency matrix structure of the graphs and update the initial feature sequences of the two graphs using the adjacency matrix.
[0052] S4: The updated feature sequences of Figure A and Figure B are concatenated to form a fused feature sequence that learns from each other. Two different continuous zigzag prediction trajectories are manually set to reconstruct the prediction path of the sequence. The sequence is then fed into the multi-path Mamba module to learn the matching matrix. Different matching matrices are output for each path. Finally, all output matching matrices are summed element-wise, averaged, and the final soft correspondence matrix is obtained through the Sinkhorn algorithm. The obtained soft correspondence matching matrix is then used to update the feature sequences of Figure A and Figure B again.
[0053] S5: The matching matrix predicted by the model is calculated using the Hungarian algorithm to obtain the final discrete Boolean matrix representing the one-to-one correspondence between the key points of Figure A and Figure B.
[0054] S6: Using the obtained Boolean matrix, calculate the loss during the model learning process through the cross-entropy loss function, perform gradient backpropagation, train and update the model weights, iterate and optimize, and continuously improve the prediction accuracy of the model.
[0055] S7: Using the obtained Boolean matrix, compare it with the correspondence of the true points recorded during sample preprocessing to calculate the matching accuracy of the model prediction. Compare it with other graph matching to verify the accuracy of the model prediction.
[0056] Step S1 specifically includes the following steps:
[0057] S11: Split the entire dataset into a training set and a test set in an 8:2 ratio. The training set is used for model training; the test set is used to evaluate model performance. During training, randomly select a category from the training set, and randomly select two images with the same number of keypoints from that category to form a pair of images to be predicted: Image A and Image B.
[0058] S12: To prevent the model from memorizing specific sequences during training, the recorded keypoint correspondences need to be shuffled. This operation increases the randomness of training, making the model more robust.
[0059] S13: During testing, 1000 image pairs are uniformly selected from each category in the test set for model prediction and the prediction accuracy is calculated. Finally, the average accuracy of all categories is calculated to obtain the test accuracy of the model on this dataset.
[0060] Step S2 specifically includes the following steps:
[0061] S21: Use the VGG16 network to extract keypoint visual features, and randomly zero out the keypoint visual features and location information in the feature dimension with a ratio of 0.3 to reduce overfitting. Design a learnable location encoder, specifically a fully connected layer network structure, which upscales the image's two-dimensional coordinates to a feature dimension that matches the visual features. After upscaling, the keypoint feature sequence and the location feature sequence are summed element-wise to obtain an enhanced feature sequence.
[0062] The optimization is performed iteratively during each training session, as shown in the following formula:
[0063] S A =MLP(F A )+MLP(P A ),
[0064] Among them, S A ∈R n×d Here, n represents the number of keypoints, d represents the feature dimension, and the position encoder is a multilayer perceptron (MLP) used to map two-dimensional coordinates into d-dimensional embedding vectors. The same MLP is applied to Figure B to generate the feature sequence S. B ∈R n×d .
[0065] S22: Input the enhanced feature sequence into the designed network for training. The training batch size is 32, the number of model layers is 3, the feature dimension inside the fully connected layer is 128, the feature dimension of the Mamba module used is 1024, the training epochs are set to 200 epochs, each epoch iterates 100 times, the initial training learning rate of the model is 0.001, and after each training epoch, it continuously decays at a rate of 0.98 until it reaches a fixed value of 0.0005 and no longer decreases.
[0066] Step S3 specifically includes the following steps:
[0067] S31: Expand the initial n×d dimensional feature sequence by expanding it in both row and column directions by a factor of n times the length of the point sequence, to obtain n 2 For a given keypoint sequence of an image, generate two distinct n-dimensional feature sequences. 2 ×d-dimensional feature sequences, each image pair contains two images, for a total of four n... 2 ×d-dimensional feature sequences.
[0068] S32: Then, the four expanded feature sequences are respectively fed into the designed directional Mamba module, such as... Figure 2 As shown, this module receives an input sequence (B, N, dim = d) and produces an output sequence (B, N, dim = 1). Specifically, the input sequence first passes through two linear projection layers. One branch enters a local convolutional layer to capture local features and enhance its understanding of the sequence data. Then, it connects to an SSM layer, and the output is merged with another branch. A non-linear activation function is applied to increase the model's expressive power. Finally, it enters the linear projection layer to obtain the output. It learns graph structure features, i.e., the adjacency matrix (predicting the probability of edges existing between its own vertices), outputting four n×n adjacency matrices: the row-predicted adjacency matrix and the column-predicted adjacency matrix for graph A, and the row-predicted adjacency matrix and the column-predicted adjacency matrix for graph B.
[0069] S33: For each graph with different row and column expansions, there are two different corresponding adjacency matrices. The elements of these two adjacency matrices are summed and averaged to form a symmetric matrix, which is the final adjacency matrix. This operation is performed simultaneously on graphs A and B, and the corresponding initial feature sequences are updated using their respective adjacency matrices. The updated feature sequences for graphs A and B are then obtained by summing the initial and positional feature sequences.
[0070] Step S4 specifically includes the following steps:
[0071] S41: Expand the row and column dimensions of the two updated feature sequences, and concatenate the expanded feature sequences along the feature dimensions to obtain n that can be used to learn the matching relationship.2 ×2d dimensional fusion feature sequence.
[0072] S42: Two zigzag continuous prediction paths are manually set, and the fused feature sequences are rearranged according to the prediction paths and sent to the multi-path Mamba module for matching matrix prediction. Each path outputs a prediction result of a matching matrix.
[0073] S42: Average each output matching matrix and pass it through the Sinkhorn algorithm (Sinkhorn algorithm parameter tau is 0.1) to obtain the final n×n soft correspondence matrix, which is the model output. Use the matching matrix to update the two feature sequences of graph A and graph B respectively. Finally, sum the updated feature sequences with the updated feature sequences of the adjacency matrix and feed them into the ReLU function for activation, which serves as the initial feature sequence for the next round.
[0074] Step S5 specifically includes the following steps: obtaining discrete matching results from the predicted soft matching matrix using the Hungarian algorithm to form a Boolean matrix, which is the final matching matrix, for subsequent loss function calculation and accuracy verification.
[0075] Step S6 specifically includes the following steps: vectorizing the predicted soft-matching matrix and the true matching matrix; calculating the model loss using the cross-entropy loss function; balancing positive and negative samples using a weight w=5; this loss is used to quantify the difference between the model's prediction and the true result; finding the optimal values of the model parameters by minimizing the loss function, thereby improving the model's prediction accuracy. During training, weighted cross-entropy is calculated as a supervision signal; model parameters are optimized through gradient backpropagation; and the model parameters are continuously adjusted to reduce the value of the loss function. The model gradually learns the inherent patterns and features of the graph feature data, thus enabling more accurate predictions.
[0076] Step S7 specifically includes the following steps: Using the obtained Boolean matrix, compare it with the correspondence between the true points recorded in the sample preprocessing stage, and calculate the ratio of the number of correctly matched point pairs to the total number of point pairs as an accuracy indicator. This method not only quantifies the accuracy of the model but also allows for systematic comparison with the prediction results of other graph matching algorithms. Recording and comparing the memory usage and number of model parameters during the inference process of different models, and comparing them with baseline methods verifies the lightweight effect, demonstrating that the lightweight design of this model achieves better results in actual inference, and can complete the same accuracy graph matching task with less resource consumption.
[0077] In summary, the advantages of the method described in this embodiment are described in the following two aspects. Regarding the graph neural network algorithm: In graph neural networks, by utilizing the relatively small memory footprint of the SSM model, both the model and computation are processed on SRAM. Compared to network structures based on attention mechanisms, the model in this embodiment significantly reduces memory usage during inference. Since attention mechanisms require numerous read operations and their attention matrix results in high memory demands, repeated copy-and-write operations are necessary. In contrast, the model in this embodiment only requires a single read-write operation. Furthermore, the model has a very low parameter count, resulting in less hardware resource consumption and making it more suitable for deployment and inference in environments with limited hardware resources. Regarding the graph matching task: The method creatively reduces the dimensionality of the graph task to a sequence prediction task and performs feature learning, effectively reducing the computational load required for model training and inference, as well as the memory consumption during inference. Simultaneously, it achieves accuracy comparable to existing methods, greatly expanding the application prospects of graph matching tasks.
[0078] Example 2
[0079] This embodiment 2 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, they implement the lightweight graph matching method described above. The method includes: acquiring image pairs to be matched; processing the acquired image pairs using a pre-trained matching model to obtain graph matching results; wherein the matching model is obtained using the model training method described in embodiment 1.
[0080] Example 3
[0081] This embodiment 3 provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions that can be executed by the processor. The processor calls the program instructions to execute the lightweight graph matching method based on the state-space model as described above. The method includes: acquiring image pairs to be matched; processing the acquired image pairs using a pre-trained matching model to obtain graph matching results; wherein the matching model is obtained using the model training method described in embodiment 1.
[0082] Example 4
[0083] This embodiment 4 provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions to implement the lightweight graph matching method based on the state-space model as described above. The method includes: acquiring image pairs to be matched; processing the acquired image pairs using a pre-trained matching model to obtain graph matching results; wherein, the matching model is obtained using the model training method described in embodiment 1.
[0084] In summary, addressing the inherent limitations of traditional network inference methods such as high memory consumption, large model parameter count, and lightweight techniques, this invention fully leverages the efficient sequence modeling capabilities and structured parameter compression features of the Mamba architecture. By utilizing the sublinear complexity of Mamba and combining it with a serialized graph structure representation method, the number of parameters in traditional graph matching models is reduced by over 90%, significantly lowering storage and computational overhead. Simultaneously, by utilizing the temporal dependency modeling capabilities of SSM, the dynamic correlation patterns between different object graph structures are automatically captured, effectively avoiding the feature information loss issues caused by traditional pruning / quantization methods. Mamba's hardware-friendly awareness mode significantly reduces memory access frequency, enabling the model to maintain efficient inference on mobile and edge devices, meeting real-time requirements. This invention demonstrates excellent applicability in resource-constrained scenarios such as mobile applications (e.g., AR scene understanding, real-time visual positioning) and industrial IoT (e.g., device topology matching). Compared to traditional solutions, this method maintains high-precision graph matching capabilities while reducing memory and computational consumption, providing a reliable solution for graph matching data processing in edge computing environments and possessing significant industrial application value.
[0085] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0086] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0087] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0088] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0089] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative effort should be included within the scope of protection of the present invention.
Claims
1. A lightweight graph matching model training method based on a state-space model, characterized in that, include: We obtained open-source datasets for multiple graph matching tasks as training sets and preprocessed the training sets: we selected photo pairs A and B with the same number of point pairs in the same category, recorded the true matching relationship of the key point sequences in the photo pairs, and then shuffled the feature sequences of the two key points. The preprocessed point-pair feature sequence is converted into a tensor data type that can be learned by the neural network. The initial key point feature sequence and position feature sequence of Figure A and Figure B are randomly set to zero respectively, and high-dimensional mapping is performed. The key point features and position features are summed element by element to form an initial feature sequence that can be updated and learned. The initial feature sequences of graphs A and B are expanded by rows and columns respectively to form four different expanded sequences. The adjacency matrix structure of the graphs is learned, and the initial feature sequences of the two graphs are updated using the adjacency matrix. The updated feature sequences of Figures A and B are concatenated to form a fused feature sequence that learns from each other. The prediction path of the sequence is reconstructed, and the matching matrix is learned. Different matching matrices are output for each path. All output matching matrices are summed element-wise, averaged, and the final soft correspondence matching matrix is calculated. The feature sequences of Figures A and B are updated again using the obtained soft correspondence matching matrix. Specifically, the initial n×d dimensional feature sequence is expanded by a factor of n in both the row and column directions to obtain an n²×d dimensional feature sequence. For the keypoint sequence of an image, two different n²×d dimensional feature sequences are generated. Each image pair contains four n²×d dimensional feature sequences from two images. The four expanded feature sequences are then used to learn the graph structure features, i.e., the adjacency matrix, to obtain the row-predicted adjacency matrix and column-predicted adjacency matrix for graph A, and the row-predicted adjacency matrix and column-predicted adjacency matrix for graph B, for a total of four n×n adjacency matrices predicted by rows and columns respectively. For each graph with different row and column expansions, there are two different corresponding adjacency matrices. The element-wise sum and average of these two adjacency matrices are performed to form a symmetric matrix, i.e., the final adjacency matrix. This operation is performed simultaneously on graph A and graph B, and the corresponding initial feature sequences are updated using their respective adjacency matrices. The element-wise sum of these initial and positional feature sequences is then performed to obtain the updated feature sequences for graph A and graph B. The predicted matching matrix is calculated using the Hungarian algorithm to obtain a discrete Boolean matrix that represents the one-to-one correspondence between the key points of Figure A and Figure B. Using the obtained Boolean matrix, the loss during the model learning process is calculated through the cross-entropy loss function, and gradient backpropagation is performed to train and update the model weights, and the process is iterated and optimized.
2. The lightweight graph matching model training method based on a state-space model according to claim 1, characterized in that, The VGG16 network is used to extract keypoint visual features, and the keypoint visual features and location information are encoded and randomly zeroed in the feature dimension. The keypoint coordinates are converted into high-dimensional location embeddings, and the location is encoded as a 2D mapping to 128D. The two input keypoint feature sequences and location feature sequences are converted into tensor types, and the keypoint feature sequences and location feature sequences are summed element-wise to obtain the enhanced feature sequence.
3. The lightweight graph matching model training method based on a state-space model according to claim 2, characterized in that, The updated two feature sequences are expanded in row and column dimensions, and then concatenated along these dimensions to obtain an n²×2d dimensional fusion feature sequence that can be used to learn matching relationships. Two zigzag continuous prediction paths are manually defined, and the fusion feature sequence is rearranged according to the prediction paths and fed into the multi-path Mamba module for matching matrix prediction. Each path outputs a prediction result for a matching matrix. The average of each output matching matrix is then passed through the Sinkhorn algorithm to obtain the final n×n soft correspondence matrix, which is the model output. The matching matrix is used to update the two feature sequences of Figure A and Figure B respectively. Finally, the updated feature sequence is summed element-wise with the feature sequence updated by the adjacency matrix and fed into the ReLU function for activation, serving as the initial feature sequence for the next round.
4. The lightweight graph matching model training method based on a state-space model according to claim 3, characterized in that, The predicted soft-match matrix and the true matching matrix are vectorized, and the model loss is calculated using the cross-entropy loss function. The positive and negative samples are balanced using a weight w=5. This loss is used to quantify the difference between the model's prediction and the true result. By minimizing the loss function, the optimal values of the model parameters are found, thereby improving the model's prediction accuracy. During training, weighted cross-entropy is calculated as a supervision signal, and the model parameters are optimized through gradient backpropagation. The model parameters are continuously adjusted to reduce the value of the loss function, and the model gradually learns the inherent rules and features of the graph feature data.
5. A lightweight graph matching method, characterized in that, include: Get the image pairs to be matched; The acquired image pairs are processed using a pre-trained matching model to obtain graph matching results; wherein the matching model is obtained using the model training method described in any one of claims 1-4.
6. A lightweight graph matching system, characterized in that, include: The acquisition module is used to acquire image pairs to be matched; The processing module is used to process the acquired image pairs using a pre-trained matching model to obtain image matching results; wherein the matching model is obtained using the model training method described in any one of claims 1-4.
7. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the lightweight graph matching method as described in claim 5.
8. A computer device, characterized in that, It includes a memory and a processor, the processor and the memory communicating with each other, the memory storing program instructions that can be executed by the processor, and the processor calling the program instructions to execute the lightweight graph matching method as described in claim 5.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions that implement the lightweight graph matching method as described in claim 5.
Citation Information
Patent Citations
Hierarchical compressed image matching method and system based on orthogonal attention mechanism
CN111783879A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1