Interpretable Image Segmentation Method Based on Graph Matching Strategy

Through the method based on the graph matching strategy, semantic component relationship diagrams are established and graph matching is performed, the problem of lack of interpretability of existing image segmentation methods is solved, and a more transparent and reliable image segmentation decision-making process is achieved.

CN119516200BActive Publication Date: 2025-05-27ZHEJIANG UNIV CITY COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510076545.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-27
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Existing image segmentation methods lack interpretability and are difficult to provide clear decision-making processes in critical applications, resulting in security and reliability issues.

Method used

An interpretable image segmentation method based on the graph matching strategy is adopted, and visual words with semantic information are extracted by selecting the image segmentation backbone network, and the semantic component relationship diagram of feature unit-level and class-level, and the segmentation result of the image is predicted through graph matching.

Benefits of technology

Enhance the transparency of model decisions, improve user trust in model predictions, reduce the risk of wrong decisions, and is suitable for applications that require high accuracy and reliability, such as medical diagnosis and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516200B_ABST
    Figure CN119516200B_ABST
Patent Text Reader

Abstract

The present invention relates to an interpretable image segmentation method based on a graph matching strategy, including: selecting an image segmentation backbone network and training the image segmentation backbone network; extracting visual words with semantic information from an input image through the image segmentation backbone network; establishing a feature unit-level semantic component relationship graph and an overall category-level semantic component relationship graph for each spatial position according to the visual words with semantic information; measuring the similarity between the feature unit-level semantic component relationship graph and the category-level semantic component relationship graph and predicting the segmentation result of the image. The beneficial effects of the present invention are as follows: through the interactive matching process between the instance-level graph and the category-level graph, the present invention transforms the inference of the deep neural network into an interpretable deductive inference process, which not only enhances the transparency of the model decision-making, but also improves the user's trust in the model prediction in key application fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and more specifically, it relates to an interpretable image segmentation method based on a graph matching strategy. Background Art

[0002] Image segmentation is a key task in the field of computer vision, aiming to group pixels in an image into regions with similar features or semantics. Starting from the success of image segmentation methods in the field of medical imaging, such as DeepLab, U-Net, etc., the application of image segmentation methods can provide richer data and deeper understanding, and has been widely used in various fields, such as autonomous driving, object recognition and tracking, virtual reality, etc.

[0003] However, although deep learning-based models have achieved good performance on challenging benchmarks, their decisions are still unclear due to the lack of interpretability. This problem may have a particularly large impact in critical applications, such as in the fields of medical diagnosis, autonomous driving, etc., which may lead to a reduction in safety.

[0004] Most interpretable artificial intelligence methods focus on classification or regression tasks. Therefore, interpretable segmentation is still considered an open problem, and there are only a few preliminary studies at the intersection of XAI segmentation. One is the symbolic semantic framework, which, together with segmentation, generates symbolic statements derived from the classification distribution. Another method generalizes the Grad-CAM method to the segmentation problem. However, both of these methods have significant drawbacks. The former requires a predefined symbolic vocabulary, while the latter may be unreliable and introduce additional biases to the results.

[0005] Existing interpretable methods for image segmentation mainly include gradient-based methods, backpropagation-based methods, and integral-based methods. Gradient-based methods generate a heatmap of the attention region by calculating the relationship between the gradient of the feature map and the class score, thus revealing the classification decision of the model. Backpropagation-based methods retain positive gradients for positively activated pixels and negative gradients for negatively activated pixels, and generate a visual heatmap by propagating these gradients back to the input image. Integral-based methods quantify the contribution of each pixel to the model output by calculating the gradient difference between the input image and a reference image. These interpretable methods mainly focus on explaining the regions, features, and context information that the model pays attention to in the image.

[0006] In addition, existing image segmentation methods mainly use convolutional neural networks as the core architecture of image segmentation methods. Through methods such as skip connections, dilated convolutions, and attention mechanisms, the segmentation accuracy is continuously improved, and they have advantages in feature extraction, spatial information retention, multi-scale feature fusion, etc. In other non-convolutional neural network architectures, traditional methods focus on using local features, connectivity, and edge information of images for segmentation. For this reason, these methods lack interpretability and comprehensibility of the model, and even cannot correct problems in case of wrong decisions, and are not applicable to some key applications that require a decision-making process. Summary of the Invention

[0007] The object of the present invention is to propose an interpretable image segmentation method based on a graphical matching strategy in view of the deficiencies of the prior art.

[0008] In a first aspect, there is provided an interpretable image segmentation method based on a graphical matching strategy, including:

[0009] S1. Select an image segmentation backbone network and train the image segmentation backbone network;

[0010] S2. Extract visual words with semantic information from the input image through the image segmentation backbone network;

[0011] S3. According to the visual words with semantic information, establish a feature unit-level semantic component relationship graph and an overall category-level semantic component relationship graph for each spatial position;

[0012] S4. Measure the similarity between the feature unit-level semantic component relationship graph and the category-level semantic component relationship graph and predict the segmentation result of the image.

[0013] Preferably, S2 includes:

[0014] S201. Given a pre-trained backbone and an input image, obtain intermediate features and discretize the intermediate features into a sequence of feature components with specific semantics; the feature components are the indexes of visual words;

[0015] S202. Map the feature components to the weighted vertices of the component relationship graph;

[0016] S203. Assign a weighted edge to each pair of vertices to represent their interaction and obtain the component relationship graph.

[0017] Preferably, S3 includes:

[0018] S301. For the discretized sequence of feature components, establish a unit-level component relationship graph for each pixel point as the center in each spatial position;

[0019] S302. Based on the number of categories in the final segmentation result, establish the corresponding number of overall category-level semantic component relationship graphs, that is, category-level semantic component relationship maps; for each overall category-level semantic component relationship graph, select several visual word indexes with the most corresponding feature components in the segmentation result as the vertices of each overall category-level semantic component relationship graph, and initialize it as a fully connected graph.

[0020] Preferably, in S4, the matcher finds the category-level semantic component relationship graph with the highest similarity to the component relationship graph corresponding to the input image in the category-level semantic component relationship map, and finally obtains the segmentation result of the image pixel by pixel; the matcher consists of a GCN module and a similarity calculation module.

[0021] Preferably, in S4, the complexity of the component relationship map is constrained, including: initializing the component relationship map with the component relationship map instances of each category on average, and deleting the edges connected to the vertices with weights lower than the given threshold from the graphs of specific categories.

[0022] Preferably, in S1, the image segmentation backbone network is Deeplabv2 or U-net.

[0023] In the second aspect, an interpretable image segmentation system based on a graph matching strategy is provided for performing any of the methods in the first aspect, including:

[0024] A selection module for selecting an image segmentation backbone network and training the image segmentation backbone network;

[0025] An extraction module for extracting visual words with semantic information from the input image through the image segmentation backbone network;

[0026] A building module for building a feature unit-level semantic component relationship graph and an overall category-level semantic component relationship graph according to the visual words with semantic information at each spatial position;

[0027] A prediction module for measuring the similarity between the feature unit-level semantic component relationship graph and the category-level semantic component relationship graph and predicting the segmentation result of the image.

[0028] In the third aspect, a computer storage medium is provided, and a computer program is stored in the computer storage medium; when the computer program runs on the computer, the computer is enabled to execute any of the methods in the first aspect.

[0029] In the fourth aspect, an electronic device is provided, including:

[0030] A memory for storing a computer program;

[0031] A processor for executing the computer program to implement the method according to any one of the first aspect.

[0032] The beneficial effects of the present invention are as follows:

[0033] 1. Through the interactive matching process between the instance-level graph and the category-level graph, the present invention transforms the inference of the deep neural network (DNN) into an interpretable deductive reasoning process. This method not only enhances the transparency of model decisions but also improves users' trust in model predictions in key application areas such as medical diagnosis and autonomous driving.

[0034] 2. By transforming the image segmentation problem into a graph matching problem, the present invention can accurately identify and interpret the classification decisions of each pixel, thereby reducing the risk of incorrect decisions in key applications. This is particularly important for applications that require extremely high accuracy and reliability, such as medical image analysis and autonomous driving systems.

[0035] 3. The present invention does not depend on a specific convolutional neural network architecture but can be used in combination with any existing deep image segmentation model (such as Deeplabv2, U-net). This compatibility and flexibility enable this method to be easily integrated into existing computer vision systems without large-scale modification of the underlying model architecture.

[0036] 4. Through the precise graph matching strategy and the use of the compositional relationship graph, the present invention can more accurately identify the details and features in the image, thereby improving the segmentation accuracy. This provides significant advantages for applications that require highly accurate image segmentation, such as precise medical diagnosis and fine target tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart of an interpretable image segmentation method based on a graph matching strategy;

[0038] Figure 2 is a schematic diagram of the inference process of an interpretable image segmentation method based on a graph matching strategy;

[0039] Figure 3 is a schematic diagram of the structure of an interpretable image segmentation system based on a graph matching strategy. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following further describes the present invention in conjunction with embodiments. The description of the following embodiments is only for helping to understand the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

[0041] Embodiment 1:

[0042] In view of problems such as the process of computational visual representation and the specific embedding of learning being opaque, the present invention designs an interpretable image segmentation method based on a graph matching strategy. The core idea of this method is to transform DNN inference into an interactive matching process between instance-level graphs and category-level graphs, thereby realizing interpretable deductive reasoning.

[0043] Specifically, as Figure 1 shown, this method includes:

[0044] S1. Select an image segmentation backbone network and train the image segmentation backbone network.

[0045] The present invention uses image segmentation based on a deep image segmentation model (such as Deeplabv2, U-net). Deeplabv2 used in the present invention is a convolutional neural network architecture for image segmentation tasks. It uses techniques such as atrous convolution, multi-scale processing, and conditional random fields to efficiently segment each pixel in an image into parts of different semantic categories. Let be the backbone network, be the RGB image, then is the output of this image backbone network, that is, the intermediate feature .

[0046] S2. Extract visual words with semantic information from the input image through the image segmentation backbone network.

[0047] Specifically, this step includes:

[0048] Given a pre-trained backbone and an input image, obtain the intermediate feature and discretize the intermediate feature into a sequence of feature components with specific semantics; the feature components are the indices of visual words.

[0049] The main purpose of feature discretization is to mine the most common patterns that appear in the dataset, and each pattern corresponds to a semantics (visual word) that can be understood by humans. However, local regions with the same semantics may exhibit different features due to scaling, rotation, or even distortion. The present invention uses a DNN backbone to extract relatively unified features instead of using traditional SIFT features. Specifically, it is by running clustering on the set of visual tokens extracted from the detection dataset to construct a visual word of size M. In addition, by replacing each element with the index of the nearest visual word, the intermediate feature X is discretized into a sequence of feature components, and the formula is:

[0050]

[0051] The characteristic component refers to the index of the visual word . In the entire set of components , the discrete sequence is represented as , where is calculated by the above formula.

[0052] S3. Establish an instance-level feature unit-level semantic component relationship graph and an overall category-level semantic component relationship graph for each spatial position.

[0053] First, define the component relationship graph and the component relationship graph spectrum mathematically:

[0054] The component relationship graph is an undirected fully connected graph , where the vertex set is the set of characteristic components of an instance or a specific category, and the edge set encodes the interaction of vertex pairs. Each vertex is assigned a non-negative weight to represent its importance. Let be the set of all vertex weights. In addition, the interaction between vertex and is quantitatively defined by a non-negative weight as the weighted adjacency matrix. The component relationship graph spectrum is defined as an abstract representation representing all features, where

[0055] the component relationship graph spectrum of a class is a set of category-level component relationship graphs, where the elements of a class have learnable vertex weights and edge weights .

[0056] S3 includes:

[0057] S301. For the discretized characteristic component sequence, it can be in one-to-one correspondence with the pixel points of the picture. Therefore, an instance-level feature unit-level semantic component relationship graph is established for each spatial position with each corresponding pixel point as the center.

[0058] It should be noted that the feature unit-level semantic component relationship graph corresponds to each input picture instance, so it is instance-level. Specifically, for the discretized characteristic component sequence , its shape is , and an unit-level component relationship graph with 5 vertices is established for each pixel point as the center , where the vertex set , is the central pixel, ​​, , are the pixels above, below, left, and right of the central pixel, respectively, and the edge set .

[0059] S302. According to the number of categories in the final segmentation result, establish the corresponding number of overall category-level semantic component relationship graphs, that is, the category-level semantic component relationship graph spectrum; for each overall category-level semantic component relationship graph, select several visual word indexes with the most corresponding feature components of this category in the segmentation result as the vertices of each overall category-level semantic component relationship graph, and initialize it as a fully connected graph.

[0060] Exemplarily, the final segmentation result contains categories, so it is necessary to establish overall category-level semantic component relationship graphs. For each category-level semantic component relationship graph, select the visual word indexes with the most corresponding feature components of this category in the segmentation result. The formula is:[[]]

[0061]

[0062] Take visual word indexes as the vertices of each category graph, and initialize it as a fully connected graph.

[0063] S4. Measure the similarity between the feature unit-level semantic component relationship graph and the category-level semantic component relationship graph and predict the segmentation result of the image.

[0064] Embodiment 2:

[0065] Based on Embodiment 1, Embodiment 2 of the present application provides a more specific interpretable image segmentation method based on a graph matching strategy, including:[[]]

[0066] S1. Select an image segmentation backbone network and train the image segmentation backbone network.

[0067] S2. Extract visual words with semantic information from the input image through the image segmentation backbone network.

[0068] S3. According to the visual words with semantic information, establish a feature unit-level semantic component relationship graph and an overall category-level semantic component relationship graph for each spatial position.

[0069] S4. Measure the similarity between the feature unit-level semantic component relationship graph and the category-level semantic component relationship graph and predict the segmentation result of the image.

[0070] In S4, after converting the input image into a feature unit-level semantic component relationship graph, the matcher finds the most similar category-level semantic component relationship graph in the category-level semantic component relationship graph atlas, and finally obtains the image segmentation result of the image pixel by pixel. The matcher consists of a GCN module and a similarity calculation module that generates the final prediction.

[0071] To feed the component relationship graph s into the GCN module, each component (i.e., the graph vertex) is assigned a trainable embedding vector, where each embedding vector is initialized from a -dimensional random vector, and these vectors are independently drawn from a multivariate Gaussian distribution .

[0072] In the GCN module, the present invention adopts , and makes a slight modification to the weighted edge: Let be the input features of all vertices in the

[0073]

[0074] layer, and its output is calculated as: where represents the non-linear activation function, represents feature normalization, is a learnable projection matrix. After passing through all layers of , the weighted average pooling layer will weight the vertex embeddings and generate a graph representation . Now for there is , and for all graphs in there is . By calculating the inner product similarity, the final prediction logarithm is defined as .

[0075] Given an instance graph and a category graph , first analyze the similarity score , where and are the weighted sums of the vertex embeddings respectively. For further discussion, let represent the -th output of the embedding vector of vertex in the layer, and represents the weight of . The original vertex embedding of , which is the same for the same component in any two graphs. Additionally, let be and the set of shared vertices in . The calculation of

[0076]

[0077] For the shallow GCN module, s can be approximated as:

[0078]

[0079] If and , the method of the present invention is equivalent to with a linear classifier.

[0080] Furthermore, the interpretability of the graph matcher with shallow GCN is expressed as: (1) the final prediction score of the class is the sum of all existing class evidences represented by vertex weights; (2) the shared neighboring nodes (local structures) connected to the shared vertex also contribute to the final prediction.

[0081] In the inference stage, the present invention uses bilinear interpolation combined with the input image size to match the size of the segmented image, while in the training stage, we reduce the resolution of the underlying ground truth segmentation to adapt to the size of the output feature map.

[0082] Furthermore, in S4, the present invention constrains the complexity of the component relationship graph to facilitate model training.

[0083] The present invention defines the complexity of the edges and vertex weights of all graphs in the component relationship graph as:

[0084]

[0085] where the function calculates the entropy of the input vector and normalizes it to the sum of components, represents the weighted edge connected to the vertex of the class in the graph . The final optimization objective is:

[0086]

[0087] where and are hyperparameters.

[0088] Each graph in the compositional relationship graph is initialized as a fully connected graph with random vertex and edge weights. However, this requires memory space to store and train the entire edge set. To alleviate this problem, the present invention initializes the compositional relationship graph by averaging the compositional relationship graph instances of each class, and deletes the edges connected to the vertices with weights lower than a given threshold from the graphs of specific classes. Such a process can not only significantly reduce the learnable parameters, but also improve the final performance.

[0089] The present invention has been experimented on Pascal VOC 2012 (PASCAL Visual Object Classes Challenge 2012). Pascal VOC2012 is a very influential dataset in the field of computer vision and is widely used for model training and evaluation of tasks such as object detection, image segmentation, and image classification. It is part of the Pascal VOC series of challenges, and the goal is to promote the research and development of image understanding.

[0090] The present invention achieves an average interaction ratio of 74.5% on the validation set of the Pascal VOC 2012 dataset, which is only 2.2% lower than Deeplabv2, as shown in Table 1. However, the interpretability of the present method is greatly improved, which helps to improve the reliability of semantic segmentation and has higher usability in key fields such as medical diagnosis and autonomous driving.

[0091] Table 1 Comparison of the effects of the present invention and other methods

[0092]

[0093] The present invention also visualizes the inference process to illustrate the interpretability of the present invention. As Figure 2 shown, first, the features of the elephant graph are discretized and the components are visualized. As Figure 2 shown in the lower left corner, it can be seen that the picture mainly includes components such as 3, 46, 107, 121, 133, 150, 198, etc. Different colors represent different components. Among them, component 198 mainly represents the elephant's body, and component 150 mainly represents the elephant's limbs. In addition, there are components representing the sky, grassland, etc. The outlines of the elephant's body and limbs are basically consistent with the visualized components. Then, an instance-level relationship graph and a class-level relationship graph are established respectively. The instance-level relationship graph is obtained from the image, and the class-level relationship graph is obtained from training. Finally, a pixel-by-pixel semantic component relationship Figure 1 pairwise matching is performed to obtain the probability of each pixel corresponding to each class, and an interpretable classification of each pixel is achieved.

[0094] It should be noted that the same or similar parts in this embodiment and Embodiment 1 can be referred to each other and will not be elaborated in this application.

[0095] Embodiment 3:

[0096] Based on Embodiments 1 and 2, Embodiment 3 of this application provides an interpretable image segmentation system based on a graph matching strategy, including:

[0097] An interpretable image segmentation system based on a graph matching strategy, including:

[0098] A selection module, configured to select an image segmentation backbone network and train the image segmentation backbone network;

[0099] An extraction module, configured to extract visual words with semantic information from an input image through the image segmentation backbone network;

[0100] A building module, configured to build a feature unit-level semantic component relationship graph and an overall category-level semantic component relationship graph according to the visual words with semantic information at each spatial position;

[0101] A prediction module, configured to measure the similarity between the feature unit-level semantic component relationship graph and the category-level semantic component relationship graph and predict the segmentation result of the image.

[0102] Specifically, the system provided in this embodiment is the system corresponding to the methods provided in Embodiments 1 and 2. Therefore, the same or similar parts in this embodiment and Embodiments 1 and 2 can be referred to each other and will not be elaborated in this application.

Claims

1. An interpretable image segmentation method based on a graph matching strategy, characterized in that: include: S1, selecting an image segmentation backbone network and training the image segmentation backbone network; S2, extracting visual words with semantic information from the input image through the image segmentation backbone network; S3, according to the visual words with semantic information, establishing a feature unit level semantic component relationship diagram and an overall category level semantic component relationship diagram by spatial position; S3 includes: S301, for the discretized feature component sequence, establish a unit-level component relationship graph at each spatial position with each pixel point as the center; S302, according to the number of categories of the final segmentation result, establish a corresponding number of overall category-level semantic component relationship graphs; for each overall category-level semantic component relationship graph, select a number of visual word indexes with the most corresponding characteristic components of the category in the segmentation result as the vertices of each overall category-level semantic component relationship graph, and initialize it as a fully connected graph; S4, measuring the similarity between the feature unit-level semantic component relationship graph and the category-level semantic component relationship graph and predicting the image segmentation result; in S4, a matcher is used to find the category-level graph with the highest similarity to the component relationship graph corresponding to the input image in the component relationship graph, and finally the image segmentation result is obtained pixel by pixel; the matcher is composed of a GCN module and a similarity calculation module.

2. The interpretable image segmentation method based on the graph matching strategy according to claim 1, characterized in that S2 include: Given a pre-trained backbone and an input image, obtain intermediate features and discretize them into a sequence of feature components with specific semantics; The feature component is an index of the visual word.

3. The interpretable image segmentation method based on the graph matching strategy according to claim 2 is characterized in that: In S1, the image segmentation backbone network is Deeplabv2 or U-net.

4. An interpretable image segmentation system based on a graph matching strategy, characterized in that: The method for executing any one of claims 1 to 3 comprises: A selection module is used to select an image segmentation backbone network and train the image segmentation backbone network; An extraction module, used to extract visual words with semantic information from an input image through an image segmentation backbone network; An establishment module is used to establish a feature unit-level semantic component relationship diagram and an overall category-level semantic component relationship diagram according to the visual words with semantic information, spatially; The prediction module is used to measure the similarity between the feature unit level semantic component relationship graph and the category level semantic component relationship graph and predict the image segmentation result.

5. A computer storage medium, characterized in that: The computer storage medium stores a computer program; when the computer program is executed on a computer, the computer executes any one of the methods described in claims 1 to 3.

6. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Image semantic segmentation method and system based on multiple thresholds

    CN112396620A

  • Remote sensing image semantic segmentation method and system based on space and semantic consistency comparative learning

    CN116935242A