Material Image Segmentation Method Based on Graph Attention with Multi-Dimensional Feature Fusion

By adopting a multi-dimensional feature fusion graph attention network in material image segmentation, the problems of insufficient accuracy and high time cost of material image segmentation in the prior art are solved, and high-precision image segmentation and efficiency improvement are achieved.

CN114708431BActive Publication Date: 2025-05-30SHANGHAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210318948.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-05-30
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

The prior art has problems of insufficient accuracy and high time cost in semantic segmentation of material images, especially when processing grayscale images, due to low regional contrast and blurred boundaries, it is difficult to effectively explore local features of the image.

Method used

A graph attention network based on multi-dimensional feature fusion is adopted. Through the combination of graph encoder, graph attention module and graph decoder, graph structure is constructed and feature fusion is carried out to enhance the network's exploration and understanding of local features of the image.

Benefits of technology

High-precision segmentation of material images is realized, the time cost and labor cost of image processing are reduced, and the segmentation accuracy and efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708431B_ABST
    Figure CN114708431B_ABST
Patent Text Reader

Abstract

The present invention discloses a material image segmentation method based on graph attention with multi-dimensional feature fusion, which can be applied to image segmentation in the field of materials science. This method first preprocesses the material images used for training, then constructs a graph attention network with multi-dimensional feature fusion, optimizes the network parameters using cross-entropy loss, and uses the trained model to predict and segment the material images; finally, it saves the processed results of the output material images. The present invention integrates graph attention with multi-dimensional feature fusion into the encoding and decoding network, improves the segmentation accuracy of the network for material images, reduces the time cost and labor cost of material image processing, and promotes the progress and development of the corresponding academic and industrial circles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision two-dimensional image analysis and processing. For two-dimensional image data, a material image segmentation method based on multi-dimensional feature fusion graph attention is proposed. The present invention can be applied to image segmentation in the field of materials science, improve the accuracy of image segmentation, reduce the time cost and labor cost of image processing, and promote the progress and development of the corresponding academic and industrial circles. Background Art

[0002] Image semantic segmentation is a problem of general concern in fields such as image processing. Semantic segmentation is to let the computer perform segmentation according to the content of the image. Segmentation is to segment different objects in the picture at the pixel level, label each pixel point in the original picture, and classify it into different labels, and the accuracy of segmentation contains the understanding of the information in the image. Material images are generally taken by advanced electron microscopes and are all single-color images, such as grayscale images. The characteristic of grayscale images is that the contrast of each region is not high. According to the light and dark presented by the material structure, the grayscale image is displayed in the form of black, white and gray. The material structure itself has characteristics such as various shape structures, small differences in texture characterization, and discontinuous or even blurred boundaries. Therefore, how to use artificial intelligence technology to quickly and accurately perform semantic segmentation on material images and extract useful information from them is one of the challenges in the field of computer vision.

[0003] There are many methods for image semantic segmentation. Among them, image semantic segmentation based on neural networks is one of the current research hotspots that has received more attention and there have been many research results. FCN (Fully convolutional network) is a classic framework for image semantic segmentation. It is trained in an end-to-end manner and uses the trained classification network for semantic segmentation; in order to restore the resolution of the image, FCN also uses deconvolution for upsampling. Compared with FCN, U-Net has a more symmetric encoding and decoding structure. The skip connection from the encoding part to the decoding part helps to restore the position information. However, since the basic module for constructing the network structure is a simple convolutional block, there is a certain degree of gradient disappearance problem, which limits the increase of the network depth; in addition, U-Net does not fully consider the connection between pixels and lacks the exploration of the dependence relationship between local features, thus affecting the accuracy of the final segmentation result. Therefore, it can be considered that how to construct a deeper and more effective network structure and optimize the network to explore more features is the key to improving the accuracy of semantic segmentation. Summary of the Invention

[0004] In order to solve the problems of the existing technology, the purpose of the present invention is to overcome the deficiencies of the existing technology, design a material image segmentation method based on multi-dimensional feature fusion graph attention, enhance the network's exploration of local features of the image, and achieve high-precision segmentation of the material image.

[0005] To achieve the above object of the invention-creation, the present invention adopts the following technical solutions:

[0006] A material image segmentation method based on graph attention with multi-dimensional feature fusion, comprising the following steps:

[0007] (1) Image preprocessing:

[0008] Adjust the original image and the annotation map for training to a unified specification respectively, and save the preprocessed image locally;

[0009] (2) Construct a network model based on graph attention:

[0010] Input the training set data into the network, optimize the model parameters of the network using cross-entropy loss, and save the trained network parameter file;

[0011] (3) Perform image prediction segmentation:

[0012] Load the trained model parameter file, input the test set data into the network, obtain the predicted segmentation result, and the segmentation result is represented by a binary image;

[0013] (4) Save the output image processing result:

[0014] Save the original image of the test set sample and the segmentation result image in the same picture.

[0015] Preferably, in the step (1), the image preprocessing includes the following steps:

[0016] (1-1) Crop and remove the part of the original image that describes the material performance data;

[0017] (1-2) Uniformly adjust the image to 512×512 pixels;

[0018] (1-3) Use a binarization algorithm to convert all annotation maps into black and white images;

[0019] (1-4) Divide and save the preprocessed image data.

[0020] Preferably, in the step (2), the graph attention module based on multi-dimensional feature fusion includes three sub-modules, namely: (a) graph encoder module, (b) graph attention module, (c) graph decoder module; construct a network model based on graph attention, use a graph encoder to construct a graph structure from the feature map; use graph convolution and graph attention to construct a graph attention module; use a graph decoder to restore the graph structure to a feature map, and the design and construction of the graph encoder include the following steps:

[0021] (2-1-1) Use the output feature map in the encoder part to adjust the dimension size of the feature map from C×H×W to C×HW;

[0022] (2-1-2) Divide the feature map into H×W nodes, and the feature dimension of each node is 1×C;

[0023] (2-1-3) Each node establishes connections in a four-neighborhood manner, that is, each central node establishes edge connections with the four nearest nodes above, below, left, and right;

[0024] (2-1-4) Establish an adjacency matrix through the graph structure to describe the connection situation between nodes;

[0025] (2-1-5) Save the established graph structure in the form of node features and an adjacency matrix.

[0026] Preferably, the graph attention module fuses graph convolution and graph attention, and the design and construction of this module include the following steps:

[0027] (2-2-1) Take the node feature matrix and the adjacency matrix as the input of the graph attention module;

[0028] (2-2-2) Use one layer of graph attention to perform multi-head attention on the input node features (1×C), learn the attention weights, and output the updated node features;

[0029] (2-2-3) Use one layer of graph convolution to perform local aggregation and feature dimensionality reduction on the input node features (1×C), and the node feature dimension of this layer is reduced from the input node feature dimension to 1 / 2×C;

[0030] (2-2-4) Use one layer of graph attention to perform multi-head attention on the input node features (1 / 2×C), learn the attention weights, and output the updated node features;

[0031] (2-2-5) Use one layer of graph convolution to perform local aggregation and feature dimensionality reduction on the input node features (1 / 2×C), and the node feature dimension of this layer is reduced from the node feature dimension output by the previous graph convolution layer to 1 / 4×C;

[0032] (2-2-6) Use one layer of graph attention to perform multi-head attention on the input node features (1 / 4×C), learn the attention weights, and output the updated node features;

[0033] (2-2-7) Use one layer of graph convolution to perform local aggregation and feature dimensionality increase on the input node features (1 / 4×C), and the node feature dimension of this layer is increased from the node feature dimension output by the previous graph convolution layer to 1 / 2×C;

[0034] (2-2-8) The output of the graph attention layer with a node feature dimension of 1 / 2×C and the output of the dimensionality increase operation of the graph convolution layer with a node feature dimension of 1 / 2×C are fused in an additive manner;

[0035] (2-2-9) Use one layer of graph convolution to perform local aggregation and feature dimensionality increase on the input node features (1 / 2×C). The node feature dimension of this layer is increased from the node feature dimension of the output of the previous graph convolution layer to 1×C;

[0036] (2-2-10) The output of the graph attention layer with a node feature dimension of 1×C and the output of the dimensionality increase operation of the graph convolution layer with a node feature dimension of 1×C are fused in an additive manner;

[0037] (2-2-11) Use the hyperparameter a to fuse the features of different branches according to a custom ratio.

[0038] Preferably, use the Resize function to construct the graph decoder module. The design of the graph decoder includes the following steps:

[0039] (2-3-1) Adjust the output dimension of the graph attention module from C×HW to C×H×W;

[0040] (2-3-2) Convert the node features with the adjusted dimension size into the feature map input to the decoder.

[0041] Preferably, in the step (2), the graph attention module uses graph convolution and graph attention layers. The implementation of graph convolution is as follows: Use the degree matrix, adjacency matrix, and node features to implement the graph convolution operation. The calculation formula is as follows:

[0042]

[0043] H l+1 is the output of the graph convolution layer, W l is the weight matrix, H l is the node feature matrix, D is the degree matrix, is the adjacency matrix plus the identity matrix, and σ is the activation function.

[0044] Preferably, for the weight coefficient in the graph attention operation, use the following formula:

[0045]

[0046] α ij is the attention coefficient, W is the weight matrix, is the feature vector of node i, is the feature vector of node j, is the weight vector, σ is the activation function, and softmax is a specific activation function.

[0047] Preferably, for the multi-head attention mechanism in the graph attention operation, the following formula is used:

[0048]

[0049] is the feature vector of the output i node, is the feature vector of the j node, W k is the weight matrix of the k-th attention head, is the attention coefficient between the i node and the j node in the k-th attention head, K is the number of attention heads, and σ is the activation function.

[0050] Preferably, for the hyperparameter a in the graph attention module, the following formula is used:

[0051]

[0052] is the sum result of the output of the graph attention layer and the output of the dimensionality increase operation of the graph convolution layer, H l+1 is the output feature of the (l + 1)-th layer of graph convolution, a is the hyperparameter, is the output of the graph attention layer.

[0053] Preferably, in step (2), when training the network model, set the number of iterations epoch to 100. Usually, when the number of iterations epoch is no more than 75, the network parameters can converge to near the optimal value. The network training includes the following steps:

[0054] (2-4-1) Input the training set images into the network;

[0055] (2-4-2) Use the Adam first-order optimization algorithm to optimize the network parameters, iteratively update the neural network weights based on the training data, and set the weight decay coefficient to alleviate the problem of model overfitting;

[0056] (2-4-3) To further obtain better network performance, set the learning rate, adopt a scheme of dynamically decreasing the learning rate to further approach the optimal value of the network parameters. When the loss value no longer decreases within a certain number of epochs, multiply the learning rate lr by the decay factor to decrease the learning rate, and save the trained model parameter file.

[0057] Preferably, in step (3), the prediction of the material image includes the following steps:

[0058] (3-1) Load the trained model parameter file;

[0059] (3-2) Input the image data into the network to obtain the predicted segmentation result;

[0060] (3-3) Save the original test set sample image and the segmentation result image in the same picture locally.

[0061] Compared with the prior art, the present invention has the following obvious and prominent substantial features and remarkable advantages:

[0062] 1. The present invention is based on a graph attention network with multi-dimensional feature fusion, which can be applied to image segmentation in the field of materials science and improve the segmentation effect. The graph attention module applied to the graph structure is combined with skip connections and transferred to the data structure in Euclidean space, alleviating the loss of spatial information while deepening the network and not changing the size of the feature map; Coarse features are learned through the graph structure representing the low-resolution feature map, and fine detail features are learned through the graph structure representing the high-resolution feature map. The graph structures with different resolutions are fused to achieve accurate learning and representation of the image, thereby realizing high-precision segmentation of the image.

[0063] 2. Considering the characteristic that the graph convolutional layer should not be too deep, the present invention proposes a hyperparameter a to control the message propagation and aggregation of each layer of graph convolution for the global node features, appropriately deepening the number of graph convolutional layers to improve the fitting ability of the network and the accuracy of semantic segmentation. Description of the Drawings

[0064] Figure 1 It is a flowchart of the operation process of the present invention.

[0065] Figure 2 It is a flowchart of the preprocessing method of the present invention. It is divided into the following steps: (1) Crop the image data into image patches that only contain the material structure; (2) Uniformly adjust all images to images with a size of 512×512 pixels; (3) Determine whether the annotation map is a black-and-white map, and use a binarization algorithm to convert non-black-and-white maps into black-and-white maps; (4) Divide and save the preprocessed image data.

[0066] Figure 3The flowchart of the image segmentation of the present invention is as follows. It is divided into the following steps: (1) Input image data, crop the original images for training and testing into image patches that only contain the material structure, preprocess the image patches, and save the preprocessed data locally; (2) Put the training set images into the encoder in the graph attention network based on multi-dimensional feature fusion to obtain a feature map representing high-level features; (3) Put the high-level feature map into the graph encoder to obtain the adjacency matrix and node features representing the graph structure; (4) Put the adjacency matrix and node features of the graph structure into the graph attention module for aggregation and update of node features; (5) Put the updated graph structure into the graph decoder to restore it to the form of a feature map; (6) Put the feature map into the decoder of the network; (7) Use the cross-entropy loss to optimize the model parameters of the network and save the trained network parameter file; (8) Load the trained model parameter file and input the test set data into the network; (9) Calculate the predicted segmentation result, and the segmentation result is represented by a binary graph. (10) Output and save the image result.

[0067] Figure 4 The structural diagram of the graph attention module of the present invention is as follows. It is divided into the following steps: (1) Input the adjacency matrix and node feature matrix representing the graph structure; (2) The input features enter the graph attention layer (GAT1) and the graph convolutional layer (GCN1) respectively to output node features; (3) The node features output by GCN1 enter the graph attention layer (GAT2) and the graph convolutional layer (GCN2) respectively to output new node features; (4) The node features output by GCN2 enter the graph attention layer (GAT3); (5) The node features output by GAT3 enter the graph convolutional layer (GCN3); (6) The result of multiplying the features output by GCN3 by (1 - a) and the result of multiplying the node features output by GAT by a are added together; (7) The added features enter the graph convolutional layer (GCN); (8) The result of multiplying the features output by GCN by (1 - a) and the result of multiplying the node features output by GAT1 by a are added together and then the final node features are output.

[0068] Figure 5 The flowchart of the graph encoder of the present invention is as follows. It is divided into the following steps: (1) Divide the input feature map into nodes and node features; (2) Connect adjacent nodes in a four-neighborhood manner; (3) Construct the adjacency matrix of the graph structure; (4) Save the adjacency matrix and node features of the graph structure. Detailed implementation manners

[0069] In order to enable those skilled in the art to better understand the scheme of the present invention, the preferred embodiments of the present invention are combined with the accompanying drawings to clarify and fully describe the technical scheme of the present invention. Obviously, the described embodiments are only part of the implementation cases of the present invention, not all implementation cases. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0070] The above scheme is further described below in conjunction with specific implementation examples. The preferred embodiments of the present invention are described in detail as follows:

[0071] Embodiment 1:

[0072] In this embodiment, a material image segmentation method based on graph attention based on multi-dimensional feature fusion is proposed. The method constructs an efficient multi-dimensional feature fusion graph attention network structure to improve the network's segmentation accuracy for image data.

[0073] The method of this embodiment uses material images to train the model, obtains model parameters for such data, and then obtains high-precision predictions for similar segmentation data other than samples, see Figure 1 , the method of this embodiment includes the following steps:

[0074] (1) Image preprocessing: adjust the original images and annotation images used for training to the same specifications, and save the preprocessed images locally;

[0075] (2) Build a network model based on graph attention: input the training set data into the network, use cross entropy loss to optimize the model parameters of the network, and save the trained network parameter file;

[0076] (3) Perform image prediction and segmentation: load the trained model parameter file, input the test set data into the network, obtain the predicted segmentation result, and the segmentation result is represented by a binary image;

[0077] (4) Save the output image processing results: save the original image of the test set sample and the segmentation result image in the same image.

[0078] The present invention is based on a material image segmentation method based on graph attention based on multi-dimensional feature fusion. First, the image used for training is preprocessed to obtain a clearer image, and the preprocessed data is saved locally; then the graph attention network is trained on the training data set using the cross entropy loss; then the trained model is used to predict the test data set, and the predicted binary graph result is saved.

[0079] Embodiment 2

[0080] This embodiment is basically the same as the first embodiment, except that:

[0081] In this embodiment, as Figure 1 shown, the image preprocessing includes the following steps:

[0082] (1-1) Crop the image data into image patches of 512×512 pixels;

[0083] (1-2) Determine whether the annotated image is a black-and-white image. For non-black-and-white images, use a binarization algorithm to convert them into black-and-white images;

[0084] (1-3) Divide and save the preprocessed image data.

[0085] In this embodiment, the graph attention module for multi-dimensional feature fusion includes three sub-modules, namely: (a) graph encoder module, (b) graph attention module, and (c) graph decoder module; The graph encoder is constructed by building a graph, and the construction of the graph encoder includes the following steps:

[0086] (2-1) Use the output feature map in the encoder part to adjust the dimension size of the feature map from C×H×W to C×HW;

[0087] (2-2) Divide the feature map into H×W nodes, and the feature dimension of each node is 1×C;

[0088] (2-3) Each node establishes connections in a four-neighborhood manner, that is, each central node establishes edge connections with the four nearest nodes above, below, left, and right;

[0089] (2-4) Establish an adjacency matrix through the graph structure to describe the connection situation between nodes;

[0090] (2-5) Save the established graph structure in the form of node features and an adjacency matrix;

[0091] In this embodiment, the graph attention network module uses graph convolution and graph attention layers for feature fusion; The construction of the graph attention module includes the following steps:

[0092] (2-6) Use the node feature matrix and the adjacency matrix as the input of the graph attention module;

[0093] (2-7) Use one layer of graph attention to perform multi-head attention on the input node features (1×C) to learn the attention weights and output the updated node features;

[0094] (2-8) Use one layer of graph convolution to perform local aggregation and feature dimensionality reduction on the input node features (1×C). The feature dimension of this layer of nodes is reduced to 1 / 2×C of the input node feature dimension;

[0095] (2-9) Use one layer of graph attention to perform multi-head attention on the input node features (1 / 2×C), learn the attention weights, and output the updated node features;

[0096] (2-10) Use one layer of graph convolution to perform local aggregation and feature dimensionality reduction on the input node features (1 / 2×C). The dimensionality of the node features in this layer is reduced from the dimensionality of the node features output by the previous graph convolution layer to 1 / 4×C;

[0097] (2-11) Use one layer of graph attention to perform multi-head attention on the input node features (1 / 4×C), learn the attention weights, and output the updated node features;

[0098] (2-12) Use one layer of graph convolution to perform local aggregation and feature dimensionality increase on the input node features (1 / 4×C). The dimensionality of the node features in this layer is increased from the dimensionality of the node features output by the previous graph convolution layer to 1 / 2×C;

[0099] (2-13) Use an addition method to fuse the output of the graph attention layer with a node feature dimensionality of 1 / 2×C and the output of the graph convolution layer dimensionality increase operation with a node feature dimensionality of 1 / 2×C;

[0100] (2-14) Use one layer of graph convolution to perform local aggregation and feature dimensionality increase on the input node features (1 / 2×C). The dimensionality of the node features in this layer is increased from the dimensionality of the node features output by the previous graph convolution layer to 1×C;

[0101] (2-15) Use an addition method to fuse the output of the graph attention layer with a node feature dimensionality of 1×C and the output of the graph convolution layer dimensionality increase operation with a node feature dimensionality of 1×C;

[0102] (2-16) Use the hyperparameter a to fuse the features of different branches according to a custom ratio;

[0103] In this embodiment, the graph attention module uses graph convolution and graph attention layers. The implementation of graph convolution is as follows:

[0104] (2-17) Implement graph convolution operations using the degree matrix, adjacency matrix, and node features. The calculation formula is as follows:

[0105]

[0106] H l+1 is the output of the graph convolution layer, W l is the weight matrix, H l is the node feature matrix, D is the degree matrix, is the adjacency matrix plus the identity matrix, and σ is the activation function;

[0107] In this embodiment, the implementation steps of graph attention are as follows:

[0108] (2-18) Calculate the attention coefficient using the node features, and the calculation formula is as follows:

[0109]

[0110] α ij is the attention coefficient, W is the weight matrix, is the feature vector of node i, is the feature vector of node j, is the weight vector, σ is the activation function, and softmax is a specific activation function;

[0111] (2-19) Update the node features using the multi-head attention mechanism, and the calculation formula is as follows:

[0112]

[0113] is the feature vector of the output node i, is the feature vector of node j, W k is the weight matrix of the k-th attention head, is the attention coefficient between node i and node j in the k-th attention head, K is the number of attention heads, and σ is the activation function.

[0114] In this example, the graph attention module uses the hyperparameter a to fuse the features of different branches, and the implementation steps are as follows:

[0115] (2-20) Use the hyperparameter a to fuse the features of different branches according to a custom ratio, and the calculation formula is as follows:

[0116] is the sum result of the output of the graph attention layer and the output of the upsampling operation of the graph convolution layer, H l+1 is the output feature of the (l+1)-th layer of graph convolution, a is the hyperparameter, is the output of the graph attention layer;

[0117] In this embodiment, the Resize function is used to construct the graph decoder module, and the construction of the graph decoder includes the following steps:

[0118] (2-21) Adjust the output dimension of the graph attention module from C×HW to C×H×W;

[0119] (2-22) Convert the node features after adjusting the dimension size into the feature map input to the decoder;

[0120] In this embodiment, when training the network model, the number of iterations epoch is set to 100. Usually, when the number of iterations epoch is no more than 75, the network parameters can converge to near the optimal value. The network training includes the following steps:

[0121] (2-23) Input the training set images into the network;

[0122] (2-24) Use the Adam first-order optimization algorithm to optimize the network parameters, iteratively update the neural network weights based on the training data, and set the weight decay coefficient to alleviate the problem of model overfitting;

[0123] (2-25) In order to further obtain better network performance, set the learning rate, and adopt a scheme of dynamically reducing the learning rate to further approach the optimal value of the network parameters. When the loss value does not decrease within a certain number of epochs, multiply the learning rate lr by the decay factor to reduce the learning rate, and save the trained model parameter file.

[0124] This embodiment is based on a graph attention network with multi-dimensional feature fusion, which can be applied to image segmentation in the field of materials science and improve the segmentation effect. The graph attention module applied to the graph structure is combined with skip connections and transferred to the data structure in Euclidean space, alleviating the loss of spatial information while deepening the network and not changing the size of the feature map; Coarse features are learned through the graph structure representing the low-resolution feature map, and fine detail features are learned through the graph structure representing the high-resolution feature map. The graph structures with different resolutions are fused to achieve accurate learning and representation of the image, thereby realizing high-precision segmentation of the image. Combining the characteristic that the graph convolutional layer should not be too deep in this embodiment, a hyperparameter a is proposed to control the message propagation and aggregation of each layer of graph convolution for global node features, appropriately deepening the number of graph convolutional layers to improve the fitting ability of the network and the accuracy of semantic segmentation.

[0125] Embodiment Three

[0126] This embodiment is basically the same as Embodiment Two, with the special feature being:

[0127] In this embodiment, the predicted segmentation of the two-dimensional image includes the following steps:

[0128] (3-1) Load the trained model parameter file;

[0129] (3-2) Input the image data into the network to obtain the predicted segmentation result;

[0130] (3-3) Place the original test set sample image and the segmentation result image on the same picture and save it locally.

[0131] Based on the above embodiments, it can be seen that Figure 2It is a flowchart of the material image segmentation method based on multi-dimensional feature fusion of graph attention in the present invention, including the following steps:

[0132] First, crop the original image for training into image patches that only contain the material structure, preprocess the image patches to obtain clearer image patches, and save the preprocessed data locally; construct a graph attention network based on multi-dimensional feature fusion, input the training set data into the network, use cross-entropy loss to optimize the model parameters of the network, and save the trained network parameter file; load the trained model parameter file, input the image data into the network, obtain the predicted segmentation result, and the segmentation result is represented by a binary image; output and save the post-processed image result. The present invention can be applied to image segmentation in the field of materials science, promoting the progress and development of various disciplinary fields. Figure 2 It is a flowchart of the preprocessing method in this embodiment. It is divided into the following steps:

[0133] (1) Crop the image data into image patches that only contain the material structure; (2) uniformly adjust all images to images of 512×512 pixels; (3) determine whether the annotation image is a black-and-white image, and use a binary algorithm to convert non-black-and-white images into black-and-white images; (4) divide and save the preprocessed image data.

[0134] Figure 3 It is a flowchart of segmenting images in this embodiment. It is divided into the following steps: (1) Input the image data, crop the original image for training and testing into image patches that only contain the material structure, preprocess the image patches, and save the preprocessed data locally; (2) put the training set images into the encoder in the graph attention network based on multi-dimensional feature fusion to obtain a feature map representing high-level features; (3) put the high-level feature map into the graph encoder to obtain the adjacency matrix and node features representing the graph structure; (4) put the adjacency matrix and node features of the graph structure into the graph attention module for aggregation and update of node features; (5) put the updated graph structure into the graph decoder to restore it to the form of a feature map; (6) put the feature map into the decoder of the network; (7) use cross-entropy loss to optimize the model parameters of the network, and save the trained network parameter file; (8) load the trained model parameter file, and input the test set data into the network; (9) calculate the predicted segmentation result, and the segmentation result is represented by a binary image. (10) Output and save the image result.

[0135] Figure 4This is the structural diagram of the graph attention module in this embodiment. It is divided into the following steps: (1) Input the adjacency matrix and node feature matrix representing the graph structure; (2) The input features enter the graph attention layer (GAT1) and the graph convolutional layer (GCN1) respectively to output node features; (3) The node features output by GCN1 enter the graph attention layer (GAT2) and the graph convolutional layer (GCN2) respectively to output new node features; (4) The node features output by GCN2 enter the graph attention layer (GAT3); (5) The node features output by GAT3 enter the graph convolutional layer (GCN3); (6) The result of multiplying the features output by GCN3 by (1 - a) and the result of multiplying the node features output by GAT by a are added together; (7) The added features enter the graph convolutional layer (GCN); (8) The result of multiplying the features output by GCN by (1 - a) and the result of multiplying the node features output by GAT1 by a are added together and then the final node features are output.

[0136] Figure 5 This is the flowchart of the graph encoder in this embodiment. It is divided into the following steps: (1) Divide the input feature map into nodes and node features; (2) Connect adjacent nodes in a four-neighborhood manner; (3) Construct the adjacency matrix of the graph structure; (4) Save the adjacency matrix and node features of the graph structure.

[0137] The material image segmentation method based on multi-dimensional feature fusion graph attention in this embodiment. It can be applied to image segmentation in the field of materials science. This method first preprocesses the training material images, then constructs a multi-dimensional feature fusion graph attention network, optimizes the network parameters using cross-entropy loss, and uses the trained model to predict and segment the material images; finally, saves the output results of the material image processing. This embodiment integrates the multi-dimensional feature fusion graph attention module into the encoding and decoding network, improves the segmentation accuracy of the network for material images, reduces the time cost and labor cost of material image processing, and promotes the progress and development of the corresponding academic and industrial circles.

[0138] The above has described the embodiments of the present invention in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments. Various changes can also be made according to the purpose of the invention of the present invention. Any changes, modifications, substitutions, combinations or simplifications made based on the spirit and principle of the technical solution of the present invention shall be equivalent replacement methods. As long as they meet the invention purpose of the present invention and do not deviate from the technical principle and inventive concept of the material image segmentation method based on multi-dimensional feature fusion graph attention of the present invention, they all belong to the protection scope of the present invention.

Claims

1. A method for material image segmentation based on graph attention with multi-dimensional feature fusion, characterized in that, it includes the following steps: (1) Image preprocessing: Clip and remove the part of the original image that describes the material performance data, adjust the original image and the annotation map for training to a unified specification respectively, and save the preprocessed image locally; (2) Construct a network model based on graph attention: Input the training set data into the network, use the cross-entropy loss to optimize the model parameters of the network, and save the trained network parameter file; (3) Perform image prediction segmentation: Load the trained model parameter file, input the test set data into the network, obtain the predicted segmentation result, and the segmentation result is represented by a binary image; (4) Save the output image processing result: Save the original image of the test set sample and the segmentation result image in the same picture; In the step (2), construct a graph attention module based on multi-dimensional feature fusion. Use a graph encoder to construct a graph structure from the feature map; use graph convolution and graph attention to construct the graph attention module; use a graph decoder to restore the graph structure to a feature map. The design and construction of the graph encoder include the following steps: (2-1-1) In the encoder part, use the output feature map to adjust the dimension size of the feature map from C×H×W to C×HW; (2-1-2) Divide the feature map into H×W nodes, and the feature dimension of each node is 1×C; (2-1-3) Each node establishes connections in a four-neighborhood manner, that is, each central node establishes edge connections with the four nearest nodes above, below, left, and right; (2-1-4) Establish an adjacency matrix through the graph structure to describe the connection situation between nodes; (2-1-5) Save the established graph structure in the form of node features and an adjacency matrix.

2. The method for material image segmentation based on graph attention with multi-dimensional feature fusion according to claim 1, characterized in that, in the step (1), the image preprocessing includes the following steps: (1-1) Clip and remove the part of the original image that describes the material performance data; (1-2) Uniformly adjust the image to 512×512 pixels; (1-3) Use a binarization algorithm to convert all annotation maps into black and white images; (1-4) Divide and save the preprocessed image data.

3. The method for material image segmentation based on graph attention with multi-dimensional feature fusion according to claim 1, characterized in that, the graph attention module fuses graph convolution and graph attention. The design and construction of this module include the following steps: (2-2-1) Use the node feature matrix and the adjacency matrix as the input of the graph attention module; (2-2-2) Use one layer of graph attention to perform multi-head attention on the input node features (1×C) to learn the attention weights and output the updated node features; (2-2-3) Use one layer of graph convolution to perform local aggregation and feature dimension reduction on the input node features (1×C). The node feature dimension of this layer is reduced to 1 / 2×C of the input node feature dimension; (2-2-4) Use one layer of graph attention to perform multi-head attention on the input node features (1 / 2×C) to learn the attention weights and output the updated node features; (2-2-5) Use one layer of graph convolution to perform local aggregation and feature dimensionality reduction on the input node features (1 / 2×C), and the dimensionality of the node features in this layer is reduced from the dimensionality of the node features output by the previous layer of graph convolution layer to 1 / 4×C; (2-2-6) Use one layer of graph attention to perform multi-head attention on the input node features (1 / 4×C), and learn the attention weights to output the updated node features; (2-2-7) Use one layer of graph convolution to perform local aggregation and feature dimensionality increase on the input node features (1 / 4×C), and the dimensionality of the node features in this layer is increased from the dimensionality of the node features output by the previous layer of graph convolution layer to 1 / 2×C; (2-2-8) Use the summation method to fuse the output of the graph attention layer with the node feature dimensionality of 1 / 2×C and the output of the dimensionality increase operation of the graph convolution layer with the node feature dimensionality of 1 / 2×C; (2-2-9) Use one layer of graph convolution to perform local aggregation and feature dimensionality increase on the input node features (1 / 2×C), and the dimensionality of the node features in this layer is increased from the dimensionality of the node features output by the previous layer of graph convolution layer to 1×C; (2-2-10) Use the summation method to fuse the output of the graph attention layer with the node feature dimensionality of 1×C and the output of the dimensionality increase operation of the graph convolution layer with the node feature dimensionality of 1×C; (2-2-11) Use the hyperparameter a to fuse the features of different branches according to a custom ratio.

4. The method for material image segmentation based on graph attention with multi-dimensional feature fusion according to claim 1, characterized in that, the design of the graph decoder includes the following steps: (2-3-1) Adjust the output dimension of the graph attention module from C×HW to C×H×W; (2-3-2) Convert the node features with the adjusted dimension size into the feature map input to the decoder.

5. The method for material image segmentation based on graph attention with multi-dimensional feature fusion according to claim 3, characterized in that, the hyperparameter a in the graph attention module uses the following formula: is the sum result of the output of the graph attention layer and the output of the dimensionality increase operation of the graph convolutional layer, H l+1 is the output feature of the (l + 1)-th layer graph convolution, a is a hyperparameter, is the output of the graph attention layer.

6. The method for material image segmentation based on graph attention with multi-dimensional feature fusion according to claim 1, characterized in that, in the step (2), the training of the network model based on graph attention includes the following steps: (2-4-1) Input the training set images into the network; (2-4-2) Use the cross-entropy loss to optimize the network model parameters; (2-4-3) Save the trained network parameter file.

7. The method for material image segmentation based on graph attention with multi-dimensional feature fusion according to claim 1, characterized in that, in the step (3), the prediction of the material image includes the following steps: (3-1) Load the trained model parameter file; (3-2) Input the image data into the network to obtain the predicted segmentation result; (3-3) Save the original image and the probability map of the test set samples on the same picture locally.

Citation Information

Patent Citations

  • Three-dimensional image segmentation method based on double-path attention coding and decoding network

    CN113643303A