Image segmentation method, training method, device, equipment and medium

By using the method of mixed attention and graph convolution in medical image segmentation, the window mechanism and attention mechanism are used to reduce the computational complexity, the problem of difficulty in dealing with abdominal images in the prior art is solved, and a more efficient and accurate abdominal multi-organ segmentation effect is achieved.

CN120125548AActive Publication Date: 2025-06-10CENT SOUTH UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510217030.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-10
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have problems such as high computational complexity and difficulty in capturing long-distance dependence and complex spatial relationships when processing abdominal images, resulting in low segmentation accuracy and efficiency.

Method used

An image segmentation method that mixes attention and graph convolution is adopted, and feature maps are divided through window mechanisms, and combined with attention mechanisms and graph representation learning methods, the local and distant dependencies are effectively learned to reduce the computational complexity.

Benefits of technology

It achieves lower computational complexity and higher segmentation accuracy, and can capture complex irregular topological structures in abdominal images more accurately, improving the effect of abdominal multi-organ segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125548A_ABST
    Figure CN120125548A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation method and device, a training method and device, equipment and a medium, and the method comprises the steps: carrying out the feature map division based on a window mechanism in the jump connection between coding and decoding, carrying out the calculation of a first feature map in which local features are fused in windows, and carrying out the calculation of a second feature map in which global features are fused between the windows, and the first feature graph and the second feature graph are spliced to obtain a fused feature graph for decoding and outputting, so that in-window and inter-window feature calculation with low overhead is adopted, the feature similarity between each pair of graph nodes is prevented from being calculated, the overall calculation complexity is reduced, and the applicability is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image segmentation method, training method, device, equipment and medium. Background Art

[0002] Abdominal organ segmentation plays a vital role in medical image processing and is an indispensable part of computer-aided diagnosis, surgical navigation, visual enhancement, radiotherapy and biomarker measurement systems. Accurate segmentation results can provide key information such as organ size, location, boundary characteristics and their spatial relationships, which is of great significance for image analysis, surgical planning, clinical decision support and follow-up. Especially in the radiotherapy of cancer and tumors, accurate segmentation of dangerous areas not only helps to improve the treatment effect, but also effectively reduces the impact on surrounding healthy tissues. However, due to the complex abdominal anatomical structure, diverse organ morphology and different lesions, accurate analysis of abdominal imaging data such as computed tomography (CT) and magnetic resonance imaging (MR) is a very challenging and time-consuming task for clinicians, and it is prone to errors. Therefore, the development of efficient and reliable abdominal organ segmentation technology has become an important direction to improve the quality and efficiency of diagnosis and treatment.

[0003] In order to improve the accuracy of image interpretation, optimize the clinical decision-making process, and enhance the quality of patient care, there is an urgent need to develop efficient and robust automatic abdominal image segmentation technology to replace the time-consuming and error-prone traditional manual delineation method. In recent years, with the advancement of deep learning technology, medical image segmentation methods based on deep learning have shown significant advantages. By learning complex features from a large amount of medical imaging data, the accuracy and efficiency of segmentation tasks have been greatly improved. However, existing mainstream models such as convolutional neural networks (CNNs) and Transformer models still face some challenges in dealing with medical image segmentation. Traditional CNNs rely on the pixel grid structure in Euclidean space and have limitations in capturing long-range dependencies and complex spatial relationships. Although Transformer has made some progress by segmenting the image into small blocks and using the self-attention mechanism to identify long-range dependencies, it is still not ideal in processing image details and edge detection. In addition, Transformer models usually have high computational complexity, and the attention mechanism used by Transformer needs to consider all other nodes when updating the current node representation, which may introduce redundant information in the feature extraction process, thereby affecting the accuracy and efficiency of the segmentation results.

[0004] Recently, graph-based modeling methods have been applied to image tasks due to their flexibility and versatility. These methods have been shown to be effective in modeling irregular structures in images and have achieved success in computer vision and medical imaging. Graph-based modeling methods can not only represent irregular topological structures, thereby flexibly modeling irregular shapes and areas with large contrast changes in medical images, but also selectively fuse features through edge connections to avoid capturing too much redundant information.

[0005] However, existing graph modeling methods usually divide the image into multiple patches and regard each patch as a graph node. Then, the edge connection relationship between nodes is established through the K-nearest neighbor algorithm. When establishing the edge connection relationship, it is necessary to calculate the feature similarity between each pair of nodes, which leads to expensive computational complexity. This limitation hinders the applicability of graph modeling methods, especially in dense prediction tasks such as image segmentation. Summary of the invention

[0006] The present application proposes an image segmentation method, training method, device, equipment and medium, which can solve one of the problems existing in the background technology.

[0007] In order to achieve the above objectives, this application adopts the following technical solutions:

[0008] In a first aspect, a training method for an image segmentation model is provided, the training method comprising:

[0009] Obtaining training data, the training data including: an image to be segmented and an image segmentation result; and

[0010] Using the training data, training the image segmentation model,

[0011] The image segmentation model includes: an encoding part, a decoding part, an output part located at the back end of the decoding part, and a jump connection part located between the encoding part and the decoding part.

[0012] The encoding part is used to obtain a multi-scale feature map from the image to be segmented;

[0013] The jump connection part is used to perform a first division on at least one scale feature map with a predetermined size window to obtain a first-level feature map; perform a second division on the first-level feature map to obtain a second-level feature map; perform self-attention calculation on the second-level feature map to obtain a first feature map that integrates local features; perform a graph convolution operation on the first-level feature map to obtain a second feature map that integrates global features; and, concatenate the first feature map and the second feature map to obtain a fused feature map to output to the decoding part.

[0014] Based on the above technical solution, in the jump connection between encoding and decoding, feature map division is performed based on the window mechanism, and then the first feature map for fusing local features within the window is calculated, and the second feature map for fusing global features between windows is calculated. The first feature map and the second feature map are concatenated to obtain a fused feature map for decoding output. In this way, low-overhead intra-window and inter-window feature calculations are used to avoid calculating the feature similarity between each pair of graph nodes, thereby reducing the overall calculation complexity and expanding the applicability.

[0015] In a possible design manner of the first aspect, the jump connection part is further used for:

[0016] The fused feature maps corresponding to the scale feature maps are concatenated in the spatial dimension with the same channel dimension to obtain a node set and a node feature map in the spatial dimension;

[0017] Using the relationship between adjacent scale nodes as edges, constructing a graph representation corresponding to the node set;

[0018] Performing a graph convolution operation on the graph representation to obtain a node feature representation reflecting the correlation between the fused feature graphs of different scales; and

[0019] The node feature representation is restored at various scales and spliced ​​with the corresponding fused feature map to obtain a spliced ​​feature map which is output to the decoding part.

[0020] In a possible design manner of the first aspect, a graph convolution operation is performed on the graph representation to obtain a node feature representation that reflects the correlation between the fused feature graphs of different scales, specifically: an edge convolution operation is performed on the node feature graph to obtain the node feature representation.

[0021] In a possible design manner of the first aspect, the encoding part includes: a first layer encoder, a second layer encoder, a third layer encoder, a fourth layer encoder and a fifth layer encoder of pooling connection, the first layer encoder inputs the image to be segmented, the decoding part includes a first layer decoder, a second layer decoder, a third layer decoder, a fourth layer decoder and a fifth layer decoder of upsampling connection, the fifth layer encoder outputs to the first layer decoder correspondingly, the jump connection part includes a first jump connection between the first layer encoder and the fifth layer decoder, a second jump connection between the second layer encoder and the fourth layer decoder, a third jump connection between the third layer encoder and the third layer decoder, and a fourth jump connection between the fourth layer encoder and the second layer decoder, the second jump connection, the third jump connection and the fourth jump connection are used to obtain a fusion feature map corresponding to the corresponding scale.

[0022] In a possible design manner of the first aspect, the image to be segmented is a computed tomography (CT) abdominal image or a magnetic resonance (MR) abdominal image.

[0023] In a second aspect, an image segmentation method is provided, the segmentation method comprising:

[0024] Obtaining an image to be processed;

[0025] The image to be processed is processed using the image segmentation model obtained through the training as described above to obtain an image processing result.

[0026] In a third aspect, a training device for an image segmentation model is provided, the training device comprising:

[0027] A first acquisition unit is used to obtain training data, wherein the training data includes: an image to be segmented and an image segmentation result; and

[0028] a training unit, configured to train the image segmentation model using the training data,

[0029] The image segmentation model includes: an encoding part, a decoding part, an output part located at the back end of the decoding part, and a jump connection part located between the encoding part and the decoding part.

[0030] The encoding part is used to obtain a multi-scale feature map from the image to be segmented;

[0031] The jump connection part is used to perform a first division on at least one scale feature map with a predetermined size window to obtain a first-level feature map; perform a second division on the first-level feature map to obtain a second-level feature map; perform self-attention calculation on the second-level feature map to obtain a first feature map that integrates local features; perform a graph convolution operation on the first-level feature map to obtain a second feature map that integrates global features; and, concatenate the first feature map and the second feature map to obtain a fused feature map to output to the decoding part.

[0032] In a fourth aspect, an image segmentation device is provided, the segmentation device comprising:

[0033] A second acquisition unit, configured to obtain an image to be processed; and

[0034] The segmentation unit is used to process the image to be processed using the image segmentation model obtained through the training as described above to obtain an image processing result.

[0035] In a fifth aspect, an electronic device is provided, the electronic device comprising: a processor, and a memory coupled to the processor,

[0036] The memory is used to store a computer program; and

[0037] The processor is used to execute the computer program stored in the memory, so that the electronic device executes the training method as described above, or executes the segmentation method as described above.

[0038] In a sixth aspect, a computer-readable storage medium is provided, characterized in that the computer-readable storage medium includes a computer program or instructions, and when the computer program or instructions are run on a computer, the computer executes the training method as described above, or executes the segmentation method as described above.

[0039] The beneficial effects of the second to sixth aspects are as described in the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0041] Figure 1 is a schematic diagram of an abdominal multi-organ segmentation network provided in an embodiment of the present application;

[0042] Figure 2 It is a schematic diagram of a window graph convolution module provided in an embodiment of the present application;

[0043] Figure 3 It is a schematic diagram of window division provided in an embodiment of the present application, in which a feature map is generated into a local feature map in a window and a window feature map by changing a feature dimension;

[0044] Figure 4 It is a schematic diagram of constructing a multi-scale feature static graph in a hierarchical architecture provided by an embodiment of the present application. According to the relationship between the upper and lower layers in the hierarchical structure, an edge from an upper layer node to a lower layer node is constructed to construct a static graph, thereby realizing modeling of hierarchical features and multi-scale features;

[0045] Figure 5 It is a schematic diagram of the abdominal multi-organ segmentation results provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0049] This embodiment will design a hybrid attention and graph convolution abdominal multi-organ segmentation method that is superior to the existing mainstream methods. In view of the high computational complexity of the existing segmentation methods based on graph neural networks during the graph construction process, this embodiment mainly divides the feature graph based on the window mechanism, and effectively learns local and remote dependencies by combining the attention mechanism with the graph representation learning method. In addition, this embodiment also designs a static graph to achieve modeling of multi-scale information in a hierarchical architecture.

[0050] In summary, this embodiment can accurately capture the complex and irregular topological structure in the abdominal image by effectively modeling medical images and multi-scale information in the hierarchical architecture to achieve better and more accurate results.

[0051] Next, Section 1 will introduce the symbolic description of this method and the corresponding technical background, Section 2 will introduce the algorithm flow proposed in this embodiment in detail, and Section 3 will state the results and advantages of the proposed algorithm.

[0052] 1. Symbols and Background

[0053] 1.1 Definition of Graph

[0054] A graph is composed of a finite number of nodes and edges connecting two nodes, which can be expressed as G = (V, E). V represents the set of nodes, and |V| = N means that the number of nodes in the graph is N, and E represents the set of edges. In graph machine learning and graph neural networks, matrices are usually used. Represents the characteristics of the node set V, d represents the dimension of the node feature; use the weighted adjacency matrix Represents the connection relationship between two nodes, and uses it to represent the edge set E. If there is an edge from node i to node j, then A i,j Set to 1 or the weight of the edge, otherwise set to 0.

[0055] 1.2K Nearest Neighbor (KNN) Composition

[0056] KNN graph construction is a graph construction method based on the K nearest neighbor algorithm, which is used to construct an image into a graph structure. Specifically, given an image of size H×W, the image is first divided into N fragments (patches). By converting each patch into a node vector We can obtain a node feature matrix X = (x 1 ,x 2 ,...,x N ). In this way, we get a set of node sets V. Then, we use the K nearest neighbor algorithm to find the K neighbors with the most similar features of each node in the feature space and add edges to them, thereby obtaining a set of edge sets E. Specifically, we need to calculate the distance between each two nodes, usually using the Euclidean distance, as shown in formula (1):

[0057]

[0058] Among them, x ik and x jk Respectively represent the values ​​of the i-th and j-th nodes on the k-th feature. In this way, the distance matrix is ​​obtained D ij represents the distance between the i-th and j-th nodes. Then, for each column of D, select the most similar Top-k neighbor nodes and set the value of the corresponding position in the column to 1, and the remaining positions to 0. In this way, the adjacency matrix is ​​obtained

[0059] Through the above process, the image structure can be converted into a graph structure, so that the image can be processed using graph-related methods.

[0060] 1.3 Graph Convolution

[0061] In traditional CNN, convolution operations are performed on regular grid data (such as images), while graph convolution is designed to process irregular data structures - graphs, where the connection relationship between nodes can be arbitrary. The goal of graph convolution is to extract features from nodes and their neighbors in the graph while maintaining the graph structure information. Graph convolution is usually performed by normalizing the adjacency matrix To aggregate the information of a node and its neighboring nodes, and transform them through linear transformation and nonlinear activation function to obtain new node feature representation. A simple graph convolution layer can be described as follows:

[0062]

[0064] here is a degree matrix, which is a diagonal matrix and its i-th diagonal element represents the degree of node i, that is, the number of edges connected to node i. is the adjacency matrix of the graph plus the identity matrix (i.e., each node is connected to itself), indicating the existence of self-loops. (l) is the trainable weight matrix of this layer, σ is the activation function, X (l) and X (l+1) Represent the node feature matrices of the current layer and the next layer respectively.

[0065] 2. Methods

[0066] This embodiment proposes a method for segmenting multiple abdominal organs by mixing attention and graph convolution. While reducing the computational complexity of the composition process, it uses the multi-scale information in the graph modeling image and the hierarchical architecture to accurately capture the complex and irregular topological structure of the abdominal region, thereby achieving the purpose of accurate segmentation of multiple abdominal organs.

[0067] 2.1 Construction of abdominal multi-organ segmentation network

[0068] For the specific task of abdominal multi-organ segmentation, we built a U-shaped architecture, such as Figure 1 It consists of a five-layer encoder, a five-layer decoder and skip connections.

[0069] Each encoder layer consists of an encoder module and a pooling layer. The encoder module consists of two 3×3 convolutional layers and one 1×1 convolutional layer. Each convolutional layer includes an activation function layer and a normalization layer. The pooling layer is used to maintain the layered architecture of the U-shaped network and mainly performs 2x downsampling.

[0070] In the skip connection part, a window graph convolution module is designed, which mainly uses the attention module and the graph module to more effectively identify basic features.

[0071] Afterwards, the hierarchical multi-scale features generated by the network are passed through multi-scale remixed patches (RGCB) to capture and fuse local and global correlations between features of different scales.

[0072] Next, information of different scales will be input into the decoder. The decoder upsamples the features from the previous decoder layer, concatenates them with the output of the same layer in the skip connection in the feature dimension, and fuses them through a 3×3 convolution layer and a 1×1 convolution layer. Each convolution layer also includes an activation function layer and a normalization layer.

[0073] Finally, the output of the decoder part passes through a segmentation head (including a 1×1 convolution layer and a sigmoid activation function) to generate pixel-level category labels to obtain the segmentation result.

[0074] 2.2 Window Graph Convolution Module

[0075] In order to reduce the complexity of the composition process, we designed a window graph convolution (WGC) module, such as Figure 2 This module introduces a window mechanism and combines the attention mechanism and graph convolution method to effectively learn local features and long-range dependencies.

[0076] (1) Window division

[0077] We first perform window division on the feature map to reduce the computational complexity of the attention mechanism and the image composition process. It should be noted that the proposed network does not perform window division in the last two layers, but directly passes Figure 1 This is mainly because the feature map that has been pooled four times in the U-shaped architecture (the size has become the original feature map) ) is already relatively small, and the computational overhead it brings is relatively low.

[0078] Specifically, for an input feature map of H×W×C and a window size of h×w, where C is the channel dimension, the input feature map is first transformed into Indicates that the window division is performed, and the h×w window is used to divide the H×W feature map, which is convenient for subsequent conversion and combination of feature dimensions to generate the feature map within the window and the window feature map. Then, the size of The local feature map in the window and the size are The window feature map of Figure 3 As shown in Figure 2. The former fuses local features through attention calculation within the local window. The latter selectively fuses global features through graph convolution operations between windows. This process can be described as

[0079] X in ,X out =window_pattern(X) (3)

[0080] Where X∈R H×W×C represents the input feature map, represents the local feature map in the window, Represents the window feature map. Among them, L=h×w C′=h×w×C.

[0081] It should be noted that H×W×C is the size of the original image, which has not been processed by the WGC module. This is mainly due to the consideration of computational overhead and effect. For the original input feature map, each point in the feature map represents a pixel, and the performance improvement brought by directly processing it is limited. At the same time, even after windowing, the time overhead of processing it is huge, so it is not processed by WGC.

[0082] (2) Self-attention calculation within a local window

[0083] After dividing the feature map into windows, we perform self-attention calculations in local windows. Limiting this calculation method to a smaller window is to reduce the global quadratic complexity. At the same time, considering that there are fewer nodes in the window, we directly model it as a fully connected graph and use the self-attention mechanism to efficiently extract local features. Why can the attention mechanism be understood as a fully connected graph? The main reason is that the attention mechanism needs to calculate the attention matrix, which is actually the feature similarity matrix between nodes. This process is consistent with the logic of adding edges in the K nearest neighbor algorithm in graph modeling. However, the difference between the two lies in the subsequent feature update method. When updating the current node, the attention mechanism needs to consider the features of all other nodes. It can be understood that the current node has an edge with each of the other nodes. This is a fully connected graph, and the graph uses an edge selection strategy to limit feature updates to the nearest few nodes instead of considering all nodes. Therefore, the graph convolution is described as "selective fusion". This is why it is mentioned in the background technology that Transformer may introduce redundant information because it needs to consider all nodes.

[0084] Specifically, for the input For the local feature map in the window, we regard dimension 1 as a batch and dimension 2 as a sequence and extract features through the Attention module. The Attention module is actually modeled in the form of a sequence, so it requires the input to be in the form of (B, N, C). In the Attention module, B is the batch and N is the length of the sequence. The module can be expressed as:

[0085] Q=W 1 ·X in ,K=W 2 ·X in ,V=W 3 ·X in

[0086]

[0087] in are learnable parameters, Q, K, and V represent queries, keys, and values, respectively. They are all based on the feature map Xin It is calculated by the linear layer. Then we transform the feature dimension to obtain the feature map that integrates the local features.

[0088] (3) Graph convolution operations between windows

[0089] The feature extraction module based on graph convolution is designed to extract global features. For the input window feature map, each window is regarded as a node, so that the node set V = {v 1 ,v 2 ,...,v N} and node feature matrix Considering that C′ is in a larger feature dimension, we first reduce the feature dimension from C′ to d through convolution. Then we use the K nearest neighbor method to calculate the feature dimension of each node v i ∈V to find the K nearest neighbors N(v i ). For all v j ∈N(v i ), add a line from v j to v i The edge ji In this way, we construct an edge set E, where e ji ∈E, thus obtaining a graph representation of the window features. Then the graph convolution layer is used to exchange the information of the nodes by aggregating the features from its neighboring nodes. The graph convolution operation can be expressed as:

[0090] G′(X)=Update(Aggregate(G,W agg ),W update ) (5)

[0091] Among them, Aggregate represents the aggregation operation, Update represents the update operation, G represents the input window feature map, G′ represents the output of the graph convolution operation, and W agg and W update are learnable parameters in aggregation and update operations.

[0092] The graph convolution process can be further described as updating the feature representation of a node by aggregating the features of neighboring nodes. Here, edge convolution is used, so the above process can be simplified to

[0093] X′ out =EdgeConv(X out ·W 4 )·W 5 (6)

[0094] Among them, EdgeConv represents edge convolution, is the window feature map, is a learnable parameter. After obtaining the updated node features Afterwards, the feature graph integrating the long-range dependency is obtained by transforming the feature dimension.

[0095] (4) Fusion of local features and long-range dependencies.

[0096] After obtaining the feature map X″ that integrates local features in ∈R H×W×C and the feature graph X″ that integrates long-range dependencies out ∈R H×W×C Afterwards, we concatenate the two feature maps along the channel dimension and fuse the local features with the long-range dependencies through convolution operations. This process can be described as

[0097] X out =Mixer(X″ in ,X″ out )=BN(Contact(X″ in ,X″ out )·W 6 ) (7)

[0098] Among them, Contact(·) represents feature concatenation along the channel dimension. is a learnable parameter, BN is a normalization layer, It is a feature map that fuses local features and long-range dependencies.

[0099] 2.3 Multi-scale remixed tiles

[0100] We explore ways to use multi-scale information in a hierarchical architecture using graph modeling and design a multi-scale remixed patch (RGCB) module that integrates local and global correlations between feature maps at different scales by leveraging both detailed features and high-level semantic features including global dependencies and local contexts.

[0101] For an input H×W×C image, multi-scale features can be obtained i is the number of layers in the network. First, it is flattened and rearranged to have the same channel dimension C, thereby obtaining the converted multi-scale features Then, the multi-scale features of all layers are concatenated in the spatial dimension (dimension 1) and regarded as nodes in the spatial dimension, thereby obtaining the node set V Bridge and node feature graph The specific calculation is as follows: Input H×W×C image, according to Figure 1 As shown, the feature maps input to this module are The five inputs, after flattening and rearranging, become Considering the first dimension as nodes, we have nodes.

[0102] Then, based on the upper and lower layer relationships in the hierarchical architecture, edges are added to the nodes of adjacent layers to construct the edge set E Bridge ,like Figure 4 Specifically, considering that the pooling operation used in the encoder will pool the upper 2×2 feature blocks to form the lower 1×1 feature blocks, we add edges from the upper feature blocks to the lower feature blocks based on this relationship, thus obtaining the edge set E Bridge and a static graph representation of the hierarchical structure G Bridge =(V Bridge ,E Bridge ). Then the constructed static graph representation and the node feature map are input into the graph convolution module represented by formula (5) to capture the correlation between feature maps of different scales. Due to the use of edge convolution, this process can be simplified as follows:

[0103] X′ Bridge =EdgeConv(X Bridge ) (8)

[0104] Among them, EdgeConv represents edge convolution, X Bridge It is the node feature map of multi-scale features.

[0105] Finally, the concatenated multi-scale features are restored to feature maps of different scales and input into the decoder of the corresponding layer. Specifically, for the node feature map of the obtained multi-scale feature We sequentially split into feature maps of different scales along the first dimension Then, after the transformation of the feature dimension, we can get Five outputs. These five outputs will be input into the decoder of the corresponding layer.

[0106] 3. Method advantages

[0107] (1) The results obtained on the public abdominal multi-organ segmentation dataset are better than the existing methods. Table 1 summarizes the experimental results of different segmentation methods on the abdominal multi-organ segmentation dataset Synapse. We use two indicators, Dice consistency loss (DSC) and Hausdorff distance (HD), to evaluate the proposed method. In addition, Table 1 records the Dice consistency loss of different methods on different organs.

[0108] (2) The proposed method can better balance performance and model complexity. Table 2 summarizes the performance, parameter count, computational complexity (FLOPs), and model training time of different segmentation methods on the abdominal multi-organ segmentation dataset Synapse. Compared with AHGNN, which is also a graph modeling, the proposed method has fewer parameters and computational complexity. Compared with other methods of the same scale, the proposed method can achieve better results.

[0109] (3) Figure 5 The qualitative comparison between our method and other segmentation methods is presented. The comparison results show that our method is more sensitive to the boundary information of organs and can capture and generate more accurate organ edge information.

[0110] Table 1 Performance comparison with representative segmentation methods

[0111]

[0112] Table 2 Comparison with representative segmentation methods in terms of performance, number of parameters, amount of computation (FLOPs), and training time

[0113]

[0114]

[0115] The present application also provides a training device for an image segmentation model, the training device comprising:

[0116] A first acquisition unit is used to obtain training data, wherein the training data includes: an image to be segmented and an image segmentation result; and

[0117] a training unit, configured to train the image segmentation model using the training data,

[0118] The image segmentation model includes: an encoding part, a decoding part, an output part located at the back end of the decoding part, and a jump connection part located between the encoding part and the decoding part.

[0119] The encoding part is used to obtain a multi-scale feature map from the image to be segmented;

[0120] The jump connection part is used to perform a first division on at least one scale feature map with a predetermined size window to obtain a first-level feature map; perform a second division on the first-level feature map to obtain a second-level feature map; perform self-attention calculation on the second-level feature map to obtain a first feature map that integrates local features; perform a graph convolution operation on the first-level feature map to obtain a second feature map that integrates global features; and, concatenate the first feature map and the second feature map to obtain a fused feature map to output to the decoding part.

[0121] The present application also provides an image segmentation device, the segmentation device comprising:

[0122] A second acquisition unit, configured to obtain an image to be processed; and

[0123] The segmentation unit is used to process the image to be processed using the image segmentation model obtained through the training as described above to obtain an image processing result.

[0124] An embodiment of the present application also provides an electronic device, comprising: a processor, and a memory coupled to the processor, wherein the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the electronic device executes a method as described in any one of the above embodiments.

[0125] The electronic device may be a computing device such as a desktop computer, a notebook, a palmtop computer, a cloud server, etc. The electronic device may include, but is not limited to, a processor and a memory.

[0126] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire device.

[0127] The memory may be used to store the computer program, and the processor implements various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.

[0128] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, at least one application required for a function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0129] The embodiment of the present application also provides a storage medium, the storage medium is a computer-readable storage medium, the computer program is stored in the computer-readable storage medium, and the computer program, when executed by the processor, can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0130] An embodiment of the present application further provides a computer program product, including: a computer program or instructions, which, when executed on a computer, enables the computer to execute any of the above-mentioned possible implementation methods.

[0131] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.

Claims

1. A training method for an image segmentation model, characterized in that: The training method comprises: Obtaining training data, the training data including: an image to be segmented and an image segmentation result; and Using the training data, training the image segmentation model, The image segmentation model includes: an encoding part, a decoding part, an output part located at the back end of the decoding part, and a jump connection part located between the encoding part and the decoding part. The encoding part is used to obtain a multi-scale feature map from the image to be segmented; The jump connection part is used to perform a first division on at least one scale feature map with a predetermined size window to obtain a first-level feature map; perform a second division on the first-level feature map to obtain a second-level feature map; perform self-attention calculation on the second-level feature map to obtain a first feature map that integrates local features; perform a graph convolution operation on the first-level feature map to obtain a second feature map that integrates global features; and, concatenate the first feature map and the second feature map to obtain a fused feature map to output to the decoding part.

2. The training method according to claim 1, characterized in that: The skip connection part is also used for: The fused feature maps corresponding to the scale feature maps are concatenated in the spatial dimension with the same channel dimension to obtain a node set and a node feature map in the spatial dimension; Using the relationship between adjacent scale nodes as edges, constructing a graph representation corresponding to the node set; Performing a graph convolution operation on the graph representation to obtain a node feature representation reflecting the correlation between the fused feature graphs of different scales; as well as The node feature representation is restored at various scales and spliced ​​with the corresponding fused feature map to obtain a spliced ​​feature map which is output to the decoding part.

3. The training method according to claim 2, characterized in that: A graph convolution operation is performed on the graph representation to obtain a node feature representation that reflects the correlation between the fused feature graphs of different scales, specifically: an edge convolution operation is performed on the node feature graph to obtain the node feature representation.

4. The training method according to claim 1, characterized in that: The encoding part includes: a first layer encoder, a second layer encoder, a third layer encoder, a fourth layer encoder and a fifth layer encoder connected by pooling, the first layer encoder inputs the image to be segmented, the decoding part includes a first layer decoder, a second layer decoder, a third layer decoder, a fourth layer decoder and a fifth layer decoder connected by upsampling, the fifth layer encoder is output to the first layer decoder accordingly, the jump connection part includes a first jump connection between the first layer encoder and the fifth layer decoder, a second jump connection between the second layer encoder and the fourth layer decoder, a third jump connection between the third layer encoder and the third layer decoder, and a fourth jump connection between the fourth layer encoder and the second layer decoder, the second jump connection, the third jump connection and the fourth jump connection are used to obtain a fusion feature map corresponding to the corresponding scale.

5. The training method according to any one of claims 1 to 4, characterized in that: The image to be segmented is a computed tomography (CT) abdominal image or a magnetic resonance (MR) abdominal image.

6. An image segmentation method, characterized in that: The segmentation method comprises: Obtaining an image to be processed; The image to be processed is processed using the image segmentation model trained according to any one of claims 1 to 5 to obtain an image processing result.

7. A training device for an image segmentation model, characterized in that: The training device comprises: A first acquisition unit is used to obtain training data, wherein the training data includes: an image to be segmented and an image segmentation result; and a training unit, configured to train the image segmentation model using the training data, The image segmentation model includes: an encoding part, a decoding part, an output part located at the back end of the decoding part, and a jump connection part located between the encoding part and the decoding part. The encoding part is used to obtain a multi-scale feature map from the image to be segmented; The jump connection part is used to perform a first division on at least one scale feature map with a predetermined size window to obtain a first-level feature map; perform a second division on the first-level feature map to obtain a second-level feature map; perform self-attention calculation on the second-level feature map to obtain a first feature map that integrates local features; perform a graph convolution operation on the first-level feature map to obtain a second feature map that integrates global features; and, concatenate the first feature map and the second feature map to obtain a fused feature map to output to the decoding part.

8. An image segmentation device, characterized in that: The segmentation device comprises: A second acquisition unit, configured to obtain an image to be processed; and A segmentation unit is used to process the image to be processed using the image segmentation model trained according to any one of claims 1 to 5 to obtain an image processing result.

9. An electronic device, characterized in that: The electronic device comprises: a processor, and a memory coupled to the processor, The memory is used to store a computer program; and The processor is used to execute the computer program stored in the memory, so that the electronic device performs the training method as described in any one of claims 1 to 5, or performs the segmentation method as described in claim 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer executes the training method according to any one of claims 1 to 5, or executes the segmentation method according to claim 6.

Citation Information

Patent Citations

  • Three-dimensional image automatic segmentation method, system, equipment and medium

    CN115880312A

  • Image segmentation method and system based on step feature fusion and attention mechanism

    CN116309622A

  • Medical image segmentation model construction method based on multi-attention fusion

    CN116309648A

  • Medical image segmentation method and system based on double-branch embedded attention mechanism

    CN116309650A

  • Image segmentation method and device, storage medium and electronic device

    CN116402996A