Multi-scale heart image segmentation method and system based on graph neural network

Through a multi-scale cardiac image segmentation method based on graph neural network, combined with an adaptive loss function and a SAM optimizer, the limitations of multimodal cardiac image segmentation in the prior art are solved, and high-precision and robust cardiac CT and MRI image segmentation are achieved.

CN120198671APending Publication Date: 2025-06-24SHANDONG NORMAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510393224.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing multimodal heart image segmentation methods rely on paired input multimodal images, and the number of cardiac-related data sets is limited, mostly unpaired images, making it difficult to achieve fast and accurate cardiac CT and MRI image segmentation.

Method used

Using a multi-scale cardiac image segmentation method based on graph neural network, a multi-scale cardiac image segmentation model is constructed, combined with an adaptive loss function and a SAM optimizer, the model performance is optimized and segmentation accuracy and robustness are improved.

Benefits of technology

It significantly improves the segmentation accuracy and robustness of cardiac CT and MRI images, and can achieve high-quality image segmentation without pairing multimodal data, simplifying the clinical image analysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198671A_ABST
    Figure CN120198671A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image segmentation, in particular to a multi-scale heart image segmentation method and system based on a graph neural network, and the method specifically comprises the following steps: collecting heart CT and MRI data based on an in-vivo clinical environment, and dividing the data into a training set and a test set; nnUNetv2 is used as a basic framework to construct a multi-scale heart image segmentation model, the model comprises five encoder layers and five decoder layers, the training set is input into the constructed model, and a segmentation result predicted by the model is obtained; designing an adaptive loss function, optimizing the performance of the multi-scale heart image segmentation model by dynamically adjusting the weights of different loss functions in the adaptive loss function, and introducing a sharpness perception minimization optimizer to construct an uncertainty training mechanism to improve the generalization performance of the multi-scale heart image segmentation model; and inputting the data in the test set into the optimized multi-scale heart image segmentation model to obtain a final predicted segmentation result. According to the method, the segmentation precision and robustness of the heart CT and MRI images can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a multi-scale cardiac image segmentation method and system based on a graph neural network. Background Art

[0002] Heart diseases are one of the most common causes of death globally. A comprehensive analysis of a patient's specific cardiac structure and motion is the basis for understanding cardiac function, early detection, and accurate treatment of cardiovascular diseases. Computed tomography (CT) and magnetic resonance imaging (MRI) can clearly reflect the structure of cardiac tissue and the morphology and function of organs, providing important supplementary information for the assessment of heart diseases. In the face of limited datasets, it is of great significance for the diagnosis and treatment of heart diseases to quickly and accurately segment the entire heart from cardiac CT and MRI volumes through multi-modal learning.

[0003] In recent years, multi-modal cardiac image segmentation methods have made significant progress in many aspects, but there is still a key limitation: existing multi-modal segmentation methods usually rely on pairwise input multi-modal images. This method requires multi-modal images from the same patient and image registration in the preprocessing stage. However, the number of cardiac-related datasets is limited and most are unpaired images.

[0004] Therefore, the present invention proposes a multi-scale cardiac image segmentation method, system, and device based on a graph neural network to solve the above problems. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention develops a multi-scale cardiac image segmentation method and system based on a graph neural network. The present invention segments cardiac images by constructing a multi-scale cardiac image segmentation model and optimizes the model through an adaptive loss function and a SAM optimizer, which can improve the segmentation accuracy and robustness of cardiac CT and MRI images.

[0006] The technical solution for the present invention to solve the technical problem is a multi-scale cardiac image segmentation method based on a graph neural network, including the following steps: S1. Data collection: Collect cardiac CT and MRI data based on the in-vivo clinical environment, and divide the collected data into a training set and a test set according to a ratio; S2. Model Construction and Training: Construct a multi-scale cardiac image segmentation model using the medical image segmentation framework nnUNetv2 as the basic framework of the model. The model includes five encoder layers and five decoder layers. Each encoder layer sequentially includes a downsampling module, a residual convolutional layer, a Mamba module, and a pooling graph neural network module. Each decoder layer sequentially includes an upsampling module, a feature fusion module, a residual convolutional layer, and a pooling graph neural network module. A semantic boundary refinement module is designed at the skip connections between the first four encoder layers and decoder layers, and an MS-SSIM_Loss loss is designed after the semantic boundary refinement module to constrain the image features. Input the data in the training set into the constructed multi-scale cardiac image segmentation model for medical image segmentation to obtain the segmentation result predicted by the model. S3. Model Optimization: It involves an adaptive loss function, which is composed of multiple different loss functions. The performance of the multi-scale cardiac image segmentation model is optimized by dynamically adjusting the weights of different loss functions, and an uncertainty training mechanism is constructed by introducing the Sharpness-Aware Minimization optimizer SAM to improve the generalization performance of the multi-scale cardiac image segmentation model. Input the data in the test set into the optimized multi-scale cardiac image segmentation model to obtain the final predicted segmentation result.

[0007] In the specific implementation manner, the data collection is as follows: The collected cardiac CT and MRI data are from the MICCAI2017 and MICCAI2024 whole heart segmentation challenges. The data in the whole heart segmentation challenges are based on the in-vivo clinical environment and cover various heart diseases. The collected data are divided into a training set and a test set according to a ratio of 6:4.

[0008] In the specific implementation manner, the training of the model is as follows: (1) Input the cardiac images in the training set into the multi-scale cardiac image segmentation model. represents any cardiac image in the training set. The cardiac image first passes through the first encoder layer, sequentially through downsampling, a residual convolutional block, a Mamba module, and a pooling visual graph neural network module, and outputs the feature , and then the feature is respectively input into the second encoder layer and the first decoder layer. (2) The feature is input into the second encoder layer, and the feature is output. Then the feature is respectively input into the third encoder layer and the second decoder layer. The feature Input to the third encoder layer, and output features , and then input the features to the fourth encoder layer and the third decoder layer respectively; Input the features to the third encoder layer, and output features , and then input the features to the fifth encoder layer and the fourth decoder layer respectively; Input the features to the fifth encoder layer, and output features ; (3) Input the features to the residual convolution block, and output features , and then input the features to the fourth upsampling module, and output features ; (4) During the process of inputting the features to the fourth decoder layer, the features are first input to the semantic boundary refinement module of the fourth layer together with the features , and output features , then the features are further input to the upsampling module together with the features to obtain features and features respectively. Concatenate the features and the features to obtain features , input the features to the fourth decoder layer, and successively pass through the residual convolution block and the pooling visual graph neural network module to output features , and then input the features to the third upsampling module, and output features ; (5) During the process of inputting the features to the third decoder layer, the features are first input to the semantic boundary refinement module of the third layer together with the features , and output features , then the features are further input to the upsampling module together with the features to obtain features and features respectively. Concatenate the features and the features to obtain features , input the features to the third decoder layer, and successively pass through the residual convolution block and the pooling visual graph neural network module to output features , and then input the feature into the second upsampling module, and output the feature ; (6) During the process of inputting the feature into the second decoder layer, the feature is first input into the semantic boundary refinement module of the second layer together with the feature , and the output feature is obtained. Then, the feature is input into the upsampling module together with the feature to respectively obtain the feature and the feature . The feature and the feature are concatenated to obtain the feature . The feature is input into the second decoder layer, and successively passes through the residual convolution block and the pooling visual graph neural network module to output the feature . Then, the feature is input into the first upsampling module, and the output feature is obtained; (7) During the process of inputting the feature into the first decoder layer, the feature is first input into the semantic boundary refinement module of the first layer, and the output feature is obtained. Then, the feature is input into the upsampling module together with the feature to respectively obtain the feature and the feature . The feature and the feature are concatenated to obtain the feature . The feature is input into the first decoder layer, and successively passes through the residual convolution block and the pooling visual graph neural network module to output the feature . Then, the feature is input into the convolution block, and the segmentation result is obtained.

[0009] In the specific implementation manner, the Mamba module is specifically as follows: Input the local features extracted by the residual convolution block into the Mamba module. The local features are respectively processed by two branches in the Mamba module, and finally merged and passed through the linear layer, and then through the reshaping operation to finally output the features of multi-level and multi-scale information. The first branch of the Mamba module successively includes a linear layer, a convolutional layer, a SiLU activation function, and a selective structured state space model SSM. The second branch successively includes a linear layer and a SiLU activation function; The Mamba module specifically extracts local features with a dimension of from the received residual convolutional block. Among them, represents the batch size, represents the height of the feature map, represents the width of the feature map, represents the depth of the feature map, represents the dimension, which is converted into a sequence by flattening the spatial dimension , is the total length of the sequence. After the serialized features are layer-normalized, they enter the dual-branch processing: The operations of the first branch are as follows: After the serialized features enter the first branch, they first pass through a linear layer projection to expand the dimension, and the output shape of the sequence becomes . Then, the sequence is processed by a convolutional layer to obtain the association of local adjacent positions, and then through the SiLU activation function to introduce non-linearity. Then, a hidden state vector is initialized through a selective structured state space model (SSM). The hidden state stores the information of the previous and subsequent time steps in the feature vector sequence, and the feature vector sequence is gradually updated through the calculation of the state equation. The calculation formula is as follows: , , Among them, represents the input state vector from the feature vector sequence at time represents the hidden state vector at time represents the updated hidden state vector at time. A, B, C, and D represent the state transition matrix, control input matrix, observation matrix, and direct transfer matrix respectively. represents the output vector; The operations of the second branch are as follows: After the serialized features enter the second branch, they also first pass through a linear layer projection to expand the dimension, and the output shape of the sequence becomes . Then, a gating signal is generated through the SiLU activation function; The outputs of the two branch paths are fused through Hadamard convolution, and then projected through a linear layer to restore the channel dimension to , and finally reshaped into the original spatial form .

[0010] In the specific implementation manner, the pooling visual graph neural network module is as follows: Transfer the features of the multi-level and multi-scale information output by the Mamba module to the pooled vision graph neural network module. The pooled vision graph neural network module includes a pooled graph neural network module P-Grapher and a feed-forward neural network FFN. Aggregate and update the graph information through graph convolution in the P-Grapher module, and transform the node features through FFN; The feature operations of the pooled vision graph neural network on multi-level and multi-scale information are as follows: Use to represent the features of the multi-level and multi-scale information output by the Mamba module, , for perform times of downsampling operations to obtain the feature map The number of downsampling times in the height, width, and depth, , where, 、 and respectively represent the feature maps Again, 、 、 , represents index of, represents whether to perform pooling operation in the height dimension during the th downsampling of the feature map , perform , do not perform ; After passing through times of downsampling, integrate the learnable position embedding to further enhance the feature map , divide the feature map into non-overlapping feature map patches, and the resolution of each feature map patch is , for each convert to a feature vector , form a sequence = , where is the sequence length, represents the feature dimension; After the P-Grapher max pooling operation, the resolution of each becomes ([[]] ), form a new sequence = , the length becomes , then convert each patch to a feature vector = , and the feature dimension also becomes ; Regarding the feature vectors of the P-Grapher max pooling operation as unordered graph nodes , for each node , find K neighboring nodes through the K-nearest neighbor algorithm to construct an edge set , and construct a graph . Then, update the node information on the graph by using the dynamic max graph convolution operation, retain the original information through residual links, and avoid gradient vanishing. After restoring the resolution through the max unpooling operation, use the FFN to further transform and non-linearly activate the features processed by P-Grapher.

[0011] In the specific implementation manner, the semantic boundary refinement SBR module is as follows: The SBR module receives the depth features output from the current encoder layer and the skip connection features output from the next decoder layer through the upsampling module of this layer. Both the depth features and the skip connection features consist of a number of pixel points, and each pixel point represents a feature point. Solve the difference information between the feature points according to the point depth features and the skip connection features at each feature, and then solve the enhanced boundary features. Then, obtain the updated skip connection features based on the enhanced boundary features and the original skip connection features, and solve the final output of the SBR module through iterative update. The calculation process is as follows: , , , where represents the index of the feature point, represents the number of iterations, represents the skip connection feature at the -th iteration at the -th feature point, represents local neighborhoods centered on , , represents the index of the feature points within, represents the skip connection feature at the -th iteration at the -th feature point, represents the -th iteration at the -th feature point and the -th feature point difference information between, represents the The depth feature at a feature point denotes the depth feature at the th feature point denotes the semantic difference between and denotes the th iteration of the th skip connection feature at the denotes the weighted coefficient of the boundary feature denotes the weighted coefficient of the skip connection feature denotes the convolutional projection function The SBR module contains a learnable boundary enhancement kernel, abbreviated as the ABE kernel. The ABE kernel is in a 3×3×3 cube structure. By fixing the values of 8 vertices to 1 and -1, a clear differential relationship is formed at the edges of the cube. The remaining 19 parameters in the ABE kernel remain learnable and are automatically adjusted during training to generate an implicit differential relationship. The boundary feature is enhanced through the ABE kernel. The calculation formula for the enhanced boundary feature is as follows: , , where and are the learnable edge operators of the skip connection feature and the depth feature at the th feature point respectively denotes the set of vertices in the cube with values fixed to -1 denotes the set of vertices in the cube with values fixed to 1 denotes the remaining learnable parameters denotes the th convolutional kernel of size at the th feature point denotes the enhanced boundary feature at the th feature point in the

[0012] In the specific implementation, the adaptive loss function is specifically as follows: Construct an adaptive loss function based on uncertainty , the adaptive loss function regulates four different loss functions, namely the metric function for evaluating the similarity of two samples , the cross-entropy loss function for measuring the difference between the predicted class distribution of the model and the true label , the focal loss function and the shape distance function , and adjust the weights of different loss functions to optimize the model; (1) Adaptive loss function The calculation formula is as follows: , where, represents the number of loss functions, represents four different loss functions to be regulated, represents the homoscedastic uncertainty parameter related to the -th loss function; (2) Metric function for evaluating the similarity between two samples ranges from 0 to 1, being 1 indicates a perfect overlap between the prediction and the ground truth, and DSC being 0 indicates no overlap. The calculation formula is as follows: , , where, represents the segmentation result output by the heart segmentation model, represents the input heart image 's true label, represents the metric function of sample similarity; (3) Measure the difference between the class distribution predicted by the model and the true label through , and calculate the cross-entropy of the logarithmic probability of each pixel category and the true label. The calculation formula is as follows: , where, represents the total number of pixel points in the image, represents the segmentation result at the -th point, represents the true label at the -th point, represents the cross-entropy loss; (4) The calculation formula is as follows: , where, represents the adjustment factor, represents the focal loss function; (5) The calculation formula is as follows: , where, represents the distance map of the target, represents the shape distance loss function; By combining a metric function for evaluating the similarity of two samples , a function for measuring the difference between the class distribution predicted by the model and the true label , and into an uncertainty-based framework, the multi-scale heart segmentation model dynamically assigns weights to each loss component; The multi-scale heart segmentation model outputs a segmentation result, calculates the loss value between the segmentation result and the true label, and then feeds it back to the heart segmentation model for parameter update.

[0013] In the specific implementation manner, the SAM optimizer is specifically as follows: The SAM optimizer is introduced to optimize the heart segmentation model, and the SAM optimizer includes a two-step iterative process; In the first step, for each mini-batch of data, the heart segmentation model calculates the loss value for evaluating the model performance at the current parameter position, and the SAM optimizer defines a perturbation constraint range under the constraint condition . Within the range , by maximizing the loss function, a perturbation is found such that the loss value under this perturbation reaches the highest. The calculation process is as follows: , where, represents the perturbation that maximizes the loss function , represents the total loss function of the model, represents the model parameters, ∙ represents the input value that makes the function maximum; In the second step, the SAM optimizer uses the maximized perturbation to calculate the gradient of the loss function. The calculation formula is as follows: , , where, represents the gradient of the loss function, represents the learning rate, represents the gradient of the model parameters , represents the current parameter at the t-th iteration, represents the updated parameter.

[0014] The present invention also provides a multi-scale heart image segmentation system based on a graph neural network, which executes a multi-scale heart image segmentation method based on a graph neural network, including the following modules: Data collection module: Collect cardiac CT and MRI data based on the existing in-vivo clinical environment, and divide the collected data into a training set and a test set according to a certain proportion. Multi-scale cardiac image segmentation module: It includes an encoder module, a decoder module, and a semantic boundary refinement module, which processes the cardiac images in the training set of the data collection module to obtain the segmentation results of the cardiac images. Optimization module: It includes an adaptive loss function calculation module and a SAM optimizer, which optimizes the multi-scale cardiac image segmentation module, adjusts the internal parameters of the multi-scale cardiac image segmentation module to the optimal, and then inputs the cardiac images in the test set of the data collection module into the optimized multi-scale cardiac image segmentation module to obtain the final segmentation results of the cardiac images.

[0015] The effects provided in the invention content are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: The present invention discloses a multi-scale cardiac image segmentation method and system based on graph neural networks. By integrating the multi-scale sequence modeling ability of the Mamba module, the global topological perception characteristics of the pooling visual graph neural network, the boundary optimization mechanism of the semantic boundary refinement module, and the collaborative optimization strategy of the adaptive loss function and the SAM optimizer, the segmentation accuracy and robustness of cardiac CT and MRI images are significantly improved.

[0016] This method has achieved an overall lead in Dice coefficient and Jaccard index on authoritative test sets such as MICCAI2017; at the same time, it supports the direct processing of unpaired multi-modal data, greatly simplifies the clinical image analysis process, can quickly process three-dimensional cardiac images in different modalities, and provides a reliable basis for the diagnosis of heart diseases.

[0017] In addition, its dynamic adaptive mechanism and generalization optimization design enable the model to maintain stable performance in grass-roots medical scenarios and under limited data conditions, effectively promoting the popularization of accurate diagnosis and treatment of heart diseases and the improvement of scientific research efficiency. Brief Description of the Drawings

[0018] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention.

[0019] Figure 1 It is a flow schematic diagram of the method in the present invention.

[0020] Figure 2 It is a comparison chart of experimental results. Detailed Embodiments

[0021] To clearly illustrate the technical features of this solution, the present invention will be elaborated in detail below through specific embodiments and in conjunction with its accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention.

[0022] Embodiment 1 A multi-scale cardiac image segmentation method based on graph neural network, comprising the following steps: S1. Data collection: Collect cardiac CT and MRI data based on the in-vivo clinical environment, and divide the collected data into a training set and a test set according to a certain ratio. S2. Model construction and training: Construct a multi-scale cardiac image segmentation model, using the medical image segmentation framework nnUNetv2 as the basic framework of the model. The model includes five encoder layers and five decoder layers. Each encoder layer sequentially includes a downsampling module, a residual convolutional layer, a Mamba module, and a pooling graph neural network module. Each decoder layer sequentially includes an upsampling module, a feature fusion module, a residual convolutional layer, and a pooling graph neural network module. And a semantic boundary refinement module is designed at the skip connection of the first four encoder layers and decoder layers, and then an MS-SSIM_Loss multi-scale structural similarity loss is designed after the semantic boundary refinement module to constrain the image features. Input the data in the training set into the constructed multi-scale cardiac image segmentation model for medical image segmentation to obtain the segmentation result predicted by the model. S3. Model optimization: Design an adaptive loss function that combines multiple different loss functions, and optimize the performance of the multi-scale cardiac image segmentation model by dynamically adjusting the weights of different loss functions. Optimize the constructed model. This loss function combines multiple different loss functions, and constructs an uncertainty training mechanism by introducing a sharpness-aware minimization optimizer to improve the generalization performance of the multi-scale cardiac image segmentation model. Input the data in the test set into the optimized multi-scale cardiac image segmentation model to obtain the final predicted segmentation result.

[0023] In the specific embodiment, the data collection is specifically as follows: The collected cardiac CT and MRI data are from the MICCAI2017 and MICCAI2024 whole heart segmentation challenges. The data in the whole heart segmentation challenges are based on the in-vivo clinical environment and cover a variety of heart diseases. The collected data are divided into a training set and a test set according to a ratio of 6:4.

[0024] In the specific embodiment, the training of the model is specifically as follows: (1) Input the cardiac images in the training set into the multi-scale cardiac image segmentation model, denote any cardiac image in the training set, the cardiac image First, it passes through the first encoder layer, successively through downsampling, residual convolutional blocks, Mamba modules, and pooling visual graph neural network modules, and outputs features , and then the features are respectively input into the second encoder layer and the first decoder layer; (2) The features are input into the second encoder layer, and the output features are obtained. Then the features are respectively input into the third encoder layer and the second decoder layer; The features are input into the third encoder layer, and the output features are obtained. Then the features are respectively input into the fourth encoder layer and the third decoder layer; The features are input into the third encoder layer, and the output features are obtained. Then the features are respectively input into the fifth encoder layer and the fourth decoder layer; The features are input into the fifth encoder layer, and the output features ; (3) The features are input into the residual convolutional block, and the output features are obtained. Then the features are input into the fourth-layer upsampling module, and the output features are obtained; (4) During the process of inputting the features into the fourth decoder layer, the features are first input into the semantic boundary refinement module of the fourth layer together with the features , and the output features are obtained. Then the features are further input into the upsampling module together with the features to respectively obtain the features and the features . The features and the features are concatenated to obtain the features . The features are input into the fourth decoder layer, successively through the residual convolutional block and the pooling visual graph neural network module, and the output features are obtained. Then the features are input into the third-layer upsampling module, and the output features are obtained; (5) During the process of inputting the features into the third decoder layer, the features are first combined with the features are jointly input into the semantic boundary refinement module of the third layer to output features , and then the features are combined with the features to jointly pass through the upsampling module to respectively obtain features and features . The features and the features are subjected to feature concatenation to obtain features . The features are input into the third decoder layer, and successively pass through the residual convolution block and the pooling visual graph neural network module to output features . Then, the features are input into the upsampling module of the second layer to output features ; (6) During the process of inputting the features into the second decoder layer, the features are first jointly input into the semantic boundary refinement module of the second layer with the features to output features . Then, the features are combined with the features to jointly pass through the upsampling module to respectively obtain features and features . The features and the features are subjected to feature concatenation to obtain features . The features are input into the second decoder layer, and successively pass through the residual convolution block and the pooling visual graph neural network module to output features . Then, the features are input into the upsampling module of the first layer to output features ; (7) During the process of inputting the features into the first decoder layer, the features are first input into the semantic boundary refinement module of the first layer to output features . Then, the features are combined with the features to jointly pass through the upsampling module to respectively obtain features and features . The features and the features are subjected to feature concatenation to obtain features . The features are input into the first decoder layer, and successively pass through the residual convolution block and the pooling visual graph neural network module to output features . Then, the features are input into the convolution block to output the segmentation result 。

[0025] In the specific implementation manner, the Mamba module is as follows: The local features extracted by the residual convolution block are input into the Mamba module. The local features are processed by two branches in the Mamba module respectively, and finally merged through a linear layer, and then through a reshaping operation to finally output features with multi-level and multi-scale information. The first branch of the Mamba module sequentially includes a linear layer, a convolutional layer, a SiLU activation function, and a selective structured state space model SSM. The second branch sequentially includes a linear layer and a SiLU activation function; The Mamba module specifically processes the local features with the dimension of extracted by the received residual convolution block, where represents the batch size, represents the height of the feature map, represents the width of the feature map, represents the depth of the feature map, represents the dimension, and it is converted into a sequence by flattening the spatial dimension , is the total length of the sequence. After the serialized features are layer-normalized, they enter the dual-branch processing: The operations of the first branch are specifically as follows: After the serialized features enter the first branch, they first pass through a linear layer for projection to expand the dimension, and the output shape of the sequence becomes , then the sequence is processed by a convolutional layer to obtain the association of local adjacent positions, and then through the SiLU activation function to introduce non-linearity, and then a hidden state vector is initialized through the selective structured state space model SSM. The hidden state stores the information of the front and back time steps in the feature vector sequence, and the feature vector sequence is gradually updated through the calculation of the state equation. The calculation formula is as follows: , , where represents the input state vector from the feature vector sequence at time represents the hidden state vector at time represents the updated hidden state vector at time, A, B, C, D respectively represent the state transition matrix, the control input matrix, the observation matrix, and the direct transfer matrix, represents the output vector; The operations of the second branch are specifically as follows: After the serialized features enter the second branch, they first go through a linear layer projection to expand the dimension, and the output shape of the sequence becomes , and then pass through the SiLU activation function to generate a gating signal; The outputs of the two branch paths are fused through Hadamard convolution, and then projected through a linear layer to restore the channel dimension to , and finally reshaped into the original spatial form .

[0026] In the specific implementation manner, the pooling visual graph neural network module is as follows: The features of the multi-level and multi-scale information output by the Mamba module are transmitted to the pooling visual graph neural network module. The pooling visual graph neural network module includes a pooling graph neural network module P-Grapher and a feed-forward neural network FFN. The graph convolution in the P-Grapher module is used to aggregate and update the graph information, and the node features are transformed through the FFN; The operations of the pooling visual graph neural network on the features of multi-level and multi-scale information are as follows: Use to represent the features of the multi-level and multi-scale information output by the Mamba module, , for perform times of downsampling operations to obtain the feature map The number of times of downsampling in the height, width, and depth, , where , and respectively represent the feature map Again, , , , represents 's index, represents the feature map in the th downsampling whether to perform pooling operation in the height dimension, perform , do not perform ; After passing through times of downsampling, integrate learnable position embeddings to further enhance the feature map , divide the feature map into non-overlapping feature map patches, and the resolution of each feature map patch is , for each convert to a feature vector , forming a sequence = , where is the sequence length, representing the feature dimension; After the P-Grapher max pooling operation, each has a resolution of ( ), forming a new sequence = , with the length becoming , and then each patch is converted into a feature vector = , and the feature dimension also becomes ; Consider the feature vectors of the P-Grapher max pooling operation as unordered graph nodes , for each node , find K neighboring nodes through the K-nearest neighbor algorithm to construct an edge set , and construct a graph . Then, on the graph , update the node information by using the dynamic max graph convolution operation, and retain the original information through the residual link to avoid gradient disappearance. After restoring the resolution through the max unpooling operation, use FFN to further transform and non-linearly activate the features processed by P-Grapher.

[0027] In the specific implementation manner, the semantic boundary refinement SBR module is as follows: The SBR module receives the depth features output from the current encoder layer and the skip connection features output from the next decoder layer after passing through the upsampling module of this layer. Both the depth features and the skip connection features are composed of several pixel points, and each pixel point represents a feature point. Solve the difference information between the feature points according to the depth features and the skip connection features at each feature point, and then solve the enhanced boundary features. Then, obtain the updated skip connection features according to the enhanced boundary features and the original skip connection features, and solve the final output of the SBR module through iterative update. The calculation process is as follows: , , , Among them, represents the index of the feature point, represents the number of iterations, represents the skip connection feature at the -th iteration and the -th feature point, represents local neighborhoods centered on , , Indicates the index of the internal feature point, indicating the th iteration and the skip connection feature at the th feature point, indicating the th iteration and the sum of the th feature point and the difference information between the th feature point, indicating the depth feature at the th feature point, indicating the depth feature at the Indicates and the semantic difference between them, indicating the th iteration and the skip connection feature at the indicating the weighting coefficient of the boundary feature, indicating the weighting coefficient of the skip connection feature, indicating the convolutional projection function; The SBR module contains a learnable boundary enhancement kernel, abbreviated as the ABE kernel. The ABE kernel is in a 3×3×3 cube structure. By fixing the values of 8 vertices to 1 and -1, clear differential relationships are formed at the edges of the cube. The remaining 19 parameters in the ABE kernel remain learnable and are automatically adjusted during training to generate implicit differential relationships. The boundary feature is enhanced through the ABE kernel. The calculation formula for the enhanced boundary feature is as follows: , , where, and are the learnable edge operators of the skip connection feature and the depth feature at the th feature point respectively, represents the set of vertices in the cube with values fixed to -1, represents the set of vertices in the cube with values fixed to 1, represents the remaining learnable parameters, indicating the th feature point and the convolutional kernel of size at that point, indicating the th iteration and the enhanced boundary feature at the

[0028] In the specific implementation manner, the adaptive loss function is specifically as follows: Construct an adaptive loss function based on uncertainty , the adaptive loss function regulates four different loss functions, namely the metric function for evaluating the similarity between two samples , the cross-entropy loss function for measuring the difference between the class distribution predicted by the model and the true label , the focal loss function and the shape distance function , and adjusts the weights of different loss functions to optimize the model; (1) The calculation formula of the adaptive loss function is as follows: , wherein, represents the number of loss functions, represents the four different loss functions regulated, represents the homoscedastic uncertainty parameter related to the th loss function; (2) The metric function for evaluating the similarity between two samples ranges from 0 to 1, being 1 indicates a perfect overlap between the prediction and the ground truth, and DSC being 0 indicates no overlap. The calculation formula is as follows: , , wherein, represents the segmentation result output by the heart segmentation model, represents the input heart image 's true label, represents the metric function of sample similarity; (3) Measure the difference between the class distribution predicted by the model and the true label through , calculate the cross-entropy of the logarithmic probability of each pixel category and the true label. The calculation formula is as follows: , wherein, represents the total number of pixel points in the image, represents the segmentation result at the th point, represents the true label at the th point, represents the cross-entropy loss; (4) The calculation formula of , Among them, represents a regulatory factor, represents the focal loss function; (5) The calculation formula of , Among them, represents the distance mapping of the target, represents the shape distance loss function; By combining the metric function used to evaluate the similarity between two samples, the function used to measure the difference between the class distribution predicted by the model and the true label, and into the uncertainty-based framework, the multi-scale cardiac segmentation model dynamically assigns weights to each loss component; The multi-scale cardiac segmentation model outputs the segmentation result, calculates the loss value between the segmentation result and the true label, and then feeds it back to the cardiac segmentation model for parameter update.

[0029] In the specific implementation manner, the SAM optimizer is specifically as follows: Introduce the SAM optimizer to optimize the cardiac segmentation model. The SAM optimizer includes a two-step iterative process; In the first step, for each mini-batch of data, the cardiac segmentation model calculates the loss value for evaluating the model performance at the current parameter position. The SAM optimizer defines a perturbation constraint range under the constraint condition , and within the range , by maximizing the loss function, find a perturbation such that the loss value under this perturbation reaches the highest. The calculation process is as follows: , Among them, represents the perturbation that maximizes the loss function , represents the total loss function of the model, represents the model parameters, ∙ represents the input value that makes the function maximum; In the second step, the SAM optimizer uses the maximized perturbation to calculate the gradient of the loss function. The calculation formula is as follows: , , Among them, represents the gradient of the loss function, represents the learning rate, represents the model parameters of the gradient, represents the current parameters at the t-th iteration, represents the updated parameters.

[0030] Example 2 The present invention also provides a multi-scale cardiac image segmentation system based on a graph neural network, which executes a multi-scale cardiac image segmentation method based on a graph neural network, including the following modules: Data collection module: Collect cardiac CT and MRI data based on the existing in-vivo clinical environment, and divide the collected data into a training set and a test set according to a certain proportion; Multi-scale cardiac image segmentation module: Includes an encoder module, a decoder module, and a semantic boundary refinement module, processes the cardiac images in the training set of the data collection module, and obtains the segmentation results of the cardiac images; Optimization module: Includes an adaptive loss function calculation module and a SAM optimizer, optimizes the multi-scale cardiac image segmentation module, adjusts the internal parameters of the multi-scale cardiac image segmentation module to the optimal, and then inputs the cardiac images in the test set of the data collection module into the optimized multi-scale cardiac image segmentation module to obtain the final segmentation results of the cardiac images.

[0031] Example 3 To prove the beneficial effects of the method in the present invention, the method in the present invention is compared with existing methods (nnFormer: non-nested Transformer medical image segmentation network, SwinUNETR: a model combining SwinTransformer with deformable window attention mechanism and UNETR, UNTER++: an improved model integrating multi-scale feature enhanced UNet and Transformer, X-shape: a symmetric medical image segmentation architecture with cross-path aggregation, Zhang et al: a generative adversarial network combining cycle consistency and shape consistency constraints, C-ViT: conditional normalization cross-modal segmentation framework) on the CT test set and MRI test set of MICCAI2017. The evaluation metrics include the Dice coefficient and the Jaccard index. The Dice coefficient calculates the ratio of the intersection area of the predicted region and the ground truth label to the average area of the two, and the value range is [0, 1]. The larger the value, the higher the overlap and the higher the segmentation accuracy; the Jaccard index calculates the ratio of the intersection area of the predicted region and the ground truth label to the union area of the two, and the value range is also [0, 1]. The larger the value, the higher the segmentation accuracy; Table 1 and Table 2 respectively show the segmentation results of the cardiac segmentation model for 7 cardiac substructures in two modalities of CT and MRI: LV (left ventricular blood cavity), Myo (left ventricular myocardium), RV (right ventricular blood cavity), LA (left atrial blood cavity), RA (right atrial blood cavity), LO (ascending aorta), PA (pulmonary artery), and WHS (whole heart); it can be seen from Table 1 and Table 2 that the cardiac segmentation model proposed by the present invention shows significant advantages in both CT and MRI dual-modal medical images, and its key segmentation accuracy indicators (Dice, Jaccard) both exceed the existing advanced models.

[0032] Table 1 Experimental comparison between the method in the present invention and the existing methods based on the CT test set of MICCAI2017 Table 2 Experimental comparison between the method in the present invention and the existing methods based on the MRI test set of MICCAI2017 To more intuitively verify the segmentation results of the method in the present invention, as Figure 2 shown, the segmentation results of the method in the present invention and the existing methods are visualized. According to Figure 2 , a qualitative evaluation of the segmentation results of the two modalities can be carried out. At the place pointed by the red arrow in the figure, it can be fully shown that the cardiac segmentation model proposed in this patent can better identify different cardiac substructures and has a more accurate segmentation of the boundaries; the place pointed by the yellow arrow is the interconnected left ventricular blood cavity and left atrial blood cavity. The model in the method of the present invention has made a more refined division of the boundaries of the two substructures in both modalities.

[0033] Taking all the results together, the present invention has achieved significant performance improvements in fitting cardiac substructures of various scales, shapes, and intensities. This shows that the method in the present invention has better image segmentation effects on medical imaging data sets with different imaging modalities (MRI and CT).

[0034] Although the specific implementation manners of the invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the present invention.

Claims

1. A multi-scale cardiac image segmentation method based on graph neural network, characterized in that: The following steps are involved: S1. Data collection: Collect cardiac CT and MRI data based on the in vivo clinical environment, and divide the collected data into training and test sets in proportion; S2. Model construction and training: Construct a multi-scale cardiac image segmentation model, using the medical image segmentation framework nnUNetv2 as the basic framework of the model. The model includes five encoder layers and five decoder layers. Each encoder layer includes a downsampling module, a residual convolution layer, a Mamba module, and a pooling graph neural network module in sequence. Each decoder layer includes an upsampling module, a feature fusion module, a residual convolution layer, and a pooling graph neural network module in sequence. A semantic boundary refinement module is designed at the jump connection of the first four encoder layers and the decoder layer, and then the MS-SSIM_Loss multi-scale structural similarity loss is designed after the semantic boundary refinement module to constrain image features. The data in the training set is input into the constructed multi-scale cardiac image segmentation model to perform medical image segmentation and obtain the segmentation results predicted by the model; S3, model optimization: involves adaptive loss function, which is composed of multiple different loss functions. The performance of the multi-scale cardiac image segmentation model is optimized by dynamically adjusting the weights of different loss functions, and the sharpness-aware minimization optimizer SAM is introduced to construct an uncertainty training mechanism to improve the generalization performance of the multi-scale cardiac image segmentation model. The data in the test set are input into the optimized multi-scale cardiac image segmentation model to obtain the final predicted segmentation result.

2. According to the multi-scale cardiac image segmentation method based on graph neural network in claim 1, it is characterized in that: The data collected are as follows: The collected cardiac CT and MRI data are from the MICCAI2017 and MICCAI2024 whole heart segmentation challenges. The data in the whole heart segmentation challenge are based on the in vivo clinical environment and cover a variety of heart diseases. The collected data are divided into training set and test set in a ratio of 6:

4.

3. The multi-scale cardiac image segmentation method based on graph neural network according to claim 2, characterized in that: The model training is as follows: (1) The heart images in the training set Input to the multi-scale cardiac image segmentation model, represents any heart image in the training set, heart image First, it passes through the first encoder layer, and then passes through downsampling, residual convolution block, Mamba module and pooling visual graph neural network module to output features. , and then the feature Input to the second encoder layer and the first decoder layer respectively; (2) Characteristics Input to the second encoder layer, output features , and then the feature Input to the third encoder layer and the second decoder layer respectively; The characteristics Input to the third encoder layer, output features , and then the feature Input to the fourth encoder layer and the third decoder layer respectively; The characteristics Input to the third encoder layer, output features , and then the feature Input to the fifth encoder layer and the fourth decoder layer respectively; The characteristics Input to the fifth encoder layer, output features ; (3) Features Input to the residual convolution block and output features , and then the feature Input to the fourth layer upsampling module, output features ; (4) Features During the input to the fourth decoder layer, the feature First and features The output features are fed into the semantic boundary refinement module of the fourth layer. , then the feature Again with characteristics Together through the upsampling module to obtain the features and Features , the feature and Features Perform feature splicing to obtain features , the feature Input to the fourth decoder layer, pass through the residual convolution block and the pooling visual graph neural network module in turn, and output the feature , and then the feature Input to the third layer upsampling module, output features ; (5) Features During the input to the third decoder layer, the feature First and features The output features are fed into the semantic boundary refinement module of the third layer. , then the feature Again with characteristics Together through the upsampling module to obtain the features and Features , the feature and Features Perform feature splicing to obtain features , the feature Input to the third decoder layer, pass through the residual convolution block and the pooling visual graph neural network module in turn, and output the feature , and then the feature Input to the second layer upsampling module and output features ; (6) Characteristics During the input to the second decoder layer, the feature First and features The two are input together to the semantic boundary refinement module of the second layer, and the output features , then the feature Again with characteristics Together through the upsampling module to obtain the features and Features , the feature and Features Perform feature splicing to obtain features , the feature Input to the second decoder layer, pass through the residual convolution block and the pooling visual graph neural network module in turn, and output the feature , and then the feature Input to the first layer upsampling module and output features ; (7) Characteristics During the input to the first decoder layer, the feature First, input to the semantic boundary refinement module of the first layer, and output feature , then the feature Again with characteristics Together through the upsampling module to obtain the features and Features , the feature and Features Perform feature splicing to obtain features , the feature Input to the first decoder layer, pass through the residual convolution block and the pooling visual graph neural network module in turn, and output the feature , and then the feature Input to the convolution block and output the segmentation result .

4. The multi-scale cardiac image segmentation method based on graph neural network according to claim 3 is characterized in that: The Mamba modules are as follows: The local features extracted by the residual convolution block are input into the Mamba module. The local features are processed separately by two branches in the Mamba module, and finally merged and passed through the linear layer. After the reshaping operation, the features of multi-level and multi-scale information are finally output. The first branch of the Mamba module includes a linear layer, a convolutional layer, a SiLU activation function, and a selective structured state space model SSM in sequence, and the second branch includes a linear layer and a SiLU activation function in sequence; The Mamba module specifically extracts the dimension of the received residual convolution block as follows: The local features of Indicates the batch size, represents the feature map height, represents the feature map width, represents the feature map depth, Represents the dimension, which is converted into a sequence by flattening the spatial dimension , is the total length of the sequence. After the serialized features are normalized by layers, they enter the dual-branch processing: The operations of the first branch are as follows: After the serialized features enter the first branch, they are first projected through a linear layer to expand the dimension, and the output shape of the sequence becomes , and then use the convolution layer to process the sequence to obtain the association between local adjacent positions, and then introduce nonlinearity through the SiLU activation function, and then initialize a hidden state vector through the selective structured state space model SSM, use the hidden state to store the information of the previous and next time steps in the feature vector sequence, and gradually update the feature vector sequence through the calculation of the state equation. The calculation formula is as follows: , , in, express The input state vector at time instant comes from the sequence of feature vectors, express The hidden state vector at time t, express The hidden state vector after the moment is updated, A, B, C, and D represent the state transfer matrix, control input matrix, observation matrix, and direct transfer matrix respectively. A vector describing the output; The operation of the second branch is as follows: After the serialized features enter the second branch, they are also projected through a linear layer to expand the dimension, and the output shape of the sequence becomes , and then generate the gating signal through the SiLU activation function; The outputs of the two branch paths are fused through Hadamard convolution, and the channel dimension is restored after linear layer projection. , and finally reshaped into the original spatial form .

5. The multi-scale cardiac image segmentation method based on graph neural network according to claim 4 is characterized in that: The pooling visual graph neural network module is as follows: The features of the multi-level and multi-scale information output by the Mamba module are passed to the pooling visual graph neural network module, which includes the pooling graph neural network module P-Grapher and the feedforward neural network FFN. The graph information is aggregated and updated through the graph convolution in the P-Grapher module, and the node features are transformed through the FFN. The pooling visual graph neural network operates on the features of multi-level and multi-scale information as follows: use Represents the characteristics of the multi-level and multi-scale information output by the Mamba module, ,right conduct Downsampling operation, get the feature map The number of times to downsample in height, width, and depth, ,in, , and Represents the feature maps Again, , , , express The index of Representation feature map In the Whether to perform pooling in the height dimension during downsampling, , not executed ; In passing After downsampling, we integrate learnable position embeddings to further enhance the feature map , the feature map Divided into non-overlapping feature patches, the resolution of each feature patch is , each Convert to feature vector , forming a sequence =[ ,in is the sequence length, Represents feature dimension; After the maximum pooling operation of P-Grapher, each The resolution becomes ( ), forming a new sequence =[ , the length becomes , and then convert each patch into a feature vector =[ , the feature dimension also becomes ; Treat the feature vector of the P-Grapher max pooling operation as an unordered graph node , for each node , find K neighboring nodes through the K nearest neighbor algorithm and build an edge set , component diagram construction diagram Then in the figure The node information is updated by using the dynamic maximum graph convolution operation, and the original information is retained through the residual link to avoid gradient disappearance. After the maximum unpooling operation restores the resolution, FFN is used to further transform and nonlinearly activate the features processed by P-Grapher.

6. The multi-scale cardiac image segmentation method based on graph neural network according to claim 5, characterized in that: The semantic boundary refinement SBR module is as follows: The SBR module receives the depth features output from the encoder layer and the jump connection features output by the next decoder layer through the upsampling module of this layer. The depth features and jump connection features are composed of several pixels, each pixel represents a feature point. The difference information between the feature points is solved according to the depth features and jump connection features at each feature point, and then the enhanced boundary features are solved. Then, the updated jump connection features are obtained according to the enhanced boundary features and the original jump connection features, and the final output of the SBR module is solved through iterative update. The calculation process is as follows: , , , in, Represents the index of the feature point, represents the number of iterations, Indicates In the iteration The skip connection features at feature points, express One The local neighborhood centered on , express The index of the internal feature point, Indicates In the iteration The skip connection features at feature points, Indicates In the iteration Feature points The feature points The difference information between Indicates The depth features at feature points, Indicates The depth features at feature points, express and The semantic difference between Indicates In the iteration The skip connection features at feature points, represents the weight coefficient of the boundary feature, represents the weight coefficient of the skip connection feature, represents the convolution projection function; The SBR module contains a learnable boundary enhancement kernel, referred to as the ABE kernel. In a 3×3×3 cube structure, the ABE kernel fixes the values ​​of 8 vertices to 1 and -1 to form a clear differential relationship on the edge of the cube. The remaining 19 parameters in the ABE kernel remain learnable and are automatically adjusted during the training process to generate implicit differential relationships. The ABE kernel is used to enhance the correction of boundary features. The calculation formula of the enhanced boundary features is as follows: , , in, and Separate A learnable edge operator that jumps between features and deep features at feature points. Represents a set of vertices in a cube whose values ​​are fixed to -1. Represents a set of vertices in a cube whose values ​​are fixed to 1. represents the remaining learnable parameters, Indicates The size of the feature point is The convolution kernel, Indicates In the iteration The enhanced boundary features at feature points.

7. The multi-scale cardiac image segmentation method based on graph neural network according to claim 6 is characterized in that: The loss function is as follows: Constructing an adaptive loss function based on uncertainty , adaptive loss function Adjust four different loss functions, which are the metric functions used to evaluate the similarity between two samples. , the cross entropy loss function used to measure the difference between the category distribution predicted by the model and the true label , focal loss function and shape distance function , and adjust the weights of different loss functions to optimize the model; (1) Adaptive loss function The calculation formula is as follows: , in, represents the number of loss functions, express Four different loss functions for regulation, Indicates Homoscedastic uncertainty parameters associated with the loss functions; (2) Metric function used to evaluate the similarity between two samples The value range is between 0 and 1. A DSC of 1 indicates perfect overlap between the prediction and the ground truth, and a DSC of 0 indicates no overlap. The calculation formula is as follows: , , in, Represents the segmentation result output by the heart segmentation model, Represents the input heart image The real label, A metric function that represents sample similarity; (3) By To measure the difference between the category distribution predicted by the model and the true label, the cross entropy between the log probability of each pixel category and the true label is calculated as follows: , in, Represents the total number of pixels in the image. Indicates The segmentation result at the point is Indicates The true label at each point, represents the cross entropy loss; (4) The calculation formula is as follows: , in, represents the adjustment factor, represents the focal loss function; (5) The calculation formula is as follows: , in, represents the distance map of the target, represents the shape distance loss function; By using the metric function used to evaluate the similarity between two samples , a function used to measure the difference between the class distribution predicted by the model and the true label , and Incorporated into the uncertainty-based framework, the multi-scale cardiac segmentation model dynamically assigns weights to each loss component; The multi-scale heart segmentation model outputs the segmentation result, calculates the loss value between the segmentation result and the true label, and then feeds it back to the heart segmentation model for parameter update.

8. The multi-scale cardiac image segmentation method based on graph neural network according to claim 7, characterized in that: The SAM optimizer is as follows: The SAM optimizer is introduced to optimize the heart segmentation model. The SAM optimizer includes a two-step iterative process; In the first step, for each mini-batch of data, the heart segmentation model calculates the loss value for evaluating the model performance at the current parameter position. The SAM optimizer calculates the loss value for evaluating the model performance at the current parameter position. Define a perturbation constraint range , in the range By maximizing the loss function, we find a perturbation that maximizes the loss value under this perturbation. The calculation process is as follows: , in, Represents the loss function Maximizing the disturbance, Represents the total loss function of the model, represents the model parameters, ∙ represents the input value that maximizes the function; In the second step, the SAM optimizer uses the maximized perturbation to calculate the gradient of the loss function. The calculation formula is as follows: , , in, represents the gradient of the loss function, represents the learning rate, Represents model parameters The gradient of represents the current parameters at the tth iteration, Indicates the updated parameters.

9. A multi-scale cardiac image segmentation system based on graph neural network, executing a multi-scale cardiac image segmentation method based on graph neural network as claimed in any one of claims 1 to 8, characterized in that: Includes the following modules: Data collection module: collects cardiac CT and MRI data based on the existing in vivo clinical environment, and divides the collected data into training sets and test sets in proportion; Multi-scale cardiac image segmentation module: including encoder module, decoder module, semantic boundary refinement module, processing cardiac images in the training set in the data collection module, and obtaining the segmentation results of cardiac images; Optimization module: includes an adaptive loss function calculation module and a SAM optimizer, which optimizes the multi-scale cardiac image segmentation module, adjusts the internal parameters of the multi-scale cardiac image segmentation module to the optimal level, and then inputs the cardiac image in the test set in the data collection module into the optimized multi-scale cardiac image segmentation module to obtain the final cardiac image segmentation result.

Citation Information

Cited By

  • Dynamic graph neural network trauma image segmentation method, electronic equipment and computer program product

    CN120543863A

  • 3D hemodynamic atlas construction method for ultrasonic examination

    CN120807797A

  • Free breathing rapid three-dimensional heart magnetic resonance imaging method and system

    CN120894508A