Convolutional neural network and graph network combined medical image segmentation method

By combining the convolutional neural network and graph network methods, texture and topological features in medical images are extracted and feature fusion is performed, which solves the problems of insufficient accuracy and poor robustness in coronary artery segmentation in traditional methods, achieving more accurate and continuous vascular segmentation results.

CN119991723APending Publication Date: 2025-05-13FUDAN UNIVERSITY

Patent Information

Application Number
CN202510067649.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional medical image segmentation methods have problems of insufficient accuracy and poor robustness when dealing with complex and variable medical images. Especially in the segmentation task of coronary artery, it is difficult to achieve accurate segmentation and continuous results relying solely on local texture features.

Method used

The method of combining convolutional neural networks and graph networks is adopted to extract the texture features and topological features of the coronary artery through a double-branch parallel encoder, and the attention mechanism module is used to fusion feature to construct a three-dimensional visual map network medical segmentation network to achieve accurate segmentation of the coronary artery.

Benefits of technology

It improves the accuracy and continuity of coronary artery segmentation, can more effectively extract three-dimensional vascular structural features, and provides more accurate visual results for the diagnosis and evaluation of cardiovascular diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991723A_ABST
    Figure CN119991723A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image segmentation, in particular to a medical image segmentation method combining a convolutional neural network and a graph network, and the method comprises the following steps: forming a data set CoronarySet by using coronary artery CT angiography three-dimensional image data collected by a computed tomography (CT) device, and randomly dividing the data set CoronarySet into a training set CoronarySet 1 and a test set CoronarySet 2; the method comprises the following steps of: segmenting data of a training set CoronarySet1 into a plurality of patches; inputting the patchings into a double-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topological feature map; constructing a feature fusion module, and fusing the texture feature maps and the topological feature maps of different layers based on an attention mechanism module to obtain fused blood vessel features; shallow layer features of a three-dimensional convolution module and fused blood vessel features are sequentially input into each layer of a segmentation decoder, a three-dimensional blood vessel structure can be extracted, and a blood vessel segmentation result is more continuous and more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and specifically to a medical image segmentation method combining a convolutional neural network with a graph network. Background Art

[0002] In the field of medical image analysis, accurate segmentation of coronary arteries is crucial for diagnosing cardiovascular diseases, assessing the extent of vascular lesions, and formulating treatment plans. Traditional medical image segmentation methods mainly rely on manual feature extraction and classic image processing algorithms, which often have problems such as insufficient accuracy and poor robustness when processing complex and changeable medical images.

[0003] With the rapid development of deep learning technology, convolutional neural networks (CNNs) have achieved remarkable results in the field of medical image segmentation due to their powerful feature extraction capabilities. However, traditional convolutional neural networks mainly focus on the local texture features of images and have limited ability to capture global topological structure information in medical images. Especially in the segmentation of blood vessels with complex topological structures such as coronary arteries, relying solely on local texture features often results in inaccurate visual assessment of coronary artery size and the segmentation results are prone to breakage, making it difficult to achieve ideal segmentation results.

[0004] To this end, we propose a medical image segmentation method that combines convolutional neural networks and graph networks to solve the above problems. Summary of the invention

[0005] The purpose of the present invention is to provide a medical image segmentation method combining a convolutional neural network and a graph network to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: a medical image segmentation method combining a convolutional neural network and a graph network, the method comprising the following steps:

[0007] The coronary artery CT angiography three-dimensional image data collected by a computed tomography (CT) device constitutes a data set CoronarySet, and the data set CoronarySet is randomly divided into a training set CoronarySet1 and a test set CoronarySet2;

[0008] The data of the training set CoronarySet1 is divided into multiple patches; the patches are input into a dual-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topology feature map;

[0009] Construct a feature fusion module, and fuse the texture feature maps and topological feature maps of different layers based on the attention mechanism module to obtain the fused vascular features;

[0010] The shallow features of the 3D convolutional module and the fused vascular features are sequentially input into each layer of the segmentation decoder, and the result of the last layer of the decoder is output to obtain the coronary artery segmentation result consistent with the input 3D image size, so as to train the 3D visual image network medical segmentation network;

[0011] After the 3D visual graph network medical segmentation network is trained, the images to be segmented in the test set CoronarySet2 are divided into patches, and each patch is input into the segmentation network to obtain the 3D segmentation results of the blood vessels of the patches. Finally, the 3D segmentation results of the blood vessels of multiple patches are spliced ​​to obtain the 3D vascular segmentation medical image of a single coronary artery CT angiography image.

[0012] Preferably, the step of collecting coronary artery CT angiography three-dimensional image data through a computed tomography (CT) device to form a data set CoronarySet, and randomly dividing the data set CoronarySet into a training set CoronarySet1 and a test set CoronarySet2 comprises: collecting a plurality of coronary artery three-dimensional angiography images for different individuals, wherein the length and width of the image sizes of different individuals are consistent, but the heights are inconsistent; splitting CoronarySet into a training set CoronarySet1 and a test set CoronarySet2 according to a certain ratio, and having a radiologist be responsible for annotating the blood vessel pixels in the training set CoronarySet1 to obtain the blood vessel labels of the training set CoronarySet1.

[0013] Preferably, the step of inputting the patches into a dual-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topology feature map includes: obtaining a dual-branch parallel encoder, wherein the dual-branch parallel encoder includes a multi-layer three-dimensional convolutional encoder and a multi-layer three-dimensional visual image encoder; inputting the patches into the multi-layer three-dimensional convolutional encoder to obtain a multi-layer coronary artery texture feature map with a smaller and smaller feature map size and an increasing number of channels; inputting the patches into the multi-layer three-dimensional visual image encoder to obtain a multi-layer coronary artery topology feature map with a smaller and smaller feature map size and an increasing number of channels.

[0014] Preferably, the step of fusing texture feature maps and topological feature maps of different layers based on the attention mechanism module to obtain fused blood vessel features comprises:

[0015] Coronary artery texture feature maps and coronary artery topological feature maps with the same length and width of feature maps and different number of channels in the multi-layer 3D convolutional encoder and the multi-layer 3D visual image encoder are respectively selected, and the two types of features corresponding to each layer are fused through the self-attention module to obtain fused vascular features, where the fused vascular features are feature maps with unchanged length and width, and the number of channels is the sum of the number of channels of the corresponding texture feature map and topological feature map.

[0016] Preferably, the shallow features of the three-dimensional convolution module and the fused vascular features are sequentially input into each layer of the segmentation decoder, and the result of the last layer of the decoder is output to obtain a coronary artery segmentation result consistent with the input three-dimensional image size, so as to train the three-dimensional visual image network medical segmentation network. The steps include: constructing a three-dimensional deconvolution decoder, wherein the three-dimensional deconvolution decoder includes a shallow deconvolution layer and a deep deconvolution layer; extracting shallow features in the three-dimensional image data based on the three-dimensional deconvolution decoder, connecting the shallow features of the three-dimensional convolution module to the deep layer of the decoder across layers, and connecting the fused vascular features to the shallow layer of the decoder across layers; based on the three-dimensional deconvolution decoder, Perform deconvolution operations layer by layer to restore the size of the feature map to the same size as the input patches, and obtain the segmented result from the last layer output of the decoder, where the segmented result is to restore the size of the feature map to the same size as the input patches to obtain the 3D vascular segmentation result of the patches; compare the segmentation result with the corresponding vascular annotation in CoronarySet1, and calculate the loss function; adjust the training weight value and random parameter value of the network through the error back propagation algorithm; repeat the training process until the predetermined number of training rounds is reached or the loss function converges to obtain the 3D visual image network medical segmentation network;

[0017] Preferably, the three-dimensional deconvolution decoder restores the size of the feature map to the same size as the input patches through layer-by-layer deconvolution operations, and the step of obtaining the segmented result from the last layer output of the decoder includes: the input of the first layer of shallow deconvolution includes two parts: the output of the feature fusion module of the last layer and the output of the feature fusion module of the previous layer; the two feature maps are spliced ​​together in the channel dimension, and the channel dimension reduction is performed through the convolution operation, and the size of the feature map is expanded by the deconvolution operation; the inputs of the remaining shallow deconvolution layers are the output of the previous deconvolution layer and the feature map output of the feature fusion layer corresponding to the length and width dimensions, and the channel dimension reduction is performed through the convolution operation, and the size of the feature map is expanded by the deconvolution operation to obtain the three-dimensional blood vessel segmentation result of the patches.

[0018] Preferably, after the "input to the first layer of shallow deconvolution includes two parts", the input to the deep deconvolution layer is also included, which is composed of the deconvolution feature map output by the previous layer and the feature map corresponding to the length and width of the feature map in the three-dimensional convolution encoder. The two feature maps are spliced ​​together in the channel dimension, and the channel dimension reduction is performed through the convolution operation. The feature map size is expanded by the deconvolution operation to obtain the three-dimensional blood vessel segmentation result of the patches.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] 1. The present invention combines a convolutional neural network with a graph network to obtain a segmentation network, which can realize the extraction of vascular pixels in the three-dimensional images of coronary artery computed tomography blood vessels, solve the problem that the general medical segmentation model is difficult to extract the three-dimensional vascular structure, and make the vascular segmentation results more continuous and accurate.

[0021] 2. In addition, the three-dimensional visual graph network encoder module proposed in the present invention can automatically complete graph node sampling and graph boundary establishment during network training, and complete the aggregation and transmission of vascular topology information in the established vascular graph network, which can provide more detailed topological information for vascular segmentation, promote the three-dimensional segmentation model's ability to segment continuous tree structures in medical images, and extract more complete vascular structure features.

[0022] 3. By segmenting the three-dimensional images of coronary artery computed tomography, the three-dimensional morphology of the blood vessels can be obtained, providing accurate and continuous visual segmentation results for evaluating the morphology and degree of blockage of vascular tissue. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] Example

[0027] See also Figure 1 The present invention provides a technical solution for a medical image segmentation method combining a convolutional neural network with a graph network: a medical image segmentation method combining a convolutional neural network with a graph network, comprising the following steps:

[0028] S1: The coronary artery CT angiography three-dimensional image data collected by a computed tomography (CT) device constitutes a data set CoronarySet, and the data set CoronarySet is randomly divided into a training set CoronarySet1 and a test set CoronarySet2;

[0029] The steps of collecting coronary artery CT angiography three-dimensional image data through a computed tomography (CT) device to form a data set CoronarySet, and randomly dividing the data set CoronarySet into a training set CoronarySet1 and a test set CoronarySet2 include: collecting a plurality of coronary artery three-dimensional angiography images for different individuals, wherein the length and width of the image sizes of different individuals are consistent, but the heights are inconsistent; splitting CoronarySet into a training set CoronarySet1 and a test set CoronarySet2 according to a certain ratio, and having a radiologist be responsible for annotating the blood vessel pixels in the training set CoronarySet1 to obtain the blood vessel labels of the training set CoronarySet1.

[0030] Specifically, a computed tomography (CT) device is used to perform coronary angiography scans on multiple different individuals. Ensure that the acquired coronary artery three-dimensional angiography images are consistent in length and width, but the height may vary due to individual differences. The acquired three-dimensional image data is preprocessed, such as denoising, correction, etc., to improve the image quality. Keep the original size of the image data unchanged, that is, the length and width remain consistent, and the height varies according to individual differences. The preprocessed coronary artery three-dimensional angiography image data is integrated into a data set named CoronarySet. The data set is split into a training set and a test set; CoronarySet is split into a training set CoronarySet1 and a test set CoronarySet2 according to a certain ratio (such as 80% for training and 20% for testing). Ensure that the split training set and test set are balanced in data distribution, that is, contain various types of coronary artery images. A radiologist annotates the vascular pixels of each image in the training set CoronarySet1. The annotation process requires accurate identification and marking of the pixel positions of the coronary arteries to form vascular labels. The blood vessel labels will be used as the ground truth for training the 3D visual image network medical segmentation network. A high-quality, accurately labeled 3D coronary artery image dataset can be constructed to provide reliable data support for training the 3D visual image network medical segmentation network.

[0031] S2: Split the data of the training set CoronarySet1 into multiple patches; input the patches into a dual-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topology feature map;

[0032] The steps of inputting the patches into the dual-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topology feature map respectively include: obtaining a dual-branch parallel encoder, wherein the dual-branch parallel encoder includes a multi-layer three-dimensional convolution encoder and a multi-layer three-dimensional visual image encoder; inputting the patches into the multi-layer three-dimensional convolution encoder to obtain a multi-layer coronary artery texture feature map with a smaller and smaller feature map size and a larger and larger number of channels; inputting the patches into the multi-layer three-dimensional visual image encoder to obtain a multi-layer coronary artery topology feature map with a smaller and smaller feature map size and a larger and larger number of channels;

[0033] S3: Construct a feature fusion module to fuse the texture feature maps and topological feature maps of different layers based on the attention mechanism module to obtain fused vascular features;

[0034] The step of fusing texture feature maps and topological feature maps of different layers based on the attention mechanism module to obtain fused blood vessel features includes:

[0035] Coronary artery texture feature maps and coronary artery topological feature maps with the same feature map length and width but different number of channels in the multi-layer 3D convolutional encoder and the multi-layer 3D visual map encoder are respectively selected. The two types of features corresponding to each layer are fused through the self-attention module to obtain fused vascular features, where the fused vascular features are feature maps with unchanged length and width, and the number of channels is the sum of the number of channels of the corresponding texture feature map and topological feature map;

[0036] Specifically, from the multi-layer 3D convolution encoder and the multi-layer 3D visual map encoder, the coronary artery texture feature map and the coronary artery topological feature map with the same length and width of the feature map but different number of channels are respectively selected. These feature maps should come from the corresponding layers so that they can be fused in the subsequent steps. For the texture feature map and the topological feature map selected in each layer, the self-attention module is applied to perform feature fusion. The self-attention module can capture the long-distance dependencies within the feature map, enhance important features and suppress irrelevant features. After being processed by the self-attention module, the texture feature map and the topological feature map are concatenated in the channel dimension. This concatenation operation keeps the length and width of the feature map unchanged, but the number of channels increases to the sum of the two. The concatenated feature map is the fused vascular feature. For each layer of the selected feature map, the above-mentioned self-attention module application and feature fusion operation are performed to obtain the fused vascular feature of the corresponding layer. The fused feature will be used in the subsequent 3D deconvolution decoder or other network components to generate the final segmentation result. The splicing operation of feature maps in the channel dimension requires that the length and width of the input feature maps must be the same, but the number of channels can be different. The number of channels of the spliced ​​feature map is the sum of the two, which helps to combine feature information from different sources (such as texture and topology). Texture feature maps and topology feature maps of different layers can be effectively fused to obtain fused vascular features containing rich information, providing strong support for subsequent segmentation tasks;

[0037] S4: Input the shallow features of the 3D convolution module and the fused vascular features into each layer of the segmentation decoder in turn, and output the result of the last layer of the decoder to obtain the coronary artery segmentation result consistent with the input 3D image size, so as to train the 3D visual image network medical segmentation network;

[0038] The shallow features of the 3D convolution module and the fused vascular features are sequentially input into each layer of the segmentation decoder, and the result of the last layer of the decoder is output to obtain the coronary artery segmentation result consistent with the input 3D image size, so as to train the 3D visual graph network medical segmentation network. The steps include: constructing a 3D deconvolution decoder, wherein the 3D deconvolution decoder includes a shallow deconvolution layer and a deep deconvolution layer; extracting shallow features from the 3D image data based on the 3D deconvolution decoder, connecting the shallow features of the 3D convolution module to the deep layer of the decoder across layers, and connecting the fused vascular features to the shallow layer of the decoder across layers; based on the 3D deconvolution decoder, by layer by layer Deconvolution operation restores the size of the feature map to the same size as the input patches, and the segmented result is obtained from the last layer output of the decoder, where the segmented result is to restore the size of the feature map to the same size as the input patches to obtain the 3D vascular segmentation result of the patches; compare the segmentation result with the corresponding vascular annotation in CoronarySet1, and calculate the loss function; adjust the training weight value and random parameter value of the network through the error back propagation algorithm; repeat the training process until the predetermined number of training rounds is reached or the loss function converges to obtain the 3D visual image network medical segmentation network;

[0039] Specifically, texture features and topological features are input into the attention mechanism module to be fused into high-dimensional features; shallow features of the three-dimensional convolution module are connected to the deep layer of the decoder across layers; the fused high-dimensional features are connected to the shallow layer of the decoder across layers; three-dimensional image data, such as coronary artery images, are collected and preprocessed to form a training set (such as CoronarySet1), and corresponding vascular annotations are prepared. The three-dimensional convolution module is used to extract shallow features from the input three-dimensional image. At the same time, texture features and topological features are obtained from other sources (such as preprocessing steps or additional sensors). Texture features and topological features are input into the attention mechanism module and fused into high-dimensional features through the attention mechanism. This fusion method can emphasize important features and suppress irrelevant features. On the other hand, the shallow features of the three-dimensional convolution module are fused with the vascular features obtained by fusion in some way (such as weighted summation or splicing). A three-dimensional deconvolution decoder is constructed, which includes a shallow deconvolution layer and a deep deconvolution layer. Cross-layer connection of the shallow features of the three-dimensional convolution module to the deep layer of the decoder helps to retain more detailed information. At the same time, connecting the fused high-dimensional features to the shallow layers of the decoder across layers can introduce more global information and help the decoder to better restore the size and details of the feature map. Based on the 3D deconvolution decoder, the size of the feature map is restored to the same size as the input 3D image (or patches) through layer-by-layer deconvolution operations. The segmented result is obtained from the output of the last layer of the decoder, which should be consistent with the size of the input 3D image and show the segmentation of the coronary artery. The segmentation result is compared with the corresponding vessel annotation in CoronarySet1, and the loss function (such as cross entropy loss, D ice coefficient loss, etc.) is calculated. The training weight value and random parameter value of the network are adjusted by the error back propagation algorithm to minimize the loss function. The training process is repeated until the predetermined number of training rounds is reached or the loss function converges. The performance of the model is evaluated using the validation set, such as calculating indicators such as accuracy, recall rate, and F1 score. According to the performance evaluation results of the validation set, the model is further optimized, such as adjusting the network structure, adding regularization terms, and using data enhancement techniques. The optimized model is deployed in practical applications for segmentation and diagnosis of coronary arteries. It can utilize the attention mechanism for feature fusion while retaining the ability of the 3D deconvolution decoder to restore the feature map size. At the same time, more global and detailed information is introduced through cross-layer connections, which helps to improve the accuracy and robustness of the segmentation results.

[0040] The training weight value is the mapping relationship between input and output in training, which can be used to calculate the output value of each neuron. During the training process, the training weight value is the parameter that the model needs to optimize. When the three-dimensional visual image network medical segmentation network is used for training in the present invention, the parameter optimization can be directly performed based on the training weight value, which will greatly reduce the training loss and improve the training efficiency.

[0041] The random parameter value is a value that is randomly added as a non-fixed value for iteration and learning in order to increase the accidental generalization during training.

[0042] The three-dimensional deconvolution decoder is based on the step of restoring the size of the feature map to the same size as the input patches through layer-by-layer deconvolution operations, and obtaining the segmented result from the last layer output of the decoder includes: the input of the first layer of the shallow deconvolution includes two parts: the output of the feature fusion module of the last layer and the output of the feature fusion module of the previous layer; the two feature maps are spliced ​​together in the channel dimension, and the channel dimension reduction is performed through the convolution operation, and the size of the feature map is expanded by the deconvolution operation; the inputs of the remaining shallow deconvolution layers are the output of the previous deconvolution layer and the feature map output of the feature fusion layer corresponding to the length and width dimensions, and the channel dimension reduction is performed through the convolution operation, and the size of the feature map is expanded by the deconvolution operation to obtain the three-dimensional blood vessel segmentation result of the patches; the input of the deep deconvolution layer is composed of the deconvolution feature map output of the previous layer and the feature map corresponding to the length and width dimensions of the feature map in the three-dimensional convolution encoder, the two feature maps are spliced ​​together in the channel dimension, and the channel dimension reduction is performed through the convolution operation, and the size of the feature map is expanded by the deconvolution operation to obtain the three-dimensional blood vessel segmentation result of the patches.

[0043] Specifically, shallow deconvolution layer processing: The first shallow deconvolution layer processing: Input: It includes two parts, one is the output of the feature fusion module of the last layer (which may contain feature fusion results from different sources or different levels), and the other is the output of the feature fusion module of the previous layer (which may be a layer in the encoder). The two feature maps are spliced ​​in the channel dimension to form a new feature map. Then, the new feature map is convolved to reduce the channel dimension (that is, feature dimension reduction). Next, the deconvolution operation is used to expand the size of the feature map to make it close to the size of the input patches. Subsequent shallow deconvolution layer processing: Input: It includes the output of the previous deconvolution layer and the feature map output of the corresponding length and width in the feature fusion layer (this feature map may come from a layer in the encoder and has been appropriately resized to match the input size of the current layer). Similar to the first shallow deconvolution layer processing, the two feature maps are spliced ​​in the channel dimension, the convolution operation is performed to reduce the channel dimension, and then the size of the feature map is expanded by the deconvolution operation.

[0044] Deep deconvolution layer processing, deep deconvolution layer input: Input: It includes two parts, one is the feature map output by the previous deconvolution layer, and the other is the feature map corresponding to the length and width of the feature map in the 3D convolution encoder (this feature map may be appropriately resized to match the input size of the current layer). These two feature maps are concatenated in the channel dimension to form a new feature map. The concatenated feature map is convolved to reduce the channel dimension, and then the size of the feature map is further enlarged by deconvolution to make it closer to the size of the input patches. This process is repeated until the last layer of the deep deconvolution layer is reached, at which time the output feature map size should be the same as the size of the input patches. The feature map output from the last layer of the decoder (that is, the last layer of the deep deconvolution layer) is the 3D vascular segmentation result of the patches. This result should be exactly the same size as the input 3D image (or patches) and show the segmentation of the coronary arteries. In the entire decoding process, the convolution operation is mainly used for feature dimensionality reduction to reduce the amount of calculation and avoid overfitting. The deconvolution operation is used to gradually expand the size of the feature map to restore it to the same size as the input patches. The concatenation operation of the feature map in the channel dimension helps to combine feature information from different levels or different sources to improve the accuracy of segmentation.

[0045] S5: After the 3D visual graph network medical segmentation network is trained, the images to be segmented in the test set CoronarySet2 are divided into patches, and each patch is input into the segmentation network to obtain the 3D vascular segmentation results of the patches. Finally, the 3D vascular segmentation results of multiple patches are spliced ​​to obtain a 3D vascular segmentation medical image of a single coronary artery CT angiography image.

[0046] The present invention combines a convolutional neural network with a graph network to obtain a segmentation network, which can realize the extraction of vascular pixels in the three-dimensional images of coronary artery computed tomography blood vessels, solve the problem that the general medical segmentation model is difficult to extract the three-dimensional vascular structure, and make the vascular segmentation results more continuous and accurate.

[0047] In addition, the three-dimensional visual graph network encoder module proposed in the present invention can automatically complete graph node sampling and graph boundary establishment during network training, and complete the aggregation and transmission of vascular topology information in the established vascular graph network, which can provide more detailed topological information for vascular segmentation, promote the three-dimensional segmentation model's ability to segment continuous tree structures in medical images, and extract more complete vascular structure features.

[0048] By segmenting the three-dimensional images of coronary artery computed tomography, the three-dimensional morphology of the blood vessels can be obtained, providing accurate and continuous visual segmentation results for evaluating the morphology and degree of blockage of these vascular tissues.

[0049] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0050] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A medical image segmentation method combining convolutional neural network and graph network, characterized in that: The following steps are involved: The coronary artery CT angiography three-dimensional image data collected by a computed tomography (CT) device constitutes a data set CoronarySet, and the data set CoronarySet is randomly divided into a training set CoronarySet1 and a test set CoronarySet2; The data of the training set CoronarySet1 is divided into multiple patches; the patches are input into a dual-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topology feature map; Construct a feature fusion module, and fuse the texture feature maps and topological feature maps of different layers based on the attention mechanism module to obtain the fused vascular features; The shallow features of the 3D convolutional module and the fused vascular features are sequentially input into each layer of the segmentation decoder, and the result of the last layer of the decoder is output to obtain the coronary artery segmentation result consistent with the input 3D image size, so as to train the 3D visual image network medical segmentation network; After the 3D visual graph network medical segmentation network is trained, the images to be segmented in the test set CoronarySet2 are divided into patches, and each patch is input into the segmentation network to obtain the 3D segmentation results of the blood vessels of the patches. Finally, the 3D segmentation results of the blood vessels of multiple patches are spliced ​​to obtain the 3D vascular segmentation medical image of a single coronary artery CT angiography image.

2. The medical image segmentation method combining a convolutional neural network and a graph network according to claim 1, characterized in that: The steps of collecting coronary artery CT angiography three-dimensional image data through a computed tomography (CT) device to form a data set CoronarySet, and randomly dividing the data set CoronarySet into a training set CoronarySet1 and a test set CoronarySet2 include: collecting a plurality of coronary artery three-dimensional angiography images for different individuals, wherein the length and width of the image sizes of different individuals are consistent, but the heights are inconsistent; splitting CoronarySet into a training set CoronarySet1 and a test set CoronarySet2 according to a certain ratio, and annotating the blood vessel pixels in the training set CoronarySet1 to obtain the blood vessel labels of the training set CoronarySet1.

3. The medical image segmentation method combining a convolutional neural network and a graph network according to claim 1, characterized in that: The steps of inputting the patches into the dual-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topology feature map include: obtaining a dual-branch parallel encoder, wherein the dual-branch parallel encoder includes a multi-layer three-dimensional convolution encoder and a multi-layer three-dimensional visual image encoder; inputting the patches into the multi-layer three-dimensional convolution encoder to obtain a multi-layer coronary artery texture feature map with a smaller and smaller feature map size and an increasing number of channels; inputting the patches into the multi-layer three-dimensional visual image encoder to obtain a multi-layer coronary artery topology feature map with a smaller and smaller feature map size and an increasing number of channels.

4. The medical image segmentation method combining a convolutional neural network and a graph network according to claim 1, characterized in that: The step of fusing texture feature maps and topological feature maps of different layers based on the attention mechanism module to obtain fused blood vessel features includes: Coronary artery texture feature maps and coronary artery topological feature maps with the same length and width of feature maps and different number of channels in the multi-layer 3D convolutional encoder and the multi-layer 3D visual image encoder are respectively selected, and the two types of features corresponding to each layer are fused through the self-attention module to obtain fused vascular features, where the fused vascular features are feature maps with unchanged length and width, and the number of channels is the sum of the number of channels of the corresponding texture feature map and topological feature map.

5. The medical image segmentation method combining a convolutional neural network and a graph network according to claim 1, characterized in that: The shallow features of the 3D convolution module and the fused vascular features are sequentially input into each layer of the segmentation decoder, and the result of the last layer of the decoder is output to obtain the coronary artery segmentation result consistent with the input 3D image size, so as to train the 3D visual graph network medical segmentation network. The steps include: constructing a 3D deconvolution decoder, wherein the 3D deconvolution decoder includes a shallow deconvolution layer and a deep deconvolution layer; extracting shallow features from the 3D image data based on the 3D deconvolution decoder, connecting the shallow features of the 3D convolution module to the deep layer of the decoder across layers, and connecting the fused vascular features to the shallow layer of the decoder across layers; based on the 3D deconvolution decoder, by layer by layer The deconvolution operation restores the size of the feature map to the same size as the input patches, and the segmented result is obtained from the last layer output of the decoder, where the segmented result is to restore the feature map size to the same size as the input patches to obtain the three-dimensional vascular segmentation result of the patches; the segmentation result is compared with the corresponding vascular annotation in CoronarySet1, and the loss function is calculated; the training weight value and random parameter value of the network are adjusted through the error back propagation algorithm; the training process is repeated until the predetermined number of training rounds is reached or the loss function converges to obtain a three-dimensional visual image network medical segmentation network.

6. The medical image segmentation method combining a convolutional neural network and a graph network according to claim 5, characterized in that: The three-dimensional deconvolution decoder is based on the step of restoring the size of the feature map to the same size as the input patches through layer-by-layer deconvolution operations, and obtaining the segmented result from the last layer output of the decoder. The input of the first layer of shallow deconvolution includes two parts: the output of the last layer feature fusion module and the output of the feature fusion module of the previous layer; the two feature maps are spliced ​​together in the channel dimension, and the channel dimension is reduced by the convolution operation, and the size of the feature map is expanded by the deconvolution operation; the inputs of the remaining shallow deconvolution layers are the output of the previous layer deconvolution and the feature map output of the feature fusion layer corresponding to the length and width, and the channel dimension is reduced by the convolution operation, and the size of the feature map is expanded by the deconvolution operation to obtain the three-dimensional blood vessel segmentation result of the patches.

7. The medical image segmentation method combining a convolutional neural network and a graph network according to claim 6, characterized in that: After "the input to the first layer of shallow deconvolution includes two parts", the input to the deep deconvolution layer is also composed of the deconvolution feature map output by the previous layer and the feature map corresponding to the length and width of the feature map in the three-dimensional convolution encoder. The two feature maps are spliced ​​together in the channel dimension, and the channel dimension is reduced by convolution operation. The feature map size is expanded by deconvolution operation to obtain the three-dimensional blood vessel segmentation results of patches.

Citation Information

Patent Citations

  • Sub-mesenteric artery blood vessel reconstruction method based on MIP sequence

    CN114897780A

  • Cross-modal double-branch complementary fusion image segmentation method and device

    CN115482241A

  • Double-branch hyperspectral image classification method based on graph convolutional neural network and attention mechanism

    CN118135306A

  • Double-coding cross fusion OCTA image segmentation method based on attention mechanism

    CN118261923A

Cited By

  • Medical image segmentation method and system based on parallel coding, variation fusion and uncertainty optimization of visual basic model

    CN120707578A

  • CT image intelligent analysis system for pneumonia auxiliary screening

    CN120953426A

  • A CT image intelligent analysis system for pneumonia auxiliary screening

    CN120953426B