Universal segmentation method applied to three-dimensional medical image with tubular structure
Through refined preprocessing, multi-scale domain prior fusion modules and combined loss function, a general segmentation method is constructed, which solves the problems of poor segmentation effect of three-dimensional medical images in tubular structures and waste of resources in the prior art, and achieves efficient and accurate segmentation effect.
Patent Information
- Application Number
- CN202510067191.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
When processing three-dimensional medical images of tubular structures, it is difficult to effectively consider their complex shapes and topological connectivity, resulting in poor segmentation effects, and serious waste of resources from mainstream methods and lack of versatility and scalability.
Through refined preprocessing processes, innovative multi-scale domain prior fusion modules and combined loss function design, and efficient training and testing strategies, a general segmentation method applied to three-dimensional medical images of tubular structures is constructed.
It significantly improves the performance and versatility of three-dimensional medical imaging segmentation of tubular structures, can efficiently and accurately process tubular structure data in many different parts and modes, and overcomes the problems of resource waste and lack of universality.
Smart Images

Figure CN119992089A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image segmentation, and more specifically to a general segmentation method for three-dimensional medical images of tubular structures. Background Art
[0002] With the rapid development of medical imaging technology, accurate imaging and analysis of tubular structures in the human body (such as blood vessels, trachea, esophagus, etc.) plays a vital role in early diagnosis of diseases, treatment planning and efficacy evaluation. However, the imaging characteristics of tubular structures in various medical imaging modalities, such as low contrast, complex and varied shapes and topological structures, make their segmentation tasks particularly complex and challenging.
[0003] In the existing technology, although deep learning-based medical image segmentation methods such as UNet, UNet++ and nnUNet have achieved remarkable results in the field of medical image processing, they still have significant defects in processing tubular structure segmentation. These methods usually fail to fully consider the complex shape and topological connectivity of tubular structures in three-dimensional space, resulting in poor segmentation results when processing tubular structures with complex shapes and many branches.
[0004] More importantly, the current mainstream method often takes the approach of designing an independent network model for each tubular target. Although this approach improves the segmentation accuracy of specific tubular structures to a certain extent, it causes a huge waste of resources. At the same time, due to the isolation of data between different tubular structures, it is difficult to achieve knowledge sharing and migration, which limits the versatility and scalability of the model. When faced with tubular structure data of various different parts and modalities, this method is obviously inadequate.
[0005] In order to overcome these limitations, researchers began to explore general medical image segmentation methods based on basic models. However, these methods still face many challenges when processing tubular structure data with diverse modalities and large scale variations. For example, although methods such as MultiTalent and UniSeg improve model performance by combining general features and task-specific features, when processing tubular structure data, the feature prompts they construct may lead to feature redundancy or even confusion due to category overlap or differences in task attributes. Although the Hermes method introduces the concepts of task prior and modality prior, in the tubular structure segmentation scenario, these methods have not been fully optimized for the complex characteristics of tubular structures, resulting in unsatisfactory segmentation results.
[0006] In addition, the existing technology for the preprocessing process of tubular structure 3D medical imaging data is often lacking in systematization and pertinence. Differences in resolution, target size, and data intensity values between different data sets have not been fully considered, which makes it difficult for the model to accurately capture the characteristic information of tubular targets during the learning process, thus affecting the segmentation accuracy.
[0007] Therefore, how to design a universal segmentation method for three-dimensional medical images of tubular structures to achieve efficient and accurate segmentation of three-dimensional medical image data of tubular structures in various parts and modalities is an urgent problem to be solved by technical personnel in this field. Summary of the invention
[0008] In view of this, the present invention provides a universal segmentation method for three-dimensional medical images of tubular structures. Through a refined preprocessing process, an innovative multi-scale domain prior fusion module and combined loss function design, as well as efficient training and testing strategies, the performance and versatility of three-dimensional medical image segmentation of tubular structures are significantly improved.
[0009] In order to achieve the above object, the present invention adopts the following technical solution:
[0010] The present invention provides a general segmentation method for three-dimensional medical images of tubular structures, comprising the following steps:
[0011] S1. Acquire multiple three-dimensional medical imaging datasets of tubular structures in different parts and modalities;
[0012] S2, preprocessing the data set to obtain a preprocessed data set;
[0013] S3, constructing a universal segmentation model for three-dimensional medical images, and optimizing the model using the preprocessed data set and the learnable data domain prior to obtain an optimized universal segmentation model for three-dimensional medical images;
[0014] S4. Input the tubular structure three-dimensional medical image to be processed and the corresponding data domain prior into the optimized three-dimensional medical image general segmentation model to obtain the corresponding medical image segmentation result.
[0015] Further, the S2 includes:
[0016] S21, cropping each data in each data set to remove the background portion where the voxel value at the edge is 0; each data includes a three-dimensional image and its corresponding annotated image;
[0017] S22, based on the voxel value of each data set after clipping, define the attributes of each data set; the attributes include: the median of the voxel value in the x, y, and z directions, the mean of the voxel value, the maximum voxel value, the minimum voxel value, the voxel value variance, the 99.5% quantile voxel value, and the 0.05% quantile voxel value;
[0018] S23, unifying the voxel values of each data set;
[0019] S24, normalizing each data set after the voxel values are unified;
[0020] S25. Set the sampling probability of the normalized data set to obtain a preprocessed data set.
[0021] Further, the S23 includes: taking the median of the voxel values in the three directions of x, y and z in each data set as the target voxel value of the data set, and resampling the voxel values of all data in each data set to the target voxel value;
[0022] Among them, the size of the i-th data in the t-th data set after resampling is It is expressed as:
[0023]
[0024] represents the original size of the i-th data in the t-th data set, represents the voxel value of the i-th data set in the t-th data set, Represents the target voxel value of the tth dataset.
[0025] Furthermore, the S24 includes: normalizing the intensity values of all data in each data set:
[0026]
[0027] in, represents the intensity value of the i-th data in the t-th data set, represents the mean intensity value in the tth data set, represents the variance of the intensity values in the tth data set.
[0028] Furthermore, before normalization, it also includes:
[0029] Based on the intensity value mean and intensity value variance in each data set, define the intensity value range of the data;
[0030] Identify outliers by intensity value range;
[0031] Outliers are replaced by interpolation, which estimates outliers by taking a weighted average of neighboring non-outliers.
[0032] Furthermore, in S25, the sampling probability prob t It is expressed as:
[0033]
[0034] Among them, n t represents the number of data in the tth data set, and T represents the number of data sets.
[0035] Furthermore, in S3, the backbone network of the three-dimensional medical image universal segmentation model is a three-dimensional U-shaped network, including an encoder and a decoder;
[0036] The encoder includes a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block and a sixth convolution block; wherein each convolution block includes two first convolution layers, and each first convolution layer includes two consecutive three-dimensional convolution layers, an instance normalization layer and a LeakyReLU activation function layer;
[0037] The decoder includes a seventh convolution block, an eighth convolution block, a ninth convolution block, a tenth convolution block and an eleventh convolution block; wherein the seventh convolution block includes three second convolution layers, and the other convolution blocks include two second convolution layers; each second convolution layer includes a three-dimensional transposed convolution layer and two consecutive three-dimensional convolution layers, an instance normalization layer and a LeakyReLU activation function layer.
[0038] Furthermore, the third convolution block, the fourth convolution block, the fifth convolution block in the encoder, and the seventh convolution block, the eighth convolution block, and the ninth convolution block in the decoder are respectively embedded with multi-scale domain prior fusion modules;
[0039] The multi-scale domain prior fusion module processes the learnable data domain prior and feature map including:
[0040] DP i =Norm(MLP(Norm(MLP(MSA(Norm(DP i-1 ),Norm(FA i-1 ))))+
[0041] DP i-1 ))+Norm(MLP(MSA(Norm(DP i-1 ),Norm(FA i-1 ))))+DP i-1
[0042] FA i=Norm(MLP(Norm(MLP(MSA(Norm(FA i-1 ),Norm(DP i ))))+FA i-1 ))
[0043] +Norm(MLP(MSA(Norm(FA i-1 ),Norm(DP i ))))+FA i-1
[0044] Among them, DP i , FA i Respectively represent the output learnable data domain prior, feature map, DP i-1 , FA i-1 They represent the learnable data domain prior and feature map of the input respectively, Norm(·) represents layer normalization, MLP(·) represents multi-layer perceptron, and MSA(·) represents multi-head self-attention mechanism.
[0045] Furthermore, in S3, the loss function L of the three-dimensional medical image general segmentation model is a combined loss function:
[0046] L seg =L DICE +L CE +L VAS +L dom
[0047] Among them, L DICE represents the DICE loss function, which is used to measure the similarity between the predicted result and the true label; L CE represents the cross entropy loss function, which is used to evaluate the performance of classification tasks; L VAS represents the topological loss function, which is used to capture the geometric non-smoothness of the vascular structure; L dom Represents the domain prediction loss function, which is used to train the model to identify the dataset domain to which the image belongs.
[0048] Furthermore, in S3, the optimization is performed using the preprocessed data set, including:
[0049] In the training phase, the AdamW optimizer is used to update the network parameters, the cosine annealing strategy is used to dynamically adjust the learning rate, and random scaling, random rotation, random gamma enhancement, and random mirroring are combined for data enhancement.
[0050] During the testing phase, the 3D image is segmented into multiple overlapping image blocks, the segmented image blocks and the corresponding data domain prior vector weights are input into the model, and the prediction results of the 3D image are obtained by merging the prediction results of all image blocks.
[0051] It can be seen from the above technical solution that compared with the prior art, the technical solution of the present invention has the following advantages:
[0052] Beneficial effects:
[0053] 1. By integrating tubular structure datasets from different imaging modalities and different human regions, a unified basic model is constructed, which not only expands the scope of application of the model but also improves the versatility of the model. Its method of integrating multi-modal, multi-regional, and multi-scale medical imaging data into one model can effectively process various types of tubular structures, overcoming the resource waste caused by designing dedicated networks for a single or a few targets in traditional methods.
[0054] 2. In the preprocessing stage, the original data was refined through a series of steps such as cropping, resampling, and normalization. This not only removed the background noise, but also unified the voxel value size of different data sets. The normalization process further improved the consistency of the data and alleviated the imbalance in prediction performance caused by differences between data sets.
[0055] 3. A multi-scale domain prior fusion module is introduced into the backbone network. This module realizes the fusion between feature maps and domain priors through the cross-attention mechanism, enabling the network to better capture the characteristics of different types of tubular structures. This helps the network to be more efficient in understanding and distinguishing different types of tubular targets, thereby improving the accuracy and generalization ability of segmentation.
[0056] 4. By combining DICE loss, cross entropy loss, topological loss and domain prediction classification loss, a comprehensive loss function framework is formed, which not only optimizes the segmentation boundary and topological connectivity, but also enhances the model's ability to distinguish different data domains. It especially considers the geometric characteristics and multi-branch structure of tubular targets, which helps to improve the model's performance in detecting small branches. At the same time, it also promotes the distinction of data domains through auxiliary tasks, thereby improving the segmentation effect of the entire framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0058] Figure 1 A flow chart of a general segmentation method for three-dimensional medical images of tubular structures provided by an embodiment of the present invention;
[0059] Figure 2A schematic diagram of a data set preprocessing process provided by an embodiment of the present invention;
[0060] Figure 3 A structural framework diagram of a universal segmentation model for 3D medical images provided in an embodiment of the present invention;
[0061] Figure 4 A schematic diagram of the data processing process of the multi-scale domain prior fusion module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] Embodiment 1;
[0064] like Figure 1 As shown, this embodiment provides a general segmentation method for tubular structure three-dimensional medical images, comprising the following steps:
[0065] S1. Acquire multiple three-dimensional medical imaging datasets of tubular structures in different parts and modalities;
[0066] S2, preprocessing the data set to obtain a preprocessed data set;
[0067] S3, constructing a universal segmentation model for three-dimensional medical images, and optimizing the model using the preprocessed data set and the learnable data domain prior to obtain an optimized universal segmentation model for three-dimensional medical images;
[0068] S4. Input the tubular structure three-dimensional medical image to be processed and the corresponding data domain prior into the optimized three-dimensional medical image general segmentation model to obtain the corresponding medical image segmentation result.
[0069] This method obtains multiple three-dimensional medical image datasets of tubular structures in different parts and modalities, performs preprocessing steps, and then constructs and optimizes a unified three-dimensional U-shaped three-dimensional medical image universal segmentation model. Finally, the optimized model is used to efficiently and accurately segment the three-dimensional medical images of tubular structures to be processed, which significantly improves the accuracy and versatility of tubular structure three-dimensional medical image segmentation.
[0070] The above steps and related technical features are further described in detail below:
[0071] In this embodiment S1, a plurality of three-dimensional medical image data sets of tubular structures in different parts and different modalities are obtained; and the data sets are divided into a training set and a test set for training and evaluating a basic model.
[0072] Specifically, seven different datasets were collected, including: 1) a CTA dataset containing renal artery vascular annotations; 2) a CT dataset containing pulmonary trachea annotations; 3) a CTA dataset containing pulmonary artery annotations; 4) a CTA dataset containing cerebral basilar artery ring annotations; 5) a TOF-MRA dataset containing cerebral basilar artery ring annotations; 6) a TOF-MRA dataset containing cerebral basilar artery ring annotations; 7) a CTA dataset containing coronary artery annotations.
[0073] It covers multiple important anatomical locations such as the renal artery, pulmonary trachea, and cerebral arterial circle, and uses imaging techniques including CTA (computed tomography angiography), CT (computed tomography), and TOF-MRA (time of flight magnetic resonance angiography). This diverse data source ensures that the model can maintain good performance under various conditions.
[0074] like Figure 2 As shown, in this embodiment S2, the data set is preprocessed to obtain a preprocessed data set; specifically, it includes:
[0075] S21, cropping each data in each data set to remove the background part with a voxel value of 0 at the edge; each data includes a three-dimensional image and its corresponding annotated image; this can reduce unnecessary calculations and enable the model to focus on meaningful information areas;
[0076] S22. Based on the voxel value of each data set after clipping, define the attributes of each data set; the attributes include: the median of the voxel value in the x, y, and z directions, the mean of the voxel value, the maximum voxel value, the minimum voxel value, the variance of the voxel value, the 99.5% quantile voxel value, and the 0.05% quantile voxel value; and record and save in detail in an attribute document to provide a basis for subsequent data standardization and feature selection;
[0077] S23, unify the voxel values of each data set; this is to ensure the consistency of spatial resolution of all data sets so that the model training is not affected by the resolution differences between different data sets;
[0078] The method comprises: taking the median of the voxel values in the three directions of x, y and z in each data set as the target voxel value of the data set, and resampling the voxel values of all data in each data set to the target voxel value;
[0079] Among them, the size of the i-th data in the t-th data set after resampling is It is expressed as:
[0080]
[0081] represents the original size of the i-th data in the t-th data set, represents the voxel value of the i-th data set in the t-th data set, Represents the target voxel value of the tth dataset.
[0082] S24, normalizing each data set after the voxel values are unified;
[0083] Include: Normalize the intensity values of all data in each dataset:
[0084]
[0085] in, represents the intensity value of the i-th data in the t-th data set, represents the mean intensity value in the tth data set, represents the variance of the intensity values in the tth data set.
[0086] Before normalizing the intensity values of all data in each dataset, also include:
[0087] Based on the intensity value mean and intensity value variance in each data set, define the intensity value range of the data;
[0088] Identify outliers by intensity value range;
[0089] Outliers are replaced by interpolation, which estimates outliers by taking a weighted average of neighboring non-outliers.
[0090] Here, outliers are estimated and replaced by using the weighted average of neighboring non-outliers, thereby maintaining the continuity and smoothness of the image. First, outliers are identified; then, weights are assigned to each non-outlier within a selected neighborhood (such as a 3x3 or 5x5 window), and the weights can be based on distance, similarity, or uniform distribution; then, an estimate is obtained by calculating the weighted average of these non-outliers and their respective weights; finally, the outlier is replaced with the estimated value. It can effectively repair local anomalies while maintaining the overall structure and details of the image.
[0091] S25. Set the sampling probability of the normalized data set to obtain a preprocessed data set, which aims to balance the training frequencies of data domains of different sizes and avoid excessive bias of network performance towards data domains with large data sizes.
[0092] Sampling probability prob t It is expressed as:
[0093]
[0094] Among them, n t represents the number of data in the tth data set, and T represents the number of data sets.
[0095] In this preprocessing stage, the quality and consistency of the three-dimensional medical imaging data of tubular structures were effectively improved by cropping to remove background noise, calculating and standardizing voxel value attributes, unifying voxel resolution, normalizing intensity values and repairing outliers, and setting reasonable sampling probabilities, laying a solid foundation for subsequent model training and optimization.
[0096] In this embodiment S3, a universal segmentation model for three-dimensional medical images is constructed, and the model is optimized using the preprocessed data set and the learnable data domain prior to obtain an optimized universal segmentation model for three-dimensional medical images;
[0097] like Figure 3 As shown, a dataset image with a size of 1×96×160×160 is input into the 3D medical image general segmentation model. The backbone network of the 3D medical image general segmentation model is a 3D U-shaped network, including an encoder and a decoder.
[0098] The encoder includes a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block and a sixth convolution block; wherein each convolution block includes two first convolution layers, each first convolution layer includes two consecutive three-dimensional convolution layers, an instance normalization layer and a LeakyReLU activation function layer; the specific information is shown in Table 1 and Table 2 below:
[0099] Table 1
[0100]
[0101] Table 2
[0102]
[0103] The decoder includes a seventh convolution block, an eighth convolution block, a ninth convolution block, a tenth convolution block and an eleventh convolution block; wherein the seventh convolution block includes three second convolution layers, and the other convolution blocks include two second convolution layers; each second convolution layer includes a three-dimensional transposed convolution layer and two consecutive three-dimensional convolution layers, an instance normalization layer and a LeakyReLU activation function layer; the specific information is shown in Table 3 below:
[0104] Table 3
[0105]
[0106]
[0107] The encoder goes deeper and deeper, extracting features through multiple convolution blocks. Each convolution block contains a double 3D convolution layer, supplemented by instance normalization and LeakyReLU activation, effectively capturing deep information of the image. The decoder reconstructs the image in reverse, using 3D transposed convolution layer upsampling, combined with continuous 3D convolution layers, normalization and activation functions, to gradually restore image details, achieving comprehensive integration and fine segmentation of features from low-level to high-level.
[0108] Furthermore, a multi-scale domain prior fusion module is introduced and embedded in different levels of the encoder and decoder to enhance the learning ability and feature expression ability of the model. Specifically, the third, fourth and fifth convolution blocks in the encoder and the seventh, eighth and ninth convolution blocks in the decoder are respectively embedded with the multi-scale domain prior fusion module;
[0109] like Figure 4 As shown, the multi-scale domain prior fusion module processes the learnable data domain prior and feature map including:
[0110] DP i =Norm(MLP(Norm(MLP(MSA(Norm(DP i-1 ),Norm(FA i-1 ))))+
[0111] DP i-1 ))+Norm(MLP(MSA(Norm(DP i-1 ),Norm(FA i-1 ))))+DP i-1
[0112] FA i =Norm(MLP(Norm(MLP(MSA(Norm(FA i-1 ),Norm(DP i ))))+FA i-1 ))
[0113] +Norm(MLP(MSA(Norm(FA i-1 ),Norm(DP i ))))+FA i-1
[0114] Among them, DP i , FA i Respectively represent the output learnable data domain prior, feature map, DP i-1 , FA i-1They represent the learnable data domain prior and feature map of the input respectively, Norm(·) represents layer normalization, MLP(·) represents multi-layer perceptron, and MSA(·) represents multi-head self-attention mechanism.
[0115] The above-mentioned "learnable data domain prior" refers to the prior knowledge related to different data domains (such as imaging modalities, anatomical parts, etc.) that can be optimized and learned during the model training process. This prior knowledge can enhance the adaptability of the model to different data domains, improve model performance, promote multi-task learning, and perform cross-attention mechanism processing with the feature map through the multi-scale domain prior fusion module to achieve the update and optimization of the feature map, thereby improving the segmentation accuracy and generalization ability of the model for tubular structure three-dimensional medical images.
[0116] Furthermore, the above processing is further explained below:
[0117] The multi-scale domain prior fusion module is located in the i-th layer of the network, and the input feature map FA i-1 , input learnable data domain prior DP i-1 ; Through layer normalization Norm and multi-layer perceptron MLP, the feature map FA i-1 Converted into a query tensor q, the learnable data domain prior DP i-1 Convert to key tensor k and value tensor v; Use the cross attention mechanism to multiply the query tensor q and value tensor v in multiple dimensions, and obtain the attention weight matrix through the normalized exponential function Softmax; Multiply the attention weight matrix and key tensor k in multiple dimensions, and obtain the updated feature representation through the multi-layer perceptron MLP and layer normalization Norm, and combine the updated feature representation with the learnable data domain prior as DP i-1 Add and output the learnable data domain prior DP i and feature map FA i .
[0118] In the above formula, the operations from obtaining the query tensor q, the key tensor k, the value tensor v to the last multi-dimensional multiplication are defined as the multi-head self-attention mechanism MSA(·).
[0119] Through the multi-head self-attention mechanism MSA(·), the model can capture more contextual information at different scales, thereby improving the understanding and segmentation accuracy of complex medical images. In addition, through multiple iterations and cross-layer information transfer, the model can more finely adjust its internal representation, making the final segmentation result more accurate and refined.
[0120] Furthermore, the loss function L of the general segmentation model of 3D medical images is a combined loss function:
[0121] L seg =LDICE +L CE +L VAS +L dom
[0122] Among them, L DICE represents the DICE loss function, which is used to measure the similarity between the predicted result and the true label; L CE represents the cross entropy loss function, which is used to evaluate the performance of classification tasks; L VAS represents the topological loss function, which is used to capture the geometric non-smoothness of the vascular structure; L dom Represents the domain prediction loss function, which is used to train the model to identify the dataset domain to which the image belongs.
[0123] Specifically, L VAS The topological characteristics of the tubular structure and the collective imbalance between different branch structures are considered, which can be expressed as:
[0124]
[0125] Among them, Vprec and Vsens represent vascular accuracy and vascular sensitivity;
[0126] The following is a further explanation of the definition process of Vprec and Vsens:
[0127] Let the input be x i , where i represents the voxel position, Y m (x i ) and G m (x i ) represent the network predicted segmentation map and the true segmentation label respectively. The network predicted skeleton map Y is obtained by skeletonizing the segmentation map or segmentation label. s (x i ) and the skeleton label G s (x i ):
[0128] Y s (x i )=Skeletonize(Y m (x i ))
[0129] G s (x i )=Skeletonize(G m (x i ))
[0130] Next, a property function D[·] is defined for the foreground point in the segmentation prediction (label), which represents the minimum distance from the foreground point to the boundary of the tubular structure; two property functions R[·] and N[·] are defined for the foreground point in the skeleton prediction (label). The former represents the minimum distance from the skeleton point to the boundary of the tubular structure, and the latter represents the square of the reciprocal of the minimum distance. Vprec and Vsens are defined as:
[0131]
[0132] In the loss function L, L DICE , L CE , L VAS Used to optimize the segmentation performance of the network framework, L dom It is used as an auxiliary function to promote the network's ability to distinguish data domains, thereby improving the network's prediction effect on vascular targets in different data sets.
[0133] The design of this combined loss function comprehensively considers the geometric complexity and topological connectivity of tubular structures, as well as the differences between different data sets. Through the collaborative optimization of multiple tasks, it not only improves the segmentation accuracy of the model for tubular structures, but also enhances the model's generalization ability for complex scenes and across data sets.
[0134] Further, the preprocessed data set is used for optimization, including:
[0135] In the training phase, the total number of training rounds is set to 1000, and each round contains 450 iterations. The batch size of each iteration is set to 2. The initial learning rate of the network is 5×10 -4 , update the network parameters through the AdamW optimizer, which can automatically adjust the learning rate according to the historical information of the gradient, thereby accelerating convergence and preventing overfitting;
[0136] The cosine annealing strategy is used to dynamically adjust the learning rate, which simulates the physical annealing process and makes the learning rate fluctuate periodically with the changes in training rounds, which helps to find the global optimal solution.
[0137] Data augmentation is performed by combining random scale transformation, random rotation, random gamma enhancement, and random mirroring; it can simulate a variety of image changes that may be encountered in actual use, allowing the model to handle more complex and diverse image inputs.
[0138] Furthermore, during the testing phase, the three-dimensional image is segmented into multiple overlapping image blocks; this not only reduces memory consumption and improves processing speed, but also utilizes the overlapping parts between image blocks to smooth the final prediction results.
[0139] The segmented image blocks and the corresponding data domain prior vector weights are input into the model; before the segmented image blocks are input into the deep learning model, the corresponding data domain prior vector weights need to be prepared. The "data domain prior vector weights" here refer to the statistical or empirical information related to the specific task, which can guide the learning process of the model and make it more focused on the features that are critical to completing the task.
[0140] The prediction result of the three-dimensional image is obtained by merging the prediction results of all image blocks; during the merging process, interpolation and splicing algorithms are used to ensure the continuity and consistency of the prediction results.
[0141] In the optimization stage, the model convergence and generalization ability are effectively promoted through sophisticated training settings and data enhancement, combined with the efficient AdamW optimizer and cosine annealing strategy. During testing, the image block segmentation and prior knowledge fusion strategy are used to ensure the accuracy and smoothness of the prediction results.
[0142] In this embodiment S4, the three-dimensional medical image of the tubular structure to be processed and the corresponding data domain prior are input into the optimized three-dimensional medical image general segmentation model to obtain the corresponding medical image segmentation result.
[0143] In this step, one or more three-dimensional medical images of tubular structures to be processed (such as blood vessels, airways, etc. obtained by CT or MRI scanning) and corresponding data domain prior information are selected.
[0144] The data domain prior information here is the imaging modality, anatomical part and patient characteristics related to the image. This information is used to guide the model on how to apply the learnable data domain prior knowledge learned during the training process to better adapt to the image data under different imaging modalities, anatomical parts and patient characteristics, so as to achieve accurate image segmentation. It can be a set of values or vectors determined based on factors such as image acquisition equipment, basic patient information, and disease type.
[0145] The specific steps include:
[0146] 1) Input preparation;
[0147] Preprocess the 3D medical images to be processed, including but not limited to cropping, voxel value unification, normalization, etc., to ensure consistency with the data processing method during training. Determine the corresponding data domain prior based on the specific situation of the image to be processed. For example, if the model considers the differences between different age groups during training, it is necessary to specify an appropriate age group as the data domain prior for the input image.
[0148] 2) Model reasoning;
[0149] The preprocessed 3D medical image and the selected data domain prior are input into the optimized 3D medical image general segmentation model. The model will output one or more segmentation results based on the input image data and data domain prior after a series of calculations (including but not limited to convolution operations, feature extraction, domain prior fusion, etc.). These results are usually expressed as one or more probability maps of the same size as the input image, and the value at each pixel represents the probability that the location belongs to a specific tubular structure.
[0150] 3) Post-processing of results;
[0151] Threshold processing can be applied to the probability map output by the model to convert the probability value into a binary segmentation mask, thereby clearly identifying the location of the tubular structure. Furthermore, morphological operations (such as dilation and erosion) can be used to improve the quality of the segmentation results, such as reducing noise and filling small holes.
[0152] This embodiment provides a general segmentation method for three-dimensional medical images of tubular structures. By collecting and preprocessing medical image data sets covering various parts and modalities, a general segmentation model based on a three-dimensional U-shaped network is constructed. The model introduces learnable data domain priors during the training process, which significantly improves the generalization ability and segmentation accuracy of the model. Specifically, the preprocessing step ensures the consistency and quality of the data. The model design combines a multi-scale domain prior fusion module and a carefully designed combined loss function, which can effectively cope with the differences between different data sets and achieve efficient and accurate segmentation of complex tubular structures. In addition, by optimizing the training strategy and the image block processing method in the test phase, the robustness of the model and the smoothness of the prediction results are further improved.
[0153] It achieves efficient and accurate segmentation of 3D medical images of tubular structures in multiple different parts and modalities, significantly improving the accuracy and versatility of tubular structure segmentation and providing strong support for medical image analysis.
[0154] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0155] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A general segmentation method for three-dimensional medical images of tubular structures, characterized in that: The following steps are involved: S1. Acquire multiple three-dimensional medical imaging datasets of tubular structures in different parts and modalities; S2, preprocessing the data set to obtain a preprocessed data set; S3, constructing a universal segmentation model for three-dimensional medical images, and optimizing the model using the preprocessed data set and the learnable data domain prior to obtain an optimized universal segmentation model for three-dimensional medical images; S4. Input the tubular structure three-dimensional medical image to be processed and the corresponding data domain prior into the optimized three-dimensional medical image general segmentation model to obtain the corresponding medical image segmentation result.
2. A general segmentation method for three-dimensional medical images of tubular structures according to claim 1, characterized in that: The S2 includes: S21, cropping each data in each data set to remove the background portion where the voxel value at the edge is 0; each data includes a three-dimensional image and its corresponding annotated image; S22, based on the voxel value of each data set after clipping, define the attributes of each data set; the attributes include: the median of the voxel value in the x, y, and z directions, the mean of the voxel value, the maximum voxel value, the minimum voxel value, the voxel value variance, the 99.5% quantile voxel value, and the 0.05% quantile voxel value; S23, unifying the voxel values of each data set; S24, normalizing each data set after the voxel values are unified; S25. Set the sampling probability of the normalized data set to obtain a preprocessed data set.
3. A general segmentation method for three-dimensional medical images of tubular structures according to claim 2, characterized in that: The S23 includes: taking the median of the voxel values in the three directions of x, y and z in each data set as the target voxel value of the data set, and resampling the voxel values of all data in each data set to the target voxel value; Among them, the size of the i-th data in the t-th data set after resampling is It is expressed as: represents the original size of the i-th data in the t-th data set, represents the voxel value of the i-th data in the t-th data set, Represents the target voxel value of the tth dataset.
4. The general segmentation method for three-dimensional medical images of tubular structures according to claim 2, characterized in that: The S24 comprises: normalizing the intensity values of all data in each data set: in, represents the intensity value of the i-th data in the t-th data set, represents the mean intensity value in the tth data set, represents the variance of the intensity values in the tth data set.
5. The general segmentation method for tubular structure three-dimensional medical images according to claim 4, characterized in that: Before normalization, it also includes: Based on the intensity value mean and intensity value variance in each data set, define the intensity value range of the data; Identify outliers through the range of intensity values; Outliers are replaced by using an interpolation method; the interpolation method estimates outliers by taking a weighted average of neighboring non-outlier values.
6. The general segmentation method for tubular structure three-dimensional medical images according to claim 2, characterized in that: In S25, the sampling probability prob t It is expressed as: Among them, n t represents the number of data in the tth data set, and T represents the number of data sets.
7. The general segmentation method for tubular structure three-dimensional medical images according to claim 1, characterized in that: In S3, the backbone network of the three-dimensional medical image universal segmentation model is a three-dimensional U-shaped network, including an encoder and a decoder; The encoder includes a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block and a sixth convolution block; wherein each convolution block includes two first convolution layers, and each first convolution layer includes two consecutive three-dimensional convolution layers, an instance normalization layer and a LeakyReLU activation function layer; The decoder includes a seventh convolution block, an eighth convolution block, a ninth convolution block, a tenth convolution block and an eleventh convolution block; wherein the seventh convolution block includes three second convolution layers, and the other convolution blocks include two second convolution layers; each second convolution layer includes a three-dimensional transposed convolution layer and two consecutive three-dimensional convolution layers, an instance normalization layer and a LeakyReLU activation function layer.
8. The general segmentation method for tubular structure three-dimensional medical images according to claim 7, characterized in that: The third convolution block, the fourth convolution block, the fifth convolution block in the encoder, and the seventh convolution block, the eighth convolution block, and the ninth convolution block in the decoder are respectively embedded with multi-scale domain prior fusion modules; The multi-scale domain prior fusion module processes the learnable data domain prior and feature map including: DP i =Norm(MLP(Norm(MLP(MSA(Norm(DP i-1 ),Norm(FA i-1 ))))+DP i-1 ))+Norm(MLP(MSA(Norm(DP i-1 ),Norm(FA i-1 ))))+DP i-1 FA i =Norm(MLP(Norm(MLP(MSA(Norm(FA i-1 ),Norm(DP i ))))+FA i-1 ))+Norm(MLP(MSA(Norm(FA i-1 ),Norm(DP i ))))+FA i-1 Among them, DP i , FA i Respectively represent the output learnable data domain prior, feature map, DP i-1 , FA i-1 They represent the learnable data domain prior and feature map of the input respectively, Norm(·) represents layer normalization, MLP(·) represents multi-layer perceptron, and MSA(·) represents multi-head self-attention mechanism.
9. The general segmentation method for tubular structure three-dimensional medical images according to claim 1, characterized in that: In S3, the loss function L of the three-dimensional medical image general segmentation model is a combined loss function: L seg =L DICE +L CE +L VAS +L dom Among them, L DICE represents the DICE loss function, which is used to measure the similarity between the predicted result and the true label; L CE represents the cross entropy loss function, which is used to evaluate the performance of classification tasks; L VAS represents the topological loss function, which is used to capture the geometric non-smoothness of the vascular structure; L dom Represents the domain prediction loss function, which is used to train the model to identify the dataset domain to which the image belongs.
10. The general segmentation method for tubular structure three-dimensional medical images according to claim 1, characterized in that: In S3, optimization is performed using the preprocessed data set, including: In the training phase, the AdamW optimizer is used to update the network parameters, the cosine annealing strategy is used to dynamically adjust the learning rate, and random scaling, random rotation, random gamma enhancement, and random mirroring are combined for data enhancement. During the testing phase, the 3D image is segmented into multiple overlapping image blocks, the segmented image blocks and the corresponding data domain prior vector weights are input into the model, and the prediction results of the 3D image are obtained by merging the prediction results of all image blocks.