Flat-scanning CT image automatic segmentation method based on deep learning model

By cascading the Deep-MSVM-UNet model, combined with the multi-scale visual state space block and UNet architecture, the problems of time-consuming and labor-intensive brain segmentation and insufficient precision are solved, and efficient and accurate brain segmentation is achieved, adapting to the complex and changeable brain structure and meeting the real-time requirements of clinical applications.

CN120765668APending Publication Date: 2025-10-10BEIJING INST OF TECH +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510810318.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-21
Filing Date
2025-06-17
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing brain region segmentation methods are time-consuming and labor-intensive, and are limited by operator experience, making it difficult to ensure the consistency and accuracy of the segmentation results. A single deep learning model lacks segmentation accuracy when processing complex and changeable brain structures.

Method used

A cascaded Deep-MSVM-UNet model is adopted, combined with multi-scale visual state space blocks and UNet architecture. Through two-stage training, the model's segmentation accuracy and adaptability to brain areas are improved. Multi-scale visual state space blocks are used to capture and aggregate multi-scale features, and linear computational complexity is used to improve computational efficiency.

Benefits of technology

It improves the accuracy and adaptability of brain region segmentation, reduces complexity, meets the real-time requirements of clinical applications, reduces doctors' workload, and improves diagnosis and research efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765668A_ABST
    Figure CN120765668A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning model-based automatic segmentation method for a plain-scan CT image, and the method comprises the steps: carrying out the region segmentation of a brain on the section image data of the brain through employing a cascading Deep-MSVM-UNet model trained in a first stage, extracting a rough brain region contour, and then carrying out the segmentation of the region of the brain through employing a Cascade Deep-MSVM-UNet model, and inputting the segmentation result of the brain region into the Deep-MSVM-UNet model trained in the second stage to carry out fine segmentation on each brain region so as to further accurately divide the boundary and the internal structure of the brain region. Accurate segmentation of the brain region is finally realized through multiple iterations and cascade processing; according to the cascade Deep-MSVM-UNet model constructed by the method, multi-scale visual state space blocks and a UNet framework are combined, multi-scale feature representation can be more effectively captured and aggregated, and meanwhile, the long-time dependency relationship between pixels is simulated, so that the brain region segmentation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and specifically relates to an automatic segmentation method for plain scan CT images based on a deep learning model, which is suitable for segmenting different brain regions from medical imaging data (NCCT). Background Art

[0002] In the field of medical image analysis, brain segmentation is a crucial task. Traditional brain segmentation methods mainly rely on manual operations or rule-based automated algorithms. These methods are not only time-consuming and labor-intensive, but also limited by the operator's experience and knowledge, making it difficult to ensure the consistency, accuracy and efficiency of the segmentation results.

[0003] The rapid development of computer technology, particularly the widespread application of deep learning techniques, has revolutionized the field of medical image segmentation. Deep learning models, particularly convolutional neural networks (CNNs), have demonstrated tremendous potential in medical image segmentation due to their powerful feature extraction and pattern recognition capabilities. However, given the complex and diverse brain structures and image data, a single deep learning model often struggles to achieve ideal segmentation results.

[0004] In recent years, the UNet model, a classic architecture in the field of medical image segmentation, has achieved accurate segmentation of fine structures in medical images with its symmetrical encoder-decoder structure and skip connections. However, the UNet model may still face the problem of insufficient segmentation accuracy when processing brain regions with complex textures and shape variations. Summary of the Invention

[0005] The purpose of the present invention is to address the above-mentioned problems existing in the prior art and provide a method for automatic segmentation of plain scan CT images based on a deep learning model.

[0006] The above-mentioned purpose of the present invention is achieved by the following technical means:

[0007] A method for automatic segmentation of plain scan CT images based on a deep learning model comprises the following steps:

[0008] Step 1: Obtain three-dimensional brain image data and corresponding annotation data from a public brain image dataset, where the annotation data includes the brain region division corresponding to the brain image data;

[0009] Step 2: performing data preprocessing on each acquired brain image data to obtain preprocessed brain image data corresponding to each brain image data;

[0010] Step 3: Slice each pre-processed brain image data into training samples. Each training sample includes an input sample, a brain segmentation label, and a brain region segmentation label. All training samples constitute a sample set, which is then divided into a training set, a validation set, and a test set.

[0011] Step 4: Build the cascaded Deep-MSVM-UNet model;

[0012] Step 5: Construct the loss function Loss of the cascaded Deep-MSVM-UNet model;

[0013] Step 6: Use the input samples and corresponding brain segmentation labels in the training set to perform the first stage training on the cascaded Deep-MSVM-UNet model trained in the first stage to obtain the final prediction results of the first training model parameters and the brain segmentation corresponding to each input sample;

[0014] Finally, the second-stage training of the cascaded Deep-MSVM-UNet model is performed using the input samples, brain region division labels, and the final prediction results of brain division to obtain the second-stage training model parameters;

[0015] Step 7: Obtain the three-dimensional brain image data to be processed, perform data preprocessing and slicing to obtain slice image data, and then input each slice image data into the cascade Deep-MSVM-UNet model to obtain the prediction results of the brain area division corresponding to each slice image data. Finally, the prediction results of the brain area division corresponding to each slice image data are spliced ​​into a three-dimensional brain area segmentation result.

[0016] The data preprocessing in step 2 mentioned above specifically includes the following steps:

[0017] Step 2.1, adjust the window level and window width of brain imaging data;

[0018] Step 2.2, augmenting brain imaging data through images;

[0019] Step 2.3: removing noise from the expanded brain image data;

[0020] Step 2.4: Crop the denoised brain image data, and normalize all the cropped brain image data to the same size by interpolation to obtain preprocessed brain image data.

[0021] The training samples in step 3 are prepared as follows:

[0022] The preprocessed brain image data is respectively axially sliced into a plurality of slice image data, and brain division labels and brain region division labels of each slice image data are obtained according to the label data of each brain image data; the slice image data is taken as an input sample, each input sample and the corresponding brain division label and brain region division label are taken as a training sample.

[0023] The cascade Deep-MSVM-UNet model in step 4 comprises I encoder modules, when i is 1, the input of the i-th encoder module is the input sample, when i is 2, 3, …, I, the input of the i-th encoder module is the output of the i-1-th encoder;

[0024] The cascade Deep-MSVM-UNet model further comprises J decoder modules, when j is 1, 2, …, J-1, the input of the j-th decoder module is the output of the j-th encoder module and the output of the j+1-th decoder module, when j is J, the input of the j-th decoder module is the output of the j-th encoder module and the output of the j+1-th encoder module, and the output of the encoder module and the input of the corresponding decoder module are connected through a shortcut connection;

[0025] The cascade Deep-MSVM-UNet model further comprises an output layer, the input of the output layer is the output of the first decoder module, and the output of the output layer is the output of the cascade Deep-MSVM-UNet model; the output layer comprises a plurality of large-kernel Patch expansion layers stacked in series;

[0026] Wherein, i and j are both serial numbers, serial number i is 1, 2, …, I, serial number j is 1, 2, …, J, and J = I-1.

[0027] In the encoder module, when i is 1, the i-th encoder module comprises an embedding layer and a visual state space block stacking module, the input of the embedding layer is the input sample, the output of the embedding layer is input into the visual state space block stacking module of the first encoder module, and the output of the visual state space block stacking module of the first encoder module is input into the second encoder module;

[0028] When i is 2, 3, …, I, the i-th encoder module comprises a Patch fusion layer and a visual state space block stacking module, the input of the Patch fusion layer of the i-th encoder module is the output of the visual state space block stacking module of the i-1-th encoder module, and the input of the visual state space block stacking module of the i-th encoder module is the output of the Patch fusion layer of the i-th encoder module; the visual state space block module comprises a plurality of visual state space blocks stacked in series.

[0029] In the decoder module, each decoder module includes a large kernel patch extension layer stacking module and a multi-scale visual state space block stacking module. The input of the multi-scale visual state space block stacking module of the j-th decoder module is the result of channel splicing operation on the output of the large kernel patch extension layer stacking module of the j-th decoder module and the output of the visual state space block stacking module of the j-th encoder module;

[0030] The multi-scale visual state space block stacking module includes a plurality of multi-scale visual state space blocks stacked in series, and the large kernel patch extension layer stacking module includes a plurality of large kernel patch extension layers stacked in series;

[0031] When j is 1, 2, ..., J-1, the input of the large kernel Patch extension layer stacking module of the j-th decoder module is the output of the multi-scale visual state space block stacking module of the j+1-th decoder module; when j is J, the input of the large kernel Patch extension layer stacking module of the j-th decoder module is the output of the visual state space block stacking module of the j+1-th encoder module.

[0032] As described above, the cascaded Deep-MSVM-UNet model includes four encoder modules, three decoder modules, and an output layer. The first encoder module, the second encoder module, and the fourth encoder module each include two visual state space blocks stacked in series. The third encoder module includes nine visual state space blocks stacked in series. The multi-scale visual state space block stacking modules of the three decoder modules each include two multi-scale visual state space blocks stacked in series. The large kernel Patch extension layer stacking modules of the three decoder modules include two large kernel Patch extension layers stacked in series, and the output layer includes four large kernel Patch extension layers stacked in series.

[0033] As mentioned above, the loss function of the cascaded Deep-MSVM-UNet model is based on the following formula:

[0034] Loss=α*Dice Loss+β*CE Loss

[0035] Where Dice Loss is the Dice coefficient loss function, CE Loss is the cross entropy loss function, α is the weight coefficient of the Dice coefficient loss function, and β is the weight coefficient of the cross entropy loss function.

[0036] The above step 6 specifically includes the following steps:

[0037] Step 6.1: Conduct the first phase of training, which includes the following steps:

[0038] Step 6.1.1, input the input samples in the training set into the cascade Deep-MSVM-UNet model trained in the first stage to obtain the first brain division prediction result output by the first decoder module, the second brain division prediction result output by the second decoder module, and the third brain division prediction result output by the third decoder module;

[0039] Step 6.1.2, calculate the loss function of the first brain division prediction result, the second brain division prediction result, and the third brain division prediction result with the brain division label respectively to obtain the first brain division loss function, the second brain division loss function, and the third brain division loss function respectively;

[0040] Step 6.1.3, finally, the first brain division loss function, the second brain division loss function, and the third brain division loss function are weighted and summed to obtain the final brain division loss function;

[0041] Step 6.1.4, use the AdamW optimizer, set the initial learning rate and the decay factor, use the cosine decay strategy for learning rate decay, minimize the loss function; use the early stopping strategy to determine whether to stop training according to the performance of the validation set to prevent overfitting; save the model parameters of the cascade Deep-MSVM-UNet model trained in the first stage after training, denoted as the first training model parameters;

[0042] Step 6.2, perform second-stage model training, specifically including the following steps:

[0043] Step 6.2.1, input the input samples in the training set into the cascade Deep-MSVM-UNet model using the first training parameters to obtain the prediction result of the brain division corresponding to each input sample;

[0044] Step 6.2.2, input the input samples in the training set and the corresponding prediction result of the brain division obtained in step 6.2.1 into the cascade Deep-MSVM-UNet model trained in the second stage to obtain the first brain region division prediction result output by the first decoder module, the second brain region division prediction result output by the second decoder module, and the third brain region division prediction result output by the second decoder module;

[0045] Step 6.2.3, calculate the loss function of the first brain region division prediction result, the second brain region division prediction result, and the third brain region division prediction result with the brain region division label to obtain the first brain region division loss function, the second brain region division loss function, and the third brain region division loss function;

[0046] Step 6.2.4: Finally, perform a weighted summation of the first brain region partition loss function, the second brain region partition loss function, and the third brain region partition loss function to obtain the final brain region partition prediction loss function.

[0047] Step 6.2.5: Use the AdamW optimizer to set the initial learning rate and decay factor, use the cosine decay strategy to decay the learning rate, and minimize the loss function; use the early stopping strategy to determine whether to stop training based on the performance of the validation set to prevent overfitting; after training is completed, save the model parameters of the cascaded Deep-MSVM-UNet model trained in the second stage and record them as the second training model parameters.

[0048] Step 7 as described above specifically includes the following steps:

[0049] Step 7.1, first perform data preprocessing in step 2 on the brain image data to be processed to obtain normalized brain image data to be processed;

[0050] Step 7.2, then executing step 3, performing axial slicing on the pre-processed brain image data to be processed, to obtain a plurality of two-dimensional slice image data to be processed;

[0051] Step 7.3: Input each slice image data to be processed into the cascaded Deep-MSVM-UNet model of the first training model parameters, and the output layer of the cascaded Deep-MSVM-UNet model of the first training model parameters outputs the predicted results of the brain partition corresponding to each slice image data to be processed;

[0052] Step 7.4: Input each slice image data to be processed and the corresponding brain segmentation prediction result into the cascade Deep-MSVM-UNet model of the second training model parameters, and the output layer of the cascade Deep-MSVM-UNet model of the second training model parameters outputs the prediction result of the brain region segmentation corresponding to each slice image data to be processed;

[0053] Step 7.5: axially splice the predicted results of brain region division corresponding to each slice image data to be processed to obtain a three-dimensional brain region segmentation result corresponding to the three-dimensional brain image data to be processed.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] (1) Improved segmentation accuracy: The cascaded Deep-MSVM-UNet model constructed by the present invention, combined with the multi-scale visual state space (MSVSS) block and the UNet architecture, can more effectively capture and aggregate multi-scale feature representations while simulating the long-term dependencies between pixels, thereby improving the accuracy of brain segmentation;

[0056] (2) Enhanced model adaptability: The cascaded Deep-MSVM-UNet model of the present invention also enhances the model's adaptability to brain regions of different sizes and shapes by introducing multi-scale visual state space blocks. This enables the model to perform well in processing complex and variable brain region structures;

[0057] (3) Improved computational efficiency: The cascaded Deep-MSVM-UNet model of the present invention adopts the linear computational complexity of the state-space model, significantly improving computational efficiency. This enables the model to process large-scale data while maintaining a high segmentation speed, meeting the real-time requirements of clinical applications;

[0058] (4) The present invention uses deep learning algorithms to automatically process brain images without manual intervention, which can significantly reduce the workload of doctors and improve the efficiency of diagnosis and research;

[0059] (5) Reduce the complexity of brain segmentation: By optimizing the model structure and algorithm design, the complexity of brain segmentation is reduced, so that the method of the present invention can more efficiently process large-scale brain imaging data and provide more timely and accurate support for clinical research and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 Flowchart of data preprocessing of the present invention;

[0061] Figure 2 Schematic diagram of the structure of the cascaded Deep-MSVM-UNet model of the present invention;

[0062] Figure 3 This is a flow chart of the first stage training of the present invention;

[0063] Figure 4 This is a flow chart of the second stage training of the present invention;

[0064] Figure 5 Schematic diagram of the overall structure of the two training processes of the present invention;

[0065] Figure 6 Schematic diagram of the overall structure of the present invention for performing brain region division (i.e., step 7) on the brain image data to be processed. DETAILED DESCRIPTION

[0066] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below with reference to the embodiments. The embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.

[0067] Example 1:

[0068] A method for automatic segmentation of plain scan CT images based on a deep learning model comprises the following steps:

[0069] Step 1: Acquire three-dimensional brain image data from a public brain image dataset (such as the BHSD public dataset, the INSTANCE public dataset, and the CQ500 public dataset). The dataset of this embodiment is an NCCT type brain image. The brain image data should include detailed annotation data of brain region division. In this embodiment, the brain region is divided into 22 regions, namely the left frontal lobe and the right frontal lobe, the left parietal lobe and the right parietal lobe, the left occipital lobe and the right occipital lobe, the left temporal lobe and the right temporal lobe, the left insula and the right insula, the left basal ganglia and the right basal ganglia, the left cerebellum and the right cerebellum, the left midbrain and the right midbrain, the left pons and the right pons, the left medulla oblongata and the right medulla oblongata, and the left thalamus and the right thalamus. These annotation data are manually annotated and generated by professional doctors.

[0070] Step 2: Preprocess the acquired brain image data to obtain preprocessed brain image data corresponding to each brain image data. Data preprocessing is a key step to ensure that the model can be effectively trained, and specifically includes the following steps:

[0071] Step 2.1, Window Level and Window Width Adjustment: Adjust the window level and window width of the brain image data to make each brain region more vivid in the image. In this embodiment, the window level is set to 40 and the window width is set to 80;

[0072] Step 2.2, Image Enhancement: Data enhancement is performed by rotating, flipping, scaling, and adding noise to expand brain image data and avoid model overfitting;

[0073] Step 2.3, denoising: Use denoising algorithms (such as median filtering and Gaussian filtering) to remove noise from the expanded brain image data and improve image quality;

[0074] Step 2.4, Image Cropping and Normalization: Crop the denoised brain image data to focus on the brain area, remove excess background, and reduce the amount of calculation; then normalize all the cropped brain image data to the same size using interpolation to obtain preprocessed brain image data. In this embodiment, the size of the preprocessed brain image data is 512x512.

[0075] Step 3: To train the model used in the present invention, the three-dimensional brain image data needs to be converted into two-dimensional brain image data. That is, each pre-processed brain image data is sliced ​​and used as a training sample. Each training sample includes an input sample, a brain segmentation label, and a brain region segmentation label. All training samples constitute a sample set, which is then divided into a training set, a validation set, and a test set. The specific steps are as follows:

[0076] Step 3.1, the pre-processed brain image data is respectively axial (horizontal slice along the z axis) sliced into multiple slice image data (slices), and the brain division label (BRAIN_MASK) and the brain region division label (MASK) of each slice image data are obtained according to the label data of each brain image data, wherein the brain division label is a brain label, and the brain region division label is a detailed brain region division corresponding to the slice image data. In this embodiment, 22 brain region division labels are included, corresponding to 22 brain regions respectively;

[0077] Step 3.2, each slice image data is taken as an input sample, each input sample and the corresponding brain division label and brain region division label are taken as a training sample; all the training samples constitute a sample set;

[0078] Step 3.3, the training samples in the sample set are divided into a training set, a validation set, and a test set according to a set proportion.

[0079] The three-dimensional brain image data in this embodiment is made into multiple two-dimensional brain image data with a size of 512x512.

[0080] The brain division label input sample is used for the first stage training of the cascaded Deep-MSVM-UNet model, and the brain region division label is used for the second stage training of the cascaded Deep-MSVM-UNet model.

[0081] Step 4, a cascaded Deep-MSVM-UNet model is constructed:

[0082] The serial number i is set to {1, 2, …, I}, and the serial number j is set to {1, 2, …, J (also I-1)};

[0083] The cascaded Deep-MSVM-UNet model includes I encoder modules, when i is 1, the input of the i-th encoder module is the input sample, when i is 2, 3, …, I, the input of the i-th encoder module is the output of the i-1-th encoder;

[0084] The cascaded Deep-MSVM-UNet model also includes J decoder modules, when j is 1, 2, …, J-1, the input of the j-th decoder module is the output of the j-th encoder module and the output of the j+1-th decoder module, when j is J, the input of the j-th decoder module is the output of the j-th encoder module and the output of the j+1-th encoder module, and the output of the encoder module and the input of the corresponding decoder module are connected through a shortcut connection.

[0085] The cascaded Deep-MSVM-UNet model also includes an output layer and multiple intermediate output layers. The input of the output layer is the output of the first decoder module, and the output of the output layer is the output of the cascaded Deep-MSVM-UNet model. Both the output layer and the intermediate output layer include multiple large kernel patch expansion layers (LKPE) stacked in series.

[0086] It should be noted that i and j in the present invention only represent serial numbers, such as the i-th encoder and the j-th encoder. When i=j, the i-th encoder and the j-th encoder are the same encoder.

[0087] For one encoder module, specifically:

[0088] When i is 1, the i-th encoder module includes an embedding layer (Patch Embedding) and a visual state space block stacking module. The input of the embedding layer (also the input of the cascaded Deep-MSVM-UNet model) is the input sample, the output of the embedding layer is input to the visual state space block stacking module of the first encoder module, and the output of the visual state space block stacking module of the first encoder module (also the output of the first encoder module) is input to the second encoder module;

[0089] When i is 2, 3, ..., 1, the i-th encoder module includes a Patch fusion layer (Patch Merging) and a visual state space block stacking module, the input of the Patch fusion layer of the i-th encoder module is the output of the visual state space block stacking module of the i-1-th encoder module, the input of the visual state space block stacking module of the i-th encoder module is the output of the Patch fusion layer of the i-th encoder module, and the output of the visual state space block stacking module of the i-th encoder module is the output of the i-th encoder module; the visual state space block module includes multiple visual state space blocks stacked in series;

[0090] For J decoder modules, specifically:

[0091] Each decoder module includes a large kernel patch expansion layer stacking module and a multi-scale visual state space block stacking module. The input of the multi-scale visual state space block stacking module of the j-th decoder module is the result of the channel concatenation operation (C) of the output of the large kernel patch expansion layer stacking module of the j-th decoder module and the output of the visual state space block stacking module of the j-th encoder module.

[0092] The multi-scale visual state space block stacking module includes a plurality of multi-scale visual state space blocks stacked in series, and the large kernel patch extension layer stacking module includes a plurality of large kernel patch extension layers stacked in series;

[0093] When j is 1, 2, ..., J-1, the input of the large kernel Patch extension layer stacking module of the j-th decoder module is the output of the multi-scale visual state space block stacking module of the j+1-th decoder module; when j is J, the input of the large kernel Patch extension layer stacking module of the j-th decoder module is the output of the visual state space block stacking module of the j+1-th encoder module.

[0094] In this embodiment, four encoder modules, three decoder modules, and one output layer are included. The first encoder module, the second encoder module, and the fourth encoder module each include two visual state space blocks stacked in series (i.e., Figure 2 The third encoder module includes nine visual state space blocks stacked in series (i.e. Figure 2 The multi-scale visual state space block stacking modules of the three decoder modules each include two multi-scale visual state space blocks stacked in series (i.e. Figure 2 The MSVSS Block×2 in the image is the Multi-Scale Vision State Space. The large-core Patch extension layer stacking module of the three decoder modules includes two large-core Patch extension layers stacked in series (i.e. Figure 2 The output layer consists of four large kernel Patch expansion layers stacked in series (i.e. Figure 2 LKPE 4X in ).

[0095] The output of the second encoder of this embodiment is also input into the intermediate output layer LKPE 8X, which is used to output the predicted results of the brain division or the predicted results of the brain region division of the second encoder; the output of the third encoder is also input into the intermediate output layer LKPE 16X, which is used to output the predicted results of the brain division or the predicted results of the brain region division of the third encoder.

[0096] The cascaded Deep-MSVM-UNet model used in the present invention adopts a U-shaped hierarchical encoder-decoder structure and adopts a shortcut connection (short-cut) between the encoder and the decoder, wherein the embedding layer is used to divide the input image (i.e., input sample) into non-overlapping blocks of size 4x4 and map the channel dimension to 96 dimensions; the visual state space block is responsible for learning the hierarchical feature representation of the input image, and the visual state space block selectively scans along the four scanning paths of the input feature map through 2D selective scanning (2D-Selective-ScanBlock, SS2D) to capture global context information and long-range dependencies; the patch fusion layer is used to downsample the input feature map; in the decoder, unlike the patch fusion layer, the large kernel patch expansion layer is used to upsample the input feature map; the multi-scale visual state space block introduces a multi-scale feed-forward network (Multi-Scale Feed-Forward In this paper, we propose a multi-scale feed-forward network (MS-FFN) with a multi-scale feed-forward network (MS-FFN). Each network uses depthwise separable convolutions with kernel sizes of 1, 3, and 5, respectively. Multi-scale feature extraction is achieved through parallel branches, effectively capturing and aggregating fine-grained multi-scale information from the contracting path and high-level semantic information from the expanding path. Finally, deep supervision is used to predict each decoder stage using the LKPE module to obtain the predicted segmentation results.

[0097] Step 5. Construct the loss function Loss of the cascaded Deep-MSVM-UNet model. The loss function Loss of the cascaded Deep-MSVM-UNet model is based on the following formula:

[0098] Loss=α*Dice Loss+β*CE Loss

[0099] Wherein, Dice Loss is the Dice coefficient loss function, CE Loss is the cross entropy loss function, α is the weight coefficient of the Dice coefficient loss function, and β is the weight coefficient of the cross entropy loss function. In this embodiment, α is 0.8 and β is 0.5. The Dice coefficient can effectively measure the accuracy of brain region segmentation by the network model and is suitable for medical image region segmentation tasks.

[0100] Step 6: Training the cascaded Deep-MSVM-UNet model includes two stages: first-stage training and second-stage training. Both the first-stage training and the second-stage training use the cascaded Deep-MSVM-UNet model. After the training is completed, the model parameters of the first-stage training and the model parameters of the second-stage training are saved respectively. Specifically, the following steps are included:

[0101] Step 6.1: Conduct the first phase of training, which includes the following steps:

[0102] Step 6.1.1. Input the input samples in the training set into the cascaded Deep-MSVM-UNet model to obtain the prediction results of the three brain partitions output by the three decoders (including the first brain partition prediction result output by the first decoder module, the second brain partition prediction result output by the second decoder module, and the third brain partition prediction result output by the third decoder module);

[0103] Step 6.1.2: Calculate the loss function of the first brain segmentation prediction result, the second brain segmentation prediction result, and the third brain segmentation prediction result with the brain segmentation label through the loss function Loss of the cascade Deep-MSVM-UNet model to obtain the loss function of the first brain segmentation prediction result, the loss function of the second brain segmentation prediction result, and the loss function of the third brain segmentation prediction result, which are respectively denoted as the first brain segmentation loss function, the second brain segmentation loss function, and the third brain segmentation loss function;

[0104] Step 6.1.3. Finally, perform a weighted summation of the first brain partition loss function, the second brain partition loss function, and the third brain partition loss function to obtain the final brain partition prediction loss function.

[0105] In this embodiment, the weight coefficients of the loss function of the first prediction result of brain segmentation, the loss function of the second prediction result of brain segmentation, and the loss function of the third prediction result of brain segmentation are 0.3, 0.6, and 1, respectively;

[0106] Step 6.1.4, set the optimizer and learning rate. In this embodiment, the AdamW optimizer is used to optimize the cascaded Deep-MSVM-UNet model trained in the first stage, and the network model parameters are gradually adjusted to minimize the loss function; the cosine decay strategy is used to decay the learning rate, and the initial learning rate and decay factor are set. In this embodiment, the initial learning rate is set to 0.001 and the decay factor is set to 0.1; the early stopping strategy is adopted during training, and whether to stop training is determined based on the performance of the validation set to prevent overfitting; after the training is completed, the model parameters of the cascaded Deep-MSVM-UNet model trained in the first stage are saved and recorded as the first training model parameters.

[0107] Step 6.2: Perform the second phase of model training, which specifically includes the following steps:

[0108] Step 6.2.1. Input the input samples in the training set into the cascaded Deep-MSVM-UNet model using the parameters of the first training model. The output layer of the cascaded Deep-MSVM-UNet model using the parameters of the first training model outputs the predicted results of the brain partition corresponding to each input sample.

[0109] Step 6.2.2: Input the input samples in the training set and the corresponding brain partition prediction results obtained in step 6.2.1 into the cascaded Deep-MSVM-UNet model trained in the second stage to obtain the three brain region prediction results output by the three decoders (including the first brain region prediction result output by the first decoder module, the second brain region prediction result output by the second decoder module, and the third brain region prediction result output by the third decoder module);

[0110] Step 6.2.3. Calculate the loss function using the concatenated Deep-MSVM-UNet model's loss function, using the first brain region division prediction result, the second brain region division prediction result, and the third brain region division prediction result as well as the brain region division label. The loss function for the first brain region division prediction result, the loss function for the second brain region division prediction result, and the loss function for the third brain region division prediction result are recorded as the first brain region division loss function, the second brain region division loss function, and the third brain region division loss function, respectively.

[0111] Step 6.2.4: Finally, perform a weighted summation of the first brain region partition loss function, the second brain region partition loss function, and the third brain region partition loss function to obtain the final brain region partition prediction loss function.

[0112] Step 6.2.5. Use the AdamW optimizer to optimize the cascaded Deep-MSVM-UNet model trained in the second stage, gradually adjust the network model parameters, and minimize the loss function; use the cosine decay strategy to decay the learning rate, set the initial learning rate and decay factor; in this embodiment, the initial learning rate is set to 0.001, and the decay factor is set to 0.1; adopt the early stopping strategy during training, and determine whether to stop training based on the performance of the validation set to prevent overfitting; after the training is completed, save the model parameters of the cascaded Deep-MSVM-UNet model trained in the second stage and record them as the second training model parameters.

[0113] After the training is completed, the model is evaluated: the effectiveness of the present invention is verified by the test set, and the Dice coefficient, precision, and recall rate are used to evaluate the cascade Deep-MSVM-UNet model of the second training model parameters.

[0114] Step 7: Obtain the 3D brain image data to be processed, perform data preprocessing and slicing to obtain slice image data, and then input each slice image data into the cascade Deep-MSVM-UNet model to obtain the prediction results of the brain region division corresponding to each slice image data. Finally, the prediction results of the brain region division corresponding to each slice image data are spliced ​​into a 3D brain region segmentation result, thereby obtaining the 3D brain region segmentation result corresponding to the 3D brain image data. The specific processing includes the following:

[0115] Step 7.1, first perform data preprocessing in step 2 on the brain image data to be processed to obtain normalized brain image data to be processed;

[0116] Step 7.2, then executing step 3, performing axial slicing on the pre-processed brain image data to be processed, to obtain a plurality of two-dimensional slice image data to be processed;

[0117] Step 7.3: Input each slice image data to be processed into the output layer of the cascaded Deep-MSVM-UNet model of the first training model parameters to output the predicted results of the brain division corresponding to each slice image data to be processed;

[0118] Step 7.4: Input each slice image data to be processed and the corresponding brain segmentation prediction result into the cascade Deep-MSVM-UNet model of the second training model parameters, and the output layer of the cascade Deep-MSVM-UNet model of the second training model parameters outputs the prediction result of the brain region segmentation corresponding to each slice image data to be processed;

[0119] Step 7.5: axially splice the predicted results of brain region division corresponding to each slice image data to be processed to obtain a three-dimensional brain region segmentation result corresponding to the three-dimensional brain image data to be processed.

[0120] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0122] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0123] It should be noted that the embodiments described in the present application are only examples to illustrate the spirit of the present application. Those skilled in the art of the present application can make various modifications or supplements to the described embodiments or replace them with similar ways, but will not deviate from the spirit of the present application or exceed the scope defined by the appended claims.

Claims

1. A method for automatic segmentation of plain scan CT images based on a deep learning model, characterized in that: The following steps are involved: Step 1: Obtain three-dimensional brain image data and corresponding annotation data from a public brain image dataset, where the annotation data includes the brain region division corresponding to the brain image data; Step 2: performing data preprocessing on each acquired brain image data to obtain preprocessed brain image data corresponding to each brain image data; Step 3: Slice each pre-processed brain image data into training samples. Each training sample includes an input sample, a brain segmentation label, and a brain region segmentation label. All training samples constitute a sample set, which is then divided into a training set, a validation set, and a test set. Step 4: Build the cascaded Deep-MSVM-UNet model; Step 5: Construct the loss function Loss of the cascaded Deep-MSVM-UNet model; Step 6: Use the input samples and corresponding brain segmentation labels in the training set to perform the first stage training on the cascaded Deep-MSVM-UNet model trained in the first stage to obtain the final prediction results of the first training model parameters and the brain segmentation corresponding to each input sample; Finally, the second-stage training of the cascaded Deep-MSVM-UNet model is performed using the input samples, brain region division labels, and the final prediction results of brain division to obtain the second-stage training model parameters; Step 7: Obtain the three-dimensional brain image data to be processed, perform data preprocessing and slicing to obtain slice image data, and then input each slice image data into the cascade Deep-MSVM-UNet model to obtain the prediction results of the brain area division corresponding to each slice image data. Finally, the prediction results of the brain area division corresponding to each slice image data are spliced ​​into a three-dimensional brain area segmentation result.

2. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 1, characterized in that: The data preprocessing in step 2 specifically includes the following steps: Step 2.1, adjust the window level and window width of brain imaging data; Step 2.2, augmenting brain imaging data through images; Step 2.3: removing noise from the expanded brain image data; Step 2.4: Crop the denoised brain image data, and normalize all the cropped brain image data to the same size by interpolation to obtain preprocessed brain image data.

3. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 1, characterized in that: The training samples in step 3 are prepared in the following way: Each preprocessed brain image data is axially sliced ​​into multiple slice image data, and the brain division label and brain region division label of each slice image data are obtained based on the annotation data of each brain image data; the slice image data is used as an input sample, and each input sample and the corresponding brain division label and brain region division label are a training sample.

4. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 1, characterized in that: The cascaded Deep-MSVM-UNet model in step 4 includes I encoder modules. When i is 1, the input of the i-th encoder module is the input sample. When i is 2, 3, ..., 1, the input of the i-th encoder module is the output of the i-1-th encoder. The cascaded Deep-MSVM-UNet model also includes J decoder modules. When j is 1, 2, …, J-1, the input of the j-th decoder module is the output of the j-th encoder module and the output of the j+1-th decoder module. When j is J, the input of the j-th decoder module is the output of the j-th encoder module and the output of the j+1-th encoder module. The output of the encoder module is connected to the input of the corresponding decoder module through a shortcut connection. The cascaded Deep-MSVM-UNet model also includes an output layer, the input of which is the output of the first decoder module, and the output of which is the output of the cascaded Deep-MSVM-UNet model; the output layer includes multiple large-kernel Patch extension layers stacked in series; Here, i and j are serial numbers, the serial number i is 1, 2, ..., I, the serial number j is 1, 2, ..., J, and J=I-1.

5. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 4, characterized in that: In the encoder module, when i is 1, the i-th encoder module includes an embedding layer and a visual state space block stacking module. The input of the embedding layer is the input sample, the output of the embedding layer is input to the visual state space block stacking module of the first encoder module, and the output of the visual state space block stacking module of the first encoder module is input to the second encoder module. When i is 2, 3, ..., 1, the i-th encoder module includes a patch fusion layer and a visual state space block stacking module, the input of the patch fusion layer of the i-th encoder module is the output of the visual state space block stacking module of the i-1-th encoder module, and the input of the visual state space block stacking module of the i-th encoder module is the output of the patch fusion layer of the i-th encoder module; The visual state space block module includes multiple visual state space blocks stacked in series.

6. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 5, characterized in that: In the decoder module, each decoder module includes a large kernel patch extension layer stacking module and a multi-scale visual state space block stacking module. The input of the multi-scale visual state space block stacking module of the j-th decoder module is the result of channel splicing operation on the output of the large kernel patch extension layer stacking module of the j-th decoder module and the output of the visual state space block stacking module of the j-th encoder module; The multi-scale visual state space block stacking module includes a plurality of multi-scale visual state space blocks stacked in series, and the large kernel patch extension layer stacking module includes a plurality of large kernel patch extension layers stacked in series; When j is 1, 2, ..., J-1, the input of the large kernel Patch extension layer stacking module of the j-th decoder module is the output of the multi-scale visual state space block stacking module of the j+1-th decoder module; when j is J, the input of the large kernel Patch extension layer stacking module of the j-th decoder module is the output of the visual state space block stacking module of the j+1-th encoder module.

7. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 6, characterized in that: The cascaded Deep-MSVM-UNet model includes four encoder modules, three decoder modules, and an output layer. The first encoder module, the second encoder module, and the fourth encoder module each include two visual state space blocks stacked in series, the third encoder module includes nine visual state space blocks stacked in series, the multi-scale visual state space block stacking modules of the three decoder modules each include two multi-scale visual state space blocks stacked in series, the large kernel Patch extension layer stacking modules of the three decoder modules include two large kernel Patch extension layers stacked in series, and the output layer includes four large kernel Patch extension layers stacked in series.

8. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 1, characterized in that: The loss function Loss of the cascaded Deep-MSVM-UNet model is based on the following formula: Loss=α*Dice Loss+β*CE Loss Where Dice Loss is the Dice coefficient loss function, CE Loss is the cross entropy loss function, α is the weight coefficient of the Dice coefficient loss function, and β is the weight coefficient of the cross entropy loss function.

9. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 1, characterized in that: The step 6 specifically includes the following steps: Step 6.1: Conduct the first phase of training, which includes the following steps: Step 6.1.

1. Input the input samples in the training set into the cascaded Deep-MSVM-UNet model trained in the first stage to obtain the first brain segmentation prediction result output by the first decoder module, the second brain segmentation prediction result output by the second decoder module, and the third brain segmentation prediction result output by the third decoder module; Step 6.1.2: Calculate the loss function for the first brain segmentation prediction result, the second brain segmentation prediction result, and the third brain segmentation prediction result with the brain segmentation label to obtain the first brain segmentation loss function, the second brain segmentation loss function, and the third brain segmentation loss function, respectively. Step 6.1.

3. Finally, perform a weighted summation of the first brain partition loss function, the second brain partition loss function, and the third brain partition loss function to obtain the final brain partition loss function. Step 6.1.4: Use the AdamW optimizer to set the initial learning rate and decay factor, and use the cosine decay strategy to decay the learning rate to minimize the loss function. Use the early stopping strategy to determine whether to stop training based on the performance of the validation set to prevent overfitting. After training is complete, save the model parameters of the cascaded Deep-MSVM-UNet model trained in the first stage and record them as the first training model parameters. Step 6.2: Perform the second phase of model training, which specifically includes the following steps: Step 6.2.

1. Input the input samples in the training set into the cascaded Deep-MSVM-UNet model using the first training parameters to obtain the predicted brain partitions corresponding to each input sample. Step 6.2.2: Input the input samples in the training set and the corresponding brain segmentation prediction results obtained in step 6.2.1 into the cascaded Deep-MSVM-UNet model trained in the second stage to obtain the first brain region segmentation prediction result output by the first decoder module, the second brain region segmentation prediction result output by the second decoder module, and the third brain region segmentation prediction result output by the second decoder module; Step 6.2.3: Calculate the loss function for the first brain region partition prediction result, the second brain region partition prediction result, and the third brain region partition prediction result with the brain region partition label to obtain the first brain region partition loss function, the second brain region partition loss function, and the third brain region partition loss function; Step 6.2.4: Finally, perform a weighted summation of the first brain region partition loss function, the second brain region partition loss function, and the third brain region partition loss function to obtain the final brain region partition prediction loss function. Step 6.2.5: Use the AdamW optimizer to set the initial learning rate and decay factor, use the cosine decay strategy to decay the learning rate, and minimize the loss function; use the early stopping strategy to determine whether to stop training based on the performance of the validation set to prevent overfitting; after training is completed, save the model parameters of the cascaded Deep-MSVM-UNet model trained in the second stage and record them as the second training model parameters.

10. The method for automatic segmentation of plain scan CT images based on a deep learning model according to claim 9, characterized in that: The step 7 specifically includes the following steps: Step 7.1, first perform data preprocessing in step 2 on the brain image data to be processed to obtain normalized brain image data to be processed; Step 7.2, then executing step 3, performing axial slicing on the pre-processed brain image data to be processed, to obtain a plurality of two-dimensional slice image data to be processed; Step 7.3: Input each slice image data to be processed into the cascaded Deep-MSVM-UNet model of the first training model parameters, and the output layer of the cascaded Deep-MSVM-UNet model of the first training model parameters outputs the predicted results of the brain partition corresponding to each slice image data to be processed; Step 7.4: Input each slice image data to be processed and the corresponding brain segmentation prediction result into the cascade Deep-MSVM-UNet model of the second training model parameters, and the output layer of the cascade Deep-MSVM-UNet model of the second training model parameters outputs the prediction result of the brain region segmentation corresponding to each slice image data to be processed; Step 7.5: axially splice the predicted results of brain region division corresponding to each slice image data to be processed to obtain a three-dimensional brain region segmentation result corresponding to the three-dimensional brain image data to be processed.

Citation Information

Cited By

  • Complex form target segmentation method based on multidimensional information guidance

    CN121236399A