Quasi-incremental multi-organ segmentation method based on joint domain representation of general and private features
By training a general characterization network in abdominal multi-organ segmentation and combining it with private features, the problem of highly specialized segmentation models in the prior art is solved, which improves the reliability and flexibility of the model and improves the segmentation performance.
Patent Information
- Application Number
- CN202310298336.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-03-24
AI Technical Summary
The prior art has the problem of highly specializing in anatomical areas and segmentation targets in abdominal multi-organ segmentation, which leads to limited application scope of the model and cannot meet the complex and changing clinical task requirements, reducing the reliability and flexibility of the model.
By training a general representation network based on self-supervised learning, the common features between the multi-anatomical regions and segmentation targets are extracted, and the private features in the downstream task class incremental multi-organ segmentation is selectively fused through the joint representation module to build a class incremental multi-organ segmentation model based on the joint domain representation of the general and private features.
It alleviates the problem of highly specialized in the segmentation model for anatomical areas and segmentation targets, improves the reliability and flexibility of the segmentation model, and improves the performance of multi-organ segmentation in the abdomen.
Smart Images

Figure CN116452515B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a quasi-incremental multi-organ segmentation method, which can be used for automatic detection and segmentation of multiple target organs in the abdominal region of a human body in a CT image. Background Art
[0002] From an anatomical point of view, the abdomen of the human body refers to the area from the diaphragm at the bottom of the chest to the pelvis, which contains many important organs of the human body, such as the liver, spleen, stomach, pancreas, kidneys, gallbladder, intestines, etc. Studies have shown that in recent years, due to the continuous deterioration of the environment and climate and the continuous change of people's lifestyles, the incidence and cancer rates of important abdominal organs have increased year by year. Therefore, accurate segmentation of multiple abdominal organs is crucial in clinical applications, including disease diagnosis, radiotherapy planning, preoperative planning, surgical navigation and other operations.
[0003] However, in the process of accurately segmenting the contours of abdominal organs, doctors need to manually outline the target organs on CT images. This process is not only time-consuming and laborious, but also the segmentation results are directly related to the doctor's anatomical knowledge and experience, resulting in suboptimal segmentation results. Therefore, it is urgent to establish a reliable computer-aided diagnosis method for automatic segmentation of multiple abdominal organs to provide a basis for imaging physicians to objectively diagnose abdominal-related diseases. With the development of computer-aided diagnosis in clinical medicine, a large number of researchers have devoted themselves to computer analysis of multiple abdominal organs using CT images.
[0004] In the existing computer-aided diagnosis of abdominal multi-organ segmentation methods, many researchers have made a lot of efforts on how to use deep learning methods to improve the accuracy of multi-organ segmentation. Gibson et al. published "Automatic Multi-Organ Segmentation on Abdominal CT with Dense V-Networks" at IEEETMI, which proposed a multi-scale DenseVNet segmentation network based on densely connected modules, and achieved accurate segmentation of 8 abdominal organs on CT images. Wang et al. published "Abdominal Multi-organ Segmentation with Organ-attention Networks and Statistical Fusion" at MIA, which proposed a two-stage reverse-connected organ attention network for automatic segmentation of abdominal organs. In the organ attention network, the network output of the first stage and the original input image are used together as the input of the second stage to reduce the interference of complex background on target segmentation and provide spatial positioning of the target area for the second stage. With the widespread application of Transformer in medical image segmentation, Chen et al. published "TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation" on CVPR, which proposed a TransUnet network that uses a hybrid encoding of CNN and Transformer. This network adds a multi-layer Transformer structure after the CNN of the encoder to process global features, and then integrates the low-level features obtained by the encoder with the high-level features obtained by the decoder, which performs well in the task of abdominal multi-organ segmentation. Tang et al. published "UNETR: Transformers for 3D Medical Image Segmentation" on IEEE, which proposed a novel network structure UNETR. This network uses pure Transformers as encoders to learn the sequence representation of the input and effectively capture global multi-scale information. At the same time, the Transformer encoder is directly connected to the decoder through jump connections of different resolutions to calculate the final semantic segmentation output. Liu et al. published "Learning Incrementally to SegmentMultiple Organs in a CT Image" on MICCAI, which proposed a category incremental learning method to make full use of the dataset with partial organ annotations to achieve accurate segmentation of abdominal multi-organs.
[0005] Although the above methods have achieved automatic segmentation of multiple abdominal organs and have achieved good results, these methods have ignored the important common characteristics between different anatomical regions and segmentation targets. The segmentation models they use are highly specialized in anatomical regions and segmentation targets. Therefore, the application scope of the model is limited and cannot meet the complex and changeable task requirements in clinical practice, thereby reducing the reliability and flexibility of the model. Summary of the invention
[0006] The purpose of the present invention is to address the deficiencies of the prior art and propose a quasi-incremental multi-organ segmentation method based on a joint domain representation of general and private features, so as to reduce the high specialization of the segmentation model on anatomical regions and segmentation targets and improve the reliability and flexibility of the segmentation model.
[0007] The technical solution to achieve the purpose of the present invention is: by training a universal representation network based on self-supervised learning, using the encoder of the network as a feature extractor to extract universal features between multiple anatomical regions and segmentation targets, and selectively fusing them with private features in the downstream task class incremental multi-organ segmentation through a joint representation module, and the implementation steps include the following:
[0008] (1) Obtaining multi-anatomical region image datasets and abdominal CT image datasets and performing preprocessing:
[0009] (1a) The CT value of the CT image is limited to the range of [-160HU, 240HU] using the window method;
[0010] (1b) Using minimum-maximum normalization, the values of all images in the multi-anatomical region image dataset and the abdominal CT image dataset are mapped to the interval [0,1];
[0011] (1c) downsampling the image of the multi-anatomical region image dataset mapped in (1b) to 192×192×192 or 192×192×96, cutting it into blocks according to a fixed size of 96×96×96, and then performing three image transformations, namely, nonlinear transformation, local pixel transformation, and random masking, on the cut images to obtain perturbed images, and the perturbed images and the cut images together form a general dataset;
[0012] (1d) All images in the abdominal CT image dataset mapped in (1b) are uniformly downsampled to 96×96×96 to obtain a segmented dataset;
[0013] (2) Divide the general data set and the segmented data set:
[0014] (2a) The general data set preprocessed in (1c) is divided into a general data training set and a general data test set in a ratio of 4:1 by using a random selection method;
[0015] (2b) The segmentation data set preprocessed in (1d) is divided into a segmentation training set and a segmentation test set in a ratio of 4:1 by a random selection method, and the segmentation training set is rotated, translated, flipped and scaled in turn to perform data augmentation to obtain an augmented segmentation training set;
[0016] (3) Construct a universal representation learning network ISRNet consisting of an encoder, a decoder, and an output head cascaded in sequence, where the encoder includes five encoding layers (E1, E2, E3, E4, E5);
[0017] (4) Based on the general data training set, the constructed ISRNet general representation learning network is iteratively trained using the back propagation method to obtain a trained general representation model;
[0018] (5) Constructing a quasi-incremental multi-organ segmentation model based on the joint domain representation of general and private features:
[0019] (5a) Establish a joint representation module DJRM consisting of a cascade of domain combination attention blocks DA and multi-scale convolution blocks DMSC;
[0020] (5b) The encoder in the universal representation model trained in step (4) is used as a feature extractor, and its five encoding layers (U1, U2, U3, U4, U5) are connected in parallel with the five encoding layers (E1, E2, E3, E4, E5) of the ISRNet encoder, and the joint representation module DJRM in (5a) is connected to the feature extractor and each encoding layer of the network encoder in turn to form a quasi-incremental multi-organ segmentation model based on the joint domain representation of universal and private features;
[0021] (6) Input the segmentation data training set into the constructed class-incremental multi-organ segmentation model, and set the segmentation loss of the model by reconstructing the background class. seg , Distillation loss Loss kd , iteratively trains it using the back-propagation method to obtain a trained incremental multi-organ segmentation model;
[0022] (7) Input the segmentation data test set into the trained class incremental multi-organ segmentation model to obtain the segmentation results of the segmentation data.
[0023] Compared with the prior art, the present invention has the following advantages:
[0024] First, the present invention trains a universal representation network based on self-supervised learning to deeply explore the important common features between different anatomical regions and segmentation targets, thereby alleviating the problem of the segmentation model being highly specialized for anatomical regions and segmentation targets.
[0025] Second, the present invention constructs a joint representation module to selectively fuse the common features between multiple anatomical regions and segmentation targets with the private features in the downstream task class incremental multi-organ segmentation, thereby improving the reliability and flexibility of the segmentation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flowchart of the implementation process of the present invention;
[0027] Figure 2 It is a framework diagram of the quasi-incremental multi-organ segmentation model constructed in the present invention;
[0028] Figure 3 for Figure 2 Schematic diagram of the joint representation module DJRM;
[0029] Figure 4 for Figure 3 Schematic diagram of the domain combination attention block DA in the joint representation module DJRM;
[0030] Figure 5 for Figure 3 Schematic diagram of the multi-scale convolution block DMSC in the joint representation module DJRM. DETAILED DESCRIPTION
[0031] The examples and effects of the present invention are further described in detail below with reference to the accompanying drawings.
[0032] See also Figure 1 The specific implementation steps of this example include the following:
[0033] Step 1: Acquire and preprocess a multi-anatomical region image dataset and an abdominal CT image dataset.
[0034] 1.1) This example obtains a multi-anatomical region image dataset from 5 different medical centers, including three anatomical regions of the chest, abdomen, and pelvis, as well as two modalities: CT and MRI;
[0035] 1.2) Obtain abdominal CT image datasets from 4 different medical centers, where the datasets from different centers contain annotated data of four different abdominal organs: liver, spleen, pancreas, and right kidney;
[0036] 1.3) The CT value of the CT image was limited to the range of [-160HU, 240HU] using the window method;
[0037] 1.4) Using minimum-maximum normalization, the values of all images in the multi-anatomical region image dataset and the abdominal CT image dataset are mapped to the interval [0,1];
[0038] 1.5) Downsampling the images of the multi-anatomical region image dataset mapped in 1.4) to 192×192×192 or 192×192×96, cutting them into blocks according to a fixed size of 96×96×96, and then sequentially performing three image transformations, namely, nonlinear transformation, local pixel transformation, and random masking, on the cut images to obtain perturbed images, and the perturbed images and the cut images together form a general dataset;
[0039] 1.6) All images in the abdominal CT image dataset mapped in 1.4) are uniformly downsampled to 96×96×96 to obtain a segmented dataset.
[0040] Step 2: Divide the general data set and the segmented data set.
[0041] 2.1) Divide the general data set preprocessed in 1.5) into a general data training set and a general data test set in a ratio of 4:1 by using a random selection method;
[0042] 2.2) The segmentation data set preprocessed in 1.6) is divided into a segmentation training set and a segmentation test set in a ratio of 4:1 by a random selection method, and the divided segmentation training set is rotated, translated, flipped and scaled in turn for data augmentation to obtain an augmented segmentation training set.
[0043] Step 3: Construct a universal representation learning network ISRNet based on self-supervised learning as the universal representation model to be trained.
[0044] The constructed ISRNet general representation learning network consists of an encoder, a decoder and an output head cascaded in sequence.
[0045] The encoder is composed of five encoding layers (E1, E2, E3, E4, E5) cascaded in sequence. The first four encoding layers (E1, E2, E3, E4) each include an Inception block, an SE channel attention block, a convolution block and a residual connection. The last encoding layer includes an Inception block, an SE channel attention block, a convolution block, a residual connection and a ViT layer, wherein:
[0046] The Inception block consists of three branches connected in parallel, each of which consists of a convolutional layer, an instance normalization layer, and a LeakyReLU activation layer. The first branch contains a convolutional layer with a kernel size of 1×1×1 and a stride of 1; the second branch contains a convolutional layer with a kernel size of 1×1×1, a stride of 1, and a convolutional layer with a kernel size of 3×3×3, a stride of 1, and a padding of 1; the third branch contains a convolutional layer with a kernel size of 1×1×1, a stride of 1, and two convolutional layers with a kernel size of 3×3×3, a stride of 1, and a padding of 1;
[0047] The SE channel attention block consists of an average pooling layer, a linear layer, a leaky relu activation layer, a linear layer and a sigmoid activation layer cascaded in sequence;
[0048] The convolution block consists of a convolution layer with a kernel size of 1×1×1 and a stride of 1, an instance normalization layer, and a LeakyReLU activation layer cascaded in sequence;
[0049] The ViT layer consists of 12×TransformerLayer layers, a Reshape layer, and a convolutional layer cascaded in sequence;
[0050] The decoder is composed of four decoding layers (D1, D2, D3, D4) cascaded in sequence, each decoding layer comprising an Inception block, an SE channel attention block, a convolution block and a residual connection;
[0051] The output head consists of a convolution layer with a convolution kernel size of 1×1×1 and a stride of 1.
[0052] Step 4: Based on the general data training set, the constructed ISRNet general representation learning network is iteratively trained using the back propagation method to obtain a trained general representation model.
[0053] 4.1) Set the total number of iterations to 200;
[0054] 4.2) Set the mean absolute error loss as the loss function of the network l1 (y,y'), which is expressed as follows:
[0055]
[0056] Where y represents the cut image, y' represents the predicted image, and N represents the total number of images;
[0057] 4.3) Input the perturbed image in the general data training set into the constructed ISRNet general representation learning network to obtain the predicted image;
[0058] 4.4) Substitute the cut-out image and predicted image in the general data training set into the loss function Loss l1 In (y,y'), the loss value of the network is calculated;
[0059] 4.5) Back-propagate the loss value obtained in 4.4) to update the network parameters and obtain the model W after preliminary training;
[0060] 4.6) The process of 4.3) to 4.5) is repeated for the model W after preliminary training until the network loss value does not decrease within 30 iterations or the total number of iterations reaches 200. Then the training is stopped to obtain the trained general representation model W'.
[0061] Step 5: Build a class-incremental multi-organ segmentation model based on the joint domain representation of general and private features.
[0062] Reference Figure 2 , the specific implementation of this step is as follows:
[0063] 5.1) Establish a joint representation module DJRM, such as Figure 3 As shown:
[0064] 5.1.1) Establish a domain combination attention block DA consisting of a domain attention branch and a general domain representation branch, such as Figure 4 As shown;
[0065] The domain attention branch consists of a global average pooling layer, a fully connected layer, and a Softmax layer in sequence;
[0066] The general domain representation branch consists of 8 parallel SEAdapter adapters, each of which consists of a global average pooling layer, a fully connected layer, a LeakyReLU activation layer, and a fully connected layer;
[0067] 5.1.2) Establish a multi-scale convolution block DMSC consisting of four convolution blocks of different scales connected in series. Each convolution block consists of a convolution layer, an instance normalization layer, and a LeakyReLU activation layer, such as Figure 5 As shown, where:
[0068] The convolution kernel size of the first convolution layer is 5×5×5, the stride is 1, and the padding is 2;
[0069] The convolution kernel size of the second convolution layer is 3×3×3, the stride is 1, and the padding is 1;
[0070] The convolution kernel size of the third and fourth convolutional layers is 1×1×1, and the stride is 1;
[0071] 5.1.3) The domain combination attention block DA and the multi-scale convolution block DMSC are cascaded to form a joint representation module DJRM;
[0072] 5.2) The encoder in the general representation model trained in step 3 is used as the feature extractor, and its five encoding layers (U1, U2, U3, U4, U5) are connected in parallel with the five encoding layers (E1, E2, E3, E4, E5) of the ISRNet encoder, and then the joint representation module DJRM in 5.1) is connected to each encoding layer of the feature extractor and the network encoder in turn to form a quasi-incremental multi-organ segmentation model based on the joint domain representation of general and private features. Its structure is shown as follows:
[0073]
[0074]
[0075] Among them, Decoder represents the ISRNet decoder, and OutputHead represents the ISRNet output head.
[0076] 5.3) Set the segmentation loss of the model by reconstructing the background class seg , Distillation loss Loss kd :
[0077] 5.3.1) The organ region to be segmented in the (1,...,t-2,t-1) stage is set as the old category, the organ region to be segmented in the current t stage is set as the new category, and the background region to be segmented in the current t stage is set as the background category;
[0078] 5.3.2) Calculate the model prediction image q at stage t t The old category and background category probability merge result Then the t stage q t The new category and background category probabilities are combined to obtain the combined result
[0079]
[0080]
[0081] Among them, the subscript i represents the index of the pixel in the image, the subscript c represents the category to which the pixel at the corresponding index position belongs, and y t-1 represents the old category corresponding to the (1,...,t-2,t-1) stage, c trepresents the new category corresponding to the current stage t, b represents the background category corresponding to the current stage t, and y t =y t-1 ∪c t represents the old or new category corresponding to the (1,...,t-1,t) stage;
[0082] 5.3.3) According to q t The old category and background category probability merge result Set segmentation loss Loss seg ;
[0083]
[0084] Among them, g t represents the label image at stage t, ε represents the smoothing coefficient;
[0085] 5.3.4) According to q t The new category and background category probability merging result Set distillation loss Loss kd ;
[0086]
[0087] Among them, q t-1 represents the pseudo-label image predicted by the model in stage t-1, and I represents the total number of image pixels;
[0088] 5.3.5) The segmentation loss Loss seg and distillation loss Loss kd Add together to form the loss function of the incremental multi-organ segmentation model.
[0089] Step 6: Use the back-propagation method to iteratively train the incremental multi-organ segmentation model.
[0090] 6.1) Set the total number of iterations to 200, the batch size to 1, the initial learning rate to 0.01, warm up the learning rate from 0 to 0.01 in the first 10 rounds, and adjust the learning rate using the cosine annealing decay strategy for the remaining 190 rounds. The optimizer is the Adam optimizer;
[0091] 6.2) Set the joint loss function to: UnionLoss = Loss seg +Loss kd ;
[0092] 6.3) Input the segmentation training set into the incremental multi-organ segmentation model in batches according to the batch size set in 6.1) to obtain the pseudo-label image of the previous stage t-1 and the predicted image of the current stage t;
[0093] 6.4) According to the pseudo-label image, predicted image and label image in the segmentation training set obtained in 6.3), the loss value of the model is calculated by the joint loss function;
[0094] 6.5) Back-propagate the loss value obtained in 6.4) to update the model parameters and obtain the preliminarily trained model M;
[0095] 6.6) Repeat the process of 6.3) to 6.5) for the initially trained model M until it is detected that the network loss value does not decrease within 30 iterations or the total number of iterations reaches 200, then stop the training and obtain the trained segmentation model M'.
[0096] Step 7: perform incremental multi-organ segmentation on the segmentation dataset.
[0097] The segmentation test set is input into the trained class-incremental multi-organ segmentation model M' to obtain the segmentation result of the test set.
[0098] 1. Simulation conditions:
[0099] The simulation experiment platform of the present invention is a Linux operating system, configured as follows CoreTMi7-12700KFCPU and NvidiaRTX3090GPU, using PyTorch deep learning framework, and the development language is Python.
[0100] The simulation experiment data of the present invention includes a multi-anatomical region image dataset and an abdominal CT image dataset, wherein the multi-anatomical region image dataset comes from an unlabeled dataset from 5 different medical centers, including three anatomical regions of the chest, abdomen, and pelvis, and two modalities of CT and MRI; the abdominal CT image dataset comes from a labeled dataset from 4 different medical centers, and the organ labels of the datasets from different centers are all labeled by professional doctors.
[0101] 2. Simulation content and result analysis:
[0102] Simulation 1, using the existing incremental multi-organ segmentation method and the method of the present invention to perform abdominal multi-organ segmentation on the above abdominal CT image data set, the results are shown in Table 1:
[0103] Table 1 Comparison of Dice (%) indicators of different methods for abdominal multi-organ segmentation
[0104]
[0105] In Table 1, Liver represents the Dice index of the liver, Spleen represents the Dice index of the spleen, Pancreas represents the Dice index of the pancreas, Rightkidney represents the Dice index of the right kidney, and Average represents the average Dice index of the four organs. The higher the Dice index, the better the organ segmentation effect.
[0106] As can be seen from Table 1, the Dice index of the segmentation of the four abdominal organs by the method of the present invention is improved compared with the existing methods, which proves that the quasi-incremental multi-organ segmentation method based on the joint domain representation of general and private features of the present invention can improve the performance of abdominal multi-organ segmentation.
[0107] Simulation 2, using the existing 3DTransUnet network and the network of the present invention to perform segmentation experiments on eight abdominal organs on the public BTCV dataset, the results are shown in Table 2:
[0108] Table 2 Comparison of Dice (%) index of joint representation module DJRM for abdominal multi-organ segmentation
[0109]
[0110] As can be seen from Table 2, after adding the DJRM module, the Dice index of spleen (Spleen), right kidney (Rightkidney), left kidney (Leftkidney), gallbladder (Gallbladder), pancreas (Pancreas), liver (Liver), and aorta (Aorta) are all improved, and the average index of the eight organs is increased by 2.58%, which proves the effectiveness of the joint characterization module DJRM proposed in the present invention.
[0111] In summary, in order to address the problem that existing methods cause high specialization in anatomical regions and segmentation targets when segmenting multiple abdominal organs, the joint representation module DJRM constructed in the present invention can selectively fuse the common features between multiple anatomical regions and segmentation targets with the private features in the downstream task-class incremental multi-organ segmentation, thereby improving the reliability and flexibility of the segmentation model.
Claims
1. A quasi-incremental multi-organ segmentation method based on joint domain representation of general and private features, characterized in that: include: (1) Obtaining multi-anatomical region image datasets and abdominal CT image datasets and performing preprocessing: (1a) The CT value of the CT image is limited to the range of [-160HU, 240HU] using the window method; (1b) Using minimum-maximum normalization, the values of all images in the multi-anatomical region image dataset and the abdominal CT image dataset are mapped to the interval [0,1]; (1c) downsampling the image of the multi-anatomical region image dataset mapped in (1b) to 192×192×192 or 192×192×96, cutting it into blocks according to a fixed size of 96×96×96, and then performing three image transformations, namely, nonlinear transformation, local pixel transformation, and random masking, on the cut images to obtain perturbed images, and the perturbed images and the cut images together form a general dataset; (1d) All images in the abdominal CT image dataset mapped in (1b) are uniformly downsampled to 96×96×96 to obtain a segmented dataset; (2) Divide the general data set and the segmented data set: (2a) The general data set preprocessed in (1c) is divided into a general data training set and a general data test set in a ratio of 4:1 by using a random selection method; (2b) The segmentation data set preprocessed in (1d) is divided into a segmentation training set and a segmentation test set in a ratio of 4:1 by a random selection method, and the segmentation training set is rotated, translated, flipped and scaled in turn to perform data augmentation to obtain an augmented segmentation training set; (3) Construct a universal representation learning network ISRNet consisting of an encoder, a decoder, and an output head cascaded in sequence, where the encoder includes five encoding layers (E1, E2, E3, E4, E5); (4) Based on the general data training set, the constructed ISRNet general representation learning network is iteratively trained using the back propagation method to obtain a trained general representation model; (5) Constructing a quasi-incremental multi-organ segmentation model based on the joint domain representation of general and private features: (5a) Establish a joint representation module DJRM consisting of a cascade of domain combination attention blocks DA and multi-scale convolution blocks DMSC; (5b) The encoder in the universal representation model trained in step (4) is used as a feature extractor, and its five encoding layers (U1, U2, U3, U4, U5) are connected in parallel with the five encoding layers (E1, E2, E3, E4, E5) of the ISRNet encoder, and the joint representation module DJRM in (5a) is connected to the feature extractor and each encoding layer of the network encoder in turn to form a quasi-incremental multi-organ segmentation model based on the joint domain representation of universal and private features; (6) Input the segmentation data training set into the constructed class-incremental multi-organ segmentation model, and set the segmentation loss of the model by reconstructing the background class. seg , Distillation loss Loss kd , iteratively trains it using the back-propagation method to obtain a trained incremental multi-organ segmentation model; (7) Input the segmentation data test set into the trained class incremental multi-organ segmentation model to obtain the segmentation results of the segmentation data.
2. The method according to claim 1, characterized in that The encoder, decoder and output head of the ISRNet general representation learning network in step (3) have the following structural parameters: The five encoding layers (E1, E2, E3, E4, E5) in the encoder are cascaded in sequence, and the first four encoding layers (E1, E2, E3, E4) each include an Inception block, an SE channel attention block, a convolution block and a residual connection, and the last encoding layer includes an Inception block, an SE channel attention block, a convolution block, a residual connection and a ViT layer. The Inception block consists of three branches connected in parallel, each of which consists of a convolutional layer, an instance normalization layer, and a LeakyReLU activation layer. The first branch contains a convolutional layer with a kernel size of 1×1×1 and a stride of 1; the second branch contains a convolutional layer with a kernel size of 1×1×1, a stride of 1, and a convolutional layer with a kernel size of 3×3×3, a stride of 1, and a padding of 1; the third branch contains a convolutional layer with a kernel size of 1×1×1, a stride of 1, and two convolutional layers with a kernel size of 3×3×3, a stride of 1, and a padding of 1; The SE channel attention block consists of an average pooling layer, a linear layer, a leaky relu activation layer, a linear layer and a sigmoid activation layer cascaded in sequence; The convolution block consists of a convolution layer with a kernel size of 1×1×1 and a stride of 1, an instance normalization layer, and a LeakyReLU activation layer cascaded in sequence; The ViT layer consists of 12×Transformer Layer, a Reshape layer and a convolutional layer cascaded in sequence; The decoder is composed of four decoding layers (D1, D2, D3, D4) cascaded in sequence, each decoding layer comprising an Inception block, an SE channel attention block, a convolution block and a residual connection; The output head consists of a convolution layer with a convolution kernel size of 1×1×1 and a stride of 1.
3. The method according to claim 1, characterized in that In (4), based on the general data training set, the back propagation method is used to iteratively train the constructed ISRNet general representation learning network, which is implemented as follows: (4a) Set the total number of iterations to 200; (4b) Set the mean absolute error loss as the loss function of the network l1 (L,y'), which is expressed as follows: Where y represents the cut image, y' represents the predicted image, and N represents the total number of images; (4c) Input the perturbed image in the general data training set into the constructed ISRNet general representation learning network to obtain the predicted image; (4d) Substitute the cut-out image and the predicted image in the general data training set into the loss function Loss l1 In (y,y'), the loss value of the network is calculated; (4e) Back-propagate the loss value obtained in (4d), update the network parameters, and obtain the model W after preliminary training; (4f) The process (4c) to (4e) is repeated for the initially trained model W until the network loss value does not decrease within 30 iterations or the total number of iterations reaches 200. The training is then stopped to obtain the trained general representation model W'.
4. The method according to claim 1, characterized in that: The structures of the domain combination attention block DA and the multi-scale convolution block DMSC constituting the joint representation module DJRM in (5a) are as follows: The domain combination attention block DA is composed of a domain attention branch connected in parallel with a general domain representation branch, wherein the domain attention branch is composed of a global average pooling layer, a fully connected layer, and a Softmax layer in sequence; the general domain representation branch is composed of 8 parallel SE Adapter adapters, each of which is composed of a global average pooling layer, a fully connected layer, a LeakyReLU activation layer, and a fully connected layer; The multi-scale convolution block DMSC is composed of four convolution blocks of different scales connected in series in sequence, each convolution block is composed of a convolution layer, an instance normalization layer and a LeakyReLU activation layer, wherein: the convolution kernel size of the first convolution layer is 5×5×5, the step size is 1, and the padding is 2; the convolution kernel size of the second convolution layer is 3×3×3, the step size is 1, and the padding is 1; the convolution kernel size of the third and fourth convolution layers is 1×1×1, and the step size is 1.
5. The method according to claim 1, characterized in that The (5b) constructs a class incremental multi-organ segmentation model based on the joint domain representation of general and private features, and the structure is represented as follows: Among them, Decoder represents the ISRNet decoder, and OutputHead represents the ISRNet output head.
6. The method according to claim 1, characterized in that In step (6), the segmentation loss Loss of the model is set by reconstructing the background class. seg , Distillation loss Loss kd , implemented as follows: (6a) The organ region to be segmented in the (1, ..., t-2, t-1) stage is set as the old category, the organ region to be segmented in the current stage t is set as the new category, and the background region to be segmented in the current stage t is set as the background category; (6b) Calculate the model prediction image q at stage t t The old category and background category probability merge result t phase q t The new category and background category probability merging result Among them, the subscript i represents the index of the pixel in the image, the subscript c represents the category to which the pixel at the corresponding index position belongs, and y t-1 represents the old category corresponding to the (1,...,t-2,t-1) stage, c t represents the new category corresponding to the current stage t, b represents the background category corresponding to the current stage t, and y t =y t-1 ∪c t represents the old or new category corresponding to the (1,...,t-1,t) stage; (6c) Set segmentation loss Loss seg ; Among them, g t represents the label image at stage t, ε represents the smoothing coefficient; (6d) Set distillation loss Loss kd ; Among them, q t-1 represents the pseudo-label image predicted by the model in stage t-1, and I represents the total number of image pixels.
7. The method according to claim 1, characterized in that In step (6), the incremental multi-organ segmentation model is iteratively trained using the back propagation method, which is implemented as follows: (6a) Set the total number of iterations to 200, the batch size to 1, the initial learning rate to 0.01, warm up the learning rate from 0 to 0.01 in the first 10 rounds, and adjust the learning rate using the cosine annealing decay strategy for the remaining 190 rounds. The optimizer is the Adam optimizer; (6b) Set the joint loss function to: UnionLoss = Loss seg +Loss kd ; (6c) inputting the segmentation training set into the incremental multi-organ segmentation model in batches according to the batch size set in (6a) to obtain the pseudo-label image of the previous stage t-1 and the predicted image of the current stage t; (6d) According to the pseudo-label image, predicted image and label image in the segmentation training set obtained in (6c), the loss value of the model is calculated by the joint loss function; (6e) Back-propagate the loss value obtained in (6d), update the model parameters, and obtain the model M after preliminary training; (6f) The process (6c) to (6e) is repeated for the initially trained model M until the network loss value does not decrease within 30 iterations or the total number of iterations reaches 200. The training is then stopped to obtain the trained segmentation model M'.
Citation Information
Patent Citations
Video processing method and device
CN114596312A
Multi-temporal liver tumor segmentation method based on multi-head cross attention conversion network
CN115330816A