A multi-organ segmentation method for abdominal CT images based on dual-scale hybrid network

By adopting a dual-scale hybrid network in multi-organ segmentation of abdominal CT images, combining the advantages of Transformer and CNN, multi-scale features are extracted and encoder features of different scales are fused, and the problems of low efficiency and poor repeatability in the prior art are solved, and more efficient and accurate multi-organ segmentation is achieved.

CN117994521BActive Publication Date: 2025-05-23HUNAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410223412.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-05-23
Estimated Expiration
2044-02-28

AI Technical Summary

Technical Problem

The prior art has problems of low efficiency and poor repeatability in multi-organ segmentation of abdominal CT images, especially when dealing with complex abdominal areas, it is difficult to accurately segment the organ boundaries.

Method used

Using a method based on a dual-scale hybrid network, the advantages of Transformer and CNN are combined with the advantages of Transformer and CNN, multi-scale features are extracted, and encoder features of different scales are gradually fused through a cascade decoder to obtain accurate segmentation results.

Benefits of technology

It improves the accuracy and efficiency of multi-organ segmentation of abdominal CT images, can better handle complex abdominal areas, and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117994521B_ABST
    Figure CN117994521B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-organ segmentation method for abdominal CT images based on a dual-scale hybrid network, which is specifically implemented as follows: (1) establishing a training data set containing the original image and its corresponding segmentation gold standard; (2) constructing a dual-scale hybrid encoding and decoding network, wherein the encoder uses dual-scale input, extracts multi-scale features of the image by making full use of the advantages of Transformer and CNN, and the decoder obtains accurate segmentation results by gradually fusing features of different scales; (4) using the training data set to train the network until a pre-set loss function converges; (5) using the trained network to test the image to be segmented to obtain the segmentation result. The present invention can fully extract local and global information of the image at different levels of the encoding end, can adapt to abdominal organs with diverse morphology and complex structure, and obtain accurate segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to a multi-organ segmentation method of abdominal CT images based on a dual-scale hybrid network. Background Art

[0002] Multi-organ segmentation of abdominal CT images plays an important role in medical image processing and analysis, and can provide important information and quantitative evaluation for surgical navigation, organ transplantation, and radiotherapy. Compared with single-organ segmentation, multi-organ segmentation is a more challenging task, which requires distinguishing whether each pixel in the image belongs to an organ and identifying which organ it belongs to. In current clinical practice, multi-organ segmentation is mainly completed by doctors manually outlining. Due to the huge number of CT slices of patients, manually outlining each organ on each CT slice is cumbersome, inefficient, and has poor repeatability. In addition, the abdomen is one of the most complex areas of the human body. Abdominal organs usually have complex structures, fuzzy boundaries, and diverse morphologies. Manual outlining is highly subjective, and its accuracy depends heavily on the experience and skills of clinicians. In recent years, high-precision automatic segmentation methods for abdominal multi-organs have attracted more and more attention.

[0003] Convolutional neural networks (CNNs) perform well in image feature extraction. Although CNN-based models, such as fully convolutional networks (FCNs) and U-Nets, have achieved considerable success, they are limited by local receptive fields and inductive biases. These CNN-based methods have difficulty in establishing dependencies between long-distance targets in images, and their segmentation performance still cannot meet clinical requirements.

[0004] In order to overcome the limitations of CNN in modeling global semantic features, Dosovitskiy et al. proposed a Transformer structure based on a multi-head self-attention mechanism (“An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020). The multi-head self-attention mechanism can learn the intrinsic connections between all input sequences to capture the long-distance dependencies of images. However, the Transformer-based network has limited ability to model and pay attention to local patterns such as boundaries and small objects. In addition, the Transformer needs to divide the input image into blocks and use the image sub-blocks as input sequences to establish long-distance dependencies. Existing methods often use single-scale input, and there are problems such as information loss and incomplete feature representation at the boundaries of image sub-blocks. Summary of the invention

[0005] In view of the shortcomings and deficiencies of the prior art, the present invention provides a multi-organ segmentation method for abdominal CT images based on a dual-scale hybrid network. The method follows the network framework of encoding and decoding, and specifically includes a dual-scale hybrid encoder and a cascade decoder. The encoder uses dual-scale input, fully utilizes the advantages of Transformer and CNN, and extracts multi-scale features. The cascade decoder will obtain accurate segmentation results by gradually fusing encoder features of different scales.

[0006] A multi-organ segmentation method for abdominal CT images based on a dual-scale hybrid network comprises the following steps:

[0007] A multi-organ segmentation method for abdominal CT images based on a dual-scale hybrid network, the specific implementation steps are as follows:

[0008] (1) Establish a training dataset A containing original images and their corresponding segmentation gold standards;

[0009] (2) Construct a dual-scale hybrid codec network, called DS-Net, which includes:

[0010] (2-a) The input image is divided into large-scale and small-scale blocks, specifically including: for an input image of size W×H, it is divided into Non-overlapping large-scale sub-blocks S L and Non-overlapping small-scale sub-blocks S S ; The r is preferably a natural number between 8 and 32;

[0011] (2-b) Construct a codec structure with skip connections as the basic network framework. The encoder contains four encoding blocks. The first, second and third encoding blocks have the same structure, each containing a two-scale Transformer block with downsampling, two local feature enhancement modules and a two-scale fusion module; the fourth encoding block contains a two-scale Transformer block, two local feature enhancement modules and a two-scale fusion module; the decoder contains four decoding blocks and a linear prediction head; the first decoding block contains an upsampling operation with a step size of s; the second, third and fourth decoding blocks have the same structure, each consisting of a concatenation operation, a 1×1 convolution and an upsampling operation with a step size of s in sequence. The output of the two-scale fusion module in the fourth coding block is used as the input of the upsampling operation in the first decoding block; the upsampling output in the first decoding block and the output of the two-scale fusion module in the third coding block are both used as the input of the splicing operation in the second decoding block; the upsampling output in the second decoding block and the output of the two-scale fusion module in the second coding block are both used as the input of the splicing operation in the third decoding block; the upsampling output in the third decoding block and the output of the two-scale fusion module in the first coding block are both used as the input of the splicing operation in the fourth decoding block; the upsampling output in the fourth decoding block is used as the input of the linear prediction head, and the output of the linear prediction head is the multi-organ segmentation result; the s is preferably an even number between 2 and 6;

[0012] (2-c) The dual-scale Transformer block with downsampling described in step (2-b), denoted as Dwon-DS-Trans block, comprises two inputs and four outputs, and the specific operation comprises: inputting the two inputs into two parallel independent branches respectively, each branch is composed of N CSWinTransformer blocks and a downsampling operation with a step length of s connected in sequence, taking the output of the Nth CSWin Transformer block in each branch and the output of the downsampling operation as the output of the dual-scale Transformer block with downsampling; the N is preferably a natural number between 3 and 8;

[0013] (2-d) The dual-scale Transformer block described in step (2-b) comprises two inputs and two outputs, and the specific operation comprises: inputting the two inputs into two parallel independent branches respectively, each branch is composed of N CSWin Transformer blocks connected in sequence, and taking the output of the Nth CSWin Transformer block in each branch as the output of the dual-scale Transformer block;

[0014] (2-e) The first, second and third coding blocks described in step (2-b) have a specific structure including: the output of the Nth CSWin Transformer block of each branch in the Dwon-DS-Trans block is respectively connected to a local feature enhancement module, and the output of the local feature enhancement module is used as the input of the two-scale fusion module, and the output of the two-scale fusion module is used as the output of the current coding block; the non-overlapping large-scale sub-blocks S in step (2-b) L and small-scale sub-block S S As the input of the first encoding block, they are respectively input into the two parallel independent branches in the Dwon-DS-Trans block; the output of the downsampling operation of the Dwon-DS-Trans block in the first encoding block is used as the input of the second encoding block, and is respectively input into the two parallel independent branches of the Dwon-DS-Trans block in the second encoding block; the output of the downsampling operation of the Dwon-DS-Trans block in the second encoding block is used as the input of the third encoding block, and is respectively input into the two parallel independent branches of the Dwon-DS-Trans block in the third encoding block; the output of the downsampling operation of the Dwon-DS-Trans block in the third encoding block is used as the input of the fourth encoding block, and is respectively input into the two parallel independent branches of the dual-scale Transformer block in the fourth encoding block;

[0015] (2-f) The fourth coding block described in step (2-b) has a specific structure comprising: the output of the Nth CSWin Transformer block of each branch in the dual-scale Transformer block is respectively connected to a local feature enhancement module, and the output of the local feature enhancement module is used as the input of the dual-scale fusion module, and the output of the dual-scale fusion module is used as the output of the fourth coding block;

[0016] (2-g) The local feature enhancement module described in step (2-b) has a specific structure comprising: for feature I, first inputting I into three independent feature extraction branches, wherein the first branch is composed of 3×3 convolution, deep convolution and GeLU activation layer connected in sequence, the second branch is composed of 3×3 convolution, GeLU activation layer, 3×3 convolution, GeLU activation layer connected in sequence, the third branch is composed of deep convolution, GeLU activation layer, deep convolution, GeLU activation layer connected in sequence, and then the outputs O of the three branches are connected. 1 , O 2 and O 3 Splicing is performed, and 3×3 convolution and layer normalization are performed on the splicing results in sequence to obtain the output of the local feature enhancement module;

[0017] (2-h) The dual-scale fusion module described in step (2-b) has two inputs and one output. The specific structure includes: for the smaller-sized input I S , first, an upsampling operation is used to restore it to the same size as the larger-sized input I L . Then, the upsampled result is concatenated with the larger-sized input I L . Then, the concatenated result is subjected to 3×3 convolution to obtain the output of the dual-scale fusion module;

[0018] (2-i) The linear prediction head described in step (2-b) is composed of a 1×1 convolution and a Softmax layer connected in sequence;

[0019] (3) Use the training dataset A to train the DS-Net until the pre-set loss function converges; preferably, a hybrid loss function combining Dice and cross-entropy is used as the objective function for network training;

[0020] (4) Use the trained network to test the image to be segmented to obtain the segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic diagram of a dual-scale hybrid encoding and decoding network structure in an embodiment of the present invention

[0022] Figure 2 is a schematic diagram of a dual-scale Transformer block structure with downsampling in an embodiment of the present invention

[0023] Figure 3 is a schematic diagram of a dual-scale Transformer block structure in an embodiment of the present invention

[0024] Figure 4 is a schematic diagram of a local feature enhancement module structure in an embodiment of the present invention

[0025] Figure 5 is a schematic diagram of a dual-scale fusion module structure in an embodiment of the present invention

[0026] Figure 6 is an example of the segmentation result in an embodiment of the present invention, Figure 6 (a) to Figure 6 (d) are four original abdominal CT images randomly selected from the FLARE dataset, Figure 6 (e) to Figure 6 (h) are the results of testing Figure 6 (a) to Figure 6 (d) using the method in Embodiment 1 of the present invention DETAILED DESCRIPTION OF THE INVENTION

[0027] Embodiment 1

[0028] A multi-organ segmentation method for abdominal CT images based on a dual-scale hybrid network is specifically implemented as follows: (1) establishing a training dataset A containing original images and their corresponding segmentation gold standards;

[0029] (2) Construct a dual-scale hybrid codec network, called DS-Net, with the following structure: Figure 1 As shown, specifically including:

[0030] (2-a) The input image is divided into large-scale and small-scale blocks, specifically including: for an input image of size W×H, it is divided into Non-overlapping large-scale sub-blocks S L and Non-overlapping small-scale sub-blocks S S ; In this embodiment, r=16 is preferred

[0031] (2-b) Construct a codec structure with skip connections as the basic network framework. The encoder contains four encoding blocks. The first, second and third encoding blocks have the same structure, each containing a two-scale Transformer block with downsampling, two local feature enhancement modules and a two-scale fusion module; the fourth encoding block contains a two-scale Transformer block, two local feature enhancement modules and a two-scale fusion module; the decoder contains four decoding blocks and a linear prediction head; the first decoding block contains an upsampling operation with a step size of s; the second, third and fourth decoding blocks have the same structure, each consisting of a concatenation operation, a 1×1 convolution and an upsampling operation with a step size of s The operations are connected in sequence; the output of the two-scale fusion module in the fourth coding block is used as the input of the upsampling operation in the first decoding block; the upsampling output in the first decoding block and the output of the two-scale fusion module in the third coding block are both used as the input of the splicing operation in the second decoding block; the upsampling output in the second decoding block and the output of the two-scale fusion module in the second coding block are both used as the input of the splicing operation in the third decoding block; the upsampling output in the third decoding block and the output of the two-scale fusion module in the first coding block are both used as the input of the splicing operation in the fourth decoding block; the upsampling output in the fourth decoding block is used as the input of the linear prediction head, and the output of the linear prediction head is the multi-organ segmentation result; this embodiment is preferred

[0032] s = 2;

[0033] (2-c) The dual-scale Transformer block with downsampling described in step (2-b) is denoted as Dwon-DS-Trans block, and its specific structure is as follows Figure 2As shown, it includes two inputs and four outputs. The specific operation includes: inputting the two inputs into two parallel independent branches respectively, each branch is composed of N CSWin Transformer blocks and a downsampling operation with a step size of s connected in sequence, taking the output of the Nth CSWin Transformer block in each branch and the output of the downsampling operation as the output of the dual-scale Transformer block with downsampling; in this embodiment, N=6 is preferred; (2-d) The dual-scale Transformer block described in step (2-b) has a specific structure as shown Figure 3 As shown, it contains two

[0034] The specific operation includes: inputting the two inputs into two parallel independent branches respectively, each branch is composed of N CSWin Transformer blocks connected in sequence, and taking the output of the Nth CSWin Transformer block in each branch as the output of the dual-scale Transformer block;

[0035] (2-e) The first, second and third coding blocks described in step (2-b) have specific structures including:

[0036] The output of the Nth CSWin Transformer block of each branch in the Dwon-DS-Trans block is connected to a local feature enhancement module, and the output of the local feature enhancement module is used as the input of the two-scale fusion module, and the output of the two-scale fusion module is used as the output of the current coding block; the non-overlapping large-scale sub-block S in step (2-b) L and small-scale sub-block S S As the input of the first encoding block, they are respectively input into the two parallel independent branches in the Dwon-DS-Trans block; the output of the downsampling operation of the Dwon-DS-Trans block in the first encoding block is used as the input of the second encoding block, and is respectively input into the two parallel independent branches of the Dwon-DS-Trans block in the second encoding block; the output of the downsampling operation of the Dwon-DS-Trans block in the second encoding block is used as the input of the third encoding block, and is respectively input into the two parallel independent branches of the Dwon-DS-Trans block in the third encoding block; the output of the downsampling operation of the Dwon-DS-Trans block in the third encoding block is used as the input of the fourth encoding block, and is respectively input into the dual-scale

[0037] Two parallel independent branches of the Transformer block;

[0038] (2-f) The fourth coding block described in step (2-b) includes: a dual-scale Transformer block

[0039] The output of the Nth CSWin Transformer block in each branch is connected to a local feature enhancement module, and the output of the local feature enhancement module is used as the input of the dual-scale fusion module.

[0040] The output of the dual-scale fusion module is used as the output of the fourth encoding block;

[0041] (2-g) The local feature enhancement module described in step (2-b) has a structure as follows: Figure 4 As shown in FIG. 1 , the specific structure includes: for feature I, I is first input into three independent feature extraction branches, wherein the first branch is composed of 3×3 convolution, deep convolution and GeLU activation layer connected in sequence, the second branch is composed of 3×3 convolution, GeLU activation layer, 3×3 convolution, GeLU activation layer connected in sequence, the third branch is composed of deep convolution, GeLU activation layer, deep convolution, GeLU activation layer connected in sequence, and then the outputs O of the three branches are connected. 1 , O 2 and O 3 Splicing is performed, and 3×3 convolution and layer normalization are performed on the splicing results in sequence to obtain the output of the local feature enhancement module;

[0042] (2-h) The dual-scale fusion module described in step (2-b) has the following structure: Figure 5 As shown, it contains two inputs and one output. The specific structure includes: for the smaller input I S , first use an upsampling operation to restore it to the same size as the larger input I L The size is the same, and then the upsampled result is combined with the larger input I L The splicing is performed, and then the splicing result is convolved 3×3 to obtain the output of the dual-scale fusion module;

[0043] (2-i) The linear prediction head described in step (2-b) is composed of a 1×1 convolution and a Softmax layer connected in sequence;

[0044] (3) Using training data set A to train DS-Net until the preset loss function converges; preferably, a hybrid loss function combining Dice and cross entropy is used as the objective function of network training;

[0045] (4) Use the trained network to test the image to be segmented and obtain the segmentation result.

[0046] Example 2

[0047] The method in Example 1 was used to experiment on the FLARE public dataset. The FLARE dataset contains 361 abdominal CT sequences from 11 medical centers. The number of slices contained in a single sequence is 80 to 281, and the slice spatial resolution is 512×512 pixels. We randomly selected 289 sequences for training and 72 sequences for testing.

[0048] Figure 6 Some 2D slice segmentation results on the test set are shown. Figure 6 (a)~ Figure 6 (d) Four original CT images randomly selected from the test data. Figure 6 (e)~ Figure 6 (h) is to adopt the method in Example 1 to Figure 6 (a)~ Figure 6 (d) The test results show that the liver, kidneys, spleen, pancreas, stomach and other organs are effectively segmented.

Claims

1. A multi-organ segmentation method for abdominal CT images based on a dual-scale hybrid network, characterized in that: The following steps are involved: (1) Establish a training dataset A containing original images and their corresponding segmentation gold standards; (2) Construct a dual-scale hybrid codec network, called DS-Net, which includes: (2-a) The input image is divided into large-scale and small-scale blocks, specifically including: for an input image of size W×H, it is divided into Non-overlapping large-scale sub-blocks S L and Non-overlapping small-scale sub-blocks S S ; (2-b) A codec structure with skip connections is constructed as the basic network framework. The encoder contains four encoding blocks. The first, second and third encoding blocks have the same structure, each containing a two-scale Transformer block with downsampling, two local feature enhancement modules and a two-scale fusion module; the fourth encoding block contains a two-scale Transformer block, two local feature enhancement modules and a two-scale fusion module; the decoder contains four decoding blocks and a linear prediction head; the first decoding block contains an upsampling operation with a step size of s; the second, third and fourth decoding blocks have the same structure, each consisting of a concatenation operation, a 1×1 convolution and a step size of s. The upsampling operations are connected in sequence; the output of the two-scale fusion module in the fourth coding block is used as the input of the upsampling operation in the first decoding block; the output of the upsampling in the first decoding block and the output of the two-scale fusion module in the third coding block are both used as the input of the splicing operation in the second decoding block; the output of the upsampling in the second decoding block and the output of the two-scale fusion module in the second coding block are both used as the input of the splicing operation in the third decoding block; the output of the upsampling in the third decoding block and the output of the two-scale fusion module in the first coding block are both used as the input of the splicing operation in the fourth decoding block; the output of the upsampling in the fourth decoding block is used as the input of the linear prediction head, and the output of the linear prediction head is the multi-organ segmentation result; (2-c) The dual-scale Transformer block with downsampling described in step (2-b), denoted as Dwon-DS-Trans block, comprises two inputs and four outputs, and the specific operations include: inputting the two inputs into two parallel independent branches respectively, each branch is composed of N CSWinTransformer blocks and a downsampling operation with a step size of s connected in sequence, taking the output of the Nth CSWin Transformer block in each branch and the output of the downsampling operation as the output of the dual-scale Transformer block with downsampling; (2-d) The dual-scale Transformer block described in step (2-b) comprises two inputs and two outputs, and the specific operation comprises: inputting the two inputs into two parallel independent branches respectively, each branch is composed of N CSWin Transformer blocks connected in sequence, and taking the output of the Nth CSWin Transformer block in each branch as the output of the dual-scale Transformer block; (2-e) The first, second and third coding blocks described in step (2-b) have specific structures including: The output of the Nth CSWin Transformer block in each branch of the Dwon-DS-Trans block is connected to a local feature enhancement module, and the output of the local feature enhancement module is used as the input of the two-scale fusion module, and the output of the two-scale fusion module is used as the output of the current encoding block; The non-overlapping large-scale sub-blocks S in step (2-b) L and small-scale sub-block S S As the input of the first encoding block, they are respectively input into the two parallel independent branches in the Dwon-DS-Trans block; the output of the downsampling operation of the Dwon-DS-Trans block in the first encoding block is used as the input of the second encoding block, and is respectively input into the two parallel independent branches of the Dwon-DS-Trans block in the second encoding block; the output of the downsampling operation of the Dwon-DS-Trans block in the second encoding block is used as the input of the third encoding block, and is respectively input into the two parallel independent branches of the Dwon-DS-Trans block in the third encoding block; the output of the downsampling operation of the Dwon-DS-Trans block in the third encoding block is used as the input of the fourth encoding block, and is respectively input into the two parallel independent branches of the dual-scale Transformer block in the fourth encoding block; (2-f) The fourth coding block described in step (2-b) has a specific structure comprising: the output of the Nth CSWinTransformer block of each branch in the dual-scale Transformer block is respectively connected to a local feature enhancement module, and the output of the local feature enhancement module is used as the input of the dual-scale fusion module, and the output of the dual-scale fusion module is used as the output of the fourth coding block; (2-g) The local feature enhancement module described in step (2-b) has a specific structure comprising: for feature I, first inputting I into three independent feature extraction branches respectively, wherein the first branch is composed of 3×3 convolution, deep convolution and GeLU activation layer connected in sequence, the second branch is composed of 3×3 convolution, GeLU activation layer, 3×3 convolution, GeLU activation layer connected in sequence, and the third branch is composed of deep convolution, GeLU activation layer, deep convolution, GeLU activation layer connected in sequence, and then the outputs O1, O2 and O3 of the three branches are spliced, and the spliced ​​results are sequentially subjected to 3×3 convolution and layer normalization to obtain the output of the local feature enhancement module; (2-h) The linear prediction head described in step (2-b) is composed of a 1×1 convolution and a Softmax layer connected in sequence; (3) Use training data set A to train DS-Net until the preset loss function converges; a hybrid loss function combining Dice and cross entropy is used as the objective function of network training; (4) Use the trained network to test the image to be segmented and obtain the segmentation result.

2. The method for multi-organ segmentation of abdominal CT images based on a dual-scale hybrid network as claimed in claim 1, characterized in that: The dual-scale fusion module described in step (2-b) includes two inputs and one output, and the specific structure includes: for the smaller input I S , first use an upsampling operation to restore it to the same size as the larger input I L The size is the same, and then the upsampled result is combined with the larger input I L The splicing is performed, and then the splicing result is convolved by 3×3 to obtain the output of the dual-scale fusion module.

3. The method for multi-organ segmentation of abdominal CT images based on a dual-scale hybrid network as claimed in claim 1, characterized in that: In step (2-a), r is a natural number between 8 and 32, in step (2-b), s is an even number between 2 and 6, and in step (2-c), N is a natural number between 3 and 8.

Citation Information

Patent Citations

  • Medical image segmentation method and device based on dual-scale encoder network, and medium

    CN116485815A

  • Abdomen multi-organ segmentation method based on Transform and U-Net fusion network

    CN116580036A