An automatic segmentation method for brain damage in premature infants
Through the encoder decoder architecture combined with the large-scale convolution module, PAFM module and CMSC module, the automatic segmentation network model of brain injury lesions in premature infants is solved, and high-precision automatic segmentation of lesions in premature infants is achieved.
Patent Information
- Application Number
- CN202510060474.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The prior art has challenges in the automatic segmentation of brain injury lesions in premature infants with lesions with lesions with lesions without fixed characteristic distribution, resulting in low diagnostic accuracy and time-consuming problems.
The automatic segmentation network model of brain injury lesions in premature infants using encoder decoder architecture combines large-scale convolution modules, PAFM modules and CMSC modules to achieve accurate segmentation of brain injury lesions in premature infants through multimodal information fusion and small lesions feature extraction.
The segmentation accuracy of brain injury lesions in premature infants is improved, surpassing other basic models in the prior art, and automatically segmenting and accurate identification of various brain injury lesions in premature infants is achieved.
Smart Images

Figure CN119991699B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic segmentation of disease lesions, and in particular to a method for automatic segmentation of lesions of premature infants with brain damage. Background Art
[0002] Brain injury in premature infants (BIPI) refers to varying degrees of cerebral ischemia, hemorrhagic and inflammatory damage in premature infants caused by various pathological factors during the perinatal period and delivery, with corresponding clinical symptoms and signs. Severe cases can lead to neurological sequelae or even direct death.
[0003] The hallmarks of premature brain injury are white matter damage and inadequate myelination, as well as neuroinflammation, axonal loss, impaired thalamocortical connectivity, and poor maturation of deep gray matter. It is generally believed that premature brain injury primarily involves white matter damage, focal necrotic lesions, or cystic lesions. These highly significant features are often evident on magnetic resonance imaging (MRI). MRI is the most effective imaging modality for detecting this disease. The diffusely abnormal signal intensity of the white matter on T1-weighted and T2-weighted imaging can more accurately detect lesions earlier. In clinical diagnosis, higher or lower values of these two weighted signals are also used to determine whether a premature infant has brain injury.
[0004] Although MRI's various imaging modalities can provide a wealth of information for clinical diagnosis and play an indispensable role in the diagnosis of brain injury lesions in premature infants, a more common clinical diagnosis is that the punctate focal necrosis lesions caused by white matter damage and the associated glial scars are scattered, multiple, and particulate, making them difficult to see on neuroimaging. Currently, the main clinical diagnostic method is to observe imaging characteristics for diagnosis, which often leads to missed diagnoses or misdiagnoses. How to ensure accurate, timely, and rapid diagnosis of premature encephalopathy is also a major concern for doctors.
[0005] Currently, MRI images are typically used by radiologists to visually classify brain lesions in premature infants. However, due to the highly heterogeneous nature and morphology of these lesions, the frequent presence of small lesions, and the irregular distribution of these lesions, diagnosis is time-consuming and can be inaccurate. Therefore, there is an urgent need to develop computer-assisted diagnosis methods based on convolutional neural networks to help radiologists automatically segment or detect relevant lesions before diagnosing the disease. Segmenting the brain of premature infants (≤1 month) based on MRI images is an extremely challenging task. The main reasons are: 1) To reduce motion artifacts, MRI scans of premature infants are short, resulting in low signal-to-noise ratio and resolution; 2) the brain is not yet myelinated, resulting in small lesions that are easily blurred and difficult to identify. In recent years, several related studies have focused on the automated segmentation of white matter in premature infant brain MRI or brain lesions. W Liu et al. analyzed brain MRI images of 63 premature infants with a gestational age of less than 37 weeks. They used deep learning neural network models such as U-Net, U-Net++, and ResU-Net to automatically segment damaged white matter in premature infants' brains. The segmented damaged white matter was then binarized and processed using digital image processing methods such as Gaussian filtering, flood filling, and connected components, ultimately obtaining clearly defined white matter regions. Hailong Li et al. used a deep CNN model to segment diffuse white matter abnormalities on T2-weighted MRI images of premature infants. They evaluated the model using cross-validation and holdout validation on a cohort of 95 premature infants and tested it on an independent cohort of 28 premature infants, achieving good results. M. Jorge Cardoso et al. proposed a new method called adaptive multimodal maximum a posteriori expectation maximization segmentation algorithm for premature infants (AdaPT), which uses an iterative relaxation strategy and incorporates a spatial uniformity term in the form of a Markov random field. Segmentation experiments were performed on a clinical cohort of 92 infants, including normal and pathological cortical gray matter, cerebellum, and ventricular volumes. These validation results were compared with widely used MRI segmentation algorithms for premature infants, and the Dice coefficient was significantly improved.
[0006] Although the above methods have achieved certain success in MRI segmentation of premature infants' brains or some damaged white matter segmentation, the diversity and non-fixed characteristic distribution of lesions in the entire field of premature infant brain injury (BIPI) lesion diagnosis still pose a huge challenge for clinical identification. Figure 1 This is an example of the difficulties we encounter in actual segmentation work. Figure 1Each row in the figure shows 4 slices of T1 images, corresponding to two patients respectively. The first and third images in each row are the original slice images, and the second and fourth images are the corresponding annotations. The first row shows white matter damage (left) and lateral ventricular cysts (right). The black circles in the original image are the corresponding lesions. The second row shows the irregular and uneven spatial distribution of encephalopathy lesions in premature infants. The black circles circle the lesions at the edge of the brain and the lesions in the middle of the brain respectively. The third row shows the similarity between the lesions and the cortical structures and the proximity of their spatial positions. The black circles in the original image are some cortical structures similar to the lesions. Figure 1 As shown in Figure 2, brain damage lesions in premature infants include not only large white matter lesions, but also some cystic lesions. Among them, white matter lesions appear as higher signals on MRI, while cystic lesions appear as lower signals on MRI. Different lesions show different characteristics on MRI. Figure 1 As shown in Figure 2, the distribution of brain injury lesions in premature infants is also very uneven throughout the brain. Some lesions are located at the edge of the brain, while others are located inside the brain. The lesions are generally irregular in shape and small in size. Figure 1 As shown, some lesions are similar to certain cortical structures in the brain, showing high signal intensity on the primary reference T1-weighted images and being spatially close. This greatly increases the difficulty of automatic model recognition, so a new, universal method is needed to achieve automatic segmentation of premature brain lesions. Summary of the Invention
[0007] The present invention aims to solve at least one of the technical problems existing in the prior art or related art.
[0008] Therefore, the purpose of the present invention is to provide a method for automatic segmentation of lesions of premature infants with brain damage.
[0009] In order to achieve the above-mentioned purpose, the technical solution of the present invention provides a method for automatic segmentation of lesions of premature brain damage, which method includes: step S1: acquiring a data set; wherein the data set is a brain MRI image of a patient with premature brain damage; the sequence of the brain MRI image includes T1WI and T2WI; the lesion regions of interest of the brain MRI image have been marked; step S2: preprocessing the data set; step S3: dividing the preprocessed data set into a training set, a test set, and a validation set according to a preset ratio; step S4: building a network model for automatic segmentation of lesions of premature brain damage; wherein the network model is an encoder-decoder architecture, and the network model includes: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group includes 2 large-scale convolution modules; the encoder part and the decoder part are both four layers; the encoder part of each layer includes 2 sub-encoders, and each sub-encoder includes: a Uxnet Block and a Downsampling Block; the decoder part of each layer includes a sub-decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; an output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; an output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the decoder of the corresponding layer in sequence through a PAFM module, a CMSC module, and a Res Block, the other output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected to the input end of the decoder of the first layer in sequence through a fusion module and a Res Block; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer each include a fusion module, a Res Block, and a transposed convolution module connected in sequence; and the other output end of the large-scale convolution module group is sequentially output through a fusion module and a Res Block for feature output, and the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module, and then through Res Block and output convolution perform channel adjustment to obtain the final lesion segmentation result; the large-scale convolution module is used to map the received preprocessed image data to a low-dimensional space to obtain the spatial potential feature representation of the corresponding image after preprocessing; each encoder of the encoder part of the first layer, the second layer, or the third layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module;Each encoder of the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the decoder of the first layer through a fusion module and a Res Block in sequence; the PAFM module is used to perform fusion enhancement of the multimodal feature map output by each encoder of the encoder part of the first layer, the second layer, or the third layer in terms of information between different modalities and within the same modality, so as to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature map into the corresponding CMSC module; the CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information can be subsequently sent to the decoder of the corresponding layer through a Res Block for decoding; the decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling through a fusion module and a Res Block in sequence, and then input the Res The feature map output by the block is upsampled through a transposed convolution module; the decoders of the second, third, and fourth layers are used to splice and fuse the original level features input by the jump connection and the deep level features output by the decoder of the previous layer through a fusion module, and then sequentially upsample the feature map output by the fusion module in the decoder of the current layer through a ResBlock and a transposed convolution module; and the large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block through a fusion module in sequence, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer are subsequently spliced and fused through a fusion module, and then channel adjustment is performed through the Res Block and the output convolution to obtain the final lesion segmentation result; step S5: input the training set into the network model for training, and adjust the hyperparameters based on the validation set; step S6: input the test set into the trained network model to obtain the automatic segmentation result of the lesion of premature brain injury corresponding to the test set.
[0010] Preferably, each PAFM module includes: an axial vision transformer module, a channel attention convolution module in parallel with the axial vision transformer module;
[0011] The workflow of the axial vision transformer module specifically includes:
[0012] Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x , Expressed as:
[0013] x=concat(encoder1(input1),encoder2(input2))(1)
[0014] In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c) , d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modes; c represents the number of channels; concat represents the concatenation operation;
[0015] Step S1.2: Embed the feature map x to x∈(d*w*h,c), input the position code pos(x) of x at the same time, and then perform layer normalization, which is expressed as:
[0016] x=LN(patchembed(x)+pos(x))(2)
[0017] In formula (2), LN represents the layer normalization operation; patchembed represents the operation of embedding layer mapping on the feature map x;
[0018] Step S1.3: Use the axial attention mechanism to perform self-attention calculation on each modality to obtain axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then, sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x′ obtained in step S2 to achieve spatial feature enhancement within each modality, which is expressed as:
[0019] z=AMHA xyz (x)+AMHA xy (x)+AMHA z (x)+x'(3)
[0020] In formula (3), AMHA xyz (x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the two-dimensional slice multi-head attention mechanism; AMHA z (x) represents the vertical multi-head attention mechanism; x′ represents the feature map x′ obtained in step S2; z represents the feature after spatial feature enhancement;
[0021] Step S1.4: The feature z after spatial feature enhancement is layer-normalized, and then processed by a multi-layer perceptron MLP. The feature processed by the multi-layer perceptron MLP and the feature z after spatial feature enhancement are summed, which is expressed as:
[0022] z=MLP(LN(z))+z(4)
[0023] In formula (4), LN represents the layer normalization operation;
[0024] The workflow of the channel attention convolution module specifically includes:
[0025] Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c), and then use 3D average pooling to aggregate the features of different channels, expressed as:
[0026] y=AvgPool3D(conv(x))(5)
[0027] In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling;
[0028] Step S2.2: Use 1x1x1 convolution to perform interactions between different modalities on y, and then pass sigmoid to obtain the attention map atty between different features, expressed as:
[0029] atty=sigmoid(conv(y))(6)
[0030] In formula (6), conv represents the convolution operation; sigmoid represents the activation function:
[0031] Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as:
[0032] y=y*atty(7)
[0033] The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c) , Then Z is fused with y obtained in step S2.3 through residual connection to obtain the output feature output, which is expressed as:
[0034] output=z+y(8)
[0035] Preferably, the workflow of each CMSC module specifically includes:
[0036] Step S3.1: The input is sequentially subjected to the conv 5x5x5 convolution operation, batch normalization operation, and activation function GELU to obtain the output x1 , Expressed as:
[0037] x1=GELU(BN(Conv5x5x5(input)))(9)
[0038] In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0039] Step S3.2: Add the input and output x1, and then perform a conv 3x3x3 convolution operation to output the result, which is expressed as:
[0040] x2=GELU(BN(Conv3x3x3(input+x1)))(10)
[0041] In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0042] Step S3.3: Add the input and output x2, and then perform a conv 1x1x1 convolution operation to output the result, which is expressed as:
[0043] x3=GELU(BN(Conv1x1x1(input+x2)))(11)
[0044] In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0045] Step S3.4: Concatenate the input, output x1, output x2, and output x3, then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform batch normalization and activation function GELU to obtain the output output. , Expressed as:
[0046] output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3))))(12)
[0047] In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
[0048] Preferably, the step S2 specifically includes: step S2.1: registering the T1WI and T2WI to the same space so that the marked lesions can be displayed at the same position of T1WI and T2WI; step S2.2: cropping the registered T1WI and T2WI to crop out the black edge areas in the registered T1WI and T2WI; step S2.3: normalizing the grayscale intensity of the cropped T1WI and T2WI; step S2.4: performing data enhancement on the grayscale intensity normalized T1WI and T2WI
[0049] Beneficial effects of the present invention:
[0050] (1) The method for automatic segmentation of brain injury lesions in premature infants provided by the present invention proposes a universal model for fully automatic segmentation of brain injury lesions in premature infants. It can automatically segment various brain injury lesions in premature infants, and its accuracy exceeds that of other basic models in the existing technology.
[0051] (2) The present invention provides a method for automatic segmentation of brain injury lesions in premature infants. By using T1-weighted and T2-weighted data, a multi-encoder is used to fuse each stage, and a new multimodal fusion module, namely the PAFM module, is proposed in combination with other modules. Experiments have shown that the PAFM module can effectively utilize multimodal information, thereby improving the segmentation accuracy of brain injury lesions in premature infants.
[0052] (3) The automatic segmentation method for premature infant brain injury lesions provided by the present invention makes the segmentation of premature infant brain injury lesions more accurate by combining a segmentation module specifically for small lesions, namely the CMSC module.
[0053] Additional aspects and advantages of the invention will become apparent from the description which follows, or may be learned by practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 shows a brain MRI image of a premature infant with brain injury according to an embodiment of the present invention;
[0055] Figure 2 A schematic flow chart showing a method for automatically segmenting lesions of premature infant brain damage according to an embodiment of the present invention is shown;
[0056] Figure 3 A schematic diagram showing the overall architecture of a network model for automatic segmentation of premature infant brain damage according to an embodiment of the present invention is shown;
[0057] Figure 4 A schematic structural diagram of a PAFM module according to an embodiment of the present invention is shown;
[0058] Figure 5 A schematic structural diagram of a CMSC module according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0059] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0060] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0061] Figure 2 FIG. 1 is a schematic flow chart showing a method for automatically segmenting premature infant brain damage lesions according to an embodiment of the present invention. Figure 2 As shown, the automatic segmentation method for premature infant brain injury lesions includes:
[0062] Step S1: Acquire a data set; the data set is brain MRI images of premature infants with brain injury;
[0063] Step S2: preprocessing the data set;
[0064] Step S3: Divide the preprocessed data set into a training set, a test set, and a validation set according to a preset ratio;
[0065] Step S4: Building a network model for automatic segmentation of premature infant brain injury lesions;
[0066] Step S5: Input the training set into the network model for training, and adjust the hyperparameters based on the validation set;
[0067] Step S6: Input the test set into the trained network model to obtain the automatic segmentation results of premature infant brain injury lesions corresponding to the test set.
[0068] In this embodiment, the brain MRI image sequence includes T1WI and T2WI. T1WI is T1-weighted imaging (T1WI for short), and T2WI is T2-weighted imaging (T2WI for short).
[0069] In this embodiment, the lesion regions of interest of the brain MRI image have been marked.
[0070] In this embodiment, if Figure 3As shown in the figure, the automatic segmentation network model of the premature infant brain injury lesion is an encoder-decoder architecture, which includes: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group includes 2 large-scale convolution modules (large kemel projection); the encoder part and the decoder part are both four layers; the encoder part of each layer includes 2 sub-encoders, each sub-encoder includes: a Uxnet Block and a Downsampling Block; the decoder part of each layer includes a decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; an output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; an output end of the encoder part of the first layer, the second layer, and the third layer are connected to the input end of the decoder of the corresponding layer in sequence through a PAFM module, a CMSC module, and a Res Block, the other output end of the encoder part of the first layer, the second layer, and the third layer are connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected in sequence through a fusion module (Concat module), a Res Block is connected to the input end of the decoder of the first layer; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer each include a fusion module (Concat module), a Res Block, and a transposed convolution module connected in sequence; and the other output end of the large-scale convolution module group is sequentially output through a fusion module (Concat module) and a Res Block for feature output, and the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module (Concat module), and then channel adjustment is performed through a Res Block and output convolution to obtain the final lesion segmentation result.
[0071] In this embodiment, the large-scale convolution module is used to map the received preprocessed image data to a low-dimensional space to obtain a spatial potential feature representation of the preprocessed image; each encoder of the encoder part of the first layer, the second layer, or the third layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module; each encoder of the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the decoder of the first layer through a fusion module and a Res Block in sequence; the PAFM module is used to perform fusion enhancement of the multimodal feature map output by each encoder of the encoder part of the first layer, the second layer, or the third layer, between different modalities and within the same modality, to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature map to the corresponding CMSC module; the CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information can be subsequently extracted through a Res Block is sent to the decoder of the corresponding layer for decoding; the decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling, and then pass it through a fusion module and a Res Block in sequence, and then upsample the feature map output by the ResBlock through a transposed convolution module; the decoders of the second, third and fourth layers are used to splice and fuse the original level features input by the decoder of the current layer and the deep level features output by the jump connection through a fusion module, and then upsample the feature map output by the fusion module in the decoder of the current layer through a Res Block and a transposed convolution module in sequence; and the large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block through a fusion module in sequence, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer are subsequently spliced and fused through a fusion module, and then channel adjustment is performed through ResBlock and output convolution to obtain the final lesion segmentation result.
[0072] In this embodiment, the CMSC module is a cascade multi-scale convolution (CMSC) module, and the PAFM module is a parallel attention fusion module (PAFM module).
[0073] In one embodiment of the present invention, each PAFM module includes: an axial vision transformer module, a channel attention convolution module in parallel with the axial vision transformer module; Figure 4 As shown in the figure, the workflow of the axial vision transformer module specifically includes:
[0074] Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x , Expressed as:
[0075] x=concat(encoder1(input1),encoder2(input2))(1)
[0076] In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c) , d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modes; c represents the number of channels; concat represents the concatenation operation;
[0077] Step S1.2: Embed the feature map x to x∈(d*w*h,c), input the position code pos(x) of x at the same time, and then perform layer normalization, which is expressed as:
[0078] x=LN(patchembed(x)+pos(x))(2)
[0079] In formula (2), LN represents the layer normalization operation; patchembed represents the operation of embedding layer mapping on the feature map x;
[0080] Step S1.3: Use the axial attention mechanism to perform self-attention calculation on each modality to obtain axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then, sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x′ obtained in step S2 to achieve spatial feature enhancement within each modality, which is expressed as:
[0081] z=AMHA xyz (x)+AMHA xy (x)+AMHA z (x)+x'(3)
[0082] In formula (3), AMHA xyz (x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the two-dimensional slice multi-head attention mechanism; AMHA z (x) represents the vertical multi-head attention mechanism; x′ represents the feature map x′ obtained in step S2; z represents the feature after spatial feature enhancement;
[0083] Step S1.4: The feature z after spatial feature enhancement is layer-normalized, and then processed by a multi-layer perceptron MLP. The feature processed by the multi-layer perceptron MLP and the feature z after spatial feature enhancement are summed, which is expressed as:
[0084] z=MLP(LN(z))+z(4)
[0085] In formula (4), LN represents the layer normalization operation;
[0086] The workflow of the channel attention convolution module specifically includes:
[0087] Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c), and then use 3D average pooling to aggregate the features of different channels, expressed as:
[0088] y=AvgPool3D(conv(x))(5)
[0089] In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling;
[0090] Step S2.2: Use 1x1x1 convolution to perform interactions between different modalities on y, and then pass sigmoid to obtain the attention map atty between different features, expressed as:
[0091] atty=sigmoid(conv(y))(6)
[0092] In formula (6), conv represents the convolution operation; sigmoid represents the activation function:
[0093] Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as:
[0094] y=y*atty(7)
[0095] The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c) , Then Z is fused with y obtained in step S2.3 through residual connection to obtain the output feature output, which is expressed as:
[0096] output=z+y(8)
[0097] In one embodiment of the present invention, Figure 5 As shown in the figure, the workflow of each CMSC module includes:
[0098] Step S3.1: The input is sequentially subjected to the convolution operation of 5x5x5, the batch normalization operation, and the activation function GELU to obtain the output x1 , Expressed as:
[0099] x1=GELU(BN(Conv5x5x5(input)))(9)
[0100] In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0101] Step S3.2: Add the input and output x1, and then perform a conv 3x3x3 convolution operation to output the result, which is expressed as:
[0102] x2=GELU(BN(Conv3x3x3(input+x1)))(10)
[0103] In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0104] Step S3.3: Add the input and output x2, and then perform a conv 1x1x1 convolution operation to output the result, which is expressed as:
[0105] x3=GELU(BN(Conv1x1x1(input+x2)))(11)
[0106] In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0107] Step S3.4: Concatenate the input, output x1, output x2, and output x3, then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform batch normalization and activation function GELU to obtain the output output. , Expressed as:
[0108] output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3))))(12)
[0109] In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
[0110] In one embodiment of the present invention, the step S2 specifically includes: step S2.1: aligning the T1WI and T2WI to the same space so that the marked lesions can be displayed at the same position of T1WI and T2WI; step S2.2: cropping the aligned T1WI and T2WI to crop out the black edge areas in the aligned T1WI and T2WI; step S2.3: normalizing the grayscale value intensity of the cropped T1WI and T2WI; step S2.4: performing data enhancement on the T1WI and T2WI after grayscale value intensity normalization.
[0111] The following is a specific embodiment to illustrate the technical solution of the present invention. The specific implementation steps of the method for automatic segmentation of premature infant brain damage are as follows:
[0112] (1) Obtaining a data set; wherein the data set is brain MRI images of premature infants with brain injury; the sequence of brain MRI images includes T1WI and T2WI; the lesion regions of interest of the brain MRI images have been marked.
[0113] The data set of this specific embodiment mainly comes from 267 patients with premature brain injury in Shengjing Hospital, including 164 male children and 103 female children. The gestational age of the children was 24 to 36 weeks, with an average gestational age of (32.5±1.5) weeks. A total of 1,785 injury lesions were extracted, and the same lesion was only counted once when displayed in different sequences.
[0114] Brain MRI images of patients with premature brain injury were obtained using a 1.5T GE Signal HDxt system scanner equipped with an 8-channel phased array head coil. Subjects were instructed to lie supine with their head first in the scanner and remain asleep throughout the examination. The MRI scan sequence included T1WI and T2WI with the following parameters: T1WI (TR = 2950 ms, TE = 24 ms, slice thickness 4 mm, field of view 22 × 22 cm, matrix 288 × 192); T2WI (TR = 5000 ms, TE = 120 ms, slice thickness 4 mm, field of view 22 × 22 cm, matrix 320 × 320).
[0115] After scanning, radiologists manually labeled lesion regions of interest (ROIs) using the SRhythmMultiLabel platform under the same window width and window level conditions. All radiologists involved in labeling received specialized training. The labeling process was divided into two phases. In the first phase, two attending physicians manually outlined visually significant lesions, while referencing other image sequences. During labeling, they identified and avoided areas of normal white matter with high water content, such as increased white matter signal on T1WI images but not low signal on T2WI sequences. In the second phase, two physicians with associate senior status or higher reviewed and verified the results of the first phase, supplementing and correcting any missed or mislabeled lesions until they reached consensus on labeling the same patient.
[0116] (2) Preprocessing the dataset. Preprocessing the dataset includes: registering T1WI and T2WI to the same space so that the marked lesions can be displayed at the same position on T1WI and T2WI; cropping the registered T1WI and T2WI to crop out the black edge areas in the registered T1WI and T2WI; normalizing the grayscale intensity of the cropped T1WI and T2WI; and performing data enhancement on the grayscale intensity normalized T1WI and T2WI.
[0117] Specifically, the original size of the T1 images we used was slightly different, and the T2 images were not exactly the same size as the T1 images. Because infants often move their heads during MRI scans, the spatial positions of the T2-weighted and T1-weighted images also differed, which caused problems during the input process into the network. To address this, we first used a rigid registration method to register the T2 images to the same space as the T1-weighted images, ensuring that the lesions annotated by the doctor would appear in the same location on both the T1-weighted and T2-weighted images.
[0118] Next, we primarily used preprocessing methods from the monai package. We first cropped the black edges of the MRI images. MRIs often have a large number of black edge pixels outside the brain. These edge pixels are considered noise for model training, affecting both the speed and effectiveness of model training. We then performed grayscale intensity normalization. The pixel values of T1-weighted and T2-weighted images are distributed within the range of (0, 2000). We linearly mapped the pixel values of T1-weighted and T2-weighted images from the interval (0, 2000) to the interval (0, 1), enabling better model convergence. Data augmentation methods, such as random image rotation and random shifting of grayscale intensity, were used to expand the data and improve the generalization of the model. After data augmentation, we resized the preprocessed images to a uniform size of (336, 336, 32) for more accurate analysis of images from different patients.
[0119] (3) The dataset is divided into 8:2 parts, 20% is used as the test set, 10% of the remaining 80% is extracted as the validation set for adjusting the hyperparameters, and the rest is used as the training set.
[0120] (4) Build a network model for automatic segmentation of premature brain injury lesions. In terms of the model, our baseline model uses an excellent medical image segmentation model 3D-UXnet, which has achieved sota results in multiple medical image segmentation tasks (including brain segmentation and organ segmentation). We migrated it to our premature brain injury lesion segmentation task and modified the model to make it more suitable for this task. We reduced the parameters of the model and removed a skip connection layer. The modules we added to the premature brain injury lesion segmentation mainly include two, namely cascaded multi-scale convolution (CMSC) and parallel attention fusion module (PAFM). The cascaded multi-scale convolution is mainly used to segment small lesions in the premature brain, such as cysts, and the parallel attention fusion module is mainly used to complete the complementary fusion of multimodal information and use multimodal information to improve the accuracy of segmentation.
[0121] The following explains why the CMSC module was added. Brain injury lesions in premature infants are generally small and have no fixed distribution pattern, often located in multiple brain regions. During the network's downsampling feature extraction process, the size of the feature map decreases as the network deepens. Feature maps generated in deeper layers are more abstract, and therefore are more likely to miss important information about small objects, affecting the detection of small brain injury lesions in premature infants. To address this challenge, we implemented a feature extraction module specifically designed for small objects.
[0122] The following explains why the PAFM module was added. Both T1-weighted and T2-weighted images play a role in the clinical diagnosis of encephalopathy in prematurity. Doctors typically use a combination of abnormal signals on these two images to determine the type or severity of encephalopathy. For example, MRI findings of leukomalacia (PVL) are low signal intensity on T1-weighted images and high signal intensity on T2-weighted images within the periventricular white matter. Signal changes in the posterior limb of the internal capsule on both T1- and T2-weighted images often indicate selective neuronal necrosis. Both T1- and T2-weighted images exhibit features consistent with encephalopathy in prematurity. Lesions that are subtle on T1-weighted images may be more pronounced on T2-weighted images, and lesions often exhibit the most prominent features on T1-weighted images. Inspired by this, we used T1-weighted data as the primary modality and T2-weighted data as an auxiliary modality to fuse the two modalities for segmentation of premature brain lesions. There are various methods for multimodal fusion of medical images, primarily classified as input-level fusion networks, hierarchical fusion networks, and decision-level networks. Because the modal features of T1-weighted and T2-weighted images corresponding to premature brain injury lesions are complex and contain a lot of similar or complementary information, we use a hierarchical fusion strategy to perform multi-level fusion, thereby obtaining multi-scale contextual features from different modalities to help achieve robust lesion segmentation. Therefore, we developed the PAFM module.
[0123] Specifically, if Figure 3As shown, the automatic segmentation network model for premature brain injury lesions is an encoder-decoder architecture, and the network model includes: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group includes 2 large-scale convolution modules; the encoder part and the decoder part are both four layers; the encoder part of each layer includes 2 sub-encoders, each sub-encoder includes: a Uxnet Block and a Downsampling Block; the decoder part of each layer includes a sub-decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; an output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; an output end of the encoder part of the first layer, the second layer, and the third layer are connected to the input end of the decoder of the corresponding layer in sequence through a PAFM module, a CMSC module, and a Res Block, the other output end of the encoder part of the first layer, the second layer, and the third layer are connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected in sequence through a fusion module, a Res Block is connected to the input end of the decoder of the first layer; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer each include a fusion module, a Res Block, and a transposed convolution module connected in sequence; and the other output end of the large-scale convolution module group outputs features through a fusion module and a Res Block in sequence, and the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module, and then channel adjustment is performed through a Res Block and output convolution to obtain the final lesion segmentation result.
[0124] The large-scale convolution module is used to map the received pre-processed image to a low-dimensional space to obtain the spatial potential feature representation of the pre-processed image; each encoder of the encoder part of the first layer, the second layer, or the third layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module; each encoder of the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map through a fusion module, a Res Block is sent to the decoder of the first layer; the PAFM module is used to perform fusion enhancement of the multimodal feature map output by each encoder of the encoder part of the first layer, the second layer, or the third layer, between different modalities and within the same modality, so as to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature map to the corresponding CMSC module; the CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information can be subsequently sent to the decoder of the corresponding layer through a ResBlock for decoding; the decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling, and then pass through a fusion module and a Res Block in sequence, and then upsample the feature map output by the Res Block through a transposed convolution module; the decoders of the second, third, and fourth layers are used to splice and fuse the original level features of the jump connection input and the deep level features output by the decoder of the previous layer through a fusion module, and then pass through a Res Block, a transposed convolution module upsampling the feature map output by the fusion module in the decoder of the current layer; and the large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block through a fusion module in sequence, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer can be spliced and fused through a fusion module, and then channel adjustment is performed through the Res Block and output convolution to obtain the final lesion segmentation result.
[0125] In general, our model uses multiple encoders to fully extract deep features from multiple modalities. Each encoder uses a 3D-UXNet encoder with reduced parameters. Our input consists of uniformly partitioned image blocks of size (64, 64, 32), primarily containing information from the T1W and T2W modalities. The input is first mapped to a low-dimensional space using large-scale convolutions to obtain a spatial latent feature representation of the image. The image is then downsampled using four UxNet blocks and a downsampling block. The multimodal feature maps obtained from the UxNet blocks and downsampling of the different modal encoders are fused using the PAFM module to integrate information across modalities and within the same modality. The fused and enhanced feature maps are then fed into a cascaded multi-scale convolution module to fully extract information about small objects (lesions). The output of the cascaded multi-scale convolutions is then fed simultaneously into the decoder through the original 3D-UXNet Res Block for decoding. The decoder uses the 3D-UXNET decoder to fuse original layer features from skip connections with deeper layer features, while also using transposed convolution for upsampling. The model learns the parameters of the transposed convolution to determine the upsampling method that best suits the dataset, resulting in superior performance compared to unpooling upsampling. The downsampled feature maps of each layer are concatenated and fused with those of the previous layer, enriching the image's multi-layer information. At the top layer, the feature maps generated by the large-scale convolutional mapping are concatenated with the final feature maps obtained from the downsampling layer, and the final output is generated through a Res Block and an output convolution block.
[0126] Furthermore, the parallel attention fusion module (PAFM) is a module created by referring to and modifying the channel attention convolution ECA and NestFormer. This fusion module is composed of an axial vision transformer module in NestFormer and an ECA channel attention convolution module in parallel. It is mainly used to calculate the spatial correlation within each modality and the correlation between different modalities. This module uses the channel attention mechanism and the intra-modal spatial attention mechanism in parallel to fully extract the information within the modality and the complementary information across modalities, and realizes the fusion operation between multimodal information to enhance the key features of the input feature map. The channel attention mechanism is mainly for extracting the correlation between modalities, and the axial vit module mainly extracts the correlation between each modality, and then fuses and outputs them.
[0127] Specifically, each PAFM module includes: an axial vision transformer module, a channel attention convolution module parallel to the axial vision transformer module; Figure 4As shown in the figure, the workflow of the axial vision transformer module specifically includes:
[0128] Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x, which is expressed as:
[0129] x=concat(encoder1(input1),encoder2(input2))(1)
[0130] In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c), d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modalities; c represents the number of channels; concat represents the concatenation operation;
[0131] Step S1.2: Embed the feature map x to x∈(d*w*h,c), input the position code pos(x) of x at the same time, and then perform layer normalization, which is expressed as:
[0132] x=LN(patchembed(x)+pos(x))(2)
[0133] In formula (2), LN represents the layer normalization operation; patchembed represents the operation of embedding layer mapping on the feature map x;
[0134] Step S1.3: Use the axial attention mechanism to perform self-attention calculation on each modality to obtain axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then, sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x′ obtained in step S2 to achieve spatial feature enhancement within each modality, which is expressed as:
[0135] z=AMHA xyz (x)+AMHA xy (x)+AMHA z (x)+x'(3)
[0136] In formula (3), AMHA xyz (x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the two-dimensional slice multi-head attention mechanism; AMHA z(x) represents the vertical multi-head attention mechanism; x′ represents the feature map x′ obtained in step S2; z represents the feature after spatial feature enhancement;
[0137] Step S1.4: The feature z after spatial feature enhancement is layer-normalized, and then processed by a multi-layer perceptron MLP. The feature processed by the multi-layer perceptron MLP and the feature z after spatial feature enhancement are summed, which is expressed as:
[0138] z=MLP(LN(z))+z(4)
[0139] In formula (4), LN represents the layer normalization operation;
[0140] In the workflow of the Axial Vision Transformer module, an attention mechanism is used to perform self-attention calculations within each modality. Since each downsampled feature map needs to be fused, the module adopts an axial attention mechanism to optimize computational efficiency and reduce model complexity. This mechanism decomposes multidimensional data into multiple axes (vertical, slice, window) and applies attention mechanisms to each of these axes. This approach significantly reduces the number of parameters and computational cost while maintaining performance.
[0141] The workflow of the channel attention convolution module specifically includes:
[0142] Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c), and then use 3D average pooling to aggregate the features of different channels, expressed as:
[0143] y=AvgPool3D(conv(x))(5)
[0144] In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling;
[0145] Step S2.2: Use 1x1x1 convolution to perform interactions between different modalities on y, and then pass sigmoid to obtain the attention map atty between different features, expressed as:
[0146] atty=sigmoid(conv(y))(6)
[0147] In formula (6), conv represents the convolution operation; sigmoid represents the activation function:
[0148] Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as:
[0149] y=y*atty(7)
[0150] The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c) , Then Z is fused with y obtained in step S2.3 through residual connection to obtain the output feature output, which is expressed as:
[0151] output=z+y(8)
[0152] In the workflow of this channel attention convolution module, we perform cross-modal channel attention calculations on features of different modalities to fully integrate features of different modalities.
[0153] Furthermore, if Figure 5 As shown, each CMSC module is specifically used for:
[0154] Step S3.1: The input is sequentially subjected to the conv 5x5x5 convolution operation, batch normalization operation, and activation function GELU to obtain the output x1 , Expressed as:
[0155] x1=GELU(BN(Conv5x5x5(input)))(9)
[0156] In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0157] Step S3.2: Add the input and output x1, and then perform a conv 3x3x3 convolution operation to output the result, which is expressed as:
[0158] x2=GELU(BN(Conv3x3x3(input+x1)))(10)
[0159] In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0160] Step S3.3: Add the input and output x2, and then perform a conv 1x1x1 convolution operation to output the result, which is expressed as:
[0161] x3=GELU(BN(Conv1x1x1(input+x2)))(11)
[0162] In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function;
[0163] Step S3.4: Concatenate the input, output x1, output x2, and output x3, then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform batch normalization and activation function GELU to obtain the output output. , Expressed as:
[0164] output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3))))(12)
[0165] In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
[0166] The cascaded multi-scale convolution (CMSC) module is a module in w-net. This module mainly extracts multi-level information of the original feature map through convolutions of different sizes (convolution kernels of length 5, 3, and 1). Through this multi-scale convolution extraction and fusion of each CMSC module, the semantic information contained in the deep layer of the feature map and the shallow semantic information can be well extracted at multiple scales. The multi-scale information is passed to the decoder through jump connections, allowing the model to focus on local details while capturing a wider range of contextual information. This combination of global and local information helps the model better understand the image, making the segmentation process more refined for lesions of premature encephalopathy and better preserving boundary information. At the same time, for lesions of different sizes, multi-scale convolution can simultaneously capture the characteristics of large and small lesions, improving the detection and recognition capabilities of the model.
[0167] (5) The training set is input into the network model for training, and the validation set is used to adjust the hyperparameters.
[0168] All experiments were conducted on an NVIDIA RTX 2080Ti GPU with 11GB of memory. Due to GPU memory limitations, we uniformly sliced the images into small patches of size (64, 64, 32) after preprocessing and data augmentation for network input. During training, we used the AdamW optimizer with a batch size of 3 and an initial learning rate of 0.001. During training, we used a cosine annealing strategy, which continuously optimizes the model's local optimal solution to find the global optimal solution. The training cycle was 150 epochs, with a minimum learning rate of 1e-5. The loss function used a combination of Dice and cross-entropy losses, with coefficients of 0.5 and 0.5, respectively, to maintain an equal contribution to the loss. All models were trained for 600 epochs. During inference, we used a sliding window inference strategy, with the window size consistent with the training patch size (64, 64, 32).
[0169] (6) Select evaluation indicators.
[0170] For the segmentation evaluation of premature infant brain injury lesions, we primarily use the Dice coefficient (DSC) and Hausdorff distance (HD) to assess the quality of our model's segmentation. The Dice coefficient is a set similarity metric, originally used to judge the similarity between two sets. In semantic segmentation, it is primarily used to measure the degree of overlap between two regions. Its value ranges from 0 to 1, with values closer to 1 indicating greater overlap and values closer to 0 indicating less overlap. We use it to assess the degree of overlap between our segmentation results and the original labels, thereby evaluating the performance of our model. The calculation formula is as follows:
[0171]
[0172] In the above formula, A and B are often pixel matrices consisting of 0 and 1 pixels. The area corresponding to the pixel value 1 represents the foreground area, and the area corresponding to the pixel value 0 represents the background area. |A| represents the total number of pixels with a value of 1 in pixel matrix A. |A∩B| is often achieved by multiplying the corresponding elements of pixel matrix A and pixel matrix B.
[0173] Hausdorff distance is also an indicator used to measure the proximity of the boundaries of two sets. In semantic segmentation, it is usually used to measure the accuracy of boundary segmentation. The closer the HD distance is to 0, the closer the boundaries of the two sets are, and the better the segmentation effect. Its calculation formula is as follows:
[0174] h(A, B)=max a∈A min b∈B ||α-b||
[0175] h(B, A)=max b∈B min a∈A ||ba||
[0176] H(A,B)=max(h(A,B),h(B,A))
[0177] ||ab|| represents the distance paradigm between the point sets a and b, and h(A, B) is called the one-way Hausdorff distance from A to B. It can be seen that the Hausdorff distance is often the larger of the one-way distances between the two point sets. A and B represent the segmented region point set and the label point set, respectively, that is, the points with a pixel value of 1. It is worth noting that in order to eliminate the impact of outliers on the evaluation effect, when calculating the one-way Hausdorff distance between the two point sets, the 95% digit value, that is, the 95% value in ascending order, is taken as the actual one-way Hausdorff distance. We call the HD distance calculated in this way HD95, and HD95 is often used in practice to evaluate the segmentation performance of the model.
[0178] (7) The test set is input into the trained network model to obtain the automatic segmentation results of the brain injury lesions of premature infants corresponding to the test set.
[0179] (8) Verify the superiority of the network model for automatic segmentation of lesions of premature infants with brain damage according to the present invention.
[0180] To compare the superiority of our proposed model for lesion segmentation in premature infants, we conducted comparisons under the same experimental conditions with several classic medical image segmentation models. These included CNN-based methods such as 3DU-Net, U-Net++, VNet, SegResNet, and Attention-UNet; methods combining CNN and VIT such as UNETR and SwinUNETR; the original 3D-UXNET; and a 3D-UXNET with reduced parameters and modified layers. We refer to this modified 3D-UXNET as mini-UXNET. The Dice coefficient (%) and HD95 (mm) described above were used as evaluation metrics. The results are shown in the table below.
[0181] Table 1Comparative experiments with different segmentation algorithms
[0182]
[0183] Our modified model demonstrated superior segmentation performance on a test dataset of premature infant brain injury lesions. Mini-UXNet, resulting from parameter reduction and layer modification, reduced its parameter count by approximately half and computational overhead by three-quarters compared to the original 3D-UXNet, while maintaining only slightly lower accuracy. This demonstrates that reducing the original model's parameters was a significant change, significantly reducing both model parameters and computational overhead while maintaining near-unchanged accuracy. All subsequent models we modified based on mini-UXNet were based on this model. Compared to the original 3D-UXNet, our model achieved an average Dice value of 0.7583 and an average HD95 value of 10.23 mm on the test set. Compared to the original 3D-UXNet baseline model without the added modules, our model achieved a 2.76% improvement in Dice and a 4.83 mm reduction in HD95, while also reducing its parameter count and computational overhead by significantly less than the original 3D-UXNet baseline. Furthermore, it achieved the best performance in both Dice and HD95 values compared to other segmentation models. HD95 is often used to measure the numerical value of the boundary between two regions and is more sensitive than the DICE value. Therefore, Dice is usually used as the main indicator and HD95 as a secondary reference indicator. At the same time, our model has 38.90M parameters and 41.79G FLOPs. Even when using multiple encoders, the number of parameters and computational complexity are still lower than UNETR, SwinUNETR, V-Net, and 3D-UXNet, and it is still a medium-sized model. Compared with the parameter-reduced 3D-UXNET, it has only 10.92M more parameters and 17.47G more computational complexity, but has achieved a significant improvement in segmentation accuracy. The improvement in the above indicators proves that our model can better extract some deep-level features, which is more beneficial for the segmentation of small lesions.
[0184] In summary, we studied the automatic segmentation of brain injury in premature infants (BIPI) lesions. We used the excellent medical image segmentation model 3D-UXNet and added a multi-scale convolutional extraction module (CMSC) to it, making it more suitable for the smaller BIPI lesions and performing fine segmentation. Furthermore, referring to the diagnostic methods used by doctors, our model also used data from two modalities, T1-weighted and T2-weighted images. We used a hierarchical modality fusion strategy and incorporated a parallel attention fusion module to fuse multimodal information, enabling the parallel extraction and fusion of intra-modal spatial attention and cross-modal attention. Experiments have demonstrated that our improved model achieves the best results in BIPI lesion segmentation compared to other models. It can better segment some difficult lesions and better assist doctors in diagnosing BIPI lesions.
[0185] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for automatic segmentation of premature infant brain lesions, characterized in that: include: Step S1: Acquire a data set; wherein the data set is brain MRI images of premature infants with brain injury; the sequence of the brain MRI images includes T1WI and T2WI; the lesion regions of interest in the brain MRI images are all marked; Step S2: preprocessing the data set; Step S3: Divide the preprocessed data set into a training set, a test set, and a validation set according to a preset ratio; Step S4: Build a network model for automatic segmentation of lesions of premature brain injury; wherein the network model is an encoder-decoder architecture, and the network model includes: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group includes two large-scale convolution modules; the encoder part and the decoder part are both four layers; the encoder part of each layer includes two sub-encoders, each sub-encoder includes: a UxnetBlock and a Downsampling Block; the decoder part of each layer includes a sub-decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; One output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; one output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the decoder of the corresponding layer in sequence through a PAFM module, a CMSC module, and a Res Block, the other output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected to the input end of the decoder of the first layer in sequence through a fusion module and a Res Block; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer each include a fusion module, a Res Block, and a transposed convolution module connected in sequence; and the other output end of the large-scale convolution module group is sequentially output through a fusion module and a Res Block for feature output, the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module, and then channel adjustment is performed through the Res Block and the output convolution to obtain the final lesion segmentation result; The large-scale convolution module is used to map the received pre-processed image data into a low-dimensional space to obtain a spatial latent feature representation of the corresponding pre-processed image; Each encoder in the encoder part of the first, second, or third layer is used to downsample the input image through a UxnetBlock and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module; Each encoder in the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the decoder of the first layer through a fusion module and a Res Block in sequence; The PAFM module is used to perform fusion enhancement of information between different modalities and within the same modality on the multimodal feature maps output by each encoder of the encoder part of the first layer, the second layer, or the third layer, so as to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature maps into the corresponding CMSC module; The CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information can be subsequently sent to the decoder of the corresponding layer through a Res Block for decoding; The decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling through a fusion module and a Res Block, and then upsample the feature map output by the Res Block through a transposed convolution module; The decoders of the second, third, and fourth layers are used to concatenate and fuse the original layer features input by the jump connection and the deep layer features output by the decoder of the previous layer through a fusion module, and then upsample the feature maps output by the fusion module in the decoder of the current layer through a Res Block and a transposed convolution module. The large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block through a fusion module in sequence, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer are subsequently spliced and fused through a fusion module, and then channel adjustment is performed through the Res Block and output convolution to obtain the final lesion segmentation result; Step S5: inputting the training set into the network model for training, and adjusting hyperparameters based on the validation set; Step S6: inputting the test set into the trained network model to obtain the automatic segmentation result of the premature infant brain injury lesion corresponding to the test set.
2. The method for automatic segmentation of premature infant brain damage lesions according to claim 1, wherein each PAFM module comprises: An axial vision transformer module and a channel attention convolution module in parallel with the axial vision transformer module; The workflow of the axial vision transformer module specifically includes: Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x , Expressed as: x=concat(encoder1(input1),encoder2(input2))(1) In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c), d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modalities; c represents the number of channels; concat represents the concatenation operation; Step S1.2: Embed the feature map x to x∈(d*w*h,c), input the position code pos(x) of x at the same time, and then perform layer normalization, which is expressed as: x=LN(patchembed(x)+pos(x))(2) In formula (2), LN represents the layer normalization operation; patchembed represents the operation of embedding layer mapping on the feature map x; Step S1.3: Use the axial attention mechanism to perform self-attention calculation on each modality to obtain axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then, sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x' obtained in step S2 to achieve spatial feature enhancement within each modality, which is expressed as: z=NOW xyz (x)+NO xy (x)+NO z (x)+x' (3) In formula (3), AMHA xyz (x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the two-dimensional slice multi-head attention mechanism; AMHA z (x) represents the vertical multi-head attention mechanism; x' represents the feature map x' obtained in step S2; z represents the feature after spatial feature enhancement; Step S1.4: The feature z after spatial feature enhancement is layer-normalized, and then processed by a multi-layer perceptron MLP. The feature processed by the multi-layer perceptron MLP and the feature z after spatial feature enhancement are summed, which is expressed as: z=MLP(LN(z))+z (4) In formula (4), LN represents the layer normalization operation; The workflow of the channel attention convolution module specifically includes: Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c) , Then use 3D average pooling to aggregate the features of different channels, expressed as: y=AvgPool3D(conv(x)) (5) In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling; Step S2.2: Use 1x1x1 convolution to perform interactions between different modalities on y, and then pass sigmoid to obtain the attention map aty between different features, expressed as: atty=sigmoid(conv(y)) (6) In formula (6), conv represents the convolution operation; sigmoid represents the activation function: Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as: y=y*atty (7) The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c) , Then Z is fused with y obtained in step S2.3 through residual connection to obtain the output feature output, which is expressed as: output=z+y (8).
3. The method for automatic segmentation of premature infant brain damage lesions according to claim 1, characterized in that: The workflow of each CMSC module includes: Step S3.1: The input is sequentially subjected to the conv 5x5x5 convolution operation, batch normalization operation, and activation function GELU to obtain the output x1 , Expressed as: x1=GELU(BN(Conv5x5x5(input)))(9) In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function; Step S3.2: Add the input and output x1, and then perform a conv 3x3x3 convolution operation to output the result, which is expressed as: x2=GELU(BN(Conv3x3x3(input+x1)))(10) In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function; Step S3.3: Add the input and output x2, and then perform a conv 1x1x1 convolution operation to output the result, which is expressed as: x3=GELU(BN(Conv1x1x1(input+x2)))(11) In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents the batch normalization operation; GELU represents the activation function; Step S3.4: Concatenate the input, output x1, output x2, and output x3, then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform batch normalization and activation function GELU to obtain the output output. , Expressed as: output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3))))(12) In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
4. The method for automatic segmentation of premature infant brain damage lesions according to any one of claims 1 to 3, characterized in that: The step S2 specifically includes: Step S2.1: registering the T1WI and T2WI images into the same space so that the marked lesions can be displayed at the same position on both T1WI and T2WI images; Step S2.2: cropping the registered T1WI and T2WI to crop out the black edge areas in the registered T1WI and T2WI; Step S2.3: Normalize the grayscale intensity of the cropped T1WI and T2WI; Step S2.4: Perform data enhancement on the T1WI and T2WI after grayscale value intensity normalization.
Citation Information
Patent Citations
Portrait segmentation network training method and device, equipment and medium
CN114549557A
Three-dimensional heart image segmentation method based on multi-scale edge perception
CN114663445A