Method for automatically segmenting focus of brain injury of premature infant
By using a network model with an encoder decoder architecture in the automatic segmentation of brain injury lesions in premature infants, combined with large-scale convolution module, PAFM module and CMSC module, the problem of accuracy and inefficiency in automatic segmentation of brain injury lesions in premature infants is solved, and high-precision lesion segmentation is achieved.
Patent Information
- Application Number
- CN202510060474.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The prior art has problems of accuracy and inefficiency in the automatic segmentation of brain injury lesions in premature infants, especially when dealing with diverse and non-fixed characteristic lesions distribution.
An automatic segmentation network model based on the encoder decoder architecture is proposed, combining large-scale convolution module, PAFM module and CMSC module to fusion of multimodal information and extract small lesions through the form of multiple encoders.
It realizes high-precision automatic segmentation of brain injury lesions in premature infants, surpasses the segmentation accuracy of the existing technology, and improves the detection ability of small lesions.
Smart Images

Figure CN119991699A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic segmentation of disease lesions, and in particular to an automatic segmentation method for lesions of premature infants with brain damage. Background Art
[0002] Brain Injury in Premature Infants (BIPI) refers to varying degrees of cerebral ischemia, hemorrhagic and inflammatory damage in premature infants caused by various pathological factors during the perinatal period and delivery, with corresponding clinical symptoms and signs. In severe cases, it can lead to neurological sequelae or even direct death.
[0003] The prominent features of premature brain injury are white matter damage and insufficient myelination, as well as neuroinflammation, neuronal axon loss, thalamocortical connection damage and poor maturation of deep gray matter. It is generally believed that premature brain injury mainly involves white matter damage, focal necrotic lesions or cystic lesions. These more significant features are often more obvious in magnetic resonance imaging (MRI). Magnetic resonance imaging is the most effective imaging method for detecting this disease. The diffuse abnormal signal intensity of white matter in T1-weighted imaging and T2-weighted imaging can detect lesions earlier and more accurately. In clinical diagnosis, the higher or lower values of these two weighted signals are also used to determine whether premature infants have brain injury.
[0004] Although MRI's various imaging methods can provide a lot of information for clinical diagnosis, it plays an indispensable role in the diagnosis of brain injury lesions in premature infants. However, in clinical diagnosis, it is more common that the punctate focal necrosis lesions caused by white matter damage and some of its related glial scars are scattered, multiple, and particulate, which are not easy to see in neuroimaging. At present, the main clinical diagnostic method is to observe the imaging characteristics for diagnosis, which often leads to artificial missed diagnosis or misdiagnosis. How to ensure accurate, timely, and rapid diagnosis of premature encephalopathy is also a big problem for doctors.
[0005] At present, for premature infants' brain injury lesions, radiologists usually observe MRI images with their naked eyes to classify the severity of the injury lesions. Due to the high morphological heterogeneity of the lesions, and the fact that they are often accompanied by small lesions and have no fixed distribution characteristics, the diagnosis takes a long time and the accuracy of the diagnosis may not be very high. Therefore, it is urgent to develop a computer-aided diagnosis method based on convolutional neural networks to help radiologists automatically segment or detect relevant lesions and then diagnose the disease. Segmenting the brain of premature infants (≤1 month) based on MRI images is an extremely challenging task. The main reasons are: ① In order to reduce motion artifacts, the MRI scanning time of premature infants' brains is relatively short, resulting in low signal-to-noise ratio and resolution; ② The brain of premature infants has not yet been myelinated, the lesions are small in size, the images are easily blurred, and the identification is difficult. In recent years, some related work has been devoted to studying the automatic segmentation of premature infant brain MRI or brain injury white matter. W Liu et al. analyzed brain MRI of 63 premature infants with gestational age <37 weeks, and used deep learning neural network models such as U-Net, U-Net++ and ResU-Net to automatically segment the damaged white matter in premature infants' brains, and binarized the segmented damaged white matter. At the same time, digital image processing methods such as Gaussian filtering, flood filling and connected components were used to process the segmented area, and finally a white matter area with clear boundaries was obtained. Hailong Li et al. used a deep CNN model to segment diffuse white matter abnormalities in T2-weighted images of premature infants' MRI. The model was evaluated on a group of 95 premature infants using cross-validation and retention validation. At the same time, the model was tested on an independent cohort of 28 premature infants, and good results were achieved. M. Jorge Cardoso et al. proposed a new method called adaptive multimodal maximum a posteriori expectation maximization segmentation algorithm for premature infants (i.e., AdaPT), which uses an iterative relaxation strategy and combines spatial uniformity terms in the form of Markov random fields. The segmentation experiments were validated on a clinical cohort of 92 infants including normal and pathological cortical gray matter, cerebellum, and ventricular volumes. These validation results were compared with widely used MRI segmentation algorithms for premature infants, and the Dice coefficient was significantly improved.
[0006] Although the above methods have achieved certain success in premature infant brain MRI segmentation or some damaged white matter segmentation, for the entire field of premature brain injury (BIPI) lesion diagnosis, the diversity and non-fixed characteristic distribution of lesions still pose a huge challenge for clinical identification. Figure 1 This is an example of the difficulties we encounter in actual segmentation work. Figure 1Each row in the figure shows slices of four T1 images, corresponding to two patients respectively. The first and third images in each row are the original slice images, and the second and fourth images are the corresponding annotations. The first row shows white matter damage (left side) and lateral ventricular cysts (right side). The black circles in the original image are the corresponding lesions. The second row shows the irregular and uneven spatial distribution of encephalopathy lesions in premature infants. The black circles circle the lesions at the edge of the brain and the lesions in the middle of the brain respectively. The third row shows the similarity between the lesions and the cortical structures and the proximity of their spatial positions. The black circles in the original image are some cortical structures similar to the lesions. Figure 1 As shown in Figure 1, brain damage lesions in premature infants include not only larger white matter lesions, but also some cystic lesions. White matter lesions show higher signals on MRI, while cystic lesions show lower signals on MRI. Different lesions show different characteristics on MRI. Figure 1 As shown in Figure 1, the distribution of brain injury lesions in premature infants is also very uneven throughout the brain. Some lesions are distributed at the edge of the brain, while others are distributed inside the brain. The lesions are generally irregular in shape and small in size. Figure 1 As shown in the figure, some lesions are similar to some cortical structures of the brain, showing higher signals on the main reference T1-weighted images and are relatively close in spatial position. These greatly increase the difficulty of automatic model recognition. Therefore, we need a new general method to achieve automatic segmentation of premature infant brain damage. Summary of the invention
[0007] The present invention aims to solve at least one of the technical problems existing in the prior art or related art.
[0008] Therefore, the object of the present invention is to provide a method for automatically segmenting lesions of premature infants with brain damage.
[0009] In order to achieve the above-mentioned purpose, the technical solution of the present invention provides a method for automatic segmentation of lesions of premature brain damage, the method comprising: step S1: acquiring a data set; wherein the data set is a brain MRI image of a patient with premature brain damage; the sequence of the brain MRI image comprises T1WI and T2WI; the lesion regions of interest of the brain MRI image have been marked; step S2: preprocessing the data set; step S3: dividing the preprocessed data set into a training set, a test set, and a validation set according to a preset ratio; step S4: building a network model for automatic segmentation of lesions of premature brain damage; wherein the network model is an encoder-decoder architecture, the network model comprises: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group comprises 2 large-scale convolution modules; the encoder part and the decoder part are both four layers; the encoder part of each layer comprises 2 sub-encoders, each sub-encoder comprises: a Uxnet Block and a Downsampling Block; the decoder part of each layer includes a sub-decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; an output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; an output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the decoder of the corresponding layer through a PAFM module, a CMSC module, and a Res Block in turn, the other output ends of the encoder part of the first layer, the second layer, and the third layer are connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected to the input end of the decoder of the first layer through a fusion module and a Res Block in turn; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer each include a fusion module, a Res Block, and a transposed convolution module connected in turn; and the other output end of the large-scale convolution module group is sequentially output through a fusion module and a Res Block for feature output, the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module, and then through Res Block Block and output convolution are used to adjust the channels to obtain the final lesion segmentation result; the large-scale convolution module is used to map the received preprocessed image data to a low-dimensional space to obtain the spatial potential feature representation of the corresponding image after preprocessing; each encoder of the encoder part of the first layer, the second layer, or the third layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module;Each encoder of the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the decoder of the first layer through a fusion module and a Res Block in sequence; the PAFM module is used to perform fusion enhancement of information between different modalities and within the same modality on the multimodal feature map output by each encoder of the encoder part of the first layer, the second layer, or the third layer, so as to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature map to the corresponding CMSC module; the CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information can be subsequently sent to the decoder of the corresponding layer through a Res Block for decoding; the decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling through a fusion module and a Res Block in sequence, and then The feature map output by the block is upsampled through a transposed convolution module; the decoders of the second, third and fourth layers are used to splice and fuse the original hierarchical features input by the jump connection and the deep hierarchical features output by the decoder of the previous layer through a fusion module, and then upsample the feature map output by the fusion module in the decoder of the current layer through a ResBlock and a transposed convolution module in turn; and the large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block through a fusion module in turn, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer are spliced and fused through a fusion module, and then the channel is adjusted through the Res Block and the output convolution to obtain the final lesion segmentation result; step S5: input the training set into the network model for training, and adjust the hyperparameters based on the validation set; step S6: input the test set into the trained network model to obtain the automatic segmentation result of the lesion of premature brain injury corresponding to the test set. ;
[0010] Preferably, each PAFM module comprises: an axial vision transformer module, a channel attention convolution module parallel to the axial vision transformer module;
[0011] The workflow of the axial vision transformer module specifically includes:
[0012] Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x, expressed as:
[0013] x=concat(encoder1(input1),encoder2(input2)) (1)
[0014] In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c), d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modes; c represents the number of channels; concat represents the concatenation operation;
[0015] Step S1.2: Map the feature map x to x∈(d*w*h,c) through an embedding layer, input the position code pos(x) of x at the same time, and then perform layer normalization, expressed as:
[0016] x=LN(patchembed(x)+pos(x)) (2)
[0017] In formula (2), LN represents the layer normalization operation; patchbed represents the operation of embedding layer mapping on the feature map x;
[0018] Step S1.3: Use the axial attention mechanism to perform self-attention calculation inside each modality to obtain the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x obtained in step S2 to achieve spatial feature enhancement inside each modality, expressed as:
[0019] z=AMHA xyz (x)+AMHA xy (x)+AMHA z (x)+x (3)
[0020] In formula (3), AMHA xyz (x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the 2D slice multi-head attention mechanism; AMHA z (x) represents the vertical multi-head attention mechanism; x represents the feature map x obtained in step S2; z represents the feature after spatial feature enhancement;
[0021] Step S1.4: perform layer normalization on the feature z after spatial feature enhancement, and then process it through a multi-layer perceptron MLP, and then sum the feature processed by the multi-layer perceptron MLP with the feature z after spatial feature enhancement, expressed as:
[0022] z=MLP(LN(z))+z (4)
[0023] In formula (4), LN represents the layer normalization operation;
[0024] The workflow of the channel attention convolution module specifically includes:
[0025] Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c), and then use 3D average pooling to aggregate the features of different channels, expressed as:
[0026] y=AvgPool3D(conv(x)) (5)
[0027] In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling;
[0028] Step S2.2: Use 1x1x1 convolution to interact between different modes of y, and then pass sigmoid to obtain the attention map atty between different features, expressed as:
[0029] atty=sigmoid(conv(y)) (6)
[0030] In formula (6), conv represents the convolution operation; sigmoid represents the activation function:
[0031] Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as:
[0032] y=y*atty (7)
[0033] The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c), and then fuse Z with y obtained in step S2.3 through the residual connection to obtain the output feature output, which is expressed as:
[0034] output=z+y (8)
[0035] Preferably, the workflow of each CMSC module specifically includes:
[0036] Step S3.1: The input is sequentially subjected to the conv 5x5x5 convolution operation, batch normalization operation, and activation function GELU to obtain the output x1, which is expressed as:
[0037] x1=GELU(BN(Conv5x5x5(input)))(9)
[0038] In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0039] Step S3.2: Add the input and output x1, and then perform the conv 3x3x3 convolution operation to output, which is expressed as:
[0040] x2=GELU(BN(Conv3x3x3(input+x1)))(10)
[0041] In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0042] Step S3.3: Add the input and output x2, and then perform the conv 1x1x1 convolution operation to output, which is expressed as:
[0043] x3=GELU(BN(Conv1x1x1(input+x2)))(11)
[0044] In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0045] Step S3.4: concatenate the input, output x1, output x2, and output x3, and then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform a batch normalization operation and activation function GELU to obtain the output output, which is expressed as:
[0046] output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3))))(12)
[0047] In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
[0048] Preferably, the step S2 specifically includes: step S2.1: registering the T1WI and T2WI in the same space so that the marked lesions can be displayed at the same position of T1WI and T2WI; step S2.2: cropping the registered T1WI and T2WI to crop out the black edge areas in the registered T1WI and T2WI; step S2.3: normalizing the grayscale intensity of the cropped T1WI and T2WI; step S2.4: performing data enhancement on the grayscale intensity normalized T1WI and T2WI
[0049] Beneficial effects of the present invention:
[0050] (1) The method for automatic segmentation of premature infant brain damage lesions provided by the present invention proposes a universal model for fully automatic segmentation of premature infant brain damage lesions. It can automatically segment various damage lesions in the premature infant's brain, and its accuracy exceeds that of other basic models in the prior art.
[0051] (2) The method for automatic segmentation of premature infant brain injury lesions provided by the present invention uses T1-weighted and T2-weighted data, uses multiple encoders to fuse each stage, and combines other modules to propose a new multimodal fusion module, namely the PAFM module. Experiments have shown that the PAFM module can effectively utilize multimodal information, thereby improving the segmentation accuracy of premature infant brain injury lesions.
[0052] (3) The automatic segmentation method for premature infant brain injury lesions provided by the present invention makes the segmentation of premature infant brain injury lesions more accurate by combining a segmentation module specifically for small lesions, namely, the CMSC module.
[0053] Additional aspects and advantages of the invention will become apparent from the following description, or may be learned by practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A brain MRI image of a premature infant with brain injury according to an embodiment of the present invention is shown;
[0055] Figure 2 A schematic flow chart showing a method for automatically segmenting premature infant brain damage lesions according to an embodiment of the present invention;
[0056] Figure 3 A schematic diagram showing the overall architecture of a network model for automatic segmentation of premature infant brain damage according to an embodiment of the present invention;
[0057] Figure 4 A schematic diagram showing the structure of a CMSC module according to an embodiment of the present invention is shown;
[0058] Figure 5 A schematic structural diagram of a PAFM module according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0059] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0060] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0061] Figure 2 FIG. 2 is a schematic flow chart showing a method for automatically segmenting premature infant brain damage lesions according to an embodiment of the present invention. Figure 2 As shown, the automatic segmentation method for premature infant brain damage lesions includes:
[0062] Step S1: Acquire a data set; the data set is brain MRI images of premature infants with brain damage;
[0063] Step S2: preprocessing the data set;
[0064] Step S3: Divide the preprocessed data set into a training set, a test set, and a validation set according to a preset ratio;
[0065] Step S4: building a network model for automatic segmentation of premature infant brain damage lesions;
[0066] Step S5: input the training set into the network model for training, and adjust the hyperparameters based on the validation set;
[0067] Step S6: Input the test set into the trained network model to obtain the automatic segmentation result of the premature infant brain injury lesion corresponding to the test set.
[0068] In this embodiment, the sequence of brain MRI images includes T1WI and T2WI. T1WI is T1 weighted imaging (T1weighted image, referred to as T1WI), and T2WI is T2 weighted imaging (T2 weighted image, referred to as T2WI).
[0069] In this embodiment, the lesion regions of interest of the brain MRI image have been marked.
[0070] In this embodiment, if Figure 3As shown in the figure, the automatic segmentation network model of the premature infant brain injury lesion is an encoder-decoder architecture, which includes: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group includes 2 large-scale convolution modules (large kemel projection); the encoder part and the decoder part are four layers; the encoder part of each layer includes 2 sub-encoders, each sub-encoder includes: a Uxnet Block and a Downsampling Block; the decoder part of each layer includes a decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; an output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; an output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the decoder of the corresponding layer through a PAFM module, a CMSC module, and a Res Block in sequence, the other output ends of the encoder parts of the first layer, the second layer, and the third layer are connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected to the decoder of the corresponding layer through a fusion module (Concat module), a Res Block in sequence. Block is connected to the input end of the decoder of the first layer; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer all include a fusion module (Concat module), a Res Block, and a transposed convolution module connected in sequence; and the other output end of the large-scale convolution module group outputs features through a fusion module (Concat module) and a Res Block in sequence, and the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module (Concat module), and then channel adjustment is performed through a Res Block and output convolution to obtain the final lesion segmentation result.
[0071] In this embodiment, the large-scale convolution module is used to map the received preprocessed image data to a low-dimensional space to obtain a spatial potential feature representation of the preprocessed image; each encoder of the encoder part of the first layer, the second layer, or the third layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module; each encoder of the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the decoder of the first layer through a fusion module and a Res Block in sequence; the PAFM module is used to perform fusion enhancement of information between different modalities and within the same modality on the multimodal feature map output by each encoder of the encoder part of the first layer, the second layer, or the third layer, so as to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature map to the corresponding CMSC module; the CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information can be subsequently extracted through a Res Block is sent to the decoder of the corresponding layer for decoding; the decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling, and then pass through a fusion module and a Res Block in sequence, and then upsample the feature map output by the ResBlock through a transposed convolution module; the decoders of the second, third and fourth layers are used to splice and fuse the original level features input by the decoder of the current layer and the deep level features output by the jump connection through a fusion module, and then upsample the feature map output by the fusion module in the decoder of the current layer through a Res Block and a transposed convolution module in sequence; and the large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block in sequence through a fusion module, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer are subsequently spliced and fused through a fusion module, and then channel adjustment is performed through ResBlock and output convolution to obtain the final lesion segmentation result.
[0072] In this embodiment, the CMSC module is a cascade multi-scale convolution (Cascade Muti-Scale Convolution, referred to as CMSC) module, and the PAFM module is a parallel attention fusion module (Parallel attention fusion module, referred to as PAFM module).
[0073] In one embodiment of the present invention, each PAFM module includes: an axial vision transformer module, a channel attention convolution module parallel to the axial vision transformer module; Figure 4 As shown, the workflow of the axial vision transformer module specifically includes:
[0074] Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x, expressed as:
[0075] x=concat(encoder1(input1),encoder2(input2))(1)
[0076] In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c), d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modes; c represents the number of channels; concat represents the concatenation operation;
[0077] Step S1.2: Map the feature map x to x∈(d*w*h,c) through an embedding layer, input the position code pos(x) of x at the same time, and then perform layer normalization, expressed as:
[0078] x=LN(patchembed(x)+pos(x)) (2)
[0079] In formula (2), LN represents the layer normalization operation; patchbed represents the operation of embedding layer mapping on the feature map x;
[0080] Step S1.3: Use the axial attention mechanism to perform self-attention calculation inside each modality to obtain the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x obtained in step S2 to achieve spatial feature enhancement inside each modality, expressed as:
[0081] z=AMHA xyz (x)+AMHA xy (x)+AMHA z (x)+x (3)
[0082] In formula (3), AMHA xyz(x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the 2D slice multi-head attention mechanism; AMHA z (x) represents the vertical multi-head attention mechanism; x represents the feature map x obtained in step S2; z represents the feature after spatial feature enhancement;
[0083] Step S1.4: perform layer normalization on the feature z after spatial feature enhancement, and then process it through a multi-layer perceptron MLP, and then sum the feature processed by the multi-layer perceptron MLP with the feature z after spatial feature enhancement, expressed as:
[0084] z=MLP(LN(z))+z (4)
[0085] In formula (4), LN represents the layer normalization operation;
[0086] The workflow of the channel attention convolution module specifically includes:
[0087] Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c), and then use 3D average pooling to aggregate the features of different channels, expressed as:
[0088] y=AvgPool3D(conv(x)) (5)
[0089] In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling;
[0090] Step S2.2: Use 1x1x1 convolution to interact between different modes of y, and then pass sigmoid to obtain the attention map atty between different features, expressed as:
[0091] atty=sigmoid(conv(y))(6)
[0092] In formula (6), conv represents the convolution operation; sigmoid represents the activation function:
[0093] Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as:
[0094] y=y*atty (7)
[0095] The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c), and then fuse Z with y obtained in step S2.3 through the residual connection to obtain the output feature output, which is expressed as:
[0096] output=z+y (8)
[0097] In one embodiment of the present invention, Figure 5 As shown in the figure, the workflow of each CMSC module includes:
[0098] Step S3.1: The input is sequentially subjected to the con5x5x5 convolution operation, batch normalization operation, and activation function GELU to obtain the output x1, which is expressed as:
[0099] x1=GELU(BN(Conv5x5x5(input))) (9)
[0100] In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0101] Step S3.2: Add the input and output x1, and then perform the conv 3x3x3 convolution operation to output, which is expressed as:
[0102] x2=GELU(BN(Conv3x3x3(input+x1))) (10)
[0103] In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0104] Step S3.3: Add the input and output x2, and then perform the conv 1x1x1 convolution operation to output, which is expressed as:
[0105] x3=GELU(BN(Conv1x1x1(input+x2))) (11)
[0106] In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0107] Step S3.4: concatenate the input, output x1, output x2, and output x3, and then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform a batch normalization operation and activation function GELU to obtain the output output, which is expressed as:
[0108] output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3))))(12)
[0109] In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
[0110] In one embodiment of the present invention, the step S2 specifically includes: step S2.1: aligning the T1WI and T2WI to the same space so that the marked lesions can be displayed at the same position of T1WI and T2WI; step S2.2: cropping the aligned T1WI and T2WI to crop out the black edge areas in the aligned T1WI and T2WI; step S2.3: normalizing the grayscale intensity of the cropped T1WI and T2WI; step S2.4: performing data enhancement on the T1WI and T2WI after the grayscale intensity normalization.
[0111] The technical solution of the present invention will be demonstrated below with a specific embodiment. The specific implementation steps of the method for automatic segmentation of premature infant brain damage lesions of the present invention are as follows:
[0112] (1) Obtaining a data set; wherein the data set is brain MRI images of premature infants with brain injury; the sequence of brain MRI images includes T1WI and T2WI; and the lesion regions of interest of the brain MRI images have been marked.
[0113] The data set of this specific embodiment mainly comes from 267 patients with premature brain injury in Shengjing Hospital, including 164 male children and 103 female children. The gestational age of the children was 24 to 36 weeks, and the average gestational age was (32.5±1.5) weeks. A total of 1,785 injury lesions were extracted, and the same lesion was only counted once when displayed in different sequences.
[0114] Brain MRI images of patients with premature brain injury were obtained using a 1.5T GE Signal HDxt system scanner equipped with an 8-channel phased array head coil. The subjects were asked to lie supine with their heads first into the scanner and remain asleep to complete the examination. The MRI scanning sequence included T1WI and T2WI, and the parameters were as follows: T1WI (TR = 2950ms, TE = 24ms, slice thickness 4mm, field of view 22×22cm, matrix 288x192); T2WI (TR = 5000ms, TE = 120ms, slice thickness 4mm, field of view 22×22cm, matrix 320x320).
[0115] After scanning, the images were manually marked with ROIs of lesions by radiologists on the SRhythmMultiLabel platform under the same window width and window position. The radiologists involved in the marking were specially trained. The marking process was divided into two stages: in the first stage, two attending physicians manually outlined the visually obvious lesion contours, while referring to other sequence images; when marking, they identified and avoided the normal white matter high water content areas with increased white matter signals on T1WI images but not low signals on T2WI sequences; in the second stage, two associate senior physicians or above reviewed and verified the markings based on the first stage, and supplemented and corrected the missed and mislabeled lesions until they reached a consensus on the results when marking the same patients.
[0116] (2) Preprocessing the data set. Preprocessing the data set specifically includes: registering T1WI and T2WI to the same space so that the marked lesions can be displayed at the same position of T1WI and T2WI; cropping the registered T1WI and T2WI to crop out the black edge areas in the registered T1WI and T2WI; normalizing the grayscale intensity of the cropped T1WI and T2WI; and performing data enhancement on the grayscale intensity normalized T1WI and T2WI.
[0117] Specifically, the original size of the T1 image we used was slightly different, and the size of the T2 image was not completely consistent with the T1 image. Since infants often shake their heads during MRI scans, the spatial positions of the T2-weighted image and the T1-weighted image are also different, which can cause problems in the input network. To solve this problem, we first used the rigid registration method to register the T2 image to the space where the T1-weighted image is located, so that the lesions marked by the doctor can be displayed in the same position of the T1-weighted image and the T2-weighted image.
[0118] Next, we mainly use the preprocessing method in the monai package. First, we crop the black edges of MRI. MRI often has a large number of black edge pixels outside the brain. These edge pixels are noise for model training, which will affect the training speed and training effect of the model. We first crop them. Then we normalize the grayscale value intensity. The pixel values of T1-weighted images and T2-weighted images are distributed in the range of (0,2000). We linearly map the pixel values of T1-weighted images and T2-weighted images from the interval of (0,2000) to the interval of (0,1), so that the model can converge better. Data enhancement methods include random rotation of images, random movement of grayscale value intensity, etc., and these methods are used to expand the data to fully improve the generalization of the model. After data enhancement, in order to analyze the images of different patients more accurately, we uniformly adjust the image size values after preprocessing to (336,336,32).
[0119] (3) The dataset is divided into 8:2 parts, 20% is used as the test set, 10% of the remaining 80% is extracted as the validation set for adjusting the hyperparameters, and the rest is used as the training set.
[0120] (4) Build a network model for automatic segmentation of premature brain injury lesions. In terms of the model, our baseline model uses an excellent medical image segmentation model 3D-UXnet, which has achieved sota results in multiple medical image segmentation tasks (including brain segmentation and organ segmentation). We migrated it to our premature brain injury lesion segmentation task and modified the model to make it more suitable for this task. We reduced the parameters of the model and removed a skip connection layer. The modules we added to the premature brain injury lesion segmentation mainly include two, namely cascaded multi-scale convolution (CMSC) and parallel attention fusion module (PAFM). The cascaded multi-scale convolution is mainly used to segment small lesions in the premature brain, such as cysts, and the parallel attention fusion module is mainly used to complete the complementary fusion of multimodal information and use multimodal information to improve the accuracy of segmentation.
[0121] The following explains why the CMSC module was added. Because the brain injury lesions of premature infants are generally small and have no fixed distribution characteristics, they are often distributed in multiple areas of the brain. In the process of extracting downsampling feature information from the network, the size of the feature map decreases as the network deepens. The feature maps generated in the deeper layers are more abstract, so it is usually easier to lose important information of some small objects, affecting the detection of small lesions of premature brain injury. To meet this challenge, we used a feature extraction module that specifically deals with small targets.
[0122] The following explains why the PAFM module was added. In the usual clinical diagnosis of premature encephalopathy, both T1-weighted images and T2-weighted images can play a certain role. Doctors usually judge the type or severity of encephalopathy based on the abnormal signals of T1-weighted images and T2-weighted images. For example, the MRI manifestation of leukomalacia (PVL) is low signal in T1W images and high signal in T2W images in the white matter around the lateral ventricle. The signal changes in the posterior limb of the internal capsule on T1 and T2 weighted images often involve selective neuronal necrosis. Both T1 and T2 weighted images have corresponding features of premature encephalopathy. Lesions that are not obvious on T1-weighted images may have more obvious features on T2-weighted images, and lesions often show the most important features on T1-weighted images. Inspired by this, we use T1 mode as the main mode and T2 mode as the auxiliary mode to fuse the data of the two modes to segment premature brain lesions. There are many methods for multimodal fusion of medical images, which are mainly divided into input-level fusion network, hierarchical fusion network and decision-level network. Since the modal features of T1-weighted images and T2-weighted images corresponding to brain injury lesions in premature infants are relatively complex and contain a lot of similar or complementary information, we use a hierarchical fusion strategy to perform multi-level fusion, so as to obtain multi-scale context features of different modalities to help achieve robust lesion segmentation. Therefore, we developed the PAFM module.
[0123] Specifically, Figure 3As shown, the automatic segmentation network model of premature brain injury lesions is an encoder-decoder architecture, and the network model includes: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group includes 2 large-scale convolution modules; the encoder part and the decoder part are both four layers; the encoder part of each layer includes 2 sub-encoders, each sub-encoder includes: a Uxnet Block and a Downsampling Block; the decoder part of each layer includes a sub-decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; an output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; an output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the decoder of the corresponding layer through a PAFM module, a CMSC module, and a Res Block in sequence, the other output ends of the encoder parts of the first layer, the second layer, and the third layer are connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected to the decoder of the corresponding layer through a fusion module, a Res Block in sequence. Block is connected to the input end of the decoder of the first layer; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer each include a fusion module, a Res Block, and a transposed convolution module connected in sequence; and the other output end of the large-scale convolution module group outputs features through a fusion module and a Res Block in sequence, and the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module, and then channel adjustment is performed through a Res Block and output convolution to obtain the final lesion segmentation result.
[0124] The large-scale convolution module is used to map the received preprocessed image to a low-dimensional space to obtain the spatial potential feature representation of the preprocessed image; each encoder of the encoder part of the first layer, the second layer, or the third layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module; each encoder of the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map through a fusion module, a Res Block is sent to the decoder of the first layer; the PAFM module is used to fuse and enhance the multimodal feature map output by each encoder of the encoder part of the first layer, the second layer, or the third layer between different modalities and the internal information of the same modality to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature map to the corresponding CMSC module; the CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information can be subsequently sent to the decoder of the corresponding layer through a ResBlock for decoding; the decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling, and then pass through a fusion module and a Res Block in turn, and then upsample the feature map output by the Res Block through a transposed convolution module; the decoders of the second, third, and fourth layers are used to splice and fuse the original level features input by the jump connection and the deep level features output by the decoder of the previous layer through a fusion module, and then pass through a Res Block in turn. Block, a transposed convolution module upsamples the feature map output by the fusion module in the decoder of the current layer; and the large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block through a fusion module in sequence, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer are subsequently spliced and fused through a fusion module, and then channel adjustment is performed through Res Block and output convolution to obtain the final lesion segmentation result.
[0125] In general, our model uses multiple encoders to fully extract deep features of multiple modalities. Each encoder uses a 3D-UXNET encoder with reduced parameters. Our input is a uniformly divided image block of size (64, 64, 32), which mainly includes information of two modalities, T1W and T2W. The input first maps the original image to a low-dimensional space through a large-scale convolution to obtain the spatial potential feature representation of the image. Then, the image is downsampled through 4 Uxnet Blocks and downsampling blocks. The uxnet blocks of different modal encoders and the multimodal feature maps obtained after downsampling are fused between different modalities and within the same modality through the PAFM module. The fused and enhanced feature maps are then input into the cascaded multi-scale convolution module to fully extract small target (lesion) information. The output of the cascaded multi-scale convolution is then sent to the decoder for decoding through the original 3D-UXNET Res Block. The decoder uses the decoder of 3D-UXNET to fuse the original level features and deep level features from the jump connection, and uses transposed convolution for upsampling. The model can obtain the upsampling method that best suits the current data set by learning the parameters in the transposed convolution, which is better than the unpooling upsampling. The downsampled feature map of each layer is spliced and fused with the feature map of the previous layer to enrich the multi-level information of the image. At the top layer, the feature map obtained by the large-scale convolution mapping and the final feature map obtained by the downsampling layer are spliced and fused, and finally output through a Res Block and an output convolution block.
[0126] Furthermore, the parallel attention fusion module (PAFM) is a module created by referring to and modifying the channel attention convolution ECA and NestFormer. The fusion module is composed of an axial vision transformer module in NestFormer and an ECA channel attention convolution module in parallel. It is mainly used to calculate the spatial correlation within each modality and the correlation between different modalities. The module uses the channel attention mechanism and the intra-modal spatial attention mechanism in parallel to fully extract the information within the modality and the complementary information across modalities, and realizes the fusion operation between multi-modal information to enhance the key features of the input feature map. The channel attention mechanism is mainly for extracting the correlation between modalities, and the axial vit module mainly extracts the correlation between each modality, and then fuses them for output.
[0127] Specifically, each PAFM module includes: an axial vision transformer module, a channel attention convolution module parallel to the axial vision transformer module; Figure 4As shown, the workflow of the axial vision transformer module specifically includes:
[0128] Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x, expressed as:
[0129] x=concat(encoder1(input1),encoder2(input2))(1)
[0130] In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c), d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modes; c represents the number of channels; concat represents the concatenation operation;
[0131] Step S1.2: Map the feature map x to x∈(d*w*h,c) through an embedding layer, input the position code pos(x) of x at the same time, and then perform layer normalization, expressed as:
[0132] x=LN(patchembed(x)+pos(x)) (2)
[0133] In formula (2), LN represents the layer normalization operation; patchbed represents the operation of embedding layer mapping on the feature map x;
[0134] Step S1.3: Use the axial attention mechanism to perform self-attention calculation inside each modality to obtain the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x obtained in step S2 to achieve spatial feature enhancement inside each modality, expressed as:
[0135] z=AMHA xyz (x)+AMHA xy (x)+AMHA z (x)+x (3)
[0136] In formula (3), AMHA xyz (x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the 2D slice multi-head attention mechanism; AMHA z(x) represents the vertical multi-head attention mechanism; x represents the feature map x obtained in step S2; z represents the feature after spatial feature enhancement;
[0137] Step S1.4: perform layer normalization on the feature z after spatial feature enhancement, and then process it through a multi-layer perceptron MLP, and then sum the feature processed by the multi-layer perceptron MLP with the feature z after spatial feature enhancement, expressed as:
[0138] z=MLP(LN(z))+z (4)
[0139] In formula (4), LN represents the layer normalization operation;
[0140] In the workflow of the axial vision transformer module, the attention mechanism is used to perform self-attention calculations within each modality. Since each downsampled feature map needs to be fused, in order to optimize computational efficiency and reduce model complexity, the module uses an axial attention mechanism, which is a method of decomposing multidimensional data into multiple axes (vertical, slice, window) and applying attention mechanisms on these axes respectively. This method significantly reduces the number of parameters and computational costs while maintaining performance.
[0141] The workflow of the channel attention convolution module specifically includes:
[0142] Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c), and then use 3D average pooling to aggregate the features of different channels, expressed as:
[0143] y=AvgPool3D(conv(x)) (5)
[0144] In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling;
[0145] Step S2.2: Use 1x1x1 convolution to interact between different modes of y, and then pass sigmoid to obtain the attention map atty between different features, expressed as:
[0146] atty=sigmoid(conv(y)) (6)
[0147] In formula (6), conv represents the convolution operation; sigmoid represents the activation function:
[0148] Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as:
[0149] y=y*atty (7)
[0150] The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c), and then fuse Z with y obtained in step S2.3 through the residual connection to obtain the output feature output, which is expressed as:
[0151] output=z+y (8)
[0152] In the workflow of this channel attention convolution module, we perform cross-modal channel attention calculations on features of different modalities to fully integrate features of different modalities.
[0153] Furthermore, if Figure 5 As shown, each CMSC module is specifically used for:
[0154] Step S3.1: The input is sequentially subjected to the conv 5x5x5 convolution operation, batch normalization operation, and activation function GELU to obtain the output x1, which is expressed as:
[0155] x1=GELU(BN(Conv5x5x5(input))) (9)
[0156] In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0157] Step S3.2: Add the input and output x1, and then perform the conv 3x3x3 convolution operation to output, which is expressed as:
[0158] x2=GELU(BN(Conv3x3x3(input+x1))) (10)
[0159] In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0160] Step S3.3: Add the input and output x2, and then perform the conv 1x1x1 convolution operation to output, which is expressed as:
[0161] x3=GELU(BN(Conv1x1x1(input+x2))) (11)
[0162] In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function;
[0163] Step S3.4: concatenate the input, output x1, output x2, and output x3, and then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform a batch normalization operation and activation function GELU to obtain the output output, which is expressed as:
[0164] output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3)))) (12)
[0165] In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
[0166] The cascaded multi-scale convolution (CMSC) module is a module in w-net. This module mainly extracts multi-level information of the original feature map through convolutions of different sizes (convolution kernels of length 5, 3, and 1). Through the multi-scale convolution extraction and fusion of each CMSC module, the semantics contained in the deep layer of the feature map and the shallow semantic information can be well extracted at multiple scales, and the multi-scale information is passed to the decoder through jump connections, so that the model can capture a wider range of contextual information while focusing on local details. This combination of global and local information helps the model better understand the image, making the segmentation process more refined for the lesions of premature encephalopathy and better retaining the boundary information. At the same time, for lesions of different sizes, multi-scale convolution can capture the characteristics of large and small lesions at the same time, improving the detection and recognition capabilities of the model.
[0167] (5) Input the training set into the network model for training, and use the validation set to adjust the hyperparameters.
[0168] All experiments were conducted on an NVIDIA RTX 2080Ti GPU with 11GB memory. Due to the limitation of GPU memory size, we uniformly cut the images into small blocks of size (64, 64, 32) after preprocessing and data augmentation and input them into the network. During training, we used the AdamW optimizer, batch_size was set to 3, the initial learning rate was set to 0.001, and the learning rate adjustment during training adopted the cosine annealing strategy, which can continuously jump out of the local optimal solution of the model to find the global optimal solution. The cycle was 150 epochs, the minimum learning rate was set to 1e-5, and the loss function used a combination of dice loss and cross entropy loss. The dice loss and cross entropy loss coefficients were 0.5 and 0.5 respectively, keeping the contribution of both to the loss the same. All models were terminated after 600 epochs. The sliding window inference strategy was used for inference, and the window size was consistent with the block size during training, which was also (64, 64, 32).
[0169] (6) Select evaluation indicators.
[0170] For the segmentation evaluation of premature infant brain injury lesions, we mainly choose Dice coefficient (DSC) and Hausdorff distance (HD) to evaluate the quality of our model segmentation. The Dice coefficient is a set similarity measurement function, which was originally used to judge the similarity of two sets. In semantic segmentation, it is mainly used to measure the overlap between two regions. Its value is a number between 0 and 1. The closer it is to 1, the greater the overlap, and the closer it is to 0, the smaller the overlap. We use it to evaluate the overlap between our segmentation results and the original labels, so as to judge the performance of our model. The calculation formula is as follows:
[0171]
[0172] In the above formula, A and B are often pixel matrices composed of 0 and 1. The area corresponding to the pixel value of 1 represents the foreground area, and the area corresponding to the pixel value of 0 represents the background area. |A| represents the sum of the number of pixels with a value of 1 in the pixel matrix A. |A∩B| is often achieved by multiplying the corresponding elements of the pixel matrix A and the pixel matrix B.
[0173] Hausdorff distance is also an indicator used to measure the proximity of the boundaries of two sets. In semantic segmentation, it is usually used to measure the accuracy of boundary segmentation. The closer the HD distance is to 0, the closer the boundaries of the two sets are, and the better the segmentation effect. Its calculation formula is as follows:
[0174] h(A,B)=max a∈A min b∈B ||ab||
[0175] h(B,A)=max b∈B min a∈A ||ba||
[0176] H(A,B)=max(h(A,B),h(B,A))
[0177] ‖ab‖ represents the distance paradigm between point sets a and b, and h(A, B) is called the one-way Hausdorff distance from A to B. It can be seen that the Hausdorff distance is often the larger value of the one-way distance between the two point sets. A and B represent the segmented area point set and the point set in the label, that is, the points with a pixel value of 1. It is worth noting that in order to eliminate the influence of outliers on the evaluation effect, when calculating the one-way Hausdorff distance between two point sets, the 95% digit value, that is, the 95% value in ascending order, is taken as the actual one-way Hausdorff distance. We call the HD distance calculated in this way HD95, and HD95 is often used in practice to judge the segmentation performance of the model.
[0178] (7) The test set is input into the trained network model to obtain the automatic segmentation results of premature infant brain damage lesions corresponding to the test set.
[0179] (8) To verify the superiority of the network model for automatic segmentation of lesions of premature infants with brain damage according to the present invention.
[0180] In order to compare the superiority of our proposed model in the segmentation of premature brain injury lesions, we compared it with several classic models of medical image segmentation under the same experimental conditions, including some CNN-based methods such as 3DU-Net, U-Net++, VNet, SegResNet, Attention-UNet, and some methods based on the combination of CNN and VIT such as UNETR, SwinUNETR, and the original 3D-UXNET and the 3D-UXNET with reduced parameters and modified layers. We call the 3D-UXNET with reduced parameters and modified layers mini-UXNET. The evaluation indicators mainly use the Dice (%) coefficient and HD95 (mm) introduced above, and the results are shown in the following table.
[0181] Table 1Comparative experiments with different segmentation algorithms
[0182]
[0183] In the test dataset of premature infant brain injury lesions we established, our modified model showed the most superior segmentation performance. For mini-UXNET, which we obtained after reducing parameters and modifying the number of layers, its parameter volume is about half of the original 3D-UXNet, and the calculation volume is reduced by three quarters. At the same time, the accuracy is only slightly lower than 3D-UXNet, which proves that reducing the parameters of the original model is a meaningful change. While keeping the original accuracy almost unchanged, the model parameters and calculation volume are greatly reduced. At the same time, our subsequent models are all modified based on mini-UXNET. Compared with the original 3D-UXNet, our model achieved an average Dice value of 0.7583 and an average HD95 value of 10.23mm on the test set. Compared with the original 3D-UXNet baseline model without adding modules, Dice has increased by 2.76%, HD95 has decreased by 4.83mm, and the parameter volume and calculation volume are also lower than the original 3D-UXNet baseline model. At the same time, compared with other segmentation models, it also achieved the best performance in both Dice value and HD95 value. HD95 is often used to measure the numerical value of the boundary between two regions and is more sensitive than the DICE value. Therefore, Dice is usually used as the main indicator and HD95 as a secondary reference indicator. At the same time, our model has 38.90M parameters and 41.79G FLOPs. When using multiple encoders, the number of parameters and the amount of calculation are still lower than UNETR, SwinUNETR, V-Net and 3D-UXNet, and it is still a medium-sized model. It has only 10.92M more parameters and 17.47G more calculations than the reduced parameter 3D-UXNET, but has achieved a significant improvement in segmentation accuracy. The improvement of the above indicators proves that our model can better extract some deep-level features, which is more beneficial for the segmentation of small lesions.
[0184] In summary, we studied the automatic segmentation of brain injury in premature infants (BIPI) lesions, using the excellent medical image segmentation model 3D-UXNet, and added a multi-scale convolutional extraction module (CMSC) on its basis to make it more suitable for the characteristics of small brain injury in premature infants (BIPI) lesions for fine segmentation. At the same time, referring to the doctor's diagnosis method, our model also uses T1-weighted image and T2-weighted image data in two modalities, using the hierarchical modality fusion strategy, and adding a parallel attention fusion module that can perform parallel extraction and fusion of modality internal spatial attention and cross-modal attention in the hierarchical fusion to fuse multimodal information. Experiments have shown that our improved model achieves the best effect in brain injury in premature infants (BIPI) lesion segmentation compared to other models, and can better segment some difficult lesions, better helping doctors complete the task of diagnosing brain injury in premature infants (BIPI) lesions.
[0185] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for automatic segmentation of premature infant brain damage lesions, characterized in that: include: Step S1: Acquire a data set; wherein the data set is a brain MRI image of a premature infant with brain injury; the sequence of the brain MRI image includes T1WI and T2WI; the lesion regions of interest of the brain MRI image have been marked; Step S2: preprocessing the data set; Step S3: Divide the preprocessed data set into a training set, a test set, and a validation set according to a preset ratio; Step S4: constructing a network model for automatic segmentation of premature infant brain injury lesions; wherein the network model is an encoder-decoder architecture, and the network model includes: a large-scale convolution module group, an encoder part, a PAFM module group, a CMSC module group and a decoder part; the large-scale convolution module group includes 2 large-scale convolution modules; the encoder part and the decoder part are both four layers; the encoder part of each layer includes 2 sub-encoders, each sub-encoder includes: a UxnetBlock and a Downsampling Block; the decoder part of each layer includes a sub-decoder; the PAFM module group includes three PAFM modules; the CMSC module group includes three CMSC modules; An output end of the large-scale convolution module group is connected to the input end of the encoder part of the first layer; an output end of the encoder part of the first layer, the second layer, and the third layer is connected to the input end of the decoder of the corresponding layer through a PAFM module, a CMSC module, and a Res Block in sequence, the other output ends of the encoder part of the first layer, the second layer, and the third layer are connected to the input end of the encoder part of the corresponding next layer, and the output end of the encoder part of the fourth layer is connected to the input end of the decoder of the first layer through a fusion module and a Res Block in sequence; the decoder of the first layer includes a transposed convolution module; the decoders of the second layer, the third layer, and the fourth layer each include a fusion module, a Res Block, and a transposed convolution module connected in sequence; and the other output end of the large-scale convolution module group is sequentially output through a fusion module and a Res Block for feature output, the feature map output by the Res Block is spliced and fused with the feature map obtained by the decoder of the fourth layer through a fusion module, and then channel adjustment is performed through Res Block and output convolution to obtain the final lesion segmentation result; The large-scale convolution module is used to map the received pre-processed image data to a low-dimensional space to obtain a spatial potential feature representation of the corresponding image after pre-processing; Each encoder of the encoder part of the first layer, the second layer, or the third layer is used to downsample the input image through a UxnetBlock and a Downsampling Block in sequence, and then send the downsampled feature map to the corresponding PAFM module; Each encoder in the encoder part of the fourth layer is used to downsample the input image through a Uxnet Block and a Downsampling Block in sequence, and then send the downsampled feature map to the decoder of the first layer through a fusion module and a Res Block in sequence; The PAFM module is used to perform fusion enhancement of information between different modalities and within the same modality on the multimodal feature map output by each encoder of the encoder part of the first layer, the second layer, or the third layer, so as to complete the complementary fusion of multimodal information, and then input the fused and enhanced feature map into the corresponding CMSC module; The CMSC module is used to receive the fused and enhanced feature map to extract the feature map corresponding to the small lesion information, so that the feature map corresponding to the small lesion information is subsequently sent to the decoder of the corresponding layer through a Res Block for decoding; The decoder of the first layer is used to directly receive the feature map output by the encoder of the fourth layer after downsampling through a fusion module and a Res Block in sequence, and then upsample the feature map output by the Res Block through a transposed convolution module; The decoders of the second, third, and fourth layers are used to concatenate and fuse the original hierarchical features input by the jump connection and the deep hierarchical features output by the decoder of the previous layer through a fusion module, and then upsample the feature maps output by the fusion module in the decoder of the current layer through a Res Block and a transposed convolution module in sequence; The large-scale convolution module is also used to send the feature map obtained by its own mapping to a Res Block through a fusion module in sequence, so that the feature map output by the Res Block and the feature map obtained by the decoder of the fourth layer are subsequently spliced and fused through a fusion module, and then channel adjustment is performed through the Res Block and the output convolution to obtain the final lesion segmentation result; Step S5: inputting the training set into the network model for training, and adjusting the hyperparameters based on the validation set; Step S6: inputting the test set into the trained network model to obtain the automatic segmentation result of the premature infant brain injury lesion corresponding to the test set.
2. According to the method for automatic segmentation of premature infant brain damage lesions according to claim 1, each PAFM module comprises: An axial vision transformer module and a channel attention convolution module in parallel with the axial vision transformer module; The workflow of the axial vision transformer module specifically includes: Step S1.1: Concatenate the feature maps output by the encoder part of each layer in the channel dimension to obtain the feature map x, expressed as: x=concat(encoder1(input1),encoder2(input2))(1) In formula (1), encoder1 and encoder2 represent the feature maps of the same layer output by using a separate encoder to process multimodal data; input1 represents the input of the first encoder of each layer; input2 represents the input of the second encoder of each layer; the output x∈(d,w,h,n*c), d represents the length of the feature block; w represents the width of the feature block; h represents the height of the feature block; n represents the number of modes; c represents the number of channels; concat represents the concatenation operation; Step S1.2: Map the feature map x to x∈(d*w*h,c) through an embedding layer, input the position code pos(x) of x at the same time, and then perform layer normalization, expressed as: x=LN(patchembed(x)+pos(x))(2) In formula (2), LN represents the layer normalization operation; patchbed represents the operation of embedding layer mapping on the feature map x; Step S1.3: Use the axial attention mechanism to perform self-attention calculation inside each modality to obtain the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction, respectively. Then sum the axial attention in the one-dimensional vertical direction, the two-dimensional slice direction, and the three-dimensional window direction with the feature map x obtained in step S2 to achieve spatial feature enhancement inside each modality, expressed as: z=NOW xyz (x)+NO xy (x)+NO z (x)+x(3) In formula (3), AMHA xyz (x) represents the 3D window multi-head attention mechanism; AMHA xy (x) represents the 2D slice multi-head attention mechanism; AMHA z (x) represents the vertical multi-head attention mechanism; x represents the feature map x obtained in step S2; z represents the feature after spatial feature enhancement; Step S1.4: perform layer normalization on the feature z after spatial feature enhancement, and then process it through a multi-layer perceptron MLP, and then sum the feature processed by the multi-layer perceptron MLP with the feature z after spatial feature enhancement, expressed as: z=MLP(LN(z))+z(4) In formula (4), LN represents the layer normalization operation; The workflow of the channel attention convolution module specifically includes: Step S2.1: Use a 1x1x1 convolutional layer to map the concatenated input x∈(d,w,h,n*c) to x∈(d,w,h,c), and then use 3D average pooling to aggregate the features of different channels, expressed as: y = AvgPool3D(conv(x))(5) In formula (5), conv represents the convolution operation; AvgPool3D represents the operation of aggregating features of different channels using 3D average pooling; y represents the output feature after aggregating features of different channels using 3D average pooling; Step S2.2: Use 1x1x1 convolution to interact between different modes of y, and then pass sigmoid to obtain the attention map atty between different features, expressed as: atty=sigmoid(conv(y))(6) In formula (6), conv represents the convolution operation; sigmoid represents the activation function: Step S2.3: Multiply y by the attention map atty to enhance the features between different channels, expressed as: y=y*atty(7) The residual connection is specifically used to fold the output of the axial vision transformer module to z∈(d,w,h,c), and then fuse Z with y obtained in step S2.3 through the residual connection to obtain the output feature output, which is expressed as: output=z+y(8) 3. The method for automatic segmentation of premature infant brain damage lesions according to claim 1, characterized in that: The workflow of each CMSC module includes: Step S3.1: The input is sequentially subjected to the conv 5x5x5 convolution operation, batch normalization operation, and activation function GELU to obtain the output x1, which is expressed as: x1=GELU(BN(Conv5x5x5(input)))(9) In formula (9), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function; Step S3.2: Add the input and output x1, and then perform the conv 3x3x3 convolution operation to output, which is expressed as: x2=GELU(BN(Conv3x3x3(input+x1)))(10) In formula (10), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function; Step S3.3: Add the input and output x2, and then perform the conv 1x1x1 convolution operation to output, which is expressed as: x3=GELU(BN(Conv1x1x1(input+x2)))(11) In formula (11), input is the feature map x∈(c,h,w,d) output by the PAFM module; BN represents batch normalization operation; GELU represents activation function; Step S3.4: concatenate the input, output x1, output x2, and output x3, and then perform a conv1x1x1 convolution operation to change the number of convolution channels, and then perform a batch normalization operation and activation function GELU to obtain the output output, which is expressed as: output=GELU(BN(Conv1x1x1(concat(input,x1,x2,x3))))(12) In formula (12), input is the feature map x∈(c,h,w,d) output by the PAFM module; concat represents the concatenation operation; BN represents the batch normalization operation; GELU represents the activation function.
4. The method for automatic segmentation of premature infant brain damage according to any one of claims 1 to 3, characterized in that: The step S2 specifically includes: Step S2.1: registering the T1WI and T2WI into the same space so that the marked lesion can be displayed at the same position of T1WI and T2WI; Step S2.2: cropping the registered T1WI and T2WI to crop out the black edge areas in the registered T1WI and T2WI; Step S2.3: normalize the grayscale intensity of the cropped T1WI and T2WI; Step S2.4: Perform data enhancement on T1WI and T2WI after grayscale value intensity normalization.
Citation Information
Patent Citations
White matter high-signal segmentation method based on multi-scale fusion and attention splitting
CN114119637A
Portrait segmentation network training method and device, equipment and medium
CN114549557A
Three-dimensional heart image segmentation method based on multi-scale edge perception
CN114663445A
Brain tumor multi-mode MRI image segmentation method based on deep learning
CN114926477A
Medical image segmentation method based on multi-scale cross-layer attention fusion network
CN117152433A
Cited By
Brain tumor image segmentation method based on lightweight multimodality
CN120689351A
Mortise broach wear state identification method
CN120995069A