Brain tissue image segmentation method based on multi-modal fusion and feature alignment

Through the brain tissue image segmentation method of multimodal fusion and feature alignment, the problems of data imbalance, difficulty in extracting multimodal information and low image contrast in infant brain MR image segmentation are solved, and efficient and accurate brain tissue segmentation is achieved, suitable for medical images with low resolution or motion artifacts.

CN120339612APending Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400110.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has problems such as data imbalance, difficulty in extracting multimodal information, and low image contrast in infant brain MR image segmentation, resulting in deviations in segmentation results and blurred boundaries, especially in the scenarios where gray matter and white matter contrast in infant brain tissue are insufficient.

Method used

The brain tissue image segmentation method is adopted with multimodal fusion and feature alignment. Single-modal boundary contour information is extracted through the Sobel boundary detection algorithm, combined with multi-scale feature extraction and cross-branch connection, a feature alignment module is introduced, and the model is trained using the mixed loss function of Dice Loss and Binary CrossEntropy to achieve end-to-end automatic segmentation.

Benefits of technology

It significantly improves the refinement ability of brain tissue boundaries, solves the problem of boundary blur caused by low image contrast, enhances the fusion and segmentation robustness of multimodal information, restores the detailed information of the original image, reduces time cost and subjective errors, and is suitable for medical images with low resolution or motion artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339612A_ABST
    Figure CN120339612A_ABST
Patent Text Reader

Abstract

The invention relates to a brain tissue image segmentation method based on multi-modal fusion and feature alignment, and belongs to the field of medical image processing. Comprising the following steps: S1, acquiring a brain tissue image data set, dividing the data set into a training set and a test set, and performing corresponding preprocessing; s2, extracting boundary contour information of the image by using a boundary detection algorithm Sobel, fusing the boundary information into the network model by using data fusion, and improving boundary refinement and learning of the network model; s3, using cross-branch connection to promote information flow between different modes, and effectively fusing and extracting multi-mode features; s4, a feature alignment module is introduced into up-sampling, and up-sampling later features with correct spatial positions and accurate boundary areas are generated; and S5, through comparison with a label, calculating loss and carrying out back propagation on the training model parameters until the model parameters converge. According to the method, the brain tissue boundary information can be enhanced under the condition that other tasks are not introduced, and the learning of the network model on the brain tissue boundary information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image processing, and relates to a method for segmenting brain tissue images with multi-modal fusion and feature alignment. Background Art

[0002] Developmental brain diseases (such as autism and mental retardation), neuropsychiatric diseases (such as depression and addiction), and neurological brain diseases (such as Alzheimer's disease and Parkinson's disease) occur relatively early, and are manifested in infancy or childhood. Autism Spectrum Disorder (ASD) is particularly obvious, but there are few early biomarkers. We often need to rely on factors such as behavioral manifestations and family history to infer whether a child has a risk of ASD, and ASD can only be diagnosed at the age of 3-4. It is estimated that about one in every 100 children in the world suffers from autism. Therefore, studying the development of the infant brain helps to evaluate the risk of future infant brain neurodevelopment and developmental neuropsychiatric diseases. An important step in studying normal and abnormal early brain development is to accurately segment infant brain MR images into different regions of interest (ROI). For example, accurately segmenting the brain into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF).

[0003] As an indispensable part of medical auxiliary diagnosis, the development of related imaging technologies in medical imaging technology is changing with each passing day. Imaging technologies such as Computed Tomography (CT), X-ray, and Magnetic Resonance Imaging (MRI) have become the mainstream medical imaging technologies, and have played an important role in disease diagnosis and treatment plan formulation. Nuclear magnetic resonance imaging has the characteristics of high contrast, multi-directional slicing, and multi-parameter imaging, and has relatively high spatial resolution. It can perform high-resolution imaging on soft tissues such as the nervous system and the brain, and is a non-invasive imaging technology that does not cause ionizing radiation damage to the human body. It is widely used in infant brain research. At present, the manual voxel annotation of MR images by professional doctors is still regarded as the gold standard in the industry. However, this work is too time-consuming and requires a lot of energy. Interferences from factors such as noise and motion artifacts also increase the difficulty of data annotation, making it non-reproducible. At the same time, due to the differences in the professional levels of doctors and the influence of subjective factors, there may be significant differences in the segmentation results of the same MR image.

[0004] With the rapid improvement of computer hardware performance and the continuous advancement of algorithm optimization, computer-aided diagnosis technology has gradually come into use, which can effectively reduce the subjectivity of decision-making and the corresponding total cost. Therefore, it is of great significance and application value to study the end-to-end automatic infant brain segmentation algorithm to assist doctors in diagnosis and improve the segmentation efficiency. In the field of computer vision, Deep Convolutional Neural Networks (DCNN) is a relatively popular model. Its convolutional layer has the advantages of translational invariance, parameter sharing to reduce the number of network parameters, and having spatial position information. Among them, the pooling layer mainly realizes the downsampling of the feature map, reducing the computational amount while increasing the receptive field of the feature map. Based on these characteristics, DCNN has a wide range of applications in the field of computer vision. In the current methods of brain tissue segmentation, the main problems are as follows: 1. Data imbalance: Medical image data is usually highly imbalanced, with a large difference in the number of pixels in different categories. Conventional feature extraction methods have the problems of small receptive fields and local feature information, which will lead to deviation in the segmentation results; 2. Difficulty in extracting multi-modal information: MRI has two modalities, T1 and T2. The imaging principles of the two modalities are different, so the advantages and disadvantages of the two images are also different (T1 is mainly used to observe anatomical structures, and T2 is used to observe lesion site information); 3. Low image contrast: Due to the particularity of the 6-9 month period, the myelin of the brain tissue gradually matures, and the change in tissue signal intensity results in a low gray contrast between GM and WM at this stage, with blurred boundaries, making it difficult to accurately predict pixels. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method for segmenting brain tissue images with multi-modal fusion and feature alignment.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A method for segmenting brain tissue images with multi-modal fusion and feature alignment specifically includes the following steps:

[0008] S1: Data preparation stage: Obtain a brain tissue image dataset, divide the dataset into a training set and a test set, and perform corresponding preprocessing;

[0009] S2: Multi-scale feature extraction stage: Use the traditional edge detection algorithm Sobel operator to extract the edge contour information of the single modality of the image, use data fusion to integrate the edge information into another modality, reduce the redundancy of feature fusion, perform multi-scale feature extraction on the other modality, and better extract features through multi-scale convolution and adaptive weights;

[0010] S3: Modal Fusion Attention Stage: Two independent multi-modal feature branches are designed, and cross-branch connections are used to facilitate the information flow between different modalities, which can effectively fuse and refine multi-modal features. Then, spatial attention is applied to each branch, and finally, the spatial information generated by each branch is averaged and fused;

[0011] S4: Upsampling Feature Alignment Stage: The spatial resolution of the feature map is increased by upsampling to help restore the detailed information of the original input. A feature alignment module is introduced, and a learnable volumetric feature deformation field is used for feature alignment to generate the upsampled late-stage features with correct spatial positions and precise boundary regions;

[0012] S5: Model Training Stage: The sum of Dice Loss and Binary CrossEntropy is used as the loss function. By comparing with the labels, the loss is calculated and the model parameters are trained by backpropagation until the model parameters converge;

[0013] Furthermore, in S1, the preprocessed brain tissue image specifically includes the following steps:

[0014] S11: Crop the brain tissue image into a size of 64×64×64.

[0015] S12: Normalize the processed brain tissue image data.

[0016] S13: Randomly extract a part of the image data with a size of 64×64×64 from the image data obtained in S12 and send it into the network model for training.

[0017] Furthermore, S2 includes the following steps:

[0018] S21: Use 3x3 convolution to extract features from the two modalities to obtain the features of the two modalities. The specific feature extraction steps are as follows:

[0019] T1 = LeakyReLU(InstanceNorm(Conv3(x1))) (1)

[0020] T2 = LeakyReLU(InstanceNorm(Conv3(x2))) (2)

[0021] Where x1 and x2 represent the inputs of modality T1 and modality T2, that is, the inputs obtained in S13, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, InstanceNorm represents instance normalization, LeakyReLU represents the activation function, T1 represents the features extracted from modality T1, and T2 represents the features extracted from modality T2.

[0022] S22: Apply the Sobel edge detection algorithm to the two modalities respectively to obtain the edge maps of the single modalities, and then perform feature extraction to obtain the edge information of the single modalities. The specific steps are as follows:

[0023] I1 = Conv3(Max(S(T1)) + Avg(S(T1))) (3)

[0024] I2 = Conv3(Max(S(T2)) + Avg(S(T2))) (4)

[0025] S1 == I1 ⊙ T1 (5)

[0026] S2 = J2 ⊙ T2 (6)

[0027] Among them, T1 and T2 represent the two modal features obtained by S21, S represents the 3D Sobel algorithm, Max and Avg respectively represent the max pooling and average pooling operations, Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, and S1 and S2 represent the outputs, that is, the edge information.

[0028] S23: Perform multi-scale feature extraction on the T1 modality of S21 and T2. The specific steps are as follows:

[0029] I1 = Conv3(T1) + Conv5(T1) + Conv7(T1) (7)

[0030] I2 = Conv3(T2) + Conv5(T1) + Conv7(T2) (8)

[0031] δ1 = F fc (F gap (I1)) (9)

[0032] δ2 = F fc (F gap (I2)) (10)

[0033]

[0034]

[0035] F1 = F13 + F15 + F17 (17)

[0036] F2 = F23 + F25 + F27 (18)

[0037] Among them, T1 and T2 are inputs, that is, the results of S21. Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1. Conv5 represents a 3D dilated convolution operation with a convolution kernel size of 3×3×3 and a dilation rate of 2. Conv7 represents a 3D dilated convolution operation with a convolution kernel size of 3×3×3 and a dilation rate of 3. F gap is global average pooling, F fc is a fully connected layer operation, SoftMax is a softmax activation function, and F1 and F2 are features extracted by a multi-scale feature extraction method.

[0038] S24: Take the features extracted by the multi-scale extraction method of S23, then take the boundary features of S22, fuse the two features, and finally perform downsampling. There are three layers in total, and the same operation is performed three times. The specific steps are as follows:

[0039] O1 = F1 + S2 (19)

[0040] O2 = F2 + S1 (20)

[0041] Among them, F1 and F2 are features extracted from two modalities obtained by S23, S1 and S2 are boundary features of two modalities obtained by S22, and O1 and O2 are features output by fusing the information of two modalities.

[0042] Furthermore, the specific steps of S3 are as follows:

[0043] S31: Take the features obtained in the S2 stage from two encodings and perform feature extraction respectively. The specific steps are as follows:

[0044] I1 = LeakyReLU(InstanceNorm(Conv3(x1))) (21)

[0045] I2 = LeakyReLU(InstanceNorm(Conv3(x2))) (22)

[0046] Among them, x1 and x2 represent the inputs of the T1 modality and the T2 modality, that is, the outputs of S24. Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1. InstanceNorm represents instance normalization, LeakyReLU represents an activation function, I1 represents the features extracted from the T1 modality, and I2 represents the features extracted from the T2 modality.

[0047] S32: Then perform cross-fusion on the two-modal data of S31. The specific steps are as follows:

[0048] F1 = I1 + x2 (23)

[0049] F2 = I2 + x1 (24)

[0050] Among them, I1 and I2 are the features obtained by S31, x2 and x1 are the output results obtained by S24, and F1 and F2 are the features of the fusion output.

[0051] S33: Perform attention mechanism calculation on the three-way data, then perform average fusion, and finally multiply with the original input features to obtain the output features. The specific steps are as follows:

[0052] T1 = Conv3(Max(F1) + Avg(F1)) ⊙ F1 (25)

[0053] T2 = Conv3(Max(F2) + Avg(F2)) ⊙ F2 (26)

[0054] T3 = Conv3(Max(x1 + x2) + Avg(x1 + x2)) ⊙ (x1 + x2) (27)

[0055]

[0056] Among them, F1 and F2 are the output results of S32, x1 and x2 are the output results of S24, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, Concat is concatenation on the channel, and O is the output result.

[0057] Furthermore, the S4 includes the following steps:

[0058] S41: First, upsample the features of the encoder, and then learn two learnable deformation fields by respectively fusing the features obtained by the upsample with the features after fusing the two modalities of the encoder. The specific steps are as follows:

[0059] DF1 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T1)))))) (29)

[0060] DF2 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T2)))))) (30)

[0061] DF3 = Conv1(Concat(DF1, Df2)) (31)

[0062] Among them, T1 and T2 are features rich in global information obtained in the encoder stage, that is, the output results of S24. LF is the input feature to be upsampled, Ups is the upsampling method, Concat is concatenation on the channel, ReLU is the ReLU activation function, GN is group normalization, Trans is the transpose operation, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, and DF3 is the obtained final feature deformation field.

[0063] S42: Perform a warping operation on the feature deformation field obtained through S41 and the original upsampled feature to obtain a feature with boundary correction and rich global information.

[0064] ALF = Warp(Ups(LF), DF3) (32)

[0065] Among them, LF represents the input feature to be upsampled, DF3 is the final feature deformation field obtained in S41, Ups is the upsampling operation, and Warp is the warping operation. Specifically, it is obtained by adding the learnable feature deformation field and the normalized grid, then performing grid sampling with the upsampled feature, and finally obtaining ALF, that is, the corrected feature map.

[0066] S43: First, splice the corrected feature map with the original upsampled feature map, then perform a convolution operation, and finally splice and perform a convolution operation with the output of the skip connection, that is, the output of S33, to finally obtain the final feature of the upsampling process.

[0067] L = Conv1(Concat(ALF, Ups(LF)) (33)

[0068] O = ReLU(GN(Conv3(Concat(I, EF)))) (34)

[0069] Among them, LF represents the input feature to be upsampled, ALF represents the corrected feature, that is, the result obtained in S42, EF is the early feature synthesized in the skip connection, that is, the output of S33, Concat is concatenation on the channel, Ups is the upsampling operation, GN is the group normalization operation, ReLU is the ReLU activation function, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, and O is the final result obtained after upsampling and feature alignment.

[0070] Furthermore, S5 includes the following steps:

[0071] The sum of DiceLoss and BinaryCrossEntropy is used as the loss function, and the formula is as follows:

[0072] L total = L DL + L BCE (35)

[0073]

[0074] Where N represents the total number of pixel points, which is equal to the number of pixels in a single image multiplied by the batchsize, p represents the predicted segmentation value, g represents the true label value, and to avoid division by zero, a smoothing term s = 10^-8 is applied in DL.

[0075] The beneficial effects of the present invention are as follows:

[0076] (1) By introducing the Sobel edge detection algorithm to extract single-modal edge contour information and combining with the multi-modal feature fusion technology, the refinement and learning ability of the brain tissue boundary are effectively strengthened. This method significantly alleviates the problem of blurred boundaries caused by low image contrast in traditional methods. Especially in the scenario where the contrast between gray matter (GM) and white matter (WM) of infant brain tissue is insufficient, it can generate segmentation results with accurate spatial positions and fine boundary regions.

[0077] (2) Design a cross-branch connection and modal fusion attention mechanism to promote the dynamic information interaction between T1 (anatomical structure) and T2 (lesion information) modalities. Through adaptive weight allocation and multi-scale feature extraction, redundant feature interference is reduced, and the complementary advantages of multi-modal data are fully integrated, solving the problem of limited single-modal information and enhancing the segmentation robustness of complex lesion regions.

[0078] (3) Introduce a learnable volume feature deformation field in the upsampling stage and combine it with a feature alignment module to accurately correct the spatial position offset of the feature map and effectively restore the detailed information of the original image. This method overcomes the problem of detail loss caused by insufficient resolution in traditional upsampling methods, especially suitable for medical images with low resolution or motion artifacts.

[0079] (4) Through multi-scale convolution and global average pooling techniques, the feature receptive field is expanded and the feature weights are adaptively adjusted to reduce the model bias caused by uneven pixel distribution of different tissue categories in medical images. Combining the hybrid loss function of Dice Loss and cross entropy further balances the class weights and enhances the segmentation sensitivity of the model to small target regions (such as cerebrospinal fluid CSF).

[0080] (5) Adopt an end-to-end network architecture design, integrate preprocessing, feature extraction, fusion, and alignment modules to achieve a fully automated segmentation process. Compared with traditional methods that rely on manual annotation, it significantly reduces the time cost and subjective error, providing an efficient and reliable solution for clinical rapid diagnosis and large-scale medical image analysis.

[0081] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings

[0082] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0083] Figure 1 It is the network structure diagram of the brain tissue segmentation model in the present invention;

[0084] Figure 2 It is the structure diagram of the multi-modal fusion module in the present invention;

[0085] Figure 3 It is the structure diagram of the multi-modal cross-attention module in the present invention;

[0086] Figure 4 It is the structure diagram of the feature alignment module in the present invention. Detailed Embodiments

[0087] The following illustrates the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0088] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; for better illustrating the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0089] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0090] Please refer to Figures 1 to 4 , Figure 1 which is the network structure diagram of the image segmentation algorithm in the present invention. The image segmentation algorithm using this model includes the following steps:

[0091] Step 1: Obtain the brain tissue image dataset, divide the dataset into a training set and a test set, and perform corresponding preprocessing. The specific steps are as follows:

[0092] Step 101: Crop the brain tissue image into a size of 64×64×64;

[0093] Step 102: Normalize the processed brain tissue image data;

[0094] Step 103: Randomly extract a part of the data with a size of 64×64×64 from the image data obtained in Step 102 and send it into the network model for training.

[0095] Step 2: Using the data obtained in Step 1, use the traditional edge detection algorithm Sobel operator to extract the boundary contour information of the single modality of the image, use multi-scale feature extraction for another modality, and through multi-scale convolution and adaptive weights, better extract features, and then use data fusion to integrate the boundary information into another modality; the specific steps are as follows

[0096] Step 201: For the result graph obtained in Step 102, first perform feature extraction on the input graph, then obtain the boundary contour graph through the Sobel edge detection algorithm, and then extract the boundary contour feature information through the boundary feature extraction module; the specific feature extraction steps are as follows:

[0097] T1 = LeakyReLU(InstanceNorm(Conv3(x1))) (1)

[0098] T2 = LeakyReLU(InstanceNorm(Conv3(x2))) (2)

[0099] I1 = Conv3(Max(S(T1)) + Avg(S(T1))) (3)

[0100] I2 = Conv3(Max(S(T2)) + Avg(S(T2))) (4)

[0101] S1 == I1 ⊙ T1 (5)

[0102] S2 = I2 ⊙ T2 (6)

[0103] Wherein, x1 and x2 represent inputs, S represents the 3D Sobel algorithm, Max and Avg represent the max pooling and average pooling operations respectively, Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, InstanceNorm represents instance normalization, LeakyReLU represents the activation function, and S1 and S2 are the output results of two modalities.

[0104] Step 202: The result graphs (T1, T2) obtained in Step 201 are subjected to image multi-scale feature extraction to obtain the original image feature information, which is fused with the boundary feature information (S1, S2) obtained in Step 201. The specific steps are as follows:

[0105] I1 = Conv3(T1) + Conv5(T1) + Conv7(T1) (7)

[0106] I2 = Conv3(T2) + Conv5(T2) + Conv7(T2) (8)

[0107] δ1 = F fc (F gap (I1)) (9)

[0108] δ2 = F fc (F gap (I2)) (10)

[0109]

[0110] F1 = F13 + F15 + F17 (17)

[0111] F2 = F23 + F25 + F27 (18)

[0112] O1 = F1 + S2 (19)

[0113] O2 = F2 + S1 (20)

[0114] Among them, T1 and T2 are inputs, that is, the results of S21. Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1. Conv5 represents a 3D dilated convolution operation with a convolution kernel size of 3×3×3 and a dilation rate of 2. Conv7 represents a 3D dilated convolution operation with a convolution kernel size of 3×3×3 and a dilation rate of 3. F gap is global average pooling, and F fc is a fully connected layer operation. SoftMax is the softmax activation function. O1 and O2 are the features output by fusing the information of the two modalities.

[0115] Step 3: Use cross-branch connections for the data in Step 2 to promote the information flow between different modalities, which can effectively fuse and refine multi-modal features. Then, perform spatial attention on each branch, and finally average and fuse the spatial information generated by each branch. The specific steps are as follows:

[0116] I1 = LeakyReLU(InstanceNorm(Conv3(x1))) (21)

[0117] I2 = LeakyReLU(InstanceNorm(Conv3(x2))) (22)

[0118] F1 = I1 + x2 (23)

[0119] F2 = I2 + x1 (24)

[0120] T1 = Conv3(Max(F1) + Avg(F1)) ⊙ F1 (25)

[0121] T2 = Conv3(Max(F2) + Avg(F2)) ⊙ F2 (26)

[0122] T3 = Conv3(Max(x1 + x2) + Avg(x1 + x2)) ⊙ (x1 + x2) (27)

[0123]

[0124] Among them, x1 and x2 represent the inputs of the T1 modality and the T2 modality. Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1. Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1. InstanceNorm represents instance normalization. LeakyReLU represents an activation function. Concat is concatenation on the channel. O is the output result.

[0125] Step 4: Upsample to increase the spatial resolution of the feature map, helping to restore the detailed information of the original input. Introduce a feature alignment module to generate the upsampled late-stage features with correct spatial positions and precise boundary regions.

[0126] Step 401: Learn from the upsampled features and the early features of the two modalities obtained by the encoder to obtain two learnable feature deformation fields, fuse them to obtain an information-rich feature deformation field, and then correct the boundaries of the upsampled features.

[0127] DF1 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T1)))))) (29)

[0128] DF2 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T2)))))) (30)

[0129] DF3 = Conv1(Concat(DF1, Df2)) (31)

[0130] ALF = Warp(Ups(LF), DF3) (32)

[0131] Among them, T1 and T2 are the features rich in global information obtained in the encoder stage, that is, the output results of Step 2. LF is used as the input feature to be upsampled. Ups is the upsampling method. Concat is concatenation on the channel. ReLU is the ReLU activation function. GN is group normalization. Trans is the transpose operation. Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1. Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1. Warp is the warping operation. Specifically, it adds the obtained learnable feature deformation field and the normalized grid, then performs grid sampling with the upsampled features, and finally obtains ALF, that is, the corrected feature map.

[0132] Step 402: Concatenate and convolve the corrected feature deformation field obtained in Step 401 with the original upsampled features, then concatenate and convolve with the features obtained from the skip connection, and finally obtain the result of the upsampled output.

[0133] I = Conv1(Concat(ALF, Ups(LF))) (33)

[0134] O = ReLU(GN(Conv3(Concat(I, EF)))) (34)

[0135] Among them, LF represents the input feature to be upsampled, ALF represents the corrected feature, that is, the result obtained in step 401, EF is the early feature synthesized by skip connection, that is, the output of step 3, Concat is the concatenation on the channel, Ups is the upsampling operation, GN is the group normalization operation, ReLU is the ReLU activation function, Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents the 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, and O is the final result obtained after upsampling and feature alignment.

[0136] Step 5: Use the sum of DiceLoss and BinaryCrossEntropy as the loss function, and the formula is as follows:

[0137] L total = L DL + L BCE (35)

[0138]

[0139] Among them, N represents the total number of pixel points, which is equal to the number of pixels in a single image multiplied by the batchsize, p represents the predicted segmentation value, g represents the true label value, and in order to avoid division by zero, a smoothing term s = 10^-8 is applied in DL.

[0140] As Figure 2 shown, the multi-modal fusion processing process in the present invention is as follows:

[0141] Step 1: Denote the input original image data as IMGT1 ∈ R B×C×W×H and IMGT2 ∈ R B×C×W×H , where B, C, W, and H represent the batch size, number of channels, width, and height of the feature map respectively. Regular convolution feature extraction is performed on them respectively to obtain two tensors T1 and T2; the expression of the processing process is as follows:

[0142] T1 = LeakyReLU(InstanceNorm(Conv3(x1))) (38)

[0143] T2 = LeakyReLU(InstanceNorm(Conv3(x2))) (39)

[0144] Among them, x1 and x2 represent the inputs of T1 modality and T2 modality, Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, InstanceNorm represents instance normalization, and LeakyReLU represents the activation function.

[0145] Step 2: Feed the features T1 and T2 obtained in Step 1 into the traditional edge detection operator sobel for convolution to obtain edge contour feature maps S1 and S2; the expression for the processing process is as follows:

[0146] I2 = Conv3(Max(S(T1)) + Avg(S(T1))) (40)

[0147] I2 = Conv3(Max(S(T1)) + Avg(S(T1))) (41)

[0148] S1 = I1 ⊙ T1 (42)

[0149] S2 = I2 ⊙ T2 (43)

[0150] Among them, T1 and T2 represent two modal tensors obtained in Step 1, S represents the 3D Sobel algorithm, Max and Avg respectively represent the max pooling and average pooling operations, and Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1.

[0151] Step 3: Perform multi-scale feature extraction on the feature maps T1 and T2 obtained in Step 1 respectively, and output features T1 feature and T2 feature ; the expression for the processing process is as follows:

[0152] I1 = Conv3(T1) + Conv5(T1) + Conv4(T1) (44)

[0153] I2 = Conv3(T2) + Conv5(T2) + Conv7(T2) (45)

[0154] δ1 = F fc (F gap (I1)) (46)

[0155] δ2 = F fc (F gap (I2)) (47)

[0156]

[0157]

[0158] T1 feature = F13 + F15 + F17 (54)

[0159] T 2feature == F23 + F25 + F27 (55)

[0160] Among them, T1 and T2 are inputs, that is, the results of S21. Conv3 represents a 3D convolution operation with a kernel size of 3×3×3 and a stride of 1. Conv5 represents a 3D dilated convolution operation with a kernel size of 3×3×3 and a dilation rate of 2. Conv7 represents a 3D dilated convolution operation with a kernel size of 3×3×3 and a dilation rate of 3. F gap is global average pooling, and F fc is a fully connected layer operation, and SoftMax is a softmax activation function.

[0161] Step 4: Interleave and fuse the features obtained in Step 3 and the features obtained in Step 2 according to different modalities to obtain image features Output with texture boundaries and rich semantic information T1 and Output T2 ; The specific steps are as follows:

[0162] Output T1 = F1 + S2 (56)

[0163] Output T2 = F2 + S1 (57)

[0164] Among them, F1 and F2 are the features extracted from two modalities obtained in Step 3, and S1 and S2 are the boundary features of two modalities obtained in Step 2.

[0165] As Figure 3 shown, the processing process of multi-modal cross-attention in the present invention is:

[0166] Step 1: The input image features are denoted as F T1 ∈R B×C×W×H and F T2 ∈R B×C×W×H , where B, C, W, and H represent the batch size, the number of channels, and the width and height of the feature map respectively. First, perform a convolution operation on the two features to extract features of different modalities; the specific steps are as follows:

[0167] I1 = LeakyReLU(InstanceNorm(Conv3(F T1 ))) (58)

[0168] I2 = LeakyReLU(InstanceNorm(Conv3(F T2 ))) (59)

[0169] Among them, F T1 and F T2The input representing the T1 modality and the T2 modality, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, InstanceNorm represents instance normalization, LeakyReLU represents an activation function, I1 represents the features extracted from the T1 modality, and I2 represents the features extracted from the T2 modality.

[0170] Step 2: Interleave and fuse the obtained features to get features, perform spatial attention operations on them respectively to obtain three-way features, perform average aggregation on the three-way features, then concatenate them with the original input features, and use convolution to restore the number of channels to obtain the final output result; the specific steps are as follows:

[0171] F1 = I1 + x2 (60)

[0172] F2 = I2 + x1 (61)

[0173] T1 = Conv3(Max(F1) + Avg(F1)) ⊙ F1 (62)

[0174] T2 = Conv3(Max(F2) + Avg(F2)) ⊙ F2 (63)

[0175] T3 = Conv3(Max(x1 + x2) + Avg(x1 + x2)) ⊙ (x1 + x2) (64)

[0176]

[0177] Among them, I1 and I2 are the features obtained in Step 1, x2 and x1 are the input results obtained in Step 1, F1 and F2 are the features of the fusion output, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, Concat is concatenation on the channels, and O is the output result.

[0178] As Figure 4 shown, the process of feature alignment in the present invention is as follows:

[0179] Step 1: First, upsample the features of the encoder, and then learn two learnable deformation fields by respectively fusing the features obtained by the upsample with the features after fusing the two modalities of the encoder, and then perform a warping operation on the features after the original upsample to obtain features with boundary correction and rich global information; the specific steps are as follows:

[0180] DF1 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T1)))))) (66)

[0181] DF2 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T2)))))) (67)

[0182] DF3 = Conv1(Concat(DF1, Ff2)) (68)

[0183] ALF = Warp(Ups(LF), DF3) (69)

[0184] Among them, T1 and T2 are features rich in global information obtained in the encoder stage, LF is the input feature to be upsampled, Ups is the upsampling method, Concat is concatenation on the channels, ReLU is the ReLU activation function, GN is group normalization, Trans is the transpose operation, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, Warp is the warping operation, specifically by adding the obtained learnable feature deformation field and the normalized grid, then performing grid sampling together with the upsampled feature, and finally obtaining ALF, that is, the corrected feature map.

[0185] Step 2: First, concatenate the corrected feature map with the original upsampled feature map, then perform a convolution operation, and finally concatenate and perform a convolution operation with the output of the skip connection to finally obtain the final feature of the upsampling process; the specific steps are as follows:

[0186] I = Conv1(Concat(ALF, Ups(LF))) (70)

[0187] O = ReLU(GN(Conv3(Concat(I, EF)))) (71)

[0188] Among them, LF represents the input feature to be upsampled, ALF represents the corrected feature, EF is the early feature synthesized in the skip connection, Concat is concatenation on the channels, Ups is the upsampling operation, GN is the group normalization operation, ReLU is the ReLU activation function, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, and O is the final result obtained after upsampling and feature alignment.

[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for segmenting brain tissue images with multimodal fusion and feature alignment, characterized in that: The method includes the following steps: S1: Data preparation stage: Obtain a brain tissue image dataset, divide the dataset into a training set and a test set, and perform corresponding preprocessing; S2: Multi-scale feature extraction stage: Use the traditional edge detection algorithm Sobel operator to extract the edge contour information of the single modality of the image, use data fusion to integrate the edge information into another modality, reduce the redundancy of feature fusion, perform multi-scale feature extraction on the other modality, and through multi-scale convolution and adaptive weights, better extract features; S3: Modal fusion attention stage: Design two independent multi-modal feature branches, use cross-branch connections to promote the information flow between different modalities, fuse and refine multi-modal features, then perform spatial attention on each branch, and finally average and fuse the spatial information generated by each branch; S4: Upsampling feature alignment stage: Increase the spatial resolution of the feature map during upsampling to help restore the details of the original input, introduce a feature alignment module, and use a learnable volumetric feature deformation field for feature alignment to generate upsampled late-stage features with correct spatial positions and precise boundary regions; S5: Model training stage: Use the sum of Dice Loss and Binary CrossEntropy as the loss function, calculate the loss by comparing with the label, and backpropagate to train the model parameters until the model parameters converge.

2. The method for segmenting brain tissue images with multi-modal fusion and feature alignment according to claim 1, characterized in that: In the above S1, the preprocessing of the brain tissue image specifically includes the following steps: S11: Crop the brain tissue image into a size of 64×64×64; S12: Normalize the processed brain tissue image data; S13: Randomly extract a part of the image data with a size of 64×64×64 from the image data obtained in S12 and send it into the network model for training.

3. The method for segmenting brain tissue images with multi-modal fusion and feature alignment according to claim 1, wherein: The above S2 includes the following steps: S21: Perform feature extraction on the two modalities using a 3x3 convolution to obtain the features of the two modalities. The specific feature extraction steps are as follows: T1 = LeakyReLU(InstanceNorm(Conv3(x1))) (1) T2 = LeakyReLU(InstanceNorm(Conv3(x2))) (2) where x1 and x2 represent the inputs of modality T1 and modality T2, that is, the inputs obtained in S13, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, InstanceNorm represents instance normalization, LeakyReLU represents an activation function, T1 represents the features extracted from modality T1, and T2 represents the features extracted from modality T2; S22: Use the Sobel edge detection algorithm for the two modalities respectively to obtain the edge maps of the single modality, and then perform feature extraction to obtain the edge information of the single modality. The specific steps are as follows: I1 = Conv3(Max(S(T1)) + Avg(S(T1))) (3) I2 = Conv3(Max(S(T2)) + Avg(S(T2))) (4) S1 == I1⊙T1 (5) S2 = I2⊙T2 (6) Among them, T1 and T2 represent two modal features obtained by S21, S represents the 3D Sobel algorithm, Max and Avg respectively represent the max pooling and average pooling operations, Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, and S1 and S2 represent the outputs, that is, boundary information; S23: Extract multi-scale features from the T1 modality and T2 of S21; the specific steps are as follows: I1 = Conv3(T1) + Conv5(T1) + Conv7(T1) (7) I2 = Conv3(T2) + Conv5(T2) + Conv7(T2) (8) δ1 = F fc (F gap (I1)) (9) δ2 = F fc (F gap (I2)) (10) F1 = F13 + F15 + F17 (17) F2 = F23 + F25 + F27 (18) Among them, T1 and T2 are inputs, that is, the results of S21. Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1. Conv5 represents a 3D dilated convolution operation with a convolution kernel size of 3×3×3 and a dilation rate of 2. Conv7 represents a 3D dilated convolution operation with a convolution kernel size of 3×3×3 and a dilation rate of 3. F gap is global average pooling, T fc is a fully connected layer operation, SoftMax is the softmax activation function, and F1 and F2 are features extracted by the multi-scale feature extraction method; S24: Take the features extracted by the multi-scale extraction method of S23, then take the boundary features of S22, fuse the two features, and finally perform downsampling. There are a total of three layers, and the same operation is performed three times; the specific steps are as follows: O1 = F1 + S2 (19) O2 = F2 + S1 (20) Among them, F1 and F2 are the features extracted from the two modalities obtained by S23, S1 and S2 are the boundary features of the two modalities obtained by S22, and O1 and O2 are the features output by fusing the two modal information.

4. The method for segmenting brain tissue images with multimodal fusion and feature alignment according to claim 1, wherein: The said S3 includes the following steps: S31: Take the features obtained in the S2 stage in the two encodings and perform feature extraction respectively; the specific steps are as follows: I1 = LeakyReLU(InstanceNorm(Conv3(x1))) (21) I2 = LeakyReLU(InstanceNorm(Conv3(x2))) (22) Among them, x1 and x2 represent the inputs of the T1 modality and T2 modality, that is, the outputs of S24, Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, InstanceNorm represents instance normalization, LeakyReLU represents the activation function, I1 represents the features extracted from the T1 modality, and I2 represents the features extracted from the T2 modality; S32: Then perform cross-fusion on the two modal data in S31; the specific steps are as follows: F1 = I1 + x2 (23) F2 = I2 + x1 (24) Among them, I1 and I2 are the features obtained by S31, x2 and x1 are the output results obtained by S24, and F1 and F2 are the features output by fusion; S33: Perform attention mechanism calculation on the three-way data, then perform average fusion, and finally multiply with the original input features to obtain the output features; the specific steps are as follows: T1 = Conv3(Max(F1) + Avg(F1)) ⊙ F1 (25) T2 = Conv3(Max(F2) + Avg(F2)) ⊙ F2 (26) T3 = Conv3(Max(x1 + x2) + Avg(x1 + x2)) ⊙ (x1 + x2) (27) Among them, F1 and F2 are the output results of S32, x1 and x2 are the output results of S24, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, Concat is concatenation on the channel, and O is the output result.

5. The method for segmenting brain tissue images with multimodal fusion and feature alignment according to claim 1, wherein: S4 includes the following steps: S41: First, upsample the features of the encoder, and then learn two learnable deformation fields from the features obtained by upsample and the features after fusing two modalities of the encoder respectively. The specific steps are as follows: DF1 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T1)))))) (29) DF2 = Trans(Conv3(ReLU(GN(Conv1(Concat(Ups(LF), T2)))))) (30) DF3 = Conv1(Concat(DF1, Df2)) (31) Among them, T1 and T2 are the features rich in global information obtained in the encoder stage, that is, the output results of S24, LF is the input feature to be upsample, Ups is the upsample method, Concat is concatenation on the channel, ReLU is the ReLU activation function, GN is group normalization, Trans is the transpose operation, Conv3 represents a 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents a 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, and DF3 is the final feature deformation field obtained; S42: Perform a warping operation on the feature deformation field obtained in S41 and the features after the original upsample to obtain features with boundary correction and rich global information; ALF = Warp(Ups(LF), DF3) (32) Among them, LF represents the input feature to be upsample, DF3 is the final feature deformation field obtained in S41, Ups is the upsample operation, Warp is the warping operation, specifically by adding the learnable feature deformation field and the normalized grid, and then performing grid sampling with the upsample features together, and finally obtaining ALF, that is, the corrected feature map; S43: First, concatenate the corrected feature map with the original feature map after upsample, then perform a convolution operation, and finally concatenate and perform a convolution operation with the output of the skip connection, that is, the output of S33, to finally obtain the final feature in the upsample process; I = Conv1(Concat(ALF, Ups(LF))) (33) O = ReLU(GN(Conv3(Concat(I, EF)))) (34) Among them, LF represents the input feature to be upsampled, ALF represents the corrected feature, that is, the result obtained in S42, EF is the early feature synthesized by the skip connection, that is, the output of S33, Concat is the concatenation on the channel, Ups is the upsampling operation, GN is the group normalization operation, ReLU is the ReLU activation function, Conv3 represents the 3D convolution operation with a convolution kernel size of 3×3×3 and a stride of 1, Conv1 represents the 3D convolution operation with a convolution kernel size of 1×1×1 and a stride of 1, and O is the final result obtained after upsampling and feature alignment.

6. The method for segmenting brain tissue images with multi-modal fusion and feature alignment according to claim 1, wherein: The S5 includes the following steps: The sum of Dice Loss and Binary CrossEntropy is used as the loss function, and the formula is as follows: L total = L DL + L BCE (35) Among them, N represents the total number of pixel points, which is equal to the number of pixels in a single image multiplied by the batchsize, p represents the predicted segmentation value, g represents the true label value, and to avoid division by zero, the smoothing term s = 10-8 in DL.