Multi-modal fusion-based nuclear magnetic resonance image auxiliary diagnosis method and system
Through individualized brain region segmentation and adaptive weighted fusion, combined with a lightweight teacher-student network model, the problems of modal sensitivity neglect and spatial error interference in multimodal magnetic resonance imaging are solved, and the diagnostic accuracy and computational efficiency of early lesions are improved.
Patent Information
- Application Number
- CN202510966226.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
AI Technical Summary
Existing multimodal magnetic resonance imaging fusion methods fail to effectively utilize the differences in brain region-specific pathological representations, resulting in the pathological signals of key brain regions being submerged by non-sensitive modal noise, reducing the accuracy of early diagnosis, and the existence of cross-modal feature redundancy and spatial misalignment interference, which affects diagnostic accuracy.
Modal lesion sensitivity weights are calculated through individualized brain region segmentation, and adaptive weighted fusion and dynamic weight adjustment are adopted, combined with edge-preserving filtering and spatial registration. A follow-up image correction mechanism is introduced, and a lightweight teacher-student network model is used for diagnosis.
It significantly improves the sensitivity and recognition accuracy of lesions in specific brain areas, enhances the detection rate and diagnostic reliability of early lesions, reduces the demand for computing resources, and is suitable for practical clinical applications.
Smart Images

Figure CN120809168A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image auxiliary diagnosis, and in particular to a nuclear magnetic resonance image auxiliary diagnosis method and system based on multi-modal fusion. BACKGROUND
[0002] Multi-modal magnetic resonance imaging (MRI) technology integrates T1-weighted structural images, T2-weighted fluid-attenuated inversion recovery images (FLAIR), diffusion-weighted imaging (DWI), and other different sequence data to provide complementary information for the diagnosis of neurodegenerative diseases (such as Alzheimer's disease and Parkinson's disease). Structural images can represent brain atrophy morphology, and functional images can reflect metabolic and microstructural disorders. The combination of the two is expected to improve the early detection rate of lesions. However, existing fusion methods have the following limitations:
[0003] 1. Insufficient use of modality complementarity, without considering the differences in pathological representation of brain regions;
[0004] The sensitivity of different brain regions to modalities varies significantly. For example, hippocampal atrophy has clear structural boundaries in T1-weighted images, while white matter lesions rely on the high signal contrast of FLAIR images, and basal ganglia iron deposition is more pronounced in DWI diffusion characteristics.
[0005] Defects in the prior art: current fusion strategies (such as weighted averaging and feature-level splicing) use a unified weight for the whole brain without distinguishing the relevance of brain regions and modalities. For example:
[0006] In bilateral hippocampal sclerosis cases, if the T1 weight is not specifically increased, it is easy to miss diagnosis due to bilateral symmetry changes;
[0007] The early microstructural changes in the substantia nigra compacta of Parkinson's disease are more sensitive in DWI, but uniform fusion dilutes the contribution of this modality.
[0008] The above problems result in the pathological signals of key brain regions being overwhelmed by non-sensitive modality noise, reducing the accuracy of early diagnosis by ≥15%.
[0009] 2. Cross-modal feature redundancy and spatial misalignment interference:
[0010] The differences in resolution and coordinate system between structural and functional images result in voxel misalignment and irrelevant feature interference during fusion:
[0011] Signals from non-target regions such as the cerebrospinal fluid artifact in T2 images can interfere with the diagnosis of cortical lesions;
[0012] Spatial shifts caused by registration errors (such as misalignment of the head-body boundary of the hippocampus) cause functional-structural feature misalignment.
[0013] Therefore, there is an urgent need for a multi-modal fusion-based magnetic resonance image assisted diagnosis method and system to solve the above problems. SUMMARY
[0014] Based on the above purpose, the application provides a multi-modal fusion-based magnetic resonance image assisted diagnosis method and system, wherein a multi-modal fusion-based magnetic resonance image assisted diagnosis method comprises:
[0015] Step 1: synchronously collecting three-dimensional T1 weighted structural images, T2 weighted fluid attenuated inversion recovery images and diffusion weighted imaging data of a subject, performing spatial registration based on the T1 weighted images and executing skull stripping and gray scale normalization;
[0016] Step 2: individualized brain region segmentation based on a brain atlas, calculating lesion sensitivity weights of each modality for each segmented brain region, and the weights are obtained by quantifying the following parameters:
[0017] T1 weighted image gray matter-white matter interface structure definition, T2 weighted fluid attenuated inversion recovery image abnormal signal distribution density, and diffusion weighted imaging microstructure disorder degree;
[0018] Step 3: extracting multi-modal image features in each brain region, adaptively weighting and fusing according to the sensitivity weights, and increasing the weight proportion of the corresponding modality when detecting that the brain lesion probability exceeds a dynamic threshold;
[0019] Step 4: inputting the whole brain region fusion features into a multi-task classifier to synchronously output disease types and clinical stages, and correcting the staging results through a feature vector change rate if there is a follow-up image.
[0020] Correspondingly, the application also provides a multi-modal fusion-based magnetic resonance image assisted diagnosis system, which comprises a memory configured to store instructions, and a processor configured to call the instructions from the memory and realize the multi-modal fusion-based magnetic resonance image assisted diagnosis method according to any of the embodiments of the application when executing the instructions.
[0021] The application has the following advantages:
[0022] 1. The application realizes adaptive weighted fusion in each brain region according to the modality sensitivity weight based on individualized brain region segmentation based on a brain atlas. By calculating the lesion sensitivity weight of each brain region in T1, T2 and DWI modalities, the application realizes the optimal fusion of specific sensitive modalities for different brain regions. The method significantly improves the sensitivity to lesions in specific brain regions, avoids the omission of key brain pathological signals caused by uniform fusion methods, and thus improves the detection rate of early lesions.
[0023] 2、The application reduces spatial registration error and misalignment interference caused by registration error to the maximum extent by taking T1 weighted structural image as a reference for spatial registration and performing skull stripping and gray scale standardization. At the same time, an edge preserving filter algorithm is adopted to eliminate the boundary blur problem that may occur in the standardization process, further ensuring the accuracy of image registration and fusion effect. In this way, the fusion of structural and functional images is more accurate, reducing image interference caused by spatial misalignment, and improving the recognition accuracy of brain lesions.
[0024] 3、The application introduces a dynamic weight adjustment mechanism. When the brain lesion probability is detected to exceed the dynamic threshold, the system automatically increases the weight of the corresponding modality, giving priority to the modality sensitive to the lesion. This mechanism effectively improves the detection ability of key brain lesions and dynamically adjusts the fusion mode of each modality according to the specific characteristics of brain lesions, avoiding blind equalization of weight settings and enhancing the accurate diagnosis of early lesions.
[0025] 4、The application introduces a follow-up image correction mechanism, which dynamically adjusts the clinical stage by calculating the Mahalanobis distance change rate. Especially when the change rate of key brain regions is inconsistent, the manual review process is started. This mechanism can effectively identify the speed and direction of lesion progression and correct the stage of the disease according to the changes in follow-up images, thereby improving the accuracy of long-term follow-up and the reliability of diagnosis.
[0026] 5、The application uses a teacher-student network model through knowledge distillation technology, migrating the knowledge of a pre-trained three-dimensional convolutional network to a lightweight student network. By freezing the convolutional layer weights of the teacher model and minimizing the probability distribution difference between the outputs of the teacher and student networks, the student network can quickly learn the feature representation of the teacher network, thereby significantly improving the computational efficiency while maintaining high diagnostic accuracy. This innovative design not only reduces the demand for computing resources, but also speeds up the diagnosis process, making the method more suitable for clinical practical application. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.
[0028] Fig. 1 The step flowchart of the method of the application;
[0029] Fig. 2 The step flowchart of the lesion sensitivity weight calculation of step 2 of the method of the application;
[0030] Fig. 3Step flow chart for building a multi-task classifier for step 4 of the method of the present application. DETAILED DESCRIPTION
[0031] The present application will be described in detail below with reference to the drawings and specific embodiments. It should be noted here that, in order to make the embodiments more detailed, the following embodiments are the best, preferred embodiments, and other alternative ways can also be implemented by those skilled in the art for some known technologies; and the drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present application.
[0032] Please refer to Figs. 1-3 , the embodiment of the present application provides a kind of nuclear magnetic resonance image auxiliary diagnosis method based on multi-modal fusion, first, in step 1, three-dimensional T1 weighted structure image, T2 weighted FLAIR image and DWI image data need to be synchronously collected to subject. After acquisition is completed, with T1 weighted image as spatial reference standard, realize the spatial alignment of the rest modal and T1 image by mutual information or image registration algorithm (such as rigid registration algorithm based on gradient descent), solve the difference of multi-modal data in resolution and coordinate system problem. After registration is completed, skull stripping is carried out by morphological method (such as skull removal technology based on threshold and region growth), and gray intensity standardization is carried out simultaneously, to ensure the consistency of different modal in intensity domain, lay a foundation for subsequent feature extraction and fusion.
[0033] In step 2, the registered image is segmented into individual brain regions using a standard brain atlas (such as Harvard-Oxford atlas). Then, for each segmented brain region, the lesion sensitivity weight of three modalities is calculated. Specifically, it includes: calculating the edge gradient or texture clarity of gray-white matter junction in T1 weighted image to reflect the structural change sensitivity; counting the density of abnormal high signal voxels in T2-FLAIR image to measure the white matter lesion recognition ability; evaluating the fluctuation range of average apparent diffusion coefficient (ADC) or anisotropy index (FA) in each brain region of DWI to quantify the degree of microstructure disorder. These weights will be important parameters for each modal fusion, to realize the brain region specific weighting strategy.
[0034] In step 3, image feature extraction techniques are used to extract structural and functional features from multi-modal images of each brain region. In this embodiment, image feature extraction techniques include three-dimensional convolution operation, gray level co-occurrence matrix, texture analysis or deep network intermediate layer output. Then, combined with the modal sensitivity weight obtained in step 2, weighted fusion is performed. In the fusion process, a dynamic threshold is set. When the lesion probability of a brain region exceeds the threshold, the weight proportion of the corresponding sensitive modality in the brain region is automatically increased, so as to enhance the expression of lesion signal and avoid dilution of key region signal by non-sensitive modality.
[0035] In step 4, the fused whole-brain features are input into a multi-task classification structure to output disease classification results and preliminary clinical staging. In the presence of follow-up images, the staging results are corrected by the rate of change of the feature vector. The specific method includes: on the basis of healthy people, the covariance matrix of the fusion features of each brain region is established, and the Mahalanobis distance increment between the follow-up image and the initial image of the target individual is calculated based on the reference, and the change rate is obtained by dividing the follow-up time interval. When the change rate is continuously higher than a certain preset threshold (such as annual change rate > 0.15), the system automatically upgrades the disease course staging; if the change trends of different key brain regions are contradictory (such as significant progression in a certain region while other regions remain stable), manual review prompts are triggered to improve the clinical reliability of the staging results.
[0036] The method provides a systematic technical solution path for the core problems of current multi-modal fusion image diagnosis, such as modal sensitivity neglect, spatial error interference, and insufficient dynamic evolution response, and has significant practical value and clinical popularization prospects.
[0037] In one possible implementation, first, a standard template image similar in age to the subject is selected from a pre-constructed image template library of healthy people, and its gray level histogram and cumulative probability distribution are extracted. The template library needs to cover different age levels to ensure the adaptability of the gray level distribution matching. Subsequently, the current image is subjected to gray level statistical analysis to obtain its gray level histogram and corresponding cumulative probability distribution function (CDF). At this time, for each gray level in the current image, the closest probability value in the CDF of the reference template is found according to its probability value in the CDF, and it is mapped to the gray level at the corresponding probability position in the reference image. This process is essentially the generation of a CDF reverse lookup mapping function for histogram matching, which can effectively adjust the current image gray level distribution to the statistical interval of the target reference template, solving the gray level shift problem caused by different image devices, scanning parameters, and tissue properties.
[0038] However, the gray level histogram transformation may introduce excessive smoothing or blurring at the tissue boundary, especially at the gray-white matter junction, the edge of the ventricle, and other positions with high anatomical structure definition requirements. Therefore, an edge-preserving filtering algorithm needs to be introduced for correction. This process calculates the gradient amplitude of each pixel as the spatial weight factor during filtering: regions with large edge gradients and prominent changes are inhibited to maintain the sharpness of the boundary; while in texture flat areas, larger range weighted smoothing is allowed to reduce gray level fluctuations and noise residues. In specific implementation, an improved bilateral filtering or guided filtering structure is often used, which embeds the gradient amplitude factor in the design of the filtering kernel, and adaptively adjusts the spatial size of the filtering kernel and the gray level weight function, so as to complete the gray level standardization correction of the whole image on the premise of local structure preservation.
[0039] The method not only realizes the gray scale uniform processing of cross-modal images, but also protects the anatomical boundary information in the processing process, and has significant application value and popularization prospect in multi-modal fusion and lesion analysis.
[0040] In one possible implementation, the individual brain region segmentation step realizes accurate brain region structure segmentation, especially the extraction of hippocampus, basal ganglia and subcortical nuclei, through a differential homeomorphic deformation algorithm and a high-precision boundary guarantee strategy.
[0041] Specifically, first, the individual T1 weighted image is registered with the standard brain atlas to ensure that the brain regions of different individuals can be correctly mapped to the anatomical standard template. In specific implementation, first, the mutual information value between the individual T1 image and the standard atlas is calculated as a measure of the similarity of the two. Mutual information is a standard for measuring the amount of information shared between two images, and a higher mutual information value indicates a stronger similarity in spatial distribution between the two.
[0042] Next, the gradient descent method is used to optimize the control point displacement of the velocity field, and the velocity field is adjusted iteratively to gradually match the individual image to the structure of the standard atlas. Gradient descent is an iterative method based on optimization idea. In each iteration, the velocity field is updated according to the current image deformation error, so as to gradually match the image to the standard atlas.
[0043] In order to avoid unreasonable deformation in the deformation process, the algorithm detects the Jacobian determinant value of the deformation field after each iteration. If the Jacobian determinant value of a region is negative, it means that the deformation of that region has an unreasonable compression or stretching phenomenon, so fluid force school needs to be applied. Fluid force school can effectively smooth and adjust the deformation field by simulating the propagation of fluid in space, avoiding unnatural deformation and ensuring the continuity and accuracy of the overall brain region structure.
[0044] In terms of boundary accuracy in segmentation, especially in the gray-white matter junction area, the gray value changes sharply in this area, which is prone to errors, so a higher resolution deformation grid needs to be used for fine modeling. High-resolution deformation grid can achieve more detailed deformation adjustment in the gray-white matter junction area, so as to more accurately identify and segment the brain region boundary. At the same time, the intensity characteristics of cerebrospinal fluid signal are used as the constraint condition of the edge. Cerebrospinal fluid appears as a low signal area in the image, which naturally forms a clear boundary between gray matter and white matter. By using this feature, the deformation algorithm can better identify the true boundary of the gray-white matter junction and avoid excessive smoothing or blurring.
[0045] By using the differential homeomorphism deformation algorithm and high-resolution grid design, combined with the signal intensity characteristics of cerebrospinal fluid, the accuracy and boundary protection ability of brain region segmentation are effectively improved, which has significant clinical application value.
[0046] In one possible implementation, first, the step of quantifying the structure clarity of the gray-white matter interface is performed by high-precision analysis of the gray-white matter interface region to evaluate the structure clarity of the region. Specifically, sampling points are set at equal intervals on the gray-white matter interface line, and a neighborhood window with a predetermined size is established centered on each sampling point. Then, the statistical dispersion of the gradient magnitude in the window is calculated as an indicator of clarity. The gradient magnitude reflects the degree of signal change in the gray-white matter interface region, and the greater the dispersion, the clearer the boundary, and vice versa. This quantification method can effectively help judge the segmentation quality of the gray-white matter interface region in the image, which is crucial for accurate brain segmentation.
[0047] Second, the step of quantifying the abnormal signal distribution density is performed by analyzing the structural changes around the ventricle to help identify possible lesion areas. Specifically, a pre-set expansion region related to anatomical structures around the ventricle is identified, and the proportion of pixels in the region that exceed a certain multiple threshold of the healthy reference value is calculated. Pixels exceeding the threshold often represent lesions or abnormal signals, and statistical analysis of these areas can reveal the presence and severity of lesions. For example, if the pixel value in a certain region is significantly higher than the normal range, it may indicate the presence of edema or tumor in that region.
[0048] Finally, the microstructure disorder degree quantification uses the fractional anisotropy map (FA map) obtained from diffusion weighted imaging (DWI) to evaluate the integrity of brain tissue microstructure. By dividing the target brain region into adaptive size sub-blocks, the average value of information entropy of each sub-block is calculated as the disorder degree indicator. Information entropy can reflect the complexity of the image region, and the greater the entropy value, the more complex or disordered the structure of the region, and the lower the entropy value, the more regular the structure of the region. This method can effectively reveal the damage to the microstructure, such as the destruction of white matter fibers or the degeneration of nerve fibers.
[0049] Through these quantification methods, combined with the evaluation of the clarity of the gray-white matter interface, the density of abnormal signals, and the degree of microstructure disorder, comprehensive and accurate brain lesion diagnosis can be provided for clinical practice, significantly improving the sensitivity and reliability of diagnosis.
[0050] In one possible implementation, first, the real-time calculation of brain region lesion probability is completed by a lightweight neural network. The network receives the multi-modal feature vectors extracted from each brain region as input, which can come from data sources such as structural images (such as T1, T2) and functional images (such as DWI, fMRI). The last layer of the network uses a sigmoid function for nonlinear transformation, and outputs a lesion probability value between 0 and 1, reflecting the likelihood of abnormality in the brain region. The parameters of the lightweight network are not trained from scratch, but are obtained by knowledge distillation from a pre-trained three-dimensional convolutional neural network. The three-dimensional convolutional network has stronger feature abstraction capability, but the computational overhead is larger. The lightweight network inherits the knowledge of the three-dimensional convolutional network, which can greatly improve the inference speed while ensuring the recognition accuracy.
[0051] Then, the weight boosting mechanism responds hierarchically according to the level of lesion probability. When the lesion probability of a brain region exceeds a preset first threshold, the system activates a linear boosting mode, that is, the weight of the modality in the fusion process is boosted in a linear function. This helps to increase the attention to high correlation modalities (such as DWI) in the suspected lesion area, thereby enhancing the recognition of abnormal structures. If the lesion probability further exceeds the second threshold, a more sensitive exponential boosting mode is activated, and the weight boosting is exponential growth, to ensure that the area with extremely high lesion risk is given sufficient attention. This hierarchical boosting mechanism allows the diagnosis process to dynamically weight key areas while maintaining overall stability.
[0052] Finally, the weights of all modalities must be normalized after each adjustment to meet the constraint that the sum of the weights is 1. This step prevents a single modality from occupying too much proportion, ensuring that different sources of image data maintain a reasonable proportion in information fusion, avoiding information bias caused by overfitting or over-reliance on a single modality.
[0053] In one possible implementation, in the feature extraction process, the morphological features and diffusion features of T2-weighted fluid-attenuated inversion recovery (FLAIR) images and diffusion-weighted imaging (DWI) images are accurately calculated and quantified to enhance the recognition ability of the lesion area and help make more accurate auxiliary diagnosis.
[0054] Specifically, first, for the morphological feature calculation of T2-weighted fluid-attenuated inversion recovery images, the lesion area is binarized. After binarization, the lesion area is subjected to multi-scale morphological closing operation. This process is achieved by first eroding and then dilating, aiming to remove small noise areas and fill small cavities in the lesion area. Through this process, the structure of the lesion area can be analyzed from multiple scale angles. The closing degree is calculated by comparing the area ratio of the lesion area before and after closing operation and converting it into a logarithmic value. The larger the closing degree, the more complete the boundary of the lesion area, and vice versa, which may have more cavities or irregular parts.
[0055] Next, for the quantification of edge irregularity, the boundary pixel sequence of the lesion area is first extracted. These boundary pixels represent the boundary between the lesion area and healthy tissue. By comparing these boundary pixels with the shape of a best-fitting ellipse, the standard distance deviation is calculated. This deviation quantifies the morphological irregularity of the lesion boundary. The more irregular the boundary, the larger the standard distance deviation, reflecting the complexity of the lesion area morphology. This feature plays an important role in identifying irregular or fuzzy boundary lesion areas, especially in complex lesions such as tumors or inflammation.
[0056] Second, for the diffusion feature calculation of diffusion-weighted imaging, a diffusion tensor ellipsoid is constructed in each voxel, which can describe the diffusion characteristics of water molecules in each direction. By calculating the length ratio of the principal axes of the ellipsoid, the diffusion eccentricity is obtained, which reflects the anisotropy of the diffusion behavior. If the eccentricity is high, it means that the diffusion in this area has obvious directionality, which is usually related to the orientation of white matter fibers. In diffusion-weighted imaging, the spatial gradient of the radial diffusion coefficient is calculated along the orientation of the white matter fibers, which can reveal the structural changes and spatial distribution characteristics of the white matter fibers. By analyzing these diffusion features, the recognition accuracy of brain lesions can be further improved, especially for the early diagnosis of neurodegenerative diseases and white matter lesions.
[0057] Through the fine processing of image data, the accuracy of magnetic resonance imaging in brain disease diagnosis can be effectively improved, providing more rich auxiliary information for clinical diagnosis.
[0058] In one possible implementation, first, in the dual-task collaborative mechanism, the classification task and the staging task are carried out in parallel and implemented using a random forest framework. The disease classification tree uses Gini impurity as the splitting criterion. Gini impurity is a standard for measuring sample purity when dividing data. Selecting Gini impurity as the splitting criterion can effectively reduce the error during classification and enhance the model's ability to distinguish different types of diseases. For the clinical staging tree, the mean square error is used as the splitting criterion. The mean square error is used to measure the difference between the predicted value and the actual value. Therefore, it can better handle continuous staging tasks and accurately predict different stages of disease development. The feature importance score output by the classification tree can quantify the contribution of each feature to the classification task, and this score is also used to screen the input features of the staging tree to ensure that the staging tree only uses features with high predictive power for disease staging, thereby improving the accuracy of the model.
[0059] Next, the follow-up data correction mechanism is another important part of this method. As time goes by, the patient's condition may change, so the follow-up data needs to be dynamically corrected. In this mechanism, the Mahalanobis distance change rate of the feature vector of the same brain region is calculated. The Mahalanobis distance is used to measure the similarity between data points. When the change rate of the feature vector of the same brain region during follow-up exceeds a certain set progression threshold, it indicates that the lesion in the region may have progressed significantly. In this case, reducing the feature weight of the brain region in the staging tree can avoid over-prediction due to disease progression, thereby making staging more accurate. If negative fluctuations occur in the follow-up data, that is, the feature value drops abnormally, this may be caused by image quality problems. Therefore, the image quality review process is triggered to further ensure the quality of the imaging data.
[0060] Through reasonable mechanism design, not only the adaptability and stability of the model are enhanced, but also the efficiency and accuracy of disease diagnosis and staging are ensured, providing clinicians with a more reliable auxiliary diagnostic tool.
[0061] In one possible implementation, the method for determining a specific multiple threshold of a healthy reference value can achieve more accurate abnormal threshold setting by combining the imaging characteristics of healthy people and the differences in disease types, thereby improving diagnostic accuracy.
[0062] Specifically, first, the relevant operations are performed in the healthy population template library. In step a, the mean and standard deviation of the pixel intensity in the periventricular region are calculated. The periventricular region is an area of physiological importance in brain imaging, and the pixel intensity in this area has a relatively stable distribution pattern in healthy people. By calculating the mean and standard deviation of this region, the range of pixel intensity in this area under normal conditions can be determined, providing basic data for subsequent identification of outliers.
[0063] In step b, the first preset percentile of the healthy population pixel intensity distribution is further calculated. The percentile represents the value of the healthy population at a certain position in the pixel intensity distribution, which helps to more accurately demarcate the boundary between normal and abnormal. The setting of this percentile can be flexibly adjusted according to actual needs. For different populations or different disease states, different percentiles can be selected to adapt to different diagnostic needs.
[0064] In step c, the intensity value exceeding the specific multiple of the percentile is set as the abnormal threshold. The setting of this multiple plays a key role in determining which pixel intensity is considered abnormal. By setting the multiple threshold, the abnormal range can be flexibly controlled to be neither too broad nor too narrow, thereby improving the early diagnosis ability of the disease.
[0065] The specific multiple is dynamically adjusted according to the type of disease, which further enhances the flexibility and adaptability of the method. For high-risk populations of neurodegenerative diseases, since the early features of these diseases are usually weak, a lower multiple threshold can make the detection more sensitive and capture possible early changes. For subjects of different ages, the adjustment range is positively correlated with the age of the subject. The older the subject, the more complex the possible brain structure changes, so the multiple of the threshold should be appropriately increased to avoid misjudgment of the normal aging process.
[0066] The determination method of the specific multiple threshold of the healthy reference value not only improves the diagnostic accuracy of magnetic resonance images, but also can be flexibly applied in different clinical environments, thereby providing doctors with a more scientific and accurate auxiliary diagnostic tool.
[0067] In one possible implementation, first, in the process of selecting a pre-trained three-dimensional residual convolutional network as the teacher model, the teacher model obtains strong feature extraction capability through pre-training. The three-dimensional residual convolutional network is a deep convolutional neural network that can effectively process three-dimensional data and performs well in processing three-dimensional images such as magnetic resonance images that have spatial information. The design of the residual network solves the problem of gradient vanishing or explosion in the training of deep networks by introducing a skip connection, thereby improving the performance of the model in complex image data.
[0068] Next, the student model is a fully connected lightweight network. Compared with the teacher model, this network is more concise in structure and has fewer parameters, so it is faster to train and more efficient to infer. Through this design, the student model can significantly reduce the consumption of computing resources while maintaining high accuracy.
[0069] The training steps of the student model include:
[0070] Freezing the convolutional layer weights of the teacher model: At this stage, the teacher model has been pre-trained and has strong feature extraction capabilities. Therefore, freezing its convolutional layer weights is to maintain its feature extraction capabilities without changing during the student model training process. This step ensures that the student model can learn the high-level features of the teacher model without destroying its existing feature representation capabilities.
[0071] Extracting the feature maps of the last convolutional layer of the teacher model: The last convolutional layer of the teacher model usually contains the most abstract representation of the input data, which can capture complex spatial information in the image. Therefore, extracting the feature maps of this layer from the teacher model is to pass these high-level features to the student model to help it obtain stronger recognition capabilities.
[0072] Spatially pooling the feature maps into vectors as input to the student network: Since the feature map is a high-dimensional spatial matrix, it needs to be converted into a vector through spatial pooling operation to input into the student model for further processing. The spatial pooling process usually includes maximum pooling or average pooling, through which the student model can obtain a concise and effective feature representation.
[0073] Minimizing the KL divergence between the probability distributions of the teacher and student network outputs: KL divergence is an index that measures the difference between two probability distributions. In the knowledge distillation process, the goal of the student model is to mimic the output probability distribution of the teacher model. By minimizing the KL divergence, the student model gradually approximates the behavior of the teacher model, enabling it to learn valuable knowledge from the teacher model while ensuring that the probability distribution of the output is as close as possible to the teacher model's output, thereby improving its performance.
[0074] By migrating the knowledge of the complex teacher model to the simplified student model, the lightweight model is finally achieved. The student model has fewer parameters and faster inference speed compared to the teacher model, making it suitable for scenarios with high computational resource requirements in practical applications.
[0075] The student model can quickly output results during inference due to its simple structure, which is particularly important in clinical diagnosis, allowing it to quickly respond to medical image data processing needs. Although the student model structure is simplified, it can effectively retain the excellent feature extraction capabilities of the teacher model through knowledge distillation, ensuring the accuracy and robustness of the diagnosis results. This approach greatly improves the performance of lightweight models in actual use.
[0076] Through knowledge distillation, the student model can effectively borrow the knowledge obtained by the teacher model trained on large-scale data without starting from scratch, which significantly reduces training time and data requirements while improving performance on small data sets.
[0077] Therefore, the knowledge distillation migration process can provide an efficient and low resource consumption solution under the premise of ensuring diagnostic accuracy, which is suitable for quickly and accurately processing magnetic resonance imaging data in a clinical environment.
[0078] Correspondingly, the embodiment of the present application also provides a magnetic resonance imaging auxiliary diagnosis system based on multi-modal fusion, comprising a memory configured to store instructions, and a processor configured to call the instructions from the memory and implement the magnetic resonance imaging auxiliary diagnosis method based on multi-modal fusion as any of the embodiments of the present application when executing the instructions.
[0079] The present application encompasses any alternative, modification, equivalent method and scheme made on the essence and scope of the present application. In order to make the public have a thorough understanding of the present application, specific details are described in the following preferred embodiments of the present application, and the present application can also be fully understood without the description of these details to those skilled in the art. In addition, in order to avoid unnecessary confusion to the essence of the present application, well-known methods, processes, procedures, elements and circuits, etc. are not described in detail.
[0080] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principle of the present application, a number of improvements and refinements can be made, which should also be regarded as the protection scope of the present application.
Claims
1. A magnetic resonance imaging-assisted diagnosis method based on multimodal fusion, characterized in that: include: Step 1: Synchronously acquire the subject's three-dimensional T1-weighted structural image, T2-weighted fluid-attenuated inversion recovery image, and diffusion-weighted imaging data. Use the T1-weighted image as a reference for spatial registration and perform skull stripping and grayscale normalization. Step 2: Based on the individualized brain region segmentation of the brain anatomical atlas, the lesion sensitivity weight of each modality is calculated for each segmented brain region. The weight is obtained by quantifying the following parameters: The clarity of the gray-white matter junction structure on T1-weighted images, the abnormal signal distribution density on T2-weighted fluid-attenuated inversion recovery images, and the degree of microstructural disorder on diffusion-weighted imaging; Step 3: Extract multimodal image features in each brain region and perform adaptive weighted fusion based on the sensitivity weight. When the probability of a brain lesion detected exceeds a dynamic threshold, increase the weight of the corresponding modality. Step 4: Input the fusion features of the entire brain region into the multi-task classifier, and simultaneously output the disease type and clinical stage. If follow-up images are available, the staging results are corrected by the feature vector change rate.
2. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 1, characterized in that: Grayscale normalization in step 1 includes: Match the standard grayscale histogram of the same age group from the pre-built healthy population template library as the reference distribution; The current image grayscale is mapped to the reference distribution interval through the cumulative distribution function transformation. The mapping function construction process is: Calculate the grayscale cumulative probability distribution of the current image and the reference image respectively, find the closest probability value in the reference cumulative distribution for each grayscale level of the current image, and map the grayscale level to the grayscale level with the corresponding probability value in the reference distribution; An edge-preserving filtering algorithm is used to eliminate the boundary blur caused by standardization. Specifically: The gradient amplitude of each pixel in the image is calculated as the spatial weight factor, the size of the filter kernel function is adaptively adjusted according to the gradient amplitude, and a weighted smoothing operation is performed on the grayscale mapped image.
3. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 1, characterized in that: The individualized brain region segmentation in step 2 includes: The standard brain anatomical atlas is mapped to individual space using a diffeomorphic deformation algorithm. The specific steps are as follows: a: Calculate the mutual information between the individual T1 image and the atlas template as the similarity measure; b: Iterative optimization of the velocity field control point displacements by gradient descent method; c: After each iteration, the Jacobian value of the deformation field is detected and a fluid dynamics correction is applied to the areas where negative values appear; The segmentation results include the hippocampus, basal ganglia, and subcortical nuclei, and their boundary accuracy is guaranteed by the following methods: A higher resolution deformable mesh is used at the gray-white matter junction, and the cerebrospinal fluid signal intensity is used as the edge constraint.
4. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 1, characterized in that: The lesion sensitivity weight calculation in step 2 includes: Quantification of gray-white matter boundary clarity: Equally spaced sampling points are set on the gray-white matter boundary line. A neighborhood window of preset size is established with each point as the center. The statistical deviation of the gradient amplitude within the window is calculated as the clarity index. Abnormal signal distribution density quantification: Identify the extended area related to the preset anatomical structure around the ventricle and count the percentage of pixels in this area that exceed a specific threshold of the healthy reference value; Quantification of microstructural disorder: Extract the fractional anisotropy map of diffusion-weighted imaging, divide the target brain area into sub-blocks of adaptive size, and calculate the average information entropy of each sub-block as the disorder index.
5. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 1, characterized in that: The dynamic weight adjustment in step 3 includes: Real-time calculation of brain lesion probability: The multimodal feature vector of the current brain region is input into a lightweight neural network. The last layer of the network uses a sigmoid function to output the probability value. The network weights are obtained by transferring knowledge distillation from a pre-trained 3D convolutional network. The execution conditions of the weight promotion mechanism include: The linear boost mode is activated when the lesion probability exceeds a first threshold; The exponential boost mode is activated when the lesion probability exceeds a second threshold; After the boosting, the weights need to be renormalized to satisfy the constraint that the sum of the weights of each modality is 1.
6. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 1, characterized in that: In the feature extraction of step 3: The morphological features of T2-weighted fluid-attenuated inversion recovery images are calculated as follows: A multi-scale morphological closing operation was performed on the binary lesion area, and the closing degree was defined as the logarithmic transformation value of the area ratio before and after the closing operation; Edge irregularity was quantified by the following steps: extracting the boundary pixel sequence of the lesion area and calculating the standard distance deviation of the sequence from the best-fit ellipse; The calculation of diffusion characteristics of diffusion-weighted imaging includes: constructing a diffusion tensor ellipsoid at the voxel level, calculating the eccentricity based on the ratio of the main axis lengths of the ellipsoid, and calculating the spatial gradient of the radial diffusion coefficient along the direction of the white matter fibers.
7. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 1, characterized in that: The multi-task classifier construction in step 4 includes: The dual-task collaborative mechanism of the random forest framework: the disease classification tree uses Gini impurity as the splitting criterion, the clinical staging tree uses mean square error as the splitting criterion, and the feature importance score output by the classification tree is used to screen the input features of the staging tree; Follow-up data correction mechanism: Calculate the rate of change of the Mahalanobis distance of the feature vectors of the same brain region. When the rate of change exceeds the progression threshold, reduce the feature weight of the brain region in the staging tree. When the rate of change fluctuates negatively, trigger the image quality review process.
8. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 4, characterized in that: Method for determining the specific multiple threshold of the health reference value: Execute in the healthy population template library: a: Calculate the mean and standard deviation of pixel intensity in each periventricular area; b: The preset percentile of the pixel intensity distribution of healthy people; c: Set the intensity value that exceeds a certain multiple of the percentile as the abnormal threshold; The specific multiple is dynamically adjusted according to the disease type: a: Use lower multiples for people at high risk of neurodegenerative diseases; b: The adjustment amplitude is positively correlated with the age of the subjects.
9. The method for MRI-assisted diagnosis based on multimodal fusion according to claim 5, characterized in that: The knowledge distillation migration process includes: The teacher model selects a pre-trained 3D residual convolutional network; The student model is a fully connected lightweight network, and its training steps are: a: Freeze the convolutional layer weights of the teacher model; b: Extract the feature map of the last convolutional layer of the teacher model; c: Pool the feature map space into a vector as the input of the student network; d: Use KL divergence to minimize the difference in probability distribution between the teacher and student network outputs.
10. A magnetic resonance imaging-assisted diagnosis system based on multimodal fusion, characterized in that: It includes a memory configured to store instructions, a processor configured to call the instructions from the memory and to implement a magnetic resonance imaging-assisted diagnosis method based on multimodal fusion as described in any one of claims 1 to 9 when executing the instructions.
Citation Information
Cited By
MRI (Magnetic Resonance Imaging) image feature extraction method for children brain development dynamic evaluation
CN121904021A
Neurodegenerative disease early auxiliary diagnosis and prediction method and system for clinical magnetic resonance imaging
CN122091166A