Intelligent classification method and system for brain tumors based on multi-modal magnetic resonance

By combining multimodal magnetic resonance imaging processing and feature fusion with visual encoders and language models, the problems of long processing time and low accuracy in brain tumor classification have been solved, achieving more accurate brain tumor classification.

CN121600329BActive Publication Date: 2026-05-12XIANGYA HOSPITAL CENT SOUTH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIANGYA HOSPITAL CENT SOUTH UNIV
Filing Date
2026-01-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies for brain tumor classification suffer from problems such as long processing time, high error rates, and low accuracy in computer classification, especially in capturing fine anatomical details.

Method used

A brain tumor intelligent classification method based on multimodal magnetic resonance imaging is adopted. By sampling and processing magnetic resonance images, anatomical details are extracted using a visual encoder, and feature fusion and projection are performed by combining a dual-layer MLP projector and a language model to achieve accurate classification.

Benefits of technology

While ensuring computational efficiency, it captures more refined anatomical details, improving the accuracy and efficiency of brain tumor classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600329B_ABST
    Figure CN121600329B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of brain tumor classification detection, in particular to a brain tumor intelligent classification method and system based on multi-modal magnetic resonance, which comprises the following steps: sampling and processing magnetic resonance images of the brain of a target object to obtain up-sampling images and down-sampling images; inputting the down-sampling images into a visual encoder to obtain global visual tokens, dividing the up-sampling images into multiple subgraphs, inputting the subgraphs into the visual encoder to obtain visual tokens; removing non-spatial patch tokens in the visual tokens to obtain anatomical detail features of the subgraphs, splicing the anatomical detail features of the subgraphs to obtain spliced features; fusing the spliced features and the global visual tokens to obtain fusion tokens, splicing the fusion tokens of the magnetic resonance images of each sequence to obtain fusion features; projecting the fusion features to obtain first projection features, inputting the first projection features into a language model to obtain a brain tumor classification label of the target object. The method can realize accurate classification of brain tumors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of brain tumor classification and detection technology, and in particular to a brain tumor intelligent classification method and system based on multimodal magnetic resonance imaging. Background Technology

[0002] Brain tumors are abnormally growing clumps of cells in or around the brain tissue. They can be classified as benign or malignant, and based on their location of origin, they can be divided into primary and secondary brain tumors. Different types of brain tumors vary greatly in their biological behavior, growth rate, invasiveness, and response to treatment. Only with accurate classification can the most appropriate treatment plan be developed.

[0003] Clinically, brain tumor classification primarily relies on manual interpretation by neuroradiologists, which is time-consuming and prone to errors. This problem is exacerbated by the global shortage of neuroradiologists, and acquiring expertise in this field requires significant time and systematic training. Therefore, there is an urgent need to design an intelligent classification and prediction system for brain tumors. Currently, computer-based brain tumor classification and prediction mainly relies on data augmentation and primary feature extraction, multi-scale feature extraction and enhancement, and feature fusion. However, this approach cannot capture more refined anatomical details while maintaining computational efficiency. For example, patent application CN 120107703 A, which uses data augmentation and primary feature extraction, multi-scale feature extraction and enhancement, and feature fusion for classification and prediction, suffers from low classification accuracy. Summary of the Invention

[0004] Therefore, it is necessary to provide a method and system for intelligent classification of brain tumors based on multimodal magnetic resonance imaging to address the aforementioned technical problems, thereby enabling accurate classification of brain tumors.

[0005] A method for intelligent classification of brain tumors based on multimodal magnetic resonance imaging, the method comprising:

[0006] S1. The magnetic resonance images of the target brain are sampled to obtain upsampled and downsampled images; the magnetic resonance images include multiple sequences of magnetic resonance images;

[0007] S2. Input the downsampled image into the visual encoder to obtain a global visual token; divide the upsampled image into multiple sub-images and input the sub-images into the visual encoder to obtain a visual token;

[0008] S3. Remove the non-spatial patch tokens from the visual tokens to obtain the anatomical detail features of each sub-image, and stitch the anatomical detail features of each sub-image to obtain the stitched features;

[0009] S4. Fuse the stitching features and the global visual token to obtain a fusion token, and stitch the fusion tokens of the magnetic resonance images of each sequence to obtain a fusion feature;

[0010] S5. Project the fused features using a dual-layer MLP projector to obtain a first projected feature. Input the first projected feature into a language model to obtain the brain tumor classification label of the target object.

[0011] The beneficial effect of this application is that it can capture more refined anatomical details while ensuring computational efficiency, thereby making the final brain tumor classification label more accurate.

[0012] In one embodiment, step S5 includes:

[0013] The fused features are projected using a dual-layer MLP projector to obtain the first projected feature, and a text instruction for classification prediction is preset.

[0014] The first projection feature and the text instruction are input into the language model to obtain the brain tumor classification label of the target object.

[0015] In this application, the first projection feature and the text instruction are simultaneously input into the language model, so that the task performed by the language model is a classification prediction task, rather than other tasks, thereby obtaining the brain tumor classification label of the target object.

[0016] In one embodiment, the language model, the visual encoder, and the two-layer MLP projector are components of a brain-oriented visual-language multimodal model. The method further includes a training process for the visual-language multimodal model, which includes:

[0017] Obtain a training set; the training set includes two-dimensional magnetic resonance images with text descriptions and multi-sequence magnetic resonance images with radiation detection reports;

[0018] Based on the two-dimensional magnetic resonance images and the multi-sequence magnetic resonance images, the dual-layer MLP projector and the anatomical detail-aware encoder in the visual-language multimodal model are trained using the gradient descent method to obtain the trained dual-layer MLP projector and the anatomical detail-aware encoder.

[0019] In this application, a gradient descent method is used to train a two-layer MLP projector and an anatomical detail-aware encoder in a visual-language multimodal model based on two-dimensional magnetic resonance images and multi-sequence magnetic resonance images. The trained two-layer MLP projector and anatomical detail-aware encoder are obtained, which can focus on the representation learning of two-dimensional magnetic resonance images, enabling the visual-language multimodal model to establish the basic ability of category recognition, and laying a solid foundation for subsequent advanced three-dimensional stereoscopic analysis.

[0020] In one embodiment, the training set further includes three-dimensional magnetic resonance images labeled with classification tags, and the method further includes:

[0021] Extract magnetic resonance images from the training set;

[0022] When the magnetic resonance images extracted from the training set are two-dimensional magnetic resonance images, the two-dimensional magnetic resonance images are input into the visual encoder to extract the first visual embedding features from the two-dimensional magnetic resonance images. The first visual embedding features are then projected through the trained two-layer MLP projector to obtain the second projection features. Based on the second projection features, the language model is used to generate a predicted text description of the two-dimensional magnetic resonance images. Based on the predicted text description of the two-dimensional magnetic resonance images and the text description, the first model loss is calculated.

[0023] When the magnetic resonance images extracted from the training set are three-dimensional magnetic resonance images or multi-sequence magnetic resonance images, the extracted magnetic resonance images are input into the visual encoder to extract the second visual embedding features from the extracted magnetic resonance images. The second visual embedding features are then projected onto the trained two-layer MLP projector to obtain the third projection features. Based on the third projection features, the language model is used to generate the prediction results of the extracted magnetic resonance images. Based on the prediction results and the true results of the extracted magnetic resonance images, the second model loss is calculated.

[0024] Based on the first model loss or the second model loss, the trained dual-layer MLP projector, LR-HR cross-resolution attention module and language model in the visual language multimodal model are trained to obtain the trained dual-layer MLP projector, LR-HR cross-resolution attention module and language model.

[0025] The training set is used to perform secondary training on the trained dual-layer MLP projector, LR-HR cross-resolution attention module, and language model to obtain a trained visual-language multimodal model.

[0026] In this application, on the one hand, while retaining high-precision slice-level perception capabilities, it can achieve smooth, efficient, and coherent knowledge transfer and fusion from two-dimensional local representations to three-dimensional global contexts, thus building a unified cross-dimensional tumor understanding capability for visual language multimodal models. On the other hand, based on the established cross-dimensional representations, the visual language multimodal model can be finely tuned through a training set, enabling the visual language multimodal model to possess advanced reasoning capabilities, achieving a leap from "being able to understand" to "being able to express and analyze".

[0027] In one embodiment, the method further includes:

[0028] Obtain the target model parameters of the visual language multimodal model in multiple training rounds, and construct intermediate visual language multimodal models based on each target model parameter, with the network architecture of each intermediate visual language multimodal model being consistent.

[0029] For the three-dimensional magnetic resonance images in the training set, each of the intermediate visual language multimodal models is used to perform classification prediction to obtain multiple predicted categories, and the category with the highest frequency among the multiple predicted categories is taken as the final predicted category of the three-dimensional magnetic resonance image.

[0030] From the predicted categories, determine the number of categories that are inconsistent with the actual categories of the three-dimensional magnetic resonance image, and calculate the confidence level of the three-dimensional magnetic resonance image based on the number of categories and the total number of predicted categories;

[0031] Based on the three-dimensional magnetic resonance image, the confidence level of the three-dimensional magnetic resonance image, and the final predicted category of the three-dimensional magnetic resonance image, a triplet is constructed;

[0032] A reliability dataset is constructed based on the triples of the three-dimensional magnetic resonance images from multiple frames.

[0033] The visual language multimodal model is optimized for parameters based on the reliability dataset, and its output format is reconstructed so that the visual language multimodal model outputs the confidence level of the brain tumor classification label.

[0034] In this application, triples are constructed based on three-dimensional magnetic resonance images, the confidence level of three-dimensional magnetic resonance images, and the final predicted category of three-dimensional magnetic resonance images. Based on the triples of multiple frames of three-dimensional magnetic resonance images, a reliability dataset is constructed, and the parameters are optimized using the reliability dataset. This allows consensus reliability knowledge to be distilled into a visual-language multimodal model, enabling simultaneous learning of classification prediction and confidence prediction.

[0035] In one embodiment, the plurality of sequences of magnetic resonance images C m The expression is:

[0036] ;

[0037] in, For magnetic resonance images of the scan plane v0 in a T1-weighted sequence, The magnetic resonance image of scan plane v0 in the T1c weighted sequence. This is the magnetic resonance image of the scan plane v3 in the s3 sequence. This is the magnetic resonance image of the scan plane v4 in the s4 sequence. The s3 sequence is the magnetic resonance image of scan plane v5 in the s5 sequence; scan plane v0 is any scan plane in the T1 weighted sequence; when the sequence includes a T2 weighted sequence and a FLAIR sequence, the s3 sequence is a T2 weighted sequence, the s4 sequence is a FLAIR sequence, scan plane v3 is any scan plane in the T2 weighted sequence, and scan plane v4 is any scan plane in the FLAIR sequence; when the sequence includes a target sequence, the s3 sequence is the target sequence, the s4 sequence is any sequence, scan plane v3 is any scan plane in the target sequence, scan plane v4 is any scan plane in any sequence, and the target sequence is either a T2 weighted sequence or a FLAIR sequence; when neither a T2 weighted sequence nor a FLAIR sequence exists in the sequence, both the s3 and s4 sequences are any sequences.

[0038] This application can fully utilize multiple sequences for classification prediction while ensuring that the computational cost is not excessive. Furthermore, it can dynamically adapt to the actual sequence and planar availability of the target object, ensuring robustness even in the event of missing sequences.

[0039] In one embodiment, the plurality of sequences constitute a sequence combination, and the fusion feature includes fusion features corresponding to the plurality of sequence combinations. Step S5 includes:

[0040] The dual-layer MLP projector is used to project the fused features corresponding to multiple sequence combinations to obtain the first projection feature of each sequence combination.

[0041] Each of the first projection features is input into the language model to obtain the initial brain tumor classification label for each sequence combination and the confidence level of each initial brain tumor classification label;

[0042] The most frequent label among the initial brain tumor classification labels is used as the brain tumor classification label of the target object, and the confidence level of the brain tumor classification label is calculated.

[0043] brain tumor classification labels confidence level The calculation formula is:

[0044] ;

[0045] Where N1 is the number of initial brain tumor classification labels that are consistent with the brain tumor classification labels, and M is the total number of initial brain tumor classification labels. For the m-th initial brain tumor classification label y m Confidence level, Let Kronecker function be used.

[0046] In this application, by outputting the current detection report of the target object, which includes the brain tumor classification label and the confidence level of the brain tumor classification label, the brain tumor classification label and the confidence level of the brain tumor classification label of the target object can be observed intuitively.

[0047] In one embodiment, the method further includes:

[0048] Output the current detection report of the target object, the current detection report including the brain tumor classification label and the confidence level of the brain tumor classification label.

[0049] In this application, by outputting the current detection report of the target object, which includes the brain tumor classification label and the confidence level of the brain tumor classification label, the brain tumor classification label and the confidence level of the brain tumor classification label of the target object can be observed intuitively.

[0050] A multimodal magnetic resonance imaging-based intelligent classification system for brain tumors, the system comprising:

[0051] A sampling subsystem is used to sample the magnetic resonance images of the target object to obtain upsampled and downsampled images; the magnetic resonance images include multiple sequences of magnetic resonance images;

[0052] The encoding subsystem is used to input the downsampled image into the visual encoder to obtain a global visual token, divide the upsampled image into multiple sub-images, and input the sub-images into the visual encoder to obtain a visual token;

[0053] A stitching subsystem is used to remove non-spatial patch tokens from the visual tokens, obtain the anatomical detail features of each sub-image, and stitch the anatomical detail features of each sub-image to obtain stitched features;

[0054] A fusion subsystem is used to fuse the stitching features and the global visual token to obtain a fusion token, and to stitch the fusion tokens of the magnetic resonance images of each sequence to obtain fusion features;

[0055] A classification subsystem is used to project the fused features using a two-layer MLP projector to obtain a first projected feature, and input the first projected feature into a language model to obtain a brain tumor classification label for the target object.

[0056] The beneficial effect of the aforementioned intelligent brain tumor classification system based on multimodal magnetic resonance is that it can capture more refined anatomical details while ensuring computational efficiency, thereby making the final brain tumor classification label more accurate. Attached Figure Description

[0057] Figure 1 This is a diagram illustrating the application environment of a brain tumor intelligent classification method based on multimodal magnetic resonance imaging in one embodiment.

[0058] Figure 2 This is a flowchart illustrating a brain tumor intelligent classification method based on multimodal magnetic resonance imaging in one embodiment.

[0059] Figure 3 This is a schematic diagram of a structure for label prediction using a visual language multimodal model in one embodiment;

[0060] Figure 4 This is a schematic diagram illustrating the process of training a visual language multimodal model in one embodiment;

[0061] Figure 5 This is a schematic diagram of a differentiated preprocessing process for different resources in one embodiment;

[0062] Figure 6 This is a structural block diagram of a brain tumor intelligent classification system based on multimodal magnetic resonance imaging in one embodiment;

[0063] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0065] The intelligent brain tumor classification method based on multimodal magnetic resonance provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 interacts with server 104 via a wired / wireless channel. A data storage system stores the data that server 104 needs to process. The server samples the magnetic resonance imaging (MRI) images of the target object's brain to obtain upsampled and downsampled images; the MRI images include multiple sequences of MRI images. The server inputs the downsampled images into a visual encoder to obtain a global visual token, divides the upsampled images into multiple sub-images, and inputs the sub-images into the visual encoder to obtain visual tokens. The server removes non-spatial patch tokens from the visual tokens to obtain anatomical detail features of each sub-image, and stitches the anatomical detail features of each sub-image to obtain a stitched feature. The server fuses the stitched feature and the global visual token to obtain a fusion token, and stitches the fusion tokens of each sequence of MRI images to obtain a fusion feature. The server projects the fusion feature using a dual-layer MLP projector to obtain a first projection feature, and inputs the first projection feature into a language model to obtain a brain tumor classification label for the target object. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, etc. Server 104 can be a single server, a server cluster consisting of multiple servers, or a cloud computing center consisting of multiple servers.

[0066] In one embodiment, such as Figure 2 As shown, a method for intelligent classification of brain tumors based on multimodal magnetic resonance imaging is provided, which is then applied to... Figure 1 Taking server 104 as an example, the following steps are included:

[0067] S1. The magnetic resonance images of the target brain are sampled to obtain upsampled and downsampled images; the magnetic resonance images include multiple sequences of magnetic resonance images;

[0068] Among them, magnetic resonance imaging refers to MRI (Magnetic Resonance Imaging).

[0069] An upsampled image is obtained by upsampling a magnetic resonance imaging (MRI) image, while a downsampled image is obtained by downsampling a MRI image. The resolution of an upsampled image is lower than that of the MRI image, and the resolution of a downsampled image is higher than that of the MRI image. For example, if the MRI image size is... The size of the downsampled image after downsampling is The size of the upsampled image is N is the number of slices in the magnetic resonance image. This represents the resolution of the magnetic resonance imaging.

[0070] S2. Input the downsampled image into the visual encoder to obtain a global visual token. Divide the upsampled image into multiple sub-images and input the sub-images into the visual encoder to obtain visual tokens.

[0071] The visual encoder is a component of the brain-oriented visual-language multimodal model, which also includes a two-layer MLP (Multilayer Perceptron) projector and a language model.

[0072] The global visual token is a low-resolution visual token. Specifically, the global visual token is the [CLS] (Classification Token) of the last layer of the visual encoder.

[0073] An upsampled image can be divided into multiple sub-images using a grid, with each sub-image not overlapping. For example, continuing from the previous example, the upsampled image... Divide into four non-overlapping subgraphs, then each subgraph is: .

[0074] The visual tokens obtained by inputting the sub-images into the visual encoder are high-resolution tokens, and each sub-image corresponds to a visual token.

[0075] S3. Remove non-spatial patch tokens from the visual tokens to obtain the anatomical detail features of each sub-image, and then stitch together the anatomical detail features of each sub-image to obtain the stitched features.

[0076] In this process, the subimage is spatially divided into several local regions. The vector representation formed by embedding each region is called a "patch token." Since these tokens correspond to specific spatial locations in the subimage, they are also called "spatial patch tokens." Removing non-spatial patch tokens from the visual tokens can be understood as retaining all spatial patch tokens from the last layer of the visual encoder. Spatial patch tokens retain the spatial layout information of the magnetic resonance imaging (MRI) image, and do not include [CLS].

[0077] S4. Combine the stitching features and the global visual token to obtain the fusion token. Combine the fusion tokens of the magnetic resonance images from each sequence to obtain the fusion features.

[0078] The fusion of concatenated features and the global visual token can be achieved through the LR-HR (Low Resolution-High Resolution) cross-attention module of the visual-language multimodal model. Specifically, the concatenated features are injected into the global visual token through the LR-HR cross-attention module, resulting in a low-resolution token that incorporates high-resolution information—the fused token. Here, the global visual token serves as the query, and the concatenated features serve as the context (key and value), outputting the final fused token. An attention mechanism is used when fusing the concatenated features and the global visual token, which effectively fuses them while maintaining computational efficiency.

[0079] The fusion tokens of the MRI images from different sequences are stitched together according to the slice dimension. For example, the MRI images include 5 sequences: a transverse image of a T1-weighted sequence, a transverse image of a T1c-weighted sequence, a sagittal image of a T1c-weighted sequence, a transverse image of a T2-weighted sequence, and a transverse image of a FLAIR sequence. Each sequence has N=32 slices, ultimately yielding a fusion feature that integrates complementary diagnostic information from multiple sequences. .

[0080] S5. Use a two-layer MLP projector to project the fused features to obtain the first projected feature. Input the first projected feature into the language model to obtain the brain tumor classification label of the target object.

[0081] The two-layer MLP projector is a tool that projects fused features onto the language space of a language model, thereby transforming the fused features into first projected features that the language model can understand. The format of the first projected features is one that the language model can comprehend. This projection step can effectively bridge the modal gap by transforming fused features into a format that the language model can process.

[0082] The language model refers to LLM (Large Language Model). For example, a hybrid causal-full attention large language model uses a pre-trained Llama-3.1-8B-Instruct model as its core. This model possesses mainstream language capabilities, a grouped query attention mechanism, and open-source scalability. It also introduces a hybrid attention strategy, applying full attention to the first projection feature to capture complex visual dependencies within the same 3D or 2D MRI image sequence, while retaining causal attention to ensure logical coherence. This hybrid attention architecture enables the language model to simultaneously model cross-modal visual associations and linguistic structural consistency, significantly improving multimodal medical inference performance.

[0083] Furthermore, the current detection report of the target object is output, which includes the brain tumor classification label of the target object.

[0084] The aforementioned intelligent brain tumor classification method based on multimodal magnetic resonance imaging can capture more refined anatomical details while ensuring computational efficiency, thereby making the final brain tumor classification label more accurate.

[0085] In one embodiment, step S5 includes:

[0086] The fused features are projected using a two-layer MLP projector to obtain the first projected feature, and a text instruction for classification prediction is preset.

[0087] The first projection features and text instructions are input into the language model to obtain the brain tumor classification label of the target object.

[0088] Furthermore, the first projection features, text instructions, and metadata of the target object are input into the language model to obtain the brain tumor classification label of the target object.

[0089] In a specific application, the first projection feature, the target object's age, gender, and the text instruction "Please make a classification prediction based on this data from a certain patient" together constitute a multimodal prompt. The multimodal prompt is input into a language model for brain tumor classification, and the brain tumor classification label of the target object is output.

[0090] In this embodiment, the first projection feature and the text instruction are simultaneously input into the language model, so that the task performed by the language model is a classification prediction task, rather than other tasks, thereby obtaining the brain tumor classification label of the target object.

[0091] In one embodiment, the language model, visual encoder, and dual-layer MLP projector are components of a brain-oriented visual-language multimodal model. The intelligent brain tumor classification method based on multimodal magnetic resonance imaging also includes a training process for the visual-language multimodal model, which includes:

[0092] Obtain the training set; the training set includes two-dimensional magnetic resonance images with text descriptions and multi-sequence magnetic resonance images with radiation detection reports;

[0093] Based on two-dimensional magnetic resonance imaging and multi-sequence magnetic resonance imaging, a two-layer MLP projector and an anatomical detail-aware encoder in a visual-language multimodal model were trained using the gradient descent method, resulting in the trained two-layer MLP projector and anatomical detail-aware encoder.

[0094] Specifically, when training the two-layer MLP projector and the anatomical detail-aware encoder in the visual-language multimodal model using gradient descent based on two-dimensional magnetic resonance imaging and multi-sequence magnetic resonance imaging, only the parameters in the two-layer MLP projector and the anatomical detail-aware encoder are optimized, while the parameters of the visual encoder and the backbone of the language model are kept frozen, which can ensure training stability.

[0095] Furthermore, the parameters of the LR-HR cross-attention module in the dual-layer MLP projector and the anatomical detail-aware encoder are optimized using gradient descent.

[0096] Two-dimensional magnetic resonance imaging (MRI) images include MRI images with textual descriptions and 2D slices extracted from multi-sequence MRI images. Further, predefined masks are used to identify and separate the 2D slices of interest from the multi-sequence MRI images. Further, textual descriptions of the 2D slices are generated using an artificial intelligence model. For example, for sample slices with tumor masks in a public dataset, 2D slices with a tumor proportion greater than 5% are selected, and textual descriptions of the 2D slices are generated using RAG (Retrieval-Augmented Generation) and an artificial intelligence model; for in-hospital scan data, neuroradiologists manually select pathologically significant slices, and corresponding textual descriptions are also generated using RAG and an artificial intelligence model.

[0097] A schematic diagram of the structure for label prediction using a visual language multimodal model is shown below. Figure 3 As shown.

[0098] In this embodiment, a gradient descent method is used to train a two-layer MLP projector and an anatomical detail-aware encoder in a visual-language multimodal model based on two-dimensional magnetic resonance images and multi-sequence magnetic resonance images. The trained two-layer MLP projector and anatomical detail-aware encoder are obtained, which can focus on the representation learning of two-dimensional magnetic resonance images, enabling the visual-language multimodal model to establish the basic ability of category recognition, and laying a solid foundation for subsequent advanced three-dimensional stereoscopic analysis.

[0099] In one embodiment, the training set also includes labeled 3D magnetic resonance images. The method further includes:

[0100] Extract magnetic resonance images from the training set;

[0101] When the magnetic resonance images extracted from the training set are two-dimensional magnetic resonance images, the two-dimensional magnetic resonance images are input into the visual encoder to extract the first visual embedding features from the two-dimensional magnetic resonance images. The first visual embedding features are then projected through a trained two-layer MLP projector to obtain the second projection features. Based on the second projection features, a language model is used to generate a predicted text description of the two-dimensional magnetic resonance images. Based on the predicted text description of the two-dimensional magnetic resonance images and the text description, the first model loss is calculated.

[0102] When the magnetic resonance images extracted from the training set are three-dimensional magnetic resonance images or multi-sequence magnetic resonance images, the extracted magnetic resonance images are input into a visual encoder to extract the second visual embedding features from the extracted magnetic resonance images. The second visual embedding features are then projected through a trained two-layer MLP projector to obtain the third projection features. Based on the third projection features, a language model is used to generate prediction results for the extracted magnetic resonance images. Based on the prediction results and the true results of the extracted magnetic resonance images, the second model loss is calculated.

[0103] Based on the first model loss or the second model loss, the trained two-layer MLP projector, LR-HR cross-resolution attention module and language model in the visual language multimodal model are trained to obtain the trained two-layer MLP projector, LR-HR cross-resolution attention module and language model.

[0104] The trained dual-layer MLP projector, LR-HR cross-resolution attention module, and language model are retrained using the training set to obtain the trained visual-language multimodal model.

[0105] In this context, when extracting magnetic resonance images from the sampled training set, the probability of extracting a two-dimensional magnetic resonance image is the first probability value, and the probability of extracting a three-dimensional magnetic resonance image or a multi-sequence magnetic resonance image is the second probability value. For example, the probability of extracting a two-dimensional magnetic resonance image is 30%, and the probability of extracting a three-dimensional magnetic resonance image or a multi-sequence magnetic resonance image is 70%.

[0106] When the magnetic resonance images extracted from the sampling training set are two-dimensional magnetic resonance images, the two-layer MLP projector, LR-HR cross-resolution attention module, and language model of the visual-language multimodal model are trained based on the first model loss; when the magnetic resonance images extracted from the sampling training set are three-dimensional magnetic resonance images or multi-sequence magnetic resonance images, the two-layer MLP projector, LR-HR cross-resolution attention module, and language model of the visual-language multimodal model are trained based on the second model loss.

[0107] When the extracted magnetic resonance image is a three-dimensional magnetic resonance image, the prediction result is the predicted label of the three-dimensional magnetic resonance image, and the actual result is the labeled classification label of the three-dimensional magnetic resonance image; when the extracted magnetic resonance image is a multi-sequence magnetic resonance image, the prediction result is the predicted radiological detection report of the multi-sequence magnetic resonance image, and the actual result is the actual radiological detection report of the multi-sequence magnetic resonance image.

[0108] When training the parameters of a language model, the LoRA parameters in the language model are trained.

[0109] When projecting the first visual embedding feature and the second visual embedding feature, the two-layer MLP projector used is a two-layer MLP projector trained using gradient descent. When training the already trained two-layer MLP projector in the visual language multimodal model based on the first model loss or the second model loss, the trained two-layer MLP projector is also a two-layer MLP projector trained using gradient descent.

[0110] When using the training set to perform secondary training on the trained dual-layer MLP projector, LR-HR cross-resolution attention module, and language model, what is being trained is the dual-layer MLP projector, LR-HR cross-resolution attention module, and language model trained based on the first model loss or the second model loss.

[0111] In one specific application, 53.14% of the training set consists of 2D MRI images with text descriptions, 38.14% consists of multi-sequence MRI images with accompanying detection reports, and 8.72% consists of 3D MRI images with pre-labeled classification tags. A flowchart illustrating the training process for the visual-language multimodal model is shown below. Figure 4 As shown.

[0112] In this embodiment, on the one hand, while retaining high-precision slice-level perception capabilities, it can achieve smooth, efficient, and coherent knowledge transfer and fusion from two-dimensional local representations to three-dimensional global contexts, thus building a unified cross-dimensional tumor understanding capability for the visual language multimodal model. On the other hand, based on the established cross-dimensional representations, the visual language multimodal model can be finely tuned through the training set, enabling the visual language multimodal model to have advanced reasoning capabilities, achieving a leap from "being able to understand" to "being able to express and analyze".

[0113] In one embodiment, the method further includes:

[0114] Obtain the target model parameters of the visual language multimodal model in multiple training rounds, and construct intermediate visual language multimodal models based on each target model parameter. The network architecture of each intermediate visual language multimodal model is consistent.

[0115] For the three-dimensional magnetic resonance images in the training set, each intermediate visual language multimodal model is used to perform classification prediction, resulting in multiple predicted categories. The category with the highest frequency among the multiple predicted categories is taken as the final predicted category of the three-dimensional magnetic resonance image.

[0116] From the predicted categories, determine the number of categories that are inconsistent with the actual categories of the 3D magnetic resonance image, and calculate the confidence level of the 3D magnetic resonance image based on the number of categories and the total number of predicted categories;

[0117] Triads are constructed based on 3D magnetic resonance images, the confidence level of 3D magnetic resonance images, and the final predicted category of 3D magnetic resonance images.

[0118] A reliability dataset is constructed based on triples of multi-frame 3D magnetic resonance images;

[0119] The parameters of the visual-language multimodal model are optimized based on a reliability dataset, and the output format of the visual-language multimodal model is reconstructed to enable the visual-language multimodal model to output the confidence of brain tumor classification labels.

[0120] The multiple training epochs can be any number of training epochs of the visual-language multimodal model, or they can be the multiple training epochs before the iteration stops. For example, if the visual-language multimodal model has a total of 1000 training epochs, then the model parameters of 10 training epochs can be randomly selected from these 1000 epochs as the target model parameters, or the model parameters of the last 10 training epochs can be used as the target model parameters.

[0121] The network architecture of the intermediate visual language multimodal model is the same as that of the visual language multimodal model. The only difference between the intermediate visual language multimodal model and the visual language multimodal model is the model parameters.

[0122] The final expression for predicting the category is: ,in, For three-dimensional magnetic resonance imaging x i The final predicted category, where N is the number of intermediate visual language multimodal models, C is the set of predicted categories, and c is the predicted category. The three-dimensional magnetic resonance image x output by the k-th intermediate visual language multimodal model i The prediction category This is an indicator function.

[0123] The formula for calculating the confidence level of a three-dimensional magnetic resonance image is: ,in, For three-dimensional magnetic resonance imaging x i Confidence level, , For three-dimensional magnetic resonance imaging x i The predicted categories are related to 3D magnetic resonance images x i The number of categories whose actual categories do not match. For three-dimensional magnetic resonance imaging x i The actual category. The larger the value, the higher the inconsistency and the lower the confidence level. The smaller the value, the higher the confidence level.

[0124] The form of a triplet is as follows: After constructing a reliability dataset based on triples from multi-frame 3D magnetic resonance imaging, consensus reliability knowledge is distilled into a visual-language multimodal model (VMM) through confidence-aware reliability fine-tuning. Specifically, the VMM's parameters are optimized based on the reliability dataset, and its output format is reconstructed to generate confidence scores for brain tumor classification labels.

[0125] The triplet is Non-real label pairs This allows for the simultaneous learning of classification prediction and confidence levels.

[0126] During the parameter optimization of the visual-language multimodal model, only the parameters of the two-layer MLP projector, the LR-HR cross-attention module, and the LoRA parameter in the language model are updated, while the visual encoder and the language model backbone remain frozen.

[0127] In this embodiment, triples are constructed based on three-dimensional magnetic resonance images, the confidence level of three-dimensional magnetic resonance images, and the final predicted category of three-dimensional magnetic resonance images. Based on the triples of multiple frames of three-dimensional magnetic resonance images, a reliability dataset is constructed, and the parameters are optimized using the reliability dataset. This allows consensus reliability knowledge to be distilled into the visual language multimodal model, enabling simultaneous learning of classification prediction and confidence prediction.

[0128] In one specific application, the training set was sourced from 36 publicly available online resources and 12 collaborating medical institutions, encompassing approximately 46,000 target subjects. To ensure comprehensive data collection, a complete brain tumor keyword system was first constructed, covering 12 major brain tumor types and all their subtypes, synonyms, and related terms. Then, combining RAG and cueing engineering techniques, the training set was uniformly converted into a standard structured format for subsequent model training. The training set included: 2D MRI images with text descriptions, multi-sequence MRI images (T1, T1c, T2, T2-FLAIR) with detection reports, and 3D MRI images with pre-labeled classification information. The sources of the training set included the following two categories:

[0129] Category 1: Structured medical imaging resources, including public databases and imaging resources on professional platforms like Radiopaedia. These repositories provide readily available, directly accessible structured imaging data (3D multiparametric MRI volumes or 2D slices). They are typically tagged with pathology information, and some also contain patient metadata or radiology reports. Examples include public databases such as BraTS2023, TCIA (The Cancer Imaging Archive), Kaggle, and Ctisus.

[0130] The second category: non-dedicated image resources. For example, the PubMed Central archive.

[0131] A schematic diagram of the differentiated preprocessing process for different resources is shown below. Figure 5 As shown, the details are as follows:

[0132] 1) Preprocessing pipeline for structured databases and pre-processed datasets. If only classification labels are labeled, the format is: {Patient ID: 3D MRI, Classification Label}; if 3D multiparametric MRI, segmentation masks, and metadata are included, functional brain regions are located based on Brainnetome Atlas and tumor masks, all available metadata (imaging modality, gender, diagnostic labels, tumor description) is fused, and a complete radiology report is generated using an artificial intelligence model and RAG. For resources such as Radiopaedia that provide multi-sequence MRI and raw reports, irrelevant information is removed, and the data is reconstructed into a unified format: {Patient ID: 3D MRI, Pathology Label, Radiology Report}.

[0133] 2) Preprocessing Pipeline for Unstructured Online Resources. For brain tumor data from platforms such as PubMed Central (a public medical center) and ImageCLEF (a medical image retrieval and evaluation platform), which often contain noisy 2D MRI slices and misaligned or inconsistent text descriptions, a complete preprocessing pipeline is designed to transform the raw heterogeneous data into a unified structured format: {Patient ID: 2D MRI slice, text description, pathology label}. For images from PubMed Central literature, 2D brain MRI slices are first retrieved using predefined brain tumor keywords, and the text mentioned in the article is associated as candidate descriptions. Then, a three-stage refinement protocol is implemented: a Faster R-CNN detector pre-trained with MedICaT is used to segment composite image sub-images and filter irrelevant images; semantic-level splitting of composite titles is performed using artificial intelligence models and professional prompts; and a one-to-one correspondence between sub-images and sub-titles is established through bounding box coordinate alignment. After final integration, a total of 24,333 independent subjects were included, of which PubMed Central contributed 23,849 cases, generating 116,296 pairs of 2D MRI image-text pairs, and ImageCLEF contributed 484 cases, generating 4,631 pairs of 2D MRI image-text pairs.

[0134] Because the number of slices in clinical 3D MRI varies significantly across devices / sequences (from 20 to hundreds of slices), SimpleITK cubic spline interpolation is used to uniformly resample each sequence to 32 slices, achieving smooth reconstruction of brain structure volume. Unlike traditional data processing pipelines (including isotropic resampling and skull dissection), the visual-language multimodal model only performs uniform slice operations, preserving the integrity of the original information to the greatest extent.

[0135] To align the classification prediction process, a specific combination of input sequence orders is defined: the main input is the coplanar T1 and T1c pair (for evaluation reinforcement); auxiliary inputs include T2 and FLAIR; multi-plane inputs incorporate T1c from different planes to provide spatial context. If a sequence is missing, it is replaced with an currently available sequence. The visual-language multimodal model uses five fixed sequences: T1 (axial / other), coplanar T1c, T2, FLAIR, and non-coplanar T1c, ensuring a clear structure and close resemblance to clinical reasoning.

[0136] Modality markers, such as [T1] and [T1c-Axial], are introduced to clarify sequence origins; task-specific cues (such as "describe the enhanced pattern" or "generate a diagnostic report") are designed to guide the language model to focus on the target. This can significantly reduce modality / task ambiguity, improve training stability, preserve original diagnostic details, and ensure that the model's inference logic closely matches actual inference logic.

[0137] In one embodiment, multiple sequences of magnetic resonance images C m The expression is:

[0138] ;

[0139] in, For magnetic resonance images of the scan plane v0 in a T1-weighted sequence, The magnetic resonance image of scan plane v0 in the T1c weighted sequence. This is the magnetic resonance image of the scan plane v3 in the s3 sequence. This is the magnetic resonance image of the scan plane v4 in the s4 sequence. The image represents the magnetic resonance imaging (MRI) image of scan plane v5 within the s5 sequence; scan plane v0 can be any scan plane in the T1-weighted sequence; when the sequence includes both T2-weighted and FLAIR sequences, s3 is a T2-weighted sequence, s4 is a FLAIR sequence, scan plane v3 is any scan plane in the T2-weighted sequence, and scan plane v4 is any scan plane in the FLAIR sequence; when the sequence includes a target sequence, s3 is the target sequence, s4 is any sequence, scan plane v3 is any scan plane in the target sequence, scan plane v4 is any scan plane in any sequence, and the target sequence is either a T2-weighted sequence or a FLAIR sequence; when neither T2-weighted nor FLAIR sequences are present in the sequence, s3 and s4 are both any sequences. Scan planes include transverse, sagittal, and coronal planes.

[0140] Furthermore, when the T1c weighted sequence includes magnetic resonance images from another scanning plane, then the s5 sequence is a T1c weighted sequence and the scanning plane v5 is another scanning plane of the T1c sequence, that is, the scanning plane v0 is inconsistent with the scanning plane v5; when the T1c weighted sequence only includes magnetic resonance images from the scanning plane v0, the s5 sequence is another available sequence, which is a sequence among multiple sequences that is inconsistent with the T1 weighted sequence, the T1c weighted sequence, the s3 sequence, and the s4 sequence.

[0141] In this embodiment, through For magnetic resonance images of the scan plane v0 in a T1-weighted sequence, The magnetic resonance image of scan plane v0 in the T1c weighted sequence. This is the magnetic resonance image of the scan plane v3 in the s3 sequence. This is the magnetic resonance image of the scan plane v4 in the s4 sequence. The s3 sequence is the magnetic resonance image of scan plane v5 in the s5 sequence; scan plane v0 is any scan plane in the T1-weighted sequence; when the sequence includes both T2-weighted and FLAIR sequences, the s3 sequence is a T2-weighted sequence, the s4 sequence is a FLAIR sequence, scan plane v3 is any scan plane in the T2-weighted sequence, and scan plane v4 is any scan plane in the FLAIR sequence; when the sequence includes the target sequence, the s3 sequence is the target sequence, the s4 sequence is any sequence, scan plane v3 is any scan plane in the target sequence, scan plane v4 is any scan plane in any sequence, and the target sequence is either a T2-weighted sequence or a FLAIR sequence; when neither a T2-weighted nor FLAIR sequence exists, both the s3 and s4 sequences are any sequences. This approach fully utilizes multiple sequences for classification and prediction while ensuring that computational costs are not excessive. Furthermore, it can dynamically adapt to the actual sequence and plane availability of the target object, ensuring robustness even in the event of missing sequences.

[0142] In one embodiment, the plurality of sequences constitute a sequence combination, and the fusion feature includes fusion features corresponding to the plurality of sequence combinations. Step S5 includes:

[0143] The dual-layer MLP projector is used to project the fused features corresponding to multiple sequence combinations to obtain the first projection feature of each sequence combination.

[0144] Each of the first projection features is input into the language model to obtain the initial brain tumor classification label for each sequence combination and the confidence level of each initial brain tumor classification label;

[0145] The most frequent label among the initial brain tumor classification labels is used as the brain tumor classification label of the target object, and the confidence level of the brain tumor classification label is calculated.

[0146] Brain tumor classification labels confidence level The calculation formula is:

[0147] ;

[0148] Where N1 is the number of initial brain tumor classification labels that are consistent with the brain tumor classification labels, and M is the total number of initial brain tumor classification labels. For the m-th initial brain tumor classification label y m Confidence level, Let Kronecker function be used.

[0149] The set of multiple sequence combinations is represented as C m Let m be the m-th sequence combination in set C, and n be the number of sequences in the sequence combination. It represents all sequence combinations of the target object, where C represents all sequence combinations that satisfy the target condition, and the target condition refers to the sequence combination in the m-th sequence combination. For magnetic resonance images of the scan plane v0 in a T1-weighted sequence, The magnetic resonance image of scan plane v0 in the T1c weighted sequence. This is the magnetic resonance image of the scan plane v3 in the s3 sequence. This is the magnetic resonance image of the scan plane v4 in the s4 sequence. The s3 sequence is the magnetic resonance image of scan plane v5 in the s5 sequence; scan plane v0 is any scan plane in the T1 weighted sequence; when the sequence includes T2 weighted sequence and FLAIR sequence, the s3 sequence is a T2 weighted sequence, the s4 sequence is a FLAIR sequence, scan plane v3 is any scan plane in the T2 weighted sequence, and scan plane v4 is any scan plane in the FLAIR sequence; when the sequence includes target sequence, the s3 sequence is the target sequence, the s4 sequence is any sequence, scan plane v3 is any scan plane in the target sequence, scan plane v4 is any scan plane in any sequence, and the target sequence is either a T2 weighted sequence or a FLAIR sequence; when neither T2 weighted sequence nor FLAIR sequence exists in the sequence, both the s3 and s4 sequences are any sequences.

[0150] Imaging examinations typically involve multi-planar image acquisition (e.g., T1-weighted and T2-weighted sequences simultaneously encompass the axial, sagittal, and coronal planes), resulting in imaging sequences far exceeding the standard five. To address this heterogeneity, multiple sequence combinations are used to predict brain tumor classification labels.

[0151] Furthermore, the metadata of each first projection feature, text instruction, and target object is input into the language model to obtain the initial brain tumor classification label for each sequence combination, the confidence level of each initial brain tumor classification label, and the expression of the current detection report. , This is a reliable visual-language multimodal model trained with confidence, where d represents the metadata of the target object, t represents the text instruction, and C represents the... m Let m and r be sequence combinations. m This is the current detection report obtained based on sequence combination m.

[0152] For the Kronecker function , if and only if y m and When consistent, Only equals 1.

[0153] In this embodiment, a dual-layer MLP projector is used to project the fusion features corresponding to multiple sequence combinations to obtain the first projected feature of each sequence combination. Each first projected feature is input into the language model to obtain the initial brain tumor classification label and the confidence level of each initial brain tumor classification label for each sequence combination. The label with the highest frequency among the initial brain tumor classification labels is taken as the brain tumor classification label of the target object, and the confidence level of the brain tumor classification label is calculated. In this way, the reliability of the brain tumor classification label can be determined based on the confidence level of the brain tumor classification label.

[0154] In one embodiment, the intelligent classification method for brain tumors based on multimodal magnetic resonance imaging further includes:

[0155] Output the current detection report for the target object, which includes the brain tumor classification label and the confidence level of the brain tumor classification label.

[0156] In one embodiment, the intelligent brain tumor classification method based on multimodal magnetic resonance further includes: outputting a current detection report of the target object, the current detection report including a brain tumor classification label, a confidence level of the brain tumor classification label, an initial brain tumor classification label consistent with the brain tumor classification label and its confidence level.

[0157] In one embodiment, the intelligent brain tumor classification method based on multimodal magnetic resonance further includes: outputting a current detection report of the target object, wherein the current detection report includes an initial brain tumor classification label, the confidence level of each initial brain tumor classification label, the brain tumor classification label, and the confidence level of the brain tumor classification label.

[0158] In this embodiment, by outputting the current detection report of the target object, which includes the brain tumor classification label and the confidence level of the brain tumor classification label, the brain tumor classification label and the confidence level of the brain tumor classification label of the target object can be observed intuitively.

[0159] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0160] Based on the same inventive concept, this application also provides a multimodal magnetic resonance imaging (MRI)-based intelligent brain tumor classification system for implementing the aforementioned multimodal magnetic resonance imaging-based intelligent brain tumor classification method. The solution provided by this system is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the multimodal magnetic resonance imaging-based intelligent brain tumor classification system provided below can be found in the limitations of the multimodal magnetic resonance imaging-based intelligent brain tumor classification method described above, and will not be repeated here.

[0161] In one embodiment, such as Figure 6 As shown, a brain tumor intelligent classification system based on multimodal magnetic resonance imaging is provided, including:

[0162] The sampling subsystem is used to sample the magnetic resonance images of the target object to obtain upsampled and downsampled images; the magnetic resonance images include multiple sequences of magnetic resonance images;

[0163] The encoding subsystem is used to input the downsampled image into the visual encoder to obtain a global visual token, divide the upsampled image into multiple sub-images, and input the sub-images into the visual encoder to obtain a visual token;

[0164] The splicing subsystem is used to remove non-spatial patch tokens from the visual tokens, obtain the anatomical detail features of each sub-image, and splice the anatomical detail features of each sub-image to obtain the spliced ​​features.

[0165] The fusion subsystem is used to fuse stitching features and global visual tokens to obtain a fusion token, and to stitch together the fusion tokens of magnetic resonance images from each sequence to obtain fusion features.

[0166] The classification subsystem is used to project the fused features using a two-layer MLP projector to obtain the first projected feature. The first projected feature is then input into the language model to obtain the brain tumor classification label of the target object.

[0167] Each subsystem in the aforementioned intelligent brain tumor classification system based on multimodal magnetic resonance imaging can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.

[0168] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores magnetic resonance images, upsampled images, downsampled images, global visual tokens, sub-images, visual tokens, anatomical detail features, stitching features, fusion tokens, fusion features, first projection features, and brain tumor classification labels. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a multimodal magnetic resonance imaging-based intelligent brain tumor classification method.

[0169] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0170] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0171] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0172] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0173] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0174] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0175] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for intelligent classification of brain tumors based on multimodal magnetic resonance imaging, characterized in that, The method includes: S1. The magnetic resonance images of the target brain are sampled to obtain upsampled and downsampled images; the magnetic resonance images include multiple sequences of magnetic resonance images; S2. Input the downsampled image into the visual encoder to obtain a global visual token; divide the upsampled image into multiple sub-images and input the sub-images into the visual encoder to obtain a visual token; S3. Retain each spatial patch token of the last layer in the visual tokens output by the visual encoder to obtain the anatomical detail features of each sub-image, and splice the anatomical detail features of each sub-image to obtain spliced ​​features. The spatial patch token is a vector representation formed by embedding the vectors of the regions into which the sub-image is divided. S4. The stitching features are injected into the global visual token through the LR-HR cross-attention module of the visual language multimodal model to obtain a fusion token that incorporates high-resolution information. The fusion tokens of the magnetic resonance images of each sequence are stitched together to obtain the fusion features. S5. Project the fused features using a dual-layer MLP projector to obtain a first projected feature. Input the first projected feature into a language model to obtain the brain tumor classification label of the target object.

2. The method according to claim 1, characterized in that, Step S5 includes: The fused features are projected using a dual-layer MLP projector to obtain the first projected feature, and a text instruction for classification prediction is preset. The first projection feature and the text instruction are input into the language model to obtain the brain tumor classification label of the target object.

3. The method according to claim 1, characterized in that, The language model, anatomical detail-aware encoder, and dual-layer MLP projector are components of a brain-oriented visual-language multimodal model. The method also includes a training process for the visual-language multimodal model, which comprises: Obtain a training set; the training set includes two-dimensional magnetic resonance images with text descriptions and multi-sequence magnetic resonance images with radiation detection reports; Based on the two-dimensional magnetic resonance images and the multi-sequence magnetic resonance images, the LR-HR cross-resolution attention module of the dual-layer MLP projector and the anatomical detail-aware encoder is trained by gradient descent, resulting in the trained dual-layer MLP projector and the LR-HR cross-resolution attention module.

4. The method according to claim 3, characterized in that, The training set also includes labeled 3D magnetic resonance images. The language model, the visual encoder, and the dual-layer MLP projector are components of a brain-oriented visual-language multimodal model. The method further includes: Extract magnetic resonance images from the training set; When the magnetic resonance images extracted from the training set are two-dimensional magnetic resonance images, the two-dimensional magnetic resonance images are input into the visual encoder to extract the first visual embedding features from the two-dimensional magnetic resonance images. The first visual embedding features are then projected through the trained two-layer MLP projector to obtain the second projection features. Based on the second projection features, the language model is used to generate a predicted text description of the two-dimensional magnetic resonance images. Based on the predicted text description of the two-dimensional magnetic resonance images and the text description, the first model loss is calculated. When the magnetic resonance images extracted from the training set are three-dimensional magnetic resonance images or multi-sequence magnetic resonance images, the extracted magnetic resonance images are input into the visual encoder to extract the second visual embedding features from the extracted magnetic resonance images. The second visual embedding features are then projected onto the trained two-layer MLP projector to obtain the third projection features. Based on the third projection features, the language model is used to generate the prediction results of the extracted magnetic resonance images. Based on the prediction results and the true results of the extracted magnetic resonance images, the second model loss is calculated. Based on the first model loss or the second model loss, the trained dual-layer MLP projector, LR-HR cross-resolution attention module and language model in the visual language multimodal model are trained to obtain the trained dual-layer MLP projector, LR-HR cross-resolution attention module and language model. The training set is used to perform secondary training on the trained dual-layer MLP projector, LR-HR cross-resolution attention module, and language model to obtain a trained visual-language multimodal model.

5. The method according to claim 4, characterized in that, The method further includes: Obtain the target model parameters of the visual language multimodal model in multiple training rounds, and construct intermediate visual language multimodal models based on each target model parameter, with the network architecture of each intermediate visual language multimodal model being consistent. For the three-dimensional magnetic resonance images in the training set, each of the intermediate visual language multimodal models is used to perform classification prediction to obtain multiple predicted categories, and the category with the highest frequency among the multiple predicted categories is taken as the final predicted category of the three-dimensional magnetic resonance image. From the predicted categories, determine the number of categories that are inconsistent with the actual categories of the three-dimensional magnetic resonance image, and calculate the confidence level of the three-dimensional magnetic resonance image based on the number of categories and the total number of predicted categories; Based on the three-dimensional magnetic resonance image, the confidence level of the three-dimensional magnetic resonance image, and the final predicted category of the three-dimensional magnetic resonance image, a triplet is constructed; A reliability dataset is constructed based on the triples of the three-dimensional magnetic resonance images from multiple frames. The visual language multimodal model is optimized for parameters based on the reliability dataset, and its output format is reconstructed so that the visual language multimodal model outputs the confidence level of the brain tumor classification label.

6. The method according to claim 1, characterized in that, The multiple sequences of magnetic resonance images C m The expression is: ; in, For magnetic resonance images of the scan plane v0 in a T1-weighted sequence, The magnetic resonance image of scan plane v0 in the T1c weighted sequence. This is the magnetic resonance image of the scan plane v3 in the s3 sequence. This is the magnetic resonance image of the scan plane v4 in the s4 sequence. The s3 sequence is the magnetic resonance image of scan plane v5 in the s5 sequence; scan plane v0 is any scan plane in the T1 weighted sequence; when the sequence includes a T2 weighted sequence and a FLAIR sequence, the s3 sequence is a T2 weighted sequence, the s4 sequence is a FLAIR sequence, scan plane v3 is any scan plane in the T2 weighted sequence, and scan plane v4 is any scan plane in the FLAIR sequence; when the sequence includes a target sequence, the s3 sequence is the target sequence, the s4 sequence is any sequence, scan plane v3 is any scan plane in the target sequence, scan plane v4 is any scan plane in any sequence, and the target sequence is either a T2 weighted sequence or a FLAIR sequence; when neither a T2 weighted sequence nor a FLAIR sequence exists in the sequence, both the s3 and s4 sequences are any sequences.

7. The method according to claim 6, characterized in that, Multiple sequences constitute a sequence combination, and the fusion feature includes fusion features corresponding to multiple sequence combinations. Step S5 includes: The dual-layer MLP projector is used to project the fused features corresponding to multiple sequence combinations to obtain the first projection feature of each sequence combination. Each of the first projection features is input into the language model to obtain the initial brain tumor classification label for each sequence combination and the confidence level of each initial brain tumor classification label; The most frequent label among the initial brain tumor classification labels is used as the brain tumor classification label of the target object, and the confidence level of the brain tumor classification label is calculated. brain tumor classification labels confidence level The calculation formula is: ; Where N1 is the number of initial brain tumor classification labels that are consistent with the brain tumor classification labels, and M is the total number of initial brain tumor classification labels. For the m-th initial brain tumor classification label y m Confidence level, Let Kronecker function be used.

8. The method according to claim 7, characterized in that, The method further includes: Output the current detection report of the target object, the current detection report including the brain tumor classification label and the confidence level of the brain tumor classification label.

9. A brain tumor intelligent classification system based on multimodal magnetic resonance imaging, characterized in that, The system includes: A sampling subsystem is used to sample the magnetic resonance images of the target object to obtain upsampled and downsampled images; the magnetic resonance images include multiple sequences of magnetic resonance images; The encoding subsystem is used to input the downsampled image into the visual encoder to obtain a global visual token, divide the upsampled image into multiple sub-images, and input the sub-images into the visual encoder to obtain a visual token; The splicing subsystem is used to retain each spatial patch token of the last layer in the visual token output by the visual encoder, obtain the anatomical detail features of each sub-image, and splice the anatomical detail features of each sub-image to obtain splicing features. The spatial patch token is a vector representation formed by embedding the vectors of the regions into which the sub-image is divided. The fusion subsystem is used to inject stitching features into the global visual token through the LR-HR cross-attention module of the visual language multimodal model to obtain a fusion token that incorporates high-resolution information, and stitch the fusion tokens of the magnetic resonance images of each sequence to obtain fusion features; A classification subsystem is used to project the fused features using a two-layer MLP projector to obtain a first projected feature, and input the first projected feature into a language model to obtain a brain tumor classification label for the target object.