Disease intelligent prediction and auxiliary diagnosis and treatment method based on multiple modes and LLM

By constructing a multimodal medical data set and combining CNN and Transformer for image feature extraction and multimodal fusion, fine-tuning the LLM generation diagnosis and treatment plan, the problem of insufficient utilization of multimodal data in the existing technology is solved, and intelligent and personalized disease prediction and auxiliary diagnosis and treatment are realized.

CN120388714APending Publication Date: 2025-07-29TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510464572.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to make full and rational use of multimodal medical data, especially image and text data, to achieve intelligent and accurate disease prediction and auxiliary diagnosis and treatment for specific disease fields. Moreover, LLM lacks customized models in the medical field, making it difficult to provide personalized support.

Method used

By constructing a multimodal medical data set, using CNN for image segmentation and feature extraction, and combining Transformer for multimodal fusion, generating disease prediction results; using public medical Q&A data and literature fine-tuning LLM, a diagnostic word vector knowledge base is constructed, and a personalized diagnosis and treatment plan is generated.

Benefits of technology

It realizes intelligent, accurate prediction and personalized diagnosis and treatment plan generation for specific diseases, improves the adaptability and expansion of auxiliary diagnosis and treatment, and is suitable for health assessment and decision-making support for patients and non-patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388714A_ABST
    Figure CN120388714A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent disease prediction and auxiliary diagnosis and treatment method based on multi-modality and LLM, and the method comprises the steps: carrying out the feature extraction of a medical image and illness condition data of a patient, and carrying out the multi-modality fusion training; secondly, a fine-tuning question and answer data set is constructed by obtaining and screening a public medical question and answer data set and medical literatures, and LLM is fine-tuned to concentrate on question and answer diagnosis tasks for specific research parts; and constructing a diagnosis and treatment large model by using question and answer data of a doctor patient, taking a consultation question and a disease prediction result of the patient as input, performing similarity matching with the text segments of the diagnosis and treatment word vector knowledge base through an Embedding model, inputting the text segment with the highest similarity into the diagnosis and treatment large model, and generating a diagnosis and treatment scheme in combination with a designed cue word. According to the method, the multi-modal medical data and the public medical text data of the patient are utilized, CNN and Transform architectures are combined with an LLM architecture, and intelligent and accurate prediction for specific diseases and personalized generation of diagnosis and treatment schemes are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medicine, and specifically to a disease intelligent prediction and auxiliary diagnosis and treatment method based on multi-modal and LLM. Background Art

[0002] With the rapid development of information technology, the application of AI (Artificial Intelligence) in the medical industry has become increasingly widespread, especially in the fields of disease prediction and auxiliary diagnosis and treatment. Most of the early AI-assisted diagnosis research focused on single-modal research, such as medical images. However, modern medicine generates a large amount of complex multi-modal medical data, including clinical texts, audio, and laboratory test results, in addition to medical images. By cleverly using the collected multi-modal medical data, a comprehensive understanding of the condition can be obtained, thus providing more accurate decision-making support for doctors. However, how to fully and reasonably utilize the existing multi-modal medical data to achieve downstream diagnosis and treatment tasks for specific disease fields or medical research contents is a difficult point in the current medical multi-modal AI research.

[0003] On the other hand, in recent years, technologies based on LLM (Large Language Model) have been widely applied in the medical field, especially in medical Q&A, treatment plan generation, etc. However, limited by the diversity of pre-training data and the need for rapid promotion of pre-trained models, most of the existing LLMs are general-purpose Q&A dialogue models, such as MedPaLM, etc., lacking customized models for specific disease fields or medical research goals, and it is difficult to provide customized accurate prediction and diagnosis and treatment support. In addition, although LLMs are good at text processing and generation tasks, their performance in image feature analysis tasks is lacking. If an LLM is used to complete the single task of accurate multi-modal disease prediction, the prediction performance is often weaker than that of CNN (Convolutional Neural Network) and Transformer architectures.

[0004] Therefore, how to fully and reasonably utilize the existing multi-modal medical data, how to cleverly combine the CNN and Transformer architectures suitable for accurate disease prediction with the LLM architecture suitable for treatment plan generation to achieve more intelligent and personalized disease prediction and auxiliary diagnosis and treatment for specific disease fields is an urgent problem to be solved currently. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the technical problem to be solved by the present invention is to provide a disease intelligent prediction and auxiliary diagnosis and treatment method based on multi-modal and LLM.

[0006] The technical solution of the present invention to solve the above technical problems is to provide a disease intelligent prediction and auxiliary diagnosis and treatment method based on multimodality and LLM, characterized in that the method includes the following steps:

[0007] Step 1: Construct a multimodal medical dataset for patients: First, collect the patient's image data and corresponding text data; the text data includes medical condition data and doctor-patient Q&A data; then preprocess the patient's image data and text data; finally, use the preprocessed image data and corresponding preprocessed text data to construct a multimodal medical dataset;

[0008] Step 2: Use the trained and evaluated CNN to perform image segmentation on the image data obtained in step 1 to obtain the target structure image; then perform image feature extraction on the target structure image to obtain the embedding vector of the image modality;

[0009] Perform feature extraction on the condition data obtained in step 1 to obtain an embedding vector for the condition data;

[0010] Step 3: Align the embedding vectors of the imaging modality with the embedding vectors of the corresponding disease data. Then, a multimodal fusion algorithm based on the Transformer architecture is used to fuse them, mining the internal correlation between the imaging modality and disease data. The disease prediction results are then obtained through the classification layer.

[0011] Step 4: Based on the public medical Q&A dataset and medical literature, a fine-tuned Q&A dataset for the research target is constructed. The LLM is then fine-tuned based on the fine-tuned Q&A dataset to obtain a fine-tuned LLM. A diagnosis and treatment word vector knowledge base for the LLM is then constructed based on the pre-processed doctor-patient Q&A data obtained in Step 1. Together with the fine-tuned LLM, this constitutes a large diagnosis and treatment model for the research target.

[0012] Step 5: Take the new user's consultation question and the disease prediction result obtained in step 3 as input, and use the Embedding model to perform similarity matching with the text segments in the diagnosis and treatment word vector knowledge base. Input the text segment with the highest similarity into the diagnosis and treatment model, and combine it with the designed prompt words to generate a diagnosis and treatment plan with accurate disease prediction.

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] (1) The present invention fully and rationally utilizes patients' multimodal medical data and public medical text data, and cleverly combines the CNN and Transformer architectures with the LLM architecture to achieve intelligent and accurate prediction and personalized generation of diagnosis and treatment plans for specific diseases, thereby improving the adaptability, efficiency and scalability of intelligent assisted diagnosis and treatment.

[0015] (2) The present invention makes full and reasonable use of existing multi-modal medical data, including: achieving accurate disease prediction based on patient imaging data and medical condition data, fine-tuning the LLM based on public medical Q&A datasets and medical literature, and combining the construction of a diagnosis and treatment word vector knowledge base based on doctor-patient Q&A data to achieve the accurate and personalized generation of diagnosis and treatment plans.

[0016] (3) According to the advantages of different model architectures, the present invention is skillfully combined. The CNN architecture is used for feature extraction, and the Transformer architecture is used for multi-modal fusion to achieve accurate disease prediction; further, the LLM architecture is used to generate diagnosis and treatment plans, and the whole process is intelligent and efficient.

[0017] (4) The beneficiaries of the present invention include not only patients but also non-patients. Patients can input their own examination data and consultation questions, and the present invention makes disease predictions for them and generates corresponding diagnosis and treatment plans; for non-patients, they can input consultation questions, and the present invention provides them with preliminary health assessments and auxiliary decision-making, showing good adaptability.

[0018] (5) By constructing multi-modal medical datasets in different disease fields, the present invention can be applied to the prediction and auxiliary diagnosis and treatment of various specific disease fields, showing good scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is the overall flowchart of the present invention;

[0020] Figure 2 is the overall logical structure diagram of the present invention;

[0021] Figure 3 is the flowchart of step 3 of the present invention;

[0022] Figure 4 is the flowchart of steps 4 to 5 of the present invention;

[0023] Figure 5 is the knee joint MRI image of Example 1 of the present invention;

[0024] Figure 6 is from Example 1 of the present invention Figure 5 is the result diagram of segmenting the anterior cruciate ligament;

[0025] Figure 7 is the receiver operating characteristic curve and the area under the receiver operating characteristic curve diagram of Example 1 of the present invention;

[0026] Figure 8 is the interface diagram of the diagnosis and treatment plan for anterior cruciate ligament disease in Example 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Specific embodiments of the present invention are given below. The specific embodiments are only used to further illustrate the present invention in detail and do not limit the protection scope of the present invention.

[0028] The present invention provides a disease intelligent prediction and auxiliary diagnosis and treatment method based on multi-modal and LLM (hereinafter referred to as the method), which is characterized in that the method includes the following steps:

[0029] Step 1: Construct a multi-modal medical data set for the patient: First, collect the patient's imaging data and corresponding text data; then preprocess the patient's imaging data and text data; finally, use the preprocessed imaging data and the corresponding preprocessed text data to construct a multi-modal medical data set; the text data includes the patient's condition data and the Q&A data between the doctor and the patient;

[0030] Preferably, in Step 1, the condition data includes the patient's basic information, clinical symptoms, and imaging reports; the Q&A data between the doctor and the patient is the communication dialogue between the doctor and the patient recorded in text form.

[0031] Preferably, in Step 1, the preprocessing of the imaging data is specifically as follows: successively perform case anonymization, data quality control, noise reduction, and standardization, then use equal-proportion scaling, and then use the method of filling with 0 pixels to unify the size of the imaging data (in this embodiment, the imaging data is unified to 21×480×480), and the resize function is not used during the unification to ensure that the texture of the image does not change during the preprocessing process.

[0032] Preferably, in Step 1, the preprocessing of the text data is specifically as follows: use text detection technology (OCR technology is used in this embodiment) to automatically recognize the text document, then save the automatically recognized text to a txt file, and then perform data cleaning and quality control.

[0033] Preferably, in Step 1, the data sets of the preprocessed imaging data and condition data are divided into their respective training sets, validation sets, and test sets, with a ratio of training set:validation set:test set = 7:2:1. The preprocessed Q&A data between the doctor and the patient is not divided into data sets.

[0034] Preferably, in Step 1, the multi-modal medical data set further includes audio data during the communication between the doctor and the patient; it is converted into Q&A data between the doctor and the patient through speech recognition technology (Kaldi toolkit is used in this embodiment), and then preprocessed, and no data set division is required.

[0035] Preferably, in step 1, the image data is MRI (Magnetic Resonance Imaging) image data and / or CT (Computed Tomography) scan images; wherein, when the image data adopts MRI image data and CT scan images, the CT scan images need to be registered and fused with the corresponding MRI image data through a medical image processing toolkit (Insight Segmentation and Registration Toolkit, ITK).

[0036] Step 2: Since the image data contains multiple tissue structures, the image data obtained in step 1 is segmented using the trained and evaluated CNN to obtain the target structure image; then, image feature extraction is performed on the target structure image to obtain the embedding vector of the image modality;

[0037] Feature extraction is performed on the condition data obtained in step 1 to obtain the embedding vector of the condition data;

[0038] Preferably, in step 2, obtaining the embedding vector of the image modality specifically is:

[0039] S21: Professional doctors make the corresponding image segmentation Mask according to the training set and validation set of the image data obtained in step 1, and then input the training set and validation set of the image data obtained in step 1 and the corresponding image segmentation Mask into the CNN (nnUnet neural network in this embodiment) for training; by continuously optimizing the loss function of the CNN to minimize its value, an image segmentation training model is obtained; then, the test set of the image data is input into the image segmentation training model to obtain the target structure image;

[0040] Preferably, in step S2.1, the loss function for CNN training uses a pixel-level cross-entropy loss function, as shown in Equation (1):

[0041]

[0042] In Equation (1), L is the total loss function; N is the total number of pixels in any slice of the image data; C is the number of categories (C = 2 in this embodiment); y ic is the probability that the i-th pixel truly belongs to category c, which is a binary value (0 or 1); P ic is the probability that the model predicts the i-th pixel belongs to category c.

[0043] The performance evaluation index for CNN training uses the Dice coefficient for multi-class segmentation, as shown in Equation (2):

[0044]

[0045] In formula (2), Dice a is the Dice coefficient calculated for class c; N is the total number of pixels in any slice of the image data; C is the number of classes (C = 2 in this embodiment); p ic is the probability that the i-th pixel predicted by the model belongs to class c; q ic is the probability that the i-th pixel truly belongs to class c, which is a binary value (0 or 1).

[0046] S22. Extract image features from the target structure image to obtain the embedding vector of the image modality.

[0047] Preferably, step S22 is specifically: unify the size of the target structure image to the size adapted to the feature extraction algorithm (in this embodiment, since the Resnet50 neural network is used, the size is unified to 21×224×224), then use the combined extraction method of multiple feature extraction algorithms to obtain various types of feature extraction vectors of the image modality, and then splice the various types of feature extraction vectors using concatenate to obtain the embedding vector of each image modality;

[0048] Preferably, in step S22, the multiple feature extraction algorithms include Fourier-based shape features, gray gradient-based texture features, and Resnet50 neural network-based depth features.

[0049] Preferably, in step 2, use a classifier (LightGBM classifier in this embodiment) to evaluate the embedding vector of the image modality obtained in step S22 (the evaluation index in this embodiment selects the area under the receiver operating characteristic curve (AUC)) to illustrate that the obtained embedding vector of the image modality performs well in the disease precise prediction task.

[0050] Preferably, in step 2, the embedding vector of the condition data is specifically obtained as follows: use the preprocessed condition data obtained in step 1 as the input, perform word segmentation using the word segmentation technology (BertTokenizer function in this embodiment), and then ensure that the lengths of all sentences are unified by specifying the padding filling parameter to convert the condition data into a format that the text feature extraction model can understand; then input the converted result into the text feature extraction model (BioBERT model in this embodiment) for feature extraction to obtain the embedding vector of each condition data.

[0051] Step 3: Align the embedding vectors of the imaging modality with the embedding vectors of the corresponding disease data. Then, a multimodal fusion algorithm based on the Transformer architecture is used to fuse them, mining the internal correlation between the imaging modality and the disease data. The disease prediction results are then obtained through the classification layer to achieve accurate disease prediction.

[0052] Preferably, in step 3, the alignment steps are as follows:

[0053] S31. Sequentially serialize the embedding vectors of the imaging modality and the condition data obtained in step 2, add classification label embedding and position embedding, and obtain the embedding matrices of the imaging modality and the condition data to meet the input sequence requirements of the subsequent multimodal fusion algorithm. (In this embodiment, the embedding vectors of the condition data are obtained using the BioBERT model. The BioBERT model and the multimodal fusion algorithm have the same Transformer architecture, so this step is not required.)

[0054] S32. Use a contrastive learning method on the embedding matrix of the imaging modality and the embedding matrix of the disease data, with the goal of minimizing the optimization loss function value, to achieve alignment and map them to the same embedding space.

[0055] Preferably, in step S32, the loss function used is the InfoNCE loss function, as shown in formula (3):

[0056]

[0057] In formula (3), μ is the embedding matrix of the disease data; υ is the embedding matrix of the imaging modality; sim(μ,v) is the similarity function between the two modalities; τ is the temperature parameter; and U is the set of embedding matrices of all imaging modalities.

[0058] Preferably, in step 3, in order to effectively fuse the information of the imaging modality and the disease data, a multimodal fusion algorithm based on the Transformer architecture is proposed. The algorithm includes two cross-attention networks, each of which consists of a 12-layer Transformer architecture with cross-attention blocks;

[0059] The specific implementation process is: embedding matrix of aligned image modality and the embedding matrix of the aligned condition data The self-attention mechanism is used to learn features, learn global dependencies within its own modality, and capture long-distance dependencies. The self-attention mechanism of Transformer is expressed as:

[0060]

[0061] In Equation (4), Q, K, and V represent the query matrix, the key matrix, and the value matrix respectively, all of which are obtained by linear transformation from the input embedding matrix, as shown in Equation (5):

[0062]

[0063] In Equation (5), is a learnable parameter matrix;

[0064] The self-attention outputs of the image modality and the disease condition data are respectively:

[0065]

[0066] Then, the self-attention output of the image modality and the self-attention output of the disease condition data are fused through a cross-attention mechanism to establish a cross-modal correlation: the self-attention output of the image modality is used as the query for the disease condition, and the self-attention output of the disease condition data is used as the query for the image. The cross-attention output of the two modalities is obtained as shown in Equation (7), and then fused (using the concatenation method) to obtain the fusion matrix H of the two modalities f , as shown in Equation (8):

[0067]

[0068] H f = Concat(H i→t , H t→i ) (8)

[0069] In Equations (7)-(8), H i→t represents the cross-attention output of the image modality under the guidance of the disease condition data, and H t→i represents the cross-attention output of the disease condition data under the guidance of the image modality; Concat(·) represents concatenation.

[0070] Preferably, in step 3, the fusion matrix H of the two modalities f is input into a classification layer composed of 4 fully connected layers to output the accurate prediction result of the disease.

[0071] Step 4: Based on the public medical Q&A dataset and medical literature, construct a fine-tuning Q&A dataset for the research objective; then, based on the fine-tuning Q&A dataset, fine-tune the LLM to obtain the fine-tuned LLM; then, based on the preprocessed Q&A data of doctors and patients obtained in step 1, construct a diagnostic word vector knowledge base for the LLM, which together with the fine-tuned LLM constitutes a diagnostic large model for the research objective;

[0072] Preferably, in step 4, constructing a fine-tuning Q&A dataset regarding the research objective specifically involves: obtaining existing publicly available medical Q&A datasets (in this embodiment, the MedQuAD and LiveQA-Med datasets are obtained), setting text screening terms according to the research objective, and screening out medical Q&A statements related to the research objective; at the same time, obtaining medical literature with authoritative explanations in the field of medical Q&A (in this embodiment, the WHO Guidelines and NICE Guidelines medical guidelines are obtained), using text search technology to screen out paragraphs related to the research objective, and constructing them into medical Q&A statements; then standardizing and formatting the above two types of medical Q&A statements to form a fine-tuning Q&A dataset regarding the research objective.

[0073] Preferably, in step 4, the LoRA method is used to fine-tune the LLM (in this embodiment, the ChatGLM3-6B pre-trained model is selected). Specifically: LoRA fine-tuning is an efficient parameter fine-tuning method. First, an additional trainable module ΔW is introduced into the pre-trained attention weight matrix. This trainable module ΔW can be decomposed into the product of a low-rank matrix A and a low-rank matrix B, reducing the number of parameters that need to be learned; then the pre-trained weights W are frozen. During forward propagation, W and ΔW are combined into a new weight W′, and W′ is continuously updated during the training process to minimize the value of the loss function, and finally the fine-tuned LLM is obtained.

[0074] Preferably, in step 4, constructing the medical diagnosis and treatment word vector knowledge base of the LLM specifically involves: segmenting the text paragraphs and unifying the text structure of the Q&A data of doctors and patients after preprocessing in step 1 to obtain a medical diagnosis and treatment knowledge base; then using an Embedding model (in this embodiment, the BioBERT embedding model) to vectorize the medical diagnosis and treatment knowledge base to obtain a medical diagnosis and treatment word vector knowledge base.

[0075] Step 5: Use the consultation questions of the new user and the disease prediction results obtained in step 3 as inputs, perform similarity matching between the Embedding model and the paragraphs in the medical diagnosis and treatment word vector knowledge base, input the paragraph with the highest similarity into the medical diagnosis and treatment large model, and combine the designed prompt (Prompt) to generate a medical diagnosis and treatment plan with accurate disease prediction.

[0076] Preferably, in step 5, the user is a patient or a non-patient (i.e., a health consultation for a person without a disease).

[0077] Preferably, in step 5, use the disease prediction results obtained in step 3 and the consultation questions of the patient as inputs, convert them into input word vectors through the same Embedding model as in step 4; then perform similarity matching comparison with the paragraphs in the medical diagnosis and treatment word vector knowledge base obtained in step 4, and input the paragraph with the highest similarity into the fine-tuned LLM.

[0078] Preferably, in step 5, the similarity calculation method is as shown in Equation (9):

[0079]

[0080] In Equation (9), v(Q) is the input word vector; v(D i ) is the word vector of the i-th passage in the medical word vector knowledge base;

[0081] v(Q)·v(D i ) represents the dot product of the two vectors; ‖v(Q)‖·‖v(D i )‖ both represent the Euclidean norm of the vector, and D * represents the passage in the medical word vector knowledge base with the highest similarity to the input word vector.

[0082] Preferably, in step 5, the prompt words include: (1) Disease prediction result presentation: The purpose of this part is to directly output the multi-modal fusion result; (2) Dynamic prompt presentation: The purpose of this part is to improve the output quality of the diagnosis and treatment plan, and the prompt is: "Please generate a comprehensive treatment plan for this patient, including drug treatment, physical therapy, surgical suggestions, and rehabilitation plan."; (3) Diagnostic plan structure presentation: The purpose of this part is to standardize the output diagnosis and treatment plan, and the final output template is: "By analyzing the imaging data and text data, you may have XXX, and the corresponding diagnosis and treatment plan includes: 1. Drug treatment plan: XXX. 2. Rehabilitation suggestions: XXX. 3. Possible surgical plan: XXX." (XXX is the content related to the research goal).

[0083] Example 1:

[0084] This example aims at the anterior cruciate ligament tissue of the knee joint to realize intelligent prediction and assisted diagnosis and treatment of anterior cruciate ligament diseases.

[0085] Step 1: Preprocess the imaging data and corresponding text data of the patient's knee joint to construct a multi-modal dataset of the patient's anterior cruciate ligament. The knee joint imaging of the patient is as Figure 5 shown.

[0086] Step 2: As Figure 5 it can be seen that the imaging data of the knee joint contains multiple tissue structures, such as the anterior cruciate ligament, posterior cruciate ligament, femur, tibia, etc. Therefore, to obtain the target structure image (i.e., the anterior cruciate ligament image), it is first necessary to use the trained and evaluated CNN for image segmentation to obtain the image of the target structure, as Figure 6 shown. The orange area represents the anterior cruciate ligament image; then, the embedding vector of the anterior cruciate ligament image modality is obtained through multi-feature joint extraction and evaluated using the area under the receiver operating characteristic curve (AUC), as Figure 7As shown, by Figure 7 it can be seen that it performs well in the precise prediction of anterior cruciate ligament diseases.

[0087] Step 3: Pass the embedding vectors of the anterior cruciate ligament imaging modality and the corresponding disease condition data (including the patient's basic information, clinical symptoms, and imaging reports of the anterior cruciate ligament) through the alignment, fusion, and classification layers to finally obtain the prediction results of anterior cruciate ligament diseases.

[0088] Step 4: Based on the publicly available medical Q&A dataset and medical literature related to the anterior cruciate ligament field, construct a fine-tuned Q&A dataset; then, based on the fine-tuned Q&A dataset, fine-tune the LLM to obtain the fine-tuned LLM; then, based on the communication dialogues between doctors and patients regarding anterior cruciate ligament diseases, construct a knowledge base of diagnostic and treatment word vectors for the LLM, which together with the fine-tuned LLM forms a diagnostic and treatment large model for the research target.

[0089] Step 5: Use the patient's consultation questions regarding anterior cruciate ligament diseases and the prediction results of anterior cruciate ligament diseases obtained in Step 3 as inputs to finally generate a diagnostic and treatment plan for anterior cruciate ligament diseases, such as Figure 8 shown. By Figure 8 it can be seen that the disease prediction and auxiliary diagnostic and treatment plan generation based on multi-modal and LLM proposed by the present invention perform well in anterior cruciate ligament diseases of the knee joint, can accurately predict the types of anterior cruciate ligament diseases of patients, and can generate personalized auxiliary diagnostic and treatment plans. The parts not described in the present invention are applicable to the prior art.

Claims

1. A disease intelligent prediction and assisted diagnosis and treatment method based on multi-modal and LLM, characterized in that, The method comprises the following steps: Step 1: Construct a multimodal medical dataset for patients: First, collect the patient's image data and corresponding text data; the text data includes medical condition data and doctor-patient Q&A data; then preprocess the patient's image data and text data; finally, use the preprocessed image data and corresponding preprocessed text data to construct a multimodal medical dataset; Step 2: Use the trained and evaluated CNN to perform image segmentation on the image data obtained in step 1 to obtain the target structure image; then perform image feature extraction on the target structure image to obtain the embedding vector of the image modality; Perform feature extraction on the condition data obtained in step 1 to obtain an embedding vector for the condition data; Step 3: Align the embedding vectors of the imaging modality with the embedding vectors of the corresponding disease data. Then, a multimodal fusion algorithm based on the Transformer architecture is used to fuse them, mining the internal correlation between the imaging modality and disease data. The disease prediction results are then obtained through the classification layer. Step 4: Based on the public medical Q&A dataset and medical literature, a fine-tuned Q&A dataset for the research target is constructed. The LLM is then fine-tuned based on the fine-tuned Q&A dataset to obtain a fine-tuned LLM. A diagnosis and treatment word vector knowledge base for the LLM is then constructed based on the pre-processed doctor-patient Q&A data obtained in Step 1. Together with the fine-tuned LLM, this constitutes a large diagnosis and treatment model for the research target. Step 5: Take the new user's consultation question and the disease prediction result obtained in step 3 as input, and use the Embedding model to perform similarity matching with the text segments in the diagnosis and treatment word vector knowledge base. Input the text segment with the highest similarity into the diagnosis and treatment model, and combine it with the designed prompt words to generate a diagnosis and treatment plan with accurate disease prediction.

2. The method for intelligent disease prediction and assisted diagnosis and treatment based on multi-modal and LLM according to claim 1, wherein, In step 1, the medical data includes basic patient information, clinical symptoms, and imaging reports; the doctor-patient Q&A data is the communication dialogue between the doctor and the patient recorded in text form; In step 1, the multimodal medical dataset also includes audio data of doctor-patient communication; it is converted into doctor-patient question-and-answer data through speech recognition technology and then preprocessed; In step 1, the image data is MRI image data and / or CT scan image; wherein, when the image data uses MRI image data and CT scan image, the CT scan image needs to be registered and fused with the corresponding MRI image data through a medical image processing toolkit.

3. The disease intelligent prediction and auxiliary diagnosis and treatment method based on multi-modal and LLM according to claim 1, wherein, In step 1, the image data is preprocessed by anonymizing the cases, performing data quality control, noise reduction, and standardization. The image data is then scaled proportionally and then padded with zero pixels to unify the size. The resize function is not applied during unification to ensure that the image texture does not change during the preprocessing process. In step 1, the preprocessing of text data is specifically: using text detection technology to automatically identify text documents, then saving the automatically identified text into a txt file, and then performing data cleaning and quality control.

4. The method for intelligent disease prediction and assisted diagnosis and treatment based on multi-modal and LLM according to claim 1, wherein, In step 2, the embedding vector of the imaging modality is obtained as follows: S21. The professional doctor makes corresponding image segmentation masks based on the training set and validation set of the image data obtained in step 1, and then inputs the training set and validation set of the image data obtained in step 1 and the corresponding image segmentation masks into the CNN for training. By continuously optimizing the loss function of the CNN to minimize its value, an image segmentation training model is obtained. Then, the test set of the image data is input into the image segmentation training model to obtain the target structure image; S22. Extract image features from the target structure image to obtain the embedding vectors of each image modality; In step 2, the specific method for obtaining the embedding vectors of the condition data is as follows: using the preprocessed condition data obtained in step 1 as the input, performing word segmentation using the word segmentation technology, and then ensuring that the lengths of all sentences are unified by specifying the padding filling parameter, converting the condition data into a format that the text feature extraction model can understand. Then, the converted result is input into the text feature extraction model for feature extraction to obtain the embedding vectors of each condition data.

5. The method for intelligent disease prediction and auxiliary diagnosis and treatment based on multi-modal and LLM according to claim 4, wherein In step S2.1, the loss function for CNN training uses the pixel-level cross-entropy loss function, as shown in Equation (1): In Equation (1), L is the total loss function; N is the total number of pixels in any one slice of the image data; C is the number of classes; y ic is the probability that the i-th pixel truly belongs to class c, which is a binary value; P ic is the probability that the model predicts the i-th pixel belongs to class c; The performance evaluation metric for CNN training uses the Dice coefficient for multi-class segmentation, as shown in Equation (2): In formula (2), Dice a is the Dice coefficient calculated for class c; N is the total number of pixels in any slice of the image data; C is the number of classes; p ic is the probability that the i-th pixel predicted by the model belongs to class c; q ic is the probability that the i-th pixel truly belongs to class c, which is a binary value; Step S22 is specifically as follows: unify the size of the target structure image to the size adapted to the feature extraction algorithm, and then use the method of jointly extracting multiple types of feature extraction vectors of the image modality by using multiple feature extraction algorithms. Then, concatenate the multiple types of feature extraction vectors using concatenate to obtain the embedding vectors of the image modality; In step S22, the multiple feature extraction algorithms include shape features based on Fourier, texture features based on gray gradient, and depth features based on the Resnet50 neural network.

6. The method for intelligent disease prediction and assisted diagnosis and treatment based on multi-modal and LLM according to claim 1, wherein In step 3, the alignment steps are as follows: S31. Serialize the embedding vectors of the image modality and the embedding vectors of the condition data obtained in step 2 in sequence, add classification token embeddings and position embeddings to obtain the embedding matrix of the image modality and the embedding matrix of the condition data to meet the input sequence requirements of the subsequent multi-modal fusion algorithm; S32. Use the contrastive learning method for the embedding matrix of the image modality and the embedding matrix of the condition data, with the goal of minimizing the optimized loss function value to achieve alignment and map them to the same embedding space; In step 3, the multi-modal fusion algorithm based on the Transformer architecture includes two cross-attention networks, and each network consists of 12 layers of the Transformer architecture with cross-attention blocks; The specific implementation process is as follows: the embedding matrix of the aligned imaging modality and the embedding matrix of the aligned condition data respectively perform feature learning through the self-attention mechanism to learn the global dependencies within their own modalities and capture long-range dependencies; the self-attention mechanism of Transformer is expressed as: In Equation (4), Q, K, and V respectively represent the query matrix, key matrix, and value matrix, all of which are obtained by linear transformation from the input embedding matrix, as shown in Equation (5): In formula (5), is a learnable parameter matrix; The self-attention outputs of the image modality and the condition data are respectively: Then, the self-attention output of the imaging modality and the self-attention output of the condition data are fused through a cross-attention mechanism to establish a cross-modal correlation relationship: the self-attention output of the imaging modality is used as the query for the condition, and the self-attention output of the condition data is used as the query for the imaging. The cross-attention output of the two modalities is shown in Equation (7), and then fused to obtain the fusion matrix H of the two modalities f , as shown in Equation (8): H f = Concat(H i→t , H t→i ) (8) In formula (7) - formula (8), H i→t represents the cross-attention output of the imaging modality guided by the disease condition data, and H t→i represents the cross-attention output of the disease condition data guided by the imaging modality; Concat(·) represents concatenation; In step 3, the fusion matrix H of the two modalities f is input into the classification layer composed of 4 fully connected layers to output the accurate prediction result of the disease.

7. The method for intelligent disease prediction and assisted diagnosis and treatment based on multi-modal and LLM according to claim 6, characterized in that In step S32, the loss function used is the InfoNCE loss function, as shown in Equation (3): In formula (3), μ is the embedding matrix of the disease condition data; v is the embedding matrix of the imaging modality; sim(μ, v) is the similarity function of the two modalities; τ is the temperature parameter; U is the set of embedding matrices of all imaging modalities.

8. The method for intelligent disease prediction and assisted diagnosis and treatment based on multi-modal and LLM according to claim 1, wherein In step 4, the specific construction of the fine-tuning Q&A dataset regarding the research objective is as follows: Obtain the existing publicly available medical Q&A dataset, set text screening words according to the research objective, and screen out the medical Q&A statements related to the research objective; Obtain the medical literature with authoritative explanatory power in the field of medical Q&A, use text search technology to screen out the paragraphs related to the research objective, and construct them into medical Q&A statements; Then standardize and format the above two types of medical Q&A statements to form the fine-tuning Q&A dataset regarding the research objective; In step 4, the LoRA method is used to fine-tune the LLM; In step 4, the specific construction of the diagnosis and treatment word vector knowledge base of the LLM is as follows: Perform text paragraph segmentation and text structure unification on the preprocessed Q&A data of doctors and patients obtained in step 1 to obtain the diagnosis and treatment knowledge base; Then use the Embedding model to vectorize the diagnosis and treatment knowledge base to obtain the diagnosis and treatment word vector knowledge base.

9. The method for intelligent disease prediction and assisted diagnosis and treatment based on multi-modal and LLM according to claim 1, wherein In step 5, the similarity calculation method is shown in formula (9): In formula (9), v(Q) is the input word vector; v(D i ) is the word vector of the i-th text segment in the medical term vector knowledge base; v(Q)·v(D i ) represents the dot product of two vectors; ‖v(Q)‖·‖v(D i )‖ both represent the Euclidean norm of the vector, and D * represents the passage in the diagnosis and treatment word vector knowledge base that has the highest similarity to the input word vector.

10. The method for intelligent disease prediction and assisted diagnosis and treatment based on multi-modal and LLM according to claim 1, characterized in that, In step 5, the prompt words include: (1) Presentation of disease prediction results; (2) Presentation of dynamic prompts; (3) Presentation of the diagnosis plan structure.