Multi-modal medical large model system and medical image auxiliary diagnosis system

By constructing a multimodal medical big model system, the existing medical image analysis models have solved the shortcomings in the coverage, interpretability and interactivity of the disease, and the efficient processing of medical text, image and attention data and the accurate output of diagnostic results are achieved.

CN120216976APending Publication Date: 2025-06-27WUHAN UNITED IMAGING HEALTHCARE SURGICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311814686.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing medical image analysis models have problems such as insufficient coverage of the disease, poor interpretability and weak interaction capabilities, which are difficult to meet the multimodal needs of front-line medical workers.

Method used

A multimodal medical big model system is provided. Through training data acquisition, model construction, training processing and result optimization modules, it builds a multimodal model that can process medical text, image and attention data, and realizes text image alignment and joint training.

Benefits of technology

It improves the overall performance of the multimodal medical model, realizes the conversion and fusion of different modal data, can automatically output more accurate diagnostic results for different types of diseases, and supports multiple rounds of Q&A and user attention interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216976A_ABST
    Figure CN120216976A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal medical large model system and a medical image auxiliary diagnosis system, and the system comprises a training data obtaining module which is used for obtaining medical text information, medical image data and attention data; the model construction module is used for constructing a text medical big model and a mixed attention medical image model to obtain an initial multi-modal medical big model; the training processing module is used for training the initial multi-modal medical large model according to the medical text information, the medical image data and the attention data to obtain a target multi-modal medical large model; and the result optimization module is used for performing optimization processing according to the diagnosis result output by the target multi-mode medical large model to obtain an automatic diagnosis result. According to the invention, accurate diagnosis results can be output for different diseases; accurate interaction can be carried out based on the attention of the user; multiple rounds of questions and answers are supported, questions of the automatic diagnosis result are explained, and the medical experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of large models, and particularly relates to a multimodal medical large model system and a medical image assisted diagnosis system. Background Art

[0002] Manpower analysis by medical professionals such as doctors is the main means for analyzing the results of medical imaging examinations such as CT (Computed Tomography) and magnetic resonance imaging. However, due to problems such as the level of doctors, misjudgments or missed judgments often occur in this manpower analysis method.

[0003] Currently, technologies for automatically analyzing medical images such as CT and magnetic resonance imaging through medical image analysis models have gradually emerged. However, the current medical image analysis models still have the following problems: Most models are classification models for single diseases, and their coverage is difficult to meet the work requirements of frontline medical workers for general screening; the interpretability of the models is relatively poor, it is difficult to give the basis and process of automatic analysis, and the credibility is difficult to guarantee in practical applications; the interaction ability of the models is relatively weak, unable to meet the multimodal demand input of frontline doctors, and it is difficult to effectively feedback the real-time attention of users such as doctors. Summary of the Invention

[0004] The technical problem to be solved by the present disclosure is to overcome the defects of the existing medical image analysis models, such as being able to handle relatively single diseases, having relatively poor interpretability, and relatively weak interaction ability, and to provide a multimodal medical large model system and a medical image assisted diagnosis system.

[0005] The present disclosure solves the above technical problems through the following technical solutions:

[0006] The present disclosure provides a multimodal medical large model system, including:

[0007] A training data acquisition module, configured to acquire medical text information, medical image data, and attention data;

[0008] A model construction module, configured to construct a text medical large model and a medical image model with hybrid attention to obtain an initial multimodal medical large model;

[0009] A training processing module, configured to train the initial multimodal medical large model according to the medical text information, the medical image data, and the attention data to obtain a target multimodal medical large model;

[0010] A result optimization module, configured to perform optimization processing on the diagnosis result output by the target multimodal medical large model to obtain an automatic diagnosis result.

[0011] Preferably, the training processing module includes:

[0012] A vectorization unit for vectorizing the medical text information, the medical image data, and the attention data to obtain corresponding text vectors, image vectors, and prompt vectors;

[0013] A training unit for inputting the text vectors, the image vectors, and the prompt vectors into the initial multi-modal medical large model for training to obtain the target multi-modal medical large model.

[0014] Preferably, the training processing module further includes:

[0015] A text-image alignment unit for aligning the text vectors, the image vectors, and the prompt vectors according to a preset text-image alignment model before inputting them into the initial multi-modal medical large model for training.

[0016] Preferably, the preset text-image alignment model adopts a CLIP (contrastive language-image pre-training model) model.

[0017] Preferably, the training unit is also used to concatenate and merge the text vectors, the image vectors, and the prompt vectors and input them into the initial multi-modal medical large model for training to obtain the target multi-modal medical large model.

[0018] Preferably, the model construction module includes:

[0019] A text model construction unit for pre-training the text medical large model, which is used to vectorize the medical text information to obtain the text vectors;

[0020] An image model construction unit for pre-training the medical image model with hybrid attention, which is used to vectorize the medical image data and the attention data to obtain the image vectors and the prompt vectors.

[0021] Preferably, the text medical large model is pre-trained using a GPT3 (an artificial intelligence model) model or a Llama2 (an open-source language model).

[0022] Preferably, the result optimization module includes:

[0023] A first test data acquisition unit for acquiring a first test medical image and a plurality of first attention data when a first test user interacts with the first test medical image in different ways;

[0024] The first test result output unit is configured to input the first test medical image and a plurality of the first attention data into the target multi-modal medical large model to output a corresponding plurality of first test diagnosis results;

[0025] The sorting unit is configured to sort a plurality of the first test diagnosis results to obtain a sorting result;

[0026] The first model update unit is configured to obtain a first evaluation result based on the sorting result, and iteratively update the target multi-modal medical large model based on the first evaluation result until the evaluation result meets a preset evaluation condition.

[0027] Preferably, the result optimization module further includes:

[0028] The sorter training unit is configured to train a sorter according to the sorting result;

[0029] The second test data acquisition unit is configured to acquire a second test medical image and a plurality of second attention data when a second test user interacts with the second test medical image in different ways;

[0030] The second test result output unit is configured to input the second test medical image and a plurality of second attention data into the target multi-modal medical large model to output a corresponding plurality of second test diagnosis results;

[0031] The evaluation unit is configured to evaluate the second test diagnosis results based on the sorter to obtain a second evaluation result;

[0032] The second model update unit is configured to iteratively update the target multi-modal medical large model based on the second evaluation result until the evaluation result meets a preset evaluation condition.

[0033] The present disclosure further provides a medical image assisted diagnosis system including the above multi-modal medical large model system, including:

[0034] The user input unit is configured to acquire medical image data corresponding to a target medical image;

[0035] The action capture unit is configured to acquire the attention focus of the user on the target medical image;

[0036] The question and answer unit is configured to ask questions to the multi-modal medical large model system according to the automatic diagnosis result;

[0037] The multi-modal medical large model system is configured to output a diagnosis result according to the medical image input by the user, and explain the diagnosis result according to the attention focus and questions of the user.

[0038] On the basis of conforming to the common knowledge in the art, various preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present disclosure.

[0039] The positive and progressive effects of the present disclosure are as follows: The present disclosure uses a text medical large model to understand medical data, a medical image model with hybrid attention to understand medical images and capture the user's attention, and improves the overall performance of the multi-modal medical large model through text-image alignment, joint training, and model optimization, realizes the conversion and fusion of different modal data, enables the target multi-modal medical large model to process multi-modal input data such as text, images, and attention, and can automatically output more accurate diagnostic results for different types of diseases, can perform more precise interactions based on the user's attention, and also supports multi-round question answering to explain doubts about the automatic diagnostic results, thereby improving the user's medical experience as a whole. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the first module of the multi-modal medical large model system according to Embodiment 1 of the present disclosure.

[0041] Figure 2 It is a schematic diagram of the second module of the multi-modal medical large model system according to Embodiment 1 of the present disclosure.

[0042] Figure 3 It is a schematic diagram of the module of the medical image assisted diagnosis system according to Embodiment 2 of the present disclosure.

[0043] Figure 4 It is a schematic diagram of the framework of the medical image assisted diagnosis system according to Embodiment 2 of the present disclosure.

[0044] Figure 5 It is an example diagram of the user attention capture method according to Embodiment 2 of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The present disclosure will be further described below by way of embodiments, but the present disclosure is not limited to the scope of the described embodiments.

[0046] Embodiment 1

[0047] This embodiment provides a multi-modal medical large model system, as Figure 1 shown, including the following modules:

[0048] A training data acquisition module 1, configured to acquire medical text information, medical image data, and attention data.

[0049] The training data acquisition module is used to acquire a large amount of training data from the entire network, including but not limited to the above-mentioned medical text information, medical image data, and attention data, etc., and this data is used in the subsequent construction, training, and optimization processes of the multi-modal medical large model.

[0050] The model construction module 2 is used to construct a text medical large model and a medical image model with hybrid attention to obtain an initial multi-modal medical large model.

[0051] The model construction module is used for pre-training to obtain a text medical large model and a medical image model with hybrid attention to obtain an initial multi-modal medical large model. Specifically, the text medical large model is trained based on medical text information, used to understand the semantic expression of medical problems, can answer medical questions, and vectorize new medical texts; the medical image model with hybrid attention is trained based on medical imaging data, and its main function is to convert medical images into image vectors and encode attention data such as mouse clicks, drawing circles, and human eye attention.

[0052] The training processing module 3 is used to train the initial multi-modal medical large model according to medical text information, medical imaging data, and attention data to obtain a target multi-modal medical large model.

[0053] On the basis of pre-training the text medical large model and the medical image model with hybrid attention to obtain an initial multi-modal medical large model, the training processing module is used to perform joint training processing on the initial multi-modal medical large model to obtain a target multi-modal medical large model. Specifically, the initial multi-modal medical large model is trained based on medical text information, medical imaging data, and attention data to adjust the parameters of the model, obtain a target multi-modal medical large model, and realize the automatic generation of diagnostic results based on medical imaging data.

[0054] The result optimization module 4 is used to perform optimization processing on the diagnostic results output by the target multi-modal medical large model to obtain an automatic diagnostic result.

[0055] After the model training is completed, in order to achieve a more accurate response to user input, the result optimization module further optimizes the model according to the diagnostic results output by the target multi-modal medical large model.

[0056] In this solution, the multi-modal medical large model system uses the text medical large model to understand medical data, uses the medical image model with hybrid attention to understand medical images and capture the user's attention, and improves the overall performance of the multi-modal medical large model through joint training and model optimization, so that the multi-modal medical large model system can automatically output more accurate diagnostic results for different types of diseases, can perform more precise interactions based on the user's attention, and can also support multi-round Q&A to explain the doubts about the automatic diagnostic results, improving the user's medical experience as a whole.

[0057] In an implementable solution, as Figure 2 shown, the training processing module 3 includes:

[0058] The vectorization unit 31 is used to perform vectorization processing on medical text information, medical image data, and attention data to obtain corresponding text vectors, image vectors, and prompt vectors.

[0059] The text medical large model is used to perform vectorization processing on medical text information to obtain text vectors; the medical image model with hybrid attention is used to perform vectorization processing on medical image data and attention data to obtain image vectors and prompt vectors.

[0060] The training unit 32 is used to input the text vectors, image vectors, and prompt vectors into the initial multi-modal medical large model for training to obtain the target multi-modal medical large model.

[0061] Input the text vectors, image vectors, and prompt vectors into the text medical large model. The text medical large model learns the diagnostic results through instruction tuning to obtain the target multi-modal medical large model, realizing the automatic generation of diagnostic results based on medical image data.

[0062] In this solution, through the vectorization unit to perform vectorization processing on the data, and then through the training unit for training processing, the target multi-modal medical large model can be obtained, realizing the automatic generation of diagnostic results.

[0063] In an implementable solution, as Figure 2 shown, the training processing module 3 further includes:

[0064] The text-image alignment unit 33 is used to perform alignment processing on the text vectors, image vectors, and prompt vectors according to the preset text-image alignment model before inputting them into the initial multi-modal medical large model for training.

[0065] The text-image alignment unit uses the preset text-image alignment model to extract medical images and corresponding descriptions from data such as electronic medical records and academic literature, and trains to obtain the target text-image alignment model to achieve the alignment of text vectors, image vectors, and prompt vectors.

[0066] In this solution, through the text-image alignment model to perform the alignment of text vectors, image vectors, and prompt vectors, realizing the conversion and fusion of different modal data, so that the target multi-modal medical large model can process multi-modal input data such as text, images, and attention to obtain diagnostic results.

[0067] In an implementable solution, the preset text-image alignment model adopts the CLIP model.

[0068] The CLIP model is a pre-trained model that can process text and images simultaneously, and is trained from unlabeled image and text data through self-supervised learning.

[0069] In this solution, text-image alignment is achieved through the CLIP model, enabling the multi-modal system to understand the semantic connection between images and text and improving the performance of the target multi-modal medical large model.

[0070] In an implementable solution, the training unit is also used to concatenate and merge the text vector, image vector, and prompt vector, and input them into the initial multi-modal medical large model for training to obtain the target multi-modal medical large model.

[0071] The text vector, image vector, and prompt vector can be concatenated and merged by means such as multiplication and addition.

[0072] In this solution, by concatenating and merging different types of vectors and inputting them into the model for training, the model can output more accurate diagnostic results for the input multi-modal data.

[0073] In an implementable solution, as Figure 2 shown, the model construction module 2 includes:

[0074] A text model construction unit 21, which is used to pre-train a text medical large model, and the text medical large model is used to vectorize medical text information to obtain a text vector.

[0075] The text medical large model is trained based on a large amount of medical text information (including but not limited to Chinese and English medical book data, medical guidance plans, hospital case data, etc.), and is used to vectorize medical text information to obtain a text vector.

[0076] An image model construction unit 22, which is used to pre-train a medical image model with hybrid attention, and the medical image model with hybrid attention is used to vectorize medical image data and attention data to obtain an image vector and a prompt vector.

[0077] The medical image model with hybrid attention is trained based on medical image data, and is used to vectorize medical image data and attention data to obtain an image vector and a prompt vector.

[0078] In this solution, the data is vectorized through the text medical large model and the medical image model with hybrid attention, enabling the data to be used for subsequent model training.

[0079] In an implementable solution, the text medical large model is pre-trained using the GPT3 model or Llama2.

[0080] The construction of the text medical large model can include but is not limited to the following two methods:

[0081] The first is to train a text medical large model from scratch. Use the GPT3 model to build an initial text medical large model with a parameter scale of 60B. Collect high-quality shared Chinese and English medical book data, medical guidance plans, hospital medical records, etc. from the whole network as training data, and train on a cluster of 1000 GPUs (graphics processors) of the A100 specification for about one month to obtain the text medical large model; The second training method is to continue training medical texts on the basis of an open-source text large model. You can choose Llama2 as the base model, and then use medical text data for continued training to obtain the text medical large model.

[0082] In this solution, the GPT3 model or the Llama2 base model is used to train the text medical large model, so that the text medical large model can process medical type text data.

[0083] In an implementable solution, as Figure 2 shown, the result optimization module 4 includes:

[0084] The first test data acquisition unit 41 is used to acquire the first test medical image and several first attention data when the first test user interacts with the first test medical image in different ways;

[0085] The first test result output unit 42 is used to input the first test medical image and several first attention data into the target multi-modal medical large model to output corresponding several first test diagnosis results;

[0086] The sorting unit 43 is used to sort several first test diagnosis results to obtain a sorting result;

[0087] The first model update unit 44 is used to obtain a first evaluation result based on the sorting result, and based on the first evaluation result, iteratively update the target multi-modal medical large model until the evaluation result meets the preset evaluation conditions.

[0088] After completing the joint training of the target multi-modal medical large model, in order to achieve a more accurate response to user input, the result optimization module further optimizes the model according to the diagnosis results output by the target multi-modal medical large model. Specifically, the first test data acquisition unit acquires the first test medical image and several first attention data when the user interacts in different ways such as mouse clicking, handwritten circling, and eye tracker; the first test result output unit inputs the above data into the target multi-modal medical large model to output corresponding several first test diagnosis results; the sorting unit sorts the results; the first model update unit optimizes and updates the target multi-modal medical large model according to the sorting result.

[0089] In this solution, optimizing and updating the model according to the diagnostic results output by the model can improve the accuracy of the model output, enabling the target multi-modal medical large model to make a more precise response to user input.

[0090] In an implementable solution, as Figure 2 shown, the result optimization module 4 further includes:

[0091] A sorter training unit 45 for training a sorter according to the sorting result;

[0092] A second test data acquisition unit 46 for acquiring a second test medical image and a number of second attention data when a second test user interacts with the second test medical image in different ways;

[0093] A second test result output unit 47 for inputting the second test medical image and a number of second attention data into the multi-modal medical large model to output a corresponding number of second test diagnostic results;

[0094] An evaluation unit 48 for evaluating the second test diagnostic results based on the sorter to obtain a second evaluation result;

[0095] A second model update unit 49 for iteratively updating the multi-modal medical large model based on the second evaluation result until the evaluation result meets the preset evaluation conditions.

[0096] Specifically, the sorter training unit is used to train a sorter according to the above sorting result; the second test data acquisition unit acquires a second test medical image and a number of second attention data when the user interacts in different ways such as mouse clicking, handwritten circling, and using an eye tracker; the second test result output unit inputs the above data into the target multi-modal medical large model to output a corresponding number of second test diagnostic results; the evaluation unit evaluates the second test diagnostic results based on the sorter to obtain a second evaluation result; the second model update unit uses reinforcement learning to perform weight iteration on the model according to the sorting result to optimize and update the target multi-modal medical large model.

[0097] In this solution, by training a sorter and using the sorter to evaluate the results output by the model, the model can be optimized and updated, further enabling the target multi-modal medical large model to make a more precise response to user input.

[0098] Next, a specific implementation manner is used to illustrate the construction process of the multi-modal medical large model system provided in this embodiment.

[0099] (1) Pre-training of the text medical large model

[0100] The pre-training of the text medical large model can include, but is not limited to, the following two methods: The first is to train the text medical large model from scratch. Use the GPT3 model to build an initial text medical large model with a parameter scale of 60B. Collect high-quality shared Chinese and English medical book data, medical guidance plans, hospital medical records, etc. from the whole network as training data. Train on a GPU cluster of 1000 A100 specifications for about a month to obtain the text medical large model. This method has a relatively high training cost. The second training method is to continue training medical texts on the basis of the open-source text large model. You can choose Llama2 as the base model and then use medical text data for continued training to obtain the text medical large model. The training cost of this method is related to the size of the training data volume and is smaller than the first method.

[0101] (2) Pre-training of the medical image model with hybrid attention

[0102] First, collect a large amount of medical examination image data and use image training models such as Vision Transformer (image conversion model) to train the image encoder. Its main function is to convert medical images into image vectors. Then train the masked encoder, mainly encoding some clicks, attention and other prompts, and summing with the image vectors according to the input and prompts. Finally, use the CLIP model to extract medical images and corresponding descriptions from electronic medical records and academic literature, and train the image-text alignment model to achieve the alignment of image encoding (the above image vectors), prompt encoding (the above prompt vectors) and text encoding (the above text vectors).

[0103] (3) Joint training of multi-modal large models

[0104] After completing the pre-training of the text medical large model and the pre-training of the medical image model, extract medical examination images, diagnostic descriptions, and diagnostic bases from the electronic medical records as joint training data. Use the image pre-training model to encode the medical examination images to obtain image vectors, and input the vectors and the set prompt templates into the text medical large model. The text medical large model generates corresponding diagnostic results. Use this process and the training data for model fine-tuning to achieve the automatic generation of diagnostic bases and diagnostic descriptions based on medical examination images.

[0105] (4) Optimization of diagnostic results based on multiple input means

[0106] After completing the joint training of the multi-modal medical large model, in order to generate a more accurate response to the user's input, the user's attention is captured through methods or devices such as mouse clicks, handwritten circles, and eye trackers. The position areas of the three types of attention, namely mouse clicks, handwritten circles, and human eye attention, are encoded, and this type of attention is encoded for prompting. It is jointly input into the target multi-modal medical large model together with image encoding and text encoding, enabling the large model to output the most correct four results. The user sorts these results and trains a sorter based on the sorting results; randomly simulate new input medical images and output diagnostic results, use the sorter to evaluate the results, and use reinforcement learning to perform weight iteration on the target multi-modal medical large model, thereby optimizing the automatic diagnostic results.

[0107] The multi-modal medical large model system provided in this embodiment uses a text medical large model to understand medical data, a medical image model with hybrid attention to understand medical images and capture the user's attention, and improves the overall performance of the multi-modal medical large model through text-image alignment, joint training, and model optimization, realizing the conversion and fusion of different modal data. This enables the target multi-modal medical large model to process multi-modal input data such as text, images, and attention, and can automatically output more accurate diagnostic results for different types of diseases, perform more precise interactions based on the user's attention, support multi-round Q&A, explain doubts about the automatic diagnostic results, and overall enhance the user's medical experience.

[0108] Embodiment 2

[0109] As Figure 3 shown, this embodiment provides a medical image assisted diagnosis system including the multi-modal medical large model system as in Embodiment 1, comprising:

[0110] A user input unit 5 for obtaining medical image data corresponding to the target medical image;

[0111] An action capture unit 6 for obtaining the user's attention focus on the target medical image;

[0112] A Q&A unit 7 for asking questions to the multi-modal medical large model system according to the automatic diagnostic results;

[0113] A multi-modal medical large model system 8 for outputting a diagnostic result according to the medical image input by the user and explaining the diagnostic result according to the user's attention focus and questions.

[0114] Figure 4 This is the overall framework of the medical image assisted diagnosis system including the multi-modal medical large model system. The multi-modal medical large model in Embodiment 1 is obtained through multi-modal large model pre-training and user feedback reinforcement training for human-machine co-integrated intelligent medical image assisted diagnosis:

[0115] (1) Automatically diagnose based on the patient's medical diagnostic images. Directly take the patient's examination images as input, and the multi-modal medical large model automatically generates diagnostic bases and results according to the medical images.

[0116] (2) Diagnostic analysis integrating attention. This system supports devices such as mouse clicks, handwritten circles, and eye trackers to capture the attention of users such as doctors. The multi-modal medical large model further interprets the image where the attention is located to achieve a human-machine integrated medical image-assisted diagnosis interaction.

[0117] (3) Medical Q&A based on the diagnostic results. The multi-modal large model naturally supports multi-round Q&A, can elaborate on the doubts of the automatic diagnosis, and improve the credibility and usability of the automatic diagnosis.

[0118] Specifically, the user input unit is used to obtain the medical image data corresponding to the target medical image (i.e., the medical examination image input in Figure 4 ); the action capture unit is used to obtain the attention focus of the user on the target medical image (i.e., the user attention capture in Figure 4 ). Taking Figure 5 as an example, the user attention (such as the area pointed by the arrow in Figure 5 ) can be obtained by means of mouse clicks, drawing circles, eye trackers, etc., which capture the attention of the human eye; the Q&A unit is used to ask questions to the multi-modal medical large model system according to the automatic diagnosis results (i.e., the user question input in Figure 4 ).

[0119] The medical image-assisted diagnosis system including the multi-modal medical large model system provided in this embodiment can automatically output more accurate diagnostic results for different types of diseases; can accept various input means and perform more precise interactions based on the user's attention; can also support multi-round Q&A, elaborate on the doubts of the automatic diagnosis results based on the thought chain of the large model, and overall improve the user's medical experience.

[0120] Although the specific implementation manners of the present disclosure have been described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these implementation manners, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A multi-modal medical large model system, characterized in that, Including: A training data acquisition module, configured to acquire medical text information, medical image data, and attention data; A model construction module, configured to construct a text medical large model and a medical image model with hybrid attention to obtain an initial multi-modal medical large model; A training processing module, configured to train the initial multi-modal medical large model according to the medical text information, the medical image data, and the attention data to obtain a target multi-modal medical large model; A result optimization module, configured to perform optimization processing on the diagnosis result output by the target multi-modal medical large model to obtain an automatic diagnosis result.

2. The multimodal medical large model system according to claim 1, wherein The training processing module includes: A vectorization unit, configured to perform vectorization processing on the medical text information, the medical image data, and the attention data to obtain corresponding text vectors, image vectors, and prompt vectors; A training unit, configured to input the text vectors, the image vectors, and the prompt vectors into the initial multi-modal medical large model for training to obtain the target multi-modal medical large model.

3. The multimodal medical large model system according to claim 2, characterized in that, The training processing module further includes: A text-image alignment unit, configured to perform alignment processing on the text vectors, the image vectors, and the prompt vectors according to a preset text-image alignment model before inputting them into the initial multi-modal medical large model for training.

4. The multimodal medical large model system according to claim 3, wherein, The preset text-image alignment model adopts a CLIP model.

5. The multimodal medical large model system according to claim 3, wherein The training unit is further configured to perform concatenation and merging processing on the text vectors, the image vectors, and the prompt vectors, and input them into the initial multi-modal medical large model for training to obtain the target multi-modal medical large model.

6. The multimodal medical large model system according to any one of claims 2-5, characterized in that, The model construction module includes: A text model construction unit, configured to pre-train the text medical large model, and the text medical large model is configured to perform vectorization on the medical text information to obtain the text vectors; An image model construction unit, configured to pre-train the medical image model with hybrid attention, and the medical image model with hybrid attention is configured to perform vectorization on the medical image data and attention data to obtain the image vectors and the prompt vectors.

7. The multimodal medical large model system according to claim 6, wherein, The text medical large model is pre-trained using a GPT3 model or Llama2.

8. The multimodal medical large model system according to claim 1, wherein The result optimization module includes: A first test data acquisition unit, configured to acquire a first test medical image, and a plurality of first attention data when a first test user interacts with the first test medical image in different ways; A first test result output unit, configured to input the first test medical image and the plurality of first attention data into the target multi-modal medical large model to output corresponding plurality of first test diagnosis results; A sorting unit, configured to sort the plurality of first test diagnosis results to obtain a sorting result; A first model update unit, configured to obtain a first evaluation result based on the sorting result, and iteratively update the target multi-modal medical large model based on the first evaluation result until the evaluation result meets a preset evaluation condition.

9. The multimodal medical large model system according to claim 8, wherein, The result optimization module further includes: A sorter training unit, configured to train a sorter according to the sorting result; A second test data acquisition unit, configured to acquire a second test medical image and a plurality of second attention data when a second test user interacts with the second test medical image in different ways; A second test result output unit, configured to input the second test medical image and the plurality of second attention data into the target multi-modal medical large model to output corresponding second test diagnosis results; An evaluation unit, configured to evaluate the second test diagnosis results based on the sorter to obtain a second evaluation result; A second model update unit, configured to iteratively update the target multi-modal medical large model based on the second evaluation result until the evaluation result meets a preset evaluation condition.

10. A medical image assisted diagnosis system comprising the multi-modal medical large model system according to any one of claims 1-9, characterized in that, Comprising: A user input unit, configured to acquire medical image data corresponding to a target medical image; An action capture unit, configured to acquire the attention focus points of a user on the target medical image; A question-and-answer unit, configured to ask questions to the multi-modal medical large model system according to an automatic diagnosis result; The multi-modal medical large model system, configured to output a diagnosis result according to the medical image input by a user, and explain the diagnosis result according to the attention focus points and questions of the user.