Multi-modal large model design method and system with capability of simultaneously processing multiple medical visual language tasks
By designing a multi-task end-to-end multimodal medical large model and utilizing multi-task medical vision datasets and a phased training strategy, we solved the problems of difficulty in acquiring medical image datasets and insufficient task generalization capabilities, and achieved efficient processing and accuracy improvement of multiple medical vision language tasks.
Patent Information
- Application Number
- CN202511094000.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Medical image datasets are difficult to obtain, the correspondence between different modalities of medical images is complex and requires rigor, there is a lack of comprehensive evaluation benchmarks, and existing multimodal large language models lack the ability to generalize and adapt to tasks in the medical field.
A multi-task end-to-end multimodal medical large model is designed, using a multi-task medical vision dataset and a phased progressive training strategy, including an image encoding module, a vocabulary embedding module, a linear projection layer module, a large language model module, and a multi-attribute expert prompt module. Expert prompts are generated through visual question answering and prediction strategies, and the MedVision-MT dataset is constructed for multi-task training and adaptive adjustment.
It has improved the task generalization capability of multimodal large language models in the medical field, enhanced the model's capture of medical visual features and semantic understanding, and is able to process multiple medical visual language tasks simultaneously, promoting the application of medical artificial intelligence in clinical practice.
Smart Images

Figure CN120597938A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of multimodal large models and medical image processing, and specifically to a multimodal large model design method and system capable of simultaneously processing multiple medical visual language tasks. Background Art
[0002] In recent years, large multimodal language models (MLLMs) such as GPT-4 and LLaVA have demonstrated outstanding performance in multiple fields. In text generation and visual understanding, they leverage strong generalization capabilities, achieving excellent generalization across diverse tasks. Current research on large multimodal language models focuses on general knowledge adaptation and model optimization. By aligning visual features and large language models with large natural images and detailed annotations, and leveraging human instructions to enhance conversational and reasoning capabilities, they excel in general tasks such as image classification and text generation, demonstrating strong cross-task adaptability. However, medical applications present distinct challenges due to the unique challenges inherent in medical imaging and semantics.
[0003] Difficulty in obtaining medical image datasets: Developing efficient multimodal large language models requires a large amount of diverse medical data to align visual features and large language models. However, the collection of medical image datasets is restricted by patient privacy protection policies and high medical costs. The number is limited and difficult to obtain. In addition, the quality of medical data is low, the image datasets have coarse data granularity, and lack information accuracy and details. Compared with the rich and fine-grained image datasets in the natural field (such as large-scale Internet datasets), the challenges are enormous.
[0004] The correspondence between different modalities of medical images is complex and requires rigorousness: the terminology and descriptions in medical image reports are professional and precise, and key terms have special semantics. They cannot be arbitrarily replaced like natural image descriptions. If important information is misunderstood, it may lead to serious consequences such as misdiagnosis. Therefore, the medical field requires multimodal large language models to not only understand images and text, but also accurately master and process medical image details to ensure that information is conveyed correctly.
[0005] The medical field lacks comprehensive evaluation benchmarks: While well-established multimodal large language model evaluation systems exist in the natural sciences, the medical field lacks comprehensive benchmarks specifically tailored to these models. Most evaluation systems are still in their infancy, lacking a comprehensive, standardized evaluation framework. While existing research on medical multimodal large language models has made initial progress in migrating original models to the medical field, these models often focus on specific tasks and lack the flexibility and adaptability to handle a wider range of biomedical tasks. This makes it difficult to efficiently execute these diverse and complex medical tasks across tasks and domains.
[0006] Recent research exploring the potential of large multimodal language models in medical applications has made some progress, but these efforts have focused on single-task solutions and lack the ability to design flexible, generalizable models for diverse biomedical tasks. The diverse and complex nature of medical tasks demands high model understanding, reasoning, and adaptability. Current large multimodal language models in medicine need further improvement in terms of task generalization and practical applicability.
[0007] In summary, while research on large multimodal language model systems in the medical field has made progress, it still faces challenges in data acquisition, cross-modal alignment accuracy, integration of specialized domain knowledge, and task generalization. To address these issues, this paper proposes a method and system for designing large multimodal models capable of simultaneously processing multiple medical visual language tasks. Summary of the Invention
[0008] The purpose of the present invention is to provide a multimodal large model design method and system that can simultaneously process multiple medical visual language tasks, break through the application bottleneck of medical scenarios, and improve task generalization and execution capabilities.
[0009] In order to achieve the above object, the present invention is implemented through the following technical solutions: The present invention provides a multimodal large model system capable of simultaneously processing multiple medical vision and language tasks. The multimodal large model system includes a multi-task end-to-end multimodal medical large model (a multi-task end-to-end multimodal medical large model proposed based on a multi-task medical vision dataset and a phased progressive training strategy). The multi-task end-to-end multimodal medical large model includes a basic module and a network structure, wherein the network structure includes a communication-connected image encoding module, a vocabulary embedding module, a linear projection layer module, a large language model module, and a multi-attribute expert prompt module. Among them, the image encoding module supports image modality input and uses the image encoding module to encode the features of the input image; the vocabulary embedding module supports text modality input and uses a vocabulary embedding module to encode the features of the input text; the linear projection layer module aligns the image modality features with the text modality features; The large language model module supports fluent text output and uses the large language model to decode the input text and image features to obtain the final text output; the multi-attribute expert prompt module selects all pathology categories as medical entities based on the data tuples of the MedVision-MT training set (benchmark dataset MedVision-MT) to generate expert knowledge.
[0010] Preferably, four attributes are designed for each medical entity, including shape [SHAPE], size [SIZE], location [LOCATION] and symptom [SYMPTOM], and are represented by A1, A2, A3 and A4 respectively. Among them, the three attributes of shape [SHAPE], size [SIZE] and location [LOCATION] represent the low-level information of each medical entity in the corresponding image, while the last attribute symptom [SYMPTOM] represents the high-level information.
[0011] Preferably, in order to obtain expert-level prompts, a large language model is used to generate the key attributes of each medical entity, and two generation strategies are adopted: visual question answering (VQA) and prediction (Pre). In the VQA solution (it is worth noting that the present invention does not limit the responses generated by the VQA solution to a single vocabulary, which can enhance the diversity of attribute values of each medical entity), the query template is designed as: "[Medical entity] [A i ] is”; while in the Pre scheme, the sentence to be predicted is: “[Medical entity]’s [A i ] is [MASK]", wherein the A i is any one of the four attributes; the response and [MASK] part of the VQA query represent the key attribute values obtained by these two generation schemes respectively.
[0012] Preferably, the two generation strategies are applied to each object to obtain an attribute pool, each entity containing eight different attribute values; the resulting expert suggestion process is summarized by the following formula: ; In the above formula: VQA means to explore the relationship between medical images and knowledge through question-answering strategy, that is, by asking questions about images to the model and obtaining answers; χ represents the input image, which is the basic data object for the model to perform visual analysis and knowledge extraction; Q represents the query question, specifically: Q A1 ,…,Q AN : Query questions designed based on medical entity attributes (such as [SHAPE], [SIZE], etc.) are used to obtain attribute information of entities in medical images through VQA strategies. N can be understood as the number of attribute-related questions.
[0013] Pre means mining attribute knowledge through prediction strategy by letting the model predict the mask ([MASK]) part in the text related to medical entity attributes; S A1 To S AN Represents sentences corresponding to different predictions. Specifically, it is a text pattern designed for prediction strategies based on medical entity attributes, including mask positions, which are used to guide the model to generate attribute-related content. N also corresponds to the number of attribute-related texts.
[0014] ExpertPrompt: The final generated expert prompt is used to input into the large language model to assist the model in understanding the expertise and features in medical images.
[0015] Template n :A template for constructing expert tips, n can represent the number or type of the template, by combining the sampled X and the medical entity (Object k ) fill in the template to form prompt text that meets the model input requirements.
[0016] Object k : Medical entities, namely pathological categories, organs and other objects involved in medical images, k is used to distinguish different medical entities.
[0017] The present invention provides a method for designing a multimodal large model capable of simultaneously processing multiple medical visual language tasks, the method comprising the following steps: Step 1): A unified dataset covering various medical vision tasks designed for multimodal large model system training, specifically including: Step 1.1) New medical vision task setting: Collect and reorganize popular medical vision datasets to form the benchmark dataset MedVision-MT. The benchmark dataset MedVision-MT includes five core tasks: medical question answering Med-VQA, radiology report generation RRG, disease localization DL, clinical diagnosis CD, and local clinical diagnosis LCD; Disease localization (DL) aims to determine the location of the disease; radiology report generation (RRG) aims to generate medical reports based on images; CD refers to diagnosing diseases based on medical images; Med-VQA focuses on answering questions based on images; local clinical diagnosis (LCD) is the inverse process of disease localization (DL), which classifies the disease after a given bounding box is obtained. Through this comprehensive and diverse multi-task setting, the benchmark dataset MedVision-MT can provide a foundation for further exploration in the field of medical MLLM. Step 1.2) Multi-task medical image dataset: Several popular medical multimodal datasets were reorganized to form a multi-task benchmark dataset MedVision-MT. Step 2): Phased adaptation training; specifically, Step 2.1) Multi-task warm-up training phase: First, the model is pre-trained on a variety of large-scale natural image datasets. Then, the model's multi-task processing capabilities are transferred from the general domain to the medical domain. During the multi-task warm-up phase, the projection matrix W is pre-trained, and the large language model is adjusted on various natural image datasets as needed to ensure that the pre-trained model can handle different tasks under human queries. Step 2.2) Biomedical information adaptation phase: After completing the multi-task warm-up, the model's ability to handle different visual tasks in the general field was initially established; then, the multi-task benchmark dataset MedVision-MT (Multi-Task Medical Dataset MedVision-MT) was used to adapt the model to the medical field.
[0018] Preferably: step 2.1) specifically includes: replacing biomedical tasks with five natural visual tasks respectively, among which image caption generation replaces radiological report generation RRG, visual question answering replaces medical question answering Med-VQA, localization replaces disease localization DL, classification replaces clinical diagnosis CD, and object classification replaces local clinical diagnosis LCD; gradually transferring multi-task capabilities from the natural domain to the medical domain, and establishing semantic connections between general vocabulary and complex biomedical terms through expert prompts to achieve more effective capture of medical image features.
[0019] Preferred: Step 1.2) The sub-datasets collected by Traditional Chinese Medicine Question Answering Med-VQA include VQA-Med-2019, VQA-Med-2021, SLAKE, VQA-RAD, and Path-VQA; the sub-datasets collected by Radiology Report Generation RRG include MIMIC-CXR and IU-X-Ray datasets; the sub-datasets collected by Clinical Diagnosis CD include NIH datasets; the sub-datasets collected by Disease Localization DL and Local Clinical Diagnosis LCD include ChestXray and MS-CXR datasets.
[0020] Preferred: In order to effectively filter out meaningful and high-quality data from the sub-dataset, further data filtering is performed on both image and text modalities, and only suitable image-text pairs are retained to form the operational dataset. ,in Represents an image-text pair (Image-TextPair), which is the basic unit of the operation dataset. Represents the text portion of an image-text pair, which contains semantic information such as medical reports, questions, and diagnosis results.
[0021] Preferably, the distribution ratios of the sub-datasets ChestXray, MS-CXR, VQA-RAD, VQA-Med-2021, SLAKE, IU-X-Ray, VQA-Med-2019, NIH, Path-VQA, and MIMIC-CXR are 5.5%, 0.88%, 1.3%, 2.2%, 5.0%, 6.6%, 7.3%, 16.3%, 26.1%, and 28.8%, respectively.
[0022] Beneficial Effects: 1) Dataset Advantages: The constructed MedVision-MT dataset covers a variety of medical vision tasks. This comprehensive and diverse task set provides a solid foundation for medical MLLM research, addressing the data scarcity issue in the medical field. Data filtering and instruction pool design improve data quality and relevance, helping the model better learn medical knowledge. 2) Training Strategy Advantages: A phased domain-adaptive training strategy, first warming up with natural domain data before migrating to the medical domain, effectively circumvents the problem of medical data scarcity and gradually improves the model's ability to handle medical tasks. The second phase trains the visual encoder and multi-attribute expert prompt generation method, enhancing the model's capture of medical visual features and semantic understanding, bridging domain gaps and improving the model's performance on medical tasks. 3) Model Advantages: The network architecture of the multi-task, end-to-end, multimodal medical large model features clear division of labor among modules. The multi-attribute expert prompt module is tailored to the specific medical domain, improving the alignment of visual features with the large language model. This enables the model to simultaneously handle multiple medical visual and language tasks, enhancing task generalization and practical application capabilities, and is expected to promote the widespread application of medical artificial intelligence in clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is one of the medical vision multi-task example sets in an embodiment of the present invention.
[0024] Figure 2 This is the second of the medical vision multi-task example sets in the embodiment of the present invention.
[0025] Figure 3 This is the third of the medical vision multi-task example sets in the embodiment of the present invention.
[0026] Figure 4 This is the fourth of the medical vision multi-task example sets in the embodiment of the present invention.
[0027] Figure 5 2 is a schematic diagram of a multi-task warm-up phase in an embodiment of the present invention.
[0028] Figure 6 2 is a schematic diagram of the biomedical information adaptation stage in an embodiment of the present invention.
[0029] Figure 7 2 is a schematic diagram of a multi-attribute expert prompt module in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0031] The present invention implements an end-to-end complete multimodal medical model, specifically as follows: 1. Construction of a unified dataset (MedVision-MT) A new medical vision task set: Popular medical vision datasets were collected and reorganized to form the MedVision-MT benchmark dataset, covering five core tasks: medical question answering (Med-VQA), radiology report generation (RRG), disease localization (DL), clinical diagnosis (CD), and local clinical diagnosis (LCD). Each task has a clear objective: DL locates the disease; RRG generates a medical report based on the image; CD diagnoses the condition based on the medical image; Med-VQA answers questions based on the image; and LCD, the inverse of DL, classifies the disease after being given a bounding box. This comprehensive and diverse task set provides a foundation for exploration in the field of medical MLLM.
[0032] Multi-task medical image dataset construction: Reorganize popular medical multimodal datasets to form MedVision-MT, collect corresponding sub-datasets for each task, filter image and text modality data, and retain appropriate image-text pairs to form operational datasets ,Five instruction pools are designed for each task to construct the instruction dataset.
[0033] 2. Phased Domain Adaptation Training Strategy Multi-task warm-up training phase: To address the problem of insufficient datasets in the medical field, the model is first preheated on a large-scale natural image dataset. Natural visual tasks (image caption generation, visual question answering, etc.) are used to replace biomedical tasks. The projection matrix W is pre-trained and a large language model is fine-tuned. The visual encoder originally trained on natural images is frozen, and multi-task capabilities are gradually transferred. Through expert prompts, semantic connections between common vocabulary and biomedical terms are established to effectively capture medical image features.
[0034] Biomedical information adaptation stage: The MedVision-MT dataset is used to adapt the model to the medical field, and the visual encoder module is further trained to enhance the perception of medical visual features. A multi-attribute expert prompt generation method is proposed to extract expert-level medical knowledge from the large language model, bridging the semantic and visual gap between the natural and medical fields.
[0035] 3. Multi-task end-to-end multimodal medical large model Basic module: Based on the commonly used TransformerBlock improvement and neural network module.
[0036] Network structure: Based on the improved structure of the classic LLaVA model, it includes an image encoding module (encoding input image features), a vocabulary embedding module (encoding input text features), a linear projection layer module (aligning image and text modal features), a large language model module (decoding text and image features and outputting text), and a multi-attribute expert prompt module (based on the MedVision-MT training set, selecting pathology categories as medical entities to generate expert knowledge, designing four attributes for each medical entity, and using VQA and Pre generation strategies to generate expert prompts, forming an attribute pool, and summarizing the expert prompt process through formulas).
[0037] Embodiment mode: 1. MedVision-MT Dataset Implementation New medical vision task definition and implementation: Medical vision-related datasets were collected, such as the VQA-Med-2019 and VQA-Med-2021 datasets for medical question answering (Med-VQA); and the MIMIC-CXR and IU-X-Ray datasets for radiology report generation (RRG). Based on the task definition, the task rules for DL (locating disease), RRG (generating medical reports), CD (diagnosis), Med-VQA (image question answering), and LCD (bounding box classification) were clarified. The datasets were annotated and organized to form the MedVision-MT dataset for subsequent model training.
[0038] like Figure 1 A specific embodiment is shown: RRG (Radiology Report Generation): The model takes a chest X-ray image as input and generates a medical report (e.g., the heart, lungs, and mediastinum are all within normal ranges). This model simulates the process of radiologists writing diagnostic reports based on images to test the model's medical text generation and image comprehension capabilities.
[0039] like Figure 2 The figure shows a specific embodiment, DL (Disease Localization): for chest X-ray images, the model outputs the answer A: the coordinates of the lesion location ({57, 46, 78, 68}, representing the bounding box information), to achieve spatial positioning of the disease in the image.
[0040] like Figure 3 A specific example is shown in CD (Clinical Diagnosis): A chest X-ray image is input and the question Q: Are any diseases found? The model outputs the answer A: Atelectasis, simulating the scenario of image-based disease determination in clinical diagnosis.
[0041] like Figure 4 A specific embodiment is shown: VQA (Medical Visual Question Answering): Given a joint X-ray image, the question Q is: Which organ system is shown in the image? The output answer is A: Musculoskeletal system. This tests the model's ability to answer medical knowledge by combining images and questions.
[0042] Implementation of multi-task medical image dataset: Integrate the collected sub-datasets, for example, integrate ChestXray and MS-CXR datasets for DL and LCD tasks. Then filter the image and text modalities, remove blurry and incomplete images and unclear and erroneous texts, retain appropriate image-text pairs, and construct the operational dataset. Five instructions with different expressions are designed for each task to form an instruction pool. For example, the instruction pool for the medical question-answering task includes instructions such as "What is the answer to the question corresponding to this medical image" and "Answer the question based on this medical image", enriching the instruction diversity of model training.
[0043] 2. Implement a phased training strategy Multi-task warm-up training: Select large-scale natural image datasets (such as the COCO dataset) and pre-train the model on them. Pre-train the projection matrix W, adjusting its parameters to achieve initial alignment between image and text features. Replace biomedical tasks with natural vision tasks, input data from image captioning tasks into the model, and train it to generate text similar to medical reports (simulating the RRG task). Also, train the model to answer medical-like questions using data from visual question-answering tasks (simulating the Med-VQA task). Freeze the visual encoder and train only the language model. This allows the model to learn to handle multiple tasks in the natural domain, preparing it for transfer to the medical domain.
[0044] like Figure 5As shown in the figure, the goal of the multi-task warm-up stage (Stage-1: Multi-TaskWarming-Up) is to let the model "practice" in the general field first, accumulate multi-task processing capabilities, and then migrate to the medical scenario. Module analysis: Pre-trainedLargeLanguageModel pre-trains large language models (such as GPT), which have the foundation for general text understanding and generation. VisualEncoder: Visual encoder, responsible for extracting image features (input natural images, output visual features). ProjectionW: Projection layer, aligns the dimensions of visual features and text features, allowing the model to integrate image and text information. TextEmbedding: Text embedding module, converts text (such as task instructions, questions) into vectors that the model can understand. Task substitution logic: Replace medical tasks with natural scene tasks as a "warm-up" to mitigate the impact of medical data scarcity: Image caption generation (captioning) replaces RRG (radiology report generation) to develop "image to text description" capabilities; general visual question answering (VQA) replaces medical VQA to build a foundation for "image-text question answering"; object localization replaces DL (disease localization) to practice "image to spatial coordinate output"; image classification replaces CD (clinical diagnosis) to master "image to category judgment"; object classification replaces LCD (local clinical diagnosis) to strengthen fine-grained classification capabilities. Training characteristics: Freeze the Visual Encoder (since it has been pre-trained on natural images), and only fine-tune the language model and projection layer to quickly transfer general multi-task capabilities.
[0045] right Figure 5Detailed Description: 1. Core Process and Module Overview: Model Input: Natural image (Image) + Text task (Text, such as Caption, VQA); Core Modules: Visual Encoder, Projection Layer (W), Text Embedding Module (Text Embedding), Pre-trained Large Language Model; Goal: The model learns the "image → text" multi-task mapping (such as image captioning and visual question answering), and accumulates general multimodal capabilities. 2. Element-by-Element Explanation: 1. Input Layer: Image + Text: Image: Refers to natural scene images (such as outdoor photos in the figure). It serves as the "visual material" for the model analysis and represents general image data (as distinct from medical imaging). Text: Contains multi-task instructions / data, corresponding to two general multimodal tasks: Task 1: Caption (Image Caption Generation): The model is required to generate a descriptive text for an image (e.g., "A group of people skateboarding in a park"), learning the "image → text description" mapping. Task 2: Visual Question Answering (VQA): Requires the model to answer questions about images (e.g., "How many people are in this image?"), learning the "image + question → answer" reasoning model. 2. Feature Encoding Layer: Visual Encoder + Text Embedding. Visual Encoder: Function: Extracts visual features from natural images (converting images into numerical feature vectors). Significance: Enables the model to "understand" image content (e.g., identifying skateboards, people, and park scenes), serving as the "visual foundation" for multimodal fusion. Text Embedding: Function: Converts "textual task instructions / data" (e.g., descriptions for image caption generation, VQA questions) into textual feature vectors. Significance: Enables the model to "understand" textual semantics (e.g., distinguishing between "generating captions" and "answering questions"), serving as the "textual foundation" for multimodal fusion. 3. Feature Alignment Layer: Projection W. Projection W: Function: Aligns the dimensions of visual features (from the Visual Encoder) and textual features (from the Text Embedding module). Significance: This solves the problem of "image feature dimension ≠ text feature dimension" so that the two can be fused and input into a large language model (for example, both are converted into 1024-dimensional vectors).4. Reasoning Core: Pre-trained Large Language Model: Function: Fusion of visual and textual features to perform multi-task reasoning (such as caption generation and question answering). Input: Fusion features (visual and textual) aligned by the ProjectionW layer. Output: Multi-task results (such as text descriptions generated from image captions and answers to video-based question answers). 5. Stage Objective: Multi-task Warm-up → Allow the model to learn foundational capabilities in image understanding, text generation, and image-text reasoning in general scenarios (natural image and text tasks). Value: This will lay the foundation for subsequent adaptation to medical scenarios. First, general multimodal logic will be mastered, and then applied to the medical field (as shown in the medical tasks in Figure 6). This approach, in other words, transfers capabilities to medical tasks (such as disease localization and report generation). This exemplifies the "general first, specialized later" approach to multimodal model training.
[0046] During the biomedical information adaptation phase, the MedVision-MT dataset was fed into the pre-trained model to train the visual encoder module and adjust its parameters to better extract medical image features (such as disease features and organ characteristics). Simultaneously, a multi-attribute expert hint generation method was initiated. Based on the training set, pathological categories (such as pneumonia and fracture) were selected as medical entities. Four attributes were designed for each entity: [SHAPE] (e.g., the shape of the pneumonia shadow), [SIZE] (the size of the shadow), [LOCATION] (the location of the shadow), and [SYMPTOM] (the corresponding symptom). A VQA strategy was used to construct the query template "What is the shape of [pneumonia]?" for the model to generate a response. A Pre strategy was used to construct the sentence "The size of [pneumonia] is [MASK]," where [MASK] is the classification label, for the model to predict the [MASK] portion. Attribute values generated by both strategies were collected to form an attribute pool for each entity. Using a formula-based expert hinting process, the information in the attribute pool was converted into expert hints and fed into the model. This enabled the model to learn medical domain expertise and feature associations, improving the accuracy and effectiveness of handling medical tasks.
[0047] like Figure 6As shown, the goal of the Biomedical Information Adaptation stage (Stage-2: Biomedical Information Adaptation) is to enable the model to "learn" medical expertise and adapt to medical scenarios, a critical step in transitioning from a "generalist" to a "medical expert." New and enhanced modules include: Multi-Attribute Expert-Prompt Module: Core innovation: Addresses the difficulty of aligning medical terminology with general knowledge. By incorporating medical entity attributes (such as disease shape, location, and symptoms), it bridges the gap between general and medical semantics. Task instruction refinement: Clarifies medical task types, such as DL tasks for locating pneumonia; RRG tasks for generating imaging reports; and other task instructions such as VQA and CD, enabling the model to accurately understand medical needs. Training features: Unlocks Visual Encoder training, allowing it to learn medical image-specific features (such as lesion texture and organ morphology); and fine-tunes using the MedVision-MT medical dataset to enhance medical knowledge.
[0048] right Figure 6 Detailed description: Figure 6 The reasoning process architecture diagram of the multimodal medical large model in the "Biomedical Information Adaptation Stage (Stage-2)" shows how the model processes medical images and language instructions to complete multi-task reasoning (such as disease localization and report generation).
[0049] Specifically: 1. Overview of core processes and modules Model input: medical images (Visual Inputs) + language instructions (Language Instruction) Core modules: Visual Encoder, Projection W, Text Embedding, Multi-Attribute Expert-Prompt Module, Pre-trained Large Language Model; Output: multi-task results (Chat-MedGen Response, including answers to DL, RRG and other tasks). 2. Detailed explanation of each element 1. Input layer: Visual Inputs + Language Instruction Visual Inputs (visual input): refers to medical images (such as the chest X-ray image in the figure), which are the "visual materials" analyzed by the model and represent real clinical data. Language Instruction (language instruction): natural language instructions for different medical tasks, corresponding to 5 types of core tasks (Figure 1- Figure 4Task system: Task 1 - DL (Disease Localization): Instruction: Require the model to locate the location of "pneumonia" in the image. Task 2 - RRG (Radiology Report Generation): Instruction: Require the model to generate a diagnostic report corresponding to the medical image. M-VQA (Medical Visual Question Answering): Example instruction: Require the model to answer questions related to medical images. CD (Clinical Diagnosis): Example instruction: Require the model to diagnose a disease based on the image. LCD (Local Clinical Diagnosis): Example instruction: Require the model to diagnose the "bounding box annotated region" in the image. 2. Feature Encoding Layer: Visual Encoder + TextEmbedding + Projection W Visual Encoder: Function: Extracts visual features from medical images (outputs a sequence, representing a visual feature vector). Purpose: Converts "medical images" into numerical features that the model can understand, capturing information such as lesions and organ morphology. Text Embedding: Function: Converts text such as "language instructions" and "expert tips" into text feature vectors (output sequences). Significance: Enables the model to understand the semantics of natural language instructions, for example, distinguishing between the tasks of "locating pneumonia" and "generating a report." Projection W (Projection Layer): Function: Aligns the dimensions of visual features (from the Visual Encoder) and text features (from Text Embedding), enabling their fusion. Significance: Addresses the issue of "image feature dimensionality ≠ text feature dimensionality" and enables multimodal information fusion. 3. Knowledge Enhancement Layer: Multi-Attribute Expert-Prompt Module: Function: Integrates the "expert tips" generated in Figure 7 to supplement the model with attribute knowledge of medical entities (such as disease shape, location, and symptoms). Significance: Bridges the gap between "general semantics" and "medical-specific semantics," enabling the model to understand medical details in images (for example, distinguishing the "shape of a pneumonia shadow" from ordinary shadows). 4. Inference Core: Pre-trained Large Language Model: Function: Fusion of "visual features, text features, and expert tips" to perform multi-task reasoning.Input: Visual features (sequence, processed by the Projection W layer), textual features (sequence, from Text Embedding), and expert hints (implied in the multi-attribute expert hint module to assist with semantic understanding); Output: Multi-task results (Chat-MedGen Response), corresponding to different tasks: DL task: Outputs the disease location (e.g., {60, 58, 70, 76}, representing the bounding box coordinates). RRG task: Outputs a medical report, such as "normal heart size, clear lungs, no focal consolidation, no pleural effusion." 5. Output layer: The multi-task result is the model's final response to the medical task, outputting the multi-task result: DL result: {60, 58, 70, 76} → disease location coordinates (e.g., pneumonia). RRG result: Normal heart size → medical imaging diagnosis report. Closed-loop logic: The output can be fed back for model optimization or directly used as "AI-assisted diagnosis recommendations" to support clinical decision-making. Figure 6 This comprehensive demonstration demonstrates the "practical workflow" for a multimodal medical large model, a solution involved in this invention: 1. Input medical images and task instructions → 2. Encode image and text features → 3. Inject medical attribute knowledge (expert prompts) → 4. Large model inference → 5. Output multi-task results (localization, reporting, etc.). Through "multimodal feature fusion + medical knowledge enhancement," the general large model is adapted to medical scenarios, solving real-world clinical tasks such as imaging diagnosis and report generation. This is the key step in transitioning models from general to medically specific.
[0050] i1 to i n : Visual feature sequence (from Visual Encoder); e1 to e n : Text embedding feature sequence (from the text embedding module Text Embedding); t1 to t n : The final fused multimodal feature sequence (input into the large language model).
[0051] 3. Implementation of Multi-task End-to-End Multimodal Medical Big Model Basic Modules and Network Architecture: Improve the basic modules based on TransformerBlock, adjusting parameters such as the attention mechanism and feedforward network to meet the needs of medical image and text processing. Build the network architecture, sequentially connecting the image encoding module (using convolutional neural networks, etc. to encode images), the vocabulary embedding module (using pre-trained word vectors to encode text), the linear projection layer module (mapping image and text features into the same dimensional space), the large language model module (e.g., fine-tuning based on an open-source large language model), and the multi-attribute expert prompt module (integrating attribute generation and prompt construction functions). Ensure communication between modules and collaborate to complete medical visual language processing tasks.
[0052] like Figure 7 As shown, the multi-attribute expert prompt module focuses on the "precision feeding" of medical knowledge. The core logic of the mechanism is to supplement the model with "attribute details" of medical entities, enabling the model to understand the mapping from "medical images to professional knowledge" and addressing the challenge of medical semantic alignment. Process breakdown: Medical entity and attribute definition: Pathology categories (such as pneumonia and fracture) are selected as medical entities. Four attributes are designed for each entity: [SHAPE], [SIZE], and [LOCATION]: These describe the "low-level visual features" of the medical entity in the image; [SYMPTOM]: This links "high-level clinical knowledge" (e.g., the image shape of pneumonia to symptoms such as cough and fever). Attribute generation strategies: A large language model (such as GPT4) is used to generate attribute values. Two approaches are available: The VQA strategy uses a template question, "What is the [attribute] of [medical entity]?" (e.g., "What is the shape of pneumonia?"), allowing the model to generate open-ended responses and preserve attribute diversity. The Pre-answer strategy uses a fill-in-the-blank template, "The [attribute] of [medical entity] is [MASK]" (e.g., "The size of pneumonia is [MASK]"), allowing the model to predict the masked content and enhance attribute accuracy. Attribute pool and expert hint construction: The attributes generated by the two strategies (each entity produces eight attribute values) are stored in an "attribute pool." Template filling is then used to combine the attributes with medical entities to generate expert hints, which are then fed into the large language model, allowing the model to learn the connection between "medical image → attribute → knowledge."
[0053] right Figure 7Detailed description: 1. Input and core object input: medical images (medical images [IMG] are represented by dotted boxes and are image identifiers), and medical entities (which can be understood as abstract labels corresponding to diseases, organs, etc., such as pneumonia and fractures). Core goal: For each medical entity, generate multi-dimensional attribute knowledge including shape (SHAPE), size (SIZE), location (LOCATION), and symptoms (SYMPTOM), build "expert tips" to feed the model, and bridge the gap between general semantics and medical professional semantics. 2. Two strategies for knowledge mining (VQA+Pre) The module uses GPT4 as a "knowledge engine" and mines four types of attributes (shape, size, location, symptoms) for medical entities through two strategies: visual question answering (VQA) and prediction (Pre): (1) Visual question answering (VQA) strategy (upper half of the process) Question template: Target, design 4 questions, covering 4 types of attributes (CLS classification tags): Q A1 : Ask shape (SHAPE); Q A2 :Ask for size (SIZE); Q A3 :Location (LOCATION);Q A4 : Ask symptoms (SYMPTOM); Knowledge output: GPT4 answers these questions, generates 4 VQA attribute results, stores them in AttributePool (attribute pool), and marks them as: A1 VQA shape, A2 VQA Size, A3 VQA Position, A4 VQA Symptoms; (2) Prediction (Pre) Strategy (the second half of the process) Fill-in-the-blank template: Design 4 sentences containing [MASK], also covering 4 types of attributes: S A1 : Fill shape; S A2 :Fill in size; S A3 :Fill in the position; S A4 :Fill in symptoms; Knowledge output: GPT4 predicts the [MASK] content, generates 4 Pre attribute results, and also stores the attribute pool, marked as: A1 Pre shape, A2 Pre Size, A3 Pre Position, A4 Pre Symptoms; 3. Expert Tip Generation Sampling (X): Randomly select one (i.e., X) from the 8 results (4 VQA + 4 Pre) in the attribute pool using a uniform distribution. Template Filling: Fill the selected X and medical entity into the preset template to generate an expert tip. Example logic: If X is A1 VQA (For example, to answer “the shape is a flake shadow”), the template might be “the shape of the medical entity is [X content], please analyze the image based on this”, ultimately forming a prompt text that allows the large model to understand the medical attributes.
[0054] Model training and testing: The constructed model is trained using the processed MedVision-MT dataset, adjusting module parameters and optimizing loss functions (such as cross-entropy loss). After training, test tasks are designed, such as inputting medical images to test the model's performance in tasks such as medical question answering, report generation, and disease localization and diagnosis. The model's multi-task processing capabilities, accuracy, and generalization capabilities in the medical field are evaluated, and the model is further optimized based on the test results.
[0055] Through the three steps of "general warm-up → medical adaptation → attribute enhancement", the entire architecture enables the multimodal model to grow from "being able to handle general graphic and text tasks" to a dedicated large model that "understands medical expertise and can solve complex medical visual language tasks". It covers scenarios such as report generation, disease location, diagnostic question and answer, and provides technical support for medical AI-assisted diagnosis.
[0056] Finally, it should be noted that the present invention is not limited to the above embodiments and may be subject to many variations. All variations that can be directly derived or imagined by a person skilled in the art from the disclosure of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A multimodal large-scale model system capable of simultaneously processing multiple medical visual language tasks, characterized by: The multimodal large model system includes a multi-task end-to-end multimodal medical large model, which includes a basic module and a network structure. The network structure includes a communication-connected image encoding module, a vocabulary embedding module, a linear projection layer module, a large language model module, and a multi-attribute expert prompt module. Among them, the image encoding module supports image modality input and uses the image encoding module to encode the features of the input image; the vocabulary embedding module supports text modality input and uses a vocabulary embedding module to encode the features of the input text; the linear projection layer module aligns the image modality features with the text modality features; The large language model module supports fluent text output, using the large language model to decode the input text and image features to obtain the final text output; the multi-attribute expert prompt module selects all pathology categories as medical entities based on the data tuples of the MedVision-MT training set to generate expert knowledge.
2. The multimodal large model system capable of simultaneously processing multiple medical visual language tasks according to claim 1, characterized in that: Four attributes are designed for each medical entity, including shape [SHAPE], size [SIZE], location [LOCATION] and symptom [SYMPTOM], which are represented by A1, A2, A3 and A4 respectively. Among them, the three attributes of shape [SHAPE], size [SIZE] and location [LOCATION] represent the low-level information of each medical entity in the corresponding image, while the last attribute symptom [SYMPTOM] represents the high-level information.
3. The multimodal large model system capable of simultaneously processing multiple medical visual language tasks according to claim 2, characterized in that: In order to obtain expert-level hints, a large language model is used to generate the key attributes of each medical entity. Two generation strategies are adopted: Visual Question Answering (VQA) and Prediction (Pre). In the VQA scheme, the query template is designed as: "[Medical Entity] [A i ] is”; while in the Pre scheme, the sentence to be predicted is: “[Medical entity]’s [A i ] is [MASK]", wherein the A i is any one of the four attributes; the response and [MASK] part of the VQA query represent the key attribute values obtained by these two generation schemes respectively.
4. The multimodal large model system capable of simultaneously processing multiple medical visual language tasks according to claim 3, characterized in that: Applying the two generation strategies to each object results in an attribute pool containing eight different attribute values for each entity. The resulting expert suggestion process is summarized by the following formula: ; In the formula, VQA represents the question-answering strategy; χ represents the input image; Q represents the query question; Pre represents the prediction strategy; S AN Indicates sentences corresponding to different predictions; Q A1 To Q AN : Query questions designed based on medical entity attributes are used to obtain attribute information of entities in medical images through VQA strategies. N is the number of attribute-related questions. Pre represents the prediction strategy, which allows the model to predict the mask part in the text related to medical entity attributes to mine attribute knowledge. S A1 To S AN Represents sentences corresponding to different predictions, where N is the number of attribute-related texts; ExpertPrompt: The final generated expert prompt, used as input to the large language model to assist the model in understanding the expertise and features in medical images; Template n : Template for building expert tips, n can represent the number or type of template; X is the sample data; Object k : Medical entities, namely pathological categories, organs and other objects involved in medical images, k is used to distinguish different medical entities.
5. A design method for a large multimodal model capable of processing multiple medical visual language tasks simultaneously, characterized by The design method includes the following steps: Step 1): A unified dataset covering various medical vision tasks designed for multimodal large model system training, specifically including: Step 1.1) New medical vision task setting: Collect and reorganize popular medical vision datasets to form the benchmark dataset MedVision-MT. The benchmark dataset MedVision-MT includes five core tasks: medical question answering Med-VQA, radiology report generation RRG, disease localization DL, clinical diagnosis CD, and local clinical diagnosis LCD; Disease localization (DL) aims to determine the location of the disease; radiology report generation (RRG) aims to generate medical reports based on images; CD refers to diagnosing diseases based on medical images; Med-VQA focuses on answering questions based on images; local clinical diagnosis (LCD) is the inverse process of disease localization (DL), which classifies the disease after a given bounding box is obtained. Through this comprehensive and diverse multi-task setting, the benchmark dataset MedVision-MT can provide a foundation for further exploration in the field of medical MLLM. Step 1.2) Multi-task medical image dataset: Several popular medical multimodal datasets are reorganized to form a multi-task benchmark dataset MedVision-MT; Step 2): Phased adaptation training; specifically including: Step 2.1) Multi-task warm-up training phase: First, the model is pre-trained on a variety of large-scale natural image datasets. Then, the model's multi-task processing capabilities are transferred from the general domain to the medical domain. During the multi-task warm-up phase, the projection matrix W is pre-trained, and the large language model is adjusted on various natural image datasets as needed to ensure that the pre-trained model can handle different tasks under human queries. Step 2.2) Biomedical information adaptation phase: After completing the multi-task warm-up, the model’s ability to handle different visual tasks in the general field is initially established; then, the multi-task benchmark dataset MedVision-MT is used to adapt the model to the medical field.
6. The method for designing a large multimodal model capable of simultaneously processing multiple medical visual language tasks according to claim 5, characterized in that: Step 2.1) specifically includes: replacing biomedical tasks with five natural vision tasks, among which image caption generation replaces radiology report generation RRG, visual question answering replaces medical question answering Med-VQA, localization replaces disease localization DL, classification replaces clinical diagnosis CD, and object classification replaces local clinical diagnosis LCD; gradually transferring multi-task capabilities from the natural domain to the medical domain, and establishing semantic connections between common vocabulary and complex biomedical terms through expert prompts to achieve more effective capture of medical image features.
7. The method for designing a large multimodal model capable of simultaneously processing multiple medical visual language tasks according to claim 5, characterized in that: Step 1.2) The sub-datasets collected by Traditional Chinese Medicine Question Answering Med-VQA include VQA-Med-2019, VQA-Med-2021, SLAKE, VQA-RAD, and Path-VQA; the sub-datasets collected by Radiology Report Generation RRG include MIMIC-CXR and IU-X-Ray datasets; the sub-datasets collected by Clinical Diagnosis CD include NIH datasets; the sub-datasets collected by Disease Localization DL and Local Clinical Diagnosis LCD include ChestXray and MS-CXR datasets.
8. The method for designing a multimodal large model capable of simultaneously processing multiple medical visual language tasks according to claim 7, characterized in that: In order to effectively filter out meaningful and high-quality data from the sub-dataset, we further filtered the data for both image and text modalities, retaining only appropriate image-text pairs to form the operational dataset. ,in represents an image-text pair, and Represents images and text respectively.
9. The method for designing a multimodal large model capable of simultaneously processing multiple medical visual language tasks according to claim 7 or 8, characterized in that: The distribution proportions of the sub-datasets ChestXray, MS-CXR, VQA-RAD, VQA-Med-2021, SLAKE, IU-X-Ray, VQA-Med-2019, NIH, Path-VQA, and MIMIC-CXR are 5.5%, 0.88%, 1.3%, 2.2%, 5.0%, 6.6%, 7.3%, 16.3%, 26.1%, and 28.8%, respectively.
Citation Information
Patent Citations
Question and answer method and device, storage medium and computer equipment
CN117235238A
Medical image automatic interpretation system based on large language model
CN118039086A
Method and system for enhancing reasoning ability of large language model in material field
CN118504682A
Medical visual question and answer method and system based on multi-task modeling
CN119202334A
Visual language feature fine alignment method for medical multi-mode large model
CN119357443A