Auxiliary diagnosis and treatment system based on large-scale language model
Through an auxiliary diagnosis and treatment system based on large-scale language models, natural language input is automatically parsed and corresponding tool components are called, which solves the complex operation of existing medical image analysis tools, and achieves a more efficient and intuitive diagnostic process.
Patent Information
- Application Number
- CN202411974160.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The existing medical image analysis tools are complex in operation, not intuitive enough, and have a high threshold for use, which makes it time-consuming for doctors to be omitted or misdiagnosed during the diagnosis process.
Using an auxiliary diagnosis and treatment system based on large-scale language models, the natural language is parsed into specific task prompts through medical imaging agents, complex problems are automatically decomposed, and the calling tool components are matched, such as report generation module, image segmentation module, lesion detection module, etc., to generate detailed step-by-step solutions.
It simplifies the operation process, lowers the threshold for use, improves diagnostic efficiency, reduces redundant operations, and significantly improves the system's execution speed and resource utilization efficiency.
Smart Images

Figure CN119943338A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image analysis, and in particular relates to an auxiliary diagnosis and treatment system based on a large-scale language model. Background Art
[0002] With the advancement of modern medicine, medical imaging plays an increasingly important role in clinical diagnosis, and imaging examinations have become a key link in disease diagnosis and treatment. However, the rapid growth of imaging data and the complexity of image processing have made doctors face huge challenges in their actual work. Doctors not only need to analyze a large number of images, but also need to write detailed diagnostic reports. This process is not only time-consuming, but also prone to omissions or misdiagnosis due to excessive workload.
[0003] The field of medical image analysis has adopted a large number of technologies based on artificial intelligence and deep learning to assist doctors in diagnosis. These technologies are mainly concentrated in image segmentation models, lesion detection models, and diagnostic report generation models.
[0004] However, the existing medical image analysis tools mostly run segmentation, detection and report generation steps independently, and in actual use, tools need to be switched manually. They also rely on complex user interfaces and command input, making the operation complex and not intuitive enough, and the threshold for use is high.
[0005] To this end, we propose an auxiliary diagnosis and treatment system based on a large-scale language model to solve the above problems. Summary of the invention
[0006] The purpose of the present invention is to solve the problems in the prior art that the operation is relatively complicated and not intuitive enough, and the usage threshold is high, and an auxiliary diagnosis and treatment system based on a large-scale language model is proposed.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] An auxiliary diagnosis and treatment system based on a large-scale language model, comprising a medical imaging agent based on a large-scale pre-trained language model, a tool component, and a storage pool module, wherein the medical imaging agent converts natural language analysis into specific task prompts, and matches and calls the tool component;
[0009] The tool component includes a report generation module, an image segmentation module, a lesion detection module, a report evaluation module and a diagnosis and treatment assistant module;
[0010] A report generation module generates medical image reports based on medical images, task prompts, and output results of tool components;
[0011] Image segmentation module, which performs boundary recognition and contour annotation on organs, tissues and lesion areas in medical images;
[0012] Lesion detection module, identifying and locating lesion areas in medical images;
[0013] Report evaluation module, which conducts in-depth analysis and self-consistency evaluation of medical imaging reports;
[0014] The diagnosis and treatment assistant module generates personalized basic diagnosis and treatment plans based on medical imaging reports;
[0015] The storage pool module is responsible for the storage and management of multimodal data and automatically saves the output of tool components for subsequent calls.
[0016] Preferably, the report generation module generates structured medical imaging reports, non-structured text dialogues and generation records based on medical images and task prompts through multimodal data fusion based on a multimodal large language model.
[0017] Preferably, the image segmentation module is based on the open source Segment Anything Model, introduces a text feature encoder, and outputs a segmentation result (Mask) based on the medical image and the medical image report.
[0018] Preferably, the lesion detection module is based on the open source Grounding DINO model, performs further model training according to the particularity of medical images and the requirements of lesion detection tasks, and outputs the detection frame of the lesion according to the medical images and medical image reports.
[0019] Preferably, the report evaluation module:
[0020] According to the medical imaging report, conduct self-consistency evaluation of disease judgment and output a self-consistency evaluation report;
[0021] According to the generation records of medical imaging reports, including token confidence information and task logs, the medical imaging reports are evaluated and the scores and rankings are output.
[0022] Preferably, the diagnosis and treatment assistant module generates a personalized basic diagnosis and treatment plan based on the large language model and the medical imaging report.
[0023] Preferably, the storage pool module is responsible for the storage and management of multimodal data, and automatically saves the output of the tool component for subsequent tasks or tool components to call.
[0024] To sum up, the technical effects and advantages of the present invention are as follows: compared with existing devices, the auxiliary diagnosis and treatment system based on large-scale language models converts the natural language analysis of user input into specific task prompts through medical imaging intelligent agents, automatically decomposes complex medical imaging problems, and matches and calls tool components to generate detailed step-by-step solutions. Compared with existing devices, it avoids the problems of complex and less intuitive operations, simplifies operations, and lowers the threshold for use.
[0025] In addition, the output results of other modules of the tool component are automatically saved through the storage pool module for flexible use in subsequent tasks or tool calls, avoiding redundant calculations and significantly improving the execution speed and resource utilization efficiency of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the structure of the present invention. DETAILED DESCRIPTION
[0027] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0028] Reference Figure 1 , an auxiliary diagnosis and treatment system based on a large-scale language model, including a medical imaging agent based on a large-scale pre-trained language model, a tool component and a storage pool module. The medical imaging agent (hereinafter referred to as the agent) converts natural language analysis into specific task prompts, and matches and calls tool components. The agent understands the doctor's input needs through natural language processing technology (Natural Language Processing, NLP). After receiving the input, the agent can automatically parse the user's intention and decompose complex diagnostic problems into multiple specific subtasks. For example, the doctor inputs "Please help me analyze the lung nodules in this CT image and generate a report", the system can understand the need to generate a diagnostic report, and output the lesion segmentation of the lung nodules.
[0029] The tool components include a report generation module, an image segmentation module, a lesion detection module, a report evaluation module, and a diagnosis and treatment assistant module;
[0030] A report generation module generates medical image reports based on medical images, task prompts, and output results of tool components;
[0031] Image segmentation module, which performs boundary recognition and contour annotation on organs, tissues and lesion areas in medical images;
[0032] Lesion detection module, identifying and locating lesion areas in medical images;
[0033] Report evaluation module, which conducts in-depth analysis and self-consistency evaluation of medical imaging reports;
[0034] The diagnosis and treatment assistant module generates personalized basic diagnosis and treatment plans based on medical imaging reports;
[0035] The storage pool module is responsible for the storage and management of multimodal data and automatically saves the output of tool components for subsequent calls.
[0036] This auxiliary diagnosis and treatment system based on a large-scale language model uses medical imaging agents to parse natural language and generate various task prompts (Prompts) to activate and coordinate tool components. These task prompts not only indicate the type of task, but also provide the necessary contextual information and execution instructions for the tool components. In the process of understanding user input, the agent generates a series of specific tags, which trigger the operation of different tools such as image segmentation, lesion detection, and report generation according to task requirements.
[0037] For example, when a doctor inputs "Please analyze the chest problems in this image in detail", the intelligent body will perform demand analysis through the big model and generate the most appropriate task requirements at this time:
[0038] Instruction 1:
[0039] <img01>{Report Generation}[vqa][Please describe what chest problems are present in this medical image.]
[0040] When the above Instruction 1 is generated by the agent, the system will automatically trigger the corresponding tool to execute. Through text parsing, the agent recognizes that the target image of the instruction is <img01>, the tool called is the report generation model (Report Generation), and the task prompt (Prompt) is "What chest problems exist in this medical image? Please explain." After the tool is executed, the report generation model returns feedback:
[0041] Feedback content:
[0042] "The left costophrenic angle becomes blunt, indicating the presence of pleural effusion."
[0043] The agent then re-enters the feedback information, the original input question and Instruction 1 into the language model to obtain the updated instructions:
[0044] Instruction 2:
[0045] <img01>{Detection}[Left costophrenic angle blunting]
[0046] After Instruction 2 is generated, the system will automatically call the corresponding lesion detection tool (Detection) according to the new task requirements. By parsing the instruction, the agent determines that the target image of this task is still <img01>, the tool is a lesion detection tool, and the prompt is "the left costophrenic angle becomes blunt." After execution, the lesion detection tool feedbacks the detection box and provides visual output, while generating a text description:
[0047] Test results:
[0048] "The left costophrenic angle is blunted {[x1,y1,x2,y2]}."
[0049] At this point, the agent inputs the above detection results together with the initial question and the questions and answers at each step into the language model. If the agent determines that the task has not been fully completed, it will continue to generate subsequent instructions, such as Instruction 3, Instruction 4, etc., and further call other tools to refine and supplement the task. If the agent determines that the task has been completed, it will output <done>Tag, indicating the end of this task.
[0050] The report generation module, based on a multimodal large language model, integrates functions such as visual question answering (VQA), disease recognition, and multi-round interactive dialogue, becoming the text generation engine of the entire system. Through multimodal data fusion, based on medical images and medical image reports, it generates structured medical image reports (containing standardized fields such as lesion location, size, type, and diagnosis results for easy storage and retrieval), non-institutional text dialogues (used to describe examination findings, diagnostic recommendations, and supplementary instructions, providing detailed explanations and text communication records), and generation records (including logs during report generation, model decision information, and task execution details to support traceability and review).
[0051] The main functions are:
[0052] Medical imaging report generation: Automatically generate structured medical imaging reports, including core contents such as imaging analysis results, diagnostic conclusions and lesion descriptions.
[0053] Visual Question Answering (VQA): Based on the image content, it answers the visual questions raised by the doctor and provides intuitive image interpretation.
[0054] Disease identification and labeling: Based on model reasoning and task module output, the disease type is automatically identified and relevant feature information is labeled.
[0055] Multi-round dialogue: Supports multiple rounds of interactive dialogue with doctors, real-time adjustment and improvement of report content to meet personalized diagnosis needs.
[0056] The image segmentation module is deeply customized based on the open source SAM (Segment Anything Model). The main improvements include the introduction of a text feature encoder to encode the text description in the medical imaging report and jointly model it with the image features. This multimodal training method enables the module to generate segmentation results for the corresponding anatomical parts or lesion areas based on the specific descriptions in the report. According to the text descriptions of medical images (such as X-rays, CT, and MRI images) and medical imaging reports, the segmentation results (Mask) are output, and the contours of the anatomical structures or lesion areas in the medical images are annotated, providing accurate image analysis results and data support for subsequent tasks such as lesion detection and report generation, ensuring the accuracy and efficiency of medical image processing.
[0057] The main functions are:
[0058] Anatomical structure segmentation: Automatically label important anatomical parts (such as heart, lungs, etc.) in medical images.
[0059] Lesion detail segmentation: Accurately identify and segment lesion areas (such as tumors, inflammatory lesions, etc.) and generate segmentation contours that meet medical standards.
[0060] The lesion detection module is optimized based on the open source Grounding DINO model. Further model training is carried out to meet the specific needs of medical images and the requirements of lesion detection tasks, so that it can achieve high-precision target detection and positioning in complex medical images. At the same time, the model also combines the text description information in the medical report to enhance the adaptability and detection accuracy of clinical scenarios, and provide key structured data support for subsequent tasks such as report generation and diagnosis and treatment assistance. The introduction of the lesion detection module has greatly improved the automation level and clinical application value of medical image analysis.
[0061] The main functions are:
[0062] Lesion area detection: Automatically identify potential lesions in medical images and mark the location and range of abnormal areas.
[0063] Multi-lesion localization: In the presence of multiple lesions, the model can detect and mark multiple areas at the same time and generate multiple detection boxes.
[0064] Report evaluation module: Based on the medical imaging report, the self-consistency evaluation of disease judgment is carried out, and the self-consistency evaluation report is output; based on the generation record of the medical imaging report, including token confidence information and task log, the medical imaging report is evaluated, and the score and ranking are output. The consistency, completeness and accuracy of the content of the medical imaging report are checked in multiple dimensions to ensure the reliability and practicality of the medical imaging report in clinical applications.
[0065] The main functions are:
[0066] Self-consistency assessment: Automatically analyze the diagnostic conclusions, lesion descriptions and related imaging data in the report to check whether the report is logically coherent and whether there are any over-reporting (redundancy) or under-reporting (omission) problems.
[0067] Disease coverage check: Evaluate the comprehensiveness of the report on identified lesions and potential diseases to ensure that key diseases are not overlooked.
[0068] Report scoring and ranking: Based on the completeness and logical consistency of the content, the report quality is comprehensively evaluated, and scores and rankings are generated to facilitate doctors' rapid screening and review.
[0069] Generation record analysis: Combined with report generation records (such as token confidence distribution, model reasoning log, etc.), the reliability and interpretability of model generation are evaluated.
[0070] The diagnosis and treatment assistant module, based on a large language model, generates personalized basic diagnosis and treatment plans according to medical imaging reports, provides clinical advice, and assists doctors and patients in making preliminary diagnosis and treatment plans.
[0071] The main functions are:
[0072] Generation of personalized diagnosis and treatment plans: Automatically generate personalized diagnosis and treatment recommendations based on the diagnostic results and lesion descriptions in the imaging report.
[0073] Multiple treatment options recommendation and explanation: Provide multiple feasible treatment options and make recommendations based on case characteristics and current clinical guidelines.
[0074] Disease management and follow-up recommendations: Generate disease management plans that include long-term management and follow-up recommendations for chronic and potentially recurrent diseases.
[0075] Risk assessment and early warning: High-risk factors identified in the report are marked and early warning information is highlighted in the diagnosis and treatment plan to facilitate doctors' decision-making.
[0076] The storage pool module adopts a dynamic data management strategy to support efficient storage and retrieval of different types of data, ensuring the smooth operation of the medical image analysis process. The storage pool module avoids redundant operations by automatically storing and updating the input and output of multiple rounds of tasks, significantly improving the execution speed and resource utilization efficiency of the overall system.
[0077] The main functions are:
[0078] Intermediate result storage, that is, automatic storage of segmentation (Mask) and lesion detection boxes, for subsequent diagnostic tool calls, not only saves storage space, but also reduces repeated tool calls.
[0079] Structured and unstructured data storage, that is, supporting the coexistence of structured medical imaging reports, unstructured interaction records, and conversation data. The support of different data formats ensures the data compatibility and flexibility of the system in various task scenarios;
[0080] Tool component result storage means storing the output results of tool components for future tasks. During the execution of tool components, the system will check whether there is reusable data in the storage pool module. If relevant data is found, the system will call it directly to avoid redundant calculations; otherwise, the system will store the results after the task is completed for future tasks.
[0081] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.< / done>
Claims
1. An auxiliary diagnosis and treatment system based on a large-scale language model, characterized in that: It includes a medical imaging agent, a tool component, and a storage pool module based on a large-scale pre-trained language model. The medical imaging agent converts natural language analysis into specific task prompts, and matches and calls tool components; The tool component includes a report generation module, an image segmentation module, a lesion detection module, a report evaluation module and a diagnosis and treatment assistant module; The report generation module generates a medical image report based on the medical image, task prompts and output results of the tool component; The image segmentation module performs boundary recognition and contour marking on organs, tissues and lesion areas in medical images; The lesion detection module identifies and locates the lesion area in the medical image; The report evaluation module performs in-depth analysis and self-consistency evaluation on the medical imaging report; The diagnosis and treatment assistant module generates a personalized basic diagnosis and treatment plan based on the medical imaging report; The storage pool module is responsible for the storage and management of multimodal data and automatically saves the output of the tool component for subsequent calls.
2. The auxiliary diagnosis and treatment system based on a large-scale language model according to claim 1, characterized in that: The report generation module generates structured medical imaging reports, non-institutionalized text dialogues, and generation records based on a multimodal large language model and multimodal data fusion according to medical images and task prompts.
3. The auxiliary diagnosis and treatment system based on a large-scale language model according to claim 2 is characterized in that: The image segmentation module is based on the open source Segment Anything Model, introduces a text feature encoder, and outputs segmentation results based on medical images and medical image reports.
4. The auxiliary diagnosis and treatment system based on a large-scale language model according to claim 2, characterized in that: The lesion detection module is based on the open source Grounding DINO model, and further performs model training according to the particularity of medical images and the requirements of lesion detection tasks, and outputs the detection frame of the lesion according to the medical images and medical image reports.
5. The auxiliary diagnosis and treatment system based on a large-scale language model according to claim 2, characterized in that: The report evaluation module: According to the medical imaging report, conduct self-consistency evaluation of disease judgment and output a self-consistency evaluation report; According to the generation records of medical imaging reports, including token confidence information and task logs, the medical imaging reports are evaluated and the scores and rankings are output.
6. The auxiliary diagnosis and treatment system based on a large-scale language model according to claim 2, characterized in that: The diagnosis and treatment assistant module generates personalized basic diagnosis and treatment plans based on the large language model and medical imaging reports.
7. The auxiliary diagnosis and treatment system based on a large-scale language model according to claim 1, characterized in that: The storage pool module is responsible for the storage and management of multimodal data, and automatically saves the output of the tool component for subsequent tasks or tool components to call.
Citation Information
Cited By
Method and device for establishing diagnosis and treatment system of digestive system disease multi-modal information
CN120998466A
Method and device for establishing a multi-modal information diagnosis and treatment system for digestive system diseases
CN120998466B