Multimodal medical imaging report generation method and system based on medical thinking chain

By simulating the human cognitive process and reconstructing the chain reasoning question-answer pairs of large language models, the diversity and accuracy problems of medical imaging report generation in existing technologies are solved, high-quality medical imaging report generation is achieved, misdiagnosis and missed diagnosis are reduced, and the diagnostic ability of the model is enhanced.

CN119418846BActive Publication Date: 2025-10-03FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411337098.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-10-03
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

Existing medical imaging report generation methods find it difficult to generate reports containing multiple heterogeneous information under a unified framework, cannot accurately locate abnormal areas in images and generate correct descriptions, lack diversity and naturalness, and large language models are prone to hallucinations, leading to misdiagnosis or missed diagnosis.

Method used

By simulating the human cognitive process, a multimodal report generation system based on the medical thought chain is constructed. A large language model is used for end-to-end semantic segmentation and chain reasoning question-answer pair reconstruction, hierarchical medical attribute question-answer pairs are extracted, and pathology inference thought chains are embedded in the model fine-tuning process to generate structured medical imaging reports.

Benefits of technology

It significantly improves the accuracy and reliability of reports, reduces hallucination problems, and can generate comprehensive medical imaging diagnosis reports, focusing on the matching of local medical features and pathological inferences, thereby enhancing the generalization performance and diagnostic capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418846B_ABST
    Figure CN119418846B_ABST
Patent Text Reader

Abstract

The present invention relates to a multimodal medical imaging report generation method and system based on a medical thinking chain, wherein the method includes the following steps: obtaining an existing original unordered and unstructured medical imaging report dataset; preprocessing the imaging reports in the dataset to obtain a high-quality dataset; processing the medical imaging reports in the high-quality dataset based on a large language model to extract hierarchical medical attribute question-answer pairs; serially reconstructing the hierarchical medical mathematical question-answer pairs into chained reasoning question-answer pairs; using the chained reasoning question-answer pairs to fine-tune an existing medical imaging report generation model, and using the fine-tuned medical imaging report generation model to generate medical imaging reports. Compared with the existing technology, the present invention has the advantages of explicitly enhancing the utilization rate of multimodal features and the recognition ability of intent understanding, improving the generalization, reliability, and accuracy of reasoning capabilities, and generating more reliable medical imaging reports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical imaging report generation, and in particular to a multimodal medical imaging report generation method and system based on a medical thinking chain. Background Art

[0002] Medical image report generation is a key research area in the field of medical artificial intelligence. It aims to leverage computer technology to automatically generate text reports corresponding to medical images, thereby reducing physician workload and improving diagnostic efficiency and accuracy. With the advancement of medical imaging technology, particularly the widespread use of imaging devices such as X-rays, CT scans, and MRIs, physicians are faced with an increasing amount of imaging data to interpret and record. Traditionally, the process of writing image reports relies on the expertise and experience of radiologists. Manual report generation is not only time-consuming but also prone to human error. Therefore, medical image report generation typically involves extracting important visual information from medical images and automatically converting this information into natural language descriptions that meet clinical needs. The core of this task lies in enabling computers to accurately understand the content of medical images and express this information in a manner consistent with medical context, thereby assisting physicians in making diagnostic and treatment decisions. Therefore, this task requires careful consideration of image feature extraction and pathological inference to ensure that the generated report truly reflects the key medical information contained in the images.

[0003] In the field of medical imaging report generation, a large number of studies have attempted to improve the accuracy and quality of report generation through various technical means. Some studies have focused on attaching text tags to medical images. Most of these studies focus on generating fully structured or semi-structured text (such as labels or templates) rather than natural text. Although these studies have made progress in specific areas, the content of the reports they generate is limited by predefined topics and cannot meet the diverse needs of doctors in real-world medical scenarios. In addition, another part of the research focuses on the application of image description generation technology. These technologies typically use an encoder-decoder architecture to extract image features through a CNN and then generate text through a RNN. Shin et al. (2016) used a CNN-RNN framework to predict chest radiograph labels (such as location and severity). However, this approach often fails to meet the requirements of generating diagnostically accurate reports when processing complex medical data. In contrast, recent research has focused on generating natural text reports. Zhang et al. (2017) proposed a semi-structured pathology report generation method, while this study is the first to propose a model for generating truly natural language reports, which are typically long and cover multiple topics. The innovation of this study lies in the combination of visual features and semantic information. Through a multi-task learning framework, a joint attention mechanism, and a hierarchical LSTM model, it can more accurately locate abnormal areas in images and generate corresponding descriptions.

[0004] In recent years, researchers have begun to combine technologies such as multi-task learning, knowledge graphs, and contrastive learning to enhance the model's learning and application of clinical features. For example, knowledge graphs are widely used to inject prior knowledge in the medical field into the model to help better identify and describe abnormalities in images. Other studies have combined multi-task learning with disease classification tasks to improve the ability to distinguish features. Although these methods have improved the performance of report generation to a certain extent, due to the lack of pathological inference analysis between local image disease features and the overall condition, the model still has deficiencies in diagnostic accuracy and handling of disease imbalances. The generated medical reports are prone to errors such as mismatches between the overall condition and local medical features.

[0005] With the development of large language models (LLMs), researchers have begun to explore the application of LLMs to medical imaging report generation tasks. Traditional report generation methods usually rely on manually annotated labels or templates. Although they can achieve automation to a certain extent, they are difficult to cope with the complex information and diverse diagnostic needs in the reports. With the advancement of deep learning and natural language processing technologies, especially the emergence of large pre-trained models such as LLaMA2-7B, these models have demonstrated strong text generation capabilities and can generate high-quality reports with appropriate prompts. However, the task of generating medical imaging reports is much more complex than the task of generating general image descriptions. Reports usually need to cover multiple structured medical observations, including detailed descriptions of normal and abnormal conditions. Traditional medical impact report generation methods have difficulty in gradually analyzing and diagnosing the condition based on different levels of medical features in the image, and have difficulty in reasoning about the condition based on the detailed features in the image.

[0006] In summary, existing medical report generation methods have the following problems:

[0007] First, a complete diagnostic report typically contains a variety of heterogeneous information (such as impressions, findings, and labels), and generating this information under a unified framework is technically challenging. Existing methods often only generate partially structured content and struggle to capture all the details in a complex report.

[0008] Secondly, existing technologies are still not ideal for locating abnormal regions in images and generating accurate descriptions for them. While traditional visual attention mechanisms can locate abnormalities in subregions of an image, they often fail to accurately identify and describe these abnormalities due to a lack of sufficient semantic information. This makes it difficult to accurately match localized image abnormalities with pathological inferences and disease diagnoses. Especially for complex or rare diseases, the generated reports may lack sufficient clinical precision, which can lead to misdiagnosis or missed diagnoses.

[0009] Furthermore, traditional structured and semi-structured medical report template generation methods often rely on predefined templates or high-frequency vocabulary, which results in a lack of diversity and naturalness in the generated reports. While this template-based generation approach can improve efficiency to a certain extent, it also limits the personalization and richness of the report. Especially for diverse patients or complex cases, the report may appear overly mechanical and fail to fully reflect the doctor's actual diagnostic thinking.

[0010] Finally, in recent years, large language models (LLMs) have been widely used in medical image report generation tasks due to their powerful language capabilities. However, they also face various significant challenges, particularly the problem of model "hallucination." This problem occurs when the model generates seemingly reasonable but actually erroneous content during the generation process. This problem is particularly serious in medical scenarios, as erroneous reports can directly affect clinical diagnosis and treatment decisions, leading to potentially serious consequences, making it difficult for the model to be truly applied in real-world medical settings. Summary of the Invention

[0011] The purpose of the present invention is to provide a multimodal medical imaging report generation method and system based on the medical thought chain, a method for constructing fine-grained instruction pairs by simulating the human cognitive process, and applying the concept of "chain of thought" (CoT) from the reasoning scenario to the training scenario, constructing chain training data containing the potential meaning of disease reasoning, and helping the model better grasp how to analyze and evaluate medical images from details to the whole.

[0012] The purpose of the present invention can be achieved by the following technical solutions:

[0013] A multimodal medical imaging report generation method based on a medical thinking chain comprises the following steps:

[0014] Obtain existing raw, unordered, and unstructured medical imaging report datasets;

[0015] Preprocess the image reports in the dataset to obtain a high-quality dataset;

[0016] Process medical imaging reports in high-quality datasets based on a large language model to extract hierarchical medical attribute question-answer pairs;

[0017] Reconstruct hierarchical medical mathematics question-answer pairs into chained reasoning question-answer pairs;

[0018] The chained reasoning question-answering data is used to fine-tune the existing medical imaging report generation model, and the fine-tuned medical imaging report generation model is used to generate medical imaging reports.

[0019] The pretreatment comprises the following steps:

[0020] Regulations for deleting data from a dataset where the resolution of medical images is lower than a preset resolution threshold;

[0021] Deleting data provisions that lack a certain portion of text content or whose length is less than a certain length;

[0022] Medical imaging reports are protected from privacy and redundancy, and content related to doctor-patient information, duplicate medical records, and content related to the patient's historical medical conditions are deleted.

[0023] The hierarchical medical attribute question-answer pairs are divided into six dimensions, which are modality, organ, size, abnormal location, symptom, and overall health status from low to high levels.

[0024] The method for extracting hierarchical medical attribute question-answer pairs is specifically as follows: based on specific prompt words, a large language model is used to perform end-to-end semantic segmentation on the original disorganized medical image report, and questions corresponding to six different dimensions, namely modality, organ, size, abnormal location, symptoms, and overall health status, are extracted. Based on each question, the information of each dimension is independently queried and retrieved in the original image report as the answer to the question, thereby hierarchizing the original disordered and unstructured medical image report text into question-answer pairs of six dimensions.

[0025] The method of serially reconstructing hierarchical medical mathematics question-answer pairs into chained reasoning question-answer pairs is specifically as follows:

[0026] For each original question except the lowest level in the extracted question-answer pair, the answers to all lower-level questions with lower levels than the original question are prefixed and concatenated with the original question in series, the question is reconstructed in a chain manner, and combined with the answer corresponding to the original question to obtain a chained reasoning question-answer pair.

[0027] A multimodal medical imaging report generation system based on medical thinking chain, including:

[0028] Raw data acquisition module: used to obtain existing raw disordered and unstructured medical imaging report data sets;

[0029] Preprocessing module: used to preprocess the image reports in the dataset to obtain high-quality datasets;

[0030] Hierarchical Question-Answer Pair Extraction Module: This module processes medical imaging reports in high-quality datasets based on a large language model and extracts hierarchical medical attribute question-answer pairs.

[0031] Chained reasoning question-answer pair reconstruction module: used to concatenate hierarchical medical mathematics question-answer pairs into chained reasoning question-answer pairs;

[0032] Model fine-tuning and report generation module: used to fine-tune the existing medical imaging report generation model using chained reasoning question-answering data, and generate medical imaging reports using the fine-tuned medical imaging report generation model.

[0033] The pre-processing module performs the following steps:

[0034] Regulations for deleting data from a dataset where the resolution of medical images is lower than a preset resolution threshold;

[0035] Deleting data provisions that lack a certain portion of text content or whose length is less than a certain length;

[0036] Medical imaging reports are protected from privacy and redundancy, and content related to doctor-patient information, duplicate medical records, and content related to the patient's historical medical conditions are deleted.

[0037] The hierarchical medical attribute question-answer pairs are divided into six dimensions, which are modality, organ, size, abnormal location, symptom, and overall health status from low to high levels.

[0038] The hierarchical question-answer pair extraction module performs the following steps: based on specific prompt words, using a large language model to perform end-to-end semantic segmentation on the original disorganized medical image report, extracting questions corresponding to six different dimensions: modality, organ, size, abnormal location, symptoms, and overall health status, and independently querying and retrieving information of each dimension in the original image report as the answer to each question, thereby hierarchically transforming the original disorganized and unstructured medical image report text into question-answer pairs of six dimensions.

[0039] The chained reasoning question answering reconstruction module performs the following steps:

[0040] For each original question except the lowest level in the extracted question-answer pair, the answers to all lower-level questions with lower levels than the original question are prefixed and concatenated with the original question in series, the question is reconstructed in a chain manner, and combined with the answer corresponding to the original question to obtain a chained reasoning question-answer pair.

[0041] This system, specifically tailored for the medical imaging field, uses a method that simulates the cognitive processes of human physicians to construct fine-grained instruction pairs. It also applies the concept of thought chaining from reasoning scenarios to training scenarios, constructing a chained dataset with implicit pathological inference and condition analysis thought processes. This helps the model learn the thinking patterns of human physicians during training, further enhancing the model's reasoning capabilities and significantly reducing the occurrence of hallucinations. This method simulates the cognitive processes of human physicians when analyzing medical images to construct more refined and structured training data. Through hierarchical semantic segmentation, the present invention transforms complex, unstructured medical reports into hierarchical question-answer pairs. These question-answer pairs represent different levels of medical information, such as the imaging modality, involved organs, abnormality location, symptoms, and overall health status. By integrating these hierarchical question-answer pairs into the model training process, the proposed method can gradually guide the model from superficial image information analysis to deeper pathological conditions, thereby improving the accuracy and reliability of the reports generated by the model.

[0042] Through the multimodal medical imaging report generation system based on the medical thinking chain proposed in the present invention, a medical imaging diagnosis report containing a variety of heterogeneous information (such as impressions, findings and labels) can be generated under a unified framework, focusing on local medical features or abnormal features in medical images, analyzing and inferring potential causes and pathological conditions, and accurately matching local abnormalities in images with pathological inferences and disease diagnoses, thereby making a diagnostic analysis of the patient's overall condition and generating a comprehensive medical imaging diagnosis report.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] (1) Unlike traditional medical image report generation tasks that require heavy image labeling work, the present invention does not require a large amount of human resource consumption and feature engineering overhead, and avoids the lack of data diversity caused by constructing image reports based on image label templates. The present invention simulates the shallow to deep processing paradigm of image information by different neurons and synapses in the human cerebral cortex, combines the thinking steps and reasoning process of human diagnosis and analysis in real medical scenarios, and directly uses a large language model to hierarchically extract and reconstruct the original unstructured and disordered medical report text content into chain-structured, hierarchical feature attribute disease analysis and reasoning chain data. Specifically, the system designs six different dimensions in medical images from shallow to deep: modality, organ, size, abnormal location, symptoms, and overall health status. It also uses powerful large language models such as ChatGPT and Gemini to perform end-to-end semantic segmentation on the original disorganized medical image reports, and extracts question-answer pairs of six different dimensions.

[0045] (2) Different from the traditional CNN-RNN framework and the existing method of completing the task of generating medical image reports based on large language models, the present invention not only overcomes the fatal defect of the traditional CNN-RNN framework that it is difficult to achieve complex and diverse free text generation, but also significantly alleviates the serious model hallucination problem that is easily generated by the existing LLMs-based methods. The existing methods are prone to ignoring the local disease attribute features in the medical image, ignoring the causal relationship and correspondence between the local features and the overall disease assessment, or misdetecting local disease attribute features that do not exist. These serious model hallucination problems will lead to missed diagnosis, misdiagnosis, and wrong diagnosis, making it difficult for the existing LLMs to be actually applied in real medical scenarios and difficult to accurately help doctors perform overall medical image assessment and auxiliary diagnosis and treatment. The present invention connects the local shallow medical attribute features of medical images with the overall deep disease assessment summary in series to construct a disease reasoning thinking chain. On the one hand, it can significantly alleviate the model hallucination problem and avoid misdiagnosis and wrong diagnosis caused by the model's fabrication. The overall disease medical report obtained based on the reasoning thinking chain is reasonable and reliable. On the other hand, it helps the model focus on local details in the image, structuring and analyzing each layer of feature attributes one by one, avoiding missing any abnormal areas in the image and preventing missed diagnoses. The system's hierarchical chain data structure effectively captures the correlations between medical feature attributes across layers, promoting feature adaptation and enhanced expressiveness between heterogeneous attributes, and providing a new learning paradigm for generating more robust chain-based disease inference and diagnosis models.

[0046] (3) Different from the traditional method of applying the reasoning stage of thought chain, this paper proposes a method to embed the concept of reasoning thought chain into training data. By simulating the human cognitive process and the information fusion strategy of neural synapses in the cerebral cortex, this method realizes the dynamic aggregation and integration of feature attributes from different levels, explicitly enhances the utilization rate of multimodal features and the recognition ability of intent understanding, and improves the generalization, reliability and accuracy of the system's reasoning ability.

[0047] (4) The present invention overcomes the drawback of the lack of high-quality datasets for the previous medical image report generation task, and novelly designs and constructs a large-scale high-quality dataset for the medical image report generation task. This dataset is based on the existing public medical image report generation dataset, and according to the image size and resolution, text content length, text sensitive information, etc., it screens out high-quality images and comprehensive report text content, and removes redundancy and doctor-patient privacy from the report content to construct a high-quality image report dataset. Based on this dataset, the key technology proposed by the present invention is applied to extract hierarchical medical attribute question-answer pairs, and serially construct chain thinking question-answer pair data that implies the process of disease reasoning and diagnosis. This overcomes the lack of diagnostic reasoning ability of the model in the previous medical image report generation task, which can infer the overall condition from the local detail features from the point to the surface, from the shallow to the deep, and enables the model to simulate the thinking and cognitive process of human doctors when analyzing medical images in real medical scenarios, gradually observe and analyze from the surface features of the image, and finally make an accurate and comprehensive assessment summary of the patient's condition in the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flow chart of the method of the present invention;

[0049] Figure 2 is a flow chart of the pre-processing steps of the present invention;

[0050] Figure 3 Schematic diagram of the hierarchical medical attribute question-answer pair extraction process of the present invention;

[0051] Figure 4 Schematic diagram of the chained reasoning question-answer pair reconstruction process of the present invention. DETAILED DESCRIPTION

[0052] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0053] The limited availability of training data and privacy issues in the medical field have greatly hindered the accuracy and generalization of models. In the task of medical imaging report generation, one way to improve the accuracy of generated content and mitigate hallucinations is to increase the robustness and generalization of the model by improving the quality and diversity of training data. However, traditional data augmentation methods, such as image flipping, rotation, scaling, and introducing synthetic noise, are not suitable for the multimodal medical field due to the high fidelity required for medical image interpretation. Any change that distorts the medical reality of the image or destroys the feature alignment between the image and text may lead to misdiagnosis or overlook critical patient-specific details. In recent years, many studies have expanded the data by directly rewriting the text using large language models, but the data diversity increased by simple rewriting is very limited.

[0054] Research in cognitive neuroscience shows that humans typically process graphic information from the surface to the deep layers, seamlessly understanding hierarchical information. However, artificial intelligence models lack this ability and must gradually analyze each small area in the image by reducing the receptive field. In addition, medical texts written by humans are usually unorganized, which complicates the process of associating visual features with unstructured free text. Therefore, the present invention applies the cognitive process of human image recognition to data organization in multimodal models. In a broad sense, the integration of textual information can help AI models simulate human cognitive processes. Based on this integration process, the present invention proposes a hierarchical text enhancement method based on a chain hierarchy of medical attribute importance. This hierarchical semantic segmentation of medical text is inspired by the cognitive steps taken by doctors when evaluating medical images and diagnosing, involving the sequential recognition and integration of information across various levels.

[0055] The cognitive process of a doctor evaluating a medical image typically begins with global information about the image, such as the modality and depicted human systems, then gradually focuses on specific organs, followed by fine-grained attributes such as the shape or location of organs or abnormalities. After thoroughly understanding this information, the human brain integrates the symptoms presented in the image, ultimately generating an unstructured image report. However, from a large-scale model perspective, medical reports are typically highly disorganized. Even for humans, quickly associating lengthy medical reports with every fine-grained semantic detail in the image is challenging. Therefore, based on this cognitive process, the present invention designs specific prompts specifically for the task of generating medical image reports. This enables powerful large-scale language models such as ChatGPT and Gemini to generate structured question-answer pairs from the original unstructured and chaotic medical report text written by humans, focusing on fine-grained medical imaging questions at different dimensions and levels. The system then simulates the thought chain paradigm and transfers this paradigm from the reasoning process to the model fine-tuning process. This involves integrating low-level information cues into high-level questions to generate new question-answer pairs and performing supervised fine-tuning (SFT) of the model. These question-answer pairs provide a more fine-grained dimensional description and are often used as high-quality training data to help models understand images, thereby alleviating the model hallucination problem.

[0056] Specifically, this embodiment provides a method for generating a multimodal medical imaging report based on a medical thinking chain, such as Figure 1 As shown, the following steps are included:

[0057] S1, obtain the existing raw unordered and unstructured medical imaging report dataset.

[0058] S2, preprocess the image reports in the dataset to obtain a high-quality dataset.

[0059] like Figure 2 As shown, the preprocessing includes the following steps:

[0060] S21, deleting data rules in the data set where the resolution of medical images is lower than a preset resolution threshold;

[0061] S22: To ensure the completeness of the report content, delete data items that lack the Finding and Impression sections or have less than 10 words in length in all medical inclusions;

[0062] S23 de-privacy and redundancy of medical imaging reports, deletes sentences related to doctor-patient information and duplicate medical records, and uses the powerful language capabilities of the large language model to filter and delete sentences related to the patient's historical medical condition.

[0063] S3 processes medical imaging reports in high-quality datasets based on a large language model and extracts hierarchical medical attribute question-answer pairs.

[0064] This embodiment divides the hierarchical medical attribute question-answer pairs into six dimensions, which are modality, organ, size, abnormal location, symptom, and overall health status from low to high levels.

[0065] In order to transform the disorganized image reports into high-quality question-answer pairs that simulate human cognition, a powerful large language model is used to perform semantic segmentation on the image reports. Figure 3 As shown in the figure, with the help of our designed prompts, we use powerful large language models such as ChatGPT and Gemini to perform end-to-end semantic segmentation on the original messy medical image report. We use large language models to perform end-to-end semantic segmentation on the original messy medical image report, extract questions corresponding to six different dimensions: modality, organ, size, abnormal location, symptoms, and overall health status. Based on each question, we independently query and retrieve information in each dimension from the original image report as the answer to the question, thereby hierarchically transforming the original disordered and unstructured medical image report text into six-dimensional question-answer pairs. The definition of the above six dimensions is based on the common focus points in medical imaging reports, and is also established to simulate the human doctor's analysis process of medical imaging reports from shallow to deep observation and analysis perspectives.

[0066] S4, reconstructs the hierarchical medical mathematics question-answer pairs into chained reasoning question-answer pairs.

[0067] In order to further simulate the human cognitive process and the diagnosis and treatment thinking steps of human doctors, this paper proposes a question-answer pair reconstruction method based on chain thinking. Figure 4 As shown in Figure 2, the method for reconstructing hierarchical medical mathematics question-answer pairs into chained reasoning question-answer pairs is as follows:

[0068] For each original question except the lowest level in the extracted question-answer pair, the answers to all lower-level questions with lower levels than the original question are prefixed and concatenated with the original question in series, the question is reconstructed in a chain manner, and combined with the answer corresponding to the original question to obtain a chained reasoning question-answer pair.

[0069] Based on the above construction process, this example collected 10,000 images with medical reports from the public medical imaging report datasets MIMIC and OpenI, and constructed 60,000 hierarchical question-answer pairs and 60,000 chained question-answer pairs. This newly constructed chained question-answer pair data, combined with the original medical images and reports, was used for supervised fine-tuning of the model.

[0070] S5, uses chain reasoning question-answering data to fine-tune the existing medical imaging report generation model, and uses the fine-tuned medical imaging report generation model to generate a medical imaging report.

[0071] This embodiment also provides a multimodal medical imaging report generation system based on a medical thinking chain, including:

[0072] Raw data acquisition module: used to obtain existing raw disordered and unstructured medical imaging report data sets;

[0073] Preprocessing module: used to preprocess the image reports in the dataset to obtain high-quality datasets;

[0074] Hierarchical Question-Answer Pair Extraction Module: This module processes medical imaging reports in high-quality datasets based on a large language model and extracts hierarchical medical attribute question-answer pairs.

[0075] Chained reasoning question-answer pair reconstruction module: used to concatenate hierarchical medical mathematics question-answer pairs into chained reasoning question-answer pairs;

[0076] Model fine-tuning and report generation module: used to fine-tune the existing medical imaging report generation model using chained reasoning question-answering data, and generate medical imaging reports using the fine-tuned medical imaging report generation model.

[0077] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0078] The multimodal medical imaging report generation system based on the medical thinking chain demonstrated in this embodiment introduces a fine-grained medical imaging attribute thinking chain into the medical imaging report generation task for the first time, effectively promoting the development of medical imaging report generation tasks in the real world, significantly improving the disease reasoning and diagnosis capabilities of LLMs in this task, and greatly alleviating the problem of medical model hallucination. At the same time, by utilizing the powerful language capabilities of LLMs, hierarchical feature attribute question-answer pairs are extracted from the original medical imaging report, and chain thinking data for etiology and disease reasoning and diagnosis is constructed. The heterogeneous modal information is preprocessed and feature extracted by combining different units, realizing the organic chain fusion of multimodal features. Finally, based on the medical thinking chain, the accuracy of multimodal medical imaging report generation is greatly enhanced, the generalization performance and accuracy of the model are improved, and the model hallucination problem is alleviated. The method proposed in the present invention can provide a complete and effective unified framework for medical imaging report generation training and reasoning, providing reliable guarantees for subsequent application in real-life human scenarios for auxiliary diagnosis and treatment.

[0079] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A multimodal medical imaging report generation method based on medical thinking chain, characterized in that: The following steps are involved: Obtain existing raw, unordered, and unstructured medical imaging report datasets; Preprocess the image reports in the dataset to obtain a high-quality dataset; Medical imaging reports in a high-quality dataset are processed based on a large language model to extract hierarchical medical attribute question-answer pairs. The hierarchical medical attribute question-answer pairs are divided into six dimensions, which are modality, organ, size, abnormal location, symptoms, and overall health status, from low to high levels. The method for extracting hierarchical medical attribute question-answer pairs is specifically as follows: based on specific prompt words, the large language model is used to perform end-to-end semantic segmentation on the original disorganized medical imaging reports, extracting questions corresponding to the six different dimensions of modality, organ, size, abnormal location, symptoms, and overall health status, and independently querying and retrieving information of each dimension in the original image report as the answer to each question, thereby hierarchically transforming the original disordered and unstructured medical imaging report text into question-answer pairs of the six dimensions. The hierarchical medical attribute question-answer pairs are serially reconstructed into chained reasoning question-answer pairs. Specifically, for each original question except the lowest level in the extracted question-answer pairs, the answers to all lower-level questions are prefixed and concatenated with the original question in series. The questions are reconstructed in a chained manner and combined with the answers corresponding to the original questions to obtain chained reasoning question-answer pairs. The existing medical imaging report generation model is fine-tuned using chained reasoning question answering, and the fine-tuned medical imaging report generation model is used to generate medical imaging reports.

2. A multimodal medical imaging report generation method based on medical thought chain according to claim 1, characterized in that: The pretreatment comprises the following steps: Regulations for deleting data from a dataset where the resolution of medical images is lower than a preset resolution threshold; Deleting data provisions that lack a certain portion of text content or whose length is less than a certain length; Medical imaging reports are protected from privacy and redundancy, and content related to doctor-patient information, duplicate medical records, and content related to the patient's historical medical conditions are deleted.

3. A multimodal medical imaging report generation system based on medical thinking chain, characterized by: include: Raw data acquisition module: used to obtain existing raw disordered and unstructured medical imaging report data sets; Preprocessing module: used to preprocess the image reports in the dataset to obtain high-quality datasets; Hierarchical Question-Answer Pair Extraction Module: This module processes medical imaging reports in a high-quality dataset based on a large language model to extract hierarchical medical attribute question-answer pairs. The hierarchical medical attribute question-answer pairs are divided into six dimensions, which are modality, organ, size, abnormal location, symptoms, and overall health status, from low to high levels. The module performs the following steps: based on specific prompt words, it uses a large language model to perform end-to-end semantic segmentation on the original disorganized medical imaging reports, extracts questions corresponding to the six different dimensions of modality, organ, size, abnormal location, symptoms, and overall health status, and independently queries and retrieves information in each dimension from the original image report as the answer to each question, thereby hierarchically transforming the original disordered and unstructured medical imaging report text into question-answer pairs of the six dimensions. Chained reasoning question-answer pair reconstruction module: used to reconstruct hierarchical medical attribute question-answer pairs in series into chained reasoning question-answer pairs; the chained reasoning question-answer pair reconstruction module performs the following steps: for each original question except the lowest level in the extracted question-answer pairs, the answers to all lower-level questions with lower levels than the original question are prefixed and concatenated with the original question in series, reconstructing the question in a chained manner, and combining it with the answer corresponding to the original question to obtain a chained reasoning question-answer pair; Model fine-tuning and report generation module: used to fine-tune the existing medical imaging report generation model using chained reasoning question-answering, and generate medical imaging reports using the fine-tuned medical imaging report generation model.

4. A multimodal medical imaging report generation system based on medical thought chain according to claim 3, characterized in that: The pre-processing module performs the following steps: Regulations for deleting data from a dataset where the resolution of medical images is lower than a preset resolution threshold; Deleting data provisions that lack a certain portion of text content or whose length is less than a certain length; Medical imaging reports are protected from privacy and redundancy, and content related to doctor-patient information, duplicate medical records, and content related to the patient's historical medical conditions are deleted.

Citation Information

Patent Citations

  • Knowledge question and answer method, device and equipment and storage medium

    CN116561278A

  • Training method and application of multi-round session type medical image analysis model

    CN116759074A