Image Report Generation Method, System and Device Based on Multi-Granularity Knowledge Fusion
By adopting a multi-grained knowledge fusion mechanism in the medical image report generation system, combining medical imaging and multi-source medical knowledge, the shortcomings in existing systems in terms of semantic accuracy and clinical relevance are solved, and more efficient diagnosis and easier to understand reporting are achieved.
Patent Information
- Application Number
- CN202510318397.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing medical image report generation system has shortcomings in the fusion of multi-grained knowledge, clinical relevance and semantic accuracy, and it is difficult to make full use of the knowledge in the medical knowledge base, resulting in poor professionalism, accuracy and interpretability of the report.
An image report generation system based on multi-grained knowledge fusion is adopted. Through the image report pre-generation unit, knowledge retrieval unit, knowledge fusion unit and final report generation unit, multiple interactions and fusions are carried out in combination with medical images, initial image report text and multi-source medical knowledge to finally generate the final image report. The system includes an improved residual network, a bidirectional Transformer encoder and a multimodal hierarchical attention mechanism for extracting multi-scale visual features, encoding text information, and realizing bidirectional interaction between visual features and text features.
It significantly improves the semantic accuracy and clinical relevance of medical imaging reports, reduces diagnostic errors, improves doctors' work efficiency, and provides patients with easier understanding reports, helping the automation and intelligence development of the medical imaging field.
Smart Images

Figure CN119851852B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to an image report generation method, system and device based on multi-granularity knowledge fusion. Background Art
[0002] In the field of medical image diagnosis and report generation, the development of automation and intelligence has become an important means to improve medical efficiency and accuracy. Currently, medical image report generation systems mainly rely on visual feature extraction and natural language processing technologies, but these systems still have many deficiencies in multi-granularity knowledge fusion, clinical relevance, and semantic accuracy. Existing systems usually extract features of medical images through a single visual encoder and use pre-trained language models to generate reports. However, when dealing with complex images and multi-modal information, this method often fails to fully utilize the rich knowledge in the medical knowledge base, resulting in a large gap in the professionalism, accuracy, and interpretability of the generated reports.
[0003] Specifically, the existing medical image report generation systems mainly face the following problems: First, a single visual feature extraction method is difficult to capture the complex details in medical images, especially in multi-scale and multi-modal image data. Second, most of the existing knowledge retrieval and fusion methods rely on static medical knowledge bases and cannot dynamically integrate the latest medical research results and knowledge in real-time data sources. This makes the generated reports lack clinical relevance and easily miss important information. In addition, there are bottlenecks in the multi-modal information fusion of existing report generation models, and they cannot effectively integrate image and text information, resulting in defects in the semantic accuracy and coherence of the generated reports. Summary of the Invention
[0004] In view of this, the present invention aims at the problems of insufficient semantic accuracy and low clinical relevance in the existing medical image analysis and report generation technologies. Therefore, the present invention adopts the following technical solutions: An image report generation system based on multi-granularity knowledge fusion, which includes:
[0005] An image report pre-generation unit, configured to receive a medical image as an input and process the medical image through an image report generation model to obtain an initial image report text;
[0006] A knowledge retrieval unit, configured to retrieve relevant medical entities and their definition knowledge from a public medical knowledge base, a locally self-built medical knowledge base, and a real-time data source based on the initial image report text to obtain multi-source medical knowledge;
[0007] A knowledge fusion unit, configured to fuse the multi-source medical knowledge to obtain a fused knowledge text;
[0008] A final report generation unit, which combines medical images, an initial image report text, and a fused knowledge text, and regenerates an image report through an image report generation model to obtain a final image report.
[0009] Preferably, the image report generation model includes a visual encoder, an adapter module, a bidirectional Transformer encoder, and a large language model; wherein:
[0010] The visual encoder extracts multi-scale visual features of medical images through an improved residual network, and the improved residual network introduces an adaptive feature fusion module on the basis of the traditional ResNet to dynamically fuse visual features at different levels, thereby enhancing the ability to capture image details;
[0011] The adapter module adapts the features extracted by the visual encoder for effective interaction with the subsequent large language model;
[0012] The bidirectional Transformer encoder performs word segmentation and embedding processing on the initial image report text, and encodes the text through the bidirectional Transformer encoder, and the bidirectional Transformer encoder introduces context-aware position encoding to capture long-distance dependencies in the text and the context semantics of medical entities;
[0013] The large language model receives inputs from the visual encoder and the text feature encoding layer in the pre-generation stage, and adopts a multi-modal hierarchical attention mechanism to realize the bidirectional interaction between visual features and text features, including visual-to-text attention and text-to-visual attention, to dynamically adjust the weight distribution of visual features and text features.
[0014] Preferably, the knowledge retrieval unit extracts medical entities in the initial image report text through entity recognition technology or deep learning technology, and performs knowledge retrieval based on these entities to obtain the definitions of medical entities related to the image and their knowledge.
[0015] Preferably, the knowledge fusion unit performs knowledge fusion in one or more of the following ways: direct splicing, knowledge text representation based on structured data to text, and large language model integration.
[0016] Preferably, the knowledge text representation based on structured data to text is realized through the following steps:
[0017] For the structured data based on the medical knowledge graph, the entities and their relationships in the knowledge graph are encoded by a graph neural network to generate a graph embedding representation, and the graph embedding representation is converted into natural language text through a graph-to-sequence model. The graph-to-sequence model adopts a multi-hop attention mechanism to capture the semantic information of multi-hop relationships in the knowledge graph and dynamically generate a coherent medical description text.
[0018] Preferably, the final report generation unit supports generating reports in different formats, including concise reports, detailed reports, professional reports, vernacular reports, and visual annotation reports.
[0019] Preferably, it further includes:
[0020] A report presentation unit, which is used to segment by combining medical images and the final image report, and mark the regions, problems, and symptoms involved in the final image report to achieve an interpretable effect of combining text and images.
[0021] Preferably, the report presentation unit includes a multi-modal alignment and segmentation module, a region-aware segmentation model, and a text-image interaction generation module; among them:
[0022] The multi-modal alignment and segmentation module realizes the precise alignment of medical images and report texts through a multi-modal alignment network;
[0023] The region-aware segmentation model precisely segments medical images through a region-aware segmentation network;
[0024] The text-image interaction generation module realizes the mutual interpretation of images and texts through a text-image interaction generator.
[0025] A method for generating an image report based on multi-granularity knowledge fusion according to the second embodiment of the present invention includes:
[0026] Receiving a medical image as input, and processing the medical image through an image report generation model to obtain an initial image report text;
[0027] Retrieving relevant medical entities and their definition knowledge from a public medical knowledge base, a locally built medical knowledge base, and real-time data sources based on the initial image report text to obtain multi-source medical knowledge;
[0028] Fusing the multi-source medical knowledge to obtain a fused knowledge text;
[0029] Combining the medical image, the initial image report text, and the fused knowledge text, and regenerating the image report through the image report generation model to obtain a final image report.
[0030] The third embodiment of the present invention also provides an imaging report generation device based on multi-granularity knowledge fusion, which includes a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement the imaging report generation method based on multi-granularity knowledge fusion as described above.
[0031] In summary, through the multi-granularity knowledge fusion mechanism, the present invention organically combines visual features, medical knowledge base knowledge, real-time auxiliary knowledge, and pre-generated content, enabling higher semantic accuracy and clinical relevance in the automated analysis and report generation of medical images. The system can significantly reduce diagnostic errors, improve doctors' work efficiency, and at the same time provide patients with more understandable reports, facilitating the development of automation and intelligence in the field of medical imaging. By dynamically updating the knowledge base, the system can always generate reports based on the latest medical knowledge, ensuring the timeliness and accuracy of the reports. In addition, the system supports the generation of reports in different formats to meet the needs of different clinical scenarios and patients, improving the flexibility and practicality of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is the overall schematic diagram of the imaging report generation system based on multi-granularity knowledge fusion according to the first embodiment of the present invention;
[0034] Figure 2 It is the working principle diagram of the imaging report generation system based on multi-granularity knowledge fusion according to the first embodiment of the present invention;
[0035] Figure 3 It is the flow schematic diagram of the imaging report generation method based on multi-granularity knowledge fusion according to the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] The present invention provides an image report generation system based on multi-granularity knowledge fusion, aiming to improve the accuracy and intelligence level of medical image data analysis and report generation. The following will combine the attached Figure 1 to the attached Figure 2 , and elaborate on the specific implementation manners of the present invention in detail.
[0038] As Figure 1 shown, in this embodiment, the image report generation system includes:
[0039] An image report pre-generation unit 10, configured to receive a medical image as input, and process the medical image through an image report generation model to obtain an initial image report text.
[0040] Specifically, in this embodiment, the image report pre-generation unit 10 is configured to process the medical image through an image report generation model to obtain an initial image report text.
[0041] In this embodiment, as Figure 2As shown, the image report generation model adopts a Multimodal Hierarchical Attention Mechanism (MHAM), and its specific structure includes a visual feature encoding layer, a text feature encoding layer, and a multimodal interaction layer. The visual feature encoding layer extracts multi-scale visual features of medical images through an improved Residual Network (ResNet). The improved Residual Network introduces an Adaptive Feature Fusion Module (AFFM) on the basis of the traditional ResNet, which is used to dynamically fuse visual features at different levels to enhance the ability to capture image details. Specifically, the adaptive feature fusion module dynamically adjusts the weights of features at each level by learning the correlation between features at different levels, so that important details in the image can be better retained when extracting high-dimensional features. The text feature encoding layer encodes the initial image report text through a bidirectional Transformer encoder. The bidirectional Transformer encoder introduces a Context-Aware Position Encoding (CAPE) on the basis of the traditional Transformer, which is used to capture long-range dependencies in the text and the context semantics of medical entities. The context-aware position encoding introduces context information into the position encoding, enabling the model to better understand medical terms and long-range dependencies in the text, thereby improving the accuracy of text encoding. The multimodal interaction layer realizes the bidirectional interaction between visual features and text features through the multimodal hierarchical attention mechanism. The multimodal hierarchical attention mechanism includes Vision-to-Text Attention (VTA) and Text-to-Vision Attention (TVA), which are used to dynamically adjust the weight distribution of visual features and text features to enhance the model's ability to fuse image and text information. Specifically, the vision-to-text attention mechanism dynamically adjusts the weight of visual features in text generation by calculating the similarity between visual features and text features, so that the generated text can more accurately reflect the details in the image. The text-to-vision attention mechanism dynamically adjusts the weight of text features in image segmentation by calculating the similarity between text features and visual features, so that image segmentation can more accurately reflect the content described in the text.
[0042] A knowledge retrieval unit 20, configured to retrieve relevant medical entities and their definition knowledge from a public medical knowledge base, a locally built medical knowledge base, and a real-time data source based on the initial image report text, so as to obtain multi-source medical knowledge.
[0043] Specifically, the knowledge retrieval unit 20 retrieves relevant medical knowledge from multi-source data based on the initial imaging report text. The multi-source data includes an open medical knowledge base, a local medical knowledge base, and real-time data sources. Among them, the open medical knowledge base includes UMLS, PubMed, ICD10, SNOMED CT, etc., which are used to retrieve entities, definitions, and their knowledge related to imaging. Specifically, UMLS is used to obtain the standardized representation of medical entities, PubMed is used to obtain the latest medical research results, ICD10 is used to obtain the international classification standard of diseases, and SNOMED CT is used to obtain the standardized representation of clinical terms and concepts. The local medical knowledge base includes medical guidelines, rare disease diagnosis and treatment guidelines, patient medical records, medical dissertations, etc., and retrieves entities, definitions, and their knowledge related to imaging through text retrieval or vector retrieval. Specifically, text retrieval obtains medical knowledge related to the initial imaging report text from the local medical knowledge base through keyword matching and semantic analysis; vector retrieval converts text and imaging features into vector representations, calculates the similarity between vectors, and thus obtains medical knowledge related to imaging features and the initial imaging report text from the local medical knowledge base. Real-time data sources include general search engines or medical engines, such as Baidu, Sogou, Bing, Google, Wanfang, CNKI, etc., which are used to obtain auxiliary knowledge related to imaging features and initial descriptions. Specifically, real-time data sources obtain auxiliary knowledge such as the latest medical research results, clinical guidelines, and patient medical records from general search engines or medical engines through web crawler technology to supplement the lack of medical knowledge in the initial imaging report text.
[0044] The knowledge fusion unit 30 is used to fuse the multi-source medical knowledge to obtain a fused knowledge text.
[0045] Specifically, the knowledge fusion unit 30 fuses multi-source medical knowledge based on artificial intelligence technology to obtain a fused knowledge text. The knowledge fusion methods include direct splicing or segmented splicing, text representation from structured data to text, and large language model integration. The direct splicing or segmented splicing method generates a fused knowledge text by splicing or segmenting the obtained knowledge units in a certain order. Specifically, the direct splicing method splices all the obtained knowledge units in order into a complete text; the segmented splicing method segments and splices the knowledge units according to the type and content of the knowledge units to improve the readability and logic of the text. The text representation from structured data to text is achieved as follows: The knowledge graph-driven text generation module encodes the entities and their relationships in the medical knowledge graph based on the structured data of the medical knowledge graph through a Graph Neural Network (GNN) to generate a graph embedding representation. The graph embedding representation is converted into natural language text through a Graph-to-Sequence (G2S) model, and the graph-to-sequence model adopts a Multi-Hop Attention Mechanism (MHAM) to capture the semantic information of multi-hop relationships in the knowledge graph and dynamically generate a coherent medical description text. Specifically, the multi-hop attention mechanism gradually captures the multi-hop relationships in the knowledge graph through multiple attention layers, thereby generating a more accurate and coherent medical description text. The structured data parsing and semantic mapping module parses non-text data such as tables, key-value pairs, and tree structures in the medical knowledge base into an intermediate representation form through a Structured Data Parser (SDP), and the intermediate representation form is converted into natural language text through a Semantic Mapping Network (SMN). The semantic mapping network adopts a Hierarchical Attention Mechanism (HAM) for semantic parsing and dynamic weight allocation of different levels of structured data to generate accurate and highly readable medical text. Specifically, the hierarchical attention mechanism gradually parses different levels of structured data through multiple attention layers, thereby generating a more accurate and highly readable medical text. The large model enhancement techniques include Domain-Adaptive Pretraining (DAP), Knowledge Injection Mechanism (KIM), and Multi-Task Joint Training (MTJT) to improve the medical accuracy and professionalism of the generated text.Domain adaptation pre-training enhances the model's understanding of medical domain terms, grammar, and semantics by performing secondary pre-training on a general large language model using a vast amount of medical text data (such as medical literature, clinical guidelines, and medical records). Specifically, domain adaptation pre-training introduces a large amount of medical text data into the general large language model, enabling the model to better understand medical domain terms, grammar, and semantics, thereby improving the medical accuracy of the generated text. The knowledge injection mechanism dynamically injects entity relationships, clinical rules, and the latest medical research findings in the medical knowledge graph into the generation process of the large language model through an external knowledge injection module (EKIM) to enhance the medical accuracy and professionalism of the generated text. Specifically, the external knowledge injection module injects entity relationships, clinical rules, and the latest medical research findings in the medical knowledge graph into the large language model in the form of vectors or text, enabling the model to refer to this external knowledge when generating text, thus generating more accurate and professional medical text. Multi-task joint training enhances the model's comprehensive understanding and application ability of medical knowledge by jointly training the large language model to complete multiple medical tasks (such as entity recognition, relation extraction, and text generation). Specifically, multi-task joint training introduces multiple medical tasks into the large language model, enabling the model to comprehensively consider various medical knowledge when generating text, thereby improving the medical accuracy and professionalism of the generated text. Finally, the fused knowledge text generated through the above knowledge fusion method is transmitted to the final report generation unit for generating the final report.
[0046] The final report generation unit 40 is configured to combine medical images, the initial image report text, and the fused knowledge text, and regenerate the image report through an image report generation model to obtain the final image report.
[0047] Specifically, the final report generation unit 40 regenerates the imaging report through the imaging report generation model based on the medical image, the initial imaging report text, the retrieved medical knowledge, or the fused knowledge text. The imaging report generation model adopts a multi-modal hierarchical attention mechanism, and its specific structure includes a visual feature encoding layer, a text feature encoding layer, and a multi-modal interaction layer. The visual feature encoding layer extracts multi-scale visual features of the medical image through an improved residual network. The improved residual network introduces an adaptive feature fusion module on the basis of the traditional residual network, which is used to dynamically fuse visual features at different levels to enhance the ability to capture image details. The text feature encoding layer encodes the initial imaging report text and the fused knowledge text through a bidirectional Transformer encoder. The bidirectional Transformer encoder introduces context-aware position encoding on the basis of the traditional Transformer, which is used to capture long-distance dependencies in the text and the context semantics of medical entities. The multi-modal interaction layer realizes the bidirectional interaction between visual features and text features through the multi-modal hierarchical attention mechanism. The multi-modal hierarchical attention mechanism includes attention from vision to text and attention from text to vision, which are used to dynamically adjust the weight distribution of visual features and text features to enhance the model's ability to fuse image and text information. Specifically, the attention mechanism from vision to text dynamically adjusts the weight of visual features in text generation by calculating the similarity between visual features and text features, so that the generated text can more accurately reflect the details in the image; the attention mechanism from text to vision dynamically adjusts the weight of text features in image segmentation by calculating the similarity between text features and visual features, so that the image segmentation can more accurately reflect the content described in the text. The final report generation unit supports the generation of reports in different formats, including concise reports, detailed reports, professional reports, plain language reports, and visualization annotation reports, to meet the needs of different clinical scenarios and patients. Specifically, concise reports are suitable for rapid diagnosis and preliminary screening, detailed reports are suitable for comprehensive diagnosis and detailed analysis, professional reports are suitable for the reference and research of medical experts, plain language reports are suitable for patient understanding and communication, and visualization annotation reports are suitable for the precise correspondence and interpretation of images and texts.
[0048] The report presentation unit 50 performs precise segmentation in combination with the medical image and the final report, and marks the involved regions, problems, and symptoms to achieve an interpretable effect with a combination of pictures and texts.
[0049] Among them, the report presentation unit 50 includes a multimodal alignment and segmentation module, a region-aware segmentation model, and an image-text interactive generation module. The multimodal alignment and segmentation module realizes the precise alignment of medical images and report texts through a multimodal alignment network (Multimodal Alignment Network, MAN). The multimodal alignment network adopts a cross-modality attention mechanism (Cross-Modality Attention Mechanism, CMAM) to dynamically capture the semantic correlation between image regions and text descriptions and generate an aligned multimodal feature representation. Specifically, the cross-modality attention mechanism realizes the precise alignment of image regions and text descriptions by calculating the similarity between medical images and report texts and dynamically adjusting the weight distribution between image regions and text descriptions. The region-aware segmentation model precisely segments medical images through a region-aware segmentation network (Region-Aware Segmentation Network, RASN). The region-aware segmentation network introduces a region attention mechanism (Region Attention Mechanism, RAM) on the basis of a traditional segmentation model to dynamically adjust the weight distribution of image segmentation according to the keywords and semantic information in the report text to enhance the segmentation accuracy of the target region. Specifically, the region attention mechanism dynamically adjusts the weight distribution of image segmentation by calculating the similarity between the keywords in the report text and the image regions, so that the image segmentation can more accurately reflect the description content in the report text. The image-text interactive generation module realizes the mutual interpretation of images and texts through an image-text interactive generator (Image-Text Interactive Generator, ITIG). The image-text interactive generator adopts a bidirectional generation strategy, including image-to-text generation and text-to-image generation. Image-to-text generation generates corresponding text descriptions based on the segmented image regions. The generation process dynamically adjusts the semantic consistency of the generated text through an image-to-text attention mechanism (Image-to-Text Attention, ITA). Specifically, the image-to-text attention mechanism dynamically adjusts the weight distribution of the generated text by calculating the similarity between the segmented image regions and the generated text, so that the generated text can more accurately describe the details in the image. Text-to-image generation generates corresponding image markers based on the descriptions in the report text. The generation process dynamically adjusts the position and shape of the image markers through a text-to-image attention mechanism (Text-to-Image Attention, TIA). Specifically, the text-to-image attention mechanism dynamically adjusts the position and shape of the image markers by calculating the similarity between the descriptions in the report text and the image markers, so that the image markers can more accurately reflect the description content in the report text.
[0050] It should be noted that in order to further improve the accuracy and real-time performance of the system, the system supports dynamic updating of the knowledge base, uses vector retrieval technology to ensure the consistency between the knowledge base and real-time auxiliary knowledge, and preferentially uses the latest knowledge sources to improve the accuracy of reports. Specifically, the knowledge base is dynamically updated by regularly obtaining the latest medical knowledge from public medical knowledge bases, local medical knowledge bases, and real-time data sources to update the content in the knowledge base. The vector retrieval technology converts medical knowledge and real-time auxiliary knowledge into vector representations, calculates the similarity between vectors, thereby ensuring the consistency between the knowledge base and real-time auxiliary knowledge. Preferentially using the latest knowledge sources sets weights or priorities to preferentially refer to the latest medical knowledge and real-time auxiliary knowledge when generating reports, thereby improving the accuracy of reports.
[0051] The system and method of the present invention will be described in detail below in combination with a specific application scenario.
[0052] Suppose a hospital needs to analyze the lung CT images of a patient and generate a report. First, the image report pre-generation unit receives the lung CT images of the patient and extracts multi-scale visual features of the images through an improved Residual Network (ResNet). The improved Residual Network introduces an Adaptive Feature Fusion Module (AFFM) on the basis of the traditional ResNet to dynamically fuse visual features at different levels and enhance the ability to capture image details. Subsequently, the initial image report text is encoded by a bidirectional Transformer encoder. The bidirectional Transformer encoder introduces Context-Aware Position Encoding (CAPE) on the basis of the traditional Transformer to capture long-distance dependencies in the text and the context semantics of medical entities. The adapter module performs adaptation processing on the features extracted by the visual feature encoding layer to effectively interact with the large language model. The large language model receives inputs from the visual feature encoding layer and the text feature encoding layer in the pre-generation stage, and uses a multi-modal hierarchical attention mechanism (MHAM) to achieve two-way interaction between visual features and text features, generating the initial image report text.
[0053] Next, the knowledge retrieval unit retrieves relevant medical entities and their definition knowledge from public medical knowledge bases, locally built medical knowledge bases, and real-time data sources based on the initial image report text. For example, it retrieves medical entities and their definition knowledge related to lung CT images from public medical knowledge bases such as UMLS and PubMed; it obtains medical entities and their definition knowledge related to the images through text retrieval or vector retrieval from locally built medical knowledge bases such as medical guidelines, rare disease diagnosis and treatment guidelines, and patient medical records; it obtains auxiliary knowledge related to image features and initial descriptions from general search engines such as Baidu and Sogou or medical engines such as Wanfang and CNKI. The knowledge retrieval unit extracts medical entities in the initial image report text through entity recognition technology or deep learning technology, and conducts knowledge retrieval based on these entities to obtain medical entities related to the images and their definition knowledge.
[0054] Then, the knowledge fusion unit fuses multi-source medical knowledge through artificial intelligence technology to generate a fused knowledge text. Specifically, the knowledge fusion unit directly splices or segmentally splices the obtained knowledge units in a direct splicing manner to form a preliminary fused knowledge text. At the same time, the knowledge fusion unit uses technologies based on deep neural networks or large language models to complete the knowledge text representation from structured data to text (Data2Text). Based on the structured data of the medical knowledge graph, the graph neural network (GNN) encodes the entities and their relationships in the knowledge graph to generate a graph embedding representation, and the graph embedding representation is converted into natural language text through the graph-to-sequence (G2S) model. The graph-to-sequence model adopts a multi-hop attention mechanism (MHAM) to capture the semantic information of multi-hop relationships in the knowledge graph and dynamically generate coherent medical description texts. In addition, through a structured data parser (SDP), non-text data such as tables, key-value pairs, and tree structures in the medical knowledge base are parsed into an intermediate representation form, and are converted into natural language text through a semantic mapping network (SMN). The semantic mapping network adopts a hierarchical attention mechanism (HAM) to perform semantic parsing and dynamic weight allocation on different levels of structured data to generate accurate and highly readable medical texts. Finally, the knowledge fusion unit uses a large language model to integrate all the obtained knowledge to generate a fused knowledge text.
[0055] The final report generation unit combines medical images, initial image report texts, and fused knowledge texts, and regenerates the image report through an image report generation model. The image report generation model includes an image report generation model constructed based on a Transformer model, a residual network model, a convolutional neural network, a recurrent neural network, a vision-language large model (VLLM), or a multimodal large model. This model adopts a multimodal hierarchical attention mechanism (MHAM), including vision-to-text attention (VTA) and text-to-vision attention (TVA), to dynamically adjust the weight distribution of visual features and text features, so as to enhance the model's ability to fuse image and text information. The final report generation unit supports generating reports in different formats, including concise reports, detailed reports, professional reports, plain-language reports, and visualization annotation reports, to meet the needs of different clinical scenarios and patients. For example, for doctors, detailed professional reports can be generated; for patients, plain-language reports can be generated for easy understanding by patients.
[0056] Finally, the report presentation unit combines medical images and the final report, performs precise segmentation, and marks the regions, problems, and symptoms involved in the image report to achieve an interpretable effect of combining text and images. The report presentation unit includes a multimodal alignment and segmentation module, a region-aware segmentation model, and a text-image interaction generation module. The multimodal alignment and segmentation module realizes the precise alignment of medical images and report texts through a multimodal alignment network (MAN). The multimodal alignment network adopts a cross-modal attention mechanism (CMAM) to dynamically capture the semantic association between image regions and text descriptions and generate an aligned multimodal feature representation. The region-aware segmentation model precisely segments medical images through a region-aware segmentation network (RASN). The region-aware segmentation network introduces a region attention mechanism (RAM) on the basis of traditional segmentation models to dynamically adjust the weight distribution of image segmentation according to the keywords and semantic information in the report text, so as to enhance the segmentation accuracy of the target region. The text-image interaction generation module realizes the mutual interpretation of images and texts through a text-image interaction generator (ITIG). The text-image interaction generator adopts a bidirectional generation strategy, including image-to-text generation and text-to-image generation, to dynamically adjust the semantic consistency of the generated text and the position and shape of image markers. Image-to-text generation generates corresponding text descriptions according to the segmented image regions. The generation process dynamically adjusts the semantic consistency of the generated text through an image-to-text attention mechanism (ITA). Text-to-image generation generates corresponding image markers according to the descriptions in the report text. The generation process dynamically adjusts the position and shape of image markers through a text-to-image attention mechanism (TIA). Through the above steps, the report presentation unit generates an interpretable medical image report combining text and images and provides it to doctors and patients.
[0057] The present invention organically combines visual features, medical knowledge base knowledge, real-time auxiliary knowledge, and pre-generated content through a multi-granularity knowledge fusion mechanism, enabling the automated analysis and report generation of medical images to have higher semantic accuracy and clinical relevance. The system can significantly reduce diagnostic errors, improve the work efficiency of doctors, and at the same time provide patients with more understandable reports, facilitating the development of automation and intelligence in the field of medical imaging. By dynamically updating the knowledge base, the system can always generate reports based on the latest medical knowledge, ensuring the timeliness and accuracy of the reports. In addition, the system supports the generation of reports in different formats to meet the needs of different clinical scenarios and patients, improving the flexibility and practicality of the system.
[0058] Please refer to Figure 3 , the second embodiment of the present invention provides an image report generation method based on multi-granularity knowledge fusion, which includes:
[0059] S201, receiving a medical image as input, and processing the medical image through an image report generation model to obtain an initial image report text;
[0060] S202, retrieving relevant medical entities and their definition knowledge from a public medical knowledge base, a locally built medical knowledge base, and real-time data sources based on the initial image report text to obtain multi-source medical knowledge;
[0061] S203, fusing the multi-source medical knowledge to obtain a fused knowledge text;
[0062] S204, combining the medical image, the initial image report text, and the fused knowledge text, and regenerating the image report through the image report generation model to obtain a final image report.
[0063] The third embodiment of the present invention further provides an image report generation device based on multi-granularity knowledge fusion, which includes a memory and a processor. The memory stores a computer program, and the computer program can be executed by the processor to implement the image report generation method based on multi-granularity knowledge fusion as described above.
[0064] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0065] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0066] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0067] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image report generation system based on multi-granularity knowledge fusion, characterized in that: include: An image report pre-generation unit, for receiving a medical image as an input, and processing the medical image through an image report generation model to obtain an initial image report text; A knowledge retrieval unit is used to retrieve relevant medical entities and their definition knowledge from public medical knowledge bases, local self-built medical knowledge bases and real-time data sources based on the initial imaging report text to obtain multi-source medical knowledge; A knowledge fusion unit, used for fusing the multi-source medical knowledge to obtain a fused knowledge text; The final report generation unit is used to combine the medical image, the initial image report text and the fusion knowledge text, and regenerate the image report through the image report generation model to obtain the final image report.
2. The imaging report generation system based on multi-granularity knowledge fusion according to claim 1, characterized in that: The image report generation model includes a visual encoder, an adapter module, a bidirectional Transformer encoder and a large language model; wherein: The visual encoder extracts multi-scale visual features of medical images through an improved residual network. The improved residual network introduces an adaptive feature fusion module based on the traditional ResNet to dynamically fuse visual features of different levels, thereby enhancing the ability to capture image details; The adapter module adapts the features extracted by the visual encoder to effectively interact with the subsequent large language model; The bidirectional Transformer encoder performs word segmentation and embedding processing on the initial imaging report text, and encodes the text through a bidirectional Transformer encoder. The bidirectional Transformer encoder introduces context-aware position encoding to capture long-distance dependencies in the text and contextual semantics of medical entities; The large language model receives inputs from the visual encoder and the text feature encoding layer in the pre-generation stage, and adopts a multimodal hierarchical attention mechanism to achieve two-way interaction between visual features and text features, including visual-to-text attention and text-to-visual attention, so as to dynamically adjust the weight distribution of visual features and text features.
3. The imaging report generation system based on multi-granularity knowledge fusion according to claim 1, characterized in that: The knowledge retrieval unit extracts medical entities from the initial imaging report text through entity recognition technology or deep learning technology, and performs knowledge retrieval based on these medical entities to obtain medical entity definitions and knowledge related to the imaging.
4. The imaging report generation system based on multi-granularity knowledge fusion according to claim 1, characterized in that: The knowledge fusion unit performs knowledge fusion in one or more of the following ways: direct concatenation, textual representation of knowledge based on structured data to text, and integration of large language models.
5. The imaging report generation system based on multi-granularity knowledge fusion as claimed in claim 4, characterized in that: The textual representation of knowledge based on structured data to text is achieved by the following steps: For structured data based on medical knowledge graph, the entities and their relationships in the knowledge graph are encoded through graph neural network to generate graph embedding representation, which is converted into natural language text through graph-to-sequence model. The graph-to-sequence model adopts multi-hop attention mechanism to capture the semantic information of multi-hop relationships in the knowledge graph and dynamically generate coherent medical description text.
6. The imaging report generation system based on multi-granularity knowledge fusion according to claim 1, characterized in that: The final report generating unit supports generating reports in different formats, including concise reports, detailed reports, professional reports, vernacular reports and visually annotated reports.
7. The imaging report generation system based on multi-granularity knowledge fusion according to claim 1, characterized in that: Also includes: The report presentation unit is used to combine the medical image and the final image report for segmentation, and mark the areas, problems and symptoms involved in the final image report to achieve an interpretable effect of combining pictures and texts.
8. The imaging report generation system based on multi-granularity knowledge fusion according to claim 7, characterized in that: The report presentation unit includes a multimodal alignment and segmentation module, a region-aware segmentation model, and a graphic-text interaction generation module; wherein: The multimodal alignment and segmentation module realizes accurate alignment of medical images and report text through a multimodal alignment network; The region-aware segmentation model accurately segments medical images through a region-aware segmentation network; The image-text interaction generation module realizes mutual interpretation between images and texts through an image-text interaction generator.
9. A method for generating an image report based on multi-granularity knowledge fusion, characterized in that: include: receiving a medical image as input, and processing the medical image through an image report generation model to obtain an initial image report text; Based on the initial imaging report text, relevant medical entities and their definition knowledge are retrieved from public medical knowledge bases, local self-built medical knowledge bases and real-time data sources to obtain multi-source medical knowledge; Fusing the multi-source medical knowledge to obtain a fused knowledge text; Combining medical images, initial image report text and fused knowledge text, the image report is regenerated through an image report generation model to obtain the final image report.
10. An imaging report generation device based on multi-granularity knowledge fusion, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement the image report generation method based on multi-granularity knowledge fusion as described in claim 9.
Citation Information
Patent Citations
Method, equipment and device for intelligently generating medical image diagnosis report
CN118888076A
Multi-modal medical examination report automatic generation method and system
CN119092032A