Radiology report generation method and apparatus, and terminal and storage medium

Through the dynamic prior network model combined with dynamic knowledge graphs and prior knowledge networks, the problem of visual and text data in the generation of radiological reports is solved, and the accuracy and quality of report generation is improved, especially when describing rare anomalies.

WO2025137892A1PCT designated stage expired Publication Date: 2025-07-03SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2023/142139
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing radiological report generation techniques cannot effectively describe rare but important anomalies, resulting in inconsistencies between visual and textual data and fail to fully cover key anomalies.

Method used

The dynamic prior network model is adopted, combining the dynamic knowledge graph network and the prior knowledge network, visual representation is obtained through the dynamic knowledge graph network, and a prior knowledge network is used to generate radiological reports, integrating prior knowledge to alleviate text data bias.

Benefits of technology

Improve the quality of radiological reports generation, better handle visual and text bias caused by limited data, enhance the model's adaptability to different situations, and generate more accurate radiological reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023142139_03072025_PF_FP_ABST
    Figure CN2023142139_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention are a radiology report generation method and apparatus, and a terminal and a storage medium. The method comprises: inputting a radiology image to be processed into a pre-trained dynamics priori network model, wherein the dynamics priori network model comprises a dynamic knowledge graph network, a prior knowledge network and a decoder; in the dynamic knowledge graph network, obtaining a dynamic knowledge graph on the basis of the radiology image, and obtaining a visual representation on the basis of the dynamic knowledge graph; acquiring prior knowledge information corresponding to the radiology image, and on the basis of the prior knowledge information and the visual representation, using the prior knowledge network to obtain a prior representation vector; and inputting the visual representation and the prior representation vector into the decoder to generate a radiology report corresponding to the radiology image. In the present invention, the radiology report is generated by combining the dynamic knowledge graph with the prior knowledge information, such that visual and textual biases caused by limited data availability can be better handled, thereby improving the report generation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Radiology report generation method, device, terminal and storage medium Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to a method, device, terminal and storage medium for generating a radiology report. Background Art

[0002] Radiological images are crucial for diagnosis and treatment in clinical practice and medical trials. However, preparing radiology reports is a time-consuming and error-prone task, especially for inexperienced radiologists. This has led to a need for automated radiology report generation. In neural network-based radiology report generation, a neural network is trained on a large dataset of radiological images and reports. Once trained, the neural network can be used to generate new reports. This system aims to reduce the workload of radiologists and promote clinical automation, which is seen as a major step forward in the application of artificial intelligence in medicine. Such automation can significantly accelerate workflows and improve the quality and standardization of healthcare. Radiology reports summarize all clinical findings and impressions obtained during a radiological examination. These reports typically contain rich information, going beyond simple disease keywords. They may also include statements of negation and uncertainty, making them comprehensive documents for medical professionals to interpret and act upon. In recent years, driven by advances in automated image description systems based on neural networks, researchers have begun exploring automated radiology report generation.

[0003] Directly applying general-purpose generative techniques to radiological images presents several specific challenges: First, image data discrepancy: Normal, disease-free radiological images comprise the vast majority of datasets. This unbalanced distribution of images, compared to abnormal radiological images, distracts the model and hinders its ability to accurately capture the characteristics of various rare abnormal regions. Second, text data discrepancy: Radiologists tend to provide comprehensive descriptions of all elements within an image in their reports, a practice that leads to an overemphasis on describing normal regions. Furthermore, sentences describing the same normal region are generally similar. Therefore, due to this bias towards describing normal regions, training on these datasets results in predominantly generating normal sentences. This limitation impacts the model's ability to effectively describe specific, important abnormalities. The widely adopted hierarchical recurrent neural networks (HRNNs) generate repetitive sentences related to normal findings, making them ineffective in describing certain rare but important abnormalities.

[0004] Therefore, existing report generation technologies are unable to comprehensively cover and accurately describe these critical but infrequently occurring abnormal areas, and are unable to address the significant text data imbalance problem, which in turn leads to inconsistencies between visual and text data when generating radiology reports.

[0005] Therefore, the existing technology has defects and needs to be improved and developed.

[0006] Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a radiology report generation method, device, terminal and storage medium in response to the above-mentioned defects of the prior art, aiming to solve the problem of inconsistency between visual and text data when generating radiology reports in the prior art.

[0008] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0009] A method for generating a radiology report, the method comprising:

[0010] Inputting the radiological image to be processed into a pre-trained dynamic prior network model, wherein the dynamic prior network model includes a dynamic knowledge graph network, a prior knowledge network and a decoder;

[0011] In the dynamic knowledge graph network, a dynamic knowledge graph is obtained according to the radiological image, and a visual representation is obtained according to the dynamic knowledge graph;

[0012] Acquiring prior knowledge information corresponding to the radiological image, and obtaining a prior representation vector using the prior knowledge network based on the prior knowledge information and the visual representation;

[0013] The visual representation and the prior representation vector are input into the decoder to generate a radiology report corresponding to the radiology image.

[0014] In one implementation, in the dynamic knowledge graph network, obtaining a dynamic knowledge graph based on the radiological image, and obtaining a visual representation based on the dynamic knowledge graph include:

[0015] In the dynamic knowledge graph network, a plurality of target standard reports matching the radiological image are obtained based on a pre-built report database;

[0016] Acquire all entities in the target standard report, and generate an image knowledge graph corresponding to the radiological image based on the entities in the target standard report;

[0017] Obtain a pre-built medical knowledge graph, and add new paths to the medical knowledge graph based on the image knowledge graph to obtain a dynamic knowledge graph;

[0018] The dynamic knowledge graph is encoded using a dynamic graph encoder in the dynamic knowledge graph network to obtain a visual representation corresponding to the radiological image.

[0019] In one implementation, obtaining prior knowledge information corresponding to the radiological image, and obtaining a prior representation vector using the prior knowledge network based on the prior knowledge information and the visual representation, includes:

[0020] Taking all the target standard reports and the medical knowledge graph as prior knowledge information;

[0021] generating a first key vector and a first value vector corresponding to all target standard reports based on a report encoder, using the visual representation as a query vector through a cross-attention mechanism, and performing a cross-attention calculation on the query vector, the first key vector, and the first value vector to obtain a first representation vector;

[0022] generating a second key vector and a second value vector corresponding to the entity in the medical knowledge graph based on the report encoder, using the visual representation as a query vector through a cross-attention mechanism, and performing a cross-attention calculation on the query vector, the second key vector, and the second value vector to obtain a second representation vector;

[0023] The first representation vector and the second representation vector are combined to obtain a priori representation vector.

[0024] In one implementation, the training step of the dynamic priori network model includes:

[0025] Acquire a training data set, the training data set comprising: one-to-one corresponding radiology training images and radiology training reports;

[0026] Constructing an initial dynamic prior network model and inputting the radiological training image into the initial dynamic prior network model, wherein the initial dynamic prior network model includes a dynamic knowledge graph network, a prior knowledge network, and a decoder;

[0027] In the dynamic knowledge graph network, a training dynamic knowledge graph is obtained according to the radiology training image, and a training visual representation is obtained according to the training dynamic knowledge graph;

[0028] Performing contrastive learning based on the trained visual representation to train the dynamic knowledge graph network;

[0029] Acquiring prior knowledge training information corresponding to the radiology training image, and obtaining a prior training representation vector using the prior knowledge network based on the prior knowledge training information and the training visual representation;

[0030] Inputting the training visual representation and the prior training representation vector into the decoder to train the prior knowledge network and the decoder;

[0031] After completing the multi-task training of the dynamic knowledge graph network, prior knowledge network and decoder, a trained dynamic prior network model is obtained.

[0032] In one implementation, in the dynamic knowledge graph network, obtaining a training dynamic knowledge graph based on the radiology training image, and obtaining a training visual representation based on the training dynamic knowledge graph include:

[0033] In the dynamic knowledge graph network, a pre-built report database is obtained, wherein the report database includes a plurality of standard reports, and the standard reports are randomly selected from the training data set;

[0034] Obtaining a report representation vector corresponding to each standard report in the report database, and obtaining an image representation vector corresponding to the radiology training image, and calculating a similarity between each report representation vector and the image representation vector;

[0035] Obtaining a plurality of training standard reports matching the radiology training images according to the respective similarities;

[0036] Obtaining all entities in the training standard report, and generating a training image knowledge graph corresponding to the radiology training image based on the entities in the training standard report;

[0037] Obtain a pre-built medical knowledge graph, and add new paths to the medical knowledge graph based on the training image knowledge graph to obtain a training dynamic knowledge graph;

[0038] The dynamic knowledge graph is encoded using a dynamic graph encoder in the dynamic knowledge graph network to obtain a training visual representation corresponding to the radiology training image.

[0039] In one implementation, performing contrastive learning based on the training visual representation to train the dynamic knowledge graph network includes:

[0040] Obtaining a pre-constructed image report contrast loss function and an image report matching loss function, and a radiology training report corresponding to the radiology training image;

[0041] Obtaining a representation vector corresponding to the radiology training report using a report encoder, and performing report image contrast learning based on the representation vector corresponding to the radiology training report and the training visual representation, taking the image report contrast loss function as a training target;

[0042] Inputting the training visual representation and the radiology training report into a multimodal encoder, performing report image pairing using the image report matching loss function as a training objective;

[0043] The dynamic knowledge graph network is trained through report image comparison learning and report image matching.

[0044] In one implementation, obtaining prior knowledge training information corresponding to the radiological training image, and obtaining a prior training representation vector using the prior knowledge network based on the prior knowledge training information and the training visual representation, includes:

[0045] Using all the training standard reports and the medical knowledge graph as prior knowledge training information;

[0046] generating a first training key vector and a first training value vector corresponding to all the training standard reports based on a report encoder, using the training visual representation as a training query vector through a cross-attention mechanism, and performing a cross-attention calculation on the training query vector, the first training key vector, and the first training value vector to obtain a first training representation vector;

[0047] generating a second key vector and a second value vector corresponding to the entity in the medical knowledge graph based on the report encoder, using the visual representation as a training query vector through a cross-attention mechanism, and performing a cross-attention calculation on the training query vector, the second key vector and the second value vector to obtain a second training representation vector;

[0048] The first training representation vector and the second training representation vector are combined to obtain a priori training representation vector.

[0049] The present invention also provides a radiology report generating device, comprising:

[0050] An input module, configured to input a radiological image to be processed into a pre-trained dynamic priori network model, wherein the dynamic priori network model includes a dynamic knowledge graph network, a priori knowledge network, and a decoder;

[0051] a dynamic network module, configured to obtain a dynamic knowledge graph based on the radiological image and obtain a visual representation based on the dynamic knowledge graph in the dynamic knowledge graph network;

[0052] a priori network module, configured to obtain priori knowledge information corresponding to the radiological image, and obtain a priori representation vector using the priori knowledge network based on the priori knowledge information and the visual representation;

[0053] A generating module is configured to input the visual representation and the prior representation vector into the decoder to generate a radiology report corresponding to the radiology image.

[0054] The present invention also provides a terminal comprising: a memory, a processor, and a radiology report generation program stored in the memory and executable on the processor, wherein the radiology report generation program implements the steps of the radiology report generation method as described above when executed by the processor.

[0055] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program can be executed to implement the steps of the radiology report generating method as described above.

[0056] The present invention provides a radiology report generation method, device, terminal and storage medium, the method comprising: inputting a radiology image to be processed into a pre-trained dynamic priori network model, the dynamic priori network model comprising a dynamic knowledge graph network, a priori knowledge network and a decoder; in the dynamic knowledge graph network, obtaining a dynamic knowledge graph based on the radiology image, and obtaining a visual representation based on the dynamic knowledge graph; obtaining prior knowledge information corresponding to the radiology image, and obtaining a priori representation vector using the prior knowledge network based on the prior knowledge information and the visual representation; inputting the visual representation and the priori representation vector into the decoder to generate a radiology report corresponding to the radiology image. The present invention introduces a dynamic knowledge graph network and a priori knowledge network into the model, thereby realizing the generation of radiology reports by combining the dynamic knowledge graph and priori knowledge information, so that the model can better handle visual and textual biases caused by limited data availability, can reduce the bias of textual data, thereby solving the problem of inconsistency between visual and textual data when generating radiology reports, and improving the quality of report generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] FIG1 is a flow chart of a preferred embodiment of a method for generating a radiology report according to the present invention.

[0058] FIG2 is a schematic diagram showing the structural principle of the dynamic priori network model of the present invention.

[0059] FIG3 is a medical knowledge graph in the present invention.

[0060] FIG4 is a functional block diagram of a preferred embodiment of a radiology report generating device according to the present invention.

[0061] FIG5 is a functional block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0063] This invention aims to generate radiology reports from images. Unlike standard image description generation tasks, radiology report generation faces more significant image and text mismatches due to limited available data and a greater reliance on prior knowledge. This invention introduces a dynamic knowledge graph and prior knowledge to enhance the intelligence of the model, and introduces image-text comparison and image-text matching modules to improve the quality of generated results. Therefore, the main goal of this invention is to improve the performance and quality of existing radiology report generation, particularly with respect to key metrics.

[0064] Please refer to Figure 1, which is a flow chart of the radiology report generation method of the present invention. As shown in Figure 1, the radiology report generation method according to the embodiment of the present invention includes:

[0065] Step S100: Input the radiological image to be processed into a pre-trained dynamic priori network model, wherein the dynamic priori network model includes a dynamic knowledge graph network, a priori knowledge network and a decoder.

[0066] The dynamic prior network model of the embodiment of the present application utilizes a dynamic knowledge graph network and a prior knowledge network, so that the dynamic knowledge graph and prior knowledge are utilized in the report generation method, that is, the graph network and medical field knowledge are utilized, thereby improving the intelligence of the model.

[0067] As shown in FIG1 , the radiology report generation method according to this embodiment further includes:

[0068] Step S200: In the dynamic knowledge graph network, a dynamic knowledge graph is obtained based on the radiological image, and a visual representation is obtained based on the dynamic knowledge graph.

[0069] In the embodiment of the present application, step S200 includes:

[0070] Step S210: obtaining, in the dynamic knowledge graph network, a number of target standard reports matching the radiological image based on a pre-built report database;

[0071] Step S220: Acquire all entities in the target standard report, and generate an image knowledge graph corresponding to the radiological image based on the entities in the target standard report;

[0072] Step S230: Obtain a pre-built medical knowledge graph, and add new paths to the medical knowledge graph based on the image knowledge graph to obtain a dynamic knowledge graph;

[0073] Step S240: Encode the dynamic knowledge graph using the dynamic graph encoder in the dynamic knowledge graph network to obtain a visual representation corresponding to the radiological image.

[0074] Specifically, the report database can be standard reports in a training dataset and can be stored in a report queue. A report representation vector corresponding to each standard report in the report database is obtained, as well as an image representation vector corresponding to the radiological image. Similarities between each report representation vector and the image representation vector are calculated, and based on the respective similarities, several target standard reports matching the radiological image are obtained. In one embodiment, the similarities are sorted in descending order, and the top K standard reports are selected as target standard reports.

[0075] The embodiments of the present application utilize a dynamic knowledge graph to better capture information and associations in radiological images, thereby improving the quality of radiological report generation; and, the embodiments of the present application establish an adaptive graph network that can flexibly process different radiological images, thereby improving the generalization ability of the model.

[0076] As shown in FIG1 , the radiology report generation method according to this embodiment further includes:

[0077] Step S300: Acquire prior knowledge information corresponding to the radiological image, and obtain a prior representation vector using the prior knowledge network based on the prior knowledge information and the visual representation.

[0078] In the embodiment of the present application, step S300 specifically includes:

[0079] Step S310: taking all the target standard reports and the medical knowledge graph as prior knowledge information;

[0080] Step S320: Generate a first key vector and a first value vector corresponding to all the target standard reports based on the report encoder, use the visual representation as a query vector through a cross-attention mechanism, perform a cross-attention calculation on the query vector, the first key vector, and the first value vector to obtain a first representation vector;

[0081] Step S330: Generate a second key vector and a second value vector corresponding to the entity in the medical knowledge graph based on the report encoder, use the visual representation as a query vector through a cross-attention mechanism, perform cross-attention calculation on the query vector, the second key vector, and the second value vector to obtain a second representation vector;

[0082] Step S340: Merge the first representation vector and the second representation vector to obtain a priori representation vector.

[0083] The embodiments of this application can integrate prior knowledge. By introducing prior knowledge, the model can better handle visual and textual biases caused by limited data availability, improving the model's adaptability to different scenarios. Furthermore, by integrating prior knowledge, biases in textual data can be mitigated, thereby better generating radiology reports relevant to abnormal areas.

[0084] As shown in FIG1 , the radiology report generation method according to this embodiment further includes:

[0085] Step S400: Input the visual representation and the prior representation vector into the decoder to generate a radiology report corresponding to the radiology image.

[0086] In an embodiment of the present application, the training steps of the dynamic priori network model include:

[0087] Step A100: Acquire a training data set, wherein the training data set includes: one-to-one corresponding radiology training images and radiology training reports;

[0088] Step A200: constructing an initial dynamic priori network model, inputting radiological training images into the initial dynamic priori network model, wherein the initial dynamic priori network model includes a dynamic knowledge graph network, a priori knowledge network, and a decoder;

[0089] Step A300: In the dynamic knowledge graph network, obtaining a training dynamic knowledge graph based on the radiology training image, and obtaining a training visual representation based on the training dynamic knowledge graph;

[0090] Step A400: performing comparative learning based on the training visual representation to train the dynamic knowledge graph network;

[0091] Step A500: obtaining prior knowledge training information corresponding to the radiology training image, and obtaining a prior training representation vector using the prior knowledge network based on the prior knowledge training information and the training visual representation;

[0092] Step A600: inputting the training visual representation and the prior training representation vector into the decoder to train the prior knowledge network and the decoder;

[0093] Step A700: After completing the multi-task training of the dynamic knowledge graph network, the prior knowledge network and the decoder, a trained dynamic prior network model is obtained.

[0094] Specifically, the embodiment of the present application proposes a unified model: the dynamic prior network model DPN (Dynamics Priori Networks For Radiology Report Generation). DPN introduces dynamic graph networks (DGN), contrastive learning (CL) and prior knowledge networks (PrKN, Prior Knowledge Networks). Specifically, the embodiment of the present application introduces a dynamic prior network model to align abnormal areas of radiological images and abnormal radiological report topics; introduces contrastive learning, including comparing and matching radiological images and radiological reports to improve visual and textual representations; introduces prior knowledge networks, including prior work experience (PrWE, Prior Working Experience) and prior medical knowledge (PrMK, Prior Medical Knowledge) extracted from the training corpus. Finally, after receiving information from the dynamic graph network and the prior knowledge network, a decoder is used to generate the final radiology report.

[0095] Regarding evaluation metrics, radiology report generation differs from natural image description generation in two key aspects: first, the accuracy of positive disease keywords in radiology image reports is crucial, compared to the equal importance of each word in natural image captions; second, evaluating report quality should emphasize matching disease keywords and their related attributes rather than counting N-gram occurrences. Experimental results on two popular radiology report datasets, IU X-Ray and MIMIC-CXR, show that the proposed method outperforms previous models in terms of language generation metrics (CIDEr). Experimental results on two popular radiology report datasets show that DPN outperforms previous models in terms of language generation metrics and produces reports that contain basic medical terms and meaningful mappings between images and reports.

[0096] For the training data set, the embodiments of the present application can use two well-validated benchmark data sets, one is the IU X-RAY data set from Indiana University, and the other is the MIMIC-CXR data set from Beth Israel Deaconess Medical Center. The IU X-RAY data set is relatively small, containing 7470 chest X-ray images and 3955 corresponding reports. In contrast, the MIMIC-CXR data set is the largest publicly available radiology data set, containing 473057 chest X-ray images and 206563 reports. The present invention focuses on generating the "findings" part of the report and excludes samples that do not have this part in IU X-RAY and MIMIC-CXR. For IU X-RAY, a data partition is set (70% for training, 10% for validation, and 20% for testing). For MIMIC-CXR, its official data partition is used.

[0097] Regarding the model architecture, generating radiology reports involves the task of converting images into text. When a radiology image is represented by I, the goal is to generate a descriptive radiology report represented as R = {y1, y2, ..., y t}. As shown in Figure 2, the dynamic prior network model introduces a dynamic knowledge graph network (DGN), contrastive learning (CL) and a prior knowledge network (PrKN). Specifically, a dynamic knowledge graph network is introduced to align abnormal areas of radiological images and abnormal topics of medical knowledge graphs; contrastive learning is introduced, including contrasting and matching the relationship between radiological images and radiological reports to improve visual and textual representations; a prior knowledge network is introduced, which includes two parts, one of which is prior work experience (PrWE) and prior medical knowledge (PrMK) extracted from the training corpus. Finally, after receiving the image representation (vector) of the dynamic knowledge graph and the representation of the prior knowledge network, the decoder is used to generate the final radiology report. The whole process is recorded as:

[0098] The embodiments of the present application utilize Transformer as the core architecture to generate smooth and accurate radiology reports.

[0099] I represents the image representation (vector). The present invention uses ResNet-152 to extract 2,048 image feature maps of size 7×7, which are then converted to 512 feature maps of size 7×7. This conversion produces

[0100] R stands for radiology report. Each examination of a patient will include multiple radiology images and a radiology report. The present invention uses SciBert to obtain the report and embed it into R. i ∈R d .

[0101] R topic The medical knowledge graph represents a medical knowledge graph. Its primary purpose is to highlight disease keywords and strengthen the connections between them. As shown in Figure 3, the medical knowledge graph includes 20 entities: normal, other findings, foreign body, cardiomegaly, scoliosis, fracture, hernia, calcification, exudate, thickening, pneumothorax, pulmonary parenchymal lesions, hypoventilation, emphysema, pneumonia, edema, atelectasis, scar, turbidity, and lesion. The present invention uses the entities in the knowledge graph, i.e., disease keywords, as input.

[0102] In the embodiment of the present application, step A300 specifically includes:

[0103] Step A310: Obtain a pre-built report database in the dynamic knowledge graph network, wherein the report database includes a plurality of standard reports, and the standard reports are randomly selected from the training data set;

[0104] Step A320: Obtain a report representation vector corresponding to each standard report in the report database, and obtain an image representation vector corresponding to the radiology training image, and calculate the similarity between each report representation vector and the image representation vector;

[0105] Step A330: obtaining a plurality of training standard reports matching the radiology training images according to the respective similarities;

[0106] Step A340: Obtain all entities in the training standard report, and generate a training image knowledge graph corresponding to the radiology training image based on the entities in the training standard report;

[0107] Step A350: Obtain a pre-built medical knowledge graph, and add new paths to the medical knowledge graph based on the training image knowledge graph to obtain a training dynamic knowledge graph;

[0108] Step A360: Encode the dynamic knowledge graph using the dynamic graph encoder in the dynamic knowledge graph network to obtain a training visual representation corresponding to the radiology training image.

[0109] Specifically, in order to align radiology images I and radiology reports R in a dynamic knowledge graph network, first, by evaluating the radiology image features f IBased on the similarity with the radiology report representation, the top T reports are retrieved from the report database, which can be a report queue formed by n standard reports randomly selected from the training dataset. For the retrieved reports, the entities in the reports are obtained, and these entities generate a knowledge graph and are stored in the form of triples (entity, relationship, entity). Each triple describes the relationship between the source entity and the target entity. This knowledge graph is used to modify the original medical knowledge graph shown in Figure 3. These modifications include adding entity nodes that are not involved in the knowledge graph into the medical knowledge graph and connecting unconnected entities. In this way, the medical knowledge graph is dynamically updated. In order to encode the dynamic medical knowledge graph, a dynamic knowledge graph encoder is introduced. The dynamic knowledge graph encoder is built based on the Transformer encoder.

[0110] This paper uses relational self-attention (RSA) to model the graph structure of dynamic graphs. Specifically, a controllable mask matrix is ​​used to model the knowledge graph, which ensures that each node in the graph can only affect its connected nodes and strengthen the relationships between them. The paper uses pre-trained SciBert as a report encoder to obtain word embeddings for entities and uses graph attention (cross-attention) to merge the dynamic knowledge graph with visual features.

[0111] In the embodiment of the present application, step A400 specifically includes:

[0112] Step A410: Obtain a pre-constructed image report contrast loss function and an image report matching loss function, as well as a radiology training report corresponding to the radiology training image;

[0113] Step A420: using a report encoder to obtain a representation vector corresponding to the radiology training report, and performing report image contrast learning based on the representation vector corresponding to the radiology training report and the training visual representation, using the image report contrast loss function as a training target;

[0114] Step A430: Input the training visual representation and the radiology training report into a multimodal encoder, and perform report image pairing using the image report matching loss function as a training target;

[0115] Step A440: Complete the training of the dynamic knowledge graph network through report image comparison learning and report image matching.

[0116] Specifically, the contrastive learning part includes the image-report contrastive loss (IRC) and the image-report matching loss (IRM), which can effectively enhance visual and textual representations. The training objective of the image-report contrastive loss (IRC) is to maximize the similarity between paired images and reports, while minimizing the similarity between unpaired text and images.

[0117] The image-report matching loss (IRM) involves a binary classification task: classifying an input image-report pair as paired or unpaired. For this task, a multimodal encoder is used to implement multimodal representations through a cross-attention mechanism. The resulting multimodal representation vector is then projected to a vector of dimension 2 through a linear layer to perform the binary classification task.

[0118] The embodiment of the present application introduces an image-text comparison module and an image-text matching module, which helps to better integrate visual and text information and improve the quality of generated results.

[0119] In the embodiment of the present application, step A500 specifically includes:

[0120] Step A510: using all the training standard reports and the medical knowledge graph as prior knowledge training information;

[0121] Step A520: Generate a first training key vector and a first training value vector corresponding to all the training standard reports based on the report encoder, use the training visual representation as a training query vector through a cross-attention mechanism, perform a cross-attention calculation on the training query vector, the first training key vector, and the first training value vector to obtain a first training representation vector;

[0122] Step A530: Generate a second key vector and a second value vector corresponding to the entity in the medical knowledge graph based on the report encoder, use the visual representation as a training query vector through a cross-attention mechanism, perform cross-attention calculation on the training query vector, the second key vector and the second value vector, and obtain a second training representation vector;

[0123] Step A540: Merge the first training representation vector and the second training representation vector to obtain a priori training representation vector.

[0124] Specifically, the prior knowledge network consists of two components: prior work experience and prior medical knowledge. For prior work experience, cross-attention is calculated by taking the visual representation as the query vector through cross-attention, and then generating the key vector and value vector from the training standard report (i.e., the extracted report). For prior medical experience, cross-attention is calculated by taking the visual representation as the query vector through cross-attention, and then generating the key and value vectors from the 20 entities of the medical knowledge graph. The two representation vectors of aggregated prior work experience and prior medical knowledge are added together with the visual representation and added to the decoder through cross-attention.

[0125] The decoder inputs the dynamic knowledge graph and the vector of the prior knowledge network to generate the final radiology report. The decoder can use the standard Transformer decoder.

[0126] Table 1 shows the performance comparison of DPN and other state-of-the-art methods on the MIMIC-CXR and IU-Xray datasets. DPN outperforms previous models in most key indicators, demonstrating the effectiveness of our method.

[0127] Table 1

[0128] The performance of the model in this embodiment is evaluated using standard natural language generation (NLG) metrics. These NLG metrics include:

[0129] First, Consistency-Based Image Description Evaluation (CIDEr): CIDEr is a novel metric for evaluating image descriptions that relies on human consistency. It quantifies the cosine similarity between n-gram TF-IDF representations in two captions, covering the range from unigram to 4-gram. The average similarity score is used as the final evaluation metric. CIDEr rewards the presence of key terms while penalizing frequently occurring terms. In the radiology report generation task, CIDEr is particularly relevant because it evaluates the model's ability to generate meaningful and informative descriptions, making it a key metric for this task.

[0130] Second, the Translation Evaluation Metric with Explicit Ranking (METEOR): METEOR considers precision and recall across the entire text corpus. It uses stemming techniques, including the Porter stemmer and the WordNet ontology, to identify synonyms within text. METEOR shows strong correlation with human judgments of translation quality, encompassing word-level, sentence-level, and sub-paragraph-level assessments.

[0131] Third, Bilingual Evaluation Underscore (BLEU): BLEU is a commonly used metric in machine translation tasks. It measures the word n-gram overlap between the model's predictions and the reference text. Notably, due to textual bias in radiology report generation datasets, even a system that simply repeats the most common sentences can achieve a respectable BLEU score. BLEU primarily emphasizes precision, focusing on the match between the n-grams in the candidate translation and the n-grams in the reference translation. It evaluates translation accuracy by measuring the word overlap between the generated text and the reference text. It is primarily used as a reference metric.

[0132] Fourth, Recall-Oriented Assistant for Summarization Evaluation (ROUGE-L): ROUGE-L calculates the ratio of the length of the longest common subsequence between machine-generated text and reference text. It evaluates the recall aspect of similarity between texts.

[0133] These NLG metrics automatically evaluate the accuracy of new models by comparing the similarities and differences between generated captions and descriptions provided by radiologists. Achieving higher scores in these metrics indicates improved model performance.

[0134] In one implementation, the experimental environment for the embodiments of this application can be conducted using the PyTorch framework, version 1.10.1, running on a high-performance computing server using the Ubuntu SMP 16.04.1 operating system. The server is equipped with a 10-core Intel Xeon Silver 4210 processor and an NVIDIA RTX8000 GPU with 48GB of memory. Python 3.6.10 was used as the programming language for these experiments.

[0135] The proposed model demonstrates excellent performance on two widely used datasets, IU-Xray and MIMIC-CXR, particularly across key metrics, demonstrating its robust performance. Advantages of the proposed model include its full utilization of dynamic knowledge graphs and prior knowledge, as well as its ability to combat visual and textual biases, resulting in excellent performance on radiology report generation tasks.

[0136] In one embodiment, as shown in FIG4 , based on the above-mentioned radiology report generation method, the present invention further provides a radiology report generation device, comprising:

[0137] An input module 100 is used to input a radiological image to be processed into a pre-trained dynamic priori network model, wherein the dynamic priori network model includes a dynamic knowledge graph network, a priori knowledge network, and a decoder;

[0138] A dynamic network module 200 is configured to obtain a dynamic knowledge graph based on the radiological image and obtain a visual representation based on the dynamic knowledge graph in the dynamic knowledge graph network;

[0139] A priori network module 300 is configured to obtain priori knowledge information corresponding to the radiological image, and obtain a priori representation vector using the priori knowledge network based on the priori knowledge information and the visual representation;

[0140] The generating module 400 is configured to input the visual representation and the prior representation vector into the decoder to generate a radiology report corresponding to the radiology image.

[0141] FIG5 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include:

[0142] Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .

[0143] When the processor 502 executes the program, the radiology report generation method provided in the above embodiment is implemented.

[0144] Furthermore, the terminal further includes:

[0145] The communication interface 503 is used for communication between the memory 501 and the processor 502 .

[0146] The memory 501 is used to store computer programs that can be run on the processor 502 .

[0147] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0148] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be interconnected via a bus to enable communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the figure shows only one line, but this does not mean that there is only one bus or only one type of bus.

[0149] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.

[0150] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0151] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned radiology report generation method when executed by a processor.

[0152] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0153] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0154] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes additional implementations in which the order shown or discussed may not be followed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by a person skilled in the art to which the embodiments of the present application belong.

[0155] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can read and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting, or otherwise processing in a suitable manner as necessary, and then storing it in a computer memory.

[0156] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0157] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0158] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0159] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

[0160] In summary, the present invention discloses a method, device, terminal and storage medium for generating a radiology report, the method comprising: inputting a radiology image to be processed into a pre-trained dynamic priori network model, the dynamic priori network model comprising a dynamic knowledge graph network, a priori knowledge network and a decoder; in the dynamic knowledge graph network, obtaining a dynamic knowledge graph based on the radiology image, and obtaining a visual representation based on the dynamic knowledge graph; obtaining prior knowledge information corresponding to the radiology image, and obtaining a priori representation vector using the prior knowledge network based on the prior knowledge information and the visual representation; inputting the visual representation and the prior representation vector into the decoder to generate a radiology report corresponding to the radiology image. The present invention introduces a dynamic knowledge graph network and a priori knowledge network into the model, thereby realizing the generation of a radiology report by combining the dynamic knowledge graph and the priori knowledge information, so that the model can better handle visual and textual biases caused by limited data availability, can reduce the bias of textual data, thereby solving the problem of inconsistency between visual and textual data when generating radiology reports, and improving the quality of report generation.

[0161] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A radiology report generation method, characterized in that, The method includes: Inputting the radiological image to be processed into a pre-trained dynamic prior network model, where the dynamic prior network model includes a dynamic knowledge graph network, a prior knowledge network, and a decoder; In the dynamic knowledge graph network, obtaining a dynamic knowledge graph according to the radiological image, and obtaining a visual representation according to the dynamic knowledge graph; Obtaining the prior knowledge information corresponding to the radiological image, and using the prior knowledge network to obtain a prior representation vector based on the prior knowledge information and the visual representation; Inputting the visual representation and the prior representation vector into the decoder to generate a radiological report corresponding to the radiological image.

2. The radiology report generation method according to claim 1, wherein, In the dynamic knowledge graph network, obtaining a dynamic knowledge graph according to the radiological image, and obtaining a visual representation according to the dynamic knowledge graph, includes: In the dynamic knowledge graph network, obtaining a number of target standard reports matching the radiological image based on a pre-constructed report database; Obtaining the entities in all the target standard reports, and generating an image knowledge graph corresponding to the radiological image based on the entities in the target standard reports; Obtaining a pre-constructed medical knowledge graph, adding new paths to the medical knowledge graph based on the image knowledge graph to obtain a dynamic knowledge graph; Encoding the dynamic knowledge graph using a dynamic graph encoder in the dynamic knowledge graph network to obtain a visual representation corresponding to the radiological image.

3. The radiology report generation method according to claim 2, wherein Obtaining the prior knowledge information corresponding to the radiological image, and using the prior knowledge network to obtain a prior representation vector based on the prior knowledge information and the visual representation, includes: Taking all the target standard reports and the medical knowledge graph as prior knowledge information; Generating a first key vector and a first value vector corresponding to all the target standard reports based on a report encoder, performing cross-attention calculation on the query vector, the first key vector, and the first value vector by using the visual representation as the query vector to obtain a first representation vector; Generating a second key vector and a second value vector corresponding to the entities in the medical knowledge graph based on the report encoder, performing cross-attention calculation on the query vector, the second key vector, and the second value vector by using the visual representation as the query vector to obtain a second representation vector; Combining the first representation vector and the second representation vector to obtain a prior representation vector.

4. The radiology report generation method according to claim 1, wherein The training steps of the dynamic prior network model include: Obtaining a training data set, where the training data set includes: radiological training images and radiological training reports in one-to-one correspondence; Constructing an initial dynamic prior network model, and inputting the radiological training images into the initial dynamic prior network model, where the initial dynamic prior network model includes a dynamic knowledge graph network, a prior knowledge network, and a decoder; In the dynamic knowledge graph network, obtaining a training dynamic knowledge graph according to the radiological training image, and obtaining a training visual representation according to the training dynamic knowledge graph; Performing contrastive learning based on the training visual representation to train the dynamic knowledge graph network; Obtain the prior knowledge training information corresponding to the radiology training image, and based on the prior knowledge training information and the training visual representation, use the prior knowledge network to obtain a prior training representation vector; Input the training visual representation and the prior training representation vector into the decoder, and train the prior knowledge network and the decoder; After completing the multi-task training of the dynamic knowledge graph network, the prior knowledge network and the decoder, obtain a trained dynamic prior network model.

5. The radiology report generation method according to claim 4, wherein In the dynamic knowledge graph network, obtain a training dynamic knowledge graph according to the radiology training image, and obtain a training visual representation according to the training dynamic knowledge graph, including: In the dynamic knowledge graph network, obtain a pre-constructed report database, where the report database includes a number of standard reports, and the standard reports are randomly selected from the training dataset; Obtain the report representation vectors corresponding to each standard report in the report database, and obtain the image representation vector corresponding to the radiology training image, and calculate the similarity between each report representation vector and the image representation vector; Obtain a number of training standard reports that match the radiology training image according to each similarity; Obtain the entities in all the training standard reports, and generate a training image knowledge graph corresponding to the radiology training image based on the entities in the training standard reports; Obtain a pre-constructed medical knowledge graph, and add new paths to the medical knowledge graph based on the training image knowledge graph to obtain a training dynamic knowledge graph; Use the dynamic graph encoder in the dynamic knowledge graph network to encode the dynamic knowledge graph to obtain a training visual representation corresponding to the radiology training image.

6. The radiology report generation method according to claim 4, wherein Perform contrastive learning based on the training visual representation to train the dynamic knowledge graph network, including: Obtain a pre-constructed image-report contrast loss function and an image-report matching loss function, and the radiology training report corresponding to the radiology training image; Use the report encoder to obtain the representation vector corresponding to the radiology training report, and based on the representation vector corresponding to the radiology training report and the training visual representation, use the image-report contrast loss function as the training target to perform report-image contrastive learning; Input the training visual representation and the radiology training report into the multi-modal encoder, and use the image-report matching loss function as the training target to perform report-image pairing; Complete the training of the dynamic knowledge graph network through report-image contrastive learning and report-image pairing.

7. The radiology report generation method according to claim 5, wherein Obtain the prior knowledge training information corresponding to the radiology training image, and based on the prior knowledge training information and the training visual representation, use the prior knowledge network to obtain a prior training representation vector, including: Use all the training standard reports and the medical knowledge graph as prior knowledge training information; Generate a first training key vector and a first training value vector corresponding to all the training standard reports based on the report encoder. Use the cross-attention mechanism to take the training visual representation as the training query vector, and perform cross-attention calculation on the training query vector, the first training key vector, and the first training value vector to obtain a first training representation vector; Generate a second key vector and a second value vector corresponding to the entities in the medical knowledge graph based on the report encoder. Use the cross-attention mechanism to take the visual representation as the training query vector, and perform cross-attention calculation on the training query vector, the second key vector, and the second value vector to obtain a second training representation vector; Merge the first training representation vector and the second training representation vector to obtain a prior training representation vector.

8. A radiology report generation device, characterized in that, The device includes: An input module for inputting the radiological image to be processed into a pre-trained dynamic prior network model, where the dynamic prior network model includes a dynamic knowledge graph network, a prior knowledge network, and a decoder; A dynamic network module for obtaining a dynamic knowledge graph based on the radiological image in the dynamic knowledge graph network, and obtaining a visual representation based on the dynamic knowledge graph; A prior network module for obtaining prior knowledge information corresponding to the radiological image, and obtaining a prior representation vector using the prior knowledge network based on the prior knowledge information and the visual representation; A generation module for inputting the visual representation and the prior representation vector into the decoder to generate a radiological report corresponding to the radiological image.

9. A terminal, characterized in that, Including: A memory, a processor, and a radiological report generation program stored on the memory and executable on the processor. When the radiological report generation program is executed by the processor, it implements the steps of the radiological report generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the radiological report generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image report generation method and device based on multi-modal fusion

    CN115331769A

  • PET / CT image report conclusion auxiliary generation method and device based on knowledge graph

    CN115910263A

  • Automatic generation of medical imaging reports based on fine grained finding labels

    US11244755B1

Cited By

  • Radiology report generation method based on comparative learning and adaptive knowledge integration

    CN120809049A

  • Electric power data blood relationship abnormity intelligent early warning method and system based on knowledge graph

    CN120910767A

  • A knowledge graph-based power data blood relationship anomaly intelligent early warning method and system

    CN120910767B

  • Digital fusion management method and system based on large medical model

    CN121171554A

  • Radiology report generation method based on enhanced bimodal features

    CN121545660A