Chest radiation medical report generation method based on mapping knowledge domain

Through the dual-path gated knowledge enhancement method based on knowledge graph, the problems of difficulty in fusion of multimodal data and insufficient prediction of rare diseases in the medical report generation model are solved, and high-quality medical report generation, especially accurate description of rare diseases are achieved, improving the fine-grainedness and interpretability of the report.

CN120376026APending Publication Date: 2025-07-25CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510438023.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing medical report generation model has difficulty in multimodal data association and fusion, rough report generation, insufficient prediction of rare diseases, and long-tail problems, resulting in inaccuracy and inefficient diagnosis.

Method used

Using a dual-path gated knowledge enhancement method based on knowledge graph, network parameters are optimized to generate high-quality medical reports by constructing a fine-grained disease-space-organ tag model, combining text, image and knowledge features.

Benefits of technology

It improves the quality and accuracy of medical reports, especially the description coverage of rare diseases, alleviates the bias generation problems caused by data bias, and improves the fine-grainedness and interpretability of reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376026A_ABST
    Figure CN120376026A_ABST
Patent Text Reader

Abstract

The invention discloses a chest radiology medical report generation method based on a knowledge graph, and the method comprises the steps: obtaining a to-be-processed chest radiology image, inputting the to-be-processed chest radiology image into a trained medical report generation network, and obtaining a chest radiology medical report. The training process of the network comprises the following steps: acquiring a chest radiology medical report set; constructing a chest radiology knowledge graph; extracting text features, image features and knowledge features from the chest radiology medical report, inputting the text features, the image features and the knowledge features into a double-path knowledge enhancement gating module, aligning the text features and the image features, aligning the image features and the knowledge features, and calculating alignment loss; fusing the aligned features to obtain enhanced fusion features; inputting the enhanced fusion features into a classifier, and calculating classification loss; the fusion features are input into a decoder, a chest radiology medical report is generated, and generation loss is calculated; and network parameters are optimized through loss back propagation. According to the invention, the medical report generation quality and accuracy can be obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a method for generating a chest radiology report based on a knowledge graph. Background Art

[0002] With the rapid development of medical imaging technology, radiologists need to process a large amount of complex imaging data every day, such as X-ray films, CT scans, MRI, etc. These imaging data are not only huge in quantity, but also the interpretation process is cumbersome, requiring doctors to have profound professional knowledge and rich clinical experience. However, due to the heavy workload and limited time of doctors, they are prone to fatigue, which may lead to misdiagnosis or missed diagnosis. Therefore, how to reduce the work pressure of doctors while improving the diagnostic efficiency and accuracy has become an important issue in the medical field. Automatically generating a disease diagnosis report for medical images can not only effectively reduce the burden on doctors, but also ensure the standardization and accuracy of the report, which is of great significance for improving the efficiency and quality of clinical diagnosis. In recent years, the breakthroughs in deep learning in the fields of computer vision and natural language processing have provided new technical means for the automatic generation of medical image reports. Traditional methods for generating medical image reports mostly rely on rules and templates, with relatively fixed content and lack of flexibility. While the end-to-end generation model based on deep learning can capture the complex mapping relationship between images and texts by learning a large amount of medical image and diagnosis report data, and then generate more natural and accurate reports.

[0003] However, although deep learning methods have made significant progress in the field of medical report generation, and the current visual language alignment methods have also achieved a certain degree of mapping between images and texts to some extent, there are still many problems. First, there are significant differences in feature representations between medical images and text data, resulting in a semantic gap when fusing cross-modal data, which affects the integration and utilization of information. Although different modal data are complementary, existing models still have difficulties in mining these internal correlations, and traditional fusion methods have not deeply explored the connections between modalities. Second, medical report generation models usually have a coarsening problem, and the generated reports lack refined disease descriptions and fine-grained classifications. Although some works have improved this situation to a certain extent by means of knowledge graphs, the existing knowledge graphs cover a limited disease spectrum. For example, the widely used chest X-ray atlas only covers 7 types of organs and 18 common diseases, resulting in the model lacking the representation ability for common lesions such as "hilar mass" and "vertebral osteoporosis" and anatomical variations such as "hilar shadow", and the current knowledge graphs do not explain several rare diseases or abnormalities such as "scar" and "interstitial opacity", limiting the clinical depth. Finally, the disease data distribution shows a long-tail effect, with few medical report samples for rare diseases and insufficient representativeness, resulting in a serious "biased generation" problem, which affects its overall diagnostic ability. For example Figure 1As shown, in all reports in the existing medical dataset, the frequency of sentences indicating normal results (no disease or abnormality) is three times that of sentences indicating the presence of at least one disease or abnormality. In addition, the number of sentences containing common diseases (appearing more than 50 times) is almost four times that of sentences containing rare diseases (appearing less than 50 times). Therefore, how to effectively integrate multi-modal data, improve the fineness and interpretability of report generation, and solve the long-tail problem are important challenges in current medical data analysis. Summary of the Invention

[0004] Objective of the present invention: To propose a medical report generation algorithm framework based on dual-path gated knowledge enhancement, and solve problems such as difficulties in associative fusion of multi-modal data and coarsening phenomenon in existing medical report generation models, especially deficiencies in rare disease prediction.

[0005] The present invention proposes a method for generating chest radiology medical reports based on a knowledge graph, and the method includes:

[0006] Obtain the chest radiology image to be processed, input it into the trained medical report generation network, and obtain the chest radiology medical report;

[0007] The training process of the medical report generation network includes:

[0008] S1: Obtain a chest radiology medical report set, and use the chest radiology medical report as a training sample;

[0009] S2: Based on the chest radiology medical report set, construct a chest radiology knowledge graph;

[0010] S3: Respectively extract text features and image features from the training samples, and extract knowledge features from the training samples based on the chest radiology knowledge graph;

[0011] S4: Input the text features, image features, and knowledge features into the dual-path knowledge enhancement gated module. One branch aligns the text features and image features and calculates the image-text alignment loss; the other branch aligns the image features and knowledge features and calculates the image-knowledge alignment loss; fuse the output results of the two branches to obtain enhanced fusion features;

[0012] S5: Input the enhanced fusion features into a classifier to obtain a classification result, and calculate the classification loss Input the fusion features into a decoder to generate a chest radiology medical report, and calculate the generation loss between the generated medical report label and the true medical report label

[0013] S6: Through the image-text alignment loss Image-knowledge alignment loss Classification loss and generation loss Backpropagate to optimize network parameters.

[0014] Advantages of the present invention: The knowledge graph, rare disease enhancement strategy, dual-path knowledge enhancement gating module, and diagnostic classification module proposed by the present invention enable the DPKRG model to explicitly model medical knowledge by constructing "disease-space-organ" labels. The generated reports have been improved in terms of text quality and disease coverage, especially the description coverage of rare diseases has been significantly increased, successfully alleviating the "bias generation" problem caused by data bias in traditional models. The present invention can significantly improve the quality and accuracy of medical report generation. Description of the drawings

[0015] Figure 1 It is a pie chart showing the proportions of normal, rare disease, and common disease diagnostic results in the existing medical report dataset;

[0016] Figure 2 It is a schematic diagram of the overall process of the embodiment of the present invention;

[0017] Figure 3 It is a flowchart of the training steps of the medical report generation network in the embodiment of the present invention;

[0018] Figure 4 It is a comparison chart of the training sample set data before and after enhancement in the embodiment of the present invention;

[0019] Figure 5 It is a schematic diagram of the disease types defined in the chest radiology knowledge graph in the embodiment of the present invention. Detailed implementation manners

[0020] The terms "first", "second", "third", "fourth", etc. in the specification, claims, and above-mentioned drawings of this application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application.

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] The embodiment of the present invention provides a method for generating chest radiology medical reports based on a knowledge graph.

[0023] Figure 2 It is a schematic diagram of the overall process of an embodiment of the present invention. Figure 2 In this embodiment of the present invention, the overall process includes a network training stage and a network application stage. In the network training stage: obtain training samples, and perform data augmentation on the training samples on both sides; use a text editor to extract text features from the training samples, and use an image editor to extract image features from the training samples; combine with a chest radiology knowledge graph to extract knowledge features from the training samples; use two parallel branches for feature alignment. One uses a cross-attention mechanism to align text features and image features, and calculates the image-text alignment loss The other uses a cross-attention mechanism to align knowledge features and image features, and calculates the image-knowledge alignment loss Input the output results of the two branches (i.e., the aligned features) into a gated unit, adjust the gating weights, perform feature fusion to obtain fused features; input the fused features into a classifier to obtain classification results, and calculate the classification loss Input the fused features into a decoder to generate a chest radiology medical report, and calculate the cross-entropy loss Through the image-text alignment loss Image-knowledge alignment loss Classification loss And cross-entropy loss Perform backpropagation to optimize the network parameters. After the network training is completed, fix the network parameters. The trained medical report generation network only includes a trained image encoder and a trained decoder. In the network application stage, obtain the chest radiograph to be processed and input it into the trained medical report generation network to generate a chest radiology medical report.

[0024] Refer to Figure 2 As shown, the method for generating a chest radiology medical report based on a knowledge graph includes: obtaining a chest radiograph to be processed, inputting it into the trained medical report generation network, and obtaining a chest radiology medical report.

[0025] Figure 3 It is a flowchart of the training steps of the medical report generation network in an embodiment of the present invention.

[0026] Refer to Figure 3 As shown, the training process of the medical report generation network includes:

[0027] S1: Obtain a chest radiology medical report set, and use the chest radiology medical report as a training sample.

[0028] Collect chest radiology medical report data from open-source platforms or hospitals. The chest radiology medical reports include chest radiology images and report texts. For example, chest radiology images include: X-ray films, CT images, MRI images, or PET-CT images, etc. Clean the chest radiology medical report data to remove personal information or sensitive information in the chest radiology medical report data, and construct a training sample set. The chest radiology medical report text includes keywords and sentences describing lesion information and clinical diagnosis. Specifically, the content of the chest radiology medical report covers the findings and impressions of medical examinations, which are composed of text descriptions written by professional radiologists during routine clinical care, and the impression is a summary of the clinical diagnosis. These contents cover the disease types and infection areas closely related to the diagnosis conclusion.

[0029] Refer to Figure 1 As shown, in the chest radiology medical report set, the proportion of rare diseases is very low, and there will be a long-tail problem of chest radiology diseases. In the embodiment of the present invention, data augmentation is performed on the chest radiology medical report data set to expand the overall distribution of rare diseases in the data set.

[0030] In the illustrated embodiment, the specific process of the data augmentation includes:

[0031] S101: Label disease tags according to the chest radiology medical report data set.

[0032] S102: Extract text word vectors from the chest radiology medical report, and construct a sentence-level key-value pool, where the key represents keywords related to the disease, and the value represents sentences including keywords or disease tags.

[0033] Specifically, use BertTokenizer to convert the text data into text word vectors. It should be noted that BertTokenizer is a tokenizer designed based on the BERT (Bidirectional Encoder Representations from Transformers) model, and is mainly used for text preprocessing in natural language processing tasks.

[0034] For example, keywords related to the disease, such as "lung enlargement"; sentences including keywords or disease tags, such as sentences including keywords or disease tags such as "lung enlargement" and "This patient has lung enlargement"; the disease tags are obtained after being labeled based on expert knowledge.

[0035] S103: Define tag count, set the tag count range, query all sentences under each disease tag in the chest radiology medical report according to the disease tags, filter out the values including the qualified disease tags, replace the current value including the disease tag with the filtered value to obtain a new value, and obtain an enhanced report dataset.

[0036] Specifically, in the embodiment of the present invention, the tag count is defined as the number of sentences including the disease tag. The higher the tag count, the more variants of the sentences describing the disease tag. In the embodiment of the present invention, the tag count range [t down , t up is defined, and tags with a tag count less than t down or greater than t up are ignored because these parts of the diseases represent too few to be general or too many to require additional distribution expansion. Subsequently, starting from the disease tag with the least unique format in this tag count range, find all sentences under each disease tag, filter out the values including the qualified disease tags (i.e., sentences including keywords or disease tags), and replace the current value including the disease tag with the filtered value to obtain a new value. In this way, repeat this operation for all reports.

[0037] To replace the sentences in the report more accurately, before replacement, we will first determine whether the sentence to be replaced contains spatial location information S. If it contains, sentences containing a subset of S need to be selected from the key-value pool, and cannot be replaced randomly. Each sentence in each chest radiology medical report to be enhanced is enhanced Q times, which can be expressed as:

[0038] Q1 = 2t up / count(D i ) - 1

[0039] Q2 = count(D i ) - 1

[0040]

[0041] In the formula, Q1 represents the expected enhancement times, count(D i ) represents the size of the sentence quantity value when the key is disease D i , Q2 represents the maximum number of enhancements that can be affected by the number of sentences under each disease tag, and count(D i ) - 1 represents skipping the sentence to be enhanced itself in the key-value pool, and Q represents the actual enhancement times.

[0042] Exemplarily, t upWhen = 50, if there are 10 different sentences under a specific disease label, the proposed enhancement strategy needs to generate 9 additional reports for each sentence separately, and a total of 10×9 = 90 reports are generated additionally for this disease label. Given that the reports may contain multiple different diseases, this enhancement process may inadvertently increase the occurrence frequencies of multiple diseases at the same time. To improve this undesirable result, the embodiments of the present invention update the statistical data of the disease labels after each round of enhancement to find the next least common disease that has not been enhanced yet.

[0043] Figure 4 This is a comparison chart of the training sample set data before and after enhancement in the embodiments of the present invention.

[0044] Refer to Figure 4 As shown, the embodiments of the present invention have successfully balanced the distribution between common diseases and rare diseases through the rare disease enhancement strategy. The overall frequency of uncommon diseases in the original dataset has increased from 34.8% to 52.9%, while the overall frequency of common diseases has decreased from 65.2% to 47.1%; effectively solving the long-tail problem.

[0045] It should be noted that as a technology that represents information in a graph structure, the knowledge graph can effectively organize and express various complex knowledge. In the medical field, the knowledge graph can integrate various knowledge sources such as medical literature, clinical guidelines, and disease diagnoses into a graph to provide more comprehensive background knowledge for automatically generating clinical reports. Existing chest radiology knowledge graphs usually only contain 18 common disease categories under 7 organs, lacking comprehensiveness. For example, many common diseases are ignored; existing knowledge graphs do not explain several rare diseases or abnormalities such as "scar" and "interstitial opacity". Thus, using the existing chest radiology knowledge graph for the chest radiology report generation network model will affect the accuracy of chest radiology reports and lack clinical depth. After analyzing the chest radiology medical report set, the embodiments of the present invention have constructed a chest radiology knowledge graph with more specific and detailed disease categories.

[0046] S2: Based on the chest radiology medical report set, construct a chest radiology knowledge graph.

[0047] For example, in the existing chest radiology knowledge graph (including 18 common disease categories under 7 organs), add more disease categories included in the radiology reports in the IU-Xray dataset and the MIMIC-CXR dataset to expand the knowledge graph.

[0048] The specific process of constructing the chest radiology knowledge graph includes:

[0049] S201: Pre-define an organ pool, a disease pool, a spatial location pool, a negative phrase pool, and a positive phrase pool, as well as synonym pools for organs and diseases respectively.

[0050] The synonym pool is a dictionary where the keys represent the potentially different written forms of an organ or a disease, and the values represent their unified written forms, avoiding interference with the model performance due to different subjective writing forms of radiologists caused by the diverse sources of the data set.

[0051] S202: Segment the text in the chest radiology report into sentence-level text.

[0052] Unify the sentence-level text to remove abbreviations (e.g., convert "n't" to "not" without indentation) and stop words and format it into lowercase form.

[0053] S203: Replace the phrases that appear in the sentence by referring to the organ synonym pool and the disease synonym pool in sequence.

[0054] Convert the sentence into the format of "unified by the data center".

[0055] S204: Capture the organs, diseases, and their spatial locations that appear in the sentence according to the organ pool, the spatial location pool, and the disease pool, and construct the sentence-level "disease - space - organ" triple tags.

[0056] Since there may be some phrases in the sentence that express a negative meaning for the disease (such as "no" and "without" in the negative phrase pool that mean "none", and "clear" and "normal" in the positive phrase pool that mean "normal"), remove the disease - space - organ tags in the sentence that contain these phrases, and in this process, splice the negative knowledge and the positive knowledge contained in the sentence together, which is uniformly called knowledge G. Among them, negative knowledge such as "none - atelectasis - right - lung", and positive knowledge such as "normal - left - lung", indicating that there is no corresponding disease risk in the sentence, guiding the model to generate correct descriptions instead of generating false positive diagnostic results.

[0057] Figure 5 This is a schematic diagram of the disease types defined in the chest radiology knowledge graph of the embodiments of the present invention.

[0058] In the illustrated embodiment, with reference to Figure 5 as shown, the entity types in the chest radiology knowledge graph include:

[0059] Disease categories, such as 127 specific disease findings like pneumonia, cardiomegaly, etc. and 1 normal label (indicating the health status);

[0060] Spatial locations, such as orientation indicators like "left", "right", and "front", etc.

[0061] Organ categories, such as 7 major organs like "lung", "heart" and "bone", and other organs. This category of other organs mainly includes pathological phenomena related to multiple organ systems but not completely belonging to a specific organ or system, or abnormalities related to other surgeries and traumas, such as "lymph node enlargement", "foreign body" and "thoracotomy", etc.

[0062] After identifying the entity types, the following semantic relationships are mainly considered in the chest radiology knowledge graph: the "disease - space - organ" triple to construct the chest radiology knowledge graph.

[0063] In the embodiment of the present invention, structured knowledge (i.e., constructing a knowledge graph) can be achieved through step S2. Structured knowledge makes a core contribution to guiding the quality of report generation. That is, by defining a knowledge graph with a wider coverage of disease types and constructing fine - grained "disease - space - organ" medical knowledge, it guides the conditional probability modeling of the decoder and restricts the clinical relevance of the generated text. Furthermore, it enhances the recognition ability of the medical report generation network for rare diseases, successfully captures rare abnormalities such as "interstitial opacity" and "scar", and effectively solves the long - tail problem.

[0064] S3: Extract text features and image features from the training samples respectively, and extract knowledge features from the training samples based on the chest radiology knowledge graph.

[0065] In the embodiment of the present invention, Vision Transformer is used as the visual encoder E f , and BERT is used as the text encoder E t , knowledge encoder E k and decoder E d .

[0066] It should be noted that the core idea of Vision Transformer (abbreviated as ViT) is to regard the input image as a series of "word embeddings", similar to the text sequence in natural language processing. Traditional image processing methods, such as convolutional neural networks, rely on convolutional operations to extract image features, while ViT breaks through this limitation and adopts the Transformer architecture to process image data. By dividing the image into multiple small patches

[88] and treating these small patches as input "words", ViT can capture local and global information in the image without a convolutional layer.

[0067] It should be noted that through the masked language modeling task, BERT has learned left - right bidirectional context representations, broken the one - way limitation of traditional autoregressive models, significantly improved the performance of text understanding, and formed a general paradigm of "masking - reasoning - generation".

[0068] Both BERT and Vision Transformer are feature extractors based on the Transformer architecture, specifically BERT in the field of natural language processing and Vision Transformer in the visual field.

[0069] Specifically, the original chest radiograph is segmented into Np small patches of size p×p and fed into the visual encoder E. f Extract image features v (i) The head token in [cls] is used to summarize visual information and represents the global visual feature. C represents the dimension, and C = 768.

[0070] Specifically, use the text encoder E t Extract text feature t (i) , t (i) ∈R M×C , M represents the number of tokens included in the chest radiology medical report, and the first token [cls] in t (i) is used to summarize the report information and represents the global report feature.

[0071] In the illustrated embodiment, an encoder based on the BERT architecture is used. First, the triple knowledge text is vectorized and then fed into BERT to extract knowledge features.

[0072] According to whether each sentence in the report appears with the corresponding label, sentence-level medical knowledge is constructed. For example, "low right lung volume, bronchovascular congestion" is constructed as "bronchovascular congestion - right - lung_low volume - right - lung" and fed into the knowledge encoder E as medical knowledge. k Obtain knowledge feature k (i) .

[0073] In the embodiment of the present invention, knowledge features are extracted based on the knowledge graph. Structured knowledge makes a core contribution to guiding the quality of report generation. That is, by defining a knowledge graph covering a wider range of disease types and constructing fine-grained "disease - space - organ" medical knowledge, the conditional probability modeling of the decoder is guided, and the clinical relevance of the generated text is constrained. The present invention can enhance the model's ability to recognize rare diseases, successfully capture rare abnormalities such as "interstitial opacity" and "scar", and effectively solve the long-tail problem.

[0074] S4: Input the text feature, image feature, and knowledge feature into the dual-path knowledge enhancement gating module. One branch aligns the text feature and the image feature and calculates the image - text alignment loss. Another branch aligns the image feature and the knowledge feature and calculates the image - knowledge alignment loss. Fuse the output results of the two branches to obtain the enhanced fusion feature;

[0075] To address the differences among multi-modal features (text features, image features, knowledge features), the present invention designs a Dual-Path Knowledge-enhanced Gated module (DPKG) for feature alignment and fusion.

[0076] In the illustrated embodiment, the Dual-Path Knowledge-enhanced Gated module (DPKG) includes two branches, each branch is provided with a cross-attention layer, and the outputs of the two branches are connected to a gated unit.

[0077] Referring to Figure 2 as shown, the text features and image features of a chest radiology report are input into the Dual-Path Knowledge-enhanced Gated module, and feature alignment is performed through the cross-attention layer in one of its branches, which is a text-guided visual enhancement branch.

[0078] S401: Input the image feature v (i) and the text feature t (i) into one branch of the Dual-Path Knowledge-enhanced Gated module, and use the cross-attention layer for interaction and context learning to obtain the text-enhanced visual feature which is specifically expressed as:

[0079]

[0080] In the formula, Q t represents the query in the transformer attention, v (i) represents the image feature, K t represents the key, t (i) represents the text feature, A t represents the attention score, (.) T represents matrix transpose, represents the text-enhanced visual feature, respectively represent learnable projection matrices.

[0081] Referring to Figure 2 as shown, the knowledge feature and image feature of a chest radiology report are input into the Dual-Path Knowledge-enhanced Gated module, and feature alignment is performed through the cross-attention layer in the other branch, which is a knowledge-guided visual enhancement branch.

[0082] S402: Input the visual feature v (i) and the knowledge feature k (i) into the other branch of the Dual-Path Knowledge-enhanced Gated module, and use the cross-attention layer for interaction and context learning to obtain the knowledge-enhanced visual feature

[0083]

[0084] In the formula, Q k represents the query in the transformer attention, v (i) represents the image feature, K k represents the key, k (i) represents the knowledge feature, A k represents the attention score, (.) T represents matrix transpose, represents the knowledge-enhanced visual feature, respectively represent learnable projection matrices.

[0085] S403: Input the image feature v (i) , the text-enhanced visual feature and the knowledge-enhanced visual feature into the gated unit together, adjust the gating weights, perform feature fusion, and obtain the enhanced fusion feature v' (i) .

[0086] The calculation formula for the gating weights is:

[0087]

[0088] In the formula, α, β, γ represent the gating weights, v g(i) represents the global feature represented by the head token [cls] of the image feature v (i) , represents the global feature represented by the head token [cls] of the text-enhanced visual feature , represents the global feature represented by the head token [cls] of the knowledge-enhanced visual feature ; [:] represents the concatenation operation and represents the weight matrix; represents the bias term, which makes the model prefer to rely on the image feature v (i) as prior knowledge, and activates the knowledge or text path when encountering difficult cases, that is, makes the initial value of the bias term b g corresponding to the α position relatively high, and set b g to = [0.5, -0.25, -0.25].

[0089] The present invention uses the gating weights α, β, γ to dynamically fuse the image feature v (i) , the report-enhanced visual feature and the knowledge-enhanced visual feature to generate a multi-modal visual feature that retains both the visual information of the image itself and incorporates text semantics and medical knowledge, that is, the enhanced fusion feature v' (i) , which is expressed as:

[0090]

[0091] In the formula, ⊙ represents element-wise multiplication (Hadamard product).

[0092] The enhanced fusion feature v' (i) will be fed into the decoder for assisting in report generation.

[0093] The present invention combines the text feature t (i) and the knowledge feature k obtained from the knowledge graph (i) , respectively, with the original image feature v (i) through the image-report alignment loss and the image-knowledge alignment loss to ensure the alignment of the two modalities of report and knowledge with the visual feature, and solve the semantic gap problem between heterogeneous modalities.

[0094] In the process of fusing the knowledge feature k obtained from the knowledge graph (i) with the image feature v (i) through the cross-attention mechanism, it is necessary to additionally ensure the information alignment between the knowledge feature and the image feature through contrastive learning, so that the model can correctly capture the correct matching relationship between the visual region and the medical entity knowledge. Image-knowledge contrastive learning can encourage positive image-knowledge pairs to have more similar representations compared to negative pairs. Specifically, first calculate the image-knowledge similarity The calculation formula is as follows:

[0095] s(v g(i) , k g(i) ) = W I (v g(i) )W G (k g(i) ) T

[0096]

[0097] In the formula, v g(i) represents the global feature represented by the head token [cls] of the image feature v (i) , k g(i) represents the global feature represented by the head token [cls] of the knowledge feature k (i) , s(v g(i) , k g(i) ) represents the similarity score between v g(i) and k g(i) , W I (.), W G (.) represent learnable projection matrices respectively, W I (.) and WG (.) is learnable, (.) T represents matrix transpose, and τ is the temperature factor.

[0098] Secondly, according to the calculated image-knowledge similarity obtain the image-knowledge alignment loss It is specifically expressed as:

[0099]

[0100] Among them, represents the cross-entropy loss, and g k (I) represents the true label of image-knowledge alignment, and g k (K) represents the true label of knowledge-image alignment, and f k (I) represents the image-knowledge similarity, and f k (K) represents the knowledge-image similarity.

[0101] Similarly, in the process of fusing the text feature t (i) with the image feature v (i) through the cross-attention mechanism, it is also necessary to ensure the information alignment between the text feature and the visual feature in the report through contrast learning, so that the model can correctly capture the matching relationship between the visual region and the phrases in the report.

[0102] The image-text alignment loss The calculation formula is:

[0103]

[0104]

[0105] Among them, represents the cross-entropy loss, and g t (I) represents the true label of image-text alignment, and g t (T) represents the true label of text-image alignment, and f k (I) represents the image-text similarity, and f k (K) represents the text-image similarity.

[0106] S5: Input the enhanced fusion feature v′ (i) into the classifier to obtain the classification result and calculate the classification loss Input the fusion feature v′ (i) into the decoder to generate a chest radiology medical report, and calculate the generation loss between the generated medical report label and the true medical report label

[0107] Input the enhanced fusion feature v′(i) An input linear classification head is used to predict the classification result of a medical image. Its true label is the labeled disease label. If it is included in the report, the corresponding binary label bit is 1, otherwise it is 0. The present invention defines the diagnostic classification loss using binary cross-entropy loss Optimize the classification result, and send the classification result into the report decoder to guide the report generation loss

[0108] The disease keywords included in the knowledge graph defined in the embodiments of the present invention directly reflect the pathological characteristics of medical images. Therefore, the present invention proposes to guide report generation by adding the diagnostic results of medical images, and defines it as a multi-label classification task, independently predicting 128 predefined disease rules in the knowledge graph. Specifically, the embodiments of the present invention use the DPKG module to process and obtain the fused feature v′ (i) , which is sent as input to a linear classification head to predict the classification result of the medical image. Its true label is the disease label y captured by the present invention in constructing the knowledge graph b,s , if it is included in the report, the corresponding binary label bit is 1, otherwise it is 0.

[0109] The present invention defines the classification loss of the diagnostic result using binary cross-entropy loss It is expressed as:

[0110]

[0111] Where B represents the batch size, S represents the number of disease label types, f(·) represents the result logit output by the disease classification branch, σ(.) represents the Softmax function, and log(.) represents the logarithmic function.

[0112] The classifier in the diagnostic classification module is closely related to the recognition of different diseases and can effectively extract and reflect the medical information contained in medical images. By analyzing the medical image features, the classifier identifies the key visual features related to specific diseases and provides important decision-making support for the subsequent report generation based on the classification results, thereby improving the accuracy and pertinence of the report generation content.

[0113] Send the fused feature v′ (i) into the Transformer-based decoder Ed(.). The entire process of the decoder layer Ed(.) can be expressed as follows:

[0114] y′ t = E d (y′1, y′2,..., y′ t-1 , v′ (i) )

[0115] Where y′t Denotes the token to be predicted at time step t.

[0116] The core training objective of the report generation task is to minimize the cross-entropy loss between the generated report Y′={y′1,y′2...,y′ t} and the ground-truth report Y={y1,y2,...,y t}:

[0117]

[0118] Through this training strategy, the model learns how to generate reports with high-quality clinical information while maintaining a language structure and semantic content that are semantically consistent with the ground-truth report.

[0119] S6: Through the image-text alignment loss The image-knowledge alignment loss The classification loss and the cross-entropy loss perform backpropagation to optimize the network parameters.

[0120] The embodiments of the present invention take the sum as the total loss function of the present invention and its calculation formula is:

[0121]

[0122] where λ1 = λ2 = λ3 = λ4 = 0.25 are hyperparameters used to balance the losses.

[0123] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, which can include: ROM, RAM, disk, or optical disc, etc.

[0124] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating a chest radiology report based on a knowledge graph, characterized in that, Including: Obtain the chest radiological image to be processed, input it into the trained medical report generation network, and obtain the chest radiology medical report; The training process of the medical report generation network includes: Obtain a chest radiology medical report set and use the chest radiology medical report as a training sample; Based on the chest radiology medical report set, construct a chest radiology knowledge graph; Extract text features and image features from the training samples respectively, and extract knowledge features from the training samples based on the chest radiology knowledge graph; Input the text features, image features, and knowledge features into a dual-path knowledge enhancement gating module. One branch aligns the text features and image features and calculates the image-text alignment loss; the other branch aligns the image features and knowledge features and calculates the image-knowledge alignment loss; fuse the output results of the two branches to obtain enhanced fusion features; Input the enhanced fusion features into the classifier to obtain the classification results, and calculate the classification loss e DCM Input the fusion features into the decoder to generate a chest radiology medical report, and calculate the generation loss l RG ; Through the image-text alignment loss l ITA 、 the image-knowledge alignment loss l IKA 、 the classification loss l DCM and the generation loss l RG Backpropagate to optimize the network parameters.

2. The method for generating a chest radiology report based on a knowledge graph according to claim 1, wherein Perform data augmentation on the chest radiology medical report set, and its specific process includes: According to the chest radiology medical report dataset, annotate disease labels; Extract text word vectors from the chest radiology medical report and construct a sentence-level key-value pool, where the key represents keywords related to diseases, and the value represents sentences including keywords or disease labels; Define label counting, set the label number interval, query all sentences under each disease label in the chest radiology medical report according to the disease label, filter the values including eligible disease labels, replace the current value including the disease label with the filtered value to obtain a new value, and obtain the enhanced report dataset.

3. The method for generating a chest radiology report based on a knowledge graph according to claim 1, wherein The specific process of constructing the chest radiology knowledge graph includes: Pre-define an organ pool, a disease pool, a spatial location pool, a negative phrase pool, and a positive phrase pool, as well as a synonym pool for each of organs and diseases; Segment the text in the chest radiology report into sentence-level text; Successively refer to the organ synonym pool and the disease synonym pool to replace the phrases that appear in the sentences; According to the organ pool, the spatial location pool, and the disease pool, capture the organs, diseases, and their spatial locations that appear in the sentences, and construct sentence-level "disease-space-organ" triple labels.

4. The method for generating a chest radiology report based on a knowledge graph according to claim 1, wherein The entity types in the chest radiology knowledge graph include disease categories, spatial locations, and organ categories, and its semantic relationship is the "disease-space-organ" triple.

5. The method for generating a chest radiology report based on a knowledge graph according to claim 1, wherein Use Vision Transformer as the visual encoder E f , use BERT as the text encoder E t , the knowledge encoder E k and the decoder E d .

6. The method for generating a chest radiology report based on a knowledge graph according to claim 1, wherein The dual-path knowledge enhancement gating module includes two branches, each branch is set with a cross-attention layer, and the outputs of the two branches are connected to a gating unit.

7. The method for generating a chest radiology report based on a knowledge graph according to claim 1, wherein Perform enhanced fusion on the text features and image features, and image features and knowledge features respectively through the dual-path knowledge enhancement gating module, and its specific process includes: Input the image features and text features into one branch of the dual-path knowledge enhancement gating module, and use the cross-attention layer for feature interaction and context learning to obtain text-enhanced visual features; Input the visual features and knowledge features into the other branch of the dual-path knowledge enhancement gating module, and use the cross-attention layer for feature interaction and context learning to obtain knowledge-enhanced visual features; Input the image features, text-enhanced visual features, and knowledge-enhanced visual features together into the gating unit, adjust the gating weights, and perform feature fusion to obtain enhanced fusion features.

Citation Information

Cited By

  • Radiology report generation method based on comparative learning and adaptive knowledge integration

    CN120809049A

  • Radiology report generation method fusing clinical semantic modulation and hyperbolic prototype classification

    CN121964039A

  • Radiology report generation method fusing clinical semantic modulation and hyperbolic prototype classification

    CN121964039B