Method for automatically generating medical image diagnosis report based on graph convolutional network

Through a graph convolutional network guided by deep disease labels, the problems of disease feature extraction and cross-modal alignment in the generation of medical image diagnostic reports are solved, and efficient and high-quality medical image diagnostic reports are achieved.

CN120299599APending Publication Date: 2025-07-11CHINA WEST NORMAL UNIVERSITY +1

Patent Information

Application Number
CN202510346985.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing medical image diagnostic report generation methods are difficult to generate high-quality, professional and rigorous diagnostic reports in terms of insufficient understanding of disease information and difficulty in aligning across modal data.

Method used

A graph convolution network based on deep disease label guidance is adopted, and through the image encoding module, text encoding module and cross-modal decoding module, combined with the graph convolution network and the Transformer model, the precise extraction and cross-modal alignment of disease characteristics are achieved to generate medical reports.

Benefits of technology

The quality and efficiency of report generation are improved, the attention to disease information is enhanced, and accurate and rigorous diagnostic reports are generated through reasonable cross-modal data fusion and alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299599A_ABST
    Figure CN120299599A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically generating a medical image diagnosis report based on a graph convolutional network, and the method achieves the extraction of specific pathological information from a medical image and the generation of an accurate diagnosis report through the combination of medical image processing and text generation technologies. Guiding image feature extraction by using a disease label, and performing targeted extraction and discretization processing on pathological information in a medical image through a GCN (Graph Convolutional Network); besides, the designed cross-modal alignment module can effectively connect the medical image, the diagnosis report and the disease label, the semantic consistency between the image and the text is enhanced, and the extra workload is reduced through the pre-constructed relation matrix. According to the method, the quality and efficiency of medical report generation are remarkably improved, a new technical path is provided for automatic generation of the medical image diagnosis report, and clinical practicability and calculation efficiency are considered at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and language generation, and in particular to an automatic generation method for a graph convolutional medical image diagnosis report guided by deep disease labels. Background Art

[0002] Medical report generation is an emerging interdisciplinary research task aimed at using professional and rigorous language to describe the lesion areas and disease severity in medical images, so as to reduce the workload of medical experts and provide auxiliary support for diagnosis. At present, medical images are widely used as important tools to assist doctors in evaluating the condition of patients and are used in the diagnosis and treatment of various diseases. In clinical practice, doctors need to provide accurate diagnosis reports for given medical images, which requires not only extensive medical expertise but also a large amount of time and manpower. To reduce the burden on doctors and improve the efficiency of patient consultations, it is particularly important and urgent to develop medical report generation methods, which has attracted extensive attention from researchers in the fields of clinical medicine and artificial intelligence.

[0003] The research on the automatic generation task of medical diagnosis reports not only involves the analysis and interpretation of medical images but also the research on discrete text data such as diagnosis reports. In previous studies, influenced by the field of machine translation, this complex machine learning training mode often uses the traditional encoding-decoding paradigm to solve. Specifically, first, an encoder mainly based on a convolutional neural network (CNN) performs feature extraction work such as detecting and identifying complex and important pathological information, lesion locations, etc. in medical images, and then the extracted high-dimensional feature vectors and the diagnosis report text are synchronously fed into a recurrent neural network (RNN) to compile and output the diagnosis report text word by word. However, although most of the existing methods adopt the traditional encoding-decoding paradigm mainly based on CNN-RNN, due to the highly similar situations of imaging methods, human tissues, and complex pathology, there are still limitations in capturing the highly complex pathological information in medical images. Therefore, the model needs stronger feature extraction capabilities to explore the pathological regions and critical conditions in medical images. Since the convolutional operation can only perform feature learning on local fields and is difficult to apply to long-distance dependent data, a large number of researchers have alleviated the local learning phenomenon by adding an attention mechanism to enhance the global representation ability of long-distance dependent data. At the same time, the matching of image feature vectors and the text content of diagnosis reports is the key to generating high-quality diagnosis reports, and the matching learning ability of medical image feature vectors and text embedding vectors is enhanced by stacking multiple layers of long short-term memory networks and attention mechanisms.

[0004] The success of Transformer in machine translation tasks has attracted extensive attention from researchers and achieved outstanding results in computer vision tasks and natural language processing tasks, providing a new option for traditional encoding-decoding architectures. Due to its powerful ability to learn the internal relationships of images and convert sequence patterns, this novel learning paradigm of Transformer can meet the needs of most artificial intelligence tasks, including medical diagnosis report generation. Compared with the traditional encoding-decoding structure based on CNN-RNN that generates text word by word when processing sequential data, Transformer can assign weights to the input text data and match the key-value pairs extracted from medical images with the queries from the text in a synchronous and parallel manner, thus solving the common memory forgetting problem in traditional recurrent neural networks and achieving the effect of adapting to high-quality text generation tasks. At the same time, as the core mechanism, the multi-head attention mechanism can learn from different subspaces and learn different potential vector representations from multiple perspectives without additional computational complexity, thereby enhancing the model's ability to extract features from images to adapt to different computer vision and natural language processing tasks.

[0005] The versatility of Transformer enables it to adapt to text generation work in most cases. However, due to the special nature of medical images, Transformer often has limitations when performing medical diagnosis report generation tasks. Specifically, for medical images, the Transformer model with the attention mechanism as the core first divides the medical image into different sub-images, and then captures the similarity between itself and the other sub-images in a self-learning manner to finally obtain network features containing the similarity relationship representation between each sub-image. However, in this process, the global features are often ignored, and it is difficult for the traditional Transformer model to capture the overall information of medical images, and such image-level features often contain important pathological information. In addition, in terms of the language pattern of the diagnosis report, the style is often rigorous and unified, and the content of the corpus is relatively single and scarce. Exploring the pathological features carried in medical images and matching the pathological information with appropriate diagnosis reports is the key to improving the model performance and generating professional and rigorous diagnosis reports. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a method for automatically generating medical image diagnosis reports based on deep disease label guidance, and the technical problem to be solved is: in the automatic generation of medical image diagnosis reports, accurately extract disease features and achieve cross-modal alignment among medical images, diagnosis texts, and disease labels, while reducing the preprocessing workload.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] A method for automatically generating a medical image diagnosis report guided by deep disease labels, comprising:

[0009] Obtain a target medical image;

[0010] Construct and train a report generation model;

[0011] Use the trained report generation model to generate a medical report corresponding to the target medical image, wherein the report generation model includes an image encoding module, a text encoding module, and a cross-modal decoding module.

[0012] The image encoding module includes an image feature extraction unit and a disease label-guided graph convolutional network.

[0013] The text encoding module includes a text memory unit;

[0014] The cross-modal decoding module includes a multi-modal fusion unit and a report generator;

[0015] The image feature extraction unit extracts features from the target medical image to generate target grid features, and performs average pooling on the target grid features to generate target global features;

[0016] Through a modified channel attention mechanism, map the target global features to target disease classification features, that is, generate independent attention maps for each disease category;

[0017] Concatenate the target disease classification features and the target global features to generate initial node features of the graph convolutional network;

[0018] Use the disease label-guided graph convolutional network to perform graph convolutional operations on the initial node features, extract specific visual feature information related to the disease, and generate target graph node state features;

[0019] The text memory unit (i.e., Figure 1 the memory bank in) stores text memory information related to disease features;

[0020] Input the target graph node state features and the text memory information into the cross-modal decoding module to generate target cross-modal features corresponding to the target medical image;

[0021] Concatenate the target cross-modal features with the target grid features to generate target fusion features;

[0022] Use the report generator to generate a medical report corresponding to the target medical image word by word from the target fusion features.

[0023] Furthermore, the text encoding module further includes a text feature extraction unit, and the training of the report generation model includes:

[0024] Obtain paired medical image-diagnosis report data in batches, where each paired data pair includes a medical image Img and a report text R corresponding to the medical image;

[0025] Use the pre-trained EfficientNet to perform preliminary visual feature extraction on the medical image Img to obtain grid features F p ;

[0026] Perform average pooling on the grid features F p to obtain global features F g ;

[0027] Use the modified channel attention mechanism to map the global features F g to disease classification features F c ;

[0028] Concatenate the disease classification features F c with the global features F g to generate initial node features F n ;

[0029] Construct a disease label-guided graph convolutional network, and perform multi-layer graph convolutional operations on the initial node features F n to extract specific visual feature information related to the disease and generate graph node state features S n ;

[0030] Use the text feature extraction unit to encode the report text R to obtain text features T;

[0031] Utilize the text features to update the memory information Mem in the text memory unit through the multi-head attention mechanism to enhance the semantic representation ability and context understanding ability of the text features;

[0032] Input the graph node state features S n and the memory information Mem in the text memory unit into the cross-modal decoding module, and perform feature fusion through the multi-head attention mechanism to obtain cross-modal features F m ;

[0033] Concatenate the cross-modal features F m with the grid features F p to obtain fused features F;

[0034] Use the report generator to generate the medical report R' word by word from the fused features F;

[0035] Optimize the report generation model by minimizing the global loss function until convergence.

[0036] Furthermore, the disease label-guided graph convolutional network consists of disease keyword nodes and central nodes. The disease keyword nodes belonging to the same organ are connected and grouped together. The disease keyword nodes are the disease classification features, and the central nodes are the global features obtained by feature extraction from the image.

[0037] Furthermore, using the text features, updating the memory information Mem in the text memory unit through the multi-head attention mechanism includes:

[0038] For the update of the memory information in the text memory unit at time t, the memory information in the text memory unit at time t - 1 is used as the query of the multi-head attention mechanism, and the text features generated from the report text obtained at time t are used as the key-value pairs of the multi-head attention mechanism, thereby performing multi-head attention calculation. The result obtained from the multi-head attention calculation is used as the update result of the memory information in the text memory unit at time t.

[0039] Furthermore, the cross-modal decoding module performs the following process:

[0040] Linearly classify the graph node state features into multiple label classification features, and the number of label classification features is the same as the number of disease classification features;

[0041] Use the label classification features as the query of the multi-head attention mechanism, and use the memory information in the text memory unit as the key-value pairs of the multi-head attention mechanism, thereby performing multi-head attention calculation;

[0042] Use the result of the multi-head attention calculation as the cross-modal feature between the image and the text.

[0043] Furthermore, generating the medical report word by word from the fusion feature is expressed by the formula:

[0044]

[0045] where θ represents the parameters of the report generator, p θ represents the probability model, p θ (Res|F) represents the probability of generating Res under the condition of F, p θ (res t |res 1:t-1 ,F) represents the probability of generating res 1:t-1 under the conditions of F and res t the probability of, res t represents the word sequence predicted at time t, res 1:t-1 represents all the word sequences predicted at all times before time t, Res represents the generated report, N Res represents the length of the generated report, F represents the fusion feature,

[0046] Furthermore, the construction of the global loss function includes:

[0047] Construct the text generation loss function L tt and the image-text alignment loss function L vt and the label classification loss function L vl , where the text generation loss function L tt represents the alignment loss between the real text and the generated text, and the image-text alignment loss function L vt represents the alignment loss between the image features and the text features, and the label classification loss function L vl represents the alignment loss between the real label and the predicted label;

[0048] Perform a weighted sum of the text generation loss function L tt , the image-text alignment loss function L vt and the label classification loss function L vl ;

[0049] Take the result of the weighted sum as the global loss function.

[0050] The present invention also provides a deep disease label-guided graph convolutional medical image diagnosis report automatic generation system, including:

[0051] A memory configured to store a computer program;

[0052] A processor configured to execute the computer program to implement the above-mentioned deep disease label-guided graph convolutional medical image diagnosis report automatic generation method.

[0053] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned deep disease label-guided graph convolutional medical image diagnosis report automatic generation method.

[0054] The beneficial effects of the present invention are:

[0055] The method in the present invention can effectively solve the problems of insufficient understanding of disease information and difficulty in cross-modal data alignment in the existing report generation methods, and while maintaining a relatively low preprocessing workload, significantly improve the quality and efficiency of report generation. By accurately extracting disease information, the attention of the model to disease information is strengthened, and through reasonable cross-modal data fusion and alignment, the understanding ability of the algorithm for different modal data is enhanced; the efficient and high-quality medical image diagnosis report generation method proposed by the present invention realizes the generation of accurate and rigorous diagnosis reports according to medical image features.

[0056] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Description of the Drawings

[0057] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings, where:

[0058] Figure 1 is the architecture diagram of the automatic generation of medical image diagnosis reports provided by an embodiment of the present invention;

[0059] Figures 2(a) and 2(b) are the result diagrams of the diagnosis reports generated from lung X-ray images under different datasets provided by an embodiment of the present invention;

[0060] Figure 3 is the result diagram of the diagnosis report generated from CT images provided by an embodiment of the present invention. Detailed Embodiments

[0061] The following will refer to the accompanying drawings to describe the preferred embodiments of the present invention in detail. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the protection scope of the present invention.

[0062] The present application proposes an automatic generation method for medical image diagnosis reports based on deep disease label guidance, constructs a new report generation architecture, which consists of a pre-trained image feature extraction network, a graph convolutional network, a text encoder, and a cross-modal decoder. Among them, the pre-trained image feature extraction network (such as EfficientNet) initially extracts global visual features and grid visual features from medical images; the graph convolutional network is guided by disease labels, encodes and discretizes the extracted visual features, extracts specific visual information related to diseases, and constructs a global representation; the text encoder uses a Transformer encoder to encode the report text and enhances the semantic representation ability and context understanding ability of the text features through a text memory unit; the cross-modal decoder then fuses the image features and text features through a multi-head attention mechanism, and then uses a Transformer decoder to decode the fused features into a probability distribution of the report text and perform cross-modal alignment, and finally generates an accurate diagnosis report.

[0063] Figure 1 is the architecture diagram of the automatic generation of medical image diagnosis reports provided by an embodiment of the present invention. In combination with Figure 1 , the automatic generation method for medical image diagnosis reports based on deep disease label guidance may include:

[0064] Obtain a target medical image;

[0065] Construct and train a report generation model;

[0066] Use the trained report generation model to generate a medical report corresponding to the target medical image, where the report generation model includes an image encoding module, a text encoding module, and a cross-modal decoding module.

[0067] The image encoding module includes an image feature extraction unit and a disease label-guided graph convolutional network.

[0068] The text encoding module includes a text memory unit;

[0069] The cross-modal decoding module includes a multi-modal fusion unit and a report generator;

[0070] The image feature extraction unit extracts features from the target medical image to generate target grid features, and performs average pooling on the target grid features to generate target global features;

[0071] Through a modified channel attention mechanism, map the target global features to target disease classification features, that is, generate independent attention maps for each disease category;

[0072] Concatenate the target disease classification features and the target global features to generate the initial node features of the graph convolutional network;

[0073] Use the disease label-guided graph convolutional network to perform graph convolutional operations on the initial node features, extract specific visual feature information related to the disease, and generate target graph node state features;

[0074] The text memory unit stores text memory information related to disease features;

[0075] Input the target graph node state features and the text memory information into the cross-modal decoding module to generate target cross-modal features corresponding to the target medical image;

[0076] Concatenate the target cross-modal features with the target grid features to generate target fusion features;

[0077] Use the report generator to generate a medical report corresponding to the target medical image word by word from the target fusion features.

[0078] The text encoding module also includes a text feature extraction unit. The training of the report generation model includes:

[0079] Obtain paired medical image-diagnostic report data in batches, where each paired data pair includes a medical image Img and a report text R corresponding to the medical image; the report text can be in any language, which can be Chinese or English, and there is no restriction here.

[0080] Use a pre-trained image feature extraction unit (such as EfficientNet, which can be used as a visual extractor) to perform preliminary visual feature extraction on the medical image Img to obtain grid features F p , that is, F p = CNN(Img), where CNN represents an image feature extraction network;

[0081] Perform average pooling on the grid features F p to obtain global features F g , that is, F g = avg_pool(F p ), where avg_pool represents the average pooling function;

[0082] Use the modified channel attention mechanism to map the global features F g to disease classification features F c , that is, F c = Cha_Att(F g ), where Cha_Att represents the modified channel attention mechanism. Through the modified channel attention mechanism, an independent attention map can be generated for each disease category. The modified channel attention mechanism is different from the traditional channel attention mechanism. The traditional channel attention mechanism calculates the attention weights in the channel dimension and weights the feature maps of each channel with the attention weights, while the modified channel attention mechanism calculates the attention weights in the spatial dimension. Specifically, first convert the global features F g into an attention map of N (i.e., the number of disease categories) classes and a reshaped global feature through 2D convolution and reshaping transpose operations, and finally multiply the attention map and the reshaped global feature matrix. It can be expressed by the formula:

[0083] Cha_Att(F g ) = softmax(Conv2d(F g ))·reshape(F g ) T , where softmax represents the normalization operation, Conv2d represents the 2D convolution operation, reshape represents the feature reshaping operation, here the length and width dimensions of the image are flattened, and the superscript T represents the matrix transpose. For example Figure 1As shown, the global features can be mapped into 14 disease classification features, namely no abnormality found (i.e., healthy), cardiomediastinal enlargement, cardiac hypertrophy, pulmonary shadow (i.e., pulmonary opacity), lung lesion, pulmonary edema, pulmonary consolidation, pneumonia, atelectasis, pneumothorax, pleural effusion, other pleural abnormalities, fracture, support device (i.e., auxiliary device);

[0084] Next, concatenate the disease classification feature F c with the global feature F g to generate the initial node feature F n , that is, F n = concat(F g , F c ), where concat represents the concatenation operation; the initial node feature F n serves as the initial node of the subsequent graph convolutional network.

[0085] Next, construct a disease label-guided graph convolutional network GCN. Perform multi-layer graph convolutional operations on the initial node feature F n through a predefined graph adjacency matrix to extract specific visual feature information related to diseases and generate the graph node state feature S n . The disease label-guided graph convolutional network GCN can be represented by a predefined graph adjacency matrix. This graph convolutional network can extract the feature states of various diseases and perform multi-label classification from the global features based on disease keywords (i.e., disease classification features). The disease label-guided graph convolutional network consists of multiple disease keyword nodes and a central node. The disease keyword nodes belonging to the same organ are connected and grouped together. The disease keyword nodes are the disease classification features, and the central node is the global feature obtained by extracting the features of the image (i.e., F g ). As Figure 1 shown, there are a total of 14 nodes. Among them, the features related to the lungs (6 of them) can be divided into the same group, and the features related to the chest cavity (3 of them) can be divided into the same group.

[0086] In some embodiments, three-layer graph convolutional operations can be used to process the node feature F n , which can be expressed by the formula:

[0087] F i+1 = GCN(F i ),

[0088] where F i represents the node feature of the i-th layer, F i+1 represents the node feature of the i + 1-th layer, and GCN represents the graph convolutional network. The specific operation can be expressed as:

[0089] R1 = ReLU(BN(Conv1d(F i))),

[0090] R2 = ReLU(D -1 / 2 (G + E N )D -1 / 2 F i W i ),

[0091] C = concat(R1, R2),

[0092] F i+1 = ReLU(BN(Conv1d(C))),

[0093] where R1, R2, and C are intermediate variables in the graph convolution process. In the first step, the initial value of F i , denoted as F 0 , is obtained by initializing the node feature F n . Conv1d represents a one-dimensional convolution operation; BN represents batch normalization operation, ReLU represents an activation function, G is a predefined graph adjacency matrix, D is the diagonal node degree matrix of matrix G, E N is an N-dimensional identity matrix, and the size of N is the same as the number of nodes in G. For example, corresponding to 14 disease classification features, N can be taken as 14 here. W i is the trainable weight of the i-th layer.

[0094] The processing of the medical diagnosis report by the text encoder may include the following process:

[0095] Use a text feature extraction unit (e.g., a Transformer-based text encoder can be used) to encode the report text R to obtain text features T, i.e., T = t_Encoder(R), where t_Encoder represents the text encoder;

[0096] In some embodiments, the text features can also be pooled to obtain a feature vector L t , which will be used for subsequent cross-modal alignment.

[0097] Utilize the text features to update the memory information Mem in the text memory unit through the multi-head attention mechanism, enhancing the semantic representation ability and context understanding ability of the text features;

[0098] For example, for the update of the memory information in the text memory unit at any t moment, the memory information in the text memory unit at t - 1 moment can be used as the query of the multi-head attention mechanism, and the text features generated from the report text obtained at t moment can be used as the key-value pair of the multi-head attention mechanism, so as to perform multi-head attention calculation, and the result obtained from the multi-head attention calculation is used as the update result of the memory information in the text memory unit at t moment.

[0099] The above process can be expressed by the formula:

[0100] Mem t = MHA(Mem t-1 , T),

[0101] where MHA represents the multi - head attention mechanism, the initial Mem can be initialized with an identity matrix of dimension D, D represents the feature size, Mem t represents the memory information in the text memory unit at time t, and Mem t-1 represents the memory information in the text memory unit at time t - 1.

[0102] MHA(Q, KV) = [Att1(Q, KV);...; Att n (Q, KV)]W O ,

[0103]

[0104] where Q and KV represent the query and key - value pair matrices respectively, d k represents the dimension of the key matrix, and W i Q , W i K and W i V are trainable parameters in the linear transformation, i corresponds to different attention heads in the multi - head attention, n represents the total number of attention heads, and W O represents the overall trainable weight. After calculating the attention, multiple attention heads Att are concatenated together, which means that when calculating the attention weights, multiple heads are concatenated into a larger vector.

[0105] To update the memory matrix and guide the forward transmission of report generation simultaneously, the multi - head attention mechanism is used to process the new and old text information. Then, the new text sequence T is used as the key - value input, and the memory matrix Mem t-1 at the previous time step is used as the query input. This setting enables the method to use MHA to process the new text information and calculate the weighted attention that combines the previous memory information. Therefore, after the memory bank is updated, it can be used as an important reference in the forward transmission process of medical report generation.

[0106] Next, the graph node state feature S n and the memory information Mem in the text memory unit are input into the cross - modal decoding module, and feature fusion is performed through the multi - head attention mechanism to obtain the cross - modal feature F m ;

[0107] The cross - modal feature Fm Concatenate with the grid feature F p to obtain the fused feature F, i.e., F = concat(F p , F m );

[0108] Use a report generator (e.g., a Transformer-based decoder) to generate the medical report R' word by word from the fused feature F;

[0109] Optimize the report generation model by minimizing the global loss function until convergence.

[0110] Among them, the cross-modal decoding module performs the following process:

[0111] Linearly classify the graph node state feature S n into multiple (e.g., 14) label classification features Vlab. The number of label classification features is the same as the number of disease classification features. These classification features Vlab will be used for visual-label alignment and report decoding;

[0112] Use the label classification features as the queries of the multi-head attention mechanism, and use the memory information in the text memory unit as the key-value pairs of the multi-head attention mechanism, so as to perform multi-head attention calculation;

[0113] Use the result of the multi-head attention calculation as the cross-modal feature between the image and the text.

[0114] The above process can be expressed by the formula as:

[0115] F m = MHA(Vlab, Mem t ).

[0116] In the present invention, the fused feature is decoded by a Transformer decoder, and a dedicated loss function is designed to achieve alignment between text-text, image-text, and image-label. From a multi-modal perspective, the optimization and fine-tuning tasks are realized.

[0117] Use a Transformer-based medical report generator to process the fused feature encoding F, and adopt an autoregressive decoding strategy to decompose the generation of the medical report into a word-by-word generation process, which can be expressed as:

[0118]

[0119] Among them, θ represents the parameters of the report generator, and p θ represents the probability model, and p θ (Res|F) represents the probability of generating Res under the condition of F, and p θ (res t |res1:t-1 , F) represents the probability of generating res 1:t-1 under the condition of F and res t , that is, the probability of generating the t-th word (corresponding to the t-th moment) given the word sequence (words predicted at all moments before the t-th moment) and the feature F. res t represents the word sequence predicted at the t-th moment, and res 1:t-1 represents the word sequence predicted at all moments before the t-th moment. Res represents the generated report, and N Res represents the length of the generated report, and F represents the fused feature

[0120] In an embodiment of the present invention, the construction of the global loss function includes:

[0121] Construct the text generation loss function L tt , the image-text alignment loss function L vt and the label classification loss function L vl . The text generation loss function L tt characterizes the alignment loss between the real text and the generated text. The image-text alignment loss function L vt characterizes the alignment loss between the image feature and the text feature. The label classification loss function L vl characterizes the alignment loss between the real label and the predicted label;

[0122] Perform a weighted sum on the text generation loss function L tt , the image-text alignment loss function L vt and the label classification loss function L vl ;

[0123] Take the weighted sum result as the global loss function.

[0124] For aligning visual (i.e., image) and text features, a triplet margin loss is adopted. By measuring the differences between the reference image (anchor), the paired text (positive sample), and the unpaired text (negative sample), the distance between the anchor and the positive sample is minimized, while the distance between the anchor and the negative sample is maximized. By applying the triplet margin loss, the paired features are pulled closer in the latent space, thus simulating the bidirectional relationship between the image and the report. For the text feature L t and the visual feature S n , it is necessary to randomly sample negative pairs from the training set and Negative pairs refer to unpaired images and texts. The alignment between the visual and text features can be expressed as:

[0125]

[0126]

[0127] Among them, d represents the cosine distance, which measures the similarity between vectors, β is a boundary parameter determined by the difference between the target image and the negative samples, B and B n are the disease labels of the input image and the negative sample image respectively, and N L represents the number of labels. The total alignment loss L between the visual feature and the text feature vt is given by the following formula:

[0128] L vt = L1 + L2,

[0129] For the alignment between the vision and the label, the method uses binary cross-entropy loss to process the classification result Lab of the GCN and the true label Vlab, which can be expressed as:

[0130]

[0131] Among them, Lab i and Vlab i represent the i-th true label and the predicted label respectively, and N L represents the number of labels.

[0132] For the alignment between the true text and the generated text, cross-entropy loss is also adopted to maximize the correctness of the generated report, which is expressed by the formula:

[0133]

[0134] Hyperparameters μ1, μ2, and μ3 are introduced to combine all the above losses to obtain the proposed final loss (i.e., the global loss function), which can be expressed as:

[0135] L = μ1L tt + μ2L vl + μ3L vt ,

[0136] where μ1, μ2, and μ3 are the weight factors between the three losses.

[0137] The present invention also provides an automatic generation system for a graph convolutional medical image diagnosis report guided by deep disease labels, including:

[0138] A memory configured to store a computer program;

[0139] A processor configured to execute the computer program to implement the automatic generation method for a graph convolutional medical image diagnosis report guided by deep disease labels as described above.

[0140] The present invention also provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the method for automatically generating a graph convolutional medical image diagnosis report based on deep disease label guidance as described above.

[0141] Figures 2(a), 2(b) and Figure 3 respectively show the report generation results of the present invention on chest X-ray images and chest CT images. Among them, the two X-ray images in Figure 2(a) are images from the MIMIC-CXR dataset, the two X-ray images in Figure 2(b) are images from the IU-Xray dataset, and "reference report" represents the real report corresponding to the image. Among them, one report for chest X-ray images corresponds to two frontal and lateral images and contains a total of four cases (only the frontal images are shown), and the ultrasound image is a chest CT slice, containing two cases. In addition, in the chest X-ray image report, bone lesions are marked in black bold format, and the underline indicates the supporting device information. The highlighted parts in green, blue, orange, red, and purple respectively correspond to the conditions correctly diagnosed by the algorithm, representing the supporting device situation, pleural effusion, lung condition, pneumothorax, and heart condition respectively. Compared with the baseline models, namely the M2KT model (see "Radiology report generation with a learned knowledge base and multi-modal alignment" published by Yang et al. in MIA2023), R2Gen (see "Generating Radiology Reports via Memory-driven Transformer" published by Chen et al. in EMNLP 2020), and CMN (see "Cross-modal Memory Networks for Radiology Report Generation" published by Chen et al. in ACL 2021), the reports generated by these three methods are placed below the method names. At the same time, as Figure 3 shown, in the chest CT image report, the inconsistent parts are marked in green, showing a high consistency between this method and the real report. It can be seen that the results of this method are closer to the real results and diagnose more pathological information.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An automatic generation method for medical image diagnosis reports based on deep disease label guidance, characterized in that including: Obtain a target medical image; Construct and train a report generation model; Use the trained report generation model to generate a medical report corresponding to the target medical image, where the report generation model includes an image encoding module, a text encoding module, and a cross-modal decoding module. The image encoding module includes an image feature extraction unit and a disease label-guided graph convolutional network. The text encoding module includes a text memory unit; The cross-modal decoding module includes a multi-modal fusion unit and a report generator; The image feature extraction unit extracts features from the target medical image to generate target grid features, and performs average pooling on the target grid features to generate target global features; Through a modified channel attention mechanism, map the target global features to target disease classification features, that is, generate independent attention maps for each disease category; Concatenate the target disease classification features and the target global features to generate the initial node features of the graph convolutional network; Use the disease label-guided graph convolutional network to perform graph convolutional operations on the initial node features, extract specific visual feature information related to the disease, and generate target graph node state features; The text memory unit stores text memory information related to disease features; Input the target graph node state features and the text memory information into the cross-modal decoding module to generate target cross-modal features corresponding to the target medical image; Concatenate the target cross-modal features with the target grid features to generate target fusion features; Use the report generator to generate a medical report corresponding to the target medical image word by word from the target fusion features.

2. The automatic generation method of a graph convolutional medical image diagnosis report guided by deep disease labels according to claim 1, wherein The text encoding module further includes a text feature extraction unit, and the training of the report generation model includes: Obtain a batch of paired medical image-diagnostic report data, and each paired data pair includes a medical image Img and a report text R corresponding to the medical image; Use the pre-trained EfficientNet to perform preliminary visual feature extraction on the medical image Img to obtain the grid feature F p ; Perform average pooling on the grid feature F p to obtain the global feature F g ; Map the global feature F using the modified channel attention mechanism g to the disease classification feature F c ; Concatenate the disease classification feature F c with the global feature F g to generate the initial node feature F n ; Construct a disease label-guided graph convolutional network to perform multi-layer graph convolutional operations on the initial node feature F through a predefined graph adjacency matrix, extract specific visual feature information related to diseases, and generate graph node state features S n n ;​ Use the text feature extraction unit to encode the report text R to obtain text features T; Use the text features to update the memory information Mem in the text memory unit through a multi-head attention mechanism to enhance the semantic representation ability and context understanding ability of the text features; Input the graph node state feature S n and the memory information Mem in the text memory unit into the cross-modal decoding module, and perform feature fusion through the multi-head attention mechanism to obtain the cross-modal feature F m ; Concatenate the cross-modal feature F m with the grid feature F p to obtain the fused feature F; Use the report generator to generate a medical report R' word by word from the fusion features F; Optimize the report generation model by minimizing the global loss function until convergence.

3. The automatic generation method of a graph convolutional medical image diagnosis report guided by a deep disease label according to claim 1 or 2, characterized in that, The disease label-guided graph convolutional network consists of disease keyword nodes and central nodes. The disease keyword nodes belonging to the same organ are connected and grouped together. The disease keyword nodes are the disease classification features, and the central node is the global feature obtained by extracting features from the image.

4. The automatic generation method of a graph convolutional medical image diagnosis report guided by deep disease labels according to claim 2, wherein Using the text features to update the memory information Mem in the text memory unit through a multi-head attention mechanism includes: For the update of the memory information in the text memory unit at time t, use the memory information in the text memory unit at time t-1 as the query of the multi-head attention mechanism, and use the text features generated from the report text obtained at time t as the key-value pair of the multi-head attention mechanism, so as to perform multi-head attention calculation, and use the result obtained from the multi-head attention calculation as the update result of the memory information in the text memory unit at time t.

5. The method for automatically generating a graph convolutional medical image diagnosis report guided by a deep disease label according to claim 1 or 2, wherein The cross-modal decoding module performs the following process: Linearly classify the graph node state features into multiple label classification features, where the number of label classification features is the same as the number of disease classification features; Use the label classification features as the queries of the multi-head attention mechanism, and use the memory information in the text memory unit as the key-value pairs of the multi-head attention mechanism, so as to perform multi-head attention calculation; Use the result of the multi-head attention calculation as the cross-modal feature between the image and the text.

6. The automatic generation method of a graph convolutional medical image diagnosis report guided by a deep disease label according to claim 1 or 2, characterized in that The formula for generating a medical report word by word from the fused features is expressed as: Among them, θ represents the parameter of the report generator, and p θ represents the probability model, and p θ (Res|F) represents the probability of generating Res under the condition of F, and p θ (res t |res 1:t-1 ,F) represents the probability of generating res 1:t-1 under the conditions of F and res t , and res t represents the word sequence predicted at time t, and res 1:t-1 represents the word sequences predicted at all times before time t. Res represents the generated report, and N Res represents the length of the generated report, and F represents the fusion feature.

7. The method for automatically generating a graph convolutional medical image diagnosis report guided by deep disease tags according to claim 2, wherein The construction of the global loss function includes: Construct the text generation loss function L tt and the image-text alignment loss function L vt and the label classification loss function L vl The text generation loss function L tt represents the alignment loss between the real text and the generated text, and the image-text alignment loss function L vt represents the alignment loss between the image features and the text features, and the label classification loss function L vl represents the alignment loss between the real label and the predicted label; Generate the loss function \(L\) for text tt 、 the image - text alignment loss function \(L\) vt and the label classification loss function \(L\) vl and perform a weighted sum; Use the weighted summation result as the global loss function.

8. An automatic generation system for medical image diagnosis reports based on deep disease label guidance, characterized in that, Include: A memory configured to store a computer program; A processor configured to execute the computer program to implement the method for automatically generating a graph convolutional medical image diagnosis report guided by deep disease labels according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for automatically generating a graph convolutional medical image diagnosis report guided by deep disease labels according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Authentication of subscriber station

    WO2001030104A1

Cited By

  • Generalized zero sample recognition method and system for generative chest X-ray image

    CN120563971A

  • Multi-modal fusion medical information processing method

    CN121054235A

  • A multi-modal fusion medical information processing method

    CN121054235B