A medical image report generation method and device fusing tag information

CN115662565BActive Publication Date: 2026-09-25CHINA THREE GORGES UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211422392.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2026-09-25
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

解决现有技术中医学影像文本报告的生成效率、精度较低的问题

Benefits of technology

[0035]另一方面,本发明公开了一种计算机设备,包括存储器、处理器以及存储在存储器上并可在处理器上运行的计算机程序,其特征在于,所述处理器执行程序时实现融合标签信息的医学影像报告生成方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662565B_ABST
    Figure CN115662565B_ABST
Patent Text Reader

Abstract

The application discloses a medical image report generation method and device fusing label information and belongs to the field of medical image processing and text generation. The method comprises the following steps: constructing a medical image report generation model; based on medical image data, extracting visual features and semantic features in the image; identifying and classifying the semantic features to obtain label features of the image; fusing the visual features and the label features to obtain fused features; and inputting the processed fused features into a text decoder to realize medical image report generation. The application accelerates the automation of the work flow, reduces the work burden of doctors, reduces the probability of error reports, and improves the quality and standardization of medical reports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing and text generation technology, and more specifically to a method and device for generating medical image reports that integrates tag information. Background Technology

[0002] The automatic generation of medical image reports aims to generate reports with the 6C characteristics (clear, correct, concise, complete, consistent, and coherent) from given medical images. Using massive amounts of diagnostic reports and medical images as the basic data source, deep learning is employed for feature extraction and analysis to generate structured diagnostic reports. This represents a novel approach combining image processing and natural language generation. Current research on automatic medical image report generation only addresses the classification and diagnostic report generation for common thoracic diseases. This paper proposes a multi-task model combining multi-label classification, object detection, and medical report generation. Its core is to predict disease labels through classification. This involves replacing the encoder and decoder networks with higher-performance ones, training additional classifiers to predict disease or medical labels, and further improving report quality. Prior knowledge is used to construct disease maps and obtain disease prediction results. However, most existing models generate reports based on visual features, and the generated reports have limitations in several evaluation metrics, resulting in low efficiency and accuracy in generating medical image text reports.

[0003] Therefore, how to provide a method and device for generating medical image reports that integrates label information is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the present invention provides a method and device for generating medical image reports by fusing tag information. Based on a medical report generation framework consisting of an encoder composed of a Transformer and MIX-MLP multi-label classification network, a co-attention mechanism, and a hierarchical LSTM decoder, the invention automatically generates medical image reports using this fusing tag information framework. This solves the problems of low efficiency and accuracy in generating medical image text reports in existing technologies.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] On one hand, this invention discloses a method for generating medical image reports by integrating tag information, comprising the following steps:

[0007] A medical image report generation model framework is constructed, which includes: an encoder, a classification module, a fusion module, and a text decoder;

[0008] Acquire medical image data, preprocess the medical image data, and then input it into the medical image report generation model framework;

[0009] The encoder extracts visual and semantic features from the image to obtain visual feature information and semantic feature information.

[0010] The semantic feature information is identified and classified by the classification module to obtain the label feature information of medical images;

[0011] The fusion module performs visual text alignment fusion on visual feature information and label feature information to obtain fused feature information.

[0012] The processed fusion feature information is input into the text decoder to generate and output a medical image report.

[0013] Preferably, the medical image report generation model framework includes: an encoder based on the Transformer model, a classification module based on a multi-label classification network of MIX-MLP, a fusion module based on a visual text alignment attention mechanism of POS-SCAN, and a text decoder based on a hierarchical LSTM network.

[0014] Preferably, the step of acquiring medical image data, preprocessing the medical image data, and inputting it into the medical image report generation model framework includes:

[0015] Acquire medical imaging data;

[0016] The medical image data is vectorized;

[0017] The vectorized medical image data is input into the medical image report generation model framework.

[0018] Preferably, the step of extracting visual and semantic features from the image through the encoder to obtain visual feature information and semantic feature information includes:

[0019] Vectorized medical image data is input into an encoder based on a Transformer model;

[0020] The encoder of the Transformer model acts as a visual and semantic feature extractor, simultaneously extracting visual and semantic features to obtain feature information.

[0021] The feature information is separated into visual feature information and semantic feature information.

[0022] The above technical solution uses a Transformer encoder as a visual and semantic feature extractor to extract two types of features simultaneously. After training, feature information is extracted from the penultimate layer, and the feature information is separated into visual features and semantic features, which are then input into the downstream module respectively.

[0023] Preferably, the step of identifying and classifying semantic feature information through the classification module to obtain label feature information of medical images includes:

[0024] The classification module of the MIX-MLP-based multi-label classification network classifies and labels semantic feature information to obtain classification and labeling results.

[0025] The Focal Loss function is introduced into the multi-label classification network of MIX-MLP to organize the classification and labeling results and obtain the label feature information of medical images.

[0026] Preferably, the step of performing visual text alignment fusion on visual feature information and label feature information through the fusion module to obtain fused feature information includes:

[0027] The fusion module based on the POS-SCAN visual text alignment attention mechanism maps visual information and semantic information from multi-label classification into the same joint semantic space to align with text information, judges the similarity between global images and text information in medical images, and obtains similarity results.

[0028] Based on the similarity results, the global image and text information in the medical images are matched at a fine-grained level to obtain fused feature information.

[0029] The above technical solution, based on the POS-SCAN visual text alignment attention mechanism, infers the global similarity between the image and the text by mapping visual information and semantic information of multi-label classification into the same joint semantic space and aligning them with the text information, thereby enabling fine-grained matching between the image and the text.

[0030] Preferably, the text decoder of the hierarchical LSTM network includes: a sentence LSTM network module and a word LSTM network module.

[0031] Preferably, the processed fusion feature information is input into the text decoder to generate and output a medical image report, including:

[0032] The LSTM network module described above generates multiple topic features from the fused feature information.

[0033] The word LSTM network module generates corresponding sentences for each topic feature;

[0034] A complete medical imaging report composed of multiple sentences is generated and output.

[0035] On the other hand, the present invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the program to implement the steps of a method for generating medical image reports with fused label information.

[0036] As can be seen from the above technical solution, compared with the prior art, this invention discloses a method and device for generating medical image reports that integrates label information. It constructs a medical image report generation model framework using three modules: an encoder based on a Transformer and MIX-MLP multi-label classification network, a co-attention mechanism, and a hierarchical LSTM decoder. This framework enables the automatic generation of medical image reports. This invention has the beneficial effects of accelerating workflow automation, reducing the workload of doctors, decreasing the probability of erroneous reports, and improving the quality and standardization of medical reports. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0038] Figure 1 This is a schematic diagram of the generation method process framework provided by the present invention;

[0039] Figure 2 This is a schematic diagram of the classification process of the classification module of the multi-label classification network based on MIX-MLP provided in an embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram illustrating the process of obtaining fused feature information using a POS-SCAN-based visual text alignment attention mechanism provided in an embodiment of the present invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] On the one hand, see appendix Figure 1As shown in the figure, this invention discloses a method for generating medical image reports by fusing tag information, including the following steps:

[0043] A medical image report generation model framework is constructed, which includes: an encoder, a classification module, a fusion module, and a text decoder;

[0044] Acquire medical imaging data, and input the preprocessed medical imaging data into the medical imaging report generation model framework;

[0045] Visual and semantic features are extracted from the image by an encoder to obtain visual feature information and semantic feature information;

[0046] The semantic feature information is identified and classified by the classification module to obtain the label feature information of medical images;

[0047] The fusion module performs visual text alignment and fusion of visual feature information and label feature information to obtain fused feature information.

[0048] The processed fusion feature information is input into the text decoder to generate a medical image report and output it.

[0049] In one specific embodiment, medical image data is acquired and then vectorized so that it can be input into the frame.

[0050] Specifically, it is processed into a 3D vector. Where C is the number of channels, and H and W represent the image height and width, respectively.

[0051] In one specific embodiment, the image vector is input into the encoder of the frame to extract visual features and label features. The specific steps are as follows:

[0052] 1) Image vectors are input into the Transformer of the framework to obtain visual features and primary semantic features, i.e., Img→f v f s′ ,in As a visual feature, These are primary semantic features.

[0053] Specifically, the image is divided into M image patches and flattened into 2D vectors. The resolution of each image patch is (P, P), the number of channels is C, and M = HW / P. 2 The number of image patches. p The vector is projected onto a D-dimensional array through a fully connected layer and then concatenated with a learnable positional encoding vector. With a one-dimensional location embedding vector carrying location information The sums are then fed into the Transformer encoder (z). l The entire encoder consists of L Transformer encoders, each containing a multi-head self-attention (MSA) network and a multi-layer perceptron (MLP) network. LayerNorm (LN) and residual connections are added before the MSA and MLP to reduce overfitting and prevent gradient vanishing. The visual and primary semantic feature vectors f... v f s′ All are outputs from the Transformer encoder, whose output vector Z = [x class ;x1;x2;…;x n ], let f v = [x1; x2; ...; x n ], f s′ =[x class ].

[0054] z0 = [x class ;i1E;i2E;…;i n E]+E pos #(1)

[0055] z′ l =MSA(LN(z) l-1 ))+z l-1 , l=1,…,L#(2)

[0056] z l =MLP(LN(z) l ))+z′ l , l=1,…,L#(3)

[0057] Z = LN(z) L )#(4)

[0058] f s′ The feature vector is output through a K-dimensional fully connected (fc) layer. Where K is the number of tag types in the dataset, D is the dimension of visual features, and D1 is the dimension of semantic features.

[0059] 2) Primary semantic information processing becomes tag information.

[0060] Specifically, primary semantic information is input into a multi-label classification network to obtain label information.

[0061] See appendix Figure 2The diagram shows the classification process of the classification module of the MIX-MLP-based multi-label classification network. The first dimension of the semantic features is processed using the ML.P network of MLP-Block; the last two dimensions of the semantic features are transposed; the second dimension of the semantic features is processed using the MLP-Block MLP network; and finally, the above steps are repeated Z times to output the label feature information.

[0062] Specifically, the classification module of the MIX-MLP-based multi-label classification network is obtained by concatenating Z MLP-Block networks, where the output of one MLP-Block network is the input of the next. Each MLP-Block consists of two MLP networks, with the first MLP network acting on... The first dimension, the second MLP network acts on The second dimension. Each MLP network contains two fully connected layers and a GELU activation function. After passing through a fully connected layer and the softmax function, we obtain... right The second dimension, namely the probability of occurrence of each tag, is used to sort the tags and select the top N tags for embedding to obtain the semantic features f. s It can be represented as:

[0063] U *,i =X *,i +W2σ(W1LayerNorm(X) *,i )#(5)

[0064] Y j,* =U j,* +W4σ(W3LayerNorm(U) J,* )#(6)

[0065]

[0066]

[0067] Among them W 1-4 Let be the parameter matrix of the MLP network, σ be the GELU activation function, i,j be the dimensions of the hidden layers of the two MLP networks (their values ​​are independent of the dimension of the feature vectors), θ be the MLP-Block layer, Z be the number of MLP-Block layers, and ζ be the topk function. Select the first N vectors after sorting.

[0068] In one specific embodiment, visual information and label information are fused into fused feature information.

[0069] See appendix Figure 3As shown, the flowchart illustrates the process of obtaining fused feature information based on the POS-SCAN visual text alignment attention mechanism. Visual features are input, their cosine similarity to text features is calculated, and the feature weights of the visual soft attention mechanism are calculated and multiplied with the visual features. Simultaneously, label feature information is input, its cosine similarity to text features is calculated, and the feature weights of the semantic soft attention mechanism are calculated and multiplied with the label features. Finally, the two vectors are concatenated and passed through a fully connected layer to output the fused feature information.

[0070] Specifically, visual and label information are input into the co-attention mechanism to obtain fused feature information. For the encoder output... The feature vectors are compared with the hidden layer states using an image-text matching mechanism to better align visual and semantic features. Specifically, f is calculated separately. v f s The hidden layer state of the LSTM network at time t-1 Cosine similarity between The method is as follows:

[0071]

[0072]

[0073]

[0074]

[0075] Where m∈[1,M], n∈[1,N], t∈[1,T], D2 is the dimension of the hidden layer state, BN is the BatchNormalization layer, which controls gradient explosion and prevents gradient vanishing and overfitting; W v W v,h It is the parameter matrix of visual similarity, W s W s,h This is the parameter matrix for semantic similarity. After standardization, the visual and semantic similarities are used to calculate the feature weights for the visual soft attention and semantic soft attention mechanisms, as shown below:

[0076]

[0077]

[0078] Where, [x] + ≡max(x,0) means taking the larger value between x and 0. Calculate the respective soft attention feature vectors using the following formulas:

[0079]

[0080]

[0081] Finally, the two vectors are concatenated through a fully connected layer W. fc Obtain the co-attention feature vector at time t Right now:

[0082]

[0083] In one specific embodiment, the fused feature information is input into the encoder network to obtain the generated text;

[0084] Specifically, the text decoder of the hierarchical LSTM network includes: sentence LSTM network module and word LSTM network module.

[0085] More specifically, the fused feature information is input into a hierarchical LSTM network to obtain a topic vector. Specifically, the feature vector output by the co-attention mechanism... As its input, and generate the corresponding topic vector. The topic vectors are input into an LSTM that generates sentences. After each topic vector is output, a stop control component determines whether to output the next topic vector. The stop control component uses the state of the previous hidden layer. With the current hidden layer state The probability p of generating the next sentence is calculated. The sentence LSTM uses the feature vector cof and the internal hidden layer state h. (t) Calculate topic vector top (t) The formula is as follows:

[0086]

[0087]

[0088]

[0089] Among them, W top,h W top,ctx W stop,t-1 W stop,t W stop,t It is a parameter matrix, LSTM1 represents an LSTM network, This represents the probability that the sentence LSTM network will generate the next sentence at step t. If p is greater than a predefined threshold, the sentence LSTM network will stop generating new topic vectors, and the word LSTM network will also stop generating words.

[0090] More specifically, the topic vector is input into a word LSTM network within a hierarchical LSTM network to generate each sentence. These sentences are then concatenated to produce the final generated report. Specifically, the word LSTM is similar to the sentence LSTM network; it is a standard LSTM network whose first and second inputs are the topic vectors generated by the word LSTM. (t) A predefined starting token is followed by a sequence of words. The hidden layer states are the same as the distribution p(y) used to predict the generated words. t |y 1:t-1 ), generating word sequences using LSTM. Then, all the generated sequences are concatenated to form the final report. The formula is as follows:

[0091]

[0092]

[0093] Among them W word,h Let v be the parameter matrix. start The starting marker is [;], which indicates concatenation. LSTM2 indicates a word LSTM network. express

[0094] In one specific embodiment, the method further includes calculating the loss between the generated report and the image report. This calculates the difference between the model-predicted text and the real sample, and through training and gradient descent, makes the model-generated text more closely resemble the real sample.

[0095] Specifically, each training sample involves multiple loss calculations. The loss is calculated at each point, and the results are summed to obtain the total loss. Each training sample is considered as a tuple (I, G, R), where I is the image, G is the Ground Truth corresponding to image I, and R is the generated report corresponding to image I, consisting of T sentences. Each sentence is composed of S... i It consists of 10 words. For each training sample (I, G, R), the model first calculates the probability distribution p of the label corresponding to its image I across all labels. tag Considering the sparsity of tag distribution, the focal loss function is used to calculate p. tag The loss is calculated based on the true label. The Focal Loss function is a loss function that handles imbalanced classification of samples, and its formula is as follows:

[0096]

[0097] Where N is the number of tags, γ is the sample difficulty adjustment factor, and α is the sample weight.

[0098] The sentence LSTM is divided into T time points. The probability distribution p' of the i-th sentence at each time point is calculated in the two states {STOP, CONTINUE}. i Finally, the topic vector is input into the word LSTM network to generate the word w. i,j Each generated word sequence has a loss calculated using the Cross Entropy (CE) loss function. The reported training loss is the sum of two cross entropy losses: the probability distribution p of the sentence number distribution. stop,i The corresponding loss sent The word distribution p of each sentence i,j The corresponding loss word Combining the three losses together yields the overall training loss:

[0099]

[0100] Where, λ tag ,λ sent ,λ word Weights are assigned to each pre-defined loss.

[0101] On the other hand, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the program to implement the steps of a method for generating medical image reports with fused label information.

[0102] As can be seen from the above technical solutions, compared with the prior art, the embodiments of the present invention disclose a method and device for generating medical image reports by fusing label information. Specifically, it is a method and device for generating medical image reports by fusing label information based on chest X-ray images, which has the following beneficial effects:

[0103] 1) This invention proposes a method for generating medical image reports from image reports, and achieves good results on the IU-XAY and MIMIC-CXR datasets. It outperforms existing models in natural language generation evaluation metrics such as BLEU, ROUGE, and METEOR.

[0104] 2) This invention proposes a method for generating medical image labels from image reports, and achieves good results on the MIMIC-CXR dataset, outperforming existing models in terms of accuracy and recall.

[0105] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0106] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating medical image reports by integrating tag information, characterized in that, Includes the following steps: A medical image report generation model framework is constructed, which includes: an encoder based on the Transformer model, a classification module based on the MIX-MLP multi-label classification network, a fusion module based on the POS-SCAN visual text alignment attention mechanism, and a text decoder based on the hierarchical LSTM network. Acquire medical image data, preprocess the medical image data, and then input it into the medical image report generation model framework; The encoder extracts visual and semantic features from the image to obtain visual feature information and semantic feature information, including: Vectorized medical image data is input into an encoder based on a Transformer model; The encoder of the Transformer model acts as a visual and semantic feature extractor, simultaneously extracting visual and semantic features to obtain feature information. The feature information is separated into visual feature information and semantic feature information; The encoder consists of L Transformer encoders, each containing a multi-head self-attention and multi-layer perceptron network; the simultaneous extraction of visual and semantic features follows the steps below: Divide the image into The image patch is flattened into a 2D vector. The resolution of each image patch is The number of channels is , The number of image patches; Projected through a fully connected layer to Dimension, then concatenate a learnable positional encoding vector. , and a one-dimensional location embedding vector carrying location information The sums are then fed into the Transformer encoder. In the process, the output vector is generated by the Transformer encoder. The formula is: ; ; ; ; In the formula, This indicates a multi-head self-attention network. This represents a multilayer perceptron network. Normalization operation; visual feature vector Primary semantic feature vector , The feature vector is obtained through a Fully connected layer output ,in This represents the number of different types of tags in the dataset. It is a dimension of visual features. It is a dimension of semantic features; The semantic feature information is identified and classified by the classification module to obtain the label feature information of medical images, including: The classification module of the MIX-MLP-based multi-label classification network classifies and labels semantic feature information to obtain classification and labeling results. Focal Loss is introduced into the multi-label classification network of MIX-MLP to organize the classification and labeling results and obtain the label feature information of medical images. The fusion module performs visual text alignment fusion on visual feature information and label feature information to obtain fused feature information, including: The fusion module based on the POS-SCAN visual text alignment attention mechanism maps visual information and semantic information from multi-label classification into the same joint semantic space to align with text information, judges the similarity between global images and text information in medical images, and obtains similarity results. Based on the similarity results, the global image and text information in the medical images are matched at a fine-grained level to obtain fused feature information; The processed fusion feature information is input into the text decoder to generate and output a medical image report.

2. The method for generating a medical image report with integrated tag information according to claim 1, characterized in that, The process of acquiring medical image data, preprocessing the medical image data, and then inputting it into the medical image report generation model framework includes: Acquire medical imaging data; The medical image data is vectorized; The vectorized medical image data is input into the medical image report generation model framework.

3. The method for generating a medical image report by fusing tag information according to claim 1, characterized in that, The text decoder of the hierarchical LSTM network includes: a sentence LSTM network module and a word LSTM network module.

4. The method for generating a medical image report with integrated tag information according to claim 3, characterized in that, The processed fusion feature information is input into the text decoder to generate and output a medical image report, including: The LSTM network module described above generates multiple topic features from the fused feature information. The word LSTM network module generates corresponding sentences for each topic feature; A complete medical imaging report composed of multiple sentences is generated and output.

5. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the medical image report generation method with fused label information as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Semantics-based medical imaging report template generation method

    CN109545302A

  • Model training method and device, equipment and medium

    CN114170482A

  • Medical image report automatic generation method based on attention mechanism

    CN115132313A