A method and device for automatically generating reports based on task perception

Through a task-aware report generation method, medical reports are split into descriptions of individual anatomical structures, and structured reports are generated using embedding vectors and multi-head decoders. This solves the problems of repeated sentences and context coherence in existing technologies and improves the accuracy and readability of reports.

CN115631826BActive Publication Date: 2025-09-09PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211156398.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2025-09-09
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

When existing technologies automatically generate medical reports, the generated paragraphs are prone to contain repeated sentences and lack contextual coherence, resulting in a decrease in report quality.

Method used

A task-aware report generation method is adopted. The original report is input into a pre-trained report generation model, and an embedding vector generator is used to generate a block embedding vector sequence. A classification embedding vector sequence is created for the anatomical structure. A shared encoder and a multi-head decoder are used to split the report into structured reports of each anatomical structure, and a description of each structure is generated separately.

Benefits of technology

It effectively avoids the occurrence of repeated sentences in paragraphs, reduces the difficulty of long text modeling, improves the accuracy and readability of reports, ensures that each decoder head only focuses on report generation of a specific structure, and reduces the interference of redundant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631826B_ABST
    Figure CN115631826B_ABST
Patent Text Reader

Abstract

The present invention provides a task-aware automatic report generation method and device, which includes: inputting the original report into a pre-trained report generation model, using an embedding vector generator to generate a block embedding vector sequence; creating a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence; inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence; inputting the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report split into each anatomical structure. The present invention uses a multi-head decoder in the report generation model to split each anatomical structure in the original report, and each decoder head only focuses on the report generation of the corresponding anatomical structure, thereby avoiding repeated sentences in the generated paragraph, reducing the difficulty of long text modeling, and improving the accuracy of report generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical analysis technology, and in particular to a method and device for automatically generating reports based on task perception. Background Art

[0002] Medical images, which depict the internal structure of anatomical regions, are often used for medical analysis. Medical reports compiled based on these images can be further used for disease diagnosis or medical research. However, due to a shortage of experienced physicians and an increasing number of patients, doctors face a significant workload of image reading and report writing, which inevitably leads to a decline in work quality. Therefore, the use of computer technology to automatically analyze images and generate diagnostic reports, achieving automated generation of medical image reports, is of great significance.

[0003] With the rapid development of deep learning technology and the emergence of large-scale medical image report generation datasets, the task of medical image captioning has garnered widespread attention in recent years. Most related work typically transfers methods from natural image captioning to the task of medical image report generation. These methods employ a similar framework, using convolutional neural networks to extract visual features from images and converting these features into a final report via a text decoder. Similar to natural captioning, these methods directly train the decoder on the complete report. Early work employed recurrent neural networks as the text decoder, but these networks struggle to model long text. Medical image reports are typically much longer than natural image captions. For example, in typical image captioning datasets, the average length of caption text is 10–15 words, while the average length of reports in the IU X-Ray dataset is approximately 30–40 words, and the recently proposed large-scale MIMIC-CXR dataset reaches 50–60 words. Consequently, more methods have opted for hierarchical long short-term memory networks for decoding. This network replaces recurrent neural networks with long short-term memory networks, alleviating the information loss problem when generating long text. Furthermore, the sentence-by-sentence generation approach reduces the length of text processed at each generation, thereby better enabling the generation of paragraph-style text such as medical reports. However, due to the lack of contextual coherence in this hierarchical text generation model, the generated paragraphs are prone to containing repeated sentences.

[0004] Therefore, the existing technology has defects and needs to be improved and developed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a task-aware automatic report generation method and device in response to the above-mentioned defects of the prior art, aiming to solve the problem in the prior art that when automatically generating reports, the generated paragraphs are prone to contain repeated sentences.

[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0007] A task-aware report automatic generation method, comprising:

[0008] Input the original report into the pre-trained report generation model and use the embedding vector generator to generate a sequence of block embedding vectors;

[0009] Create a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence;

[0010] Inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence;

[0011] The hidden state sequence and the classification identification sequence are input into a multi-head decoder to obtain a structured report divided into various anatomical structures.

[0012] In one implementation, inputting the original report into a pre-trained report generation model and generating a block embedding vector sequence using an embedding vector generator includes:

[0013] Inputting the original report into a pre-trained report generation model, and using a visual extractor based on a convolutional neural network to extract image visual features of the original report to obtain an image visual feature vector sequence;

[0014] Linear projection processing is performed on each image visual feature vector in the image visual feature vector sequence to obtain a block embedding vector sequence.

[0015] In one implementation, the block embedding vector sequence and the classification embedding vector sequence are input into a shared encoder to obtain a hidden state sequence and a classification identification sequence, including:

[0016] Inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder;

[0017] The classification embedding vector sequence corresponds to the output classification identification, and a multi-layer perceptron is used to supervise each output classification identification;

[0018] The anatomical structure information in the block embedding vector sequence is extracted according to each output classification identifier to obtain a hidden state sequence and a classification identifier sequence.

[0019] In one implementation, the hidden state sequence and the classification identification sequence are input into a multi-head decoder to obtain a structured report divided into individual anatomical structures, including:

[0020] Inputting the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain corresponding anatomical structure reports output by each decoder head;

[0021] The various anatomical structure reports are spliced ​​together according to a preset splicing order to obtain a structured report.

[0022] In one implementation, the step of training the report generation model includes:

[0023] Acquire a training data set, wherein the training data set includes an original training report;

[0024] Preprocessing the original training report to obtain a structured training report, and using the structured training report as a reference report;

[0025] Inputting the original training report into an initial report generation model, performing generation task training and classification task training on the initial report generation model to obtain a structured generation report;

[0026] When the total loss function of the generation task and the classification task reaches a stable state, the training is completed and the trained report generation model is obtained;

[0027] The initial report generation model is a CNN-Transformer model.

[0028] In one implementation, preprocessing the original training report to obtain a structured training report, and using the structured training report as a reference report, includes:

[0029] extracting report keywords from the original training report;

[0030] Obtaining a chest X-ray knowledge graph, classifying different sentences into different anatomical structures according to the report keywords and the chest X-ray knowledge graph, and obtaining a structured training report;

[0031] The structured training report is used as a reference report.

[0032] In one implementation, the initial report generation model includes a convolutional neural network-based visual extractor and a linear projection; the original training report is input into the initial report generation model, and the initial report generation model is trained for a generation task and a classification task to obtain a structured generation report, including:

[0033] Extracting the image visual features in the original training report using a visual extractor based on a convolutional neural network to obtain an image visual feature vector sequence corresponding to the original training report;

[0034] Using linear projection, the dimension of each image visual feature vector in the image visual feature vector sequence corresponding to the original training report is reduced to 512, so as to obtain a block embedding vector sequence corresponding to the original training report;

[0035] Creating a corresponding classification embedding vector for each anatomical structure in the original training report to obtain a classification embedding vector sequence corresponding to the original training report;

[0036] Inputting the block embedding vector sequence and the classification embedding vector sequence of the original training report into a shared encoder to obtain the hidden state sequence and the classification identification sequence of the original training report;

[0037] Inputting the hidden state sequence and classification identification sequence of the original training report into a multi-head decoder to obtain the corresponding anatomical structure training report output by each decoder head;

[0038] Each anatomical structure training report is spliced ​​in a preset splicing order to obtain a structured generation report.

[0039] In one implementation, the calculation formula of the total loss function is:

[0040] Among them, the is the loss function of the generation task, is the loss function of the classification task, and λ is a hyperparameter used to adjust the loss ratio between the generation task and the classification task;

[0041] The calculation formula of the loss function of the generation task is:

[0042] Wherein, m is the number of anatomical structures, l i is the length of the reference report for the ith anatomical structure, the r ij The jth word in the reference report of the i-th anatomical structure is represented by y ij The jth word in the structured generated report representing the i-th anatomical structure;

[0043] The calculation formula of the loss function of the classification task is:

[0044] Among them, the g i is the category label, and the i To predict the results;

[0045] The category labels include a normal label and an abnormal label for each anatomical structure.

[0046] In one implementation, the step of obtaining the splicing order includes:

[0047] Get the sentence sequence set S of the original training report in the training data set = {s1, s2, ..., s N}, wherein N is the number of original training reports in the training dataset;

[0048] Construct a set of all possible statement sequences T = {t1, t2, ..., t M}, where M is the number of possible statement sequences;

[0049] Comparing all possible sentence sequences with the sentence sequences corresponding to the training dataset to obtain the degree of similarity of the overall sentence order between all possible sentence sequences and the sentence sequences of the training dataset;

[0050] The splicing order is obtained according to the similarity level.

[0051] In one implementation, the block embedding vector sequence and the classification embedding vector sequence of the original training report are input into a shared encoder to obtain the hidden state sequence and the classification identification sequence of the original training report, including:

[0052] Inputting the block embedding vector sequence and the classification embedding vector sequence of the original training report into a shared encoder;

[0053] The classification embedding vector sequence of the original training report corresponds to the output classification identifier, and a multi-layer perceptron is used to supervise each output classification identifier;

[0054] The anatomical structure information in the block embedding vector sequence of the original training report is extracted according to each output classification identifier to obtain the hidden state sequence and classification identifier sequence of the original training report.

[0055] The present invention also provides a task-aware report automatic generation device, which includes:

[0056] An input module, configured to input the original report into a pre-trained report generation model and generate a block embedding vector sequence using an embedding vector generator;

[0057] A creation module is used to create a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence;

[0058] An encoding module, configured to input the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence;

[0059] The report generation module is used to input the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report divided into various anatomical structures.

[0060] The present invention also provides a terminal, which includes: a memory, a processor, and a task-awareness-based report automatic generation program stored in the memory and runnable on the processor. When the task-awareness-based report automatic generation program is executed by the processor, the steps of the task-awareness-based report automatic generation method described above are implemented.

[0061] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the task-awareness-based automatic report generation method as described above.

[0062] The present invention provides a task-aware automatic report generation method and device, the method comprising: inputting the original report into a pre-trained report generation model, generating a block embedding vector sequence using an embedding vector generator; creating a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence; inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence; inputting the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report split into individual anatomical structures. The present invention utilizes a multi-head decoder in the report generation model to split the individual anatomical structures in the original report, and each decoder head only focuses on generating a report for the corresponding anatomical structure, thereby avoiding repeated sentences in the generated paragraphs and greatly reducing the length of text that each decoder head needs to process, reducing the difficulty of long text modeling, and improving the accuracy of report generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a flowchart of a preferred embodiment of the task-aware report automatic generation method of the present invention.

[0064] Figure 2 This is the principle block diagram of the task distillation module and the task-aware report generation module.

[0065] Figure 3 This is the principle block diagram of the classification identification module.

[0066] Figure 4 It is a functional principle block diagram of a preferred embodiment of the task-aware report automatic generation device in the present invention.

[0067] Figure 5 It is a functional principle block diagram of the terminal in the present invention. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0069] With the rapid development of deep learning technology and the emergence of large-scale medical image report generation datasets, the task of medical image captioning has received widespread attention in recent years. Most related work typically transfers methods from natural image captioning to the task of medical image report generation. These methods generally adopt a similar framework, using convolutional neural networks to extract visual features from images and converting these features into a final report via a text decoder. Similar to natural captioning, these methods directly train the decoder on the complete report. Early work used recurrent neural networks as the text decoder, but these networks have difficulty modeling long text. Medical image reports are typically much longer than natural image captions. For example, in general image captioning datasets, the average length of caption text is 10-15 words, while the average length of reports in the IU X-Ray dataset is approximately 30-40 words, and the recently proposed large-scale MIMIC-CXR dataset reports 50-60 words. Consequently, more methods have chosen hierarchical long short-term memory networks for decoding. This network replaces recurrent neural networks with long short-term memory networks, alleviating the information loss problem when generating long text. Its sentence-by-sentence generation approach also reduces the length of text processed at each generation, improving the generation of paragraph-style text such as medical reports. However, due to the lack of contextual coherence in this hierarchical text generation model, the generated paragraphs are prone to containing repeated sentences.

[0070] Some recent studies have used the Transformer model to solve the problem of difficult generation of medical image reports due to the length of the text.

[0071] While utilizing more effective models can effectively address the difficulty of modeling long text, these studies have overlooked the fundamental differences between medical report generation and natural image captioning. In fact, medical reports are more organized than natural image captioning, suggesting that medical image report generation is a task-aware problem. Specifically, medical images of the same body part and modality should be described from the same perspective (typically based on anatomical structures). For example, chest X-ray images are typically described from the perspectives of the heart, thorax, bones, lungs, air cavities, arteries, and external devices. Furthermore, unlike natural image captioning, which primarily describes regions of interest, medical image reports attempt to include as much information as possible from the image. However, the reports in the table also exhibit incomplete content. Consequently, models trained directly on such datasets may overlook the description of certain structures, resulting in incomplete reports. This can potentially overlook unusual structures, impacting their usability in clinical practice. Reports in existing datasets also exhibit varying ordering of different structures, a consequence of the varying writing habits of different physicians. Although the order of statements in a diagnostic report doesn't actually change the information it conveys, using these differently written reports directly for network training inevitably increases the difficulty for the complete report generation network to understand the text. Furthermore, the reports generated by this trained network are relatively unreadable.

[0072] This example proposes a task-aware structured report generation method for generating accurate diagnostic reports. This method reduces the difficulty of the task by breaking down the complete report into its structure and generating descriptions of different anatomical structures. Furthermore, by incorporating the prediction of anatomical abnormalities, this algorithm constructs a multi-task learning model, further improving the quality of the generated reports.

[0073] See Figure 1 , Figure 1 This is a flow chart of the task-aware report automatic generation method in the present invention. Figure 1 As shown, the task-aware report automatic generation method according to an embodiment of the present invention includes the following steps:

[0074] Step S100: Input the original report into the pre-trained report generation model, and use the embedding vector generator to generate a block embedding vector sequence.

[0075] Specifically, the original report includes an image and a brief text description. This embodiment uses a pre-trained report generation model and an embedding vector generator to convert features in the image into a block embedding vector sequence.

[0076] In one implementation, step S100 specifically includes:

[0077] Step S110: Input the original report into a pre-trained report generation model, and use a visual extractor based on a convolutional neural network to extract image visual features of the original report to obtain an image visual feature vector sequence;

[0078] Step S120 : performing linear projection processing on each image visual feature vector in the image visual feature vector sequence to obtain a block embedding vector sequence.

[0079] Specifically, the embedding vector generator is composed of a convolutional neural network based visual extractor f V and a trainable linear projection f P For the input original report (such as radiological image X), the calculation process is: {e1,e2,…,e n}=f P (f V (X)). The convolutional neural network-based visual extractor can be a ResNet-101 model. In this algorithm framework, the last fully connected layer of the convolutional neural network is removed, and the feature map output by the last convolutional layer is adaptively averaged pooled to obtain an output of a fixed size of 7×7×2048. It is regarded as a sequence of image visual feature vectors with a length of 49 and a dimension of 2048. Each vector represents the visual information of a certain area of ​​the image. Subsequently, linear projection reduces the dimension of each image visual feature vector to 512, and finally obtains a block embedding vector sequence {e1,e2,…,e n}.

[0080] like Figure 1 As shown, the task-aware report automatic generation method further includes the following steps:

[0081] Step S200: Create a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence.

[0082] Specifically, a classification embedding vector is created for each structure, with the same dimension as the block embedding vector. The classification embedding vector is learnable and is input into the encoder of the Transformer model together with the block embedding vector sequence.

[0083] In one embodiment, step S200 specifically includes:

[0084] Step S210: input the block embedding vector sequence and the classification embedding vector sequence into a shared encoder;

[0085] Step S220: The classification embedding vector sequence corresponds to an output classification identifier, and a multi-layer perceptron is used to supervise each output classification identifier;

[0086] Step S230: extract the anatomical structure information in the block embedding vector sequence according to each output classification identifier to obtain a hidden state sequence and a classification identifier sequence.

[0087] That is, this embodiment adds a multi-layer perceptron for supervision after each output classification identification, so that it can extract information about the specific structure in the block embedding vector.

[0088] like Figure 1 As shown, the task-aware report automatic generation method further includes the following steps:

[0089] Step S300: Input the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence.

[0090] Specifically, the task-aware multi-head Transformer consists of a shared encoder and multiple separate decoders. The shared encoder is a stack of standard Transformer encoder blocks, which transforms the block embedding vector into a hidden state. The shared encoder leverages the Transformer's ability to capture global information to extract more effective visual feature information.

[0091] like Figure 1 As shown, the task-aware report automatic generation method further includes the following steps:

[0092] Step S400: Input the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report divided into various anatomical structures.

[0093] Specifically, each head of the multi-head decoder is only responsible for generating a description of a specific anatomical structure. This ensures that structures with fewer descriptions are not neglected, and also reduces the difficulty for the generative model to understand complex reports.

[0094] In one implementation, step S400 specifically includes:

[0095] Step S410: input the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a corresponding anatomical structure report output by each decoder head;

[0096] Step S420: splice the anatomical structure reports according to a preset splicing order to obtain a structured report.

[0097] Specifically, each decoder head is a modified version of the Transformer decoder structure, using relational memory and memory-driven conditional layers to effectively simulate doctors' clinical writing and generate professional descriptions to obtain corresponding anatomical structure reports.

[0098] During the training process, the present invention proposes a task-aware report generation algorithm framework. First, based on prior knowledge, the complete report is decomposed into descriptions of multiple anatomical structures, and the decomposed report is called a structured report. Different from the previous method of directly generating a complete report, the task-aware report generation algorithm generates descriptions of different structures in the task-aware report through different decoders. In such a framework, on the one hand, the length of text that each decoder needs to process is greatly reduced, which reduces the difficulty of modeling long texts. On the other hand, each decoder only focuses on the generation of reports of specific structures, so during training, it can pay attention to the information of the corresponding position in the image and reduce the interference of redundant information. In addition, the task-aware report generation method proposed by the present invention adopts a method of generating reports of different structures separately, so there is no influence of the order between structures, which also improves the accuracy of the generated report to a certain extent.

[0099] Based on the task-aware report generation algorithm framework, the present invention constructs a classification task. By first judging whether there are abnormalities in different structures, and then generating them separately based on this semantic information and combined with image information, the modeling difficulty of the medical image report generation task is further reduced.

[0100] In one implementation, the step of training the report generation model includes:

[0101] Step S10: obtaining a training data set, wherein the training data set includes an original training report;

[0102] Step S20: pre-processing the original training report to obtain a structured training report, and using the structured training report as a reference report;

[0103] Step S30: input the original training report into the initial report generation model, perform generation task training and classification task training on the initial report generation model, and obtain a structured generation report;

[0104] Step S40: When the total loss function of the generation task and the classification task reaches stability, the training is completed and a trained report generation model is obtained.

[0105] The initial report generation model is a CNN-Transformer model.

[0106] Specifically, this embodiment is based on the basic CNN-Transformer architecture. The model framework of the present invention can be divided into three parts: task distillation module, task-aware report generation module, and classification identification module. Figure 2 and Figure 3 The task distillation module operates on the original training reports in the training dataset, splitting them into different reports. The report generation module generates reports for different anatomical structures. The classification and identification module utilizes information from different parts to assist in report generation.

[0107] In one embodiment, step S20 specifically includes:

[0108] Step S21, extracting report keywords from the original training report;

[0109] Step S22: Obtain a chest X-ray knowledge graph, classify different sentences into different anatomical structures according to the report keywords and the chest X-ray knowledge graph, and obtain a structured training report;

[0110] Step S23: Use the structured training report as a reference report.

[0111] Specifically, since the core idea of ​​this embodiment is to generate reports for different anatomical structures, and the original reports are complete paragraphs, it is necessary to first decompose the original training reports in the training dataset to obtain structural reports. The task distillation module extracts report keywords and classifies different sentences into different structures based on the chest X-ray knowledge graph, thereby obtaining structured training reports.

[0112] In one implementation, the initial report generation model includes a convolutional neural network-based visual extractor and a linear projection. Step S30 specifically includes:

[0113] Step S31: extracting the image visual features in the original training report using a visual extractor based on a convolutional neural network to obtain a sequence of image visual feature vectors corresponding to the original training report;

[0114] Step S32: using linear projection to reduce the dimension of each image visual feature vector in the image visual feature vector sequence corresponding to the original training report to 512, to obtain a block embedding vector sequence corresponding to the original training report;

[0115] Step S33: creating a corresponding classification embedding vector for each anatomical structure in the original training report, to obtain a classification embedding vector sequence corresponding to the original training report;

[0116] Step S34: input the block embedding vector sequence and the classification embedding vector sequence of the original training report into a shared encoder to obtain the hidden state sequence and classification identification sequence of the original training report;

[0117] Step S35: input the hidden state sequence and classification identification sequence of the original training report into a multi-head decoder to obtain the corresponding anatomical structure training report output by each decoder head;

[0118] Step S36: splice the various anatomical structure training reports according to a preset splicing order to obtain a structured generation report.

[0119] The original training report includes original training images, such as chest X-ray images, which are first input into the task-aware report generation module. Similar to a general encoder-decoder structure, the encoder portion of the task-aware report generation module is used to extract the visual representation of the image, while the decoder portion is composed of multiple decoder heads, each of which uses the visual representation of the image to generate reports of different structures. This greatly reduces the length of text that each decoder needs to process, making it easier to model long text. Furthermore, each decoder focuses only on generating reports of a specific structure. Therefore, during training, it can focus on information at the corresponding position in the image, reducing interference from redundant information.

[0120] Specifically, the report generation module consists of an embedding vector generator module and a task-aware multi-head Transformer network, where each head corresponds to a description of a specific structure. The embedding vector generator module consists of a convolutional neural network-based visual extractor f V and a trainable linear projection f P For the input original training report (such as radioactive image X), the calculation process is: {e1,e2,…,e n}=f P (f V (X)).

[0121] In this embodiment, the ResNet-101 model is selected to extract visual features, that is, the visual extractor based on convolutional neural network can be a ResNet-101 model. In this algorithm framework, the last fully connected layer of the convolutional neural network is removed, and the feature map output by the last convolutional layer is adaptively averaged pooled to obtain an output of a fixed size of 7×7×2048, and it is regarded as a sequence of image visual feature vectors with a length of 49 and a dimension of 2048. Each vector represents the visual information of a certain area of ​​the image. Subsequently, linear projection reduces the dimension of each image visual feature vector to 512, and finally obtains a block embedding vector sequence {e1,e2,…,e n}, which will serve as the source input for the subsequent task-aware multi-head Transformer network.

[0122] The task-aware multi-head Transformer consists of a shared encoder and multiple separate decoders. The shared encoder is a stack of standard Transformer encoder blocks, which is used to transform the block embedding vector into a hidden state. This process can be expressed as: {h1,h2,…,h n}=f E ({e1,e2,…,e n}); where f E represents a shared encoder, {h1,h2,…,h n} represents the hidden state sequence. The role of the shared encoder is to use the Transformer's ability to capture global information and extract more effective visual feature information.

[0123] Unlike existing literature that generates an entire report at once, each head of the multi-head decoder is only responsible for generating a description of a specific anatomical structure. This not only prevents structures with fewer descriptions from being neglected, but also reduces the difficulty of the generative model in understanding complex reports. Specifically, each head is a modified version of the Transformer decoder structure, in which relational memory and memory-driven conditional layers are used to effectively simulate doctors' clinical writing and generate professional descriptions. Task-specific description Y = {y1, y2, ..., y m The generation of} can be expressed as: in Represents the Transformer decoder of the i-th head.

[0124] When writing radiology reports, doctors typically perform a pre-determination approach, first checking whether the area in question is abnormal and then reporting whether it is normal or abnormal. Inspired by this mechanism, a multi-task model combining multiple classification and generation techniques has been proposed for medical report generation. This multi-task model has proven effective in medical report generation, producing reports with more accurate diagnoses.

[0125] In addition, existing methods usually input the global image feature vector extracted by the convolutional neural network into different multi-layer perceptrons (MLP) to implement multiple binary classification tasks. This approach does not take into account that the image information required for different classification tasks is different. For example, the diagnosis of bone diseases does not require information about the heart. In addition, a simple convolutional neural network cannot establish long-range modeling. To this end, the implementation of the classification identification module of the present invention is based on the encoder side of the Transformer model, while utilizing the characteristics of the convolutional neural network's inductive bias and the global inductive modeling capability of the Transformer model. Its structure is as follows: Figure 3 shown.

[0126] First, a classification embedding vector is created for each structure, with the same dimension as the block embedding vector. The classification embedding vector is learnable and is input into the encoder of the Transformer model together with the block embedding vector sequence. The calculation process is as follows:

[0127] {h1,h2,…,h n},{t1,t2,…,t m}=f E ({e1,e2,…,e n},{c1,c2,…,c m});

[0128] Among them, {c1,c2,…,c m} represents the classification embedding vector sequence, {t1,t2,…,t m} represents the output classification token sequence.

[0129] The classification identification sequence and the hidden state sequence are input to the decoder to provide high-level features for abnormal areas. The report generation process can be modified as follows:

[0130] {h1,h2,…,h n},{t1,t2,…,t m}=f D ({e1,e2,…,e n},{c1,c2,…,c m}); where f D Represents a decoder.

[0131] In one implementation, the calculation formula of the total loss function is: Among them, the is the loss function of the generation task, is the loss function of the classification task, and λ is a hyperparameter used to adjust the loss ratio between the generation task and the classification task.

[0132] The loss function of the generation task is the superposition of the cross entropy loss functions generated by different anatomical structures. The calculation formula of the loss function of the generation task is: Wherein, m is the number of anatomical structures, l i is the length of the reference report for the ith anatomical structure, the r ij The jth word in the reference report of the i-th anatomical structure is represented by y ij Represents the jth word in the structured generated report for the i-th anatomical structure.

[0133] The loss function of the classification task is the superposition of the loss functions of multiple binary classification tasks. The calculation formula of the loss function of the classification task is: Among them, the g i is the category label, and the i is the prediction result; the category label includes a normal label and an abnormal label for each anatomical structure.

[0134] Specifically, to provide labels for classification training, this example extracts labels for common diseases and combines them with keywords associated with common, non-abnormal symptoms to derive normal and abnormal labels for each anatomical structure. To simulate the doctor's process of diagnosing and then writing a report, this example proposes a classification identification module to monitor the presence of abnormalities in different structures and assist the report generation module in producing reports with more accurate diagnoses.

[0135] In one embodiment, the step of obtaining the splicing order includes:

[0136] Get the sentence sequence set S of the original training report in the training data set = {s1, s2, ..., s N}, wherein N is the number of original training reports in the training dataset;

[0137] Construct a set of all possible statement sequences T = {t1, t2, ..., t M}, where M is the number of possible statement sequences;

[0138] Comparing all possible sentence sequences with the sentence sequences corresponding to the training dataset to obtain the degree of similarity of the overall sentence order between all possible sentence sequences and the sentence sequences of the training dataset;

[0139] The splicing order is obtained according to the similarity level.

[0140] Specifically, the report finally generated by this embodiment is at the structural level. In order to facilitate the calculation of corresponding indicators with the reference report and comparison with other algorithms, this embodiment splices reports of different structures in a certain order to obtain the final report. Although the order of statements in the diagnostic report does not affect the overall semantics of the report in practice, the ROUGE-L evaluation indicator is calculated based on the longest common subsequence between the reference report and the generated report, and will be affected by the order of statements. To this end, this embodiment simply designs a method based on the longest common subsequence to find a suitable statement order as much as possible.

[0141] First, we obtain the set of all reported sentence sequences S = {s1, s2, ..., s N}, where N is the number of original training reports in the training dataset. Then construct a set of all possible sentence sequences T = {t1, t2, ..., t M}, where M is the number of possible sentence sequences, and M = m!. For any possible sentence sequence, we compare it with all sentence sequences in the training set to determine its similarity to the overall sentence order of the training set.

[0142] The calculation formula of the similarity is: The LCS function is used to find the length of the longest common subsequence between two sequences. In this way, this embodiment can find a sentence sequence that is closest to the overall sentence order of the training set. Finally, t pos This is the splicing order finally adopted in this embodiment.

[0143] In one implementation, step S34 specifically includes: inputting the block embedding vector sequence and classification embedding vector sequence of the original training report into a shared encoder; the classification embedding vector sequence of the original training report corresponds to the output classification identifier, and a multi-layer perceptron is used to supervise each output classification identifier; according to each output classification identifier, the anatomical structure information in the block embedding vector sequence of the original training report is extracted to obtain the hidden state sequence and classification identifier sequence of the original training report. That is, this embodiment adds a multi-layer perceptron for supervision after each output classification identifier, so that it can extract information about specific structures in the block embedding vector. The process is expressed as: i =MLP i (t i ), i∈1,2,…,m; where MLP stands for multilayer perceptron, and different structures use separate multilayer perceptrons. i is the final generated binary classification probability vector, that is, the classification identification sequence.

[0144] Compared to previous methods, this embodiment improves the supervision objective. Existing literature primarily utilizes two types of supervision: one is the prediction of disease labels, and the other is the supervision of keywords extracted from reports. Disease labels are extracted using a tool called Chexpert Labeler, which includes 12 common diseases found on chest X-ray images, resulting in disease labels and no-abnormality labels. Accurately predicting disease labels is generally difficult, and due to the non-standard nature of report datasets, the network struggles to match disease labels with report content. Keyword supervision, however, presents the following challenges: Keywords extracted from reports primarily include structure, disease type, and corresponding descriptors. Structural keywords, such as "Volume," typically appear in most reports, so supervision of these keywords does not improve report quality. Corresponding descriptors include "Normal" and "Left," and the network cannot determine which structure these descriptors describe based on these keywords. The algorithm presented here, however, supervises the presence of abnormalities in different structures. Compared to directly predicting the disease, this determination is simpler and allows the network to focus on the corresponding image regions. At the same time, the supervision of whether different structures have anomalies is more consistent with the way the algorithm of the present invention generates different structures separately, and the information obtained by the classification can directly guide the generation network.

[0145] Further, if Figure 4 As shown, based on the above-mentioned task-awareness-based automatic report generation method, the present invention also provides a task-awareness-based automatic report generation device, including:

[0146] An input module 100 is configured to input an original report into a pre-trained report generation model and generate a sequence of block embedding vectors using an embedding vector generator;

[0147] A creation module 200 is used to create a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence;

[0148] An encoding module 300 is configured to input the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence;

[0149] The report generation module 400 is configured to input the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report divided into various anatomical structures.

[0150] like Figure 5As shown, the present invention also provides a terminal, comprising: a memory 20, a processor 10, and a task-aware report automatic generation program 30 stored on the memory 20 and executable on the processor 10, wherein the task-aware report automatic generation program 30 implements the steps of the task-aware report automatic generation method described above when executed by the processor 10.

[0151] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program can be executed to implement the steps of the task-awareness-based automatic report generation method as described above.

[0152] In summary, the present invention discloses a task-aware automatic report generation method and device, the method comprising: inputting the original report into a pre-trained report generation model, generating a block embedding vector sequence using an embedding vector generator; creating corresponding classification embedding vectors for each anatomical structure in the original report to obtain a classification embedding vector sequence; inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence; inputting the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report split into individual anatomical structures. The present invention utilizes a multi-head decoder in the report generation model to split the individual anatomical structures in the original report, and each decoder head only focuses on the report generation of the corresponding anatomical structure, thereby avoiding repeated sentences in the generated paragraphs and greatly reducing the length of text that each decoder head needs to process, reducing the difficulty of long text modeling, and improving the accuracy of report generation.

[0153] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A task-aware report automatic generation method, characterized in that: include: Inputting an original report including an image and a brief text description into a pre-trained report generation model and generating a block embedding vector sequence using an embedding vector generator; Creating a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence, wherein the dimension of the classification embedding vector is the same as the dimension of the block embedding vector; Inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence; the task-aware multi-head Transformer consists of a shared encoder and multiple separate decoders, wherein the shared encoder is stacked by standard Transformer encoder blocks to convert the block embedding vector into a hidden state; Inputting the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report divided into individual anatomical structures; each head of the multi-head decoder is only responsible for generating a description of a specific anatomical structure; The method of inputting the original report into a pre-trained report generation model and using an embedding vector generator to generate a block embedding vector sequence includes: Inputting the original report into a pre-trained report generation model, and using a visual extractor based on a convolutional neural network to extract image visual features of the original report to obtain an image visual feature vector sequence; Performing linear projection processing on each image visual feature vector in the image visual feature vector sequence to obtain a block embedding vector sequence; Inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence, including: Inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder; The classification embedding vector sequence corresponds to the output classification identification, and a multi-layer perceptron is used to supervise each output classification identification; Extracting anatomical structure information from the block embedding vector sequence according to each output classification identifier to obtain a hidden state sequence and a classification identifier sequence; The hidden state sequence and the classification identification sequence are input into a multi-head decoder to obtain a structured report divided into various anatomical structures, including: Inputting the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain corresponding anatomical structure reports output by each decoder head; The various anatomical structure reports are spliced ​​together according to a preset splicing order to obtain a structured report.

2. The task-aware report automatic generation method according to claim 1, characterized in that: The training steps of the report generation model include: Acquire a training data set, wherein the training data set includes an original training report; Preprocessing the original training report to obtain a structured training report, and using the structured training report as a reference report; Inputting the original training report into an initial report generation model, performing generation task training and classification task training on the initial report generation model to obtain a structured generation report; When the total loss function of the generation task and the classification task reaches a stable state, the training is completed and the trained report generation model is obtained; The initial report generation model is a CNN-Transformer model.

3. The task-aware report automatic generation method according to claim 2, characterized in that: Preprocessing the original training report to obtain a structured training report, and using the structured training report as a reference report, including: extracting report keywords from the original training report; Obtaining a chest X-ray knowledge graph, classifying different sentences into different anatomical structures according to the report keywords and the chest X-ray knowledge graph, and obtaining a structured training report; The structured training report is used as a reference report.

4. The task-aware report automatic generation method according to claim 2, characterized in that: The initial report generation model includes a convolutional neural network-based visual extractor and a linear projection; the original training report is input into the initial report generation model, and the initial report generation model is trained for a generation task and a classification task to obtain a structured generation report, including: Extracting the image visual features in the original training report using a visual extractor based on a convolutional neural network to obtain an image visual feature vector sequence corresponding to the original training report; Using linear projection, the dimension of each image visual feature vector in the image visual feature vector sequence corresponding to the original training report is reduced to 512, so as to obtain a block embedding vector sequence corresponding to the original training report; Creating a corresponding classification embedding vector for each anatomical structure in the original training report to obtain a classification embedding vector sequence corresponding to the original training report; Inputting the block embedding vector sequence and the classification embedding vector sequence of the original training report into a shared encoder to obtain the hidden state sequence and the classification identification sequence of the original training report; Inputting the hidden state sequence and classification identification sequence of the original training report into a multi-head decoder to obtain the corresponding anatomical structure training report output by each decoder head; Each anatomical structure training report is spliced ​​in a preset splicing order to obtain a structured generation report.

5. The task-aware report automatic generation method according to claim 2, characterized in that: The calculation formula of the total loss function is: ; Among them, the is the loss function of the generation task, is the loss function of the classification task, is a hyperparameter used to adjust the loss ratio between the generation task and the classification task; The calculation formula of the loss function of the generation task is: ; Among them, the is the number of anatomical structures, It is The length of the reference report for each anatomical structure, Indicates the Reference report on anatomical structures words, the Indicates the The first words; The calculation formula of the loss function of the classification task is: ; Among them, the is the category label, To predict the results; The category labels include a normal label and an abnormal label for each anatomical structure.

6. The task-aware report automatic generation method according to claim 4, characterized in that: The step of obtaining the splicing order includes: Get the sentence sequence set of the original training report in the training dataset , wherein the is the number of original training reports in the training dataset; Construct the set of all possible statement sequences , wherein the is the number of possible statement sequences; Comparing all possible sentence sequences with the sentence sequences corresponding to the training dataset to obtain the degree of similarity of the overall sentence order between all possible sentence sequences and the sentence sequences of the training dataset; The splicing order is obtained according to the similarity level.

7. The task-aware report automatic generation method according to claim 6, characterized in that: Inputting the block embedding vector sequence and the classification embedding vector sequence of the original training report into a shared encoder to obtain the hidden state sequence and the classification identification sequence of the original training report, including: Inputting the block embedding vector sequence and the classification embedding vector sequence of the original training report into a shared encoder; The classification embedding vector sequence of the original training report corresponds to the output classification identifier, and a multi-layer perceptron is used to supervise each output classification identifier; The anatomical structure information in the block embedding vector sequence of the original training report is extracted according to each output classification identifier to obtain the hidden state sequence and classification identifier sequence of the original training report.

8. A task-aware report automatic generation device, characterized in that: include: an input module, configured to input an original report including an image and a brief text description into a pre-trained report generation model and generate a block embedding vector sequence using an embedding vector generator; A creation module is used to create a corresponding classification embedding vector for each anatomical structure in the original report to obtain a classification embedding vector sequence, where the dimension of the classification embedding vector is the same as the dimension of the block embedding vector; An encoding module is configured to input the block embedding vector sequence and the category embedding vector sequence into a shared encoder to obtain a hidden state sequence and a category identification sequence; the task-aware multi-head Transformer consists of a shared encoder and multiple separate decoders, wherein the shared encoder is stacked by standard Transformer encoder blocks to convert the block embedding vector into a hidden state; a report generation module, configured to input the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain a structured report divided into individual anatomical structures; each head of the multi-head decoder is only responsible for generating a description of a specific anatomical structure; The method of inputting the original report into a pre-trained report generation model and using an embedding vector generator to generate a block embedding vector sequence includes: Inputting the original report into a pre-trained report generation model, and using a visual extractor based on a convolutional neural network to extract image visual features of the original report to obtain an image visual feature vector sequence; Performing linear projection processing on each image visual feature vector in the image visual feature vector sequence to obtain a block embedding vector sequence; Inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder to obtain a hidden state sequence and a classification identification sequence, including: Inputting the block embedding vector sequence and the classification embedding vector sequence into a shared encoder; The classification embedding vector sequence corresponds to the output classification identification, and a multi-layer perceptron is used to supervise each output classification identification; Extracting anatomical structure information from the block embedding vector sequence according to each output classification identifier to obtain a hidden state sequence and a classification identifier sequence; The hidden state sequence and the classification identification sequence are input into a multi-head decoder to obtain a structured report divided into various anatomical structures, including: Inputting the hidden state sequence and the classification identification sequence into a multi-head decoder to obtain corresponding anatomical structure reports output by each decoder head; The various anatomical structure reports are spliced ​​together according to a preset splicing order to obtain a structured report.

9. A terminal, characterized in that: include: A memory, a processor, and a task-aware report automatic generation program stored in the memory and executable on the processor, wherein the task-aware report automatic generation program, when executed by the processor, implements the steps of the task-aware report automatic generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the task-awareness-based automatic report generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Zero-sample picture classification method based on auto-encoder

    CN112487193A

  • Brain CT medical report generation method based on hierarchical self-attention sequence coding

    CN112614561A