Training Method and System for Radiology Report Generation Model Based on Multimodal Contrastive Learning

Through the multimodal contrast learning method, the visual characteristics of medical images and the semantic characteristics of text are integrated, and the problem of semantic inconsistency in the generation of radiological reports in the prior art is solved, achieving higher clinical accuracy and consistency.

CN115293128BActive Publication Date: 2025-05-30SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210931458.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-05-30
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

The prior art has problems in the generation of radiological reports, such as inconsistent visual representation and text semantic representation of medical images, resulting in semantic inconsistency in the generated reports and affecting clinical accuracy.

Method used

The multimodal contrast learning method is adopted, through self-supervised representation learning, the image encoder and sentence encoder are learned, and the visual and semantic features are fused, and the Impression and Findings parts of the radiological report are generated recursively.

Benefits of technology

Improves consistency between images and text, enhances medical semantic coherence and accuracy of reports, and reduces errors in generating reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293128B_ABST
    Figure CN115293128B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method and system for a radiology report generation model based on multi-modal contrastive learning, including: adopting a self-supervised representation learning method, learning sentence representations based on a contrastive learning-based sentence-level training strategy, and obtaining image representations through bidirectional contrastive learning between paired images and texts; then embedding the learned image encoder and sentence encoder into the radiology report generation model, generating the Impression part of the report through an encoding-decoding process, and recursively generating the Findings part of the report by fusing visual features and semantic features. The training method and system for the radiology report generation model based on multi-modal contrastive learning provided by the present invention optimize the representations of images and texts, realize the radiology report generation task, can assist radiologists in making diagnoses, and improve the accuracy in the diagnosis process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for training a radiology report generation model based on multi-modal contrastive learning, belonging to the fields of computer and medicine. Background Art

[0002] Medical images, such as radiology images, are widely used for disease diagnosis. The reading and understanding of medical images are usually carried out by professional medical staff. They identify the normal and abnormal areas in the examined images by analyzing the images and use the learned medical knowledge and accumulated work experience to write radiology reports. However, due to factors such as the lack of knowledge, incorrect reasoning, shortage of personnel, and excessive workload of radiologists, the generated reports may be incorrect. Therefore, the automatic generation of radiology reports has become an attractive research direction in artificial intelligence and clinical medicine to reduce the workload of radiologists and minimize the occurrence of errors.

[0003] Imaging studies are usually accompanied by radiology reports, which record the observations of radiologists in daily clinical care. The paragraphs consisting of the Impression and Findings sections represent the most direct transcriptions of the imaging studies. Among them, the Impression section is a conclusive diagnosis and can be regarded as the conclusion or topic sentence of the report, while the Findings section is a paragraph composed of multiple structural sentences, and each sentence focuses on the specific medical observations of a certain specific area in the radiology image. These sentences are usually longer and more complex than the captions in the standard image caption dataset. Therefore, many existing image caption models are not directly applicable to this task, which requires special solutions.

[0004] In recent years, many works have explored the automatic generation of radiology reports and proposed many new ideas based on traditional generation methods for improvement. However, the existing generation methods generally have several problems: (1) Most previous studies used CNN encoders pre-trained based on ImageNet, and this pre-trained model is not applicable to medical images; (2) The quality of the sentence representations obtained by the LSTM-based or BERT-based sentence encoders used in previous studies is not high; (3) The generated reports are semantically inconsistent in the medical field, and the optimization of the clinical accuracy of the generated reports is achieved at the cost of reducing the performance of other indicators to a certain extent. Summary of the Invention

[0005] The objective of the present invention is to: achieve the radiology report generation task, improve the visual representation of the extracted medical images and the semantic representation of the text, and enhance the consistency between the image and the text.

[0006] To achieve the above object, a technical solution of the present invention is to provide a training method for a radiology report generation model based on multi-modal contrast learning, which is characterized by including the following steps:

[0007] S100. Sample data collection:

[0008] Collect radiology images and text data, and transmit the image data and the corresponding text data to the sample database, where each text data includes a conclusive diagnostic statement Impression and a detailed description paragraph Findings;

[0009] S200. Multi-modal contrast learning:

[0010] Adopt a self-supervised representation learning method to learn an image encoder for visual features in image data and a sentence encoder for extracting semantic features in text data based on the sample database, where:

[0011] The sentence encoder learns sentence representations through a sentence-level training strategy based on contrast learning;

[0012] The image encoder learns image representations through a bidirectional contrast learning objective between paired image data and text data;

[0013] S300. Radiology report generation:

[0014] By fusing the visual features extracted by the image encoder and the semantic features extracted by the sentence encoder, recursively generate the diagnostic statement Impression and the description paragraph Findings of the radiology report, specifically including the following steps:

[0015] S301. The Impression generation module generates a single diagnostic statement Impression based on the encoder-decoder framework, specifically including the following steps:

[0016] The image encoder extracts visual features from the input image, and then sends them to the sentence decoder in the Impression part to generate the entire sentence word by word as the diagnostic statement Impression;

[0017] S302. The Findings generation module fuses the visual features of the image and the semantic features of the sentence, generates sentences in a loop, and finally generates a long paragraph containing multiple structural statements, specifically including the following steps:

[0018] In the Findings generation module, to make the generated sentences focus on describing different image regions, based on the attention framework, the visual features output by the image encoder and the semantic features of the previous sentence output by the sentence encoder are input into a fully connected layer, and then fed into the SoftMax layer to obtain weighted visual features. The weighted visual features are fed into the Findings partial sentence decoder, and the encoding of the sentence is obtained by the Findings partial sentence decoder. The encoding of the sentence is used as the encoding of the previous sentence and input into the sentence encoder, and then input into the sentence encoder to obtain new weighted visual features. This process is repeated continuously until the Findings partial sentence decoder generates an empty sentence, indicating that the generation of the description paragraph Findings has been completed;

[0019] Among them, the encoding of the previous sentence and the visual features of the image are combined through the weighted visual features to guide the generation of the next sentence. The attention network used to calculate the weighted visual representation is defined as:

[0020] V w = Attention(v, s)

[0021] Among them, V w is the weighted visual representation to be obtained; attention() is the attention function; v = {v 1 , v 2 , …, v k}, are the image features learned by the image encoder, and each feature v i is a representation of dimension d v , corresponding to a part of the image; represents the encoding of the previous sentence, and d s is the dimension of the semantic features;

[0022] Given the visual features and the encoding of the previous sentence, an attention distribution is generated over K regions of the image through a single attention allocator, represented by a single-layer neural network and a SoftMax function:

[0023]

[0024] α = softmax(z)

[0025] In the formula: is a vector with all elements set to 1; are the parameters of the attention network; is the weight of the feature in v in the attention pair (v, s).

[0026] The weighted visual representation V based on the attention distribution w is obtained by the following:

[0027]

[0028] where α i is the i-th dimensional element in α;

[0029] S400, Report Analysis and Evaluation:

[0030] Use evaluation metrics to evaluate the generated radiology reports:

[0031] S500, Result Output:

[0032] Combine the separately generated diagnostic statement Impression and the descriptive paragraph Findings into a complete radiology report for output, and at the same time use multiple evaluation metrics to evaluate the output result.

[0033] Preferably, in step S100, the image data and text data in the sample database are in a one-to-one correspondence, and include a training set, a test set, and a validation set.

[0034] Preferably, in step S200, learning the sentence representation through a sentence-level training strategy based on contrastive learning includes the following steps:

[0035] For the same sentence in a set of sentences, a series of different enhanced versions of the sentence representation are obtained as positive examples in the training set, test set, or validation set by implementing different data augmentation methods, and other sentences are used as negative examples;

[0036] When training the sentence encoder, by maximizing the consistency between the sentence representations of different enhanced versions of the same sample while keeping the sentence vectors of different samples as far apart as possible, a sentence encoder is established and a semantic sentence embedding is constructed.

[0037] Preferably, in step S200, take a set of sentences x i represents the i-th sentence, m represents the total number of sentences in the set ; apply two different data augmentation methods f() and f′() to the sentence x i to generate two different versions of the sentence embedding e i , e′ i :

[0038] e i = f(x i )

[0039] e′ i= f′(x i )

[0040] where e i , L is the length of the sentence embedding and D is the hidden dimension of the sentence embedding;

[0041] Then, the sentence embeddings e i , e′ i are encoded to obtain sentence representations h i , h′ i .

[0042] For a mini-batch of N sentences, the training objective i of sentence x is as follows:

[0043]

[0044] where τ is a temperature hyperparameter and sim() is the cosine similarity, then

[0045] The final contrastive loss is the average of all N within-batch losses:

[0046] Preferably, in step S200, the image representation is learned through a bi-directional contrastive learning objective between paired images and texts, including the following steps:

[0047] The image encoder is learned through paired inputs (X v , X s ), where X v represents an image or a set of images, and X s represents a sentence sequence describing the imaging information in X v ; for each input image X v and each input sentence X d , they are converted into a fixed-dimensional vector h v and h s by an image encoder f v () and a sentence encoder f s (); then, the representations h v , h s of the two modalities are projected from their encoder spaces to the same D-dimensional space through projection functions g v () and g s () for contrastive learning;

[0048] For N input pairs (X v , X s) to obtain the corresponding N pairs of representations (v, s), where:

[0049] v = g v (f v (X v ))

[0050] s = g s (f s (X s ))

[0051] In the formula, v,

[0052] Using (v i , s i ) to represent the i-th pair of representations, its training objective includes two loss functions: image-text contrast loss and text-image contrast loss

[0053]

[0054]

[0055] The final training loss is a weighted combination of the two losses:

[0056]

[0057] In the formula, λ is a scalar weight;

[0058] By maximizing the consistency between image-text representation pairs, an image encoder is learned that maps an image to a vector of a fixed dimension.

[0059] Preferably, in step S301, the Impression part sentence decoder adopts an LSTM-based method to generate a word at each time step to produce a caption according to the context vector, the previous hidden state, and the previously generated word;

[0060] The initial hidden state and cell state of the LSTM are set to zero, and the hidden state h at time t t is modeled as:

[0061] h t = LSTM(x t , h t-1 , m t-1 )

[0062] where, x t is the input vector, and m t-1 is the memory cell vector at time t-1.

[0063] The visual feature vector is used as the initial input of the LSTM to predict the first word of the sentence; before being input into the LSTM, a fully connected layer converts the visual feature vector output by the image encoder into the same dimension as the word embedding. The LSTM extracts visual features from the last convolutional layer and then generates the entire sentence word by word.

[0064] Preferably, the step S400 includes:

[0065] S401. Evaluate the generated report using BLEU and its variants;

[0066] S402. Evaluate the generated report using METEOR;

[0067] S403. Evaluate the generated report using ROUGE;

[0068] S404. Evaluate the generated report using CIDEr.

[0069] Another technical solution of the present invention is to provide a training system for a radiology report generation model based on multi-modal contrast learning, which is characterized by including:

[0070] A sample database for storing the collected radiology images and text data, wherein each text data includes a conclusive diagnostic statement Impression and a detailed description paragraph Findings;

[0071] A multi-modal contrast learning module that uses a self-supervised representation learning method to learn an image encoder for visual features in the image data and a sentence encoder for extracting semantic features in the text data based on the sample database, wherein:

[0072] The sentence encoder learns sentence representations through a sentence-level training strategy based on contrast learning;

[0073] The image encoder learns image representations through a bidirectional contrast learning objective between paired image data and text data;

[0074] A radiology report generation module that recursively generates the diagnostic statement Impression and the description paragraph Findings of the radiology report by fusing the visual features extracted by the image encoder and the semantic features extracted by the sentence encoder;

[0075] The radiology report generation module further includes an Impression generation module and a Findings generation module, wherein:

[0076] The Impression generation module generates a single diagnostic statement Impression based on the encoder-decoder framework. The implementation of the Impression generation module includes the following steps:

[0077] The image encoder extracts visual features from the input image and then feeds them into the Impression partial sentence decoder to generate the entire sentence word by word as the diagnostic statement Impression;

[0078] The Findings generation module fuses the visual features of the image and the semantic features of the sentence to generate sentences in a loop, and finally generates a long paragraph containing multiple structural statements. The implementation of the Findings generation module includes the following steps:

[0079] In the Findings generation module, in order to make the generated sentences focus on describing different image regions, based on the attention framework, the visual features output by the image encoder and the semantic features of the previous sentence output by the sentence encoder are input into a fully connected layer, and then fed into the SoftMax layer to obtain weighted visual features. The weighted visual features are fed into the Findings partial sentence decoder, and the encoding of the sentence is obtained by the Findings partial sentence decoder. The encoding of the sentence is used as the encoding of the previous sentence and input into the sentence encoder, and then input into the sentence encoder to obtain new weighted visual features. This process is repeated continuously until the Findings partial sentence decoder generates an empty sentence, indicating that the generation of the Findings part has been completed;

[0080] Among them, the encoding of the previous sentence and the visual features of the image are combined through the weighted visual features to guide the generation of the next sentence. The attention network used to calculate the weighted visual representation is defined as:

[0081] V w = Attention(v, s)

[0082] Among them, V w is the weighted visual representation to be obtained; Attention() is the attention function; v = {v 1 , v 2 , …, v k}, is the image feature learned by the image encoder, and each feature v i is a representation of dimension d v , corresponding to a part of the image; represents the encoding of the previous sentence, and d s is the dimension of the semantic feature;

[0083] In view of the visual features The encoding of the previous sentence generates an attention distribution over K regions of the image through a single attention allocator, which is represented by a single-layer neural network and a SoftMax function:

[0084]

[0085] α = softmax(z)

[0086] where: is a vector with all elements set to 1; are the parameters of the attention network; is the weight of the feature in v in the attention pair (v, s).

[0087] Based on the attention distribution, the weighted visual representation V w is obtained as follows:

[0088]

[0089] where α i is the i-th dimensional element in α;

[0090] The report analysis and evaluation module uses evaluation metrics to evaluate the generated radiology report;

[0091] The result output module is used to combine the separately generated diagnostic statement Impression and the description paragraph Findings into a complete radiology report for output, and at the same time use multiple evaluation metrics to evaluate the output result.

[0092] A radiology report generation model training system according to claim 8, characterized in that, in the radiology report generation module, the image encoder is modeled as a CNN; the sentence encoder is modeled as a model obtained by fine-tuning in a BERT language model, and the semantic representation of the sentence is generated by an average pooling layer.

[0093] The structure of the present invention is reasonably designed, self-supervised learning is carried out using a sample database, radiology images are sent into the system, and finally corresponding radiology reports are generated to assist radiologists in making diagnoses, greatly improving the accuracy in the diagnosis process.

[0094] In summary, compared with the prior art, the present invention has at least the following advantages:

[0095] (1) The present invention proposes a recursive model based on multi-modal contrastive learning for generating radiology reports. This model combines the visual features of medical images and the semantic features of sentences, and generates the Impression and Findings parts of radiology reports through a recursive network respectively;

[0096] (2) The present invention proposes a model pre-training method based on multi-modal contrastive learning to improve the expressiveness of visual and text representations;

[0097] (3) The present invention performs bidirectional contrastive learning with paired medical images and reports to pre-train the image encoder, enabling it to effectively extract visual representations and improve the consistency between image data and text data.

[0098] (4) The present invention builds a sentence encoder based on a contrastive learning-based sentence-level training objective, enabling it to construct semantically coherent sentence embeddings for text representation;

[0099] (5) The training method and system of a radiology report generation model based on multi-modal contrastive learning proposed by the present invention can effectively provide interpretable reasons, perform self-supervised learning using a sample database, input radiology images into the system, and finally generate corresponding radiology reports to assist radiologists in making decisions, greatly improving the accuracy in the diagnosis process. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 is the overall framework diagram of the present invention;

[0101] Figure 2 is the sample data of an embodiment of the present invention;

[0102] Figure 3 is the framework diagram of multi-modal contrastive learning in the present invention;

[0103] Figure 4 is the schematic diagram of multi-modal feature fusion in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0104] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0105] Such as Figure 1As shown, an embodiment of the present invention proposes a training method for a radiology report generation model based on multi-modal contrast learning. By embedding the self-supervised representation learning method of multi-modal contrast learning into the radiology report generation model, a radiology report corresponding to each radiology image is generated, including the following steps:

[0106] S100. Sample data collection:

[0107] Collect radiology images of the human chest through a nuclear magnetic resonance instrument, and transmit the image data and the corresponding text data to the sample database. Specifically, referring to Figure 2 , the image data and text data in the sample database are in a one-to-one correspondence, and include a training set, a test set, and a validation set. The text data is a radiology report, including a conclusive diagnostic statement "Impression" and a detailed description paragraph "Findings".

[0108] S200. Multi-modal contrast learning:

[0109] Adopt a self-supervised representation learning method to learn an image encoder and a sentence encoder based on the sample database, which are respectively used to extract visual features in the image and semantic features in the sentence. As Figure 3 shown, it specifically includes the following steps:

[0110] S201. Learn sentence representations through a contrast learning-based sentence-level training strategy.

[0111] Take a set of sentences x i to represent the i-th sentence, and m represents the total number of sentences in the set . Apply two different data augmentation methods f() and f′() to the sentence x i to generate two different versions of sentence embeddings e i , e′ i :

[0112] e i = f(x i )

[0113] e′ i = f′(x i )

[0114] In the formula, e i , L is the length of the sentence embedding, and D is the hidden dimension of the sentence embedding.

[0115] Then, the sentence embeddings e i , e′ i are encoded to obtain sentence representations h i , h′i 。

[0116] Therefore, for the same sentence, by implementing different data augmentation methods, such as truncation, deletion, etc., the present invention obtains a series of different embeddings as "positive examples", and uses other sentences in the same batch as "negative examples". Then, for a mini-batch of N sentences, sentence x i The training objective is as follows:

[0117]

[0118] where τ is a temperature hyperparameter and sim() is the cosine similarity, then there is

[0119] The numerator in the above formula represents the cosine similarity between h i , h' i , where h' i is a positive example; and the denominator represents the sum of the cosine similarities between h j , h' j , where h' j includes all positive and negative examples.

[0120] The final contrastive loss is the average of all N within-batch losses:

[0121]

[0122] By maximizing the consistency between different augmented versions of the same sample while keeping the sentence vectors of different samples as far apart as possible, the present invention establishes a sentence encoder and constructs semantic sentence embeddings.

[0123] S202. Learn image representations through a bi-directional contrastive learning objective between paired images and texts.

[0124] Learn an image encoder through paired inputs (X v , X s ), where X v represents an image or a set of images, and X s represents a sentence sequence describing the imaging information in X v . For each input image X v and each input sentence X s , they are converted into a fixed-dimensional vector h v () and a sentence encoder f s () into a fixed-dimensional vector h v and h s . Then, the representations h v, h s Projected through the projection functions g v () and g s () from their encoder spaces into the same D-dimensional space for contrastive learning.

[0125] Therefore, for N input pairs (x v , X s ), the corresponding N representation pairs (v, s) can be obtained, where:

[0126] v = g v (f v (X v ))

[0127] s = g s (f s (x s ))

[0128] In the formula, v,

[0129] Use (v i , s i ) to represent the i-th pair of representation pairs, and train using the same medical dataset as the downstream task. The present invention uses Info-NCE as the loss function, which is a contrastive loss function for self-supervised learning. The training objective of the i-th pair of representation pairs (v i , s i ) includes two loss functions: image-text contrast loss and text-image contrast loss

[0130]

[0131]

[0132] The final training loss is a weighted combination of the two losses:

[0133]

[0134] In the formula, λ is a scalar weight.

[0135] By maximizing the consistency between image-text representation pairs, the present invention learns an image encoder that maps images to a fixed-dimensional vector.

[0136] S203. Embed the image encoder and sentence encoder learned through the above steps into the radiology report generation model.

[0137] S300. Radiology report generation:

[0138] By fusing the visual features extracted by the image encoder and the semantic features extracted by the sentence encoder, recursively generate the Impression part and the Findings part of the radiology report, specifically including the following steps:

[0139] S301. The Impression generation module generates a single conclusive diagnostic statement Impression based on a simple encoder-decoder framework.

[0140] Specifically, first, the image encoder extracts visual features from the input image, and then sends them to the sentence decoder to generate the entire sentence word by word.

[0141] The purpose of the image encoder is to automatically extract visual features from the image, map the image into a context vector, which serves as the visual input for all subsequent modules. This vector is obtained through multi-modal contrastive pre-training. Specifically, the image encoder is parameterized as a fully connected layer, and the visual features are extracted from the last convolutional layer. Then, the visual features are sent to the sentence decoder to generate the Impression part.

[0142] The present invention adopts an LSTM-based method to generate a word at each time step to produce the title according to the context vector, the previous hidden state, and the previously generated word. The initial hidden state and cell state of the LSTM are set to zero, and the hidden state h at time t t is modeled as:

[0143] h t = LSTM(x t , h t-1 , m t-1 )

[0144] where x t is the input vector, and m t-1 is the memory cell vector at time t-1.

[0145] The visual feature vector is used as the initial input of the LSTM to predict the first word of the sentence. However, before inputting into the LSTM, a fully connected layer is required to convert the visual feature vector into the same dimension as the word embedding. Then, the entire sentence can be generated word by word.

[0146] S302. The Findings generation module fuses the visual features of the image and the semantic features of the sentence, generates sentences in a loop, and finally generates a long paragraph containing multiple structural statements.

[0147] Specifically, the multi-modal feature fusion refers to Figure 4 , Figure 4The upper part generates branches of visual features. The image encoder obtained through contrastive pre-training is modeled as a CNN for extracting visual representations from input images. Figure 4 The lower part represents the branch for generating semantic features. The sentence encoder obtained through contrastive pre-training is modeled as a model fine-tuned in a language model similar to BERT, and the semantic representation of the sentence is generated through an average pooling layer.

[0148] To make the generated sentences focus on describing different image regions, based on the attention framework, the visual features of the image and the semantic features of the text are input into a fully connected layer and then fed into the SoftMax layer to obtain weighted visual features.

[0149] The attention network for calculating the weighted visual representation is defined as:

[0150] V w = Attention(v, s)

[0151] where V w is the weighted visual representation to be obtained; Attention() is the attention function; v = {v 1 , v 2 , …, v k}, are the image features learned by the image encoder, and each feature v i is a representation of dimension d v , corresponding to a part of the image; represents the encoding of the previous sentence, and d s is the dimension of the semantic feature.

[0152] Given the visual features and the encoding of the previous sentence, the present invention generates an attention distribution over K regions of the image through a single attention allocator, represented by a single-layer neural network and a SoftMax function:

[0153]

[0154] α = softmax(z)

[0155] where: is a vector with all elements set to 1; are the parameters of the attention network; is the weight of the feature in v in the attention pair (, s).

[0156] Based on the attention distribution, the weighted visual representation V w can be obtained in the following way:

[0157]

[0158] where α i is the i-th dimensional element in α.

[0159] Now, the input to the sentence decoder is the weighted visual representation, so that the decoder will pay attention to specific regions of the image in order to generate sentences describing different image regions. The encoding of the previously learned sentence and the visual features of the image are combined to guide the generation of the next sentence. This process is repeated until an empty sentence is generated, indicating that the generation of the Findings section is complete. In this way, as different sentences are generated, the model can focus on different regions of the image according to the context of the previous sentences and ensure the coherence and consistency of the medical semantics of the generated report.

[0160] S400. Report analysis and evaluation:

[0161] Four common evaluation metrics are used to evaluate the generated radiology reports. The larger the value of the evaluation metric, the better the performance of the radiology report generation model. The specific steps are as follows:

[0162] S401. Evaluate the generated report using BLEU and its variants.

[0163] BLEU is a method for automatically evaluating machine translation. Its general idea is accuracy and can be further divided into many variants according to "n-gram". Four common metrics are BLEU-1, BLEU-2, BLEU-3, and BLEU-4, where n-gram refers to the number of consecutive words being n.

[0164] S402. Evaluate the generated report using METEOR.

[0165] METEOR is an automatic metric for machine translation evaluation, which has a better correlation with human judgment. Different from BLEU, METEOR considers both the accuracy and recall rate based on the entire corpus and obtains the final metric.

[0166] S403. Evaluate the generated report using ROUGE.

[0167] ROUGE is designed to measure the quality of abstracts. It measures the "similarity" between the automatically generated abstract and the reference abstract and calculates the corresponding score.

[0168] S404. Evaluate the generated report using CIDEr.

[0169] CIDEr is a consensus-based image description evaluation that calculates the cosine similarity between the reference caption and the caption generated by the model as the score.

[0170] S500, Result Output:

[0171] The separately generated Impression part and Findings part are combined into a complete radiology report for output, and at the same time, multiple evaluation indicators are used to evaluate the output result.

[0172] The result output module is signal-connected to a display screen and a printer. By setting that the result output module is signal-connected to a display screen and a printer, the screen display and document printing of the diagnostic report are realized, which is convenient for medical staff to analyze the report.

[0173] In this embodiment, the radiological image is obtained by a nuclear magnetic resonance instrument. The principle of the nuclear magnetic resonance instrument is to place the human body in a special magnetic field, excite the hydrogen nuclei in the human body with radio frequency pulses, cause the hydrogen nuclei to resonate and absorb energy. After stopping the radio frequency pulses, the hydrogen nuclei emit radio signals at a specific frequency and release the absorbed energy, which is collected by a receiver outside the body and processed by an electronic computer to obtain an image. The amount of information provided by the image generated by nuclear magnetic resonance is not only greater than that of many other imaging techniques in medical imaging, but also different from the existing imaging techniques. Therefore, it has great potential advantages for the diagnosis of diseases.

[0174] In this embodiment, the result output module is signal-connected to a display screen and a printer. By setting that the result output module is signal-connected to a display screen and a printer, the screen display and document printing of the diagnostic report are realized, which is convenient for medical staff to analyze the report.

Claims

1. A training method for a radiology report generation model based on multi-modal contrast learning, characterized in that, it includes the following steps: S100. Sample data collection: Collect radiology images and text data, and transmit the image data and the corresponding text data to the sample database. Among them, each text data contains a conclusive diagnostic statement Impression and a detailed description paragraph Findings; S200. Multi-modal contrast learning: Adopt a self-supervised representation learning method to learn an image encoder for visual features in the image data and a sentence encoder for extracting semantic features in the text data based on the sample database, where: The sentence encoder learns sentence representations through a contrastive learning-based sentence-level training strategy; The image encoder learns image representations through a bidirectional contrastive learning objective between paired image data and text data; S300. Radiology report generation: By fusing the visual features extracted by the image encoder and the semantic features extracted by the sentence encoder, recursively generate the diagnostic statement Impression and the description paragraph Findings of the radiology report, specifically including the following steps: S301. The Impression generation module generates a single diagnostic statement Impression based on the encoder-decoder framework, specifically including the following steps: The image encoder extracts visual features from the input image, and then sends them to the sentence decoder in the Impression part to generate the entire sentence word by word as the diagnostic statement Impression; S302. The Findings generation module fuses the visual features of the image and the semantic features of the sentence, generates sentences in a loop, and finally generates a long paragraph containing multiple structural statements, specifically including the following steps: In the Findings generation module, in order to make the generated sentences focus on describing different image regions, based on the attention framework, the visual features output by the image encoder and the semantic features of the previous sentence output by the sentence encoder are input into a fully connected layer, and then sent to the SoftMax layer to obtain weighted visual features. The weighted visual features are sent to the sentence decoder in the Findings part. The sentence decoder in the Findings part obtains the encoding of the sentence, and takes the encoding of the sentence as the encoding of the previous sentence and inputs it into the sentence encoder. Input the sentence encoder to obtain new weighted visual features. This process is repeated continuously until the sentence decoder in the Findings part generates an empty sentence, which indicates that the generation of the description paragraph Findings has been completed; Among them, the encoding of the previous sentence and the visual features of the image are combined through the weighted visual features to guide the generation of the next sentence. The attention network for calculating the weighted visual representation is defined as: V w = Attention(v, s) Among them, V w is the weighted visual representation to be obtained; Attention() is the attention function; v = {v 1 , v 2 , …, v k}, are the image features learned by the image encoder, and each feature v i is a representation of dimension d v , corresponding to a part of the image; represents the encoding of the previous sentence, and d s is the dimension of the semantic feature; Given visual features and the encoding of the previous sentence, an attention distribution is generated over K regions of the image by a single attention allocator, represented by a single-layer neural network and a SoftMax function: α = softmax(z) where: is a vector with all elements set to 1; are the parameters of the attention network; is the weight of the feature in v in the attention pair (v, s); Weighted visual representation V based on the attention distribution w is obtained by the following: where α i is the i-th dimensional element in α; S400. Report analysis and evaluation: Use evaluation metrics to evaluate the generated radiology reports: S500. Result output: Combine the separately generated diagnostic statement Impression and the descriptive paragraph Findings into a complete radiology report for output, and at the same time use multiple evaluation metrics to evaluate the output results.

2. A method for training a radiology report generation model based on multi-modal contrast learning as claimed in claim 1, characterized in that, in step S100, the image data and text data in the sample database are in a one-to-one correspondence relationship, and include a training set, a test set, and a validation set.

3. A method for training a radiology report generation model based on multi-modal contrast learning as claimed in claim 2, characterized in that, in step S200, learning sentence representations through a contrastive learning-based sentence-level training strategy includes the following steps: For the same sentence in a set of sentences, a series of different enhanced versions of the sentence representations are obtained by implementing different data enhancement methods as positive examples in the training set, test set, or validation set, and other sentences are used as negative examples; When training the sentence encoder, by maximizing the consistency between the sentence representations of different enhanced versions of the same sample, while keeping the sentence vectors of different samples as far apart as possible, a sentence encoder is established and a semantic sentence embedding is constructed.

4. A method for training a radiology report generation model based on multi-modal contrast learning as claimed in claim 3, characterized in that, In step S200, take a set of sentences x i represents the i-th sentence, and m represents the total number of sentences in the set Apply two different data augmentation methods f() and f′() to the sentence x i to generate two different versions of sentence embeddings e i and e′ i : e i = f(x i ) e′ i = f′(x i ) wherein, L is the length of the sentence embedding, and D is the hidden dimension of the sentence embedding; Then, the sentence embeddings e i and e' i are encoded to obtain sentence representations h i and h' i; For a small batch of N sentences, sentence x i the training objective l i is as follows: where τ is a temperature hyperparameter and sim() is the cosine similarity, then we have Final contrastive loss is the average of all N within-batch losses:

5. A method for training a radiology report generation model based on multi-modal contrast learning as claimed in claim 1, characterized in that, in step S200, learning image representations through a bidirectional contrastive learning objective between paired images and texts includes the following steps: Through paired input (X v ,X s ) learns an image encoder, where X v represents an image or a group of images, X s Representative Description X v In the sentence sequence of imaging information; for each input image X v And for each input sentence X s , which are encoded by an image encoder f v () and a sentence encoder f s () is converted into a fixed-dimensional vector h v and h s ; Then, the two modal representations h v 、h s By projecting the function g v () and g s () project from their encoder space to the same D-dimensional space for contrastive learning; For N input pairs (X v , X s ), the corresponding N representation pairs (v, s) are obtained, where: v = g v (f v (X v )) s = g s (f s (X s )) In the formula, Denoted by (v i , s i ), the i-th pair of representation pairs has training objectives including two loss functions: the image-text contrast loss and the text-image contrast loss Final training loss is a weighted combination of two losses: where λ is a scalar weight; By maximizing the consistency between image-text representation pairs, an image encoder is learned, and this image encoder maps an image to a vector of a fixed dimension.

6. A method for training a radiology report generation model based on multi-modal contrast learning as claimed in claim 1, characterized in that, in step S301, the sentence decoder for the Impression part adopts an LSTM-based method to generate a word at each time step according to the context vector, the previous hidden state, and the previously generated word to produce a caption; The initial hidden state and cell state of the LSTM are set to zero, and the hidden state h at time t t is modeled as: h t = LSTM(x t , h t-1 , m t-1 ) where x t is the input vector, and m t-1 is the memory cell vector at time t - 1; The visual feature vector is used as the initial input of the LSTM to predict the first word of the sentence; before inputting into the LSTM, a fully connected layer converts the visual feature vector output by the image encoder into the same dimension as the word embedding. The LSTM extracts visual features from the last convolutional layer and then generates the entire sentence word by word.

7. A method for training a radiology report generation model based on multi-modal contrast learning as claimed in claim 1, characterized in that, said step S400 includes: S401. Evaluate the generated report using BLEU and its variants; S402. Evaluate the generated report using METEOR; S403. Evaluate the generated report using ROUGE; S404. Evaluate the generated report using CIDEr.

8. A training system for a radiology report generation model based on multi-modal contrastive learning, characterized in that, it includes: A sample database for storing the collected radiology images and text data, where each text data contains a conclusive diagnostic statement Impression and a detailed descriptive paragraph Findings; A multi-modal contrastive learning module that uses a self-supervised representation learning method to learn an image encoder for visual features in the image data and a sentence encoder for extracting semantic features in the text data based on the sample database, where: The sentence encoder learns sentence representations through a contrastive learning-based sentence-level training strategy; The image encoder learns image representations through a bidirectional contrastive learning objective between paired image data and text data; A radiology report generation module that recursively generates the diagnostic statement Impression and the descriptive paragraph Findings of the radiology report by fusing the visual features extracted by the image encoder and the semantic features extracted by the sentence encoder; The radiology report generation module further includes an Impression generation module and a Findings generation module, where: The Impression generation module generates a single diagnostic statement Impression based on an encoder-decoder framework. The implementation of the Impression generation module includes the following steps: The image encoder extracts visual features from the input image and then sends them to the Impression part sentence decoder to generate the entire sentence word by word as the diagnostic statement Impression; The Findings generation module fuses the visual features of the image and the semantic features of the sentence to generate sentences cyclically and finally generates a long paragraph containing multiple structural statements. The implementation of the Findings generation module includes the following steps: In the Findings generation module, in order to make the generated sentences focus on describing different image regions, based on the attention framework, the visual features output by the image encoder and the semantic features of the previous sentence output by the sentence encoder are input into a fully connected layer and then sent to the SoftMax layer to obtain weighted visual features. The weighted visual features are sent to the Findings part sentence decoder, and the encoding of the sentence is obtained by the Findings part sentence decoder. The encoding of the sentence is used as the encoding of the previous sentence and input into the sentence encoder to obtain new weighted visual features. This process is repeated continuously until the Findings part sentence decoder generates an empty sentence, indicating that the generation of the Findings part is completed; Among them, the encoding of the previous sentence and the visual features of the image are combined through the weighted visual features to guide the generation of the next sentence. The attention network for calculating the weighted visual representation is defined as: V w = Attention(v, s) Among them, V w is the weighted visual representation to be obtained; Attention() is the attention function; v = {v 1 , v 2 , …, v k}, is the image feature learned by the image encoder, and each feature v i is a representation of dimension d v , corresponding to a part of the image; represents the encoding of the previous sentence, and d s is the dimension of the semantic feature; Given the visual features and the encoding of the previous sentence, an attention distribution is generated over K regions of the image by a single attention allocator, represented by a single-layer neural network and a SoftMax function: α = softmax(z) In the formula: is a vector with all elements set to 1; are the parameters of the attention network; is the weight of the feature in v in the attention pair (v, s); Based on the attention distribution, the weighted visual representation V w is obtained by the following: where α i is the i-th dimensional element in α; A report analysis and evaluation module that uses evaluation metrics to evaluate the generated radiology report; A result output module is used to combine the separately generated diagnostic statement Impression and the descriptive paragraph Findings into a complete radiology report for output, and at the same time, multiple evaluation indicators are used to evaluate the output result.

9. A training system for a radiology report generation model based on multi-modal contrastive learning as claimed in claim 8, wherein, in the radiology report generation module, the image encoder is modeled as a CNN; the sentence encoder is modeled as a model obtained by fine-tuning in a BERT language model, and the semantic representation of the sentence is generated through an average pooling layer.

Citation Information

Patent Citations

  • Training method of medical image report generation model and image report generation method

    CN112992308A

  • Semantic coding method of long-short-term memory network based on attention distraction

    CN113033189A