Closed-loop calibration method, device and equipment for improving generation quality of medical image report
Through the closed-loop calibration method, the interaction between the graphical text model and the literary model is used to calculate the loss and perform parameter calibration, which solves the problem of insufficient semantic consistency between the text and images of medical imaging reports and improves the accuracy of the report.
Patent Information
- Application Number
- CN202510354077.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
The text of medical image report generated by the prior art is insufficient in semantic consistency with the original image, resulting in low text accuracy.
The closed-loop calibration method is used to calculate text loss and image loss through the interaction between the graphical text model and the textual graphics model, and parameter calibration is performed based on these losses to improve the alignment of the image and text features.
Through the closed-loop calibration mechanism, the quality of medical image reports is gradually enhanced, the alignment between images and text features is improved, thereby improving the accuracy of the report.
Smart Images

Figure CN120220946A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a closed-loop calibration method, device and equipment for improving the quality of medical image report generation, and belongs to the fields of medical image processing and natural language processing. Background Art
[0002] Medical images are important tools for modern clinical diagnosis and can intuitively display the internal structure of patients. However, traditional medical image analysis relies on the manual interpretation of radiologists, which has problems such as low efficiency, strong subjectivity, and the risk of human error. In recent years, with the progress of artificial intelligence, especially deep learning technology, automated medical image analysis and diagnosis have become a research hotspot. However, the text generated by existing technologies has insufficient semantic consistency with the original image, resulting in low text accuracy. Summary of the Invention
[0003] In view of this, the present application provides a closed-loop calibration method, device and equipment for improving the quality of medical image report generation. The embodiments of the present application solve the technical problem of low text accuracy in related technologies.
[0004] The first aspect of the embodiments of the present application discloses a closed-loop calibration method for improving the quality of medical image report generation, and the method includes:
[0005] Generating a first text report based on a first image through an image-to-text model, where the first image is a real image;
[0006] Calculating the text loss between the first text report and a second text report, where the second text report is the real report of the first image;
[0007] Generating a second image based on the first text report through a text-to-image model;
[0008] Calculating the image loss between the first image and the second image;
[0009] Determining the image-to-text model loss based on the text loss and the image loss;
[0010] Determining the text-to-image model loss based on the image loss;
[0011] Performing parameter calibration based on the image-to-text model loss and the text-to-image model loss.
[0012] Further, the image-to-text model includes:
[0013] An image feature extractor for obtaining image features based on an image, where the image features represent lesion labels;
[0014] A text encoder for obtaining medical history information features based on medical history information;
[0015] A decoder for obtaining a medical report based on the image features, the medical history information features, and the lesion labels after matrix splicing.
[0016] Furthermore, the text-to-image model includes:
[0017] A noise addition module for adding noise to an image to a preset depth;
[0018] An image feature extractor for obtaining noise image features based on the noise-added image;
[0019] A text encoder for obtaining comprehensive text information features based on a medical report, medical history information, and the noise addition depth;
[0020] A noise predictor for obtaining predicted noise based on the fused noise image features and the comprehensive text information features;
[0021] A denoising module for obtaining a medical image based on the noise-added image and the predicted noise.
[0022] Furthermore, pre-training the image feature extractor includes:
[0023] Training a disease classification model on a dataset to learn the feature representations of various diseases;
[0024] Transferring the features learned by the disease classification model to the imaging report task and fine-tuning the disease classification model to obtain the pre-trained image feature extractor.
[0025] Furthermore, determining the text loss includes:
[0026] Calculating a report-level loss based on the real report and the generated report generated by the image-to-text model;
[0027] Dividing the real report at the sentence level to obtain the sentence set of the real report as the candidate set of matching sentences, and taking the sentences with lesion labels as the key matching sentences;
[0028] Dividing the generated report at the sentence level to obtain the sentence set of the generated report and taking the sentences with lesion labels as the key generated sentences;
[0029] Matching the sentence set of the generated report with the candidate set of matching sentences, calculating the cross-entropy loss based on the two sentences with the highest matching degree, multiplying the loss of the key sentences by a preset multiple, and adding up the cross-entropy losses of all sentences to obtain the text matching loss;
[0030] Adding the text matching loss and the key sentence loss to obtain the sentence-level text loss between the real report and the generated report.
[0031] The weighted sum of the report-level loss and the sentence-level text loss is obtained to get the final text loss.
[0032] Furthermore, determining the image loss includes:
[0033] Based on the trained object detection model, detecting the lesion features in the real image and the generated image;
[0034] Based on the lesion features, according to the pixel-level loss formula, the perceptual loss formula, and the key sentence loss, calculating the key feature loss of the image;
[0035] Using the perceptual loss formula to measure the overall similarity between the real image and the generated image, obtaining the perceptual loss of the overall image;
[0036] The weighted sum of the key feature loss and the perceptual loss is obtained to get the final image loss.
[0037] Furthermore, the pixel-level loss formula is as follows:
[0038]
[0039] where I real is the real image, I pred is the generated image, H is the height of the image, W is the width of the image, and C is the number of color channels;
[0040] The perceptual loss formula is as follows:
[0041]
[0042] where F i represents the feature extraction function of the i-th layer of feature extraction, and l is the number of layers of the feature extraction network;
[0043] The key feature loss is as follows:
[0044] L img-key = β1·L pixel-key + β2·L perceptual-key
[0045] where β1 and β2 are weight coefficients.
[0046] The second aspect of the embodiments of the present application discloses a closed-loop calibration device for improving the quality of medical image report generation. The device includes:
[0047] A first generation module, configured to generate a first text report based on a first image through a graph-to-text model, where the first image is a real image;
[0048] A first calculation module, configured to calculate the text loss between the first text report and the second text report, where the second text report is the ground truth report of the first image;
[0049] A second generation module, configured to generate a second image based on the first text report through a text-to-image model;
[0050] A second calculation module, configured to calculate the image loss between the first image and the second image;
[0051] A first determination module, configured to determine the image-to-text model loss based on the text loss and the image loss;
[0052] A second determination module, configured to determine the text-to-image model loss based on the image loss;
[0053] An adjustment module, configured to perform parameter adjustment based on the image-to-text model loss and the text-to-image model loss.
[0054] A third aspect of the embodiments of the present application discloses a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls a processor of the device where it is located to execute the closed-loop adjustment method of the above embodiments.
[0055] A fourth aspect of the embodiments of the present application discloses an electronic device, where the electronic device includes: one or more processors; a storage device, configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors execute the closed-loop adjustment method of the above embodiments.
[0056] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0057] An embodiment of the present application provides a closed-loop calibration method, device, and equipment for improving the quality of medical image report generation. The closed-loop calibration method includes: generating a first text report based on a first image through an image-to-text model, where the first image is a real image; calculating a text loss between the first text report and a second text report, where the second text report is the real report of the first image; generating a second image based on the first text report through a text-to-image model; calculating an image loss between the first image and the second image; determining an image-to-text model loss based on the text loss and the image loss; determining a text-to-image model loss based on the image loss; and performing parameter calibration based on the image-to-text model loss and the text-to-image model loss. This embodiment proposes a closed-loop calibration mechanism that aligns image and text features through a process from image to text (image-to-text) and then to image (text-to-image), thereby improving the quality of medical image reports. Specifically, based on the same medical image report dataset, an image-to-text model and a text-to-image model are respectively trained, and through the interactive fine-tuning of the two models, the alignment of image and text features is gradually enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0059] Figure 1 It is a schematic flowchart of a closed-loop calibration method provided by an embodiment of the present application.
[0060] Figure 2 It is a schematic flowchart of a method for automatically generating a medical image report provided by an embodiment of the present application.
[0061] Figure 3 It is a schematic diagram of the processing of a disease classification model provided by an embodiment of the present application.
[0062] Figure 4 It is a schematic framework diagram of an image-to-text model provided by an embodiment of the present application.
[0063] Figure 5 It is a schematic framework diagram of a text-to-image model provided by an embodiment of the present application.
[0064] Figure 6 It is a schematic framework diagram of an image-to-text + text-to-image model provided by an embodiment of the present application.
[0065] Figure 7 It is a schematic flowchart of text loss calculation provided by an embodiment of the present application.
[0066] Figure 8 A flowchart of an image loss calculation provided by an embodiment of the present application.
[0067] Figure 9 A framework diagram of a closed-loop calibration provided by an embodiment of the present application.
[0068] Figure 10 A flowchart of a closed-loop calibration mechanism provided by an embodiment of the present application.
[0069] Figure 11 A flowchart of a classification model training provided by an embodiment of the present application.
[0070] Figure 12 A flowchart of a text-to-image model training provided by an embodiment of the present application.
[0071] Figure 13 A flowchart of an image-to-text model training provided by an embodiment of the present application.
[0072] Figure 14 A schematic framework diagram of a closed-loop calibration device provided by an embodiment of the present application. Detailed implementation manners
[0073] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0074] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0075] Embodiment 1:
[0076] Figure 1 A schematic flowchart of a closed-loop calibration method provided by an embodiment of the present application, asFigure 1 As shown, the method may include the following steps:
[0077] S101 Generate a first text report based on a first image through an image-to-text model, where the first image is a real image.
[0078] S102 Calculate the text loss between the first text report and a second text report, where the second text report is the real report of the first image.
[0079] S103 Generate a second image based on the first text report through a text-to-image model.
[0080] S104 Calculate the image loss between the first image and the second image.
[0081] S105 Determine the image-to-text model loss based on the text loss and the image loss.
[0082] S106 Determine the text-to-image model loss based on the image loss.
[0083] S107 Perform parameter calibration based on the image-to-text model loss and the text-to-image model loss.
[0084] This embodiment proposes a closed-loop calibration mechanism. Through the process from image to text (image-to-text) and then to image (text-to-image), the image and text features are aligned, thereby improving the quality of medical image reports. Specifically, based on the same medical image report dataset, an image-to-text model and a text-to-image model are respectively trained, and through the interactive fine-tuning of the two models, the alignment of image and text features is gradually enhanced.
[0085] It should be noted that this embodiment takes the diagnostic report of chest X-ray images as an example. In practical applications, it can be extended to various medical images, such as CT, MRI, or ultrasound images, etc., and is applicable to the generation and analysis of diagnostic reports in different medical image fields. The real image is taken by a medical device, and the real report is obtained by one or more doctors diagnosing and writing.
[0086] To better illustrate the solution of this embodiment, as Figure 2 shown, the flow overview of the entire technical solution is as follows:
[0087] First, train an image feature extractor to obtain its pre-trained model; then train an image-to-text model and fine-tune the image feature extractor during the training process to obtain the pre-trained model of image-to-text; then train a text-to-image model to obtain its pre-trained model; then set loss functions for the image-to-text model and the text-to-image model respectively for subsequent fine-tuning; finally, adopt a closed-loop calibration mechanism to fine-tune these two pre-trained models to align the image semantic features and text semantic features, thereby improving the quality of the generated image reports.
[0088] Specifically, the closed-loop calibration mechanism is as follows Figure 9 As shown, first, an image report is generated through the image-to-text model. The generated report is compared with the original report to obtain text differences. Then, a medical image is generated through the text-to-image model. The generated image is compared with the original image to obtain image differences. Finally, based on the text differences and image differences, the parameters of the image-to-text model and the text-to-image model are updated.
[0089] Based on Figure 9 the closed-loop calibration mechanism shown, the two pre-trained models are fine-tuned. The specific process is as follows Figure 10 as shown
[0090] 1) After generating text using the image-to-text model, calculate the loss between the two generated texts and the real text to obtain the text loss L text .
[0091] 2) Input the text generated by the image-to-text model into the text-to-image model to obtain the generated image generated by the text-to-image model.
[0092] 3) Calculate the loss between the generated image and the real image to obtain the image loss L image .
[0093] 4) Use the weighted sum of the text loss L text and the image loss L image as the image-to-text model loss L img_to_text , and use the image loss L image as the text-to-image model loss L text-to-img .
[0094] 5) Finally, use a strategy to update and iterate the parameters of the image-to-text model and the text-to-image model. For example, update the parameters of the image-to-text model and the text-to-image model simultaneously during each training, or first update the parameters of the image-to-text model until convergence, then update the parameters of the text-to-image model until convergence, and then update the parameters of the image-to-text model until convergence, and so on.
[0095] In this embodiment, the loss function L img_to_text of the image-to-text model uses the weighted sum of the text loss L text and the image loss L image as follows
[0096] L img_to_text = λ1·L text + λ2·L image
[0097] where λ1 and λ2 are weight coefficients used to balance the text loss and the image loss.
[0098] In the closed-loop mechanism, since the text input to the text-to-image model is not real text but the text generated by the image-to-text model, the image loss used to update the text-to-image model should calculate the loss L' with the lesion features mentioned in the generated text as the key features. image The loss function L of the text-to-image model text-to-img adopts the image loss L' image , as shown in the following formula:
[0099] L text-to-img = L' image
[0100] Furthermore, as Figure 4 shown, the image-to-text model includes:
[0101] An image feature extractor for obtaining image features based on the image, where the image features represent lesion labels.
[0102] A text encoder for obtaining medical history information features based on the medical history information.
[0103] A decoder for obtaining a medical report based on the image features, the medical history information features, and the lesion labels after matrix concatenation.
[0104] The processing logic of the image-to-text model includes:
[0105] a. Using the image feature extractor to encode the medical image to obtain a feature representation of the image.
[0106] b. Using the text encoder to encode the patient's medical history information to obtain a feature representation of the medical history information.
[0107] c. Predicting the lesion labels through the image feature representation to obtain all the lesion labels.
[0108] d. Fusing the image feature representation and the text feature representation to obtain a fused feature representation of the image and the text.
[0109] e. Inputting the fused feature representation and the lesion labels into the decoder model to obtain the medical report generated by the decoder model.
[0110] Furthermore, as Figure 5 shown, the text-to-image model includes:
[0111] A noise addition module for adding noise to the image to a preset depth.
[0112] An image feature extractor for obtaining noise image features based on the noise-added image.
[0113] A text encoder for obtaining a comprehensive text information feature based on the medical report, the medical history information, and the noise addition depth.
[0114] A noise predictor, configured to obtain predicted noise based on the fused noise image features and the comprehensive text information features.
[0115] A denoising module, configured to obtain a medical image based on the noisy image and the predicted noise.
[0116] The processing logic of the text-to-image model includes:
[0117] A. Add noise to the image N times until the image becomes a pure noise image.
[0118] B. Use an image feature extractor to encode the noise image with a noise addition depth of step to obtain a feature representation of the noise image.
[0119] C. Use a text encoder to encode information such as medical reports, medical history information, and noise addition depth to obtain a comprehensive feature representation of a series of text information.
[0120] D. Fuse the image feature representation and the text feature representation to obtain a fused feature representation of the image and the text.
[0121] E. Input the fused feature representation into the noise predictor to obtain predicted noise.
[0122] F. Subtract the predicted noise from the noise image with a noise addition depth of step to obtain the denoised image, that is, the medical image with a noise addition depth of (step - 1).
[0123] G. Repeat the above steps until step is 0 to obtain the finally generated medical image.
[0124] In some embodiments, the image-to-text model adopts an encoder-decoder architecture (Encoder-Decoder), where the image encoder adopts a Swin-Transformer model, the text encoder adopts a bidirectional encoding model based on transformers (bidirectional encoder representations from transformers, BERT), and the decoder adopts a DistilGPT-2 model; the text-to-image model adopts a diffusion model (Diffusion Model); the object detection model adopts YOLOv3.
[0125] In some embodiments, the processing logic of the image-to-text + text-to-image model is as Figure 6Shown as follows: In the first step, the lesion label obtained from the image-to-text model, the semantic features of the medical history information, the text features of the medical image report generated by the decoder, and the noisy image with a noise addition depth of step are input into the noise predictor to obtain the predicted noise. Then, the noisy image is subtracted by the predicted noise to obtain the denoised image, that is, the medical image with a noise addition depth of (step - 1). In the second step, continue to denoise the noisy image according to the denoising steps of the first step until the noise addition depth is 0, and the generated medical image is obtained.
[0126] Furthermore, pre-train the image feature extractor, including: training a disease classification model on a dataset to learn the feature representations of various diseases; migrating the features learned by the disease classification model to the medical image report task, and fine-tuning the disease classification model to obtain the pre-trained image feature extractor.
[0127] Specifically, the pre-training strategy: train a disease classification model on a dataset to learn the feature representations of various diseases, strengthen the recognition and discrimination of the features of various diseased areas, and then migrate these features to the medical image report task to fine-tune the image feature extractor (Fine-tuning). As Figure 3 shown, use the image feature extractor to encode the medical image to obtain the feature representation of the image; predict the lesion label through the image feature representation to obtain all the lesion labels.
[0128] In some embodiments, during the training process, the binary cross-entropy loss (Binary Cross-Entropy, BCE Loss) is used as the loss function. The output of each category is regarded as an independent binary classification problem, and the loss is calculated as follows:
[0129]
[0130] where N is the number of samples, C is the number of categories, y ij is the true label (0 or 1) of the i-th sample in the j-th category, is the predicted probability of the i-th sample in the j-th category.
[0131] Further, determining the text loss includes: calculating a report-level loss based on the ground truth report and the generated report generated by the image-to-text model; dividing the ground truth report at the sentence level to obtain a sentence set of the ground truth report, which is used as a candidate set of matching sentences, and using the sentences with lesion labels as key matching sentences; dividing the generated report at the sentence level to obtain a sentence set of the generated report, and using the sentences with lesion labels as key generated sentences; matching the sentence set of the generated report with the candidate set of matching sentences, calculating the cross-entropy loss based on the two sentences with the highest matching degree, multiplying the loss of the key sentences by a preset multiple, adding up the cross-entropy losses of all sentences to obtain the text matching loss; adding the text matching loss and the key sentence loss to obtain the sentence-level text loss between the ground truth report and the generated report; and performing a weighted sum of the report-level loss and the sentence-level text loss to obtain the final text loss.
[0132] It should be noted that the cross-entropy loss is one of the most commonly used loss functions, especially in natural language generation tasks. It measures the gap between the predicted distribution and the true word distribution at each time step, and its formula is as follows:
[0133]
[0134] where L report is the report loss, T is the report length, and p(y t |y < t, x) represents the probability of predicting the next word y t given the previous word y < t and the input sequence x.
[0135] However, the similarity measurement of medical image reports is different from the semantic similarity measurement of general long texts because most sentences from the same report are highly independent, that is, some sentences are context-independent. Therefore, dividing the highly independent sentences in the report and comparing the similarity between sentences can more accurately reflect the similarity between the ground truth report and the generated report. Secondly, regarding the accuracy of the image report, it mainly examines whether all lesion labels are included and whether the description of the lesion characteristics is accurate, rather than strictly requiring the generated text to be exactly the same as the ground truth text. Therefore, using the sentences with lesion labels as key sentences mainly examines the accuracy of the key sentences. The calculation process is as Figure 7 shown.
[0136] S11 Calculate the report-level loss L report .
[0137] S12 Divide the ground truth report at the sentence level to obtain a sentence set of the ground truth report, which is used as a candidate set of matching sentences, and use the sentences with lesion labels as key matching sentences.
[0138] S13 generates a report by dividing it at the sentence level, obtains a sentence set of the generated report, and uses the sentences with lesion labels as key generated sentences.
[0139] S14 matches each sentence in the generated sentence set with the candidate set of matching sentences, calculates the cross-entropy loss between the candidate sentence with the highest matching degree and this sentence, then multiplies the loss of the key generated sentence by an appropriate multiple, and then sums up the cross-entropy losses of all sentences to obtain the text matching loss L. match 。
[0140] S15 determines whether there are key sentences that have not been matched. If so, a relatively large loss is given, and a loss corresponding to the number of unmatched important sentences is given according to the corresponding multiple, and it is named the key sentence loss L. key 。
[0141] S16 adds the text matching loss L match and the key sentence loss L text-key to obtain the sentence-level text loss L between the real report and the generated report. sentence For example: L sentence = L match + a·L text-key 。
[0142] S17 performs a weighted sum of the report-level loss L report and the sentence-level loss L sentence to obtain the final text loss L text For example: L text = α1·L report + α2·L sentence 。
[0143] Among them, a is a positive number greater than 1, aiming to increase the contribution of the key sentence loss to the total loss and make the model pay more attention to the key sentences; α1, α2 are weight coefficients used to balance the text losses at the report level and the sentence level.
[0144] Further, determining the image loss includes: based on the trained object detection model, detecting the lesion features in the real image and the generated image; based on the lesion features, calculating the key feature loss of the image according to the pixel-level loss formula, the perceptual loss formula, and the key sentence loss; using the perceptual loss formula to measure the overall similarity between the real image and the generated image to obtain the perceptual loss of the overall image; performing a weighted sum of the key feature loss and the perceptual loss to obtain the final image loss.
[0145] It should be noted that Pixel-wise Loss and Perceptual Loss are two widely used image generation loss functions, which provide different advantages in different aspects.
[0146] The pixel-level loss is simple and intuitive to calculate and easy to implement, which can ensure that the generated image is as close as possible to the real image at the pixel level. However, it has the following disadvantages:
[0147] 1. It tends to produce smooth results, which may lead to the generated image lacking details or sharpness;
[0148] 2. It is not conducive to capturing high-level features (such as textures and structures) because the pixel-level loss focuses on local rather than global consistency;
[0149] 3. It may cause a "blurring" effect, especially in high-resolution image generation tasks.
[0150] The perceptual loss pays more attention to the perceptual quality of the image, can better maintain the structure and texture of the generated image, can avoid the over-smoothing problem brought by the pixel-level loss, and is conducive to generating clearer and more realistic images, but it requires additional feature extraction layers, increasing the computational overhead.
[0151] Generally speaking, the pixel-level loss measures the difference between two images at the pixel level, while the perceptual loss is a method to measure the similarity between the generated image and the real image. This method can capture the structural information of the image, rather than just the difference at the pixel level. Therefore, the pixel-level loss and the perceptual loss are combined to balance the local details and the overall perceptual quality of the image.
[0152] Furthermore, the formula of the pixel-level loss is as follows:
[0153]
[0154] where, I real is the real image, I pred is the generated image, H is the height of the image, W is the width of the image, and C is the number of color channels;
[0155] The formula of the perceptual loss is as follows:
[0156]
[0157] where, F i represents the feature extraction function of the i-th layer of feature extraction, and l is the number of layers of the feature extraction network.
[0158] Similar to the similarity metric for medical image reports, the quality of the generated medical image mainly depends on whether the lesion area is similar to that of the real image, and it is not necessary to ensure a high similarity everywhere with the real image. Therefore, the lesion features mentioned in the real text are used as key features, and the similarity between the lesion areas of the generated image and the real image is mainly examined. The calculation process is as follows Figure 8 shown
[0159] S21 uses an object detection model to detect the lesion features in the real image and the generated image.
[0160] S22 calculates the loss between the lesion areas of the two images according to the pixel-level loss and the perceptual loss, and obtains the key feature loss L of the image img-key , such as: L img-key =β1·L pixel-key +β2·L perceptual-key .
[0161] S23 uses the perceptual loss to measure the similarity between the entire images, and obtains the perceptual loss L of the overall image perceptual .
[0162] S24 finally combines the two as the final image loss L image , such as: L image =b·L img-key +L perceptual .
[0163] Among them, β1 and β2 are weight coefficients used to balance the pixel-level loss and the perceptual loss; b is a positive number greater than 1, aiming to increase the contribution of the key feature loss to the total loss and make the model pay more attention to the key features of the image.
[0164] Figure 3 The training flow chart of the classification model shown in Figure 11 is shown in Figure 4 The training flow chart of the image-to-text model shown in Figure 12 is shown in Figure 5 The training flow chart of the text-to-image model shown in Figure 13 is shown in
[0165] Example 2:
[0166] Figure 14 The framework schematic diagram of a closed-loop calibration device provided by an embodiment of the present application is shown in Figure 13 shown, and the device may include the following modules:
[0167] The first generation module 1401 is used to generate a first text report based on the first image through the image-to-text model, and the first image is a real image.
[0168] The first calculation module 1402 is configured to calculate the text loss between the first text report and the second text report, where the second text report is the ground truth report of the first image.
[0169] The second generation module 1403 is configured to generate a second image based on the first text report through a text-to-image model.
[0170] The second calculation module 1404 is configured to calculate the image loss between the first image and the second image.
[0171] The first determination module 1405 is configured to determine the image-to-text model loss based on the text loss and the image loss.
[0172] The second determination module 1406 is configured to determine the text-to-image model loss based on the image loss.
[0173] The calibration module 1407 is configured to perform parameter calibration based on the image-to-text model loss and the text-to-image model loss.
[0174] Furthermore, the image-to-text model includes:
[0175] An image feature extractor, configured to obtain image features based on an image, where the image features represent lesion labels;
[0176] A text encoder, configured to obtain medical history information features based on medical history information;
[0177] A decoder, configured to obtain a medical report based on the image features, the medical history information features, and the lesion labels after matrix concatenation.
[0178] Furthermore, the text-to-image model includes:
[0179] A noise addition module, configured to add noise to an image to a preset depth;
[0180] An image feature extractor, configured to obtain noise image features based on the noise-added image;
[0181] A text encoder, configured to obtain comprehensive text information features based on a medical report, medical history information, and the noise addition depth;
[0182] A noise predictor, configured to obtain predicted noise based on the fused noise image features and the comprehensive text information features;
[0183] A denoising module, configured to obtain a medical image based on the noise-added image and the predicted noise.
[0184] Furthermore, pre-training the image feature extractor includes:
[0185] Train a disease classification model on a dataset to learn the feature representations of various diseases;
[0186] Transfer the features learned by the disease classification model to the image report task, and fine-tune the disease classification model to obtain the pre-trained image feature extractor.
[0187] Furthermore, determining the text loss includes:
[0188] Calculate the report-level loss based on the real report and the generated report generated by the text-to-image model;
[0189] Divide the real report at the sentence level to obtain the sentence set of the real report, which is used as the candidate set of matching sentences, and use the sentences with lesion labels as the key matching sentences;
[0190] Divide the generated report at the sentence level to obtain the sentence set of the generated report, and use the sentences with lesion labels as the key generated sentences;
[0191] Match the sentence set of the generated report with the candidate set of matching sentences, calculate the cross-entropy loss based on the two sentences with the highest matching degree, multiply the loss of the key sentences by a preset multiple, and add up the cross-entropy losses of all sentences to obtain the text matching loss;
[0192] Add the text matching loss and the key sentence loss to obtain the sentence-level text loss between the real report and the generated report;
[0193] Perform a weighted sum of the report-level loss and the sentence-level text loss to obtain the final text loss.
[0194] Furthermore, determining the image loss includes:
[0195] Based on the trained object detection model, detect the lesion features in the real image and the generated image;
[0196] Based on the lesion features, calculate the key feature loss of the image according to the pixel-level loss formula, the perceptual loss formula, and the key sentence loss;
[0197] Use the perceptual loss formula to measure the overall similarity between the real image and the generated image to obtain the perceptual loss of the overall image;
[0198] Perform a weighted sum of the key feature loss and the perceptual loss to obtain the final image loss.
[0199] Furthermore, the pixel-level loss formula is as follows:
[0200]
[0201] Among them, I real is the real image, I pred is the generated image, H is the height of the image, W is the width of the image, and C is the number of color channels;
[0202] The perceptual loss formula is as follows:
[0203]
[0204] Among them, F i represents the feature extraction function of the i-th layer of feature extraction, and l is the number of layers of the feature extraction network;
[0205] The key feature loss is as follows:
[0206] L img-key = β1·L pixel-key + β2·L perceptual-key
[0207] Among them, β1 and β2 are weight coefficients.
[0208] Example 3:
[0209] An embodiment of the present application further provides an electronic device, including: a memory storing an executable program; a processor for running the program, wherein when the program runs, it executes the methods in the various embodiments of the present invention.
[0210] The above-mentioned memory may refer to a device inside a computer for storing data and programs, which may include a memory, a hard disk, etc. Among them, the memory can be used for temporarily storing running programs and data, and the hard disk can be used for long-term storing programs and data. The memory can be used to enable a computer to read and write data and execute programs; the above-mentioned processor can be responsible for executing instructions in a computer program and performing data processing, and can be responsible for controlling and executing various operations, including arithmetic operations, logical operations, data transmission, etc.
[0211] Example 4:
[0212] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present invention.
[0213] The above-mentioned computer storage medium may refer to a medium in a computer memory for storing a certain discontinuous physical quantity. The computer storage medium mainly includes semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser discs, etc.; the stored program included in the computer-readable storage medium can be a set of instructions that a computer can recognize and execute, running on an electronic computer, and is an information tool to meet people's certain needs.
[0214] Example 5:
[0215] An embodiment of the present application also provides a computer program product, including a computer program, which implements the methods in various embodiments of the present invention when executed by a processor.
[0216] The above computer program product may refer to a software program that has been written, tested, and released, and can run on a computer or other devices. The computer program product may include application programs, operating systems, tool software, etc., and is used to implement specific functions or solve specific problems.
[0217] Example 6:
[0218] An embodiment of the present application also provides a computer program product, including a non-volatile computer-readable storage medium, which is used to store a computer program, and the computer program implements the methods in various embodiments of the present invention when executed by a processor.
[0219] The above non-volatile computer-readable storage medium may refer to a medium for storing data. The non-volatile computer-readable storage medium can keep data from being lost when powered off, and can be used to store data for long-term preservation, such as operating systems, application programs, and user files. The non-volatile storage medium may include hard disk drives, solid-state drives, optical discs, and flash storage devices, etc.
[0220] Example 7:
[0221] An embodiment of the present application also provides a computer program, which implements the methods in various embodiments of the present invention when executed by a processor.
[0222] The above computer program may refer to a set of instructions, which are used to tell a computer to perform specific tasks or operations. The computer program can be written by a programmer using a specific programming language, and may include algorithms, data structures, logic, and control flows, etc. The computer program can be used for various purposes, including application software, operating systems, etc.
[0223] In the above embodiments of the present invention, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0224] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0225] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0226] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0227] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs and other various media that can store program codes.
[0228] To sum up, the advantages of the technical solution of the present invention are:
[0229] 1. Since most sentences from the same report are highly independent, different from general long texts that need to consider the context and there is no consideration of order, that is to say, the image report does not need to be calculated at the report level, and it is more accurate to calculate its similarity at the sentence level. The loss function of the report adopts the form of separately calculating and then summing independent sentences, which can more accurately reflect the quality of the generated text.
[0230] 2. Regarding the accuracy of the imaging report, it mainly examines whether all lesion labels are included and whether the description of lesion characteristics is accurate, rather than strictly requiring the generated text to be exactly the same as the real text. Therefore, taking the sentences with lesion labels as key sentences and mainly examining the accuracy of key sentences can better reflect the quality of the generated report.
[0231] 3. Similar to the judgment of the accuracy of medical imaging reports, the generated medical images do not need to be highly similar to the real images everywhere. The judgment of their quality mainly depends on whether the lesion areas are similar to those of the real images. Therefore, taking the lesion characteristics mentioned in the text as key features and mainly examining the similarity between the generated images and the lesion areas of the real images can more accurately evaluate the quality of image generation.
[0232] 4. Generate an imaging report through the image-to-text model, compare the original report with this report to obtain text differences; then generate medical images through the text-to-image model, compare the original image with this image to obtain image differences; finally, update the parameters of the two models through these two differences in text and images. Fine-tuning the pre-trained image-to-text model and text-to-image model through this closed-loop mechanism helps align the text semantic features with the image semantic features, thereby improving the quality of automatically generated medical imaging reports.
[0233] 5. The overall framework of the present invention has high versatility and can be directly applied to other application scenarios. For example, in image description, this framework can capture the key features in the image, align the semantic features of the image and the text, and describe the main events in the image more accurately; in art creation, this framework can also capture the key sentences in the text description, and this technology can be used to quickly generate creative concept maps or sketches that meet the user's ideas.
[0234] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A closed-loop calibration method for improving the quality of medical image report generation, characterized in that: include: Generate a first text report based on a first image by using an image-to-text model, where the first image is a real image; calculating a text loss between the first text report and a second text report, the second text report being a true report of the first image; generating a second image based on the first text report by using a text graph model; calculating an image loss between the first image and the second image; Determine a graph-to-text model loss based on the text loss and the image loss; Determining a Vincent graph model loss based on the image loss; Parameter calibration is performed based on the graph-to-text model loss and the text-to-graph model loss.
2. The closed-loop calibration method according to claim 1, characterized in that: The graph-to-text model includes: An image feature extractor, used for obtaining image features based on the image, wherein the image features represent lesion labels; A text encoder, used for obtaining medical history information features based on the medical history information; A decoder is used to obtain a medical report based on the image features after matrix splicing, the medical history information features, and the lesion label.
3. The closed-loop calibration method according to claim 1, characterized in that: The Wensheng graph model includes: A noise adding module, used to add noise to the image to a preset depth; An image feature extractor, used for obtaining noise image features based on the image after adding noise; A text encoder, used to obtain comprehensive features of text information based on medical reports, medical history information, and noise depth; A noise predictor, used for obtaining predicted noise based on the fused noise image features and the text information comprehensive features; A denoising module is used to obtain a medical image based on the noisy image and the predicted noise.
4. The closed-loop calibration method according to claim 2 or 3, characterized in that: Pre-training the image feature extractor, comprising: Train the disease classification model on the data set and learn the characteristic representations of various diseases; The features learned by the disease classification model are transferred to the image reporting task, and the disease classification model is fine-tuned to obtain the pre-trained image feature extractor.
5. The closed-loop calibration method according to claim 1, characterized in that: Determining the text loss includes: Calculate report-level loss based on the real report and the generated report generated by the graph-generated model; Dividing the real report based on sentence level to obtain a sentence set of the real report as a candidate set of matching sentences, and taking sentences with lesion labels as key matching sentences; Dividing the generated report based on sentence levels to obtain a sentence set of the generated report, and taking sentences with lesion labels as key generated sentences; Matching the sentence set of the generated report with the candidate set of matching sentences, calculating the cross entropy loss based on the two sentences with the highest matching degree, multiplying the loss of the key sentence by a preset multiple, and adding the cross entropy losses of all sentences to obtain the text matching loss; Adding the text matching loss and the key sentence loss to obtain the sentence-level text loss of the real report and the generated report; The report-level loss and the sentence-level text loss are weighted and summed to obtain the final text loss.
6. The closed-loop calibration method according to claim 1, characterized in that: Determining the image loss includes: Detect lesion features in real images and generated images based on the trained target detection model; Based on the lesion features, the key feature loss of the image is calculated according to the pixel-level loss formula, the perceptual loss formula and the key sentence loss; Using a perceptual loss formula to measure the overall similarity between the real image and the generated image, and obtaining the perceptual loss of the overall image; The key feature loss and the perceptual loss are weighted and summed to obtain the final image loss.
7. The closed-loop calibration method according to claim 6, characterized in that: The pixel-level loss formula is as follows: Among them, I real is a real image, I pred is the generated image, H is the height of the image, W is the width of the image, and C is the number of color channels; The perceptual loss formula is as follows: Among them, F i represents the feature extraction function of the i-th layer of feature extraction, and l is the number of layers of the feature extraction network; The key feature loss is as follows: L img-key =β1·L pixel-key +β2·L perceptual-key Among them, β1, β2 are weight coefficients.
8. A closed-loop calibration device for improving the quality of medical image report generation, characterized in that: include: A first generating module, configured to generate a first text report based on a first image by using an image-to-text model, wherein the first image is a real image; A first calculation module, configured to calculate a text loss between the first text report and a second text report, wherein the second text report is a true report of the first image; A second generating module, configured to generate a second image based on the first text report by using a text graph model; A second calculation module, used for calculating the image loss between the first image and the second image; A first determination module, configured to determine a loss of an image-to-text model based on the text loss and the image loss; A second determination module, configured to determine a Vincent graph model loss based on the image loss; A calibration module is used to perform parameter calibration based on the graph-to-text model loss and the text-to-graph model loss.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the closed-loop calibration method according to any one of claims 1 to 7 is executed in a processor of a device where the program is controlled.
10. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors execute the closed-loop calibration method described in any one of claims 1 to 7.