Model training method, medical image report generation method and device, and storage medium
Patent Information
- Application Number
- CN202610992217.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-29
Smart Images

Figure CN122842835A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a model training method, a medical image report generation method and apparatus, and a storage medium thereof. Background Technology
[0002] Radiology Report Generation (RRG) is an interdisciplinary research field that uses artificial intelligence (AI) to automatically generate diagnostic reports from X-ray images. Its core objective is to alleviate the workload of radiologists and improve the consistency of diagnostic efficiency and report quality by leveraging AI technology.
[0003] To improve the transparency of model reasoning, Chain-of-Thought (CoT) was integrated into the Multimodal Large Language Model (MLLM) framework. CoT breaks down complex problems into interpretable intermediate steps, enabling the model to follow a clearer logical path when solving problems. Summary of the Invention
[0004] The inventors noted that in related technologies, CoT can generate clear logical paths through chain reasoning, but this chain reasoning lacks a clear connection with visual evidence, thus leading to model illusion, i.e., MLLM outputs diagnostic reports that seem reasonable but lack clinical basis.
[0005] Accordingly, this disclosure provides a model training method and apparatus, as well as a medical image report generation method and apparatus. This disclosure effectively suppresses model hallucinations and improves the accuracy of the generated diagnostic reports by associating the generation of diagnostic reports with visual evidence of the corresponding anatomical regions.
[0006] In a first aspect of this disclosure, a model training method is provided, executed by a model training device, comprising: generating an initial sample set, wherein each sample in the initial sample set includes an X-ray image, an X-ray report corresponding to the X-ray image, and bounding box information of each anatomical region in the X-ray image; processing the initial sample set using a first large language model to generate prompt words based on the thought chain CoT, obtaining CoT report samples, wherein the CoT report samples include the X-ray image, the X-ray report, and an inference chain associated with the bounding boxes of each anatomical region; supervising fine-tuning the multimodal large language model using the CoT report samples to obtain an intermediate model, wherein the intermediate model has the ability to locate anatomical structures in an input X-ray image and generate an analysis report based on the location results; and training the intermediate model using reinforcement learning based on a target reward value to obtain a MedCoT-XR model for generating X-ray reports based on multimodal thought chains.
[0007] In some embodiments, training the intermediate model using reinforcement learning includes: determining a visual localization-based inference reward value based on the bounding box of the target anatomical region predicted by the intermediate model in each inference step and the actual bounding box of the target anatomical region; determining a medical fact consistency reward value based on the generated report output by the intermediate model for each sample and the real report corresponding to each sample; determining a target reward value based on the visual localization-based inference reward value and the medical fact consistency reward value; and adjusting the parameters of the intermediate model based on the target reward value.
[0008] In some embodiments, determining the target reward value includes: normalizing the visual positioning-based inference reward value to obtain a first inference reward value; normalizing the medical fact consistency reward value to obtain a second inference reward value; and calculating a weighted sum of the first inference reward value and the second inference reward value to obtain the target reward value.
[0009] In some embodiments, determining the inference reward value based on visual localization includes: calculating the intersection-union ratio (IUR) of the bounding box of the target anatomical region predicted by the intermediate model in each inference step and the actual bounding box of the target anatomical region to obtain multiple IURs; and calculating the average of the multiple IURs to obtain the inference reward value based on visual localization.
[0010] In some embodiments, determining the medical fact consistency reward value includes: using the RadGraph F1 scoring function, extracting clinical entities and their relationship graphs from the generated report output by the intermediate model for each sample and the real report corresponding to each sample, and calculating the entity-relationship F1 score between the generated report and the real report as the medical fact consistency reward value.
[0011] In some embodiments, the inference chain includes multiple inference steps, wherein each inference step includes inference information and bounding box information of the anatomical region corresponding to the inference information.
[0012] In some embodiments, the CoT-generated prompt includes command information for indicating that each inference step is bound to the bounding box information of the corresponding anatomical region.
[0013] In some embodiments, the CoT-generated prompts may further include the execution order of the plurality of inference steps.
[0014] In some embodiments, generating the initial sample set includes: generating a first sample using CT report samples and CT volume data samples corresponding to the CT report samples, wherein the first sample includes a synthetic X-ray report, a synthetic X-ray image, and bounding box information of each anatomical region in the synthetic X-ray image; training a bounding box detection model using the synthetic X-ray image and the bounding box information of each anatomical region in the synthetic X-ray image; processing the X-ray image samples using the bounding box detection model to obtain the bounding box information of each anatomical region in the X-ray image samples; generating a second sample using the X-ray image samples, the X-ray report samples corresponding to the X-ray image samples, and the bounding box information of each anatomical region in the X-ray image samples; and obtaining the initial sample set based on the first sample and the second sample.
[0015] In some embodiments, generating the first sample includes: processing the CT report sample based on the report rewriting prompts using a second language model to obtain the synthetic X-ray report; performing planar projection on the CT volume data sample to obtain the synthetic X-ray image; using the anatomical region segmentation and annotation information in the CT volume data sample to obtain the bounding box information of each anatomical region in the synthetic X-ray image; and generating the first sample based on the synthetic X-ray report, the synthetic X-ray image, and the bounding box information of each anatomical region in the synthetic X-ray image.
[0016] In some embodiments, the report rewriting prompt includes command information for deleting X-ray-invisible and inferable content from the synthetic X-ray report.
[0017] In a second aspect of this disclosure, a model training apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute instructions stored in the memory to implement the model training method as described in any of the above embodiments.
[0018] In a third aspect of this disclosure, a method for generating a medical image report is provided, comprising: acquiring an X-ray image to be tested; processing the X-ray image to be tested using a MedCoT-XR model based on prompts generated by the X-ray report to obtain an X-ray report, wherein the MedCoT-XR model is trained using the model training method described in any of the above embodiments; wherein processing the X-ray image to be tested using the MedCoT-XR model based on prompts generated by the X-ray report includes: sequentially inferring each anatomical region in the X-ray image to be tested according to the anatomical region examination order indicated by the prompts generated by the X-ray report; locating each anatomical region during the inference process to obtain bounding box information of each anatomical region; analyzing the region determined by the bounding box information of each anatomical region to output the analysis result of each anatomical region; and generating the X-ray report based on the analysis results of all anatomical regions in the X-ray image to be tested.
[0019] In a fourth aspect of this disclosure, a medical image report generation apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute instructions stored in the memory to implement the medical image report generation method as described in any of the above embodiments.
[0020] In a fifth aspect of this disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method as described in any of the above embodiments.
[0021] In a sixth aspect of this disclosure, a computer program product is provided, including computer instructions, wherein the computer instructions, when executed by a processor, implement the method as described in any of the above embodiments.
[0022] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a schematic flowchart of a model training method according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram illustrating the process of generating an initial sample set according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the structure of a model training apparatus according to an embodiment of the present disclosure; Figure 4 This is a flowchart illustrating a medical image report generation method according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of a medical image report generation device according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of an X-ray image to be processed according to an embodiment of the present disclosure; Figure 7 This is a schematic diagram of an X-ray image of a marked rib region according to an embodiment of the present disclosure; Figure 8 This is a schematic diagram of an X-ray image of a marked cardiac region according to an embodiment of the present disclosure. Detailed Implementation
[0025] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0026] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0027] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0028] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0029] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0030] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, no further discussion is required in subsequent drawings.
[0031] Figure 1 is a schematic flowchart of a model training method according to an embodiment of the present disclosure. In some embodiments, the following model training method is executed by a model training apparatus, and includes steps 11 to 14.
[0032] In step 11, an initial sample set is generated, where each sample in the initial sample set includes an X-ray image, an X-ray report corresponding to the X-ray image, and bounding box (bbox for short) information of each anatomical region in the X-ray image.
[0033] That is, each sample is a triplet including <X-ray image, X-ray report, box>.
[0034] In some embodiments, the step of generating the initial sample set includes steps S101-S105.
[0035] S101. Using CT (Computed Tomography) report samples and CT volume data samples corresponding to the CT report samples to generate a first sample, wherein the first sample includes a generated X-ray report, a generated X-ray image, and bounding box information of each anatomical region in the generated X-ray image.
[0036] In some embodiments, the step of generating the first sample includes steps S201-S204.
[0037] S201. Using an LLM (Large Language Model) to process CT report samples according to report rewriting prompts, to obtain a Generated Xray Report.
[0038] In some embodiments, the report rewriting prompt includes command information for deleting X-ray invisible content and un-inferable content in the generated X-ray report.
[0039] For example, when generating a synthetic X-ray report, detailed contents such as bronchial subsegments and bronchial vessels in the CT report sample are deleted to improve the accuracy of the X-ray report output by the model.
[0040] S202. Performing planar projection on the CT volume data sample to obtain a Generated Xray.
[0041] For example, by performing 2D planar projection such as DRR (Digitally Reconstructed Radiograph) projection on CT volume data samples, a synthesized X-ray image is obtained.
[0042] S203, using the anatomical region segmentation annotation information in the CT volume data sample to obtain bounding box information of each anatomical region in the synthesized X-ray image.
[0043] It should be noted here that since the CT volume data itself has 3D anatomical region segmentation annotation information, when performing projection, the projection mask is used to map the 3D segmentation annotation information of each anatomical region to the 2D planar projection, thereby obtaining bounding box information of each anatomical region in the synthesized X-ray image.
[0044] S204, generating a first sample according to the synthesized X-ray report, the synthesized X-ray image, and the bounding box information of each anatomical region in the synthesized X-ray image.
[0045] That is, the first sample is a triplet comprising <synthesized X-ray image, synthesized X-ray report, bbox>.
[0046] S102, training a bounding box detection model using the synthesized X-ray images and the bounding box information of each anatomical region in the synthesized X-ray images.
[0047] It should be noted that the trained bounding box detection model has the capability to detect the bounding boxes of each anatomical region in X-ray images.
[0048] For example, the bounding box detection model may comprise a YOLO (You Only Look Once) model or other models suitable for bounding box detection.
[0049] S103, processing X-ray image samples with the bounding box detection model to obtain bounding box information of each anatomical region in the X-ray image samples.
[0050] S104, generating a second sample using the X-ray image samples, the X-ray report samples corresponding to the X-ray image samples, and the bounding box information of each anatomical region in the X-ray image samples.
[0051] That is, the second sample is a triplet comprising <X-ray image sample, X-ray report sample, bbox>.
[0052] S105, obtaining an initial sample set according to the first samples and the second samples.
[0053] It should be noted that, since the number of sample data in the X-ray modality is relatively small, sample data in the X-ray modality is generated by using sample data in the CT modality, which helps to complete the model training.
[0054] The process of generating the initial sample set is explained in detail below.
[0055] like Figure 2 As shown, in the CT modality, paired CT report samples and CT volume data samples are acquired. Using LLM to rewrite prompts based on the report, the CT report samples are processed to obtain a synthetic X-ray report.
[0056] In some embodiments, the report rewriting prompt includes command information for deleting X-ray-invisible and inferable content from the composite X-ray report.
[0057] Next, the CT body data samples are subjected to planar projection to obtain a synthetic X-ray image 21. Then, using a projection mask 22, the 3D segmentation and annotation information of each anatomical region is mapped onto the planar projection, thereby obtaining the bounding box information of each anatomical region in the synthetic X-ray image, such as... Figure 2 Image 23 is shown in the image.
[0058] Therefore, the first sample is generated based on the synthetic X-ray report, the synthetic X-ray image, and the bounding box information of each anatomical region in the synthetic X-ray image.
[0059] In addition, the YOLO model is trained using synthetic X-ray images and bounding box information of each anatomical region in the synthetic X-ray images, so that the YOLO model has the ability to detect the bounding boxes of each anatomical region in the X-ray images.
[0060] In the X-ray modality, paired X-ray image samples 24 and X-ray report samples were acquired. The trained YOLO model was used to process the X-ray image samples 24 to obtain the bounding box information of each anatomical region in the X-ray image samples, such as... Figure 2 Image 25 is shown in the image.
[0061] Therefore, a second sample is generated using the X-ray image sample, the X-ray report sample corresponding to the X-ray image sample, and the bounding box information of each anatomical region in the X-ray image sample.
[0062] Finally, an initial sample set is generated based on the first and second samples obtained.
[0063] In step 12, the large language model is used to generate prompt words based on CoT, and the initial sample set is processed to obtain CoT report samples, which include X-ray images, X-ray reports, and inference chains associated with the bounding boxes of each anatomical region.
[0064] In some embodiments, the large language model used here can be a large language model of the GPT (Generative Pre-trained Transformer) class, in order to better generate inference chains.
[0065] In some embodiments, the inference chain includes multiple inference steps, each including inference information and bounding box information of the anatomical region corresponding to the inference information. This achieves fine alignment between each inference step and visual evidence, thereby facilitating the association of the generated diagnostic report with visual evidence of the corresponding anatomical region.
[0066] For example, a sample CoT report can be shown in Table 1.
[0067] Table 1
[0068] In some embodiments, the CoT-generated prompts include command information indicating how to bind each reasoning step to the bounding box information of the corresponding anatomical region. This enables fine alignment between each reasoning step and visual evidence.
[0069] In some embodiments, the CoT-generated prompts may also include the execution order of multiple inference steps.
[0070] It's important to note that when interpreting X-ray images, doctors typically examine the various anatomical regions in a specific order. This is because abnormalities in different anatomical regions are often interconnected; therefore, examining them in a sequential manner helps doctors establish a chain of causal reasoning.
[0071] For example, the corresponding examination sequence for a chest X-ray image is: (a) soft tissue and bony thorax, (b) lung fields, (c) hilum and mediastinum, (d) diaphragm and costophrenic angle, (e) cardiac silhouette and great vessels, and (f) foreign bodies in the chest.
[0072] In some embodiments, CoT can generate prompts as shown in Table 2.
[0073] Table 2
[0074] For example, in one sample, the bounding box information of each anatomical region in the X-ray image is shown in Table 3.
[0075] Table 3
[0076] For example, in this sample, the X-ray reports corresponding to the X-ray images are shown in Table 4.
[0077] Table 4
[0078] Next, the large language model is used to generate prompt words based on CoT, and the initial sample set is processed to obtain CoT report samples, as shown in Table 5.
[0079] Table 5
[0080] In step 13, the multimodal large language model is subjected to supervised fine-tuning (SFT) using CoT report samples to obtain an intermediate model.
[0081] For example, a multimodal large language model can be a model such as Qwen2.5-VL-7B-Instruct or Qwen2.5-VL-32B-Instruct.
[0082] It should be noted that the intermediate model, which is fine-tuned under supervision, has the ability to locate anatomical structures in the input X-ray image and generate an analysis report based on the localization results.
[0083] In other words, the intermediate model, after supervised fine-tuning, already possesses the ability to localize before describing, thus effectively reducing the occurrence of model illusions. Furthermore, since each analysis step is associated with a specific image region, the reliability and credibility of the audit trail are effectively improved.
[0084] In some embodiments, the process steps for supervised fine-tuning are as follows.
[0085] 1) Constructing training data: using Figure 2 The illustrated process generates multiple CoT samples (e.g., 8K), each sample structured as <real X-ray image, prompt word, target output report>, where in the target output report, <think>Including <bbox> Step-by-step reasoning of the token< / bbox> < / think> <answer> FINDINGS & IMPRESSION< / answer> sequence.
[0086] 2) Constructing the input: The image token (the visual encoder output is mapped to the hidden layer dimension of the LLM through the projection layer) is concatenated with the text prompt token to form the input sequence of the LLM.
[0087] 3) Constructing a supervised target: Apply the cross-entropy loss of standard next-token prediction to each token on the target output sequence (only supervise the target output part, the loss of the cue words and image token parts is masked to 0).
[0088] 4) Construct training hyperparameters: For example, construct multiple epochs, such as 3 epochs. For optimizer, learning rate, batch size, warmup parameters, etc., refer to the Qwen2.5-VL official command for fine-tuning configuration.
[0089] 5) Training effect: Through supervised fine-tuning, the model is able to locate anatomical structures in input X-ray images and generate analysis reports based on the localization results.
[0090] In step 14, the intermediate model is trained using reinforcement learning based on the target reward value to obtain the MedCoT-XR model.
[0091] In some embodiments, the step of training the intermediate model using reinforcement learning includes steps S301-S304.
[0092] S301. Based on the bounding box of the target anatomical region predicted by the intermediate model in each inference step and the actual bounding box of the target anatomical region, determine the ground reasoning reward value based on visual localization.
[0093] In some embodiments, multiple intersection-over-union (IoU) ratios are obtained by calculating the intersection-over-union ratio (IoU) between the bounding boxes of the target anatomical region predicted by the intermediate model in each inference step and the actual bounding boxes of the target anatomical region. Next, the average of these multiple IoU ratios is calculated to obtain the inference reward value based on visual localization.
[0094] For example, the inference reward value based on visual location. As shown in formula (1).
[0095] (1)
[0096] In formula (1), N is the total number of reasoning steps in a thought chain. Let be the bounding box of the target anatomical region predicted in the i-th inference step. Let be the actual bounding box of the target anatomical region in the i-th reasoning step. This is the function for calculating the intersection-union ratio.
[0097] It should be noted here that the reasoning reward value based on visual positioning... The closer the value is to 1, the higher the overlap between the area corresponding to the bounding box of the measured target anatomical region and the area corresponding to the actual bounding box of the target anatomical region, meaning the more accurate the visual localization result predicted by the model.
[0098] It should also be noted here that by introducing a reasoning reward value based on visual location... This can help encourage models to conduct more honest and step-by-step analyses.
[0099] S302. Based on the generated report output by the intermediate model for each sample and the corresponding real report for each sample, determine the Medical Factual Consistency Reward.
[0100] In some embodiments, the RadGraph F1 scoring function is used to extract clinical entities and their relationship graphs from the generated report output by the intermediate model for each sample and the real report corresponding to each sample, and to calculate the entity-relationship F1 score between the generated report and the real report as the medical fact consistency reward value.
[0101] For example, the medical fact consistency reward value As shown in formula (2).
[0102] (2)
[0103] In formula (2), The report generated for the i-th sample, For the true report of the i-th sample, The RadGraph F1 scoring function first extracts clinical entities (such as diseases, anatomical locations, attributes, etc.) and their relationship graphs from the report using Radgraph, and then calculates the F1 score of the entity-relationship graph between the predicted graph and the ground truth graph.
[0104] It should be noted here that this is achieved by introducing a reward value for consistency of medical facts. This can encourage models to optimize for factual consistency at the entity and relation levels, resulting in reports that are not only better worded but also have greater clinical applicability.
[0105] S303. Determine the target reward value based on the reasoning reward value based on visual positioning and the medical fact consistency reward value.
[0106] In some embodiments, the step of determining the target reward value may include steps S401-S403.
[0107] S401. Normalize the reasoning reward value based on visual positioning to obtain the first reasoning reward value.
[0108] In some embodiments, the mean and variance of the visual location-based inference reward values obtained from the current training batch are normalized.
[0109] For example, formula (3) can be used to evaluate the reward value for visual positioning-based inference. Normalization is performed to obtain the first reasoning reward value. .
[0110] (3)
[0111] In formula (3), To utilize the average of the visual localization-based inference reward values obtained from the current training batch, To utilize the variance of the visual location-based inference reward values obtained from the current training batch, To avoid small constants with denominators of 0 (e.g., ).
[0112] S402. Normalize the medical fact consistency reward value to obtain the second reasoning reward value.
[0113] In some embodiments, the mean and variance of the medical fact consistency reward values obtained from the current training batch are normalized.
[0114] For example, formula (4) can be used to assign a reward value for consistency of medical facts. Normalization is performed to obtain the second reasoning reward value. .
[0115] (4)
[0116] In formula (4), To utilize the average of the medical fact consistency reward values obtained from the current training batch, To utilize the variance of the medical fact consistency reward values obtained from the current training batch, To avoid small constants with denominators of 0 (e.g., ).
[0117] It should be noted here that, due to the reasoning reward value based on visual positioning... Consistency with medical facts reward value The ranges and variances of the two inference reward values are not the same, which can lead to the gradient of one of the inference reward values dominating, resulting in biased updates and unstable training results. To address this issue, this disclosure separately addresses the inference reward values based on visual localization. Consistency with medical facts reward value Normalization is performed to effectively suppress oscillations during training, thereby achieving stable convergence.
[0118] S403. Calculate the weighted sum of the first reasoning reward value and the second reasoning reward value to obtain the target reward value.
[0119] For example, target reward value As shown in formula (5).
[0120] (5)
[0121] In formula (5), The first reasoning reward value The weights, The second reasoning reward value The weights. and The conditions of formula (6) must be met.
[0122] (6)
[0123] S304. Adjust the parameters of the intermediate model according to the target reward value.
[0124] In some embodiments, the GRPO (Group Relative Policy Optimization) algorithm can be used for policy optimization based on the target reward value.
[0125] In some embodiments, the steps for policy optimization using the GRPO algorithm are as follows.
[0126] 1) Sampling Phase: For each input X-ray image and prompt word, the current policy model... (i.e., MedCoT-XR under training) samples a set (G, e.g., G=8) of candidate outputs { ,..., }, each All are complete. <think> ... <bbox> ... < / bbox> ... < / think> <answer> ... < / answer> "sequence.
[0127] 2) Reward Calculation Phase: For each The reasoning reward value based on visual positioning is calculated using the formula (1) above. The medical fact consistency reward value is calculated using the above formula (2). Next, the normalized inference reward value is calculated using the above formula (3). The normalized inference reward value is calculated using the formula (4) above. Finally, the target reward value is calculated using the formula (5) above. .
[0128] It should be noted that the normalization process described above maps reward values of different scales (IoU values are usually large and RadGraph F1 is small) and different variances to a unified dimensionless scale, thereby preventing one reward value from "dominating" another in the gradient and ensuring training stability and multi-objective balance.
[0129] 3) Dominance function: It should be noted that GRPO uses the within-group mean as the baseline and does not require an additional Critic network.
[0130] 4) Policy update loss (GRPO / PPO-clip form): (7) in, Importance sampling ratio, Set the PPO cutoff threshold (e.g., 0.2). The coefficients of the KL regularization term, This serves as a reference model for freezing after the SFT phase, used to constrain strategies from deviating from clinically reasonable outputs.
[0131] 5) Training cycle: The trainable parameters of MedCoT-XR (visual projection layer and all parameters of LLM) are updated by backpropagation using L(θ) as the loss. The RL phase is trained for 1 epoch.
[0132] It should be noted that by jointly optimizing the reasoning reward value based on visual positioning and the medical fact consistency reward value, continuous improvements can be made in both factual accuracy and visual positioning, thereby making the learning process more reliable.
[0133] It should also be noted that although the intermediate model, after supervised fine-tuning, already possesses the ability to locate before describing, its accuracy is still limited. Therefore, by using a target reward value generated based on inference reward values based on visual localization and medical fact consistency reward values to train the intermediate model through reinforcement learning, the resulting MedCoT-XR model can output more accurate X-ray reports.
[0134] Figure 3 This is a schematic diagram of the structure of a model training device according to an embodiment of the present disclosure.
[0135] like Figure 3 As shown, the model training device 30 can be represented in the form of a general computing device. The model training device 30 includes a memory 31, a processor 32, and a bus 33 connecting different system components.
[0136] The memory 31 may include, for example, system memory, non-volatile storage media, etc. System memory may store, for example, an operating system, application programs, a boot loader, and other programs. System memory may include volatile storage media, such as random access memory (RAM) and / or cache memory. Non-volatile storage media may store, for example, instructions for a corresponding embodiment of an executing model training method. Non-volatile storage media include, but are not limited to, disk storage, optical storage, flash memory, etc.
[0137] Processor 32 can be implemented using a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic devices, discrete hardware components such as discrete gates or transistors. Accordingly, each module, such as the acquisition module, calculation module, and adjustment module, can be implemented by executing instructions in the central processing unit (CPU) running memory to perform the corresponding steps, or by implementing dedicated circuitry to perform the corresponding steps.
[0138] For example, processor 32 is configured for memory-based instruction execution implementation such as Figure 1 The method involved in any of the embodiments.
[0139] Bus 33 can use any of the various bus architectures. For example, bus architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, and the Peripheral Component Interconnect (PCI) bus.
[0140] The interfaces 34, 35, and 36 of the model training device 30, as well as the memory 31 and processor 32, can be connected via bus 33. Input / output interface 34 provides a connection interface for input / output devices such as a monitor, mouse, and keyboard. Network interface 35 provides a connection interface for various networked devices. Storage interface 36 provides a connection interface for external storage devices such as floppy disks, USB flash drives, and SD cards.
[0141] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations thereof, can be implemented by computer-readable program instructions.
[0142] These computer-readable program instructions are provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device to produce a machine, such that execution of the instructions by the processor produces means for implementing the functions specified in one or more boxes of the flowchart and / or block diagram.
[0143] These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions cause a computer to work in a particular manner to produce an article of manufacture, including instructions that implement the functions specified in one or more boxes in a flowchart and / or block diagram.
[0144] This disclosure may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0145] This disclosure also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement... Figure 1 The method involved in any of the embodiments.
[0146] This disclosure also provides a computer program product, including computer instructions, wherein the computer instructions, when executed by a processor, implement as follows: Figure 1 The method involved in any of the embodiments.
[0147] Figure 4 This is a schematic flowchart of a medical image report generation method according to one embodiment of the present disclosure. In some embodiments, the following medical image report generation method is performed by a medical image report generation device, including steps 41-42.
[0148] In step 41, the X-ray image to be detected is acquired.
[0149] In step 42, the MedCoT-XR model is used to generate prompts based on the X-ray report, and the X-ray image to be detected is processed to obtain the X-ray report.
[0150] It should be noted here that the MedCoT-XR model utilizes... Figure 1 The model is trained using the model training method described in any of the embodiments.
[0151] In some embodiments, the MedCoT-XR model generates prompts based on the X-ray report, and the steps for processing the X-ray image to be detected include the following steps S501-505.
[0152] S501. Following the order of anatomical region examination indicated by the prompt words in the X-ray report generation, reason sequentially for each anatomical region in the X-ray image to be examined.
[0153] S502. During the reasoning process for each anatomical region, each anatomical region is located to obtain the bounding box information of each anatomical region. S503. Analyze the regions determined by the bounding box information of each anatomical region to output the analysis results for each anatomical region; S504. Generate an X-ray report based on the analysis results of all anatomical regions in the X-ray image to be detected.
[0154] Figure 5 This is a schematic diagram of the structure of a medical image report generation device according to an embodiment of the present disclosure.
[0155] like Figure 5 As shown, the medical image report generation device includes a memory 51, a processor 52, a bus 53, an input / output interface 54, a network interface 55, and a storage interface 56. Figure 5 and Figure 3 The difference is that, in Figure 5 In the illustrated embodiment, processor 52 is configured to execute instructions stored in memory 51 as follows: Figure 4 The method involved in any of the embodiments.
[0156] This disclosure also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement... Figure 4 The method involved in any of the embodiments.
[0157] This disclosure also provides a computer program product, including computer instructions, wherein the computer instructions, when executed by a processor, implement as follows: Figure 4 The method involved in any of the embodiments.
[0158] The following specific examples illustrate the medical image report generation method disclosed herein.
[0159] Figure 6 This is a schematic diagram of an X-ray image to be processed according to an embodiment of this disclosure. Figure 6 As shown, the X-ray image to be processed is a chest X-ray image. This X-ray image is input into the MedCoT-XR model, which performs the examinations according to a predetermined examination sequence.
[0160] Step 1: Examination of soft tissues and the bony thoracic cavity.
[0161] For example, such as Figure 7 As shown, this is for <bbox> Ribs: [0, 236, 1048, 896]< / bbox> The entire rib area is scanned to examine the soft tissues of the chest wall, the integrity of the ribs, and the presence of fractures or deformities in the bony thorax. Figure 7 The minimum bounding box of the global rib region is displayed in the image.
[0162] Step 2: Examination of lung fields.
[0163] For example, here we use <bbox> lung:[...]< / bbox>(If necessary, further subdivide into left lung, right lung, and each lobe bbox), compare lung markings, tracheal shadows, and the presence or absence of exudative shadows / nodules / masses / pneumothorax and other abnormalities side by side and lobe by lobe.
[0164] Step 3: Examination of the hilum and mediastinum.
[0165] For example, here we use <bbox> Mediastinum: [...]< / bbox> (If applicable) or mediastinal contour bbox, assess hilar density, morphology, mediastinal width, and displacement.
[0166] Step 4: Diaphragm and Costophrenic Angles Examination.
[0167] For example, here we use <bbox> Diaphragm: [...]< / bbox> (Or examine the left and right diaphragms, costophrenic angles, and observe the position and outline of the diaphragm to see if they are smooth and if the costophrenic angles are sharp (obtuse angles indicate fluid accumulation).
[0168] Step 5: Cardiac Silhouette and Great Vessels Examination.
[0169] For example, such as Figure 8 As shown, here we use <bbox> Heart: [440,480,756,686]< / bbox> Calculate the cardiothoracic ratio and assess cardiac silhouette morphology and the course of the great vessels. Figure 8 The minimum bounding box of the heart shadow region is displayed in the image.
[0170] Step 6: Foreign Bodies Inspection.
[0171] For example, check for foreign objects such as metallic shadows, catheters, pacemakers, and stents within the aforementioned bbox areas or in a full-image scan.
[0172] After completing the above steps, the MedCoT-XR model will... <answer> ...< / answer> The system summarizes and outputs the final two diagnostic reports, FINDINGS and IMPRESSION, to generate a complete X-ray report.
[0173] pass Figure 6 , Figure 7 and Figure 8 It is evident that the MedCoT-XR model effectively simulates the workflow of a radiologist: it first locates specific anatomical structures and then generates corresponding text analysis. This approach effectively enhances the interpretability of the reasoning and significantly reduces the risk of the model exhibiting hallucinations.
[0174] The performance of the MedCoT-XR model provided in this publication will be analyzed below using a test dataset.
[0175] CheXpert Plus and IU X-Ray are used as benchmark datasets here. These datasets are widely used in radiology report generation and can reflect both report quality and clinical consistency.
[0176] This section compares the performance of dedicated commercial models and open-source medical multimodal models. Dedicated commercial models include GPT-4.1, GPT-5, and Doubao Seed 1.6. Open-source models are grouped by parameter size. Models with fewer than 10 billion parameters include MedGemma 4B, Qwen2.5-VL 7B, HuatuoGPT-V 7B, and Lingshu 7B. Models with more than 10 billion parameters include MedPlib 14B, MedGemma 27B, Qwen2.5-VL 32B, Lingshu32B, HealthGPT 14B and 32B, and HuatuoGPT-V 34B.
[0177] To enable comparison with existing models within the corresponding parameter range, MedCoT-XR 7B was generated based on Qwen2.5-VL-7B-Instruct, and MedCoT-XR 32B was generated based on Qwen2.5-VL-32B-Instruct. MedCoT-XR 7B and MedCoT-XR 32B were then trained for 3 SFT epochs and 1 RL epoch using 8000 CoT samples.
[0178] The comparison results are shown in Table 6. In Table 6, ROUGE-L, CIDEr, and RaTEScore (RaTE) were used as automatic evaluation indicators. ROUGE-L and CIDEr were used to measure the text overlap and descriptive consistency between the generated report and the actual report, while RaTE was used to evaluate medical semantic consistency and clinical rationality.
[0179] Table 6
[0180] As shown in Table 6, among open-source models with fewer than 10 billion parameters, MedCoT-XR 7B demonstrates strong competitiveness, outperforming almost all existing models in the IU Xray benchmark, with an average performance 5.56% higher than the second-best model. The relative improvement for ROUGE-L increased by 2.24% (from 44.52 to 45.52), for CIDEr by 6.64% (from 200.47 to 213.79), and for RaTE by 7.81% (from 61.99 to 66.83). These results demonstrate that even with limited parameters, the solution provided in this disclosure can effectively improve the quality of medical report generation, exhibiting strong parameter efficiency and generalization ability.
[0181] Furthermore, among open-source models with over 10 billion parameters, MedCoT-XR 32B achieved the best performance in the CheXpert Plus benchmark, significantly outperforming the second-ranked Lingshu 32B model. Specifically, the ROUGE-L metric improved by 2.72% (from 25.29 to 28.01), the CIDEr metric improved by 18.94% (from 77.42 to 96.36), and the RaTE metric improved by 2.31% (from 48.73 to 51.04). These results demonstrate that, at large-scale parameter settings, the models provided in this disclosure offer significant advantages in report completeness, language quality, and medical consistency.
[0182] Furthermore, an ablation study was conducted on the model provided in this disclosure. The results of the ablation study are shown in Table 7.
[0183] Table 7
[0184] As shown in Table 7, "MedCoT-XR 7B w SFT" refers to a model that only performs supervised fine-tuning but not reinforcement learning. "MedCoT-XR 7B w SFT" refers to a model that only performs supervised fine-tuning but not reinforcement learning. "MedCoT-XR7B w Ground Reasoning Reward" refers to a model that performs both supervised fine-tuning and reinforcement learning, but the reinforcement learning only uses visual location-based reasoning rewards. "MedCoT-XR 7B w Medical Factual ConsistencyReward" refers to a model that performs supervised fine-tuning and reinforcement learning, but the reinforcement learning only uses medical factual consistency rewards. "MedCoT-XR 7B w GRPO" refers to a model that performs supervised fine-tuning and reinforcement learning, and the reinforcement learning uses both visual location-based reasoning rewards and medical factual consistency rewards.
[0185] As shown in Table 7, compared with the model that only uses supervised fine-tuning, the model that adds a visual location-based inference reward value can improve performance, the model that adds a medical fact consistency reward can further improve performance, and the model that adds both visual location-based inference reward value and medical fact consistency reward value has the best performance.
[0186] This disclosure unifies anatomical region localization and semantic interpretation, while incorporating visual localization-based inference rewards and medical factual consistency rewards during reinforcement learning. This effectively reduces model hallucinations, enhances factual evidence, and improves interpretability. Extensive experiments in radiology report generation benchmarks demonstrate that the model can generate evidence-based and clinically reliable reports.
[0187] In some embodiments, the functional units described above may be implemented as general-purpose processors, programmable logic controllers (PLCs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any suitable combination thereof for performing the functions described herein.
[0188] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0189] The description in this disclosure is provided for illustrative and descriptive purposes only and is not intended to be exhaustive or to limit the disclosure to its forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of this disclosure and to enable those skilled in the art to understand this disclosure and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A model training method, executed by a model training device, comprising: Generate an initial sample set, wherein each sample in the initial sample set includes an X-ray image, an X-ray report corresponding to the X-ray image, and bounding box information of each anatomical region in the X-ray image; The first language model is used to generate prompt words based on the thought chain CoT, and the initial sample set is processed to obtain CoT report samples, wherein the CoT report samples include the X-ray image, the X-ray report, and the inference chain associated with the bounding box of each anatomical region; The CoT report samples are used to supervise and fine-tune the multimodal large language model to obtain an intermediate model, wherein the intermediate model has the ability to locate anatomical structures in the input X-ray image and generate an analysis report based on the localization results. The intermediate model is trained using reinforcement learning based on the target reward value to obtain the MedCoT-XR model for X-ray report generation based on a multimodal thinking chain.
2. The model training method according to claim 1, wherein, The reinforcement learning training of the intermediate model includes: The visual localization-based inference reward value is determined based on the bounding box of the target anatomical region predicted by the intermediate model in each inference step and the actual bounding box of the target anatomical region. Based on the generated report output by the intermediate model for each sample and the actual report corresponding to each sample, a medical fact consistency reward value is determined. The target reward value is determined based on the visual positioning-based reasoning reward value and the medical fact consistency reward value. The parameters of the intermediate model are adjusted based on the target reward value.
3. The model training method according to claim 2, wherein, The determination of the target reward value includes: The inference reward value based on visual positioning is normalized to obtain the first inference reward value; The medical fact consistency reward value is normalized to obtain the second reasoning reward value; The target reward value is obtained by calculating the weighted sum of the first inference reward value and the second inference reward value.
4. The model training method according to claim 2, wherein, The determination of the reasoning reward value based on visual positioning includes: The intersection-union ratio (IUU) of the bounding box of the target anatomical region predicted by the intermediate model in each inference step and the actual bounding box of the target anatomical region is calculated to obtain multiple IUU ratios. The average value of the multiple intersection-union ratios is calculated to obtain the inference reward value based on visual localization.
5. The model training method according to claim 2, wherein, The determination of the medical fact consistency reward value includes: Using the RadGraph F1 scoring function, clinical entities and their relationship graphs are extracted from the generated report output by the intermediate model for each sample and the real report corresponding to each sample. The entity-relationship F1 score between the generated report and the real report is calculated as the medical fact consistency reward value.
6. The model training method according to claim 1, wherein, The inference chain includes multiple inference steps, wherein each inference step includes inference information and bounding box information of the anatomical region corresponding to the inference information.
7. The model training method according to claim 6, wherein, The CoT-generated prompts include command information indicating how to bind each inference step to the bounding box information of the corresponding anatomical region.
8. The model training method according to claim 7, wherein, The CoT-generated prompts also include the execution order of the multiple inference steps.
9. The model training method according to any one of claims 1-8, wherein, The generation of the initial sample set includes: A first sample is generated using a CT report sample and a CT body data sample corresponding to the CT report sample, wherein the first sample includes a synthetic X-ray report, a synthetic X-ray image, and bounding box information of each anatomical region in the synthetic X-ray image; Using the synthesized X-ray image and the bounding box information of each anatomical region in the synthesized X-ray image, a bounding box detection model is trained; The bounding box detection model is used to process the X-ray image samples to obtain the bounding box information of each anatomical region in the X-ray image samples; A second sample is generated using the X-ray image sample, the X-ray report sample corresponding to the X-ray image sample, and the bounding box information of each anatomical region in the X-ray image sample; The initial sample set is obtained based on the first sample and the second sample.
10. The model training method according to claim 9, wherein, The generation of the first sample includes: The CT report sample is processed using the second language model based on the report rewriting prompts to obtain the synthetic X-ray report; The composite X-ray image is obtained by performing a planar projection on the CT body data sample; Using the anatomical region segmentation and annotation information in the CT body data sample, the bounding box information of each anatomical region in the synthesized X-ray image is obtained; The first sample is generated based on the synthetic X-ray report, the synthetic X-ray image, and the bounding box information of each anatomical region in the synthetic X-ray image.
11. The model training method according to claim 10, wherein, The report rewriting prompt includes command information for deleting X-ray-invisible and inferable content from the synthesized X-ray report.
12. A model training device, comprising: Memory; A processor, coupled to a memory, is configured to implement the model training method as described in any one of claims 1-11 based on memory-stored instruction execution.
13. A method for generating a medical image report, comprising: Acquire the X-ray image to be detected; The MedCoT-XR model is used to generate prompts based on the X-ray report, and the X-ray image to be detected is processed to obtain the X-ray report. The MedCoT-XR model is trained using the model training method described in any one of claims 1-11. The step of using the MedCoT-XR model to generate prompts based on the X-ray report and processing the X-ray image to be detected includes: Following the order of anatomical region examination indicated by the X-ray report generation prompts, reasoning is performed sequentially on each anatomical region in the X-ray image to be examined. During the reasoning process for each anatomical region, each anatomical region is located to obtain the bounding box information of each anatomical region; The region defined by the bounding box information of each anatomical region is analyzed to output the analysis results for each anatomical region; The X-ray report is generated based on the analysis results of all anatomical regions in the X-ray image to be detected.
14. A medical image report generation device, comprising: Memory; A processor, coupled to a memory, is configured to execute instructions stored in the memory to implement the medical image report generation method as described in claim 13.
15. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-11 and 13.
16. A computer program product comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the method as described in any one of claims 1-11, 13.