Medical image report generation method and system, storage medium and equipment
Through the method of deeply fusion of multimodal information, the model is fine-tuned using pre-made data sets and set instructions to generate structured medical image reports, solving the problems of insufficient multimodal dynamic fusion, lack of medical normativeness and weak interpretability in the prior art, and achieving efficient and accurate report generation.
Patent Information
- Application Number
- CN202510368095.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing medical image report generation technology has problems such as insufficient multimodal dynamic fusion, lack of medical normativeness and weak interpretability, which makes it difficult to guarantee the accuracy and efficiency of generating reports.
By deeply fusion of multimodal information, the model is fine-tuned using pre-made data sets and set instructions to generate structured medical image reports. The specific steps include acquiring and preprocessing medical image data, filtering out chest X-ray orthotopic images, and using the report generation model to generate structured reports. The model takes LLaVA-OneVision-7B-Qwen2.5 as the core architecture, and fine-tuning the visual encoder and language model in stages to ensure the structure and accuracy of the generated reports.
The deep fusion of multimodal information is achieved. The generated reports strictly follow medical specifications and have high-quality structured output, which improves the efficiency and accuracy of report generation, enhances the explicit correlation of "image area-text description-medical entity", and reduces the difficulty of clinical verification.
Smart Images

Figure CN120220947A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and specifically to a method, a system, a storage medium and a device for generating a medical image report. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Digital Radiography (DR) is a mature imaging method in the medical field. In the traditional process of generating a DR image report, an experienced radiologist needs to manually write it according to the DR results, and its accuracy highly depends on the doctor's clinical experience and professional level. The efficiency of this method is difficult to guarantee.
[0004] Currently, medical image report generation technologies are mainly based on two types of methods: One is to optimize the vision-language alignment by using a contrastive loss function by comparing the visual features of healthy and diseased examples and the similarity of their corresponding text descriptions. However, only by simply enhancing single-modal features (images or texts) through comparison, cross-modal deep interaction and fusion are not achieved, resulting in insufficient spatial correlation between visual features and semantic descriptions in the generated report. For example, it is impossible to accurately locate the image area corresponding to "atelectasis of the right lower lung". The other is report generation based on cross-modal fusion of medical history and images, which combines the patient's medical history text and image features and generates a report through a Long Short-Term Memory (LSTM) network or an attention mechanism. The reports generated in this way lack guidance on the clinical diagnosis path (for example, do not follow the standard structure of "examination technique → main findings → impression"), resulting in logical confusion or omission of key information.
[0005] In summary, the current medical image report generation technologies generally have the following problems.
[0006] (1) Insufficient multi-modal dynamic fusion: The interaction mechanism between vision, text and the knowledge base is rigid, and the fusion strategy cannot be dynamically adjusted according to lesion characteristics.
[0007] (2) Lack of medical standardization: Since the generated report is a free text report rather than a structured report, it does not strictly follow clinical guidelines (such as the ACR structured template), and the accuracy and logic of terms are difficult to meet the requirements of the radiology department.
[0008] (3) Weak interpretability: The lack of an explicit association between "image area - text description - medical entity" makes it difficult for clinical verification. Summary of the Invention
[0009] To solve the technical problems existing in the above-mentioned background technology, the present invention provides a method, system, storage medium and device for generating medical imaging reports, which are applied in the process of generating chest X-ray reports. By deeply integrating multi-modal information, strictly following medical norms, and supporting structured output, the efficiency is improved.
[0010] To achieve the above object, the present invention adopts the following technical solutions: The first aspect of the present invention provides a method for generating a medical imaging report, including the following steps: Obtain medical imaging data, and after preprocessing and classification, screen out the frontal chest X-ray images; Take the screened frontal chest X-ray images and the problems or information to be queried related to the images as the input of the report generation model, and obtain a structured report for the input images; Among them, the report generation model is fine-tuned using a pre-made data set and set instructions to achieve the generation of a structured report. The data set includes image-text pairs, specifically: frontal chest X-ray images, "answers" related to the images, and structured report texts corresponding to the "answers"; Among them, the set instructions include: "input image" - "answer" - "response", where "response" is "[description & impression]", and "description" is: the morphological characteristics, state characteristics, position characteristics, size characteristics and edge characteristics of each tissue in the chest in the input image, and "impression" is: the findings, preliminary diagnosis opinions and subsequent treatment methods proposed based on the content in the "description".
[0011] Furthermore, the data set also includes text-lesion location pairs, specifically: frontal chest X-ray images, lesion location annotation bounding boxes corresponding to the images, coordinate information of the annotation bounding boxes, and lesion category information corresponding to the annotation; Among them, when fine-tuning the report generation model using text-lesion location pairs, the set instructions are: "input image" - "answer" - "response", where "answer" has "lesion category information" and "lesion location information", and "response" is "[bounding box coordinates]".
[0012] Furthermore, the preprocessing includes: If the label of the medical imaging data carries window width and window level information, then perform standardization and normalization processing according to the window width and window level to obtain a single-channel grayscale image; If the window width and window level information are not carried, then convert the original medical imaging data into a single-channel grayscale image according to the HU value of the image and the set threshold range.
[0013] Further, the classification is specifically as follows: A classifier is constructed through the ChestNet network structure to distinguish the frontal chest X-ray images from other images in the medical image data, and the selected frontal chest X-ray images are used as the input of the report generation model.
[0014] Further, the report generation model uses the LLaVA-OneVision-7B-Qwen2.5 model as the core architecture and is fine-tuned through a pre-made dataset and set instructions. In the instructions, the format of the text information included in "response" follows the clinical guidelines.
[0015] Further, "answer" in the instructions includes questions related to the input image or information to be queried based on the input image. "Answer" is pre-made according to the dialogue format in the real scenario to obtain a dialogue instruction dataset.
[0016] Further, the fine-tuning process of the report generation model is a phased fine-tuning, specifically as follows: In the first stage, the SigLIP visual encoder is frozen, the weights of the SigLIP visual encoder are kept unchanged, and the Qwen-2.5 language model is unfrozen. Only the parameters of the Qwen-2.5 language model are updated during the fine-tuning process. In the second stage, the SigLI visual encoder is unfrozen and fine-tuned layer by layer through different learning rates. Starting from the last layer, the visual encoder is unfrozen layer by layer forward. In the third stage, the SigLIP vision and the Qwen-2.5 language model are jointly fine-tuned. By adjusting the weights of the visual encoder and the language model simultaneously, the visual encoder is allowed to locate the target area, and the language model is allowed to output accurate position information.
[0017] The second aspect of the present invention provides a system required to implement the above method, including: A preprocessing module, configured to: obtain medical image data and preprocess it; An image classification module, configured to: classify and process the preprocessed medical image data to screen out the frontal chest X-ray images; A report generation module, configured to: use the screened frontal chest X-ray images and the questions or information to be queried related to the images as the input of the report generation model to obtain a structured report for the input images; Among them, the report generation model is fine-tuned using a pre-made dataset and set instructions to achieve the generation of a structured report. The dataset includes image-text pairs, specifically: frontal chest X-ray images, "answer" related to the images, and structured report texts corresponding to "answer"; Among them, the set instructions include: "input image" - "answer" - "response", where "response" is "[description & impression]", "description" is: the morphological features, state features, position features, size features and edge features of each tissue in the chest in the input image, and "impression" is: the findings, preliminary diagnosis opinions and subsequent treatment methods proposed based on the content in the "description".
[0018] The third aspect of the present invention provides a computer-readable storage medium.
[0019] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the method for generating a medical image report as described above.
[0020] The fourth aspect of the present invention provides a computer device.
[0021] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for generating a medical image report as described above.
[0022] Compared with the prior art, the above one or more technical solutions have the following beneficial effects: 1. Most of the reports generated by traditional technologies are free-text reports rather than structured reports. Due to the lack of clinical diagnosis path guidance and not following the standard structure of "examination technique → main findings → impression", free-text reports lead to logical confusion or omission of key information, making it difficult to accurately locate some lesion positions. In contrast, this solution uses the computing power of a computer to understand the association between image and text information, changing the process of simulating medical diagnosis in traditional technologies into a process of learning the association between image and text information, and can generate high-quality structured reports, thus improving efficiency.
[0023] 2. The report generation model utilizes the powerful text and image processing capabilities of the existing model, enabling the model to fully understand the image context and its associated clinical information. Since the set instructions and pre-made data sets are used for training, the training cost is lower than that of traditional technologies. Therefore, the generation of structured reports can be achieved through fine-tuning.
[0024] 3. The pre-made dataset contains conversations, structured report generation, visual localization annotation boxes, and relevant clinical information corresponding to the annotation boxes. In addition to being able to generate reports for the input images in the form of question and answer, it can also annotate the lesion categories and lesion locations corresponding to the reports in the input images, enhancing the spatial correlation between visual features and semantic descriptions in the generated reports, achieving an explicit association of "image region - text description - medical entity", and greatly reducing the difficulty of clinical verification.
[0025] 4. During model fine-tuning, since the text data in the made dataset is already "structured" and has followed the relevant requirements of the medical industry, the model can learn the correlation between text information and the correlation with image information, thereby fully understanding the image context and its associated clinical information, and finally generating high-quality structured reports. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0027] Figure 1 It is a schematic diagram of the overall process of generating a medical image report provided by one or more embodiments of the present invention; Figure 2 It is a schematic diagram of the report generation process for a chest X-ray image provided by one or more embodiments of the present invention; Figure 3 It is a schematic diagram of the architecture of the ChestNet classifier provided by one or more embodiments of the present invention; Figure 4 It is a schematic diagram of the architecture of the medical image report generation system provided by one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The present invention will be further described below in conjunction with the drawings and embodiments.
[0029] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0030] It should be noted that the terms used in the following embodiments are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0031] The following embodiments provide a method, system, storage medium, and device for generating a medical imaging report, which are applied in the process of generating a chest X-ray report. By deeply integrating multi-modal information and realizing the display association between the image region and the text description, a structured chest X-ray report that meets medical specifications is generated to improve efficiency.
[0032] Embodiment 1: As Figure 1 shown, the method for generating a medical imaging report includes the following steps: Obtain medical imaging data, and after preprocessing and classification, screen out the frontal chest X-ray images; Use the screened frontal chest X-ray images and the problems or information to be queried related to the images as the input of the report generation model, and obtain a structured report for the input images; Among them, the report generation model is fine-tuned using a pre-made data set and set instructions to achieve the generation of a structured report. The data set includes image-text pairs, specifically: frontal chest X-ray images, "answers" related to the images, and structured report texts corresponding to the "answers"; Among them, the instruction for fine-tuning the report generation model is: "input image" - "answer" - "response", where "response" is "[description & impression]", and "description" is: the morphological features, state features, position features, size features, and edge features of each tissue in the chest in the input image, and "impression" is: the findings, preliminary diagnosis opinions, and subsequent treatment methods proposed based on the content in the "description".
[0033] The data set also includes text-lesion location pairs, specifically: frontal chest X-ray images, lesion location annotation bounding boxes corresponding to the images, coordinate information of the annotation bounding boxes, and lesion category information corresponding to the annotation; Among them, the instruction for fine-tuning the report generation model using text-lesion location pairs is: "input image" - "answer" - "response", where "answer" has "lesion category information" and "lesion location information", and "response" is "[bounding box coordinates]".
[0034] In this embodiment, "answer" refers to a question related to a chest X-ray image or information to be queried.
[0035] Perform data preprocessing on the input DR image. This includes DR image window processing, standardizing and normalizing the Dicom format data according to window width and window level to ensure the consistency and comparability of image data, and converting the normalized data into [0, 255] natural image data to obtain image information.
[0036] Construct a chest X-ray PA classifier. By constructing a classifier to distinguish chest X-ray PA images from other images (such as chest X-ray lateral images, abdominal images, joint images, etc.), select chest X-ray PA images as the input for the multi-modal vision-language large model.
[0037] Construct a report generation model, that is, construct a chest X-ray multi-modal vision-language large model. In this embodiment, Chest-DRVLM (Chest DR Vision Language Model, Chest-DRVLM) is selected, with the LLaVA-OneVision-7B-Qwen2.5 model as the core architecture. The language model Qwen-2.5 not only has an increase in the number of parameters, but also has significant improvements in the training dataset and optimization algorithm, and can better generate structured and accurate medical reports. In Chest-DRVLM, the input is a chest X-ray image and "answer", and the output is a structured report containing detailed findings descriptions and diagnostic conclusions, and the bounding boxes (Bboxes) of the lesions are output, and the positions of these lesions are visually displayed on the image to ensure the explicit association of "image region - text description - medical entity", and the accuracy of the output report can be judged more intuitively.
[0038] Adopting the above method significantly improves the efficiency and accuracy of chest X-ray image analysis and report generation. It can automatically extract key information from chest X-ray images and generate structured medical reports, reducing the time and workload of doctors writing reports manually. Using the advanced Qwen-2.5 language model, the system can generate high-quality structured reports that meet medical specifications, provide detailed findings descriptions and diagnostic conclusions, and enhance the reliability and practicality of the reports. In the inference stage, the visualization of the structured report and the bounding boxes (Bboxes) of the lesions enhances the spatial correlation between the visual features and semantic descriptions in the generated report, realizes the explicit association of "image region - text description - medical entity", and greatly reduces the difficulty of clinical verification.
[0039] This embodiment is based on a multi-modal vision-language model, and uses chest X-ray image data to obtain a structured report. The process of report generation is as Figure 2 shown, including the following steps: Step S101: The preprocessing stage mainly includes window processing for chest X-ray images.
[0040] Specifically, if the window width and window level information is carried in the Dicom data tag, according to the standardization and normalization processing of the window width and window level, the normalized data is converted into [0, 255] natural image data, which is a single-channel grayscale image, in order to obtain clear chest X-ray image information. The specific window width and window level values come from the window width and window level values carried in the Dicom data.
[0041] If the DR image data carries the window width and window level information, the window processing is as follows: ; ; ; where ww is the window width; wl is the window level; img_min is the lower limit of the window width for clearly viewing the couch; img_max is the upper limit of the window width for clearly viewing the couch, is the CT image to which it belongs, is the image after window processing; is to take the largest integer not greater than . "Couch" refers to the examination couch or scanning couch on which the patient lies in the imaging device.
[0042] If the window width and window level information does not appear in the Dicom data tag, the HU value range of the image is selected as the values between the 0.05th and 0.95th percentiles. The HU values of all pixels are statistically counted , and the values of the 0.05th and 0.95th percentiles are determined, denoted as and , and the pixel values are linearly mapped to the range of [0, 255] and converted into a single-channel grayscale image.
[0043] If the DR image data does not carry the window width and window level information, the window processing is as follows: ; After the window processing of the image, all chest X-ray images are adjusted to a unified size for subsequent input into the classification model.
[0044] Step S102: Input the chest X-ray images obtained in Step S101 into the trained chest X-ray posteroanterior image classifier to output the category of the DR image: chest X-ray posteroanterior image category or other DR image category. The selected chest X-ray images are used as the input to the chest X-ray multi-modal vision-language large model, and the selected other DR images are deleted and a prompt is given that the image is other DR image.
[0045] Among them, the training process of the frontal chest X-ray image classifier is as follows Figure 3 shown, including the following steps: Step S1021: Make a chest X-ray image dataset and store it in the frontal chest X-ray image folder and other image folders respectively.
[0046] Step S1022: Construct a classifier module to build a frontal chest X-ray image classifier for screening frontal chest X-ray images and other images. The input is the chest X-ray images generated by the preprocessing module, and the output is the category of the images.
[0047] Specifically, construct the ChestNet network structure, whose network structure includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and a multi-scale dynamic feature weighting mechanism layer, and the training loss function is the cross-entropy loss function.
[0048] Input layer: All input images are uniformly adjusted to 224×224 pixels, single-channel grayscale images.
[0049] Convolutional layer 1 (Conv1): This layer contains 32 filters, each filter is 5×5 in size, the stride is 1, and the padding method is SAME. The activation function uses ReLU. After the convolution operation, the output size is 224×224×32.
[0050] Max pooling layer 1 (Pool1): Use a 2×2 pooling window and a max pooling operation with a stride of 2 to reduce the feature map of the previous layer by half, and the output size becomes 112×112×32.
[0051] Multi-scale dynamic feature weighting mechanism layer: At this stage, introduce the multi-scale dynamic feature weighting mechanism (Multi-Scale Dynamic Feature Weighting Mechanism, MSD-FWM). Specifically, use convolutional kernels of different sizes (such as 3×3, 5×5, and 7×7) to perform convolution operations on the input feature map to extract features of different scales. Assume that the number of filters for each scale is 64, then multiple feature maps of different scales are output, and the output size of each scale is 112×112×64. For each scale of feature map, first obtain the global information through global average pooling , then generate dynamic weights through a multi-layer perceptron (MLP) , and use the dynamic weights to weight each scale of feature map to obtain the weighted feature map , and finally combine the weighted feature maps Stitch and fuse by channel. In the case of three scales, the output size after stitching is 112×112×192.
[0052] The global information among them is: ; The dynamic weights among them are: ; The weighted feature maps among them are: ; Convolutional layer 2 (Conv2): It is set with 64 filters, the filter size is 3×3, the stride is 1, the padding method is SAME, and ReLU is used as the activation function. The input comes from the feature weighted fusion layer, with a size of 112×112×192, and the output size after convolution operation is 112×112×64.
[0053] Max pooling layer 2 (Pool2): Applies a 2×2 pooling window and max pooling operation with a stride of 2 to further compress the feature map size to 56×56×64.
[0054] Convolutional layer 3 (Conv3): Configures 128 filters, each filter size is 3×3, the stride is 1, the padding method is SAME, and ReLU is used as the activation function. The output size is 56×56×128.
[0055] Global average pooling layer: Averages the spatial dimensions of each feature map into 1 value, and the output size is 128.
[0056] Fully connected layer: The final fully connected layer has 2 output units for the binary classification task: frontal chest X-ray images and other images. The activation function selects Softmax, and the output size is 2.
[0057] The basic architecture of ChestNet, after integrating the multi-scale dynamic feature weighting mechanism (MSD-FWM), significantly enhances the model's ability to capture complex patterns, especially performing excellently in medical image analysis fields such as chest X-ray image classification. This structure ensures that the importance of each feature can be adaptively adjusted at different stages of the network, improving the robustness and generalization ability of the model.
[0058] Step S103: After passing the chest X-ray image data obtained in Step S101 through the classifier in Step S102, determine whether it is a frontal chest X-ray image. If so, input the chest X-ray image data and the relevant "answer" into the trained chest X-ray multi-modal vision-language large model module to output a structured report of the chest X-ray image and the bounding box (Bbox) of the lesion.
[0059] Furthermore, the chest X-ray multi-modal vision-language large model is fine-tuned using Chest-DRVLM. The fine-tuning process includes: Step S1031: The chest X-ray multi-modal vision-language large model module constructs the chest X-ray multi-modal vision-language large model Chest-DRVLM with the LLaVA-OneVision-7B-Qwen2.5 model as the core architecture. In the LLaVA-OneVision-7B-Qwen2.5 model, the language model is upgraded to Qwen-2.5, integrated with the visual encoder SigLIP and the 2-layer MLP cross-modal adapter, and the initialized model and model fine-tuning are performed according to the constructed large-scale chest X-ray instruction dataset to ensure that the model can fully understand the image context and its associated clinical information during training and generate high-quality structured reports.
[0060] Among them, during the Chest-DRVLM fine-tuning process, in order to ensure that Chest-DRVLM can fully understand the image context and its associated clinical information and generate high-quality structured reports, this solution constructs a large-scale structured report dataset including dialogue, structured report generation, and visual localization Bbox boxes. In the constructed large-scale structured report dataset, the report text corresponding to each DR image must follow clinical guidelines, such as the ACR structured template can be referred to. The images and structured reports in this dataset come from cooperative hospitals and have obtained the permission of the subjects. The DR images are in dicom format, and each image corresponds to a structured report written by a doctor. The structured report must contain two parts, namely the description part and the impression part, which detail all the findings observed from the DR image, including normal and abnormal conditions. To achieve visual localization, according to the lesion content given in the report, radiologists are invited to annotate the Bbox annotation boxes.
[0061] Furthermore, the process of constructing the large-scale chest X-ray structured report dataset includes the following steps: Step S1032: The dataset contains a large number of image-text pairs. The main purpose is to input the chest X-ray image and "answer" into Chest-DRVLM, and the output is a structured report.
[0062] When fine-tuning Chest-DRVLM, on the one hand, the report generated by the model for "answer" needs to meet the information format included in the structured report to give a complete prompt and ensure that Chest-DRVLM outputs a structured report rather than a free-text report; On the other hand, the design of the text (i.e., "response") needs to follow clinical guidelines such as referring to the ACR structured template. In the description part, all findings observed from the chest X-ray image should be detailedly recorded, including normal and abnormal conditions. In the impression part, the main findings should be clearly expressed, the most significant problems should be concisely summarized, and a diagnostic opinion should be provided.
[0063] In the text design of the dataset, the content of "description" should include all or several of the following key points: (1) Conditions in the lungs: Evaluate the transparency of the lung fields, whether there is an increase or decrease in density, and describe its location, shape and other characteristics. Whether there are nodules, masses or other abnormal shadows, and describe their size, shape, status, edge characteristics, etc. For example, the report describes that the lung markings are naturally distributed, the hilar structures are normal, the bilateral lung markings are increased, thickened, blurred and disordered, multiple patchy high densities are seen, a high-density patchy and cord-like shadow is visible in the upper right lung field, with high density and clear edges, etc.
[0064] (2) Conditions of the skeletal structure: For example, the report describes that the bilateral bony thoraxes are symmetric, etc.
[0065] (3) Conditions of the mediastinum and trachea: For example, the report describes that the mediastinum and trachea are centered, no mediastinal shift is seen, etc.
[0066] (4) Conditions of the heart and great vessels: For example, the report describes that the cardiac silhouette is slightly enlarged, the cardiac silhouette is normal in shape, aortic calcification shadows, etc.
[0067] (5) Conditions of the pleura and diaphragm: For example, the report describes that the bilateral diaphragmatic surfaces and costophrenic angles are slightly blurred, the left costophrenic angle is blunted, etc.
[0068] In the text design, the "impression" part includes all or several of the following key points: (1) Main findings or problems: For example, the impression in the report is that the bilateral lung markings are increased, the bilateral lung markings are thickened, a round nodule about 2 cm in size is seen in the upper right lobe of the right lung, with smooth edges, etc.
[0069] (2) Preliminary diagnostic opinions: For example, the impression in the report is bilateral pneumonia, a small amount of fluid in the left thoracic cavity, etc.
[0070] (3) Suggested follow-up steps: For example, the impression in the report is to combine with CT examination, etc.
[0071] According to the above requirements, the instructions in the structured dataset are designed as follows: "Input image" - "answer" - "response", where "response" is "[Description & Impression]", "Description" is: the morphological features, state features, position features, size features, and edge features of each tissue in the chest in the input image, and "Impression" is: the findings, preliminary diagnosis opinions, and subsequent treatment methods proposed in sequence according to the content in "Description".
[0072] For example, it can be: Input: Chest X-ray image.
[0073] answer: "Please generate a structured report based on this chest X-ray image."
[0074] response: "[Description: The bilateral bony thorax is symmetric, the lung markings in both lungs are increased and disordered, patchy blurred shadows are seen in both lungs, the trachea is in the middle, the mediastinum is not displaced, the cardiac silhouette is slightly enlarged, and the left costophrenic angle is blunted. Impression: Inflammation of both lungs, please combine with clinical manifestations, the cardiac silhouette is slightly enlarged, and there is a small amount of pleural effusion in the left chest cavity]". The dataset also contains a large number of text-lesion location pairs. The main purpose is to input a query about the location of a specific lesion in Chest-DRVLM and output the corresponding bounding box coordinates. Specifically: for each image, select the lesion location and mark the bounding box coordinates . Normalize the coordinates to floating-point values between [0, 1]. In the Chest-DRVLM model, the categories of specific lesions include infiltration, increased density shadow, thickened texture, nodule, mass, interstitial pneumonia, postoperative change, pleural effusion, atelectasis, aortic calcification, enlarged cardiac silhouette, aortic dilation categories. In the design of this part of the instructions, the content of answer needs to include the category of the specific lesion and the location where the lesion appears, such as the two parts of information "pleural effusion" and "location", so as to output the location information of the specified lesion.
[0075] According to the above lesion categories, the instructions are designed as follows: "answer" - "response", where "answer" includes "the category of the lesion" and "the location where the lesion appears", and "response" is "[Bounding box coordinates]". "answer" refers to the question related to the chest X-ray image or the information to be queried. When the text in "question" or "information" contains "the category of the lesion" and "the location where the lesion appears", the model outputs the corresponding "response", that is, the coordinates of the bounding box are output.
[0076] For example, it can be: answer: "Please indicate the location of the pleural effusion area on this chest X-ray." response: "[Bounding box coordinates]".
[0077] The above instruction design brings out the image and text processing capabilities of the Chest-DRVLM model. When the model is trained using a data set and designed instructions, it is easier to make the model reach an ideal state. Compared with traditional methods, the cost of training the model is lower, and the model can be trained by fine-tuning the model.
[0078] In practical applications, considering that the format of the doctor's "answer" is not unique, such as "Please generate a structured report based on this chest X-ray", it is necessary to collect the dialogue format of doctors in real scenarios in the instruction dataset and create a dialogue instruction dataset covering various types of questions and answers. Then the instruction can be designed as: answer: "Generate structured reports." response: "[Description: The chest is symmetrical. The texture of both lungs runs naturally. The hilar structure is normal. The heart shadow is normal. The two diaphragmatic surfaces are smooth, and the costophrenic angles are sharp. Impression: No obvious abnormalities are found in the chest]" answer: "Please specify the location of the increased density shadow." response: "[bounding box coordinates]".
[0079] Furthermore, the specific steps in the Chest-DRVLM fine-tuning stage are as follows: Step S1033: During the Chest-DRVLM training and fine-tuning phase, since the language model is upgraded to Qwen-2.5, the model needs to be initialized using the Qwen-2.5 pre-trained weights, and the trained SigLIP visual encoder and MLP adapter parameters.
[0080] The Chest-DRVLM framework uses an autoregressive loss function for model training. During the fine-tuning process, the fine-tuning strategy is staged fine-tuning. The first stage freezes the SigLIP visual encoder and fine-tunes the Qwen-2.5 language part. The purpose is to quickly adjust the language model to adapt to new task instructions, such as ansewr instructions containing specific words 'structured report', 'location' and response instructions containing words 'description', 'impression', '[bounding box coordinates]'.
[0081] The technical principle of the first stage is as follows: In this stage, freeze the SigLIP visual encoder, keep the weights of the SigLIP visual encoder unchanged, unfreeze the Qwen-2.5 language model, and only update the parameters of the Qwen-2.5 language model during fine-tuning. The existing powerful natural language processing capabilities of Qwen-2.5 can be utilized to quickly make the model familiar with new instruction sets specific to chest DR images (e.g., "structured report" and "location"). Since the visual feature extraction part remains unchanged, this ensures that the understanding of existing visual information is not disrupted, while enabling the language model to better understand the new task requirements.
[0082] In the second stage, gradually unfreeze and fine-tune the SigLI visual encoder, starting from the last layer and moving forward layer by layer. And during fine-tuning, set different learning rates for different layers, with a lower learning rate for the layers closer to the input and a higher learning rate for the layers closer to the output to ensure that the model does not become unstable due to excessive updates; and use the self-attention mechanism to enhance the attention of the visual encoder to specific lesion areas in the DR image to improve the accuracy.
[0083] Specifically, the learning rate settings for different layers are as follows: ; where, is the learning rate of the th layer, is the base learning rate, is the th layer's number of parameters.
[0084] Specifically, the self-attention mechanism is as follows: ; where, represents the attention score of the i th visual token to the jth visual token, is the normalized attention weight.
[0085] In the third stage, jointly fine-tune the SigLIP vision and Qwen-2.5 language models. By simultaneously adjusting the weights of the visual encoder and the language model, ensure a higher semantic consistency between the visual features extracted from the image and the text generated by the language model. For the output task of bounding boxes (bboxes) containing specific lesions, joint fine-tuning allows the visual encoder to more accurately locate the target area and enables the language model to output precise location information, reducing the error of information transmission between the visual encoder and the language model when only fine-tuning one model in the first and second stages, thereby improving the overall performance.
[0086] Among them, during the fine-tuning process, autoregressive loss modeling is used to adjust the model parameters through backpropagation, enabling the model to better predict the next target token, thereby generating coherent and accurate answers. The autoregressive loss formula is as follows: ; Among them, represents the target token generated by the expression model, represents the visual token, represents the question token, represents at the current token the sequence of target tokens generated by the model before, and L represents the number of tokens in the target sequence generated by the model. represents given the visual token and the question token when the model generates the target token the conditional probability.
[0087] In traditional technologies, whether it is by comparing "healthy and diseased examples" or by combining the patient's medical history text and imaging features to generate reports, the computing power of the computer is used to try to simulate the doctor's diagnosis process. However, the doctor's diagnosis process is complex, resulting in most of the generated reports being free-text reports. Only the information related to the lesion is found, and the obtained reports are logically unconnected. Therefore, there is a lack of clinical diagnosis path guidance and the standard structure of "examination technique → main findings → impression" is not followed, leading to logical chaos or omission of key information, and thus it is difficult to accurately locate some lesion positions. For example, in the example in the background technology, "it is impossible to accurately locate the image area corresponding to 'atelectasis of the right lower lung'".
[0088] In this solution, by utilizing the powerful text and image processing capabilities of the existing large model and through model fine-tuning, the model can fully understand the image context and its associated clinical information during the training process. By using the computing power of the computer to understand the association between image and text information, the process of simulating medical diagnosis in traditional technologies is changed to a process of learning the association between image and text information, thereby generating high-quality reports; Meanwhile, during model fine-tuning, since the text data in the prepared dataset is already "structured" and has followed the relevant requirements of the medical industry, the model can learn the association between text information and the association with image information, thereby fully understanding the image context and its associated clinical information, and finally generating high-quality structured reports.
[0089] Embodiment 2: As Figure 4 shown, a medical image report generation system includes: A preprocessing module, configured to: obtain medical image data and perform preprocessing; An image classification module, configured to: perform classification processing on the preprocessed medical image data, and screen out the frontal chest X-ray images; A report generation module, configured to: use the screened frontal chest X-ray images and the problems or information to be queried related to the images as the input of a report generation model, and obtain a structured report for the input images; Among them, the report generation model is fine-tuned using a pre-made dataset and set instructions to achieve the generation of a structured report. The dataset includes image-text pairs, specifically: frontal chest X-ray images, "answers" related to the images, and structured report texts corresponding to the "answers"; Among them, the set instructions include: "input image" - "answer" - "response", where "response" is "[description & impression]", and "description" is: the morphological characteristics, state characteristics, position characteristics, size characteristics, and edge characteristics of each tissue in the chest in the input image, and "impression" is: the findings, preliminary diagnosis opinions, and subsequent treatment methods proposed based on the content in the "description".
[0090] Embodiment 3: This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the method for generating a medical image report as described in Embodiment 1 above.
[0091] Embodiment 4: This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for generating a medical image report as described in Embodiment 1 above.
[0092] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0093] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for generating a medical imaging report, characterized in that: The following steps are involved: Obtain medical imaging data, pre-process and classify, and filter out chest X-ray positive images; The screened chest X-ray positive image and questions or information to be queried related to the image are used as inputs of the report generation model to obtain a structured report for the input image; The report generation model is fine-tuned using a pre-made dataset and set instructions to achieve the generation of structured reports. The dataset includes image-text pairs, specifically: chest X-ray images, "answers" related to the images, and structured report text corresponding to the "answers"; The set instructions include: "input image" - "answer" - "response", where "response" is "[description & impression]", where "description" is: the morphological characteristics, state characteristics, position characteristics, size characteristics and edge characteristics of each chest tissue in the input image, and "impression" is: the findings, preliminary diagnosis opinions and subsequent processing methods proposed based on the content in the "description".
2. The method for generating a medical imaging report according to claim 1, wherein: The dataset also includes text-lesion location pairs, specifically: chest X-ray positive image, lesion location annotated bounding box corresponding to the image, coordinate information of the annotated bounding box, and lesion category information corresponding to the annotation; Among them, when the fine-tuning report generation model uses text-lesion location pair fine-tuning, the set instructions are: "input image"-"answer"-"response", where "answer" has "lesion category information" and "lesion location information", and "response" is "[bounding box coordinates]".
3. The method for generating a medical imaging report according to claim 1, wherein: The preprocessing is specifically as follows: If the label of the medical imaging data carries the window width and window position information, the single-channel grayscale image is obtained by standardizing and normalizing the window width and window position; If the window width and window position information is not carried, the original medical image data is converted into a single-channel grayscale image according to the HU value of the image and the set threshold range.
4. The method for generating a medical imaging report according to claim 1, wherein: The classification is specifically as follows: a classifier is constructed through a ChestNet network structure to distinguish chest X-ray images from other images in medical imaging data, and the selected chest X-ray images are used as input to a report generation model.
5. The method for generating a medical imaging report according to claim 1, wherein: The report generation model uses the LLaVA-OneVision-7B-Qwen2.5 model as the core architecture and is fine-tuned using pre-made datasets and set instructions. In the instructions, the format of the text information contained in "response" follows clinical guidelines.
6. The method for generating a medical imaging report according to claim 1, wherein: The "answer" in the instruction includes questions related to the input image, or information to be queried based on the input image. The "answer" is pre-made according to the dialogue format in the real scene to obtain a dialogue instruction dataset.
7. The method for generating a medical imaging report according to claim 1, wherein: The fine-tuning process of the report generation model is a stage-by-stage fine-tuning process, specifically: In the first stage, the SigLIP visual encoder is frozen, the weight of the SigLIP visual encoder is kept unchanged, the Qwen-2.5 language model is unfrozen, and only the parameters of the Qwen-2.5 language model are updated during fine-tuning; In the second stage, the SigLI visual encoder is unfrozen and fine-tuned layer by layer using different learning rates, starting from the last layer and unfreezing the visual encoder layer by layer. In the third stage, the SigLIP vision and Qwen-2.5 language models are jointly fine-tuned. By simultaneously adjusting the weights of the visual encoder and the language model, the visual encoder is allowed to locate the target area and the language model outputs precise location information.
8. A system for generating a medical imaging report, characterized in that: include: The preprocessing module is configured to: acquire and preprocess medical imaging data; The image classification module is configured to: classify and process the pre-processed medical image data, and screen out chest X-ray positive images; The report generation module is configured to: use the screened chest X-ray positive image and the questions or information to be queried related to the image as inputs of the report generation model to obtain a structured report for the input image; The report generation model is fine-tuned using a pre-made dataset and set instructions to achieve the generation of structured reports. The dataset includes image-text pairs, specifically: chest X-ray images, "answers" related to the images, and structured report text corresponding to the "answers"; The set instructions include: "input image" - "answer" - "response", where "response" is "[description & impression]", where "description" is: the morphological characteristics, state characteristics, position characteristics, size characteristics and edge characteristics of each chest tissue in the input image, and "impression" is: the findings, preliminary diagnosis opinions and subsequent processing methods proposed based on the content in the "description".
9. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the steps in the method for generating a medical imaging report as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method for generating a medical imaging report according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Medical image report generation method, system and device and storage medium
CN121439069A