Chinese Medical Report Generation Method, System and Terminal Based on Pre-trained Large Model
By combining a pre-trained model of medical vision encoder and a large language model, hallucinations problems in the generation of medical imaging reports are solved, more accurate lesion description and diagnostic suggestions are achieved, and the quality and efficiency of reports are improved.
Patent Information
- Application Number
- CN202411308793.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-19
AI Technical Summary
There are hallucinations in existing medical image report generation techniques, leading to the risk of misdiagnosis or misdiagnosis, especially when multimodal large models generate text that does not match the content of the input image or is logically unreasonable.
The Chinese medical report generation method based on pre-trained large models is adopted, combined with medical vision encoder and medical large language model, image feature extraction and report matching are performed through a similar report search module, and image alignment training and instruction fine-tuning are used to optimize model performance to reduce hallucination problems.
It improves the accuracy and efficiency of medical reports, ensures the accuracy of lesion description and diagnostic suggestions, reduces hallucinations when generating reports, and improves the credibility and practicality of reports.
Smart Images

Figure CN119207694B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, system and terminal for generating Chinese medical reports based on a pre-trained large model. Background Art
[0002] With the rapid development of medical technology, medical images play an increasingly crucial role in clinical diagnosis. As an important basis for doctors' diagnosis and treatment, the accuracy and detail of medical image reports directly affect medical quality and patient safety.
[0003] Currently, the technology for automatically generating medical image reports is mainly based on deep learning and multi-modal large model methods. Multi-modal large models usually consist of a visual encoder and a large language model (LLM), providing an effective way for image-to-text generation. These models combine image information and language information to achieve the automatic conversion of medical images into diagnostic reports. However, multi-modal large models may have hallucination problems when generating text, that is, generating descriptions that do not match the input image content or are logically unreasonable. This problem is particularly serious in medical report generation because it may lead to misdiagnosis or missed diagnosis, bringing uncontrollable risks to patients. This is the deficiency of the prior art. Summary of the Invention
[0004] To solve the deficiencies of the prior art, the present invention provides a method, system and terminal for generating Chinese medical reports based on a pre-trained large model. The pre-trained large model can fully understand the complex content and lesion characteristics in medical images, reducing the hallucination problem when generating text.
[0005] In a first aspect, the present invention provides a method for generating a Chinese medical report based on a pre-trained large model, the method comprising: S1, obtaining an X-ray medical image of the report to be generated;
[0006] S2, obtaining the requirements for generating a Chinese medical report input by the user;
[0007] S3, importing the obtained X-ray medical image and the requirements for generating a Chinese medical report into a pre-trained large model to obtain a Chinese medical report of the X-ray medical image; the large model includes a medical visual encoder, a medical large language model and a similar report retrieval module;
[0008] The medical visual encoder is used to encode the input X-ray medical image to obtain visual features of the X-ray medical image, and the visual features include global features and local features;
[0009] The similar report retrieval module is used to retrieve the medical reports corresponding to the X-ray medical images similar to the X-ray medical image according to the global features of the X-ray medical image obtained by encoding in the preset image memory and the medical report files to be queried, and input the retrieved medical reports into the medical large language model; the image memory stores the global feature vectors of historical X-ray medical images and their corresponding index numbers; the medical report files to be queried store the medical reports corresponding to the historical X-ray medical images and the index numbers of the global feature vectors of the historical X-ray medical images;
[0010] The medical large language model is used to generate the Chinese medical report of the X-ray medical image according to the visual features of the X-ray medical image obtained by the medical visual encoder, and according to the input Chinese medical report generation requirements and the retrieved medical reports.
[0011] Further, the input X-ray medical image is encoded to obtain the visual features of the X-ray medical image, including;
[0012] The input X-ray medical image is segmented into N preset non-overlapping image blocks of a fixed size to obtain an image block sequence ;
[0013] The image block sequence is mapped to a vector space of dimension through the fully connected layer in the medical visual encoder, to obtain an image block embedding sequence ; ;
[0014] An image classification embedding is added to the head of the image block embedding sequence , and then a position encoding is added to each image block in the image block embedding sequence with the image classification embedding added to the head and the image classification embedding , to obtain the image block encoding sequence of the X-ray medical image ; ;
[0015] The image block encoding sequence is imported into the first layer of the Transformer encoder in the medical visual encoder, and the output of each layer is used as the input of the next layer of the Transformer encoder for feature extraction layer by layer. After encoding by the last layer of the Transformer encoder, the visual features of the X-ray medical image are output ;
[0016] Among them, represents the X-ray medical image, represents the th image block obtained by segmenting the acquired X-ray medical image, , Indicates the height of the th image patch, Indicates the th image patch width, Indicates the th number of channels of the image patch, , Indicates the number of image patches obtained by segmenting the acquired X-ray medical image, Indicates the fully connected layer described, Indicates the dimension of the vector space described, Indicates image classification embedding, Indicates image position encoding, Indicates the image patch encoding sequence; Indicates the visual features of the X-ray medical image, Indicates the global features of the X-ray medical image, Indicates the local features of the X-ray medical image; Indicates the Transformer encoder.
[0017] Furthermore, before the medical large language model generates the Chinese medical report of the X-ray medical image based on the visual features of the X-ray medical image obtained by the medical vision encoder and according to the input Chinese medical report generation requirements and the retrieved medical reports, it further includes: performing visual embedding processing on the visual features of the X-ray medical image, including:
[0018] Through the trained projection layer Map the local features of the X-ray medical image to the same dimension as the word embedding space of the medical large language model to obtain the visual embedding of the X-ray medical image , ;
[0019] The training method of the projection layer includes:
[0020] Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and the obtained general data to obtain the projection layer sample training set;
[0021] Image-text alignment training: Use the data in the projection layer sample training set Perform data conversion using the multi-modal template, and use the converted data for the projection layer Perform iterative training, and utilize the partial data in the converted mixed data and the loss function to calculate the loss value of the projection layer After that, use the backpropagation algorithm to calculate the gradient of the parameters of the projection layer and use the gradient descent method to update the parameters of the projection layer until the preset number of iterative training times is completed or the calculated loss value of the projection layer no longer decreases, and complete the training of the projection layer to obtain a trained projection layer ;
[0022] Among them, represents the visual features of the images of the data in the projection layer sample training set, and represents the text description of the images of the data in the projection layer sample training set.
[0023] Furthermore, a method for constructing a medical large language model includes:
[0024] Create a Qwen-7b model as the initial model and perform model initialization;
[0025] A training method for a medical large language model includes:
[0026] Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and the obtained general data to obtain a model sample training set;
[0027] Instruction fine-tuning training: After importing the data in the model sample training set into the initial model, use the multi-modal template to perform data conversion, and then utilize the partial data in the converted data and the loss function to calculate the loss value of the medical large language model, use the backpropagation algorithm to calculate the gradient of the medical large language model, and then use the low-rank adaptation fine-tuning method to update the model parameters of other linear layers except the lm_head layer of the medical large language model until the preset number of iterative training times is completed or the calculated loss value of the medical large language model no longer decreases, and complete the training of the medical large language model to obtain a trained medical large language model;
[0028] Among them, represents the visual features of the images of the data in the model sample training set, represents the text descriptions of the images of the data in the model sample training set, represents the model role and capabilities of the medical large language model, represents the user label, represents the report generation requirements input by the user to the medical large language model, represents the assistant label, represents the stop marker indicating the end of the medical large language model's response to the report generation requirements input by the user to the medical large language model.
[0029] Furthermore, in step S3, retrieving the medical reports corresponding to the X-ray medical images similar to the X-ray medical image in the pre-set image memory storage and the medical report files to be queried according to the global features of the encoded X-ray medical images includes:
[0030] Normalize the global features of the X-ray medical images encoded by the medical visual encoder to obtain the global feature vectors of the X-ray medical images;
[0031] Calculate the square of the Euclidean distance between the processed global feature vectors of the X-ray medical images and each feature vector in the pre-set image memory storage, and sort each feature vector in the image memory storage in ascending order according to the calculated square of the Euclidean distance. Select the feature vector in the image memory storage with the smallest square of the Euclidean distance calculated from the global feature vectors of the X-ray medical images, and denote it as the feature vector ; the historical X-ray medical image corresponding to the selected feature vector in the image memory storage is the X-ray medical image similar to the X-ray medical image;
[0032] Find the medical report corresponding to the feature vector in the pre-set medical report files to be queried. The found medical report is the medical report of the X-ray medical image similar to the X-ray medical image.
[0033] Furthermore, the construction method of the pre-set image memory storage includes:
[0034] Collect historical X-ray medical images;
[0035] Use the medical visual encoder to encode the collected historical X-ray medical images to obtain the global features of each historical X-ray medical image;
[0036] Normalize the global features of each obtained historical X-ray medical image to obtain the global feature vector of each historical X-ray medical image;
[0037] Set an index number for the global feature vector of each obtained historical X-ray medical image;
[0038] Store the global feature vector of each historical X-ray medical image and its corresponding index number in a pre-set memory to obtain a pre-set image memory;
[0039] The construction method of the pre-set medical report file to be queried includes:
[0040] Collect the medical reports corresponding to the historical X-ray medical images and store them and the index numbers of the global feature vectors of their corresponding historical X-ray medical images in a pre-set file to obtain a pre-set medical report file to be queried.
[0041] In a second aspect, the present invention provides a Chinese medical report generation system based on a pre-trained large model. The system includes an image acquisition module, an instruction acquisition module, and a report generation module;
[0042] The image acquisition module is used to acquire the X-ray medical image for which a report is to be generated;
[0043] The instruction acquisition module is used to acquire the Chinese medical report generation requirements input by the user;
[0044] The report generation module is used to import the acquired X-ray medical image and the Chinese medical report generation requirements input by the user into a pre-trained large model to obtain the Chinese medical report of the X-ray medical image; the large model includes a medical vision encoder, a medical large language model, and a similar report retrieval module;
[0045] The medical vision encoder is used to encode the input X-ray medical image to obtain the visual features of the X-ray medical image, and the visual features include global features and local features;
[0046] The similar report retrieval module is used to retrieve the medical reports corresponding to the X-ray medical images similar to the input X-ray medical image in a pre-set image memory and a medical report file to be queried according to the global features of the X-ray medical image obtained by encoding, and input the retrieved medical reports into the medical large language model; the image memory stores the global feature vectors of historical X-ray medical images and their corresponding index numbers; the medical report file to be queried stores the medical reports corresponding to the historical X-ray medical images and the index numbers of the global feature vectors of the historical X-ray medical images;
[0047] The medical large language model is used to generate a Chinese medical report for the X-ray medical image according to the visual features of the X-ray medical image obtained by the medical visual encoder and according to the generation requirements of the input Chinese medical report and the retrieved medical report.
[0048] Further, in the report generation module, retrieve the medical report corresponding to the X-ray medical image similar to the X-ray medical image in the pre-set image memory storage and the medical report files to be queried according to the global features of the X-ray medical image obtained by encoding, including:
[0049] Normalize the global features of the X-ray medical image encoded by the medical visual encoder to obtain the global feature vector of the X-ray medical image;
[0050] Calculate the square of the Euclidean distance between the processed global feature vector of the X-ray medical image and each feature vector in the pre-set image memory storage, and sort each feature vector in the image memory storage in ascending order according to the calculated square of the Euclidean distance, and select the feature vector in the image memory storage with the smallest square of the Euclidean distance calculated with the global feature vector of the X-ray medical image, denoted as the feature vector ; the selected feature vector in the image memory storage The corresponding historical X-ray medical image is the X-ray medical image similar to the X-ray medical image;
[0051] Find the medical report corresponding to the feature vector in the pre-set medical report files to be queried, and the found medical report is the medical report of the X-ray medical image similar to the X-ray medical image.
[0052] Further, the construction method of the pre-set image memory storage includes:
[0053] Collect historical X-ray medical images;
[0054] Use the medical visual encoder to encode the collected historical X-ray medical images to obtain the global features of each historical X-ray medical image;
[0055] Normalize the global features of each obtained historical X-ray medical image to obtain the global feature vector of each historical X-ray medical image;
[0056] Set an index number for the global feature vector of each obtained historical X-ray medical image;
[0057] Store the global feature vector of each historical X-ray medical image and its corresponding index number into the pre-set memory storage to obtain the pre-set image memory storage;
[0058] The construction method of the pre-set medical report file to be queried includes:
[0059] Collect the medical reports corresponding to the historical X-ray medical images, and store their index numbers of the global feature vectors corresponding to the historical X-ray medical images in a pre-set file to obtain the pre-set medical report file to be queried.
[0060] In a third aspect, the present invention provides a terminal, which includes a memory and a processor;
[0061] The memory is used to store the Chinese medical report automatic generation program;
[0062] The processor is used to implement the Chinese medical report generation method based on the pre-trained large model according to any one of the first aspects when executing the Chinese medical report automatic generation program.
[0063] From the above technical solutions, it can be seen that the present invention has the following advantages:
[0064] The present invention constructs a large model, sets similar report retrieval, and through the image feature processing technology that combines the medical large language model and the medical visual encoder, realizes the accurate matching of medical images and their corresponding reports. The medical reports obtained through retrieval optimize the performance of the medical large language model in the inference stage, further reduce the hallucination problem generated by the medical large language model, improve the accuracy of the generated medical reports in lesion description and diagnostic suggestions, and effectively improve the quality and efficiency of medical reports.
[0065] The present invention extracts the global features and local features of X-ray medical images through the medical visual encoder in the constructed large model, improving the accuracy of lesion feature recognition. Secondly, the medical large language model in the large model constructed by the present invention, on the basis of the Qwen-7b model, performs text-image alignment training and instruction fine-tuning through the constructed mixed data sample training set, improving the accuracy, robustness and stability of the large model's understanding of the visual features of X-ray medical images.
[0066] In addition, the present invention constructs a mixed data sample training set containing rich text-image correspondence relationships, and performs text-image alignment training on the set projection layer, which not only enhances the projection layer's understanding ability of the visual features of X-ray images, but also promotes the deep integration between the medical large language model of image information and text description, enabling the generated medical reports to more accurately reflect the lesion conditions in the images.
[0067] Moreover, the present invention constructs a mixed data sample training set containing rich graphic-text correspondence relationships to perform instruction fine-tuning training on a medical large language model, guiding the model to learn how to generate standardized, accurate, and logical medical reports based on image features, including detailed lesion descriptions, precise diagnostic suggestions, etc. In addition, the rank adaptation fine-tuning method is adopted in the present invention to perform efficient parameter fine-tuning on the medical large language model, which not only effectively reduces the number of parameters to be fine-tuned and significantly reduces the consumption of computing resources of the medical large language model, but also ensures that the fine-tuned medical large language model can effectively adapt to downstream tasks. This process effectively alleviates the hallucination problem (i.e., generating text inconsistent with the image content) that may occur when the model generates reports, and improves the credibility and practicality of the reports.
[0068] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Figure 1 It is a schematic flowchart of an embodiment of the method for generating a Chinese medical report based on a pre-trained large model according to the present invention;
[0071] Figure 2 It is a schematic block diagram of an embodiment of the system for generating a Chinese medical report based on a pre-trained large model according to the present invention;
[0072] Figure 3 It is a schematic structural diagram of an embodiment of the terminal according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0075] Such asFigure 1 As shown in the figure, the present invention provides a method for generating a Chinese medical report based on a pre-trained large model, and the method includes:
[0076] S1. Obtain the X-ray medical image of the report to be generated;
[0077] S2. Obtain the requirements for generating a Chinese medical report input by the user;
[0078] S3. Import the obtained X-ray medical image and the requirements for generating a Chinese medical report input by the user into a pre-trained large model to obtain the Chinese medical report of the X-ray medical image; the large model includes a medical vision encoder, a medical large language model, and a similar report retrieval module;
[0079] The medical vision encoder is used to encode the input X-ray medical image to obtain the visual features of the X-ray medical image, and the visual features include global features and local features;
[0080] The similar report retrieval module is used to retrieve the medical reports corresponding to the X-ray medical images similar to the input X-ray medical image in a pre-set image memory and a medical report file to be queried according to the global features of the X-ray medical image obtained by encoding, and input the retrieved medical reports into the medical large language model; the global feature vectors of historical X-ray medical images and their corresponding index numbers are stored in the image memory; the medical reports corresponding to historical X-ray medical images and the index numbers of the global feature vectors of historical X-ray medical images are stored in the medical report file to be queried;
[0081] The medical large language model is used to generate the Chinese medical report of the X-ray medical image according to the visual features of the X-ray medical image obtained by the medical vision encoder, and according to the requirements for generating a Chinese medical report input and the retrieved medical reports.
[0082] For the convenience of understanding the present invention, the principle of the method for generating a Chinese medical report based on a pre-trained large model of the present invention is hereinafter further described in combination with the process of the method for generating a Chinese medical report based on a pre-trained large model in the embodiments.
[0083] Specifically, the method for generating a Chinese medical report based on a pre-trained large model includes:
[0084] Step 110. Obtain the X-ray medical image of the report to be generated.
[0085] Specifically, the X-ray medical images taken by the doctor are initially judged. It is checked whether the taken X-ray medical images are clear and whether they contain the required diagnostic information. If the quality of the taken X-ray medical images meets the requirements, the doctor will upload the taken X-ray medical images to the hospital's public image storage system. Download the X-ray medical images for which a medical report needs to be generated from the public image storage system.
[0086] Step 120: Obtain the Chinese medical report generation requirements input by the user.
[0087] Step 130: Import the obtained X-ray medical images and the Chinese medical report generation requirements input by the user into a pre-trained large model to obtain the Chinese medical report of the X-ray medical images; the large model includes a medical vision encoder, a medical large language model, and a similar report retrieval module;
[0088] The medical vision encoder is used to encode the input X-ray medical images to obtain the visual features of the X-ray medical images, and the visual features include global features and local features;
[0089] The similar report retrieval module is used to retrieve the medical reports corresponding to the X-ray medical images similar to the input X-ray medical images in a pre-set image memory storage and the medical report files to be queried according to the global features of the X-ray medical images obtained by encoding, and input the retrieved medical reports into the medical large language model; the global feature vectors of historical X-ray medical images and their corresponding index numbers are stored in the image memory storage; the medical reports corresponding to historical X-ray medical images and the index numbers of the global feature vectors of historical X-ray medical images are stored in the medical report files to be queried;
[0090] The medical large language model is used to generate the Chinese medical report of the X-ray medical images according to the visual features of the X-ray medical images obtained by the medical vision encoder, and according to the input Chinese medical report generation requirements and the retrieved medical reports.
[0091] It should be noted that after inputting the Chinese medical report generation requirements and the retrieved medical reports into the medical large language model, it is necessary to first convert the Chinese medical report generation requirements and the retrieved medical reports into instruction embedding extraction and similar report embedding extraction through the tokenizer in the medical large language model, and use the converted instruction embedding and similar report embedding for the trained medical large language model to generate the Chinese medical report of the X-ray medical images.
[0092] Specifically, the medical vision encoder is a vision Transformer model encoder using the ViT architecture, and in this embodiment, the vision encoder of BiomedCLIP is adopted. The medical large language model uses a pre-trained large language model (LLM) with rich medical knowledge and strong Chinese processing capabilities. In this embodiment, Taiyi-LLM is used as the medical LLM component. Taiyi-LLM, that is, the "Taiyi" biomedical large model, is a Chinese-English bilingual biomedical large model based on multi-task instruction fine-tuning developed by the Information Retrieval Research Laboratory of the School of Computer Science, Dalian University of Technology, with Qwen-7b as the base.
[0093] Specifically, the medical vision encoder is used to encode the input X-ray medical image to obtain the visual features of the X-ray medical image. The visual features include global features and local features, including:
[0094] Segment the input X-ray medical image into N pre-set non-overlapping image patches of a fixed size. After segmentation, each image patch is a one-dimensional vector, obtaining an image patch sequence ; It should be noted that the segmented image patches are squares with equal length and width.
[0095] Pass the image patch sequence through the fully connected layer in the medical vision encoder to map the image patch sequence to a vector space with a dimension of to obtain an image patch embedding sequence ;
[0096] Add an image classification embedding to the head of the image patch embedding sequence , and then add a position encoding to each image patch in the image patch embedding sequence with the image classification embedding added to the head added to the head, obtaining the image patch encoding sequence of the X-ray medical image ; ; ;
[0097] Import the image patch encoding sequence into the first layer of the Transformer encoder in the medical vision encoder, take the output of each layer as the input of the next layer of the Transformer encoder, and perform feature extraction layer by layer. After encoding by the last layer of the Transformer encoder, output the visual features of the X-ray medical image ;
[0098] Among them, represents the X-ray medical image, represents the An image patch, , represents the height of the th image patch, represents the width of the th image patch, , represents the number of image patches obtained by segmenting the X-ray medical image, represents the said fully connected layer, represents the dimension of the said vector space, represents the image classification embedding, represents the image position encoding, represents the image patch encoding sequence; represents the visual features of the X-ray medical image, represents the global features of the X-ray medical image, represents the local features of the X-ray medical image; represents the Transformer encoder.
[0099] Specifically, in this paper, the visual encoder of BiomedCLIP is selected to extract the visual features of X-ray images. BiomedCLIP is a vision-language model in the biomedical field trained on the PMC-15M dataset. Since its visual encoder already has sufficient ability to extract medical image features, the model weights are frozen throughout the training process.
[0100] As an embodiment of the present invention, before the medical large language model generates the Chinese medical report of the X-ray medical image based on the visual features of the X-ray medical image obtained from the medical visual encoder and according to the input Chinese medical report generation requirements and the retrieved medical reports, it further includes: performing visual embedding processing on the visual features of the X-ray medical image:
[0101] Through the trained projection layer map the local features of the input X-ray medical image to the same dimension as the word embedding space of the medical large language model to obtain the visual embedding of the X-ray medical image , .
[0102] Specifically, the training method of the projection layer includes:
[0103] Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and the obtained general data to obtain a projection layer sample training set;
[0104] Image-text alignment training: Use the data in the projection layer sample training set with a multi-modal template for data conversion, and use the converted data to iteratively train the projection layer Calculate the loss value of the projection layer using part of the data and the loss function in the converted mixed data Then use the backpropagation algorithm to calculate the gradient of the projection layer parameters, and use the gradient descent method to update the projection layer parameters until the preset number of iterative training times is completed or the calculated loss value of the projection layer no longer decreases, and complete the projection layer training to obtain a trained projection layer ; ;
[0105] Among them, represents the visual features of the image of the data in the projection layer sample training set, represents the text description of the image of the data in the projection layer sample training set.
[0106] It should be noted that the projection layer is an MLP neural network, and the projection layer is used to project the local features of X-ray medical images onto the same dimension as the word embedding space of the medical large language model, so that the local features of X-ray medical images can be used to guide the generation of Chinese medical reports for this X-ray medical image.
[0107] Specifically, the MIMIC-CXR dataset is the largest publicly available dataset in the field of medical report generation. The Vision-FLAN dataset is a dataset that integrates 191 Visual Question Answering (VQA) tasks from 101 open-source datasets. In the Vision-FLAN dataset, each visual question answering task is accompanied by instructions written by experts. The Vision-FLAN dataset contains VFLAN-Caption data and VFLAN-Instruct data. VFLAN-Caption data represents an image caption dataset, and VFLAN-Instruct data represents a visual question answering task instruction dataset.
[0108] The MIMIC-CXR dataset is in English, and the medical reports in the MIMIC-CXR dataset are unstructured data. The MIMIC-CXR dataset used in this example is obtained from the XrayGPT project, which processes the unstructured medical reports through the GPT3.5 model to generate a concise and coherent medical abstract that is more in line with human language habits. In addition, in this example, the MIMIC-CXR data of the XrayGPT project is translated into Chinese using Tencent Translate service to obtain the Chinese MIMIC-CXR dataset for the research on Chinese medical report generation tasks.
[0109] The Vision-FLAN dataset is the Chinese version obtained from the ALLaVA project. In this example, the data in the obtained Vision-FLAN dataset is preprocessed to obtain VFLAN-Caption data and VFLAN-Instruct data, specifically including: deleting the data in the Vision-FLAN dataset that does not conform to the model training format, the data with garbled characters, and the data with meaningless symbols.
[0110] Exemplarily, a method for constructing a medical large language model includes:
[0111] Create a Qwen-7b model as the initial model and perform model initialization.
[0112] Exemplarily, a method for training a medical large language model includes:
[0113] Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and the obtained general data to obtain a model sample training set;
[0114] Instruction fine-tuning training: After importing the data in the model sample training set into the initial model, use a multimodal template to perform data conversion on the imported data, and then use part of the data and the loss function in the converted data to calculate the loss value of the medical large language model, use the backpropagation algorithm to calculate the gradient of the medical large language model, and then use the low-rank adaptation fine-tuning method to update the model parameters of other linear layers except the lm_head layer of the medical large language model until the preset number of iterative training times is completed or the calculated loss value of the medical large language model no longer decreases, complete the training of the medical large language model, and obtain a trained medical large language model;
[0115] Among them, represents the visual features of the images of the data in the model sample training set, represents the text description of the images of the data in the model sample training set, represents the model role and capabilities of the medical large language model, represents the user label, represents the report generation requirements input by the user to the medical large language model, represents the assistant label, represents the stop token indicating the end of the medical large language model's answer to the report generation requirements input by the user to the medical large language model.
[0116] Specifically, the principle of the low-rank adaptation fine-tuning method is as follows:
[0117] Assume that the weight matrix of a certain layer in the model is , and the low-rank adaptation fine-tuning method approximately represents the incremental parameter matrix to be fine-tuned as two matrices and with fewer parameters, and the parameter update process can be regarded as ;
[0118] At this time, the number of parameters for full-parameter fine-tuning changes from the original to becomes , the number of parameters to be updated will be significantly reduced. During fine-tuning, keep frozen, and initialize the matrix with a random Gaussian distribution and initialize the matrix with a zero matrix This can ensure that gradually learns new knowledge during training to ensure the stability of the entire process. Therefore, given the input , for the output , its forward calculation becomes . During inference, can be directly added to the original parameter without introducing additional computational latency. If it is necessary to switch the additional matrix to adapt to different application scenarios, it can also be simply achieved by subtracting and then adding a different .
[0119] Among them, and are the dimensions of the pre-trained model weights, is the rank of the incremental matrix, and .
[0120] Specifically, during the training process of the medical large language model, the model sample training set includes X-ray medical images, their corresponding Chinese medical reports, and their corresponding Chinese medical report generation requirements.
[0121] Exemplarily, the construction method of the pre-set image memory storage includes:
[0122] Collect historical X-ray medical images;
[0123] Use a medical visual encoder to encode the collected historical X-ray medical images to obtain the global features of each historical X-ray medical image;
[0124] Perform normalization processing on the global features of each obtained historical X-ray medical image to obtain the global feature vector of each historical X-ray medical image;
[0125] Set an index number for the global feature vector of each obtained historical X-ray medical image to obtain the feature vector to be queried, and gather all the obtained feature vectors to be queried to obtain a feature vector sequence , , and store the global feature vector of each historical X-ray medical image and its corresponding index number, that is, the feature vector sequence into the pre-set memory to obtain the pre-set image memory storage.
[0126] The construction method of the pre-set medical report file to be queried includes:
[0127] Collect the medical reports corresponding to the historical X-ray medical images, and pool them with the index numbers of the global feature vectors of their corresponding historical X-ray medical images to obtain a medical report sequence , , store the medical report and the index number of the global feature vector of its corresponding historical X-ray medical image, that is, the medical report sequence in a pre-set file to obtain a pre-set medical report file to be queried.
[0128] Among them, represents the feature vector to be queried with the number , represents the medical report with the number , represents the number of historical X-ray medical images obtained.
[0129] Specifically, in step 130, retrieve the medical reports corresponding to the X-ray medical images similar to the X-ray medical image in the pre-set image memory and the medical report file to be queried according to the global features of the encoded X-ray medical images, including:
[0130] Normalize the global features of the X-ray medical images encoded by the medical visual encoder to obtain the global feature vector of the X-ray medical images;
[0131] Calculate the square of the Euclidean distance between the processed global feature vector of the X-ray medical image and each feature vector in the pre-set image memory, and sort each feature vector in the image memory in ascending order of the calculated square of the Euclidean distance. Select the feature vector in the image memory with the smallest square of the Euclidean distance calculated from the global feature vector of the X-ray medical image, and denote it as the feature vector ; the historical X-ray medical image corresponding to the selected feature vector in the image memory is the X-ray medical image similar to the X-ray medical image;
[0132] Find the medical report corresponding to the feature vector in the pre-set medical report file to be queried. The found medical report is the medical report of the X-ray medical image similar to the X-ray medical image.
[0133] Specifically, in this embodiment, the IndexFlatL2 structure of the Faiss library is used to retrieve the feature vectors of the X-ray images to be queried in the image memory according to the processed feature vectors. The Euclidean distance method is adopted. According to the formula , calculate the square of the Euclidean distance between the global feature vector of the processed X-ray medical image and each feature vector in the pre-set image memory, and sort each feature vector in the image memory in ascending order according to the calculated square of the Euclidean distance. Since the smaller the square of the Euclidean distance between two vectors, the greater the similarity between the images corresponding to the two vectors, so select the feature vector in the image memory with the smallest square of the Euclidean distance calculated with the global feature vector of the X-ray medical image, and denote it as the feature vector ; The selected feature vector The corresponding X-ray medical image is the X-ray medical image similar to the said X-ray medical image. According to the selected feature vector 's index number, retrieve the pre-set medical report file to be queried, and obtain the medical report with the same index number as the feature vector to be queried . The retrieved medical report is the medical report of the X-ray medical image similar to the said X-ray medical image, and send the obtained medical report to the medical large language model.
[0134] As an embodiment of the present invention, as Figure 2 shown, the present invention also provides a Chinese medical report generation system based on a pre-trained large model. The system 200 includes: an image acquisition module 210, an instruction acquisition module 220, and a report generation module 230;
[0135] The image acquisition module 210 is used to acquire the X-ray medical image of the report to be generated;
[0136] The instruction acquisition module 220 is used to acquire the Chinese medical report generation requirements input by the user;
[0137] The report generation module 230 is used to import the acquired X-ray medical image and the Chinese medical report generation requirements input by the user into the pre-trained large model to obtain the Chinese medical report of the X-ray medical image; The large model includes a medical vision encoder 231, a medical large language model 232, and a similar report retrieval module 233;
[0138] The medical vision encoder 231 is used to encode the input X-ray medical image to obtain the visual features of the X-ray medical image, and the visual features include global features and local features;
[0139] The similar report retrieval module 233 is used to retrieve the medical reports corresponding to the X-ray medical images similar to the X-ray medical image according to the global features of the X-ray medical image obtained by encoding in the pre-set image memory and the medical report files to be queried, and input the retrieved medical reports into the medical large language model 232; the image memory stores the global feature vectors of historical X-ray medical images and their corresponding index numbers; the medical report files to be queried store the medical reports corresponding to the historical X-ray medical images and the index numbers of the global feature vectors of the historical X-ray medical images;
[0140] The medical large language model 232 is used to generate the Chinese medical report of the X-ray medical image according to the visual features of the X-ray medical image obtained by the medical vision encoder 231, and according to the requirements for generating the Chinese medical report input and the retrieved medical report.
[0141] Further, in the report generation module 230, the input X-ray medical image is encoded to obtain the visual features of the encoded X-ray medical image, including:
[0142] The input X-ray medical image is segmented into N pre-set non-overlapping image blocks of a fixed size to obtain an image block sequence ;
[0143] Through the fully connected layer in the medical vision encoder 231 Map the image block sequence To a vector space of dimension To obtain an image block embedding sequence ;
[0144] Add an image classification embedding to the head of the image block embedding sequence After that, for each image block in the image block embedding sequence with an image classification embedding added to the head Add a position encoding To obtain the image block encoding sequence of the X-ray medical image ; ;
[0145] Import the image block encoding sequence Into the first layer of the Transformer encoder in the medical vision encoder 231, use the output of each layer as the input of the next layer of the Transformer encoder, perform feature extraction layer by layer, and after encoding by the last layer of the Transformer encoder, output the visual features of the X-ray medical image ;
[0146] Among them, Represents the X-ray medical image, Indicates the th image patch obtained from the segmentation of X-ray medical images, , Indicates the height of the th image patch, Indicates the width of the th image patch, Indicates the number of channels of the th image patch, , Indicates the number of image patches obtained by segmenting the X-ray medical images, Indicates the fully connected layer mentioned above, Indicates the dimension of the vector space mentioned above, Indicates the image classification embedding, Indicates the image position encoding, Indicates the image patch encoding sequence; Indicates the visual features of the X-ray medical images, Indicates the global features of the X-ray medical images, Indicates the local features of the X-ray medical images; Indicates the Transformer encoder.
[0147] Furthermore, before the medical large language model 232 generates the Chinese medical report of the X-ray medical image based on the visual features of the X-ray medical image obtained by the medical vision encoder 231 and according to the input requirements for generating Chinese medical reports and the retrieved medical reports, it further includes: performing visual embedding processing on the visual features of the X-ray medical image:
[0148] Through the trained projection layer map the input local features of the X-ray medical image to the same dimension as the word embedding space of the medical large language model 232 to obtain the visual embedding of the X-ray medical image , ;
[0149] The training method of the projection layer includes:
[0150] Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and the obtained general data to obtain the sample training set for the projection layer;
[0151] Text and image alignment training: Use the data in the projection layer sample training set with a multi-modal template for data conversion, and use the converted data to iteratively train the projection layer Calculate the loss value of the projection layer using part of the data and the loss function in the converted mixed data Calculate the gradient of the projection layer parameters using the backpropagation algorithm, and update the projection layer parameters using the gradient descent method until the preset number of iterative training times is completed or the calculated loss value of the projection layer no longer decreases, and complete the training of the projection layer to obtain a trained projection layer ; ; ;
[0152] Among them, represents the visual features of the images of the data in the projection layer sample training set, represents the text description of the images of the data in the projection layer sample training set.
[0153] Furthermore, a method for constructing a medical large language model 232 includes:
[0154] Create a Qwen-7b model as the initial model and perform model initialization;
[0155] A training method for a medical large language model 232 includes:
[0156] Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and the obtained general data to obtain a model sample training set;
[0157] Instruction fine-tuning training: After importing the data in the model sample training set into the initial model, use the imported data with a multi-modal template for data conversion, and then use the Calculate the loss value of the medical large language model 232 for part of the data and the loss function, calculate the gradient of the medical large language model 232 using the backpropagation algorithm, and then use the low-rank adaptation fine-tuning method to update the model parameters of other linear layers except the lm_head layer of the medical large language model 232 until the preset number of iterative training times is completed or the calculated loss value of the medical large language model 232 no longer decreases, completing the training of the medical large language model 232 and obtaining the trained medical large language model 232;
[0158] Among them, represents the visual feature of the image of the data in the model sample training set, represents the text description of the image of the data in the model sample training set, represents the model role and the capabilities of the medical large language model, represents the user label, represents the report generation requirement input by the user to the medical large language model, represents the assistant label, represents the stop mark indicating the end of the medical large language model's response to the report generation requirement input by the user to the medical large language model.
[0159] Exemplarily, in the report generation module 230, retrieving the medical report corresponding to the X-ray medical image similar to the encoded X-ray medical image from the pre-set image memory and the medical report file to be queried includes:
[0160] Normalize the global feature of the X-ray medical image encoded by the medical visual encoder 231 to obtain the global feature vector of the X-ray medical image;
[0161] Calculate the square of the Euclidean distance between the processed global feature vector of the X-ray medical image and each feature vector in the pre-set image memory, and sort each feature vector in the image memory in ascending order according to the calculated square of the Euclidean distance. Since the smaller the square of the Euclidean distance between two vectors, the greater the similarity between the images corresponding to the two vectors, so select the feature vector in the image memory with the smallest square of the Euclidean distance calculated with the global feature vector of the X-ray medical image, denoted as feature vector ; the selected feature vector in the image memory The corresponding historical X-ray medical image is the X-ray medical image similar to the X-ray medical image;
[0162] Search for the feature vector in the pre-set medical report file to be queried The corresponding medical report, and the found medical report is the medical report of the X-ray medical image similar to the X-ray medical image mentioned above.
[0163] Exemplarily, the method for constructing the pre-set image memory includes:
[0164] Collect historical X-ray medical images;
[0165] Use a medical vision encoder 231 to encode the collected historical X-ray medical images to obtain the global features of each historical X-ray medical image;
[0166] Perform normalization processing on the global features of each obtained historical X-ray medical image to obtain the global feature vector of each historical X-ray medical image;
[0167] Set an index number for the global feature vector of each obtained historical X-ray medical image;
[0168] Store the global feature vector of each historical X-ray medical image and its corresponding index number into a pre-set memory to obtain the pre-set image memory.
[0169] Exemplarily, the method for constructing the pre-set medical report file to be queried includes:
[0170] Collect the medical reports corresponding to the historical X-ray medical images, and store them and the index numbers of the global feature vectors of their corresponding historical X-ray medical images into a pre-set file to obtain the pre-set medical report file to be queried.
[0171] Figure 3 FIG. 28 is a schematic structural diagram of a terminal 300 provided by an embodiment of the present invention. The terminal 300 can be used to execute the Chinese medical report generation system based on a pre-trained large model provided by the embodiment of the present invention.
[0172] Among them, the terminal 300 may include: a processor 310, a memory 320, and a communication module 330. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0173] Among them, the memory 320 can be used to store the execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. When the execution instructions in the memory 320 are executed by the processor 310, the terminal 300 can execute some or all of the steps in the above system embodiments.
[0174] The processor 310 is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 320, and calling data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor can be composed of an integrated circuit (Integrated Circuit, abbreviated as IC). For example, it can be composed of a single packaged IC, or composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 310 may only include a central processing unit (Central Processing Unit, abbreviated as CPU). In the embodiment of the present invention, the CPU can be a single arithmetic core or include multiple arithmetic cores.
[0175] The communication module 330 is used to establish a communication channel, so that the storage terminal can communicate with other terminals. Receive user data sent by other terminals or send user data to other terminals.
[0176] The present invention constructs a large model and sets up similar report retrieval. By combining the image feature processing technology of a medical large language model and a medical vision encoder, it realizes the accurate matching of medical images and their corresponding reports. The medical reports obtained through retrieval optimize the performance of the medical large language model in the inference stage, further reduce the hallucination problem generated by the medical large language model, improve the accuracy of the generated medical reports in lesion description and diagnostic suggestions, and effectively improve the quality and efficiency of medical reports.
[0177] The present invention extracts the global features and local features of X-ray medical images through the medical vision encoder in the constructed large model, improving the accuracy of lesion feature recognition. Secondly, based on the Qwen-7b model, the medical large language model in the large model constructed by the present invention is trained for image-text alignment and instruction fine-tuning through the constructed mixed data sample training set, improving the accuracy, robustness and stability of the large model's understanding of the visual features of X-ray medical images.
[0178] In addition, the present invention constructs a mixed data sample training set containing rich graphic-text correspondence relationships to perform graphic-text alignment training on the set projection layer, which not only enhances the projection layer's ability to understand the visual features of X-ray images, but also promotes the in-depth integration of the medical large language model between image information and text descriptions, enabling the generated medical reports to more accurately reflect the pathological conditions in the images.
[0179] Moreover, the present invention constructs a mixed data sample training set containing rich graphic-text correspondence relationships to perform instruction fine-tuning training on the medical large language model, guiding the model to learn how to generate standardized, accurate, and logical medical reports based on image features, including detailed pathological descriptions, precise diagnostic suggestions, etc. And in the present invention, the rank adaptation fine-tuning method is used to perform efficient parameter fine-tuning on the medical large language model, which not only effectively reduces the number of parameters for fine-tuning and significantly reduces the consumption of computing resources of the medical large language model, but also ensures that the fine-tuned medical large language model can effectively adapt to downstream tasks. This process effectively alleviates the hallucination problem (i.e., generating text inconsistent with the image content) that may occur when the model generates reports, improving the credibility and practicality of the reports.
[0180] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect.
[0181] The technical effects that can be achieved by this embodiment can be referred to the descriptions above, and will not be elaborated here.
[0182] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store program codes, and includes several instructions to enable a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the system described in each embodiment of the present invention.
[0183] For the same or similar parts between the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the system embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the system embodiments.
[0184] In several embodiments provided by the present invention, it should be understood that the disclosed systems and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or module can be in an electrical, mechanical or other form.
[0185] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0186] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0187] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person familiar with the technical field of the present invention can easily think of changes or substitutions within the technical scope disclosed by the present invention, and they should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A Chinese medical report generation method based on a pre-trained large model, characterized in that S1. Obtain the X-ray medical image of the report to be generated; S2. Obtain the Chinese medical report generation requirements input by the user; S3. Import the obtained X-ray medical image and the Chinese medical report generation requirements into a pre-trained large model to obtain the Chinese medical report of the X-ray medical image; the large model includes a medical vision encoder, a medical large language model, and a similar report retrieval module; The medical vision encoder is used to encode the input X-ray medical image to obtain the visual features of the X-ray medical image, and the visual features include global features and local features; The similar report retrieval module is used to retrieve the medical reports corresponding to the X-ray medical images similar to the X-ray medical image in a pre-set image memory and the medical report files to be queried according to the global features of the X-ray medical image obtained by encoding, and input the retrieved medical reports into the medical large language model; the global feature vectors of historical X-ray medical images and their corresponding index numbers are stored in the image memory; the medical reports corresponding to historical X-ray medical images and the index numbers of the global feature vectors of historical X-ray medical images are stored in the medical report files to be queried; The medical large language model is used to generate the Chinese medical report of the X-ray medical image according to the visual features of the X-ray medical image obtained by the medical vision encoder, and according to the input Chinese medical report generation requirements and the retrieved medical reports; The construction method of the medical large language model includes: Create a Qwen-7b model as the initial model and perform model initialization; The training method of the medical large language model includes: Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and the obtained general data to obtain the model sample training set; Instruction fine-tuning training: After importing the data in the model sample training set into the initial model, the imported data is converted using a multimodal template, and then the partial data in the converted data and the loss function are used to calculate the loss value of the medical large language model. The gradient of the medical large language model is calculated using the backpropagation algorithm, and then the low-rank adaptation fine-tuning method is used to update the model parameters of other linear layers except the lm_head layer of the medical large language model until the preset number of iterative training times is completed or the calculated loss value of the medical large language model no longer decreases, completing the training of the medical large language model and obtaining the trained medical large language model; Among them, represents the visual features of the images of the data in the model sample training set, represents the text description of the images of the data in the model sample training set, represents the model role and capabilities of the medical large language model, represents the user label, represents the report generation requirements input by the user to the medical large language model, represents the assistant label, represents the stop marker indicating the end of the medical large language model's response to the report generation requirements input by the user to the medical large language model.
2. The method for generating a Chinese medical report based on a pre-trained large model according to claim 1, wherein Encode the input X-ray medical image to obtain the visual features of the X-ray medical image, including; Segment the input X-ray medical image into N non-overlapping image patches of a preset fixed size to obtain an image patch sequence ; The sequence of image patches is mapped to a vector space of dimension through the fully connected layer in the medical vision encoder, obtaining a sequence of image patch embeddings ; Add the image classification embedding to the head of the image patch embedding sequence , and then add the image classification embedding to the head For each image patch in the image patch embedding sequence with the image classification embedding added to the head Add positional encoding , to obtain the image patch encoding sequence of the X-ray medical image ; Encode the image block sequence Import it into the first layer of the Transformer encoder in the medical vision encoder. Use the output of each layer as the input of the next layer of the Transformer encoder, and perform feature extraction layer by layer. After encoding by the last layer of the Transformer encoder, output the visual features of the X-ray medical image ; Among them, represents the X-ray medical image, represents the th image patch obtained by segmenting the acquired X-ray medical image, , represents the th height of the image patch, represents the th width of the image patch, represents the th number of channels of the image patch, , represents the number of image patches obtained by segmenting the acquired X-ray medical image, represents the said fully connected layer, represents the dimension of the said vector space, represents the image classification embedding, represents the image position encoding, represents the image patch encoding sequence; represents the visual feature of the X-ray medical image, represents the global feature of the X-ray medical image, represents the local feature of the X-ray medical image; represents the Transformer encoder.
3. The Chinese medical report generation method based on a pre-trained large model according to claim 1, wherein Before the medical large language model generates the Chinese medical report of the X-ray medical image according to the visual features of the X-ray medical image obtained by the medical vision encoder, and according to the input Chinese medical report generation requirements and the retrieved medical reports, it also includes: performing visual embedding processing on the visual features of the X-ray medical image: Through the trained projection layer Map the local features of the X-ray medical image to the same dimension as the word embedding space of the medical large language model to obtain the visual embedding of the X-ray medical image , ; The projection layer The training method includes: Construct a sample training set: Obtain the data in the MIMIC-CXR dataset as in-domain data, obtain the data in the Vision-FLAN dataset, preprocess the obtained data in the Vision-FLAN dataset to obtain VFLAN-Caption data and VFLAN-Instruct data, use the VFLAN-Caption data and VFLAN-Instruct data as general data, and pool the obtained in-domain data and general data to obtain a projection layer sample training set; Text-image alignment training: Use the data in the projection layer sample training set with a multi-modal template for data conversion, and use the converted data to iteratively train the projection layer Calculate the loss value of the projection layer using a part of the data and the loss function in the converted mixed data Then use the backpropagation algorithm to calculate the gradient of the projection layer parameters, and use the gradient descent method to update the projection layer parameters until the preset number of iterative training times is completed or the calculated loss value of the projection layer no longer decreases, completing the projection layer training to obtain a trained projection layer ; Among them, represents the visual features of the images of the data in the projection layer sample training set, represents the text description of the images of the data in the projection layer sample training set.
4. The Chinese medical report generation method based on a pre-trained large model according to claim 1, wherein, In step S3, retrieving the medical report corresponding to the X-ray medical image similar to the X-ray medical image based on the global feature of the encoded X-ray medical image in a preset image memory and a medical report file to be queried includes: Normalize the global feature of the X-ray medical image encoded by the medical vision encoder to obtain the global feature vector of the X-ray medical image; Calculate the square of the Euclidean distance between the global feature vector of the processed X-ray medical image and each feature vector in the pre-set image memory, and sort each feature vector in the image memory in ascending order according to the calculated square of the Euclidean distance. Select the feature vector in the image memory with the smallest square of the Euclidean distance calculated from the global feature vector of the X-ray medical image in the sorting, and denote it as the feature vector. ; The selected feature vector in the image memory The corresponding historical X-ray medical image is the X-ray medical image similar to the X-ray medical image; Search for the eigenvector in the pre-set medical report file to be queried The corresponding medical report, and the found medical report is the medical report of the X-ray medical image similar to the X-ray medical image 5. The Chinese medical report generation method based on a pre-trained large model according to claim 4, wherein The construction method of the preset image memory includes: Collect historical X-ray medical images; Use the medical vision encoder to encode the collected historical X-ray medical images to obtain the global feature of each historical X-ray medical image; Normalize the obtained global feature of each historical X-ray medical image to obtain the global feature vector of each historical X-ray medical image; Set an index number for the global feature vector of each obtained historical X-ray medical image; Store the global feature vector of each historical X-ray medical image and its corresponding index number in a preset memory to obtain a preset image memory; The construction method of the preset medical report file to be queried includes: Collect the medical reports corresponding to the historical X-ray medical images, and store them and the index numbers of the global feature vectors of their corresponding historical X-ray medical images in a preset file to obtain a preset medical report file to be queried.
6. A Chinese medical report generation system based on a pre-trained large model, characterized in that, The system includes an image acquisition module, an instruction acquisition module, and a report generation module; The image acquisition module is used to acquire the X-ray medical image for which a report is to be generated; The instruction acquisition module is used to acquire the Chinese medical report generation requirement input by the user; The report generation module is used to import the acquired X-ray medical image and the Chinese medical report generation requirement input by the user into a pre-trained large model to obtain the Chinese medical report of the X-ray medical image; the large model includes a medical vision encoder, a medical large language model, and a similar report retrieval module; The medical vision encoder is used to encode the input X-ray medical image to obtain the visual features of the X-ray medical image, and the visual features include global features and local features; The similar report retrieval module is used to retrieve the medical reports corresponding to the X-ray medical images similar to the X-ray medical image according to the global features of the X-ray medical image obtained by encoding in the pre-set image memory and the medical report files to be queried, and input the retrieved medical reports into the medical large language model; the image memory stores the global feature vectors of historical X-ray medical images and their corresponding index numbers; the medical report files to be queried store the medical reports corresponding to the historical X-ray medical images and the index numbers of the global feature vectors of the historical X-ray medical images; The medical large language model is used to generate the Chinese medical report of the X-ray medical image according to the visual features of the X-ray medical image obtained by the medical vision encoder, and according to the requirements for generating the Chinese medical report input and the retrieved medical report.
7. The Chinese medical report generation system based on a pre-trained large model according to claim 6, wherein In the report generation module, retrieving the medical reports corresponding to the X-ray medical images similar to the X-ray medical image according to the global features of the X-ray medical image obtained by encoding in the pre-set image memory and the medical report files to be queried includes: Normalize the global features of the X-ray medical image encoded by the medical vision encoder to obtain the global feature vector of the X-ray medical image; Calculate the square of the Euclidean distance between the global feature vector of the processed X-ray medical image and each feature vector in the pre-set image memory, and sort each feature vector in the image memory in ascending order according to the calculated square of the Euclidean distance. Select the feature vector in the image memory with the smallest square of the Euclidean distance calculated from the global feature vector of the X-ray medical image in the sorting, and denote it as the feature vector ; The selected feature vector in the image memory The corresponding historical X-ray medical image is the X-ray medical image similar to the X-ray medical image; Search for the feature vector in the pre-set medical report file to be queried The corresponding medical report, and the found medical report is the medical report of the X-ray medical image similar to the X-ray medical image described 8. The Chinese medical report generation system based on a pre-trained large model according to claim 7, wherein, The construction method of the pre-set image memory includes: Collect historical X-ray medical images; Use the medical vision encoder to encode the collected historical X-ray medical images to obtain the global features of each historical X-ray medical image; Normalize the global features of each obtained historical X-ray medical image to obtain the global feature vector of each historical X-ray medical image; Set an index number for the global feature vector of each obtained historical X-ray medical image; Store the global feature vector of each historical X-ray medical image and its corresponding index number into the pre-set memory to obtain the pre-set image memory; The construction method of the pre-set medical report file to be queried includes: Collect the medical reports corresponding to the historical X-ray medical images, and store them and the index numbers of the global feature vectors of their corresponding historical X-ray medical images into the pre-set file to obtain the pre-set medical report file to be queried.
9. A terminal, characterized in that, The terminal includes a memory and a processor; The memory is used to store the Chinese medical report generation program based on the pre-trained large model; The processor is used to implement the Chinese medical report generation method based on the pre-trained large model according to any one of claims 1-5 when executing the Chinese medical report generation program based on the pre-trained large model.
Citation Information
Patent Citations
Method and device for generating medical image report
CN117352121A
Diagnosis report generation system based on medical image and disease attribute description pair
CN118447994A