Medical image analysis method and device of visual language model, and storage medium
By constructing a data set containing medical images and medical imaging diagnostic reports, using visual language models to extract elements of disease entities embedded vectors and train the model, the problem of high false alarm rate in medical image analysis is solved, and higher accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510318097.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing universal visual language models are difficult to accurately identify pathological areas when processing complex medical images, resulting in high false positives, increasing medical burden and potentially harming patients.
The data set containing medical images and medical imaging diagnostic reports are constructed, and the element embedding vectors of disease entities are extracted using visual language models. Through cosine similarity calculation and maximum pooling processing, the visual language model is trained to improve the similarity matching of image and text features, constrain the cross-entropy loss of prediction mask layer and real mask layer, and ensure that the model focuses on disease-related areas.
It improves the accuracy of the model in medical image analysis, reduces the false alarm rate, improves the accuracy of image classification and retrieval, and reduces unnecessary medical procedures.
Smart Images

Figure CN120259695A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual language model medical image analysis, and particularly to a medical image analysis method, device and storage medium for a visual language model. Background Art
[0002] Visual language models have been applied to many tasks, such as image retrieval, captioning, and object localization. Among them, contrastive language-image pre-training models use natural language prompts to interpret visual data and show excellent performance. Contrastive language-image pre-training models can handle various tasks without the need for specialized fine-tuning. However, it has limitations when applied to highly specialized fields, such as medical imaging.
[0003] In the medical field, accurate anomaly detection is crucial for early diagnosis and treatment. However, general models like contrastive language-image pre-training models often struggle to process intricate medical images that contain subtle features unique to medical imaging, which are crucial for identifying pathological regions. To improve performance in the biomedical field, various variants of CLIP, such as BiomedCLIP and MedCLIP-SAMv2, have been proposed for the medical field to enhance performance in the biomedical field. Compared with ordinary models, they show a remarkable enhancement in medical reasoning, but the problem of false positives still prevails. These false positives can lead to unnecessary medical procedures, increase the burden on the healthcare system and may harm patients. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a medical image analysis method, device and storage medium for a visual language model.
[0005] In a first aspect, the present invention provides a medical image analysis method for a visual language model, including:
[0006] Using a medical imaging diagnosis report containing medical images and image diagnosis content and an image feature description library of disease types to construct a data set containing disease entities, images and image mask layers, wherein one image corresponds to at least one disease entity, and the disease entity includes the attributes of the disease, the onset location, the affected organ, the category of the disease and the image feature description of the disease entity;
[0007] Extracting non-repeating elements of all disease entities within the entire scope of the medical imaging diagnosis report as reference elements, and using a first text encoder to extract reference element embedding vectors of all reference elements;
[0008] For each image in the dataset, use the first text encoder to extract the element embedding vectors of each element in its disease entities, and calculate the cosine similarity between each element embedding vector and all reference element embedding vectors; perform max pooling on the cosine similarity matrix of all disease entities of any image, and project it into an image label with a dimension of 1×the number of reference elements;
[0009] Concatenate the image labels of all images to obtain a similarity matrix between all images and all reference elements in the medical image diagnosis report;
[0010] Construct and train a vision-language model, the vision-language model includes an image encoder, an image decoder, and a second text encoder to be trained. The image encoder extracts the initial image features of all images in the medical image diagnosis report. Each group of initial image features is input into the image decoder on the one hand, and the image decoder maps the initial image features into a predicted mask layer. Each group of initial image features is input into the mapping layer on the other hand, and the mapping layer obtains image features based on the initial image features; the second text encoder extracts text features from the disease entities, calculates the similarity between the text features and the image features, and obtains a predicted similarity matrix; during training, constrain the sum of the cross-entropy loss between the predicted mask layer and the true mask layer in the dataset, the cross-entropy loss and the mean square error loss between the predicted similarity matrix and the actual similarity matrix to be minimized;
[0011] During application, take the trained image encoder and the second text encoder to perform medical image analysis tasks. The medical image analysis tasks include: medical image classification from images to disease types, or retrieval matching of images according to disease descriptions.
[0012] Furthermore, the process of using the medical image diagnosis report containing medical images and image diagnosis content and the image feature description library of disease types to construct a dataset containing disease entities, images, and mask layers includes:
[0013] Use a pre-trained semantic segmentation model to extract the mask layer and the image description within the mask layer area from the image. The mask layer includes the mask layer of the abnormal area in the image;
[0014] Use the disease entity extraction prompt words to control the large language model to extract disease entities from the image diagnosis content and the image feature descriptions of disease types; the image feature description library of disease types contains the image feature descriptions of various diseases;
[0015] The disease entities obtained from the image diagnosis content include: the attributes of the disease, the onset location, the affected organ, and the category of the disease. The disease entities obtained from the image feature descriptions of disease types include the category of the disease, the disease image color description, and the disease impact form description;
[0016] Using the disease category to match the image feature description of the disease to the disease entity obtained from the image diagnosis content as a complete disease entity;
[0017] Combining the image, the disease entity corresponding to the image, and the mask layer as a dataset.
[0018] Furthermore, the disease entity extraction prompt words for controlling the large language model to extract disease entities from the image feature description include:
[0019] Limiting the large language model to use medical knowledge to extract disease-related information sentence by sentence from the image feature description as disease entities;
[0020] Limiting the format of each extracted disease entity: {disease category}{disease image color description}{disease image morphology description}, where when any image feature description lacks any information, a mask is used to replace the missing information;
[0021] Limiting the separator symbol between disease entities.
[0022] Furthermore, the disease entity extraction prompt words for controlling the large language model to extract disease entities from the image diagnosis content include:
[0023] Limiting the large language model to use medical knowledge to extract disease-related information sentence by sentence from the provided medical image diagnosis report as disease entities;
[0024] Limiting the format of each extracted disease entity: {disease attribute}{onset location}{affected organ}{disease category}, where the disease attribute is the severity or stage and scope of the disease; where when any image diagnosis content lacks any information, a mask is used to replace the missing information;
[0025] Limiting that when the negative situation of the disease is mentioned in the medical image diagnosis report, a negative word is added before the disease entity;
[0026] Limiting to extract all disease entities from each sentence of the image diagnosis content;
[0027] Limiting to ignore words unrelated to the disease description when extracting all disease entities;
[0028] Limiting the separator symbol between disease entities.
[0029] Furthermore, for large language models lacking capabilities in the medical field, before using them to extract disease entities, the large language models are fine-tuned. The fine-tuning dataset includes: prompt words, example image diagnostic content, example imaging feature descriptions, and example large language model responses. The cross-entropy between the predicted response and the actual response of the large language model is used as the fine-tuning loss function, and fine-tuning training is carried out with the aim of minimizing the fine-tuning loss function.
[0030] Furthermore, the image encoder of the vision language model uses a Resnet model, and the model structure of the image decoder is symmetric to that of the image encoder.
[0031] Furthermore, the second text encoder uses a CLIP text model; the image encoder extracts the initial image features of all images in the medical imaging diagnosis report. On the one hand, each group of initial image features is input into the image decoder, and the image decoder maps the initial image features into a predicted mask layer. On the other hand, each group of initial image features is input into the mapping layer, and the mapping layer obtains image features based on the initial image features; the disease entities corresponding to the images in the dataset are tokenized and padded through the CLIP word processor, converted into tokens, and then input into the second text encoder to extract the pooled output. Then, the pooled output is adjusted in dimension through the projection layer to obtain text features. After that, the image and text features are L2-normalized to obtain image embedding vectors and text embedding vectors respectively. Finally, the similarity matrix between the image embedding vector and the text embedding vector is calculated, multiplied by a learnable similarity temperature parameter, to obtain the predicted similarity matrix.
[0032] Furthermore, the image to be classified is input into the image encoder to obtain an image embedding vector, all elements of the disease entity are input into the second text encoder to obtain the text embedding vector of all elements, and the similarity matrix between the image embedding vector and the text embedding vector is obtained;
[0033] For the image, retrieve the top-k elements that are most similar to it;
[0034] Calculate the overlapping quantity between the top-k elements in the retrieval result and each true disease entity, and select the one with the largest overlapping quantity as the predicted disease entity to achieve the image classification task;
[0035] The retrieval text is input into the second text encoder to obtain the text embedding vector of all elements, and the image is input into the image encoder to obtain the image embedding vector;
[0036] For each retrieval text, retrieve the top-k images that are most similar to it;
[0037] Statistically retrieve the number of overlaps between the images in the retrieval results and the images in the real medical imaging diagnosis report, and select the one with the largest number of overlaps as the predicted medical imaging diagnosis report to achieve the image retrieval task.
[0038] The present invention provides a medical image analysis device for a vision language model, including: at least one processing unit, the processing unit is connected to a storage unit through a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the medical image analysis method of the vision language model is implemented.
[0039] The present invention provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the medical image analysis method of the vision language model is implemented.
[0040] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:
[0041] This application constructs a dataset containing disease entities, images, and image mask layers by using medical imaging diagnosis reports that include medical images and image diagnosis content, and an image feature description library of disease types. Among them, one image corresponds to at least one disease entity, and the disease entity includes the attributes of the disease, the onset location, the affected organ, the category of the disease, and the image feature description of the disease entity; all non-repeating elements of all disease entities within the scope of all medical imaging diagnosis reports are extracted as reference elements, and a reference element embedding vector of all reference elements is extracted by using a first text encoder; for each image in the dataset, an element embedding vector of each element in its disease entity is extracted by using the first text encoder, and the cosine similarity between each element embedding vector and all reference element embedding vectors is calculated; the cosine similarity matrix of all disease entities of any image is subjected to max pooling processing and projected into an image label with a dimension of 1×the number of reference elements; the image labels of all images are concatenated to obtain a similarity matrix between all images and all reference elements in the medical imaging diagnosis report; a vision-language model is constructed and trained, and the vision-language model includes an image encoder to be trained, an image decoder, and a second text encoder. The image encoder extracts the initial image features of all images in the medical imaging diagnosis report. On the one hand, each group of initial image features is input into the image decoder, and the image decoder maps the initial image features into a predicted mask layer. On the one hand, each group of initial image features is input into a mapping layer, and the mapping layer obtains image features based on the initial image features; the second text encoder extracts text features from the disease entity, calculates the similarity between the text features and the image features, and obtains a predicted similarity matrix; during training, the sum of the cross-entropy loss between the predicted mask layer and the true mask layer in the dataset, the cross-entropy loss and the mean square error loss between the predicted similarity matrix and the actual similarity matrix is constrained to be the smallest; by means of the mask layer and the constrained image encoder and image decoder, the image features involved in the image encoder retain the information related to the disease impact feature description in the image information during the compression process. Because ultimately, the image decoder can restore the mask layer of the image from the initial image features, and the mask layer is the mask of different disease imaging regions, corresponding to different imaging feature descriptions. Through this constraint, the influence of the normal region in the image having similar imaging features to the disease features on the model can be avoided as much as possible. The cross-entropy loss and the mean square error loss between the predicted similarity matrix and the actual similarity matrix are used to constrain the entire training process, so that the image encoder and the second text encoder can extract features that support establishing the connection between the two from the image and the disease entity in a coordinated manner. The image encoder and the second text encoder constrain each other to focus the attention of the vision-language model on the image regions related to the disease and the disease entity elements.
[0042] The vision-language model architecture and training method of this application can effectively improve the accuracy of the model in processing images and texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings herein are incorporated into and form a part of the specification, showing embodiments in accordance with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a flowchart of a medical image analysis method for a visual language model provided by an embodiment of the present invention;
[0046] Figure 2 It is a schematic diagram of the process of preparing a data set provided by an embodiment of the present invention;
[0047] Figure 3 It is a flowchart of constructing a data set including disease entities, images, and image mask layers by using a medical image diagnosis report containing medical images and image diagnosis content and an image feature description library of disease types provided by an embodiment of the present invention;
[0048] Figure 4 It is a schematic diagram of a visual language model provided by an embodiment of the present invention;
[0049] Figure 5 It is a schematic diagram of a medical image analysis device of a visual language model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0051] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising said element.
[0052] Embodiment 1
[0053] As Figure 1 shown, the medical image analysis method of the visual language model includes:
[0054] S100, constructing a data set including disease entities, images and image mask layers by using a medical image diagnosis report containing medical images and image diagnosis content and an image feature description library of disease types, wherein one image corresponds to at least one disease entity, and the disease entity includes the attributes of the disease, the onset location, the affected organ, the category of the disease and the image feature description of the disease entity; As Figure 2 and Figure 3 shown, the process includes:
[0055] S101, extracting the mask layer and the image description within the mask layer area from the image by using a pre-trained semantic segmentation model, and the mask layer includes the mask layer of the abnormal area in the image; The semantic segmentation model can be an existing medical semantic segmentation model based on Unet.
[0056] S102, using the disease entity extraction prompt to control the large language model to extract disease entities from the image diagnosis content and the image feature description of the disease type.
[0057] The image feature description library of the disease types contains the image feature descriptions of various diseases; such as ascites: due to fluid accumulation in the abdominal cavity, ultrasound examination may show elevated diaphragm or presence of a fluid wave. Pulmonary embolism: due to infarction or wedge-shaped opacity, the lung area may appear darker; it may also cause pulmonary artery dilation. Pleural effusion: fluid accumulation in the pleural cavity, usually presenting as uniform density in the lower lobe of the lung or a crescent sign along the chest wall, indicating fluid in the pleural cavity. Pulmonary nodule: small round opacity in the lung parenchyma, the size and number may vary. Pneumoperitoneum: gas accumulation in the abdominal cavity, which may accumulate under the diaphragm and present as a sharp lucent line on an upright X-ray. Pulmonary fibrosis: reticular markings, honeycombing changes and possible traction bronchiectasis, indicating lung tissue stiffness and scarring. Infectious process: depending on the type and extent of the infection, there may be opacity, consolidation or cavities ranging from focal to diffuse. Bronchiectasis: abnormally dilated and thickened bronchi, which may appear as circular or tubular structures on an X-ray. Cyst: fluid-filled space, which may appear as a round, well-defined lucency or opacity, depending on its content and wall thickness. Abscess: usually presenting as a round or oval density, often with a fluid level, indicating pus in a newly formed or existing cavity. Embolism: obstruction in a blood vessel, which may not be directly visible on an X-ray but may cause an increase or decrease in lung density. Pneumomediastinum: free air in the mediastinal space, presenting as linear or striated lucency outlining the mediastinal structures. Granuloma: small, local, round areas of different image densities formed due to past inflammation or infection. Lymphadenopathy: enlarged lymph nodes, usually presenting as round opacities in the mediastinal or hilar regions. Subcutaneous emphysema: air in the subcutaneous tissue, presenting as striated or bubbly lucency under the skin, usually around the neck or chest wall. Hemorrhage: depending on its location, it may cause an increase in image density of blood in the lung tissue, and if in a body cavity, there may be an air-fluid level. Infiltration: areas of diffuse increase in image density in the lung, indicating inflammation, infection or other processes affecting the lung parenchyma. Aneurysm: local dilation of a blood vessel, which may appear as an abnormal round or oval shadow, but is more clearly visible on CT or MRI.
[0058] Specifically, the disease entity extraction prompt for controlling the large language model to extract disease entities from image diagnosis content includes: restricting the large language model to use medical knowledge to extract disease-related information sentence by sentence from the provided medical imaging diagnosis report as disease entities; restricting the format of each extracted disease entity: {attribute of the disease}{location of onset}{organ of onset}{category of the disease}, where the attribute of the disease is the severity or stage, scope of the disease; when any information is missing in any image diagnosis content, use a mask to replace the missing information; when the medical imaging diagnosis report mentions a negative situation of the disease, add a negative word before the disease entity; restricting to extract all disease entities from each sentence of the image diagnosis content; restricting to ignore words irrelevant to the disease description when extracting all disease entities; restricting the separation symbol between disease entities.
[0059] An example of a prompt is as follows:
[0060] "Now you need to use your medical knowledge to help me. I will give you some medical imaging diagnosis reports, and you need to extract disease-related information sentence by sentence from these medical imaging diagnosis reports as disease entities. When extracting information, you need to follow the following rules:
[0061] Each extracted disease entity conforms to the format: {attribute of the disease}{location of onset}{organ of onset}{category of the disease}, and the attribute of the disease can be the severity or stage, scope of the disease. For example, for the image diagnosis content "There is a small area of nodular opacity aggregation in the right lower lung lobe.", the extracted disease-related information should be "{small area}{right lower}{lung lobe}{nodular opacity}".
[0062] Some image diagnosis content may lack some information. In this case, you need to use a mask to replace the missing information. For example, if the sentence is "Inflammation of the right lung", then the extracted disease-related information should be {mask}{right}{lung}{inflammation}.
[0063] If the medical imaging diagnosis report mentions a negative situation of the disease, for example, the image diagnosis content is "There is no pneumothorax or left pleural effusion.", then the extracted disease-related information should be "{no}{mask}{mask}{pneumothorax}" and "{no}{left}{pleura}{effusion}".
[0064] A sentence in the image diagnosis content may contain multiple disease descriptions. You need to find them and extract the disease-related information.
[0065] Ignore words irrelevant to the disease description and there is no disease-related information, so you don't need to extract information from this sentence.
[0066] Separate the information with commas."
[0067] The image diagnosis content input into the large language model is as follows: Findings: The heart size is normal. The mediastinal and hilar contours are normal. The pulmonary vessels are normal. There are diffuse interstitial prominences. No pleural effusion or pneumothorax is seen. There is no evidence of pulmonary edema. Impression: Diffuse, mild interstitial abnormalities may represent interstitial lung disease or chronic heart failure. There is no pulmonary edema.
[0068] The response of the large language model is as follows: {Diffuse} {mask} {interstitial} {prominence}, {No} {mask} {mask} {pleural} {pleural effusion}, {No} {mask} {mask} {mask} {pneumothorax}, {No} {mask} {mask} {pulmonary} {pulmonary edema}.
[0069] The disease entity extraction prompts for controlling the large language model to extract disease entities from the description of imaging features include:
[0070] Limit the large language model to use medical knowledge to extract disease-related information from the description of imaging features sentence by sentence as disease entities;
[0071] Limit the format of each extracted disease entity: {Category of disease} {Color description of disease image} {Morphological description of disease image}, where when any information in the description of any imaging feature is missing, use a mask to replace the missing information;
[0072] Limit the delimiter between disease entities.
[0073] For a large language model lacking medical domain capabilities, before using it to extract disease entities, fine-tune the large language model. The fine-tuning dataset includes: prompts, example image diagnosis content, example descriptions of imaging features, and example responses of the large language model. Use the cross-entropy between the predicted response and the actual response of the large language model as the fine-tuning loss function, and perform fine-tuning training with the aim of minimizing the fine-tuning loss function.
[0074] Take the fine-tuning data for extracting disease entities from the image diagnosis content as an example:
[0075] Prompt: "Now you need to use your medical knowledge to help me. I will give you some medical imaging diagnosis reports, and you need to extract disease-related information from these medical imaging diagnosis reports sentence by sentence as entities. When extracting information, you need to follow the following rules:
[0076] Each extracted entity conforms to the format: {Attribute of disease} {Location of onset} {Organ of onset} {Category of disease}. For example, for the image diagnosis content "There is a small area of nodular opacity aggregation in the right lower lung lobe.", the disease-related information extracted should be "{Small area} {Right lower} {Lung lobe} {Nodular opacity}".
[0077] Some image diagnosis content may lack certain information. In this case, you need to use masks instead. For example, if the sentence is "Inflammation of the right lung", then the disease-related information extracted should be {mask}{right}{lung}{inflammation}.
[0078] If the medical imaging diagnosis report mentions the negative situation of a disease. For example, the image diagnosis content is "There is no pneumothorax or left pleural effusion.", then the disease-related information extracted should be "{no}{mask}{mask}{chest}{pneumothorax}" and "{no}{mask}{left}{pleural}{effusion}".
[0079] A sentence in the image diagnosis content may contain multiple disease descriptions. You need to find them and extract the disease-related information.
[0080] Ignore the words irrelevant to the disease description and there is no disease-related information, so you don't need to extract information from this sentence.
[0081] Separate the information with commas.
[0082] Example image diagnosis content: "Findings: Posteroanterior and lateral chest radiographs show low lung volumes, resulting in crowded bronchovascular markings. There is a small amount of bilateral pleural effusion and adjacent atelectasis. The prominence of interstitial markings has not changed compared with previous examinations, which may reflect chronic interstitial lung disease. Superimposed mild pulmonary edema cannot be excluded. There is no pneumothorax. The cardiac silhouette and hilar contours are stable. Impression: The prominence of interstitial lung markings has not changed compared with previous examinations, which may reflect chronic interstitial lung disease. Superimposed mild pulmonary edema cannot be excluded."
[0083] Example large language model response: "{low}{mask}{lung}{lung volume}, {small amount}{bilateral}{pleural}{pleural effusion}, {adjacent}{mask}{lung}{atelectasis}, {no change compared with previous examinations}{mask}{interstitial}{prominence of interstitial markings}, {possibly}{mask}{lung}{chronic interstitial lung disease}, {mild}{mask}{lung}{pulmonary edema}."
[0084] Taking the sum of the cross-entropy loss and mean squared error loss between the large language model prediction response and the actual response as the fine-tuning loss function, and aiming at minimizing the fine-tuning loss function, LlamaFactory is used for fine-tuning.
[0085] The disease entities obtained from the image diagnosis content include: the attributes of the disease, the onset location, the affected organ, and the category of the disease. The disease entities obtained from the imaging feature descriptions of the disease type include the category of the disease, the disease imaging color description, and the disease impact morphology description.
[0086] By means of extracting disease entities through a large language model, convert medical imaging reports and disease imaging descriptions in text form into a standardized disease entity format, and support the processing of negation and missing values.
[0087] S103, use the category of the disease to match the imaging feature description of the disease to the disease entity obtained from the image diagnosis content to obtain a complete disease entity.
[0088] S104, combine the image, the disease entity corresponding to the image, and the mask layer as a data set.
[0089] S200, extract non-repeated elements of all disease entities within the entire medical imaging diagnosis report as reference elements, and use the first text encoder to extract reference element embedding vectors of all reference elements; the first text encoder is the tokenizer of a pre-trained language model, such as the tokenizer of the BERT model.
[0090] S300, for each image in the data set, use the first text encoder to extract the element embedding vector of each element in its disease entity, and calculate the cosine similarity between each element embedding vector and all reference element embedding vectors;
[0091] S400, perform max pooling on the cosine similarity matrix of all disease entities of an arbitrary image, and project it into an image label with a dimension of 1×the number of reference elements; the max pooling selects the maximum similarity value for each column of the cosine similarity matrix and projects it into an image label with a dimension of 1×the number of reference elements.
[0092] Concatenate the image labels of all images to obtain a similarity matrix between all images and all reference elements in the medical imaging diagnosis report; this similarity matrix is the true similarity matrix between all images and all reference elements.
[0093] S500, construct and train a vision-language model, such as Figure 4 as shown, the vision-language model includes an image encoder, an image decoder, and a second text encoder to be trained. The image encoder of the vision-language model uses a Resnet model, and the model structure of the image decoder is symmetric to that of the image encoder; the second text encoder uses a CLIP text model.
[0094] Training stage: The image encoder extracts the initial image features of all images in the medical imaging diagnosis report. On the one hand, each group of initial image features is input into the image decoder, and the image decoder maps the initial image features into a predicted mask layer. On the other hand, each group of initial image features is input into the mapping layer, and the mapping layer obtains image features based on the initial image features; the disease entities corresponding to the images in the dataset are tokenized and padded through the CLIP word processor, converted into tokens, and then input into the second text encoder to extract the pooled output. Then, the pooled output is adjusted in dimension through the projection layer to obtain text features. After that, the image and text features are L2-normalized to obtain image embedding vectors and text embedding vectors respectively. Finally, the similarity matrix between the image embedding vector and the text embedding vector is calculated, and multiplied by a learnable similarity temperature parameter to obtain the predicted similarity matrix. During training, the sum of the cross-entropy loss between the predicted mask layer and the true mask layer in the dataset, the cross-entropy loss and the mean squared error loss between the predicted similarity matrix and the actual similarity matrix is minimized.
[0095] By means of the mask layer and the constrained image encoder and image decoder, the image features involved in the image encoder retain the information related to the disease impact feature description in the image information during the compression process. Because ultimately, the image decoder can restore the mask layer of the image from the initial image features. Since the mask layer is the mask of different disease imaging regions, corresponding to different imaging feature descriptions. Through this constraint, it is possible to avoid the influence of the normal region in the image having similar imaging features to the disease features on the model as much as possible.
[0096] The cross-entropy loss and the mean squared error loss between the predicted similarity matrix and the actual similarity matrix are used to constrain the entire training process, so that the image encoder and the second text encoder can extract the features that support establishing the connection between the two from the image and the disease entity in a coordinated manner. The image encoder and the second text encoder constrain each other to focus the visual language model's attention on the disease-related image regions and disease entity elements.
[0097] S600, during application, the trained image encoder and the second text encoder are used for medical image analysis tasks. The medical image analysis tasks include: medical image classification from images to disease types, or retrieval matching of images according to disease descriptions.
[0098] The image to be classified is input into the image encoder to obtain an image embedding vector, all elements of the disease entity are input into the second text encoder to obtain the text embedding vector of all elements, and the similarity matrix between the image embedding vector and the text embedding vector is obtained;
[0099] For the image, retrieve the top-k elements that are most similar to it;
[0100] Calculate the number of overlaps between the top-k elements in the retrieval results and each true disease entity, and select the one with the largest number of overlaps as the predicted disease entity to implement the image classification task;
[0101] Input the retrieval text into the second text encoder to obtain the text embedding vectors of all elements, and input the image into the image encoder to obtain the image embedding vectors;
[0102] For each retrieval text, retrieve the top-k images that are most similar to it;
[0103] Count the number of overlaps between the images in the retrieval results and the images in the true medical image diagnosis report, and select the one with the largest number of overlaps as the predicted medical image diagnosis report to implement the image retrieval task.
[0104] Example 2
[0105] Refer to Figure 5 As shown, the embodiment of the present invention provides a medical image analysis device for a vision language model, including: at least one processing unit, the processing unit is connected to a storage unit through a bus unit, and the storage unit is used as a computer-readable storage medium and can be used to store software programs, computer-executable programs, and modules, such as the software programs, computer-executable programs, and modules corresponding to a medical image analysis method of a vision language model in the embodiment of the present invention. The processing unit realizes the above-mentioned medical image analysis method of a vision language model by running the software programs, computer-executable programs, and modules stored in the storage unit, including:
[0106] Use a medical image diagnosis report containing medical images and image diagnosis content and an image feature description library of disease types to construct a data set containing disease entities, images, and image mask layers. Among them, one image corresponds to at least one disease entity, and the disease entity includes the attributes of the disease, the onset location, the onset organ, the category of the disease, and the image feature description of the disease entity;
[0107] Extract the non-repeated elements of all disease entities within the scope of all medical image diagnosis reports as reference elements, and use the first text encoder to extract the reference element embedding vectors of all reference elements;
[0108] For each image in the data set, use the first text encoder to extract the element embedding vectors of each element in its disease entity, and calculate the cosine similarity between each element embedding vector and all reference element embedding vectors;
[0109] Perform max pooling on the cosine similarity matrix of all disease entities in an arbitrary image, and project it into an image label with a dimension of 1×the number of reference elements; splice the image labels of all images to obtain a similarity matrix between all images and all reference elements in the medical image diagnosis report;
[0110] Construct and train a vision-language model, which includes an image encoder to be trained, an image decoder, and a second text encoder. The image encoder extracts the initial image features of all images in the medical image diagnosis report. Each group of initial image features is input into the image decoder on the one hand, and the image decoder maps the initial image features into a predicted mask layer. Each group of initial image features is input into the mapping layer on the other hand, and the mapping layer obtains image features based on the initial image features; the second text encoder extracts text features from disease entities, calculates the similarity between the text features and the image features, and obtains a predicted similarity matrix; during training, constrain the sum of the cross-entropy loss between the predicted mask layer and the true mask layer in the dataset, the cross-entropy loss and the mean square error loss between the predicted similarity matrix and the actual similarity matrix to be minimized;
[0111] During application, take the trained image encoder and the second text encoder to perform medical image analysis tasks, and the medical image analysis tasks include: medical image classification from images to disease types, or retrieval matching of images by disease descriptions.
[0112] Of course, the computer program stored in the storage unit of the medical image analysis device of a vision-language model provided by the embodiments of the present invention is not limited to the method operations described above, and can also execute related operations in the medical image analysis method of a vision-language model provided by any embodiment of the present invention.
[0113] Embodiment 3
[0114] The embodiments of the present invention provide a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the medical image analysis method of the vision-language model, including:
[0115] Use a medical image diagnosis report containing medical images and image diagnosis content and an image feature description library of disease types to construct a dataset containing disease entities, images, and image mask layers. Among them, one image corresponds to at least one disease entity, and the disease entity includes the attributes of the disease, the onset location, the onset organ, the category of the disease, and the image feature description of the disease entity;
[0116] Extract non-repeated elements of all disease entities within the scope of all medical image diagnosis reports as reference elements, and use the first text encoder to extract the reference element embedding vectors of all reference elements;
[0117] For each image in the dataset, use the first text encoder to extract the element embedding vectors of each element in its disease entity, and calculate the cosine similarity between each element embedding vector and all reference element embedding vectors;
[0118] Perform max-pooling on the cosine similarity matrix of all disease entities of any image and project it into an image label with a dimension of 1×the number of reference elements; splice the image labels of all images to obtain the similarity matrix between all images and all reference elements in the medical image diagnosis report;
[0119] Construct and train a vision-language model, which includes an image encoder to be trained, an image decoder, and a second text encoder. The image encoder extracts the initial image features of all images in the medical image diagnosis report. Each group of initial image features is input into the image decoder on the one hand. The image decoder maps the initial image features into a predicted mask layer. Each group of initial image features is input into the mapping layer on the other hand. The mapping layer obtains image features based on the initial image features. The second text encoder extracts text features from the disease entity and calculates the similarity between the text features and the image features to obtain a predicted similarity matrix. During training, constrain the sum of the cross-entropy loss between the predicted mask layer and the true mask layer in the dataset, the cross-entropy loss and the mean square error loss between the predicted similarity matrix and the actual similarity matrix to be minimized;
[0120] During application, take the trained image encoder and the second text encoder to perform medical image analysis tasks, and the medical image analysis tasks include: medical image classification from images to disease types, or retrieval matching of images by disease description.
[0121] The computer-readable storage medium provided by the embodiments of the present invention stores computer programs that are not limited to the method operations described above, and can also execute related operations in a medical image analysis method of a vision-language model provided by any embodiment of the present invention.
[0122] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of structures or units can be in electrical, mechanical or other forms.
[0123] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0125] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for medical image analysis of a visual language model, characterized in that, Including: Construct a dataset containing disease entities, images, and image mask layers by using medical imaging diagnosis reports that include medical images and image diagnosis content and an image feature description library of disease types. Among them, one image corresponds to at least one disease entity, and the disease entity includes the attributes of the disease, the onset location, the affected organ, the category of the disease, and the image feature description of the disease entity; Extract the non-repeating elements of all disease entities within the scope of all medical imaging diagnosis reports as reference elements, and use the first text encoder to extract the reference element embedding vectors of all reference elements; For each image in the dataset, use the first text encoder to extract the element embedding vectors of each element in its disease entity, and calculate the cosine similarity between each element embedding vector and all reference element embedding vectors; Perform max pooling on the cosine similarity matrix of all disease entities of any image and project it into an image label with a dimension of 1×the number of reference elements; splice the image labels of all images to obtain the similarity matrix between all images and all reference elements in the medical imaging diagnosis report; Construct and train a vision-language model, which includes an image encoder to be trained, an image decoder, and a second text encoder. The image encoder extracts the initial image features of all images in the medical imaging diagnosis report. On the one hand, each group of initial image features is input into the image decoder, and the image decoder maps the initial image features into a predicted mask layer. On the one hand, each group of initial image features is input into the mapping layer, and the mapping layer obtains image features based on the initial image features; the second text encoder extracts text features from the disease entity and calculates the similarity between the text features and the image features to obtain a predicted similarity matrix; during training, constrain the sum of the cross-entropy loss between the predicted mask layer and the true mask layer in the dataset, the cross-entropy loss and the mean square error loss between the predicted similarity matrix and the actual similarity matrix to be the smallest; During application, take the trained image encoder and the second text encoder to perform medical image analysis tasks. The medical image analysis tasks include: medical image classification from images to disease types, or retrieval matching of images according to disease descriptions.
2. The medical image analysis method of the visual language model according to claim 1, wherein The process of constructing a dataset containing disease entities, images, and mask layers by using medical imaging diagnosis reports that include medical images and image diagnosis content and an image feature description library of disease types includes: Use a pre-trained semantic segmentation model to extract the mask layer and the image description within the mask layer area from the image. The mask layer includes the mask layer of the abnormal area in the image; Use the disease entity extraction prompt to control the large language model to extract disease entities from the image diagnosis content and the image feature description of the disease type; the image feature description library of the disease type includes the image feature descriptions of various diseases; The disease entities obtained from the image diagnosis content include: the attributes of the disease, the onset location, the affected organ, the category of the disease, and the disease entities obtained from the image feature description of the disease type include the category of the disease, the disease image color description, and the disease impact form description; Using the category of diseases to match the image feature descriptions of diseases to the disease entities obtained from the image diagnosis content as complete disease entities; Combining the image, the corresponding disease entity of the image, and the mask layer as a data set.
3. The medical image analysis method of the visual language model according to claim 2, wherein The disease entity extraction prompts for controlling the large language model to extract disease entities from the image feature descriptions include: Limiting the large language model to extract disease-related information sentence by sentence from the image feature descriptions using medical knowledge as disease entities; Limiting the format of each extracted disease entity: {category of disease}{disease image color description}{disease image morphology description}, where when any image feature description lacks any information, a mask is used to replace the missing information; Limiting the separator symbol between disease entities.
4. The medical image analysis method of the visual language model according to claim 2, wherein The disease entity extraction prompts for controlling the large language model to extract disease entities from the image diagnosis content include: Limiting the large language model to extract disease-related information sentence by sentence from the provided medical image diagnosis report using medical knowledge as disease entities; Limiting the format of each extracted disease entity: {attribute of disease}{onset location}{affected organ}{category of disease}, where the attribute of the disease is the severity or stage of development, scope; where when any image diagnosis content lacks any information, a mask is used to replace the missing information; Limiting that when the negative situation of the disease is mentioned in the medical image diagnosis report, a negative word is added before the disease entity; Limiting to extract all disease entities from each sentence of the image diagnosis content; Limiting to ignore words unrelated to the disease description when extracting all disease entities; Limiting the separator symbol between disease entities.
5. The medical image analysis method of the visual language model according to claim 2, wherein For a large language model lacking medical domain capabilities, before using it to extract disease entities, fine-tune the large language model. The fine-tuning data set includes: prompts, example image diagnosis content, example image feature descriptions, and example large language model responses. Use the cross-entropy between the large language model's predicted response and the actual response as the fine-tuning loss function, and perform fine-tuning training with the aim of minimizing the fine-tuning loss function.
6. The medical image analysis method of the visual language model according to claim 1, characterized in that, The image encoder of the vision language model uses a Resnet model, and the model structure of the image decoder is symmetric to that of the image encoder.
7. The method for medical image analysis of the visual language model according to claim 1, wherein, The second text encoder uses a CLIP text model; the image encoder extracts the initial image features of all images in the medical image diagnosis report. On the one hand, each group of initial image features is input into the image decoder, and the image decoder maps the initial image features into a predicted mask layer. On the other hand, each group of initial image features is input into the mapping layer, and the mapping layer obtains image features based on the initial image features; The disease entities corresponding to the images in the data set are tokenized and padded through a CLIP word processor, converted into tokens, and then input into the second text encoder to extract the pooled output. Then, the pooled output is adjusted in dimension through a projection layer to obtain text features; after that, the image and text features are L2-normalized to obtain image embedding vectors and text embedding vectors respectively; finally, the similarity matrix between the image embedding vectors and the text embedding vectors is calculated and multiplied by a learnable similarity temperature parameter to obtain the predicted similarity matrix.
8. The medical image analysis method of the visual language model according to claim 1, wherein Input the image to be classified into the image encoder to obtain an image embedding vector, input all elements of the disease entity into the second text encoder to obtain a text embedding vector of all elements, and obtain a similarity matrix between the image embedding vector and the text embedding vector; For an image, retrieve the top-k elements that are most similar to it; Calculate the overlapping quantity between the top-k elements in the retrieval result and each true disease entity, and select the one with the largest overlapping quantity as the predicted disease entity to implement the image classification task; Input the retrieval text into the second text encoder to obtain a text embedding vector of all elements, and input the image into the image encoder to obtain an image embedding vector; For each retrieval text, retrieve the top-k images that are most similar to it; Count the overlapping quantity between the images in the retrieval result and the images in the true medical image diagnosis report, and select the one with the largest overlapping quantity as the predicted medical image diagnosis report to implement the image retrieval task.
9. A medical image analysis device for a vision language model, characterized in that, Including: At least one processing unit, the processing unit is connected to the storage unit through the bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the medical image analysis method of the vision language model as described in any one of claims 1-8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the medical image analysis method of the vision language model as described in any one of claims 1-8 is implemented.
Citation Information
Cited By
Quantitative characterization and intelligent analysis method for surrounding rock structure based on deep learning
CN121073886A
A Deep Learning-Based Quantitative Characterization and Intelligent Analysis Method for Surrounding Rock Structures
CN121073886B
CT report direct preference optimization method and device based on visual attention mask
CN122091065A