A diagnostic report generation system based on medical image and disease attribute description pairs
By designing a diagnostic report generation system based on medical image and disease attribute description pairs, using joint training text encoder and image encoder, the problem of medical staff's accuracy and efficiency in obtaining diagnostic reports through image analysis is solved, and more efficient and accurate diagnostic report generation is achieved.
Patent Information
- Application Number
- CN202410654310.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-05-24
AI Technical Summary
The method of obtaining diagnostic reports through image analysis by traditional Chinese medicine personnel in the prior art is problem of accuracy and low efficiency.
Design a diagnostic report generation system based on medical image and disease attribute description pairs, including a data set acquisition module, an information extractor, an image knowledge injector, a text encoder, an image encoder, and a diagnostic report generation module. Generate accurate diagnostic reports by jointly training the text encoder and image encoder.
It improves the accuracy and efficiency of diagnostic reports, reduces the possibility of misdiagnosis and missed diagnosis, and reduces the work pressure of medical staff.
Smart Images

Figure CN118447994B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing, and in particular relates to a diagnosis report generation system based on medical images and disease attribute description pairs. Background Art
[0002] Multimodal contrastive learning of images and text is a machine learning technique designed to understand and process data that contains both image and text information. It enables the model to capture cross-modal semantic similarities by learning the correspondence between image content and related text descriptions. Contrastive learning encodes paired image and text inputs and optimizes the contrastive loss function so that matching image and text pairs are closer in the representation space, while mismatched image and text pairs are farther away from each other in the representation space. In the inference phase, the model generates the corresponding embedding vector based on the input modality and finds the closest match in the learned multimodal embedding space.
[0003] The development and application of pre-trained models for medical images and text has become an important progress in the field of medical artificial intelligence. These models can effectively understand and parse medical images and their corresponding text descriptions by training them with a large number of naturally paired medical reports and image datasets. This self-supervised learning method enables the model to automatically learn from a large number of unlabeled medical images and texts without relying on manual annotation.
[0004] In clinical applications, medical staff usually need to obtain a diagnosis report through a series of image analysis, which is a time-consuming and error-prone process. The medical image and text pre-training model can automatically extract the most relevant disease descriptions from images and generate accurate diagnosis reports, thereby effectively reducing the workload of medical staff, improving diagnostic efficiency and accuracy, and reducing the possibility of misdiagnosis and missed diagnosis. Summary of the invention
[0005] The purpose of the present invention is to solve the problem of low accuracy and efficiency in obtaining diagnosis reports by medical personnel analyzing images, and to propose a diagnosis report generation system based on medical images and disease attribute description pairs.
[0006] The technical solution adopted by the present invention to solve the above technical problems is:
[0007] A diagnosis report generation system based on medical images and disease attribute description pairs, the system comprising a data set acquisition module, an information extractor, an image knowledge injector, a text encoder, an image encoder and a diagnosis report generation module;
[0008] The data set acquisition module is used to acquire a data set consisting of medical images and corresponding diagnosis reports;
[0009] The information extractor is used to identify and extract disease information Tf conforming to a preset format from each diagnosis report of the data set according to the prompt;
[0010] The image knowledge injector is used to extract the image attribute description of each medical image in the data set, that is, to obtain the image attribute description Te corresponding to the disease information in each diagnosis report;
[0011] The text encoder and image encoder are jointly trained using the disease information and image attribute description of each medical image-diagnosis report sample pair in the dataset, as well as each medical image;
[0012] The text encoder is used to encode the disease information and image attribute description of each medical image-diagnosis report sample pair in the dataset to obtain the text features of each medical image-diagnosis report sample pair;
[0013] The image encoder is used to encode each medical image in the dataset into image features;
[0014] The diagnosis report generation module inputs the medical image for which the diagnosis report is to be generated into a trained image encoder, filters out the image attribute description that is most similar to the image features of the medical image for which the diagnosis report is to be generated from the image attribute description repository, and generates a diagnosis report based on the disease information corresponding to the filtered image attribute description.
[0015] Furthermore, the information extractor identifies and extracts disease information conforming to a preset format from the diagnosis report according to the prompt, using a large language model.
[0016] Furthermore, the format of the disease information Tf is disease severity+disease location+disease category.
[0017] Furthermore, the image attribute description includes shape, texture and color of the diseased area image.
[0018] Furthermore, the image encoder is a ResNet model or a ConvNext model, and the text encoder is a Transformer model.
[0019] Furthermore, the text encoder and the image encoder are jointly trained, and the specific process of the joint training is:
[0020] The disease information and image attribute description of each medical image-diagnosis report sample pair are input into a text encoder to obtain the text features of each medical image-diagnosis report sample pair, and the semantic similarity matrix s is calculated according to the text features of each medical image-diagnosis report sample pair;
[0021] Each medical image in the dataset is input into the image encoder to obtain the image features of each medical image. The similarity matrix y between the image and the text is calculated based on the text features of each medical image-diagnosis report sample pair and the image features generated by the image encoder.
[0022] Calculate the cross entropy loss CELoss and mean square loss MSELoss between the semantic similarity matrix s and the similarity matrix y, and then calculate the total loss function based on the cross entropy loss CELoss and mean square loss MSELoss;
[0023] The training is stopped when the total loss function converges to obtain the trained text encoder and image encoder.
[0024] Furthermore, the calculation method of the semantic similarity matrix s is:
[0025] Step A1: For a data set including N medical images and diagnostic reports corresponding to the N medical images, the number of disease information Tf contained in the t-th diagnostic report is recorded as c t , t=1, 2, ..., N, the total number of disease information Tf contained in N diagnosis reports is recorded as M, that is,
[0026] And each disease information Tf and the corresponding image attribute description constitute a structured label, that is, a total of M structured labels are obtained;
[0027] Step A2: For any structured label corresponding to the t-th medical image, calculate the similarity between the structured label and each of the M structured labels; similarly, after processing each structured label corresponding to the t-th medical image, obtain the similarity sub-matrix sub_matrix corresponding to the t-th medical image, and the dimension of the similarity sub-matrix sub_matrix is c t ×M;
[0028] sub_matrix i,j =cos(E sim (l i ), E sim (l j ))
[0029] Among them, l i represents the i-th structured label in the structured label corresponding to the t-th medical image, i = 1, 2, 3, ..., c t , l j represents the jth structured tag among all structured tags, j = 1, 2, 3, ..., M, E sim (·) means converting text into feature vector, cos(E sim(l i ), E sim (l j )) represents the calculation of the eigenvector E sim (l i ) and the eigenvector E sim (l j ), cos(E sim (l i ), E sim (l j )) is the element in the i-th row and j-th column of the similarity sub-matrix sub_matrix;
[0030] Step A3: Select the maximum value in each column of the similarity submatrix obtained in step A2 to obtain a similarity vector of size 1×M:
[0031] sub_matrix * =maxpool(sub_matrix)
[0032] Among them, maxpool is the maximum value pooling;
[0033] Step A4: After executing steps A2 and A3 for each medical image, a semantic similarity matrix s with a dimension of N×M is formed using the similarity vectors corresponding to each medical image.
[0034] Furthermore, the calculation method of the similarity matrix y between the image and the text is:
[0035] Step B1: Calculate I i and L j :
[0036] I i =f img (E img (x img,i ))
[0037] L j =f txt (E txt (x txt,j ))
[0038] Among them, x img,i represents the i-th medical image input to the image encoder, E img represents the image encoder, x txt,j represents the jth structured label of the input text encoder, E txt represents the text encoder, f img and f txt It is used to map the image features encoded by the image encoder and the text features encoded by the text encoder into the same encoding space.i represents the mapped image features corresponding to the i-th medical image, L j Represents the mapped text features corresponding to the j-th structured tag;
[0039] Step B2: Calculate the image and text similarity matrix:
[0040]
[0041] The superscript T represents the transpose, ||·|| represents the 2-norm, and y i,j It is the element in the i-th row and j-th column in the image and text similarity matrix, i = 1, 2, 3, ..., N.
[0042] Furthermore, the cross entropy loss CELoss is:
[0043]
[0044] Among them, s i,i Represents the element in the i-th row and j-th column of the semantic similarity matrix s;
[0045] The mean square loss MSELoss is:
[0046]
[0047] The total loss function is calculated based on the cross entropy loss CELoss and the mean square loss MSELoss, specifically:
[0048]
[0049] in, is the total loss function.
[0050] Furthermore, the image attribute description that is most similar to the image features of the medical image to be used to generate the diagnosis report is selected from the image attribute description repository, specifically:
[0051] P(I∈C|T)=Sim(E img (I), E txt (T))
[0052] Among them, E img (I) represents the image features of the medical image I for which the diagnosis report is to be generated after the image encoder outputs it, C represents the disease information in the structured label, T represents the image attribute description corresponding to C, and E txt (T) represents the text features output by T after the text encoder, Sim represents the similarity function, and P(I∈C|T) represents the similarity between the medical image I for which the diagnosis report is to be generated and the image attribute description T;
[0053] T* = argmax T∈V (P(I∈C|T))
[0054] Among them, T * is the most similar image attribute description filtered out, and V is the image attribute description repository.
[0055] The beneficial effects of the present invention are:
[0056] The present invention uses an advanced large language model to perform in-depth disease information extraction on diagnostic reports. Compared with traditional manual extraction methods, the cost of the present invention is extremely low, almost negligible, and while improving the quality of disease information extraction, it can also effectively reduce errors. In addition, by injecting knowledge from the medical field, the present invention further enhances the model's ability to understand the visual attributes of different disease categories, and can successfully apply this knowledge to the learning of unknown categories. The present invention further leverages the advantages of disease information extracted from real scenes and medical field knowledge, and constructs a semantic similarity matrix that accurately describes the correlation between images and diseases to achieve more accurate comparative learning. As an efficient and accurate medical vision-language pre-training method, the present invention significantly improves the efficiency and accuracy of medical image analysis tasks by utilizing naturally paired medical images and diagnostic report data, thereby improving the accuracy and efficiency of diagnostic report generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a workflow diagram of the information extractor based on the large language model;
[0058] Figure 2 It is the workflow diagram of the image knowledge injector;
[0059] Figure 3 It is a training flowchart of the image-text multimodal model. DETAILED DESCRIPTION
[0060] Specific implementation method 1: Combination Figure 1 and Figure 2 The present embodiment is described. The present embodiment describes a diagnosis report generation system based on medical images and disease attribute description pairs, the system comprising a data set acquisition module, an information extractor, an image knowledge injector, a text encoder, an image encoder and a diagnosis report generation module;
[0061] The data set acquisition module is used to acquire a data set consisting of medical images and corresponding diagnosis reports;
[0062] The information extractor is used to identify and extract disease information Tf conforming to a preset format from each diagnosis report of the data set according to the prompt;
[0063] The present invention establishes a medical field knowledge base for implicitly modeling the relationship between diseases. The knowledge base includes the relationship between diseases (such as mutual relationships, complications, etc.), the relationship between diseases and symptoms, the relationship between diseases and treatment methods, etc. Diseases are regarded as nodes in a graph neural network (GNN), and the relationship between diseases is regarded as edges, so as to construct a disease relationship graph to implicitly model the relationship between diseases. During the training process, the model understands the complex relationships between diseases, such as complication relationships, alternative diagnoses, etc. by learning the disease relationship graph. Therefore, when extracting disease information, in addition to the extracted disease information coming from the content of the diagnosis report, part of the information can also come from the modeled medical field knowledge base;
[0064] The image knowledge injector is used to extract the image attribute description of each medical image in the data set, that is, to obtain the image attribute description Te corresponding to the disease information in each diagnosis report;
[0065] The text encoder and image encoder are jointly trained using the disease information and image attribute description of each medical image-diagnosis report sample pair in the dataset, as well as each medical image;
[0066] The text encoder is used to encode the disease information and image attribute description of each medical image-diagnosis report sample pair in the dataset to obtain the text features of each medical image-diagnosis report sample pair;
[0067] The image encoder is used to encode each medical image in the dataset into image features;
[0068] The diagnosis report generation module inputs the medical image for which the diagnosis report is to be generated into a trained image encoder, filters out the image attribute description that is most similar to the image features of the medical image for which the diagnosis report is to be generated from the image attribute description repository, and generates a diagnosis report based on the disease information corresponding to the filtered image attribute description.
[0069] Specific implementation method 2: This implementation method is different from specific implementation method 1 in that the information extractor identifies and extracts disease information that conforms to a preset format from the diagnosis report according to prompts, and uses a large language model.
[0070] The other steps and parameters are the same as those in the first embodiment.
[0071] The large language model that can be used in this embodiment includes but is not limited to LLM. By inputting prompt words into the large language model (telling the model what information to output and the output format) and providing the model with some output examples, the large language model can complete the corresponding task according to the requirements in the prompt words, that is, after the diagnosis report is input into the large language model, the large language model can obtain output that meets the requirements.
[0072] Specific implementation method three: This implementation method is different from specific implementation methods one or two in that the format of the disease information Tf is disease severity+disease location+disease category.
[0073] The other steps and parameters are the same as those in the first or second embodiment.
[0074] Not every diagnosis report can extract complete triplet information. Triplet missing disease severity or disease location can be retained, while triplet missing disease category will be removed.
[0075] Specific implementation method 4: This implementation method is different from any one of specific implementation methods 1 to 3 in that the image attribute description includes the shape, texture and color of the disease area image.
[0076] The other steps and parameters are the same as those in Specific Embodiments 1 to 3.
[0077] Specific implementation five: This implementation differs from any one of specific implementations one to four in that the image encoder is a ResNet model or a ConvNext model, and the text encoder is a Transformer model.
[0078] The other steps and parameters are the same as those in Specific Embodiments 1 to 4.
[0079] The present invention selects a model that has been pre-trained on a large-scale natural image dataset (such as ImageNet) from existing deep learning models. Commonly used models include ResNet, Vision Transformer, SwinTransformer, etc. According to the characteristics of medical images and task requirements, necessary adjustments are made to the pre-trained model. During training, the weights of the convolutional layer of the pre-trained model are frozen, and only a small number of parameters in the fully connected layer are fine-tuned to learn the specific features of the medical image. Finally, the model performance is evaluated on an independent medical image validation set, and the training strategy is adjusted according to the evaluation results until the predetermined performance indicators are met. The present invention uses a model pre-trained based on natural images for transfer learning, which can accelerate the training process.
[0080] Specific implementation method six: Combination Figure 3 This embodiment is different from the first to fifth embodiments in that the text encoder and the image encoder are jointly trained, and the specific process of the joint training is as follows:
[0081] The disease information and image attribute description of each medical image-diagnosis report sample pair are input into a text encoder to obtain the text features of each medical image-diagnosis report sample pair, and the semantic similarity matrix s is calculated according to the text features of each medical image-diagnosis report sample pair;
[0082] Each medical image in the dataset is input into the image encoder to obtain the image features of each medical image. The similarity matrix y between the image and the text is calculated based on the text features of each medical image-diagnosis report sample pair and the image features generated by the image encoder.
[0083] Calculate the cross entropy loss CELoss and mean square loss MSELoss between the semantic similarity matrix s and the similarity matrix y, and then calculate the total loss function based on the cross entropy loss CELoss and mean square loss MSELoss;
[0084] The training is stopped when the total loss function converges to obtain the trained text encoder and image encoder.
[0085] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 5.
[0086] Specific implementation method 7: This implementation method is different from any one of specific implementation methods 1 to 6 in that the calculation method of the semantic similarity matrix s is:
[0087] Step A1: For a data set including N medical images and diagnostic reports corresponding to the N medical images, the number of disease information Tf contained in the t-th diagnostic report is recorded as c t , t=1, 2, ..., N, the total number of disease information Tf contained in N diagnosis reports is recorded as M, that is,
[0088] And each disease information Tf and the corresponding image attribute description constitute a structured label (for an image including multiple disease information Tf, multiple disease information corresponds to the image attribute description of this image), that is, a total of M structured labels are obtained;
[0089] Step A2: For any structured label corresponding to the t-th medical image, calculate the similarity between the structured label and each of the M structured labels; similarly, after processing each structured label corresponding to the t-th medical image, obtain the similarity sub-matrix sub_matrix corresponding to the t-th medical image, and the dimension of the similarity sub-matrix sub_matrix is c t ×M;
[0090] sub_matrix i,j =cos(E sim (l i ), E sim (l j ))
[0091] Among them, l irepresents the i-th structured label in the structured label corresponding to the t-th medical image, i = 1, 2, 3, ..., c t , l j represents the jth structured tag among all structured tags, j = 1, 2, 3, ..., M, E sim (·) means converting text into feature vector, cos(E sim (l i ), E sim (l j )) represents the calculation of the eigenvector E sim (l i ) and the eigenvector E sim (l j ), cos(E sim (l i ), E sim (l j )) is the element in the i-th row and j-th column of the similarity sub-matrix sub_matrix;
[0092] Step A3: Select the maximum value in each column of the similarity submatrix obtained in step A2 to obtain a similarity vector of size 1×M:
[0093] sub_matrix * =maxpool(sub_matrix)
[0094] Among them, maxpool is the maximum value pooling, that is, compressing the matrix into a one-dimensional vector by retaining only the maximum value in one column;
[0095] Step A4: After executing steps A2 and A3 for each medical image, a semantic similarity matrix s with a dimension of N×M is formed using the similarity vectors corresponding to each medical image.
[0096] The other steps and parameters are the same as those in Specific Embodiments 1 to 6.
[0097] The text labels are strongly correlated with the images, so for example, the similarity between the image of "severe pneumonia" and the text of "mild pneumonia" can be obtained by calculating the similarity between the text of "severe pneumonia" and the text of "mild pneumonia".
[0098] The semantic similarity matrix calculated by the present invention is used to replace the multi-hot label matrix commonly used in contrastive learning to enhance the alignment of images and texts. By replacing the binary values ({0, 1}) in the label matrix with continuous values ([0, 1]), a fine-grained similarity measure between images and disease labels is provided, thereby achieving more accurate visual-language alignment, improving model performance, and improving the accuracy of diagnostic report generation.
[0099] Specific implementation eight: This implementation differs from specific implementations one to seven in that the method for calculating the similarity matrix y between the image and the text is:
[0100] Step B1: Calculate I i and L j :
[0101] I i =f img (E img (x img,i ))
[0102] L j =f txt (E txt (x txt,j ))
[0103] Among them, x img,i represents the i-th medical image input to the image encoder, E img represents the image encoder, x txt,j represents the jth structured label of the input text encoder, E txt represents the text encoder, f img and f txt It is used to map the image features encoded by the image encoder and the text features encoded by the text encoder into the same encoding space. i represents the mapped image features corresponding to the i-th medical image, L j Represents the mapped text features corresponding to the j-th structured tag;
[0104] Step B2: Calculate the image and text similarity matrix:
[0105]
[0106] The superscript T represents the transpose, ||·|| represents the 2-norm, and y i,j It is the element in the i-th row and j-th column in the image and text similarity matrix, i = 1, 2, 3, ..., N.
[0107] The other steps and parameters are the same as those in Specific Embodiments 1 to 7.
[0108] Specific implementation method 9: This implementation method is different from any one of specific implementation methods 1 to 8 in that the cross entropy loss CELoss is:
[0109]
[0110] Among them, s i,i Represents the element in the i-th row and j-th column of the semantic similarity matrix s;
[0111] The mean square loss MSELoss is:
[0112]
[0113] The total loss function is calculated based on the cross entropy loss CELoss and the mean square loss MSELoss, specifically:
[0114]
[0115] in, is the total loss function.
[0116] The other steps and parameters are the same as those in Specific Embodiments 1 to 8.
[0117] Calculating the total loss function can improve the performance of the model on the classification task and improve the alignment between the prediction and the label. By optimizing the loss function designed by the present invention, the model can learn richer and more diverse features, thereby improving the quality and generalization ability of feature expression, and further improving the accuracy of diagnostic report generation.
[0118] Specific implementation method 10: This implementation method is different from any one of specific implementation methods 1 to 9 in that the image attribute description that is most similar to the image features of the medical image to be used to generate the diagnosis report is selected from the image attribute description repository, specifically:
[0119] P(I∈C|T)=Sim(E img (I), E txt (T))
[0120] Among them, E img (I) represents the image features of the medical image I for which the diagnosis report is to be generated after the image encoder outputs it, C represents the disease information in the structured label, T represents the image attribute description corresponding to C (that is, T and C come from the same structured label), E txt (T) represents the text features output by T after the text encoder, Sim represents the similarity function (i.e., the feature vector E img (I) and E txt (T) After doing the inner product, we get the cosine similarity. The calculation process of the similarity function is equivalent to first converting E img (I) and E txt (T) is mapped to the same coding space, and then the method of step B2 is used to calculate the similarity of the mapped features), P(I∈C|T) represents the similarity between the medical image I for which the diagnosis report is to be generated and the image attribute description T;
[0121] T * = argmax T∈V (P(I∈C|T))
[0122] Among them, T * is the most similar image attribute description filtered out, and V is the image attribute description repository.
[0123] The other steps and parameters are the same as those in Specific Embodiments 1 to 9.
[0124] The above calculation examples of the present invention are only used to explain the calculation model and calculation process of the present invention in detail, and are not intended to limit the implementation methods of the present invention. For ordinary technicians in the relevant field, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation methods here. All obvious changes or modifications derived from the technical solution of the present invention are still within the scope of protection of the present invention.
Claims
1. A diagnostic report generation system based on medical images and disease attribute description pairs, characterized in that: The system includes a data set acquisition module, an information extractor, an image knowledge injector, a text encoder, an image encoder, and a diagnosis report generation module; The data set acquisition module is used to acquire a data set consisting of medical images and corresponding diagnosis reports; The information extractor is used to identify and extract disease information Tf conforming to a preset format from each diagnosis report of the data set according to the prompt; The image knowledge injector is used to extract the image attribute description of each medical image in the data set, that is, to obtain the image attribute description Te corresponding to the disease information in each diagnosis report; The text encoder and image encoder are jointly trained using the disease information and image attribute description of each medical image-diagnosis report sample pair in the dataset, as well as each medical image; The text encoder is used to encode the disease information and image attribute description of each medical image-diagnosis report sample pair in the dataset to obtain the text features of each medical image-diagnosis report sample pair; The image encoder is used to encode each medical image in the dataset into image features; The text encoder and the image encoder are jointly trained, and the specific process of the joint training is as follows: The disease information and image attribute description of each medical image-diagnosis report sample pair are input into a text encoder to obtain the text features of each medical image-diagnosis report sample pair, and the semantic similarity matrix s is calculated according to the text features of each medical image-diagnosis report sample pair; The calculation method of the semantic similarity matrix s is: Step A1: For a data set including N medical images and diagnostic reports corresponding to the N medical images, the number of disease information Tf contained in the t-th diagnostic report is recorded as c t , t=1,2,...,N, the total number of disease information Tf contained in N diagnosis reports is recorded as M, that is, And each disease information Tf and the corresponding image attribute description constitute a structured label, that is, a total of M structured labels are obtained; Step A2: For any structured label corresponding to the t-th medical image, calculate the similarity between the structured label and each of the M structured labels; similarly, after processing each structured label corresponding to the t-th medical image, obtain the similarity sub-matrix sub_matrix corresponding to the t-th medical image, and the dimension of the similarity sub-matrix sub_matrix is c t ×M; sub_matrix i,j =cos(E sim (l i ),AND sim (l j )) Among them, l i represents the i-th structured label in the structured label corresponding to the t-th medical image, i = 1, 2, 3, ..., c t , l j Represents the jth structured tag among all structured tags, j = 1, 2, 3, ..., M, E sim (·) means converting text into feature vector, cos (E sim (l i ),E sim (l j )) represents the calculation of the eigenvector E sim (l i ) and the eigenvector E sim (l j ), cos (E sim (l i ),E sim (l j )) is the element in the i-th row and j-th column of the similarity sub-matrix sub_matrix; Step A3: Select the maximum value in each column of the similarity submatrix obtained in step A2 to obtain a similarity vector of size 1×M: sub_matrix * =maxpool(sub_matrix) Among them, maxpool is the maximum value pooling; Step A4: After executing steps A2 and A3 for each medical image, a semantic similarity matrix s with a dimension of N × M is formed using the similarity vectors corresponding to each medical image; Each medical image in the dataset is input into the image encoder to obtain the image features of each medical image. The similarity matrix y between the image and the text is calculated based on the text features of each medical image-diagnosis report sample pair and the image features generated by the image encoder. The calculation method of the similarity matrix y between the image and the text is: Step B1: Calculate I i and L j : I i =f img (E img (x img,i )) L j =f txt (E txt (x txt,j )) Among them, x img,i represents the i-th medical image input to the image encoder, E img represents the image encoder, x txt,j represents the jth structured label of the input text encoder, E txt represents the text encoder, f img and f txt It is used to map the image features encoded by the image encoder and the text features encoded by the text encoder into the same encoding space. i represents the mapped image features corresponding to the i-th medical image, L j Represents the mapped text features corresponding to the j-th structured tag; Step B2: Calculate the image and text similarity matrix: The superscript T represents the transpose, ||·|| represents the 2-norm, and y i,j is the element in the i-th row and j-th column of the image and text similarity matrix, i = 1, 2, 3, ..., N; Calculate the cross entropy loss CELoss and mean square loss MSELoss between the semantic similarity matrix s and the similarity matrix y, and then calculate the total loss function based on the cross entropy loss CELoss and mean square loss MSELoss; The cross entropy loss CELoss is: Among them, s i,j Represents the element in the i-th row and j-th column of the semantic similarity matrix s; The mean square loss MSELoss is: The total loss function is calculated based on the cross entropy loss CELoss and the mean square loss MSELoss, specifically: in, is the total loss function; The training is stopped until the total loss function converges, and the trained text encoder and image encoder are obtained; The diagnosis report generation module inputs the medical image for which the diagnosis report is to be generated into a trained image encoder, filters out the image attribute description that is most similar to the image features of the medical image for which the diagnosis report is to be generated from the image attribute description repository, and generates a diagnosis report based on the disease information corresponding to the filtered image attribute description.
2. A diagnostic report generation system based on medical images and disease attribute description pairs according to claim 1, characterized in that: The information extractor identifies and extracts disease information that conforms to a preset format from the diagnosis report according to prompts, using a large language model.
3. A diagnostic report generation system based on medical images and disease attribute description pairs according to claim 1, characterized in that: The format of the disease information Tf is disease severity+disease location+disease category.
4. A diagnostic report generation system based on medical images and disease attribute description pairs according to claim 1, characterized in that: The image attribute description includes the shape, texture and color of the diseased area image.
5. A diagnostic report generation system based on medical images and disease attribute description pairs according to claim 1, characterized in that: The image encoder is a ResNet model or a ConvNext model, and the text encoder is a Transformer model.
6. A diagnostic report generation system based on medical images and disease attribute description pairs according to claim 1, characterized in that: The method of selecting the image attribute description that is most similar to the image features of the medical image to be used to generate the diagnosis report from the image attribute description repository is as follows: P(I∈C|T)=Sim(E img (I),E txt (T)) Among them, E img (I) represents the image features of the medical image I for which the diagnosis report is to be generated after the image encoder outputs it, C represents the disease information in the structured label, T represents the image attribute description corresponding to C, and E txt (T) represents the text features output by T after the text encoder, Sim represents the similarity function, and P(I∈C|T) represents the similarity between the medical image I for which the diagnosis report is to be generated and the image attribute description T; T * =argmax T∈V (P(I∈C|T)) Among them, T * is the most similar image attribute description filtered out, and V is the image attribute description repository.
Citation Information
Patent Citations
Stomach diagnosis report generation system based on self-supervised joint learning
CN116884561A
Visual-semantic representation learning via multi-modal contrastive training
US20220284321A1