Medical image diagnosis report automatic generation method based on modal alignment network architecture
By adopting a modal alignment network architecture method in the medical image report generation model, cross-modal interaction, modal information enhancement and dynamic knowledge graphs are used to solve the problems of incomplete modal interaction, difficulty in alignment and data deviation, and more efficient report generation and better performance are achieved.
Patent Information
- Application Number
- CN202510376761.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
AI Technical Summary
The existing medical image report generation model has problems such as incomplete modal interaction, difficulty in modal alignment, high complexity in report retrieval and calculation, and insufficient data deviation processing.
Using a modal alignment network architecture method, modal interaction is promoted through cross-modal interaction and modal information enhancement, modal alignment is guided, clustered information is used to reduce retrieval complexity, and data deviation problem is alleviated through comprehensive utilization of dynamic knowledge graphs and cross-modal matrix.
It effectively enhances the interaction between multimodals, reduces the computational complexity of report retrieval, improves the quality of report generation, and significantly improves performance on public data sets.
Smart Images

Figure CN120220948A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image analysis, and particularly relates to an automatic medical image diagnosis report generation method based on a modal alignment network architecture. Background Art
[0002] Medical images are widely used in disease diagnosis. However, in the traditional clinical process, medical experts need to carefully analyze the images and then write reports, which is time-consuming and error-prone. The automatic medical image report generation technology can quickly generate reports, assist doctors in diagnosis, reduce their work burden, and save medical resources, so it has attracted much attention from researchers.
[0003] Different from traditional image description tasks, medical reports are longer, have more sentences, and have more complex language and semantic patterns. Moreover, there are data deviation problems in common datasets. Most training samples are normal cases, and the proportion of abnormal regions is small. Even in pathological cases, most descriptions focus on normal findings. At the same time, existing research mostly focuses on learning single-modal features and ignores cross-modal interactions, but cross-modal interactions are crucial for dealing with complex semantic relationships between images and texts. Previous research has tried to adopt the self-attention mechanism in the encoder-decoder architecture to alleviate this problem, but still cannot fully capture complex cross-modal patterns. These problems pose great challenges to cross-modal alignment and feature learning.
[0004] Most existing encoder-decoder architectures, in order to fully capture complex cross-modal patterns, will use an automatic label generation tool to splice and cluster the image and report data by category, pre-construct a cross-modal prototype matrix, and interact the input image and report with the cross-modal matrix by category in subsequent training. Although the model has made great progress in capturing complex cross-modal patterns, it is still limited by the dataset and cannot fully utilize prior knowledge to enhance cross-modal learning, and at the same time, it does not well alleviate the data deviation problem.
[0005] In order to make full use of prior knowledge and alleviate data deviation, and at the same time solve the limitation problem of the knowledge graph when dealing with entities not existing in the graph, the prior art has a solution that adopts the method of a dynamic knowledge graph. During the model training process, entities and relationships are dynamically extracted from similar reports and updated into the knowledge graph, and the graph encoding is used to enhance the image features. Although the model provided by this prior solution can capture complex cross-modal patterns to a certain extent and promote modal alignment, this alignment is implicit and the constraint on modal alignment is very weak. At the same time, a large number of similarity calculations are required to obtain similar reports, and there are also inaccuracies in obtaining similar reports only by cosine similarity calculation. Summary of the Invention
[0006] 1. Technical Problems to be Solved by the Invention
[0007] In view of the problems existing in the existing report generation models, such as incomplete modal interaction, difficult modal alignment, high computational complexity of report retrieval, and insufficient handling of data bias, the present invention provides an automatic medical image diagnosis report generation method based on a modal alignment network architecture. The present invention promotes modal interaction through cross-modal interaction and modal information enhancement, guides modal alignment, greatly reduces the retrieval complexity with clustering information, and comprehensively utilizes dynamic atlases and cross-modal matrices to alleviate the data bias problem.
[0008] 2. Technical Solution
[0009] To achieve the above object, the technical solution provided by the present invention is as follows:
[0010] An automatic medical image diagnosis report generation method based on a modal alignment network architecture of the present invention includes the following steps:
[0011] Step 1: Obtain a basic data set from the Internet, perform data preprocessing on the medical images and corresponding reports in the data set, extract disease category labels, and construct a knowledge graph;
[0012] Step 2: Input the preprocessed image I and report T into a text feature extractor and a visual feature extractor respectively to obtain a set of image features I and text features T. At the same time, input the entity relationship K in the knowledge graph n into the text feature extractor to obtain entity relationship text features rK n ; Based on the disease category labels, pre-construct a cross-modal matrix by splicing the image features I, text features T, and corresponding entity relationship features rK n and use the K-Means clustering algorithm to generate the initial value of the cross-modal matrix;
[0013] Step 3: Extract the visual features of the medical images, calculate the similarity based on the cluster centers obtained by clustering, and select the top K most similar report data from within the cluster;
[0014] Step 4: Dynamically update the knowledge graph according to the selected similar reports, encode the knowledge graph through a graph encoder, and use the attention mechanism to generate enhanced visual features;
[0015] Step 5: Interact the obtained enhanced visual features and the corresponding reports with the cross-modal matrix, and obtain visual feature responses and report feature responses;
[0016] Step 6: Splice the enhanced visual features with the visual feature responses, and splice the text features with the report feature responses to obtain fused visual features and text features;
[0017] Step 7: Input the fused visual features into an encoder to obtain an intermediate representation; use the intermediate representation and the fused text features together as the input to the decoder to generate a diagnostic report.
[0018] 3. Beneficial effects
[0019] Adopting the technical solution provided by the present invention, compared with the existing well-known technologies, it has the following remarkable effects:
[0020] (1) For the automatic generation method of medical image diagnosis reports based on the modal alignment network architecture of the present invention, to solve the problems of insufficient cross-modal interaction and insufficient modal alignment, while using image and report data, the knowledge contained in the report is extracted, and the three types of image, report, and corresponding knowledge are fused to pre-construct a cross-modal interaction module; to solve the problem that retrieving similar reports requires a large amount of calculation and the retrieval is inaccurate only relying on cosine similarity, and at the same time to fully connect each module, the present invention uses clustering information, allows the image to calculate the similarity with the cluster center after clustering, and then selects similar reports from the most similar clusters. Finally, to further enhance modal interaction and alignment and alleviate the data deviation problem, the present invention uses dynamic knowledge graph encoding to enhance visual features and allows the enhanced visual features to participate in subsequent tasks. By retrieving reports using clustering information based on the dynamic knowledge graph, the amount of calculation is greatly reduced, and at the same time, the update of the knowledge graph is further optimized, improving the quality of report generation.
[0021] (2) The automatic generation method of medical image diagnosis reports based on the modal alignment network architecture of the present invention further utilizes prior knowledge and interacts in a way of enhancing image features, effectively enhancing the full interaction between multiple modalities; the experimental results show that the model has achieved better results in natural language generation metrics and image description metrics on the public datasets IU-Xray and MIMIC-CXR compared with previous studies, indicating the superiority and effectiveness of the present invention. Brief description of the drawings
[0022] Figure 1 It is the model structure diagram of the automatic generation method of medical image diagnosis reports based on the modal alignment network architecture of the present invention.
[0023] Figure 2 It is the flowchart of the automatic generation method of medical image diagnosis reports based on the modal alignment network architecture of the present invention. Detailed implementation manners
[0024] To further understand the content of the present invention, the present invention will be described in detail in combination with the drawings and embodiments.
[0025] Embodiment 1
[0026] Combined with the drawings, the steps of the automatic generation method of medical image diagnosis reports based on the modal alignment network architecture of this embodiment are as follows:
[0027] Step 1: Data Preprocessing and Knowledge Extraction
[0028] Obtain the basic dataset IU-XRay from the Internet, perform statistics and collection on the dataset, remove the data with missing images, and record the corresponding relationship between each report T and the image I in the annotation file.
[0029] Read the basic data generated by the automatic report, and perform the following processing on each group of reports and medical image data:
[0030] Step 1.1: Use the CheXbert tool to annotate the report to obtain the disease categories with corresponding one-hot encoding
[0031] For each report T in the basic data:
[0032]
[0033] Among them, N t represents the total number of words in this report, w i represents the i-th word, and use the open-source label prediction tool CheXbert to obtain its corresponding disease category k with one-hot encoding:
[0034] k = {y1, y2,..., y N}, y i ∈(0, 1) (2)
[0035] Among them, 0 means that this disease category does not exist, 1 means that it exists; N represents the number of categories. In this embodiment, according to 14 common chest imaging categories, N = 14 is taken.
[0036] Step 1.2: Construct a knowledge graph
[0037] Use the entity extraction tool stanza to first extract all entities e in the report T, and use these entities e to obtain the entity relationship K from the RadGraph knowledge graph n , and persistently store the extracted relationship. The storage method is to store it in the form of <e s , rs, e o > triple, where e s is the subjective entity, e o is the objective entity, rs represents the relationship between entities, and is modified by three types: "implies", "modifies", and "is located in", such as "pleural#modify#lungs".
[0038] Step 1.3: Medical Image Preprocessing
[0039] For the medical image corresponding to report T, first scale it to a fixed size, such as [256, 256], then randomly crop it to a fixed size, such as [224, 224], and finally convert it into a tensor and perform normalization processing to obtain the preliminarily processed image I.
[0040] Step 2: Pre-construct the cross-modal matrix
[0041] Input the preprocessed image I and report T into the text feature extractor and visual feature extractor respectively to obtain a set of image features I and text features T. At the same time, input the entity relationship K n into the text feature extractor to obtain the entity relationship text feature rK of the entity relationship n . Among them, the visual feature extractor is constructed based on the pre-trained ResNet101 model, and the text feature extractor is constructed based on the pre-trained Bert model.
[0042] Specifically, to pre-construct the cross-modal matrix, it is necessary to collect reports, corresponding images, and knowledge according to disease categories, and cluster the collected data to obtain the initial value of the matrix. The specific steps for pre-constructing the cross-modal matrix are as follows:
[0043] Step 2.1: Extract multi-modal features
[0044] Use the pre-trained ResNet-101 model as the visual feature extractor to perform visual extraction on the preprocessed image I to obtain the image feature I. The image feature is a global feature with a dimension of [1, C1], where C1 represents the number of extracted channels, which is 2048. Similarly, use the pre-trained Bert model as the text feature extractor to extract the global text feature T with a dimension of [1, C2], and at the same time extract the entity relationship text feature rK of the entity relationship K n with a dimension of [1, C3], and both C2 and C3 are 768. n Step 2.2: Cluster cross-modal data
[0045] First, according to the disease category k corresponding to the data, collect all the image features I and text features T into the visual feature set
[0046] and the text feature set respectively, where u represents the sample pair, and y
[0047]
[0048] = 1 indicates that the sample u belongs to the disease category k. u,k
[0049] Secondly, according to the disease category k, combine the visual feature set with the text feature set The corresponding image feature I and text feature T in it, and the corresponding entity relationship feature rK n The three are jointly concatenated to obtain a set R of cross-modal vectors r k , where:
[0050] r = Concat(I, T, rK n ) (5)
[0051] Finally, use the K-Means clustering algorithm to cluster each cross-modal vector set R k to form N p clusters, and take the average value PM of the features within each cluster, where:
[0052]
[0053]
[0054] where, f KMeans represents the K-Means clustering algorithm, is the i-th cluster in disease category k, taking N p = 20, represents the total number of samples in the i-th cluster in disease category k, and finally uses the average value PM of the features within the cluster as the initial value of the cross-modal matrix.
[0055] Step 3: Extract the visual features of the medical image, calculate the similarity based on the cluster centers obtained by clustering, and select the top K most similar report data from within the cluster. The specific steps are as follows:
[0056] Step 3.1: The model regards the medical image automatic report generation task as a sequence-to-sequence process. For the given image I, use the visual feature extractor described in Step 2.1 to extract visual features, and use the features before the final average pooling as the image feature I r (At this time H, W, and C represent the height, width, and number of channels of the image, which are 224, 224, and 512 respectively), and then concatenate the rows of the image features to obtain the final visual feature representation I v :
[0057]
[0058] where, N v = H × W, v i represents the regional feature at the i-th position of I v .
[0059] Step 3.2: For the given image feature I v , calculate the similarity simG according to all cluster centers, that is, the average value PM(k, i) of the features within the cluster,
[0060] simG = I v ·R PM,T (9)
[0061] Among them, R PM,T represents the reported text feature T corresponding to the cross-modal vector r that is closest to the Euclidean distance of the in-cluster feature average value PM(k, i) recorded during the K-Means clustering process. Before calculation, the image feature I v needs to be mapped to the same dimension as the text feature T using a linear layer.
[0062] Step 3.3: Select the top 1 most similar cluster and calculate the similarity simT with the corresponding text feature T in the cluster
[0063] simT = I v ·R G,k (10)
[0064] Among them, R G,k represents the set of all text features T in a certain cluster of disease category k. Finally, select the top K most similar data according to simT. In this embodiment, specifically select the top 3 most similar data according to the similarity.
[0065] Step 4: Update the entities and entity relationships of the selected similar reports to the knowledge graph, encode the knowledge graph through a graph encoder, and use the attention mechanism to obtain enhanced image features and generate enhanced visual features.
[0066] Step 4.1: According to the entities e and entity relationships K of the top K most similar report data (3 data in this embodiment) obtained in Step 3.3 n , update the entities to the pre-constructed knowledge graph, and set the relationship between the updated entity nodes and the original nodes in the adjacency matrix AM to 1.
[0067] Step 4.2: Use a graph encoder constructed based on the Transformer encoder to obtain the encoding result f G ,
[0068] e rsa = LN(MMHA(f e , AM)+ f e ) (11)
[0069] f G = LN(FFN(e rsa )+ e rsa ) (12)
[0070] Among them, LN represents the normalization operation, MMHA represents the masked multi-head attention, f eIt is a node representation, which consists of the word embeddings of entities extracted using Bert and the hierarchical encoding used to distinguish whether the node is a root node, an organ node, or a disease keyword node. FFN is a feed-forward network.
[0071] Step 4.3: Project the encoding result f G onto the same dimension as the regional image feature v i to obtain f Gv . After that, use the regional image feature v i as the query, the projected encoding result f Gv as the key and value, and calculate the guided attention ega i for each regional image feature v i .
[0072]
[0073] where d k represents the embedding dimension.
[0074] Step 4.4: Calculate the enhanced image feature of each region i according to the guided attention ega to obtain the final enhanced visual feature
[0075]
[0076] Step 5: Interact the obtained enhanced visual feature with the corresponding report T and the cross-modal matrix to obtain a response. The specific steps are as follows:
[0077] Step 5.1: Pass the report T through an embedding layer to obtain the embedding of the report T w
[0078]
[0079] Step 5.2: Determine the vector pv in the cross-modal matrix that needs to be interacted according to the label disease category k of the current data
[0080] pv = {PM(k)|y k = 1} (17)
[0081] Step 5.3: To reduce noise, use a learnable matrix W pv to map pv to C p dimensions
[0082] p = pv · W pv (18)
[0083] Step 5.4: Use two learnable matrices W vw and Wp Enhance the visual feature sequence Report the embedded sequence tw i , cross-modal vector p i Project them onto the same dimension
[0084]
[0085] Step 5.5: There may be a large number of irrelevant vectors in the selected cross-modal vectors. Use similarity to calculate the top γ most similar vectors. The implementation steps are as follows:
[0086] Step 5.5.1: Calculate the similarity between the cross-modal vector and the visual sequence and the similarity with the report sequence
[0087]
[0088] Step 5.5.2: According to the calculated similarity, calculate the similarity weight of the visual sequence The similarity weight of the report sequence Select the top γ most similar vectors according to the weight size.
[0089]
[0090] Step 5.6: Use a learnable matrix W e Project the selected most similar cross-modal vectors, and use the weights obtained by Equation (21) for weighting to obtain the visual feature response and the report feature response
[0091]
[0092] where and respectively represent the j-th vector in the set of the most similar cross-modal vectors corresponding to the i-th image region and word.
[0093] Step 6: Concatenate the unimodal features and their response modalities, that is, the enhanced visual feature sequence concatenate with the visual feature response r v concatenate the report text embedded sequence tw with the report feature response r w concatenate, and then fuse through a fully connected layer to obtain the fused visual feature l v and the fused text feature l w
[0094]
[0095] Step 7: The encoder-decoder of Transformer is used to generate a report, and the steps are as follows:
[0096] Step 7.1: Send the fused visual feature l v into the encoder to obtain an intermediate representation:
[0097]
[0098] Step 7.2: Use the intermediate representation and the fused text feature l w together as the decoder input to predict the word P at the current time step t t
[0099]
[0100] Step 7.3: Repeat Step 7.1 and Step 7.2 until a complete report is generated.
[0101] Step 8: Use the Adam optimizer to optimize the model. Set the learning rate of the visual extractor to 1e -3 , and the learning rate of the encoder-decoder to 2e -3 . The learning rate decays by 0.8 for each iteration, and the batch size is uniformly set to 16.
[0102] This embodiment promotes modal interaction with cross-modal interaction and modal information enhancement, guides modal alignment, greatly reduces the retrieval complexity with clustering information, and comprehensively utilizes the dynamic graph and cross-modal matrix to alleviate the data bias problem. Tables 1 and 2 respectively show the performance comparison between the present invention and other methods on the IU-XRay dataset and the MIMIC-CXR dataset. It can be seen that the present invention has made significant progress in CIDEr, indicating that the model can accurately grasp the overall semantic consistency and generate text similarity.
[0103] Table 1 Performance Comparison between the Present Invention and Other Methods on the IU-XRay Dataset
[0104]
[0105] Table 2 Performance Comparison between the Present Invention and Other Methods on the MIMIC-CXR Dataset
[0106]
[0107] The above schematically describes the present invention and its implementation manners. This description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A method for automatically generating medical imaging diagnostic reports based on a modality alignment network architecture, characterized in that: The following steps are involved: Step 1: Obtain basic data sets from the Internet, preprocess the medical images and corresponding reports in the data sets, extract disease category labels and construct a knowledge graph; Step 2: Input the preprocessed image I and report T into the text feature extractor and visual feature extractor respectively to obtain a set of image features I and text features T, and at the same time, the entity relationship K in the knowledge graph n Input text feature extractor to obtain entity relationship text feature rK n ; Based on the disease category label, by splicing image features I, text features T and corresponding entity relationship features rK n Pre-construct the cross-modal matrix and use the K-Means clustering algorithm to generate the initial value of the cross-modal matrix; Step 3: Extract the visual features of the medical image, calculate the similarity based on the cluster centers obtained by clustering, and select the most similar top K report data from the cluster; Step 4: Dynamically update the knowledge graph based on the selected similarity reports, encode the knowledge graph through the graph encoder, and use the attention mechanism to generate enhanced visual features; Step 5: Interact the obtained enhanced visual features and the corresponding reports with the cross-modal matrix, and obtain visual feature responses and report feature responses; Step 6: Concatenate the enhanced visual features with the visual feature responses, and concatenate the text features with the report feature responses to obtain fused visual features and text features; Step 7: Send the fused visual features into the encoder to obtain an intermediate representation; use the intermediate representation and the fused text features as decoder input to generate a diagnosis report.
2. According to claim 1, a method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture is characterized by: The step 1 comprises: Step 1.1: Use the label prediction tool to annotate the medical report and obtain the corresponding one-hot encoded disease category k; Step 1.2: Use entity extraction tools to extract entities from the report and build a knowledge graph in the form of triples based on the RadGraph knowledge graph; Step 1.3: Perform resizing, random cropping and normalization preprocessing on the medical image to obtain the preprocessed image I.
3. The method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to claim 1, characterized in that: In step 2, the visual feature extractor is built based on the pre-trained ResNet101 model, and the text feature extractor is built based on the pre-trained Bert model.
4. The method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to claim 3, characterized in that: In step 2, according to the disease category k, all image features I and text features T are collected into the visual feature set and text feature set In the above example, the visual feature set is divided into With text feature set The corresponding image features I and text features T in the n The three are concatenated together to obtain the set R of cross-modal vectors r k ; Use K-Means clustering algorithm to cluster each cross-modal vector set R k Clustering is performed to form N p clusters, and take the average PM of the features within each cluster, and use the average PM of the features within the cluster as the initial value of the cross-modal matrix.
5. A method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to any one of claims 1 to 4, characterized in that: The step 3 comprises: Step 3.1: Use the visual feature extractor to extract the visual features of the medical image, and then splice the visual feature images to obtain the final visual feature I v ; Step 3.2: For the final visual feature I v , calculate the similarity simG according to the average value of the features in the cluster PM(k, i), and select the most similar cluster; Step 3.3: Calculate the similarity simT between visual features and text features within the selected cluster and select the top K similarity report data.
6. The method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to claim 5, characterized in that: In step 3.2, the calculation formula of the similarity simG is as follows: simG=I v ·R PM,T Among them, R PM,T It represents the text feature T reported by the cross-modal vector r with the closest Euclidean distance to the mean value PM(k, i) of the cluster features recorded in the K-Means clustering process. Before calculation, the image feature I v Use a linear layer to map it to the same dimension as the text feature T.
7. The method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to claim 5, characterized in that: The step 4 comprises: Step 4.1: Based on the entities and entity relationships of the most similar K reported data, update the entities to the pre-built knowledge graph, and set the relationship between the updated entity nodes and the original nodes to 1 in the adjacency matrix AM; Step 4.2: Use the Transformer-based graph encoder to obtain the encoding result f G ; Step 4.3: The encoding result f G Projected to the region image feature v i The same dimension gets f Gv , and then the regional image feature v i As a query, the projected encoding result f Gv As the key and value, calculate the image feature v of each region i Directing attention i ; Step 4.4: According to the guidance of attention eGa i , calculate the enhanced image features of each region To obtain the final enhanced visual features 8. The method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to claim 7, characterized in that: The step 5 comprises: Step 5.1: Pass the report T through an embedding layer to obtain the report embedding T w ; Step 5.2: Determine the vector pv in the cross-modal matrix that needs to interact based on the labeled disease category k of the current data; Step 5.3: Use a learnable matrix W pv Mapping pv to C p Dimension; Step 5.4: Using two learnable matrices W vw and W p The visual feature sequence will be enhanced Report embedded sequence tw i , cross-modal vector p i Projected into the same dimension; Step 5.5: Use similarity to calculate the first γ most similar cross-modal vectors; Step 5.6: Using a learnable matrix W e The most similar cross-modal vector is projected and weighted with the obtained weights to obtain the visual feature response. Response with report characteristics 9. The method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to claim 8, characterized in that: The step 5.5 comprises: Step 5.5.1: Calculate the similarity between the cross-modal vector and the visual sequence Similarity to reported sequences Step 5.5.2: Calculate the visual sequence similarity weight based on the calculated similarity Report sequence similarity weights Select the first γ most similar vectors by weight.
10. The method for automatically generating a medical imaging diagnosis report based on a modality alignment network architecture according to claim 8, characterized in that: Step 6: The enhanced visual feature sequence With r v Splice, report text embedding sequence tw and r w splicing, and then fusion through the fully connected layer to obtain the fused visual feature l v And the fused text features l w .
Citation Information
Cited By
Inplausible medical multi-modal retrieval enhancement generation method, system, equipment and medium
CN121834023A