A method for generating radiology reports based on semantic alignment
By building a radiological report generation network, cross-modal information interaction and semantic alignment are achieved, the problem of time-consuming and poor quality of radiological report generation in the prior art is solved, and the accuracy and quality of the report are improved.
Patent Information
- Application Number
- CN202510065209.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing radiological report generation methods are time-consuming and labor-intensive, and the generated reports are poor, especially in the accuracy of semantic expression and overall expression.
Using a radiological report generation method based on semantic alignment, a radiological report generation network is constructed, which includes a multi-scale feature extraction module, a cross-modal semantic alignment module and a Transformer report generation module to realize cross-modal information interaction and semantic alignment.
Improved the accuracy and quality of radiological report generation, especially in terms of individual characters, phrase combinations, grammatical and semantic coherence, and generated reports are closer to the logical and semantic integrity of the actual radiological report samples.
Smart Images

Figure CN119517277B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radiology report generation, and particularly relates to a radiology report generation method based on semantic alignment. Background Art
[0002] In modern medical diagnosis, radiology reports are important bases for doctors to make diagnostic and treatment decisions. However, the existing traditional method for generating radiology reports is to write diagnostic reports after manually interpreting medical images. This method is very time-consuming and laborious, and there may even be problems with inaccurate writing of radiology reports. Therefore, the method of generating radiology reports based on neural networks has attracted more and more attention from researchers. However, radiology report generation is a very challenging task. Different from traditional image description tasks, this task requires detailed and professional descriptions of radiology images (i.e., X-ray images). Moreover, the semantic relationship between the professional terms in radiology reports and their corresponding image regions is often more complex.
[0003] Currently, the method of generating radiology reports based on neural networks widely adopts an encoder-decoder architecture. Some existing methods for generating radiology reports based on neural networks often focus on learning single-modal features and ignore the importance of cross-modal interaction. Although there are also some methods for generating radiology reports based on cross-modal alignment in the prior art, the radiology reports generated by these methods generally have poor quality, such as poor semantic expression and inaccurate overall expression in the generated radiology reports. For this reason, the present application designs a radiology report generation method based on semantic alignment. Summary of the Invention
[0004] In view of the above-mentioned defects and deficiencies in the prior art, the present application proposes a radiology report generation method based on semantic alignment.
[0005] To achieve the above-mentioned invention purpose, the present invention adopts the following technical solutions:
[0006] A radiology report generation method based on semantic alignment, comprising the following steps:
[0007] S1. Construct a radiology report generation network, which includes a multi-scale feature extraction module, a cross-modal semantic alignment module, and a Transformer report generation module connected in sequence; the multi-scale feature extraction module includes a convolutional neural network and a multi-scale local sparse attention module connected in sequence; the cross-modal semantic alignment module includes an embedding layer, a multi-head attention module, and a gated fusion module connected to the output end of the embedding layer. The input ends of the multi-head attention module and the gated fusion module are also respectively connected to the output end of the multi-scale local sparse attention module in the multi-scale feature extraction module;
[0008] In this application, a convolutional neural network is used to extract the features of all patches in a radiological image, obtaining an image region feature map ;
[0009] The multi-scale local sparse attention module is used to enhance the features of each patch in the image region feature map to obtain a single-modal image embedding containing multi-scale semantic information E (I) ; The embedding layer is used to perform an embedding mapping operation on the radiological report sample to obtain a single-modal text embedding E (R) ;
[0010] The multi-head attention module calculates the similarity between the two based on all the vectors of the learnable matrix M and the single-modal embedding E , and then only selects the first M vectors in the learnable matrix μ that are most similar to the single-modal embedding E for information interaction, and then performs weighted sum feature fusion based on the first M vectors in the selected learnable matrix μ to complete the calculation of attention and obtain a single-modal fusion representation containing cross-modal information ; Among them, when the single-modal embedding E is the single-modal image embedding E (I) , the single-modal fusion representation containing cross-modal information obtained is the single-modal fusion image representation , when the single-modal embedding E is the single-modal text embedding E (R) , the single-modal fusion representation containing cross-modal information obtained is the single-modal fusion text representation ;
[0011] The gated fusion module is used to fuse the single-modal embedding E and the single-modal fusion representation to obtain the final single-modal feature F. When the gated fusion module fuses the single-modal image embedding E (I) and the single-modal fusion image representation , the final image feature is obtained. When the gated fusion module fuses the single-modal text embedding E (R) and the single-modal fusion text representation , the final text feature is obtained;
[0012] S2. Construct the total loss of the radiology report generation network :
[0013] In this application, the character sequence formed by all the predicted characters in the order of acquisition is used as the predicted report sequence generated by the radiology report generation network model, and the radiology report sample sequence corresponding to the character sequence is . The cross-entropy loss between the predicted report sequence and the corresponding radiology report sample sequence as well as the multi-label contrast loss constitute the total loss of the radiology report generation network ;
[0014] S3. Train the radiology report generation network based on the training set and the total loss of the radiology report generation network to obtain the radiology report generation network model;
[0015] S4. Input the radiology image for which the radiology report is to be generated into the radiology report generation network model, and perform one forward propagation to obtain the predicted radiology report.
[0016] Preferably, in step S1, the convolutional neural network is a pre-trained DenseNet121 network, DenseNet121 network, DenseNet network or ResNet101 network.
[0017] Preferably, in step S1, the multi-scale local sparse attention module includes four linear layers, and the inputs of three of the linear layers are all image region feature maps . After the image region feature map undergoes linear transformation through three linear layers, the query, key and value are obtained respectively; then, the query, key and value are divided into several attention heads respectively, and then, according to the number of attention heads H head the query is split into H head pieces Q h of sub-vectors, the key is split into H head pieces K h of sub-vectors, and the value is split into H head pieces V h of sub-vectors; then, through a sliding window operation, from K h sub-vectors and V hSelect some patches from the sub-vectors to calculate the scaled dot-product attention, and obtain the weighted output ; Then, all the outputs are concatenated using a Concat layer, and then feature fusion is performed through the last linear layer to obtain the image region features containing rich multi-scale semantic information , Then, the image region features containing multi-scale semantic information are flattened in the spatial dimension to obtain a unimodal image embedding containing rich multi-scale semantic information .
[0018] Preferably, in step S1, through a sliding window operation from K h sub-vectors and V h sub-vectors, select some patches to calculate the scaled dot-product attention, which specifically includes the following steps: from K h sub-vectors and V h sub-vectors, with each patch as the center, select the patches around it (i.e., the patch at the center) by setting the hyperparameters of the sliding window operation, namely the convolution kernel size k and the dilation rate r, and obtain the new key and the new value ; In this application, different attention heads are set with different dilation rates, and different attention heads can obtain the semantic information of radiological images at different scales; Then, calculate Q h the similarity between the sub-vector and the new key , then divide by the scaling factor to obtain the scaled similarity, and the scaled similarity value is converted into attention weights through a softmax operation; Then, multiply the attention weights by the new value to obtain the weighted output , and the weighted output is the image feature containing local semantic information, realizing the calculation of the scaled dot-product attention.
[0019] Preferably, in step S1, the gated fusion module includes two linear layers. The input end of the first linear layer is respectively connected to the output end of the multi-scale local sparse attention module and the output end of the embedding layer. The input end of the second linear layer is connected to the output end of the multi-head attention module. The output ends of the two linear layers are both connected to the first Add layer. The output end of the first Add layer is connected to the Sigmoid activation layer. The output end of the Sigmoid activation layer is respectively connected to the first element-wise multiplication module and the second element-wise multiplication module. The input end of the first element-wise multiplication module is also connected to the input end of the first linear layer. The input ends of the first linear layer and the second linear layer are both connected to the input end of the second Add layer. The output end of the second Add layer is sequentially connected to the tanh activation layer and the second element-wise multiplication module. The output end of the first element-wise multiplication module and the input end of the second linear layer are both connected to the input end of the third Add layer. The output end of the third Add layer and the output end of the second element-wise multiplication module are both connected to the fourth Add layer.
[0020] Preferably, in step S1, the Transformer report generation module includes an encoding unit, a decoding unit, and a prediction head connected in sequence. The encoding unit and the decoding unit are also both connected to the output end of the gated fusion module in the cross-modal semantic alignment module. Among them, the encoding unit is used to encode the final image features output by the gated fusion module to obtain the encoded image representation L ; the decoding unit is used to perform a decoding operation on the image representation L and the final text features output by the current gated fusion module to output a decoded representation; the prediction head is used to perform a linear transformation on the decoded representation output by the decoding unit and apply the Softmax operation to output the character with the highest probability, and this character is the predicted next character.
[0021] Preferably, in step S1, the encoding unit includes three Transformer encoder layers connected in sequence. Each Transformer encoder layer contains a multi-head self-attention mechanism and a feed-forward neural network. The parameters of each Transformer encoder layer are set to 8 attention heads and a hidden state of 512 dimensions.
[0022] Preferably, in step S1, the decoding unit includes three Transformer decoder layers connected in sequence. Each Transformer decoder layer contains two parts: a multi-head self-attention mechanism and a feed-forward neural network. The parameters of each Transformer decoder layer are set to 8 attention heads and a hidden state of 512 dimensions.
[0023] Preferably, in step S1, the prediction head includes a linear layer and a Softmax layer connected in sequence.
[0024] Preferably, in step S2, calculate the multi-label contrast loss When calculating, first, use the CheXpert tool to generate pseudo-labels for each X-ray image-text pair sample used for training. Since each X-ray image-text pair sample may contain multiple disease labels, the pseudo-labels set in this application are represented as one-hot vectors of multiple labels. Take the current X-ray image sample as the image anchor, and then select the sample with the highest cosine similarity to the anchor from text samples that are completely different from the labels of the current X-ray image sample as the text negative sample. Then, take the current text sample (this current text sample belongs to an X-ray image-text pair sample with the X-ray image sample mentioned above) as the text anchor, and then select the sample with the highest cosine similarity to the anchor from X-ray image samples that are completely different from the labels of the current text sample as the image negative sample; then calculate the multi-label contrast loss based on the determined negative samples.
[0025] Preferably, in step S3, input the images in the training set and their corresponding radiology report samples into the radiology report generation network for forward propagation, and calculate the total loss of the radiology report generation network and perform backpropagation under the guidance of the total loss to update the weight parameters of the radiology report generation network and complete one epoch of the training process. Repeat the iterative training to obtain the radiology report generation network model.
[0026] Preferably, in step S3, the training set is the training set in the IU-Xray dataset or the training set in the MIMIC-CXR dataset; when training the radiology report generation network based on the training set in the IU-Xray dataset, input the paired images of the IU-Xray dataset into the radiology report generation network, and set μ in the top μ most similar vectors to 64; when training the radiology report generation network based on the training set in the MIMIC-CXR dataset, input a single image of the MIMIC-CXR dataset into the radiology report generation network, and set μ in the top μ most similar vectors to 128.
[0027] Compared with the prior art, the beneficial technical effects of this application are as follows:
[0028] The multi-scale local sparse attention module constructed in this application can effectively enhance the features of each patch in the image region feature map and obtain a single-modal image embedding containing multi-scale semantic information E (I) ; the multi-head attention module constructed in this application processes the embeddings from two modalities (i.e., single-modal text embeddings E(R) and single-modal image embedding E (I) ), which enables the radiology report generation network described in this application to utilize the multi-head attention module to achieve cross-modal information interaction, and then use the gated fusion module to obtain single-modal features with fused multi-modal information; in addition, the multi-label contrast loss constructed in this application can also establish a closer semantic connection between the features of the two modalities of image features and text features, realizing supervised cross-modal information fusion and semantic alignment. Through testing, it can be seen that the radiology report generation method described in this application shows stronger accuracy in the matching of single characters, can generate more content that is consistent with the single characters of the reference text, and improves the generation quality at the single-character level; the radiology report generation method described in this application also performs better in bigram matching, can generate more accurate phrase combinations, and enhances the expression ability of the text at the phrase level; the radiology report generation method described in this application is optimized in trigram matching, and the generated text is more coherent in grammar and semantics, reflecting higher phrase matching accuracy; the method of this application performs well in quadrigram matching, and the generated text has stronger advantages in higher-level phrase collocations and semantic consistency, further improving the quality of the generated report; the text generated by the radiology report generation method described in this application is better in semantic expression and the accuracy and flexibility of the overall expression; the radiology report text generated by the radiology report generation method described in this application is closer to the radiology report sample text in terms of sentence structure and coherence, and can better maintain the overall logic and semantic integrity of the radiology report; in summary, the quality of the report generated by the radiology report generation method described in this application is relatively high. Description of the Drawings
[0029] Figure 1 is a schematic diagram of the network structure of the radiology report generation network of the present invention;
[0030] Figure 2 is Figure 1 a schematic diagram of the network structure of the multi-scale local sparse attention module in
[0031] Figure 3 is Figure 1 a schematic diagram of the network structure of the gated fusion module in Detailed Embodiments
[0032] Example 1:
[0033] A radiology report generation method based on semantic alignment includes the following steps:
[0034] S1. Construct an end-to-end radiology report generation network. The network structure of the radiology report generation network is as Figure 1As shown, the radiology report generation network includes a multi-scale feature extraction module, a cross-modal semantic alignment module, and a Transformer report generation module connected in sequence. Among them, the multi-scale feature extraction module includes a convolutional neural network and a multi-scale local sparse attention module connected in sequence. In this application, the cross-modal semantic alignment module includes an embedding layer, a multi-head attention module, and a gated fusion module connected to the output end of the embedding layer. The input ends of the multi-head attention module and the gated fusion module are also connected to the output end of the multi-scale local sparse attention module in the multi-scale feature extraction module. In this application, the Transformer report generation module includes an encoding unit, a decoding unit, and a prediction head connected in sequence. The encoding unit and the decoding unit are also connected to the output end of the gated fusion module in the cross-modal semantic alignment module. The encoding unit includes three Transformer encoder layers connected in sequence, and the decoding unit includes three Transformer decoder layers connected in sequence.
[0035] In this application, the convolutional neural network in the multi-scale feature extraction module is used to extract features from radiology images (X-ray images in this embodiment) to obtain an image region feature map , and the image region feature map contains the features of all patches in the X-ray image . The convolutional neural network in this application can be a pre-trained DenseNet121 network, DenseNet121 network, DenseNet network, or ResNet101 network. In this embodiment, the pre-trained DenseNet121 network is used, and the structure and function of this pre-trained DenseNet121 network are the same as those of the DenseNet121 network structure disclosed in the paper "Densely connected convolutional networks".
[0036] In this application, the multi-scale local sparse attention module uses each patch in the image region feature map output by the convolutional neural network as a query, and sparsely selects key regions (such as anatomical regions) as keys and values by setting the dilation rate r and the convolutional kernel size k in a window centered on each patch, and then calculates self-attention on these key regions to obtain a single-modal image embedding E (I)。In this application, in the window centered on each patch, key regions (such as anatomical regions) are sparsely selected as keys and values by setting the dilation rate r and the convolution kernel size k, and then self-attention is calculated on these key regions, enabling the multi-scale local sparse attention module to calculate self-attention on the key regions of different-sized regions in the radiological image.
[0037] In this application, the embedding layer in the cross-modal semantic alignment module is used to perform an embedding mapping operation on the radiological report sample to obtain the text features of the radiological report sample, resulting in a single-modal text embedding. E (R) ;
[0038] The multi-head attention module first calculates the similarity between all vectors of the learnable matrix M and the single-modal image embedding E (I) , and then only selects the first M vectors in the learnable matrix that are most similar to the single-modal image embedding μ for information interaction. Then, based on the first M vectors in the selected learnable matrix μ , weighted sum and feature fusion are performed to complete the calculation of attention, obtaining a single-modal fused image representation containing cross-modal information ; The multi-head attention module can also calculate the similarity between all vectors of the learnable matrix M and the single-modal text embedding E (R) , and then only selects the first M vectors in the learnable matrix that are most similar to the single-modal text embedding μ for information interaction. Then, based on the first M vectors in the selected learnable matrix μ , weighted sum and feature fusion are performed to complete the calculation of attention, obtaining a single-modal fused text representation containing cross-modal information ; In this application, the feature dimensions of the single-modal image embedding E (I) and the single-modal text embedding E (R) are the same. The learnable matrix M is a randomly initialized matrix with a size of 2048× d , d represents the feature dimension of the single-modal image embedding E (I) ;
[0039] The multi-head attention module adopted in this application has the same structure as the multi-head self-attention mechanism module in the decoder part of the Transformer structure in the paper "Attention is all you need". However, the multi-head attention module in this application has a completely different function from the multi-head self-attention mechanism module in the decoder part of the Transformer structure in the paper "Attention is all you need". Specifically: the input of the multi-head attention mechanism module in the decoder part of the Transformer structure in the paper "Attention is all you need" is a single-modal sequence. Through the linear transformation of queries, keys, and values, attention scores are calculated, and the input information is weighted and aggregated according to the attention scores to capture the dependencies within the sequence; while the multi-head attention module in this application processes the embeddings from two modalities (i.e., single-modal text embeddings and single-modal image embeddings ) by introducing a shared learnable matrix M, which enables the radiology report generation network described in this application to utilize the multi-head attention module to achieve cross-modal information interaction, and then use the gated fusion module to obtain single-modal features with fused multi-modal information; in addition, when the multi-head attention module in this application performs attention calculation, since only the first M vectors in the learnable matrix that are most similar to the single-modal text embeddings μ are selected for information interaction, and also only the first M vectors in the learnable matrix that are most similar to the single-modal image embeddings μ are selected for information interaction, rather than using all the vectors in the learnable matrix M , the above settings can effectively ensure that some basic vectors in the learnable matrix M matrix space will not be updated excessively, effectively avoiding the problems of overfitting or excessive attention to certain features that may be caused by the interaction between vectors at all positions;
[0040] The gated fusion module is used to fuse the single-modal image embeddings E (I) and the single-modal fused image representations to obtain the final image features ; the gated fusion module is also used to fuse the single-modal text embeddings E (R) and the single-modal fused text representations to obtain the final text features .
[0041] In this application, the encoding unit in the Transformer report generation module is used to encode the final image features output by the gated fusion module in the cross-modal alignment module (the final image features including the final image features of N patches, denoted as , where is the final image feature of the Nth patch) to obtain the encoded image representation L ( L = l1, l2, l3, ..., lN , where lN represents the image representation of the Nth patch);
[0042] The decoding unit is used to perform a decoding operation on the encoded image representation L and the final text features output by the current gated fusion module (the final text features including the final text features of T-1 characters, the final text features , where represents the final text feature of the (T-1)th character) to output a decoded representation;
[0043] The prediction head is used to perform a linear transformation on the decoded representation output by the decoding unit and apply the Softmax operation to obtain the probability distribution of each character in the dictionary (in this application, the dictionary is composed of the characters in the true reports in the training set), and output the character with the highest probability, which is the predicted next character. In this application, the prediction head includes a linear layer and a Softmax layer connected in sequence.
[0044] In this embodiment, the next character is predicted according to all the characters in the current radiology report using the radiology report generation network described in this application. The initial content of the radiology report is a preset sentence marker character. That is, at the beginning, the embedding feature of the preset sentence marker character is obtained using the embedding layer of the radiology report generation network of this application. Then, the embedding feature of the preset sentence marker character is input into the multi-head attention module. After obtaining the next character using the radiology report generation network of this application, the embedding features of the preset sentence marker character and the next character of the preset sentence marker character in the current radiology report are input into the multi-head attention module, and the next character of the next character is obtained using the radiology report generation network of this application. This process is repeated until the preset end character is obtained, and the predicted radiology report corresponding to the X-ray image is obtained.
[0045] In this application, the network structure of the multi-scale local sparse attention module, as Figure 2 shown, includes four linear layers, where the inputs of three linear layers are all the image region feature maps The image region feature map After linear transformation through three linear layers, Query, Key, and Value are obtained respectively, and the calculation methods of Query, Key, and Value are shown in Equation (1).
[0046] (1)
[0047] In Equation (1), Q represents Query, K represents Key, and V represents Value. 、 、 All represent trainable parameters. represents the image region feature map.
[0048] To obtain semantic information under different regional sizes of radiological images, the present application further divides Query, Key, and Value into several attention heads respectively. Then, according to the number of attention heads H head split Query into H head sub-vectors, split Key into Q h sub-vectors, split Value into H head sub-vectors, K h sub-vectors, H head sub-vectors, V h sub-vectors. Q h sub-vectors, K h sub-vectors and V h The feature dimensions of sub-vectors all become , where represents Q h sub-vectors, K h sub-vectors and V h sub-vectors, Q h sub-vectors, K h sub-vectors and V h sub-vectors have the same feature dimension size. represents the feature dimension of Query, Key, and Value. The feature dimension sizes of Query, Key, and Value are the same. In the present application, Hhead is an integer divisible by .
[0049] Then, in order to calculate the local sparse self-attention, the present application selects some patches from the K h sub-vector and V h sub-vector to calculate the scaled dot-product attention. Specifically: Since images usually have spatial structures, in the present application, from the K h sub-vector and V h sub-vector, with each patch as the center, the patches around it (i.e., the patch at the center) are selected by setting the hyperparameters of the sliding window operation, namely the convolution kernel size k and the dilation rate r, to obtain the new key and the new value ; Different dilation rates are set for different attention heads in the present application. In this embodiment, the number of attention heads is four, and the dilation rates r of the four attention heads are respectively set to 1, 2, 3, and 4, and k is uniformly set to 3. Then, different attention heads can obtain semantic information of radiological images at different scales;
[0050] Then, calculate the Q h sub-vector and the similarity between the new key , then divide by the scaling factor to obtain the scaled similarity. The scaled similarity value is converted into attention weights through the softmax operation; then, multiply the attention weights by the new value to obtain the weighted output , as shown in Equation (2). The weighted output is the image feature containing local semantic information.
[0051] (2)
[0052] In Equation (2), T represents the transpose operation, represents Q h sub-vector, K h sub-vector, and V h the feature dimensions of the sub-vector, Q h sub-vector, K h sub-vector, and V h sub-vector have the same feature dimension size.
[0053] In this application, first calculate Q h the similarity between the sub-vector and the new key and then divide it by the scaling factor to obtain the scaled similarity. The scaled similarity value is converted into an attention weight through a softmax operation; then, the attention weight is multiplied by the new value The above method realizes the calculation of scaled dot-product attention.
[0054] Since different dilation rates are set for different attention heads in this application, different attention heads can obtain semantic information of radiological images at different scales. All the outputs of this application are concatenated using a Concat layer and then feature fusion is performed through the last linear layer. The above method can effectively enhance the features of each patch. After feature fusion is performed through the last linear layer, the obtained image region features are image region features containing rich multi-scale semantic information. Then, the image region features containing multi-scale semantic information are flattened in the spatial dimension (the spatial dimension size is H×W) to obtain a single-modal image embedding , where N represents the number of patches, N the size of is and represents the feature dimension of the single-modal image embedding The size of the feature dimension represented by is equal to the size of the feature dimension represented by Since the image region features are image region features containing rich multi-scale semantic information, the single-modal image embedding obtained by flattening the image region features also contains rich multi-scale semantic information. That is to say, the single-modal image embedding
[0055] In this application, after the single-modal embedding E is input into the multi-head attention module, the multi-head attention module first obtains the attention scores in the attention operation; then, through the attention scores A find the top M vectors in the learnable matrix E that are most similar to each visual character in the single-modal embedding μ to obtain a new attention score and a new value embedding , the new attention score is calculated as shown in Equation (3):
[0056] (3)
[0057] In Equation (3), E represents the unimodal embedding, that is, the unimodal embedding E can represent the unimodal image embedding E (I) or the unimodal text embedding E (R) , W Q , W K both represent trainable parameters, T represents the transpose operation, d represents the unimodal embedding E 's feature dimension;
[0058] The calculation method of the new value embedding is shown in Equation (4):
[0059] (4)
[0060] In Equation (4), M represents the learnable matrix, W V represents the trainable parameter;
[0061] Based on the new attention score and the new value embedding calculate the attention to obtain the unimodal fusion representation , as shown in Equation (5):
[0062] (5)
[0063] In Equation (5), the unimodal fusion representation represents the unimodal fusion image representation or the unimodal fusion text representation . When the new attention score is calculated based on the unimodal image embedding E (I) , then the unimodal fusion representation calculated based on this new attention score represents the unimodal fusion image representation ; when the new attention score is calculated based on the unimodal text embedding E (R) , then the unimodal fusion representation The calculated single-modal fusion representation represents the single-modal fusion text representation .
[0064] In this application, the network structure of the gating fusion module is as Figure 3 shown Figure 3 in , represents the element-wise multiplication module. The gating fusion module includes two linear layers. The input end of the first linear layer is respectively connected to the output end of the multi-scale local sparse attention module and the output end of the embedding layer. The input end of the second linear layer is connected to the output end of the multi-head attention module. The output ends of the two linear layers are both connected to the first Add layer. The output end of the first Add layer is connected to the Sigmoid activation layer. The output end of the Sigmoid activation layer is respectively connected to the first element-wise multiplication module and the second element-wise multiplication module. The input end of the first element-wise multiplication module is also connected to the input end of the first linear layer; the input ends of the first linear layer and the second linear layer are both connected to the input end of the second Add layer. The output end of the second Add layer is sequentially connected to the tanh activation layer and the second element-wise multiplication module; the output end of the first element-wise multiplication module and the input end of the second linear layer are both connected to the input end of the third Add layer. The output end of the third Add layer and the output end of the second element-wise multiplication module are both connected to the fourth Add layer.
[0065] In this application, the functions of each sub-module in the gating fusion module are as follows:
[0066] In the gating fusion module, the first linear layer is used to perform a linear transformation operation on the input single-modal embedding E (the single-modal embedding E is the single-modal image embedding output by the multi-scale local sparse attention module E (I) or the single-modal text embedding output by the embedding layer E (R) ); the second linear layer is used to perform a linear transformation operation on the single-modal fusion representation output by the multi-head attention module (the single-modal fusion representation is the single-modal fusion image representation or the single-modal fusion text representation ); the settings of the first linear layer and the second linear layer can facilitate the single-modal embedding E and the single-modal fusion representation Map to the same space to facilitate the first Add layer to perform feature addition for information integration; the first Add layer is used to add the features output by the first linear layer and the features output by the second linear layer to obtain unimodal features with fused multimodal information; the Sigmoid activation layer is used to perform a non-linear compression operation on the features output by the first Add layer to obtain the weights of different features; the first element-wise multiplication module is used to perform element-wise multiplication on the unimodal embedding E and the weights output by the Sigmoid activation layer to obtain unimodal features with the characteristics of weighted fusion information; the second Add layer is used to perform feature addition directly on the unimodal fusion representation and the unimodal embedding E ; the tanh activation layer is used to perform a non-linear transformation operation on the features output by the second Add layer to obtain unimodal features; the second element-wise multiplication module is used to perform element-wise multiplication on the unimodal features output by the tanh activation layer and the weights output by the Sigmoid activation layer to obtain unimodal features with adjusted weights; the third Add layer is used to add the features output by the unimodal embedding E and the features output by the first element-wise multiplication module to obtain unimodal features with rich cross-modal information; the fourth Add layer is used to add the features output by the second Add layer and the features output by the third Add layer to obtain the final unimodal features with in-depth information interaction F .
[0067] In this application, since the learnable matrix M needs to select the top E (R) most similar vectors in the unimodal text embedding μ for information interaction, and also needs to select the top M most similar vectors in the learnable matrix E (I) in the unimodal image embedding μ for information interaction. During the process of training the radiology report generation network, under the supervision of the total loss and the update of network parameters, the learnable matrix M contains the cross-modal information of the unimodal text embedding and the unimodal image embedding. Therefore, the multi-head attention module is based on the learnable matrix M and the unimodal image embedding E (I) to obtain the unimodal fusion image representation which contains cross-modal information. The unimodal fusion image representation and the unimodal image embedding E (I) are respectively subjected to linear transformation operations by the linear layer and then the first Add layer is used for feature addition. The obtained unimodal features are the unimodal features with fused multimodal information;
[0068] In this application, the final unimodal feature F is calculated as shown in Equation (6):
[0069] (6)
[0070] In Equation (6), F represents the final unimodal feature, and the final unimodal feature F can represent the final image feature F (I) or the final text feature F (R) , E represents the unimodal embedding, and the unimodal embedding E can represent the unimodal image embedding E (I) or the unimodal text embedding E (R) , represents the unimodal fusion representation, and the unimodal fusion representation can represent the final image feature F (I) or the final text feature F (R) , represents element-wise multiplication, and tanh represents the activation function. When the unimodal embedding E is the unimodal image embedding E (I) , and the unimodal fusion representation is the unimodal fusion image representation , the gated fusion module outputs the final image feature F (I) ; when the unimodal embedding E is the unimodal text embedding E (R) , and the unimodal fusion representation is the unimodal fusion text representation , the gated fusion module outputs the final text feature F (R) .
[0071] In Equation (6), G represents the weight output by the Sigmoid activation layer, and the calculation method of G is as shown in Equation (7):
[0072] (7)
[0073] In Equation (7), σ is the sigmoid function, represents the unimodal fusion representation, and the unimodal fusion representation can represent the unimodal fusion image representation or the unimodal fusion text representation , E represents a unimodal embedding, and the unimodal embedding E can represent a unimodal image embedding E (I) or a unimodal text embedding E (R) , where both W1 and W2 represent trainable parameters; in Equation (7), the unimodal embedding E is a unimodal image embedding E (I) when, the unimodal fusion representation is the unimodal fusion image representation ; when the unimodal embedding E is a unimodal text embedding E (R) when, the unimodal fusion representation is the unimodal fusion text representation .
[0074] In this application, the encoding unit includes three sequentially connected Transformer encoder layers. Each Transformer encoder layer contains a multi-head self-attention mechanism and a feed-forward neural network. The parameters of each Transformer encoder layer are set to 8 attention heads and a hidden state of 512 dimensions; the decoding unit includes three sequentially connected Transformer decoder layers. Each Transformer decoder layer contains two parts: a multi-head self-attention mechanism and a feed-forward neural network. The parameters of each Transformer decoder layer are set to 8 attention heads and a hidden state of 512 dimensions.
[0075] S2. Construct the total loss of the radiology report generation network :
[0076] In this application, the character sequence formed by all the predicted characters in the order of acquisition is used as the predicted report sequence generated by the radiology report generation network model, and the corresponding radiology report sample sequence is , the predicted report sequence and the corresponding radiology report sample sequence The cross-entropy loss between and the multi-label contrast loss The loss constitutes the total loss of the radiology report generation network , the total loss , the cross-entropy loss and the multi-label contrast loss The relationship between them is shown in Equation (8):
[0077] (8)
[0078] The multi-label contrastive loss constructed in this application , aims to establish a closer semantic connection between the features of two modalities, namely image features and text features, and achieve supervised cross-modal information fusion and semantic alignment. In this application, the multi-label contrastive loss is calculated as shown in Equation (9);
[0079] (9)
[0080] In Equation (9), represents the cosine similarity, is the boundary value, which takes the value of 0.7 in this embodiment, represents the final image feature of the image anchor, represents the final image feature of the image negative sample, represents the final text feature of the text anchor, represents the final text feature of the text negative sample.
[0081] When this application calculates the multi-label contrastive loss , first, the CheXpert tool (from "Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison.") is used to generate pseudo-labels for each X-ray image-text pair sample used for training. Since each X-ray image-text pair sample may contain multiple disease labels, the pseudo-labels set in this application are represented as one-hot vectors of multiple labels. The current X-ray image sample is used as the image anchor, and then the sample with the highest cosine similarity to the anchor is selected from the text samples that are completely different from the labels of the current X-ray image sample as the text negative sample. Then, the current text sample (this current text sample belongs to an X-ray image-text pair sample with the X-ray image sample mentioned above) is used as the text anchor, and then the sample with the highest cosine similarity to the anchor is selected from the X-ray image samples that are completely different from the labels of the current text sample as the image negative sample; then, based on the determined negative samples, the multi-label contrastive loss is calculated using Equation (9).
[0082] In Equation (9) of this application, the calculation method of the cross-entropy loss is shown in Equation (10):
[0083] (10)
[0084] In Equation (10), N r represents the number of characters in the report, represents the i-th character in the radiology report sample, represents the i-th character predicted by the model.
[0085] S3. Train the radiology report generation network based on the training set in the IU-Xray dataset and the total loss of the radiology report generation network to obtain a radiology report generation network model; specifically:
[0086] Input the paired images in the IU-Xray dataset and their corresponding radiology report samples into the radiology report generation network for forward propagation to calculate the total loss of the radiology report generation network , and perform backpropagation under the guidance of the total loss to update the weight parameters of the radiology report generation network, complete one epoch of the training process. After iterating 50 epochs of the training process, the training of the radiology report generation network is completed to obtain a radiology report generation network model. In this application, inputting the paired images in the IU-Xray dataset and their corresponding radiology report samples into the radiology report generation network to train the radiology report generation network is a conventional training method. For example, the paper "Cross-modal memory networks for radiology report generation" discloses this training method.
[0087] In this embodiment, the training process of the radiology report generation network uses the Adam optimizer to optimize the loss gradient. During the training process, the learning rate of the pre-trained DenseNet121 network in the radiology report generation network is set to 1e-4, and the learning rate of all trainable parts except the pre-trained DenseNet121 network in the radiology report generation network is set to 5e-4. μ in the μ most similar vectors is set to 64.
[0088] S4. Input the X-ray image for which a radiology report is to be generated into the radiology report generation network model obtained in step S3 for one forward propagation to obtain a predicted radiology report.
[0089] Embodiment 2:
[0090] The difference between Embodiment 2 and Embodiment 1 is that in step S3, the radiology report generation network is trained based on the training set in the MIMIC-CXR dataset and the total loss of the radiology report generation network to obtain a radiology report generation network model; specifically:
[0091] Input a single image from the MIMIC-CXR dataset and its corresponding radiology report sample into the radiology report generation network for forward propagation to calculate the total loss of the radiology report generation network , and perform backpropagation under the guidance of the total loss to update the weight parameters of the radiology report generation network, completing one epoch of the training process. After iterating 30 epochs of the training process, the training of the radiology report generation network is completed, and the radiology report generation network model is obtained. In this application, inputting a single image from the MIMIC-CXR dataset and its corresponding radiology report sample into the radiology report generation network to train the radiology report generation network is also a conventional training method. For example, the paper "Cross-modal memory networks for radiology report generation" discloses this training method.
[0092] In this embodiment, the training process of the radiology report generation network uses the Adam optimizer to optimize the loss gradient. During the training process, the learning rate of the pre-trained DenseNet121 network in the radiology report generation network is set to 5e-5, and the learning rate of all trainable parts except the pre-trained DenseNet121 network in the radiology report generation network is set to 1e-4. μ in the μ most similar vectors is set to 128.
[0093] Test 1:
[0094] To verify the excellent effects of the radiology report generation method described in the present invention compared with other radiology report generation methods, this application uses the test set in the IU-Xray dataset and the radiology report generation network model obtained in Example 1 to test six existing radiology report generation methods such as the CoATT method (from "On the automatic generation of medical imaging reports"), the HRGR method (from "Hybrid retrieval-generation reinforced agent for medical image report generation"), the CMAS-RL method (from "Show, describe and conclude: On exploiting the structure information of chest X-ray reports"), the SENTSAT+KG method (from "When radiology report generation meets knowledge graph"), the CMCL method (from "Competence-based multimodal curriculum learning for medical report generation"), and the R2GenCMN method (from "Cross-modal memory networks for radiology report generation"). The test results are shown in Table 1.
[0095] Table 1 Test results of different methods on the test sets in the IU-Xray dataset and the test set in the MIMIC-CXR dataset
[0096]
[0097] Test 2:
[0098] To verify the excellent effects of the radiology report generation method described in the present invention compared with other radiology report generation methods, this application uses the test set in the MIMIC-CXR dataset and the radiology report generation network model obtained in Example 2 to test six existing radiology report generation methods such as the ST method (from "A hierarchical approach for generating descriptive image paragraphs"), the ATT2IN method (from "Self-critical sequence training for image captioning"), the ADAATT method (from "Knowing when to look: Adaptive attention via a visual sentinel for image captioning"), the TopDown method (from "Bottom-up and top-down attention for image captioning and visual question answering"), the CMCL method (from "Competence-based multimodal curriculum learning for medical report generation"), and the R2GenCMN method (from "Cross-modal memory networks for radiology report generation"). The test results are shown in Table 2.
[0099] Table 2 Test results of different methods on the test set in the MIMIC-CXR dataset and the test set in the MIMIC-CXR dataset
[0100]
[0101] In Tables 1 and 2, BLEU-1, BLEU-2, BLEU-3, BLEU-4, ROUGE-L, and METEOR are commonly used evaluation indicators in text generation, and the OURS method refers to the radiology report generation method proposed in this application.
[0102] In Table 1 and Table 2, the BLEU metric is used to measure the lexical matching degree between the text generated by the model and the reference text. The higher the value, the stronger the consistency of the generated text in terms of vocabulary and phrases. BLEU-1 to BLEU-4 represent the matching situations of single characters, bigrams, trigrams, and quadgrams respectively, with the level gradually increasing. The higher the value, the better the exact matching degree between the generated text and the reference text. The METEOR metric comprehensively considers precision and recall, can capture the semantic similarity between the generated text and the reference text, and also introduces factors such as synonym matching and word order. The higher the value, the higher the accuracy and flexibility of the generated text in semantic expression. The ROUGE-L metric measures the structural similarity between the generated text and the reference text, and is calculated based on the longest common subsequence. The higher the value, the closer the overall structure and coherence of the generated text at the sentence level are to the reference text. Generally speaking, these metrics evaluate the quality of the generated text from the lexical, semantic, and structural levels respectively, providing multi-angle references for evaluating the generation effect of the model. Ours refers to the radiology report generation method based on semantic alignment described in this application. In this application, an English word is regarded as a character.
[0103] Compared with other existing radiology report generation methods disclosed in this application, the existing R2GENCMN method is closer to the radiology report generation method described in this application. Therefore, this application focuses on comparing the test results of the R2GENCMN method with the test results of the radiology report generation method described in this application, as follows:
[0104] As can be seen from Table 1, when testing based on the test set in the IU-Xray dataset:
[0105] The BLEU-1 metric obtained by the radiology report generation method described in this application is 8.46% higher than the BLEU-1 metric obtained by the R2GENCMN method. This shows that the radiology report generation method described in this application shows stronger accuracy in single-character matching, can generate more content that is consistent with the single characters of the reference text, and improves the generation quality at the single-character level;
[0106] The BLEU-2 metric obtained by the radiology report generation method described in this application is 14.97% higher than the BLEU-2 metric obtained by the R2GENCMN method. This shows that the radiology report generation method described in this application performs better in bigram matching, can generate more precise phrase combinations, and enhances the expression ability of the text (i.e., the radiology report) at the phrase level;
[0107] The BLEU-3 index obtained by the radiology report generation method described in this application has increased by 15.38% compared to the BLEU-3 index obtained by the R2GENCMN method. This indicates that the radiology report generation method described in this application has been optimized in triple matching, and the generated text (i.e., the radiology report) is more coherent in grammar and semantics, demonstrating a higher phrase matching accuracy;
[0108] The BLEU-4 index obtained by the radiology report generation method described in this application has increased by 10.97% compared to the BLEU-4 index obtained by the R2GENCMN method. This shows that the method described in this application performs excellently in quadruple matching, and the generated text (i.e., the radiology report) has stronger advantages in higher-level phrase collocations and semantic consistency, further improving the quality of the generated report;
[0109] The METEOR index obtained by the radiology report generation method described in this application has increased by 19.79% compared to the METEOR index obtained by the R2GENCMN method. This indicates that the text (i.e., the radiology report) generated by the radiology report generation method described in this application is better in semantic expression and the accuracy and flexibility of the overall expression;
[0110] The ROUGE-L index obtained by the radiology report generation method described in this application has increased by 6.98% compared to the ROUGE-L index obtained by the R2GENCMN method. This shows that the text (i.e., the radiology report) generated by the radiology report generation method described in this application is closer to the radiology report sample text in terms of sentence structure and coherence, and can better maintain the overall logic and semantic integrity of the radiology report.
[0111] The MIMIC-CXR dataset used in this application is a large dataset, and its scale is almost 70 times that of the IU-Xray dataset. The data complexity and label noise are significantly increased compared to the IU-Xray dataset. As can be seen from Table 2, when this application is tested based on the test set in the MIMIC-CXR dataset,
[0112] The BLEU-1 index obtained by the radiology report generation method described in this application has increased by 0.88% compared to the BLEU-1 index obtained by the R2GENCMN method. This indicates that when testing the test set in the MIMIC-CXR dataset (a large dataset), the radiology report generation method described in this application shows stronger accuracy in single-character matching, can generate more content that is consistent with the reference text at the single-character level, and improves the generation quality at the single-character level;
[0113] The BLEU-2 index obtained by the radiology report generation method described in this application has increased by 1.88% compared to the BLEU-2 index obtained by the R2GENCMN method. This indicates that when testing on the test set of the MIMIC-CXR dataset (a large dataset), the radiology report generation method described in this application performs better in bigram matching, can generate more precise phrase combinations, and enhances the expression ability of the text (i.e., the radiology report) at the phrase level;
[0114] The BLEU-3 index obtained by the radiology report generation method described in this application has increased by 2.78% compared to the BLEU-3 index obtained by the R2GENCMN method. This indicates that when testing on the test set of the MIMIC-CXR dataset (a large dataset), the radiology report generation method described in this application is optimized in trigram matching, and the generated text (i.e., the radiology report) is more coherent in grammar and semantics, reflecting a higher phrase matching accuracy;
[0115] The BLEU-4 index obtained by the radiology report generation method described in this application has increased by 3.85% compared to the BLEU-4 index obtained by the R2GENCMN method. This indicates that when testing on the test set of the MIMIC-CXR dataset (a large dataset), the method of this application performs well in quadruple matching, and the generated text (i.e., the radiology report) has stronger advantages in higher-level phrase collocations and semantic consistency, further improving the quality of the generated report;
[0116] The METEOR index obtained by the radiology report generation method described in this application has increased by 4.44% compared to the METEOR index obtained by the R2GENCMN method. This indicates that when testing on the test set of the MIMIC-CXR dataset (a large dataset), the text (i.e., the radiology report) generated by the radiology report generation method described in this application is better in semantic expression and the accuracy and flexibility of the overall expression;
[0117] The ROUGE-L index obtained by the radiology report generation method described in this application has increased by 2.53% compared to the ROUGE-L index obtained by the R2GENCMN method. This indicates that when testing on the test set of the MIMIC-CXR dataset (a large dataset), the text (i.e., the radiology report) generated by the radiology report generation method described in this application is closer to the radiology report sample text in terms of sentence structure and coherence, and can better maintain the overall logic and semantic integrity of the radiology report.
Claims
1. A method for generating radiology reports based on semantic alignment, characterized in that: The following steps are involved: S1. Constructing a radiology report generation network, the radiology report generation network includes a multi-scale feature extraction module, a cross-modal semantic alignment module and a Transformer report generation module connected in sequence; the multi-scale feature extraction module includes a convolutional neural network and a multi-scale local sparse attention module connected in sequence; The cross-modal semantic alignment module includes an embedding layer and a multi-head attention module and a gated fusion module connected to the embedding layer, and the input ends of the multi-head attention module and the gated fusion module are also connected to the output end of the multi-scale local sparse attention module; Convolutional neural network is used to extract the features of all patches in radiological images and obtain image region feature maps. ; Multi-scale local sparse attention module for enhancing image region feature maps The features of each patch in , obtain a unimodal image embedding containing multi-scale semantic information E (I) ; The embedding layer is used to perform embedding mapping operations on the radiology report samples to obtain unimodal text embedding E (R) ; Multi-head attention module based on learnable matrix M and unimodal embedding E Perform attention calculation to obtain a unimodal fusion representation containing cross-modal information , the unimodal embedding E Embedding for unimodal images E (I) Or unimodal text embedding E (R) , the single modality fusion representation Single modality fusion image representation Or unimodal fusion text representation ; The gated fusion module is used to embed the single modality into E and single modality fusion representation Fusion is performed to obtain the final unimodal feature F, which is the final image feature Or the final text feature ; The gated fusion module includes two linear layers, the input end of the first linear layer is respectively connected to the output end of the multi-scale local sparse attention module and the output end of the embedding layer, the input end of the second linear layer is connected to the output end of the multi-head attention module, the output ends of the two linear layers are both connected to the first Add layer, the output end of the first Add layer is connected to the Sigmoid activation layer, the output end of the Sigmoid activation layer is respectively connected to the first element-by-element multiplication module and the second element-by-element multiplication module, and the input end of the first element-by-element multiplication module is also connected to the input end of the first linear layer; the input end of the first linear layer and the input end of the second linear layer are both connected to the input end of the second Add layer, and the output end of the second Add layer is sequentially connected to the tanh activation layer and the second element-by-element multiplication module; the output end of the first element-by-element multiplication module and the input end of the second linear layer are both connected to the input end of the third Add layer, and the output end of the third Add layer and the output end of the second element-by-element multiplication module are both connected to the fourth Add layer; S2. Total loss of building a radiology report generation network , where the prediction report sequence And the corresponding radiology report sample sequence The cross entropy loss between And multi-label contrast loss The loss constitutes the total loss of the radiology report generation network ; Computing multi-label contrastive loss When training, use the CheXpert tool to generate pseudo labels for each X-ray image-text pair sample used for training. The pseudo labels are represented as one-hot vectors of multiple labels. The current X-ray image sample is used as the image anchor point. From the text samples with completely different labels from the current X-ray image sample, the sample with the highest cosine similarity to the anchor point is selected as the text negative sample. The current text sample is used as the text anchor point. From the X-ray image samples with completely different labels from the current text sample, the sample with the highest cosine similarity to the anchor point is selected as the image negative sample. Calculate the multi-label contrast loss based on the determined negative samples. Multi-label contrast loss The calculation method of is shown in formula (9); (9) In formula (9), represents cosine similarity, α is the boundary value, which is 0.
7. The final image features representing the image anchor points, The final image features representing the negative samples of the image, The final text feature representing the text anchor, The final text features representing the text negative samples; S3, training the radiology report generation network based on the training set and the total loss of the radiology report generation network to obtain a radiology report generation network model; S4. Input the radiological image for which the radiological report is to be generated into the radiological report generation network model, forward propagate once, and obtain the predicted radiological report.
2. The method for generating radiology reports based on semantic alignment according to claim 1, characterized in that: In step S1, the multi-scale local sparse attention module includes four linear layers, and the inputs of three linear layers are image region feature maps , image region feature map After linear transformation through three linear layers, we get query, key and value respectively. Then, we divide query, key and value into several attention heads respectively. Then, according to the number of attention heads, Split the query into indivual Subvectors, split the key into indivual Subvectors, splitting values into indivual sub-vector; and then through the sliding window operation from Subvector and Select some patches in the sub-vector to calculate the scaled dot product attention and get the weighted output ; Then, all the output The Concat layer is used for splicing, and then the feature fusion is performed through the last linear layer to obtain the image region features containing rich multi-scale semantic information. , and then, the image region features containing multi-scale semantic information Flattening in the spatial dimension to obtain a unimodal image embedding containing rich multi-scale semantic information .
3. The method for generating radiology reports based on semantic alignment according to claim 2, characterized in that: In step S1, a sliding window operation is performed from Subvector and Selecting some patches in the sub-vector for calculating the scaled dot product attention includes the following steps: Subvector and In the sub-vector, each patch is centered and the surrounding patches are selected by setting the hyperparameters of the sliding window operation, namely the convolution kernel size k and the expansion rate r, and a new key is obtained. and the new value ; Different attention heads are set with different expansion rates; then, calculate Subvectors and new keys The similarity between them is then divided by the scaling factor to obtain the scaled similarity. The scaled similarity value is converted into an attention weight through a softmax operation; then, the attention weight is combined with the new value Multiply them together to get the weighted output , the weighted output That is, the image features containing local semantic information are used to calculate the scaled dot product attention.
4. The method for generating radiology reports based on semantic alignment according to claim 1, characterized in that: In step S1, the Transformer report generation module includes a coding unit, a decoding unit and a prediction head connected in sequence, and the coding unit and the decoding unit are also connected to the output end of the gated fusion module in the cross-modal semantic alignment module; wherein the coding unit is used to output the final image features of the gated fusion module. Encode to obtain the encoded image representation L; the decoding unit is used to decode the image representation L and the final text features output by the current gated fusion module Perform a decoding operation and output a decoded representation; the prediction head is used to perform a linear transformation on the decoded representation output by the decoding unit and apply a Softmax operation to output the character with the highest probability, which is the predicted next character.
5. The method for generating radiology reports based on semantic alignment according to claim 4, characterized in that: In step S1, the encoding unit includes three sequentially connected Transformer encoder layers, each of which contains a multi-head self-attention mechanism and a feedforward neural network. The parameters of each Transformer encoder layer are set to 8 attention heads and 512-dimensional hidden states.
6. The method for generating radiology reports based on semantic alignment according to claim 4, characterized in that: In step S1, the decoding unit includes three sequentially connected Transformer decoder layers, each of which includes two parts: a multi-head self-attention mechanism and a feedforward neural network. The parameters of each Transformer decoder layer are set to 8 attention heads and 512-dimensional hidden states.
7. The method for generating radiology reports based on semantic alignment according to claim 1, characterized in that: In step S3, the images in the training set and their corresponding radiology report samples are input into the radiology report generation network for forward propagation, and the total loss of the radiology report generation network is calculated. , and in the total loss Back propagation is performed under the guidance of , the weight parameters of the radiology report generation network are updated, the training process of one epoch is completed, and the iterative training is repeated to obtain the radiology report generation network model.
8. The method for generating radiology reports based on semantic alignment according to claim 7, characterized in that: In step S3, the training set is the training set in the IU-Xray dataset or the training set in the MIMIC-CXR dataset; when the radiology report generation network is trained based on the training set in the IU-Xray dataset, the paired images of the IU-Xray dataset are input into the radiology report generation network, and μ in the most similar first μ vectors is set to 64; when the radiology report generation network is trained based on the training set in the MIMIC-CXR dataset, a single image of the MIMIC-CXR dataset is input into the radiology report generation network, and μ in the most similar first μ vectors is set to 128.
Citation Information
Patent Citations
Multi-modal text abstract system based on dependence gating fusion mechanism
CN113609285A
X-ray chest radiography diagnosis report generation method based on multi-task multi-mode deep learning
CN115223678A