Implementation Method of an Automatic Generation System for Fetal Heart Ultrasound Image Diagnostic Reports
By applying attention mechanism and contrast learning in fetal heart ultrasound images, the problems of unclear image texture and high noise are solved, and more accurate diagnostic report generation is achieved, improving the accuracy and efficiency of diagnosis.
Patent Information
- Application Number
- CN202210210339.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-04
AI Technical Summary
Fetal heart ultrasound images have the characteristics of unclear texture, high noise, and single background information, which makes it challenging to generate diagnostic reports. Especially when multiple targets exist in the image and the shape is not fixed, it is difficult to accurately judge the organs, and the prior art is difficult to effectively obtain useful information in the image.
The attention mechanism is used to automatically focus on the key areas of the image, and the effect of the attention mechanism is enhanced through multiple interactions, combined with contrast learning, and to enhance the representation ability of the image, thereby generating a longer diagnostic report.
By automatically focusing on the key areas of the image and enhancing attention mechanism, useful information in the image can be obtained more accurately, the accuracy of generation of diagnostic reports can be improved, the loss of context information can be reduced, and the doctor can effectively assist in diagnosis.
Smart Images

Figure CN114664404B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image topic generation in artificial intelligence, and relates to a method for implementing an automatic generation system for fetal cardiac ultrasound image diagnostic reports. Background Art
[0002] Today, with the increasing population, the physical health of newborns is very important to people. Many fetuses die at birth, and most of them die from congenital heart diseases. Therefore, the diagnosis of the fetal heart is crucial. Doctors can understand the health problems of the fetus through ultrasound images during pregnancy and write corresponding diagnostic reports. For experienced doctors, writing diagnostic reports is a very boring task. For inexperienced doctors, such work is not only boring but also error-prone. Therefore, automatically generating text for fetal cardiac ultrasound images is a great convenience, which can not only reduce the unnecessary workload of doctors, reduce the probability of errors of inexperienced doctors, but also reduce the waiting time of patients.
[0003] Due to the characteristics of unclear texture, high noise, and single background information in fetal cardiac ultrasound images, generating diagnostic reports for fetal cardiac ultrasound images poses a great challenge. When there are multiple targets in a fetal cardiac ultrasound image, if the shapes are not fixed, due to the different instrument angles during the imaging of fetal cardiac ultrasound images, the size and shape of the same organ may not be the same in different images, and the shapes of different organs may be the same in the same image, making it difficult to determine whether they are the same organ. Compared with the concise description of natural images, many diagnostic reports in fetal cardiac ultrasound images are relatively long, tend to be templated, and most fetal cardiac ultrasound images are for a small number of disease categories, with small differences between images. Simple supervised learning methods cannot well fit the true distribution of the data, so that useful information in the images cannot be well obtained. Summary of the Invention
[0004] To solve the above problems, a method for implementing an automatic generation system for fetal cardiac ultrasound image diagnostic reports is proposed. This method uses an attention mechanism to focus on key regions in the image and enhances the attention through multiple interactions, which is beneficial to the model's ability to identify key regions in the image and generate long diagnostic reports. At the same time, contrastive learning is used to increase the between-class differences and decrease the within-class differences of the images, enhance the representation ability of the images, better obtain useful information in the images, and thus improve the performance of the overall model. Finally, the model is encapsulated into the system, and doctors can perform auxiliary diagnosis through a simple interactive interface.
[0005] The present invention aims to solve the above problems in the prior art and proposes a method for implementing an automatic generation system for fetal cardiac ultrasound image diagnostic reports, including the following steps:
[0006] 1) After performing two types of data augmentation on the input fetal heart ultrasound image, it is encoded into a feature representation, and the contrast loss is calculated based on the features.
[0007] 2) Based on the image encoding features obtained in step 1), the global image features and local attention features are calculated using the attention mechanism.
[0008] 3) Combine the local attention features obtained in step 2), the word vectors of the true sentence corresponding to the image, and the context vector of the previous time step of the decoder to obtain new features.
[0009] 4) Interact the features obtained in step 3) with the hidden state of the decoder, and input the interacted input and hidden state into the decoder LSTM.
[0010] 5) Input the global image features calculated in step 2) and the hidden state of the current time step generated by the decoder into the attention block together. Calculate new features using the GLU activation function for the features generated by the attention block and the hidden state of the decoder, which is called context information. Use the context information to predict and generate words, and calculate the cross-entropy loss based on the true words and the generated words.
[0011] 6) Loop through steps 1) to 5) to train the model.
[0012] 7) Package the model trained in step 6) into an automatic fetal heart ultrasound image diagnosis report generation system.
[0013] The present invention has the following beneficial technical effects:
[0014] The method proposed by the present invention can automatically focus on the key areas of the image using the attention mechanism, make full use of the high-order interaction information between the image and sentence modalities, and use a multi-order interaction method to enhance the attention, reducing the loss of context information, so as to solve the problem of some diagnosis reports being too long. For fetal heart ultrasound images, most of them are images for a small number of disease categories, with small differences between images, and the boundaries between inter-class differences and intra-class differences are not clear. Using the idea of contrastive learning, a contrast loss is constructed for training, making samples with high similarity close and samples with low similarity far away, thereby improving the encoder's representation ability of the image and the accuracy of the overall model in generating sentences. It can more effectively assist doctors in diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the overall framework diagram of the present invention;
[0016] Figure 2 is the schematic diagram of constructing the contrast loss;
[0017] Figure 3It is a diagram of the operation of the decoder at each time step.
[0018] Figure 4 It is a schematic diagram of the data preprocessing of the automatic fetal cardiac ultrasound image diagnosis report generation system.
[0019] Figure 5 It is a schematic diagram of the diagnostic ultrasound image of the automatic fetal cardiac ultrasound image diagnosis report generation system.
[0020] Figure 6 It is a complete diagram of the diagnostic report generated by the automatic fetal cardiac ultrasound image diagnosis report generation system.
[0021] Figure 7 It is a flowchart of the automatic fetal cardiac ultrasound image diagnosis report generation system. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0023] As Figure 1 shown, a method for implementing an automatic fetal cardiac ultrasound image diagnosis report generation system includes:
[0024] 1), Construct a fetal cardiac ultrasound image data set, and perform necessary preprocessing operations on a batch of fetal cardiac ultrasound images for training.
[0025] 2), Encode the image into a feature representation.
[0026] 3), As Figure 2 shown, construct and calculate the contrast loss.
[0027] 4), According to the image encoding representation obtained in step 2), use the attention mechanism to calculate the global image feature and the local attention feature.
[0028] 5), Combine the local attention feature obtained in step 4), the word vector of the image corresponding real sentence, and the context vector of the previous time step of the decoder to obtain a new feature.
[0029] 6), As Figure 3 shown, interact the feature obtained in step 5) with the hidden state of the decoder to reduce the loss of context information and better model the long-distance semantics. Input the interacted input and hidden state into the decoder LSTM.
[0030] 7) The global image features calculated in step 4) and the hidden state of the current time step generated by the decoder are fed into the attention block together to better utilize the high-order interaction between image information and semantic information. The features generated by the attention block and the hidden state of the decoder are used to calculate new features using the GLU activation function, which are called context information. And the context information is used to predict and generate words.
[0031] 8) Combine the contrastive loss in step 3) with the cross-entropy loss between the predicted words and the true words in step 7) to train the model.
[0032] 9) Package the trained model into an automatic fetal heart ultrasound image diagnosis report generation system. The system mainly includes two functions, namely image enhancement and report generation. When using the system, first selectively preprocess the image, and then click the report generation button of the system. The system will use model inference to generate a diagnosis report corresponding to the fetal heart ultrasound image, which is displayed in the lower right corner of the system, and the report and image can be manually saved.
[0033] 10) The system usage includes: as Figure 4 shown, first, the image can be selectively preprocessed. After preprocessing, as Figure 5 shown, input the image, click to generate a report, and the system will perform inference according to steps 2), 4), 5), 6), and 7), and display the generated report in the box in the lower right corner of the system. As Figure 5 shown, the generated diagnosis report and image are saved together into a doc document for the user to print. As Figure 6 shown, the generated diagnosis report result and image are saved together into a file. Figure 7 It is a flowchart for using the system to generate a diagnosis report corresponding to the image.
[0034] Furthermore, in step 1), two data augmentation methods are first used to augment the images with a batch size of B for training to obtain new images {f aug1 (I 1 ),..., f aug1 (I B ), f aug2 (I 1 ),..., f aug2 (I B )} using two different data augmentation methods. The data augmentation methods f aug1 and f aug2 include methods such as random rotation, horizontal flipping, and vertical flipping.
[0035] Further, the specific operation of step 2) is as follows: Take the enhanced image obtained in step 1) as a new batch of images, with the batch size doubled. Then input these images into the encoder, and the encoder uses the residual network Resnet-101:
[0036] V i = Resnet101(I i )
[0037] Obtain the encoded features of 2B images Among them, the image feature indicates that N groups of features are obtained after image encoding, and v k represents the k-th group of features of the image, represents the dimension of each group of image features, and D v is the dimension size.
[0038] Further, the specific details of step 3) are as follows: For each image V in the image features with a batch size of 2B i , use global average pooling operation, and then use a fully connected layer to map it to new image features:
[0039]
[0040] Among them indicates that the dimension size of the feature is n, and fc represents the fully connected layer. For the features (z i , z j ) after two different data augmentations of the same image, calculate the contrastive loss using the following formula:
[0041]
[0042] Among them, τ is the hyperparameter temperature coefficient, which can adjust the attention to difficult samples. The smaller the temperature coefficient, the more easily the original sample and the most similar negative sample can be separated. sim(z i , z j ) is the cosine similarity of the two feature vectors z i , z j , and the calculation formula is as follows:
[0043]
[0044] respectively represent the k-th value of the image feature vectors z i , z j .
[0045] The total contrastive loss function is the sum of the contrastive loss functions of this batch of images:
[0046]
[0047] Among them, \(k\in[1,2,\cdots,B]\), the \(k\)-th image and the \((k + B)\)-th image are two different images obtained by data augmentation of the same image. \(l\) represents the loss function.
[0048] Furthermore, starting from step 4), all subsequent descriptions are for one image, and the operations on other images in a batch of images are the same. For a batch of image features in an image feature \(V\) i processing is performed to obtain the initial query \(Q\) of the attention block (0) , key \(K\) (0) , value \(V\) (0) , the value of \(V\) is as shown in step 2): i
[0049]
[0050] \(K\) (0) \(=\) \(V\) i
[0051] \(V\) (0) \(=\) \(V\) i
[0052] Then, the image attention feature is calculated using the attention mechanism:
[0053]
[0054] Among them, \(M\) is the number of stacked attention blocks. The calculation process \(F\) X-Linear \((K, V, Q)\) of each attention block is calculated as follows:
[0055]
[0056]
[0057] \(\beta\) s \(=\) softmax(\(B\) s )
[0058]
[0059]
[0060]
[0061] Among them, \(W\) k , \(W\) b , \(W\) e , \(W\) v , They are all embedding matrices, k i represents the i-th key, The query Q and each key k i The joint bilinear query-key representation between them, σ is the activation function, B s is the transformed bilinear query-key representation, is the i-th element of B s and and represent the dimension of the embedding, β s is the distribution of B s and is the i-th element of β s and is The globally computed channel descriptor, β c is the attention distribution over the channels, v i is the i-th value of the value sequence V, ⊙ represents element-wise multiplication;
[0062] The calculation formula for stacking attention blocks is as follows:
[0063]
[0064]
[0065]
[0066] where represents the embedding matrix, m = {1, 2, 3,..., M + 1}, and M represents the number of stacked attention modules. represents the feature obtained after stacking m attention blocks, represents the initial key sequence K (0) After stacking m attention blocks, K (m) is the i-th element of represents the initial value sequence V (0) After stacking m attention blocks, V (m) is the i-th element of
[0067] After stacking M attention blocks, the global image feature v global and the local attention feature v att are as follows:
[0068]
[0069]
[0070] where is the embedding matrix and D is the dimension size.
[0071] Further, the specific details in step 5) are as follows: for the current time step t of the decoder LSTM, according to the local attention feature v calculated in step 4) att and the word vector at the current time step, the input of the decoder LSTM at the current time step is:
[0072] x t = [v att + c′ t-1 , e t
[0073] Further, the calculation process of the multiple interactions between the input and the hidden state in step 6) is as follows:
[0074]
[0075]
[0076]
[0077]
[0078] where e t is the word vector at the current time step, ⊙ represents element-wise multiplication, and c′ t-1 is the context information of the previous time step. is the embedding matrix, D x , D h is the dimension of the input feature x t and the hidden state h t-1 , x t and h t-1 are the current input feature and the hidden state of the LSTM at the previous time step respectively, is equivalent to x t , is equivalent to h t-1 . When t is 0, that is, at the first time step, h -1 , c′ -1 are initialized. The input feature is multiplied element-wise with the embedding of the latest calculated hidden state to obtain a new feature each time, and the hidden state is multiplied element-wise with the embedding of the latest calculated feature to obtain a new hidden state each time. and are the final feature and hidden state after calculation.
[0079] The input feature and hidden state after multiple interactions are input into the decoder LSTM:
[0080]
[0081] where c t Represents the cell state of the LSTM at the t-th time step.
[0082] Further, the specific operation of step 7) is as follows: Use the hidden state h of the decoder at the current time step t as the query of the attention block, and the global image feature v calculated in step 4) global as both the key and the value at the same time, and input them into the attention block to calculate a new feature Adjust using a fully connected layer the dimension of and use the GLU activation function together with the hidden state of the current time step to obtain the context information c' of the current time step t :
[0083] c' t = GLU([W de F X-Linear (v global , v global , h t ), h t )
[0084] where W de is the embedding matrix, and F X-Linear represents the calculation of an X-Linear attention block.
[0085] Use the context information to predict the distribution of the output word vector at the current time step, that is, the probability of each word in the generated vector:
[0086] w t = Softmax(W |Σ| c' t )
[0087] where W |Σ| is the embedding matrix, and |Σ| is the size of the vocabulary. Finally, directly take the word with the highest probability as the output of the current time step.
[0088] During the working process of the decoder LSTM, the word vector at the first time step represents a " <start>The special character of " is represented by the word vector generated at the last time step <end>”. Repeat the process of generating words in steps 4) to 7) until “ <end>"So far, the sentence generation is completed.
[0089] The cross-entropy loss function for training is as follows:
[0090]
[0091] Among them, represents the true sentence composed of the previous t - 1 true words to generate the true word at the current time step probability. T represents the length of the true sentence.
[0092] Furthermore, the specific details of step 8) are as follows: Combining the loss functions of step 3) and step 7), construct the total loss function for training the entire model. Since the importance of the two losses in training may vary, the combined total loss function is as follows:
[0093] L all = αL c + βL(θ)
[0094] where the hyperparameters α and β are the weights of the two losses in training.
[0095] Furthermore, in step 9), the trained model is encapsulated into the fetal heart ultrasound image diagnosis report automatic generation system. When the user clicks to generate a report, the operations performed by the system are as follows: First, use Resnet101 as the encoder to extract a set of regional features of the image, obtain the global feature and the attention feature through the attention module. The attention feature is fused with the word embedding vector and the context vector of the previous time step at each time step as the input of the decoder. At each time step, the input interacts with the hidden state of the previous time step multiple times to better obtain the information of the previous time steps. The fused feature and the new hidden state are input into the LSTM decoder. The output of the LSTM at each time step and the global feature of the image are calculated through the attention module and use the GLU activation function to obtain the context information. The context information is decoded into the predicted word through the fully connected layer. Until the predicted word is the special character " <end>Up to this point, the reasoning process ends. After the reasoning process ends, all the generated words are combined into sentences and displayed on the system interface.
[0096] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.< / end> < / end> < / end> < / start>
Claims
1. An implementation method of an automatic generation system for fetal heart ultrasound image diagnosis reports, characterized in that, it includes the following steps: 1) Perform two types of data augmentation on the input fetal heart ultrasound images and encode them into feature representations, and calculate the contrast loss according to the features, where k ∈ [1, 2,..., B], the k-th image and the (k + B)-th image are two different data-augmented images of the same image, l represents the loss function, and B represents the number of images; 2) According to the image encoding features obtained in step 1), use the attention mechanism to calculate the global image features and local attention features; 3) Combine the local attention features obtained in step 2), the word vectors of the real sentence corresponding to the image, and the context vector of the previous time step of the decoder to obtain a new input feature, x t = [v att + c t ′ -1 , e t , where v att represents the local attention feature, e t is the word vector of the current time step, and c t ′ -1 is the context information of the previous time step; 4) Interact the input features obtained in step 3) with the hidden state of the decoder, and input the new input features and new hidden state obtained after the interaction into the decoder LSTM; interact the input x multiple times t with the hidden state h of the previous time step t-1 The calculation process is as follows: Among them, ⊙ represents element-wise product, is the embedding matrix, D x , D h is the dimension of the input feature x t and the hidden state h t-1 ; x t and h t-1 are the current input feature and the hidden state of the previous time step of the LSTM respectively, is equal to x t , is equal to h t-1 ; After the above calculations, the obtained and are the final feature and hidden state after calculation; The input feature and the hidden state after multiple interactions are input into the decoder LSTM to participate in decoding; 5) Input the global image features calculated in step 2) and the hidden state of the current time step generated by the decoder into the attention block together, and use the GLU activation function to calculate new features from the features generated by the attention block and the hidden state of the decoder, which is called context information. Use the context information to predict and generate words, and calculate the cross-entropy loss according to the real words and the generated words; Take the hidden state h of the current time step of the decoder t as the query of the attention block, and the global image feature v global as both the key and the value at the same time, and input them into the attention block to calculate a new feature Adjust using a fully connected layer After the dimension of, use the GLU activation function together with the hidden state of the current time step to obtain the context information c t ′, and use the context information to predict the distribution of the output word vector of the current time step; c t ′ = GLU([W de F X-Linear (v global , v global , h t ), h t ) Among them, W de is the embedding matrix, and F X-Linear represents the calculation of an X-Linear attention block; The calculating the cross-entropy loss according to the real words and the generated words Among them, represents the true sentence formed by the previous t - 1 true words to generate the true word at the current time step with probability, where T represents the length of the true sentence; 6) Repeat steps 1) to 5) to train the model; 7) Package the model trained in step 6) into the automatic generation system for fetal heart ultrasound image diagnosis reports.
2. The implementation method of an automatic generation system for fetal heart ultrasound image diagnosis reports according to claim 1, characterized in that: The performing two types of data augmentation on the input fetal heart ultrasound images and encoding them into feature representations in step 1) includes: For a batch of images with a batch size of B during training Two different data augmentation methods are used to augment the images to obtain new images {f aug1 (I 1 ),..., f aug1 (I B ), f aug2 (I 1 ),..., f aug2 (I B )}. The data augmentation methods f aug1 and f aug2 include random rotation, horizontal flipping, and vertical flipping. Then, the obtained new batch of images, with the batch size doubled, are fed into the encoder, and the encoder uses the residual network Resnet-101: V i = Resnet101(I i ) Obtain the encoded features of 2B images Among them, the image features indicate that N groups of features are obtained after image encoding, v k represents the k-th group of features of the image, represents the dimension of each group of image features, D v is the dimension size.
3. The implementation method of an automatic generation system for fetal heart ultrasound image diagnosis reports according to claim 2, characterized in that: Step 2) Calculate the global image feature v global and the local attention feature v att , specifically: The calculation formula for stacking attention blocks is as follows: Among them represents the embedding matrix, where \(m = \{1, 2, 3, \cdots, M + 1\}\), and \(M\) represents the number of stacked attention modules represents the feature obtained after stacking \(m\) attention blocks represents the initial key sequence \(K\) (0) After stacking \(m\) attention blocks, \(K\) is obtained (m) The \(i\)-th element of represents the initial value sequence \(V\) (0) After stacking \(m\) attention blocks, \(V\) is obtained (m) The \(i\)-th element of; After stacking M superimposed attention blocks, the global image feature v is obtained global and the local attention feature v att as follows: Among them, is the embedding matrix, and D is the dimension size.
4. The implementation method of an automatic generation system for fetal heart ultrasound image diagnosis reports according to any one of claims 1-3, characterized in that: Construct the total loss function for training the entire model as: Among them, the hyperparameters α and β are the weights of the two losses during training. On the left is the total contrastive loss \(k\in[1,2,\cdots,B]\), where the \(k\)-th image and the \((k + B)\)-th image are two different data augmentations of the same image. On the right is the cross-entropy loss. represents the true sentence composed of the first \(t - 1\) true words to generate the true word at the current time step probability, and \(T\) represents the length of the true sentence.
5. The implementation method of an automatic generation system for fetal heart ultrasound image diagnosis reports according to claim 1, characterized in that: In step 7), the trained model is packaged into the system. The system mainly includes two functions, namely image enhancement and report generation. When using the system, first selectively preprocess the image, and then click the report generation button of the system. The system will use model inference to generate a diagnosis report corresponding to the fetal heart ultrasound image, which is displayed in the lower right corner of the system, and the report and image can be saved manually.