Intelligent report interpretation method and system based on deep learning
The image and text features in the physical examination report are extracted through deep learning technology, and the feature weights are dynamically allocated using the attention fusion layer, and interpreted in combination with the knowledge graph, which solves the problem of subjective differences and inefficiency in traditional manual interpretation, and achieves a more efficient and accurate interpretation of the physical examination report.
Patent Information
- Application Number
- CN202510299297.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The interpretation methods of traditional physical examination reports rely on manual labor, have subjective differences and are inefficient, making it difficult to cope with a large number of reporting needs, affecting the quality of medical services.
Using a deep learning-based intelligent report interpretation method, the features of physical examination items are extracted through image encoder and text encoder, and the feature weights are dynamically allocated using the attention fusion layer, and interpreted in combination with the knowledge graph.
It improves the accuracy and efficiency of interpretation of physical examination reports, can better capture the details and deep semantics in multimodal data, and improves the quality of medical decision-making.
Smart Images

Figure CN120144783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly to an intelligent report interpretation method and system based on deep learning. Background Art
[0002] Traditional physical examination report interpretation methods usually involve doctors or nurses with certain medical knowledge interpreting the physical examination report item by item to analyze the patient's health status. Subjective judgments by doctors may lead to differences in interpretation results, reducing the reliability of the physical examination report. In addition, manual interpretation is inefficient and difficult to meet the demand for a large number of reports, affecting the quality of medical services.
[0003] Therefore, it is necessary to provide a solution that can improve the accuracy and efficiency of physical examination report interpretation. Summary of the Invention
[0004] To solve the above problems, the present invention provides an intelligent report interpretation method and system based on deep learning, which can improve the accuracy and efficiency of physical examination report interpretation.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] On the one hand, an embodiment of the present invention provides an intelligent report interpretation method based on deep learning, and the method includes the following steps:
[0007] Obtain a physical examination report containing multiple physical examination items, and identify the text information and image information in the physical examination report;
[0008] Input the text information and image information into a report recognition model including an image encoder, a text encoder, an attention fusion layer, and a classifier. Extract the image features of the physical examination items through the image encoder, and extract the text features of the physical examination items through the text encoder; Dynamically allocate the weights of the image features and text features through the attention fusion layer. After aligning the weighted image features to the dimension of the text features, splice them with the weighted text features to obtain fused features; Classify the fused features through the classifier to identify the concerned items in the physical examination report, and output the health categories corresponding to the concerned items;
[0009] Input the concerned items and the corresponding health categories into the knowledge graph to obtain the interpretation result of the physical examination report.
[0010] Optionally, the report recognition model is trained in the following manner:
[0011] Construct a multimodal model with the same architecture as the report recognition model;
[0012] Obtain a sample set, input the sample set into the multi-modal model, freeze all layers of the image encoder and the text encoder, iterate the multi-modal model, and gradually adjust the learning rate according to the number of iterations and the first loss value of the multi-modal model;
[0013] After the training parameters of the multi-modal model reach the warm-up steps, layer by layer unfreeze the image encoder and the text encoder according to the second loss value of the multi-modal model, and gradually adjust the learning rate according to the number of iterations and the second loss value of the multi-modal model until the second loss value of the multi-modal model is lower than the second loss threshold, obtaining a report recognition model; wherein, the input data of the sample set is a sample report containing text and image information, and the output data is the items of interest and the corresponding health categories in the sample report.
[0014] Optionally, the dynamic allocation of the weights of the image features and the text features through the attention fusion layer includes:
[0015] Calculate the attention weight α of the text features to the image features through cross-attention ij , and the attention weight β of the image feature V to the text feature T kl :
[0016] Use the Softmax function to normalize the attention weight α ij and the attention weight β kl to obtain the image weight and the text weight.
[0017] Optionally, the attention weight of the text features to the image features is calculated by the following formula:
[0018]
[0019] wherein, T is the text feature, V is the image feature, α ij represents the attention weight of the i-th text token to the j-th image patch, V is the image feature matrix, T is the text feature matrix, Q v is the mapping matrix from the image feature to the text feature query, K t is the mapping matrix from the text feature to the image feature key value, V ik is the mapping matrix from the image feature to the text feature query, Q v , K t are learnable matrices, k is the index variable of the image patch, d v is the dimension of the image feature, d t is the dimension of the text feature.
[0020] Optionally, the attention weight of the image feature V to the text feature T is calculated by the following formula:
[0021]
[0022] where β kl represents the attention weight of the k-th image patch to the l-th text token, and Q t is the mapping matrix from text features to image feature queries, and K v is the mapping matrix from image features to text feature key values, and Q t , K v are learnable matrices, and m is the index variable of the text token.
[0023] Optionally, the calculation formula for the learning rate is:
[0024] cut
[0025] lr(t) new = lr(t)·δ;
[0026]
[0027] where t is the current training step, lr(t) new is the learning rate for the t-th iteration, lr(t) is the baseline learning rate for the t-th iteration, δ and γ are adjustment factors, 0 < δ < 1, 0 < γ < 1, cut is the number of iterations when the loss value decrease amplitude is continuously lower than the amplitude threshold, lr base is the initial learning rate, w s is the warm-up step, t s is the total number of training steps, L1 represents the first loss value, and L2 represents the second loss value.
[0028] Optionally, the first loss value of the multimodal model is calculated by the following formula:
[0029] L1 = λ·L align +(1 - λ)·L class ;
[0030]
[0031] where L1 align represents the cross-attention alignment loss, A i,i represents the cross-attention matrix of text features to image features, A i,i = Softmax(α ij ); N ′ is the batch size, N is the number of patches of image features, M is the number of tokens of text features, L1 class represents the first classification loss value, Y ih represents the one-hot encoding of the true label, is the probability that the i-th sample in the multimodal model training belongs to class h.
[0032] Optionally, the second loss value of the multimodal model is calculated by the following formula:
[0033] L2 = -L align + L2 class + L2 KL ;
[0034]
[0035] where L2 represents the second loss value, L2 var represents the variance penalty value, L2 class represents the second classification loss value, L2 KL represents the KL divergence value; q represents the variational distribution of the parameter θ, θ represents the hyperparameter of the model, q(θ) represents the posterior distribution of the parameter θ, represents the variance parameter, E q(θ) represents the expected value of the parameter θ, Z (y=h) represents the indicator function, p(y = h|x,θ) represents the probability that the input fused feature x belongs to the healthy class h, y represents the true label, H represents the total number of healthy classes, σ 2 represents the variance of the posterior distribution q(θ), p(θ) represents the prior distribution of the parameter θ, represents the variance of the prior distribution p(θ).
[0036] On the other hand, an embodiment of the present invention provides an intelligent report interpretation system based on deep learning, including:
[0037] At least one processor;
[0038] At least one memory for storing at least one program;
[0039] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0040] The beneficial effects of the present invention are as follows: The present invention discloses an intelligent report interpretation method and system based on deep learning. The present invention extracts the image features of physical examination items through an image encoder and extracts the text features of physical examination items through a text encoder; through an attention fusion layer, the weights of the image features and text features are dynamically allocated. By dynamically allocating weights, more attention can be paid to key information, thereby improving the accuracy and efficiency of report interpretation. Through multimodal fusion, not only can details be captured, but also deep semantic meanings can be understood to achieve accurate interpretation and improve the quality of medical decision-making. After aligning the weighted image features to the dimension of the text features, they are concatenated with the weighted text features to obtain fused features; through a classifier, the fused features are classified to identify the items of concern in the physical examination report, and the health categories corresponding to the items of concern are output. The present invention significantly improves the accuracy and efficiency of report interpretation through multimodal feature fusion and loss optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 FIG. is a schematic flowchart of an intelligent report interpretation method based on deep learning according to an embodiment of the present invention;
[0043] Figure 2 FIG. is a schematic structural diagram of an intelligent report interpretation system based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following will clearly and completely describe the concept, specific structure and technical effects generated by the present invention disclosed in combination with the embodiments and the drawings, so as to fully understand the purpose, solution and effects of the present invention disclosed. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0045] Refer to Figure 1 as Figure 1 shown, an intelligent report interpretation method based on deep learning provided by an embodiment of the present invention includes the following steps:
[0046] S100, obtain a physical examination report containing multiple physical examination items, and identify the text information and image information in the physical examination report;
[0047] Specifically, the physical examination report can be uploaded and submitted in formats such as scanned images and PDFs. After receiving the physical examination report uploaded by the user, OCR is used to recognize the text information in the physical examination report, keywords are used to match the test names in the text information to obtain the result data corresponding to the item names. At the same time, image recognition technology is used to extract the key indicators in the image information to ensure the comprehensiveness and accuracy of the data. Then, the obtained data is preprocessed, including data cleaning, format unification, and feature extraction, to prepare for the input of the subsequent deep learning model. After the standardized data (text information and image information) in the physical examination report, it is convenient to use the deep learning model to perform feature learning and pattern recognition on the data.
[0048] S200, input the text information and image information into a report recognition model that includes an image encoder, a text encoder, an attention fusion layer, and a classifier. Extract the image features of the physical examination items through the image encoder, and extract the text features of the physical examination items through the text encoder; dynamically allocate the weights of the image features and text features through the attention fusion layer. After aligning the weighted image features to the dimension of the text features, splice them with the weighted text features to obtain the fusion features; perform classification processing on the fusion features through the classifier to identify the items of concern in the physical examination report and output the health category corresponding to the items of concern.
[0049] Specifically, the multi-modal model adopts a dual-branch parallel encoding and cross-modal attention fusion architecture. The image encoder is based on the ViT model and is used to extract the global semantic features of the image information; the text encoder is based on the pre-trained language model of the BERT model and is used to capture the context semantic information of the text information; the fusion layer of cross-modal attention realizes feature interaction between modalities through dynamic weight allocation; input the fused features into the classifier to output the final health category. The health category is divided into five levels, namely normal, mild abnormality, moderate abnormality, severe abnormality, and critical. According to the classification results, the system automatically marks the physical examination items that need to be focused on. The items of concern include the physical examination items below the preset health category. After forming the items of concern by combining multiple key physical examination items, match the corresponding health category through the classifier. For example, if the blood pressure index and cholesterol index show moderate abnormality, then mark the blood pressure index and cholesterol index as the physical examination items that need to be focused on.
[0050] S300, input the items of concern and the corresponding health category into the knowledge graph to obtain the interpretation result of the physical examination report.
[0051] Specifically, the knowledge graph is constructed based on a medical knowledge base, capable of associating project information with disease knowledge and providing accurate interpretations. The interpretation results include disease probabilities, influencing factors, and preventive measures, helping users comprehensively understand their own health status and take effective measures for health management. Exemplarily, taking blood pressure indicators and cholesterol indicators as the concerned projects and associating them with cardiovascular diseases, the interpretation results show a moderate abnormal risk, suggesting regular monitoring, dietary adjustment, and increased exercise to prevent the occurrence of cardiovascular diseases. Through the intelligent report interpretation method, users can not only quickly master the key information of the physical examination but also obtain personalized health guidance, improving the efficiency of health management. The present invention effectively integrates deep learning technology and expert knowledge to provide accurate and convenient health services for users.
[0052] As an improvement of the above embodiment, the report recognition model is trained in the following manner:
[0053] Construct a multi-modal model with the same architecture as the report recognition model;
[0054] The report recognition model is obtained based on the training of the multi-modal model. The multi-modal model includes an image encoder, a text encoder, an attention fusion layer, and a classifier. The attention fusion layer is used to achieve feature interaction between text information and image information by adopting dynamic weight allocation; the classifier is used to output multiple health categories;
[0055] Obtain a sample set, input the sample set into the multi-modal model, freeze all layers of the image encoder and the text encoder, iterate the multi-modal model, and gradually adjust the learning rate according to the number of iterations and the first loss value of the multi-modal model;
[0056] After the training parameters of the multi-modal model reach the warm-up steps, layer by layer unfreeze the image encoder and the text encoder according to the second loss value of the multi-modal model, and gradually adjust the learning rate according to the number of iterations and the second loss value of the multi-modal model until the second loss value of the multi-modal model is lower than the second loss threshold, obtaining the report recognition model; wherein, the input data of the sample set is a sample report containing text and image information, and the output data is the concerned projects and corresponding health categories in the sample report.
[0057] It should be noted that in the warm-up stage, freeze all layers of the image encoder and the text encoder, and iterate to train the cross-attention layer and the classifier until the loss value of the multi-modal model is lower than the first loss threshold; in the unfreeze stage, iterate to train the image encoder, the text encoder, the attention fusion layer, and the classifier, gradually unfreeze the image encoder and the text encoder, and gradually unfreeze each layer of the image encoder and the text encoder in the order from high to low level by level. When the loss value of the multi-modal model is lower than the second loss threshold, complete the training of the multi-modal model and obtain the report recognition model.
[0058] Specifically, the image encoder uses the ViT model, and the text encoder uses the BERT model. The image features are extracted by the ViT model, and the text features are extracted by the BERT model. If the text features contain keywords related to the health category (such as the test results), the text weight automatically increases, and more attention is paid to the text semantics. If there are feature areas strongly related to the health category (such as lesion areas) in the image, the image weight increases. In this way, dynamic weights are adaptively assigned to images and texts, improving the robustness and generalization of cross-modal tasks. The hyperparameters of the initialization multimodal model include: setting the pre-training weights of the ViT model and the BERT model, setting the number of attention heads of the cross-attention layer; the input dimension of the classifier is the total dimension of the concatenation of the output features of the ViT model and the BERT model, and the output dimension is the total number of categories of the health category; the setting of the pre-training weights of the ViT model refers to ViTB / 16, and the setting of the pre-training weights of the BERT model refers to BERTbase, retaining its underlying feature extraction capabilities. Only the parameters of the cross-attention layer and the parameters of the classifier are fine-tuned to reduce training time and computational cost. By loading the pre-trained ViT model and BERT model, the parameters of all layers are frozen to prevent them from being updated during training.
[0059] The ViT model normalizes image information to 256×256 pixels and normalizes it to the interval [0,1]. It introduces local window attention through the Transformer architecture of the ViT model to capture global features while retaining local details. The BERT model performs word segmentation and word embedding on text information to generate serialized input.
[0060] The cross-attention layer captures the correlation between modalities and enhances feature fusion; the classifier accurately determines the health level, derives the focus items and corresponding health categories, and improves prediction accuracy. The multi-stage training strategy optimizes model performance and ensures efficient and stable health assessment.
[0061] Through a learnable attention weight matrix, the contribution of features between modalities is adaptively allocated (images have higher weights in early layers, and text has increased weights in later layers), significantly improving the accuracy of health category recognition.
[0062] An adaptive learning strategy is adopted to dynamically adjust model parameters to adapt to different data distributions, further enhancing the generalization and robustness of the model, ensuring that stable and reliable interpretation results can be provided in different scenarios, achieving efficient fusion of the model on multimodal data, and significantly improving the accuracy of health assessment.
[0063] As an improvement of the above embodiment, the method of dynamically allocating weights of image features and text features through the attention fusion layer includes:
[0064] The attention weight of the text feature T on the image feature V is calculated by cross attention. The calculation formula is:
[0065]
[0066] where α ij represents the attention weight of the i-th text token to the j-th image patch, V is the image feature matrix, T is the text feature matrix, and Q v is the mapping matrix from image features to text feature queries, and K t is the mapping matrix from text features to image feature key values, and V ik is the mapping matrix from image features to text feature queries, and Q v , K t are learnable matrices, k is the index variable of the image patch, d v is the dimension of the image features, and d t is the dimension of the text features.
[0067] It should be noted that Q v , K t are learnable weight matrices used to project the concatenated image and text features into the dynamic weight space, learn the interaction pattern between image and text features, and dynamically assign weights If the dimension of the image features d v = 512 and the dimension of the text features d t = 512, then Q v , K t are matrices of 512×512. V is the image feature matrix with a shape of (N, d v ), T is the text feature matrix with a shape of (M, d t ), N is the number of image patches, and M is the number of text feature tokens.
[0068] The attention weight of the image feature V to the text feature T is calculated through cross-attention, and the calculation formula is:
[0069]
[0070] where β kl represents the attention weight of the k-th image patch to the l-th text token, Q t is the mapping matrix from text features to image feature queries, K v is the mapping matrix from image features to text feature key values, and Q t , K v are learnable matrices, and m is the index variable of the text token.
[0071] It should be noted that d v is the dimension of the image features, and d tis the text feature dimension.
[0072] The Softmax function is used to normalize the attention weights α ij and the attention weight β kl to obtain the image weight and the text weight.
[0073] Through the cross-attention mechanism, the model can more comprehensively understand the complex relationship between images and texts, improving the performance of multi-modal tasks. Experiments show that this method is effective in tasks such as image captioning and question answering, verifying its effectiveness. Further optimizing the model parameters can improve the cross-modal information fusion accuracy and enhance the model's generalization ability.
[0074] As an improvement of the above embodiment, the calculation formula of the learning rate is:
[0075] lr(t) new = lr(t)·δ cut ;
[0076]
[0077] where t is the current training step, lr(t) new is the learning rate of the t-th iteration, lr(t) is the base learning rate of the t-th iteration, δ and γ are adjustment factors, 0 < δ < 1, 0 < γ < 1, cut is the number of iterations when the loss value decrease amplitude is continuously lower than the amplitude threshold, lr base is the initial learning rate, w s is the warm-up step, t s is the total number of training steps, L1 represents the first loss value, and L2 represents the second loss value.
[0078] Specifically, the warm-up step is set to 1000 and the total number of training steps is 50000. The base learning rate of the ViT model and the BERT model is 1×10 6 (fine-tuning after unfreezing), the base learning rate of the cross-attention layer and the classifier is 1×10 4 , and the default learning rate of other parameters is 1×10 5 . The larger the loss value, the larger the learning rate; if the loss value decrease amplitude is less than the amplitude threshold, the learning rate is reduced; by dynamically adjusting the learning rate, the model converges quickly in the initial stage of training and is finely optimized in the later stage, effectively improving the overall performance and generalization ability.
[0079] In some embodiments, the first loss value of the multi-modal model is calculated by the following formula:
[0080] L1 = λ1·L align +(1 - λ1)·L class ;
[0081]
[0082] Among them, L1 align represents the cross-attention alignment loss, and A i,i represents the cross-attention matrix of text features to image features, and A i,i = Softmax(α ij ); N ′ is the batch size, N is the number of patches of image features, M is the number of tokens of text features, and L1 class represents the first classification loss value, and Y ih represents the one-hot encoding of the true label (One-Hot encoding, when the i-th sample belongs to class h, Y ih is 1), is the probability that the i-th sample of the multi-modal model training belongs to class h; λ1 is the weight coefficient, 0 < λ1 < 1.
[0083] It should be noted that by freezing the pre-trained encoder, iterating the multi-modal model, and dynamically adjusting the learning rate, the optimization requirements of feature extraction and multi-modal alignment can be effectively balanced. Through the dynamic adjustment strategy, the multi-modal model can adaptively balance the convergence speed and accuracy during the training process, so as to show stronger robustness and adaptability in complex tasks.
[0084] In some embodiments, the second loss value of the multi-modal model is calculated by the following formula:
[0085] L2 = L2 var + L2 class + L2 KL ;
[0086]
[0087] Among them, L2 represents the second loss value, L2 var represents the variance penalty value, L2 class represents the second classification loss value, L2 KL represents the KL divergence value; q represents the variational distribution of parameter θ, θ represents the hyperparameter of the model, q(θ) represents the posterior distribution of parameter θ, represents the variance parameter, E q(θ) represents the expected value of parameter θ, Z (y=h) represents the indicator function, p(y = h|x,θ) represents the probability that the input fused feature x belongs to the healthy class h, y represents the true label, H represents the total number of healthy classes, and σ 2 represents the variance of the posterior distribution q(θ), and p(θ) represents the prior distribution of parameter θ, represents the variance of the prior distribution p(θ); λ2, λ3, and λ4 are all weight coefficients with values ranging from 0 to 1, and λ2 + λ3 + λ4 = 1.
[0088] It should be noted that x represents the input fused feature, and y represents the true health level label of the input sample; when the true label y is equal to the health category h, Z (y=h) = 1; otherwise Z (y=h) = 0. p(y = h|x,θ) represents the probability that the input fused feature x belongs to the category h (h = 1, 2,..., 5 correspond to normal, mildly abnormal, moderately abnormal, severely abnormal, and critical respectively). represents the model's estimate of the output uncertainty, and all categories share the same variance.
[0089] The second loss value of the multi-modal model includes a variance penalty value, a second classification loss value, and a KL divergence value; the variance penalty value is used to encourage the variance of the model prediction not to be too large to avoid overconfidence; when is close to the variance of the true data distribution, it will reduce the weight of the classification error value, reflecting the impact of uncertainty on the prediction. The constraint parameter θ follows the prior distribution p(θ) to prevent overfitting. By combining the variance penalty value, the second classification loss value, and the KL divergence value, the classification accuracy and the uncertainty of the multi-modal model are balanced.
[0090] During the training process, by continuously adjusting the weight coefficients λ2, λ3, and λ4, the overall loss function is optimized to improve the prediction accuracy. At the same time, monitor the changes in the variance parameters to ensure that the model's estimate of uncertainty is reasonable and avoid overconfidence or overconservatism. Finally, efficient fusion of multi-modal data and accurate health level classification are achieved. It can not only accurately identify the health level but also effectively quantify the uncertainty of the prediction.
[0091] Reference Figure 2 , the embodiment of the present invention also provides an intelligent report interpretation system based on deep learning, including:
[0092] at least one processor;
[0093] at least one memory for storing at least one program;
[0094] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0095] The content in the above method embodiments is applicable to this embodiment. The functions specifically implemented in this embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments, and will not be repeated here.
[0096] Although the description of the present disclosure has been quite detailed and has particularly described several of the described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but rather should be regarded as effectively covering the intended scope of the present disclosure by reference to the appended claims, considering the prior art to provide a broad interpretation of these claims. In addition, the present disclosure is described above in terms of embodiments foreseeable by the inventors for the purpose of providing a useful description, and non-substantive modifications to the present disclosure that are not currently foreseeable may still represent equivalent modifications of the present disclosure.
Claims
1. An intelligent report interpretation method based on deep learning, characterized in that: The method comprises the following steps: Obtain a physical examination report containing multiple physical examination items, and identify text information and image information in the physical examination report; Input text information and image information into a report recognition model including an image encoder, a text encoder, an attention fusion layer, and a classifier; extract image features of the physical examination items through the image encoder, and extract text features of the physical examination items through the text encoder; dynamically assign weights of image features and text features through the attention fusion layer, align the weighted image features to the dimensions of the text features, and then concatenate them with the weighted text features to obtain fused features; classify the fused features through the classifier, identify the items of concern in the physical examination report, and output the health category corresponding to the items of concern; Input the focus items and the health categories corresponding to the focus items into the knowledge graph to obtain the interpretation results of the physical examination report.
2. The method according to claim 1, characterized in that The report recognition model is trained in the following way: Construct a multimodal model consistent with the reported recognition model architecture; Obtain a sample set, input the sample set into the multimodal model, freeze all layers of the image encoder and the text encoder, iterate the multimodal model, and gradually adjust the learning rate according to the number of iterations and the first loss value of the multimodal model; After the training parameters of the multimodal model reach the number of preheating steps, the image encoder and the text encoder are unfrozen layer by layer according to the second loss value of the multimodal model, and the learning rate is gradually adjusted according to the number of iterations and the second loss value of the multimodal model until the second loss value of the multimodal model is lower than the second loss threshold, thereby obtaining a report recognition model; wherein the input data of the sample set is a sample report containing graphic and text information, and the output data is the focus items and corresponding health categories in the sample report.
3. The method according to claim 2, characterized in that The method of dynamically allocating weights of image features and text features through the attention fusion layer includes: Calculate the attention weight α of text features on image features through cross attention ij , and the attention weight β of the image feature V on the text feature T kl : Use Softmax function to adjust the attention weight α ij and the attention weight β kl Normalize and get image weight and text weight.
4. The method according to claim 3, characterized in that The attention weight of the text feature to the image feature is calculated by the following formula: Among them, T is the text feature, V is the image feature, α ij represents the attention weight of the i-th text token to the j-th image patch, V is the image feature matrix, T is the text feature matrix, Q v is the mapping matrix from image features to text feature queries, K t is the mapping matrix from text features to image feature keys, V ik is the mapping matrix from image features to text feature queries, Q v , K t is a learnable matrix, k is the index variable of the image patch, d v is the image feature dimension, d t is the text feature dimension.
5. The method according to claim 4, characterized in that The attention weight of the image feature V to the text feature T is calculated by the following formula: Among them, β kl represents the attention weight of the k-th image patch to the l-th text token, Q t is the mapping matrix from text features to image feature queries, K v is the mapping matrix from image features to text feature keys, Q t , K v is a learnable matrix, and m is the index variable of the text token.
6. The method according to claim 2, characterized in that The learning rate is calculated as follows: cut lr(t) new =lr(t)·δ; Among them, t is the current training step number, lr(t) new is the learning rate of the tth iteration, lr(t) is the baseline learning rate of the tth iteration, δ and γ are adjustment factors, 0<δ<1, 0<γ<1, cut is the number of iterations in which the loss value decreases continuously below the amplitude threshold, lr base is the initial learning rate, w s is the number of warm-up steps, t s is the total number of training steps, L1 represents the first loss value, and L2 represents the second loss value.
7. The method according to claim 6, characterized in that The first loss value of the multimodal model is calculated by the following formula: L1⼝λ·L align +(1-λ)·L class ; Among them, L1 align represents the cross attention alignment loss, A i,i A represents the cross attention matrix of text features to image features. i,i =Softmax(α ij );N ′ is the batch size, N is the number of patches of image features, M is the number of tokens of text features, L1 class Represents the first classification loss value, Y ih represents a single-bit efficient encoding of the true label, The probability that the i-th sample belongs to class h for training the multimodal model.
8. The method according to claim 6, characterized in that The second loss value of the multimodal model is calculated by the following formula: <h2 style=";text-align:left;direction:ltr">L2=-L<h2 style=";text-align:left;direction:ltr"> align <h2 style=";text-align:left;direction:ltr"> +L2<h2 style=";text-align:left;direction:ltr"> class <h2 style=";text-align:left;direction:ltr"> +L2<h2 style=";text-align:left;direction:ltr"> KL <h2 style=";text-align:left;direction:ltr"> ; Among them, L2 represents the second loss value, L2 var represents the variance penalty value, L2 class Represents the second classification loss value, L2 KL represents the KL divergence value; q represents the variational distribution of parameter θ, θ represents the hyperparameter of the model, and q(θ) represents the posterior distribution of parameter θ. represents the variance parameter, E q(θ) represents the expected value of parameter θ, Z (y=h) represents the indicator function, p(y=h|x,θ) represents the probability that the input fusion feature x belongs to the healthy category h, y represents the true label, H represents the total number of healthy categories, σ 2 represents the variance of the posterior distribution q(θ), p(θ) represents the prior distribution of the parameter θ, represents the variance of the prior distribution p(θ).
9. An intelligent report interpretation system based on deep learning, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Traditional Chinese medicine knowledge question-answering method fusing knowledge graph and multi-modal dialogue model
CN117851571A
Multi-modal teaching knowledge graph construction method based on medical image report
CN118468996A
Cited By
Report interpretation method and system
CN121052260A
Report interpretation method and system
CN121052260B