Mammary gland X-ray photography pre-examination risk assessment method based on Internet hospital
By introducing a multi-head attention mechanism and deep generation network in breast cancer risk assessment, the problem of insufficient multimodal data correlation capture in the prior art is solved, and higher accuracy risk assessment and better explanatory evaluation results are achieved.
Patent Information
- Application Number
- CN202510068044.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-13
AI Technical Summary
Existing breast cancer risk assessment methods cannot effectively capture the potential correlation between multimodal data, resulting in reduced feature fusion accuracy and limited accuracy and reliability of evaluation results.
By introducing a multi-head attention mechanism, dynamically allocate the weights of text features, numerical features and image features, perform feature fusion, and use a deep generation network to model the data potential spatially to generate potential distributions and risk distributions.
The accuracy of feature fusion is improved, high-quality potential spatial representations are generated, and the accuracy of risk distribution modeling before mammogram is improved, and the interpretability of evaluation results is improved.
Smart Images

Figure CN120148836A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical and health information processing, and specifically provides a risk assessment method before breast X-ray examination based on an Internet hospital. Background Art
[0002] Breast cancer is one of the most common malignant tumors among women globally, and early screening is of great significance for reducing the mortality rate of breast cancer. Breast X-ray photography is a widely used breast cancer screening method at present, which can effectively detect early lesions. The accuracy and efficiency of breast X-ray photography screening largely depend on the personalized risk assessment of patients. Existing risk assessment methods usually rely on static unimodal data and cannot capture the complex associations of multimodal data in patients' health information, so the accuracy and reliability of the assessment results are limited.
[0003] The current breast cancer risk assessment methods face the following technical problems: Breast cancer risk assessment involves multimodal health information such as text data, numerical data, and image data. Existing technologies usually adopt simple weighted sum and concatenation operations when fusing multimodal data, which cannot effectively capture the potential associations between different modal data, thus affecting the accuracy of feature fusion and leading to a decrease in the accuracy of risk assessment results.
[0004] In practical applications, there are significant differences in the quality and quantity of health data among different patients. Especially for high-risk patients, there is less data and some health information is missing. The imbalance and sparsity of data distribution make traditional risk assessment methods unable to generate reliable personalized risk prediction models, and the assessment results are prone to deviation.
[0005] Breast cancer risk assessment requires potential space modeling based on multimodal data and sampling of latent variables. However, the sampling methods of existing technologies are inefficient and prone to falling into local optima, resulting in a decrease in the accuracy and dynamic update ability of risk distribution modeling.
[0006] Existing breast cancer risk assessment methods usually output a single risk score, lacking a probabilistic representation of the score result and an association analysis with the specific health conditions of patients, resulting in poor interpretability of the assessment results and limiting the understanding and application of the assessment results by doctors and patients.
[0007] Therefore, those skilled in the art provide a risk assessment method before breast X-ray examination based on an Internet hospital to solve the above-mentioned problems. Summary of the Invention
[0008] Aiming at the deficiencies of the existing technology, the present invention provides a risk assessment method before breast X-ray examination based on an Internet hospital to solve the problems raised in the above background art.
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: A risk assessment method before mammography based on an Internet hospital, including: Step 1: Obtain the multi-modal health data of the patient through the Internet hospital platform. The multi-modal health data includes text data, numerical data, and image data. The text data includes the patient's living habits and family medical history information. The numerical data includes age, body mass index, menstrual history, and reproductive history. The image data includes mammary gland density images; Step 2: Perform preliminary processing on the multi-modal health data obtained in Step 1. The text data uses an embedding model to generate text feature vectors. The numerical data generates standardized features through standardization processing. The image data extracts image features through a deep neural network to obtain text features, numerical features, and image features; Step 3: Perform feature fusion based on the text features, numerical features, and image features obtained in Step 2. Generate a unified feature representation through feature fusion operations. When performing feature fusion, an attention mechanism is introduced to calculate the weights of the text features, numerical features, and image features, and generate a unified feature representation. The unified feature representation will be used as the input data for subsequent risk modeling; Step 4: Based on the unified feature representation in Step 3, perform latent space modeling on the data through a deep generative network to generate a latent distribution. Specifically, the unified feature representation is input into an encoder to generate the mean parameter and variance parameter of the latent distribution. Based on the mean parameter and variance parameter of the latent distribution, a latent distribution is constructed to represent the latent risk features before mammography; Step 5: Based on the latent distribution generated in Step 4, sample the latent distribution to obtain latent variables. The latent variables are the expressions of the unified representation of the text features, numerical features, and image features captured in the latent distribution in the latent space. The latent variables will be used for further modeling of the risk distribution and risk assessment; Step 6: Introduce a dynamic optimization process for the latent variables generated in Step 5 to optimize the distribution of the latent variables. Specifically, sampling techniques are used to obtain optimized latent variables in the latent distribution, and based on the optimized latent variables, a risk distribution before mammography is generated. The risk distribution includes the probability mean and probability variance of the patient's risk; Step 7: Model the risk distribution generated in Step 6, and classify the patient's risk based on the probability mean and probability variance of the risk distribution. Specifically, it is divided into low risk, medium risk, and high risk. The risk classification result is associated with the feature representation corresponding to the latent variable, and is used to guide the generation of the next screening recommendation; Step 8: Combine the risk grading results in Step 6 and the feature representations generated during the potential space modeling process to generate a risk assessment result associated with the patient's personalized features. The risk assessment result includes risk distribution, risk grading, and screening recommendations. When the patient submits new data, Step 7 dynamically updates the risk distribution and adjusts the risk assessment result in real time.
[0010] Preferably, the calculation formula of the attention mechanism for the feature fusion operation in Step 3 is: , where, is the query matrix, is the key matrix, is the value matrix, is the dimension of the key matrix, represents the dot product between the query matrix and the key matrix.
[0011] Preferably, the final unified feature representation calculated by the attention mechanism is: , where, represents the text feature vector, represents the normalized numerical feature, represents the image feature vector, represents the unified feature after feature fusion.
[0012] Preferably, the construction of the potential distribution in Step 4 is based on a variational autoencoder, and the optimization objective is the variational lower bound: , where, is the objective function, represents the approximate posterior distribution, is the prior distribution of the latent variable, represents the latent variable generating the observed data probability, is the divergence, represents the expectation of under the approximate posterior distribution with respect to
[0013] Preferably, the prior distribution is assumed to be a multi-dimensional independent Gaussian distribution , and the probability density function is: , where, is the latent variable vector, is the dimension of the latent variable, is * Identity matrix of denotes the squared Euclidean norm of
[0014] Preferably, the dynamic optimization process in step 6 is implemented by the Hamiltonian Monte Carlo method, and the sampling steps include: , , , where is the momentum variable, is the latent variable, is the step size, is the potential energy gradient, denotes the potential energy of the latent distribution in the latent space, denotes the value of the momentum variable at time step t, is the half-step update result of the momentum variable at the current time step t, denotes the value of the latent variable at time step t, denotes at time step the updated value of the latent variable at the moment.
[0015] Preferably, the potential energy function is defined as: , where denotes the reconstructed distribution, denotes the prior of the latent distribution.
[0016] Preferably, the probability mean and the probability variance of the risk distribution in step 7 are used for risk grading, and the specific grading is as follows: Low risk: < 20%, < 5%; Medium risk: 20% ≤ ≤ 50%; High risk: > 50%.
[0017] Preferably, the risk grading result is used to generate personalized screening suggestions, where the screening suggestions are represented by the following formula: , where denotes the risk mean, denotes the risk standard deviation, denotes the personalized conditional characteristics of the patient.
[0018] Preferably, when the patient submits new data through the Internet hospital platform, the risk assessment result is dynamically updated through the following steps: a. Update the unified feature representation in step 3 based on the submitted new data, and re-fuse the text feature vector, standardized numerical feature, and image feature generated from the new data into a new unified feature representation; b. Input the updated unified feature representation into the deep generative network in step 4, and recalculate the mean parameter and variance parameter of the latent distribution; c. Based on the updated latent distribution, according to the sampling and dynamic optimization processes defined in steps 5 and 6, regenerate the latent variables and optimize the latent variable distribution; d. Use the risk distribution modeling method in step 7 to update the mean and variance of the patient's risk distribution probability, re-perform risk grading, and generate personalized screening recommendations based on the new risk grading. The updated risk assessment result includes the latest risk distribution, risk grading, and screening recommendations.
[0019] The present invention provides a risk assessment method for breast mammography before examination based on an Internet hospital, which has the following beneficial effects: 1. By introducing a multi-head attention mechanism to dynamically allocate weights to text features, numerical features, and image features, the present invention effectively captures the potential associations between different modal data and generates a unified feature representation, resulting in improved feature fusion accuracy and providing high-quality input for subsequent latent space modeling.
[0020] 2. By using a variational autoencoder to perform latent space modeling on the unified feature representation, generating the mean parameter and variance parameter of the latent distribution, the present invention realizes the probabilistic representation of multi-modal features and the adaptability to data distribution changes, resulting in the effect of stably capturing breast cancer-related risk features in the case of data sparsity and imbalance.
[0021] 3. By adopting the Hamiltonian Monte Carlo method to dynamically optimize the distribution of latent variables and combining the potential energy gradient to calculate the sampling path in the latent space, the present invention realizes efficient sampling in the latent distribution and optimization of latent variables, resulting in the effect of improving the accuracy of risk distribution modeling and assessment accuracy for breast mammography before examination.
[0022] 4. By combining the probability mean and variance of the risk distribution to generate risk grading and generating personalized screening recommendations based on the risk grading and patient-specific conditions, the present invention realizes the intuitive expression and interpretation of the patient's risk assessment result, resulting in the effect of improving the interpretability of the assessment result and guiding patients and doctors to formulate screening and intervention strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flowchart of the present invention. Detailed implementation manners
[0024] To enable those skilled in the art to understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] The present invention will be described in detail below with reference to the accompanying drawings: Embodiment: Please refer to the attached Figure 1 , an embodiment of the present invention provides a risk assessment method before mammography based on an Internet hospital, including: Step 1: Obtain multi-modal health data of a patient through an Internet hospital platform. The multi-modal health data includes text data, numerical data, and image data. The text data includes the patient's living habits and family medical history information. The numerical data includes age, body mass index, menstrual history, and reproductive history. The image data includes mammary gland density images; Step 2: Perform preliminary processing on the multi-modal health data obtained in Step 1. The text data uses an embedding model to generate text feature vectors. The numerical data generates standardized features through standardization processing. The image data extracts image features through a deep neural network to obtain text features, numerical features, and image features; Step 3: Perform feature fusion based on the text features, numerical features, and image features obtained in Step 2. A unified feature representation is generated through a feature fusion operation. An attention mechanism is introduced during feature fusion to calculate the weights of the text features, numerical features, and image features to generate a unified feature representation. The unified feature representation will be used as the input data for subsequent risk modeling; Step 4: Based on the unified feature representation in Step 3, perform latent space modeling on the data through a deep generative network to generate a latent distribution, specifically including: inputting the unified feature representation into an encoder to generate the mean parameter and variance parameter of the latent distribution, and constructing a latent distribution based on the mean parameter and variance parameter of the latent distribution to represent the latent risk features before mammography; Step 5: Based on the latent distribution generated in Step 4, sample the latent distribution to obtain latent variables. The latent variables are the expressions of the unified representation of the text features, numerical features, and image features captured in the latent distribution in the latent space. The latent variables will be used for further modeling of the risk distribution and risk assessment; Step 6: By introducing a dynamic optimization process for the latent variables generated in Step 5, optimize the distribution of the latent variables. Specifically, obtain the optimized latent variables in the latent distribution through sampling techniques, and generate a risk distribution before mammography based on the optimized latent variables. The risk distribution includes the probability mean and probability variance of the patient's risk. Step 7: Model the risk distribution generated in Step 6, and classify the patients according to the probability mean and probability variance of the risk distribution. Specifically, classify them into low risk, medium risk, and high risk. The risk classification result is associated with the feature representation corresponding to the latent variables, and is used to guide the generation of the next screening recommendation. Step 8: Combine the risk classification result in Step 6 and the feature representation generated in the latent space modeling process to generate a risk assessment result associated with the patient's personalized characteristics. The risk assessment result includes the risk distribution, risk classification, and screening recommendation. When the patient submits new data, Step 7 dynamically updates the risk distribution and adjusts the risk assessment result in real time.
[0026] Benefits of Step 1: Integrate the patient's text data, numerical data, and image data through the Internet hospital platform, comprehensively cover the patient's health information, provide a multi-dimensional data basis for risk assessment, solve the problem of insufficient single-modal data, capture the patient's complex health characteristics, and improve the comprehensiveness and accuracy of the assessment. Benefits of Step 2: Use the embedding model, standardization processing, and deep neural network to extract features from text, numerical, and image data respectively, and unify them into text feature vectors, standardized features, and image features. Solve the problem of inconsistent heterogeneous data formats through multi-modal data feature extraction technology, and provide a standardized input for feature fusion. Benefits of Step 3: By introducing the attention mechanism, dynamically allocate the weights of text, numerical, and image features to generate a unified feature representation, which can effectively capture the potential correlation between different modal data, improve the accuracy of feature fusion, and provide high-quality input data for subsequent latent space modeling. Benefits of Step 4: Based on the unified feature representation, use the deep generative network to perform latent space modeling on the data, generate the mean parameter and variance parameter of the latent distribution, and solve the problems of uneven patient data distribution and data sparsity through the modeling method of probability distribution, improving the adaptability of risk assessment to complex data distributions. Benefits of Step 5: By sampling the latent distribution, generate latent variables, capture the expression of the patient's multi-modal characteristics in the latent space, extract the key features in the latent space, and use them for further modeling and risk assessment, improving the characterization ability of the assessment model for complex data. Benefits of Step 6: By introducing a dynamic optimization process and sampling techniques, the latent variable distribution is optimized and an optimized risk distribution for mammography examinations is generated, including the probability mean and probability variance of the risks, improving the sampling efficiency and accuracy of the latent distribution and ensuring the precision and stability of the risk assessment results; Benefits of Step 7: Based on the probability mean and variance of the risk distribution, patients are risk-graded, which can transform the complex probability distribution into an easily understandable grading result, laying a foundation for generating subsequent screening recommendations and increasing the practicality of the assessment results; Benefits of Step 8: Combining the risk grading results and the feature representations generated by latent space modeling, personalized risk assessment results are output, including the risk distribution, risk grading, and screening recommendations, supporting real-time updates of the risk distribution after the patient submits new data. Through a dynamic update mechanism, it adapts to changes in the patient's health status, improving the timeliness and personalization of the assessment results.
[0027] The calculation formula of the attention mechanism for the feature fusion operation in Step 3 is: , where is the query matrix, is the key matrix, is the value matrix, is the dimension of the key matrix, represents the dot product between the query matrix and the key matrix.
[0028] The final unified feature representation calculated by the attention mechanism is: , where represents the text feature vector, represents the normalized numerical feature, represents the image feature vector, represents the unified feature after feature fusion.
[0029] The calculation formula of the attention mechanism realizes dynamic weight allocation for different modality features. Among them, the correlation between modalities is calculated through the dot product operation , and the weight distribution is normalized using the function, which can dynamically evaluate the importance of different features and avoid information loss caused by simple weighted methods.
[0030] The attention mechanism can capture the potential relationships between text, numerical, and image features through the dot product calculation of the query matrix and the key matrix , effectively reflecting the semantic relevance between multi-modal features and solving the problem of insufficient isolated modeling of single-modal features.
[0031] By dynamically fusing different modal features, the generated As the input data for subsequent latent space modeling, it can improve the accuracy of feature fusion, provide high-quality input for latent distribution modeling, and thus improve the overall accuracy of risk assessment.
[0032] The attention mechanism dynamically adapts the weight distribution between features during the feature fusion process. Even if the quality of a certain modal data is low and noisy, the modal features can supplement information through weight adjustment, thereby improving the robustness of the final unified feature representation and ensuring the stability of subsequent modeling.
[0033] The construction of the latent distribution in Step 4 is based on a variational autoencoder, and the optimization objective is the variational lower bound: , where, is the objective function, represents the approximate posterior distribution, is the prior distribution of the latent variable, represents the latent variable generating the probability of the observed data , is the divergence, represents the expectation of under the approximate posterior distribution .
[0034] The prior distribution is assumed to be a multi-dimensional independent Gaussian distribution , and the probability density function is: , where, is the latent variable vector, is the dimension of the latent variable, is * the identity matrix of, represents the squared Euclidean norm of.
[0035] By optimizing the variational lower bound, the variational autoencoder can effectively reconstruct the input data, learn the complex multi-modal data distribution, capture the latent feature relationships in the data, and provide an accurate latent space representation for the risk assessment before breast mammography.
[0036] Through the concise mathematical properties of the Gaussian distribution, the latent distribution can efficiently describe the complex features in multi-modal data. In addition, the differentiability and multi-dimensional expansion ability of the Gaussian distribution have good adaptability to the changes in the data distribution. Especially in the case of sparse and uneven data distributions, reliable modeling can be performed through the latent distribution.
[0037] Probabilistic representations can capture the uncertainty of data, address the deficiencies of traditional single-point estimation methods in dealing with data noise and incompleteness, and provide more robust features for subsequent risk distribution modeling.
[0038] Through the high-dimensional representation of latent variables, variational autoencoders can compress the input multi-modal features into the latent space, capturing the global representation of the data and the interaction relationships between modalities. The multi-dimensional Gaussian assumption of the latent distribution provides theoretical support for high-dimensional data processing, avoiding the curse of dimensionality problem caused by directly processing the original multi-modal data.
[0039] The dynamic optimization process in step 6 is implemented by the Hamiltonian Monte Carlo method, and the sampling steps include: , , , where, is the momentum variable, is the latent variable, is the step size, is the potential energy gradient, represents the potential energy of the latent distribution in the latent space, represents the value of the momentum variable at time step t, is the half-step update result of the momentum variable at the current time step t, represents the value of the latent variable at time step t, represents at time step the updated value of the latent variable at time.
[0040] The potential energy function is defined as: , where, represents the reconstruction distribution, represents the prior of the latent distribution.
[0041] In the dynamic optimization process, the Hamiltonian Monte Carlo method is adopted. By introducing the momentum variable and the potential energy function and combining the gradient information, efficient sampling is carried out.
[0042] The momentum variable drives the latent variable to move along the gradient direction in the latent space, avoiding the random walk phenomenon in traditional sampling methods, and significantly improving the sampling efficiency and the quality of latent distribution optimization.
[0043] The Hamiltonian Monte Carlo uses the gradient information of the potential energy function to guide the latent variable.
[0044] Gradient information can indicate the optimal direction in the latent space, guiding the sampling path of latent variables away from local optima and ultimately approaching the global distribution, thereby improving the accuracy of risk distribution modeling.
[0045] By introducing a momentum variable, Hamiltonian Monte Carlo simulates a physical dynamics system in a high-dimensional latent space, enabling sampling points to move rapidly along the inertial direction and efficiently explore the global distribution of the latent space, which is particularly suitable for the dynamic optimization of high-dimensional latent variables.
[0046] During the sampling process, Hamiltonian Monte Carlo uses the combined effect of the potential energy function and the momentum variable to accurately model the distribution of latent variables. The distribution of sampling points is close to the true latent distribution, avoiding evaluation biases caused by low-quality sampling. By generating the risk distribution of mammography examinations through the optimized latent variables, including the mean and variance of risk probabilities, the reliability and accuracy of risk assessment are further improved.
[0047] The mean probability of the risk distribution in step 7 and the probability variance are used for risk grading, specifically divided as follows: Low risk: <20%, <5%; Medium risk: 20% ≤ ≤ 50%; High risk: >50%.
[0048] The risk grading results are used to generate personalized screening recommendations, where the screening recommendations are represented by the following formula: , where, represents the risk mean, represents the risk standard deviation, represents the personalized conditional characteristics of the patient.
[0049] The grading criteria are based on the central tendency and uncertainty quantification of risks, which can clearly distinguish patients with different risk levels and make the evaluation results more instructive. The refined risk grading provides an intuitive risk level division for doctors and patients, facilitating the rational allocation of screening resources.
[0050] Based on the traditional single risk value assessment, the reliability of risk grading is significantly improved, especially in the case of high risks or large data noise, providing a more evidence-based evaluation result for screening decisions.
[0051] Personalized screening recommendations through risk grading can concentrate the allocation of screening resources to high-risk patient groups, avoiding waste of resources on over-screening of low-risk patients. At the same time, the appropriate adjustment of the screening frequency for medium-risk patients and the combination of multiple screening methods contribute to maximizing the screening benefits under limited medical resources.
[0052] When a patient submits new data through the Internet hospital platform, the risk assessment results are dynamically updated through the following steps: a. Update the unified feature representation in step 3 based on the submitted new data, and re-fuse the text feature vector, standardized numerical feature, and image feature generated from the new data into a new unified feature representation; b. Input the updated unified feature representation into the deep generative network in step 4 to recalculate the mean parameter and variance parameter of the latent distribution; c. Based on the updated latent distribution, regenerate the latent variables and optimize the latent variable distribution according to the sampling and dynamic optimization processes defined in steps 5 and 6; d. Use the risk distribution modeling method in step 7 to update the mean and variance of the risk distribution probability of the patient, re-perform risk grading, and generate personalized screening recommendations based on the new risk grading. The updated risk assessment results include the latest risk distribution, risk grading, and screening recommendations.
[0053] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A risk assessment method before breast X-ray examination based on Internet hospital, characterized in that: include: Step 1: Obtain the patient's multimodal health data through the Internet hospital platform. The multimodal health data includes text data, numerical data and image data. The text data includes the patient's living habits and family medical history information, the numerical data includes age, body mass index, menstrual history and reproductive history, and the image data includes breast density images; Step 2: Perform preliminary processing on the multimodal health data obtained in step 1. The text data generates text feature vectors using the embedding model, the numerical data generates standardized features through standardization, and the image data extracts image features through a deep neural network to obtain text features, numerical features, and image features; Step 3: Based on the text features, numerical features and image features obtained in step 2, feature fusion is performed to generate a unified feature representation. During feature fusion, an attention mechanism is introduced to calculate the weights of text features, numerical features and image features to generate a unified feature representation. The unified feature representation will be used as input data for subsequent risk modeling. Step 4: Based on the unified feature representation in step 3, a deep generative network is used to perform latent space modeling on the data to generate a latent distribution, specifically including: inputting the unified feature representation into an encoder to generate a mean parameter and a variance parameter of the latent distribution, and based on the mean parameter and the variance parameter of the latent distribution, constructing a latent distribution for representing the potential risk characteristics before the mammography examination; Step 5: Based on the latent distribution generated in step 4, the latent distribution is sampled to obtain latent variables. The latent variables are expressed in the latent space through the unified representation of text features, numerical features, and image features captured in the latent distribution. The latent variables will be used for further modeling and risk assessment of risk distribution. Step 6, optimizing the distribution of the latent variables by introducing a dynamic optimization process into the latent variables generated in step 5, specifically obtaining the optimized latent variables in the latent distribution by a sampling technique, and generating a risk distribution before the mammography examination based on the optimized latent variables, the risk distribution including a probability mean and a probability variance of the patient's risk; Step 7: Model the risk distribution generated in step 6, and classify the patients into low risk, medium risk, and high risk according to the probability mean and probability variance of the risk distribution. The risk classification results are associated with the feature representation corresponding to the latent variable to guide the generation of the next screening recommendation; Step 8: Combine the risk grading results in step 6 and the feature representation generated in the latent space modeling process to generate a risk assessment result associated with the patient's personalized characteristics. The risk assessment result includes risk distribution, risk grading, and screening recommendations. When the patient submits new data, step 7 dynamically updates the risk distribution and adjusts the risk assessment result in real time.
2. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 1, characterized in that: The calculation formula of the attention mechanism of the feature fusion operation in step 3 is: , in, is the query matrix, is the key matrix, is the value matrix, is the dimension of the key matrix, Represents the dot product between the query matrix and the key matrix.
3. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 2, characterized in that: The final unified feature representation calculated by the attention mechanism is: , in, represents the text feature vector, represents the normalized numerical feature, represents the image feature vector, Represents the unified feature after feature fusion.
4. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 1, characterized in that: The construction of the potential distribution in step 4 is based on the variational autoencoder, and the optimization target is the variational lower bound: , in, is the objective function, represents the approximate posterior distribution, is the prior distribution of the latent variable, Represents latent variables Generate observation data The probability of is the divergence, Represents the approximate posterior distribution Next pair expected value.
5. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 4, characterized in that: The prior distribution Assuming multi-dimensional independent Gaussian distribution , the probability density function is: , in, is the latent variable vector, is the dimension of the latent variable, for * The identity matrix of express The square of the Euclidean norm of .
6. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 1, characterized in that: The dynamic optimization process in step 6 is implemented by the Hamiltonian Monte Carlo method, and the sampling steps include: , , , in, is the momentum variable, is a latent variable, is the step length, is the potential energy gradient, represents the potential energy of the potential distribution in the latent space, represents the value of the momentum variable at time step t, is the half-step update result of the momentum variable at the current time step t, represents the value of the latent variable at time step t, Indicates that at time step The value of the latent variable after the moment is updated.
7. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 6, characterized in that: The potential energy function is defined as: , in, represents the reconstructed distribution, Represents a prior on the underlying distribution.
8. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 1, characterized in that: The probability mean of the risk distribution in step 7 and probability variance Used for risk classification, specifically divided into: Low risk: <20%, <5%; Medium risk: 20% ≤ ≤50%; High risk: >50%.
9. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 8, characterized in that: The risk grading results are used to generate personalized screening recommendations, where the screening recommendations are expressed by the following formula: , in, represents the risk mean, represents the risk standard deviation, Represents the patient's personalized condition characteristics.
10. The method for risk assessment before breast X-ray examination based on Internet hospital according to claim 1, characterized in that: When the patient submits new data through the Internet hospital platform, the risk assessment results are dynamically updated through the following steps: a. Update the unified feature representation in step 3 based on the submitted new data, and re-integrate the text feature vector, standardized numerical features, and image features generated by the new data into a new unified feature representation; b. Input the updated unified feature representation into the deep generative network in step 4 and recalculate the mean parameter and variance parameter of the potential distribution; c. Based on the updated latent distribution, regenerate the latent variables and optimize the latent variable distribution according to the sampling and dynamic optimization process defined in steps 5 and 6; d. Using the risk distribution modeling method in step 7, update the patient's risk distribution probability mean and probability variance, re-risk grade, and generate personalized screening recommendations based on the new risk grade. The updated risk assessment results include the latest risk distribution, risk grade, and screening recommendations.