Big data-based personalized hypertension risk assessment method and system
By constructing a time-series feature matrix and a risk propagation graph, and utilizing Granger causality tests and propagation entropy calculations, the problems of insufficient personalization and accuracy in traditional hypertension assessment methods are solved, achieving precision and timeliness in personalized hypertension risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE AFFILIATED HOSPITAL OF GUIZHOU MEDICAL UNIV
- Filing Date
- 2026-01-13
- Publication Date
- 2026-05-29
Smart Images

Figure CN122117356A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical and health monitoring and prevention technology, specifically relating to a personalized hypertension risk assessment method and system based on big data. Background Technology
[0002] Hypertension is a common chronic disease affecting approximately one billion people worldwide. It can lead to cardiovascular and cerebrovascular diseases and is one of the leading causes of death and disability globally. With the development of big data and artificial intelligence technologies, it has become possible to accurately assess individual hypertension risk using massive amounts of health data, which is of great significance for the early prevention and intervention of hypertension.
[0003] Traditional hypertension risk assessment primarily relies on statistical models, such as the Framingham Risk Score system, which predicts hypertension risk by analyzing static indicators such as age, gender, blood pressure, and blood lipids. In recent years, with the widespread adoption of wearable devices and health management applications, dynamic monitoring data of individual physiological indicators and behavioral habits have accumulated, providing a new data foundation for accurate hypertension risk assessment. Summary of the Invention
[0004] This invention provides a personalized hypertension risk assessment method and system based on big data, which can solve the problems in the prior art.
[0005] A first aspect of this invention provides a personalized hypertension risk assessment method based on big data, comprising: Acquire physiological and behavioral data of the target individual and identify multiple risk factors; Continuous monitoring curves are constructed for each risk factor. Step change points are detected by cumulative sum control chart algorithm. The rate of change and duration of change of indicators before and after the step change point are extracted as step features and a time series feature matrix is constructed. Based on the time series data of each risk factor in the time series feature matrix, the unidirectional time series dependency relationship between risk factors is identified through Granger causality test, a directed acyclic risk propagation graph is constructed, and the information transmission strength of each directed edge is calculated by the transfer entropy to obtain the edge weight matrix. The risk propagation graph of the target individual is matched with the risk propagation graphs of each historical individual. Reference individuals are selected by using a graph kernel function, and the cascading amplification law parameters along the propagation path are extracted to construct a multi-step propagation dynamics equation. Combined with the edge weight matrix, a personalized risk propagation model is obtained. The step characteristics of the target individual at the current moment are input into the personalized risk propagation model, and the risk signal is simulated to propagate in multiple stages along the directed edges of the risk propagation graph. The cumulative risk intensity propagated to the hypertension onset node is calculated to obtain the personalized hypertension risk assessment value.
[0006] Continuous monitoring curves are constructed for each risk factor. Step change points are detected using cumulative sum control chart algorithms. The rate of change and duration of change before and after the step change point are extracted as step features. A time-series feature matrix is then constructed, including: Continuous monitoring curves are constructed for each risk factor according to a preset sampling frequency; For the continuous monitoring curve, the cumulative deviation sequence is calculated by the cumulative sum control chart algorithm, and the cumulative deviation sequence is decomposed into multiple scale components by wavelet transform. Local extreme points are detected on each scale component, and the local extreme points are projected and fused on the time axis. The time points that appear on multiple scale components at the same time are marked as step change points. Extract the rate of change of the index and the duration of change within a preset time window before and after each step change point as step features, and calculate the fractal dimension difference of the monitoring curves before and after the step change point. The fractal dimension difference represents the risk factor and the degree of transformation from an ordered state to a chaotic state. Combine the rate of change of the index, the duration of change and the fractal dimension difference to form a step feature vector. The step feature vectors of each risk factor are aligned and arranged according to the time dimension to construct a time series feature matrix.
[0007] Based on the time-series data of each risk factor in the aforementioned time-series feature matrix, the unidirectional time-series dependencies between risk factors are identified through Granger causality tests. A directed acyclic risk propagation graph is constructed, and the information transmission strength of each directed edge is calculated using propagation entropy to obtain the edge weight matrix, including: Based on the time series data of each risk factor, the predictive contribution of the first risk factor to the second risk factor is calculated by Granger causality test. When it exceeds the significance threshold, the one-way time series dependency from the first risk factor to the second risk factor is determined. The Granger causality test is performed on all risk factor pairs to extract all one-way time series dependencies. Each risk factor is used as a node, and the unidirectional temporal dependency is used as a directed edge to construct an initial directed graph. Loop detection is performed on the initial directed graph. When a loop is detected, the time lag order between risk factor pairs corresponding to each directed edge in the loop is calculated, and the directed edge with the largest time lag order is deleted to obtain a directed acyclic risk propagation graph. For each directed edge in the risk propagation graph, its time-series data is symbolically encoded to obtain the source symbol sequence and the target symbol sequence. The relative entropy between their joint probability distribution and the marginal probability distribution is calculated. The relative entropy is used as the information transmission strength of the directed edge. The edges are then arranged in a matrix according to the index positions of the starting and ending nodes of the directed edges to obtain the edge weight matrix.
[0008] For each directed edge in the risk propagation graph, its temporal data is symbolically encoded to obtain a source symbol sequence and a target symbol sequence. The relative entropy between their joint probability distribution and the marginal probability distribution is calculated, and this relative entropy is used as the information transmission strength of the directed edge, including: For each directed edge, the time series data of its starting risk factor is used as the source time series data, and the time series data of its ending risk factor is used as the target time series data. The source time series data and target time series data are divided into multiple time periods. The data distribution characteristics in each time period are calculated, the quantile threshold of symbolic encoding is determined, and the time series data in each time period are mapped to discrete symbols to form source symbol sequences and target symbol sequences. The source symbol sequence and the target symbol sequence are aligned with a fixed time lag step. Continuous symbol substrings are extracted from the source symbol sequence by sliding with a preset embedding dimension. At the same time, a single symbol in the target symbol sequence at the corresponding time is extracted. The joint occurrence frequency of all symbol substrings and single symbols is counted and normalized to obtain the joint probability distribution. The frequency of edge occurrences of source symbol substrings and target single symbols is counted separately, and the edge probability distribution is obtained by normalization. The ratio of the product of the joint probability distribution and the edge probability distribution is calculated and the logarithm is taken. The relative entropy is obtained by multiplying the product of the joint probability distribution and summing the logarithms. This relative entropy is used as the information transmission strength of each directed edge.
[0009] The risk propagation graph of the target individual is matched with the risk propagation graphs of each historical individual using graph isomorphism matching. Reference individuals are selected using a graph kernel function, and parameters of the cascading amplification pattern along the propagation path are extracted. A multi-step propagation dynamics equation is constructed, and combined with the edge weight matrix, a personalized risk propagation model is obtained, including: The risk propagation graph of the target individual and the risk propagation graph of each historical individual are topologically encoded to generate corresponding node degree distribution sequences and adjacency relationship matrices. Based on the node degree distribution sequence and the adjacency matrix, the graph kernel function value between the target individual and each historical individual is calculated. The graph kernel function value is obtained by weighted summation of the local subgraph structure similarity of each node. Historical individuals whose graph kernel function value exceeds the similarity threshold are selected as reference individuals. For the risk propagation graph of the reference individual, the directed propagation path from the initial risk factor to the hypertension onset node is extracted, and the stepwise amplification coefficient is calculated. The stepwise amplification coefficient is the product of the edge weight between adjacent nodes on the propagation path and the cumulative risk intensity of the preceding sequence. Based on the stepwise amplification coefficient of all directed propagation paths of each reference individual, the cascade amplification law parameter is calculated. Based on the cascade amplification law parameters, a multi-step propagation dynamic equation describing the cumulative enhancement mechanism of risk signals during multi-level propagation is constructed, and associated with the risk propagation graph and edge weight matrix of the target individual to obtain a personalized risk propagation model.
[0010] Based on the cascade amplification law parameters, a multi-step propagation dynamics equation describing the cumulative enhancement mechanism of risk signals during multi-stage propagation is constructed. This equation is then correlated with the risk propagation graph and edge weight matrix of the target individual to obtain a personalized risk propagation model, including: The parameters of the cascade amplification law are fitted with a function to obtain the amplification coefficient function between the step amplification coefficient and the propagation level, and a multi-step propagation dynamics equation is constructed. The multi-step propagation dynamics equation sets the current observation value of the risk factor as the initial risk intensity, and performs iterative calculations according to the propagation level order, accumulating the calculations up to the propagation level where the hypertension onset node is located. The directed edges in the risk propagation graph of the target individual are divided into multiple propagation levels, and the edge weights of the directed edges are extracted. The average value is calculated as the level propagation efficiency, and multiplied with the output value of the amplification coefficient function at the corresponding level to obtain the personalized amplification coefficient. Combined with the multi-step propagation dynamics equation, a personalized risk propagation model is obtained.
[0011] The step characteristics of the target individual at the current moment are input into the personalized risk propagation model, and the risk signal is simulated to propagate through multiple levels along the directed edges of the risk propagation graph. The cumulative risk intensity propagated to the hypertension onset node is calculated to obtain a personalized hypertension risk assessment value, including: The step characteristics of the target individual at the current moment are mapped to the corresponding risk factor nodes in the risk propagation graph to obtain the initial risk intensity value of each source node; According to the topological order of the risk propagation graph, starting from each source node, the upstream risk signal received by each intermediate node is calculated sequentially, and the upstream risk signal is weighted and combined with the corresponding step-by-step amplification coefficient to obtain the propagation risk intensity value of the intermediate node. The post-propagation risk intensity values of all propagation paths along the risk propagation map reaching the hypertension onset node are nonlinearly fused. Based on the path length and the number of nodes on each propagation path, a differentiated weight is applied to the post-propagation risk intensity values to obtain the cumulative risk intensity of the hypertension onset node. The cumulative risk intensity is mapped to a preset risk assessment quantification range to obtain a personalized hypertension risk assessment value.
[0012] A second aspect of this invention provides a personalized hypertension risk assessment system based on big data, comprising: The first unit is used to acquire physiological and behavioral data of the target individual and identify multiple risk factors. The second unit is used to construct continuous monitoring curves for each risk factor, detect step change points through cumulative sum control chart algorithms, extract the rate of change and duration of change of indicators before and after the step change point as step features, and construct a time series feature matrix. The third unit is used to identify the unidirectional temporal dependencies between risk factors based on the temporal data of each risk factor in the temporal feature matrix through Granger causality test, construct a directed acyclic risk propagation graph, and calculate the information transmission strength of each directed edge through transfer entropy to obtain the edge weight matrix. The fourth unit is used to perform graph isomorphic matching between the risk propagation map of the target individual and the risk propagation maps of each historical individual, filter out reference individuals through graph kernel functions, extract the cascade amplification law parameters along the propagation path, construct a multi-step propagation dynamics equation, and combine it with the edge weight matrix to obtain a personalized risk propagation model. The fifth unit is used to input the step characteristics of the target individual at the current moment into the personalized risk propagation model, simulate the risk signal to propagate in multiple stages along the directed edges of the risk propagation graph, calculate the cumulative risk intensity propagated to the hypertension onset node, and obtain the personalized hypertension risk assessment value.
[0013] A third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0014] Fourth aspect of the present invention, A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention provides a personalized hypertension risk assessment method based on big data. By constructing a time-series feature matrix and a risk propagation graph, it achieves accurate assessment of hypertension risk. This method can dynamically capture the step change characteristics of various physiological indicators and behavioral habit data, effectively identify the time-series dependencies between risk factors, and thus more accurately reflect an individual's actual health status.
[0016] 2. This invention constructs a directed acyclic risk propagation graph and its edge weight matrix through Granger causality testing and propagation entropy calculation, which can objectively quantify the propagation strength of risk signals. Reference individuals selected based on graph isomorphism matching technology provide a more targeted risk assessment basis for target individuals, making the assessment results more closely match individual characteristics and avoiding the general assessment bias found in traditional methods.
[0017] 3. This invention calculates the cumulative risk intensity by simulating the multi-stage propagation process of risk signals on a propagation graph, providing a more reliable basis for medical decision-making. This assessment method based on propagation dynamics not only considers the independent effects of each risk factor but also their complex interactions and cascade amplification effects, improving the timeliness and accuracy of hypertension risk warnings and facilitating early intervention and personalized prevention of the disease. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the personalized hypertension risk assessment method based on big data, as described in an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the construction process of a directed acyclic risk propagation graph and its edge weight matrix. Detailed Implementation
[0019] Example 1: Please refer to the appendix Figure 1-2 This invention provides a personalized hypertension risk assessment method and system based on big data. The technical solution of this invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0020] Figure 1 This is a flowchart illustrating the personalized hypertension risk assessment method based on big data according to an embodiment of the present invention. Figure 1 As shown, the overall method specifically includes: Step 1: Obtain physiological indicator data and behavioral habit data of the target individual, and identify multiple risk factors; Step 2: Construct continuous monitoring curves for each risk factor, detect step change points using the cumulative sum control chart algorithm, extract the rate of change and duration of change of indicators before and after the step change point as step features, and construct a time series feature matrix. Step 3: Based on the time series data of each risk factor in the time series feature matrix, the unidirectional time series dependency relationship between risk factors is identified through Granger causality test, a directed acyclic risk propagation graph is constructed, and the information transmission strength of each directed edge is calculated by the propagation entropy to obtain the edge weight matrix. Step 4: Perform graph isomorphic matching between the risk propagation graph of the target individual and the risk propagation graphs of each historical individual. Select reference individuals through graph kernel functions and extract the cascading amplification law parameters along the propagation path. Construct a multi-step propagation dynamics equation and combine it with the edge weight matrix to obtain a personalized risk propagation model. Step 5: Input the step characteristics of the target individual at the current moment into the personalized risk propagation model, and simulate the risk signal to propagate in multiple stages along the directed edges of the risk propagation graph. Calculate the cumulative risk intensity propagated to the hypertension onset node to obtain the personalized hypertension risk assessment value.
[0021] In an optional implementation, for step 1: constructing continuous monitoring curves for each risk factor, detecting step change points using a cumulative sum control chart algorithm, extracting the rate of change and duration of change of the indicators before and after the step change point as step features, and constructing a time series feature matrix, including: Continuous monitoring curves are constructed for each risk factor according to a preset sampling frequency; For the continuous monitoring curve, the cumulative deviation sequence is calculated by the cumulative sum control chart algorithm, and the cumulative deviation sequence is decomposed into multiple scale components by wavelet transform. Local extreme points are detected on each scale component, and the local extreme points are projected and fused on the time axis. The time points that appear on multiple scale components at the same time are marked as step change points. Extract the rate of change of the index and the duration of change within a preset time window before and after each step change point as step features, and calculate the fractal dimension difference of the monitoring curves before and after the step change point. The fractal dimension difference represents the risk factor and the degree of transformation from an ordered state to a chaotic state. Combine the rate of change of the index, the duration of change and the fractal dimension difference to form a step feature vector. The step feature vectors of each risk factor are aligned and arranged according to the time dimension to construct a time series feature matrix.
[0022] In the specific implementation of this step, when constructing continuous monitoring curves for each risk factor, it is necessary to determine the sampling frequency of the risk factor. For example, for blood pressure indicators, a sampling frequency of once every 4 hours can be used; for heart rate variability indicators, a sampling frequency of once per hour can be used. The collected data points are connected in chronological order to form a continuous monitoring curve. Assuming that a patient's systolic blood pressure is sampled at a frequency of 6 hours and continuously monitored for 14 days, 56 data points will be obtained, forming a complete monitoring curve. By analyzing these data, the patient's blood pressure control, diurnal pattern, and response to drug treatment can be assessed. At the same time, combined with the patient's dietary habits, exercise frequency, medication adherence, and other behavioral data, multiple risk factors affecting blood pressure fluctuations can be identified.
[0023] For the constructed continuous monitoring curve, the cumulative sum control chart algorithm is applied to detect step change points. This process first calculates the difference between the original curve and its mean, and then accumulates these differences to form a cumulative deviation sequence. For example, the original monitoring sequence is [5.2, 5.4, 5.3, 5.8, 6.2, 6.4, 6.3, 6.5], with a mean of 5.89. The calculated deviation sequence is [-0.69, -0.49, -0.59, -0.09, 0.31, 0.51, 0.41, 0.61]. After accumulation, the cumulative deviation sequence is [-0.69, -1.18, -1.77, -1.86, -1.55, -1.04, -0.63, -0.02].
[0024] Wavelet transform can be applied to the cumulative deviation sequence using the discrete wavelet transform method, which decomposes the sequence into multiple scale components. In practical applications, 3 to 5 decomposition levels are usually selected. For example, using the db4 wavelet basis to decompose the cumulative deviation sequence into 4 levels yields 4 detail coefficient sequences and 1 approximate coefficient sequence. For each decomposed coefficient sequence, local extrema are detected. Local extrema are points that achieve the maximum or minimum value within their left and right neighborhoods. In practice, the neighborhood size can be set to 3 to 5 data points. For the above sample data, local extrema were detected at the 5th and 6th points in the first level of detail coefficients, indicating that there is a step change at this point.
[0025] The local extrema detected on each scale component are projected onto the time axis. If a point at the same time is detected as a local extrema on multiple scale components, it is marked as a step change point. In practice, a threshold can be set requiring that the point appear simultaneously on at least two or three scale components. For the example above, if the fifth point is detected as a local extrema on three scale components, it is marked as a step change point, indicating that the risk factor has changed significantly at that time.
[0026] In a practical application case, the blood pressure data of a hypertensive patient over 30 consecutive days was analyzed. The data sampling frequency was once every 6 hours, and a total of 120 data points were obtained. Using the above algorithm, a significant step change point was detected on the 8th and 22nd days. The step change point on the 8th day corresponded to the patient starting to take a new antihypertensive drug, which caused the blood pressure to drop from an average of 145 / 90 mmHg to 130 / 85 mmHg. The step change point on the 22nd day corresponded to the patient's blood pressure rising from an average of 130 / 85 mmHg to 140 / 88 mmHg due to increased work pressure.
[0027] For the detected step change point, extract the features before and after it, and determine the size of the feature extraction window, which is usually 10 to 20 data points before and after the step point. Calculate the rate of change of the index before and after the step, that is, the percentage difference between the mean in the window after the step and the mean in the window before the step. For example, if the mean of the 10 points before the step is 5.3 and the mean of the 10 points after the step is 6.4, then the rate of change is (6.4-5.3) / 5.3×100%=20.75%.
[0028] The duration of change refers to the length of time after a step change that the new level remains stable. It can be determined by detecting the stability of the curve after the step change. If the standard deviation of the data points after the step change is lower than a preset threshold (such as 50% of the standard deviation of the original sequence) and lasts for a certain period of time (such as 30 data points), the change is considered stable, and this duration is the duration of change. For example, if the curve after the step change remains stable within 40 data points and the sampling frequency is 10 minutes, the duration of change is 400 minutes.
[0029] To calculate the fractal dimension of the monitoring curves before and after a step change, the box counting method can be used. The curves within the time window before and after the step change are divided into grids of different scales. The number of non-empty grids at each scale is counted, and the fractal dimension is determined by the logarithmic relationship. For example, if the fractal dimension of the curve before the step change is 1.25 and after the step change it is 1.48, then the difference in fractal dimension is 0.23, indicating that the system complexity has increased and the risk factor has changed from a relatively ordered state to a more chaotic state.
[0030] The extracted rate of change, duration of change, and fractal dimension difference are combined to form a step feature vector. For example, for a detected step point, its feature vector can be represented as [20.75%, 400 minutes, 0.23]. For multiple risk factors, step feature vectors are extracted separately and arranged in alignment according to the time dimension to construct a time series feature matrix. Assuming that 5 risk factors are monitored and 3 step change points are detected for each factor, the size of the constructed time series feature matrix is 5×3×3, where 5 represents the number of risk factors, the first 3 represents the number of step points for each factor, and the second 3 represents the feature dimension of each step point.
[0031] This time-series feature matrix captures the key change characteristics of risk factors over time, and can be used for subsequent tasks such as risk warning model training, risk correlation analysis, and abnormal pattern recognition, thereby improving the accuracy and timeliness of risk monitoring.
[0032] Figure 2 This is a schematic diagram illustrating the construction process of a directed acyclic risk propagation graph and its edge weight matrix.
[0033] In an optional implementation, step 3 specifically involves: based on the time-series data of each risk factor in the time-series feature matrix, identifying the unidirectional time-series dependencies between risk factors through Granger causality tests, constructing a directed acyclic risk propagation graph, and calculating the information transmission strength of each directed edge through propagation entropy to obtain the edge weight matrix.
[0034] The above steps include: calculating the predictive contribution of the first risk factor to the second risk factor based on the time series data of each risk factor through Granger causality test; when it exceeds the significance threshold, determining the one-way time series dependency from the first risk factor to the second risk factor; performing the Granger causality test on all risk factor pairs to extract all one-way time series dependencies. Each risk factor is used as a node, and the unidirectional temporal dependency is used as a directed edge to construct an initial directed graph. Loop detection is performed on the initial directed graph. When a loop is detected, the time lag order between risk factor pairs corresponding to each directed edge in the loop is calculated, and the directed edge with the largest time lag order is deleted to obtain a directed acyclic risk propagation graph. For each directed edge in the risk propagation graph, its time-series data is symbolically encoded to obtain the source symbol sequence and the target symbol sequence. The relative entropy between their joint probability distribution and the marginal probability distribution is calculated. The relative entropy is used as the information transmission strength of the directed edge. The edges are then arranged in a matrix according to the index positions of the starting and ending nodes of the directed edges to obtain the edge weight matrix.
[0035] In this embodiment, a feature matrix containing time-series data of multiple risk factors needs to be obtained. The rows of the matrix represent different time points, the columns represent different risk factors, and the element values in the matrix represent the risk factor values at the corresponding time points. For example, in the case of hypertension risk analysis, observation data of 10 risk factors, including systolic blood pressure, diastolic blood pressure, heart rate, body mass index, blood glucose level, blood lipid level, sodium intake, daily exercise volume, stress index, and sleep quality, can be obtained over the past 24 weeks to form a 24×10 time-series feature matrix.
[0036] Based on the acquired time-series feature matrix, Granger causality tests are performed on all risk factor pairs to identify unidirectional time-series dependencies between risk factors. The core idea of Granger causality tests is to evaluate the predictive contribution of a risk factor to a target risk factor by comparing the performance of predictive models that include and do not include historical data of a specific risk factor. In specific implementation, a maximum lag order p is set. This parameter indicates that the furthest historical time point is considered when predicting the risk factor value at the current time point. For hypertension monitoring data, if the data collection frequency is once a week, the lag order p can be set to 4 to 8, indicating that historical data from the previous 4 to 8 weeks are considered.
[0037] For each pair of risk factors, a Granger causality test is performed. Assuming there are m risk factors in the time series feature matrix, m×(m-1) tests are required to examine each one-way dependency. For risk factors i and j, two regression models are constructed: a restricted model and an unrestricted model. The restricted model uses only the historical data of risk factor j itself to predict its current value; the unrestricted model uses the historical data of both risk factor j and risk factor i to predict the current value of risk factor j. Both models adopt a linear autoregressive structure, and the least squares method is used to estimate the model parameters.
[0038] After the model is constructed, the residual sums of squares for the restricted and unrestricted models are calculated and denoted as RSSR and RSSU, respectively. Based on these two residual sums of squares, the F-statistic is calculated. The formula for the F-statistic is: (RSSR-RSSU) / p divided by RSSU / (n-2p-1), where n is the length of the time series. The F-statistic follows an F-distribution with (p, n-2p-1) degrees of freedom. The calculated F-statistic is compared with the critical value at a preset significance level α (usually taken as 0.05 or 0.01). If the F-statistic is greater than the critical value, the null hypothesis is rejected, and it is considered that risk factor i has a Granger causal relationship with risk factor j.
[0039] In practical applications, it is also necessary to solve the problem of choosing the lag order p. Too large a p will lead to too many model parameters and increase the risk of overfitting; too small a p will fail to capture sufficiently long-term dependencies. To solve this problem, the information criterion method is used to automatically select the optimal lag order. Specifically, for each pair of risk factors, different lag orders (from 1 to the maximum value pmax) are tried, and the corresponding Bayesian information criterion (BIC) value is calculated. The lag order that minimizes the BIC value is selected as the optimal lag order for that pair of risk factors.
[0040] To improve the robustness of the test results, a multiple test correction mechanism was introduced. Since m×(m-1) hypothesis tests were performed simultaneously, there was a problem of increased false positive results due to multiple comparisons. The Bonferroni correction method was adopted, which divides the significance level α by the total number of tests to obtain the adjusted significance level α'. Granger causality was only confirmed when the p-value corresponding to the F statistic was less than the adjusted significance level α'.
[0041] It also implements the function of causal strength quantification, which measures the influence of risk factor i on risk factor j by calculating the predictive contribution. The predictive contribution is defined as (RSSR-RSSU) / RSSR, which represents the relative reduction in the prediction error of risk factor j after adding the historical data of risk factor i. The value range of the predictive contribution is 0 to 1. The larger the value, the greater the predictive contribution of risk factor i to risk factor j.
[0042] To illustrate this step, in a specific application example, monitoring data from a group of hypertensive patients were analyzed, including five risk factors: systolic blood pressure, diastolic blood pressure, heart rate, sodium intake, and daily average exercise. Weekly averages were continuously monitored for 24 weeks, forming a 24×5 time-series characteristic matrix. Using a lag order of p=4 and a significance level of α=0.05, 5×4=20 Granger causality tests were performed. The results showed that the F-statistic for sodium intake on systolic blood pressure was 4.82, corresponding to a p-value of 0.008, which was less than the significance level. With a level of 0.05 and a predicted contribution of 0.31, a unidirectional temporal dependency from sodium intake to systolic blood pressure was confirmed. Similarly, a unidirectional dependency from daily exercise volume to diastolic blood pressure (F-statistic = 3.95, p-value = 0.017, predicted contribution = 0.26) and a unidirectional dependency from heart rate to systolic blood pressure (F-statistic = 3.12, p-value = 0.042, predicted contribution = 0.19) were also detected. Using risk factors as nodes and unidirectional temporal dependencies as directed edges, an initial directed graph was constructed.
[0043] Loop detection is performed on the initial directed graph. A depth-first search algorithm is used to traverse all nodes in the graph to determine if a loop exists. When a loop is detected, it needs to be broken to ensure the acyclicity of the risk propagation graph. For each directed edge in the loop, the time lag order of the corresponding risk factor pair is calculated. The time lag order refers to the minimum lag period in Granger causality tests where the historical data of the source node has predictive power for the current value of the target node. For example, if the prediction of risk factor 1 for risk factor 4 requires data from at least 3 periods ago, then the lag order of this edge is 3. The directed edge with the largest lag order in the loop is found and deleted. For example, in the loop formed by risk factor 1→4→7→1, if the lag orders of the three edges are 3, 2, and 4 respectively, then the edge with a lag order of 4 (the edge from risk factor 7 to risk factor 1) is deleted.
[0044] Perform the above operation on all cycles in the initial directed graph until there are no cycles in the graph, and obtain a directed acyclic risk propagation graph. For example, if you need to delete 3 edges and keep 12 directed edges to form an acyclic graph.
[0045] To quantify the information transmission intensity of directed edges in the risk propagation graph, the propagation entropy of each directed edge is calculated. For each directed edge in the risk propagation graph, the time series data of the source node and the target node are extracted and symbol encoding is performed. Symbol encoding converts continuous time series data into discrete symbol sequences to calculate the probability distribution. A simple encoding method is to encode growth as "1", decrease as "0", and no change as "2" according to the data change trend. For example, the original time series data of risk factor 1 [3.2, 3.4, 3.3, 3.6, 3.5] is encoded into the symbol sequence [1, 0, 1, 0].
[0046] Based on the encoded symbol sequence, the joint probability distribution and marginal probability distribution of the source and target symbol sequences are calculated. For example, the frequency of both the source and target sequences being "1" is used as the joint probability P(source = "1", target = "1"). The difference between the conditional entropy of the target sequence and its own entropy, known from the source sequence, is calculated as the transfer entropy, which is used as the information transfer strength of the directed edge. For example, the transfer entropy from risk factor 1 to risk factor 4 is calculated to be 0.28, indicating that risk factor 1 provides 0.28 of information for predicting risk factor 4.
[0047] The calculated transfer entropy values are arranged in a matrix according to the index positions of the start and end nodes of the directed edges to obtain the edge weight matrix. This matrix is an n×n square matrix, where n is the number of risk factors. The element in the i-th row and j-th column of the matrix represents the transfer entropy value from risk factor i to risk factor j. If there is no directed edge between two factors, the value at the corresponding position is 0. For example, in the case of 10 risk factors, a 10×10 edge weight matrix is obtained. The 12 non-zero elements in the matrix correspond to the information transfer intensity of the 12 directed edges in the risk propagation graph.
[0048] The risk transmission network constructed using the above methods not only reveals the dependency structure between risk factors but also quantifies the intensity of information transmission, providing an important basis for risk management and decision-making.
[0049] In an optional implementation, step 33: for each directed edge in the risk propagation graph, its time-series data is symbolically encoded to obtain a source symbol sequence and a target symbol sequence, and the relative entropy between their joint probability distribution and the marginal probability distribution is calculated. The relative entropy is used as the information transmission strength of the directed edge, including: For each directed edge, the time series data of its starting risk factor is used as the source time series data, and the time series data of its ending risk factor is used as the target time series data. The source time series data and target time series data are divided into multiple time periods. The data distribution characteristics in each time period are calculated, the quantile threshold of symbolic encoding is determined, and the time series data in each time period are mapped to discrete symbols to form source symbol sequences and target symbol sequences. The source symbol sequence and the target symbol sequence are aligned with a fixed time lag step. Continuous symbol substrings are extracted from the source symbol sequence by sliding with a preset embedding dimension. At the same time, a single symbol in the target symbol sequence at the corresponding time is extracted. The joint occurrence frequency of all symbol substrings and single symbols is counted and normalized to obtain the joint probability distribution. The frequency of edge occurrences of source symbol substrings and target single symbols is counted separately, and the edge probability distribution is obtained by normalization. The ratio of the product of the joint probability distribution and the edge probability distribution is calculated and the logarithm is taken. The relative entropy is obtained by multiplying the product of the joint probability distribution and summing the logarithms. This relative entropy is used as the information transmission strength of each directed edge.
[0050] In this specific embodiment, the risk propagation graph consists of multiple risk factor nodes and directed connections between them. For each directed edge in the graph, the time series data of the starting risk factor must be identified as the source time series data, and the time series data of the ending risk factor must be identified as the target time series data. For example, if there is a directed edge from factor A to factor B, then the time series data of A is taken as the source time series data, and the time series data of B is taken as the target time series data. These time series data are usually represented as continuously sampled numerical sequences.
[0051] To calculate the intensity of information transmission, the source and target time series data are divided into several time periods. Within each time period, the data distribution characteristics are calculated, including statistical measures such as mean, variance, and quartiles. Based on these statistical characteristics, the quantile thresholds for symbolic coding are determined. The commonly used coding method is ternary coding, which divides the data range into three regions based on the 33.3% and 66.7% quantiles, represented by the symbols '0', '1', and '2', respectively. Divisible (median) or quartile coding can also be used, and the choice can be made flexibly according to the required analytical precision.
[0052] Time series data within each time period are mapped to discrete symbols. For example, if the ternary digits of blood pressure in a certain time period are [105.3, 120.8], then data below 105.3 is mapped to '0', data between 105.3 and 120.8 is mapped to '1', and data above 120.8 is mapped to '2'. In this way, the source time series data and the target time series data are converted into source symbol sequences and target symbol sequences, respectively.
[0053] Considering the time lag effect in risk propagation, the source symbol sequence and the target symbol sequence are aligned with a fixed time lag step. In financial risk propagation analysis, a lag step of 1-5 days is commonly used. This embodiment uses a lag step of 2 days, assuming that the impact of the source risk factor on the target risk factor will become apparent after 2 days. Continuous symbol substrings are extracted from the source symbol sequence using a preset embedding dimension, and a single symbol from the target symbol sequence at the corresponding time is extracted simultaneously. The embedding dimension represents the length of the historical state considered; in this embodiment, it is set to 3, meaning that 3 consecutive symbols are extracted from the source sequence each time to form a substring.
[0054] Suppose a source symbol sequence segment is "01201" and a target symbol sequence segment is "12012", with a lag step size of 2 and an embedding dimension of 3. Extract the symbol substring "012" from positions 0-2 in the source sequence and the symbol "0" from position 2 in the target sequence to form a joint observation ("012", "0"). Extract the symbol substring "120" from positions 1-3 in the source sequence and the symbol "1" from position 3 in the target sequence to form another joint observation ("120", "1"); and so on.
[0055] The frequency of joint occurrence of all symbol substrings and single symbols is counted. If ternary encoding is used and the embedding dimension is 3, there are 27 combinations of source symbol substrings (33) and 3 values of target single symbols. A total of 81 joint states need to be counted. These frequencies are normalized (divided by the total number of observations) to obtain the joint probability distribution P(x, y), where x represents the source symbol substring and y represents the target single symbol.
[0056] The edge occurrence frequencies of the source symbol substring and the target single symbol are counted separately and normalized to obtain the edge probability distributions P(x) and P(y). For example, P(x="012") is obtained by counting the number of observations of all substrings with "012" as the source symbol and dividing by the total number of observations; similarly, P(y="0") is calculated.
[0057] Calculate the ratio of the product of the joint probability distribution and the marginal probability distribution, and take the logarithm. Specifically, for each joint state (x, y), calculate log(P(x, y) / (P(x)×P(y))). Multiply these logarithmic values by the corresponding joint probability P(x, y) and sum them to obtain the relative entropy value. This relative entropy is the information transmission strength of the directed edge.
[0058] In practical applications, if the calculated information transmission strength is 0.25, it indicates that there is a significant information flow from the source risk factor to the target risk factor, which needs to be closely monitored. In contrast, if the information transmission strength is only 0.05, it indicates that the impact of this propagation path is weak. By comparing the information transmission strength of different directed edges, risk managers can identify the key propagation paths in the risk network and formulate targeted risk control measures.
[0059] In an optional implementation, for step 4: performing graph isomorphic matching between the risk propagation map of the target individual and the risk propagation maps of each historical individual, filtering out reference individuals using a graph kernel function, extracting their cascading amplification law parameters along the propagation path, constructing a multi-step propagation dynamics equation, and combining it with the edge weight matrix to obtain a personalized risk propagation model, including: The risk propagation graph of the target individual and the risk propagation graph of each historical individual are topologically encoded to generate corresponding node degree distribution sequences and adjacency relationship matrices. Based on the node degree distribution sequence and the adjacency matrix, the graph kernel function value between the target individual and each historical individual is calculated. The graph kernel function value is obtained by weighted summation of the local subgraph structure similarity of each node. Historical individuals whose graph kernel function value exceeds the similarity threshold are selected as reference individuals. For the risk propagation graph of the reference individual, the directed propagation path from the initial risk factor to the hypertension onset node is extracted, and the stepwise amplification coefficient is calculated. The stepwise amplification coefficient is the product of the edge weight between adjacent nodes on the propagation path and the cumulative risk intensity of the preceding sequence. Based on the stepwise amplification coefficient of all directed propagation paths of each reference individual, the cascade amplification law parameter is calculated. Based on the cascade amplification law parameters, a multi-step propagation dynamic equation describing the cumulative enhancement mechanism of risk signals during multi-level propagation is constructed, and associated with the risk propagation graph and edge weight matrix of the target individual to obtain a personalized risk propagation model.
[0060] In this embodiment, when performing topological encoding on the risk propagation graph of a target individual, it is necessary to collect the individual's physiological data and lifestyle information, including risk factors such as blood pressure, blood lipids, body mass index, smoking habits, and dietary structure. These risk factors are used as nodes in the graph, and the influence relationships between risk factors are used as edges. A unique identifier is assigned to each node, recording its position and connectivity in the graph. The degree of each node, i.e., the number of edges connected to that node, is calculated to obtain the node degree distribution sequence. At the same time, an adjacency matrix is constructed, where elements indicate whether there is a connection between nodes (1 for existence, 0 for non-existence). For example, for a risk propagation graph containing 5 nodes, its node degree distribution sequence is [3, 2, 4, 1, 2], indicating that the 5 nodes have 3, 2, 4, 1, and 2 connecting edges, respectively.
[0061] The risk propagation graphs of historical individuals are encoded in the same way to construct a database. This database contains a large number of individual risk propagation graphs and their topological coding information for individuals with known hypertension development trajectories. For example, the node degree distribution sequence of a certain historical individual is [4, 2, 3, 1, 2], and its adjacency matrix records the connections between each node.
[0062] When calculating the graph kernel function value between the target individual and each historical individual, a node-matching-based graph kernel method is used. For the target individual graph G1 and the historical individual graph G2, the local subgraph structure of each node is extracted, including the star subgraph composed of the central node and its first-order neighbors. The similarity of the local subgraph structures of each node in G1 and G2 is compared. The similarity calculation considers factors such as the difference in node degree and the difference in degree distribution of neighboring nodes. The similarity of all node pairs is weighted and summed to obtain the overall graph kernel function value. A similarity threshold of 0.85 is set, and historical individuals with graph kernel function values exceeding this threshold are selected as reference individuals. In practical applications, about 50 reference individuals are selected from 1000 historical individuals. These individuals have a high similarity to the target individual in terms of risk propagation structure.
[0063] When extracting the directed propagation path of the risk propagation map of a reference individual, the initial risk factor node (such as smoking, high-salt diet, etc.) and the endpoint node (onset of hypertension) are determined. The breadth-first search algorithm is used to identify all paths from the initial node to the endpoint node. For example, a reference individual may have two main propagation paths: "high-salt diet → increased blood volume → increased cardiac load → hypertension" and "smoking → vascular endothelial damage → decreased vascular elasticity → hypertension".
[0064] When calculating the stepwise amplification factor, the influence strength between adjacent nodes on each propagation path is analyzed. The edge weight represents the degree of direct influence of the upstream node on the downstream node, and the value ranges from 0 to 1. The preceding cumulative risk intensity represents the cumulative effect of the risk in the propagation process. The stepwise amplification factor between adjacent nodes is equal to the product of the edge weight and the preceding cumulative risk intensity. For example, in the propagation segment of "high salt diet → increased blood volume", if the edge weight is 0.7 and the preceding cumulative risk intensity is 1.2, then the stepwise amplification factor is 0.84.
[0065] Based on the propagation paths and step-by-step amplification coefficients of all reference individuals, cascading amplification parameters are calculated. These parameters describe how risk accumulates and amplifies along a specific path. The cascading amplification parameters include the base amplification rate, the accumulation coefficient, and the path specificity coefficient. The base amplification rate describes the natural growth rate of risk when no other factors intervene; the accumulation coefficient describes the synergistic effect of multiple risk factors; and the path specificity coefficient describes the unique properties of different propagation paths.
[0066] When constructing the multi-step propagation dynamics equation, the risk propagation process is simulated with discrete time steps. The equation describes how the risk state of each node at each time step is affected by the state of its upstream nodes. The equation incorporates cascade amplification parameters, enabling the model to capture the nonlinear growth characteristics of risk during the propagation process. For example, the risk state of a node will increase exponentially with the accumulation of upstream risks, rather than a simple linear superposition.
[0067] By linking the multi-step propagation dynamics equation with the risk propagation graph and edge weight matrix of the target individual, a personalized risk propagation model is formed. The elements in the edge weight matrix represent the influence intensity between nodes in the risk propagation graph of the target individual. Customized according to the actual physiological data and lifestyle of the target individual, the final personalized risk propagation model can simulate the propagation path and intensity changes of specific risk factors in the target individual and predict the cumulative process of hypertension risk.
[0068] During the validation phase, the control group used the traditional risk scoring method, while the experimental group used the personalized risk transmission model constructed using this method. One hundred subjects were followed up for three years. Results showed that this method achieved an accuracy rate of 92% in identifying high-risk individuals, a 9 percentage point improvement compared to the control group's 83%, particularly demonstrating superior performance in cases with multiple coexisting risk factors.
[0069] In an optional implementation, for step 44: based on the cascade amplification law parameters, a multi-step propagation dynamics equation describing the cumulative enhancement mechanism of the risk signal during multi-stage propagation is constructed, and associated with the risk propagation graph and edge weight matrix of the target individual to obtain a personalized risk propagation model, including: The parameters of the cascade amplification law are fitted with a function to obtain the amplification coefficient function between the step amplification coefficient and the propagation level, and a multi-step propagation dynamics equation is constructed. The multi-step propagation dynamics equation sets the current observation value of the risk factor as the initial risk intensity, and performs iterative calculations according to the propagation level order, accumulating the calculations up to the propagation level where the hypertension onset node is located. The directed edges in the risk propagation graph of the target individual are divided into multiple propagation levels, and the edge weights of the directed edges are extracted. The average value is calculated as the level propagation efficiency, and multiplied with the output value of the amplification coefficient function at the corresponding level to obtain the personalized amplification coefficient. Combined with the multi-step propagation dynamics equation, a personalized risk propagation model is obtained.
[0070] In this embodiment, it is necessary to obtain cascade amplification law parameters. These parameters describe the amplification characteristics of risk signals at different propagation levels. Specifically, through clinical data analysis, quantitative parameters of the stepwise amplification phenomenon of risk factors such as blood glucose, blood lipids, and age from the initial level to the hypertension onset level can be extracted. For example, the analysis of 5-year follow-up data of 1,000 subjects shows that blood glucose, as a risk factor, has an average amplification coefficient of 1.2 at the first level, 1.5 at the second level, 1.9 at the third level, and 2.4 at the fourth level. This stepwise increasing amplification coefficient indicates that the risk signal exhibits non-linear growth characteristics during propagation.
[0071] The obtained cascade amplification law parameters are fitted with functions to establish the mathematical relationship between the amplification coefficient and the propagation level. A polynomial fitting method can be used to fit the discrete amplification coefficient data points into a continuous function. Taking the blood glucose factor as an example, by fitting the amplification coefficients of the above four levels, the amplification coefficient function can be obtained. This function can be expressed as a polynomial expression of the propagation level, so that the corresponding amplification coefficient value can be calculated for any level k.
[0072] A multi-step propagation dynamics equation is constructed to describe how risk signals accumulate and amplify step by step along the propagation path. In the equation, the current observed value of the risk factor is set as the initial risk intensity. Taking the blood glucose value of the target individual of 6.1 mmol / L as an example, it is used as the initial risk intensity input into the equation. The multi-step propagation dynamics equation calculates the evolution of risk intensity in the order of propagation levels through iteration. Each iteration represents the risk signal passing through a propagation level. The risk intensity will be adjusted according to the amplification coefficient corresponding to that level. The iterative calculation continues until the propagation level where the hypertension onset node is located is reached.
[0073] Obtain the risk propagation graph and edge weight matrix for the target individual. The risk propagation graph is a directed graph where nodes represent physiological indicators or pathological states, and directed edges represent risk propagation paths. For example, a simplified risk propagation graph contains a propagation path such as "blood glucose" → "insulin resistance" → "endothelial dysfunction" → "blood pressure regulation abnormality" → "hypertension". The edge weight matrix stores the weight value of each directed edge, reflecting the efficiency of risk propagation along that path. These weights can be determined by analyzing the target individual's historical medical data and family medical history.
[0074] The directed edges in the risk propagation graph are divided into multiple propagation levels. The division method is based on the shortest path analysis from the initial risk factor node to the hypertension onset node. For example, for the simplified path mentioned above, "blood glucose" → "insulin resistance" can be divided into the first level, "insulin resistance" → "endothelial dysfunction" into the second level, and so on. For each propagation level, the edge weights of all directed edges in that level are extracted, and their average value is calculated as the propagation efficiency of that level. Taking the target individual data as an example, the average edge weight is 0.75 for the first level, 0.82 for the second level, 0.88 for the third level, and 0.92 for the fourth level.
[0075] Calculate the personalized amplification factor. Multiply the propagation efficiency of each propagation level by the output value of the amplification factor function at the corresponding level to obtain the personalized amplification factor. This step combines the general amplification law with the individual-specific propagation efficiency, realizing the personalized adjustment of the risk propagation model. For example, if the output of the amplification factor function at the first level is 1.2 and the propagation efficiency at that level is 0.75, then the personalized amplification factor is 0.9.
[0076] By substituting the personalized amplification factor into the multi-step propagation dynamics equation, a complete personalized risk propagation model is constructed. This model can accurately predict the cumulative risk intensity at the hypertension onset point after the risk signal has been propagated through multiple levels, based on the initial risk factor values of the target individual. For the target individual in the aforementioned example with an initial blood glucose value of 6.1 mmol / L, the model can calculate that the cumulative risk intensity at the hypertension onset point after four levels of propagation is 15.8, which is much higher than the risk threshold of 10, indicating that the individual has a high risk of developing hypertension.
[0077] The personalized risk transmission model constructed using the above methods can accurately characterize the cumulative amplification effect of risk signals in the multi-level transmission process, providing an effective tool for the early identification and precise prevention of hypertension risk. Compared with traditional scoring methods, this model considers the dynamic process of risk transmission and individual differences, and its prediction accuracy is improved, providing strong support for precision medicine and personalized health management.
[0078] In an optional implementation, for step 5: inputting the step characteristics of the target individual at the current moment into the personalized risk propagation model, simulating the multi-level propagation of the risk signal along the directed edges of the risk propagation graph, calculating the cumulative risk intensity propagated to the hypertension onset node, and obtaining a personalized hypertension risk assessment value, including: The step characteristics of the target individual at the current moment are mapped to the corresponding risk factor nodes in the risk propagation graph to obtain the initial risk intensity value of each source node; According to the topological order of the risk propagation graph, starting from each source node, the upstream risk signal received by each intermediate node is calculated sequentially, and the upstream risk signal is weighted and combined with the corresponding step-by-step amplification coefficient to obtain the propagation risk intensity value of the intermediate node. The post-propagation risk intensity values of all propagation paths along the risk propagation map reaching the hypertension onset node are nonlinearly fused. Based on the path length and the number of nodes on each propagation path, a differentiated weight is applied to the post-propagation risk intensity values to obtain the cumulative risk intensity of the hypertension onset node. The cumulative risk intensity is mapped to a preset risk assessment quantification range to obtain a personalized hypertension risk assessment value.
[0079] In practical applications, the step characteristics of a target individual at the current moment are obtained, including multi-dimensional characteristics such as age, gender, blood pressure level, body mass index, blood glucose level, blood lipid level, smoking status, alcohol consumption, exercise frequency, dietary habits, and family medical history. For example, for a 45-year-old male with a body mass index of 28.5, systolic blood pressure of 135 mmHg, diastolic blood pressure of 88 mmHg, fasting blood glucose of 6.2 mmol / L, triglycerides of 1.9 mmol / L, low-density lipoprotein of 3.6 mmol / L, a smoking habit (20 cigarettes per day), occasional alcohol consumption, exercise twice a week, a high-salt diet, and a father with hypertension, these characteristics constitute the step characteristics of the target individual at the current moment.
[0080] These step features are mapped to the corresponding risk factor nodes in a pre-constructed risk propagation graph. The risk propagation graph is a directed acyclic graph consisting of multiple risk factor nodes and directed edges between them. The mapping process assigns an initial risk intensity value to each source node based on the feature type and value range. For example, in the case above, the initial risk intensity value of the age node is set to 0.65 (because 45 years old is in the middle-aged range), the body mass index node is set to 0.75 (because 28.5 is in the overweight range), the systolic blood pressure node is set to 0.6 (because 135 mmHg is in the high normal range), the smoking node is set to 0.8 (because 20 cigarettes a day is heavy smoking), and the dietary habits node is set to 0.7 (because a high-salt diet increases the risk of hypertension), etc.
[0081] Risk propagation calculations are performed according to the topological order of the risk propagation graph. This topological order ensures that the risk intensities of all predecessor nodes have been calculated before the risk intensity of a given node is calculated. Starting from each source node, the risk signal propagates along the directed edges to downstream nodes. For each intermediate node, the risk signals received from all its upstream nodes are calculated and weighted together with the corresponding amplification coefficients. For example, an obese node receives risk signals from three upstream nodes: body mass index, dietary habits, and exercise frequency. These three risk signals are multiplied by their respective amplification coefficients (e.g., 0.6, 0.5, and 0.4), resulting in a post-propagation risk intensity value of 0.75 × 0.6 + 0.7 × 0.5 + 0.4 × 0.4 = 0.88. Similarly, a node with abnormal blood sugar receives risk signals from both dietary habits and obesity nodes, resulting in a post-propagation risk intensity value of 0.72.
[0082] The risk propagation network is calculated for each node in the topological order until the risk signal spreads throughout the entire network. This risk signal propagates to the hypertension-causing node via multiple different paths, each contributing differently to the final hypertension risk. For example, obesity can affect insulin resistance, thereby affecting vascular endothelial function and ultimately leading to hypertension; smoking can directly affect vasoconstriction and indirectly affect the occurrence of hypertension through oxidative stress.
[0083] After the risk signals of all paths are calculated, the risk intensity values of all propagation paths leading to the hypertension onset node are nonlinearly fused. The fusion process considers the path length and the number of nodes on the path, and applies differentiated weights to the risk intensity of different propagation paths. For example, the shorter the path, the greater the weight; the greater the influence of key nodes on the path (such as blood pressure level, kidney function, etc.). Through this weighting method, the cumulative risk intensity of the hypertension onset node is calculated. For the above case, it is assumed that the cumulative risk intensity value of the hypertension onset node calculated through each path is 0.82.
[0084] The calculated cumulative risk intensity value is mapped to a preset risk assessment quantification range to obtain a personalized hypertension risk assessment value. For example, the cumulative risk intensity of 0-1 is mapped to a five-level risk level: 0-0.2 is low risk, 0.2-0.4 is low-to-medium risk, 0.4-0.6 is medium risk, 0.6-0.8 is medium-to-high risk, and 0.8-1 is high risk. Based on the cumulative risk intensity of 0.82 calculated in the above case, the hypertension risk assessment result of this individual is "high risk", indicating that the individual has a high risk of developing hypertension in the next five years.
[0085] Finally, based on the risk assessment results, personalized intervention recommendations can be generated for the target individual, such as weight loss, increased exercise frequency, smoking cessation and alcohol limitation, reduced salt intake, and regular blood pressure monitoring, in order to reduce the risk of hypertension. This assessment method based on the risk propagation model can comprehensively consider the complex interactions and nonlinear relationships between multiple risk factors and provide a more accurate personalized hypertension risk assessment.
[0086] This invention provides a personalized hypertension risk assessment system based on big data, comprising: The first unit is used to acquire physiological and behavioral data of the target individual and identify multiple risk factors. The second unit is used to construct continuous monitoring curves for each risk factor, detect step change points through cumulative sum control chart algorithms, extract the rate of change and duration of change of indicators before and after the step change point as step features, and construct a time series feature matrix. The third unit is used to identify the unidirectional temporal dependencies between risk factors based on the temporal data of each risk factor in the temporal feature matrix through Granger causality test, construct a directed acyclic risk propagation graph, and calculate the information transmission strength of each directed edge through transfer entropy to obtain the edge weight matrix. The fourth unit is used to perform graph isomorphic matching between the risk propagation map of the target individual and the risk propagation maps of each historical individual, filter out reference individuals through graph kernel functions, extract the cascade amplification law parameters along the propagation path, construct a multi-step propagation dynamics equation, and combine it with the edge weight matrix to obtain a personalized risk propagation model. The fifth unit is used to input the step characteristics of the target individual at the current moment into the personalized risk propagation model, simulate the risk signal to propagate in multiple stages along the directed edges of the risk propagation graph, calculate the cumulative risk intensity propagated to the hypertension onset node, and obtain the personalized hypertension risk assessment value.
[0087] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0088] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0089] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A personalized hypertension risk assessment method based on big data, characterized in that, Includes the following steps: S1. Obtain physiological indicator data and behavioral habit data of the target individual, and identify multiple risk factors; S2. Construct continuous monitoring curves for each risk factor, detect step change points through cumulative sum control chart algorithm, extract the rate of change and duration of change of indicators before and after the step change point as step features, and construct time series feature matrix. S3. Based on the time series data of each risk factor in the time series feature matrix, the unidirectional time series dependency relationship between risk factors is identified through Granger causality test, a directed acyclic risk propagation graph is constructed, and the information transmission strength of each directed edge is calculated by the transfer entropy to obtain the edge weight matrix. S4. Perform graph isomorphic matching between the risk propagation graph of the target individual and the risk propagation graph of each historical individual, filter out reference individuals through graph kernel function, extract the cascade amplification law parameters along the propagation path, construct a multi-step propagation dynamics equation, and combine it with the edge weight matrix to obtain a personalized risk propagation model. S5. Input the step characteristics of the target individual at the current moment into the personalized risk propagation model, and simulate the risk signal to propagate in multiple stages along the directed edges of the risk propagation graph. Calculate the cumulative risk intensity propagated to the hypertension onset node to obtain a personalized hypertension risk assessment value.
2. The personalized hypertension risk assessment method based on big data according to claim 1, characterized in that, S2 includes the following steps: S21. Construct continuous monitoring curves for each risk factor according to a preset sampling frequency; S22. For the continuous monitoring curve, the cumulative deviation sequence is calculated by the cumulative sum control chart algorithm, and the cumulative deviation sequence is decomposed into multiple scale components by wavelet transform. Local extreme points are detected on each scale component, and the local extreme points are projected and fused on the time axis. The time points that appear on multiple scale components at the same time are marked as step change points. S23. Extract the rate of change of the index and the duration of change within a preset time window before and after each step change point as step features, and calculate the fractal dimension difference of the monitoring curves before and after the step change point. The fractal dimension difference represents the risk factor and the degree of transformation from an ordered state to a chaotic state. Combine the rate of change of the index, the duration of change and the fractal dimension difference to form a step feature vector. S24. Align and arrange the step feature vectors of each risk factor according to the time dimension to construct a time series feature matrix.
3. The personalized hypertension risk assessment method based on big data according to claim 1, characterized in that, S3 includes the following steps: S31. Based on the time series data of each risk factor, calculate the predictive contribution of the first risk factor to the second risk factor through the Granger causality test. When it exceeds the significance threshold, determine the one-way time series dependency from the first risk factor to the second risk factor. Perform the Granger causality test on all risk factor pairs to extract all one-way time series dependencies. S32. Using each risk factor as a node and the unidirectional time-series dependency as a directed edge, construct an initial directed graph and perform loop detection on the initial directed graph. When a loop is detected, calculate the time lag order between risk factor pairs corresponding to each directed edge in the loop, delete the directed edge with the largest time lag order, and obtain a directed acyclic risk propagation graph. S33. For each directed edge in the risk propagation graph, its time-series data is symbolically encoded to obtain the source symbol sequence and the target symbol sequence. The relative entropy between their joint probability distribution and the marginal probability distribution is calculated. The relative entropy is used as the information transmission strength of the directed edge. The edges are then arranged in a matrix according to the index positions of the starting and ending nodes of the directed edges to obtain the edge weight matrix.
4. The personalized hypertension risk assessment method based on big data according to claim 3, characterized in that, S33 specifically includes the following steps: S331. For each directed edge, the time series data of its starting risk factor is used as the source time series data, and the time series data of its ending risk factor is used as the target time series data. S332. Divide the source time series data and target time series data into multiple time periods, calculate the data distribution characteristics in each time period, determine the quantile threshold of symbolic encoding, and map the time series data in each time period into discrete symbols to form a source symbol sequence and a target symbol sequence. S333. Align the source symbol sequence and the target symbol sequence with a fixed time lag step, slide to extract continuous symbol substrings in the source symbol sequence with a preset embedding dimension, and simultaneously extract individual symbols in the target symbol sequence at the corresponding time. Calculate the joint occurrence frequency of all symbol substrings and individual symbols, and normalize to obtain the joint probability distribution. S334. Count the frequency of edge occurrence of source symbol substring and target single symbol respectively, and normalize to obtain edge probability distribution. Calculate the ratio of the product of joint probability distribution and edge probability distribution and take the logarithm. Multiply it with joint probability distribution and sum to obtain relative entropy. Use it as the information transmission strength of each directed edge.
5. The personalized hypertension risk assessment method based on big data according to claim 1, characterized in that, S4 specifically includes the following steps: S41. Perform topological structure encoding on the risk propagation graph of the target individual and the risk propagation graph of each historical individual to generate the corresponding node degree distribution sequence and adjacency matrix; S42. Based on the node degree distribution sequence and the adjacency matrix, calculate the graph kernel function value between the target individual and each historical individual. The graph kernel function value is obtained by weighted summation of the local subgraph structure similarity of each node. Select historical individuals whose graph kernel function value exceeds the similarity threshold as reference individuals. S43. For the risk propagation graph of the reference individual, extract the directed propagation path from the initial risk factor to the hypertension onset node, and calculate the step-by-step amplification coefficient. The step-by-step amplification coefficient is the product of the edge weight between adjacent nodes on the propagation path and the previous cumulative risk intensity. Calculate the cascade amplification law parameter based on the step-by-step amplification coefficient of all directed propagation paths of each reference individual. S44. Based on the cascade amplification law parameters, construct a multi-step propagation dynamic equation describing the cumulative enhancement mechanism of risk signals during multi-level propagation, and associate it with the risk propagation graph and edge weight matrix of the target individual to obtain a personalized risk propagation model.
6. The personalized hypertension risk assessment method based on big data according to claim 5, characterized in that, S44 specifically includes the following steps: S441. The cascade amplification law parameters are fitted with a function to obtain the amplification coefficient function between the step amplification coefficient and the propagation level, and a multi-step propagation dynamics equation is constructed. The multi-step propagation dynamics equation sets the current observation value of the risk factor as the initial risk intensity, and performs iterative calculation according to the propagation level order, accumulating the calculation to the propagation level where the hypertension onset node is located. S442. Divide the directed edges in the risk propagation graph of the target individual into multiple propagation levels and extract the edge weights of the directed edges. Calculate their average value as the level propagation efficiency and multiply it with the output value of the amplification coefficient function at the corresponding level to obtain the personalized amplification coefficient. Combine this with the multi-step propagation dynamics equation to obtain the personalized risk propagation model.
7. The personalized hypertension risk assessment method based on big data according to claim 1, characterized in that, S5 specifically includes the following steps: S51. Map the step characteristics of the target individual at the current moment to the corresponding risk factor node in the risk propagation graph to obtain the initial risk intensity value of each source node. S52. According to the topological order of the risk propagation diagram, starting from each source node, calculate the upstream risk signal received by each intermediate node in sequence, and combine the upstream risk signal with the corresponding step-by-step amplification coefficient to obtain the propagation risk intensity value of the intermediate node. S53. The post-propagation risk intensity values of all propagation paths along the risk propagation map that reach the hypertension onset node are nonlinearly fused. Based on the path length and the number of nodes on each propagation path, a differentiated weight is applied to the post-propagation risk intensity values to obtain the cumulative risk intensity of the hypertension onset node. S54. Map the cumulative risk intensity to a preset risk assessment quantification range to obtain a personalized hypertension risk assessment value.
8. A personalized hypertension risk assessment system based on big data, used to implement the method as described in any one of claims 1-7, characterized in that, include: The first unit is used to acquire physiological and behavioral data of the target individual and identify multiple risk factors. The second unit is used to construct continuous monitoring curves for each risk factor, detect step change points through cumulative sum control chart algorithms, extract the rate of change and duration of change of indicators before and after the step change point as step features, and construct a time series feature matrix. The third unit is used to identify the unidirectional temporal dependencies between risk factors based on the temporal data of each risk factor in the temporal feature matrix through Granger causality test, construct a directed acyclic risk propagation graph, and calculate the information transmission strength of each directed edge through transfer entropy to obtain the edge weight matrix. The fourth unit is used to perform graph isomorphic matching between the risk propagation map of the target individual and the risk propagation maps of each historical individual, filter out reference individuals through graph kernel functions, extract the cascade amplification law parameters along the propagation path, construct a multi-step propagation dynamics equation, and combine it with the edge weight matrix to obtain a personalized risk propagation model. The fifth unit is used to input the step characteristics of the target individual at the current moment into the personalized risk propagation model, simulate the risk signal to propagate in multiple stages along the directed edges of the risk propagation graph, calculate the cumulative risk intensity propagated to the hypertension onset node, and obtain the personalized hypertension risk assessment value.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.