Diabetic nephropathy risk identification method and system based on big data analysis

Through big data analysis and screening and interactive feature extraction, virtual sample sets are established and time series changes are identified, which solves the problems of insufficient screening of key indicators in diabetic nephropathy risk identification and uncaptured complex interactive relationships, and achieves more accurate risk assessment and dynamic prediction.

CN120015324APending Publication Date: 2025-05-16ZHU XIANYI MEMORIAL HOSPITAL OF TIANJIN MEDICAL UNIV (TIANJIN MEDICAL UNIV METABOLIC DISEASE HOSPITAL TIANJIN METABOLIC DISEASE PREVENTION CENT)
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510177797.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems in the identification of diabetic nephropathy risk, insufficient screening of key indicator characteristics, failure to fully capture complex interactions, insufficient sample imbalance and data distribution deviation, insufficient non-stationarity error identification in time series data analysis, weak dynamic prediction ability, and inaccurate quantification of data weights in the time-lapse risk assessment time.

Method used

Through big data analysis, screen the key indicator characteristics of risk patients, extract the interaction characteristics of risk indicators, establish a virtual sample set of risk patients, identify the key change rate and distribution shear points in the time series, divide the time series stable interval, extract segmented feature interval data, analyze the change trend and interactive changes, calculate dynamic predictive values, conduct risk contribution analysis and global risk quantification.

Benefits of technology

It significantly enhances the coverage and sample balance of data, improves the accuracy of time series data analysis, enhances the sensitivity and accuracy of dynamic prediction, and ensures the comprehensiveness and reliability of risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015324A_ABST
    Figure CN120015324A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical health big data analysis, in particular to a diabetic nephropathy risk identification method and system based on big data analysis, and the method comprises the following steps: screening key index features of risk patients through big data analysis based on health risk data of diabetic nephropathy patients; association of key indexes of the risk patients is analyzed, risk index interaction features are extracted, risk feature constraint conditions are determined, virtual sample parameters are identified through multi-dimensional data association, and a virtual sample set of the risk patients is established. According to the method, interaction characteristic values are extracted by analyzing key indexes of risk patients, high-risk index association is accurately captured, coverage and balance are enhanced based on multi-dimensional data association, non-stationary influence is solved by combining time sequence key change rate and shear point identification, and accurate matching is realized by analyzing offset rate and change trend. And the overall risk level is quantitatively evaluated by integrating interval data weighting, so that the risk identification comprehensiveness and the result reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical health big data analysis, and in particular to a method and system for identifying diabetic nephropathy risk based on big data analysis. Background Art

[0002] The field of medical and health big data analysis technology includes technologies for collecting, storing, analyzing and utilizing large-scale health-related data. The core content of this technology is to reveal potential health risk factors and predict the health development trends of individuals or groups by processing and analyzing large amounts of health data from various sources, thereby providing decision support for disease prevention and health management. The field of medical and health big data analysis technology includes data collection, data preprocessing, feature extraction, model building and risk assessment, etc., involving accurate analysis and scientific modeling of different forms of health data to support the identification of health risks and the implementation of interventions.

[0003] Among them, the diabetic nephropathy risk identification method refers to a method of conducting specific analysis on large-scale data related to diabetic nephropathy to discover the risk factors that cause the disease and evaluate the risk of disease for individuals or groups. This method mainly targets the health data of diabetic patients, and through analyzing key indicators such as blood sugar levels, renal function indicators and related time series data, it uses computerized modeling technology and feature extraction algorithms to conduct in-depth analysis of the data. This type of method is based on big data technology, and achieves the assessment of diabetic nephropathy risk through precise pattern recognition and feature matching of patients' health data.

[0004] Existing technologies are insufficient in feature screening of key indicators in health data analysis, making it difficult to fully capture the complex interactive relationships between high-risk indicators, resulting in missing details in the feature construction process. The processing of sample data is mostly focused on direct analysis of raw data, and fails to fully address the problems of sample imbalance and data distribution bias. This deficiency limits the coverage and representativeness of the data. In terms of time series data analysis, existing technologies have not fully addressed the errors caused by non-stationarity, resulting in inaccurate identification of distribution shear points, affecting the extraction of segmented data features. In the dynamic prediction link, existing technologies have weak capabilities to capture the changing trends and offset rates of time series parameters, resulting in insufficient sensitivity and accuracy in the prediction results. Existing technologies lack precise quantification of interval data weights in comprehensive risk assessment, making it difficult to accurately identify the patient's overall risk level, affecting the comprehensiveness and reliability of risk identification. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a method and system for identifying diabetic nephropathy risk based on big data analysis.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: a method for identifying diabetic nephropathy risk based on big data analysis, comprising the following steps:

[0007] S1: Based on the health risk data of patients with diabetic nephropathy, the key indicator characteristics of risk patients are screened through big data analysis, the correlation of key indicators of risk patients is analyzed, the interactive characteristics of risk indicators are extracted, the risk characteristic constraints are determined, the virtual sample parameters are identified through multidimensional data association, and a virtual sample set of risk patients is established;

[0008] S2: Based on the virtual sample set of risk patients, extract the time series data of blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen of the patients, identify the key change rate in the sequence by combining big data analysis, extract the fluctuation position and determine the distribution shear point, divide the time series into stable intervals, extract the data feature set in each interval, and obtain the segmented feature interval data;

[0009] S3: Based on the segmented feature interval data, analyze the change trend of each interval, analyze the interactive changes and offset rates of the parameters, calculate the interval dynamic prediction value according to the interval distribution law, expand the potential features through the step-by-step analysis of the risk indicators, and establish a dynamic feature risk prediction data set;

[0010] S4: Analyze the risk contribution of each interval through the dynamic characteristic risk prediction data set, weight the allocation ratio of the segmented characteristic interval weights, integrate the interval data and perform global risk quantification, evaluate the patient's overall risk based on the interval characteristics, and obtain the risk assessment results for diabetic nephropathy patients.

[0011] As a further solution of the present invention, the step of acquiring the risk indicator interaction feature is specifically as follows:

[0012] S111: Based on the health risk data of patients with diabetic nephropathy, perform structured analysis of the data, including classifying and summarizing the multi-dimensional health indicators of patients, grouping by type, calculating each statistical characteristic for the grouped indicators, and obtaining statistical data of the grouped indicators;

[0013] S112: Based on the statistical data of the grouping indicators, performing difference analysis on the key indicators of the screened risk patients, using an inductive strategy to evaluate the degree of correlation between the indicators, identifying and marking the correlation between the key risk indicators, and obtaining key indicator correlation information;

[0014] S113: Based on the key indicator association information, the interaction features between potential risk indicators are extracted, the extraction results are prioritized according to feature complementarity and risk gain value, and the interaction features with real-time significance are screened to obtain risk indicator interaction features.

[0015] As a further solution of the present invention, the steps of obtaining the risk patient virtual sample set are specifically as follows:

[0016] S121: extracting risk indicators associated with diabetic nephropathy from the multidimensional data according to the risk indicator interaction characteristics, analyzing the interaction characteristics between each indicator, quantifying the interaction level through the correlation coefficient and the dynamic change trend, setting constraints in combination with the dependency relationship between the characteristics and the weight distribution value, and obtaining preliminary risk characteristic constraint conditions;

[0017] S122: Calling the preliminary risk feature constraint condition to perform interactive calculation with the patient risk data, and fitting the dynamic change rate of the multi-dimensional interactive feature using the formula:

[0018]

[0019] Calculate and obtain multi-dimensional interaction constraints;

[0020] Among them, F s represents the value of multidimensional interaction constraints, w i represents the weight of the i-th dimension feature, T i represents the theoretical constraint value, R i represents the real-time calculated value, k is the adjustment factor used to optimize the feature fit, and n represents the total dimension of the interactive feature;

[0021] S123: Apply the multidimensional interactive constraint conditions to virtual sample parameter screening, determine whether the sample parameter combination meets the constraint conditions, use dimension-by-dimensional parameter recursive matching to perform parameter constraint combination, and establish a virtual sample set of risk patients by analyzing the multidimensional data association and dynamic trend fitting of the parameter combination.

[0022] As a further solution of the present invention, the step of obtaining the segmented feature interval data is specifically as follows:

[0023] S211: Based on the risk patient virtual sample set, extract the patient's blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen time series data, analyze the corresponding change rate at each time point in the time series, identify the key change rate in the time series by comparing the points where the change rate exceeds the key threshold, and generate a preliminary change rate set;

[0024] S212: Locate the position of the fluctuation in the time series according to the preliminary set of change rates, and identify the shear point in the data by combining the change rate with the data distribution difference, using the formula:

[0025]

[0026] Calculate the distribution weight value of the shear point and generate the distribution shear point of the fluctuation position;

[0027] Among them, C k Represents the distribution weight value of the shear point, V j represents the value of the fluctuation point, μ is the mean of the time series, δ j is the weight coefficient of the fluctuation point, and m is the total number of fluctuation points;

[0028] S213: Segment the time series according to the distribution shear points of the fluctuation position, identify the stable interval of the segment and extract the corresponding data feature set, analyze the segment features in combination with the mean, variance and change rate distribution of the data features in the interval, verify the rationality of the interval division by feature matching, and obtain the segment feature interval data.

[0029] As a further solution of the present invention, the step of obtaining the interval dynamic prediction value is specifically as follows:

[0030] S311: calling the segmented feature interval data, extracting the mean, variance and change rate of each interval, analyzing the time change trend of the parameters in the interval, identifying the parameter change curve through trend fitting, and generating a change trend analysis result;

[0031] S312: Analyze the interaction relationship and offset rate between parameters according to the change trend analysis result, using the formula:

[0032]

[0033] Calculate the offset weight value of the interval and generate the interactive change and offset rate distribution results;

[0034] Among them, P k represents the offset weight value of the interval, λ u is the weight coefficient of the uth parameter, M u is the real-time value of the parameter, A u is the theoretical reference value, q is the total number of parameters;

[0035] S313: According to the interactive change and the offset rate distribution results, a prediction model is constructed through the interval distribution law, the offset weight value and the trend curve fitting result are analyzed, and the prediction demand of each time period in the interval is identified to obtain the interval dynamic prediction value.

[0036] As a further solution of the present invention, the steps of acquiring the dynamic feature risk prediction data set are specifically:

[0037] S321: Based on the interval dynamic prediction value, extract the risk indicator data within the differentiated time interval, group and process according to the time series, analyze the data structure and match the start and end points, detect the indicator value change trend to obtain the data sequence relationship, mark and classify the sequence relationship, and obtain the time interval indicator data;

[0038] S322: Based on the time interval indicator data, extract interval indicator change feature values, analyze the change amplitudes of adjacent points in the time series, extract trend distribution data, perform comparative judgment and multi-interval correlation analysis on the trend distribution data, and obtain feature change parameter data;

[0039] S323: Based on the feature change parameter data, extract the core data of feature distribution relationship, rearrange them according to the frequency of occurrence of time series and the change pattern of related indicators, classify and integrate the arranged feature data as a whole, optimize the feature distribution characteristics and data integrity, and establish a dynamic feature risk prediction data set.

[0040] As a further solution of the present invention, the steps for obtaining the risk assessment results of diabetic nephropathy patients are specifically as follows:

[0041] S411: Based on the dynamic characteristic risk prediction data set, extract the risk value of each interval, analyze the interval characteristic value distribution and statistical data characteristics, quantify the interval characteristic value difference and adjust the range, and generate an adjusted interval characteristic value;

[0042] S412: Using the adjusted interval characteristic value and the corresponding weight ratio, extract the risk contribution and analyze the parameter correspondence, evaluate the proportional relationship between the interval characteristic value and the risk contribution, using the formula:

[0043] Z = (a·b+c·d) e ;

[0044] Calculate the risk contribution value of the interval;

[0045] Among them, Z represents the risk contribution value of the interval, a represents the adjusted interval characteristic value, b represents the weight ratio, c represents the risk assessment value, d represents the dynamic factor associated with the characteristic value, and e is the adjustment coefficient;

[0046] S413: weighted integration is performed on the risk contribution values ​​of the intervals, a global risk quantification index is identified, the risk contribution values ​​are called to adjust the interval characteristic parameters through the global weight, the overall risk value of the patient is evaluated, and a risk assessment result of the diabetic nephropathy patient is obtained.

[0047] A diabetic nephropathy risk identification system based on big data analysis, the diabetic nephropathy risk identification system based on big data analysis is used to execute the above-mentioned diabetic nephropathy risk identification method based on big data analysis, the system comprising:

[0048] The health indicator association analysis module is based on the health risk data of diabetic nephropathy patients, and interactively analyzes the blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen of the health risk data, identifies the correlation between indicators, analyzes the parameter interaction intensity based on the distribution trend, identifies the intensity indicator combination, and classifies each group of fluctuation ranges, extracts the interaction characteristics, and obtains the key interaction feature combination data;

[0049] The risk feature partitioning module divides the fluctuation range of the parameter combination into multiple partitions based on the key interactive feature combination data, calculates the parameter offset value and trend deviation within the partition, analyzes the offset value combination characteristics, collects the trend deviation distribution, integrates them into a virtual feature set, and establishes a virtual patient feature parameter set;

[0050] The feature dynamic analysis module divides the time series data into segments based on the virtual patient feature parameter set, analyzes the data fluctuation range and change rate in each segment, analyzes the dynamic characteristics of the parameters, extracts the change mode and offset structure, integrates the segment analysis and dynamic performance, and obtains the feature area offset analysis results;

[0051] The risk contribution quantification module extracts the dynamic characteristics and offset structure of the parameters in the region based on the characteristic region offset analysis results, evaluates the global characteristic offset range and contribution value distribution, integrates the contribution of regional parameters to the global dynamics, quantitatively analyzes the risk level, aggregates the distribution data, and generates risk assessment results for patients with diabetic nephropathy.

[0052] Compared with the prior art, the advantages and positive effects of the present invention are:

[0053] In the present invention, by analyzing the key indicators of risk patients, screening features and extracting interactive eigenvalues, it is possible to accurately capture the correlation changes between high-risk indicators, thereby providing an accurate basis for feature construction. Based on the correlation calculation of multidimensional data and the identification of virtual sample parameters, the data distribution range is expanded, and the data coverage and sample balance are significantly enhanced. In the process of extracting time series features, combined with the identification of key change rates and distribution shear points, the influence of the non-stationarity of time series data on the analysis accuracy is effectively solved. Through the detailed analysis of the characteristic intervals of segmented data, the identification and integration capabilities of segmented characteristic parameters are optimized, providing a more stable input for subsequent dynamic prediction. In dynamic prediction, by analyzing the offset rate and change trend of interval parameters, accurate matching of dynamic features is achieved, and the sensitivity and accuracy of risk contribution identification are enhanced. The weighted integration and global quantification of data from each interval are combined to effectively evaluate the overall risk level of the patient, significantly improving the comprehensiveness of risk identification and the reliability of results. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of the workflow of the present invention;

[0055] Figure 2 is a flow chart of the risk indicator interaction feature in the present invention;

[0056] Figure 3 A flowchart of a virtual sample set of risk patients in the present invention;

[0057] Figure 4 It is a flow chart of segmented feature interval data in the present invention;

[0058] Figure 5 It is a flow chart of interval dynamic prediction value in the present invention;

[0059] Figure 6 It is a flow chart of the dynamic feature risk prediction data set in the present invention;

[0060] Figure 7 This is a flow chart of the risk assessment results for patients with diabetic nephropathy in the present invention. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0062] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0063] Embodiment 1

[0064] See also Figure 1 The present invention provides a technical solution: a method for identifying diabetic nephropathy risk based on big data analysis, comprising the following steps:

[0065] S1: Based on the health risk data of patients with diabetic nephropathy, the key indicator characteristics of risk patients are screened through big data analysis, the correlation of key indicators of risk patients is analyzed, the interactive characteristics of risk indicators are extracted, the risk characteristic constraints are determined, the virtual sample parameters are identified through multidimensional data association, and a virtual sample set of risk patients is established;

[0066] S2: Based on the virtual sample set of risk patients, extract the time series data of blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen of patients, combine big data analysis to identify the key change rate in the sequence, extract the fluctuation position and determine the distribution shear point, divide the time series into stable intervals, extract the data feature set in each interval, and obtain the segmented feature interval data;

[0067] S3: Based on the segmented feature interval data, analyze the change trend of each interval, analyze the interactive changes and offset rates of the parameters, calculate the interval dynamic prediction value according to the interval distribution law, expand the potential features through the step-by-step analysis of the risk indicators, and establish a dynamic feature risk prediction data set;

[0068] S4: Analyze the risk contribution of each interval through the dynamic characteristic risk prediction data set, weight the distribution ratio of the segmented characteristic interval weights, integrate the interval data and perform global risk quantification, evaluate the patient's overall risk based on the interval characteristics, and obtain the risk assessment results for patients with diabetic nephropathy.

[0069] The interactive characteristics of risk indicators are specifically associated characteristic values, interactive weight values, and offset ratio values. The virtual sample set of risk patients includes virtual sample condition parameters, virtual sample characteristic ranges, and virtual sample distribution structures. The segmented characteristic interval data includes time series interval characteristic values, data fluctuation values ​​within the interval, and interval shear point parameters. The interval dynamic prediction values ​​are specifically dynamic change trend values, interval risk calculation values, and interval weight influence values. The dynamic characteristic risk prediction data set includes risk prediction characteristic values, interval risk weight values, and dynamic risk distribution values. The risk assessment results of diabetic nephropathy patients include overall risk level, key risk indicator values, and risk distribution characteristics.

[0070] See also Figure 2 , the steps for obtaining the interactive features of risk indicators are as follows:

[0071] S111: Based on the health risk data of patients with diabetic nephropathy, perform structured analysis of the data, including classifying and summarizing the multi-dimensional health indicators of patients, grouping by type, calculating each statistical characteristic for the grouped indicators, and obtaining statistical data of the grouped indicators;

[0072] First, the health indicators are grouped according to their measurement dimensions, including physiological data, pathological data, and behavioral data. After grouping, the basic statistical characteristics such as mean, median, and standard deviation are calculated for the indicators of each dimension, such as blood sugar level, renal function indicators (creatinine, urine protein), blood pressure, and weight change rate. Combined with the data distribution, quantiles are used to divide the health risk groups, such as low risk, medium risk, and high risk. The skewness coefficient and peak coefficient of the data are calculated for each group to clarify the central trend and discrete characteristics of health risks. By calculating the covariance matrix between different indicators, the intrinsic connection between the grouped indicators is evaluated, and the existing indicator redundancy is screened to ensure that the statistical results are more representative. Finally, each group of data is structured into a multi-level form, including the group name, indicator statistical characteristics, and the correlation between statistical characteristics, to form group indicator statistical data.

[0073] S112: Based on the statistical data of grouped indicators, differential analysis is performed on the key indicators of the screened risk patients, an inductive strategy is used to evaluate the degree of correlation between the indicators, the correlation between the key risk indicators is identified and marked, and the correlation information between the key indicators is obtained;

[0074] Combined with group statistical data, indicators closely related to the risk of diabetic nephropathy, such as fasting blood glucose level, creatinine concentration, and urine protein content, were selected for inter-group difference analysis. The difference between the mean and variance of patients with different risk groups was calculated to determine the contribution of each indicator to health risks. The cross-comparison method was used to evaluate the correlation between indicators. By comparing the overlapping areas of indicator distribution, the indicator pairs with linkage relationships were marked. For example, whether there was a significant correlation between fasting blood glucose and urine protein levels was evaluated. The linear or nonlinear relationship between indicators was clarified by constructing a correlation matrix, and the most influential correlation paths were extracted. After the evaluation was completed, the paths were further classified and labeled, and the primary and secondary correlations were distinguished according to the correlation strength to form key indicator correlation information.

[0075] S113: based on the key indicator correlation information, extract the interaction features between potential risk indicators, prioritize the extraction results according to feature complementarity and risk gain value, screen the interaction features with real-time significance, and obtain the risk indicator interaction features;

[0076] Extract the interaction features between key indicators, determine the importance of the features by evaluating the gain value of each group of indicators on the patient's health risk when interacting, take fasting blood glucose and urine protein content as an example, calculate the risk gain value under the interaction of the two by constructing a bivariate distribution diagram, and compare and analyze it with the risk contribution value of a single indicator to identify whether there is a risk superposition effect, further combine with complementary analysis, eliminate the interaction features with low risk coverage or information redundancy, sort the features with real-time significance by priority, introduce the time dimension to analyze the dynamic changes of the interaction features, ensure that the screened interaction features can reflect the changes in risk status in a timely manner, integrate the high-priority risk indicator interaction features into risk indicator interaction features, and provide support for subsequent risk assessment and intervention strategies.

[0077] See also Figure 3 , the specific steps for obtaining the virtual sample set of risk patients are:

[0078] S121: extract risk indicators associated with diabetic nephropathy from multidimensional data according to the interaction characteristics of risk indicators, analyze the interaction characteristics between each indicator, quantify the interaction level through the correlation coefficient and the dynamic change trend, set constraints based on the dependency relationship between the characteristics and the weight distribution value, and obtain preliminary risk characteristic constraint conditions;

[0079] For the risk indicators, the mean, variance, extreme value and distribution trend of each indicator are calculated through statistical feature extraction methods, the results are matched with the correlation matrix of the target disease, and the features with higher significance are screened as preliminary data features. The interaction characteristics among the indicators are further analyzed, and the evaluation function of the feature dependency relationship is established based on the dynamic volatility and distribution overlap rate of the interaction characteristics. The dependency relationship is used to calculate the significance scores under different indicator combinations and generate weight allocation values. The risk feature constraints are set in combination with the feature significance and weight results to generate preliminary risk feature constraints.

[0080] S122: Call the preliminary risk feature constraint conditions to interact with the patient risk data, and fit the dynamic change rate of the multi-dimensional interactive features using the formula:

[0081]

[0082] Calculate and obtain multi-dimensional interaction constraints;

[0083] Among them, F s represents the value of multidimensional interaction constraints, w i Represents the weight of the i-th dimension feature, T i represents the theoretical constraint value, R i represents the real-time calculated value, k is the adjustment factor used to optimize the feature fit, and n represents the total dimension of the interactive feature;

[0084] The benefit of the formula is that through the weight parameter w i Differentiated impacts are given to risk characteristics of different dimensions, and the adjustment factor k is introduced to improve the feature fit, which can more accurately measure the stability of multi-dimensional interactive characteristics;

[0085] Suppose there are four risk characteristic dimensions n=4, where the weight parameter w of each dimension is i The specific values ​​are obtained by calculating the significance scores and normalizing them. The specific values ​​are w1=0.3, w2=0.25, w3=0.2, w4=0.25, and the theoretical constraint value T i Compared with the actual calculated value R i Based on the average value and current fluctuation value of each dimension feature in the monitoring data, let T =

[0086] {0.8, 0.7, 0.6, 0.9}, R = {0.75, 0.72, 0.62, 0.87}, adjustment factor k = 2;

[0087] Substituting into the formula:

[0088]

[0089] The result shows that the stability value of the interaction feature is 0.034, indicating that the stability between the current risk features is relatively high, and the multidimensional interaction constraints can be further used for virtual sample construction.

[0090] S123: applying multidimensional interactive constraint conditions to virtual sample parameter screening, determining whether the sample parameter combination satisfies the constraint conditions, performing parameter constraint combination by dimension-by-dimension parameter recursive matching, and establishing a risk patient virtual sample set by analyzing multidimensional data association and dynamic trend fitting of the parameter combination;

[0091] Firstly, the patient data related to the risk of diabetic nephropathy were extracted, and the risk characteristic values ​​and dynamic change rates of the patient data were quantified according to the data type. The hierarchical screening method was used to gradually eliminate the parameter combinations with large deviations from the interactive constraints. The similarity between the feature combinations of each dimension and the target constraints was calculated through the feature matching function. The recursive method was used to accurately match the parameters dimension by dimension. The rationality of the sample was verified through the dynamic trend fitting analysis between the multidimensional feature data, and a virtual sample set of risk patients was established.

[0092] See also Figure 4 , the specific steps for obtaining segmented feature interval data are:

[0093] S211: Based on the virtual sample set of risk patients, extract the time series data of blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen of the patients, analyze the corresponding change rate at each time point in the time series, identify the key change rate in the time series by comparing the points where the change rate exceeds the key threshold, and generate a preliminary set of change rates;

[0094] A complete time series data set is constructed through the detection values ​​recorded hourly. For each data type, the rate of change between adjacent time points is calculated. The rate of change is defined as the difference between the values ​​of two consecutive time points divided by the absolute value of the value of the previous time point. The rate of change at each time point is compared with a reasonable significance threshold, and the time points exceeding the threshold are marked. The data fluctuation amplitude within a certain range before and after each marked point is further analyzed to calculate the slope of the continuous change rate, mark the time interval of the significant change rate points and detect the points. Combined with the patient's historical trend data, some missed time points are included in the change rate set. By checking the distribution of the change rate set, it is ensured that it covers all key change positions. At the same time, meaningless fluctuation points or noise points are deleted, and finally a preliminary change rate set is generated to ensure that the basic data for the calculation of subsequent shear points is comprehensive and reliable.

[0095] S212: Locate the position of fluctuations in the time series based on the preliminary set of change rates, and identify the shear points in the data by combining the change rate with the data distribution difference, using the formula:

[0096]

[0097] Calculate the distribution weight value of the shear point and generate the distribution shear point of the fluctuation position;

[0098] Among them, C k Represents the distribution weight value of the shear point, V j represents the value of the fluctuation point, μ is the mean of the time series, δ j is the weight coefficient of the fluctuation point, and m is the total number of fluctuation points;

[0099] The benefit of the formula is that by introducing the weight coefficient δ of the fluctuation point j , reflects the actual influence of each fluctuation point on the distribution weight, and further optimizes the precise calculation of the shear point by combining the comparison of the mean μ, so that the formula can improve the accuracy and stability of shear point identification in complex distributions;

[0100] The set of fluctuation point values ​​in the time series is set to V = {1.2, 0.9, 1.1, 1.3, 0.8}, where the set of fluctuation point weight coefficients is δ = {1.5, 1.2, 1.8, 1.4, 1.3}, and the mean μ of the calculated time series is 1.06;

[0101] Substitute into the formula to calculate:

[0102]

[0103] The result shows that the distribution weight value of the shear point in the time series is 0.0124, which means that in the existing set of fluctuation points, the distribution influence of the main shear points is relatively low and needs to be combined with further distribution analysis and processing.

[0104] S213: segmenting the time series according to the distribution shear points of the fluctuation position, identifying the stable interval of the segment and extracting the corresponding data feature set, analyzing the segment features in combination with the mean, variance and change rate distribution of the data features in the interval, verifying the rationality of the interval division by feature matching, and obtaining the segment feature interval data;

[0105] Retrieve the time point of each fluctuation point and the data set within a certain range before and after it one by one, divide it into multiple time periods, calculate the ratio of the data standard deviation to the mean of each time period, judge the stability of the time period within a reasonable range with the ratio, extract the time period with higher stability, further analyze the distribution of data features in the segmented interval, including features such as mean, rate of change and variance, match them by combining the time distribution pattern of the target feature, analyze the degree of fit between each feature point in the time period and the dynamic trend of the target feature, generate a set of feature values ​​for each segmented interval based on the degree of fit, and finally obtain the segmented feature interval data to ensure that the characteristics of each interval can fully cover the dynamic information related to the risk, and provide detailed and reliable data support for subsequent prediction and analysis.

[0106] See also Figure 5 , the specific steps for obtaining the interval dynamic prediction value are:

[0107] S311: calling the segmented feature interval data, extracting the mean, variance and change rate of each interval, analyzing the time change trend of the parameters in the interval, identifying the parameter change curve through trend fitting, and generating the change trend analysis result;

[0108] The extraction results are mapped to the time series for further analysis. For the mean of each interval, the sum of all data points in the interval is calculated and then divided by the number of data points. At the same time, the variance value in the interval is calculated by combining the square of the deviation between the data point and the mean to measure the fluctuation range and stability of the data distribution in the interval. For the change rate of each interval, the values ​​of continuous time points are selected, and the change rate is calculated as the difference between two adjacent time points divided by the absolute value of the previous time point. The maximum value of all change rates in the interval is selected as the maximum change rate of the interval. After sorting out the above indicators of all intervals, the temporal change characteristics of each interval parameter are analyzed through trend fitting. The trend curve fitting results of the data points are used to calculate the slope of the trend and mark the turning points. For the marking of the turning points, they are verified by combining the difference in the dynamic change rates of the previous and next intervals. Only the turning points that meet the threshold standards are retained to obtain the change trend analysis results to ensure the accuracy and completeness of the interval parameter change trend analysis.

[0109] S312: According to the change trend analysis results, the interaction relationship and offset rate between the parameters are analyzed, and the formula is used:

[0110]

[0111] Calculate the offset weight value of the interval and generate the interactive change and offset rate distribution results;

[0112] Among them, P k represents the offset weight value of the interval, λ u is the weight coefficient of the uth parameter, M u is the real-time value of the parameter, A u is the theoretical reference value, q is the total number of parameters;

[0113] The benefit of the formula is that by introducing the parameter weight coefficient λ u , so that the influence of each parameter can be adjusted dynamically, combined with the actual value M u Compared with the theoretical value A u The difference analysis can realize the comprehensive evaluation of the interactive changes of multi-dimensional parameters within the interval, and significantly improve the calculation accuracy of the offset weight in complex systems;

[0114] Assume that an interval contains three parameters, the actual measurement value set is M = {1.2, 0.9, 1.0}, the theoretical reference value set is A = {1.0, 1.0, 1.0}, and the weight parameter set is λ = {1.3, 1.1, 1.2}. First, calculate the offset value of each parameter;

[0115] Calculate item by item through the formula:

[0116]

[0117] The result shows that the offset weight value of the current interval is 0.35, which means that the offset degree of each parameter in the overall trend is moderate. This value can be used for subsequent dynamic prediction value calculation and adjustment.

[0118] S313: According to the interactive change and the offset rate distribution results, a prediction model is constructed through the interval distribution law, the offset weight value and the trend curve fitting result are analyzed, and the prediction demand of each time period in the interval is identified to obtain the interval dynamic prediction value;

[0119] First, the offset weight value and trend analysis result of the interval are extracted, and the dynamic fitting curve is obtained by combining the two item by item. The dynamic fitting curve depends on the time distribution law of each interval. By analyzing the relationship between the trend line slope, inflection point and change rate in the time period, the fitting parameters of the time distribution law are comprehensively obtained. At the same time, the time weight and the proportion of historical reference data in the prediction model are adjusted to make the dynamic prediction value more in line with the actual distribution characteristics of the interval. In each interval dynamic prediction value, the prediction curve fitting result is compared with the actual value of the current interval, and the weight parameters of the prediction model are updated to make it more in line with the prediction needs of subsequent time periods. Finally, the dynamic prediction values ​​of each interval are integrated to generate interval dynamic prediction values ​​that can reflect the overall time series trend, providing reliable support for further analysis.

[0120] See also Figure 6 ,The specific steps for obtaining the dynamic feature risk prediction dataset are:

[0121] S321: Based on the interval dynamic prediction value, extract the risk indicator data within the differentiated time interval, group and process according to the time series, analyze the data structure and match the start and end points, detect the indicator value change trend to obtain the data sequence relationship, mark and classify the sequence relationship, and obtain the time interval indicator data;

[0122] For dynamic prediction values, indicator data are extracted from various time intervals such as daily, weekly, and monthly, and grouped and processed by time series. The risk indicator data of each time interval is rearranged into a time series matrix, and the starting and ending data of each time series are analyzed one by one to determine the integrity of the data record. For each group of time series, the change rate and fluctuation range between each point are calculated, and the intervals with abnormal changes are screened out. The growth or decline trend of the indicator is analyzed with trend lines. In the process of data analysis, the sequence relationship is marked as three types of continuous rise, continuous decline, or fluctuation interval according to time progression, and the sequence is classified into slow, medium, or fast fluctuation types according to the change rate. Through sequence relationship marking and classification, the time interval indicator data is obtained, and the dynamic characteristics of the indicators in each interval are clarified, laying the foundation for further analysis of the time interval risk pattern.

[0123] S322: Based on the time interval indicator data, extract the interval indicator change characteristic value, analyze the change range of adjacent points in the time series, extract the trend distribution data, compare and judge the trend distribution data and perform multi-interval correlation analysis to obtain the characteristic change parameter data;

[0124] Extract change characteristic values ​​from time interval indicator data, analyze the change amplitudes of adjacent data points in each time series one by one, and use this to calculate the average change rate and maximum fluctuation amplitude within the interval, analyze the trend distribution data of the time series, and use the sliding window method to compare the similarity of trends in each interval segment by segment, detect the continuity of indicator changes in continuous intervals, and classify trend data in different intervals based on the direction and amplitude of change, extract significant trend features, such as monotonically increasing intervals, multi-peak fluctuation intervals, etc., and calculate the duration and cumulative impact value of interval features. Combined with the correlation analysis of multiple intervals, identify the linkage effect and time series characteristics of indicator changes in different time periods, and finally extract feature change parameter data that can characterize key change features to evaluate the importance of interval features in risk identification.

[0125] S323: Based on the feature change parameter data, extract the core data of feature distribution relationship, rearrange them according to the time series occurrence frequency and the change pattern of related indicators, classify and integrate the arranged feature data as a whole, optimize the feature distribution characteristics and data integrity, and establish a dynamic feature risk prediction data set;

[0126] The distribution relationship of each feature data is analyzed, and high-frequency features are screened through frequency statistics. The distribution density in the time series and the change correlation between adjacent features are calculated. According to the indicator change pattern, the feature data is re-arranged to ensure that high-frequency features and low-frequency features with strong correlation can be arranged adjacent to each other to form a complete feature sequence. The arranged feature data are classified and integrated, and multiple groups are established according to trend features divided by different time intervals. The distribution characteristics of each group of features are optimized, duplicate or incomplete data records are eliminated, and missing values ​​between data are filled to improve the integrity and consistency of the overall data. A dynamic feature risk prediction data set is established to provide accurate and comprehensive data support for subsequent model construction and risk assessment.

[0127] See also Figure 7 , the specific steps for obtaining risk assessment results for patients with diabetic nephropathy are as follows:

[0128] S411: Based on the dynamic characteristic risk prediction data set, extract the risk value of each interval, analyze the interval characteristic value distribution and statistical data characteristics, quantify the interval characteristic value difference and adjust the range, and generate the adjusted interval characteristic value;

[0129] First, it is necessary to check the integrity of the characteristic data of each interval. For missing data, fill it in through the interpolation method of adjacent values ​​of the interval to ensure the validity of all interval data. Then, remove the outliers that exceed the reasonable range. The range is determined by the upper and lower quartiles of the data distribution and the range calculated by 1.5 times the interquartile range. Next, call the eigenvalues ​​of each interval in the data set, calculate the minimum, maximum and mean values ​​respectively, and use them as the preliminary representation of the interval characteristics. Based on the eigenvalues ​​of the preliminary representation, perform normalization adjustment on the eigenvalue range, compress each interval eigenvalue to the interval between 0 and 1 through the normalization formula, and call the normalized eigenvalue and its corresponding weight parameter at the same time. By comparing the interval eigenvalue weight relationship table, match the corresponding relationship between the two item by item to generate the adjusted interval eigenvalue. Finally, store the eigenvalue in a way that is associated with the original weight ratio to provide basic data support for subsequent analysis.

[0130] S412: Using the adjusted interval eigenvalues ​​and corresponding weight ratios, extract the risk contribution and analyze the corresponding relationship between the parameters, evaluate the proportional relationship between the interval eigenvalues ​​and the risk contribution, and use the formula:

[0131] Z = (a·b+c·d) e ;

[0132] Calculate the risk contribution value of the interval;

[0133] Among them, Z represents the risk contribution value of the interval, a represents the adjusted interval characteristic value, b represents the weight ratio, c represents the risk assessment value, d represents the dynamic factor associated with the characteristic value, and e is the adjustment coefficient;

[0134] The benefit of the formula is that it effectively handles the imbalance of risk values ​​in different intervals by introducing dynamic factors and adjustment coefficients, making the formula flexible to adapt to the risk distribution characteristics of different intervals, and dynamically optimizes the risk contribution calculation by adjusting the factor parameters, thus improving the applicability of the calculation results;

[0135] a=0.85 is the normalized interval eigenvalue, calculated using the normalization formula;

[0136] b = 0.7 is the weight ratio, which is obtained based on the matching relationship between the eigenvalue and the weight ratio;

[0137] c = 0.9 is the risk assessment value, which is obtained by calculating the weighted mean of the interval data;

[0138] d = 1.2 is the dynamic factor, which is calculated based on the ratio of the change amplitude of the interval eigenvalue to the standard deviation;

[0139] e=2 is the adjustment coefficient, which is obtained by adjusting the dynamic adaptability of the formula through experiments;

[0140] Calculation process:

[0141] Substitute into the formula Z = (0.85 0.7 + 0.9 1.2) 2 ;

[0142] Calculated Z = (0.595 + 1.08) 2 =1.675 2 =2.805625;

[0143] The results show that the risk contribution value of each interval can more accurately reflect the actual risk distribution of different intervals through the calculation of dynamic weights and adjustment factors, providing an accurate basis for further risk integration.

[0144] S413: weighted integration of the risk contribution values ​​of the intervals, identification of global risk quantification indicators, calling the risk contribution values ​​to adjust the interval characteristic parameters through the global weight, assessing the overall risk value of the patient, and obtaining the risk assessment results of the diabetic nephropathy patient;

[0145] First, the risk contribution values ​​of all intervals are called from the data set and arranged in order of size. The global weight proportion of each risk interval is divided based on the interval risk contribution value. The weights of interval parameters are redistributed by adjusting the global sorting of risk contribution values. The intervals with large risk value fluctuations are set as high-weight segments. At the same time, the proportion of the sorted interval parameters in the overall risk contribution is calculated respectively, and the proportion is adjusted to a standardized form through a global normalization formula to eliminate the impact of different interval characteristic values ​​on the total risk assessment. Then, the standardized risk values ​​of each interval are redistributed according to the global proportion. By comparing the changes in risk contribution values ​​and interval proportions, the interval weight fluctuation range is further corrected. Combined with the adjusted interval characteristic weights, the global risk assessment quantitative indicators are re-integrated and calculated. Finally, the indicators are integrated into the patient's overall risk assessment formula to generate the overall risk assessment results for patients with diabetic nephropathy.

[0146] The diabetic nephropathy risk identification system based on big data analysis is used to execute the above-mentioned diabetic nephropathy risk identification method based on big data analysis, and the system includes:

[0147] The health indicator association analysis module is based on the health risk data of diabetic nephropathy patients, and interactively analyzes the blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen of the health risk data, identifies the correlation between indicators, analyzes the parameter interaction intensity based on the distribution trend, identifies the intensity indicator combination, and classifies each group of fluctuation ranges, extracts the interaction characteristics, and obtains the key interaction feature combination data;

[0148] The risk feature partitioning module divides the fluctuation range of the parameter combination into multiple partitions based on the key interactive feature combination data, calculates the parameter offset value and trend deviation within the partition, analyzes the offset value combination characteristics, collects the trend deviation distribution, integrates them into a virtual feature set, and establishes a virtual patient feature parameter set;

[0149] The feature dynamic analysis module divides the time series data into segments based on the virtual patient feature parameter set, analyzes the data fluctuation range and change rate in each segment, analyzes the dynamic characteristics of the parameters, extracts the change pattern and offset structure, integrates segment analysis and dynamic performance, and obtains the feature area offset analysis results;

[0150] The risk contribution quantification module is based on the characteristic region offset analysis results, extracts the dynamic characteristics and offset structure of the parameters in the region, evaluates the global characteristic offset range and contribution value distribution, integrates the contribution of regional parameters to the global dynamics, quantitatively analyzes the risk level, aggregates the distribution data, and generates risk assessment results for patients with diabetic nephropathy.

[0151] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A method for identifying diabetic nephropathy risk based on big data analysis, characterized in that: The following steps are involved: S1: Based on the health risk data of patients with diabetic nephropathy, the key indicator characteristics of risk patients are screened through big data analysis, the correlation of key indicators of risk patients is analyzed, the interactive characteristics of risk indicators are extracted, the risk characteristic constraints are determined, the virtual sample parameters are identified through multidimensional data association, and a virtual sample set of risk patients is established; S2: Based on the virtual sample set of risk patients, extract the time series data of blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen of the patients, identify the key change rate in the sequence by combining big data analysis, extract the fluctuation position and determine the distribution shear point, divide the time series into stable intervals, extract the data feature set in each interval, and obtain the segmented feature interval data; S3: Based on the segmented feature interval data, analyze the change trend of each interval, analyze the interactive changes and offset rates of the parameters, calculate the interval dynamic prediction value according to the interval distribution law, expand the potential features through the step-by-step analysis of the risk indicators, and establish a dynamic feature risk prediction data set; S4: Analyze the risk contribution of each interval through the dynamic characteristic risk prediction data set, weight the allocation ratio of the segmented characteristic interval weights, integrate the interval data and perform global risk quantification, evaluate the patient's overall risk based on the interval characteristics, and obtain the risk assessment results for diabetic nephropathy patients.

2. The method for identifying diabetic nephropathy risk based on big data analysis according to claim 1, characterized in that: The steps for obtaining the risk indicator interaction characteristics are specifically as follows: S111: Based on the health risk data of patients with diabetic nephropathy, perform structured analysis of the data, including classifying and summarizing the multi-dimensional health indicators of patients, grouping by type, calculating each statistical characteristic for the grouped indicators, and obtaining statistical data of the grouped indicators; S112: Based on the statistical data of the grouping indicators, performing difference analysis on the key indicators of the screened risk patients, using an inductive strategy to evaluate the degree of correlation between the indicators, identifying and marking the correlation between the key risk indicators, and obtaining key indicator correlation information; S113: Based on the key indicator association information, the interaction features between potential risk indicators are extracted, the extraction results are prioritized according to feature complementarity and risk gain value, and the interaction features with real-time significance are screened to obtain risk indicator interaction features.

3. The method for identifying diabetic nephropathy risk based on big data analysis according to claim 2, characterized in that: The steps for obtaining the risk patient virtual sample set are specifically as follows: S121: extracting risk indicators associated with diabetic nephropathy from the multidimensional data according to the risk indicator interaction characteristics, analyzing the interaction characteristics between each indicator, quantifying the interaction level through the correlation coefficient and the dynamic change trend, setting constraints in combination with the dependency relationship between the characteristics and the weight distribution value, and obtaining preliminary risk characteristic constraint conditions; S122: Calling the preliminary risk feature constraint condition to perform interactive calculation with the patient risk data, and fitting the dynamic change rate of the multi-dimensional interactive feature using the formula: Calculate and obtain multi-dimensional interaction constraints; Among them, F s represents the value of multidimensional interaction constraints, w i represents the weight of the i-th dimension feature, T i represents the theoretical constraint value, R i represents the real-time calculated value, k is the adjustment factor used to optimize the feature fit, and n represents the total dimension of the interactive feature; S123: Apply the multidimensional interactive constraint conditions to virtual sample parameter screening, determine whether the sample parameter combination meets the constraint conditions, use dimension-by-dimensional parameter recursive matching to perform parameter constraint combination, and establish a virtual sample set of risk patients by analyzing the multidimensional data association and dynamic trend fitting of the parameter combination.

4. The method for identifying diabetic nephropathy risk based on big data analysis according to claim 3, characterized in that: The steps for obtaining the segmented feature interval data are specifically as follows: S211: Based on the risk patient virtual sample set, extract the patient's blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen time series data, analyze the corresponding change rate at each time point in the time series, identify the key change rate in the time series by comparing the points where the change rate exceeds the key threshold, and generate a preliminary change rate set; S212: Locate the position of the fluctuation in the time series according to the preliminary set of change rates, and identify the shear point in the data by combining the change rate with the data distribution difference, using the formula: Calculate the distribution weight value of the shear point and generate the distribution shear point of the fluctuation position; Among them, C k Represents the distribution weight value of the shear point, V j represents the value of the fluctuation point, μ is the mean of the time series, δ j is the weight coefficient of the fluctuation point, and m is the total number of fluctuation points; S213: Segment the time series according to the distribution shear points of the fluctuation position, identify the stable interval of the segment and extract the corresponding data feature set, analyze the segment features in combination with the mean, variance and change rate distribution of the data features in the interval, verify the rationality of the interval division by feature matching, and obtain the segment feature interval data.

5. The method for identifying diabetic nephropathy risk based on big data analysis according to claim 4, characterized in that: The steps for obtaining the interval dynamic prediction value are specifically as follows: S311: calling the segmented feature interval data, extracting the mean, variance and change rate of each interval, analyzing the time change trend of the parameters in the interval, identifying the parameter change curve through trend fitting, and generating a change trend analysis result; S312: Analyze the interaction relationship and offset rate between parameters according to the change trend analysis result, using the formula: Calculate the offset weight value of the interval and generate the interactive change and offset rate distribution results; Among them, P k represents the offset weight value of the interval, λ u is the weight coefficient of the uth parameter, M u is the real-time value of the parameter, A u is the theoretical reference value, q is the total number of parameters; S313: According to the interactive change and the offset rate distribution results, a prediction model is constructed through the interval distribution law, the offset weight value and the trend curve fitting result are analyzed, and the prediction demand of each time period in the interval is identified to obtain the interval dynamic prediction value.

6. The method for identifying diabetic nephropathy risk based on big data analysis according to claim 5, characterized in that: The steps for obtaining the dynamic feature risk prediction data set are specifically as follows: S321: Based on the interval dynamic prediction value, extract the risk indicator data within the differentiated time interval, group and process according to the time series, analyze the data structure and match the start and end points, detect the indicator value change trend to obtain the data sequence relationship, mark and classify the sequence relationship, and obtain the time interval indicator data; S322: Based on the time interval indicator data, extract interval indicator change feature values, analyze the change amplitudes of adjacent points in the time series, extract trend distribution data, perform comparative judgment and multi-interval correlation analysis on the trend distribution data, and obtain feature change parameter data; S323: Based on the feature change parameter data, extract the core data of feature distribution relationship, rearrange them according to the frequency of occurrence of time series and the change pattern of related indicators, classify and integrate the arranged feature data as a whole, optimize the feature distribution characteristics and data integrity, and establish a dynamic feature risk prediction data set.

7. The method for identifying diabetic nephropathy risk based on big data analysis according to claim 6, characterized in that: The steps for obtaining the risk assessment results of diabetic nephropathy patients are specifically as follows: S411: Based on the dynamic characteristic risk prediction data set, extract the risk value of each interval, analyze the interval characteristic value distribution and statistical data characteristics, quantify the interval characteristic value difference and adjust the range, and generate an adjusted interval characteristic value; S412: Using the adjusted interval characteristic value and the corresponding weight ratio, extract the risk contribution and analyze the parameter correspondence, evaluate the proportional relationship between the interval characteristic value and the risk contribution, using the formula: Z=(a·b+c·d) e ; Calculate the risk contribution value of the interval; Among them, Z represents the risk contribution value of the interval, a represents the adjusted interval characteristic value, b represents the weight ratio, c represents the risk assessment value, d represents the dynamic factor associated with the characteristic value, and e is the adjustment coefficient; S413: weighted integration is performed on the risk contribution values ​​of the intervals, a global risk quantification index is identified, the risk contribution values ​​are called to adjust the interval characteristic parameters through the global weight, the overall risk value of the patient is evaluated, and a risk assessment result of the diabetic nephropathy patient is obtained.

8. A diabetic nephropathy risk identification system based on big data analysis, characterized in that: According to any one of claims 1 to 7, the method for identifying diabetic nephropathy risk based on big data analysis comprises: The health indicator association analysis module is based on the health risk data of diabetic nephropathy patients, and interactively analyzes the blood sugar, blood pressure, urine microalbumin, and blood urea nitrogen of the health risk data, identifies the correlation between indicators, analyzes the parameter interaction intensity based on the distribution trend, identifies the intensity indicator combination, and classifies each group of fluctuation ranges, extracts the interaction characteristics, and obtains the key interaction feature combination data; The risk feature partitioning module divides the fluctuation range of the parameter combination into multiple partitions based on the key interactive feature combination data, calculates the parameter offset value and trend deviation within the partition, analyzes the offset value combination characteristics, collects the trend deviation distribution, integrates them into a virtual feature set, and establishes a virtual patient feature parameter set; The feature dynamic analysis module divides the time series data into segments based on the virtual patient feature parameter set, analyzes the data fluctuation range and change rate in each segment, analyzes the dynamic characteristics of the parameters, extracts the change mode and offset structure, integrates the segment analysis and dynamic performance, and obtains the feature area offset analysis results; The risk contribution quantification module extracts the dynamic characteristics and offset structure of the parameters in the region based on the characteristic region offset analysis results, evaluates the global characteristic offset range and contribution value distribution, integrates the contribution of regional parameters to the global dynamics, quantitatively analyzes the risk level, aggregates the distribution data, and generates risk assessment results for patients with diabetic nephropathy.

Citation Information

Cited By

  • Data-driven pathogenic microorganism target analysis method and system

    CN120256901A

  • Clinical test data intelligent analysis method and system

    CN120299591A

  • Multi-modal data fusion diabetes risk prediction system after acute pancreatitis

    CN121171617A

  • Diabetic nephropathy risk assessment method and model

    CN121354934A

  • A model for assessing risk of diabetic nephropathy

    CN121354934B