A method and system for generating risk assessment data for chronic kidney disease

CN122575671APending Publication Date: 2026-08-14SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]本发明的目的是针对现有技术所存在的慢性肾脏病风险评估手段存在的检测成本高、操作流程复杂等缺陷现,提供一种基于易获取的体检数据和共病病史数据的慢性肾脏病风险评估数据生成方法及系统,以提升慢性肾脏病风险评估普及性和有效性

Benefits of technology

1、本发明临床操作简单,只需要简单运用根据大样本训练出的评分模型,即可通过易于获取的体检数据快速评估受检者患慢性肾脏病的风险;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575671A_ABST
    Figure CN122575671A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for generating risk assessment data for chronic kidney disease. The method includes: a data acquisition step, acquiring physical examination data and comorbidity history characteristic data; the physical examination data includes at least one first data point, and the comorbidity history data includes at least one second data point; a data screening step, using a pre-established random forest model to calculate the contribution of each first data point and each second data point, obtaining the corresponding contribution, and screening based on the contribution to determine at least one important data point; a data determination step, constructing a logistic regression model based on the important data point, outputting the probability value and advantage ratio of each important data point, and determining at least one significant data point based on the probability value and advantage ratio; a data scoring step, performing data transformation processing based on the advantage ratio of the significant data point to obtain a score value for the significant data point; and a risk assessment step, establishing a chronic kidney disease risk assessment model based on the score values ​​of the significant data point, and obtaining risk assessment data from the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing technology, and in particular to a method and system for generating risk assessment data for chronic kidney disease. Background Technology

[0002] With socio-economic development and improved living standards, unhealthy lifestyles such as staying up late and overeating have become increasingly common. Coupled with the high incidence of underlying diseases like hypertension, diabetes, and gout, the incidence of chronic kidney disease (CKD) is showing a significant upward trend. Clinical studies have shown that most underlying diseases eventually affect the structure and function of the kidneys, leading to specific types of CKD such as hypertensive nephropathy, diabetic nephropathy, and gouty nephropathy. At the same time, with advancements in diagnostic technology and the implementation of various general screening programs and targeted screening of high-risk groups, the detection range of CKD is continuously expanding, further confirming its high prevalence.

[0003] However, current public awareness, intervention, and control rates of chronic kidney disease (CKD) are all low, resulting in most patients being diagnosed only when they develop obvious clinical symptoms, have suffered substantial kidney damage, or have even progressed to end-stage renal disease. This situation causes patients to miss the optimal treatment window, seriously threatening their lives and health, and significantly increasing the medical and economic burden on individuals and society. Therefore, early screening and risk intervention for CKD have become critical issues that urgently need to be addressed in the medical field.

[0004] Clinical practice has confirmed that early detection of chronic kidney disease (CKD) and increasing the detection rate of undiagnosed CKD are core prerequisites for improving the effectiveness of risk intervention and enhancing prevention and treatment outcomes. Currently, the diagnosis of CKD largely relies on methods such as postprandial blood glucose testing and glycated hemoglobin (HbA1c) testing. These methods have drawbacks such as high testing costs and relatively complex procedures, making them unsuitable for early screening of large populations and hindering the efficient detection of undiagnosed patients.

[0005] In recent years, artificial intelligence (AI) technology has developed rapidly and penetrated widely into the healthcare field, providing a new technological approach for disease risk assessment. Research has found a complex intrinsic relationship between the risk of chronic kidney disease (CKD) and several easily accessible indicators, such as an individual's basic information, dietary habits, and physical activity level. Based on this, analyzing this easily accessible information using statistical models holds promise for achieving efficient and low-cost assessment of CKD risk. Given the limitations of existing CKD screening methods and the advantages of AI and statistical models in medical risk assessment, there is an urgent need to develop a CKD risk analysis scheme based on easily accessible individual information to compensate for existing technological deficiencies and improve the accessibility and effectiveness of early CKD screening. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing chronic kidney disease risk assessment methods, such as high detection costs and complex operation procedures, by providing a method and system for generating chronic kidney disease risk assessment data based on easily accessible physical examination data and comorbid medical history data, so as to improve the accessibility and effectiveness of chronic kidney disease risk assessment.

[0007] To achieve the above objectives, a first aspect of the present invention provides a method for generating chronic kidney disease risk assessment data, comprising: The data acquisition step involves acquiring physical examination data and comorbidity history characteristic data; wherein the physical examination data includes at least one first data, and the comorbidity history data includes at least one second data. The data filtering step involves using a pre-established random forest model to calculate the contribution of each first data point and each second data point, obtaining the contribution corresponding to each first data point and each second data point respectively, and filtering based on the contribution of the first data point and the second data point to determine at least one important data point. The data determination step involves constructing a logistic regression model based on the at least one important data point, outputting the probability value and odds ratio of each important data point, and determining at least one significant data point based on the probability value and odds ratio of each important data point. The data scoring step involves performing data transformation processing based on the advantage ratio of the significant data to obtain the score value of the significant data. The risk assessment step involves establishing a chronic kidney disease risk assessment model based on the score values ​​of the significant data, and calculating risk assessment data based on the chronic kidney disease risk assessment model.

[0008] Preferably, the first data includes gender, age, brachial pulse wave velocity, smoking data, alcohol consumption data, body mass index, heart rate, systolic blood pressure, diastolic blood pressure, creatinine, estimated glomerular filtration rate, ankle-brachial index, uric acid, total cholesterol, triglycerides, low-density lipoprotein, high-density lipoprotein, fasting blood glucose, glycated hemoglobin, C-reactive protein, or urinary microalbumin-creatinine ratio.

[0009] Preferably, the second data is hypertension, hyperuricemia, hyperlipidemia, diabetes, heart disease, autoimmune disease, or kidney disease.

[0010] Preferably, after acquiring physical examination data and comorbid medical history characteristics data, the method further includes: The first data and / or the second data are preprocessed using a preset data preprocessing method.

[0011] Preferably, the data preprocessing of the first data and / or the second data using a preset data preprocessing method specifically includes: The first data and / or the second data are converted into binary classification data according to a preset threshold.

[0012] Preferably, the first data and the second data include multiple types of variables. The step of using a pre-established random forest model to calculate the contribution of each of the first data and each of the second data to obtain the contribution corresponding to each of the first data and each of the second data specifically includes: Random sampling with replacement is performed on each of the first data and each of the second data to form multiple bootstrap samples and out-of-bag data corresponding to each bootstrap sample; The data outside the bag is classified to obtain the first voting score corresponding to the self-service sample for each data outside the bag; Select the first type of variable from the multiple types of variables, and randomly change the order of the values ​​of the first type of variable in the data outside each bag to form a second test sample; The random forest model is used to classify the second test sample to obtain the second voting score for each bootstrap sample; The importance of each first data point and each second data point is calculated based on the first voting score and the second voting score, and the contribution corresponding to each first data point and each second data point is obtained respectively.

[0013] Preferably, constructing a logistic regression model based on the at least one important data point specifically includes: The binary dependent variable y and m independent variables x1, x2, ..., xm that influence the value of y. m Given m independent variables, the conditional probability p(y=1|x1,x2,……,x) of y=1 is... m The constructed logistic regression model is expressed as follows: Where β0 is a constant term, β1,β2,……,β m These are partial regression coefficients.

[0014] Preferably, the odds ratio of the significant data satisfies: All other things being equal, the first risk factor Two different exposure levels and The natural logarithm of the dominance ratio is: in, , These represent the odds ratio and odds coefficient after multivariate adjustment, respectively. , They represent when The probability of disease occurrence when the values ​​are c0 and c1. Represented as: .

[0015] Preferably, establishing a chronic kidney disease risk assessment model based on the score values ​​of the significant data specifically includes: The advantage ratio of each significant data point is converted according to a preset conversion method to obtain the score value of each significant data point. The chronic kidney disease risk assessment model is as follows: in, The values ​​of the binary variable after transforming each significant data point are: These are the scores for each significant data point.

[0016] A second aspect of the present invention provides a chronic kidney disease risk assessment data generation system, the system comprising: A data acquisition unit is used to acquire physical examination data and comorbidity history characteristic data; wherein, the physical examination data includes at least one first data, and the comorbidity history data includes at least one second data; The data filtering unit is used to calculate the contribution of each first data and each second data using a pre-established random forest model, to obtain the contribution corresponding to each first data and each second data respectively, and to filter at least one important data based on the contribution of the first data and the second data. A data determination unit is used to construct a logistic regression model based on the at least one important data, output the probability value and odds ratio of each important data, and determine at least one significant data based on the probability value and odds ratio of each important data. A data scoring unit is used to perform data transformation processing based on the dominance ratio of the significant data to obtain a score value for the significant data. The risk assessment unit is used to establish a chronic kidney disease risk assessment model based on the score values ​​of the significant data, and to perform calculations based on the chronic kidney disease risk assessment model to obtain risk assessment data.

[0017] This invention provides a method and system for generating chronic kidney disease risk assessment data. Based on readily available physical examination data and comorbid medical history data, it generates risk assessment data for chronic kidney disease. Compared to existing technologies, this invention offers at least the following beneficial technical effects: 1. The present invention is simple to operate in clinical practice. It only requires the simple application of a scoring model trained on a large sample to quickly assess the risk of chronic kidney disease in examinees using easily accessible physical examination data. 2. The present invention collects a sufficiently large sample size of subjects and conducts a wide range of tests, including various comorbid medical history characteristics, blood routine and urine routine test data, and comprehensively predicts the risk of chronic kidney disease in subjects by combining multiple characteristic variables; 3. This invention combines random forest model and logistic regression model to perform multiple screening of feature variables, which can better remove redundant features and improve the accuracy of the model. Attached Figure Description

[0018] Figure 1 A flowchart illustrating a method for generating risk assessment data for chronic kidney disease, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of chronic kidney disease risk assessment data derived from a first dataset, provided in an embodiment of the present invention. Figure 3 A block diagram of a chronic kidney disease risk assessment data generation system provided in an embodiment of the present invention; Figure 4 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0020] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0021] The present invention provides a method and system for generating risk assessment data for chronic kidney disease, which applies artificial intelligence and statistical modeling technology to the field of risk assessment for chronic kidney disease. This changes the traditional screening method that relies on complex detection methods and enables the generation of low-cost, efficient and accurate assessment data for the risk of chronic kidney disease.

[0022] Figure 1 A flowchart of a method for generating risk assessment data for chronic kidney disease provided by an embodiment of the present invention is shown below. Figure 1 The technical solution of the present invention will be described with reference to specific embodiments.

[0023] like Figure 1As shown in the figure, an embodiment of the present invention provides a method for generating chronic kidney disease risk assessment data, which includes the following steps: Data acquisition step S100: Acquire physical examination data and comorbid medical history characteristics data.

[0024] The physical examination data includes at least one primary data point, and the comorbid medical history data includes at least one secondary data point.

[0025] Specifically, in this embodiment of the invention, the physical examination data and comorbidity history characteristics data are obtained from a proprietary clinical database. Specifically, they include multiple physical examination data and comorbidity history characteristics of 90,035 people from a certain hospital between 2009 and 2022, totaling 111,942 times. The physical examination data and comorbidity history characteristics data obtained in this invention have been obtained after preliminary screening by professionals based on clinical practice.

[0026] In a preferred embodiment of the present invention, the acquired physical examination data includes at least one of the following first data: gender, age, brachial-ankle pulse wave velocity (bapwv), smoking data, alcohol consumption data, body mass index (BMI), heart rate, systolic blood pressure, diastolic blood pressure, creatinine, estimated glomerular filtration rate, ankle-brachial index (ABI), uric acid, total cholesterol, triglycerides, low-density lipoprotein, high-density lipoprotein, fasting blood glucose, glycated hemoglobin, C-reactive protein, or urinary microalbumin-to-creatinine ratio, etc. The acquired comorbid medical history characteristics data include at least one of the following second data: hypertension, hyperuricemia, hyperlipidemia, diabetes, heart disease, autoimmune disease, or kidney disease, etc.

[0027] In a preferred embodiment of the present invention, the first data and the second data can be expanded or deleted according to the development of medical science and technology, so as to optimize the risk assessment data generation method provided by the embodiments of the present invention.

[0028] In a preferred embodiment of the present invention, after acquiring physical examination data and comorbidity history characteristic data, a preset data preprocessing method is used to preprocess the first data and / or the second data. The preset data preprocessing method is pre-installed in this solution before implementing the chronic kidney disease risk assessment data generation method of the present invention, and includes multiple data preprocessing methods to process different types of physical examination data and comorbidity history characteristic data. Specific data processing methods include: The following first or second data points that can be preprocessed using binary classification are converted into binary classification data according to their respective preset thresholds, for example: Hypertension: If diastolic blood pressure > 90 or systolic blood pressure > 140, the subject is considered to have hypertension, and the variable is assigned a value of 1; otherwise, it is assigned a value of 0. Here, 90 is regarded as the preset threshold for diastolic blood pressure, and 140 is the preset threshold for systolic blood pressure.

[0029] Diabetes: If fasting blood glucose > 7.1 mmol / L or glycated hemoglobin > 6.5%, the subject is considered to have hypertension, and this variable is assigned a value of 1; otherwise, it is assigned a value of 0. Here, 7.1 mmol / L is considered the preset threshold for fasting blood glucose, and 6.5% is considered the preset threshold for glycated hemoglobin.

[0030] Chronic kidney disease: If the urinary albumin / creatinine ratio (UACR) is ≥37 mg / g or creatinine is ≥110 μmol / L, or if a clinician diagnoses the patient with kidney disease or chronic kidney disease, the subject is considered to have kidney disease, and this variable is assigned a value of 1; otherwise, it is assigned a value of 0. Here, 37 mg / g is considered the preset threshold for UACR, and 110 μmol / L is considered the preset threshold for creatinine.

[0031] Hyperlipidemia: If triglycerides > 1.7 mmol / L, LDL cholesterol > 3.4 mmol / L, or total cholesterol > 5.7 mmol / L, or if diagnosed with hyperlipidemia by a physician, the subject is considered to have hyperlipidemia, and this variable is assigned a value of 1; otherwise, it is assigned a value of 0. Here, 1.7 mmol / L is considered the preset threshold for triglycerides, 3.4 mmol / L for LDL cholesterol, and 5.7 mmol / L for total cholesterol.

[0032] High uric acid: If uric acid > 420 μmol / L or a doctor diagnoses high uric acid or gout, the subject is considered to have high uric acid, and this variable is assigned a value of 1; otherwise, it is assigned a value of 0. 420 μmol / L is considered the preset threshold for uric acid.

[0033] BAPWV: Determining the abnormal threshold of BAPWV is relatively complex and requires a comprehensive assessment based on the subject's gender, age, and BAPWV, as shown in Table 1 below: Table 1: For example, in a specific example of the present invention, a subject, male, 35 years old, has a bapwv value of 1487 cm / s, which exceeds the corresponding preset threshold of 1304. Therefore, the subject's bapwv is considered abnormal, and the variable is assigned a value of 1. Otherwise, it is assigned a value of 0.

[0034] For other first and second data that cannot be preprocessed using the binary search method, preprocessing is performed according to their respective preset thresholds to obtain their respective preprocessed first and second data.

[0035] In the data filtering step S200, a pre-established random forest model is used to calculate the contribution of each first data point and each second data point, obtaining the contribution corresponding to each first data point and each second data point respectively. Based on the contribution of the first data point and the second data point, filtering is performed to determine at least one important data point.

[0036] Specifically, in this embodiment of the invention, the random forest model is pre-established, and the acquired first and second data include multiple types of variables. The pre-established random forest model is used to calculate the contribution of each first and second data, and the contribution corresponding to each first and second data is obtained, specifically including: Step S201: Randomly sample each first data point and each second data point with replacement to form multiple bootstrap samples and out-of-bag data corresponding to each bootstrap sample.

[0037] Step S202: Classify the data outside the bag to obtain the first voting score corresponding to the self-service sample for each data outside the bag.

[0038] Step S203: Select the first type of variable from the multiple types of variables, and randomly change the order of the values ​​of the first type of variable in the data outside each bag to form the second test sample.

[0039] Step S204: Use a random forest model to classify the second test samples and obtain the second voting score for each bootstrap sample.

[0040] Step S205: Calculate the importance of each first data point and each second data point based on the first voting score and the second voting score, and obtain the contribution corresponding to each first data point and each second data point respectively.

[0041] After calculating the contribution of each first data point and each second data point through the above steps, sort the first data points and each second data point in descending order of contribution. Then, determine the first number of first data points and second data points as important data points from the sorted first data points and each second data point according to the first number, where the first number is greater than or equal to 1 and the first number is an integer.

[0042] In the data determination step S300, a logistic regression model is constructed based on at least one important data point, outputting the probability value and odds ratio of each important data point, and at least one significant data point is determined based on the probability value and odds ratio of each important data point.

[0043] Specifically, in a preferred embodiment of the present invention, let the dependent variable y be a binary variable, with values ​​of y=1 (positive result: onset, effective, death, etc.) or y=0 (negative result: no onset, ineffective, survival, etc.), and the m independent variables affecting the value of y be x1, x2, ..., xm. m For example, the dependent variable y=1 represents having chronic kidney disease, and y=0 represents not having chronic kidney disease; the independent variables are gender, age (years), and whether or not one has hypertension, etc. Given the conditional probability p=p(y=1|x1,x2,……,xm) of a positive outcome under the influence of m independent variables (i.e., exposure factors), the Logistic regression model can be expressed as: (Equation 1) Perform a logit transformation on (Equation 1), i.e., use the following transformation (Equation 2): (Equation 2) The logistic regression model can then be expressed in the following linear form: (Equation 3) In summary, the steps for building a logistic regression model based on important data include: a binary dependent variable y and m independent variables x1, x2, ..., xn that influence the value of y. m Given m independent variables, the conditional probability p(y=1|x1,x2,……,x) of y=1 is... m The constructed logistic regression model is expressed as (Equation 1).

[0044] After establishing the logistic regression model, the probability value and advantage ratio of each important data point are output. Then, according to the second number, a second number of significant data points are determined from the first number of important data points based on the probability value and advantage ratio of each important data point. The second number is greater than or equal to 1 and is an integer.

[0045] In a preferred embodiment of the present invention, the odds ratio of significant data satisfies the following condition: All other things being equal, the first risk factor Two different exposure levels and The natural logarithm of the dominance ratio is: (Equation 4) in, , These represent the odds ratio and odds coefficient after multivariate adjustment, respectively. , They represent when The probability of disease occurrence when the values ​​are c0 and c1. Represented as: (Equation 5) The logistic regression model of this invention was constructed through the above steps, and at least one significant data point was identified. The number of significant data points is determined by a second quantity, and the relationship between the second quantity and the first quantity satisfies that the first quantity is greater than or equal to the second quantity.

[0046] In the data scoring step S400, the data is transformed based on the advantage ratio of the significant data to obtain the score value of the significant data.

[0047] Specifically, in this embodiment of the invention, the advantage ratio of significant data is transformed using a preset data transformation method to obtain the score value of significant data. For example, in a specific example of this embodiment, the obtained advantage ratio (OR) values ​​of each variable are transformed into variable scores, with 1 point assigned for every 0.5, and 1 point also assigned for values ​​greater than 0.25 but less than 0.5. That is to say, the advantage ratio of significant data is transformed into the score value of each variable by assigning 1 point for every 0.5 points.

[0048] In risk assessment step S500, a chronic kidney disease risk assessment model is established based on the score values ​​of significant data, and risk assessment data is obtained by calculation based on the chronic kidney disease risk assessment model.

[0049] Specifically, in a preferred embodiment of the present invention, the chronic kidney disease risk assessment model established based on the score values ​​of significant data is as follows: (Equation 6) in, For risk assessment data, The values ​​of the binary variable after transforming each significant data point are: These are the scores for each significant data point.

[0050] In a preferred embodiment of the present invention, the risk assessment data is divided into four levels according to the score: low risk (≤5), medium risk (6~10), high risk (11~15), and very high risk (≥16).

[0051] The above steps S100-S500 establish a method for generating chronic kidney disease risk assessment data provided by this invention. Figure 2This is a schematic diagram of chronic kidney disease (CKD) risk assessment data derived from a first dataset provided in an embodiment of the present invention. It illustrates CKD risk assessment data derived from the first dataset, which includes physical examination data and comorbid medical history characteristics data obtained in step S100. The darker color represents the first dataset; the lighter color represents the validation dataset. The overall incidence of CKD in the first dataset is 5.21%, and the assessed incidence of CKD increases with increasing risk score. In the validation dataset of 34,811 individuals, the model demonstrated good discriminative ability; the increased risk score was again strongly correlated with CKD, as shown in the figure, with low and high risk scores ranging from 2% to 37%, respectively.

[0052] The above provides a detailed description of a method for generating risk assessment data for chronic kidney disease provided by this invention. Embodiments of this invention also provide a system for generating risk assessment data for chronic kidney disease to implement this data generation method. The following sections will refer to the appendix... Figure 3 A detailed introduction to the system is provided.

[0053] Figure 3 This is a block diagram of a chronic kidney disease risk assessment data generation system provided in an embodiment of the present invention. The chronic kidney disease risk assessment data generation system 10000 provided by the present invention includes: The data acquisition unit 10001 is used to acquire physical examination data and comorbidity history characteristic data; wherein, the physical examination data includes at least one first data and the comorbidity history data includes at least one second data.

[0054] The data filtering unit 10002 is used to calculate the contribution of each first data and each second data using a pre-established random forest model, to obtain the contribution corresponding to each first data and each second data respectively, and to filter at least one important data based on the contribution of the first data and the second data.

[0055] The data determination unit 10003 is used to construct a logistic regression model based on at least one important data point, output the probability value and advantage ratio of each important data point, and determine at least one significant data point based on the probability value and advantage ratio of each important data point.

[0056] Data scoring unit 10004 is used to perform data transformation processing based on the advantage ratio of significant data to obtain the score value of significant data.

[0057] Risk assessment unit 10005 is used to establish a chronic kidney disease risk assessment model based on the score values ​​of significant data, and to calculate risk assessment data based on the chronic kidney disease risk assessment model.

[0058] The specific implementation steps of each unit above have been described in detail in the introduction of the method for generating chronic kidney disease risk assessment data provided by this invention, and will not be repeated here.

[0059] This invention provides a method and system for generating risk assessment data for chronic kidney disease. Based on readily available physical examination data and comorbid medical history data, it generates risk assessment data for chronic kidney disease. Compared to existing technologies, this invention offers at least the following beneficial technical effects: 1. The present invention is simple to operate in clinical practice. It only requires the simple application of a scoring model trained on a large sample to quickly assess the risk of chronic kidney disease in examinees using easily accessible physical examination data. 2. The present invention collects a sufficiently large sample size of subjects and conducts a wide range of tests, including various comorbid medical history characteristics, blood routine and urine routine test data, and comprehensively predicts the risk of chronic kidney disease in subjects by combining multiple characteristic variables; 3. This invention combines random forest model and logistic regression model to perform multiple screening of feature variables, which can better remove redundant features and improve the accuracy of the model.

[0060] Based on the same concept, the present invention also provides a non-transitory computer-readable storage medium. Figure 4 This is a structural block diagram of an electronic device provided in an embodiment of the present invention, such as... Figure 4 As shown, the non-transitory computer-readable storage medium can be a server, which may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the steps of the chronic kidney disease risk scoring method.

[0061] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0062] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0063] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0064] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for generating risk assessment data for chronic kidney disease, characterized in that, The method includes: The data acquisition step involves acquiring physical examination data and comorbidity history characteristic data; wherein the physical examination data includes at least one first data, and the comorbidity history data includes at least one second data. The data filtering step involves using a pre-established random forest model to calculate the contribution of each first data point and each second data point, obtaining the contribution corresponding to each first data point and each second data point respectively, and filtering based on the contribution of the first data point and the second data point to determine at least one important data point. The data determination step involves constructing a logistic regression model based on the at least one important data point, outputting the probability value and odds ratio of each important data point, and determining at least one significant data point based on the probability value and odds ratio of each important data point. The data scoring step involves performing data transformation processing based on the advantage ratio of the significant data to obtain the score value of the significant data. The risk assessment step involves establishing a chronic kidney disease risk assessment model based on the score values ​​of the significant data, and calculating risk assessment data based on the chronic kidney disease risk assessment model.

2. The method for generating chronic kidney disease risk assessment data according to claim 1, characterized in that, The first data includes gender, age, brachial pulse wave velocity, smoking data, alcohol consumption data, body mass index, heart rate, systolic blood pressure, diastolic blood pressure, creatinine, estimated glomerular filtration rate, ankle-brachial index, uric acid, total cholesterol, triglycerides, low-density lipoprotein, high-density lipoprotein, fasting blood glucose, glycated hemoglobin, C-reactive protein or urinary microalbumin-creatinine ratio.

3. The method for generating chronic kidney disease risk assessment data according to claim 1, characterized in that, The second set of data includes hypertension, hyperuricemia, hyperlipidemia, diabetes, heart disease, autoimmune diseases, or kidney disease.

4. The method for generating chronic kidney disease risk assessment data according to claim 1, characterized in that, After acquiring physical examination data and comorbid medical history characteristics data, the method further includes: The first data and / or the second data are preprocessed using a preset data preprocessing method.

5. The method for generating chronic kidney disease risk assessment data according to claim 4, characterized in that, The specific steps of preprocessing the first data and / or the second data using a preset data preprocessing method include: The first data and / or the second data are converted into binary classification data according to a preset threshold.

6. The method for generating chronic kidney disease risk assessment data according to claim 1, characterized in that, The first and second data include multiple types of variables. The step of using a pre-established random forest model to calculate the contribution of each of the first and second data to obtain the contribution corresponding to each of the first and second data specifically includes: Random sampling with replacement is performed on each of the first data and each of the second data to form multiple bootstrap samples and out-of-bag data corresponding to each bootstrap sample; The data outside the bag is classified to obtain the first voting score corresponding to the self-service sample for each data outside the bag; Select the first type of variable from the multiple types of variables, and randomly change the order of the values ​​of the first type of variable in the data outside each bag to form a second test sample; The random forest model is used to classify the second test sample to obtain the second voting score for each bootstrap sample; The importance of each first data point and each second data point is calculated based on the first voting score and the second voting score, and the contribution corresponding to each first data point and each second data point is obtained respectively.

7. The method for generating chronic kidney disease risk assessment data according to claim 1, characterized in that, The construction of the logistic regression model based on the at least one important data specifically includes: The binary dependent variable y and m independent variables x1, x2, ..., xm that influence the value of y. m Given m independent variables, the conditional probability p(y=1|x1,x2,……,x) of y=1 is... m The constructed logistic regression model is expressed as follows: Where β0 is a constant term, β1,β2,……,β m These are partial regression coefficients.

8. The method for generating chronic kidney disease risk assessment data according to claim 1, characterized in that, The advantage ratio of the significant data satisfies: All other things being equal, the first risk factor Two different exposure levels and The natural logarithm of the dominance ratio is: in, , These represent the odds ratio and odds coefficient after multivariate adjustment, respectively. , They represent when The probability of disease occurrence when the values ​​are c0 and c1. Represented as: 。 9. The method for generating chronic kidney disease risk assessment data according to claim 1, characterized in that, The establishment of a chronic kidney disease risk assessment model based on the score values ​​of the aforementioned significant data specifically includes: The advantage ratio of each significant data point is converted according to a preset conversion method to obtain the score value of each significant data point. The chronic kidney disease risk assessment model is as follows: in, The values ​​of the binary variable after transforming each significant data point are: These are the scores for each significant data point.

10. A system for generating risk assessment data for chronic kidney disease, characterized in that, The system includes: A data acquisition unit is used to acquire physical examination data and comorbidity history characteristic data; wherein, the physical examination data includes at least one first data, and the comorbidity history data includes at least one second data; The data filtering unit is used to calculate the contribution of each first data and each second data using a pre-established random forest model, to obtain the contribution corresponding to each first data and each second data respectively, and to filter at least one important data based on the contribution of the first data and the second data. A data determination unit is used to construct a logistic regression model based on the at least one important data, output the probability value and odds ratio of each important data, and determine at least one significant data based on the probability value and odds ratio of each important data. A data scoring unit is used to perform data transformation processing based on the dominance ratio of the significant data to obtain a score value for the significant data. The risk assessment unit is used to establish a chronic kidney disease risk assessment model based on the score values ​​of the significant data, and to perform calculations based on the chronic kidney disease risk assessment model to obtain risk assessment data.