Chronic kidney disease risk assessment method and chronic kidney disease risk assessment system

A genetic and clinical data-driven approach using machine learning models enhances chronic kidney disease risk assessment, addressing population-specific genetic variations for early detection and management.

JP2026517963APending Publication Date: 2026-06-02ホンミェンチェ

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ホンミェンチェ
Filing Date
2023-12-08
Publication Date
2026-06-02

Smart Images

  • Figure 2026517963000001_ABST
    Figure 2026517963000001_ABST
Patent Text Reader

Abstract

This disclosure provides a method for assessing the risk of chronic kidney disease. A method for assessing chronic kidney disease risk comprises the steps of: providing a reference database; providing a subject's nucleic acid sample and biological dataset; performing a gene testing step; performing a risk score calculation step; and performing a model establishment step, wherein a machine learning algorithm is used to train and converge multiple reference polygene risk score data, multiple reference clinical data, and multiple reference gene marker data from the reference database to obtain an analytical model; and performing a data analysis step, wherein the analytical model is used to analyze the subject's polygene risk score, clinical data, and gene marker data to obtain risk analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for analyzing medical information and a system therefor. In particular, the present disclosure relates to a method for evaluating the risk of chronic kidney disease and a system for evaluating the risk of chronic kidney disease.

Background Art

[0002] Chronic kidney disease (CKD) is a chronic disease that is spreading worldwide. Chronic kidney disease and end-stage kidney disease are increasing in morbidity due to the lack of early prediction and reliable treatment methods, and have become a global public health problem. According to the statistical results of the World Health Organization (WHO) and published literature, about 5% to 10% of the world's population suffers from chronic kidney disease, and the global prevalence of chronic kidney disease across all ages increased by 29.3% from 1990 to 2017.

[0003] The treatment and management of chronic kidney disease require huge medical resources and financial expenditures, which have imposed a heavy burden on individuals and society. Therefore, the early detection of chronic kidney disease and the establishment of an early risk assessment method contribute to the early prevention and treatment of chronic kidney disease and lead to the reduction of the burden on individuals and society.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Currently, the risk prediction models used for chronic kidney disease mainly predict the risk that a patient will develop end-stage kidney disease within 2 to 5 years based on the evaluation of clinical factors (such as age, gender, estimated glomerular filtration rate (eGFR), and the concentrations of albumin, calcium, phosphate, bicarbonate, etc.). However, recent studies have revealed that chronic kidney disease is associated with genetic factors, and that the genetic background of different populations affects the association between genes and diseases. Therefore, there is a need for a method for predicting the risk of chronic kidney disease in individuals with different genetic backgrounds. [Means for solving the problem]

[0005] According to one embodiment of the present disclosure, a method for assessing chronic kidney disease risk includes the step of providing a reference database, wherein the reference database includes multiple target gene data and a reference ontology dataset, and the reference ontology dataset includes multiple reference polygene risk score data, multiple reference clinical data and multiple reference gene markers (genetic The method includes: a step of including marker) data; a step of providing a nucleic acid sample and a biological dataset of a subject, wherein the biological dataset includes clinical data and at least one gene marker data; a step of performing a gene testing step, wherein a plurality of single nucleotide polymorphisms (SNPs) of the nucleic acid sample are simultaneously detected by a nucleic acid testing method, and the plurality of SNPs are compared with the plurality of target gene data to obtain a genotype determination result; a step of performing a risk score calculation step, wherein the genotype determination result is calculated by a polygene risk score evaluation method to obtain a polygene risk score; a step of performing a model establishment step, wherein a machine learning algorithm is used to train the plurality of reference polygene risk score data, the plurality of reference clinical data and the plurality of reference gene marker data to achieve convergence and obtain an analysis model; and a step of performing a data analysis step, wherein the analysis model is used to analyze the polygene risk score, the clinical data and the at least one gene marker data to obtain a risk analysis result, and the risk analysis result is used to evaluate the age at which the subject suffers from chronic kidney disease.

[0006] According to other forms of the present disclosure, a chronic kidney disease risk assessment system is used to store a subject's biological dataset and a reference database, the biological dataset comprising nucleic acid samples, clinical data and at least one gene marker data, the reference database comprising multiple target gene data and a reference ontology dataset, and the reference ontology dataset comprising multiple reference polygene risk score data, multiple reference clinical data and multiple reference gene marker data, and a non-transient machine-readable medium signal-connected to the non-transient machine-readable medium to detect multiple single nucleotide polymorphisms of the nucleic acid sample and obtain a genotyping result, and the multiple single nucleotide polymorphisms and The system comprises a gene testing device used to compare the aforementioned multiple target gene data, and a processor signal-connected to the gene testing device and the non-transient machine-readable medium, which includes polygene risk score software and a risk assessment program, wherein the genotype determination result is calculated by the polygene risk score software to obtain a polygene risk score, the risk assessment program includes an analysis model, and the polygene risk score, the clinical data and the at least one gene marker data are analyzed by the analysis model to obtain a risk analysis result, and the risk analysis result is used to assess the age at which the subject suffers from chronic kidney disease. [Brief explanation of the drawing]

[0007] This disclosure can be more fully understood by reading the following detailed description of the embodiments while referring to the attached drawings below. [Figure 1] Figure 1 is a flowchart illustrating a method for assessing chronic kidney disease risk according to the first embodiment of this disclosure. [Figure 2] Figure 2 is a flowchart illustrating a method for assessing chronic kidney disease risk according to a second embodiment of this disclosure. [Figure 3] Figure 3 is a block diagram showing a chronic kidney disease risk assessment system according to a third embodiment of the present disclosure. [Figure 4]Figure 4 shows the risk analysis results obtained from the chronic kidney disease risk assessment system of this disclosure. [Figure 5] Figure 5 shows the results of the risk analysis for the East Asian population using the chronic kidney disease risk assessment system described herein. [Figure 6] Figure 6 shows the results of the risk analysis for the European population using the chronic kidney disease risk assessment system described herein. [Figure 7] Figure 7 shows the discriminant statistics results for East Asian and European populations on the test dataset using the analysis model described herein. [Figure 8] Figure 8 shows a forest plot comparison of odds ratios (ORs) for chronic kidney disease risk in groups with different polygenic risk scores. [Modes for carrying out the invention]

[0008] This disclosure is further illustrated by the following specific examples to facilitate the full use and implementation of the disclosure by those skilled in the art without excessive interpretation or experimentation. However, these practical details are not necessary, although they are used to illustrate how to carry out the materials and methods of this disclosure.

[0009] [Chronic kidney disease risk assessment method in this disclosure] Referring to Figure 1, Figure 1 is a flowchart showing a chronic kidney disease risk assessment method 100 according to a first embodiment of the present disclosure. The chronic kidney disease risk assessment method 100 includes steps 110, 120, 130, 140, 150, and 160.

[0010] In step 110, a reference database is provided, the reference database including multiple target gene data and a reference ontology dataset, and the reference ontology dataset includes multiple reference polygene risk score data, multiple reference clinical data and multiple reference gene marker data.

[0011] In particular, the multiple target gene data includes, but is not limited to, multiple gene data related to chronic kidney disease, including ASCC3 gene data, CDC1 gene data, F12 gene data, FAM47E gene data, FBXO22 gene data, HCRTR2 gene data, KNG1 gene data, LRP2 gene data, RAI14 gene data, NRG4 gene data, PAX8 gene data, PDILT gene data, SIM1 gene data, STC1 gene data, TINAG gene data, UBE2Q2 gene data, and WDR72 gene data.

[0012] The reference ontology dataset may be a dataset obtained from the results of patient diagnoses, examinations, and data analyses in the subject's electronic medical record (EMR) in a hospital, and the multiple reference polygene risk score data, multiple reference clinical data, and multiple reference gene marker data used in this disclosure are further described below.

[0013] In particular, the multiple reference polygene risk score data is obtained from the results of genotype imputation in the subject's multiple genetic data, one of which is obtained from the results of genetic testing, and each of the multiple reference polygene risk score data is a composite score calculated from the genotypes of each of the subject's multiple loci and the weights corresponding to the genotypes of those loci.

[0014] Multiple reference clinical data include multiple reference age data, multiple reference sex data, multiple reference family medical history data, multiple reference personal medical history data, multiple reference drug use history data, and multiple reference biochemical test data of the subject, each of which may be blood test data, and the blood test data includes reference estimated glomerular filtration rate (eGFR) data, reference albumin concentration data, reference calcium concentration data, reference phosphate concentration data, and reference bicarbonate concentration data.

[0015] Multiple reference gene marker data may include multiple molecular gene marker data, and multiple molecular gene marker data may include multiple reference RNA-level data and multiple reference protein-level data of a subject, and multiple reference RNA-level data and multiple reference protein-level data may correspond to the RNA-level and protein-level data of multiple target gene data. Therefore, this contributes to improving the evaluation accuracy of the analysis model using multiple reference gene marker data.

[0016] In step 120, a nucleic acid sample and a biological dataset of a subject are provided, the biological dataset including clinical data and data for at least one gene marker. In particular, the subject may be an inpatient or an outpatient, and the nucleic acid sample may be blood, urine, saliva, or a combination thereof, but is not limited to these. The clinical data may include age data, sex data, family medical history data, personal medical history data, drug use history data, and biochemical test data, and the biochemical test data may include estimated glomerular filtration rate data, albumin concentration data, calcium concentration data, phosphate concentration data, and bicarbonate concentration data. The clinical data may further include lifestyle data, which may include meal frequency data, smoking habit data, and alcohol use data, but is not limited to these. The gene marker data may include multiple RNA-level data and multiple protein-level data, and the multiple RNA-level data and multiple protein-level data may correspond to the RNA-level and protein-level data of multiple target gene data.

[0017] In step 130, a gene testing step is performed to simultaneously detect multiple single nucleotide polymorphisms (SNPs) in a nucleic acid sample using a nucleic acid testing method, and a genotype determination result is obtained by comparing the multiple SNPs with multiple target gene data. Specifically, the nucleic acid testing method may be a polymerase chain reaction (PCR) method, a gene chip analysis method, or a next-generation sequencing (NGS) method, but this disclosure is not limited to these.

[0018] In Project 140, a risk score calculation process is executed, and a genotype determination result is calculated by a polygenic risk score evaluation method to obtain a polygenic risk score. In particular, the polygenic risk score evaluation method may be executed by PRSice-2 software, and the polygenic risk score is calculated by PRSice-2 software based on the following formula:

Number

[0019] Note that since the details of the calculation of the polygenic risk score are well-known in the art, the detailed process will not be described again in this specification.

[0020] In Project 150, a model establishment process is executed, and a plurality of reference polygenic risk score data, a plurality of reference clinical data, and a plurality of reference gene marker data are trained by a machine learning algorithm to achieve convergence and obtain an analysis model. In particular, in the chronic kidney disease risk assessment method 100, the machine learning algorithm may be a linear regression algorithm, a logistic regression algorithm, a decision tree algorithm, or a random forest algorithm, but the present disclosure is not limited thereto.

[0021] In step 160, a data analysis step is performed, and the polygene risk score, clinical data and at least one genetic marker data are analyzed using an analysis model to obtain risk analysis results, and the risk analysis results are used to assess the age at which the subject has chronic kidney disease. Specifically, the risk analysis results include a cumulative incidence curve of chronic kidney disease across polygene risk categories, the polygene risk categories may be grouped based on the ranking results of the polygene risk score, and the cumulative incidence curve may be used to assess the age at which the subject has chronic kidney disease.

[0022] Therefore, the analysis model analyzes the subject's polygene risk score, clinical data, and genetic marker data. The analysis model is established by training it with multiple reference polygene risk score data, multiple reference clinical data, and multiple reference genetic marker data from a reference ontology dataset. Thus, the chronic kidney disease risk assessment method 100 of this disclosure can rapidly and accurately output the risk analysis results for the subject. Therefore, the chronic kidney disease risk assessment method 100 of this disclosure has excellent clinical application potential to contribute to the diagnosis of chronic kidney disease and the design of subsequent medical plans for the subject.

[0023] Referring to Figure 2, Figure 2 is a flowchart showing a chronic kidney disease risk assessment method 200 according to a second embodiment of the present disclosure. The chronic kidney disease risk assessment method 200 includes steps 210, 220, 230, 240, 250, 260 and 270, and since steps 210, 220, 230, 240, 250 and 260 are the same as steps 110, 120, 130, 140, 150 and 160 in Figure 1, the same details between these steps will not be described again.

[0024] In step 270, a disease detection step is performed. Specifically, based on the risk analysis results obtained from the analytical model, periodic follow-up examinations are performed on the subject to check whether the subject's biochemical indicators are abnormal, thereby contributing to the confirmation of whether the subject has early chronic kidney disease. The biochemical indicators may include, but are not limited to, urine protein levels, blood pressure, creatinine concentration, calcium concentration, and potassium concentration.

[0025] Therefore, by regularly examining the biochemical indicators of subjects, chronic kidney disease can be treated and managed earlier when its symptoms are detected early, and the effectiveness of treatment and quality of life of the subjects can be improved. Thus, the chronic kidney disease risk assessment method 200 has excellent potential for clinical application.

[0026] [The chronic kidney disease risk assessment system described in this disclosure] Referring to Figure 3, Figure 3 is a block diagram showing a chronic kidney disease risk assessment system 300 according to a third embodiment of the present disclosure. The chronic kidney disease risk assessment system 300 includes a non-transient machine-readable medium 310, a genetic testing equipment 320, and a processor 330.

[0027] A non-transient machine-readable medium 310 is used to store a subject's biological dataset and a reference database, the biological dataset including nucleic acid samples, clinical data and at least one gene marker data, the reference database including multiple target gene data and a reference ontology dataset, and the reference ontology dataset including multiple reference polygene risk score data, multiple reference clinical data and multiple reference gene marker data.

[0028] Specifically, nucleic acid samples may be blood, urine, saliva, or a combination thereof. Clinical data may include age data, sex data, family medical history data, personal medical history data, drug use history data, and biochemical test data, and biochemical test data may include estimated glomerular filtration rate data, albumin concentration data, calcium concentration data, phosphate concentration data, and bicarbonate concentration data. Clinical data may further include lifestyle data, which may include, but are not limited to, meal frequency data, smoking habit data, and alcohol use data. Gene marker data may include multiple RNA-level data and multiple protein-level data, and the multiple RNA-level data and multiple protein-level data may correspond to the RNA-level and protein-level data of multiple target gene data.

[0029] Multiple target gene data include, but are not limited to, ASCC3 gene data, CDC1 gene data, F12 gene data, FAM47E gene data, FBXO22 gene data, HCRTR2 gene data, KNG1 gene data, LRP2 gene data, RAI14 gene data, NRG4 gene data, PAX8 gene data, PDILT gene data, SIM1 gene data, STC1 gene data, TINAG gene data, UBE2Q2 gene data, and WDR72 gene data.

[0030] The reference ontology dataset may be obtained from patient diagnoses in the electronic medical records of hospital subjects, laboratory tests, and data analysis results, and the multiple reference polygene risk score data may be obtained from the results of genotype imputation in the subject's multiple genetic data, one of which may be obtained from the results of genetic testing, and each of the multiple reference polygene risk score data may be a composite score calculated from the genotypes of each of the subject's multiple loci and the weights corresponding to the genotypes of those multiple loci.

[0031] Multiple reference clinical data include multiple reference age data, multiple reference sex data, multiple reference family medical history data, multiple reference personal medical history data, multiple reference drug use history data, and multiple reference biochemical test data of the subject, each of which may be blood test data, and the blood test data includes reference estimated glomerular filtration rate data, reference albumin concentration data, reference calcium concentration data, reference phosphate concentration data, and reference bicarbonate concentration data.

[0032] Multiple reference gene marker data may include multiple molecular gene marker data, and multiple molecular gene marker data may include multiple reference RNA-level data and multiple reference protein-level data of a subject, and multiple reference RNA-level data and multiple reference protein-level data may correspond to the RNA-level and protein-level data of multiple target gene data. Therefore, this contributes to improving the evaluation accuracy of the analysis model using multiple reference gene marker data.

[0033] The gene testing equipment 320 is signal-connected to a non-transient machine-readable medium 310 and is used to detect multiple single nucleotide polymorphisms (SNPs) in a nucleic acid sample and to compare the multiple SNPs with multiple target gene data in order to obtain a genotype determination result. Specifically, the multiple SNPs may be detected simultaneously by polymerase chain reaction, gene chip analysis, or next-generation sequencing.

[0034] The processor 330 is signal-connected to the genetic testing equipment 320 and a non-transient machine-readable medium 310, and includes polygene risk score software 340 and a risk assessment program 350, and the genotype determination result is calculated by the polygene risk score software 340 to obtain a polygene risk score. Specifically, the polygene risk score software 340 may be PRSice-2 software. The details of the calculation of the polygene risk score have been explained above, so the detailed process will not be explained again.

[0035] The risk assessment program 350 includes an analysis model 351, and polygene risk scores, clinical data, and at least one gene marker data are analyzed by the analysis model 351 to obtain risk analysis results, which are used to assess the age at which a subject will have chronic kidney disease. In particular, the analysis model is obtained by training and converging multiple reference polygene risk score data, multiple reference clinical data, and multiple reference gene marker data using a machine learning algorithm, and the machine learning algorithm may be a linear regression algorithm, a logistic regression algorithm, a decision tree algorithm, or a random forest algorithm.

[0036] Therefore, by analyzing the subject's polygene risk score, clinical data, and genetic marker data using the analysis model 351 of this disclosure, the chronic kidney disease risk assessment system 300 of this disclosure can rapidly and accurately output the subject's risk analysis results for designing the subject's subsequent medical plan, and thus the chronic kidney disease risk assessment system 300 of this disclosure has excellent potential for clinical application.

[0037] [Example] I. Reference Databases The reference database used in this disclosure is obtained from the electronic medical records of hospital subjects, and the subjects include healthy individuals and patients with chronic kidney disease.

[0038] The following examples demonstrate the evaluation of the accuracy of the chronic kidney disease risk assessment system and the chronic kidney disease risk assessment method of the Disclosure by applying the chronic kidney disease risk assessment method of the Disclosure to the chronic kidney disease risk assessment system and the chronic kidney disease risk assessment method of the Disclosure. The chronic kidney disease risk assessment system of the Disclosure may be the chronic kidney disease risk assessment system 300 as described above, and the chronic kidney disease risk assessment method of the Disclosure may be the chronic kidney disease risk assessment method 100 or the chronic kidney disease risk assessment method 200. Thus, the details of the chronic kidney disease risk assessment system and the chronic kidney disease risk assessment method are described above and will not be explained again.

[0039] II. Results 1. Risk analysis results of this disclosure Refer to Figure 4 and Table 1. Figure 4 shows the risk analysis results obtained from the chronic kidney disease risk assessment system of this disclosure. Table 1 shows the estimated age at which subjects classified into different groups in Figure 4 have chronic kidney disease.

[0040] [Table 1]

[0041] As shown in Figure 4, the risk analysis results obtained from the chronic kidney disease risk assessment system include cumulative incidence curves of chronic kidney disease (CKD) across polygene risk categories. The polygene risk categories are grouped based on the ranking results of the polygene risk score, and in Figure 4, the polygene risk categories are divided into six groups. The six groups of polygene risk categories are the top 2% of polygene risk scores (Group 1), the top 2% to 10% of polygene risk scores (Group 2), the top 10% to 20% of polygene risk scores (Group 3), the middle 20% to 60% of polygene risk scores (Group 4), the bottom 20% to 40% of polygene risk scores (Group 5), and the bottom 20% of polygene risk scores (Group 6), and the p-value of the log-rank test is less than 0.0001.

[0042] As shown in Figure 4 and Table 1, subjects classified as Group 1 are more likely to develop chronic kidney disease at age 64.8, and subjects classified as Group 6 are more likely to develop chronic kidney disease at age 89.9. Therefore, the risk analysis results obtained from the analytical model of this disclosure can be used to assess the age at which subjects develop chronic kidney disease, thereby enabling ultra-early risk prediction of chronic kidney disease.

[0043] 2. Performance of the analysis model described herein in assessing the risk of chronic kidney disease in East Asian and European populations. To evaluate the performance of the analytical model of the chronic kidney disease risk assessment system disclosed herein, the accuracy of the analytical model of the chronic kidney disease risk assessment system disclosed herein will be confirmed for East Asian and European populations using internal and external test datasets. Specifically, the internal test dataset will be collected from the East Asian Biobank, and the external test dataset will be collected from the European Biobank. The East Asian population will be analyzed using the biological dataset of the East Asian population, and the European population will be analyzed using the biological dataset of the European population.

[0044] Refer to Figures 5 to 7. Figure 5 shows the risk analysis results for the East Asian population using the chronic kidney disease risk assessment system of this disclosure. Figure 6 shows the risk analysis results for the European population using the chronic kidney disease risk assessment system of this disclosure. Figure 7 shows the discriminant statistics results for the East Asian and European populations on the test dataset using the analysis model of this disclosure. In particular, in Figure 7, the area under the receiver operating characteristic curve (hereinafter abbreviated as "AUC") for the East Asian population is 0.880 (0.875 to 0.885), and the AUC for the European population is 0.788 (0.783 to 0.793).

[0045] As shown in Figures 5 and 6, the risk analysis results obtained from the chronic kidney disease risk assessment system include cumulative incidence curves across the chronic kidney disease polygene risk categories, and the polygene risk categories are grouped based on the ranking results of the polygene risk scores. In Figure 5, the polygene risk categories are divided into three groups. The three groups of polygene risk categories in Figure 5 are the top 20% of polygene risk scores, the median 40% to 60% of polygene risk scores, and the bottom 20% of polygene risk scores, and the p-value of the log-rank test is less than 0.0001. In Figure 6, the polygene risk categories are also divided into three groups. The three groups of polygene risk categories in Figure 6 are the top 3% of polygene risk scores, the median of polygene risk scores, and the bottom 3% of polygene risk scores, and the p-value of the log-rank test is also less than 0.0001.

[0046] As further shown in Figures 5 and 6, the chronic kidney disease risk assessment system and chronic kidney disease risk assessment method of this disclosure may be used to assess the age at which a subject has chronic kidney disease, and the subject may have an East Asian or European genetic background.

[0047] As shown in Figure 7, when evaluating East Asian and European populations, the AUC of the chronic kidney disease risk assessment system and method of this disclosure exceeds 0.7. In particular, an AUC equal to 0.5 indicates that the model lacks class separation ability. As the AUC continues to improve, the class separation ability of the assessment model is getting better and better. Compared to the AUC of the European population, the AUC of the East Asian population evaluated by the chronic kidney disease risk assessment system and method of this disclosure can reach 0.880, which demonstrates the accuracy of this disclosure in evaluating chronic kidney disease in East Asian populations.

[0048] 3. Influence of polygenic risk scores on the odds ratio (OR) of chronic kidney disease Referring to Figure 8, which shows a forest plot comparison of odds ratios (ORs) for chronic kidney disease risk in groups with different polygene risk scores. Specifically, the polygene risk categories are divided into five groups: the bottom 20% of polygene risk scores (Group 1), the bottom 20% to 39% of polygene risk scores (Group 2), the median 41% to 60% of polygene risk scores (Group 3), the top 21% to 40% of polygene risk scores (Group 4), and the top 20% of polygene risk scores (Group 5). Group 1 is used as the reference group. "95% CI" represents the 95% confidence interval.

[0049] As shown in Figure 8, the risk of developing chronic kidney disease in subjects classified as Group 5 is 23 times higher than in subjects classified as Group 1.

[0050] As shown in the results above, the chronic kidney disease risk assessment method and system of this disclosure may be used to assess the age at which a subject may have chronic kidney disease, thereby achieving very early risk prediction of chronic kidney disease. Furthermore, the chronic kidney disease risk assessment method and system of this disclosure may be used to assess the risk of chronic kidney disease in subjects with different genetic backgrounds, and the chronic kidney disease risk assessment method and system of this disclosure have excellent assessment accuracy. In addition, the impact of family history of chronic kidney disease on individual risk can be assessed by the chronic kidney disease risk assessment method and system of this disclosure, improving the accuracy of chronic kidney disease assessment. Moreover, this disclosure has a broad sample collection range based on East Asian genes with the largest sample volume, and can provide improved and accurate risk assessment and prediction for chronic kidney disease patients in Asia, thereby improving renal function and reducing the burden of healthcare costs in the East Asian population. Therefore, contributing to the diagnosis of chronic kidney disease and the design of subsequent medical treatment plans for subjects, the chronic kidney disease risk assessment method and chronic kidney disease risk assessment system of this disclosure have excellent assessment accuracy and significant clinical application potential and commercial value.

[0051] While this disclosure has been described in considerable detail with reference to its specific embodiments, other embodiments are also possible. Therefore, the spirit and scope of the appended claims should not be limited to the description of the embodiments contained herein.

[0052] It will be apparent to those skilled in the art that various modifications and variations can be made to the structure of this disclosure without departing from the scope or spirit of this disclosure. Therefore, this disclosure is intended to cover modifications and variations of this disclosure to the extent that they fall within the scope of the appended claims.

Claims

1. A step of providing a reference database, wherein the reference database includes multiple target gene data and a reference ontology dataset, and the reference ontology dataset includes multiple reference polygene risk score data, multiple reference clinical data and multiple reference gene marker data. A step of providing a nucleic acid sample and a biological dataset of a subject, wherein the biological dataset includes clinical data and at least one genetic marker data. A step of performing a gene testing process, comprising: simultaneously detecting multiple single nucleotide polymorphisms (SNPs) of a nucleic acid sample using a nucleic acid testing method, and comparing the multiple SNPs with the multiple target gene data to obtain a genotype determination result; A step of performing a risk score calculation process, comprising the step of calculating the genotype determination result using a polygene risk score evaluation method to obtain a polygene risk score, A step of performing a model establishment process, comprising training the plurality of reference polygene risk score data, the plurality of reference clinical data, and the plurality of reference gene marker data using a machine learning algorithm to achieve convergence and obtain an analytical model, A step of performing a data analysis process, comprising: analyzing the polygene risk score, the clinical data and the at least one gene marker data using the analysis model to obtain a risk analysis result, and using the risk analysis result to evaluate the age at which the subject suffers from chronic kidney disease; A chronic kidney disease risk assessment method that includes the following features.

2. The method according to claim 1, wherein the plurality of target gene data include ASCC3 gene data, CDC1 gene data, F12 gene data, FAM47E gene data, FBXO22 gene data, HCRTR2 gene data, KNG1 gene data, LRP2 gene data, RAI14 gene data, NRG4 gene data, PAX8 gene data, PDILT gene data, SIM1 gene data, STC1 gene data, TINAG gene data, UBE2Q2 gene data, and WDR72 gene data.

3. The method according to claim 1, wherein the nucleic acid sample is blood, urine, saliva, or a combination thereof.

4. The method according to claim 1, wherein the nucleic acid testing method is a polymerase chain reaction (PCR) method, a gene chip analysis method, or a next-generation sequencing (NGS) method.

5. The method for evaluating the polygene risk score is the method according to claim 1, which is performed by PRSice-2 software.

6. The method according to claim 1, wherein the machine learning algorithm is a linear regression algorithm, a logistic regression algorithm, a decision tree algorithm, or a random forest algorithm.

7. The method according to claim 1, wherein the clinical data includes age data, sex data, family medical history data, personal medical history data, drug use history data, and biochemical test data.

8. The method according to claim 7, wherein the biochemical test data includes estimated glomerular filtration rate (eGFR) data, albumin concentration data, calcium concentration data, phosphate concentration data, and bicarbonate concentration data.

9. Used to store a subject's biological dataset and reference database, wherein the biological dataset includes nucleic acid samples, clinical data and at least one gene marker data, the reference database includes multiple target gene data and a reference ontology dataset, and the reference ontology dataset includes multiple reference polygene risk score data, multiple reference clinical data and multiple reference gene marker data in a non-transient machine-readable medium. A genetic testing device is used to detect multiple single nucleotide polymorphisms in a nucleic acid sample and compare the multiple single nucleotide polymorphisms with the multiple target gene data, and to obtain a genotype determination result, which is signal-connected to the aforementioned non-transient machine-readable medium. A processor is signal-connected to the aforementioned gene testing equipment and the aforementioned non-transient machine-readable medium, and includes polygene risk score software and a risk assessment program, wherein the genotype determination result is calculated by the polygene risk score software to obtain a polygene risk score. Equipped with, The risk assessment program includes an analytical model, wherein the polygene risk score, the clinical data, and the at least one gene marker data are analyzed by the analytical model to obtain a risk analysis result, and the risk analysis result is used to assess the age at which the subject suffers from chronic kidney disease.

10. The chronic kidney disease risk assessment system according to claim 9, wherein the plurality of target gene data include ASCC3 gene data, CDC1 gene data, F12 gene data, FAM47E gene data, FBXO22 gene data, HCRTR2 gene data, KNG1 gene data, LRP2 gene data, RAI14 gene data, NRG4 gene data, PAX8 gene data, PDILT gene data, SIM1 gene data, STC1 gene data, TINAG gene data, UBE2Q2 gene data, and WDR72 gene data.

11. The chronic kidney disease risk assessment system according to claim 9, wherein the nucleic acid sample is blood, urine, saliva, or a combination thereof.

12. The chronic kidney disease risk assessment system according to claim 9, wherein the plurality of single nucleotide polymorphisms are simultaneously detected by polymerase chain reaction, gene chip analysis, or next-generation sequencing.

13. The chronic kidney disease risk assessment system according to claim 9, wherein the multi-gene risk score software is PRSice-2 software.

14. The chronic kidney disease risk assessment system according to claim 9, obtained by training the analysis model with a machine learning algorithm to converge the plurality of reference polygene risk score data, the plurality of reference clinical data, and the plurality of reference gene marker data.

15. The chronic kidney disease risk assessment system according to claim 14, wherein the machine learning algorithm is a linear regression algorithm, a logistic regression algorithm, a decision tree algorithm, or a random forest algorithm.

16. The chronic kidney disease risk assessment system according to claim 9, wherein the clinical data includes age data, sex data, family medical history data, personal medical history data, drug use history data, and biochemical test data.

17. The chronic kidney disease risk assessment system according to claim 16, wherein the biochemical test data includes estimated glomerular filtration rate data, albumin concentration data, calcium concentration data, phosphate concentration data, and bicarbonate concentration data.