Three-gene combination, primer probe set and system for qPCR (quantitative polymerase chain reaction) rapid screening of tuberculosis progression risk and application of three-gene combination, primer probe set and system

The rapid qPCR screening technology using a combination of APOL1, C1QC, and SOCS1 genes addresses the shortcomings in performance and stability of existing technologies for predicting the risk of tuberculosis progression. It achieves efficient and reliable risk prediction in adult populations and across multiple time windows, making it suitable for tuberculosis screening and individualized intervention.

CN121065332APending Publication Date: 2025-12-05SHANGHAI PUBLIC HEALTH CLINICAL CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511536500.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing multi-gene blood transcriptome signatures are inadequate in predicting the risk of latent tuberculosis infection in adults progressing to active tuberculosis, failing to meet the performance standards of the World Health Organization, especially over long time windows.

Method used

By using a combination of three biomarkers—APOL1, C1QC, and SOCS1—along with a constrained parametric scoring and population/platform recalibration process, we can predict the risk of latent tuberculosis infection progressing to active tuberculosis within a predetermined time window using qPCR rapid screening technology.

Benefits of technology

It maintains stable performance in adult populations and across multiple time windows, achieving high sensitivity and specificity in risk stratification, meeting or exceeding WHO's predictive performance standards, and is suitable for screening and individualized intervention decisions in different populations, while reducing testing costs and implementation difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121065332A_ABST
    Figure CN121065332A_ABST
Patent Text Reader

Abstract

The invention discloses a three-gene combination, a primer probe set and a system for qPCR (quantitative polymerase chain reaction) rapid screening of tuberculosis progression risk and application of the three-gene combination, the primer probe set and the system. The method comprises the following steps: firstly, screening out candidate genes APOL1, C1QC and SOCS1 which are stable across time windows and have high prediction performance; then a three-gene combination scoring you model (PRI3) composed of APOL1, C1QC and SOCS1 is constructed and used for predicting the risk that a latent tuberculosis infected individual progresses into active tuberculosis within a set time window; the method comprises the following steps of: firstly, selecting a reference model (RISK6, Sweeney3, Gliddon4 and BATF2), then performing performance verification in an independent verification queue, and comparing with the existing reference model (RISK6, Sweeney3, Gliddon4 and BATF2), and the result shows that the scoring model disclosed by the invention can keep stable performance in adult population and multiple time windows (such as 0-6 months, 0-12 months and 0-24 months). The scheme of the invention is suitable for crowd screening, disease early warning and individualized intervention decision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of molecular diagnostics and bioinformatics, in particular to a three-gene combination, primer probe set, system and application for rapid screening of tuberculosis progression risk qPCR, for predicting the risk of latent tuberculosis infected individuals progressing to active tuberculosis within a given time window. BACKGROUND

[0002] Tuberculosis (TB) is caused by Mycobacterium tuberculosis and remains one of the single-pathogen related diseases with the highest global death burden. Most infected individuals are in a latent state, and only a portion of individuals will progress to active tuberculosis at some time in the future; the transition from latency to disease often goes through an asymptomatic subclinical stage, which is the key window for intervention and prevention. Therefore, how to identify individuals at high risk of progression before symptoms appear has always been a core technical problem in the prevention and control of tuberculosis.

[0003] In recent years, risk prediction models based on blood transcriptome data have been widely studied. For example, RISK6 is a six-gene blood transcriptome signature that shows good short-term prediction ability in adolescents, and other models such as BATF2, Sweeney3 and Gliddon4 also have certain performance. However, the prediction performance of these models in adult populations or in longer time windows (e.g., 12 months, 24 months) is significantly reduced, making it difficult to continuously meet the minimum performance standards for tuberculosis prediction products established by the World Health Organization (WHO), i.e., the sensitivity and specificity of predicting disease within 12 months should not be less than 75%.

[0004] Independent assessments and the views of experts in the field indicate that the overall performance of the host transcriptome response signature dominated by interferon-stimulated genes (ISG) may have reached its peak; and further believe that under the existing conventional research design and method paradigm, it is difficult to obtain better diagnostic / prediction tools by continuing to develop signature discovery. In this technical background, it is not expected to significantly surpass the existing route, which also raises higher requirements for new technical paths and implementation methods.

[0005] Based on the above status quo, there is an urgent need for a risk prediction scheme that is compact in structure, strong in information complementarity, portable across platforms and populations, and easy to re-standardize: while keeping the detection throughput and cost under control, it can achieve stable and reusable risk stratification in adult populations and multiple time windows. To this end, a three-gene complementary combination is adopted, together with a constrained parameterized scoring and population / platform re-standardization process, providing an implementable path for breaking through the existing paradigm; under the premise that the performance has been considered to have reached its peak in the industry, if significant performance improvement can be achieved on this path, it reflects the non-triviality and technical innovation value of the present application scheme. SUMMARY

[0006] The object of the present application is to solve the problems of insufficient prediction performance and stability of existing multi-gene blood transcriptome signatures (such as RISK6, Gliddon4, Sweeney3, BATF2) in predicting the risk of latent tuberculosis infection developing into active tuberculosis, and to provide a three-gene combination, a primer probe set and a system for rapid screening of tuberculosis progression risk qPCR, which is used to predict the risk of latent tuberculosis infection individuals developing into active tuberculosis within a given time window, and is suitable for screening, disease early warning and individualized intervention decision-making of different populations.

[0007] To achieve the above object, the present application adopts the following technical solutions: In a first aspect, the present application provides a biomarker for predicting the risk of tuberculosis progression, which is selected from at least one of APOL1, C1QC and SOCS1.

[0008] In some embodiments, the predicted subject is aged 18 years or older.

[0009] In some embodiments, the time window period of the prediction is any window period within 24 months, including 0-6 months, 0-7 months, 0-8 months, 0-9 months, 0-10 months, 0-11 months, 0-12 months, 0-13 months, 0-14 months, 0-15 months, 0-16 months, 0-17 months, 0-18 months, 0-19 months, 0-20 months, 0-21 months, 0-22 months, 0-23 months, 0-24 months.

[0010] In a second aspect, the present application provides use of the biomarker or its detection reagent according to the first aspect in the preparation of a product for predicting or evaluating the risk of tuberculosis progression.

[0011] In some embodiments, the product includes a kit, a detection chip, a system or a device.

[0012] Preferably, the detection reagent comprises a reagent for detecting the biomarker at the genetic level and / or the protein level.

[0013] Preferably, the detection reagent is a reagent for one or more detection techniques or methods selected from the group consisting of enzyme-linked immunosorbent assay, immunofluorescence method, radioimmunoassay method, immunoprecipitation method, immunoblotting method, high performance liquid chromatography method, capillary gel electrophoresis method, near infrared spectroscopy method, mass spectrometry method, immunochemiluminescence method, colloidal gold immunotechnology, fluorescent immunochromatography technology, surface plasmon resonance technology, immuno-PCR technology and biotin-avidin technology.

[0014] In a third aspect, the present application provides a kit for predicting the risk of developing tuberculosis, the kit comprising at least reagents for detecting the expression level of the biomarker as described in the first aspect of the present application.

[0015] In some embodiments, the kit comprises at least detection reagents for specifically detecting the mRNA expression level of the biomarker, the detection reagents comprising qPCR primer probes as shown in SEQ ID NOs: 1-9.

[0016] In some embodiments, the kit further comprises qPCR primer probes for detecting the mRNA expression level of the reference gene HPRT1, the sequences of which are shown in SEQ ID NOs: 10-12.

[0017] In some embodiments, the kit is used for detection according to the following method: first, the expression level of the biomarker is detected by the reagents in the kit, and then the expression level data obtained by detection is substituted into the pre-constructed scoring model S = a·E(APOL1) + b·E(C1QC) + c·E(SOCS1) to calculate the score S, and then the risk prediction result of the subject developing active tuberculosis is obtained by comparing the score S with the pre-obtained threshold value, wherein E represents the expression level of the corresponding gene; the coefficients a, b, and c are determined by optimization tools or algorithms (such as genetic algorithms) under the constraints of regularity and stability by training data within a preset time window, so as to ensure the optimal discrimination (such as the area under the AUC curve).

[0018] In some embodiments, the threshold value is obtained by existing clinical data sets according to a set calculation rule (such as Youden index method and / or objective constraint method) and can be re-determined by newly added clinical data sets.

[0019] In some embodiments, the kit is used for detecting the blood sample of the subject.

[0020] In a fourth aspect, the present application provides a system for predicting the risk of developing tuberculosis, the system comprising a sample collection device, a sample detection device, and a prediction device; wherein: The sample collection device is configured as a device for collecting a blood sample of a subject; the sample detection device is a device capable of detecting the expression amount of the biomarker gene or protein as described in the first aspect in the blood sample; the prediction device comprises a data input and processing module and a prediction module, the data input and processing module is configured to obtain the data detected by the sample detection device, and the data is standardized (including but not limited to logarithmic conversion, quantile normalization, Z-score standardization); the prediction module is configured to input the data into a pre-constructed scoring model: S = a·E(APOL1) + b·E(C1QC) + c·E(SOCS1), and output the prediction result according to the value S of the scoring model; wherein E represents the expression amount of the corresponding gene; the coefficients a, b and c are determined by optimization tools or algorithms (such as genetic algorithm) under the constraints of regularization and stability by training data in a preset time window, so as to ensure the optimal discrimination (such as the area under the AUC curve).

[0021] Compared with the prior art, the present application has the following beneficial effects: The present application provides a risk scoring scheme based on three complementary genes without relying on complex black box algorithms, which can maintain stable performance in adult population and multiple time windows (such as 0-6, 0-12, 0-24 months), and realize reliable migration and rapid landing through standardized threshold setting and population / platform re-calibration process; At the same time, the qPCR primer probe set and SOP / software module are matched to reduce the implementation and maintenance cost, improve the interpretability and reproducibility, and thus achieve or exceed the target performance interval proposed by WHO and meet the application requirements of grassroots and field. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 .Overall flowchart for predicting the risk of tuberculosis.

[0023] Figure 2 . Comparison of the performance of candidate gene signature in predicting the risk of tuberculosis in different time intervals and populations: Receiver operating characteristic (ROC) curves comparing the three-gene classifier PRI3 (composed of APOL1, C1QC, SOCS1) to existing benchmark models (RISK6, Gliddon4, Sweeney3, BATF2) in two age strata (12-84 years old overall population and 18-84 years old adult population); model performance is shown for three prediction time windows (0-6 months (A), 0-12 months (B), 0-24 months (C) pre-disease onset) in the figure, based on a pooled analysis of multiple external validation cohorts (ACS, London, Leicester), with 95% confidence intervals shown as shaded bands, and AUC values (95% CI) listed in the legend box.

[0024] Figure 3 Comparison of three-gene combination performance in predicting risk of tuberculosis in different age strata for 0-12 months and 0-24 months pre-disease onset prediction windows: Receiver operating characteristic (ROC) curves showing the prediction performance of the three-gene classifier PRI3 (composed of APOL1, C1QC, SOCS1) compared to existing benchmark models (RISK6, Sweeney3, Gliddon4, BATF2) in four age groups (12-17 years old, 18-25 years old, 26-35 years old, and 36-84 years old) for two time windows (0-12 months and 0-24 months pre-disease onset); all analyses are based on a pooled analysis of external validation cohorts (ACS, London, Leicester), with ROC curves shown in the figure, and 95% confidence intervals shown as shaded areas, with AUC values (95% CI) and DeLong test p-values compared to benchmark models listed in the legend box. DETAILED DESCRIPTION

[0025] In order to make the present application more obvious and easy to understand, the preferred embodiments are described in detail below with the help of the accompanying drawings.

[0026] The application firstly obtains a blood transcriptome expression matrix in a public dataset GC6-74; secondly performs whole genome screening to find high-prediction-performance candidate genes stable across time windows; then constructs a three-gene combination model (PRI3) composed of APOL1, C1QC and SOCS1 for predicting the risk of individuals with latent tuberculosis infection progressing to active tuberculosis within a given time window; then performs performance verification in independent verification cohorts (ACS, London, Leicester) and compares with existing benchmark models (RISK6, Sweeney3, Gliddon4, BATF2); finally realizes risk stratification, clinical prevention intervention and low-cost detection application. The application scheme is suitable for population screening, disease early warning and individualized intervention decision-making. The overall process of the application is shown in Figure 1 and sequentially includes: data preprocessing and time alignment, candidate screening and three-gene complementary combination determination, constrained parameterized score learning, threshold setting and population / platform re-calibration, external verification and application deployment.

[0027] wherein the risk score is represented by a constrained parameterized linear function; the coefficients thereof are automatically determined under given constraints according to training data, and can be re-calibrated according to target populations and detection platforms to ensure the portability and consistency across platforms and populations. The application is provided with qPCR primer probe sets, internal participation and quality control requirements, and software modules for data standardization, score calculation and report generation. The technical points and alternative implementation modes of each step are described one by one below.

[0028] Embodiment The embodiment provides a screening method of a three-gene combination prediction factor and a construction method of a score model, and specifically includes the following steps. Step 1: Data processing and time alignment Natural population cohort data with a clear diagnosis time point are collected to obtain peripheral blood transcriptome expression profiles of progressors and non-progressors at multiple sampling time points. The original data are standardized (such as logarithmic conversion, quantile normalization, Z-score standardization), and the following operations are completed: 1) time alignment is performed with the diagnosis date as the zero point to generate time_to_diagnosis and time window labels (such as 0-6, 0-12, 0-24 months); 2) age stratification labels (such as 12-17, 18-25, 26-35, 36-84 years old) are recorded; 3) Ct / intensity metrics are mapped to unified expression E(·) according to the platform; 4) routine quality control and outlier detection (including well consistency and batch record) are performed. To avoid mixing nonspecific signals after onset, samples on the diagnosis day (time_to_diagnosis = 0) are excluded, and subsequent modeling and evaluation are always based on pre-defined time windows.

[0029] Step 2: Candidate screening and three-gene combination determination With GC6-74 (8-60 years old) as the discovery set, after time alignment of samples by diagnosis date, the classification performance (AUC) of each gene pair on progression outcome was calculated in each of the four preset time windows (0-6, 0-12, 0-18, 0-24 months) one by one, and a candidate ranking was formed in each window accordingly. The gene set was screened based on the stability of the ranking across windows as the main criterion (e.g. high frequency of appearance and small ranking fluctuations in the middle of the ranking across windows). Finally, APOL1, C1QC, SOCS1 were determined as the three-gene combination in GC6-74 (listed in alphabetical order, not representing the order of importance). Subsequently, the three-gene combination was externally validated and evaluated in independent cohorts such as ACS, London, Leicester, etc.

[0030] Step 3: Score model construction Based on the three genes APOL1, C1QC, SOCS1 determined in Step 2, a linear risk score (PRI3, Progression Risk Index 3) was constructed: S = a·E(APOL1) + b·E(C1QC) + c·E(SOCS1), where E(·) is the standardized expression obtained by the unified process. The coefficients a, b, c are automatically determined by the public optimization method under the constraints of regularization and stability to improve the discrimination (such as AUC) by the training data in the preset time window; the specific numerical value is not limited. The model is used for external validation and threshold re-calibration after cross-validation. In some specific embodiments, we use genetic algorithm (R package GA::ga()) to optimize the linear combination coefficients of the three genes, with the AUC of the 0-24 month time window in the training set as the objective function for iterative search. The final optimal linear expression is: PROGRESS3 = E ( SOCS1 ) -0.5× [ E ( C1QC ) -2 ×E( APOL1 ) ], where E ( · ) represents the log2 CPM expression of each gene in each sample, which is standardized by Z-score.

[0031] Step 4: Threshold determination and local re-calibration (two-stage strategy) To adapt to different populations and detection platforms, the invention adopts a two-stage threshold strategy when deployed, and can set thresholds according to preset time windows (0-6, 0-12, 0-18, 0-24 months) and age stratification.

[0032] Stage 1: Initialization of threshold (θ0). With no or only a small amount of local follow-up data, a temporary threshold θ0 is obtained from the external validation cohort (e.g. ACS, London, Leicester direct merge without cross-cohort batch correction to simulate real deployment differences) according to pre-defined rules for starting a pilot run: (1) Youden Index method: maximize (sensitivity + specificity - 1) on validation data; (2) Target constraint method: find threshold under a given clinical target (e.g. minimum 75% sensitivity / 75% specificity or optimal 90% sensitivity / 90% specificity of WHO TPP); Stage 2: Local re-calibration (θ * ). With rolling increase of follow-up cohort, update threshold on local calibration set according to the same pre-defined rules: 1) Obtain local labeled samples and complete standardization (e.g. z-score; if qPCR / ΔCt platform, use linear or quantile mapping to achieve scale alignment) consistent with training protocol; 2) Recalculate threshold with Youden Index or Target constraint method to obtain θ * (Updateable by time window and age group respectively); 3) Lock θ * as the running threshold and trigger re-calculation according to periodic or sample size threshold (e.g. ≥N new follow-up per period). When the actual score S of a subject exceeds the running threshold θ * , the prediction result is that the subject has a higher risk of developing active tuberculosis; when the actual score S of a subject is lower than the running threshold θ * , the prediction result is that the subject has a lower risk of developing active tuberculosis.

[0033] Note: Threshold is a recalibratable parameter, not limited to a fixed value; it can be updated with changes in population distribution, detection platform or clinical target.

[0034] Step 5: Adaptability verification across age groups and different time windows The model is evaluated for its generalizability in the following external independent validation data sets (the number of progressors / non-progressors in parentheses): ACS cohort (Adolescent Cohort Study, 12-18 years old, high tuberculosis prevalence area, 110 / 245) London cohort (GEO: GSE94438, adult population, low-burden area, 9 / 351) Leicester cohort (GEO: GSE107995, adult population, low-burden area, 35 / 38) The stability and consistency of the model in different populations in the real world were verified through stratified evaluation of different time windows and age groups. In the external independent population cohort (including adolescent population in high TB burden areas and adult population in low burden areas), the performance of the model in different age groups (12-17 years old, 18-25 years old, 26-35 years old, 36-84 years old) and different time windows (6, 12, 24 months) was evaluated, and ROC curve, sensitivity, specificity and AUC were used as main indicators to test whether the model meets the prediction performance standards set by the World Health Organization (WHO), including the minimum standard (75% sensitivity / 75% specificity) and the optimal standard (90% sensitivity / 90% specificity).

[0035] As Figure 2 With Figure 3 shown, the three-gene model PRI3 described in the application performs better than the existing benchmark models (including RISK6, Sweeney3, Gliddon4 and BATF2) in multiple age groups and time windows, especially showing unique advantages in the 0-6 month and 0-24 month windows. Figure 2 The results in Table 1 show that PRI3 reaches the optimal performance standard set by WHO (sensitivity 90%, specificity 90%) in the 0-6 month window of the 18-84 year old adult population; in the 0-12 month and 0-24 month two longer prediction windows, PRI3 also reaches the minimum performance standard of WHO (75% sensitivity / 75% specificity), and the rest of the comparison models do not meet the standard.

[0036] Figure 3 The results in Table 1 are mainly as follows: In the 0-12 month prediction window of the 18-25 year old population, PRI3 reaches the optimal performance standard of WHO (90% sensitivity / 90% specificity).

[0037] In the 0-12 month window of the 26-35 year old population, PRI3 performs best, better than all comparison models, close to the optimal threshold of WHO.

[0038] In the 0-24 month window of the 36-84 year old population, PRI3 is the only signature that reaches the optimal performance standard of WHO (90% sensitivity / 90% specificity).

[0039] In the 12-17 year old adolescent group, the prediction performance of all candidate models is limited, reflecting that the current strategy is more suitable for adult population.

[0040] Actual application and deployment form: The constructed scoring model (three-gene combination PRI3) can be directly quantitatively determined by a conventional molecular detection method such as qPCR without relying on a complex model reasoning process due to its simple structure. The scoring result can be used for disease prevention screening, intervention decision-making, dynamic monitoring of healthy population and the like, and is particularly suitable for risk control of tuberculosis aggregation infection in collective environments such as schools, construction sites and prisons. The sequence of the qPCR primer probe set is shown in the following table:

[0041] Representative threshold values and performance: In the externally combined verification set (ACS, London, Leicester) after unified standardization processing (for example, z-score), compared with existing benchmark signatures (such as RISK6, Sweeney3, Gliddon4 and BATF2), the three-gene model PRI3 proposed in the application shows stable discriminability in different age groups and prediction windows, and reaches or exceeds the performance target set by WHO in key population and window.

[0042] (1) 18-84 years old, 0-24 months window: PRI3 reaches the minimum performance standard of WHO (≥75% sensitivity / specificity) in this population and window, and is superior to the control signature in balancing sensitivity and specificity (see Table 1 and the 18-84 years old panel in Figure 1 ).

[0043] (2) 36-84 years old, 0-24 months window: In the older adult population, PRI3 reaches the optimal performance standard of WHO (≥90% sensitivity / specificity), and is the best model in comprehensive performance (see Table 1 and the 36-84 years old panel in Figure 2 ).

[0044] (3) 18-25 years old, 0-12 months window: In the young adult population, PRI3 also reaches the optimal performance standard of WHO (≥90% sensitivity / specificity), showing an advantage in the short-term prediction window (see Table 1 and the 18-25 years old panel in Figure 2 ).

[0045] Note: The threshold values shown in the above results are exemplary parameters, and in actual deployment, the threshold value θ0 can be started, and updated to θ * (see "Step 4: Threshold setting and re-calibration").

[0046] Table 1. Performance comparison with existing signatures (RISK6, Gliddon4, Sweeney3, BATF2) (same population and time window, uniform threshold rule)

[0047] The above merely describes the preferred embodiments of the present application, and is not intended to limit the present application in any form or in essence. It should be noted that, for those skilled in the art, without departing from the present application, a number of improvements and supplements can also be made, which should also be considered as the protection scope of the present application.

Claims

1. A biomarker for predicting the risk of progression of tuberculosis, characterized in that, The biomarker is selected from at least one of APOL1, C1QC and SOCS1.

2. Use of the biomarker or the detection reagent thereof of claim 1 in the preparation of a product for predicting or evaluating the risk of progression of tuberculosis.

3. Use according to claim 2, wherein the compound is ###0002### The product comprises a kit, a detection chip, a system or a device.

4. The use according to claim 2, wherein the compound is ###0002### The detection reagent comprises a reagent for detecting the biomarker at the gene level and / or the protein level.

5. The use according to claim 4, wherein the compound is ###0002### The detection reagent is a reagent for one or more detection techniques or methods selected from the group consisting of enzyme-linked immunosorbent assay, immunofluorescence method, radioimmunoassay method, co-immunoprecipitation method, immunoblotting method, high-performance liquid chromatography method, capillary gel electrophoresis method, near-infrared spectroscopy method, mass spectrometry method, immunochemiluminescence method, colloidal gold immunological technology, fluorescent immunochromatographic technology, surface plasmon resonance technology, immuno-PCR technology and biotin-avidin technology.

6. A kit for predicting the risk of progression of tuberculosis, characterized in that, The kit at least comprises a reagent for detecting the expression amount of the biomarker as described in claim 1 at the gene or protein level.

7. The kit of claim 6, wherein The kit at least comprises a detection reagent for specifically detecting the mRNA expression amount of the biomarker, which comprises a qPCR primer probe set as shown in SEQ ID NOs: 1-9.

8. The kit of claim 7, wherein The kit further comprises a qPCR primer probe set for detecting the expression amount of the reference gene HPRT1 mRNA, the sequence of which is shown in SEQ ID NOs: 10-12.

9. The kit of claim 6, wherein The kit is detected according to the following method: first, the expression amount of the biomarker is detected by the reagent in the kit, then the expression amount data obtained by detection is substituted into the pre-constructed scoring model: S = a·E(APOL1) + b·E(C1QC) + c·E(SOCS1) to calculate the score S, and then the risk prediction result of the subject developing active tuberculosis is obtained by comparing the score S with the pre-obtained threshold value, wherein E represents the expression amount of the corresponding gene; the coefficients a, b and c are determined by optimization tools or algorithms under the constraints of regularity and stability by training data within a preset time window, so as to ensure the optimal discrimination.

10. A system for predicting the risk of progression of tuberculosis, comprising a sample collection device, a sample detection device and a prediction device; wherein: The sample collection device is configured as a device for collecting a blood sample of a subject; the sample detection device is a device capable of detecting the expression amount of the biomarker gene or protein in the blood sample according to claim 1; the prediction device comprises a data acquisition and processing module and a prediction module, the data acquisition and processing module is configured to acquire the data detected by the sample detection device, and to standardize the data; the prediction module is configured to input the data obtained by the data acquisition and processing module into a pre-constructed scoring model: S = a·E(APOL1) + b·E(C1QC) + c·E(SOCS1), and output the prediction result according to the output value S of the scoring model; wherein E represents the expression amount of the corresponding gene; the coefficients a, b and c are determined by optimization tools or algorithms under the constraints of regularity and stability within a preset time window by training data, so as to ensure the optimal discrimination.