Non-ETP ALL scoring model based on transcription factor characteristics and construction method thereof

By constructing a transcription factor-based non-ETP ALL scoring model, and utilizing nine selected transcription factors and machine learning algorithms, the accuracy and stability issues of predicting non-early precursor T-cell acute lymphoblastic leukemia were resolved, achieving effective prognostic prediction for non-ETP ALL patients.

CN121459910APending Publication Date: 2026-02-03INST OF MEDICAL BIOLOGY CHINESE ACAD OF MEDICAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511617493.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing technologies lack effective multi-gene prognostic prediction models for non-early precursor T-cell acute lymphoblastic leukemia (non-ETP ALL), resulting in unstable predictions and insufficient accuracy, especially under heterogeneous conditions with different data platforms and sample sources.

Method used

A non-ETP ALL scoring model based on transcription factor characteristics was constructed. Nine key transcription factors (ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852, and ZNF765) were selected and a scoring model was constructed using machine learning algorithms to predict prognosis.

Benefits of technology

The model improved the stability and sensitivity of prognostic prediction for non-ETP ALL patients. The overall survival of high-scoring patients was significantly reduced. The model performed well in the validation set and had good reliability and predictive effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459910A_ABST
    Figure CN121459910A_ABST
Patent Text Reader

Abstract

The invention provides a non-ETP ALL scoring model based on transcription factor characteristics and a construction method thereof, and belongs to the technical field of biomedicine. The invention provides a biomarker combination for non-ETP ALL prognostic prediction based on TFs (Transfer Frameworks) characteristics. The biomarker combination comprises ZNF578, FOXO6, KCNIP3, ZNF474, BPTF (Bovine Particle Transfer Function), NHLH1, ZNF112, ZNF852 and ZNF765. The method comprises the following steps: firstly collecting next-generation sequencing data of a patient and randomly distributing a training set and a verification set, then carrying out Cox and LASSO analysis on genes in the training set, then carrying out re-screening through a plurality of machine learning algorithms so as to screen out TFs, and constructing a prognosis prediction model by taking the TFs as features. The model constructed by the method is good in performance and reliable in prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical technology, specifically relating to a non-ETP ALL scoring model based on transcription factor characteristics and its construction method. Background Technology

[0002] Acute T-cell acute lymphoblastic leukemia (T-ALL) is a highly malignant hematologic malignancy, accounting for approximately 15% of ALL cases in children and 25% in adults. Abnormally proliferating naive T cells accumulate in the thymus, bone marrow, peripheral blood, and lymphoid tissues, causing severe clinical symptoms and complications. Due to its high heterogeneity, tendency to invade the central nervous system, and insensitivity to conventional chemotherapy, T-ALL is considered a subtype with a poor prognosis. In terms of risk stratification, T-ALL is divided into early T-cell precursor ALL (ETP-ALL) and non-ETP ALL based on immunophenotype. The former is a high-risk subtype with an extremely poor prognosis, while the latter accounts for a higher proportion of patients. Despite this, non-ETP patients still experience drug resistance and relapse. The main reason for this is the extreme complexity of the pathogenesis and molecular genetic map of T-ALL, and the lack of a widely accepted molecular prognostic model in clinical practice. Therefore, how to stratify patients by risk based on the molecular biological characteristics of T-ALL and accurately predict their prognosis is a current hot topic and challenge in clinical and basic research.

[0003] With the accumulation of genomic data, prognostic models based on multi-gene expression have been established for some tumors. For example, early studies constructed expression profiles of 70 genes that can be used to distinguish between breast cancer patients with good and poor prognoses. In recent years, researchers have begun to focus on the role of transcription factors (TFs) in tumor prognosis. For example, a model that screened nine key TFs from differentially expressed TFs showed good predictive efficacy for recurrence-free survival in breast cancer. Therefore, constructing prognostic models using large-scale gene or specific functional gene combinations (such as TFs, immune genes, etc.) has become a trend in future diagnosis, treatment, and prognostic assessment.

[0004] In recent years, next-generation RNA-seq-based model development has significantly improved model stability and reproducibility. These feature models typically utilize transcriptome data and employ machine learning methods such as univariate Cox regression and LASSO to screen key genes from hundreds or thousands of candidate genes and construct risk scoring systems. Compared to single-gene biomarkers, multi-gene risk models can significantly improve prognostic accuracy. The construction of multi-gene feature models often requires strategies such as cross-validation to avoid overfitting. Common evaluation methods include multiple cross-validation and calibration curve analysis. However, most current multi-gene models specific to particular cancer types and treatment contexts have limited applicability to other tumors. Therefore, different tumor types often require the development of their own specific models.

[0005] Transcription factors (TFs), as core molecules regulating gene expression, play a crucial role in T cell development and leukemia development. Numerous studies have shown that abnormal expression of several key transcription factors in T-ALL patients is closely related to prognosis. For example, approximately 60% of T-ALL patients exhibit abnormally high expression of TAL1 or LYL1. These two helical-loop-helical TFs promote leukemia development by inhibiting E protein. In mouse models, E2A deficiency or overexpression of TAL1 / LYL1 can lead to T-ALL-like leukemia. Furthermore, mutations or abnormal expression of T cell development-related TFs such as c-MYC, IKZF1, GATA3, TCF7, and LEF1, as well as key molecules in the NOTCH signaling pathway, are frequently observed in different T-ALL subtypes. For instance, NOTCH1 signaling is activated in most T-ALL cases, and its downstream TFs, such as HES1, are involved in tumor cell proliferation and apoptosis. The TF expression profiles differ among different subtypes of T-ALL: mature T-ALL often exhibits enrichment in TAL / LMO gene rearrangements, cortical T-ALL expresses TLX1 / 3, while ETP-ALL is accompanied by high expression of HOXA groups and MEF2C. These results indicate that TF expression patterns are highly correlated with the biological characteristics and clinical prognosis of T-ALL, making them potential prognostic biomarkers.

[0006] Despite improvements in the overall prognosis of T-ALL patients, research on multigene prognostic models for T-ALL remains relatively insufficient. In adult T-ALL, the available molecular prognostic biomarkers, besides MRD, are very limited. Although a six-gene prognostic model based on the TARGET database can effectively predict overall survival in ALL patients, this model only applies to mixed ALL patient populations. To date, there is still a lack of specific predictive models for T-ALL, especially non-ETP ALL. Furthermore, due to the heterogeneity of sample sources and data platforms, existing models often exhibit instability across different cohorts. Therefore, constructing a multigene prognostic model for non-ETP ALL patients is both theoretically and practically necessary and urgent. Summary of the Invention

[0007] This invention provides a non-ETP ALL scoring model based on transcription factor characteristics and its construction method. The prognostic prediction model for non-ETP ALL constructed using nine selected TFs has good predictive stability and high sensitivity.

[0008] This invention provides a combination of biomarkers for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on transcription factor characteristics, including ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852 and ZNF765.

[0009] The present invention also provides the application of a reagent for detecting the above-mentioned combination of biomarkers in the preparation of a kit for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia.

[0010] The present invention also provides a kit for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on transcription factor characteristics, comprising reagents for detecting the transcriptional expression levels of transcription factors contained in the above-mentioned combination of biomarkers.

[0011] The present invention also provides a method for constructing a scoring model for predicting the prognosis of non-early precursor T acute lymphoblastic leukemia based on transcription factor characteristics, comprising the following steps: (1) collecting gene expression data of non-early precursor T acute lymphoblastic leukemia and randomly dividing it into n folds, where n is a positive integer ≥2, with one part serving as the training set and the other part serving as the validation set; (2) LASSO and Cox analyses were used on genes in the training set to obtain initial screening transcription factors; (3) The obtained initial screening transcription factors are processed by several machine learning algorithms, the number of gene hits in the training set of each fold is counted, and 9 transcription factors with a hit count ≥10 are collected to construct the prognostic scoring model for the non-early precursor T acute lymphoblastic leukemia; the 9 transcription factors include ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852 and ZNF765.

[0012] In a preferred embodiment of the present invention, the gene expression data in step (1) includes the expression information of 1639 transcription factors downloaded from the human transcription factor library.

[0013] In a preferred embodiment of the present invention, in step (1), the collected gene expression data is randomly divided into 4 folds, of which 3 folds are randomly selected as the training set and the remaining 1 fold is used as the validation set. Thus, each fold serves as a training set, forming a 4×4 fold cross-nested validation queue.

[0014] In a preferred embodiment of the present invention, the initial screening in step (2) includes screening for prognostic relevance in patients with non-early precursor T acute lymphoblastic leukemia to identify key transcription factors with p < 0.05.

[0015] In a preferred embodiment of the present invention, the machine learning algorithm in step (3) includes: support vector machine recursive feature elimination, random forest, random survival forest, stability selection, extreme gradient boosting, adaptive enhancement and generalized additive model.

[0016] In a preferred embodiment of the present invention, the formula for the scoring model in step (3) is: ; Where z gene This represents the expression level of a gene in the sample after standardization by the mean and standard deviation of the training set.

[0017] The present invention also provides a system or device for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia, comprising a scoring model for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on the combined characteristics of the above-mentioned biomarkers and constructed using the above-mentioned construction method.

[0018] Beneficial Effects: This invention provides a combination of biomarkers for prognostic prediction of non-ETP ALL based on TF features, including ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852, and ZNF765. This invention also constructs a prognostic prediction scoring model for non-ETP ALL based on the nine selected TFs. First, data is collected and training and validation sets are randomly assigned. Then, Cox and LASSO analyses are performed on genes in the training set. A machine learning algorithm is then used for further screening to select the nine TFs, and a prognostic prediction model is constructed using these nine TFs as features. The 9-TF model constructed in this invention performs well in the 4-fold training set, with a significantly reduced overall survival (OS) in high-scoring patients. Validation in the validation set shows that the 9-TF predicts well in all validation sets, and the prediction results have good reliability. Attached Figure Description

[0019] Figure 1 The graph shows the performance of the 9-TF model in the training set. The graph shows the KM survival curves of the 9-TF model in the 4-fold training set. Red: high score and high risk; blue: low score and low risk. Figure 2 The graph shows the performance of the 9-TF model in the validation set. The graph shows the KM survival curves of patients with the 9-TF model in the 4-fold validation set. Red: high score and high risk; blue: low score and low risk. Figure 3 The graph shows the performance comparison results of 9-TF compared to random gene models of different sizes. The line graph shows the change in the win rate of 9-TF model outperforming random gene model as the size of random gene models increases. The model's superiority or inferiority is determined by the C-index of the corresponding model in multiple Cox analysis. The horizontal axis represents the number of genes included in the random gene model; the vertical axis represents the win rate of 9-TF model relative to random gene model. The win rate statistics are based on a comparison constructed from 1000 random gene models. Detailed Implementation

[0020] This invention provides a combination of biomarkers for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on transcription factor characteristics, including ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852 and ZNF765.

[0021] In this embodiment of the invention, a cohort of 1362 patients was collected from literature materials, including 863 non-ETP patients, 168 near-ETP patients, 110 ETP patients, and the remaining patients were excluded from the cohort due to unclear subtyping. The gene expression data from the literature described in this invention can be obtained from Synapse (https: / / doi.org / 10.7303 / syn54032669), accession number syn54032669.

[0022] This invention uses all known human transcription factors (TFs) as a basis for screening. A total of 1639 TFs were downloaded from the Human Transcription Factor Library (https: / / humantfs.ccbr.utoronto.ca / ) and screened based on this. Finally, 9 biomarkers that can be used as prognostic predictors of non-ETP ALL were selected. The relationships between the 9 TFs as biomarkers and their HUGO (HGNC) numbers are ZNF578 (HGNC:26449.), FOXO6 (HGNC:24814.), KCNIP3 (HGNC:15523.), ZNF474 (HGNC:23245.), BPTF (HGNC:3581.), NHLH1 (HGNC:7817.), ZNF112 (HGNC:12892.), ZNF852 (HGNC:27713.), and ZNF765 (HGNC:25092.).

[0023] The present invention also provides the application of a reagent for detecting the above-mentioned combination of biomarkers in the preparation of a kit for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia.

[0024] This invention uses the expression level of the biomarker combination as the independent variable and the prognosis of non-ETP ALL patients as the variable. A corresponding model can be constructed through data analysis and machine learning algorithms, thereby predicting the prognosis of corresponding non-ETP ALL patients based on the expression level of the biomarker combination.

[0025] The present invention also provides a kit for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on transcription factor characteristics, comprising reagents for detecting the transcriptional expression levels of transcription factors contained in the above-mentioned combination of biomarkers.

[0026] In this invention, next-generation sequencing (RNA-seq) data is integrated into the 9-TF risk analysis protocol, and the specific process is as follows: 1. Input and Data Compliance 1.1 Data type: RNA-seq raw sequences (FASTQ, paired or single-ended), or equivalent gene-level expression matrices (TPM recommended).

[0027] 1.2 Metadata: Unique sample ID, library construction strategy (stranded / unstranded), species and reference version (GRCh38), library batch, sequencing platform.

[0028] 1.3 Compliance Requirements: Human samples must comply with ethical approval and data export / privacy compliance; they must be de-identified upon submission (ID mapping is stored by the data controller).

[0029] 1.4 Minimum sequencing depth (recommended): ≥2×10⁻⁶ per sample 7 Successful segment matching; Q30 ≥ 85%.

[0030] 2. Raw sequence preprocessing and quality control 2.1 De-connection and quality trimming: Use fastp or cutadapt to remove connectors and low-quality bases.

[0031] 2.2 Preliminary quality control: Q20 / Q30, GC%, read length distribution, adapter contamination, repetition rate, rRNA content, etc. were summarized using FastQC / MultiQC.

[0032] 2.3 Acceptable thresholds (example): Unique alignment rate ≥ 70% (alignment method) or effective mapping rate ≥ 65% (quasi-alignment method); rRNA ratio < 20%; duplication rate < 60%. Unacceptable samples are marked and enter the review / retest process.

[0033] 3. Quantitative Expression and Annotation 3.1 References and annotations: Fixed reference genome and annotation version (e.g., GENCODE / Ensembl).

[0034] 3.2 Quantitative Procedure (Example of Alignment Method): STAR two-pass alignment (matching strand characteristics with sequencing mode), featureCounts count and TPM calculation.

[0035] 3.3 Gene ID unification: Using Ensembl ID as the primary key and HGNC symbol as the alias; establish a synonym mapping table.

[0036] 3.4 Expression matrix format: Form a “sample × gene” TPM matrix (row = sample, column = gene; UTF-8, tab-separated, first column is sample ID).

[0037] 4. Model gene set and variable notation 4.1 Target gene set (9-TF): ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852, ZNF765.

[0038] 4.2 Variable Definitions and Text Formulas (Corresponding to the Schematic Diagram): Original Expression Level: E(g) (Unit: TPM).

[0039] Normalized Expression Level: s(g) = (E(g) - μ(g)) / σ(g), where μ(g) and σ(g) are fixed as the mean and standard deviation of the training set.

[0040] Linear Predictor (Risk Score Coordinate): LP = Σβ(g)·s(g) (summation over all g).

[0041] Partial Risk Score: PH = e LP .

[0042] 4.3 Design Note: Standardization always uses the μ and σ of the training set to avoid information leakage and maintain comparability across cohorts; β(g) is the fixed coefficient of the final 9-TF model.

[0043] 5. Standardization and Missing Value Handling 5.1 Saving Standardization Parameters: During the training phase, output standardizer_mu_sigma.tsv (columns include GENE, μ, σ, from the training set).

[0044] 5.2 Standardization of New Samples: Only use μ(g) and σ(g) in 5.1 to calculate s(g); do not re-estimate the mean or standard deviation on new data.

[0045] 5.3 Handling Missing Genes: If ≤1 gene is missing among the 9 genes, set the missing gene's s(g) = 0 (equivalent to filling with the training set mean); if ≥2 genes are missing, mark as "not assessable" and enter the exception list.

[0046] 6. Risk Scoring and Threshold Application 6.1 Score Calculation: First calculate LP according to 4.2, then calculate PH = e LP .

[0047] 6.2 Threshold Source: The globally frozen threshold file output / final_thresholds.json during the training phase, which contains the field "LP_threshold" (the median threshold of LP).

[0048] 6.3 Grouping Rule: High = (LP ≥ LP threshold); Low = (LP < LP threshold). The LP threshold described in this invention is 0.000113.

[0049] 6.4 Consistency: LP and PH are only monotonic transformations, and the grouping should be consistent; if inconsistent, an audit needs to be triggered.

[0050] 7. Interface and File Specifications 7.1 Input files: gexp.tsv (sample × gene TPM; first column ID); annot.tsv (containing OS and OS.status for auditing if necessary); standardizer_mu_sigma.tsv (can be placed in output / ); final_9gene_coxph_params.tsv (coefficient β(g), or directly embedded in the implementation); final_thresholds.json (fixed thresholds generated during training).

[0051] 7.2 Output files: risk_groups_global.tsv / .json (ID, LP, PH, risk_group, threshold, method, model_version, ref_annot_version); and audit logs (running parameters, threshold source files, missing gene list, abnormal sample list).

[0052] 8. Information Leakage Control and Methods School Verification 8.1 Key points for preventing leakage: μ, σ, β, and threshold are all derived from the training set and frozen during the application phase; for new samples, only s(g), LP, PH, and risk_group are calculated, and no re-estimation or parameter tuning is performed on new data.

[0053] 8.2 Consistency check: For the same batch of data, the grouping based on LP and the grouping based on PH should be completely consistent; otherwise, record and verify.

[0054] 8.3 Robustness handling: For extreme values ​​(such as TPM exceeding the upper tail quantile), winsorization (e.g., 99.5%) can be performed before standardization and recorded in the log.

[0055] 8.4 External Validation: Verify the reproducibility of KM, HR, and C-index in a separate queue using the same standardizer_mu_sigma.tsv and final_thresholds.json.

[0056] 9. Software and Implementation Environment (Textual Description) 9.1 Implementation Recommendations: In the Python environment, use pandas for table reading and writing and index alignment, and numpy for matrix and type conversion; scikit-learn for standardization and cross-validation stratification; lifelines for Cox / LASSO-Cox model training and coefficient freezing; joblib for parallel grid evaluation; tqdm / psutil for progress and resource monitoring; and pyarrow to improve columnar I / O (version and key parameters are fixed in the example).

[0057] 10. Method Claims Skeleton (Text Formula) 10.1 Clean reads were obtained from the subject's RNA-seq data after adapter removal and quality trimming; 10.2 Quantitative expression was performed according to the preset reference annotations to obtain the TPM matrix unified with the Ensembl stable ID; 10.3 Extract the expression of ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852, and ZNF765; 10.4 The expression of gene 9 was normalized using μ and σ determined by the training set to obtain s(g); 10.5 Calculate the linear predictor based on the fixed coefficient β(g): LP = Σβ(g)·s(g); 10.6 Calculate the partial risk score: PH=e LP ; 10.7 Compare LP and / or PH with a predetermined and fixed threshold, and output a high-risk or low-risk judgment result; In step 10.4 and step 10.7, μ, σ, and threshold are determined by the training set and remain unchanged during application.

[0058] The present invention also provides a method for constructing a scoring model for predicting the prognosis of non-early precursor T acute lymphoblastic leukemia based on transcription factor characteristics, comprising the following steps: (1) collecting gene expression data of non-early precursor T acute lymphoblastic leukemia and randomly dividing it into n folds, where n is a positive integer ≥2, with one part serving as the training set and the other part serving as the validation set; (2) Single-gene Cox and LASSO analyses were used on the training set to obtain initial screening transcription factors; (3) The obtained initial screening transcription factors are processed by several machine learning algorithms, the number of gene hits in the training set of each fold is counted, and 9 transcription factors with a hit count ≥10 are collected to construct the prognostic scoring model for the non-early precursor T acute lymphoblastic leukemia; the 9 transcription factors include ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852 and ZNF765.

[0059] The patient cohorts and gene expression data used in this embodiment of the invention were obtained from the literature. The gene expression data was obtained from Synapse (https: / / doi.org / 10.7303 / syn54032669), accession number syn54032669. A total of 1362 patients were collected, including 863 non-ETP patients, 168 near-ETP patients, 110 ETP patients, and the remaining patients were excluded from the cohort due to unclear subtypes.

[0060] This invention uses TFs as a feature. In the embodiments, all TFs in the currently known transcription factor library are used as the basis for screening. For example, the list information of all 1639 human TFs was downloaded from the Human Transcription Factor Library (https: / / humantfs.ccbr.utoronto.ca / ).

[0061] To avoid validation set leakage and ensure the accuracy of model evaluation, this invention strictly adheres to a 4×4 fold nested cross-validation design, specifically including outer and inner cross-validation. The outer cross-validation involves randomly dividing all non-ETP ALL patients into four folds, with three folds used as the training set and the remaining fold as the validation set; this is repeated four times, changing the validation set each time to ensure each patient appears in only one validation set. The inner cross-validation involves performing a 4-fold internal cross-validation within each outer training set for hyperparameter selection and feature screening, thereby strictly preventing validation set contamination. In one embodiment, the data partitioning is implemented using Python to construct and manage queues, with a random seed set to 42 to ensure the randomness and repeatability of data partitioning; the expression matrix and clinical data are stored separately to ensure a clear and reproducible analysis process.

[0062] This invention utilizes single-gene Cox analysis to screen for their prognostic relevance in non-ETP ALL patients, ultimately identifying key transfer factors (TFs) with p < 0.05 in the 4-fold training set. Subsequently, the LASSO-Cox algorithm is used to initially reduce the TF model, reducing the number of TFs to less than 100. Based on the data output by LASSO, especially the inflection points of the C-index and selected genes, preliminary models for each fold training set are obtained.

[0063] Considering the requirements of the events per variable (EPV) criterion and the discussions on model size in recent studies, this invention compresses the model to below 10 based on the total number of deaths in the patient cohort (88 patients). Therefore, this invention also selects machine learning algorithms for model simplification, namely: Support Vector Machine Recursive Feature Elimination (SVM-RFE), Random Forest (RF), Random Survival Forest (RSF), Stability Selection, Extreme Gradient Boosting (XGBoost), Adaptive Boosting (AdaBoost), and Generalized Additive Model (GAM). Subsequently, these seven different machine learning methods are used to evaluate each initial screening list to further simplify the data model of this invention. Furthermore, this invention calculates the number of times each TF passes the machine learning tests; for example, if a TF passes all six machine learning tests simultaneously, it is recorded as 6 times; if it passes only three machine learning tests, it is recorded as 3 times. For each fold of the training set, the number of gene hits was counted, and the top 20 genes were ranked according to the number of hits. Finally, a 9-TF model was constructed from the TFs with ≥10 hits, including ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852, and ZNF765. The 9-TF model described in this invention performed well on the 4-fold training set, and the overall survival (OS) of high-scoring patients was significantly reduced.

[0064] The formula for the scoring model described in this invention is: ; Where z gene This represents the expression level of a gene in the sample after standardization by the mean and standard deviation of the training set.

[0065] The present invention also provides a system for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia, including a scoring model for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on the combined characteristics of the above-mentioned biomarkers and constructed using the above-mentioned construction method.

[0066] To further illustrate the present invention, the following detailed description, in conjunction with embodiments, of a non-ETP ALL scoring model based on transcription factor characteristics and its construction method provided by the present invention, should not be construed as limiting the scope of protection of the present invention.

[0067] Example 1 1. Materials and Software Program Description The patient cohorts and gene expression data used in this embodiment of the invention were obtained from the literature. The gene expression data was obtained from Synapse (https: / / doi.org / 10.7303 / syn54032669), accession number syn54032669. The entire cohort included 1362 patients, of which 863 were non-ETP patients, 168 were near-ETP patients, 110 were ETP patients, and the remaining patients were removed from the cohort due to unclear subtypes.

[0068] This invention involves all 1639 human transcription factors (TFs), whose information was downloaded from the Human Transcription Factor Library (https: / / humantfs.ccbr.utoronto.ca / ).

[0069] To avoid validation set leakage and ensure the accuracy of model evaluation, the model construction in this invention strictly follows a 4×4 fold nested cross-validation design.

[0070] This process was implemented in a Python 3.13 environment, with the following dependencies and versions: pandas 2.2.2 (table reading, writing, joining / filtering / transposing), numpy 1.26.4 (numerical matrices and type conversion), and scikit-learn 1.4.2 (feature standardization and cross-validation). First, pandas was used to read the clinical annotations and typing tables, generating typing fields according to the rule of "retaining only a single positive typing," and establishing aligned clinical data based on sample IDs. Then, pandas was used to read the expression matrix, filtering out Non-ETP individuals based on typing, and performing a strict inner join between the sample IDs and the clinical table to ensure a one-to-one correspondence. Subsequently, scikit-learn was used to perform Z-score standardization of column vectors within each fold. Four-fold cross-validation was constructed using scikit-learn on all Non-ETP samples, randomly dividing them into four parts with a fixed random seed of 42. Training / validation sample IDs were provided for each fold, thus creating corresponding "standardized expression matrix subsets" and "clinical subsets." Finally, four sets of training / validation data and their standardized features were formed, which can be directly used for downstream modeling and validation, facilitating reproducibility and auditing.

[0071] The specific process is as follows: Outer cross-validation: All non-ETP ALL patients are randomly divided into 4 folds, with 3 folds used as the training set and the remaining 1 fold as the validation set; this is repeated 4 times, with the validation set changed each time to ensure that each patient appears in the validation set only once. Inner cross-validation: Each outer training set is then subjected to 4-fold inner cross-validation for hyperparameter selection and feature screening, thereby strictly preventing validation set contamination.

[0072] Python was used to construct and manage the queues, ensuring the randomness and reproducibility of data partitioning. Expression matrices and clinical data were stored separately to ensure the analysis process was clear and reproducible.

[0073] The univariate Cox proportional hazards model is a statistical tool used in survival analysis to study the relationship between the timing of a specific event and factors (covariates).

[0074] This process was implemented in a Python 3.13 environment, and depends on pandas 2.2.2 (table reading and writing, index alignment, missing value handling), numpy 1.26.4 (numerical arrays and type conversion), and lifelines 0.28.0 (Cox proportional hazards model). The specific steps are as follows: First, load the data from the "training expression matrix and corresponding clinical table (training set only)" generated in the previous stage and the provided transcription factor gene list, and complete the configuration of the running parameters (significance threshold α=0.05); then process each fold: for the current fold, find the intersection of the training expression and clinical data by sample ID to achieve strict alignment, retain only the Non-ETP training samples and report the sample size and number of features; based on the transcription factor list, perform column name intersection filtering on the features, indicating the number of missing genes and retaining the effective gene subset; perform univariate Cox regression on the transcription factor set of this fold: construct a columnar design matrix for each gene containing survival time, outcome events, and the expression of that gene, delete samples with missing values ​​and skip genes with "no variance in expression" in this fold; use lifelines.CoxPHFitter, with parameters penalizer=0.0, l 1_ratio=0.0 (i.e., no regularization, pure univariate test), the rest remain at default (proportional risk hypothesis test follows lifelines' default behavior, confidence interval is 95%); extract and summarize statistics such as coef, HR=exp(coef), se, z, p, CI 95% upper and lower limits for successfully converged genes to generate a complete result table for that fold, and extract significant genes according to p<0.05 and sort them in ascending order of p value to form a significant list; if a fold has no valid results, record it and skip the subsequent screening and export; the entire process only uses the training set to complete the initial feature screening, never touching the validation set to avoid information leakage; after processing four folds, obtain "single gene Cox statistical summary for each fold" and "significant gene list for each fold (p<0.05)", and output the total time and result directory location, as the candidate feature base for subsequent feature shrinkage (such as LASSO) and modeling.

[0075] This invention uses a total of 7 machine learning algorithms, namely Support Vector Machine Recursive Feature Elimination (SVM-RFE), Random Forest (RF), Random Survival Forest (RSF), Stability Selection, Extreme Gradient Boosting (XGBoost), Adaptive Boosting (AdaBoost), and Generalized Additive Model (GAM).

[0076] I. Operating Environment and General Configuration Software environment: Implemented under Python 3.13.

[0077] Dependencies and their uses: (a) pandas 2.2.2: Reading and writing table data, index alignment, handling missing / constant columns, batch loading of training data for each fold and frequency statistics; (b) NumPy 1.26.4: Matrix Operations and Type Conversion; (c) scikit-learn 1.4.2: outer 4-fold partitioning, feature standardization, several classification / feature selection models and metrics; (d) lifelines 0.28.0: Cox proportional hazards model fit with L1 regularization (LASSO-Cox), parameter l1_ratio=1.0; Randomness: The random seed is fixed at 42 throughout the entire process; II. Sample In / Out Criteria and Data Alignment Clinical annotation and classification reading: Read the complete clinical table and classification information, retain only cases with "single ETP subtype" and select Non-ETP individuals as subjects for subsequent analysis.

[0078] Expression matrix alignment: The full transcriptome expression matrix was read and aligned with the clinical table by sample ID to obtain an analysis matrix of "sample × gene"; then Z-score normalization was performed on each column.

[0079] III. Outer 4-fold Survival Modeling Process Based on LASSO-Cox (Initial Feature Screening) Objective: Under strict information isolation, this study utilizes an L1-regularized Cox model within an outer 4-fold framework to select penalty strength and acquire sparse features, producing candidate feature sets and optimal penalty parameters for each fold. Outer fold division: An outer 4-fold scheme is adopted.

[0080] Candidate feature determination: Within each outer fold, only the training samples of that fold are used, and the intersection with the previously obtained "significant gene list for that fold" is taken to limit candidate features; unmatched genes are removed. Survival outcome readiness: Event columns and follow-up time columns are extracted from the clinical table and mapped one-to-one with the training feature matrix of the current fold.

[0081] Inner layer parameter tuning: Under 5-fold cross-validation, the optimal penalizer is searched for on 200 penalty intensities of a log-uniform grid, and the average validation C-index is used as the criterion for selection.

[0082] Model refitting and sparsity control: Refit LASSO-Cox on all training data in this fold using the optimal penalizer. If the number of non-zero coefficient features obtained is >100, the penalty strength is increased proportionally (e.g., ×1.1) and the fitting is repeated until the number of non-zero features is ≤100 or a preset upper bound is reached.

[0083] Output and Recording: For each fold, output (i) the selected feature set, (ii) the corresponding optimal penalty intensity, and record (iii) the penalty-C-index tuning curve.

[0084] Information leakage control: Each layer's validation set is used only for inner-layer performance evaluation and penalty selection, and does not participate in any form of variable screening or threshold setting; outer-layer validation data is not used for feature discovery or parameter determination within this layer, thus eliminating information leakage from the process design.

[0085] IV. Robustness Verification and Consensus Feature Generation Based on Seven Methods (Multi-Algorithm-Based Re-screening) Objective: To improve the robustness of results and form a consensus feature across folds and methods by independently verifying the importance of features using multiple complementary algorithms on top of the initial LASSO-Cox screening.

[0086] Input ready: Load the four-fold training expression / clinical data and the list of candidate features screened by LASSO for each fold at once; perform column screening for each fold feature according to the list and then run StandardScaler again.

[0087] Algorithm set and parameters (running independently fold by fold): Algorithm set and parameters: In each fold, only the training subset of that fold is used to independently complete feature ranking and model selection. First, SVM-RFE is used as the base learner. Ten candidate points are selected at equal intervals within the logarithmic scale range of the penalty coefficient C (1×10⁻³~1×10¹). The ROC-AUC of the training set is used for evaluation, and recursive feature elimination is used to rank the outputs. Then, a random forest is trained, using a splitting strategy of bootstrap sampling and sampling by the square root of the feature number. The number of trees is increased from 50 to 3000, and expansion is based on the continuous improvement of the out-of-bag (OOB) AUC. Training stops early if there is no improvement after five consecutive iterations. For AdaBoost, a small grid search with three-fold cross-validation is used. The optimal combination of iterations {200, 500, 1000} and learning rates {0.05, 0.1, 0.2} is selected based on ROC-AUC. When using the gradient boosting tree classifier of XGBoost 2.0.3, the learning rate is set to 0.01, and the number of iterations is 500. The system employs a tree model, and after training, extracts feature importance based on gain and / or split weights. For the survival task, a Random Survival Forest (RSF) is used, scored using the C-index. For stability selection, a logistic regression model with L1 regularization is used as the base learner, with 200 subsampling iterations (randomly selecting approximately 70% of individuals each time). The frequency of "non-zero coefficients" is accumulated and ranked accordingly. Finally, a generalized additive model (GAM) is used to characterize potential nonlinear effects, and after training, importance is measured and ranked by the sum of the absolute values ​​of the coefficients of each feature spline. Standardization and parameter tuning are performed within each iteration, and validation / test subsets are used only for evaluation to avoid information leakage. The importance and ranking of the outputs from each iteration are then robustly aggregated between iterations.

[0088] Unified output specifications: Rank the best output and the corresponding score for each of the above 7 methods in each fold; if a method fails to fit in a certain fold, record it as empty and retain the alarm information.

[0089] Global frequency statistics and consensus threshold: Collect feature names of all folds and all methods and perform frequency statistics; set a frequency threshold ≥3 (i.e., appear at least 3 times in 28 results) to identify "consensus features" as the final refined set.

[0090] Randomness control: The random number seed related to modeling is fixed at 0; the random seed related to resolution consistency is fixed at 42.

[0091] Main output: (a) The selected feature set and optimal penalty intensity of the outer 4-fold LASSO-Cox; (b) Top-20 features and scores of the seven methods in each fold; (c) Global frequency statistics table and consensus feature list of “frequency ≥ 3”; (d) Performance curves and resource consumption records corresponding to the above.

[0092] VI. Key Points for Information Leakage Prevention and Reproducibility Strict fold isolation: Any feature selection, threshold setting, and parameter selection are based solely on the training subset of that fold and its inner validation layer; outer validation samples do not participate in selection or parameter tuning.

[0093] Fixed decision-making order: first sample input, output and alignment → standardization → candidate set merging → inner layer parameter tuning → outer layer refitting → output and recording, avoiding cross-contamination of data and parameters.

[0094] Randomness is traceable: random seeds (42 and 0) are fixed in each of the steps of division, modeling and permutation, and written in the log.

[0095] Irreversible data transformation: Standardized parameters (mean, standard deviation) are strictly fitted to the training set and applied to the corresponding validation / test data, realizing a one-way flow of training to application.

[0096] 2. Construction of the 9-TF prognostic model First, training and validation sets were constructed using a 4×4 nested cross-validation method, and no content from the validation sets was used until the model was completed. Then, information on 1639 TFs was downloaded from the human TF database, and their prognostic relevance in non-ETP ALL patients was screened using single-gene Cox analysis. Finally, key TFs with p<0.05 were selected from the 4-fold training set.

[0097] Subsequently, LASSO-Cox was used to perform the first contraction of the transcription factor set, controlling the number of candidate TFs to within 100. Based on the penalty path of LASSO and the performance of in-fold validation, especially the inflection point of the curve between the C-index and the number of selected genes, a preliminary model was determined for each fold. The specific process was as follows: First, the complete clinical annotations and typing were read, retaining only the single ETP typing and screening non-ETP individuals that met the criteria; then, the full expression matrix was read, and after inner-joining with the clinical table by sample ID, Z-score normalization was performed on each column of "sample × gene". Then, an outer four-fold cross-validation was established and run independently for each fold: in each fold, only the intersection of the training samples of that fold and the list of significant genes of that fold was used to determine candidate features, and survival subjects were generated based on the clinical table. Five-fold cross-validation is used within each fold. The L1 penalty strength (ranging from 0.001 to 0.05) is searched for among 200 candidate points on a logarithmic scale, and the optimal value is selected based on the average validation C-index. This penalty is then used to refit the LASSO-Cox model on all training data for that fold, resulting in a set of features with non-zero coefficients. If the number of selected features exceeds 100, the penalty is gradually increased by a ratio of 1.1, and the fitting is repeated until the upper limit is met or the safety boundary is reached. Each fold outputs the selected features and their corresponding penalty strength, and records key points of the penalty-C-index curve. A fixed random seed of 42 is maintained throughout the process, and validation / test samples are not allowed to participate in the selection or parameter tuning in any way, thus ensuring no information leakage and providing a reproducible set of candidate features for subsequent multi-model refinement.

[0098] Considering the requirements of the events per variable (EPV) criterion and the discussions on model size in recent studies, based on the total number of deaths (88) in the patient cohort according to this invention, this invention believes it is best to compress the model to below 10. Given the complexity of high-dimensional gene expression data, recent studies have shown that integrating multiple machine learning algorithms can significantly improve the robustness and accuracy of feature selection for various diseases, including tumors. This invention selects seven machine learning algorithms for model simplification: SVM-RFE, RF, RSF, AdaBoost, XGBoost, GAM, and stability selection. Subsequently, these seven different machine learning methods are used to evaluate each of the aforementioned initial screening lists to further refine the preliminary model. Furthermore, the number of times each TF passes the machine learning tests is calculated; for example, if a TF passes all six machine learning tests, it is counted as 6 times; if it passes only three machine learning tests, it is counted as 3 times. Gene hit counts are performed on the training set for each fold, and the top 20 genes are ranked according to their hit counts. Finally, a 9-TF model was constructed using TFs with ≥10 hits, including ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852, and ZNF765. The 9-TF model described in this invention performed well on the 4-fold training set, and the overall survival (OS) of high-scoring patients was significantly reduced. Figure 1 ).

[0099] Calculation formula: Standardized definition (Z-score) (make ) Linear predictor (LP) Danger function (Cox model) Risk ratio (relative to baseline hazard) Single gene hazard ratio and 95% confidence interval Interpretation of coefficient direction (Elevated expression increases risk); (Elevated expression indicates reduced risk) Final model (9-TF) linear predictor Furthermore, in making specific judgments, a final value ≥ 0.000113 is considered high risk; while a value < 0.000113 is considered low risk.

[0100] 3. Model prediction capability verification 3.1 Selection of Analysis Software Package Clinical and genome-wide expression data of non-ETP ALL patients were integrated, and a Cox regression model was constructed using Lifelines, which was used to calculate the risk score and model consistency index. Gene expression was standardized using scikit-learn, and the false discovery rate (FDR) was calculated using statsmodels.

[0101] 3.2 Implementation Steps This step aims to evaluate the prognostic predictive power distribution of randomized gene models. Based on non-ETP type T-ALL patients, different numbers of gene sets N were randomly selected to form "randomized models." Cox survival regression was constructed, and the results were compared with a pre-defined TF gene model. The C-index (consistency index) of this randomized gene model was recorded. The randomization and model evaluation process was repeated 1000 times for each size N to obtain a large number of randomized model results. Finally, detailed results from each simulation were saved as a tabular file, and the performance of the randomized models under different gene set sizes was statistically summarized.

[0102] 3.3 Parameter Settings: The main parameters involved in the randomized model simulation include: 1000 simulations for each gene set size to ensure statistical representativeness; and a list of randomized gene set sizes covering multiple values ​​from small to large to facilitate direct comparison with the corresponding TF model. Sufficient random sampling and model construction provide a reference for evaluating the significance of the TF model.

[0103] This invention validates the predictive performance of the 9-TF model in a 4-fold validation set. The results show that the 9-TF model performs well in all validation sets, with a significant reduction in overall survival (OS) for high-scoring patients. Figure 2 Therefore, the members of this 9-TF were validated and integrated, and all non-ETP ALL patients were integrated for final parameter calculation.

[0104] 4. Comparison of win rates between TF model and random model: 4.1 Selection of analysis software package: The simulation results were processed using Pandas, simple calculations were performed using NumPy, and statsmodels was used to help determine the significance correction value of the baseline model.

[0105] 4.2 Implementation steps: The predictive power of the established TF model was compared with that of a randomized genetic model to calculate the win rate of the TF model relative to the randomized model. First, the p-values ​​obtained from the univariate Cox regression of the baseline model were extracted from previous analysis results. Next, for each simulation result of the randomized model, it was compared with the baseline TF model one by one.

[0106] The comparison criterion is set as follows: whether the C-index of the randomized model is greater than the p-value of the baseline model. In practice, this invention iterates through each simulation, accumulates the number of times the baseline model wins, and calculates the win rate. Win rate = number of times the baseline model wins / number of effective simulations. This ratio reflects the frequency with which the baseline TF model outperforms the randomized model when constructing a model with the same number of randomly selected genes. Finally, the win rate results of the TF model under different gene set sizes are saved, and the summary statistics are output.

[0107] 4.3 Parameter Settings: In the win rate comparison process, the C-index of the baseline model is taken from the actual calculation results of its respective model; the performance index of the random model is obtained successively from the simulation results. This invention performs 1000 random simulations to calculate the win rate, thereby ensuring the stability of the calculated win rate. The win rate results are used to illustrate the extent to which a model of the same scale obtained by randomly selecting genes can rival a specific TF model, indirectly verifying the significance and reliability of the selected TF model.

[0108] The results showed that the 9-TF model maintained an extremely high win rate (>99%) at the same magnitude. Figure 3 Furthermore, when compared with a randomized model of 20 genes, the 9-TF model still had a win rate of over 70%, indicating that the 9-TF model has high reliability.

[0109] Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, and not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.

Claims

1. A combination of biomarkers for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on transcription factor characteristics, characterized in that, Including ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852 and ZNF765.

2. The use of a reagent for detecting the combination of biomarkers of claim 1 in the preparation of a kit for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia.

3. A kit for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia based on transcription factor characteristics, characterized in that, Including reagents for detecting the transcriptional expression levels of transcription factors contained in the biomarker combination of claim 1.

4. A method for constructing a scoring model for prognostic prediction of non-early precursor T-cell acute lymphoblastic leukemia based on transcription factor characteristics, characterized in that, Includes the following steps: (1) Collect gene expression data of patients with non-early precursor T acute lymphoblastic leukemia and randomly divide them into n folds, where n is a positive integer ≥2. One part is used as the training set and the other part is used as the validation set. (2) Cox and LASSO analyses were used on genes in the training set to obtain initial screening transcription factors; (3) The obtained initial screening transcription factors are processed by several machine learning algorithms, the number of gene hits in the training set of each fold is counted, and 9 transcription factors with a hit count ≥10 are collected to construct the prognostic scoring model for the non-early precursor T acute lymphoblastic leukemia; the 9 transcription factors include ZNF578, FOXO6, KCNIP3, ZNF474, BPTF, NHLH1, ZNF112, ZNF852 and ZNF765.

5. The construction method according to claim 4, characterized in that, The gene expression data in step (1) includes the expression information of 1639 transcription factors downloaded from the human transcription factor library.

6. The construction method according to claim 4 or 5, characterized in that, In step (1), the collected gene expression data is randomly divided into 4 folds, of which 3 folds are selected as the training set and the remaining 1 fold is used as the validation set. Each fold serves as a training set, forming a 4×4 fold cross-nested validation queue.

7. The construction method according to claim 4, characterized in that, The initial screening in step (2) includes screening for prognostic relevance in patients with non-early precursor T acute lymphoblastic leukemia and screening out key transcription factors with p<0.

05.

8. The construction method according to claim 4, characterized in that, The machine learning algorithms described in step (3) include: support vector machine recursive feature elimination, random forest, random survival forest, stability selection, extreme gradient boosting, adaptive enhancement, and generalized additive model.

9. The construction method according to claim 4, characterized in that, The formula for the scoring model in step (3) is: ; Where z gene This represents the expression level of a gene in the sample after standardization by the mean and standard deviation of the training set.

10. A system for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia, characterized in that, This includes a scoring model for predicting the prognosis of non-early precursor T-cell acute lymphoblastic leukemia, constructed using the construction method described in any one of claims 4 to 9, based on the combination characteristics of the biomarkers described in claim 1.