Construction method and equipment of predictive cell aging model, medium and program product

By constructing the PreCSenM model and utilizing multi-gene combinations and machine learning algorithms, the accuracy problem of cell senescence detection in existing technologies has been solved, enabling precise assessment and therapeutic target discovery in cancer research and improving treatment outcomes.

CN120853690APending Publication Date: 2025-10-28INSTITUTE OF BASIC MEDICAL SCIENCES CHINESE ACADEMY OF MEDICAL SCIENCES

Patent Information

Application Number
CN202510981811.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

The lack of accurate cell senescence detection markers in existing technologies makes it difficult to accurately characterize senescence levels in cancer research, affecting treatment efficacy and specificity.

Method used

A predictive cell senescence model, PreCSenM, was constructed. By integrating multiple aging characteristic gene sets and gene scoring algorithms, key aging gene sets were identified, and a cell senescence scoring system was constructed using a machine learning model to achieve accurate prediction of the aging status of tissue samples and screening of potential therapeutic drugs.

Benefits of technology

Robust aging prediction was achieved under different data types and cell types, the correlation between cellular aging levels and clinical characteristics was revealed, potential therapeutic targets such as AP-1 regulatory factors were discovered, and the accuracy and effectiveness of cancer treatment were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853690A_ABST
    Figure CN120853690A_ABST
Patent Text Reader

Abstract

The invention provides a construction method of a predictive cell senescence model, a method for predicting the senescence state of a tissue sample based on the senescence model, a method for screening potential therapeutic drugs, equipment, a medium and a program product, and relates to the field of intelligent medical treatment. The model construction method comprises the following steps: acquiring a training set sample expression profile data set; identifying a key senescence gene set from the data set by using a feature selection algorithm; inputting the key senescence gene set into a machine learning model to fit a prediction model, and determining an optimal hyper-parameter to obtain a cell senescence model containing the weight of a single gene in the key senescence gene set; the cell senescence model is a senescence score obtained by calculating the sum of the product of the expression quantity of a single gene and the regression coefficient thereof. The cell senescence model, namely PreCSenM, is constructed by integrating a plurality of senescence characteristic gene sets and a gene scoring algorithm, the accuracy in CS evaluation is superior to that of 10 existing methods, and the application of CS from biological research to clinical scenes is also realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent healthcare, and more specifically, to a method, apparatus, medium, and program product for constructing a predictive cellular senescence model. Background Technology

[0002] Cellular senescence (CS), characterized by irreversible cell cycle arrest, involves various biological processes, development, and age-related diseases, accompanied by macromolecular changes and supersecretion and pro-inflammatory phenotypes. Senescent cells play diverse roles in physiological and pathological processes. Recent research has deepened our understanding of the protective function of senescent cells against cancer. In particular, CS provides a natural barrier to tumorigenesis by interfering with the epigenome to inhibit cell division and growth. This process can eliminate damaged cells or potential cancer cells and stimulate the immune system to attack cancer cells, thus exerting a tumor-suppressive effect. However, the accumulation and retention of senescent cells can have detrimental effects, contributing to aging and age-related diseases, including cancer. Senescent cells secrete cytokines in a paracrine or endocrine manner, potentially inducing the formation of a chronic inflammatory microenvironment, thereby promoting tumorigenesis. Therefore, CS has potential beneficial effects on cancer treatment and is crucial for improving treatment efficacy and specificity.

[0003] However, the lack of universal biomarkers for detecting senescent cells hinders the accurate characterization of aging levels. Currently, researchers are attempting to move beyond single senescence-associated β-galactosidase (SA-β-gal) activity assays by combining multiple traditional biomarkers to define the degree of senescence in normal cells. Single classic biomarkers (CDKN1A and CDKN2A) and five publicly available sets of senescence-specific genes (SenMayo, Casella, Fridman, Purcell, and Hernandez) have been reported. Cancer cells often carry mutations in senescence-related genes such as CDKN1A and CDKN2A, but these biomarkers are identified in normal cells; cancer, as a complex disease, is characterized by high inter-tumor and intra-tumor heterogeneity. Therefore, conventional senescence biomarkers based on non-cancer cells may lack accuracy or applicability in cancer research. Ideally, biomarkers should maintain stable expression within tumor tissues and across different tumor types to ensure the reliability of patient test results. Against this backdrop, employing a multi-gene combination detection strategy may become an effective way to analyze the aging state in malignant phenotypes. Summary of the Invention

[0004] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention provides a method for constructing a predictive cellular senescence model, as well as a method for predicting the senescence status of tissue samples based on this senescence model and a method for screening potential therapeutic drugs. The method of this invention constructs a cellular senescence model, PreCSenM, by integrating multiple senescence characteristic gene sets and gene scoring algorithms. Its accuracy in CS assessment surpasses that of 10 existing methods, and it also enables the application of CS from biological research to clinical scenarios, which is impossible with other methods.

[0005] The first aspect of this application discloses a method for constructing a predictive cellular senescence model, the method comprising:

[0006] S101, Obtain the training set sample expression profile dataset, which includes normal samples from senescent and non-senescent cell lines;

[0007] S102, using the feature selection algorithm (Boruta algorithm) to identify the set of key aging genes from the training set sample expression profile dataset;

[0008] S103, the key aging gene set is input into the machine learning model to fit the prediction model, and the optimal hyperparameters are determined based on the LOOCV framework to obtain a cell aging model containing the weights of individual genes in the key aging gene set; the cell aging model calculates the aging score by the sum of the products of the expression level of a single gene and its regression coefficient, and the training set samples are divided into high group and low group according to the threshold set by the aging score.

[0009] A second aspect of this application discloses a method for predicting the aging state of a tissue sample, the method comprising:

[0010] S201, to obtain the expression levels of key senescence gene sets in normal or tumor tissue samples;

[0011] S202, input the expression level into the cell senescence model constructed by the method described in the first aspect of this application, calculate the senescence score and divide the tissue sample into a high group or a low group according to the senescence score; if it is a high group, output an auxiliary prediction result that the tissue sample is in a senescent state with a high probability; if it is a low group, output an auxiliary prediction result that the tissue sample is in a senescent state with a low probability.

[0012] A third aspect of this application discloses a method for screening potential therapeutic drugs, the method comprising:

[0013] S301, Obtain the genomic dataset of tumor samples containing survival data and the drug to be tested;

[0014] S302, input the dataset into the cell senescence model constructed by the method described in the first aspect of this application and divide the samples into high-group or low-group;

[0015] S303, compare the high and low groups to screen out the first differentially expressed gene; evaluate the correlation between the first differentially expressed gene and survival time to screen out the second differentially expressed gene; optimize the second differentially expressed gene to obtain the target gene with a positive variable importance value;

[0016] S304. The feature matching analysis method is used to process the test drugs and target genes, and the feature matching score is calculated; the preliminary ranking results of the test drugs are obtained based on the feature matching score.

[0017] S305, Rank clustering analysis based on order statistics is used to analyze the preliminary ranking results and obtain the final ranking results;

[0018] S306, potential therapeutic drugs are obtained based on the final sorting results.

[0019] In some embodiments, the feature matching analysis method includes any one or more of the following: KS test, extreme value summation, and reverse gene expression scoring.

[0020] A fourth aspect of this application discloses a computer device, the device comprising: a memory and a processor; the memory being used to store a computer program; and the processor executing the computer program to implement the steps of the above-described method.

[0021] The fifth aspect of this application discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0022] The sixth aspect of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0023] This application has the following beneficial effects:

[0024] 1. This application innovatively proposes the "cellular senescence score," a novel quantitative indicator. By screening a key senescence gene set through a training set, it discovers that each gene has a different weight, thereby constructing the PreCSenM model to provide a more robust tool for predicting senescence. PreCSenM is currently the most robust and accurate senescence prediction method, outperforming methods including SenCID (…). Figure 8 All published tools, including SenCID in d, include methods that use GSVA to calculate the activities of two gene sets (positive and negative CS genes) in a single sample, with the CS score defined as the difference between positive and negative CS-related activities. Figure 8The reason for this difference is that the GSVAScore method, which treats positive and negative CS genes as two independent individuals to directly calculate differences, contains redundant genes or treats all genes as equally important. In contrast, this application treats each gene as an independent individual, resulting in higher accuracy.

[0025] Furthermore, the predictive performance of PreCSenM is not affected by cell type (cancer cells vs. non-cancer cells), data type (microarrays vs. RNA sequencing), or senescence-induced mode (replication vs. stress-induced).

[0026] 2. This application further confirms the clinical translational value of PreCSenM in cancer, further transforming CS assessment into a clinically meaningful application. This application systematically analyzes the relationship between CS scores and clinical characteristics (…). Figure 9 The study explored associations between CS scores and survival outcomes across transcriptomics, single-cell transcriptomics, genomics, proteomics, and epigenetics. Notably, we observed a strong correlation between CS scores and survival outcomes in 1,837 patients across 10 datasets, while no significant association was observed with the other 10 aging prediction tools. Figure 9 These data suggest that the CS score holds promise as a reliable predictive biomarker for clinical prognosis.

[0027] 3. This application also proposes a data-driven method for screening aging-related drugs and experimentally validates the roles of HDACi and the core transcription factor FOSB in pro-aging. To our knowledge, this is the first comprehensive framework that systematically utilizes cellular senescence to drive biological discovery.

[0028] In summary, this application constructs a predictive cellular senescence model (PreCSenM). This model integrates multiple machine learning algorithms to perform multi-platform transcriptome data analysis on 230 cell lines from 11 different senescent tissue types, constructing a consistent cellular senescence-related gene signature (CSGS). (CSGS refers to a specific set of genes whose expression patterns are closely related to the cellular senescence process. Cellular senescence is a cellular state in which cells cease division and lose their normal function; this phenomenon is associated with various age-related diseases and the overall aging process of organisms.) This establishes a quantitative system for calculating CS levels. Based on this model, the correlation between cellular senescence levels, molecular characteristics, and clinical prognosis is revealed, and the effect of small histone deacetylase inhibitors (HDACis) on CS levels in lung adenocarcinoma (LUAD) is further investigated. Through transcriptome and epigenomic analysis of HDACis-treated cells, activator protein 1 was identified. 1. The AP-1 family member FOSB can regulate cancer sclerotherapy (CS). The integrated analysis method established in this study provides a machine learning framework for accurately assessing cancer CS levels, and also suggests that AP-1 regulators may serve as potential therapeutic targets. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the method flow provided in the first aspect of the present invention;

[0031] Figure 2 This is a schematic diagram of the method flow provided in the second aspect of the present invention;

[0032] Figure 3 This is a schematic diagram of the method flow provided in the third aspect of the present invention;

[0033] Figure 4 This is a schematic diagram of a computer device provided in an embodiment of the present invention;

[0034] Figure 5This is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present invention;

[0035] Figure 6 This is a schematic diagram of the storage medium provided in an embodiment of the present invention;

[0036] Figure 7 This is an overview of the development and application of the PreCSenM model provided in the embodiments of the present invention; Figure 7 'a' is a schematic diagram illustrating the development method of the PreCSenM computational model and its application in assessing CS levels and LUAD research. Figure 7 b is a two-step workflow: (1) Experimental data organization and preprocessing: including screening and organizing 888 publicly available experimental datasets from RNA sequencing and microarray data of normal cell lines and tumor cell lines / samples, performing quality control indicators, and calculating the z-score of each experiment; (2) Aging feature extraction and model construction and verification: using the Boruta machine learning model to distinguish between aging and non-aging responses and identify key aging-related genes; then applying 10 machine learning models to build classifiers and evaluating model performance through multiple indicators;

[0037] Figure 8 This is a performance evaluation of PreCSenM provided in this embodiment of the invention on diverse datasets and algorithms. Figure 8 'a' is an enrichment network composed of clusters of the core ontology with 5 color codes, visualized using Cytoscape software. Figure 8 b shows the ROC curves and corresponding AUC values ​​of ten machine learning models across three datasets: left, training set (containing samples from 13 normal aging tissue types); middle, test set (combined with the training set, covering 519 samples); right, tumor set (containing samples from more than 10 cancer types). The AUC value is the average of the true positive rate (sensitivity) relative to different false positive rates (1-specificity). Figure 8 c represents a comparison of the performance metrics of 10 prediction models across 3 datasets, including accuracy, precision, recall, F1 score, AUC, and AUPRC. The overall performance of the 10 machine learning models is evaluated using a weighted average of these 7 metrics. Figure 8 d represents a comparison of the classification performance of PreCSenM with single classic markers (CDKN1A, CDKN2A), published aging features, and algorithms on three aging datasets, using weighted values ​​as the evaluation criterion. Abbreviations: LR (Logistic Regression); LMM (Linear Mixed Effects Model); ANN (Artificial Neural Network); Lasso (LASSO Regression); SVM (Support Vector Machine); XGBoost (Limited Gradient Boosting); Ridge Regression; RF (Random Forest); Enet (Elastic Network Regression); PLSR (Partial Least Squares Regression).

[0038] Figure 9 The PreCSenM-derived CS score provided in this embodiment of the invention is significantly correlated with clinical indicators of LUAD patients. Figure 9 a) Comparison of CS scores between primary tumor (blue) and adjacent normal tissue (gray). Data were selected from 10 cancer types in the TCGA database (each containing >35 pairs of normal and tumor samples), and sorted in descending order of tumor CS score. Significance: ****P<0.0001, **P<0.01 (Wilcoxon test). Figure 9 b represents the Kaplan-Meier OS curve for the TCGA-LUAD cohort (log-rank test: P = 0.022, n = 539) based on CS score grouping. Figure 9 c represents the Kaplan-Meier OS curve for GSE30219 (P<0.0001, n=293) based on CS score grouping. Figure 9 d represents the Kaplan-Meier OS curve for GSE41271 (P = 0.0098, n = 182) based on CS score grouping. Figure 9 e represents the Kaplan-Meier OS curve for GSE11969 (P = 0.0013, n = 90) based on CS score grouping. Figure 9 f represents the Kaplan-Meier OS curve for GSE42127 (P = 0.0057, n = 133) based on CS score grouping. Figure 9 g represents the Kaplan-Meier OS curve for GSE68465 (P = 0.00035, n = 178) based on CS score grouping. Figure 9 h represents the Kaplan-Meier OS curve for GSE101929 (P = 0.0015, n = 32) based on CS score grouping. Figure 9 i represents the Kaplan-Meier OS curve for GSE3141 (P = 0.037, n = 58) based on CS score grouping. Figure 9 j represents the Kaplan-Meier OS curve for GSE31210 (P = 0.025, n = 226) based on CS score grouping. Figure 9 k represents the Kaplan-Meier OS curve based on CS score grouping for GSE37745 (P = 0.035, n = 106). Figure 9 l represents the Kaplan-Meier OS curve based on CS score grouping in the meta-cohort integration analysis (P<0.0001, n=1,837). Figure 9m represents the Cox proportional hazards model and survival analysis forest plot of the CS assessment method in the TCGA-LUAD data, showing the OS hazard ratios (HR) and 95% confidence intervals (CI) for the CS group. Significance is defined as log-rank P < 0.05. Figure 9 n represents the association between CS score and AJCC and TNM stages in TCGA-LUAD patients. AJCC stage: Stage I / II vs. Stage III / IV (n = 421 vs. 110); T stage: T1 vs. T2 / T3 / T4 (n = 176 vs. 360); N stage: N0 / N1 vs. N2 (n = 447 vs. 74); M stage: M0 vs. M1 (n = 365 vs. 25). Figure 9 o is a box plot showing the gene expression levels of TP53 (left) and KRAS (right) in the CS group based on TCGA-LUAD data. Wilcoxon test was used for intergroup comparisons, and p-values ​​are reported.

[0039] Figure 10 The PreCSenM-derived CS score provided in this embodiment of the invention reveals the multi-omics molecular characteristics of LUAD patients and malignant cells. Figure 10 'a' represents the Oncoprint map of gene mutation status in the LUAD-related aging pathway. Each column represents one sample, with blue boxes indicating non-synonymous mutations, and the upper bands displaying the patient's CS score. Significance was assessed using Fisher's exact test. Figure 10 b represents a comparison of mutation numbers across different pathways and subgroups. P-values ​​were calculated using Fisher's exact test. Figure 10 c represents the correlation between SCNA and CS scores in LUAD (based on Spearman correlation analysis using rank). Figure 10 Figure d shows the distribution of methylation β values ​​in the two aging subgroups in the LUAD database using box plots and ridge plots. In the ridge plot, the dots represent the median values ​​for each group. Intergroup comparisons were performed using the Wilcoxon test, and p-values ​​were labeled. Figure 10 e represents the Spearman correlation between CS score and pathway activity (calculated via Gene Set Variation Analysis (GSVA) ​​and expressed as a standardized enrichment score (NES)) in TCGA-LUAD patients. Dot size indicates significance, color represents positive (green) / negative (yellow) correlated pathways, and gray indicates no statistical significance. Figure 10 f is designed for LUAD scRNA-seq data to reveal the relationship between different CS states (high CS / low CS) of malignant cells and the other two major cell types (immune cells and stromal cells). Figure 10 g is a single-cell level UMAP dimensionality reduction map, colored according to the three major lung cancer cell types, with each point representing one cell. Figure 10h is a violin plot of CS scores for different T stages of malignant cells in LUAD single-cell data. The Wilcoxon test was used for intergroup comparisons and the P-values ​​were labeled. Figure 10 i is a bubble plot showing significant ligand-receptor interactions between high / low CS malignant cells and neighboring cells. The x-axis represents ligand / receptor expressing cells, and the y-axis represents ligand / receptor molecules. Bubble size indicates the importance of the interaction, and color represents the average expression level of interacting cells (based on CellPhoneDB permutation test). Significance criteria: average expression > 0.2 and P < 0.05. Figure 10 j represents the Reactome pathway enrichment analysis of differentially expressed genes in high / low CS malignant cells. The size of the dots indicates the significance of enrichment; green indicates upregulated pathways in the high CS group, and yellow indicates downregulated pathways.

[0040] Figure 11 This invention provides an exploration of a PreCSenM-driven data-guided tumor treatment strategy. Figure 11 Part a presents a step-by-step process for constructing optimal drug screening features based on the TCGA-LUAD dataset. The upper part generates query features through differential expression analysis of high and low CS groups, univariate Cox proportional hazards regression, and random survival forest screening. The lower part identifies CS subgroup-specific candidate drugs based on LINCS data using feature matching methods. Figure 11 b represents the drug prediction ranking results, highlighting the top 10 potential therapeutic drugs. Figure 11 c represents the two-dimensional and three-dimensional chemical structures of TSA (top) and SAHA (bottom). Figure 11 d represents a comparison of the activity distribution of TSA with three other classes of drugs (chemotherapeutic drugs, targeted anticancer drugs, and non-tumor drugs) in the LUAD cell line. The median IC50 values ​​for each drug were calculated using the PRISM dataset.

[0041] Figure 12 This invention provides the verification of HDACi-induced aging and the construction of an aging transcription program. Figure 12 a) HDACi-based RNA-seq and ATAC-seq data of A549 cells were used to analyze aging characteristics and screen core TFs (left); combined with the TCGA-LUAD cohort, the association between core TFs and patient clinical characteristics was established (right). Figure 12 b shows the survival curves (96 hours) of A549 cells treated with TSA and SAHA (concentrations of 300 nM and 500 nM, respectively). Figure 12 c represents the PreCSenM assessment of senescence levels in A549 cells treated with TSA (RNA-seq data). Differences between groups were analyzed using the Student t-test. Data are expressed as mean ± standard deviation. Figure 12d is a volcano plot showing differentially expressed genes (DEGs) compared to the control group at different time points (6, 12, 24, and 48 hours): green = upregulated (FDR < 0.05, log2 [fold change] ≥ 1), yellow = downregulated (log2 [fold change] ≤ –1). The number of CSGS-related genes and DEGs is labeled. Figure 12 e represents the GO biological process enrichment analysis of differentially expressed genes (top 5 significant entries), with the size of the dots representing the number of genes and the color indicating the significance of enrichment. Figure 12 f represents the transcription factor enrichment analysis based on RcisTarget, which screens for significantly enriched TFs at each time point (AP-1 family TFs are highlighted in the 12-hour treatment group). Figure 12 g and Figure 12 h represents representative SA-β-gal staining images (g) and positive cell quantification (h) of A549 cells after treatment with TSA (300 nM) and SAHA (500 nM) for different time periods. n = 6 biologically independent samples, with 3 random regions analyzed in each group; scale bar = 100 μm; data are expressed as mean ± standard deviation. **** P<0.0001. Figure 12 i and Figure 12 j represents representative images of p21 immunofluorescence staining in A549 cells of the TSA / SAHA treatment group (i; p21 = green, DAPI = blue) and quantitative data of positive cells (j). n = 6 biologically independent samples, scale bar = 25 μm; data are expressed as mean ± standard deviation. ** P<0.01, * P<0.05. Detailed Implementation

[0042] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0043] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] Figure 1 This is a schematic flowchart of a method for constructing a predictive cellular senescence model provided by an embodiment of the present invention. Specifically, the method includes:

[0046] S101, Obtain the training set sample expression profile dataset, which includes normal samples from senescent and non-senescent cell lines;

[0047] In some embodiments, the senescent cell lines in the normal sample expression profile dataset are naturally senescent; in some more specific embodiments, senescence-inducing conditions include: replicative senescence, DNA damage induction, gene editing induction, oncogene activation, and oxidative stress. In some embodiments, the method further includes: standardizing the expression profile dataset to obtain a standardized expression profile dataset; the expression profile dataset includes RNA-seq and microarray datasets.

[0048] S102, using a feature selection algorithm (Boruta algorithm) to identify key aging gene sets from the training set sample expression profile dataset;

[0049] S103, input the key aging gene set into the machine learning model to fit the prediction model, and determine the optimal hyperparameters based on the LOOCV framework to obtain a cell aging model containing the weights of individual genes in the key aging gene set; the cell aging model calculates the aging score by the sum of the products of the expression level of a single gene and its regression coefficient, and divides the training set samples into high and low groups according to the threshold set based on the aging score.

[0050] In some embodiments, the cell senescence model is specifically: Score=E1×β1+E2×β2+···+Ei×βi; where Ei is the expression level of a single gene in the key senescence gene set, βi is the regression coefficient of the corresponding gene, and Score is the senescence score;

[0051] In some embodiments, the machine learning model includes: linear mixed effects model, logistic regression, ridge regression, Lasso regression, elastic network regression, random forest, extreme gradient boosting, support vector machine, artificial neural network, and partial least squares regression;

[0052] Figure 2 This is a schematic flowchart of a method for predicting the aging state of tissue samples according to an embodiment of the present invention. Specifically, the method includes:

[0053] S201, to obtain the expression levels of key senescence gene sets in normal or tumor tissue samples;

[0054] In some embodiments, the key aging gene set includes: LBR, CCDC8, TMPO, ZNF22, HNRNPF, DAZAP1, PTMA, KHDRBS1, RBBP4, RPL22L1, N4BP2, CBX3, SMC4, OSTM1, BCAP31, RAB6B, HAGH, and LRPAP1.

[0055] In some embodiments, the key aging gene set further includes: PHB2, USP1, MCUB, CNOT6, PARP1, RSBN1L, APPL1, SNRPB2, CPSF6, PXMP2, RBM15B, ZNF300, TRIM59, LMNB1, LCORL, CCHCR1, KCTD15, CNTLN, RBIS, ATP6AP1, ABCA3, PLA2G15, GLRB, CYB5R1, ASAH1, NIPAL3, IGFBP7, DNAJB4, CAV2, SLC9A7, GPR155, PPP2R5B, PSG9, and PIK3IP1;

[0056] In some embodiments, the key aging gene set further includes: HMGB1, HMGB3, CDCA7, OLFML3, SFPQ, SMC6, GPSM2, MIOS, SET, MCM7, HNRNPR, TGFB1I1, PCOLCE, CTDSPL2, PFAS, GFPT2, TAF5, SRSF1, RFXAP, RNF138, CEP57, PHGDH, POLA1, EIF2S3, TRA2B, MRPL1, BLMH, PSIP1, HDGF, RPL14, HEPH, NLE1, ID4, H2AZ2, ILF3, RFXANK, DHFR, NTAQ1, BNC2, LRRC40, AD SL, CNTRL, CEP135, RPL29, KDM1A, PPP1CA, PRPF38A, THG1L, RANBP1, INTS7, MAN2B2, SYNGR1, NPDC1, SEMA4F, CHST7, BSDC1, KCTD4, YPEL5, TMEM30A, SLC9A1, KIAA0930, PERP, MICA, RETSAT, LZTS3, MXRA7, H2AC6, RRM2B, SLC17A5, FRY, PTCHD4, TMEM9B, TM7SF3, SLC31A2, SLC25A45, DNASE1L1, CCND1, ITM2B, SLC1A1;

[0057] In some embodiments, the key aging gene set also includes AMPH, SOX11, ALMS1, CPXM1, IMPDH2, TRIM28, DFFB, MRPL48, RCC2, NUP54, NUP160, CPNE2, FBXO45, MRPL30, TAF1B, ARHGAP32, PAXIP1, AGPAT5, FANCE, RPS7, SENP7, OLFML2B, EFS, NUP85, TMA16, TEAD2, NFIB, TM2D1, EDEM2, TMTC1, FUCA1, IFI6, TMEM217, SEC11C, SPR, RAB11FIP5, ZNF219, CLPTM1L, H2BC21, SPATA18, DEDD2, DNAJC5, SLC35F5, SERINC1, CHMP5, MFSD11, GNPTG, and MARCHF2.

[0058] In some embodiments, the key aging gene set further includes: DMAC2, UTP20, PAXBP1, PELI2, EML4, EFNB3, C5, SRSF7, PTBP2, RFX3, NEMP2, C1QBP, HNRNPA2B1, RBM10, LSM8, THOC1, MAD2L2, ZNF518A, SMIM15, FH, ALKBH2, UTP11, FANCL, RBMS3, NID1, BNC1, TIA1, ABCB7. MDM4, GTF3C4, U2SURP, DERA, KCTD3, CHD1, RABAC1, KLHL21, SLC25A20, COX14, NFE2L1, MFSD3, AMPD3, DYNLT3, PL CB4, RNF11, ENPP4, SLC4A11, MIEF2, MAP3K7CL, CREG1, ZDHHC24, BPGM, CDKN1A, ADCY6, RNF14, MYOZ2, AK3, HSPB8.

[0059] S202, Input the expression level into the cell senescence model constructed by the method of the first aspect of this application, calculate the senescence score and divide the tissue sample into a high group or a low group according to the senescence score; If it is a high group, output the auxiliary prediction result that the tissue sample is in a senescent state with a high probability; if it is a low group, output the auxiliary prediction result that the tissue sample is in a senescent state with a low probability.

[0060] In some embodiments, if the tissue sample is tumor tissue, the method further includes: predicting the overall survival of the corresponding subject and / or predicting the tumor purity of the tissue sample based on whether the tissue sample belongs to a high or low group; specifically, if it is a high group, outputting an auxiliary prediction result indicating a shorter overall survival of the tissue sample; if it is a low group, outputting an auxiliary prediction result indicating a longer overall survival of the tissue sample; if it is a high group, outputting an auxiliary prediction result indicating a lower tumor purity of the tissue sample; if it is a low group, outputting an auxiliary prediction result indicating a higher tumor purity of the tissue sample.

[0061] In some more specific embodiments, tumor tissue types include: brain cancer, breast cancer, leukemia, liposarcoma, liver cancer, lung cancer, melanoma, osteosarcoma, pancreatic cancer, and colorectal cancer.

[0062] In some embodiments, this application also discloses a method for predicting the prognosis of LUAD using a predictive model, the method comprising:

[0063] Obtain the expression levels of key aging gene sets of the subjects; input the gene set expression levels into the aging model constructed by the method of the first aspect of this application, calculate the aging score, and divide the subjects into high or low groups according to the aging score;

[0064] The molecular characteristics of the subjects are predicted based on whether they belong to a high or low group. These molecular characteristics are specifically included at the genomic, epigenetic, transcriptomic, and single-cell levels: (Genomic level) If the subject is in a high group, an auxiliary prediction result indicating lower genomic variability is output; if in a low group, an auxiliary prediction result indicating higher genomic variability is output. (Epigenetic level) If the subject is in a high group, an auxiliary prediction result indicating lower DNA methylation levels of the differentially methylated probes is output; if in a low group, an auxiliary prediction result indicating higher DNA methylation levels of the differentially methylated probes is output. (Transcriptomic level) If the subject is in a high group, an auxiliary prediction result indicating suppressed proliferation and activated inflammatory responses is output. (Single-cell level) If the subject is in a high group, an auxiliary prediction result indicating that the malignant cells can activate immune regulatory signals and have a good prognosis is output.

[0065] In some embodiments, the terms “subject” or “test subject” or “sample to be tested” as used herein refer to any animal (e.g., a mammal), including but not limited to humans, non-human primates, rodents, etc., which will become the recipient of a particular treatment. Generally, the terms “subject” and “patient” are used interchangeably herein when referring to human subjects. Preferably, the subject is a human. In some embodiments, the sample to be tested is a patient clinically used for prognostic assessment.

[0066] Figure 3This is a schematic flowchart of a screening method for potential therapeutic drugs provided in an embodiment of the present invention. Specifically, the method includes:

[0067] S301, Obtain a tumor sample genomic dataset containing survival data and the drug to be tested; S302, Input the dataset into the cell senescence model constructed by the method of the first aspect of this application and divide the samples into high-group or low-group; S303, Compare the high-group and low-group to screen out the first differentially expressed gene; Evaluate the correlation between the first differentially expressed gene and survival time to screen out the second differentially expressed gene; Optimize the second differentially expressed gene using an algorithm to obtain the target gene with a positive variable importance value; S304, Process the drug to be tested and the target gene using a feature matching analysis method and calculate the feature matching score; Obtain the preliminary ranking result of the drug to be tested based on the feature matching score; S305, Analyze the preliminary ranking result based on the rank aggregation analysis method of the order statistic to obtain the final ranking result; S306, Obtain potential therapeutic drugs based on the final ranking result.

[0068] In some embodiments, the above-mentioned auxiliary prediction results include, but are not limited to, paper or electronic reports. These results are obtained by intelligent machines based on the relevant data of the subjects and are only used as a reference for medical staff, and are not considered as the final diagnosis results of the subjects.

[0069] In some embodiments, the threshold is obtained through training on training set samples. It can be a specific threshold or an interval range. The specific form is not specifically limited in this embodiment.

[0070] Figure 4 This is a schematic diagram of a computer device provided in an embodiment of the present invention, such as... Figure 4 As shown, device 2000 may include: one or more processors 2010 and one or more memories 2020; wherein the memories store computer-readable code that, when run by one or more processors, can execute the methods described above.

[0071] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 or ARM architecture.

[0072] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0073] For example, the method or apparatus according to embodiments of this disclosure can also be used by means of Figure 5 The architecture of the computing device 3000 shown is used for implementation. For example... Figure 5 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 5 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 5 One or more components in the computing device shown.

[0074] This invention also includes a computer-readable storage medium, such as... Figure 6The diagram illustrates a storage medium 4000 provided in an embodiment of the present invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method described above according to embodiments of the present disclosure can be performed. The computer-readable storage medium in the embodiments of the present disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0075] This disclosure also provides a computer program product or system, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0076] In some embodiments, this embodiment also discloses a system for constructing a predictive cellular senescence model, the system comprising:

[0077] The training set sample acquisition module is used or configured to acquire the training set sample expression profile dataset, which includes normal samples from senescent and non-senescent cell lines.

[0078] A key aging gene set extraction module is used or configured to identify key aging gene sets from the training set sample expression profile dataset using a feature selection algorithm (Boruta algorithm);

[0079] The cell senescence model training module is used or configured to input the key senescence gene set into the machine learning model to fit the prediction model, and determine the optimal hyperparameters based on the LOOCV framework to obtain a cell senescence model containing the weights of individual genes in the key senescence gene set. The cell senescence model calculates the senescence score by summing the products of the expression level of a single gene and its regression coefficient, and divides the training set samples into high and low groups according to the senescence score by setting a threshold.

[0080] In some embodiments, this embodiment also discloses a system for predicting the aging state of tissue samples, the system comprising:

[0081] A module for obtaining the expression levels of key senescent gene sets is used or configured to obtain the expression levels of key senescent gene sets in normal or tumor tissue samples.

[0082] The grouping module is used or configured to input the expression level into the cell senescence model constructed by the method of the first aspect of this application, calculate the senescence score and divide the tissue sample into a high group or a low group according to the senescence score; if it is a high group, it outputs an auxiliary prediction result that the tissue sample is in a senescent state with a high probability; if it is a low group, it outputs an auxiliary prediction result that the tissue sample is in a senescent state with a low probability.

[0083] Optional key aging gene sets include: LBR, CCDC8, TMPO, ZNF22, HNRNPF, DAZAP1, PTMA, KHDRBS1, RBBP4, RPL22L1, N4BP2, CBX3, SMC4, OSTM1, BCAP31, RAB6B, HAGH, and LRPAP1.

[0084] In some embodiments, this embodiment also discloses a screening system for potential therapeutic drugs, the system comprising:

[0085] The dataset and drug acquisition module is used or configured to acquire a tumor sample genomic dataset containing survival data and the drug to be tested.

[0086] A sample grouping module is used or configured to divide samples into high or low groups when inputting a dataset into a cell senescence model constructed by the method of the first aspect of this application;

[0087] The target gene screening module is used or configured to screen for the first differentially expressed gene by comparing high and low groups; evaluate the correlation between the first differentially expressed gene and survival time to screen for the second differentially expressed gene; and optimize the second differentially expressed gene to obtain the target gene with a positive variable importance value.

[0088] The feature matching score calculation module is used or configured to process the test drug and target gene using feature matching analysis methods, calculate the feature matching score, and obtain the preliminary ranking results of the test drug based on the feature matching score.

[0089] The sorting result output module is used or configured to analyze the preliminary sorting results using the rank clustering analysis method based on order statistics to obtain the final sorting results.

[0090] A potential therapeutic drug screening module, used or configured to obtain potential therapeutic drugs based on the final ranking results.

[0091] Example 1:

[0092] CS cell line culture experiments: This method relies on a large amount of well-validated CS experimental data in the literature. Dataset selection must meet the following requirements: ① Include RNA-seq and microarray datasets; ② Use standardized single-channel microarrays, excluding custom microarrays to ensure annotation compatibility; ③ The dataset must fully record CS events; ④ Include at least two control and paired treatment groups (raw or processed data are acceptable); ⑤ Use existing software packages for data processing; ⑥ Data with fewer than 16,000 genes will not be used as training data.

[0093] A comprehensive search was conducted on transcriptome data related to CS in cell lines or tissues. Using the European Molecular Biology Laboratory (EMBLEBI) Gene Expression Omnibus Database (GEO) and the ArrayExpress database, a total of 888 cell lines from 68 publicly available datasets were retrieved. Of these, 519 samples represented normal cell types, including lung fibroblasts, foreskin fibroblasts, dermal fibroblasts, lung epithelial cells, mesenchymal stem cells, umbilical cord endothelial cells, ocular epithelial cells, myoepithelial cells, keratinocytes, prostate stromal cells, vascular smooth muscle cells, and nervous system cells. 369 samples came from more than 10 cancer types, including brain cancer, breast cancer, leukemia, liposarcoma, liver cancer, lung cancer, melanoma, osteosarcoma, pancreatic cancer, and colorectal cancer. These selected datasets covered various aging-induced conditions, such as replicative aging, DNA damage, gene editing, oncogenes, and oxidation.

[0094] Publicly available multi-omics data and LUAD resources: RNA sequencing data (TPM), somatic mutation data (mutation annotation format), DNA methylation data (Illumina Human Methylation 450K-array), and relevant clinical information from the TCGA-LUAD project were downloaded using the R package TCGAbiolinks (version 2.18.0). Nine independent LUAD public transcriptome datasets and complete clinical data were retrieved from GEO, including GSE30219, GSE41271, GSE11969, GSE42127, GSE68465, GSE101929, GSE3141, GSE31210, and GSE37745. The GSE117570 dataset was used to determine the CS index using LUAD RNA-seq data at the single-cell level.

[0095] Cell processing for RNA-seq and ATAC-seq experiments: TSA and SAHA were purchased from MedChemExpress (New Jersey, USA) and dissolved in dimethyl sulfoxide (DMSO; Solarbio, Beijing, China) to achieve a storage concentration of 10 mM. For RNA-seq and ATAC-seq, A549 cells were cultured at 2 × 10⁻⁶ cells / cells. 5 Cells were seeded at a density of 100 cells / well in six-well plates and exposed to growth medium containing TSA (300 nM) or SAHA (500 nM). Under control conditions, cells were cultured in growth medium containing 0.0033% DMSO (same concentration as TSA) or 0.005% DMSO (same concentration as SAHA) for 48 hours without any further treatment. Cells were exposed to the above concentrations of TSA or SAHA for 6, 12, 24, and 48 hours. Each condition was repeated to obtain samples for RNA-seq and ATAC-seq.

[0096] RNA-seq library preparation and sequencing: Total RNA was extracted from A549 cells using the TRIzol method and treated with RNase-free DNase I (Takara) according to the manufacturer's instructions. Paired-end read cDNA libraries were constructed and sequenced on a NovaSeq 6000 system.

[0097] RNA Sequencing and Microarray Processing: Raw RNA-seq data (in FastQ file format) was downloaded from SRA using prefetch and fastq-dump in the SRA Toolkit (version 2.11.0), or the samples were sequenced using the NovaSeq 6000 sequencing system. Quality control was performed using FastQC software (version v0.11.9), discarding low-quality reads with an average quality score <20. Trimming of adaptor sequences and low-quality bases was performed using trim-galore (version 0.6.7). Aligned sample sequences were aligned to the GRCh38 genome using STAR (version 2.5.2a). Gene expression (TPM) was quantified using RSEM (version v1.3.3). Protein-coding genes were identified using the BiomaRtR software package (version 2.46.3). After excluding non-coding genes, 19,425 protein-coding genes were included in subsequent analyses. Log2 transformation was performed on normalized gene expression. Differentially expressed genes (DEGs) between the treatment and control groups under each condition were compared using the DESeq2 package (version 1.30.1) in R. GO analysis was performed using the ClusterProfiler package (version 3.18.1) in R to investigate the potential biological functions of DEGs.

[0098] The raw microarray data were retrieved from the GEO database, which was generated using various platforms, including Affymetrix GPL570 (Human Genome U133Plus 2.0 array), GPL96 (Human Genome U133A array), GPL6244 (Human Genome 1.0ST array), GPL5175 (Human Exon 1.0ST array), GPL7015 and GPL6884 (Human WG-6v3.0 expression bead array), Illumina GPL10558 (Human HT-12V4.0 expression bead array), Agilent GPL6480 (whole human genome microarray 4x44KG4112F) and Agilent GPL13497 (whole human genome microarray 4x44Kv2). Raw data from the Affymetrix platform were quality controlled and normalized using the Robust MultichipAverage method implemented in the oligonucleotide package (version 1.54.1), and data processing was performed on the Illumina and Agilent platforms using the limma package (version 3.46.0). Probe annotations were matched against HUGO Gene Nomenclature Committee (HGNC) symbols using available annotation packages or GPL platform files for the datasets. Probes that did not match HGNC symbols were removed, and the average intensity of probes matching the same HGNC symbol was calculated.

[0099] All training samples were combined for the machine learning process, and batch effects were eliminated using the ComBat algorithm from the sva package (version 3.38.0). Gene expression in all samples was readjusted using z-score transformation.

[0100] ATAC-seq processing: Low-quality reads and adaptor sequences were removed from the raw sequence reads using Trim_galore (version 0.6.7). The cleaned data was aligned to the human reference genome (hg38) using Bowtie2 (version 2.3.4.2). Uniquely mapped paired reads were retained, and the SAM files were sorted and converted to BAM format using Samtools (version 1.3.1). Subsequently, peak calls were performed using MACS2 (version 2.2.7.1) to identify significant peaks in each sample. Peak regions in the genome were annotated using the ChIPseeker R package (version 1.26.2). The Integrative GenomicsViewer tool was used to convert the bigwig files for viewing in DeepTools (version 3.5.1). To identify DAR, the treatment and control groups were compared in different cell lines using DESeq2 (version 1.30.1), with a threshold of P < 0.05.

[0101] Machine Learning Ensemble Approaches for Feature Generation, Model Building, and Validation: To develop consistent aging features and predictive models, 11 machine learning algorithms were integrated to improve the accuracy and stability of the results. These algorithms include Boruta, LMM, LR, Ridge, Lasso, Enet, RF, XGBoost, SVM, ANN, and PLSR. After generating aging features, model development and validation were performed as follows: First, approximately 70% of normal aging cell lines were randomly assigned for training based on dataset selection criteria to identify feature genes and build a robust predictive model; this process was repeated 10 times. The remaining 30% of normal aging cell lines were used as the test set, while tumor cell lines or tissues were used as the validation set. Second, the Boruta feature selection algorithm (an ensemble learning model for feature selection) was applied to the training dataset to identify a precise consistent aging feature gene set (CSGS). Third, Metascape was used to perform functional enrichment analysis on genes associated with aging features. Subsequently, ten binary classification algorithms (excluding Boruta) were applied to the aging features generated in the previous step to fit the predictive model, and the optimal hyperparameters were determined based on the LOOCV framework in the training set. All models were tested on the remaining test and tumor datasets, exhibiting different aging-inducing patterns. To ensure robustness, model evaluation was repeated at least ten times using randomly split training and test datasets. Finally, accuracy, precision, recall, F1 score, area under the receiver operating characteristic (ROC) curve, and area under the precision-recall curve (AUPRC) were evaluated for each model across all training, test, and tumor datasets. The model with the highest weighted metric based on sample size was considered the best. Specifically... Figure 7 As shown in b. Figure 7 a represents the development method of the computational model and its application in assessing CS levels and LUAD research.

[0102] After rigorous preprocessing, data integration, and machine learning-based feature extraction, 236 genes were selected for constructing a consistent CSGS. This marker is significantly enriched in aging-related biological processes, including cell cycle regulation, mRNA processing, RNA splicing, cellular responses to DNA damage stimuli, and mRNA metabolism. Figure 8 a) These genes include several widely validated aging-related genes (such as CDKN1A, CCND1, LMNB1, HMGB1, and MCM7), genes involved in mRNA processing and splicing (such as SNRPB2, SRSF1, and SRSF7), interferon-responsive genes (IFI6), and histone variants (such as H2AC6, H2AZ2, and H2BC21). In summary, CSGS constructed based on multi-platform data can comprehensively reflect the typical characteristics of CS.

[0103] Model Description: A total of 11 models were implemented using R packages, including one feature-based model and 10 binary classifiers. These packages include lmerTest (version 3.1.3), glmnet (version 4.1.2), caret (version 6.0.88), randomForest (version 4.6.14), Boruta (version 7.0.0), xgboost (version 1.5.0.2), e1071 (version 1.7.6), neuralnet (version 1.44.2), and pls (version 2.8.0). The Boruta feature selection algorithm was used to identify key genes related to CS. We utilized the Boruta feature selection algorithm, specifically designed for ensemble learning models, to improve the accuracy of the results. The response variable was the aging state, where 0 represented non-aging and 1 represented aging. This method generated a compact and high-quality set of aging-related feature genes for further analysis. Then, important aging-related feature genes were used as input variables to construct a predictive binary classification model.

[0104] Ten classification methods, including LMM, were evaluated using CSGS, and the optimal parameters were selected using LOOCV on the training dataset. Except for Enet and PLSR, all models showed consistent performance metrics across the training, test, and tumor datasets. This indicates that the models have good generalization ability to unknown data without overfitting the training set, and exhibit robust performance on both the test and tumor datasets. Figure 8 b, Figure 8 c). Among them, the LR model showed the highest weighting index (0.95) in the training, testing, and tumor datasets, and was selected as the optimal model. Figure 8 c). Based on this, we developed PreCSenM, which integrates CSGS with the optimal LR model coefficients, and can accurately distinguish between senescent and non-senescent cells in normal and tumor environments.

[0105] The LOOCV framework is implemented to identify the optimal hyperparameters for each model on the training set, thereby constructing a best-performing classifier. The dataset is divided into n parts, where n is the number of samples in the training set. In each iteration, n-1 parts are used as the training set to build the model, and the remaining parts are used as the validation set for model evaluation (n = number of training samples). This process is repeated n times until each part has been used as validation data once. For prediction models, the optimal marker gene is used as an input variable in algorithms such as Ridge, Lasso, Enet, RF, XGBoost, SVM, ANN, and PLSR. Conversely, in LMM and LR models, each gene is used as a separate input variable. The aging score in LMM and LR models is calculated by multiplying the gene expression (Ei) of the best marker gene in each sample by its respective coefficient (βi), as follows: core = E1×β1 + E2×β2 + ... + Ei×βi; where Ei is the expression level of a single gene in the key aging gene set, βi is the regression coefficient of the corresponding gene, and the score is the aging score. The data is scaled to a mean of 0 and a standard deviation of 1, and min-max normalization is performed to map the CS probability score of each sample to the range of 0-1 to account for differences in gene expression feature intensity and to compare the relative scores of different samples simultaneously. A higher score indicates a higher probability that the sample is in an aging state. The default cutoff value is 0.5 (median of the mean score), which can be adjusted according to actual needs.

[0106] Model Performance Evaluation: Model performance is evaluated using the following metrics: accuracy, precision, recall, F1 score, ROC curve AUC, and AUPRC. All metrics and charts are generated based on the ROCR package (version 1.0.11) or the confusion matrix. Accuracy is the proportion of true positive and true negative predictions out of all predictions, calculated using the following formula: Where FP represents the number of false positives, TP represents the number of true positives, and TN represents the number of true negatives. Precision is the proportion of true positives among samples predicted as positive, calculated using the following formula: Recall (sensitivity) is the proportion of true positive samples that are correctly predicted, and it is calculated using the following formula: Where FN is the false negative number. The F1 score is the harmonic mean of the combined precision and recall, calculated as follows: The ROC curve reflects the relationship between the true positive rate and the false positive rate at all possible thresholds. The AUPRC curve evaluates the predictive performance of the model based on the relationship between precision (positive predictive value) and recall (true positive rate).

[0107] PreCSenM outperforms existing aging prediction methods: To evaluate the performance of PreCSenM in predicting senescent cells, we compared it with the following methods: single classic biomarkers (CDKN1A and CDKN2A), five publicly available aging feature sets (SenMayo, Casella, Fridman, Purcell, and Hernandez), and three aging gene scoring algorithms (GSVAscore, SenCID, and SENCAN). We compared the area under the receiver operating characteristic (AUC) of the single biomarker, aging feature set, gene scoring algorithm, and PreCSenM method on the training, test, and tumor datasets. Figure 8 d). The results show that PreCSenM exhibits superior performance across all queues, with a weighted AUC of 0.97, significantly outperforming all other methods and algorithms. Figure 8 d).

[0108] To validate the applicability of PreCSenM under different conditions and sample types, the single-cell data analysis capability was evaluated using the annotated senescence-state scRNA-seq dataset (GSE115301). Experimental datasets from 10 human and mouse cancer models (including tumor samples / cell lines with clearly defined senescence states) were included. Cancer and normal tissue data for 10 cancer types (sample size ≥ 35) were selected from the TCGA (The Cancer Genome Atlas) database. Differences in CS scores between groups were analyzed using the Mann-Whitney U test (Wilcoxon test, for non-normal data) or t-test (for normal data), and the corresponding p-values ​​were determined.

[0109] Validating CS scores across various datasets: Scores were measured using the PreCSenM method to validate the applicability of CS scores to various conditions and data types. First, to assess the applicability of PreCSenM to scRNA-seq data analysis, we used a publicly available single-cell transcriptomics dataset reporting senescence status (GSE115301). Additionally, we selected TCGA data containing normal and cancer samples across 10 cancer types (more than 35 samples), consistent with previous results. To assess statistical differences in CS scores between groups, a two-sided Mann-Whitney U test (Wilcoxon test) was performed on the non-normally distributed data, and p-values ​​were determined.

[0110] CS Score and Clinical Predictors: To elucidate the relationship between CS score, overall survival (OS), and clinical characteristics in patients with LUAD, survival analyses were performed on 10 independent datasets, including the TCGA and GEO databases. All samples were divided into high-CS and low-CS groups, and the optimal cutoff value was determined using the "surv_cutpoint" algorithm in the R package Survminer (version 0.4.9). Kaplan-Meier survival curves were then used to assess survival differences between the two groups, and log-rank tests were used to determine statistical significance (P < 0.05). Survival analyses were performed using R survival analysis (version 3.1.12) and the Survminer package. A comprehensive analysis explored the correlation between CS score and various clinical characteristics (e.g., TNM stage, AJCC stage, and vital status) in the two different CS groups.

[0111] Single-cell RNA sequencing data analysis: Gene expression matrices for LUAD were obtained by extracting scRNA-seq data from lung cancer (GSE117570). These matrices were then preprocessed using standard analysis procedures in MAESTRO to select a high-quality cell subset. Cells with fewer than 1000 unique reads, fewer than 500 expressed genes, or mitochondrial genome UMIs > 5% were excluded. The remaining cells were used for downstream expression quantification and cluster analysis. Based on prior knowledge of marker genes, the cell clusters obtained from scRNA-seq were annotated into three distinct major cell types: immune cells, stromal cells, and malignant cells.

[0112] The CS score of scRNA-seq data was calculated by multiplying the gene expression levels of 236 CSGS in each cell by a coefficient. Different cell types and tumor T stages were assessed to determine their senescence status. The median score was used as the default cutoff value for two groups of different malignant cells; a higher score indicated a greater likelihood of senescence.

[0113] To identify senescence-related DEGs between high-CS and low-CS groups in malignant cells, the FindMarkers function in the R package Seurat (version 4.0.1) was based on the Wilcoxon rank-sum test. Genes were selected when the p-value was <0.05 and the mean difference after log2 transformation between the two groups was >0.15. The biopathological enrichment of DEGs in scRNA-seq data was analyzed using the Reactome database. To investigate intercellular communication networks via ligand-receptor interactions, counting matrices of four different cell types (immune cells, stromal cells, highly cytotoxic malignant cells, and low-cytotoxic malignant cells) were used as input to CellPhoneDB (version 1.1.0). Permutation tests were performed between cell types based on the mean gene expression of the ligand in one cell type and the corresponding receptor in another to identify potentially important interaction pairs. Pairs with mean expression levels >0.2 and p-values ​​<0.05 were designated as significant predicted interaction pairs.

[0114] Clinical applications of PreCSenM:

[0115] 1. PreCSenM CS score as a potential predictor of clinical outcomes in LUAD;

[0116] Existing evidence suggests that in most cancer types, tumor tissue exhibits lower levels of senescence compared to normal tissue. We applied the CS score to a cohort from the TCGA database, containing over 35 samples across 10 cancer types. The results showed that, except for thyroid cancer (THCA) tissue (consistent with previous studies), normal tissue generally exhibited higher levels of senescence than tumor tissue. Figure 9 a). This indicates that the PreCSenM classifier can accurately detect the aging state of different tumor tissues.

[0117] Lung cancer has the lowest 5-year survival rate among all cancers and is the leading cause of cancer-related death. Currently, there is a large amount of publicly available multi-omics data on lung cancer (LUAD). Among 10 cancer types, LUAD showed the highest CS level (…). Figure 9 a). To investigate the relationship between CS score and clinicopathological features, we collected 10 LUAD datasets (a total of 1,837 patients). Based on the optimal cutoff value determined independently for each dataset (e.g., the TCGA cohort), patients were divided into high CS and low CS groups. Survival analysis showed that, across all datasets, the overall survival (OS) of patients in the high CS group was significantly better than that of the low CS group. Figure 9 bk). Meta-cohort analysis integrating all 1,837 samples showed the same trend (P<0.0001); Figure 9Furthermore, in our survival comparison analysis using multiple CS assessment methods on the TCGA-LUAD dataset, we found that PreCSenM consistently outperformed other methods, producing the most statistically significant results, highlighting its robustness and clinical translation potential. Figure 9 High CS levels are also significantly associated with several clinicopathological features, including survival status, AJCC (American Joint Committee on Cancer) stage, and TNM stage. Figure 9 Tumors with higher CS levels typically exhibit lower tumor purity, which may be driven by increased stromal or immune cell infiltration in the senescent tumor microenvironment. Furthermore, TP53 and KRAS gene expression levels were significantly lower in the high-CS group than in the low-CS group. Figure 9 These results suggest that the CS score may serve as an independent biomarker for predicting the prognosis of LUAD.

[0118] 2. The association between LUAD multi-omics molecular features and CS;

[0119] To elucidate the role of CS in LUAD, we integrated multi-omics analysis based on TCGA-LUAD data to explore the association between CS score and molecular characteristics. At the genomic level, by comparing the mutation profiles of key aging-related genes in patients with high and low CS, we found that the mutation burden of most genes was significantly higher in the low CS group than in the high CS group. Figure 10 (a, b) Among them, the mutation frequency of TP53, a core regulator of the cell cycle / TP53 pathway, showed the most significant difference (59% in the low CS group vs. 37% in the high CS group; P = 6.27e-08), while RB1 mutations were also significantly enriched in the low CS group (8% vs. 2%; P = 3.75e-04). These results suggest that TP53 mutations are highly prevalent in patients with low CS, consistent with their clinically aggressive tumor behavior. Further assessment using somatic copy-number alteration (SCNA) scores (including focal, chromosome arm, and total chromosome-wide variations) revealed that the overall SCNA score in LUAD patients significantly decreased with increasing CS scores (R² = -0.35, P = 2.3e-16). Figure 10 c). Furthermore, the CS score was significantly negatively correlated with tumor mutation burden (TMB), aneuploidy score (AS), and microsatellite instability (MSI) score. These results indicate that LUAD patients with high CS have significantly reduced genomic variation, suggesting that tumorigenesis is accompanied by more dramatic genomic alterations, and that cancer cell senescence may suppress these variations.

[0120] At the epigenetic level, analysis of aging-related epigenetic changes in TCGA-LUAD patients revealed that the DNA methylation level of differentially methylated probes (DMPs) was generally lower in the high CS group. Figure 10 d). Furthermore, DMPs with significant signal intensity (such as the CSGS member gene CPXM1) were negatively correlated with aging levels. These results suggest that patients with low aging levels exhibit hypermethylation characteristics, and hypermethylation is a marker of malignant cancer phenotypes.

[0121] At the transcriptomic level, gene set variation analysis (GSVA) ​​was used to calculate the activity of landmark pathways, revealing that 34 out of 50 pathways were associated with aging-related transcriptomic features of LUAD. Among these, the activities of proliferation-related gene sets (such as MYC targets, G2M checkpoints, and E2F targets), as well as pathways involving oxidative phosphorylation, mTORC1 signaling, DNA repair, and unfolded protein responses, were significantly negatively correlated with CS scores. Figure 10 e). Inflammatory response pathways (such as IL2-STAT5 signaling, IL6-JAK-STAT3 signaling, NFκB-mediated TNFα signaling, and interferon-γ response), as well as Kras signaling, the p53 pathway, and apoptosis, were significantly positively correlated with CS scores. Figure 10 e). In summary, age-related pathway alterations suggest that proliferative processes are suppressed while inflammatory responses are activated in aging patients.

[0122] At the single-cell level, single-cell analysis based on the MAESTRO standard procedure showed that cells in the LUAD microenvironment could be clustered into three major categories: immune cells, stromal cells, and malignant cells. Figure 10 f). CS levels in malignant cells were the lowest among the three cell types. Figure 10 g), and the CS level of malignant cells in T1 phase was significantly higher than that in T2 phase ( Figure 10 h), consistent with previous results ( Figure 9 Further analysis of cell-cell interactions revealed that specific ligand-receptor pairs (such as HLA-DPA1-TNFSF9, LGALS9-SORL1, TGFB1-TGFBR1, and COPA-P2RY6) between immune cells and high-CS malignant cells have pro-inflammatory effects, while ligand-receptor pairs (such as C5AR1-RPS19, NR3C1-CXCL8, and BSG-PPIA) between immune cells and low-CS malignant cells promote cancer progression and metastasis. Figure 10i). Pathway enrichment analysis of differentially expressed genes (DEGs) showed significant enrichment of immune-related pathways (such as cytokine signaling, adaptive immune system, interferon signaling, and PD-1 signaling) in high-CS malignant cells. Figure 10 j). Overall, the results indicate that high levels of CS malignant cells are associated with favorable clinical outcomes by activating immune regulatory signals.

[0123] Example 2:

[0124] Identification of aging-related drugs: To advance the research and development of drugs related to lung cancer cell aging, we systematically screened potential therapeutic drugs according to the following process ( Figure 11 a). First, a step-by-step screening process was constructed based on the TCGA-LUAD cohort to obtain gene features related to CS and survival: 6,717 DEGs (corrected P-value < 0.05 and |FC fold change| > 2) were retained as LUAD aging features by comparing high / low CS groups. Second, univariate Cox regression analysis was used to assess the prognostic correlation between these DEGs and OS, and 205 significantly related genes were screened out. Subsequently, to optimize the size of the feature gene set, a random survival forest (RSF) model was constructed (1000 trees were generated each time, repeated 1000 times). Finally, 182 genes with positive variable importance values ​​were included. Figure 11 a).

[0125] Based on these 182 genes, we employed three complementary methods—Kolmogorov-Smirnov (KS) test, eXtreme Sum (XSum), and Reverse Gene Expression Score (RGES)—to perform feature matching analysis in the A549 lung cancer cell line database of the LINCS (Library of Integrated Network-Based Cellular Signatures) project. Figure 11 a). The KS method categorizes input genes into discrete positive and negative regulators. Using all compound spectra as a reference, the positive (ES) ratio based on maximum deviation is estimated. pos ) and negative (ES) neg Enrichment score (ES). The KS score is defined as: KS score = ES pos -ES negIn this study, both positive and negative regulators had KS scores of 0. The KS method is the most widely used method for feature matching, especially in CMap research. Similar to the KS method, the XSum method divides the input genes into two gene sets. In short, we calculated the change in the reference drug features relative to the positive and negative regulators (Sum). pos and Sum neg The sum of all elements. XSum is defined as: XSum = Sum pos -Sum neg The RGES method is an improvement on the original KS method and has recently been shown to perform better in drug prediction. RGES is an indicator that describes the negative correlation between gene expression profiles of a specific disease and its treatment, defined as: RGES = ES pos -ES neg Regardless of direction. Based on the ranking results of the above methods, a robust drug ranking is obtained using the rank aggregation analysis method based on order statistics proposed by Stuart et al. The rank aggregation score (RAS) represents the probability used to generate the final ranking list of potential drugs; drugs with higher RAS indicate stronger reversal efficacy in tumors. RAS is defined as: RAS = -log10 (probability score); through the integration method based on order statistics ( Figure 11 b) Perform sorting and aggregation analysis on the results to obtain robust drug prediction results.

[0126] Notably, HDACis such as trichostatin-a (TSA) and vorinostat (SAHA) exhibited the most significant anti-aging activity among all candidate drugs. Figure 11 (b, c) Given the extensive research data on TSA in the LUAD cell line, we focused on comparing the IC50 values ​​of TSA with other drug classes (chemotherapeutic drugs, targeted anticancer drugs, and non-tumor drugs). The results showed that the IC50 value of TSA was significantly lower than that of other drug classes ( Figure 11 d) indicates that the LUAD cell line has higher sensitivity to TSA treatment.

[0127] 1. HDACi induces senescence in LUAD cells

[0128] To verify whether HDACi can induce senescence in LUAD cells, we performed RNA-seq on A549 cells treated with TSA at different time points. Figure 12 a). Notably, with prolonged treatment time, HDACi significantly reduced cancer cell viability ( Figure 12 b). Assessing aging levels using CS scores calculated using the PreCSenM model revealed a sustained increase in CS scores during the 6–12 hour treatment period. Figure 12 c). Differential gene expression analysis showed that the number of DEGs peaked after 12 hours of treatment. Figure 12 d). Gene ontology (GO) analysis further revealed that at this stage, DEGs were significantly enriched in aging-related biological functions, including cell cycle arrest, G2 / M phase transition, and histone modifications. Figure 12 e).

[0129] To confirm whether HDACi treatment induces senescence in A549 cells, we measured the activity of the classic senescence marker SA-β-gal. The results were consistent with the CS score. Figure 12 f) After treatment with 300 nM TSA or 500 nM SAHA for 6–12 hours, the number of SA-β-gal positive cells increased significantly. Figure 12 g). Time-dynamic analysis showed that the pro-aging effects of TSA and SAHA tended to stabilize after 12 hours of treatment. Figure 12 h). Given that p21 plays a crucial role in cancer cell senescence by regulating the cell cycle, we examined the effect of HDACi on p21 protein levels in the nucleus of A549 cells. Immunofluorescence analysis showed that treatment with 300 nM TSA or 500 nM SAHA for 12 hours significantly upregulated p21 protein expression (h). Figure 12 The above results indicate that HDACi treatment for 12 hours can effectively induce senescence in A549 cells.

[0130] 2. The AP-1 family is involved in enhancing HDACi-induced aging.

[0131] Epigenetic mechanisms play a crucial role in the initiation and progression of cerebral senescence (CS). To reveal the dynamic changes in chromatin accessibility during HDACi-induced senescence, we performed time-series transposase-accessible chromatin sequencing (ATAC-seq) analysis on TSA-treated A549 cells. The promoter region of the 12-hour drug-treated group showed a higher ATAC peak signal than the control group, indicating that chromatin accessibility gradually increased at 12 hours. Since the epigenetic status of the promoter region affects the expression patterns of downstream genes, we further integrated differential accessibility region (DAR) analysis using RNA-seq and ATAC-seq to conduct a comprehensive study of DEGs. The results showed that the chromatin accessibility of CS-related upregulated genes in the promoter region was significantly higher than that of downregulated genes.

[0132] When inferring potential regulatory transcription factors (TFs) using DARs and DEGs, we found that different DEGs corresponded to different TFs at different time points. The TFs specific to 12 hours all belonged to the AP-1 family, including FOS, JUN, JUNB, FOSL1, and FOSL2. Notably, the top 10 TFs obtained from DAR enrichment analysis (such as JUNB, FRA2, ATF3, FOSL2, FOS, and AP-1) also belonged to the AP-1 family. Furthermore, pioneer transcription factors may play a crucial role in the aging process by binding to tightly bound chromatin regions, opening chromatin structures, and recruiting other TFs. Based on ATAC-seq data, we used protein interaction quantitation (PIQ) to calculate chromatin dependence (CD) and the chromatin opening index (COI), identifying AP-1 as a potential pioneer factor for binding to newly opened chromatin regions during HDACi treatment. In summary, transcriptomic and epigenetic analyses consistently demonstrate that the AP-1 family significantly enhances regulatory activity during HDACi-induced aging.

[0133] 3. FOSB-mediated aging reprogramming plays a key role in lung cancer.

[0134] To assess the relationship between the regulatory activity of AP-1 family TFs and the survival status of TCGA-LUAD patients, we performed regulator activity analysis on 11 members of the AP-1 family. Survival analysis showed that among the 11 candidate regulators of the AP-1 family, only the FOSB regulator phenotype was significantly associated with overall survival (OS), and patients with the activated FOSB regulator phenotype had a better survival prognosis (P = 0.026). Simultaneously, FOSB regulator activity was positively correlated with favorable clinicopathological features (e.g., AJCC tumor stage) and CS score. Furthermore, Human Protein Atlas (HPA) immunohistochemical results showed that FOSB protein expression levels in lung tumor tissues were lower than in normal tissues. Based on CPTAC proteogenomic data, we validated FOSB protein expression levels in 213 LUAD patient samples. The results showed that FOSB protein abundance in LUAD tissues was significantly lower than in normal tissues (P < 2.22e-16).

[0135] To investigate the function of FOSB in aging, we knocked down FOSB expression in A549 cells using siRNA. RT-qPCR and Western blotting confirmed a significant decrease in FOSB expression. Compared to the control group, treatment with TSA or SAHA for 12 hours significantly increased SA-β-gal activity, but this increase was significantly inhibited by FOSB knockdown. These results indicate that FOSB knockdown can attenuate HDACi-induced aging characteristics, suggesting that FOSB may act as a novel aging-driving molecule involved in the development and progression of LUAD.

[0136] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0137] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.

Claims

1. A method for constructing a predictive cellular senescence model, characterized in that, The method includes: S101, Obtain the training set sample expression profile dataset, which includes normal samples from senescent and non-senescent cell lines; S102, using a feature selection algorithm to identify a set of key aging genes from the training set sample expression profile dataset; S103, the key aging gene set is input into the machine learning model to fit the prediction model, and the optimal hyperparameters are determined based on the LOOCV framework to obtain a cell aging model containing the weights of individual genes in the key aging gene set; the cell aging model calculates the aging score by the sum of the products of the expression level of a single gene and its regression coefficient, and the training set samples are divided into high group and low group according to the threshold set by the aging score.

2. The method for constructing a predictive cellular senescence model according to claim 1, characterized in that, The cell senescence model is specifically as follows: Where Ei is the expression level of a single gene in the key aging gene set, βi is the regression coefficient of the corresponding gene, and Score is the aging score. The senescent cell lines in the normal sample expression profile dataset are naturally senescent; senescence-inducing conditions include: replicative senescence, DNA damage induction, gene editing induction, oncogene activation, and oxidative stress.

3. The method for constructing a predictive cellular senescence model according to claim 1, characterized in that, The machine learning model includes any one or more of the following: linear mixed effects model, logistic regression, ridge regression, Lasso regression, elastic network regression, random forest, extreme gradient boosting, support vector machine, artificial neural network and partial least squares regression.

4. A method for predicting the aging state of a tissue sample, characterized in that, The method includes; S201, to obtain the expression levels of key senescence gene sets in normal or tumor tissue samples; S202, input the expression level into the cell senescence model constructed by the method of any one of claims 1-3, calculate the senescence score and divide the tissue sample into a high group or a low group according to the senescence score; if it is a high group, output an auxiliary prediction result that the tissue sample is in a senescent state with a high probability; if it is a low group, output an auxiliary prediction result that the tissue sample is in a senescent state with a low probability. The key aging gene set includes: LBR, CCDC8, TMPO, ZNF22, HNRNPF, DAZAP1, PTMA, KHDRBS1, RBBP4, RPL22L1, N4BP2, CBX3, SMC4, OSTM1, BCAP31, RAB6B, HAGH, and LRPAP1.

5. The method for predicting the aging state of tissue samples according to claim 4, characterized in that, If the tissue sample is tumor tissue, the method further includes: predicting the overall survival of the corresponding subject and / or predicting the tumor purity of the tissue sample based on whether the tissue sample belongs to a high or low group; specifically, if it is a high group, outputting an auxiliary prediction result indicating a shorter overall survival of the tissue sample; if it is a low group, outputting an auxiliary prediction result indicating a longer overall survival of the tissue sample; if it is a high group, outputting an auxiliary prediction result indicating a lower tumor purity of the tissue sample; if it is a low group, outputting an auxiliary prediction result indicating a higher tumor purity of the tissue sample. Optionally, the tumor tissue types include: brain cancer, breast cancer, leukemia, liposarcoma, liver cancer, lung cancer, melanoma, osteosarcoma, pancreatic cancer, and colorectal cancer.

6. A method for screening potential therapeutic drugs, characterized in that, The method includes; S301, Obtain the genomic dataset of tumor samples containing survival data and the drug to be tested; S302, input the dataset into the cell senescence model constructed by the method of any one of claims 1-3 and divide the samples into high-group or low-group; S303, by comparing the high and low groups, the first differentially expressed gene was selected; The correlation between the first differentially expressed gene and survival time was assessed, and the second differentially expressed gene was screened out. Optimize the second differentially expressed gene to obtain the target gene with a positive variable importance value; S304, use feature matching analysis to process the drug to be tested and the target gene, and calculate the feature matching score; The preliminary ranking of the drugs to be tested is obtained based on the feature matching score; S305, Rank clustering analysis based on order statistics is used to analyze the preliminary ranking results and obtain the final ranking results; S306, potential therapeutic drugs are obtained based on the final sorting results.

7. The method for screening potential therapeutic agents according to claim 6, characterized in that, The feature matching analysis method includes any one or more of the following: KS test, extreme value summation and reverse gene expression score.

8. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Composition for evaluating senescence characteristics of cells in tissue and application of composition

    CN114507732A

  • Intelligent combined old medicine selection method

    CN114765067A

  • Construction method and application of senile breast cancer senescence scoring model

    CN117746983A

  • Intelligent tumor microenvironment analysis and drug screening method based on multi-modal data

    CN120319504A

Cited By

  • Pharmacodynamic prediction method, device and kit based on expression profile of few genes

    CN115905898A

  • Method, device and kit for predicting efficacy based on expression profile of small number of genes

    CN115905898B

  • Cell state rapid evaluation method and application thereof

    CN121483363A