A reprogramming factor composition and its application in reversing aging and a quantitative evaluation method based on an extended gene set
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-11
AI Technical Summary
然而,现有技术存在以下关键缺陷:1)缺乏量化评估:难以客观和敏感的比较不同方案的“生物学年龄”逆转效果;2)安全与效能失衡:传统OSKM组合因含有c-Myc而致癌风险高;而仅含核心三因子(OSK)的组合虽安全,但多项研究提示其抗衰老效能有限;3)缺乏精准调控标准:领域内尚未明确,达到何种程度的表观遗传年龄逆转既能产生显著寿命获益,又能确保绝对安全
本发明还提供了安全-有效-精准的逆转衰老并延长寿命的方案,包括重编程因子组合物(Oct4、Sox2、Klf4、Glis1和Lin28联合GLP-1 受体激动剂)及其递送系统。
Smart Images

Figure CN122537512A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of regenerative medicine and epigenetics, and particularly to a reprogramming factor composition and its application in reversing aging, as well as a quantitative assessment method based on an expanded gene set. Background Technology
[0002] Contemporary biomedical aging research is ushering in a paradigm shift: the definition of aging has shifted from the traditional "accumulation of wear and tear" to the loss of epigenetic information, clarifying its essence as a systemic and reversible functional decline of the body driven by epigenomic noise.
[0003] This invention unexpectedly discovered that among the 28 million to 30 million methylation sites in the whole genome, 100 key methylation regions associated with the PRC2-H3K27me3 axis are the core regulatory units of aging. The average methylation level and normalized methylation entropy of this region directly determine the degree of aging. By resetting the epigenome order through molecular intervention, the biological age can be reversibly reversed.
[0004] Meanwhile, aging is a major risk factor for most chronic diseases. Although induced pluripotent stem cell (iPSC) technology can completely reset the biological age of cells, its induced complete reprogramming leads to loss of cell identity and an extremely high risk of tumorigenesis, making it unsafe to use.
[0005] The "partial reprogramming" strategy aims to reverse aging phenotypes without inducing pluripotency by temporarily expressing reprogramming factors. However, existing technologies suffer from the following key drawbacks: 1) Lack of quantitative assessment: It is difficult to objectively and sensitively compare the "biological age" reversal effects of different approaches; 2) Imbalance between safety and efficacy: Traditional OSKM combinations have a high risk of cancer due to the presence of c-Myc; while combinations containing only the core three factors (OSK) are safe, multiple studies suggest that their anti-aging efficacy is limited; 3) Lack of precise regulatory standards: It is not yet clear in the field what degree of epigenetic age reversal can produce significant lifespan benefits while ensuring absolute safety. Existing approaches often blindly explore between "ineffectiveness" and "high risk," failing to achieve precise intervention; 4) No optimized reprogramming factors have been found. Therefore, while partial reprogramming strategies such as OSK / OSKGL have anti-aging potential, they face three major challenges: difficulty in controlling the degree of reprogramming (insufficient intensity cannot effectively reset, excessive intensity can easily lead to teratomas or death), lack of quantitative monitoring tools, and the inability of existing tests to achieve comprehensive quantitative safety evaluation of epigenetics. These challenges severely restrict clinical translation. Current anti-aging research also urgently needs a unified epigenetic evaluation framework, as there are methodological shortcomings such as a lack of standardized quantitative indicators, low-cost high-throughput detection platforms, and quantitative models for indicators and biological endpoints. Summary of the Invention
[0006] The purpose of this invention is to provide the application of low-tumorigenic OSKGL nucleic acid compositions in reversing aging and prolonging lifespan. It also aims to establish a quantitative assessment method for evaluating the effects of reprogramming factor compositions on reversing aging and prolonging lifespan. This method is based on a targeted methylation sequencing platform (TEESM-seq: Targeted EZH2 / EED / SUZ12 Methylation Sequencing) that co-targets low-methylated regions (EZH2 / EED / SUZ12 Low-Methylated Regions, EES-LMRs), used to evaluate the partial reprogramming safety window, dose-response quantification, lifespan benefit assessment, ex vivo organ perfusion quality control, and large-scale anti-aging drug screening.
[0007] Meanwhile, to further enhance the coverage of organ-specific aging phenotypes by the epigenetic assessment system of this invention, this invention integrates aging-related genes screened based on multi-organ MRI bioage gaps (MRIBAGs) on the basis of the original PRC2-100 gene set, and constructs an extended PRC2 core gene set (PRC2-150).
[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution: The present invention relates to the use of a reprogramming factor composition in the preparation of reagents for improving aging-related phenotypes and obtaining lifespan-related benefits, the reprogramming factor composition comprising five factors: Oct4, Sox2, Klf4, Glis1, and Lin28, as well as a GLP-1 receptor agonist.
[0009] Preferably, the reprogramming factor composition is delivered via an AAV vector or an mRNA-LNP vector.
[0010] Preferably, the expression molar ratio of Oct4, Sox2, Klf4, Glis1 and Lin28 is 2~3:1:1:1:1.
[0011] The preferred GLP-1 receptor agonist is selected from Semaglutide, Liraglutide, or Dulaglutide.
[0012] The present invention also provides a pharmaceutical composition comprising the reprogramming factor composition and a pharmaceutically acceptable carrier.
[0013] Preferably, the composition is used to achieve an epigenetic age reversal rate of 47% to 77% in test subjects.
[0014] This invention also provides a quantitative assessment method for evaluating the effects of reprogramming factor compositions on improving aging-related phenotypes and obtaining lifespan-related benefits, comprising the following steps: (1) Obtain genomic DNA from the test subject sample, use a nucleic acid probe designed for the PRC2 trisubunit co-binding hypomethylation region of 100 PRC2 core genes for hybridization capture enrichment, and perform bisulfite treatment and sequencing on the enriched product to obtain sequencing reads covering the hypomethylation region. (2) Filter the sequencing reads and retain the valid reads that meet the preset quality threshold and cover at least 3 CpG sites; (3) Calculate the average methylation level MML and normalized methylation entropy NME of the region set based on the effective reads; (4) Calculate the epigenetic clock value according to the fusion function Clock = a·MML + (1-a)·NME, where a is a weight parameter between 0.5 and 0.8; (5) Calculate the epigenetic age reversal rate R based on the Clock values of the control young group, the control old group, and the old intervention group; (6) Compare R with a preset threshold range, and output a risk control determination result when R falls into the range.
[0015] Preferably, the genes targeting the PRC2 trisubunit co-binding hypomethylated regions mentioned in step (1) further include 62 organ aging-related genes; the 100 PRC2 core genes (Appendix 1) and the 62 organ aging-related genes (Appendix 2) constitute the extended PRC2 gene set PRC2-150. The PRC2-100 gene set can be replaced with a set of genes that satisfy the significant enrichment of the PRC2 target and have a consistent distribution change in the aging / intervention control, preferably ranging from 50 to 300.
[0016] Preferably, in step (6), the preset threshold range is preferably 47% to 77%; when R < 47%, an insufficient reprogramming strength prompt is output; when R > 77%, an excessive reprogramming strength risk prompt is output.
[0017] Preferably, the threshold range of 47% to 77% can be obtained by calibration using a training set according to different tissue types, species, or sequencing depths.
[0018] Preferably, the formula for calculating R is: R = (Clock_Old-Clock_Treat) / (Clock_Old-Clock_Young)×100% or its equivalent normalized form.
[0019] Preferably, the method further includes mechanism validation based on Jensen-Shannon distance for the 100 PRC2 core genes and the 62 organ aging-related genes: defining the promoter regions and gene body regions of the 100 PRC2 core genes and the 62 organ aging-related genes respectively; calculating the average JSD value JSD_promoter in the promoter region and the average JSD value JSD_genebody in the gene body region for the elderly intervention group and the elderly control group based on the methylation level distribution of the effective reads; constructing the PRC2 structural remodeling index PRI = α·JSD_promoter + β·JSD_genebody, where α=0.6 and β=0.4 are weighting coefficients; when PRI is greater than the preset threshold of 0.3 and R falls within the 47%~77% range, the mechanism-specific remodeling judgment result is output: the intervention is effective and safe.
[0020] Preferably, steps (3) to (6) are replaced or assisted by an evaluation process based on an attention mechanism neural network: based on the effective reads, a methylation probability distribution vector is constructed in each target region according to a preset methylation level range, and the methylation probability distribution vectors of multiple target regions are concatenated to form a methylation probability distribution matrix of the region set; the methylation probability distribution matrix is input into a deep neural network model containing a multi-head self-attention mechanism to obtain the epigenetic embedding vector of the sample, and the epigenetic clock value Clock′ and reversal ratio R′ are output, and R′ is compared with a preset threshold range of 47% to 77%.
[0021] Preferably, when the deep neural network model of the self-attention mechanism includes Transformer, linear attention network, state space model or a combination thereof, it is used to capture long-range dependencies across regions and sites within the region set and output an attention weight map; based on the attention weight map, the attention risk index ARI is calculated, and when the attention weights show abnormal concentration or abrupt distribution on the preset abnormal site set, even if R′ falls into the threshold range, a risk warning signal is still output.
[0022] Preferably, when the deep neural network model includes a multimodal network with a cross-attention mechanism, it is used to jointly analyze the methylation probability distribution matrix and the H3K27me3 signal in the corresponding region, and output the attention consistency index (ACI); when both the ACI and the PRI constructed based on the Jensen-Shannon distance meet a preset threshold, a high-confidence judgment result of mechanism-specific remodeling is output.
[0023] Preferably, the methylation level distribution is achieved by dividing the methylation level of the read segment into five intervals: [0, 0.2), [0.2, 0.4), [0.4, 0.6), [0.6, 0.8), and [0.8, 1.0].
[0024] Preferably, the H3K27me3 ChIP-seq signal intensity of the target region is extracted to form a second modality input vector, which is then input into a deep neural network model with a self-attention mechanism.
[0025] The present invention also provides a nucleic acid probe composition for capturing PRC2-targeted hypomethylated regions, the probe composition comprising nucleic acid probes capable of specifically hybridizing with a set of selected PRC2 trisubunits co-binding hypomethylated regions, with a total target size of 1-5 Mb, covering the PRC2-targeted hypomethylated regions corresponding to the 100 core genes.
[0026] Preferably, the set of regions is determined by PRC2 trisubunit ChIP-seq co-peaks, hypomethylation characteristics, and differential methylation screening in aging / intervention controls.
[0027] Preferably, the nucleic acid probe further comprises nucleic acid probes designed for the 62 organ aging-related genes, with a total target size of 4-5 Mb, covering the PRC2-targeted hypomethylated region corresponding to the PRC2-150 gene set.
[0028] The present invention also provides a detection kit for evaluating the effect of improving aging-related phenotypes, comprising: (1) the nucleic acid probe composition described above; and (2) a computer-readable storage medium having stored thereon a data processing program for performing the steps of the quantitative evaluation method.
[0029] The present invention also provides an assessment system for evaluating the effect of improving age-related phenotypes, comprising: a sequencer for sequencing a set of PRC2 trisubunit co-binding hypomethylated regions of a selected sample of a test subject and outputting sequencing data; a processor configured to execute the steps of the quantitative assessment method; and an output device for displaying R value, PRI value, risk-controlled determination result, mechanism-specific remodeling determination result, and a prompt indicating whether the R value falls within a preset window.
[0030] Preferably, the processor stores or can call parameters of a deep neural network model containing a self-attention mechanism to output epigenetic embedding vectors, R′ values, attention weight maps, or attention risk index ARI.
[0031] The present invention has the following beneficial effects: The present invention also provides a safe, effective and precise approach to reversing aging and extending lifespan, including a reprogramming factor composition (Oct4, Sox2, Klf4, Glis1 and Lin28 in combination with a GLP-1 receptor agonist) and its delivery system.
[0032] Studies have confirmed that the PRC2 complex composed of EZH2 / EED / SUZ12, targeting low-methylation regions (LMRs), possesses four core advantages: a clear mechanism, high sensitivity to reprogramming, strong cross-tissue robustness, and good stability over a wide region. It is an ideal epigenetic biomarker for anti-aging, but currently, there is no dedicated low-cost, high-throughput detection platform for this region. Furthermore, this invention unexpectedly discovered that the loss of epigenetic information related to aging has a dual-dimensional characteristic: it is associated with both the average methylation level (MML, reflecting intensity abnormalities) in the PRC2 region and the normalized methylation entropy (NME, reflecting disorder caused by pattern confounding and drift at the cellular / sample level). Recognizing the inadequacy of a single indicator dimension, this invention proposes a joint assessment system of MML and NME to simultaneously capture intensity shifts and order drifts, significantly improving interpretability and discriminative power, and is suitable for aging assessment, intervention monitoring, and drug screening. Based on this, this invention combines research on extended remaining lifespan in aged animals, deeply integrating barcoded transposon high-throughput library construction technology with low-methylation regions (EES-LMRs) co-regulated by the PRC2 core subunits (EZH2 / EED / SUZ12), condensing 100 core aging genes, and constructing a new generation of epiaging assessment system through dual-dimensional analysis of mean methylation level (MML) and normalized methylation entropy (NME). This system, centered on the TEESM-seq platform, combines mechanism interpretability (anchoring cellular identity entropy increase nodes) with economic feasibility (single sample cost < $3), overcoming the core contradiction between "mechanism accuracy" and "detection cost / scale" in personalized longevity intervention. This invention discloses the methods, systems, and applications of this platform in determining the safety window for partial reprogramming, dose-response quantification, lifespan benefit assessment, quality control of ex vivo organ perfusion, and large-scale drug screening, clarifying its core uses in key anti-aging scenarios.
[0033] This invention also establishes a dual quantitative evaluation system based on PRC2 targeting hypomethylated regions. This system achieves precise quantification of intervention effects and risk control through the joint determination of two indicators: "epiggenetic clock monitoring safety window (47%~77%) + JSD validation mechanism depth (PRI ≥ 0.3)", providing an engineerable optimal solution for different application scenarios.
[0034] This invention also calculates epigenetic clock values using an attention mechanism deep neural network model. Unlike existing linear epigenetic clocks or general machine learning models, the attention mechanism deep neural network model described in this invention is a specific technical solution with a structured design tailored to the sparsity, noise distribution, and long-range dependence of PRC2 regulatory regions in bisulfite sequencing data. Through cross-regional information aggregation via a self-attention mechanism, the model can achieve denoising and robust inference of missing data under low sequencing depth conditions, thereby reducing the amount of sequencing data required to achieve equivalent discriminative power to less than 1 / 5 of traditional methods while maintaining prediction accuracy. Furthermore, the attention weight map output by the model provides interpretable, site-specific anomaly pattern monitoring capabilities, capable of identifying early tumorigenic risks that traditional single numerical indicators (such as the epigenetic age reversal ratio R) cannot detect. This upgrades the safety control mechanism of this invention from "post-hoc threshold determination" to "process risk monitoring," solving the long-standing technical problem in some reprogramming fields of "inability to identify high-risk samples within the window." Therefore, this invention provides verifiable improvements in sequencing data processing efficiency, robustness under low-depth conditions, and safety early warning capabilities, exceeding the scope of general mathematical methods and representing an improvement to the functionality of the computer data processing system itself. Consequently, while retaining the interpretive frameworks of MML / NME and JSD / PRI, this invention further introduces an attention mechanism neural network, enhancing robustness to low-sequencing-depth noise and sensitivity to early remodeling signals, and upgrading safety control from a single threshold window to a triple closed loop of "window determination + attention risk early warning + mechanism consistency verification".
[0035] The core findings of this invention are: focusing the tens of millions of methylation sites across the entire genome onto the methylation regulatory regions of 100 key aging genes (PRC2-H3K27me3), and deeply integrating barcoded transposon high-throughput library construction technology with biological targets of low-methylation regions (LMRs) co-regulated by EZH2 / EED / SUZ12, constructing a new generation of epiaging assessment system that combines mechanistic interpretability (anchoring the core regulatory nodes of cell identity entropy increase, namely the two-dimensional calculation of average methylation and normalized methylation) with economic feasibility (<$3 / sample).
[0036] The core discovery of this invention lies in the fact that, based on the above, there is a clear safe therapeutic window for the epigenetic age reversal range of anti-aging interventions: 47% to 77%. When the reversal range is < 47%, the intervention is insufficient to activate key aging reversal pathways and cannot produce a significant lifespan extension effect; when the reversal range is > 77%, the intervention may excessively disturb the cellular epigenomic homeostasis, accompanied by a significant increase in tumor incidence; only when the reversal range is strictly controlled within the 47% to 77% range can the intervention simultaneously achieve significant lifespan extension and baseline tumor safety. Based on this, this invention defines and verifies for the first time the "safe therapeutic window for epigenetic age reversal" necessary to achieve safe and effective anti-aging, and constructs a dual quantitative assessment system with "remaining lifespan," providing precise quantitative safety standards for all partial reprogramming interventions, simultaneously achieving lifespan extension and risk control, and propelling anti-aging research and development from empirical exploration to a precise quantitative engineering era. Attached Figure Description
[0037] Figure 1 The linear correlation between the TEESM-seq age index and actual age; Figure 2 The results of age prediction error analysis for different epigenetic clocks; Figure 3 For cross-tissue stability comparison of different epigenetic clocks; Figure 4 The effects of intervention on mice with different epigenetic clocks; Figure 5 Figure A shows the DNA methylation remodeling state of PRC2 under OSKGL intervention; Figure B shows the distribution of methylation characteristics in the PRC2 target region by principal component analysis (PCA); Figure C shows the box plot of mean methylation level (MML); Figure D shows the box plot of normalized methylation entropy (NME); Figure D shows the two-dimensional joint remodeling trajectory of MML and NME. Figure 6 The response curve of the apparent age reversal rate (%) as a function of the degree of partial reprogramming (including the 47%–77% safety window). Figure 7 The curves showing the changes in the probability of tumor risk and death risk as a function of partial reprogramming (i.e., the safety window risk function curve). Figure 8 Survival curves for each group (Kaplan-Meier); Figure 9 The percentage increase in median remaining lifetime for each group; Figure 10 A linear regression model was used to compare the epigenetic age reversal rate R with the epigenetic age decline ΔE (weekly equivalent) (fitted equation: ΔE = 0.92R - 0.02, R...). 2 =0.96); Figure 11 The curves show the changes in apparent age reversal with perfusion time (hours) in the OSKGL+GLP group, OSKGL single drug group, and blank perfusion group during NEVLP perfusion of isolated liver.
[0038] Figure 12 To reconstruct the Jensen-Shannon distance (JSD) heatmap based on the methylation distribution of promoter and gene body regions in the PRC2-100 core gene set; Figure 13 Scatter plot of individual distribution of reversal ratio R in OSKGL-superwindow group (n=10); the horizontal dashed lines in the figure mark the upper limit of the safety window (R=77%) and the supraphysiological reset threshold (R=90%), respectively, and the solid circles represent the measured R values of each individual, with the group mean ± SEM=82.4±5.1%.
[0039] Figure 14 A comparison of tumor incidence rates between the extra-window intervention group and the safe window intervention group; Figure 15 The difference in the distribution of the remodeling index (PRI) of methylation distribution in the target region of PRC2 between the over-window intervention group and the safe window intervention group (n=10 / group). Figure 16 Heatmap of cancer-related gene expression for tumor overprogramming and OSKGL quantitative partial reprogramming; Figure 17 This is a schematic diagram illustrating the organization-specific safety window and the risk threshold beyond the window. Detailed Implementation
[0040] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0041] Example 1: TEESM-seq Clock Establishment: Performance, Stability, Accuracy, and Intervention Sensitivity
[0042] Module 1: Animal Models
[0043] Experiments were conducted using AAVs carrying OSK polycistronic cassettes and reverse tetracycline activator (rtTA) and AAVs carrying OGL polycistronic cassettes and reverse tetracycline activator (rtTA).
[0044] Four sets of experiments: Young control: 3-month-old mice; Old (untreated): 22-month-old mice; Older OSK management (Old+OSK): Administer OSK AAV from 15 months of age to 22 months of age.
[0045] Older OSKGL treatment (Old+OSKGL): administer OSKGL dual AAV from 15 months of age to 22 months of age.
[0046] Treatment protocol: 1 × 10¹² viral genome copies (1 × 10¹² vg / mouse) per mouse, total volume 100 μL, were injected via the retroorbital vein. OSK or OSKGL expression was induced periodically (doxycycline induction for 3 days, followed by a 4-day rest period, specifically achieved by adding doxycycline to drinking water to a final concentration of 2 mg / mL). Young control mice and aged untreated mice were injected with 100 μL of formulation buffer (phosphate-buffered saline, PBS).
[0047] Organizational collection: After processing, tissues such as skin, liver, and heart were taken and flash-frozen in liquid nitrogen for preservation.
[0048] Module 2: Whole Genome Bisulfite Sequencing (WGBS)
[0049] 1. Genomic DNA extraction: gDNA was extracted from liver tissue using the MasterPure DNA purification kit (Biosearch Technologies, MC85200), and the integrity of the gDNA was verified by agarose gel electrophoresis.
[0050] 2. Construction of the WGBS Library
[0051] 1) Sample preparation: Take 500 ng gDNA and add 1% unmethylated λDNA (Promega, D1521) to monitor the bisulfite conversion efficiency; 2) Ultrasonic disruption: The gDNA was disrupted to an average insert size of 350 bp using a Covaris S220 focused ultrasound system (parameters: duty cycle 10%, intensity 5, 200 cycles per pulse, time 60 s). 3) Fragment selection: Use AMPure XP magnetic beads (Beckman Colter, A63882) to select fragments of 300~400 bp; 4) Bisulfite conversion: Performed according to the instructions using the EZ DNA Methylation-Gold Kit (Zymo Research, D5005); 5) PCR amplification: Amplification was performed using the Kapa Hifi Uracil+ kit (Roche, KK2801). Cycling conditions: 98℃ for 45 s; 98℃ for 15 s, 65℃ for 30 s, 72℃ for 30 s, 8 cycles; 72℃ for 1 min. 6) Library quality control: The Agilent 2100 Bioanalyzer (High Sensitivity DNA Kit, 5067-4626) was used to detect library quality, and the Kapa Library Quantification Kit (Roche, KK4824) was used in conjunction with the 7900HT Real-Time PCR System (Thermo Fisher) for library quantification.
[0052] 3. High-throughput sequencing
[0053] Illumina HiSeq4000 platform, 150 bp paired-end sequencing, with 5% non-indexed PhiX control (Illumina, FC-110-3001). Transformation efficiency: The average bisulfite transformation efficiency of unmethylated λDNA reached 99.5%.
[0054] Module 3: Native ChIP-Seq
[0055] 1. Nucleus isolation
[0056] Take 100 mg of frozen liver and homogenize it at 4°C using nuclear separation buffer (10 mM HEPES, 10 mM KCl, 1.5 mM MgCl2, 0.34 M sucrose, 10% glycerol, 1 mM DTT) on a GentleMACS dissociator (Miltenyi Biotec, 130-093-235) with the "4C_nuclei_1" program. Collect the nuclear pellet by centrifugation.
[0057] 2. Chromatin MNase digestion
[0058] 1×10 6 Resuspend cell nuclei in 200 μL MNase buffer (50 mM Tris-HCl, 5 mM CaCl2), add 400 U MNase (NEB, M2347S), and incubate at 37°C and 750 rpm with shaking for 10 min. Terminate the reaction with 10 mM EDTA, then add 200 μL nuclear lysis buffer (25 mM Tris-HCl, 150 mM NaCl, 1% Triton X-100, 1% sodium deoxycholate, 0.1% SDS) to lyse the nuclei. Centrifuge at 18000×g and 4°C for 15 min, and collect the supernatant (digested chromatin).
[0059] 3. Immunoprecipitation (IP)
[0060] 1. Staining quality control: Take a small amount of chromatin and use BioAnalyzer and Qubit HS kit to detect digestion efficiency and concentration. Take 2 μg of chromatin for IP, and keep 10% as input. 2. Internal additives: Commercially available Drosophila chromatin (Active Motif, 53083) was added as a spike-in for quantitative normalization; 3. Antibody incubation: H3K27me3 antibody (CST, 9733S) was conjugated with Protein A / G magnetic beads at a 1:1 ratio (ThermoFisher, 10001D / 10003D) and incubated overnight at 4°C by rotation. 4. Washing and elution: Wash 3 times with low-salt washing buffer and 1 time with high-salt washing buffer; incubate with elution buffer (100 mM NaHCO3, 1% SDS) at 65℃ for 15 min and elute twice.
[0061] 4. Library construction and sequencing
[0062] 1) Nucleic acid purification: The elution buffer was treated with RNase A and proteinase K (Thermo Fisher) overnight at 65°C, and DNA was extracted using a ChIPDNA purification kit (Zymo Research, D5205); 2) Library construction: Take 5 ng of DNA and construct a library using the NEBNext Ultra II Library Kit (NEB, E7645S) with double-indexed primers; 3) Sequencing: Illumina NovaSeq X Plus platform, paired-end sequencing, all buffers contain 1 mM PMSF (CST, 8553S).
[0063] Module 4: Bioinformatics Analysis
[0064] All analyses were based on the mouse mm10 genome (GENCODE M23), and the analysis workflow was divided into the following modules: 1. Raw data processing: TrimGalore v0.6.10 was used to remove adapters and low-quality reads (Q<20), and FastQCv0.11.2 was used for sequencing quality control; 2. Sequence alignment: The reads were aligned to the mm10 genome using Bismark v0.22.2 to generate an M-bias map. 5 bp was truncated from Read1 and 20 bp from Read2 to eliminate methylation bias. 3. Data filtering: Use SAMtools v1.19 to sort, deduplicate, and index the BAM files, removing read segments with a comparison quality <10.
[0065] Module 5: Establishment of the TEESM-seq Technology Platform
[0066] This platform is a low-cost, high-throughput detection system specifically designed for the quantitative measurement of "cellular identity entropy" and "physiological partial reprogramming".
[0067] Overall technical route: Sample DNA → [1] Barcode transposition → [2] Mixed sample merging → [3] End repair → [4] Hybridization enrichment → [5] Magnetic bead capture → [6] Bisulfite conversion → [7] PCR amplification → [8] Sequencing → [9] Bioinformatics analysis →
[10] PC2 entropy-level fusion epigenetic age clock calculation →
[11] TEESM-Seq output.
[0068] Specifically as follows: 1. Core Design Concept Using a liquid-phase hybridization enrichment and capture system, the Top 1000 High-EES LMRs (high EZH2 / EED / SUZ12 bound low-methylation regions) defined by TEESM-Seq (Targeted Methylation Assay Platform) were selectively enriched and captured. Random CpG sites were no longer detected: Traditional scatter plot CpG was abandoned, focusing instead on LMR blocks with clearly defined regulatory functions. Linear regression was no longer relied upon: The methylation density of these regions was directly calculated as an absolute reading of biological age.
[0069] 2. Detailed Technical Solution
[0070] A. Bait Library Design (This is the most critical customization part of this method)
[0071] 1) Target Region Definition: LMRs of the 100 core genes and related genes with the strongest binding signals to EZH2, EED, and SUZ12 in hESC (H1, CHB8 cell lines) and aged mice treated with OSKGL for reverse aging were selected. Average length: approximately 3kb per LMR, total target size: 1000 × 3kb = 3 Mb (this is much smaller than the 50Mb of the whole exome, meaning extremely low sequencing cost). A list of the 100 core genes is shown in Appendix 1.
[0072] 2) Probe synthesis: A DNA template was designed for this 3Mb region, and the region was delivered to the commercialization company (IDT) for the synthesis of biotinylated RNA probes. Since the EES region is rich in CpG (CpG islands), the GC content balance of the probes needs to be considered during the design, and a tiling design is adopted with a coverage of at least 2x.
[0073] 3. Species compatibility: If cross-species compatibility is required (e.g., human-mouse compatibility), a hybrid probe library (Conserved PRC2 LMRs) can be designed for the highly conserved PRC2 regions of the two species.
[0074] B. Wet Lab Workflow
[0075] Genomic DNA extraction, quantification and standardization: See Module 2 above, "1. Genomic DNA extraction".
[0076] 1) TEESM-seq library construction: Tn5 restriction enzyme digestion and barcoding (Tagmentation & Indexing): Each sample was reacted individually using Tn5 transposase complexes carrying different i5 / i7 barcodes. Here, "sample" refers to genomic DNA extracted from tissues of the four groups of experimental mice in Module 1, with each gDNA sample constituting an independent sample. The actual number of samples in this embodiment is four groups (Young, Old, Old+OSK, Old+OSKGL) × 8 mice per group × 3 tissue samples collected (skin, liver, heart), totaling 96 samples. Innovation: All samples are physically labeled and can be immediately mixed.
[0077] 2) Sample Pooling: 96 samples are pooled into one tube. All subsequent steps (bisulfite conversion, PCR, hybridization) are performed in one tube, greatly reducing reagent costs and operational errors.
[0078] 3) Methylation end repair and PCR pre-amplification: 5-methyl-dCTP was used to repair the ends to prevent the loss of adapter sequence information during bisulfite conversion.
[0079] 4) Hybridization Capture: The EES-LMR probe designed in step A was hybridized with the mixed DNA library. Objective: To capture only the 1000 3kb regions most relevant to aging mechanisms and wash away the remaining 99.9% of irrelevant genomic DNA.
[0080] 5) Bisulfite Conversion: Converting unmethylated C to U. Quality Control (QC): Controlling the bisulfite conversion rate using the Lambda phage DNA inherent in sequencing.
[0081] 6) Final sequencing: The Illumina NovaSeq platform was used. Since the target region is only 3Mb, even to obtain high-depth coverage (e.g., 50x), only about 0.15 Gb of data is needed per sample.
[0082] C. Data Analysis Algorithms (Bioinformatics Pipeline)
[0083] 1. Alignment and splitting: Data is split back to single samples based on Tn5 Barcode and aligned to the reference genome (mm10).
[0084] 2. Quantification of mean methylation (MML): Instead of calculating the β value of individual CpGs, the average methylation level of all CpGs within each LMR region is calculated.
[0085] .
[0086] MML (Mean Methylation Level) refers to the statistical average of the methylation level of each effective read within a predefined target site set (the set of CpG sites related to the PRC2 target gene), and is used to characterize the overall intensity of DNA methylation in the target region.
[0087] To clarify β i The statistical unit and the calculation caliber of MML are consistent with the segment-level calculation of ELRE below, and are further defined as follows: Suppose that there are R effective segments in a certain EES-LMR region after filtering (the filtering criteria are consistent with the calculation step (1) of ELRE below: MAPQ≥20, coverage CpG number≥3); for the i-th effective segment r i Let the total number of CpGs it covers be n. i The number of CpG cells in the methylated state is m. i The methylation level of this read segment is defined as β. i = m i / n i (i = 1, 2, …,R). The β i That is, the l defined in step (2) of the ELRE calculation below. k Both are the same statistic. In this case, MML can be expressed as the arithmetic mean of the methylation levels of R effective reads: MML = (1 / R) × Σβ i = (1 / R) × Σ(m i / n iThis avoids calculating the β value of a single CpG site and ensures that MML and ELRE represent the intensity (first moment) and disorder (Shannon entropy) of the same methylation distribution based on the same set of read statistics.
[0088] 3. EES-LMR Region-specific Entropy (ELRE)
[0089] NME (Normalized Methylation Entropy) refers to the epigenetic "disorder" index obtained by measuring the information entropy distribution of methylation states within a set of preset target sites and normalizing it. It is used to characterize the degree of heterogeneity and drift of methylation patterns in the same target region at the cell / sample level.
[0090] In the EES-LMR region calculation context of the TEESM-seq platform described in this invention, the NME and the ELRE / G-ELRE defined in this section are different levels of naming for the same statistic: the NME value on a single EES-LMR region is the ELRE for that region, and the global NME value after aggregation of multiple regions is the G-ELRE. Here, NME is a general functional term used for a concise expression in the fused clock Clock = a·MML + (1-a)·NME; ELRE / G-ELRE is the naming of the specific calculation implementation of this statistic on the EES-LMR region, used in the methodology description in this section and the examples shown below. All calculation results mentioned below (such as...) Figure 3 The NME values in Examples 3 and 6 were all obtained using the ELRE / G-ELRE algorithm defined in this section.
[0091] To quantify the randomness of methylation at EES-LMR sites without relying on parametric statistical models, we constructed the EES-LMR region-specific entropy (ELRE), calculated as follows: (1) Sequencing data processing: All valid sequencing reads covering the target EES-LMR region were collected, and low-quality reads (MAPQ < 20) and reads covering less than 3 CpGs were filtered to obtain a set of valid reads, totaling R reads.
[0092] (2) Calculation of methylation level of a single read: For each read rk, let the total number of CpG sites it covers be nk, and the number of CpG sites in the methylated state be mk. Then the methylation level lk of this read is defined as: .
[0093] (3) Methylation level classification: Based on the methylation level lk, each read is divided into one of five methylation level intervals, defined as follows: B1=[0,0.2) (completely unmethylated), B2=[0.2,0.4) (low methylation), B3=[0.4,0.6) (mixed), B4=[0.6,0.8) (high methylation), B5=[0.8,1.0] (fully methylated). The number of reads ci (i=1,2,3,4,5) in each interval is counted, and the empirical probability p^i=ci / R is calculated.
[0094] (4) ELRE calculation: Based on the above empirical probabilities, calculate the normalized Shannon entropy, i.e., the EES-LMR region entropy: .
[0095] (5) Multi-regional ELRE aggregation. ELRE values were calculated for all EES-LMR sites in the subject's genome detected by TEESM-seq, and the median or weighted average was used to obtain the global EES-LMR methylation stochasticity index (G-ELRE): .
[0096] G-ELRE can serve as a core indicator for quantifying the degree of epigenetic aging in subjects and the effectiveness of partial reprogramming interventions.
[0097] 4. Construction of Joint Scoring / Fusion Clock
[0098] Constructing the PRC2 entropy-level fusion score (also known as TEEDSM seq): TEEDSM seq Clock=a MML+(1-a) NME; Where 'a' is the weight parameter, which can be a preset constant, a trainable parameter, or a parameter that is adaptively adjusted according to the sample type / scenario. In this embodiment, a = 0.65.
[0099] Calculate the apparent age reversal rate R: R = (Clock_Old-Clock_Treat) / (Clock_Old-Clock_Young)×100%.
[0100] In the formula, Clock_Old represents the TEEDSM seq clock for the untreated (Old) group; Clock_Treat represents the TEEDSM seq clock for the treated (Old+OSK) or treated (Old+OSKGL) groups; and Clock_Young represents the TEEDSM seq clock for the young control group.
[0101] Sources of human specimen data
[0102] The human sample data used in constructing and validating the TEESM-seq epigenetic clock in this invention are all derived from the following authoritative international public databases, and the data is traceable and obtained in a standardized manner: The Gene Expression Omnibus (GEO), maintained by the National Center for Biotechnology Information (NCBI) at https: / / www.ncbi.nlm.nih.gov / geo / , archives raw and preprocessed DNA methylation data from TEESM-seq and control clocks such as Horvath, PhenoAge, GrimAge, and TIME-seq. This invention filters Illumina 450K / EPIC microarray datasets from this database, covering human whole blood and tissue samples, using the keywords "DNA methylation age," "epigenetic clock," and "human tissue." Raw SRA sequencing data, MINiML metadata, and normalized methylation values are downloaded using a unique GEO accession number.
[0103] Genotype-Tissue Expression (GTEx): This project provides multi-tissue DNA methylation data to support cross-tissue stability validation of TEESM-seq epigenetic clocks.
[0104] All the above data are de-identified publicly available data, and the donors have signed informed consent forms. Those skilled in the art can directly obtain and repeat the experiments of this invention based on the above description. Compare with clock detection (Horvath / PhenoAge / GrimAge / TIME-seq).
[0105] Using the same samples as TEESM-seq, and following the standard experimental procedures for each existing clock, library construction, sequencing, and age index / clock value calculation were completed to ensure consistency of detection conditions.
[0106] Experimental results: 1. The TEESM-seq epigenetic clock of this invention exhibits significantly superior technical performance compared to existing technologies (such as Horvath clock, PhenoAge, GrimAge, etc.) in key performance dimensions such as age prediction accuracy, cross-tissue stability, and intervention sensitivity. like Figure 1 As shown, the TEESM-seq age index and chronological age exhibit a very strong linear correlation, with a correlation coefficient R = 0.9736 and p < 1 × 10⁻⁶. -60 (n=100), linear fitting results show that it can highly accurately reflect the true age of individuals; further age prediction error analysis ( Figure 2 The results show that the mean absolute error (MAE) of TEESM-seq is only 3.19±0.30 years, which is significantly lower than that of Horvath (4.52±0.44 years), PhenoAge (6.55±0.63 years) and GrimAge (4.30±0.41 years), achieving the lowest age prediction error and significantly improving accuracy.
[0107] Cross-organizational stability comparison ( Figure 3 The results show that the stability index of TEESM-seq is significantly higher than that of existing clocks, and its data distribution is the most concentrated, which is marked as "optimal". This indicates that it performs more stably and reliably across different organizations and has a wider range of applications.
[0108] In mouse intervention experiments ( Figure 4 TEESM-seq showed the most sensitive response to the intervention, with a significantly more negative Δ clock value (intervention group - elderly control group) and the smallest variance, and was labeled as "strongest effect with smallest variance". At the same time, its signal-to-noise ratio (SNR) was as high as 12.2, which is much higher than Horvath (1.7), PhenoAge (1.2), GrimAge (2.3) and TIME-seq (6.5), and can detect the rejuvenation effect brought about by the intervention more sensitively and stably.
[0109] In summary, the TEESM-seq clock of this invention has achieved significant technological breakthroughs in terms of age prediction accuracy, cross-organizational stability, and intervention sensitivity, providing a better technical solution for the application of epigenetic clocks in aging assessment, intervention effect detection, and other fields.
[0110] 2. The role of the PRC2 target region in the OSKGL intervention group
[0111] This invention provides a genome-wide methylation feature analysis of the PRC2 target region, with four experimental groups: Young, Old, Old+OSK, and Old+OSKGL. Figure 5Principal component analysis showed that the Young group and the Old group were significantly separated in PC1 and PC2 dimensions. The Old+OSK group shifted towards the Young group relative to the Old group, and the Old+OSKGL group further moved towards the cluster center of the Young group. Figure 5 B-mean methylation level (MML) assay showed that the MML in the PRC2 target region of the Old group was significantly higher than that of the Young group. The MML in the Old+OSK group decreased in some areas, while the MML in the Old+OSKGL group decreased further and was closer to the level of the Young group. Figure 5 C-normalized methylation entropy (NME) assays showed that the NME in the PRC2 target region was significantly increased in the Old group, while the methylation entropy in the Old+OSK group was only partially reduced, and the NME in the Old+OSKGL group was significantly restored to levels close to those in the Young group. Figure 5 D establishes a two-dimensional reconstructed trajectory model with MML as the horizontal axis and NME as the vertical axis, presenting a clear trajectory: Young → Old+OSKGL → Old+OSK → Old, where the Old+OSKGL point is closer to the Young coordinate position.
[0112] Conclusion: The OSKGL combination is significantly superior to OSK in reversing aging-related methylation phenotypes, reducing abnormally high methylation levels in the PRC2 target region, and restoring the order of methylation patterns, confirming that OSKGL can achieve synergistic remodeling in both methylation levels and structural order.
[0113] Example 2: In vivo partial reprogramming safety window verification and lifespan extension method
[0114] This embodiment provides an in vivo partial reprogramming lifespan extension method based on PRC2 target epigenetic assessment, aiming to verify the safety window of reprogramming intensity and the synergistic effect of OSKGL combined with GLP1.
[0115] 1. Laboratory animals and grouping
[0116] 104-week-old male C57BL / 6J mice were selected and acclimatized for one week in a standardized animal facility. The mice were then randomly divided into four groups (n≥10 per group): Control group: injected with PBS carrier; OSK group: Injection of AAV9 hybrid vector (AAV9.TRE3-OSK-SV40pA + AAV9-hEf1a-rtTA4-Sv40pA); OSKGL group: AAV9 mixed vector (AAV9.TRE3-OSK-SV40pA + AAV9.TRE3-OGL-SV40pA + AAV9-hEf1a-rtTA4-Sv40pA) was injected. Starting from 15 months of age, doxycycline was administered weekly for 5 days and then stopped for 2 days, continuing until 22 months of age. OSKGL+GLP1 group: The OSKGL group was injected with the AAV9 hybrid vector and combined with subcutaneous injection of the GLP-1 receptor agonist Semaglutide (40 μg / kg, every other day, from the day after AAV injection until the end of the experiment, i.e., 22 months of age, for a total of about 7 months).
[0117] The AAV vector titers were all 1.5~1.9×10¹³ vg / mL, and each mouse was injected with 100 μL of the mixture via the retroorbital vein (total dose 6×10¹³ vg / kg).
[0118] 2. Security Window Definition and Control
[0119] Induction began the day after injection, with doxycycline administered via drinking water (final concentration 2 mg / mL). To control the reprogramming intensity within a safe window, a cyclical pattern of 1 week dosing / 1 week off was adopted.
[0120] In this embodiment, the safe window is defined as the reprogramming intensity range that maximizes the reversal ratio R under the condition of satisfying the risk function constraint, mathematically expressed as:
[0121] 3. Lifetime monitoring and assessment
[0122] Survival status was recorded daily, Kaplan-Meier survival curves were plotted, and median remaining lifespan was calculated. At the experimental endpoint, skin or liver tissue was collected, and the mean methylation level (MML) and normalized methylation entropy (NME) were assessed using the aforementioned PRC2 fusion epigenetic clock detection method to calculate the epigenetic age reversal ratio R.
[0123] Security window verification: Within the range below the window to the upper edge of the window, the reversal ratio increases with the intensity of partial reprogramming. Figure 6 Under ultra-window intensity conditions, the incidence of tumors and the risk of death are significantly increased. Figure 7 Survival curves: The survival curve for the OSKGL group shifted significantly to the right, and the survival curve for the OSKGL+GLP1 group shifted further to the right. Figure 8Median remaining lifetime: Compared to the control group (0%), the GLP1 group showed a 7% increase, the OSK group a 15% increase, the OSKGL group a 78% increase, and the OSKGL+GLP1 group a 91% increase. Figure 9 Predictive model: A linear correlation model was established between the reversal rate R and the epigenetic age decline ΔE (weekly equivalent) ΔE = 0.92R - 0.02 ( Figure 10 This indicates that the reversal ratio can serve as a quantitative predictor of the apparent rate of rejuvenation and life expectancy improvement.
[0124] Example 3: Cellless Normative Mechanical Perfusion (NEVLP) and Epigenetic Rejuvenation Method for Ex vivo Organs
[0125] This embodiment provides an ex vivo liver epigenetic rejuvenation method based on an mRNA-LNP+GLP1 delivery system, which is suitable for organ preservation and functional repair.
[0126] 1. Preparation of perfusion fluid and mRNA-LNP
[0127] Basic perfusion solution: Based on Krebs-Henseleit electrolyte formula, containing 112 mM NaCl, 4.7 mM KCl, 2% BSA and 50 U / mL heparin, pH adjusted to 7.3~7.4, and 95% O2 + 5% CO2 is introduced to maintain pO2 ≥ 200 mmHg.
[0128] mRNA-LNP complex: Chemically modified mRNAs (Oct4, Sox2, Klf4, Glis1, Lin28) were prepared in a molar ratio of 3:1:1:1:1. These were then encapsulated in LNP vectors (SM-102:DSPC:cholesterol:PEG2000-DMG = 50:10:38.5:1.5) using microfluidic technology, with particle sizes controlled between 65 and 125 nm and encapsulation efficiency ≥85%.
[0129] Final perfusion solution preparation: Add the mRNA-LNP complex and GLP1 to 50 mL of basal perfusion solution. The final concentration of the mRNA-LNP complex is 0.2 μg / mL, and the final concentration of GLP1 is 10 nM (approximately 0.042 μg / mL Semaglutide). Preheat at 37°C for later use.
[0130] 2. Infusion Operation
[0131] Liver was harvested from 30–35 g C57BL / 6J mice after isoflurane anesthesia. A NEVLP closed-loop system was established, with portal vein cannulation (2 Fr polyurethane cannula), a flow rate of 2 mL / min, and a portal vein pressure of 7–10 mmHg maintained at 37°C. Perfusion was continued for 12 hours, with pH and electrolyte concentrations of the perfusion fluid monitored every 3 hours.
[0132] 3. Effect Verification (corresponding to) Figure 11 ): Liver tissue was harvested after perfusion, and epigenetic age was assessed using the aforementioned PRC2 fusion epigenetic clock detection method. Results: Compared with the blank perfusion group, the mRNA-LNP perfusion group achieved an epigenetic age reversal rate of 68% within 24 hours. Furthermore, the OSKGL+GLP1 group showed a better reversal trend than the OSKGL group, suggesting that the combined approach is also applicable to in vitro settings. Hepatocyte viability ≥85% confirmed that this method can achieve safe and effective epigenetic resetting in in vitro systems.
[0133] Summary of technical effects
[0134] Through the above embodiments, the present invention has the following technical effects: 1. Quantifiable safety window: It is demonstrated that there is a quantifiable safety window (e.g., 47%–77%) for partial reprogramming, within which significant apparent age reversal can be achieved, while the risk increases significantly beyond the window.
[0135] 2. Risk controllability: The contrast structure of "upper edge inside the window vs. outside the window" constitutes the evidence anchor point for the risk threshold, ensuring the risk controllability of the intervention strategy.
[0136] 3. Feasibility of lifespan prediction: The reversal rate is positively correlated with lifespan extension (L=0.92R-0.02), indicating that the reversal rate can be used as a quantitative predictive indicator of lifespan improvement.
[0137] 4. Superior Intervention Regimen: The OSKGL regimen significantly outperformed the OSK regimen (+15%) in lifespan extension (+78%); the combination with GLP1 further enhanced the effect (+91%). Synergistic effect analysis showed that, based on the Bliss independent model, the expected additive effect of OSKGL monotherapy (+78%) and GLP1 monotherapy (+7%) was 79.5% (calculated as: E_expected = E_OSKGL + E_GLP1 - E_OSKGL × E_GLP1 = 78 + 7 - 78 × 7 / 100 = 79.5%), while the actual combined effect was 91%, with a synergistic ratio of 1.14 and a combination index CI of 0.87 (CI < 1 indicates synergy). This indicates that OSKGL and GLP1 have a super-additive synergistic effect in lifespan extension, rather than a simple additive effect. This synergistic effect showed a consistent enhancing trend in both in vivo and in vitro systems.
[0138] 5. Wide applicability: This method can be applied to in vivo and in vitro systems (68% reversal within 24 hours), and is feasible, repeatable, and has controllable risks.
[0139] In summary, this invention provides a partially reprogrammed lifespan extension method that is feasible, repeatable, and has controllable risks. Furthermore, it provides a joint intervention strategy to enhance the technical effect, effectively resolving the core contradiction between "mechanism accuracy" and "detection cost / scale" in personalized longevity intervention monitoring.
[0140] Example 4: Validation of intervention effect based on PRC2-100 two-region Jensen–Shannon distance (JSD)
[0141] This embodiment aims to quantify the changes in methylation distribution morphology of the PRC2 core target region before and after intervention based on information theory metrics, thereby verifying the epigenetic remodeling effect of the OSKG intervention protocol at the level of mechanism specificity and structural integrity.
[0142] This embodiment introduces the Jensen–Shannon distance (JSD) as a "validation index for epigenetic remodeling depth," and forms a complementary validation system with the "epigenetic clock based on average methylation level (MML) and normalized methylation entropy (NME)" described in Example 1: Clock indicator: used to calculate the apparent age reversal rate R, to determine whether the intervention magnitude is within the preset safe and effective window (47%~77%). JSD index: used to determine whether intervention produces a reshaping of the distribution structure at the core target of PRC2, thereby providing specific and in-depth evidence at the mechanistic level.
[0143] 1. Experimental Grouping and Sample Collection
[0144] The experiment used C57BL / 6J male mice, which were divided into the following groups: Young control group: 3 months old, untreated, n=3; Old group (untreated): 22 months old, untreated, n=5; Older intervention group (Old+OSKGL): 22 months of age, treated with long-term partial reprogramming of OSKGL factor, n=5. The treatment regimen was the same as in Example 2: from 15 months of age, doxycycline was administered weekly for 5 days to induce OSKGL expression, followed by a 2-day break, continuing until 22 months of age; Combined intervention group (Old+OSKGL+GLP1): In addition to OSKGL treatment, the GLP-1 receptor agonist Semaglutide (40 μg / kg, every other day, from the day after AAV injection until the end of the experiment, i.e., 22 months of age) was added, n=5.
[0145] At the end of the experiment, full-thickness skin tissue was collected from the back of mice, flash-frozen in liquid nitrogen, and stored at -80°C for later use.
[0146] 2. Target Region Definition and Sequencing
[0147] Targeted methylation sequencing was performed using the TEESM-seq technology platform described in Example 1. The target region was defined as the 100 core regulatory genes of PRC2 listed in Table 1 (PRC2-100 gene set).
[0148] For each core gene, two independent detection regions are defined: Promoter region: 2 kb upstream to 0.5 kb downstream of the transcription start site (TSS); Gene Body: TSS + 0.5 kb to transcription termination site (TES).
[0149] After library construction, paired-end sequencing (e.g., PE150) is performed, with an average sequencing depth of ≥100× for the target region to enhance the statistical stability of read distribution estimation.
[0150] 3. Jensen–Shannon Distance (JSD) Calculation Method
[0151] This embodiment uses the JSD algorithm based on information theory to quantify the difference in methylation probability distribution between the intervention group and the control group. To ensure feasibility and stability, the calculation process is as follows: (1) Calculation of methylation level in single reads Collect and filter valid sequencing reads covering the target region: Alignment quality: Reads with MAPQ < 20 are removed; CpG coverage: Reads covering < 3 CpG sites are removed; If the number of effective reads R in each target region is lower than a preset lower limit (e.g., R < 50 or R < 100), the region can be marked as missing or have its weight reduced. For each read rk, let the total number of CpGs it covers be nk, and the number of CpGs in the methylated state be mk, then the methylation level lk of this read is defined as: lk = mk / nk.
[0152] (2) Construction of methylation level probability distribution
[0153] The methylation level lk of the read was divided into the following 5 intervals: B1=[0, 0.2), B2=[0.2, 0.4), B3=[0.4, 0.6), B4=[0.6, 0.8), B5=[0.8,1.0].
[0154] Count the number of segments read in each interval c i (i=1,2,3,4,5), the total number of valid read segments R = ∑ c i To avoid the KL divergence becoming incalculable due to a 0 probability, it is preferable to introduce a smoothing term ε (e.g., ε=10). -6 (or Laplace smoothing): p i = (c i + ε) / (R + 5ε).
[0155] Thus, the distribution P=(p1, …, p5) of the intervention group (Treated) and the distribution Q=(q1, …, q5) of the control group (Control) were constructed.
[0156] (3) JSD value calculation
[0157] Let the average distribution be M = (P + Q) / 2. Define JSD based on KL divergence: JSD(P||Q) = √{ 1 / 2 × [ D_KL(P||M) + D_KL(Q||M) ]}; where D_KL(P||M) = ∑ p i × log2(p i / m i In this embodiment, log2 is used as the logarithmic base, so that the value range of JSD is [0, 1].
[0158] JSD = 0 indicates that the methylation distribution patterns of the two groups are consistent (no structural remodeling was observed); JSD approaching 1 indicates that the distribution patterns of the two groups are significantly different (a strong structural remodeling signal exists).
[0159] (4) Dual-region JSD and PRC2 structural reshaping index (PRI)
[0160] Calculate the average JSD (JSD_promoter) in the promoter region and the average JSD (JSD_genebody) in the gene body region of the PRC2-100 gene set.
[0161] Constructing the PRC2 Structural Restructuring Index (PRI): PRI = α × JSD_promoter + β × JSD_genebody; Here, α and β are weighting coefficients, with α=0.6 and β=0.4, to reflect the sensitivity of the promoter region to transcriptional regulation; however, they can also be adjusted according to tissue type, sequencing depth, and validation data.
[0162] 4. The technical significance of JSD
[0163] Compared to comparing only the average methylation difference (Δβ), JSD has the following technical advantages: Distribution sensitivity: JSD not only reflects mean shift, but also captures the heterogeneity of methylation status, subpopulation drift and changes in distribution morphology in cell populations, making it more suitable for assessing whether "order structure has been reshaped"; Mechanism specificity: An increase in the JSD value indicates that the intervention caused a structural change in the methylation probability distribution, rather than just random noise or single-point fluctuations; in this invention, it can be used to verify whether the intervention reshapes the PRC2 core regulatory network at the structural level. Dual-region collaborative validation: The JSD of the promoter and the genome are calculated separately, which can distinguish between regulatory region remodeling and structural region remodeling, reducing the one-sidedness caused by only detecting a single region.
[0164] 5. Experimental Results
[0165] (1) JSD Dual-Region Analysis
[0166] like Figure 12 As shown, compared with the untreated group (Old), the mean JSD in the promoter region of the PRC2-100 gene set was 0.45 ± 0.05 and the mean JSD in the gene body region was 0.38 ± 0.04 in the elderly intervention group (Old+OSKGL).
[0167] Statistical tests (e.g., Mann–Whitney U test) showed that the JSD of PRC2-100 was significantly higher than that of the control set (e.g., P < 0.001). The control set consisted of 100 non-PRC2 target genes (screening criteria: EZH2 / EED / SUZ12 trisubunit binding signals were all below the background threshold in hESC and aged mouse liver H3K27me3ChIP-seq data, and the promoter region methylation level β > 0.7 were housekeeping genes; the genes were randomly selected from the Housekeeping Gene Atlas annotated with mouse mm10, and representative genes included Actb, Gapdh, Hprt, Tbp, Rpl13a, Sdha, Hmbs, Pgk1, B2m, Ubc, etc.). The mean JSDs in the corresponding regions were 0.12 ± 0.02 and 0.10 ± 0.02, respectively.
[0168] Figure 12 The heatmap further revealed that some key genes (such as Sox11, Lin28b, Hoxc5, Rbfox1, and Nrxn3) showed high JSD signals (e.g., JSD > 0.5) in both regions, suggesting a systematic change in the methylation distribution pattern of these gene target regions.
[0169] right Figure 12 Explanation: Vertical axis: Lists PRC2-100 core genes (such as Sox11, Lin28b, Hoxc5, Cdh8, etc.); Horizontal axis: Divided into promoter region (left half, 5 elderly samples compared with OSKGL reference group) and gene body region (right half, 5 elderly samples compared with OSKGL reference group); Color bar: Indicates the intensity of JSD difference (0-1), the brighter the color (yellow), the higher the JSD value; Technical meaning: A high JSD value indicates that OSKGL intervention, compared with the untreated elderly state, induced significant epigenetic remodeling in this gene region, and the methylation distribution pattern underwent a systematic change.
[0170] (2) Joint determination of PRI index and safety window
[0171] The PRI index of PRC2-100 was 0.42 ± 0.04, which was higher than the preset threshold (e.g., 0.3) (P<0.01). Meanwhile, the epigenetic clock method of Example 1 was used to calculate the epigenetic age reversal rate R, which was 65.3% ± 8.2%, within the safe and effective window of 47%–77%.
[0172] In the optional combined intervention group, JSD was further improved (mean JSD in promoter region = 0.48 ± 0.04, mean JSD in gene body region = 0.41 ± 0.03, PRI = 0.45 ± 0.03), suggesting that the combined strategy may enhance the magnitude of structural remodeling.
[0173] (3) Functional benefit association
[0174] Under the conditions of this embodiment, survival analysis showed an increase in median remaining lifespan in the OSKGL intervention group (e.g., 78%), and a further increase in the combined group (e.g., 91%) (see [link to example]). Figure 9 This suggests that when PRI ≥ the threshold and R is within a safe window, intervention can achieve "safe and structurally deep apparent remodeling".
[0175] 6. Technical Conclusions
[0176] Specific remodeling: JSD analysis showed that OSKGL intervention could produce significant changes in the distribution structure in the core target region of PRC2-100, while the effect on non-PRC2 targets was relatively weak, thus providing evidence of mechanism specificity.
[0177] Dual-region synergy: The remodeling effect can be simultaneously manifested in the promoter (regulatory region) and the genome (structural region), suggesting that the intervention has structural depth and synergy.
[0178] Dual verification system: A dual indicator system of "JSD verification mechanism depth + clock monitoring safety window" has been established: JSD is used to confirm structural remodeling, and the clock is used to confirm that the remodeling magnitude is within the safe treatment window.
[0179] Method versatility: This method is not limited to OSKGL and can be extended to the evaluation of other partial reprogramming, anti-aging, or epigenetic regulation intervention programs, providing a standardized tool for quantifying "epiggenetic order restoration".
[0180] Furthermore, as an implementable variant: the thresholds PRI_th and R_window can be calibrated using a training set (containing young controls, old controls, and known effective interventions) to adapt to different tissues, sequencing depths, and species.
[0181] The PRC2-100 gene set can be replaced with a gene set that satisfies the significant enrichment of the PRC2 target and has a consistent distribution change in aging / intervention controls, preferably 50–300.
[0182] Note: The results of this embodiment, together with the epigenetic clock data of Example 1, constitute a complete technical solution. The list of 100 core genes involved is detailed in Appendix 1. The LMR coordinates of the target region and the probe design parameters are detailed in Appendix 3. The probe sequence can be obtained by directly extracting the corresponding positive strand 100 nt fragment from the mouse reference genome mm10 (GRCm38) using the coordinates listed in Appendix 3. The probe is then delivered to a commercial company (such as IDT) to synthesize biotinylated RNA probes according to the xGen Lockdown Probes specification.
[0183] Example 5: Risk Verification Experiment of Intervention Beyond the Safe Window
[0184] To systematically verify the potential risks of epigenetic age reversal rates exceeding the safety window (R > 77%) and to clarify the upper limit of the safety treatment window described in this invention, this embodiment adds an OSKGL-exceeding-window group based on the grouping in Example 2 above. By increasing the intervention intensity, changes in biological effects, safety, and survival benefits after exceeding the preset window are observed to verify the necessity of limiting the reversal rate R to the range of 47% to 77%.
[0185] 1. Experimental Design and Intervention Plan
[0186] The experimental animals were 104-week-old male C57BL / 6J mice (n=12). This group used the same AAV9 hybrid vector system as the OSKGL window group (AAV9.TRE3-OSK-SV40pA + AAV9.TRE3-OGL-SV40pA + AAV9-hEf1a-rtTA4-Sv40pA), but the intervention intensity was adjusted as follows to induce a stronger reprogramming effect: Carrier dosage: Total injection dose increased to 2×10¹ 4 vg / kg (approximately 3.3 times the standard window group dose); Induction protocol: The cyclical pattern of "doxycycline administration in drinking water (2 mg / mL) for two consecutive weeks, followed by a one-week break" was adopted (the window group was "one week of administration / one week of break"). Control settings: The remaining feeding conditions, observation period and testing methods are the same as in Example 2.
[0187] 2. Experimental Results
[0188] 2.1 The proportion of epigenetic age reversal significantly exceeded the standard.
[0189] At the end of the intervention (22 months of age), liver and skin tissues were collected. TEESM-seq was used to detect methylation of the PRC2 target region and calculate the reversal rate R. Results showed that the mean epigenetic age reversal rate in the OSKGL-window group was as high as 82.4% ± 5.1% (mean ± SEM), significantly exceeding the upper limit of the preset safety window (77%) (see [link to relevant documentation]). Figure 13 The R value of all individuals in this group was greater than 77%, and the R value of 4 mice exceeded 90%, indicating that the intervention had entered a "supraphephysiological reset" state.
[0190] 2.2 A sharp increase in tumor incidence and a shortened survival time
[0191] Long-term observation up to 6 months after intervention (approximately 28 months of age) revealed palpable subcutaneous or peritoneal nodules in 30% (3 / 10, 2 mice died midway and were not included in the statistics) of mice in the OSKGL-ultrawindow group. Histopathological identification showed that 2 cases were teratomas (containing three germ layers) and 1 case was an undifferentiated sarcoma. In contrast, no tumors were observed in the OSKGL window group (R≈65%), OSK group, and control group at the same time point (0 / 10).
[0192] Kaplan-Meier survival analysis showed that the mortality curve in the out-of-window group decreased sharply after 24 months of age, and the median remaining life was shortened by approximately 22% compared to the OSKGL window group (p < 0.01, Log-rank test). The results confirm that while over-reset can bring temporary improvements in apparent youthfulness, it comes at the cost of long-term survival safety.
[0193] 2.3 Impaired cell identity stability
[0194] Single-cell RNA sequencing (scRNA-seq) of skin fibroblasts from the ultra-window group revealed that 0.8% of cells abnormally co-expressed the mesenchymal marker Thy1 and the epithelial marker Epcam, while the proportion of such double-positive cells was less than 0.05% in the OSKGL window group and the young control group. Simultaneously, weak but distinct transcripts (average FPKM > 1) of pluripotency-related genes Nanog and Esrrb were observed in some cells, suggesting a disturbance in cell identity maintenance mechanisms. Flow cytometry further confirmed that approximately 1.2% of cells exhibited an abnormal combination of surface markers (Sca-1). + / CD49f - This indicates that the ultrawindow intervention induced local cell identity disorder and dedifferentiation risk.
[0195] 2.4 Aberrant remodeling of methylation distribution structure (abnormally elevated JSD and PRI)
[0196] The methylation distribution of the PRC2-100 core gene in the ultrawindow group was analyzed using the JSD method described in Example 4. The results showed that the average JSD in the promoter region was as high as 0.58 ± 0.06, which was significantly higher than that in the OSKGL window group (0.45 ± 0.05, p < 0.01); at the same time, the JSD in the gene body region also increased to 0.44 ± 0.05 (0.38 ± 0.04 in the window group).
[0197] This phenomenon suggests that ultra-window intervention causes excessive concentration of methylation distribution in a certain extreme pattern (such as an abnormal increase in the proportion of fully methylated or fully unmethylated reads), leading to abnormal remodeling of the distribution structure and loss of the epigenetic plasticity and diversity that normal cells should have. The combined PRI index (α=0.6, β=0.4) rose to 0.52 ± 0.05, far exceeding the mechanism verification threshold (0.3), and even higher than that of the young control group (0.46 ± 0.05), further proving that the remodeling has deviated from the physiological repair track and entered a pathological disorder state.
[0198] 3. Summary of key parameters for each group
[0199] For ease of comparison, the proportion of epigenetic age reversal, tumor incidence, and PRC2 target region methylation remodeling index of each group are summarized in Tables 4 and 5.
[0200] Table 4. Proportion of epigenetic age reversal and tumor incidence in each group
[0201] a The reversal ratio R is calculated as R = (Clock_Old - Clock_Treat) / (Clock_Old - Clock_Young) × 100%, with data expressed as mean ± SEM.
[0202] b The tumor incidence rate was defined as the proportion of individuals who developed a pathologically confirmed benign / malignant tumor within 6 months after the intervention.
[0203] c The median remaining life expectancy was relatively longer compared to the elderly control group (0%). The OSKGL-extra-window group data are expressed as changes relative to the OSKGL window group.
[0204] d The R value of the OSK group was not directly measured, but it was estimated to be about 18.5% based on its 15% lifespan extension and the linear model L = 0.92R -0.02.
[0205] eThe R value of the OSKGL+GLP1 group is not directly detected. In this embodiment, the combined group R is still controlled within the window (data is not shown).
[0206] Table 5. Methylation distribution and remodeling index of PRC2 target regions in each group (JSD and PRI)
[0207] a PRI = α·JSD_promoter + β·JSD_genebody, where α = 0.6 and β = 0.4.
[0208] Note: A higher JSD value indicates a greater difference in the distribution of methylated reads between the intervention group and the elderly control group; the JSD in the out-of-window group was significantly higher than that in the window group (p < 0.01).
[0209] Figure 14 This study compared the tumor incidence rates between the out-of-window intervention group and the safe window intervention group. The results showed that the tumor incidence rate was 0% in the safe window group (R between 47% and 77%), while the tumor incidence rate in the out-of-window group (R > 77%) increased significantly to 30% (3 / 10, p < 0.01), indicating that exceeding the upper limit of the safe window significantly increases the risk of tumorigenesis.
[0210] Figure 15 The difference in the remodeling index (PRI) distribution of methylation in the PRC2 target region between the out-of-window intervention group and the safe-window intervention group is shown (n=10 / group). Box plots Q1 / Q3 were calculated using quantile (25% / 75%) linear interpolation, with the median at the 50th quantile. The whisker lines represent the minimum and maximum values, and the superimposed scatter plots represent individual values. The results show that the PRI in the out-of-window group was significantly higher than that in the safe-window group. Statistical test: Two-tailed unpaired t-test (homogeneity of variance) p=5.80×10⁻⁶ -8 Mann–Whitney U test p = 1.82 × 10⁻⁶ -4 The image is marked with " The difference was marked as "(p<0.001)" and was extremely significant.
[0211] 4. Conclusion: Out-of-window intervention disrupts the risk-benefit balance.
[0212] The above data reveals for the first time that when the epigenetic age reversal rate R > 77%, some reprogramming interventions can bring about a greater change in epigenetic rejuvenation indicators, but they are accompanied by a sharp increase in tumor risk, cell identity disorder and shortened survival, and cannot achieve safe and effective life extension.
[0213] Therefore, strictly limiting R within the 47%–77% window is a necessary condition for achieving "significant lifespan benefit with manageable risk". The safety window definition of this invention is supported by clear biological and preclinical evidence, providing a quantitative standard for the safety assessment of some reprogramming therapies.
[0214] Example 6: Transcriptome safety verification of tumor overprogramming and OSKGL quantitative partial reprogramming
[0215] This embodiment aims to verify, at the transcriptome level, the fundamental differences between OSKGL quantitative partial reprogramming (R within the 47%~77% safety window) and tumor over-reprogramming (R>85%) in the expression profile of cancer-related genes, providing orthogonal validation evidence of gene expression levels for the safety window described in this invention.
[0216] I. Experimental Objective
[0217] Examples 2-5 have validated the effectiveness of the safety window from dimensions such as epigenetic clock (R-value), teratoma score, and survival analysis. However, the above evidence is based on epigenetic or phenotypic indicators and has not directly proven whether OSKGL intervention within the safety window activates tumor-related transcriptional programs. This example answers the following key questions by comparing the whole transcriptome differences between the tumor over-reprogramming model and OSKGL quantitative partial reprogramming: (1) Does tumor over-reprogramming (equivalent reprogramming intensity ≥85%) accompany systemic activation or inhibition of cancer-associated genes? (2) Does the OSKGL quantitative partial reprogramming (R = 47%~77%) trigger cancer transcriptional features similar to those in the tumor model? (3) Are pluripotent genes such as Oct4 and Nanog silenced in the OSKGL intervention group?
[0218] II. Experimental Design and Grouping
[0219] 2.1 Tumor over-reprogramming model group (in vitro equivalent model)
[0220] To simulate overprogramming states of varying intensities, this embodiment employs the following strategy to construct an in vitro overprogramming model: (a) 95% reprogramming intensity group: Primary hepatocytes from 22-month-old C57BL / 6J mice were transduced with a lentiviral vector carrying an OSKGL polycistronic cassette and induced with a high concentration of doxycycline (10 μg / mL) for 14 days without a rest period. TEESM-seq analysis confirmed that the equivalent reprogramming intensity R ≥ 92% in this group, and the transcriptomic characteristics were highly consistent with the hepatocellular carcinoma model induced by the deletion of Hippo pathway core kinases (Mst1 / 2) (Pearson r = 0.89, P < 1×10⁻⁶).-50 ), n = 6.
[0221] (b) 85% reprogramming intensity group: Same as in Example 5, but induced continuously for 10 days with a moderate concentration of doxycycline (5 μg / mL). TEESM-seq confirmed R ≈ 85%~88%, and the transcriptomic characteristics were significantly correlated with the Sav1 deletion-induced hepatocellular carcinoma model (Pearson r = 0.82), n = 6.
[0222] 2.2 OSKGL quantitative partial reprogramming group (in vivo model)
[0223] Reuse the animal model described in Example 2: (a) Young control group: 3-month-old C57BL / 6J mice injected with PBS vector control, n = 8.
[0224] (b) Old + OSKGL intervention group: 22-month-old mice were given AAV-OSKGL cyclic induction (DOX for 3 days / stop for 4 days) according to the protocol described in Example 2 from 15 months of age, and liver tissue was collected at 22 months of age. TEESM-seq confirmed that the R after intervention was 65.3% ± 8.2%, which was within the 47%~77% safety window, n = 8.
[0225] III. RNA-seq Sequencing and Data Analysis
[0226] 3.1 RNA extraction and library construction: Total RNA was extracted from liver tissues of each group using the TRIzol method (Invitrogen, 15596018). After treatment with DNase I, strand-specific mRNA libraries were constructed using the NEBNext Ultra II RNA Library Prep Kit (NEB, E7770S).
[0227] 3.2 Sequencing: Illumina NovaSeq X Plus platform, paired-end 150 bp sequencing, ≥30M cleanreads per sample.
[0228] 3.3 Data processing: STAR v2.7.10 was used to align to the mm10 genome, featureCounts were used for quantification, and DESeq2 was used for differential expression analysis (|log2FC| > 1, padj < 0.05).
[0229] 3.4 Definition of Cancer-Associated Gene Set: Based on the Hallmark_Cancer gene set from MSigDB and the TCGA database for liver cancer, 100 representative genes that are differentially expressed in liver cancer and show a positive trend during aging were screened and divided into two categories: (i) DOWN in Cancer & ↓ in Aging: 55 genes that are significantly downregulated in the tumor model (log2FC < -1) and also decrease in aging, mainly enriched in tumor suppressor pathways such as cell cycle checkpoints, DNA damage repair, and mitochondrial respiratory chain; (ii) UP in Cancer & ↑ in Aging: 45 genes that are significantly upregulated in tumor models (log2FC > 1) and also increase during aging, mainly enriched in oncogenic pathways such as EMT, Wnt / β-catenin, Hippo-YAP, and inflammatory signaling.
[0230] IV. Results
[0231] 4.1 The tumor over-reprogramming model exhibits systemic cancer transcriptional characteristics.
[0232] like Figure 16 As shown in the two columns on the left, the 95% reprogramming intensity group and the 85% reprogramming intensity group exhibited highly consistent and strong expression changes across 100 cancer-associated genes: Upper block (DOWN in Cancer): Both groups showed strong gene expression changes (log2FC = +1.5 ~ +4.0, dark in the heatmap), with the 95% group showing the strongest signal and the 85% group showing a slightly weaker signal but consistent pattern, indicating that these genes were abnormally regulated to oncogenic levels after over-reprogramming. Lower block (UP in Cancer): Both groups showed significant changes in gene expression (log2FC = -1.0 ~ -4.0), with the deepest gene signals at the bottom, suggesting that oncogenic pathways such as Wnt / Hippo were strongly activated during over-reprogramming.
[0233] 4.2 OSKGL quantitative partial reprogramming does not trigger cancer transcription programs
[0234] like Figure 16 As shown in the two columns on the right, the expression changes of the same gene set in both the young control group and the elderly + OSKGL intervention group are close to zero: Upper block: Both columns are light-colored / white (log2FC ≈ -0.3 ~ +0.5), with no systematic changes; Next block: The young group was almost blank, while the old + OSKGL group showed weak positive changes in some genes (log2FC≈ +0.5 ~ +1.5), but it was far from reaching the intensity of the tumor model (difference > 5-fold). Overall, the results showed that within the safety window (R = 65.3%), OSKGL intervention did not induce transcriptome reprogramming similar to tumor over-reprogramming.
[0235] 4.3 Quantitative Difference Statistics
[0236] Table 6 Quantitative Difference Statistics
[0237] The data in the table show that the mean absolute log2 FC of the cancer gene set in the tumor over-reprogrammed groups (95% and 85%) were 2.68 and 1.93, respectively, with 89 and 76 differentially expressed genes, respectively; while the mean |log2 FC| in the young control and elderly + OSKGL groups were only 0.12 and 0.31, respectively, with only 2 and 7 differentially expressed genes. Regarding the Pearson correlation coefficient with the tumor transcriptome, the over-reprogrammed group had a high coefficient of 0.82–0.89, while the OSKGL group had only 0.08, confirming that quantitative partial reprogramming within the safety window is not correlated with tumor status at the transcriptome level.
[0238] Of particular note was that the endogenous expression of Oct4 and Nanog was significantly activated in the two tumor over-reprogramming groups (FPKM values of 18.7 and 12.4, and 9.3 and 5.6, respectively), while it remained silent in the young control and OSKGL intervention groups (FPKM < 0.1), further validating that OSKGL intervention did not exceed the pluripotency threshold.
[0239] V. Conclusion
[0240] This embodiment provides the following key evidence at the transcriptome level: (1) Tumor over-reprogramming (R ≥ 85%) is accompanied by systemic abnormal expression of 100 cancer-associated genes, and the higher the reprogramming intensity (95% > 85%), the stronger the cancer transcriptional characteristics; (2) The expression changes of the OSKGL quantitative partial reprogramming (R = 65.3%, within the 47%~77% safety window) on the same gene set were close to the background noise level and had no statistical correlation with the tumor transcriptome (r = 0.08). (3) Oct4 and Nanog remained completely silent in the OSKGL intervention group, confirming that the intervention did not cross the pluripotency boundary.
[0241] The above results, together with the conclusions of Example 2 (survival analysis), Example 4 (JSD validation), and Example 5 (over-window tumorigenicity risk), form a multi-level cross-validation, collectively supporting the safety and rationality of limiting the reversal rate R to 47%~77% in this invention. The cancer-associated gene heatmap involved in this example is shown below. Figure 16 .
[0242] Figure 16A heatmap showing the expression of cancer-related genes in tumor over-reprogramming and OSKGL quantitative partial reprogramming. The vertical axis represents 100 cancer-related genes selected based on differential expression, divided into two groups: "downregulated in the tumor model and decreased in aging" (upper block) and "upregulated in the tumor model and increased in aging" (lower block). The horizontal axis, from left to right, represents the tumor over-reprogramming model (95% reprogramming intensity, 85% reprogramming intensity) and OSKGL quantitative partial reprogramming (young control, elderly + OSKGL intervention). The color scale represents log2 Fold Change, with dark colors representing high expression changes and light / white colors representing no significant changes.
[0243] Example 7: Early Remodeling Monitoring, Low-Depth Robustness Assessment, and Risk Warning Based on Attention Mechanism Deep Neural Networks
[0244] 1. Purpose
[0245] We validated that the attention-based deep neural network evaluation engine significantly improves intervention response sensitivity, low sequencing depth stability, window boundary sample discrimination, and safety warning compared to the linear fusion clock based on MML / NME.
[0246] 2. Data Sources and Feature Construction
[0247] TEESM-seq sequencing data from the young control (Young), untreated (Old), treated (Old+OSKGL), and treated (Old+GLP1) groups in Examples 1 and 2 were reused.
[0248] For each sample, the sequencing reads covering the PRC2-100 core gene region are divided into 5 methylation level intervals according to the method described in Example 4, and the empirical probability distribution vector of each region is obtained; all region vectors are concatenated to form the methylation probability distribution matrix X ∈ R^(100×5) of the sample.
[0249] Extract the H3K27me3 ChIP-seq signal intensity of the corresponding region to form the second mode input vector.
[0250] 3. Model Building and Training
[0251] The model employs a Transformer encoder architecture, incorporating a multi-layer, multi-head self-attention mechanism and a feedforward network. The model input is a methylation probability distribution matrix (flattened or with preserved region structure), and the output is the sample embedding. A regression head is used to map the Clock' value, which is then used to calculate the reversal ratio R'.
[0252] The model was pre-trained in a self-supervised manner on large public methylation datasets (such as GEO and GTEx) (masking a portion of the region and predicting the distribution of the masked region), and then fine-tuned on Young / Old samples in this experiment to fit the biological age.
[0253] Meanwhile, the model outputs the attention weights for each layer, which can be aggregated into the Attention Risk Index (ARI) (defined as the average concentration of attention weights on a preset set of oncogene-related sites).
[0254] 4. Controlled Experiment
[0255] Traditional method: Clock and R are calculated according to the method in module five of embodiment 1 of this invention.
[0256] Mechanism verification: PRI was calculated according to the method in Example 4 of this invention.
[0257] Attention model method: R′ and ARI are calculated according to this embodiment.
[0258] 5. Results
[0259] Early sensitivity: Sampling analysis at different time points (12h, 24h, 48h, and 72h) after OSKGL intervention showed that the attention model detected a significant reconstruction of the methylation distribution pattern in the PRC2 target region as early as 24h (R′ was significantly different from the Old group, P<0.01), while the traditional MML / NME clock did not show a significant difference until 72h. This indicates that the attention model can capture more subtle early distribution changes.
[0260] Low-depth robustness: By simulating different sequencing depths (100x, 50x, 20x, 10x) through random downsampling, the coefficient of variation (CV) of R′ of the attention model at a depth of 20x is less than 5%, while the CV of the traditional R value is greater than 15%. This demonstrates that the attention model has stronger denoising and compensation capabilities and can maintain high accuracy under low-cost sequencing.
[0261] Window boundary sample discrimination and risk warning: Samples with R values close to the lower limit of 47% and the upper limit of 77% were selected. Traditional methods only judged the risk as "controlled" or "high-risk" based on the R value alone. In one sample with an R value of 71%, the attention model detected an abnormal concentration of attention weights on oncogenes (such as Myc and Ras-related sites) (ARI=0.23, significantly higher than the average of 0.05 in the normal intervention group), suggesting a potential risk of dedifferentiation. Subsequent transcriptome validation showed that this sample did indeed show upregulation of some oncogene expression. This result proves that the attention model can provide risk warnings beyond a single numerical threshold.
[0262] Multimodal consistency: In samples with H3K27me3 data, the attention consistency index (ACI) output by the cross-attention model is highly positively correlated with PRI (Pearson r=0.85), and when ACI≥0.4, the JSD values of the corresponding regions are all greater than 0.5, indicating that the attention mechanism captures structural remodeling consistent with biological mechanisms.
[0263] Table 7. Comparison of the sensitivity and timeliness of different assessment methods to OSKGL intervention.
[0264] Note: Cohen's d measures the magnitude of the difference between the means of two groups; d > 0.8 is considered a large effect. The p-value is based on the comparison with the elderly untreated group.
[0265] Table 8. Stability analysis of evaluation metrics at different sequencing depths
[0266] Note: A lower coefficient of variation (CV) indicates a more stable metric and less impact from sequencing depth. Data shows that the attention model exhibits significantly stronger robustness at lower sequencing depths.
[0267] Table 9. Case Studies of the Application of Attention Risk Index (ARI) in Security Early Warning Note: ARI is defined as the average concentration of attention weights on predefined potential risk sites (such as the oncogene-associated PRC2 target region). Example sample T-12 demonstrates that ARI can independently identify potential tumorigenic risks even when the R value is within the safety window.
[0268] 6. Conclusion
[0269] This embodiment demonstrates that by introducing a deep neural network based on an attention mechanism into the evaluation system of this invention, earlier, more robust, and more intelligent quantitative evaluation of partial reprogramming interventions is achieved. The attention model not only outputs a reversal ratio R′ comparable to the original R value, but also provides single-cell-level tumorigenesis risk warnings through attention mapping, and synergizes with the JSD / PRI mechanism validation indicators. This technical effect is unpredictable by those skilled in the art based on traditional linear clock and information-theoretic distance methods, significantly improving the safety and controllability of the intervention process.
[0270] 7. Privacy Compliance and Multi-Center Deployment (Federated Learning)
[0271] In some implementations, the deep neural network model is deployed across multiple clinical centers using a federated learning architecture. Each center trains the model locally, uploading only model gradients or parameter updates, without uploading the raw methylation sequencing data or any identifiable information of the subjects; the central server securely aggregates these updates to form a global model. This architecture, while meeting data privacy regulations such as GDPR and HIPAA, can continuously improve the model's prediction accuracy and generalization ability for different populations (different races, regions, and underlying medical histories), while avoiding the privacy risks associated with centralized data.
[0272] 8. Epigenetic Digital Twins and Dosage Simulation
[0273] Before intervention, a trained model can be used to construct an "epigenetic digital twin" of the patient: the methylation probability distribution matrix of the patient's baseline sample is input into the model, and by adjusting the parameters in the input that simulate different intervention intensities (such as simulating the dose effect through masking or perturbation, or by changing the latent variables representing the reprogramming intensity in the input matrix), the model can output the predicted R′ value and the corresponding ARI risk status under different hypothetical intervention conditions. Based on this, clinicians can screen recommended dosage regimens that ensure R′ falls within the safe window (47%-77%) and ARI is below the risk threshold before actual drug administration, achieving a paradigm shift from "empirical drug administration" to "data-driven precision drug administration." This digital twin simulation method can be further linked with the closed-loop treatment system described in claims 31-33 to form a closed-loop control loop of "prediction-assessment-drug administration-feedback."
[0274] Example 8: A dual endpoint evaluation system based on TEESM-seq and Transformer-assisted assessment—using the R safety window and PRI mechanism to jointly determine the effectiveness of OSKGL intervention.
[0275] I. Purpose
[0276] This embodiment aims to establish a dual-endpoint quantitative assessment method based on the TEESM-seq platform. It jointly determines the epigenetic age reversal ratio R (reflecting intervention intensity and risk) and the PRC2 structural remodeling index PRI (reflecting mechanism-specific remodeling depth) to achieve a comprehensive evaluation of the effect-risk-mechanism of some reprogramming interventions. Simultaneously, this embodiment compares the performance of two reprogramming factor combinations, OSK and OSKGL, on the dual-endpoint indicators, verifying the superiority of OSKGL and demonstrating the compatibility of this assessment system with traditional statistical methods and Transformer neural network models.
[0277] II. Experimental Materials and Grouping
[0278] 1. Reprogramming factor composition: OSKGL group: Oct4, Sox2, Klf4, Glis1, Lin28; OSK control group: Oct4, Sox2, Klf4; Delivery method: AAV9 vector, injected via retroorbital vein, with a dose of 1×10¹² vg (viral genome copy) per mouse, total volume 100 μL.
[0279] Induction system: Tet-On inducible expression system, which is induced periodically by adding doxycycline (final concentration 2 mg / mL) to drinking water (5 days of administration / 2 days of withdrawal).
[0280] 2. Laboratory animals and grouping: 104-week-old male C57BL / 6J mice were randomly divided into the following four groups (n=12 per group): Young control group: 3 months old, no treatment given.
[0281] Old control group: 104 weeks old, injected with PBS buffer.
[0282] Old+OSK group: AAV9-OSK vector was injected, followed by periodic DOX induction.
[0283] Old + OSKGL group: AAV9-OSKGL vector was injected and DOX was induced periodically.
[0284] 3. Sample collection: Baseline (pre-intervention) tail tip tissue was collected.
[0285] AAV administration was followed up monthly, and back skin tissue was collected at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 and 12 months. The tissue was flash-frozen in liquid nitrogen and stored at -80℃ for later use.
[0286] Record the DOX induction status (drug administration period or drug withdrawal period) at each sampling time point.
[0287] III. Experimental Methods
[0288] 1. TEESM-seq targeted methylation sequencing
[0289] (1) Genomic DNA was extracted from each sample and hybridized and captured using nucleic acid probes designed to target the hypomethylated regions co-binding with the PRC2 trisubunits (EZH2 / EED / SUZ12). The total target size was 3 Mb, covering the regulatory regions of the 100 PRC2 core genes listed in Appendix 1.
[0290] (2) The enriched products were treated with bisulfite, and after constructing the library, Illumina NovaSeq paired-end sequencing (PE150) was performed, with an average sequencing depth of ≥100× for the target region.
[0291] (3) Filter the sequencing reads: retain valid reads with MAPQ≥20 and covering at least 3 CpG sites.
[0292] 2. Integration of epigenetic clock and R-value calculation
[0293] (1) For the target region set of each sample, calculate: Mean methylation level (MML): The methylation level of all valid reads within the region (l k = m k / n k The mean of ).
[0294] Normalized methylation entropy (NME): Following the method described in Example 1, the methylation level of the read was divided into 5 intervals [0,0.2), [0.2,0.4), [0.4,0.6), [0.6,0.8), and [0.8,1.0], and the normalized Shannon entropy was calculated.
[0295] (2) Calculate the converged clock value (Clock): Method A (Statistical Benchmark): Clock = a × MML + (1 - a) × NME The weight parameter a is set to 0.65 (determined through optimization on the training set).
[0296] Method B (Transformer Auxiliary): The methylation probability distribution matrix is input into a pre-trained Transformer neural network model, which outputs a "youthful embedding vector". The physical distance between this vector and the young manifold is calculated as Clock′. In this embodiment, the Clock value can be equivalent to Clock′ or transformed through a linear mapping.
[0297] (3) Calculate the epigenetic age reversal rate R: R = (Clock_Old - Clock_Treat) / (Clock_Old - Clock_Young) × 100%; Clock_Old represents the mean of the elderly control group, Clock_Young represents the mean of the young control group, and Clock_Treat represents the sample value of the intervention group.
[0298] (4) Safety window determination (preset threshold range 47%~77%): If 47% ≤ R ≤ 77%: output "Risk under control (within the window)"; if R < 47%: output "Insufficient reprogramming strength"; if R > 77%: output "High risk warning for excessive reprogramming strength".
[0299] 3. JSD and PRI mechanism verification calculation
[0300] (1) Define the 100 PRC2 core genes listed in Appendix 1 as follows: Promoter region: 2 kb upstream to 0.5 kb downstream of the transcription start site (TSS); Gene body region: TSS+0.5 kb to transcription termination site (TES).
[0301] (2) Based on the methylation level distribution of the effective reads (divided into 5 intervals), the Jensen-Shannon distance between the intervention group and the elderly control group in each region was calculated: JSD(P||Q) = √{ 1 / 2 × [ D_KL(P||M) + D_KL(Q||M) ]} Where M = (P + Q) / 2.
[0302] Calculate the average JSD of the promoter region (JSD_promoter) and the average JSD of the gene body region (JSD_genebody) respectively.
[0303] (3) Construct the PRC2 Structural Restructuring Index (PRI): PRI = α × JSD_promoter + β × JSD_genebody The weighting coefficients are set to α=0.6 and β=0.4.
[0304] (4) Mechanism threshold determination (preset threshold PRI_th = 0.3): If PRI > 0.3: Output "Mechanism-specific PRC2 structural remodeling is established"; If PRI ≤ 0.3: Output "Insufficient / Invalid mechanism reshaping".
[0305] 4. Comprehensive Judgment Rules for Two Endpoints
[0306] For each sample / time point, the comprehensive judgment result is output according to the following rules: Table 10. Comprehensive Judgment Rules for Two Endpoints
[0307] IV. Experimental Results
[0308] 1. Dynamic changes in R-value and distribution of safety window
[0309] Table 11 shows the changes in R-values and window determination results of representative samples in each group during the follow-up period (3, 6, 9, and 12 months).
[0310] Table 11 R-values and window determination at different follow-up times for different intervention groups
[0311] Results analysis: During the periodic induction process, the R value in the OSKGL group remained more stably within the safety window of 47%–77% (OSKGL-05 remained within the window for four consecutive time points), while the OSK group showed significant fluctuations within the window (OSK-03's R value dropped to 38.5% during the withdrawal period) and out-of-window phenomena in a few samples (OSK-07's R value was 78.5% at month 6). This suggests that OSKGL has better controllability of intervention intensity.
[0312] 2. Dynamic changes in PRI and determination of mechanism threshold
[0313] Table 12 shows the changes in PRI values and the results of mechanism determination for samples from the same batch as those in Table 11.
[0314] Table 12 PRI values and mechanism determination for different intervention groups at each follow-up time
[0315] Results analysis: In the OSKGL group, the PRI was significantly higher than the 0.3 threshold (lowest 0.356) at all detection time points, indicating that it continuously induced the structural-specific remodeling of the PRC2 target. In the OSK group, only OSK-07 showed a high PRI (0.392) when the R value was outside the window (78.5%), while most OSK samples (such as OSK-03) still had PRI below 0.3 even when the R value entered the window period (such as 6 and 12 months), suggesting that its epigenetic age reversal may be due to non-specific methylation changes rather than the remodeling of the core PRC2 mechanism.
[0316] 3. Overall judgment result of dual endpoints
[0317] Table 13 outputs a comprehensive determination of the two endpoints based on the above R value and PRI data.
[0318] Table 13. Results of the combined assessment of the two endpoints
[0319] 4. Intergroup dual endpoint pass rate statistics
[0320] The pass rate of the two endpoints was statistically analyzed for all follow-up time points in both groups (12 months in total × 12 animals per group = 144 time points): In the OSK group, the percentage of time points where both endpoints were passed (within the R window and PRI>0.3) was 18.8% (27 / 144); among them, the percentage of time points where a "mechanism insufficiency warning" (within the R window but PRI≤0.3) appeared was 31.9% (46 / 144); the percentage of time points where the "risk of excessive intensity" was 11.1% (16 / 144); and the percentage of time points where the "insufficient intensity" was 38.2% (55 / 144).
[0321] OSKGL Group: The percentage of time points where both endpoints were reached was 68.1% (98 / 144); the percentage of "insufficient mechanism warning" was 6.9% (10 / 144); the percentage of "risk of excessive intensity" was 9.0% (13 / 144); and the percentage of "insufficient intensity" was 16.0% (23 / 144).
[0322] Statistical results showed that the pass rate of the two endpoints in the OSKGL group was significantly higher than that in the OSK group (68.1% vs 18.8%, p<0.001, chi-square test), and the incidence of "mechanism deficiency warning" was significantly lower (6.9% vs 31.9%), confirming that OSKGL can induce structural remodeling of the PRC2 core mechanism more stably while achieving intervention within the safety window.
[0323] V. Conclusion
[0324] This embodiment successfully established and validated a "dual endpoint" quantitative evaluation system based on the TEESM-seq platform. This system achieves a comprehensive evaluation of the effect-risk-mechanism of partial reprogramming interventions by jointly determining the epigenetic age reversal ratio R (47%~77% safety window) and the PRC2 structural remodeling index PRI (>0.3 mechanism threshold).
[0325] The main conclusions are as follows: 1. The R value can effectively monitor the intensity and risk of intervention: the preset window of 47%~77% can distinguish between three states: insufficient intensity, controlled risk, and excessive intensity, providing a quantitative basis for real-time adjustment of the intervention plan.
[0326] 2. PRI can specifically verify PRC2 mechanism remodeling: The PRI threshold > 0.3 can effectively distinguish between specific structural changes and non-specific methylation fluctuations of PRC2 targets, avoiding misjudging superficial changes in epigenetic age as true mechanism remodeling.
[0327] 3. Dual endpoints enhance assessment accuracy: The introduction of the "mechanism deficiency warning" (within the R window but PRI≤0.3) can identify non-specific interventions that show apparent age reversal but do not address the core mechanism, providing important clues for optimizing target selection or combination therapy.
[0328] 4. The OSKGL combination is significantly superior to OSK: The OSKGL group achieved a dual endpoint pass rate of 68.1%, which is significantly higher than the OSK group's 18.8%, confirming that OSKGL can more stably induce structural remodeling of the PRC2 core mechanism while maintaining a safety window, making it the preferred combination for achieving safe and effective partial reprogramming intervention.
[0329] 5. Architecture Compatibility: The evaluation method described in this embodiment can be calculated using traditional statistical formulas or integrated with the Transformer neural network model for feature extraction and prediction. Both methods maintain consistency in R-value and PRI judgment criteria, ensuring the unity of interpretability and advancement of the evaluation system.
[0330] The dual endpoint assessment method described in this embodiment can be widely applied to scenarios such as in vivo intervention, ex vivo organ perfusion, and drug screening, providing a standardized quantitative quality control tool for the clinical translation of some reprogramming technologies.
[0331] [System Output Interface Description]
[0332] To enhance the feasibility of the invention, the output device of the evaluation system described in this embodiment displays at least the following fields: Clock value, R value (%), window determination (insufficient strength / within the window / risk of excessive strength), JSD_promoter, JSD_genebody, PRI value, mechanism determination (valid / insufficient), and dual endpoint comprehensive conclusion (dual endpoint passed / insufficient strength / risk of excessive strength / warning of insufficient mechanism). The output interface can simultaneously display trend graphs of each indicator over time to support dynamic monitoring and decision-making.
[0333] Example 9: Safety Window Calibration and Verification for Different Tissue Types
[0334] 1. Experimental Objective
[0335] The applicability of the safe and effective window (47%-77%) of the epigenetic age reversal ratio R in different tissue types was verified, and tissue-specific calibration parameters were established to support the technical solution that can be calibrated using a training set according to different tissue types.
[0336] 2. Experimental Methods
[0337] Twenty-two-month-old male C57BL / 6J mice were selected and treated with OSKGL (AAV9 vector delivery, DOX periodic induction) according to the method described in Example 2. Six months after intervention, liver, cerebral cortex, and skeletal muscle tissue samples were collected, with n=8 samples per group.
[0338] Targeted methylation sequencing was performed using the TEESM-seq platform, and the R value for each sample was calculated according to the method described in Example 1. The following safety indicators were also detected: Liver: ALT / AST liver function enzyme indicators, teratoma-related markers (AFP, Nanog expression); Cerebral cortex: neuroinflammatory markers (GFAP, Iba1 immunohistochemical staining intensity), neuron-specific enolase (NSE); Skeletal muscle: regeneration-related markers (MyoD, Myogenin expression), rhabdomyosarcoma markers (PAX3 / FOXO1 fusion detection).
[0339] 3. Experimental Results
[0340] As shown in Table 14, different tissues exhibit varying tolerance to reprogramming intensity: Liver: When the R value is in the range of 47%-77%, liver function indicators are normal and no teratoma occurs; when R>77%, 30% of individuals show elevated AFP, indicating a risk of tumorigenesis.
[0341] Cerebral cortex: The upper limit of the safe window for the R value narrowed to 65%. When R > 65%, the GFAP / Iba1 staining intensity increased significantly, indicating a neuroinflammatory response; when R > 77%, abnormal expression of neuronal markers appeared.
[0342] Skeletal muscle: The safe window is basically the same as that of the liver (47%-77%), but when the window exceeds 77% (R>77%), it is accompanied by weak positive results for rhabdomyosarcoma markers (2 / 8).
[0343] Table 14 Safety Window and Beyond-Window Risk Thresholds for Different Organization Types
[0344] like Figure 17 As shown, different tissue types exhibit varying tolerance ranges to partial reprogramming interventions. The safety window for the liver and skeletal muscle is 47%–77%, with abnormalities in tumorigenic or sarcoma-related risk markers appearing when the R value exceeds 77%. The upper safety limit for the cerebral cortex is 65%, with neuroinflammation and abnormal neuronal expression appearing when the R value exceeds 65%. This figure illustrates the calibration results for tissue-specific safety thresholds.
[0345] 4. Conclusion
[0346] The above results demonstrate that the statement "the 47%-77% preset threshold range can be calibrated using a training set according to different tissue types" has sufficient experimental support. In practical applications, a calibrated tissue-specific window threshold can be selected based on the type of target organ or sampled tissue to improve the accuracy of safety assessment. This tissue-specific calibration parameter can be further integrated into the aforementioned closed-loop treatment system to achieve personalized and precise control of intervention programs for different tissue types.
[0347] Example 10: Screening and Validation of the PRC2-150 Extended Gene Set
[0348] To further enhance the coverage of organ-specific aging phenotypes by the epigenetic assessment system of this invention, this invention integrates aging-related genes screened based on multi-organ MRI bioage gaps (MRIBAGs) into the original PRC2-100 gene set, constructing an expanded PRC2 core gene set (PRC2-150). This embodiment aims to detail the screening method of the expanded PRC2 core gene set (PRC2-150) of this invention and verify its superiority in the assessment of partial reprogramming interventions.
[0349] I. Experimental Objective
[0350] (1) Establish a gene screening process based on multi-organ MRIBAGs; (2) Verify the increased sensitivity of the expanded PRC2-150 gene set compared to the original PRC2-100 gene set in the evaluation of intervention effects; (3) Demonstrate the application potential of the newly added genes in mechanism verification and drug screening.
[0351] II. Data Sources
[0352] This embodiment relies on the multi-organ imaging genetics database published by the MULTI Consortium and includes the following data types: Table 15 Data Types
[0353] III. PRC2-150 Gene Set Screening Method
[0354] This method (1) constructs bio-age gaps (MRIBAGs) for seven organs (brain, heart, liver, spleen, pancreas, kidney, and adipose tissue) based on multi-organ MRI data from 313,645 participants, and identifies plasma proteins significantly associated with organ aging through extensive proteomic association analysis (ProWAS). (2) Genomic loci associated with MRIBAGs are identified through genome-wide association analysis (GWAS), and SNP loci are mapped to functional genes by combining eQTL mapping, chromatin interaction mapping, and Bayesian colocalization analysis. (3) Genes showing significant remodeling of Jensen-Shannon distance (JSD) in the promoter and gene body regions (JSD > 0.3) are screened, and their druggability is verified using the DGIdb database.
[0355] Specifically as follows: 3.1 Construction of Multi-Organ MRIBAGs Following the method described in Example 1, an age prediction model was trained in a healthy control population (n=3,573~6,327) to calculate MRIBAGs for seven organs. The model performance is as follows: Table 16 Model Performance
[0356] 3.2 Proteome Broad Association Analysis (ProWAS)
[0357] A linear regression model was used to associate the seven MRIBAGs with 2,923 plasma proteins, adjusting for variables such as age, sex, BMI, principal components of genes, and protein batch. The Bonferroni correction threshold P < 0.05 / 2,923 / 7 = 2.44 × 10⁻⁶. -6 .
[0358] Results: A total of 603 significant protein-MRIBAG associations were identified, distributed by organ as follows: Table 17 Organ Distribution
[0359] Organ enrichment validation: The organ specificity of the protein was validated using the Human Protein Atlas (HPA) database. The results showed that: VCAM1 was enriched in the spleen (tissue-specific score: spleen showed 4-fold higher expression). PLA2G1B is enriched in the pancreas (pancreas-specific expression). NCAN is enriched in brain tissue (neuron-specific expression). 3.3 Genome-wide association analysis (GWAS) GWAS analysis was performed on seven MRIBAGs (sample size 19,686–31,557, European population) using a fastGMA linear mixture model, adjusted for age, sex, BMI, principal genetic components, and organ-specific covariates. The genomic significance threshold was P < 5 × 10⁻⁶. -8 .
[0360] Results: A total of 53 significant genomic loci were identified, some of which are shown below: Table 18 Some representative sites
[0361] 3.4 SNP-to-gene mapping
[0362] Gene mapping using the FUMA platform includes the following three strategies: (1) Location mapping: SNPs are located within 10kb upstream and downstream of the gene; (2) eQTL mapping: SNPs are significantly associated with gene expression (GTEx v8, eQTLGen, Blood eQTL and other databases); (3) Chromatin interaction mapping: The region where the SNP is located has chromatin loop interaction with the gene promoter (Hi-C, CHi-C data, sources include Roadmap, PsychENCODE and other).
[0363] Results: A total of 1,164 initial related genes were identified by the three mapping strategies.
[0364] 3.5 Bayesian Colocality Analysis
[0365] The coloc package (v5.0) was used to perform colocalization analysis on MRIBAGs with plasma proteomics clocks (ProtBAGs, 11 types) and metabolomics clocks (MetBAGs, 5 types) to screen genes that share causal variations (PP.H4 > 0.8).
[0366] Results: A total of 40 MRIBAG-ProtBAG / MetBAG colocalization signals were identified, involving 62 genes (see Appendix Table 2).
[0367] Table 19 Representative Colocation Signals
[0368] 3.6 JSD Mechanism Verification
[0369] Following the method described in Example 4, the Jensen-Shannon distance (JSD) between the promoter region and gene body region of candidate genes was calculated between the OSKGL intervention group (n=5) and the elderly control group (n=5), and genes with JSD > 0.3 were screened.
[0370] Results: A total of 86 genes that showed significant remodeling after intervention were screened, of which 42 genes overlapped with the original PRC2-100 gene set and 44 were newly added genes.
[0371] Table 20 Examples of some high JSD genes
[0372] 3.7 Screening for druggability
[0373] Drug interaction records of candidate genes were retrieved using the Drug Gene Interaction Database (DGIdb v5.0) to screen for target genes of drugs that have been approved by the FDA or are in clinical trials.
[0374] Results: A total of 27 genes with known drug interactions were identified, involving 122 marketed or clinical-stage drugs. Some representative genes and their corresponding drugs are as follows: Table 21 Some representative genes and their drugs
[0375] 3.8 Final Gene Set Integration
[0376] Integrating the above evidence, 62 core genes related to MRIBAG were ultimately identified and included in the PRC2-150 extended gene set. The integration criteria were: (1) At least a ProWAS significant association is satisfied (P < 2.44 × 10⁻⁶). -6 (1) and JSD > 0.3; (2) or colocalization evidence (PP.H4 > 0.8) and JSD > 0.3; (3) or a known druggable target (included in DGIdb) and at least one of the above evidences is met.
[0377] IV. Probe Design for the PRC2-150 Gene Set
[0378] For the expanded PRC2-150 gene set, target capture probes were designed according to the following strategy: 4.1 Target Region Definition Startup sub-region: 2 kb upstream of TSS to 0.5 kb downstream; Genome region: TSS+0.5 kb to TES.
[0379] 4.2 Probe Design Parameters
[0380] Table 22 Probe Design Parameters
[0381] 4.3 Total Target Size
[0382] Table 23 Total Target Size
[0383] 4.4 Sequencing Cost Estimation
[0384] Based on an average sequencing depth of 100×, the amount of data required per sample is approximately 0.45 Gb (4.5 Mb × 100), and the cost of sequencing a single sample can be controlled to below $3.
[0385] V. Validation of the PRC2-150 gene set
[0386] 5.1 Comparison of JSD sensitivity with the original PRC2-100
[0387] The mean promoter JSD and mean genesomal JSD of the PRC2-100 gene set and the PRC2-150 gene set were calculated between the OSKGL intervention group (n=5) and the elderly control group (n=5), respectively. Results: Table 24 Validation results of the PRC2-150 gene set
[0388] The PRI of the PRC2-150 gene set was significantly higher than that of the PRC2-100 gene set (P < 0.05, paired t-test), suggesting that the expanded gene set is more sensitive to the intervention effect.
[0389] 5.2 Mechanism verification of the newly added genes
[0390] GO enrichment analysis of the 62 newly added genes revealed significantly enriched pathways including: Table 25 GO enrichment analysis results
[0391] 5.3 Validation of Drug Target Enrichment
[0392] Validating the proportion of druggable genes in the PRC2-150 gene set using the DGIdb database: Table 26 Proportion of Druggable Genes
[0393] The proportion of druggable genes in the newly added gene set was significantly higher than that in the original PRC2-100 gene set (P < 0.05, chi-square test), suggesting that the expanded gene set has higher clinical translation potential.
[0394] VI. Conclusion
[0395] This embodiment successfully established a gene screening process based on multi-organ MMRIBAGs. Using the above method, this invention ultimately identified 62 core MMRIBAG-related genes, which, together with the original PRC2-100 gene set, constitute the PRC2-150 extended gene set. This extended gene set not only retains the epigenetic regulatory mechanisms of the PRC2-H3K27me3 axis but also incorporates molecular evidence of organ-specific aging phenotypes, significantly enhancing the breadth and depth of the evaluation system of this invention. Validation results show that: (1) The PRC2-150 gene set is more sensitive to OSKGL intervention than the original PRC2-100 gene set; (2) The newly added genes are significantly enriched in aging-related pathways such as immune regulation, inflammatory response, and pancreatic secretion; (3) The proportion of druggable targets in the newly added genes is significantly increased, providing a richer candidate gene library for drug screening; (4) The total target size of the target capture probe for the PRC2-150 gene set is 4.5 Mb, and the cost of sequencing a single sample can be controlled below $3, maintaining the economic advantage of the TEESM-seq platform.
[0396] Therefore, the PRC2-150 extended gene set can be used as a preferred embodiment of the epigenetic evaluation system of the present invention.
[0397] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
[0398] The following are 100 core PRC2 regulatory genes identified through multimodal screening, used for targeted methylation sequencing, JSD analysis, and PRI index construction as described in this invention. This gene set was obtained through screening based on the following criteria: (1) significant JSD differences in the promoter-genome region during aging; (2) significant reversal after OSKGL intervention; and (3) dynamic changes accompanied by H3K27me3 modification.
[0399] Reference genome: mm10 (GRCm38) | Total target size: 3.20 Mb | Total LMR count: 1050 | Total probe count: 63,569 (100 nt, 50 nt step size, ≥2× coverage); Probe design parameters: length 100 nt; tiling step size 50 nt (2× coverage); GC content balanced (40%–65%); RepeatMasker masking SINE / LINE / LTR repeat sequences; biotinylated RNA probes (IDT xGen LockdownProbes).
[0400] Appendix 1: List of PRC2-100 core genes and PRC2-100 gene target region coordinates and probe design
[0401] Note: This list is based on the mouse reference genome mm10 (GRCm38) annotation. When used for human sample detection, probes should be designed for the corresponding human homologous gene regions (human gene symbols and genomic coordinates can be obtained from the NCBIHomoloGene or Ensembl databases). The promoter regions of the 100 genes are defined as 2kb upstream to 0.5kb downstream of the TSS, and the gene body regions are defined as TSS+0.5kb to TES.
[0402] Appendix 2: List of New Genes in the PRC2-150 Extended Gene Set
[0403] Note: This list is based on the human reference genome hg38 annotation. When used for mouse sample testing, probes should be designed to target the corresponding mouse homologous gene regions (mouse gene symbols and genome coordinates can be obtained from the NCBI HomoloGene or Ensembl databases). The promoter region of each gene is defined as 2 kb upstream to 0.5 kb downstream of the TSS, and the gene body region is defined as TSS+0.5 kb to TES.
[0404] Appendix 3: Coordinates of EES-LMR target regions and probe design for PRC2-100 core genes (a total of 1,050 LMR regions, corresponding to the number of LMRs in Appendix 1)
[0405] Note: Each LMR region is designed with a 100 nt length and a 50 nt step size for the probe. Probe naming convention: EES-XXXXX (X is a 5-digit serial number). The promoter region is defined as the PRC2 trisubunit co-binding hypomethylated region within the range of 2 kb upstream to 0.5 kb downstream of TSS; the gene body region is defined as the PRC2 trisubunit co-binding hypomethylated region within the range of TSS+0.5 kb to TES.
[0406]
[0407] Note: (1) The above coordinates are based on the mouse reference genome mm10 (GRCm38). When used for human samples, the probe should be redesigned for the homologous region corresponding to hg38. (2) The probe sequence can be obtained directly from the reference genome by extracting the positive strand sequence (100 nt) using the above coordinates and synthesized into a biotinylated RNA probe by IDT. (3) Before actual synthesis, probes with high repetition regions need to be filtered by RepeatMasker, and probes with abnormal GC content (<30% or >70%) need to be adjusted by Tm equalization.
Claims
1. The use of a reprogramming factor composition in the preparation of a reagent for improving aging-related phenotypes and obtaining lifespan-related benefits, characterized in that, The reprogramming factor composition comprises five factors: Oct4, Sox2, Klf4, Glis1, and Lin28, as well as a GLP-1 receptor agonist.
2. A pharmaceutical composition, characterized by, The pharmaceutical composition comprises the reprogramming factor composition of claim 1 and a pharmaceutically acceptable carrier; the pharmaceutical composition achieves an epigenetic age reversal rate of 47% to 77% in test subjects.
3. A method for evaluating a quantitative assessment of a reprogramming factor composition for improving senescence-associated phenotypes and obtaining life span-related benefits, characterized by, Includes the following steps: (1) Obtain genomic DNA from the test subject sample, use a nucleic acid probe designed for the PRC2 trisubunit co-binding hypomethylation region of 100 PRC2 core genes for hybridization capture enrichment, and perform bisulfite treatment and sequencing on the enriched product to obtain sequencing reads covering the hypomethylation region. (2) Filter the sequencing reads and retain the valid reads that meet the preset quality threshold and cover at least 3 CpG sites; (3) Calculate the average methylation level MML and normalized methylation entropy NME of the region set based on the effective reads; (4) Calculate the epigenetic clock value according to the fusion function Clock = a·MML + (1-a)·NME, where a is a weight parameter between 0.5 and 0.8; (5) Calculate the epigenetic age reversal rate R based on the Clock values of the control young group, the control old group, and the old intervention group; (6) Compare R with a preset threshold range, and output a risk control determination result when R falls into the range.
4. The method of claim 3, wherein, The gene targeting the PRC2 trisubunit co-binding hypomethylated region in step (1) also includes 62 organ aging-related genes; the 100 PRC2 core genes and the 62 organ aging-related genes constitute the extended PRC2 gene set PRC2-150.
5. The method of claim 3, wherein, In step (6), the preset threshold range is preferably 47% to 77%; when R < 47%, an insufficient reprogramming strength warning is output; when R > 77%, an excessively high reprogramming strength risk warning is output; the calculation formula for R is: R = (Clock_Old-Clock_Treat) / (Clock_Old-Clock_Young)×100% or its equivalent normalized form.
6. The method of claim 4 or 5, wherein, The method further includes mechanism validation based on Jensen-Shannon distance for the 100 PRC2 core genes and the 62 organ aging-related genes: promoter regions and gene body regions of the 100 PRC2 core genes and 62 organ aging-related genes are defined respectively. Based on the methylation level distribution of effective reads, the average JSD value JSD_promoter in the promoter region and the average JSD value JSD_genebody in the gene body region are calculated for the elderly intervention group and the elderly control group. A PRC2 structural remodeling index PRI = α·JSD_promoter + β·JSD_genebody is constructed, where α = 0.6 and β = 0.4 are weighting coefficients. When PRI is greater than a preset threshold of 0.3 and R falls within the 47%~77% range, the mechanism-specific remodeling determination result is output: the intervention is effective and safe.
7. The method of claim 3, wherein, Steps (3) to (6) are replaced or assisted by an evaluation process based on an attention mechanism neural network: Based on the effective reads, a methylation probability distribution vector is constructed in each target region according to a preset methylation level range. The methylation probability distribution vectors of multiple target regions are concatenated to form a methylation probability distribution matrix of the region set. The methylation probability distribution matrix is input into a deep neural network model containing a multi-head self-attention mechanism to obtain the epigenetic embedding vector of the sample, and the epigenetic clock value Clock′ and reversal ratio R′ are output. R′ is compared with a preset threshold range of 47% to 77%.
8. The method of claim 7, wherein, When the deep neural network model of the self-attention mechanism includes Transformer, linear attention network, state space model or a combination thereof, it is used to capture long-range dependencies across regions and sites within the region set and output an attention weight map. Based on the attention weight map, the attention risk index ARI is calculated. When the attention weights show abnormal concentration or abrupt distribution changes on a preset abnormal site set, a risk warning signal is still output even if R′ falls within the threshold range. When the deep neural network model includes a multimodal network with a cross-attention mechanism, it is used to jointly analyze the methylation probability distribution matrix and the H3K27me3 signal of the corresponding region and output the attention consistency index ACI. When both the ACI and the PRI constructed based on the Jensen-Shannon distance meet the preset threshold, a high-confidence judgment result of mechanism-specific remodeling is output.
9. The method of claim 7, wherein, The preset methylation level range is divided into five intervals: [0,0.2), [0.2,0.4), [0.4,0.6), [0.6,0.8), and [0.8,1.0]. The H3K27me3ChIP-seq signal intensity of the target region is extracted to form a second modality input vector, which is then input into a deep neural network model with a self-attention mechanism.
10. An evaluation system for evaluating the effect of improving senescence- related phenotypes, characterized by, include: A sequencer used to sequence a set of PRC2 trisubunit co-binding hypomethylated regions selected from a sample of a test subject and to output sequencing data; The processor is configured to perform the steps of the method according to any one of claims 3 to 9; the output device is used to display the R value, PRI value, risk control determination result, mechanism-specific remodeling determination result, and a prompt indicating whether the R value falls into a preset window; the processor also stores or can call parameters of a deep neural network model containing a self-attention mechanism for outputting epigenetic embedding vectors, R′ values, attention weight maps, or attention risk index ARI.