Multi-omics data fusion analysis method for EDTA (Ethylene Diamine Tetraacetic Acid) induced cell differentiation in dental pulp blood supply reconstruction
By employing a multi-omics data fusion analysis method, a dynamic multi-layered regulatory network induced by EDTA in dental pulp revascularization was constructed. This solved the problem of the hierarchical regulation of dental pulp stem cell differentiation into dentin, which is difficult to analyze in existing technologies. It enabled the systematic analysis of the dynamic patterns after EDTA treatment, the identification of key regulatory modules and pathways, and support for dental pulp functional reconstruction.
Patent Information
- Application Number
- CN202511434589.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-09
AI Technical Summary
Existing studies are unable to systematically elucidate the transcriptional-protein-metabolic hierarchical regulatory relationships and dynamic evolution of dentin differentiation in dental pulp stem cells (DPSCs) after EDTA treatment. They mostly focus on single molecules or single mic levels and lack systematic analysis.
Using a multi-omics data fusion analysis method, transcriptomic, proteomic, and metabolomic data were collected simultaneously at multiple time points by setting up EDTA treatment and control groups. A dynamic multi-layer regulatory network was constructed to identify key regulatory modules and predict key pathways, thereby achieving a systematic analysis of hierarchical regulatory mechanisms.
Breaking through the limitations of single-dimensional research, this study systematically analyzes the dynamic patterns of EDTA-induced differentiation of dental pulp stem cells, identifies key regulatory modules and pathways, and provides a scientific basis for dental pulp revascularization.
Smart Images

Figure CN121306268A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of dental pulp revascularization technology, and more specifically, relates to a multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization. Background Technology
[0002] While traditional root canal treatment can control infection after pulp tissue damage or infection, it can lead to loss of pulp function, particularly negatively impacting the development of young permanent teeth. Pulp revascularization techniques, by inducing the proliferation and differentiation of dental pulp stem cells (DPSCs) to form functional pulp-like tissue, have become an important alternative to traditional treatments. Ethylenediaminetetraacetic acid (EDTA), a commonly used root canal irrigator, not only removes smears and inhibits bacteria, but recent studies have also found that it can promote the differentiation of DPSCs into odontoblasts by regulating the extracellular microenvironment (such as releasing growth factors and altering calcium ion concentration). However, its specific regulatory mechanism is not yet fully understood.
[0003] Odontogenic differentiation of DPSCs is a complex and dynamic process involving the synergistic effects of multiple genes, proteins, and metabolites: at the transcriptional level, the expression of key odontogenic differentiation genes (such as DSPP and DMP1) is regulated by transcription factors (such as Runx2 and Osterix); at the protein level, protein activation and modification (such as phosphorylation) of signaling pathways (such as Wnt and BMP) mediate signal transduction; at the metabolic level, changes in energy metabolism (such as glycolysis and oxidative phosphorylation) and small molecule metabolites (such as amino acids and lipids) provide the material and energy basis for differentiation. However, existing studies mostly focus on single molecular or single-mic levels, making it difficult to systematically analyze the hierarchical regulatory relationships and dynamic evolution of "transcription-protein-metabolism" after EDTA treatment. Summary of the Invention
[0004] This invention provides a multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization, aiming to solve the technical problem that current methods, which focus on single molecules or single omics levels, are unable to systematically analyze the hierarchical regulatory relationship and dynamic evolution law of "transcription-protein-metabolism" after EDTA treatment.
[0005] A multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization includes the following steps: S1. Based on primary human dental pulp stem cells, EDTA treatment group and control group were set up, and samples were collected synchronously at multiple time points to generate raw omics data of transcriptome, proteome and metabolome. Quantitative phenotypic data that are spatiotemporally synchronized with the omics data were generated through quantitative detection methods. S2. The raw omics data of transcriptomics, proteomics and metabolomics are standardized respectively. Then, differential molecules between the EDTA treatment group and the control group are screened by time series differential analysis. The time response pattern of molecules is identified by dynamic trend analysis. Molecules strongly correlated with core phenotypes are screened by phenotypic association analysis. A list of differential molecules with spatiotemporal dynamic characteristics and phenotypic association and time series data are output. S3. Based on the obtained differential analysis list and time series data, a dynamic multilayer regulatory network is constructed, including intralayer networks composed of transcriptional, protein, and metabolic layers, as well as interlayer connections. The intralayer and interlayer networks at each time point are then integrated to form a time series network snapshot, resulting in a dynamic multilayer biological network model. S4. Based on dynamic multilayer biological network models, quantitative phenotypic data, and molecular time series data, the pivotal nodes, bottleneck nodes, and core modules at each time point are identified through network topology. The temporal variation patterns of nodes and modules and their correlation with odontogenic differentiation phenotypes are analyzed through dynamic trajectory analysis. Core regulatory modules with topological importance, dynamic synchronization, and functional enrichment are screened out. Based on cross-level pathway construction and validation, the key regulatory pathways for EDTA-induced differentiation of dental pulp stem cells are predicted.
[0006] This invention, by setting up EDAT treatment and control groups and multiple time points, simultaneously collects multi-omics raw data and quantitative phenotypic data, providing a multi-level data foundation for spatiotemporal matching to analyze dynamic patterns. Then, through standardization and differential analysis, it screens out differentially expressed molecules that realize corresponding patterns and are strongly correlated with phenotypes, focusing on core regulatory candidates and avoiding the limitations of single-omics studies. A dynamic multi-layered regulatory network containing transcriptional, protein, and metabolic layers and their interlayer connections is constructed, and time snapshots are integrated to directly characterize the hierarchical regulatory relationship of transcription-protein-metabolism. Finally, network topology analysis identifies key nodes and modules, and dynamic trajectory analysis tracks their temporal evolution, ultimately screening core regulatory modules and predicting cross-level key pathways. This achieves a systematic analysis from multi-omics data to hierarchical regulatory mechanisms and dynamic patterns, effectively overcoming the limitations of single-dimensional research.
[0007] Preferably, in step S1, samples are collected at 6 time points, covering the early response, mid-term signal transduction and late differentiation stages of EDTA stimulation.
[0008] Preferably, the standardization process in step S2 includes: Transcriptome: After FastQC quality check and Trimmomatic cleaning of low-quality sequences, the transcriptome was aligned with the reference genome by STAR / HISAT2, and a raw counting matrix was generated using featureCounts. Finally, it was normalized to a transcriptome expression matrix using the TMM method. Proteome: Peptides were identified by Mascot, abundance was calculated using LFQ, missing values were filtered out and imputed with KNN, and then normalized by VSN and batch effect correction to obtain the proteome abundance matrix. Metabolomics: Peak information was extracted and aligned using XCMS, missing peaks were filtered after metabolite annotation, and the peak intensity matrix was obtained by semi-minimum interpolation and generalization.
[0009] Preferably, the standardization process in step S2 includes the following steps: Time series differential analysis: A linear model was used to analyze the differences between the EDTA treatment group and the control group, where time nodes were used as ordered factors to retain time series information; the significance of the interaction term was calculated by F test or t test, and after FDR correction, molecules whose corrected p-values met the predetermined threshold and whose fold change met the predetermined fold change threshold were selected to obtain a list of differentially expressed genes, proteins and metabolites, i.e., a list of differentially expressed molecules. Time-based dynamic trend analysis: Based on the list of differential molecules, the standardized expression / abundance values at all times are extracted. First, unsupervised clustering is used to cluster the time series data. Then, based on the predefined dynamic pattern, the corresponding mathematical model is selected for fitting. Finally, by calculating the goodness of fit, molecules with a goodness of fit greater than the threshold are screened to obtain a subset of molecules with significant dynamic patterns. Phenotypic strong association analysis: Based on a subset of molecules with significant dynamic patterns and quantitative phenotypic data, the time series expression values of molecules in the EDTA-treated group are extracted, the Spearman correlation coefficient with the core phenotypic index is calculated, and molecules with correlation coefficients greater than the correlation threshold and p-values less than the predetermined value are selected. In this way, molecules that simultaneously satisfy differential analysis, dynamic trends and phenotypic association are obtained. Based on this, a list of differentially expressed molecules with spatiotemporal dynamic characteristics and phenotypic association and time series data are obtained.
[0010] Preferably, the construction of the intra-layer network includes the following steps: Intralayer networks for the transcriptional, protein, and metabolic layers were constructed based on a list of differentially expressed molecules with spatiotemporal dynamics and phenotypic associations, and time-series data. The transcription layer uses differentially expressed transcription factors and target genes as nodes. The regulatory relationship between transcription factors and target genes is obtained using the TRRUST and ENCODE databases, and directed edges are constructed using Cytoscape or igraph packages. The protein layer uses differentially expressed proteins as nodes and constructs undirected or directed interaction edges based on high-confidence PPI data from the STRING database. The metabolic layer uses differentially metabolites as nodes and obtains the reaction relationship between enzymes and metabolites by mapping the KEGG pathway, thus constructing a metabolic reaction network.
[0011] Preferably, the inter-layer connections are constructed with intra-layer network nodes as the core: Gene identifiers are mapped to protein identifiers through the UniProt database, forming directed edges corresponding to genes and proteins; Based on KEGG enzyme numbering, proteins and metabolites are linked to construct directed edges for enzyme-metabolite catalysis. By mining literature or using the HMDB database, we can obtain the regulatory relationships between metabolites and proteins / transcription factors, and form directed edges for metabolic feedback regulation.
[0012] Preferably, step S4 includes the following steps: Network topology analysis: Based on a dynamic multilayer biological network model using time series, the degree of nodes in the network is calculated at each time point. A threshold is determined based on the average degree and the label difference, and nodes with a degree higher than the threshold are designated as hub nodes. Similarly, a threshold is determined based on the average betweenness centrality and the label difference, and nodes with betweenness centrality higher than the threshold are designated as bottleneck nodes. The Louvain algorithm is then used to identify the community structure, selecting the top 10% of modules with a degree greater than the threshold as core modules, and recording the constituent nodes and functional enrichment results of the core modules. Dynamic trajectory analysis: Based on the obtained hub nodes, bottleneck nodes, and core modules, combined with quantitative phenotypic data and molecular time series data, the dynamic characteristics of nodes and modules are analyzed to obtain node dynamic data and module stability results; Core control module selection: Based on node dynamic data and module stability results, core control modules are selected through multi-dimensional screening. Key regulatory pathway prediction: Based on the screened core regulatory modules, key transcription factors are selected from the core modules of the transcriptome as starting points to construct cross-level pathways of transcription factors, protein kinases, metabolic enzymes and metabolites. The pathways satisfy the condition that the weight of the edges is greater than a preset threshold and the time span is less than or equal to the maximum time span. Then, the top k candidate pathways are selected by comprehensive pathway scoring. After verification by phenotypic relevance, literature support and dynamic robustness, a third set of regulatory pathways covering at least two omics levels and including hub / bottleneck nodes is output.
[0013] Preferably, the dynamic trajectory analysis includes the following steps: Time series curves of expression levels of pivot nodes and bottleneck nodes were plotted, trends were analyzed by linear regression, and Pearson correlation coefficients with odontogenic differentiation phenotypes were calculated. For the core module, plot the expression trend of the internal nodes, and evaluate the dynamic changes in the module's stability through the Pearson correlation coefficient of the nodes; The weight data of the edges within the core module are extracted, the weight change curve is plotted, and the edges with significant changes are identified by t-test. Based on this, the dynamic change curve of the nodes, the correlation coefficient with the phenotype, the stability trend of the module, and the statistical results of the significant changes in edge weights are obtained.
[0014] Preferably, the selection of the core control module includes the following steps: Calculate the normalized topological importance of nodes at all time points, and retain nodes whose normalized topological importance is greater than a threshold; Calculate the dynamic temporal regularization distance between node expression level and phenotype, and retain nodes whose dynamic temporal regularization distance is less than the coefficient multiplied by the maximum dynamic temporal regularization distance; Calculate the weighted correlation coefficient between the average trajectory and phenotype of each module, and retain modules whose weighted correlation coefficient is greater than or equal to a preset threshold. Modules enriched in tooth differentiation-related pathways were screened using hypergeometric verification and corrected p-values.
[0015] Preferably, the step of screening key transcription factors from the core transcriptome module as a starting point includes: Key transcription factors that simultaneously meet the following conditions are selected: their expression level after EDTA treatment is less than a predetermined threshold with the corresponding p value, and the change is less than a predetermined multiple; secondly, the number of downstream target genes they regulate is greater than a predetermined number, and their correlation with the odontogenic differentiation phenotype is greater than a predetermined threshold with the corresponding p value being less than a predetermined threshold. If multiple key transcription factors that simultaneously meet the above conditions are selected, the key transcription factor that regulates the most downstream target genes, has the highest correlation, or has the largest change in expression level should be chosen as the starting point.
[0016] The beneficial effects of this invention include: This invention, by setting up EDAT treatment and control groups and multiple time points, simultaneously collects multi-omics raw data and quantitative phenotypic data, providing a multi-level data foundation for spatiotemporal matching to analyze dynamic patterns. Then, through standardization and differential analysis, it screens out differentially expressed molecules that realize corresponding patterns and are strongly correlated with phenotypes, focusing on core regulatory candidates and avoiding the limitations of single-omics studies. A dynamic multi-layered regulatory network containing transcriptional, protein, and metabolic layers and their interlayer connections is constructed, and time snapshots are integrated to directly characterize the hierarchical regulatory relationship of transcription-protein-metabolism. Finally, network topology analysis identifies key nodes and modules, and dynamic trajectory analysis tracks their temporal evolution, ultimately screening core regulatory modules and predicting cross-level key pathways. This achieves a systematic analysis from multi-omics data to hierarchical regulatory mechanisms and dynamic patterns, effectively overcoming the limitations of single-dimensional research. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A simplified block diagram illustrating the overall steps provided in this embodiment of the invention.
[0019] Figure 2 This is an example diagram of a multilayer control network provided in an embodiment of the present invention.
[0020] Figure 3 A simplified block diagram of step S4 provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.
[0022] See Figure 1 As shown, the multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization includes the following steps: S1. Dynamic Sample Design and Multi-omics Data Generation Primary human dental pulp stem cells (DPSCs) were obtained and the DPSCs were evenly seeded into multiple culture containers (such as multi-well plates). Two experimental groups were set up: EDTA treatment group (culture medium containing 17% EDTA) and control group (culture medium containing an equal amount of PBS). Sampling time points are set, for example, 0 hours, 6 hours, 24 hours, 48 hours, 72 hours, and 7 days, covering the early response, mid-term signal transduction, and late differentiation stages of EDAT stimulation; Spatiotemporally synchronized sample collection was performed at each preset time point using the following method: At each preset time point, the processing of the corresponding culture well is terminated synchronously to ensure that all samples at the same time point reflect the cell state at the same moment. Samples were distributed from parallel culture wells at the same time point and in the same treatment group to different downstream analytical directions: A portion of the samples are used for omics analysis (RNA-seq, proteomics, metabolomics), and need to be processed immediately with the corresponding reagents and stored in a standardized manner. The corresponding reagent processing includes RNA Later to preserve RNA, RIPA to lyse proteins or methanol to quench metabolites, and storage conditions are -80℃ or liquid nitrogen storage. Another portion of the samples are used for phenotypic validation, which requires that the samples come from the same batch of parallel cultures as the omics samples. Phenotypic validation includes immunofluorescence fixation, RNA / protein extraction by qPCR / WB, or live cell function detection. All samples are uniformly labeled, including treatment group, time point, and biological duplicate number, to ensure that each sample can be traced back to a specific culture well and time point, and data collection is completed based on this. After sample collection, standardized tests were performed on the samples for omics analysis to generate raw data for transcriptomics, proteomics, and metabolomics, including: Transcriptome data generation: The quality of the preserved RNA samples was tested (RIN value ≥7.0), and sequencing libraries were constructed by polyA enrichment or rRNA removal. Paired-end sequencing was performed using the Illumina platform (≥30M reads / sample), and the raw sequences in FASTQ format were finally produced. Proteome data generation: After quantification of protein samples, trypsin digestion was performed, and peptides were separated by liquid chromatography. Data-dependent acquisition was then performed using a high-resolution mass spectrometer to produce RAW format mass spectra (including peptide molecular weight and fragment information). Metabolomics data generation: Metabolite extracts were reconstituted and centrifuged, separated by liquid chromatography, and then subjected to full scan and MS / MS fragmentation detection by mass spectrometry to generate RAW format metabolomics mass spectrometry data.
[0023] In parallel with the generation of multi-omics data (generating raw data for transcriptomics, proteomics, and metabolomics), quantitative phenotypic data are generated using traditional experimental methods based on the collected phenotypic test samples. This data is used for the subsequent validation and correlation of omics results, as detailed below: The mRNA levels of key genes for odontogenic differentiation (DSPP, DMP1, etc.) were quantified by qPCR, and the expression and modification (such as phosphorylation) of the corresponding proteins were detected by Western blot or immunofluorescence. Alizarin red staining was used to quantify the area of mineralized nodules (reflecting mineralization capacity), CCK-8 / EdU was used to detect cell proliferation, and Annexin V / PI was used to detect apoptosis. All phenotypic data were quantified (e.g., gray value, positive rate, absorbance), and the number of experimental replicates and coefficient of variation were recorded. The quantitative phenotypic data generated is from the same source as the sample data generated by the multi-omics data, and the two are synchronized in the time dimension, that is, the phenotypic data at the same time point corresponds to the omics data; the generated phenotypic data is not only the "validation standard" of the omics data, but also the key basis for subsequent dynamic trajectory association analysis. After the omics and phenotypic data are generated, all data are summarized, namely, the original omics data (transcriptomics FASTQ file, proteomics / metabolomics RAW file), the quantitative phenotypic data (qPCR relative expression level, WB gray value, mineralization area percentage, etc.), and the metadata (sample quality information such as RNA RIN value, experimental conditions, number of repetitions, etc.).
[0024] This embodiment is designed with multiple time points to cover the entire life cycle dynamics, avoiding the omission of key molecular events due to missing time points. Furthermore, spatiotemporal synchronous acquisition ensures that the sources of omics data and phenotypic data are consistent, eliminating errors caused by sample heterogeneity from the source. In addition, biological significance is anchored in advance through phenotypic data, preventing omics data from becoming a collection of numbers without biological associations.
[0025] S2. Single-omics data processing and spatiotemporal correlation feature selection In step S1, raw omics data and quantitative phenotypic data were obtained, including raw omics data such as transcriptome FASTQ files, proteome RAW files, and metabolome RAW files. For transcriptome data, the FastQC tool was used to check the quality distribution of the original FASTQ file, and then the Trimmomatic tool was used to remove low-quality bases, adapter sequences, and excessively short reads. Low-quality bases were defined as Phredscores <20 or 30, and excessively short reads were reads less than 50 bp in length. Based on this, a high-quality cleaned FASTQ file was obtained. Based on the cleaned high-quality FASTQ file, the cleaned reads are aligned to the reference genome using STAR or HISAT2 tools to generate alignment result files in SAM / BAM format. Based on the alignment results file, featureCounts is used to count the raw count of each gene, and the gene type and gene ID are specified to obtain the raw count matrix for each sample. The TMM (Trimmed Mean of M-values) method was then used to calculate the normalization factor, converting the original counts into normalized expression values, and thus obtaining the labeled transcriptome expression matrix.
[0026] For proteomic data, the Mascot search engine was used to identify peptides and proteins based on the Uniprot database. The false discovery rate (FDR) was set to 1% to ensure the accuracy of the identification results, and a list of identified peptides and proteins was obtained based on this. Based on the obtained peptide and protein lists, the relative abundance of proteins was calculated using the LFQ (Label-Free Quantification) method to generate a protein abundance matrix for each sample. Proteins missing in more than 50% of the samples were then filtered out (i.e., a protein missing in more than 50% of the samples was removed from the dataset), and the remaining missing values were imputed using the KNN (k-Nearest Neighbor) algorithm to obtain a complete protein abundance matrix. Finally, the protein abundance data were standardized using the VSN (variance-stabilizing normalization) method, and robust regression was used to adjust the data to eliminate batch effects, resulting in a standardized proteome abundance matrix.
[0027] For metabolomics data, XCMS was used to extract peak information of metabolites, and parameters (e.g., --ppm 0.01 --sn 3 --bw 0.1) were set for peak alignment to obtain an aligned peak intensity matrix. Then, based on the exact mass (m / z) and retention time, combined with secondary mass spectrometry information, metabolites were annotated, with the annotation criteria being an exact mass deviation of ±5 ppm and a retention time deviation of ±0.5 minutes. Based on this, a list of annotated metabolites and their corresponding peak intensity matrices were obtained. Peaks missing in >50% of the samples were then filtered out (i.e., if a peak is missing in more than 50% of the samples, it is removed from the dataset), and semi-minimum interpolation was used to obtain a complete metabolite peak intensity matrix. Finally, the summation normalization method was used to normalize the peak intensity of each sample to a relative proportion to obtain a normalized metabolomics peak intensity matrix.
[0028] Time-series difference analysis: To analyze the differences between the EDTA treatment group and the control group, and considering time dependence, the following linear model was used: ; In the formula: This represents the normalized expression / abundance of the i-th group (EDTA treatment group or control group), the j-th time point, and the k-th sample; This represents the global mean. Indicate the effect of group i; Indicates the time point effect (0h, 6h, 24h, 48h, 72h, 7d); This represents the interaction effect between the i-th group and time point j; Let represent the residual term of the i-th group, the j-th time point, and the k-th sample; The transcriptome was implemented using DESeq2 to achieve the above linear model, and the proteome and metabolome were implemented using limma to achieve the above linear model; and in this embodiment, the time points (0h, 6h, 24h, 48h, 72h, 7d) were treated as ordered factor variables to preserve the order information of the time series. The interaction term for each analysis was calculated using the F-test or t-test. The statistical significance was determined, and the fold change and correction were calculated using the FDR test. value; If the correction The value satisfies or And the change in the difference multiple satisfies or If the value is significantly different, it is considered a molecule with a significant difference; the corrected p-value is used to control the false positive rate. Based on this, we obtain a list of differentially expressed genes, a list of differentially expressed proteins, and a list of differentially expressed metabolites.
[0029] Time-based dynamic trend analysis: Standard expression / abundance values of differentially expressed molecules at all time points in the EDTA-treated and control groups were extracted from the list of differentially expressed molecules. Unsupervised clustering methods (such as K-means, Mfuzz, etc.) were used to cluster the time series data of differentially expressed molecules. Based on the clustering, the cluster ID of each molecule and the center or representative time series of the cluster were obtained. In dynamic pattern recognition based on clustering results, this embodiment first predefines the following dynamic pattern types: Early transient response: Molecules change significantly at early time points (e.g., 6h or 24h) and then recover; Continuous change: The molecules begin to change continuously from an early time point until the end of the experiment; Peak response: The molecule reaches a peak value at a certain time point, and then decreases; Hysteresis response: Molecules show significant changes only at later time points (such as 72 hours or 7 days); Based on predefined pattern types, select the corresponding fitting mathematical model, for example: Early transient responses use double exponential functions or piecewise linear models; continuous changes use linear regression or exponential growth models; kurtotic responses use Gaussian or quadratic functions; and hysteretic responses use delayed linear or exponential models. For each molecule's time series data, the selected fitting model described above is used for fitting, and the goodness of fit is calculated: ; In the formula: This represents the expression value of molecule m at time t; Indicates the fitted value; This represents the mean of the numerator m; Indicates goodness of fit; Based on the above calculation of goodness of fit, the goodness of fit for each mode is obtained, and the mode corresponding to the highest goodness of fit value is selected as the mode type of the molecule; and a goodness of fit threshold such as 0.7 is set; if the goodness of fit is less than 0.7, it is considered that there is no significant mode; based on this, a subset of molecules with significant dynamic modes is obtained.
[0030] Phenotypic strong association analysis was performed based on a list of differentially expressed molecules. First, the standardized expression / abundance values of molecules at each time point in the EDTA-treated group were extracted from the list. This ensured that the core phenotypic index data corresponded to the molecular data at each time point, meaning that each time point had corresponding molecular expression and phenotypic index values. The core phenotypic index data referred to were those measured in the experiment, such as transcriptome data, expression levels of DSPP and DMP-1, mineralization degree, calcified nodule formation, or mineralization-related indicators. For each differentially expressed molecule, the Spearman correlation coefficient between its expression level at each time point in the EDTA group and the core phenotypic indices was calculated. : In the formula: The value represents the difference between the rank of the expression value of molecule m at time point t and the rank of the phenotypic index x at time point t; T represents the total number of time points. The p-value is calculated based on the t-test to test whether the correlation coefficient is significantly different from zero. ; In the formula: t represents the test statistic; T represents the total number of time points; based on the t-value and degrees of freedom Find the corresponding p value; Based on the calculated correlation coefficients and test statistics, molecules strongly correlated with the core phenotypic indicators were selected. The specific selection criteria are as follows: Correlation strength: Significance: p-value less than 0.05; Based on the above screening criteria, molecules that meet the conditions are selected; and based on the molecules that simultaneously meet the criteria of differential analysis, temporal dynamic trends and phenotypic associations, a list of differentially expressed molecules with spatiotemporal dynamic characteristics and phenotypic associations and time series data are obtained.
[0031] In this embodiment, through standardization and multi-layer screening, the core regulatory molecules are precisely focused. By introducing an interaction term linear model through time series difference analysis, compared with single time point difference analysis, it can better capture the time-dependent differences of EDTA treatment and avoid missing molecules that become increasingly significant over time. Secondly, dynamic trend analysis classifies differential molecules according to their time behavior through aggregation and pattern fitting, revealing the temporal patterns of molecule participation in differentiation. Phenotypic strong association analysis further screens out core molecules directly related to the odontogenic phenotype from the differential molecules, reducing data redundancy.
[0032] S3. Construct a dynamic multi-level control network Based on the list of differentially expressed molecules with spatiotemporal dynamic characteristics and phenotypic associations and time series data obtained in step S2, a dynamic multilayer regulatory network is constructed, including intralayer network construction, interlayer network construction and temporal network integration. See Figure 2 As shown, the intralayer network construction includes a transcription layer, a protein layer, and a metabolic layer; The transcription layer uses differential polarity and differentially expressed transcription factors as nodes. The regulatory relationship between transcription factors and target genes is obtained using the TRRUST and ENCODE databases. Then, directed edges are constructed using Cytoscape or the igraph package in R. The protein layer uses differentially expressed proteins as nodes. In this embodiment, PPI data is downloaded from the STRING database, high-confidence interactions are screened, and undirected or directed edges (based on functional or physical interactions) are constructed. The metabolic layer uses differential metabolites as nodes to map metabolites to KEGG pathways, obtain the reaction relationship between enzymes and metabolites, and construct a metabolic reaction network.
[0033] The interlayer network connections include gene-protein correspondence, enzyme-metabolite catalysis, and metabolic feedback regulation; The gene-protein mapping uses the UniProt database to map gene identifiers (such as Entrez ID) to protein identifiers (such as UniProt ID) and construct directed edges. The enzyme-metabolite catalysis uses KEGG enzyme numbering to link proteins (enzymes) and metabolites, constructing directed edges; The metabolic feedback regulation is constructed by obtaining the regulatory relationship between metabolites and proteins or transcription factors based on literature mining or databases (such as HMDB) and constructing directed edges.
[0034] The time network is integrated by combining transcription, protein, metabolic layers, and interlayer connections at each time point to form an independent network snapshot, which is then stored as a graph database. Based on this, a dynamic multilayer biological network model of time series is obtained, including nodes, intralayer edges, interlayer edges, and their time point states.
[0035] In this embodiment, the limitations of a single omics are overcome by integrating multiple layers of networks and embedding time dimensions. The intra-layer network captures the regulatory logic within each omics, ensuring the biological rationality of intra-layer relationships. Inter-layer connections break down omics barriers, solving the problem that a single omics cannot reflect cross-layer regulation. Furthermore, the time-based network integration upgrades the static network into a dynamic model, capturing network structure changes at different differentiation stages and avoiding mechanism misjudgments caused by ignoring the time dimension.
[0036] S4. Network Topology and Dynamic Trajectory Analysis See Figure 3 As shown, network topology analysis: Based on a dynamic multilayer biological network model with time series, the degree of all nodes in the network at each time point is calculated, i.e., the number of connections between each node and other nodes. The average degree of all nodes is then calculated. and standard deviation And calculate the threshold based on the mean and standard deviation. , where k represents the coefficient, which is a positive integer; nodes with a degree higher than the threshold are marked as hub nodes; The bottleneck node is then determined by the betweenness centrality of the nodes: the betweenness centrality of all nodes is calculated, and then the mean and standard deviation of the betweenness centrality of all nodes are calculated; a threshold is calculated based on the mean and standard deviation, and the threshold is calculated in the same way as the threshold of the degree. Nodes with betweenness centrality higher than the threshold are marked as bottleneck nodes. The clarity of dynamic multilayer biological network models is measured using the Louvain algorithm, a community detection algorithm. ; In the formula: Q represents the modularity, and the larger the value, the clearer the community structure; This represents the adjacency matrix, indicating whether there is an edge connecting nodes i and j; and represents the degree of node i and node j; m represents the total number of edges in the network; This indicates an indicator function; it is 1 if node i and node j are in the same community, and 0 otherwise.
[0037] Based on the calculated modularity, the modularity is sorted from high to low, and the top 10% of communities are selected as core modules. Alternatively, a modularity threshold (0.3 or 0.4) can be set, and those exceeding the threshold are designated as core modules. The constituent nodes of each core module and their functional enrichment results are recorded. Based on this, we obtain the list of hub nodes, bottleneck nodes, core modules in the network, and their constituent nodes at each point in time.
[0038] Dynamic trajectory analysis: This mainly analyzes the expression or abundance changes of pivot nodes and bottleneck nodes in time series, and identifies the correlation between their dynamic behavior and odontogenic differentiation phenotype. Specifically, it includes the following steps: Obtain node expression levels or abundance data at each time point, including the expression levels of differentially expressed genes, proteins, and metabolites; use the Pandas library to read the data and the Matplotlib library to plot time series curves; observe the changing trends over time by plotting the expression curves of hub nodes and bottleneck nodes. Then, linear regression is used to analyze the trend of each node in the time series. An example linear regression model is as follows: ; In the formula: The expression level of a node is represented by x; x represents a time point. Indicates the intercept; Indicates the slope; Indicates the error term; Then, the correlation between the dynamic changes of nodes and the phenotypes of odontogenic components (such as DSPP expression level and mineralization degree) was calculated using the Pearson correlation coefficient: based on this, the dynamic change curves of key nodes and the correlation coefficients between nodes and odontogenic differentiation phenotypes were obtained.
[0039] Based on the core modules identified by the community detection algorithm, the node list of each module is extracted, and the expression level time series change graph of all nodes within each core module is plotted, including the expression level change curve of all nodes within the module. For each core module, the Pearson correlation coefficient of all node pairs within the module is calculated, and a threshold (e.g., 0.7) is set based on the biological background and statistical significance as a stability criterion. If the average correlation coefficient of the nodes within the module is higher than the threshold, the module is considered stable during that time period; otherwise, the module is considered to have low stability. The above calculation is repeated for each time point to evaluate the dynamic changes in module stability, and a curve showing the change in module stability over time is plotted to demonstrate the trend of module stability.
[0040] Weight data for all edges within the core module of a multilayer biological network model is extracted, including intra-layer and inter-layer edges. A time-series curve of the weight change for each edge is plotted to demonstrate the trend of weight variation over time. Statistical methods (t-test) are then used to identify significant changes in edge weights at specific time points. ; In the formula: and These are the averages for two different time periods; and These represent the sample variances for the two time periods, respectively. and Indicates the number of samples in two time periods; Then, based on the calculated t-value and the degrees of freedom (using conventional techniques), the p-value is obtained. The p-value is compared with a set threshold. If it is less than the threshold, the difference is considered significant; otherwise, it is not significant. Based on this, the number and proportion of edges that have changed significantly are statistically analyzed.
[0041] Core control module selection: First, select nodes with stable high topological importance at multiple time points. For each node i, calculate its importance at all time points. Normalized topological importance: In the formula: This indicates the degree centrality of node i at time t; This indicates the betweenness centrality of node i at time t; This represents the maximum degree value of all nodes at time point t; This represents the maximum value of all node intermediate values at time point t; and Indicates the weighting coefficient. and The sum of is 1; Indicates the total number of time points; reserve The node, in this embodiment, has a threshold. The value is 0.8, meaning its importance consistently exceeds the 80th percentile. Recalculate node expression levels Dynamic time-normalized distance from phenotypic indicators: ; In the formula: Indicates the timing alignment path; and These represent the point-in-time indexes for nodes and tables, respectively. Indicates a point in time The node expression level; Indicates a point in time Phenotypic indicators; This represents the dynamic time-normalized distance between node expression levels and phenotypic metrics; In computation-based normalized topological importance For all nodes within module M, calculate the weighted correlation coefficient between the module's average trajectory and the phenotype: ; In the formula: The Pearson correlation coefficient represents the relationship between the node representation vector and the phenotypic vector. reserve The nodes, among which , Indicates the threshold; reserve The nodes, among which , Indicates the threshold; reserve The module, in which , representing the threshold; Finally, the machine uses the hypergeometric test calculation module to enrich the p-value of pathway P, and performs functional enrichment screening based on this: ; In the formula: N represents the total number of nodes in the network; n represents the number of nodes in module M; This represents the number of relevant nodes for path P in the entire network; The p-value represents the hypergeometric test result; K represents the number of nodes in module M belonging to path P; k represents the strong sum variable, from... arrive , representing the value of the number of nodes in the module that actually belong to path P; Suppose there are m pathways. Sort the original p-values of the m pathways in ascending order, and calculate the corrected p-value for the i-th p-value after sorting: ; In the formula: This represents the original p-value at the i-th position after sorting; m represents the total number of tests, i.e., the total number of pathways; i represents the sorting position of the current p-value. If a correction value is greater than a subsequent correction value, then the minimum subsequent value is taken: This ensures that the corrected sequence is non-decreasing; and the upper limit of all correction p-values is 1. like ,in If the value is 0.05, the corresponding module will be retained; Finally, based on topological importance screening, dynamic trajectory synchronization screening, cross-level correlation screening, and functional enrichment screening, the final output module is obtained, which satisfies the following conditions: containing at least one TF, one metabolic enzyme, one metabolite, and an average node average within the module. It was greater than 0.8, had an MSI greater than 0.7 with DSPP expression, and was significantly enriched in tooth differentiation-related pathways.
[0042] Secondly, this invention introduces time-weighted average (NTII) to address the node importance drift problem in dynamic networks, and uses DTW distance to solve the temporal phase drift problem, while MSI weighted correlation highlights the contributions of core nodes. Furthermore, it automatically filters out modules with signal transduction to metabolic execution capabilities through cross-level correlation dimension screening, excluding single-mathematics-dominated local modules. Based on this, biological prior knowledge is transformed into computable mathematical constraints, and temporal patterns are captured through dynamic network analysis. Finally, the anchoring function mechanism is verified through statistical validation. This combination of prior guidance and data-driven approach ensures interpretability of results while maximizing the mining of data value.
[0043] Key regulatory pathway prediction: Based on the above steps, a dynamic multilayer biological network model was obtained, and core regulatory modules and node sets were screened. Based on this, key regulatory pathways were predicted, and the specific steps are as follows: Cross-level pathway construction: Key transcription factors were selected from the core regulatory modules of the transcriptome as the starting point, and priority was given to key transcription factors that simultaneously meet the following conditions: their expression level changed significantly after EDTA treatment (p value < 0.05 and the change range ≥ 1.5-fold or 2-fold), they regulated more than 50 downstream target genes, and they were highly correlated with odontogenic differentiation phenotype (|r| ≥ 0.5 and p value < 0.05). Based on the above screening criteria, if multiple key transcription factors (TFs) that simultaneously meet the above criteria are selected, the key transcription factor that regulates the most downstream target genes, has the highest correlation, or has the largest change in expression level should be selected as the starting point. After selecting the starting point, construct a three-level progressive search path: First jump: TF → protein kinase, only retaining the experimentally verified TF-kinase regulatory relationship (PhosphoSitePlus database). Second jump: Kinase → Metabolic enzyme, select metabolic enzymes directly phosphorylated by kinases (verified using the STRING database); The third step: metabolic enzymes → metabolic products, the direct products catalyzed by metabolic enzymes (KEGG reaction pathway). The search path satisfies the following constraints: ; In the formula: This represents the edge weight at time point t; Indicates the weight threshold; Indicates the time span of the path; Indicates the maximum allowed time span; This indicates an indicator function that returns 1 if the condition is met, and 0 otherwise. An initial set of paths is obtained through path search, and then a comprehensive path score is calculated based on this initial set of paths. ; In the formula: Represents a node Betweenness centrality; Represents the total expression level of path nodes With phenotype The Pearson correlation coefficient; Indicates the number of hops across levels in the path; This represents the attenuation coefficient; n represents the number of nodes in the path. This represents the expression level of node v at time point t; Based on the comprehensive path score, all initial paths are sorted in descending order, and the top k paths (top 10 paths) are selected as candidate key control paths. Then, path verification is performed based on the selected candidate key control paths. The total expression level of path P at time t is calculated based on the expression levels of nodes on each path at time t. ; In the formula: This represents the total expression level of path P at time point t; Represents a node The expression level at time point t; n represents the number of nodes contained in path P; Then, the Pearson correlation coefficient was used to assess the correlation between pathway expression levels and phenotypes at time point t: ; In the formula: and These represent the mean values of path expression levels and phenotypes in the time series, respectively. The phenotype at time point t is represented by T; T represents the number of time points. Based on this, paths meeting the criteria were selected: paths with a correlation coefficient with the phenotype greater than or equal to 0.7 were retained, i.e., the Pearson correlation coefficient. Paths with an absolute value greater than or equal to 0.7, and whose relevance is... If the value is less than 0.01, pathways that are not significantly related to the phenotype are removed. Based on this, the first set of key regulatory pathways is obtained. Then, the regulatory relationships in the first set of key regulatory pathways are verified based on the support of the literature to see if they are biologically supported, as follows: For each edge in the path, find its literature support in PubMed or other biomedical databases, and calculate the literature support score for each edge based on this: ; In the formula: Representing an edge The degree of literature support; Indicates the number of relevant literature reports; The document support of path P is obtained by averaging the document support of all edges in path P. ; ; In the formula: Represents the set of edges in path P; Based on this, paths with literature support greater than a predetermined threshold (5) are retained to ensure that the regulatory relationship of the path is supported by sufficient evidence in the literature; based on this, a second set of key regulatory paths is obtained, and then a dynamic robustness assessment is performed based on the second set of regulatory paths: Suppose that in a time series, the adjacency matrix of path P at time point t is... It describes the connectivity between nodes in a path and evaluates the stability of the path by calculating the changes in the path adjacency matrix at different time points. An example of the steps is as follows: Between time t and t+1, the change in the path adjacency matrix is measured by the Frobenius norm, with the specific expression as follows: ; In the formula: Represents nodes in the path and Edge weights at time t; Represents nodes in the path and Edge weights at time t+1; This represents the adjacency matrix of path P at time point t+1; Indicates the number of nodes; The stability of the path is then obtained by calculating the changes in the adjacency matrix of the path at multiple time points, and the stability is evaluated based on the following expression: ; In the formula: This indicates an indicator function, which states that if the change in the adjacency matrix is less than a threshold... If the value is 0.2, then the path is considered stable in the time series. Preserving dynamic robustness Paths exceeding a threshold (e.g., 0.8) are used to obtain the third set of regulatory paths; The final set of third regulatory pathways includes each pathway that: covers at least two omics levels (e.g., transcription, protein, metabolites); and contains at least one hub or bottleneck node.
[0044] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization, characterized in that, Includes the following steps: S1. Based on primary human dental pulp stem cells, EDTA treatment group and control group were set up, and samples were collected synchronously at multiple time points to generate raw omics data of transcriptome, proteome and metabolome. Quantitative phenotypic data that are spatiotemporally synchronized with the omics data were generated through quantitative detection methods. S2. The raw omics data of transcriptomics, proteomics and metabolomics are standardized respectively. Then, differential molecules between the EDTA treatment group and the control group are screened by time series differential analysis. The time response pattern of molecules is identified by dynamic trend analysis. Molecules strongly correlated with core phenotypes are screened by phenotypic association analysis. A list of differential molecules with spatiotemporal dynamic characteristics and phenotypic association and time series data are output. S3. Based on the obtained differential analysis list and time series data, a dynamic multilayer regulatory network is constructed, including intralayer networks composed of transcriptional, protein, and metabolic layers, as well as interlayer connections. The intralayer and interlayer networks at each time point are then integrated to form a time series network snapshot, resulting in a dynamic multilayer biological network model. S4. Based on dynamic multilayer biological network models, quantitative phenotypic data, and molecular time series data, the pivotal nodes, bottleneck nodes, and core modules at each time point are identified through network topology. The temporal variation patterns of nodes and modules and their correlation with odontogenic differentiation phenotypes are analyzed through dynamic trajectory analysis. Core regulatory modules with topological importance, dynamic synchronization, and functional enrichment are screened out. Based on cross-level pathway construction and validation, the key regulatory pathways for EDTA-induced differentiation of dental pulp stem cells are predicted.
2. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 1, characterized in that, In step S1, samples are collected at six time points, covering the early response, mid-term signal transduction, and late differentiation stages of EDTA stimulation.
3. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 1, characterized in that, The standardization process in step S2 includes: Transcriptome: After FastQC quality checks and Trimmomatic cleaning of low-quality sequences, the transcriptome was aligned to the reference genome using STAR / HISAT2, and a raw counting matrix was generated using featureCounts. Finally, it was normalized to a transcriptome expression matrix using the TMM method. Proteome: Peptides were identified by Mascot, abundance was calculated using LFQ, missing values were filtered out and imputed with KNN, and then normalized by VSN and batch effect correction to obtain the proteome abundance matrix. Metabolomics: Peak information was extracted and aligned using XCMS, missing peaks were filtered after metabolite annotation, and the peak intensity matrix was obtained by semi-minimum interpolation and generalization.
4. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 1, characterized in that, The standardization process in step S2 includes the following steps: Time series differential analysis: A linear model was used to analyze the differences between the EDTA treatment group and the control group, where time nodes were used as ordered factors to retain time series information; the significance of the interaction term was calculated by F test or t test, and after FDR correction, molecules whose corrected p-values met the predetermined threshold and whose fold change met the predetermined fold change threshold were selected to obtain a list of differentially expressed genes, proteins and metabolites, i.e., a list of differentially expressed molecules. Time-based dynamic trend analysis: Based on the list of differential molecules, the standardized expression / abundance values at all times are extracted. First, unsupervised clustering is used to cluster the time series data. Then, based on the predefined dynamic pattern, the corresponding mathematical model is selected for fitting. Finally, by calculating the goodness of fit, molecules with a goodness of fit greater than the threshold are screened to obtain a subset of molecules with significant dynamic patterns. Phenotypic strong association analysis: Based on a subset of molecules with significant dynamic patterns and quantitative phenotypic data, the time series expression values of molecules in the EDTA-treated group are extracted, the Spearman correlation coefficient with the core phenotypic index is calculated, and molecules with correlation coefficients greater than the correlation threshold and p-values less than the predetermined value are selected. In this way, molecules that simultaneously satisfy differential analysis, dynamic trends and phenotypic association are obtained. Based on this, a list of differentially expressed molecules with spatiotemporal dynamic characteristics and phenotypic association and time series data are obtained.
5. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 1, characterized in that, The construction of the intra-layer network includes the following steps: Intralayer networks for the transcriptional, protein, and metabolic layers were constructed based on a list of differentially expressed molecules with spatiotemporal dynamics and phenotypic associations, and time-series data. The transcription layer uses differentially expressed transcription factors and target genes as nodes. The regulatory relationship between transcription factors and target genes is obtained using the TRRUST and ENCODE databases, and directed edges are constructed using Cytoscape or igraph packages. The protein layer uses differentially expressed proteins as nodes and constructs undirected or directed interaction edges based on high-confidence PPI data from the STRING database. The metabolic layer uses differentially metabolites as nodes and obtains the reaction relationship between enzymes and metabolites by mapping the KEGG pathway, thus constructing a metabolic reaction network.
6. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 5, characterized in that, The inter-layer connections are built around intra-layer network nodes: Gene identifiers are mapped to protein identifiers through the UniProt database, forming directed edges corresponding to genes and proteins; Based on KEGG enzyme numbering, proteins and metabolites are linked to construct directed edges for enzyme-metabolite catalysis. By mining literature or using the HMDB database, we can obtain the regulatory relationships between metabolites and proteins / transcription factors, and form directed edges for metabolic feedback regulation.
7. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 1, characterized in that, Step S4 includes the following steps: Network topology analysis: Based on a dynamic multilayer biological network model using time series, the degree of nodes in the network is calculated at each time point. A threshold is determined based on the average degree and the label difference, and nodes with a degree higher than the threshold are designated as hub nodes. Similarly, a threshold is determined based on the average betweenness centrality and the label difference, and nodes with betweenness centrality higher than the threshold are designated as bottleneck nodes. The Louvain algorithm is then used to identify the community structure, selecting the top 10% of modules with a degree greater than the threshold as core modules, and recording the constituent nodes and functional enrichment results of the core modules. Dynamic trajectory analysis: Based on the obtained hub nodes, bottleneck nodes, and core modules, combined with quantitative phenotypic data and molecular time series data, the dynamic characteristics of nodes and modules are analyzed to obtain node dynamic data and module stability results; Core control module selection: Based on node dynamic data and module stability results, core control modules are selected through multi-dimensional screening. Key regulatory pathway prediction: Based on the screened core regulatory modules, key transcription factors are selected from the core modules of the transcriptome as starting points to construct cross-level pathways of transcription factors, protein kinases, metabolic enzymes and metabolites. The pathways satisfy the condition that the weight of the edges is greater than a preset threshold and the time span is less than or equal to the maximum time span. Then, the top k candidate pathways are selected by comprehensive pathway scoring. After verification by phenotypic relevance, literature support and dynamic robustness, a third set of regulatory pathways covering at least two omics levels and including hub / bottleneck nodes is output.
8. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 7, characterized in that, The dynamic trajectory analysis includes the following steps: Time series curves of expression levels of pivot nodes and bottleneck nodes were plotted, trends were analyzed by linear regression, and Pearson correlation coefficients with odontogenic differentiation phenotypes were calculated. For the core module, plot the expression trend of the internal nodes, and evaluate the dynamic changes in the module's stability through the Pearson correlation coefficient of the nodes; The weight data of the edges within the core module are extracted, the weight change curve is plotted, and the edges with significant changes are identified by t-test. Based on this, the dynamic change curve of the nodes, the correlation coefficient with the phenotype, the stability trend of the module, and the statistical results of the significant changes in edge weights are obtained.
9. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 7, characterized in that, The screening of the core control module includes the following steps: Calculate the normalized topological importance of nodes at all time points, and retain nodes whose normalized topological importance is greater than a threshold; Calculate the dynamic temporal regularization distance between node expression level and phenotype, and retain nodes whose dynamic temporal regularization distance is less than the coefficient multiplied by the maximum dynamic temporal regularization distance; Calculate the weighted correlation coefficient between the average trajectory and phenotype of each module, and retain modules whose weighted correlation coefficient is greater than or equal to a preset threshold. Modules enriched in tooth differentiation-related pathways were screened using hypergeometric verification and corrected p-values.
10. The multi-omics data fusion analysis method for EDTA-induced cell differentiation in dental pulp revascularization according to claim 7, characterized in that, The starting point for screening key transcription factors from the core modules of the transcriptome includes: Key transcription factors that simultaneously meet the following conditions are selected: their expression level after EDTA treatment is less than a predetermined threshold with the corresponding p value, and the change is less than a predetermined multiple; secondly, the number of downstream target genes they regulate is greater than a predetermined number, and their correlation with the odontogenic differentiation phenotype is greater than a predetermined threshold with the corresponding p value being less than a predetermined threshold. If multiple key transcription factors that simultaneously meet the above conditions are selected, the key transcription factor that regulates the most downstream target genes, has the highest correlation, or has the largest change in expression level should be chosen as the starting point.
Citation Information
Cited By
Bioinformatics analysis method for identifying cerebral hemorrhage prognosis related gene target
CN121601020A