Method for constructing miRNA regulation and control network
By constructing a third-order regulatory tensor and a hierarchical Markov state transition model of the miRNA regulatory network, the problem of inaccurate identification of epigenetic reprogramming and post-transcriptional regulatory interactions in existing technologies was solved, achieving dynamic feature preservation of the liver cancer regulatory network and improving the accuracy of key path identification.
Patent Information
- Application Number
- CN202511124775.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-18
AI Technical Summary
Existing miRNA regulatory network construction methods fail to accurately reflect the interaction between epigenetic reprogramming and post-transcriptional regulation in complex diseases such as liver cancer, resulting in inaccurate identification of key regulatory pathways and a lack of quantitative modeling of the synergistic/competitive relationship between epigenetic and ceRNA mechanisms.
Using a three-order regulatory tensor construction method, combined with EZH2 binding site data and H3K27me3 epigenetic modification data, a mechanism-hierarchical Markov state transition model was constructed through mechanism identification and labeling. Biological credibility-weighted fusion was then performed to generate a fused Markov state diagram, analyze high-frequency transfer pathways, and construct a miRNA regulatory network.
Independent quantitative characterization of epigenetic and ceRNA regulatory mechanisms has been achieved, improving the accuracy of key regulatory pathway identification, providing reliable biomarkers for targeted therapy of liver cancer, and ensuring that dynamic networks accurately capture key state transitions in disease progression.
Smart Images

Figure CN120977403A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, and in particular to a method for constructing an miRNA regulatory network. BACKGROUND
[0002] In recent years, the method for constructing an miRNA regulatory network based on multi-omics data has become an important direction of non-coding RNA research. In the prior art, a typical representative is the "ceRNA network integration analysis method" proposed by Zhou et al., which constructs a static ceRNA interaction network by integrating lncRNA-miRNA-mRNA expression correlation, miRNA binding site prediction and co-expression analysis. This method uses a Bayesian network model to identify potential ceRNA regulatory relationships and has been applied in tumor research such as liver cancer.
[0003] The existing method has obvious limitations in mechanism analysis. Due to the lack of quantitative modeling of the synergistic / competitive relationship between epigenetic mechanisms (such as EZH2-mediated H3K27me3 modification) and ceRNA mechanisms, the dynamic characteristics of the network are not adequately represented. In particular, in complex diseases such as liver cancer, this single mechanism modeling approach cannot accurately reflect the interaction between epigenetic reprogramming and post-transcriptional regulation. Existing network models use an equal-weight fusion strategy and cannot distinguish the contribution of different mechanisms based on biological evidence such as CLIP-seq verified binding strength, which affects the accuracy of identifying key regulatory pathways. SUMMARY
[0004] In view of the above existing problems, the present application is proposed.
[0005] Therefore, the present application provides a method for constructing an miRNA regulatory network to solve the problem of inaccurate identification of key regulatory pathways in liver cancer caused by insufficient quantitative modeling of the synergistic / competitive relationship between epigenetic regulation and ceRNA regulation.
[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a method for constructing an miRNA regulatory network, comprising, Collecting liver cancer regulatory multi-omics data and preprocessing to construct a three-order regulatory tensor; Performing mechanism identification and label annotation on the three-order regulatory tensor to obtain a mechanism-labeled regulatory tensor; Constructing a mechanism-layered Markov state transition model according to the mechanism-labeled regulatory tensor and performing weighted fusion according to biological credibility to generate a fused Markov state graph; Analyzing high-frequency transition paths in the fused Markov state transition graph and converting them into a graph structure form of regulatory relationship to obtain a regulatory edge set; The regulatory edge set is structured by a graph construction to obtain the miRNA regulatory network.
[0007] As a preferred scheme of the method for constructing the miRNA regulatory network, the liver cancer regulatory multi-omics data comprises lncRNA expression data, miRNA expression data, mRNA expression data, EZH2 binding site data and H3K27me3 epigenetic modification data. The preprocessing comprises missing value filling, normalization, sample consistency alignment and expression matrix format standardization.
[0008] As a preferred scheme of the method for constructing the miRNA regulatory network, the method comprises the following steps: Each of the lncRNA expression data, the miRNA expression data and the mRNA expression data in the preprocessed liver cancer regulatory multi-omics data is combined to generate a ternary combination. The ternary combination, the EZH2 binding site data and the H3K27me3 epigenetic modification data are combined to calculate the joint regulatory strength and the mechanism clue score of each group of ternary combinations, and a regulatory score matrix is obtained. A three-order regulatory tensor with the ternary combination as the three-dimensional coordinate axis is constructed based on the regulatory score matrix.
[0009] As a preferred scheme of the method for constructing the miRNA regulatory network, the method comprises the following steps: Based on the EZH2 binding site data and the H3K27me3 epigenetic modification data in the preprocessed liver cancer regulatory multi-omics data, an epigenetic mechanism confidence score of the ternary combination is calculated. A Pearson correlation coefficient of the lncRNA expression data and the miRNA expression data is calculated, and a ceRNA mechanism confidence score is calculated. According to the relative size relationship between the epigenetic mechanism confidence score and the ceRNA mechanism confidence score, a mechanism label is annotated to the ternary combination and integrated into the three-order regulatory tensor to generate a mechanism-annotated regulatory tensor.
[0010] As a preferred scheme of the method for constructing the miRNA regulatory network, the method comprises the following steps: The ternary combinations annotated with the epigenetic mechanism and the ternary combinations annotated with the ceRNA mechanism are separated from the mechanism-annotated regulatory tensor to generate an epigenetic subset and a ceRNA subset. According to the joint regulation intensity of the three-element combination in the apparent subset, the expression level state is divided to obtain an apparent state space; According to the joint regulation intensity of the three-element combination in the ceRNA subset, the expression level state is divided to obtain a ceRNA state space; Based on the preprocessed liver cancer regulation multiomics data, state transition frequencies of the apparent state space and the ceRNA state space are counted respectively to obtain an apparent transition matrix and a ceRNA transition matrix; Based on the Markov chain state transition probability system, the apparent transition matrix and the ceRNA transition matrix are subjected to weighted linear combination to obtain a mechanism hierarchical Markov state transition model.
[0011] As a preferred scheme of the method for constructing the miRNA regulation network, the mechanism is weighted and fused according to biological credibility to generate a fused Markov state diagram, and the specific steps are as follows, Based on the mechanism hierarchical Markov state transition model, weight calculation is performed through regression analysis and statistical test to obtain apparent transition matrix fusion weights and ceRNA transition matrix fusion weights; The apparent transition matrix fusion weights and the ceRNA transition matrix fusion weights are subjected to linear transformation with the corresponding apparent transition matrix and the ceRNA transition matrix respectively through matrix scalar product operation to obtain a weighted apparent transition matrix and a weighted ceRNA transition matrix; The weighted apparent transition matrix and the weighted ceRNA transition matrix are subjected to probability space fusion through matrix element level addition operation to generate a fused Markov state diagram.
[0012] As a preferred scheme of the method for constructing the miRNA regulation network, the mechanism is weighted and fused according to biological credibility to generate a fused Markov state diagram, and the specific steps are as follows, High-frequency state transition paths are extracted from the fused Markov state transition diagram; The high-frequency state transition paths are subjected to molecular expression trend analysis to identify a regulation relationship chain, and subjected to binary relationship conversion and edge weight assignment to obtain a regulation edge set.
[0013] As a preferred scheme of the method for constructing the miRNA regulation network, the mechanism is weighted and fused according to biological credibility to generate a fused Markov state diagram, and the specific steps are as follows, The edge weights of the regulation edge set are subjected to attribute annotation to obtain a weighted edge attribute set; The lncRNA expression data nodes, miRNA expression data nodes and mRNA expression data nodes in the regulation edge set are classified and stored to obtain a structured node set; The topological relationship of the directed edge relationship of the structured node set is generated to obtain a network topology of the directed connection relationship; The weighted edge attribute set, the structured node set and the network topology are integrated to obtain the miRNA regulation network.
[0014] In a second aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the computer program, when executed by the processor, implements any step of the method for constructing a miRNA regulation network according to the first aspect of the present application.
[0015] In a third aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any step of the method for constructing a miRNA regulation network according to the first aspect of the present application.
[0016] The present application has the following beneficial effects: by constructing a mechanism hierarchical Markov state transition model, independent quantitative characterization of epigenetic regulation and ceRNA regulation mechanism is realized, and the dynamic characteristics of different regulation modes are completely retained; by generating a fusion Markov state diagram, the two mechanisms are organically integrated based on a weight fusion algorithm of biological evidence, and a unified dynamic network with clear biological interpretation is established. By mechanism-specific modeling, the identification accuracy of key regulation paths is significantly improved, and more reliable biomarkers are provided for liver cancer targeted therapy; the evidence-driven fusion method established not only ensures that the dynamic network accurately captures the key state transitions in disease progression, but also provides a generalizable technical paradigm for cross-mechanism regulation network research, and exhibits unique value in pathogenesis research and precision medicine applications. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0018] Fig. 1 The flowchart of the method for constructing a miRNA regulation network.
[0019] Fig. 2 The flowchart of generating a mechanism-labeled regulation tensor.
[0020] Fig. 3A flowchart for constructing the mechanism layered Markov model.
[0021] Fig. 4 A flowchart for generating the fused Markov state graph. DETAILED DESCRIPTION
[0022] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0023] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the concept of the present application, so the present application is not limited to the specific embodiments disclosed below.
[0024] Secondly, the "one embodiment" or "embodiment" referred to herein means that a specific feature, structure or characteristic can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is separate or alternative to other embodiments.
[0025] Reference Figs. 1-4 For one embodiment of the present application, the embodiment provides a method for constructing an miRNA regulatory network, comprising the following steps: S1, collecting liver cancer regulatory multi-omics data and pre-processing.
[0026] S1.1, the liver cancer regulatory multi-omics data includes lncRNA expression data, miRNA expression data, mRNA expression data, EZH2 binding site data and H3K27me3 epigenetic modification data.
[0027] S1.2, the pre-processing includes missing value filling, normalization, sample consistency alignment and expression matrix format standardization.
[0028] Specifically, the k-nearest neighbor algorithm is used to fill in the missing values, in the lncRNA expression data, miRNA expression data and mRNA expression data, for each missing sample, select k=10 nearest neighbor samples in the feature space, fill in the missing site with the median of the neighborhood samples, and the missing values of the EZH2 binding site data and the H3K27me3 epigenetic modification data are replaced with zero values of the same omics sample; The lncRNA expression data, miRNA expression data and mRNA expression data are executed TPM standardization, and the original reads are converted into per million transcript count. The EZH2 binding site data adopts RPKM standardization, and the H3K27me3 epigenetic modification data uses the standardization method of per kilobase per million mapped reads. The lncRNA expression data, miRNA expression data, mRNA expression data, EZH2 binding site data and H3K27me3 epigenetic modification data are ensured to come from the same batch of patient samples by sample ID matching, and the data with unmatched sample numbers are excluded; The lncRNA expression data, miRNA expression data and mRNA expression data are uniformly converted into a matrix format with rows representing genes and columns representing samples, and the sample numbers are in the TCGA standard format. The EZH2 binding site data and H3K27me3 epigenetic modification data are converted into BED format to store genomic coordinate information.
[0029] S2, construct a three-order regulation tensor with ternary combination as three-dimensional coordinate axes.
[0030] S2.1, combine each set of lncRNA expression data, miRNA expression data and mRNA expression data in the preprocessed liver cancer regulatory multi-omics data to generate ternary combinations.
[0031] Specifically, in the preprocessed liver cancer regulatory multi-omics data, lncRNA expression data, miRNA expression data and mRNA expression data are extracted. The lncRNA expression data, miRNA expression data and mRNA expression data of each sample are matched according to gene names, and genes with expression greater than 1 TPM are screened. The lncRNA-miRNA-mRNA expression value ternary combination in each sample that matches successfully is taken as a ternary combination, and recorded as a structured data unit of lncRNA name, miRNA name, mRNA name, lncRNA expression value, miRNA expression value, and mRNA expression value. The ternary combination with missing values is excluded to generate ternary combinations.
[0032] S2.2, combine the ternary combinations, EZH2 binding site data and H3K27me3 epigenetic modification data, calculate the joint regulation strength and mechanism clue score of each set of ternary combinations, and obtain the regulation score matrix.
[0033] Specifically, for each ternary combination, lncRNA expression, miRNA expression, and mRNA expression values were extracted. The negative correlation coefficient between lncRNA and mRNA expression values was calculated as the baseline value for the joint regulatory strength. Using genomic coordinate matching, the peak ChIP-seq signal intensity overlapping with the transcription start site of lncRNA expression values was located in the EZH2 binding site data. Simultaneously, the modification level reading overlapping with the mRNA promoter region (within an exemplary 2000 bp upstream of the transcription start site) was located in the H3K27me3 epigenetic modification data. The epigenetic regulatory strength was obtained by normalizing the EZH2 binding site data signal intensity and the H3K27me3 modification level through a product operation. The combined regulatory strength was linearly combined with the baseline value and the epigenetic regulatory strength at an exemplary weight ratio of 0.6:0.4. Mechanism cue scores were calculated using the normalized geometric mean of the EZH2 binding site signal intensity and the H3K27me3 epigenetic modification data. The combined regulatory strength and mechanism cue score for each ternary combination were then entered into the corresponding positions in the matrix, with rows corresponding to the lncRNA expression data-miRNA expression data-mRNA expression data combination numbers, and columns representing the combined regulatory strength value and the mechanism cue score value, respectively, to generate a regulatory score matrix. It should be noted that the expression for calculating the negative correlation coefficient between lncRNA expression value and mRNA expression value is as follows: ; in, The negative correlation coefficient between lncRNA expression value and mRNA expression value is an example value (-1, 1). It is the first The expression values of lncRNA in each sample It is the arithmetic mean of lncRNA expression values across all samples. It is the arithmetic mean of mRNA expression levels across all samples. It is the first mRNA expression levels in each sample It is the index variable of the sample.
[0034] S2.3 Construct a third-order control tensor with ternary combinations as three-dimensional coordinate axes based on the control scoring matrix.
[0035] Specifically, the names of the lncRNAs, the names of the miRNAs and the names of the mRNAs in all the triple combinations in the regulation score matrix are extracted as the indexes of the three coordinate axes, respectively. The joint regulation intensity values are filled in the positions of the tensor with the name of the lncRNA as the first dimension, the name of the miRNA as the second dimension and the name of the mRNA as the third dimension, and the mechanism clue score values are filled in the second channel of the same position. The positions of the triple combinations that are not matched to the regulation score matrix are filled with an exemplary default value 0. The dimension structure of the third-order regulation tensor is: the first dimension is the total number of the names of the lncRNAs, the second dimension is the total number of the names of the miRNAs, the third dimension is the total number of the names of the mRNAs, and the fourth dimension is fixed as 2 characteristic values (the joint regulation intensity and the mechanism clue score). The zero-value placeholder area is compressed using a sparse storage format, and the third-order regulation tensor is generated.
[0036] S3, mechanism identification and label annotation are performed on the third-order regulation tensor to obtain a mechanism-labeled regulation tensor.
[0037] S3.1, based on the EZH2 binding site data and the H3K27me3 epigenetic modification data in the preprocessed liver cancer regulation multi-omics data, the epigenetic mechanism confidence score of the triple combination is calculated.
[0038] Specifically, the ChIP-seq peak signal intensity in the EZH2 binding site data that overlaps with the transcription start site of the lncRNA expression data in the triple combination is extracted, and an exemplary maximum peak in a ±500bp window is selected. In the H3K27me3 epigenetic modification data, the modification level reads that overlap with the promoter region of the mRNA expression data in the triple combination are extracted, and an exemplary average reads coverage in the range of 2000bp upstream of the transcription start site is counted. The signal intensity of the EZH2 binding site data and the H3K27me3 epigenetic modification data are weighted and summed according to an exemplary ratio of 0.4:0.6, and the weighted sum result is normalized to the range of 0-1 as the epigenetic mechanism confidence score.
[0039] S3.2, the Pearson correlation coefficient of the lncRNA expression data and the miRNA expression data is calculated, and the ceRNA mechanism confidence score is calculated.
[0040] Specifically, for the lncRNA expression data and the miRNA expression data in each triad combination, the TPM normalized lncRNA expression data and miRNA expression data sequences are extracted in sample order. The ratio of the covariance of the lncRNA expression data and the miRNA expression data to the standard deviation of each is calculated to obtain the Pearson correlation coefficient. The sequence complementarity of the lncRNA expression data and the miRNA expression data is checked, and the seed region (2-8th position) is exemplarily required to be completely matched. The Pearson correlation coefficient and the complementarity matching result are weighted and summed in a ratio of 0.7:0.3 according to the exemplified weight, and the Pearson correlation coefficient is taken as an absolute value. The weighted sum result is mapped to the range of 0-1 through a Sigmoid function, serving as a ceRNA mechanism confidence score; It should be noted that the expression for calculating the ratio of the covariance of the lncRNA expression data and the miRNA expression data to the standard deviation of each is: ; wherein, is the ratio of the covariance of the lncRNA expression data and the miRNA expression data to the standard deviation of each, is the covariance of the lncRNA expression data and the miRNA expression data , and is the standard deviation of the lncRNA expression data, is the standard deviation of the miRNA expression data.
[0041] S3.3, according to the relative size relationship between the epigenetic mechanism confidence score and the ceRNA mechanism confidence score, a mechanism label is labeled and integrated into the three-order regulation tensor to generate a mechanism-labeled regulation tensor.
[0042] Specifically, the epigenetic mechanism confidence score and the ceRNA mechanism confidence score of each triad combination are compared. When the epigenetic mechanism confidence score is greater than the ceRNA mechanism confidence score, the mechanism label is labeled as "epigenetic mechanism"; otherwise, it is labeled as "ceRNA mechanism". The labeling result is written in the fourth dimension of the triad combination position of the three-order regulation tensor in string format, and "Epigenetic" and "ceRNA" are exemplarily used as label values. The triad combinations that do not reach the exemplified confidence of 0.5 are labeled with the "Unknown" label. The three-order regulation tensor written in string format includes the original three-dimensional structure and the newly added mechanism label dimension, generating a mechanism-labeled regulation tensor.
[0043] S4, according to the mechanism-labeled regulation tensor, a mechanism hierarchical Markov state transition model is constructed.
[0044] S4.1, separate epigenetic mechanism-annotated tri-combinations and ceRNA mechanism-annotated tri-combinations from the mechanism-annotated regulatory tensor to generate an epigenetic subset and a ceRNA subset.
[0045] Specifically, traverse the mechanism label dimension of each tri-combination in the mechanism-annotated regulatory tensor, filter tri-combinations with a label value of "Epigenetic" and corresponding joint regulation intensity and mechanism clue score data, and store them as an epigenetic subset; filter tri-combinations with a label value of "ceRNA" and related data, and store them as a ceRNA subset. Tri-combinations with a label value of "Unknown" are discarded. The epigenetic subset and the ceRNA subset maintain the original three-order tensor data structure, and exemplary retain the lncRNA-miRNA-mRNA three-dimensional coordinate mapping relationship and the joint regulation intensity and mechanism clue score values.
[0046] S4.2, divide the expression level state according to the joint regulation intensity of the tri-combinations in the epigenetic subset to obtain an epigenetic state space.
[0047] Specifically, extract the joint regulation intensity of all tri-combinations in the epigenetic subset, and divide the numerical values into exemplary 3 levels of high, medium and low expression states using the K-means clustering algorithm. The cluster centers are initialized using the percentile method, and the 30th, 60th and 90th percentiles are exemplary. According to the clustering results, each tri-combination is mapped to the corresponding expression level state, and the expression level state is stored as a discrete value, exemplary assigned as 1 = low expression, 2 = medium expression, and 3 = high expression. A two-dimensional matrix containing all tri-combinations and expression state labels is obtained, with rows corresponding to tri-combination indices and columns storing state label values, generating an epigenetic state space.
[0048] S4.3, divide the expression level state according to the joint regulation intensity of the tri-combinations in the ceRNA subset to obtain a ceRNA state space.
[0049] Specifically, extract the joint regulation intensity of all tri-combinations in the ceRNA subset, and divide the expression level state using the tertile method, exemplary setting the first tertile = 0.3 quantile and the second tertile = 0.6 quantile. Tri-combinations with a joint regulation intensity less than the first tertile are classified as low expression, tri-combinations with a joint regulation intensity between the first and second tertiles are classified as medium expression, and tri-combinations with a joint regulation intensity greater than the second tertile are classified as high expression. The joint regulation intensity of each tri-combination is mapped to a discrete state label, exemplary coded as 1 = low expression, 2 = medium expression, and 3 = high expression. A two-dimensional matrix storing the mapping relationship between tri-combination indices and state labels is obtained, with rows corresponding to tri-combination numbers and columns storing state code values, generating a ceRNA state space.
[0050] S4.4, based on the pre-processed liver cancer regulatory multi-omics data, the state transition frequency statistics of the epigenetic state space and the ceRNA state space are performed respectively, and the epigenetic transition matrix and the ceRNA transition matrix are obtained.
[0051] Specifically, in the time sequence samples of liver cancer regulatory multi-omics data, the expression state change of each three-element combination in the epigenetic state space at adjacent time points is tracked. The state transition event frequency is recorded. For example, the number of times of transition from state 1 (low expression) to state 2 (medium expression) is counted, and a 3x3 epigenetic state space transition frequency matrix is obtained. The transition frequency matrix row represents the starting state, the list represents the target state, and the cell value is the cumulative frequency of the corresponding transition. The same operation is performed on the ceRNA state space, and the state transition frequency is independently counted to obtain the ceRNA state space transition frequency matrix. The epigenetic state space transition frequency matrix and the ceRNA state space transition frequency matrix are subjected to row standardization operation respectively, and the zero value is processed by Laplace smoothing, and finally the epigenetic transition matrix and the ceRNA transition matrix are generated; It should be noted that the process of processing zero value by Laplace smoothing is: in the epigenetic state space transition frequency matrix and the ceRNA state space transition frequency matrix, the original frequency value of each cell is increased by an example smoothing parameter of 0.001. The adjusted frequency value is converted to a probability value by row normalization operation, and the sum of all adjusted frequency values in each row is controlled to be 1, eliminating the influence of zero probability.
[0052] S4.5, based on the Markov chain state transition probability system, the epigenetic transition matrix and the ceRNA transition matrix are subjected to weighted linear combination, and the mechanism hierarchical Markov state transition model is obtained.
[0053] Specifically, based on the Markov chain state transition probability system, the epigenetic mechanism weight is set to 0.6 and the ceRNA mechanism weight is set to 0.4. Under the framework of Markov state transition probability, the corresponding state transition probability values in the epigenetic transition matrix and the ceRNA transition matrix are subjected to weight influence respectively, the epigenetic transition matrix probability value is subjected to scalar multiplication operation with 0.6, and the ceRNA transition matrix probability value is subjected to scalar multiplication operation with 0.4. The values of the same starting state and target state position in the weighted epigenetic transition matrix and the ceRNA transition matrix are added to generate a fusion transition probability matrix conforming to the Markov chain probability characteristics. The fusion transition probability matrix is verified to satisfy the Markov chain state transition probability constraint condition, and the probability sum of each row is verified to be in the interval of 0.999999 to 1.000001. The verified fusion transition probability matrix is integrated with the state space encoding information to form a Markov state transition model with hierarchical characteristics of epigenetic mechanism and ceRNA mechanism.
[0054] S5, the mechanisms are fused according to biological credibility to generate a fused Markov state diagram.
[0055] S5.1, based on the mechanism hierarchical Markov state transition model, weight calculation is performed through regression analysis and statistical test to obtain apparent transition matrix fusion weight and ceRNA transition matrix fusion weight.
[0056] Specifically, in the mechanism hierarchical Markov state transition model, the historical state transition records of the epigenetic transition matrix and the ceRNA transition matrix are extracted, the least square method is used to fit the epigenetic mechanism weight regression curve, and the fitting result with a determination coefficient greater than 0.8 is exemplarily selected. For the ceRNA transition matrix, the miRNA-lncRNA interaction data verified by CLIP-seq is used, and the statistical significance of the ceRNA mechanism weight is evaluated by t-test, and the p value is exemplarily required to be less than 0.05. The epigenetic mechanism weight obtained by regression analysis and the ceRNA mechanism weight determined by statistical test are normalized so that the sum of the epigenetic mechanism weight and the ceRNA mechanism weight is 1, and finally the apparent transition matrix fusion weight and the ceRNA transition matrix fusion weight are output.
[0057] S5.2, the apparent transition matrix fusion weight and the ceRNA transition matrix fusion weight are respectively linearly transformed with the corresponding apparent transition matrix and ceRNA transition matrix through matrix scalar product operation to obtain a weighted apparent transition matrix and a weighted ceRNA transition matrix.
[0058] Specifically, the apparent transition matrix fusion weight and the ceRNA transition matrix fusion weight are extracted, and the apparent transition matrix fusion weight is exemplarily 0.6 and the ceRNA transition matrix fusion weight is exemplarily 0.4. The scalar product operation of each state transition probability value in the apparent transition matrix with the apparent transition matrix fusion weight exemplarily 0.6 is performed to generate a weighted apparent transition matrix. The scalar product operation of each state transition probability value in the ceRNA transition matrix with the ceRNA transition matrix fusion weight exemplarily 0.4 is performed to generate a weighted ceRNA transition matrix. It should be noted that the expression of the weighted apparent transition matrix is: ; wherein, is the weighted apparent transition matrix, is the apparent transition matrix fusion weight exemplarily 0.6, is the apparent transition matrix, is the product operation symbol; It should be noted that the expression of the weighted ceRNA transition matrix is: ; wherein, is a weighted ceRNA transition matrix, is a ceRNA transition matrix fusion weight exemplary 0.4, is a ceRNA transition matrix.
[0059] S5.3, the weighted apparent transition matrix and the weighted ceRNA transition matrix are fused in the probability space by matrix element level addition operation, and a fused Markov state diagram is generated.
[0060] Specifically, the numerical values of each same coordinate position in the weighted apparent transition matrix and the weighted ceRNA transition matrix are subjected to a probability merging operation, and the weighted probability values of the corresponding positions of the weighted apparent transition matrix and the weighted ceRNA transition matrix are merged to generate a fused probability value, forming a new 3x3 fused probability matrix. Verify whether the sum of all transition probability values of each row of the fused probability matrix is within the exemplary allowable range of 0.999 to 1.001, and eliminate abnormal rows that do not meet the conditions. Match the fused probability matrix that passes the verification with the expression level state (low=1, medium=2, high=3) to obtain a directed graph structure: the nodes represent the low / medium / high expression states, the directed edges store the fused probability values as the regulation weights, the node attributes label the corresponding expression level states, and finally a complete fused Markov state diagram is generated.
[0061] S6, parse the high-frequency transition path in the fused Markov state transition diagram, and convert it into a graph structure form of the regulation relationship to obtain a regulation edge set.
[0062] S6.1, extract the high-frequency state transition path from the fused Markov state transition diagram.
[0063] Specifically, all state transition paths in the fused Markov state diagram are traversed, and paths with a regulation weight greater than an exemplary 0.7 are selected as candidate paths. The candidate paths are verified: the epigenetic mechanism path must satisfy the condition that the EZH2 binding site data signal strength is greater than an exemplary 5 and the H3K27me3 modification level is greater than an exemplary 3; the ceRNA mechanism path must satisfy the condition that the absolute value of the Pearson correlation coefficient is greater than an exemplary 0.5. Eliminate paths that do not meet any of the above conditions. The paths that pass the verification are arranged in descending order of regulation weight, and the top exemplary 10% of the paths are retained as high-frequency state transition paths, and the remaining paths are marked as invalid paths and not output.
[0064] S6.2, molecular expression trend analysis is performed on the high-frequency state transition path, the regulation relationship chain is identified, and binary relationship conversion and edge weight assignment are performed to obtain a regulation edge set.
[0065] Specifically, the expression level change patterns of the lncRNA expression data, the miRNA expression data and the mRNA expression data in the high-frequency state transition path are analyzed. When a pattern of lncRNA expression data expression decrease accompanied by miRNA expression data expression increase and mRNA expression data expression decrease is detected, it is determined that there is an lncRNA expression data→miRNA expression data→mRNA expression data regulatory relationship chain. The lncRNA expression data→miRNA expression data→mRNA expression data regulatory relationship chain is split into two binary regulatory relationships of lncRNA expression data→miRNA expression data and miRNA expression data→mRNA expression data. The regulatory weight of the high-frequency state transition path is taken as the edge weight basic value, an exemplary correction coefficient 0.8 is added to the lncRNA expression data→miRNA expression data edge weight, and an exemplary correction coefficient 1.2 is added to the miRNA expression data→mRNA expression data edge weight to generate the final regulatory edge weight. All verified binary regulatory relationships and regulatory edge weights are stored in a structured list to form a regulatory edge set.
[0066] S7, constructing a structured graph for the regulatory edge set to obtain a miRNA regulatory network.
[0067] S7.1, attribute labeling the edge weight of the regulatory edge set to obtain a weighted edge attribute set.
[0068] Specifically, the weight value of each binary regulatory relationship in the regulatory edge set is extracted, and the final regulatory weight value is calculated based on the edge weight correction coefficient (lncRNA expression data→miRNA expression data edge 0.8, miRNA expression data→mRNA expression data edge 1.2). The regulatory weight value is standardized and mapped to the exemplary interval 0-1 range. An attribute label is added to each binary regulatory relationship to mark the regulatory type (epigenetic / ceRNA mechanism) and the regulatory weight value of each binary regulatory relationship, and stored in a structured key-value pair format (such as {"regulatory type": "ceRNA", "weight": 0.75}). All annotated edge attribute information is sorted according to the three-element combination index to generate a weighted edge attribute set.
[0069] S7.2, classifying and storing the lncRNA expression data nodes, the miRNA expression data nodes and the mRNA expression data nodes in the regulatory edge set to obtain a structured node set.
[0070] Specifically, all lncRNA expression data nodes, miRNA expression data nodes and mRNA expression data nodes in the regulatory edge set are traversed, the lncRNA expression data nodes are classified as long-chain non-coding RNA type, the miRNA expression data nodes are classified as microRNA type, and the mRNA expression data nodes are classified as messenger RNA type. A unique identifier is assigned to each node, and the nodes are numbered in the format of "LNC_number", "MIR_number" and "MRNA_number". The node basic attributes, including the types, expression mean values and regulatory correlation degrees of the lncRNA expression data nodes, miRNA expression data nodes and mRNA expression data nodes, are stored. The classified lncRNA expression data nodes, miRNA expression data nodes and mRNA expression data nodes are grouped by type and stored in a structured database table. The table structure includes the lncRNA expression data node, miRNA expression data node and mRNA expression data node ID, type, expression mean value and correlation degree fields, and a structured node set is generated.
[0071] S7.3, Topological relationship generation is performed on the directed edge relationship of the structured node set to obtain the network topology structure of the directed connection relationship.
[0072] Specifically, all binary regulatory relationships in the regulatory edge set are traversed, and edges containing lncRNA expression data node pointing to miRNA expression data node and edges containing miRNA expression data node pointing to mRNA expression data node are extracted. According to the matching relationship of the lncRNA expression data nodes, miRNA expression data nodes and mRNA expression data nodes in the structured node set, the directed connection of the lncRNA expression data node to the miRNA expression data node and the directed connection of the miRNA expression data node to the mRNA expression data node are established. Each directed connection is assigned a corresponding regulatory weight value in the regulatory edge set, and the exemplary regulatory weight range is 0.1 to 0.9. The integrity of the connection relationship is checked to ensure that the source node and the target node of all directed connections exist in the structured node set, and abnormal connections that do not meet the conditions are removed. The verified directed connections and regulatory weight values are stored as an adjacency list structure to obtain the network topology structure of the lncRNA expression data node→miRNA expression data node and miRNA expression data node→mRNA expression data node directed connection relationship.
[0073] S7.4, The weighted edge attribute set, the structured node set and the network topology structure are integrated to obtain the miRNA regulatory network.
[0074] Specifically, the regulation type in the weighted edge attribute set is matched with the regulation weight information, the node attribute information in the structured node set, and the directed connection relationship in the network topology structure. The mapping relationship between the weighted edge attribute set and the network topology structure is established based on the lncRNA expression data node, the miRNA expression data node, and the mRNA expression data node ID, so as to ensure that the regulation weight value and the regulation type of each directed connection are correctly corresponded. The node attribute information in the structured node set is integrated into the network topology structure, and the expression average and the regulation correlation degree attribute are added to the ncRNA expression data node, the miRNA expression data node, and the mRNA expression data node. The data consistency is verified, and it is required that the source node and the target node of the directed connection relationship of the lncRNA expression data node→the miRNA expression data node and the miRNA expression data node→the mRNA expression data node must exist in the node set, and all node attributes and regulation edge weights are completely matched. Finally, the miRNA regulation network containing the node attribute, the regulation edge weight, and the topology connection relationship is generated.
[0075] The embodiment also provides a computer device suitable for the case of the method for constructing the miRNA regulation network, including a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the method for constructing the miRNA regulation network proposed in the above embodiment.
[0076] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to perform wired or wireless communication with an external terminal. The wireless communication can be achieved through WIFI, an operator network, NFC (near field communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, a trackball, or a touchpad arranged on the shell of the computer device. In addition, the input device can be an external keyboard, a touchpad, or a mouse, etc.
[0077] The embodiment also provides a storage medium on which a computer program is stored, the program being executed by a processor to implement the method for constructing a miRNA regulatory network proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.
[0078] To sum up, the application realizes independent quantitative characterization of epigenetic regulation and ceRNA regulation mechanism by constructing a mechanism hierarchical Markov state transition model, so that the dynamic characteristics of different regulation modes are completely retained; a unified dynamic network with clear biological interpretability is established by generating a fusion Markov state diagram and organically integrating the double mechanisms based on a weight fusion algorithm of biological evidence. The mechanism-specific modeling significantly improves the identification accuracy of key regulation paths, and provides more reliable biomarkers for liver cancer targeted therapy; the evidence-driven fusion method ensures that the dynamic network accurately captures key state transitions in disease progression, and provides a generalizable technical paradigm for cross-mechanism regulation network research, and exhibits unique value in pathogenesis research and precision medical applications.
[0079] It should be noted that the above embodiments are only used to illustrate the technical solutions of the application rather than limit the application. Although the application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the application, and all of them should be covered in the scope of the claims of the application.
Claims
1. A method for constructing a miRNA regulatory network, characterized in that: include, We collected multi-omics data on liver cancer regulation, preprocessed them, and constructed a third-order regulatory tensor. Mechanism identification and labeling are performed on the third-order regulation tensor to obtain the mechanism-labeled regulation tensor; Based on the mechanism-labeled regulatory tensor, a mechanism-hierarchical Markov state transition model was constructed, and the model was weighted and fused according to biological credibility to generate a fused Markov state diagram. The high-frequency transition paths in the fused Markov state transition graph are analyzed and converted into the control relationships in the form of a graph structure, resulting in a set of control edges; A structured graph was constructed from the set of regulatory edges to obtain the miRNA regulatory network.
2. The method for constructing a miRNA regulatory network as described in claim 1, characterized in that: The multi-omics data on liver cancer regulation includes lncRNA expression data, miRNA expression data, mRNA expression data, EZH2 binding site data, and H3K27me3 epigenetic modification data. The preprocessing includes missing value imputation, normalization, sample consistency alignment, and representation matrix format standardization.
3. The method for constructing a miRNA regulatory network as described in claim 2, characterized in that: The specific steps for constructing the third-order control tensor are as follows: The lncRNA expression data, miRNA expression data and mRNA expression data of each group in the preprocessed multi-omics data on liver cancer regulation were combined to generate a ternary combination; By combining the trigonometric combination, EZH2 binding site data, and H3K27me3 epigenetic modification data, the joint regulatory intensity and mechanism cue score of each trigonometric combination were calculated to obtain the regulatory score matrix. A third-order control tensor with ternary combinations as three-dimensional coordinate axes is constructed based on the control score matrix.
4. The method for constructing a miRNA regulatory network as described in claim 3, characterized in that: The mechanism identification and labeling of the third-order modulation tensor to obtain the mechanism-labeled modulation tensor are performed as follows: Based on EZH2 binding site data and H3K27me3 epigenetic modification data from preprocessed hepatocellular carcinoma regulatory multi-omics data, the confidence score of epigenetic mechanism of the ternary combination was calculated. Calculate the Pearson correlation coefficient between lncRNA expression data and miRNA expression data, and perform confidence score calculation for ceRNA mechanism; Based on the relative magnitudes of the confidence scores for epigenetic mechanisms and ceRNA mechanisms, the mechanism tags are used as ternary combinations and integrated into the third-order regulatory tensor to generate a mechanism-labeled regulatory tensor.
5. The method for constructing a miRNA regulatory network as described in claim 4, characterized in that: The mechanism-based annotation and control tensor are used to construct a hierarchical Markov state transition model. The specific steps are as follows: Separate epigenetic mechanism-labeled triplets and ceRNA mechanism-labeled triplets from the mechanism-labeled regulatory tensor to generate epigenetic subsets and ceRNA subsets; Based on the joint regulatory intensity of the ternary combination in the epigenetic subset, the expression level states are divided to obtain the epigenetic state space; Based on the joint regulatory strength of the ternary combinations in the ceRNA subset, the expression level states are divided to obtain the ceRNA state space; Based on preprocessed multi-omics data on liver cancer regulation, the state transition frequency of the epigenetic state space and ceRNA state space was statistically analyzed to obtain the epigenetic transition matrix and ceRNA transition matrix. Based on the Markov chain state transition probability system, a weighted linear combination of the apparent transition matrix and the ceRNA transition matrix is performed to obtain a mechanism-hierarchical Markov state transition model.
6. The method for constructing a miRNA regulatory network as described in claim 5, characterized in that: The mechanisms are weighted and fused according to their biological credibility to generate a fused Markov state diagram. The specific steps are as follows: Based on the mechanism-hierarchical Markov state transition model, the weights are calculated through regression analysis and statistical tests to obtain the fusion weights of the apparent transition matrix and the fusion weights of the ceRNA transition matrix. The apparent transfer matrix fusion weights and ceRNA transfer matrix fusion weights are linearly transformed with their respective apparent transfer matrices and ceRNA transfer matrices by matrix scalar product operation, resulting in weighted apparent transfer matrices and weighted ceRNA transfer matrices. The weighted apparent transition matrix and the weighted ceRNA transition matrix are probabilistically fused by matrix element-level addition operations to generate a fused Markov state diagram.
7. The method for constructing a miRNA regulatory network as described in claim 6, characterized in that: The analytical fusion of high-frequency transition paths in the Markov state transition graph and its conversion into a graph structure of control relationships yields a set of control edges. The specific steps are as follows: Extracting high-frequency state transition paths from the fused Markov state transition graph; Molecular expression trend analysis was performed on high-frequency state transition pathways to identify regulatory relationship chains. Binary relationship transformation and edge weight assignment were then performed to obtain the set of regulatory edges.
8. The method for constructing a miRNA regulatory network as described in claim 7, characterized in that: The structured graph construction of the regulatory edge set yields the miRNA regulatory network. The specific steps are as follows: The edge weights of the control edge set are labeled with attributes to obtain the weighted edge attribute set; The lncRNA expression data nodes, miRNA expression data nodes, and mRNA expression data nodes in the regulatory edge set are classified and stored to obtain a structured node set; Topological relationships are generated from the directed edge relationships of the structured node set to obtain the network topology structure with directed connections; By integrating the weighted edge attribute set, the structured node set, and the network topology, a miRNA regulatory network is obtained.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method for constructing a miRNA regulatory network as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method for constructing a miRNA regulatory network as described in any one of claims 1 to 8.