Gene abnormality regulation and control detection method based on regulation and control pathway disturbance analysis
By introducing gene dynamic influence factors and pathway context-related factors, the pathway perturbation scoring method was optimized, solving the problem that existing technologies cannot identify abnormal pathway regulatory patterns dominated by high-influence genes. This enabled highly sensitive and accurate identification of pathway perturbations in complex diseases, improving the ability to study disease molecular mechanisms and discover targets.
Patent Information
- Application Number
- CN202511054381.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies are unable to effectively identify aberrant regulatory patterns in key pathways dominated by a few high-influence genes when assessing gene pathway perturbations, which hinders a deeper understanding of the molecular mechanisms of diseases and precise intervention strategies.
By introducing dynamic gene influence factors and pathway contextual factors, the pathway perturbation scoring method is optimized. Combining the regulatory position and expression coordination of genes in the pathway topology, dynamic influence factors and contextual optimization factors are constructed to improve the ability to identify high-influence genes.
It significantly improves the sensitivity and accuracy of gene set variation analysis methods in identifying pathway perturbations in complex diseases, enhances the ability to identify pathway perturbations dominated by a few key genes, and improves the biological interpretability and downstream clinical application value of pathway analysis results.
Smart Images

Figure CN120913644A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and in particular to a method for detecting abnormal gene regulation based on the analysis of regulatory pathway perturbations. Background Technology
[0002] In life science research, especially in oncology, understanding anomalous changes in gene expression regulatory networks is crucial for revealing the molecular mechanisms of disease development. Genes do not function independently but rather work together to form complex biological pathways (signaling pathways, metabolic pathways, etc.) to perform specific cellular functions. Analyzing changes in activity at the entire pathway level, rather than focusing solely on changes in the expression of individual genes, provides a more systematic and biologically meaningful perspective.
[0003] Existing technologies utilize high-throughput transcriptome sequencing ( Gene set variation analysis (GFA) is a commonly used strategy for gene expression profiling data generated by technologies such as gene set variation analysis. ) is one of the representative methods. It is an unsupervised, unbiased computational method that can calculate the enrichment score of each predefined gene set (pathway) for each sample. Specifically, By comparing the empirical cumulative distribution functions of genes within and outside the pathway in the overall gene expression ranking of this sample (…), ), and utilize similar The test statistic quantifies whether genes within a pathway tend to concentrate at high or low expression levels, thus obtaining the relative activity or perturbation level of that pathway in the sample. This is achieved by comparing samples from different groups (tumor samples and normal control samples). By analyzing fractional distributions, researchers can identify pathways that undergo significant changes in activity during disease states. These methods are widely used in exploring sample heterogeneity, associated pathway activity, and clinical phenotypes because they can assess pathway activity in each sample.
[0004] Regulatory abnormalities in biological systems are often complex. In many diseases, especially cancer, key phenotypic changes may be primarily caused by dramatic and functionally defined abnormalities in the expression of a few core driver genes, such as specific oncogenes or tumor suppressor genes.
[0005] Existing technology In the evaluation of pathway perturbation, its core computing logic is based on the overall rank distribution comparison of the expression levels of all genes in the pathway. This scoring mechanism based on the collective behavior of genes in the set can reflect the overall trend of pathway activity, but it inherently averages or considers the same impact on the distribution of the contribution of all genes in the pathway. Specifically, when the abnormal disturbance of a pathway is mainly caused by a few key genes (i.e. high-impact genes) in the pathway, which have extremely significant expression changes and show high consistency between samples, The calculated enrichment score mixes and dilutes the strong signal of these key genes with the weak, insignificant, or even opposite signals of a large number of genes in the pathway. Therefore, the standard The method lacks sensitivity to accurately identify and prioritize such strongly driven pathway abnormal regulation patterns by a few high-impact genes. The method cannot effectively distinguish whether the pathway disturbance is caused by the universal moderate changes of a large number of genes or the strong changes of a few key genes, which directly leads researchers to ignore the biological pathways whose functional disorder is actually closely related to the abnormal activity of specific key driving genes (which are often important potential therapeutic targets or core of disease mechanism), and hinders the in-depth understanding of disease molecular mechanisms and the development of precise intervention strategies. Therefore, it is urgent to propose a new pathway disturbance evaluation method to overcome the above-mentioned defects of the prior art and improve the identification ability of key pathway abnormal regulation events dominated by high-impact genes. SUMMARY
[0006] Therefore, the present application aims to provide a gene abnormal regulation detection method based on regulatory pathway disturbance analysis to improve the identification ability of key pathway abnormal regulation events dominated by high-impact genes.
[0007] To achieve the above-mentioned purpose, the technical solution of the present application is as follows:
[0008] The gene abnormal regulation detection method based on regulatory pathway disturbance analysis comprises the following steps:
[0009] Step S1: Collecting biological sample and pathway structure data;
[0010] Step S2: Obtaining a weighted optimized pathway disturbance score through gene dynamic behavior and pathway context information;
[0011] Step S2.1: Obtaining a gene dynamic impact factor through differential expression significance and sample consistency;
[0012] Step S2.2: Obtaining a pathway-specific context correlation optimization factor through pathway topological structure and expression coordination;
[0013] Step S2.3: Optimize the GSVA pathway perturbation score by integrating contribution weights;
[0014] Step S3: Identify abnormal regulatory pathways based on optimized scoring.
[0015] Furthermore, according to step S2.1, obtaining the gene dynamic influence factor through differential expression significance and inter-sample consistency specifically includes:
[0016] A standardized gene expression profile matrix is obtained. Component-based differential expression analysis is performed on the standardized gene expression profile matrix to obtain the t-statistic and two-tailed p-values of the genes. The absolute value of the t-statistics is used to assess the gene variation amplitude characteristics. A negative logarithmic transformation is applied to the two-tailed p-values to obtain the gene significance characteristics. The coefficient of variation of gene expression values is used to assess gene consistency characteristics, and expression consistency analysis is performed on these consistency characteristics to obtain the gene expression consistency measure factor. Dimensional weights are evaluated for the gene variation amplitude characteristics, significance characteristics, and expression consistency measure factor to obtain the feature weights of each of the three feature dimensions. Finally, the influence of the gene is comprehensively evaluated using the feature weights of each of the three feature dimensions to obtain the dynamic influence factor of the gene.
[0017] Furthermore, based on the aforementioned method of evaluating gene expression values using the coefficient of variation to obtain gene consistency characteristics, and by performing expression consistency analysis on these characteristics, gene expression consistency metrics are obtained, including:
[0018] The formula for calculating the gene expression consistency measure factor is:
[0019] ;
[0020] in, Indicates gene Expression consistency measure factor; Indicates gene In the sample group The coefficient of variation in expression; It represents a very small positive number.
[0021] Furthermore, based on the aforementioned dimensional weighting assessment of gene variation magnitude characteristics, significance characteristics, and expression consistency measurement factors, the feature weights of each of the three feature dimensions are obtained. Then, the gene's influence is comprehensively assessed using the feature weights of each of the three feature dimensions to obtain the dynamic influence factor of the gene, specifically including:
[0022] The formula for calculating the weights of the saliency feature dimension is:
[0023] ;
[0024] The calculation formula of the weight of the amplitude feature dimension is:
[0025] ;
[0026] The calculation formula of the weight of the consistency feature dimension is:
[0027] ;
[0028] wherein, represents the weight of the significance feature dimension, i.e. the dimension weight of this statistical feature; represents the weight of the amplitude feature dimension, i.e. the dimension weight of this statistical feature; represents the weight of the consistency feature dimension, i.e. the dimension weight of the consistency measure factor; represents the variance of the normalized significance feature in the whole gene set; represents the variance of the normalized amplitude feature in the whole gene set; represents the variance of the normalized consistency measure factor in the whole gene set; represents a very small positive number;
[0029] The calculation formula of the single dynamic influence factor is:
[0030] ;
[0031] wherein, represents the dynamic influence factor of the gene ; represents an index symbol, representing a specific gene in the genome; respectively represent the data weight of the significance feature, the amplitude feature and the consistency feature; represents the significance feature of the gene after standardization processing; represents the amplitude feature of the gene after standardization processing; represents the consistency feature of the gene after standardization processing.
[0032] Further, the pathway-specific context association optimization factor is obtained by the pathway topology and expression coordination according to the step S2.2, specifically comprising:
[0033] The betweenness centrality evaluation and the degree centrality evaluation of the gene in the pathway are obtained; the betweenness centrality evaluation and the degree centrality evaluation of the gene in the pathway are standardized to obtain the standardized betweenness centrality evaluation and the standardized degree centrality evaluation of the gene in the pathway; the weight of the betweenness centrality and the weight of the degree centrality in the pathway are obtained by evaluating the weight of the standardized betweenness centrality variance and the standardized degree centrality variance of the gene in the pathway; the comprehensive topological importance evaluation of the gene in the pathway is obtained by fusing the standardized betweenness centrality evaluation and the standardized degree centrality evaluation of the gene according to the weight of the betweenness centrality and the weight of the degree centrality in the pathway; the expression profile vector of the gene and the expression profile vector of all neighbor genes in the pathway where the gene is located are obtained; the local expression coordination evaluation of the gene is obtained by evaluating the correlation between the expression profile vector of the gene and the expression profile vector of all neighbor genes in the pathway where the gene is located; the context correlation optimization factor of the gene in the pathway is obtained by comprehensively analyzing the comprehensive topological importance evaluation of the gene in the pathway and the local expression coordination evaluation of the gene.
[0034] Further, according to the weight of the betweenness centrality and the weight of the degree centrality in the pathway are obtained by evaluating the weight of the standardized betweenness centrality variance and the standardized degree centrality variance of the gene in the pathway, comprising:
[0035] The weight of the standardized betweenness centrality of the gene is calculated according to the following formula:
[0036] ;
[0037] Wherein, denotes the weight of the standardized betweenness centrality in the pathway; denotes the variance of the standardized betweenness centrality values of all genes in the pathway; denotes the variance of the standardized degree centrality values of all genes in the pathway; denotes a very small positive number; The weight of the standardized degree centrality of the gene is calculated according to the following formula:
[0038]
[0039] ;
[0040] Wherein, denotes the weight of the standardized degree centrality in the pathway; denotes the variance of the standardized betweenness centrality values of all genes in the pathway; denotes the variance of the standardized degree centrality values of all genes in the pathway; denotes a very small positive number.
[0041] Further, according to the fusion analysis of the standardized betweenness centrality evaluation and the standardized degree centrality evaluation of the gene in the pathway by the betweenness centrality weight and the degree centrality weight in the pathway, a comprehensive topological importance evaluation of the gene in the pathway is obtained, and the comprehensive topological importance evaluation specifically comprises:
[0042] The comprehensive topological importance evaluation formula is:
[0043] ;
[0044] wherein, denotes a comprehensive topological importance score of the gene in the pathway ; denotes a weight of the standardized betweenness centrality in the pathway ; denotes a standardized betweenness centrality value of the gene in the pathway ; denotes a weight of the standardized degree centrality in the pathway ; denotes a standardized degree centrality value of the gene in the pathway .
[0045] Further, according to the correlation evaluation between the expression profile vector of the gene and the expression profile vectors of all neighbor genes of the gene in the pathway, a local expression coordination evaluation of the gene is obtained, and the local expression coordination evaluation comprises:
[0046] The calculation formula of the local expression coordination evaluation of the gene between the expression profile vector of the gene and the average expression profile vector of all neighbor genes of the gene in the pathway is:
[0047] ;
[0048] wherein, denotes the local expression coordination evaluation of the gene, that is, the correlation coefficient between the expression profile vector of the gene in the pathway and the average expression profile vector of all neighbor genes of the gene in the pathway ; denotes the expression profile vector of the gene in the sample , denotes the number of neighbor genes of the gene in the pathway ; denotes the neighbor gene set of the gene in the pathway . represents the Pearson correlation coefficient calculation function.
[0049] Further, according to the comprehensive analysis of the comprehensive topological importance evaluation of the gene in the pathway and the local expression coordination evaluation of the gene, a context correlation optimization factor of the gene in the pathway is obtained, including:
[0050] The calculation formula of the context correlation optimization factor of the gene in the pathway is:
[0051] ;
[0052] wherein, represents the context correlation optimization factor of the gene in the pathway ; represents the comprehensive topological importance score of the gene in the pathway ; represents the local expression coordination evaluation of the gene, that is, the correlation coefficient between the expression profile vector of the gene in the pathway and the average expression profile vector of all neighbor genes of the gene in the pathway .
[0053] Further, according to the step S2.3, the GSVA pathway perturbation score is optimized by the comprehensive contribution weight, specifically including:
[0054] The standardized gene expression profile matrix, the set biological pathway set and the member gene set in each pathway are obtained; the comprehensive contribution weight is evaluated by the dynamic influence factor of each gene and the context correlation factor of the pathway in which the gene is located, and the calculation formula of the comprehensive contribution weight is:
[0055] ;
[0056] wherein, represents the comprehensive contribution weight of the gene in the pathway , represents the dynamic influence factor of the gene ; represents the context correlation optimization factor of the gene in the pathway ;
[0057] When the pathway perturbation score calculation is performed, the GSVA pathway perturbation score is optimized by the following steps:
[0058] Step 1: all genes in the sample are sorted in descending order of expression amount;
[0059] Step 2: calculate the empirical cumulative distribution function of the genes in the pathway in the sorted list and the empirical cumulative distribution function of the genes outside the pathway;
[0060] Step 3: calculate the original difference value for each sorting position;
[0061] Step 4: if the gene corresponding to the sorting position is in the pathway, the final difference value is the result of weighting adjustment of the original difference value by the comprehensive contribution weight; if the gene corresponding to the sorting position is not in the pathway, the original difference value is taken as the final difference value;
[0062] Step 5: in the difference value sequence formed by the final difference values, the maximum positive deviation and the maximum negative deviation are obtained, and the pathway disturbance score is calculated;
[0063] Step 6: repeat the above steps to obtain a pathway disturbance score matrix of all pathways in the sample set, which is used to represent the optimized pathway disturbance degree.
[0064] Compared with the prior art, the present application has the following advantages:
[0065] The gene abnormal regulation detection method based on regulatory pathway disturbance analysis has the advantages that
[0066] The present application significantly improves the sensitivity and accuracy of gene set variation analysis method in complex disease pathway disturbance identification by introducing dynamic influence factor at gene level and pathway context correlation factor. By integrating the differential expression significance, change amplitude of genes in tumor and normal samples, and the consistency characteristics within the tumor population, the dynamic influence factor is constructed, which effectively identifies high-influence genes that are stable and significantly abnormal in the disease background, overcoming the signal dilution problem caused by averaging in traditional methods. Further, combined with the regulatory position of genes in the pathway topology structure and the coordination of adjacent genes in the expression behavior, the pathway-specific context optimization factor is constructed to finely modulate the weight contribution in different pathways, thereby improving the response ability of the pathway disturbance score to the structural core genes and functional consistency signals. Through the two optimization steps, the present application effectively enhances the recognition ability of a small number of key gene dominant pathway disturbance, improves the biological interpretation and downstream clinical application value of the pathway analysis results, and is suitable for molecular mechanism research and target discovery of heterogeneous diseases such as tumors. BRIEF DESCRIPTION OF DRAWINGS
[0067] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments illustrated in the drawings are provided to explain the present application and should not be considered as an improper limitation on the present application. In the drawings:
[0068] Figure 1A method flow chart of a gene abnormal regulation detection method based on regulation pathway disturbance analysis according to an embodiment of the present application. DETAILED DESCRIPTION
[0069] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0070] In the description of the present application, it should be noted that the terms "upper", "lower", "inner", "back" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0071] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0072] Reference Figure 1 is a method flow chart of a gene abnormal regulation detection method based on regulation pathway disturbance analysis provided by the first embodiment of the present application, as Figure 1 shown, a gene abnormal regulation detection method based on regulation pathway disturbance analysis can include:
[0073] S1, collecting biological sample and pathway structure data.
[0074] In the present embodiment, first, the original sequencing reads of the tissue samples of non-small cell lung cancer patients and the paired or adjacent normal tissue samples adjacent to the cancer are obtained through public data sets; at the same time, the pre-defined biological pathway annotation information is obtained, which contains a series of known biological pathways and lists the member gene set contained in each pathway. Topological structure information capable of describing known interactions or regulation relationships between genes inside the pathway also needs to be obtained or constructed; in the present embodiment, the pathway definition and the corresponding standard topological structure information are obtained from the public database. Next, the necessary bioinformatics processing procedures are performed on the obtained original sequencing data, including quality control, read alignment, gene expression quantification and standardization, and finally a standardized gene expression profile matrix is obtained, in which the rows represent genes and the columns represent samples (including tumor samples and normal samples).
[0075] S2, obtaining a weighted and optimized pathway disturbance score through gene dynamic behavior and pathway context information.
[0076] In assessing pathway perturbation by existing techniques such as Gene Set Variation Analysis, its core mechanism relies on comparing the relative position distribution of genes within a pathway and outside of it in the overall ordering of gene expression profiles of individual samples. This rank-based and empirical cumulative distribution function comparison approach aims to capture the overall up- or down-trend of pathway activity. A key technical limitation is that this approach implicitly averages the contribution of all genes within a pathway, and the final enrichment score mainly reflects the degree of distribution shift of the gene set as a whole, without effectively distinguishing whether this shift is driven by a large number of genes with small but consistent changes, or a few genes with dramatic and statistically significant changes. In the biological mechanisms of complex diseases such as non-small cell lung cancer, the latter case, i.e. the dramatic dysregulation of a few high-impact genes, is more indicative of key pathogenic events or potential therapeutic targets.
[0077] This "signal dilution" effect is an inherent property of the set distribution comparison approach. Even if a gene has a huge change in expression (a key cancer gene is activated by tens of times), and this change is statistically highly reliable, but there are enough other genes within the pathway that have small or inconsistent changes, the strong signal of this key gene will be significantly weakened when calculating the overall distribution shift. In addition, by mainly focusing on the relative position of genes in the ordered list, the magnitude of the change in the marginal contribution of the final enrichment score for those genes already at the extreme of the ordered list (whether due to high expression or low expression) is relatively limited. At the same time, this approach does not consider the stability and consistency of gene expression changes across different disease samples. A gene that shows similar and significant expression abnormalities in all or most tumor samples is usually more representative of stable and universally significant regulatory abnormalities in the disease state than a gene that only changes dramatically in a few samples but behaves normally in others.
[0078] S2.1, obtain gene dynamic impact factor by differential expression significance and inter-sample consistency.
[0079] To overcome the above limitations and improve the recognition ability of pathway perturbation dominated by high-impact genes, first focus on quantifying the dynamic impact of each gene itself. This dynamic impact is not based on its static role within a certain pathway, but is defined according to its actual expression behavior characteristics in the specific biological scenario being studied. The dynamic impact of a gene should reflect the magnitude, statistical significance and consistency of its expression changes in the disease sample population.
[0080] Specifically, first based on the standardized gene expression matrix, two sets of differential expression analysis is adopted to obtain the statistical indicators of each gene, including statistical quantity and corresponding two-tailed value, respectively used to describe the amplitude of expression difference and statistical significance, for quantifying statistical significance, the negative logarithmic transformation of value is obtained , the larger the value indicates the more significant the difference; for quantifying the amplitude of change, the absolute value of statistical quantity is used , which considers both effect size and data variability, and is more robust than simple fold change.
[0081] In order to further reflect the consistency of gene expression behavior in the disease population, the expression coefficient of variation in the tumor sample group is calculated, and the calculation formula of the expression coefficient of variation in the tumor sample group is:
[0082]
[0083] Among them, represents the expression coefficient of variation of gene in the tumor sample group ; represents the standard deviation of the expression value of gene in the tumor sample set ; represents the average expression value of gene in the tumor sample set ; represents a very small positive number to prevent the denominator from being , and in this embodiment .
[0084] After obtaining the expression coefficient of variation of the gene, the expression consistency measurement factor can be constructed accordingly, and the calculation formula of the expression consistency measurement factor is:
[0085]
[0086] Among them, represents the expression consistency measurement factor of gene ; represents the expression coefficient of variation of gene in the tumor sample group ; represents a very small positive number to prevent the denominator from being , and in this embodiment .
[0087] The above three indicators ( ) together constitute the original feature set for evaluating the dynamic influence of genes, which preliminarily characterize the abnormal expression behavior of genes from the aspects of reliability, intensity and universality.
[0088] Secondly, since the three original features extracted above have different dimensions and numerical ranges, direct combination will lead to results biased towards features with larger numerical ranges. To ensure comparability of each feature in the subsequent integration process, the three features need to be standardized. In this embodiment, the standardization method is used to process the three features in the set of all genes successfully obtaining statistical quantities in the differential expression analysis ), to obtain standardized significant features, standardized amplitude features and standardized consistency features.
[0089] Thirdly, the contribution of different feature dimensions (significance, amplitude and consistency) to the final determination of whether a gene is high-influence is not uniform, and the relative importance depends on the characteristics of the specific data set. This application adopts a data-driven weighting strategy to calculate the variance of each standardized feature in the entire gene set. The size of the variance reflects the ability of the feature to distinguish different gene behaviors or the information content in the current data set to some extent. Accordingly, the weight of each of the three different features is evaluated by the proportion of the variance of each feature to the total variance. The weight of the significant feature dimension is calculated by the formula:
[0090]
[0091] The weight of the amplitude feature dimension is calculated by the formula:
[0092]
[0093] The weight of the consistency feature dimension is calculated by the formula:
[0094]
[0095] wherein, represents the weight of the significant feature dimension, i.e. the dimension weight of this statistical feature; represents the weight of the amplitude feature dimension, i.e. the dimension weight of this statistical feature; represents the weight of the consistency feature dimension, i.e. the dimension weight of the consistency measure factor; represents the variance of the standardized significant feature in the entire gene set; represents the variance of the standardized amplitude feature in the entire gene set; denotes the variance of the normalized consistency metric factor in the whole gene set; denotes a very small positive number to prevent the denominator from being In this embodiment, the value of .
[0096] It should be noted that this way of determining feature weights based on the statistical characteristics of the data itself enhances the adaptive ability of the method to different data sets, making the evaluation of the dynamic influence of genes more in line with the actual data distribution.
[0097] Finally, the multiple features after standardization and weighting processing are integrated to obtain a single dynamic influence factor, and the calculation formula of the single dynamic influence factor is:
[0098]
[0099] wherein, denotes the dynamic influence factor of gene ; denotes the index symbol, representing a specific gene in the genome; respectively denote the data weights of the significance feature, the amplitude feature and the consistency feature; denotes the significance feature of gene after standardization processing; denotes the amplitude feature of gene after standardization processing; denotes the consistency feature of gene after standardization processing.
[0100] S2.2, obtaining a pathway-specific context association optimization factor by coordinating the pathway topology structure and the expression.
[0101] Obtaining the dynamic influence factor initially addresses the problem of high-influence gene signals being diluted due to existing technologies failing to fully consider individual gene expression behavior. However, the dynamic influence factor is a gene-centric measure, assessing the gene's own dynamic characteristics without relating it to its role and behavioral patterns within a specific biological pathway. A gene may have a high influence factor (i.e., its own changes are significant and consistent), but if it occupies a peripheral position in the pathway's topology, or if its expression changes contradict or are unrelated to the changing trends of its direct functional partners within the pathway, its actual contribution to driving overall dysfunction of that specific pathway is not as large as the dynamic influence factor suggests. Specifically, amplifying gene signals solely based on the dynamic influence factor can enhance the influence of genes that, while exhibiting dramatic changes, are not core or exhibit uncoordinated behavior within the functional regulatory network of the examined pathway, thus leading to biased judgments about the source of pathway disturbances.
[0102] Therefore, based on the dynamic influence factors of genes, it is necessary to further evaluate specific context-related optimization factors by considering the structural importance and functional coordination of genes within specific pathways. When evaluating context-related optimization factors, the first consideration is the topological importance and local expression coordination of genes within the pathway network structure. For the comprehensive topological importance of genes within a pathway, the betweenness centrality and degree centrality of genes within the pathway are first obtained. Betweenness centrality quantifies the frequency with which a gene acts as a bridge for information transmission between other gene pairs within the pathway, reflecting the gene's criticality in controlling the information flow of the pathway. Degree centrality quantifies the number of directly connected neighbors of a gene, reflecting the breadth of its local connections. Since different pathways have different sizes and connection patterns, directly comparing the original centrality values across pathways has limited significance. Therefore, it is necessary to evaluate these two original centrality values within the gene set of each pathway. Standardization is performed to obtain standardized betweenness centrality and standardized degree centrality. Specifically, if all genes within a pathway share the same topological feature value, then the standardized value for that feature is set to 0 for all genes within the pathway. This makes all the values in Within the interval, the assessment of topological importance is ensured to be compared within the context of the pathway itself.
[0103] Secondly, within the same pathway, the relative contributions of betweenness centrality and degree centrality to evaluating gene importance may differ. Therefore, it is necessary to further evaluate the weights of the standardized topological features of all genes within the pathway, namely, the standardized betweenness centrality and the standardized degree centrality. The formula for calculating the weight of the standardized betweenness centrality of a gene is as follows:
[0104]
[0105] wherein, denotes the weight of the normalized betweenness centrality in the pathway ; denotes the variance of the normalized betweenness centrality values of all genes in the pathway ; denotes the variance of the normalized degree centrality values of all genes in the pathway ; denotes a very small positive number used to prevent division by zero , which is set to in the present embodiment.
[0106] The calculation formula of the weight of the normalized degree centrality of the gene is:
[0107]
[0108] wherein, denotes the weight of the normalized degree centrality in the pathway ; denotes the variance of the normalized betweenness centrality values of all genes in the pathway ; denotes the variance of the normalized degree centrality values of all genes in the pathway ; denotes a very small positive number used to prevent division by zero , which is set to in the present embodiment.
[0109] After obtaining the weight of the normalized betweenness centrality and the weight of the normalized degree centrality in the pathway, the standardized topological features can be fused into a single comprehensive topological importance evaluation through the obtained weights, and the comprehensive topological importance evaluation formula is:
[0110]
[0111] wherein, denotes the comprehensive topological importance score of the gene in the pathway ; denotes the weight of the normalized betweenness centrality in the pathway ; denotes the normalized betweenness centrality value of the gene in the pathway ; denotes the weight of the normalized degree centrality in the pathway ; denotes the normalized degree centrality value of the gene in the pathway The standardized centrality value within.
[0112] After obtaining the importance score of genes in the pathway based purely on their structural location, in addition to static topological structure, whether the gene expression changes are coordinated with the state of its functional partners within the pathway is also an important aspect of measuring its contextual relevance. Based on this, the local expression coordination is evaluated by the correlation coefficient between the gene's expression spectrum vector and the average expression spectrum vector of all its neighboring genes in the pathway. The formula for calculating the correlation coefficient between the gene's expression spectrum vector and the average expression spectrum vector of all its neighboring genes in the pathway is as follows:
[0113]
[0114] in, Indicates gene In the pathway The expression spectral vector in the pathway and its relationship with the pathway The Pearson correlation coefficient between the average expression profile vectors of all neighboring genes in the sample; Indicates gene In the sample The spectral vector on the expression, Indicates gene In the pathway The number of neighboring genes; Indicates gene In the pathway The set of neighboring genes in; This represents the function for calculating the Pearson correlation coefficient.
[0115] It should be noted that if genes In the pathway There are no neighbors in the middle (| ), or genes The expression profile of the tumor sample or the average expression profile of its neighbors If there is no variation (i.e., the standard deviation required to calculate the correlation coefficient is zero), then... Defined as This indicates that effective expressive coordination cannot be assessed or does not exist. The value range is within Between them, genes were quantified. The synchronicity of changes in the environments of their direct interactions is considered. Positive correlation indicates synergistic changes, negative correlation indicates antagonistic changes, and near-zero correlation indicates relative independence. This step introduces dynamic functional-level information, supplementing the limitations of static topological analysis.
[0116] After obtaining the comprehensive topological importance assessment and local expression coordination assessment of the gene within the pathway, the contextual relevance optimization factor can be evaluated using these assessments. The formula for calculating the contextual relevance optimization factor is as follows:
[0117]
[0118] in, Indicates gene In the pathway Contextualization optimization factor; Indicates gene In the pathway The overall topological importance score within the area; Indicates gene In the pathway The expression spectral vector in the pathway and its relationship with the pathway The Pearson correlation coefficient between the average expression profile vectors of all neighboring genes.
[0119] It should be noted that, in order to ensure that the coordination of local gene expression in the pathway regulates the context association optimization factor (positive correlation enhances, negative correlation weakens), and to ensure that the factor has a positive value, the coordination of local expression is... Perform an exponential transformation so that the context-related optimization factor will The correlation is mapped to approximately The positive interval is calculated, and finally, the topological score (plus 1 to ensure the basic contribution is not less than 1) is multiplied by the coordination regulatory factor to obtain the context-related optimization factor. This method aims to reward genes that have both high topological importance and high coordination with their pathway neighbors; these genes will receive the highest reward. Value. Genes with low topological importance or inconsistent expression (especially negative correlations) have... The value will then decrease accordingly. Ensuring topological importance itself provides a fundamental contribution, while It then amplifies or reduces the expression based on dynamic relational relationships.
[0120] S2.3 Optimize the GSVA pathway perturbation score by integrating contribution weights.
[0121] After obtaining the dynamic influence factor and context-related optimization factor of the gene, a comprehensive contribution weight assessment can be performed using both. This optimization factor is then applied to the core enrichment score assessment of the standard gene set variation analysis algorithm to generate an optimized score. In this embodiment, through... calculate This optimization is achieved by weighting and modulating the sample statistics process.
[0122] In practice, for a single sample and a single pathway According to the standard Process, for samples genes Based on its original standardized expression level Sort the genes and calculate the empirical cumulative distribution function of genes within the pathway. The empirical cumulative distribution function of extra-pathway genes Access Path Each gene Overall contribution weight In comparison and To find the maximum deviation (i.e., to calculate) When using sample statistics, introduce weights. Modulation: Modulate each position in the sorted list The calculated original difference Weighting is applied. If the position... Corresponding genes It belongs to the pathway Then the weighted difference of its contribution is adjusted to If genes Not a pathway Then the weighted difference remains as Based on these weighted differences Find the maximum positive deviation and the maximum negative deviation, and follow... The rules are used to calculate the final enrichment score.
[0123] The above weighting will be applied. The enrichment score obtained by calculating the sample statistics is denoted as the optimization score. score This weighting method ensures that weights with high overall contribution are assigned. The genes (i.e., high-influence, highly associated genes) affect their location. Differences have a greater influence, thus being enhanced in the calculation of the final enrichment score. This is achieved by introducing... Modulation core computation can more accurately reflect pathway perturbation signals dominated by a few key genes. Final output optimization. Rating Matrix .
[0124] S3 is based on optimized scoring to identify abnormal regulatory pathways.
[0125] Through weighted optimization The optimization obtained by the method Rating Matrix (in Representative pathway In the sample The enrichment score obtained through optimization calculation was compared with different sample groups (tumor samples). Compared with normal samples The differences in scores between pathway enrichment score matrices were used to identify biological pathways that exhibit significant aberrant regulation in disease states. This step employed a standard statistical comparison framework from the field of bioinformatics for downstream analysis of pathway enrichment score matrices.
[0126] The specific implementation is as follows: For optimization Rating Matrix Each pathway in Extract it from the tumor sample group All the rating values constitute a vector and in the normal sample group All the rating values constitute a vector To assess the pathway To determine whether there is a statistically significant difference in the perturbation levels between the two groups, a non-parametric method is used. Rank-sum test comparison and The distribution. By analyzing each pathway Performing this test yields an original... value.
[0127] Considering the problem of multiple checks, we adopt Method to calculate the error detection rate ( ) for the original The value is corrected. A significance threshold is set. In this embodiment, a significance threshold is set. .Will Pathways with values less than this threshold The pathway was identified as having statistically significant perturbation differences between tumor and normal samples.
[0128] The final output is a list of abnormal regulatory pathways that meet the screening criteria. This list is an optimization based on this invention. Rating Matrix The filtering is performed. Due to the matrix... ratings In its calculation process, comprehensive weights have already been used. exist The core enrichment score calculation process enhances the signal contribution of high-influence, highly associated genes; therefore, conventional statistical methods are applied to the matrix. The obtained differential pathway identification results, compared with those obtained by applying the same statistical methods to standard... The calculated score diluted with the key signal can more sensitively and accurately identify abnormal regulation pathways closely related to disease mechanisms driven by a few key genes, embodying the improvement of the present application over the prior art in capturing the ability of a specific type of pathway disturbance. The method effectively improves the ability of capturing a specific type of pathway disturbance.
[0129] The above merely describes preferred embodiments of the present application and is not used to limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for detecting abnormal regulation of genes based on regulatory pathway perturbation analysis, characterized by, The method comprises the following steps: Step S1: collecting biological samples and pathway structure data; Step S2: obtaining a weighted optimized pathway perturbation score through gene dynamic behavior and pathway context information; Step S2.1: obtaining a gene dynamic influence factor through differential expression significance and sample consistency; Step S2.2: obtaining a pathway-specific context correlation optimization factor through pathway topology structure and expression coordination; Step S2.3: optimizing a GSVA pathway perturbation score through comprehensive contribution weight; Step S3: identifying an abnormal regulation pathway based on the optimized score.
2. The method for detecting abnormal gene regulation based on regulatory pathway perturbation analysis according to claim 1, wherein, According to the step S2.1, the gene dynamic influence factor is obtained through differential expression significance and sample consistency, and specifically comprises the following steps: A standardized gene expression profile matrix is obtained; t statistics and a two-tailed p value of genes are obtained through component differential expression analysis on the standardized gene expression profile matrix; a change amplitude feature of the genes is obtained through absolute value evaluation on the t statistics of the genes; a significance feature of the genes is obtained through negative logarithmic transformation on the two-tailed p value of the genes; a consistency feature of the genes is obtained through coefficient of variation evaluation on the expression values of the genes, and an expression consistency measurement factor of the genes is obtained through expression consistency analysis on the consistency feature of the genes; feature weights of three feature dimensions are obtained through dimension weight evaluation on the change amplitude feature, the significance feature and the expression consistency measurement factor of the genes, and a dynamic influence factor of the genes is obtained through influence comprehensive evaluation on the genes by using the feature weights of the three feature dimensions.
3. The method for detecting abnormal gene regulation based on regulatory pathway perturbation analysis according to claim 2, wherein, According to the step S2.1, the gene dynamic influence factor is obtained through differential expression significance and sample consistency, and specifically comprises the following steps: The calculation formula of the expression consistency measurement factor of the genes is: ; wherein, represents a gene consistency measure factor for expression of; represents a gene coefficient of variation in expression of; in a sample group represents a very small positive number.
4. The method for detecting abnormal gene regulation based on regulatory pathway perturbation analysis according to claim 3, wherein, According to the step S2.1, the gene dynamic influence factor is obtained through differential expression significance and sample consistency, and specifically comprises the following steps: The calculation formula of the weight of the significance feature dimension is: ; The calculation formula of the weight of the amplitude feature dimension is: ; The calculation formula of the weight of the consistency feature dimension is: ; wherein, denotes the weight of the significance feature dimension, i.e. the dimension weight of this statistical feature; denotes the weight of the amplitude feature dimension, i.e. the dimension weight of this statistical feature; denotes the weight of the consistency feature dimension, i.e. the dimension weight of the consistency measure factor; denotes the variance of the normalized significance feature over the entire set of genes; denotes the variance of the normalized amplitude feature over the entire set of genes; denotes the variance of the normalized consistency measure factor over the entire set of genes; denotes a very small positive number; The calculation formula of a single dynamic influence factor is: ; wherein, denotes the dynamic influence factor of a gene ; denotes an index symbol representing a specific gene in a genome; denote data weights of the significance feature, the amplitude feature and the consistency feature, respectively; denotes the significance feature of a gene after normalization; denotes the amplitude feature of a gene after normalization; denotes the consistency feature of a gene after normalization.
5. The gene abnormal regulation detection method based on regulatory pathway perturbation analysis according to claim 1, characterized in that, According to the step S2.2, the pathway-specific context correlation optimization factor is obtained through pathway topology structure and expression coordination, and specifically comprises the following steps: The calculation formula of the weight of the significance feature dimension is: The calculation formula of the weight of the amplitude feature dimension is: The calculation formula of the weight of the consistency feature dimension is: The calculation formula of a single dynamic influence factor is: According to the step S2.2, the pathway-specific context correlation optimization factor is obtained through pathway topology structure and expression coordination, and specifically comprises the following steps: The betweenness centrality evaluation and the degree centrality evaluation of the gene in the pathway are obtained; the betweenness centrality evaluation and the degree centrality evaluation of the gene in the pathway are standardized to obtain the standardized betweenness centrality evaluation and the standardized degree centrality evaluation of the gene in the pathway; the weight of the betweenness centrality and the weight of the degree centrality in the pathway are obtained by evaluating the weight of the variance of the standardized betweenness centrality and the variance of the standardized degree centrality of the gene in the pathway; the comprehensive topological importance evaluation of the gene in the pathway is obtained by fusing the standardized betweenness centrality evaluation and the standardized degree centrality evaluation of the gene through the weight of the betweenness centrality and the weight of the degree centrality in the pathway; the expression profile vector of the gene and the expression profile vector of all neighbor genes of the gene in the pathway are obtained; the local expression coordination evaluation of the gene is obtained by evaluating the correlation between the expression profile vector of the gene and the expression profile vector of all neighbor genes of the gene in the pathway; the context correlation optimization factor of the gene in the pathway is obtained by comprehensively analyzing the comprehensive topological importance evaluation of the gene in the pathway and the local expression coordination evaluation of the gene.
6. The method of claim 1, wherein the method comprises: According to the weight evaluation of the variance of the standardized betweenness centrality and the variance of the standardized degree centrality of the gene in the pathway, the weight of the betweenness centrality and the weight of the degree centrality in the pathway are obtained, which comprises: The weight calculation formula of the standardized betweenness centrality of the gene is: ; wherein, represents the weight of the normalized betweenness centrality in the pathway represents the variance of the normalized betweenness centrality values of all genes in the pathway represents the variance of the normalized degree centrality values of all genes in the pathway represents a very small positive number; The weight calculation formula of the standardized degree centrality of the gene is: ; wherein, denotes the weight of the normalized degree centrality in the pathway denotes the weight of the normalized betweenness centrality in the pathway denotes the variance of the normalized betweenness centrality values of all genes in the pathway denotes the variance of the normalized betweenness centrality values of all genes in the pathway denotes the variance of the normalized degree centrality values of all genes in the pathway denotes the variance of the normalized degree centrality values of all genes in the pathway denotes a very small positive number.
7. The method of claim 1, wherein the method is based on analysis of perturbation of a regulatory pathway. According to the fusion analysis of the standardized betweenness centrality evaluation and the standardized degree centrality evaluation of the gene through the weight of the betweenness centrality and the weight of the degree centrality in the pathway, the comprehensive topological importance evaluation of the gene in the pathway is obtained, which specifically comprises: The comprehensive topological importance evaluation formula is: ; wherein, represents the gene in the pathway with the integrated topological importance score; represents the weight of the normalized betweenness centrality in the pathway represents the normalized betweenness centrality value of the gene in the pathway represents the weight of the normalized degree centrality in the pathway represents the normalized degree centrality value of the gene in the pathway 8. The method for detecting abnormal gene regulation based on regulatory pathway perturbation analysis according to claim 1, wherein, According to the correlation evaluation between the expression profile vector of the gene and the expression profile vector of all neighbor genes of the gene in the pathway, the local expression coordination evaluation of the gene is obtained, which comprises: The local expression coordination evaluation calculation formula between the expression profile vector of the gene and the average expression profile vector of all neighbor genes of the gene in the pathway is: ; wherein, represents the local expression coordination assessment of a gene, i.e. the Pearson correlation coefficient between the expression profile vector of the gene in the pathway and the average expression profile vector of all its neighbor genes in the pathway ; represents the expression profile vector of a gene in a sample , represents the number of neighbor genes of a gene in a pathway ; represents the set of neighbor genes of a gene in a pathway ; represents the Pearson correlation coefficient computation function.
9. The method for detecting abnormal gene regulation based on regulatory pathway perturbation analysis according to claim 1, wherein, According to the comprehensive analysis of the comprehensive topological importance evaluation of the gene in the pathway and the local expression coordination evaluation of the gene, the context correlation optimization factor of the gene in the pathway is obtained, which comprises: The context correlation optimization factor calculation formula of the gene in the pathway is: ; wherein, represents the genes in the pathway in the context of the optimization factor; represents the genes in the pathway in the integrated topological importance score; represents the local expression coordination assessment of a gene, i.e. the gene in the pathway the correlation coefficient between the expression profile vector of a gene in the pathway and the average expression profile vector of all its neighboring genes in the pathway.
10. The method of claim 1, wherein the method is based on analysis of perturbation of a regulatory pathway. According to the step S2.3, the GSVA pathway perturbation score is optimized by the comprehensive contribution weight, which specifically comprises: The standardized gene expression profile matrix, the set biological pathway set and the member gene set in each pathway are obtained; the comprehensive contribution weight is evaluated by the dynamic influence factor and the context correlation factor of each gene and the pathway in which the gene is located, and the comprehensive contribution weight calculation formula is: ; wherein, represents the dynamic influence factor of the gene represents the contextual association optimization factor of the gene in the pathway represents the dynamic influence factor of the gene in the pathway represents the contextual association optimization factor of the gene in the pathway in the pathway When the pathway perturbation score calculation is performed, the GSVA pathway perturbation score is optimized by the following steps: Step 1: all genes in the sample are sorted in descending order of expression; Step 2: the empirical cumulative distribution function of the gene in the pathway and the empirical cumulative distribution function of the gene outside the pathway are calculated; Step 3: the original difference value is calculated for each sorting position; Step 4: If the gene corresponding to the ranking position is in the pathway, the final difference value is the result of weighting and adjusting the original difference value of the gene by the comprehensive contribution weight; if the gene corresponding to the ranking position is not in the pathway, the final difference value is the original difference value of the gene; Step 5: In the difference value sequence formed by the final difference values, the maximum positive deviation and the maximum negative deviation are obtained, and the pathway perturbation score is calculated; Step 6: Repeat the above steps to obtain a pathway perturbation score matrix of all pathways in the sample set, which is used to represent the optimized pathway perturbation degree.
Citation Information
Patent Citations
Method for identifying gene pathways based on GSA
CN107203704A
Method of inferring gene pathway activity
CN108763862A
Identification of early diagnosis markers of lung adenocarcinoma based on co-expression similarity, and constructing method of risk prediction model
CN109841281A
Key node identification method of disease marker expression regulation and control network
CN120260692A
Method for producing improved results for applications which directly or indirectly utilize gene expression assay results
US20070059685A1