Gene abnormal regulation detection method based on regulatory pathway perturbation analysis
Patent Information
- Application Number
- CN202511054381.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-07-30
AI Technical Summary
因此,标准方法对于准确识别并优先排序这类由少数高影响力基因强烈驱动的通路异常调控模式的敏感度不足
[0065]本发明所述的一种基于调控通路扰动分析的基因异常调控检测方法,通过设有
Smart Images

Figure CN120913644B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and in particular to a method for detecting abnormal gene regulation based on the analysis of regulatory pathway perturbations. Background Technology
[0002] In life science research, especially in oncology, understanding anomalous changes in gene expression regulatory networks is crucial for revealing the molecular mechanisms of disease development. Genes do not function independently but rather work together to form complex biological pathways (signaling pathways, metabolic pathways, etc.) to perform specific cellular functions. Analyzing changes in activity at the entire pathway level, rather than focusing solely on changes in the expression of individual genes, provides a more systematic and biologically meaningful perspective.
[0003] Existing technologies utilize high-throughput transcriptome sequencing ( Gene set variation analysis (GFA) is a commonly used strategy for gene expression profiling data generated by technologies such as gene set variation analysis. ) is one of the representative methods. It is an unsupervised, unbiased computational method that can calculate the enrichment score of each predefined gene set (pathway) for each sample. Specifically, By comparing the empirical cumulative distribution functions of genes within and outside the pathway in the overall gene expression ranking of this sample (…), ), and utilize similar The test statistic quantifies whether genes within a pathway tend to concentrate at high or low expression levels, thus obtaining the relative activity or perturbation level of that pathway in the sample. This is achieved by comparing samples from different groups (tumor samples and normal control samples). By analyzing fractional distributions, researchers can identify pathways that undergo significant changes in activity during disease states. These methods are widely used in exploring sample heterogeneity, associated pathway activity, and clinical phenotypes because they can assess pathway activity in each sample.
[0004] Regulatory abnormalities in biological systems are often complex. In many diseases, especially cancer, key phenotypic changes may be primarily caused by dramatic and functionally defined abnormalities in the expression of a few core driver genes, such as specific oncogenes or tumor suppressor genes.
[0005] Existing technology When assessing pathway perturbations, the core computational logic is based on comparing the overall rank distribution of expression levels of all genes within the pathway. While this scoring mechanism, based on the behavior of gene groups within a set, can reflect the overall trend of pathway activity, it inherently averages the contributions of all genes within the pathway or treats them as having an equal impact on the distribution. Specifically, when a pathway's anomalous perturbation is primarily dominated by a few key genes (i.e., high-influence genes) exhibiting extremely large, statistically significant, and highly consistent expression changes across samples, The calculated enrichment fractions mix and dilute the strong signals from these key genes with a large number of other genes within the pathway that show weak, insignificant, or even opposite signals. Therefore, the standard... The current method lacks sensitivity in accurately identifying and prioritizing abnormal pathway regulatory patterns strongly driven by a few high-impact genes. It cannot effectively distinguish whether pathway perturbations are caused by generalized moderate changes in a large number of genes or by potent changes in a few key genes. This directly leads researchers to overlook biological pathways whose overall gene set changes are not prominent, but whose dysfunction is closely related to the abnormal activity of specific key driver genes (which are often important potential therapeutic targets or the core of disease mechanisms). This hinders a deeper understanding of the molecular mechanisms of diseases and the development of precise intervention strategies. Therefore, there is an urgent need to propose a new method for assessing pathway perturbations to overcome the above-mentioned shortcomings of existing technologies and improve the ability to identify abnormal regulatory events in key pathways dominated by high-impact genes. Summary of the Invention
[0006] In view of this, the present invention aims to propose a gene abnormal regulation detection method based on regulatory pathway perturbation analysis, so as to improve the ability to identify abnormal regulatory events in key pathways dominated by high-influence genes.
[0007] To achieve the above objectives, the technical solution of the present invention is implemented as follows:
[0008] A gene regulation abnormality detection method based on regulatory pathway perturbation analysis, the method comprising the following steps:
[0009] Step S1: Collect biological samples and pathway structure data;
[0010] Step S2: Obtain a weighted and optimized pathway perturbation score by analyzing gene dynamics and pathway context information;
[0011] Step S2.1: Obtain the dynamic influence factor of genes by measuring the significance of differential expression and the consistency between samples;
[0012] Step S2.2: Obtain pathway-specific contextual optimization factors through pathway topology and expression coordination;
[0013] Step S2.3: Optimize the GSVA pathway perturbation score by integrating contribution weights;
[0014] Step S3: Identify abnormal regulatory pathways based on optimized scoring.
[0015] Furthermore, according to step S2.1, obtaining the gene dynamic influence factor through differential expression significance and inter-sample consistency specifically includes:
[0016] A standardized gene expression profile matrix is obtained. Component-based differential expression analysis is performed on the standardized gene expression profile matrix to obtain the t-statistic and two-tailed p-values of the genes. The absolute value of the t-statistics is used to assess the gene variation amplitude characteristics. A negative logarithmic transformation is applied to the two-tailed p-values to obtain the gene significance characteristics. The coefficient of variation of gene expression values is used to assess gene consistency characteristics, and expression consistency analysis is performed on these consistency characteristics to obtain the gene expression consistency measure factor. Dimensional weights are evaluated for the gene variation amplitude characteristics, significance characteristics, and expression consistency measure factor to obtain the feature weights of each of the three feature dimensions. Finally, the influence of the gene is comprehensively evaluated using the feature weights of each of the three feature dimensions to obtain the dynamic influence factor of the gene.
[0017] Furthermore, based on the aforementioned method of evaluating gene expression values using the coefficient of variation to obtain gene consistency characteristics, and by performing expression consistency analysis on these characteristics, gene expression consistency metrics are obtained, including:
[0018] The formula for calculating the gene expression consistency measure factor is:
[0019] ;
[0020] in, Indicates gene Expression consistency measure factor; Indicates gene In the sample group The coefficient of variation in expression; It represents a very small positive number.
[0021] Furthermore, based on the aforementioned dimensional weighting assessment of gene variation magnitude characteristics, significance characteristics, and expression consistency measurement factors, the feature weights of each of the three feature dimensions are obtained. Then, the gene's influence is comprehensively assessed using the feature weights of each of the three feature dimensions to obtain the dynamic influence factor of the gene, specifically including:
[0022] The formula for calculating the weights of the saliency feature dimension is:
[0023] ;
[0024] The formula for calculating the weights of the magnitude feature dimension is:
[0025] ;
[0026] The formula for calculating the weight of the consistency feature dimension is:
[0027] ;
[0028] in, The weights of the saliency feature dimensions, i.e. The dimensional weights of this statistical feature; The weights representing the magnitude feature dimension are... The dimensional weights of this statistical feature; The weights representing the consistency feature dimensions, i.e. Dimensional weights of consistency metrics; This represents the variance of the standardized salient features across the entire gene set. This represents the variance of the standardized amplitude feature across the entire gene set. This represents the variance of the standardized consistency measure across the entire gene set. Represents a very small positive number;
[0029] The formula for calculating a single dynamic influence factor is:
[0030] ;
[0031] in, Indicates gene Dynamic influence factors; This represents an index symbol, indicating a specific gene in the genome; These represent the data weights for significance, magnitude, and consistency features, respectively. Indicates gene Significant features after standardization; Indicates gene Amplitude characteristics after standardization; Indicates gene Consistency characteristics after standardization.
[0032] Furthermore, according to step S2.2, obtaining pathway-specific contextual optimization factors through pathway topology and expression coordination specifically includes:
[0033] The study obtains betweenness centrality and degree centrality assessments of genes within pathways; standardizes these assessments to obtain standardized betweenness centrality and degree centrality assessments; weights are applied to the standardized betweenness centrality variance and standardized degree centrality variance of genes within pathways to obtain betweenness centrality weights and degree centrality weights; a fusion analysis of these weights yields a comprehensive topological importance assessment of genes within pathways; the expression profile vectors of genes and their neighboring genes within the pathway are obtained; correlation assessments are performed between these vectors to obtain a local expression coordination assessment of genes; and a comprehensive analysis of the comprehensive topological importance assessment and the local expression coordination assessment of genes within pathways yields a contextual relevance optimization factor for genes within pathways.
[0034] Furthermore, based on the aforementioned weighting of the normalized betweenness centrality variance and normalized degree centrality variance of genes within the pathway, the betweenness centrality weight and degree centrality weight within the pathway are obtained, including:
[0035] The formula for calculating the weight of the normalized betweenness centrality of a gene is:
[0036] ;
[0037] in, Indicates pathway The weight of the normalized betweenness centrality in the equation; Indicates pathway The variance of the standardized betweenness centrality values of all genes within the group; Indicates pathway The variance of the normalized centrality values of all genes within the group; Represents a very small positive number;
[0038] The formula for calculating the weights of gene normalization centrality is:
[0039] ;
[0040] in, Indicates pathway The weights of standardization centrality in the equation; Indicates pathway The variance of the standardized betweenness centrality values of all genes within the group; Indicates pathway The variance of the normalized centrality values of all genes within the group; It represents a very small positive number.
[0041] Furthermore, based on the aforementioned fusion analysis of the normalized betweenness centrality assessment and normalized degree centrality assessment of genes through the betweenness centrality weight and degree centrality weight within the pathway, a comprehensive topological importance assessment of genes within the pathway is obtained, specifically including:
[0042] The formula for evaluating the overall topological importance is:
[0043] ;
[0044] in, Indicates gene In the pathway The overall topological importance score within the area; Indicates pathway The weight of the normalized betweenness centrality in the equation; Indicates gene In the pathway The normalized betweenness centrality value within; Indicates pathway The weights of standardization centrality in the equation; Indicates gene In the pathway The standardized centrality value within.
[0045] Furthermore, based on the aforementioned method of assessing the correlation between the gene's expression profile vector and the expression profile vectors of all neighboring genes in the pathway, an assessment of the gene's local expression coordination is obtained, including:
[0046] The formula for calculating the local expression coordination of a gene between its expression profile vector and the average expression profile vector of all its neighboring genes in the pathway is as follows:
[0047] ;
[0048] in, This refers to the assessment of the local expression coordination of genes, i.e., genes. In the pathway The expression spectral vector in the pathway and its relationship with the pathway The correlation coefficient between the average expression profile vectors of all neighboring genes in the sample; Indicates gene In the sample The spectral vector on the expression, Indicates gene In the pathway The number of neighboring genes; Indicates gene In the pathway The set of neighboring genes in; This represents the function for calculating the Pearson correlation coefficient.
[0049] Furthermore, based on the comprehensive analysis of gene topological importance assessment and gene local expression coordination assessment within the pathway, contextual association optimization factors for genes within the pathway are obtained, including:
[0050] The formula for calculating the contextual optimization factor of genes within a pathway is as follows:
[0051] ;
[0052] in, Indicates gene In the pathway Contextualization optimization factor; Indicates gene In the pathway The overall topological importance score within the area; This refers to the assessment of the local expression coordination of genes, i.e., genes. In the pathway The expression spectral vector in the pathway and its relationship with the pathway The correlation coefficient between the average expression profile vectors of all neighboring genes.
[0053] Furthermore, according to step S2.3, optimizing the GSVA pathway perturbation score through comprehensive contribution weighting specifically includes:
[0054] Obtain a standardized gene expression profile matrix and a defined set of biological pathways, along with the set of member genes within each pathway. Assess the overall contribution weight of each gene by evaluating its dynamic influence factors and contextual factors within its pathway. The formula for calculating the overall contribution weight is as follows:
[0055] ;
[0056] in, Indicates gene In the pathway The overall contribution weight within, Indicates gene Dynamic influence factors; Indicates gene In the pathway Contextualization optimization factor;
[0057] When calculating the GSVA path disturbance score, the following steps are used to optimize the GSVA path disturbance score:
[0058] Step 1: Sort all genes in the sample in descending order of expression level;
[0059] Step 2: Calculate the empirical cumulative distribution function of genes within the pathway and the empirical cumulative distribution function of genes outside the pathway in the sorting list;
[0060] Step 3: Calculate the original difference for each sorting position;
[0061] Step 4: If the gene corresponding to the sorting position is within the pathway, the original difference is used as the final difference after weighting by the comprehensive contribution weight; if the gene corresponding to the sorting position is not within the pathway, the original difference is used as the final difference.
[0062] Step 5: In the difference sequence formed by the final difference, obtain the maximum positive deviation and the maximum negative deviation, and calculate the path disturbance score;
[0063] Step 6: Repeat the above steps to obtain the pathway perturbation score matrix for all pathways in the sample set, which is used to represent the degree of optimized pathway perturbation.
[0064] Compared with the prior art, the present invention has the following advantages:
[0065] The present invention discloses a gene abnormality regulation detection method based on regulatory pathway perturbation analysis, which uses a method with...
[0066] This invention significantly improves the sensitivity and accuracy of gene set variation analysis methods in identifying pathway perturbations in complex diseases by introducing dynamic influence factors and pathway context-related factors at the gene level. By integrating the significance, magnitude, and consistency of gene differential expression in tumor and normal samples, as well as within the tumor population, a dynamic influence factor is constructed. This effectively identifies high-influence genes that exhibit stable and significantly abnormal expression in the disease context, overcoming the signal dilution problem caused by averaging in traditional methods. Furthermore, by combining the regulatory role of genes in pathway topology and their coordination with neighboring genes in expression behavior, a pathway-specific context optimization factor is constructed to finely modulate the weight contributions in different pathways, thereby improving the responsiveness of pathway perturbation scores to structural core genes and functional consistency signals. Through these two optimization steps, this invention effectively enhances the ability to identify pathway perturbations dominated by a few key genes, improves the biological interpretability and downstream clinical application value of pathway analysis results, and is applicable to molecular mechanism research and target discovery in heterogeneous diseases such as tumors. Attached Figure Description
[0067] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0068] Figure 1This is a flowchart of a gene abnormal regulation detection method based on regulatory pathway perturbation analysis, as described in an embodiment of the present invention. Detailed Implementation
[0069] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0070] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "back," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0071] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0072] See Figure 1 This is a flowchart of a gene abnormal regulation detection method based on regulatory pathway perturbation analysis provided in Embodiment 1 of the present invention, as follows: Figure 1 As shown, a gene abnormality regulation detection method based on regulatory pathway perturbation analysis may include:
[0073] S1 collects biological samples and pathway structure data.
[0074] In this embodiment, raw sequencing reads of tissue samples from non-small cell lung cancer patients and their paired or adjacent normal tissue samples are first obtained from publicly available datasets. Simultaneously, predefined biological pathway annotation information is acquired, which includes a series of known biological pathways and lists the set of member genes contained in each pathway. Furthermore, topological structure information capable of describing known interactions or regulatory relationships between genes within a pathway needs to be acquired or constructed. In this embodiment, both the pathway definitions and corresponding standard topological structure information are obtained from publicly available datasets. Database. Next, the acquired raw sequencing data undergoes necessary bioinformatics processing, including quality control, read alignment, gene expression quantification and standardization, ultimately resulting in a standardized gene expression profile matrix, where rows represent genes and columns represent samples (including tumor samples and normal samples).
[0075] S2 obtains a weighted and optimized pathway perturbation score through gene dynamic behavior and pathway context information.
[0076] When assessing pathway perturbations using existing techniques such as gene set variation analysis, the core mechanism relies on comparing the relative positional distribution of genes within and outside the pathway within the overall gene expression profile of a single sample. This method, based on rank and empirical cumulative distribution function comparisons, aims to capture the overall upregulation or downregulation trend of pathway activity. A key technical limitation is that this method implicitly averages the contributions of all genes within the pathway. The final enrichment score primarily reflects the distributional shift of the gene set as a whole, but it cannot effectively distinguish whether this shift is driven by small but consistent changes in a large number of genes, or by large, statistically significant changes in a few genes within the pathway. In the biological mechanisms of complex diseases such as non-small cell lung cancer, the latter case—the dramatic dysregulation of a few high-influence genes—is more indicative of key pathogenic events or potential therapeutic targets.
[0077] This "signal dilution" effect is This is an inherent property of set distribution comparison methods. Even if a gene's expression level changes dramatically (a key oncogene is activated tens of times), and this change is statistically highly reliable, the strong signal of this key gene will be significantly weakened when calculating the overall distribution shift, provided that a sufficient number of other genes within the pathway show little or no change in the same direction. Furthermore, This method primarily focuses on the relative position of genes in the ranking list. For genes already at the extremes of the ranking list (whether due to high or low expression), the magnitude of their changes has a relatively limited marginal contribution to the final enrichment score. Furthermore, this method does not consider the stability and consistency of gene expression changes across different disease samples. A gene exhibiting similar and significant expression abnormalities in all or most tumor samples is generally more representative of a stable and universally significant regulatory abnormality in that disease state than a gene that shows drastic changes only in a few samples but is normal in others.
[0078] S2.1, obtain the dynamic influence factor of genes by differential expression significance and consistency between samples.
[0079] To overcome the aforementioned limitations and improve the ability to identify pathway perturbations dominated by high-influence genes, the first focus is on quantifying the dynamic influence of each gene itself. This dynamic influence is not based on its static role within a pathway, but rather on its actual expression behavior characteristics exhibited in the specific biological context under study. The dynamic influence of a gene should comprehensively reflect the magnitude of its expression changes, statistical significance, and consistency across disease sample populations.
[0080] Specifically, firstly, based on the standardized gene expression matrix, differential expression analysis between the two groups was used to obtain statistical indicators for each gene, including... Statistics and corresponding two-tailed statistics The values are used to describe the magnitude and statistical significance of the difference, respectively. To quantify statistical significance, [the following values are used for...]. The value is obtained by performing a negative logarithmic transformation. The larger the value, the more significant the difference; to quantify the magnitude of change, a more significant value is used. absolute value of the statistic This indicator takes into account both effect size and data variability, making it more robust than simply considering fold changes.
[0081] To further reflect the consistency of gene expression behavior in the disease population, the coefficient of variation of gene expression in the tumor sample group was calculated. The formula for calculating the coefficient of variation of gene expression in the tumor sample group is as follows:
[0082]
[0083] in, Indicates gene In tumor sample group The coefficient of variation in expression; Indicates gene In tumor sample collection The standard deviation of the expression values in the data; Indicates gene In tumor sample collection Mean expression value in; Representing extremely small positive numbers is used to prevent the denominator from being 0. In this embodiment, it is set .
[0084] After obtaining the gene expression variation coefficient, an expression consistency measure factor can be constructed based on it. The formula for calculating the expression consistency measure factor is as follows:
[0085]
[0086] in, Indicates gene Expression consistency measure factor; Indicates gene In tumor sample group The coefficient of variation in expression; Representing extremely small positive numbers is used to prevent the denominator from being 0. In this embodiment, it is set .
[0087] The three indicators mentioned above ( These features together constitute the original feature set for assessing the dynamic influence of genes, and provide a preliminary characterization of abnormal gene expression behavior from the perspectives of reliability, strength, and universality.
[0088] Secondly, since the three original features extracted above have different dimensions and numerical ranges, directly combining them will lead to a result biased towards the feature with the larger numerical range. To ensure the comparability of each feature in the subsequent integration process, it is necessary to standardize the three features. In this embodiment, standardization is adopted. The standardized method was used to perform differential expression analysis and successfully obtained statistics for three features of the set of all genes. The standardization saliency feature, standardization magnitude feature, and standardization consistency feature are processed to obtain the standardization saliency feature, standardization magnitude feature, and standardization consistency feature.
[0089] Furthermore, different feature dimensions (saliency, magnitude, consistency) do not contribute equally to the final determination of whether a gene is highly influential, and this relative importance depends on the characteristics of the specific dataset. This invention employs a data-driven weighting strategy to calculate the variance of each standardized feature in the entire gene set. The magnitude of the variance reflects, to some extent, the ability or information content of that feature to distinguish different gene behaviors in the current dataset. Accordingly, the weight of each of the three different features is evaluated by the proportion of the variance of each feature to the total variance. The formula for calculating the weight of the saliency feature dimension is as follows:
[0090]
[0091] The formula for calculating the weights of the magnitude feature dimension is:
[0092]
[0093] The formula for calculating the weight of the consistency feature dimension is:
[0094]
[0095] in, The weights of the saliency feature dimensions, i.e. The dimensional weights of this statistical feature; The weights representing the magnitude feature dimension are... The dimensional weights of this statistical feature; The weights representing the consistency feature dimensions, i.e. Dimensional weights of consistency metrics; This represents the variance of the standardized salient features across the entire gene set. This represents the variance of the standardized amplitude feature across the entire gene set. This represents the variance of the standardized consistency measure across the entire gene set. Representing extremely small positive numbers is used to prevent the denominator from being 0. In this embodiment, it is set .
[0096] It should be noted that this method of determining feature weights based on the statistical characteristics of the data itself enhances the adaptability of this method to different datasets, making the assessment of the dynamic influence of genes more consistent with the actual data distribution.
[0097] Finally, the standardized and weighted features are comprehensively evaluated to obtain a single dynamic influence factor. The formula for calculating the single dynamic influence factor is as follows:
[0098]
[0099] in, Indicates gene Dynamic influence factors; This represents an index symbol, indicating a specific gene in the genome; These represent the data weights for significance, magnitude, and consistency features, respectively. Indicates gene Significant features after standardization; Indicates gene Amplitude characteristics after standardization; Indicates gene Consistency characteristics after standardization.
[0100] S2.2, obtain pathway-specific contextual optimization factors through pathway topology and expression coordination.
[0101] Obtaining the dynamic influence factor initially addresses the problem of high-influence gene signals being diluted due to existing technologies failing to fully consider individual gene expression behavior. However, the dynamic influence factor is a gene-centric measure, assessing the gene's own dynamic characteristics without relating it to its role and behavioral patterns within a specific biological pathway. A gene may have a high influence factor (i.e., its own changes are significant and consistent), but if it occupies a peripheral position in the pathway's topology, or if its expression changes contradict or are unrelated to the changing trends of its direct functional partners within the pathway, its actual contribution to driving overall dysfunction of that specific pathway is not as large as the dynamic influence factor suggests. Specifically, amplifying gene signals solely based on the dynamic influence factor can enhance the influence of genes that, while exhibiting dramatic changes, are not core or exhibit uncoordinated behavior within the functional regulatory network of the examined pathway, thus leading to biased judgments about the source of pathway disturbances.
[0102] Therefore, based on the dynamic influence factors of genes, it is necessary to further evaluate specific context-related optimization factors by considering the structural importance and functional coordination of genes within specific pathways. When evaluating context-related optimization factors, the first consideration is the topological importance and local expression coordination of genes within the pathway network structure. For the comprehensive topological importance of genes within a pathway, the betweenness centrality and degree centrality of genes within the pathway are first obtained. Betweenness centrality quantifies the frequency with which a gene acts as a bridge for information transmission between other gene pairs within the pathway, reflecting the gene's criticality in controlling the information flow of the pathway. Degree centrality quantifies the number of directly connected neighbors of a gene, reflecting the breadth of its local connections. Since different pathways have different sizes and connection patterns, directly comparing the original centrality values across pathways has limited significance. Therefore, it is necessary to evaluate these two original centrality values within the gene set of each pathway. Standardization is performed to obtain standardized betweenness centrality and standardized degree centrality. Specifically, if all genes within a pathway share the same topological feature value, then the standardized value for that feature is set to 0 for all genes within the pathway. This makes all the values in Within the interval, the assessment of topological importance is ensured to be compared within the context of the pathway itself.
[0103] Secondly, within the same pathway, the relative contributions of betweenness centrality and degree centrality to evaluating gene importance may differ. Therefore, it is necessary to further evaluate the weights of the standardized topological features of all genes within the pathway, namely, the standardized betweenness centrality and the standardized degree centrality. The formula for calculating the weight of the standardized betweenness centrality of a gene is as follows:
[0104]
[0105] in, Indicates pathway The weight of the normalized betweenness centrality in the equation; Indicates pathway The variance of the standardized betweenness centrality values of all genes within the group; Indicates pathway The variance of the normalized centrality values of all genes within the group; Representing extremely small positive numbers is used to prevent the denominator from being 0. In this embodiment, it is set .
[0106] The formula for calculating the weights of gene normalization centrality is:
[0107]
[0108] in, Indicates pathway The weights of standardization centrality in the equation; Indicates pathway The variance of the standardized betweenness centrality values of all genes within the group; Indicates pathway The variance of the normalized centrality values of all genes within the group; Representing extremely small positive numbers is used to prevent the denominator from being 0. In this embodiment, it is set .
[0109] After obtaining the weights of the normalized betweenness centrality and the normalized degree centrality in the pathway, the normalized topological features can be fused into a single comprehensive topological importance assessment using the obtained weights. The formula for the comprehensive topological importance assessment is as follows:
[0110]
[0111] in, Indicates gene In the pathway The overall topological importance score within the area; Indicates pathway The weight of the normalized betweenness centrality in the equation; Indicates gene In the pathway The normalized betweenness centrality value within; Indicates pathway The weights of standardization centrality in the equation; Indicates gene In the pathway The standardized centrality value within.
[0112] After obtaining the importance score of genes in the pathway based purely on their structural location, in addition to static topological structure, whether the gene's expression changes are coordinated with the state of its functional partners within the pathway is also an important aspect of measuring its contextual relevance. Based on this, the local expression coordination is evaluated by using the correlation coefficient between the gene's expression spectrum vector and the average expression spectrum vector of all its neighboring genes in the pathway. The formula for calculating the correlation coefficient between the gene's expression spectrum vector and the average expression spectrum vector of all its neighboring genes in the pathway is as follows:
[0113]
[0114] in, Indicates gene In the pathway The expression spectral vector in the pathway and its relationship with the pathway The Pearson correlation coefficient between the average expression profile vectors of all neighboring genes in the sample; Indicates gene In the sample The spectral vector on the expression, Indicates gene In the pathway The number of neighboring genes; Indicates gene In the pathway The set of neighboring genes in; This represents the function for calculating the Pearson correlation coefficient.
[0115] It should be noted that if genes In the pathway There are no neighbors in the middle (| ), or genes The expression profile of the tumor sample or the average expression profile of its neighbors If there is no variation (i.e., the standard deviation required to calculate the correlation coefficient is zero), then... Defined as This indicates that effective expressive coordination cannot be assessed or does not exist. The value range is within Between them, genes were quantified. The synchronicity of changes in the environments of their direct interactions is considered. Positive correlation indicates synergistic changes, negative correlation indicates antagonistic changes, and near-zero correlation indicates relative independence. This step introduces dynamic functional-level information, supplementing the limitations of static topological analysis.
[0116] After obtaining the comprehensive topological importance assessment and local expression coordination assessment of the gene within the pathway, the contextual relevance optimization factor can be evaluated using these assessments. The formula for calculating the contextual relevance optimization factor is as follows:
[0117]
[0118] in, Indicates gene In the pathway Contextualization optimization factor; Indicates gene In the pathway The overall topological importance score within the area; Indicates gene In the pathway The expression spectral vector in the pathway and its relationship with the pathway The Pearson correlation coefficient between the average expression profile vectors of all neighboring genes.
[0119] It should be noted that, in order to ensure that the coordination of local gene expression in the pathway regulates the context association optimization factor (positive correlation enhances, negative correlation weakens), and to ensure that the factor has a positive value, the coordination of local expression is... Perform an exponential transformation so that the context-related optimization factor will The correlation is mapped to approximately The positive interval is calculated, and finally, the topological score (plus 1 to ensure the basic contribution is not less than 1) is multiplied by the coordination regulator to obtain the context-related optimization factor. This method aims to reward genes that have both high topological importance and high coordination with their pathway neighbors; these genes will receive the highest reward. Value. Genes with low topological importance or inconsistent expression (especially negative correlations) have... The value will then decrease accordingly. Ensuring topological importance itself provides a fundamental contribution, while It then amplifies or reduces the expression based on dynamic relational relationships.
[0120] S2.3 Optimize the GSVA pathway perturbation score by integrating contribution weights.
[0121] After obtaining the dynamic influence factor and context-related optimization factor of the gene, a comprehensive contribution weight assessment can be performed using both. This optimization factor is then applied to the core enrichment score assessment of the standard gene set variation analysis algorithm to generate an optimized score. In this embodiment, through... calculate This optimization is achieved by weighting and modulating the sample statistics process.
[0122] In practice, for a single sample and a single pathway According to the standard Process, for samples genes Based on its original standardized expression level Sort the genes and calculate the empirical cumulative distribution function of genes within the pathway. The empirical cumulative distribution function of extra-pathway genes Access Path Each gene Overall contribution weight In comparison and To find the maximum deviation (i.e., to calculate) When using sample statistics, introduce weights. Modulation: Modulate each position in the sorted list The calculated original difference Weighting is applied. If the position... Corresponding genes It belongs to the pathway Then the weighted difference of its contribution is adjusted to If genes Not a pathway Then the weighted difference remains as Based on these weighted differences Find the maximum positive deviation and the maximum negative deviation, and follow... The rules are used to calculate the final enrichment score.
[0123] The above weighting will be applied. The enrichment score obtained by calculating the sample statistics is denoted as the optimization score. score This weighting method ensures that weights with high overall contribution are assigned. The genes (i.e., high-influence, highly associated genes) affect their location. Differences have a greater influence, thus being enhanced in the calculation of the final enrichment score. This is achieved by introducing... Modulation core computation can more accurately reflect pathway perturbation signals dominated by a few key genes. Final output optimization. Rating Matrix .
[0124] S3 is based on optimized scoring to identify abnormal regulatory pathways.
[0125] Through weighted optimization The optimization obtained by the method Rating Matrix (in Representative pathway In the sample The enrichment score obtained through optimization calculation was compared with different sample groups (tumor samples). Compared with normal samples The differences in scores between pathway enrichment score matrices were used to identify biological pathways that exhibit significant aberrant regulation in disease states. This step employed a standard statistical comparison framework from the field of bioinformatics for downstream analysis of pathway enrichment score matrices.
[0126] The specific implementation is as follows: For optimization Rating Matrix Each pathway in Extract it from the tumor sample group All the rating values constitute a vector and in the normal sample group All the rating values constitute a vector To assess the pathway To determine whether there is a statistically significant difference in the perturbation levels between the two groups, a non-parametric method is used. Rank-sum test comparison and The distribution. By analyzing each pathway Performing this test yields an original... value.
[0127] Considering the problem of multiple checks, we adopt Method to calculate the error detection rate ( ) to the original The value is corrected. A significance threshold is set. In this embodiment, a significance threshold is set. .Will Pathways with values less than this threshold The pathway was identified as having statistically significant perturbation differences between tumor and normal samples.
[0128] The final output is a list of abnormal regulatory pathways that meet the screening criteria. This list is an optimization based on this invention. Rating Matrix The filtering is performed. Due to the matrix... ratings In its calculation process, comprehensive weights have already been used. exist The core enrichment score calculation process enhances the signal contribution of high-influence, highly associated genes; therefore, conventional statistical methods are applied to the matrix. The obtained differential pathway identification results, compared with those obtained by applying the same statistical methods to standard... The calculated, diluted scores of key signals will enable more sensitive and accurate identification of aberrant regulatory pathways driven by a few key genes and closely related to disease mechanisms, demonstrating the advantages of this invention over existing methods. This represents an effective improvement in the method's ability to capture specific types of pathway perturbations.
[0129] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting gene abnormal regulation based on regulatory pathway perturbation analysis, characterized in that, The method includes the following steps: Step S1: Collect biological samples and pathway structure data; Step S2: Obtain a weighted and optimized pathway perturbation score by analyzing gene dynamics and pathway context information; Step S2.1: Obtain the dynamic influence factor of genes by measuring the significance of differential expression and the consistency between samples; Step S2.2: Obtain pathway-specific contextual optimization factors through pathway topology and expression coordination; Step S2.3: Optimize the GSVA pathway perturbation score by integrating contribution weights; Step S3: Identify abnormal regulatory pathways based on optimized scoring; Step S2.1, which obtains the dynamic influence factor of genes through differential expression significance and inter-sample consistency, specifically includes: obtaining a standardized gene expression profile matrix; performing component differential expression analysis on the standardized gene expression profile matrix to obtain the t-statistic and two-tailed p-value of genes; obtaining the gene variation amplitude characteristics by evaluating the absolute value of the gene t-statistic; obtaining the gene significance characteristics by performing negative logarithmic transformation on the gene's two-tailed p-value; obtaining the gene consistency characteristics by evaluating the coefficient of variation of gene expression values, and obtaining the gene expression consistency measurement factor by performing expression consistency analysis on the gene consistency characteristics; and obtaining the feature weights of each of the three feature dimensions by evaluating the dimensional weights of the gene variation amplitude characteristics, significance characteristics, and expression consistency measurement factor, and then comprehensively evaluating the gene's influence using the feature weights of each of the three feature dimensions to obtain the dynamic influence factor of the gene. Step S2.2, which obtains pathway-specific contextualization optimization factors through pathway topology and expression coordination, specifically includes: obtaining betweenness centrality and degree centrality assessments of genes within the pathway; standardizing these assessments to obtain standardized betweenness centrality and degree centrality assessments; weighting the standardized betweenness centrality variance and degree centrality variance of genes within the pathway to obtain betweenness centrality weights and degree centrality weights; performing a fusion analysis of the standardized betweenness centrality and degree centrality assessments of genes within the pathway using these weights to obtain a comprehensive topological importance assessment of genes within the pathway; obtaining the expression spectrum vector of the gene and the expression spectrum vectors of all neighboring genes in the pathway; evaluating the correlation between the gene's expression spectrum vector and the expression spectrum vectors of all neighboring genes in the pathway to obtain a local expression coordination assessment of the gene; and comprehensively analyzing the comprehensive topological importance assessment and the local expression coordination assessment of genes within the pathway to obtain contextualization optimization factors for genes within the pathway. Step S2.3 optimizes the GSVA pathway perturbation score by integrating contribution weights, specifically including: Obtain a standardized gene expression profile matrix and a defined set of biological pathways, along with the set of member genes within each pathway. Assess the overall contribution weight of each gene by evaluating its dynamic influence factors and contextual factors within its pathway. The formula for calculating the overall contribution weight is as follows: ; in, Indicates gene In the pathway The overall contribution weight within, Indicates gene Dynamic influence factors; Indicates gene In the pathway Contextualization optimization factor; When calculating the GSVA path disturbance score, the following steps are used to optimize the GSVA path disturbance score: Step 1: Sort all genes in the sample in descending order of expression level; Step 2: Calculate the empirical cumulative distribution function of genes within the pathway and the empirical cumulative distribution function of genes outside the pathway in the sorting list; Step 3: Calculate the original difference for each sorting position; Step 4: If the gene corresponding to the sorting position is within the pathway, the original difference is used as the final difference after weighting by the comprehensive contribution weight; if the gene corresponding to the sorting position is not within the pathway, the original difference is used as the final difference. Step 5: In the difference sequence formed by the final difference, obtain the maximum positive deviation and the maximum negative deviation, and calculate the path disturbance score; Step 6: Repeat the above steps to obtain the pathway perturbation score matrix for all pathways in the sample set, which is used to represent the degree of optimized pathway perturbation.
2. The gene abnormal regulation detection method based on regulatory pathway perturbation analysis according to claim 1, characterized in that, Based on the aforementioned method of evaluating gene expression values using the coefficient of variation, gene consistency characteristics are obtained. Furthermore, by performing expression consistency analysis on these characteristics, gene expression consistency metrics are derived, including: The formula for calculating the gene expression consistency measure factor is: ; in, Indicates gene Expression consistency measure factor; Indicates gene In the sample group The coefficient of variation in expression; It represents a very small positive number.
3. The gene abnormal regulation detection method based on regulatory pathway perturbation analysis according to claim 2, characterized in that, Based on the aforementioned method of evaluating the dimensional weights of gene variation magnitude characteristics, significance characteristics, and expression consistency measurement factors, the feature weights of each of the three feature dimensions are obtained. Then, the overall influence of the gene is comprehensively evaluated using these three feature weights to obtain the dynamic influence factor of the gene, specifically including: The formula for calculating the weights of the saliency feature dimension is: ; The formula for calculating the weights of the amplitude feature dimension is: ; The formula for calculating the weight of the consistency feature dimension is: ; in, The weights of the saliency feature dimensions, i.e. The dimensional weights of this statistical feature; The weights representing the magnitude feature dimension are... The dimensional weights of this statistical feature; The weights representing the consistency feature dimensions, i.e. Dimensional weights of consistency metrics; This represents the variance of the standardized salient features across the entire gene set. This represents the variance of the standardized amplitude feature across the entire gene set. This represents the variance of the standardized consistency measure across the entire gene set. Represents a very small positive number; The formula for calculating a single dynamic influence factor is: ; in, Indicates gene Dynamic influence factors; This represents an index symbol, indicating a specific gene in the genome; These represent the data weights for significance, magnitude, and consistency features, respectively. Indicates gene Significant features after standardization; Indicates gene Amplitude characteristics after standardization; Indicates gene Consistency characteristics after standardization.
4. The gene abnormal regulation detection method based on regulatory pathway perturbation analysis according to claim 1, characterized in that, Based on the aforementioned method of weighting the normalized betweenness centrality variance and normalized degree centrality variance of genes within the pathway, the betweenness centrality weight and degree centrality weight within the pathway are obtained, including: The formula for calculating the weight of the normalized betweenness centrality of a gene is: ; in, Indicates pathway The weight of the normalized betweenness centrality in the equation; Indicates pathway The variance of the standardized betweenness centrality values of all genes within the group; Indicates pathway The variance of the normalized centrality values of all genes within the group; Represents a very small positive number; The formula for calculating the weights of gene normalization centrality is: ; in, Indicates pathway The weights of standardization centrality in the equation; Indicates pathway The variance of the standardized betweenness centrality values of all genes within the group; Indicates pathway The variance of the normalized centrality values of all genes within the group; It represents a very small positive number.
5. The gene abnormal regulation detection method based on regulatory pathway perturbation analysis according to claim 1, characterized in that, Based on the aforementioned fusion analysis of the normalized betweenness centrality assessment and normalized degree centrality assessment of genes using betweenness centrality weights and degree centrality weights within the pathway, a comprehensive topological importance assessment of genes within the pathway is obtained, specifically including: The formula for evaluating the overall topological importance is: ; in, Indicates gene In the pathway The overall topological importance score within the area; Indicates pathway The weight of the normalized betweenness centrality in the equation; Indicates gene In the pathway The normalized betweenness centrality value within; Indicates pathway The weights of standardization centrality in the equation; Indicates gene In the pathway The standardized centrality value within.
6. The gene abnormal regulation detection method based on regulatory pathway perturbation analysis according to claim 1, characterized in that, The assessment of local gene expression coordination is obtained by evaluating the correlation between the gene's expression profile vector and the expression profile vectors of all neighboring genes in the pathway, including: The formula for calculating the local expression coordination of a gene between its expression profile vector and the average expression profile vector of all its neighboring genes in the pathway is as follows: ; in, This refers to the assessment of the local expression coordination of genes, i.e., genes. In the pathway The expression spectral vector in the pathway and its relationship with the pathway The correlation coefficient between the average expression profile vectors of all neighboring genes in the sample; Indicates gene In the sample The spectral vector on the expression, Indicates gene In the pathway The number of neighboring genes; Indicates gene In the pathway The set of neighboring genes in; This represents the function for calculating the Pearson correlation coefficient.
7. The gene abnormal regulation detection method based on regulatory pathway perturbation analysis according to claim 1, characterized in that, Based on the comprehensive analysis of gene topological importance assessment and gene local expression coordination assessment within the pathway, contextual association optimization factors for genes within the pathway are obtained, including: The formula for calculating the contextual optimization factor of genes within a pathway is as follows: ; in, Indicates gene In the pathway Contextualization optimization factor; Indicates gene In the pathway The overall topological importance score within the area; This refers to the assessment of the local expression coordination of genes, i.e., genes. In the pathway The expression spectral vector in the pathway and its relationship with the pathway The correlation coefficient between the average expression profile vectors of all neighboring genes.
Citation Information
Patent Citations
Method for identifying gene pathways based on GSA
CN107203704A
Method of inferring gene pathway activity
CN108763862A