A pathogenic gene identification method based on expression abnormality and related device

By calculating the tail enrichment statistics and significance p-values ​​of genes in both tail directions, pathogenic genes are identified, solving the problem of low accuracy in pathogenic gene identification in existing technologies and achieving higher identification accuracy and information richness.

CN122369600APending Publication Date: 2026-07-10CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-04-15
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing gene expression differential analysis methods cannot effectively detect tail shifts and extreme value clustering when identifying pathogenic genes, resulting in low identification accuracy.

Method used

By calculating the tail enrichment statistics of the target gene in two tail directions, the optimal tail aggregation direction is determined. Combined with the significance p-value and the degree of tail offset, the abnormal gene expression characteristics of the pathogenic gene are identified, and the pathogenic gene is identified.

Benefits of technology

This improved the accuracy and information richness of pathogenic gene identification, enabling a thorough analysis of pathogenic genes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369600A_ABST
    Figure CN122369600A_ABST
Patent Text Reader

Abstract

This application relates to the field of gene recognition technology, and provides a method and related equipment for identifying pathogenic genes based on abnormal expression. The method includes: for each target gene, calculating tail enrichment statistics in two tail directions based on the grouping tag values ​​of all signal vectors to characterize the abnormal expression of the target gene at the low or high expression end; determining the optimal tail enrichment direction from the two tail directions based on the two tail enrichment statistics; calculating the tail boundary shift degree of the target gene according to the optimal tail enrichment direction; calculating the significance p-value of each target gene; identifying multiple pathogenic genes from all target genes based on all significant p-values; and performing pathogenicity analysis based on the tail boundary shift degree and optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification result. The method of this application can improve the accuracy of pathogenic gene identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene recognition technology, and in particular to a method and related equipment for identifying pathogenic genes based on abnormal expression. Background Technology

[0002] The development of complex diseases (such as neurodegenerative diseases) is often accompanied by abnormal changes in gene expression. Existing differential analysis methods for expression data (such as edgeR, DESeq2 package for differential analysis of count data, and limma, linear models for microarray data) are widely used in practical research. They mainly use statistical tests on genes through differences in mean between groups, variance structure, or linear model frameworks to screen candidate genes associated with disease states.

[0003] However, in real biological data, disease-related anomalies do not always manifest as a stable shift in the overall mean, but may instead appear as tail shifts, clustering of extreme values, or alterations in the tail structure of the distribution, occurring only in a subset of samples. Traditional methods often focus more on the overall distribution center or differences in population parameters in statistical modeling, thus lacking sensitivity to these "tail-enriched" signals. This leads to problems such as missed detections, decreased identification stability, or limited interpretability of results, thereby hindering the further development of pathogenic gene identification technology and resulting in low accuracy in pathogenic gene identification. Summary of the Invention

[0004] This application provides a method and related equipment for identifying pathogenic genes based on abnormal expression, which can solve the problem of low accuracy in identifying pathogenic genes.

[0005] In a first aspect, this application provides a method for identifying pathogenic genes based on abnormal expression, the method comprising: Obtain grouping tag values ​​for multiple signal vectors; each signal vector contains multiple target genes, and the grouping tag values ​​are used to describe whether the signal vector is pathogenic. For each target gene, the tail enrichment statistics of the target gene in two tail directions are calculated based on the grouping tag values ​​of all signaling vectors. Based on the two tail enrichment statistics, the optimal tail enrichment direction is determined from the two tail directions. The optimal tail aggregation direction is used to characterize the abnormal expression features of the target gene. Based on the optimal tail enrichment direction of each target gene, the degree of tail boundary shift of the target gene is calculated; the degree of tail boundary shift is used to describe the degree of tail boundary shift when the target gene is on pathogenic signal vectors and non-pathogenic signal vectors. Calculate the significance p-value for each target gene; the significance p-value is used to describe the degree of significance of the tail enrichment phenomenon of the target gene. Multiple pathogenic genes were identified from all target genes based on all significant p-values. Pathogenicity analysis was then performed based on the degree of tail boundary shift and the optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results.

[0006] Optionally, the tail enrichment statistics of the target gene in both tail directions are calculated based on the grouping tag values ​​of all signal vectors, including: For each tail direction, based on the grouped tag values ​​of all signal vectors, the cumulative tag deviation of the target gene at each sorting position in the tail direction is calculated, and all cumulative tag deviations are normalized to obtain the tail enrichment statistics of the target gene in the tail direction.

[0007] Optionally, based on the grouping tag values ​​of all signal vectors, calculate the cumulative tag deviation of the target gene at each sorting position in the tail direction, including: Through the formula:

[0008] Calculate the first The target gene is located at the tail end. Next sort position Cumulative deviation of labels at the location ; in, Indicates the first The target gene is located at the tail end. Next sort position At that location, the corresponding group tag value of the signal carrier. This represents the mean of the group labels. Indicates the direction of enrichment at the low expression level. Indicates the direction of high expression enrichment. Indicates the number of sorting positions. , This indicates the number of target genes.

[0009] Optionally, the cumulative deviation of all tags is normalized to obtain the tail enrichment statistics of the target gene in the tail direction, including: Through the formula:

[0010] Calculate the first The target gene is located at the tail end. Tail enrichment statistics ; in, This represents the normalization factor.

[0011] Optionally, based on two tail enrichment statistics, the optimal tail enrichment direction is determined from two tail directions, including: The tail direction corresponding to the one with the largest median of the two tail enrichment statistics is taken as the optimal tail enrichment direction for the target gene.

[0012] Optionally, based on the optimal tail enrichment direction of each target gene, the degree of tail boundary offset of the target gene is calculated, including: Through the formula:

[0013] Calculate the first The degree of tail boundary offset of each target gene ; in, Indicates the first Optimal tail enrichment direction for target genes enrichment direction for low expression At that time, the average expression value of the pathogenic signaling vector in the optimal tail enrichment direction, Indicates the first Optimal tail enrichment direction for target genes For high expression enrichment direction At that time, the average expression value of the pathogenic signaling vector in the optimal tail enrichment direction, Indicates the first The target gene has a low expression threshold value at the non-pathogenic signaling vector. Indicates the first The boundary value of the high expression end of the target gene in the non-pathogenic signaling vector. Indicates the direction of enrichment at the low expression level. The tail sample set of the pathogenic signaling vector. Indicates the direction of high expression enrichment. The tail sample set of the pathogenic signaling vector.

[0014] Optionally, calculate the significance p-value for each target gene, including: Through the formula:

[0015] Calculate the first The significance p-value of each target gene ; in, Indicates the number of label replacements. Indicates the number of target genes. This indicates the number of statistics that are not less than the observed statistic of the gene in the joint pooling empirical null distribution:

[0016] in, Indicates the first Under the condition of tag replacement, the first The permutation observation statistics for each target gene Indicates the first Tail enrichment statistics for the optimal tail enrichment direction of a target gene.

[0017] Optionally, multiple pathogenic genes can be identified from all target genes based on all significant p-values, including: For each target gene, determine whether the significance p-value of the target gene is less than the preset significance threshold. If so, the target gene is regarded as a disease gene.

[0018] Secondly, this application provides a pathogenic gene identification device based on abnormal expression, comprising: The acquisition module is used to acquire group tag values ​​of multiple signal vectors; each signal vector includes multiple target genes, and the group tag values ​​are used to describe whether the signal vector is pathogenic. The determination module is used to calculate the tail enrichment statistics of the target gene in two tail directions for each target gene based on the grouping tag values ​​of all signal vectors, and to determine the optimal tail enrichment direction from the two tail directions based on the two tail enrichment statistics; the optimal tail aggregation direction is used to characterize the abnormal expression features of the target gene. The first calculation module is used to calculate the degree of tail boundary offset of the target gene based on the optimal tail enrichment direction of each target gene; the degree of tail boundary offset is used to describe the degree of tail boundary offset when the target gene is on pathogenic signal vectors and non-pathogenic signal vectors. The second calculation module is used to calculate the significance p-value for each target gene; the significance p-value is used to describe the degree of significance of the tail enrichment phenomenon of the target gene. The identification module is used to identify multiple pathogenic genes from all target genes based on all significant p-values, and to perform pathogenicity analysis based on the degree of tail boundary offset and the optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results.

[0019] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described method for identifying pathogenic genes based on abnormal expression.

[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for identifying pathogenic genes based on abnormal expression.

[0021] The above-mentioned solution in this application has the following beneficial effects: In the embodiments of this application, grouping tag values ​​of multiple signal vectors are obtained. Then, for each target gene, tail enrichment statistics of the target gene in two tail directions are calculated based on the grouping tag values ​​of all signal vectors. Based on the two tail enrichment statistics, the optimal tail enrichment direction is determined from the two tail directions. Then, based on the optimal tail enrichment direction of each target gene, the tail boundary offset degree of the target gene is calculated. Then, the significance p-value of each target gene is calculated. Finally, based on all significance p-values, multiple pathogenic genes are determined from all target genes. Pathogenicity analysis is performed based on the tail boundary offset degree and the optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results. Specifically, the detection and sorting of target genes from two tail directions, rather than from a single direction, improves the identification ability and interpretability of tail enrichment information. The offset degree and significance p-value of the target gene are used to describe the offset degree and significance of the target gene, respectively, which improves the information richness of pathogenic gene identification, achieves a full analysis of pathogenic genes, and effectively improves the accuracy of pathogenic gene identification.

[0022] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating a method for identifying pathogenic genes based on abnormal expression, provided in an embodiment of this application; Figure 2 A schematic diagram of the structure of a pathogenic gene recognition device based on abnormal expression provided in an embodiment of this application; Figure 3This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0025] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0026] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0027] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0028] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0029] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0031] To address the issue of poor accuracy in existing pathogenic gene identification methods, this application provides a method for identifying pathogenic genes based on abnormal expression. This method acquires grouping tag values ​​from multiple signal vectors, then calculates tail enrichment statistics for each target gene in two tail directions based on the grouping tag values ​​of all signal vectors. Based on these statistics, the optimal tail enrichment direction is determined. The tail boundary shift of each target gene is then calculated according to its optimal tail enrichment direction. Finally, a significance p-value is calculated for each target gene. Based on all significance p-values, multiple pathogenic genes are identified from all target genes. Pathogenicity analysis is then performed based on the tail boundary shift and optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results. The method of detecting and sorting target genes from two tail directions, rather than from a single direction, improves the ability to identify tail enrichment information and the interpretability of results. The degree of shift and significance of target genes are expressed by boundary offset and significance p-value, respectively, which improves the information richness of pathogenic gene identification, enables full analysis of pathogenic genes, and effectively improves the accuracy of pathogenic gene identification.

[0032] The pathogenic gene identification method provided in this application will be described exemplarily below.

[0033] like Figure 1 As shown, the pathogenic gene identification method provided in this application includes the following steps: Step 11: Obtain the group tag values ​​of multiple signal carriers.

[0034] Each signal vector contains multiple target genes. The grouping tag value describes whether the signal vector is pathogenic; for example, if the signal vector is pathogenic, the grouping tag value is +1; if the signal vector is not pathogenic, the grouping tag value is -1. Target genes are the genes whose pathogenicity needs to be identified, and signal vectors are RNA, etc. At least one pathogenic signal vector must be present.

[0035] It should be noted that the selection of signal vectors can be based on the disease to be analyzed. For example, if the disease to be analyzed is a neurodegenerative disease, RNA from patients who do not have the disease or RNA from patients who have the disease can be selected as signal vectors. Then, the grouping label value of the signal vector can be determined according to whether the source of the signal vector is the disease.

[0036] For example, the `generateSyntheticData` function in `compcodeR` is used to generate baseline data for differentially expressed gene counts. The baseline data is a negative binomial distribution count, with 50 samples per group (50 samples each for the experimental group (pathogenic) and the control group (non-pathogenic). The total number of genes is set to 1000, the number of differentially expressed genes is set to 0, and the sequencing depth parameter is set to 1×10⁻⁶. 7 .

[0037] Based on the baseline data, perturbation was applied to construct tail-end enrichment simulation data. "Perturbation" refers to applying a directional transformation to a subset of tail-end samples selected according to a preset ratio in the experimental group without changing the expression distribution of the control group samples. This transformation causes the subset to cross the corresponding tail-end boundary of the control group in the low-end or high-end direction and form a tail-end enrichment pattern with controllable intensity. Specifically, when the tail-end direction is right, the maximum expression value of the control group samples is used as the high-expression boundary. The top k samples with the largest expression values ​​in the experimental group are selected as the tail-end sample subset. This subset is first translated to outside the boundary, and then scaled and adjusted using the boundary as a fulcrum so that the average expression value of the subset reaches the preset intensity. When the tail-end direction is left, the minimum expression value of the control group samples is used as the low-expression boundary. The top k samples with the smallest expression values ​​in the experimental group are selected as the tail-end sample subset, and the same translation and scaling methods are used to construct a tail-end enrichment pattern in the low-end direction. The size k of the tail sample subset is determined by the number of tail samples n (in this implementation, n is set to 50) and the tail sample proportion level p, i.e., k = n × p, where p is 30%, 40%, and 50%, respectively. The perturbation intensity level r of the tail enrichment perturbation is 0.1, 0.3, and 0.5. The perturbation is applied by scaling to make the mean of the tail sample subset shift by a multiple of 1 ± r relative to the boundary of the control group. Five sets of data are generated for each experimental design (combination of tail direction, tail sample proportion, and perturbation intensity) to obtain the gene expression data matrix and the corresponding sample group label information. The gene expression data matrix is ​​arranged with genes as rows and samples as columns. The matrix elements are the gene expression count values ​​in the samples (where the count values ​​corresponding to the tail samples are the counts after perturbation according to the aforementioned perturbation method). The sample group label information is used to identify the group to which each column of samples belongs, including the group labels of experimental group and control group. The experimental group is marked as +1, and the control group is marked as -1, and they correspond one-to-one with the column order of the gene expression data matrix.

[0038] For example: For a specific gene g, if its expression value sequence in 4 samples (denoted as sample 1 to sample 4 in column order) is ( The corresponding label vector is If the value is (-1, -1, +1, +1), then the established correspondence is: , where -1 represents the control group sample and +1 represents the experimental group sample.

[0039] Read the gene expression data matrix obtained in the above steps, determine the group to which each sample belongs according to the column order of the matrix, and establish the correspondence between sample columns and group labels; The experimental group samples were uniformly coded as +1, and the control group samples were uniformly coded as -1. The coding results are arranged sequentially according to the column order of the gene expression data matrix to form a label vector y, so that the length of the label vector y is consistent with the total number of samples and corresponds one-to-one with the samples in each column of the gene expression data matrix.

[0040] Under the above simulated data conditions, the number of samples in each group is 50, with 50 samples in the experimental group and 50 samples in the control group. Therefore, the length of the label vector y is 100. When the first 50 columns of the gene expression data matrix are the control group and the last 50 columns are the experimental group, the label vector y can be represented as an encoding vector with the first 50 elements being -1 and the last 50 elements being +1.

[0041] Step 12: For each target gene, calculate the tail enrichment statistics of the target gene in the two tail directions based on the grouping tag values ​​of all signal vectors, and determine the optimal tail enrichment direction from the two tail directions based on the two tail enrichment statistics.

[0042] The aforementioned tail enrichment statistics describe the degree of clustering of the target gene in the tail direction. A higher tail enrichment statistic indicates a higher degree of clustering of the target gene in the corresponding tail direction. The two tail directions are the low-expression enrichment direction and the high-expression enrichment direction. The low-expression enrichment direction represents the lowest expression end of the gene in the signal vector, and the high-expression enrichment direction represents the highest expression end of the gene in the signal vector. The optimal tail clustering direction is used to characterize the abnormal expression features of the target gene.

[0043] In some embodiments of this application, the steps of calculating the tail enrichment statistics of the target gene in two tail directions based on the grouping tag values ​​of all signal vectors, and determining the optimal tail enrichment direction from the two tail directions based on the two tail enrichment statistics, include: The first step is to calculate the cumulative deviation of the target gene at each sorting position in the tail direction for each tail direction based on the grouping tag values ​​of all signal vectors, and then normalize all the cumulative deviations of the tags to obtain the tail enrichment statistics of the target gene in the tail direction.

[0044] Specifically, through the formula:

[0045] Calculate the first The target gene is located at the tail end. Next sort position Cumulative deviation of labels at the location .

[0046] in, Indicates the first The target gene is located at the tail end. Next sort position At that location, the corresponding group tag value of the signal carrier. This represents the mean of the group labels. Indicates the direction of enrichment at the low expression level. Indicates the direction of high expression enrichment. Indicates the number of sorting positions. , This indicates the number of target genes.

[0047] Through the formula:

[0048] Calculate the first The target gene is located at the tail end. Tail enrichment statistics .

[0049] in, This represents the normalization factor.

[0050] For example, cumulative label deviation is used to characterize the cumulative shift trend of experimental group labels relative to the population mean during the ranking process. When there are tied values ​​in the sample expression values, a secondary ranking rule that is independent of the sample grouping and repeatable is used to break up the tied values, so as to reduce the influence of the original sample arrangement order on the calculation results of the tail enrichment statistic and improve the stability and robustness of the statistic calculation.

[0051] The second step is to determine the optimal tail enrichment direction from the two tail enrichment statistics.

[0052] Specifically, the tail direction corresponding to the one with the largest median of the two tail enrichment statistics is taken as the optimal tail enrichment direction for the target gene.

[0053] Step 13: Calculate the tail boundary offset of the target gene based on the optimal tail enrichment direction of each target gene.

[0054] The above-mentioned tail boundary offset is used to describe the degree of tail boundary offset when the target gene is on pathogenic signal vectors and non-pathogenic signal vectors. The larger the value, the greater the offset.

[0055] Specifically, through the formula:

[0056] Calculate the first The degree of tail boundary offset of each target gene .

[0057] in, Indicates the first Optimal tail enrichment direction for target genes enrichment direction for low expression At that time, the average expression value of the pathogenic signaling vector in the optimal tail enrichment direction, Indicates the first Optimal tail enrichment direction for target genes For high expression enrichment direction At that time, the average expression value of the pathogenic signaling vector in the optimal tail enrichment direction, Indicates the first The target gene has a low expression threshold value at the non-pathogenic signaling vector. Indicates the first The boundary value of the high expression end of the target gene in the non-pathogenic signaling vector. Indicates the direction of enrichment at the low expression level. The tail sample set of the pathogenic signaling vector. Indicates the direction of high expression enrichment. The tail sample set of the pathogenic signaling vector.

[0058] For example, the meaning of the above formula is as follows: When the optimal tail enrichment direction is the low-end direction, the minimum expression value of the control group sample is used as the tail boundary of the control group, and samples in the experimental group that are smaller than the tail boundary of the control group are selected as the low-end tail sample set; when the low-end tail sample set is not empty, its average expression value is calculated. When the optimal tail enrichment direction is the high-end direction, the maximum value of the expression value of the control group sample is used as the tail boundary of the control group. Samples in the experimental group that are greater than the tail boundary of the control group are selected as the high-end tail sample set. When the high-end tail sample set is not empty, its average expression value is calculated. The degree of tail boundary offset is calculated based on the average expression value of the tail samples in the experimental group and the tail boundary value of the control group. It is preferred to use the logarithmic ratio with constant offset for calculation. In this application, the offset constant is 1. When there are no experimental group tail samples located outside the tail boundary of the control group in the corresponding optimal direction, the degree of tail boundary offset of the gene is recorded as 0.

[0059] Step 14: Calculate the significance p-value for each target gene.

[0060] The above significance p-value is used to describe the significance of the tail enrichment phenomenon of the target gene; the larger the value, the greater the significance.

[0061] Specifically, through the formula:

[0062] Calculate the first The significance p-value of each target gene .

[0063] in, Indicates the number of label replacements. Indicates the number of target genes. This indicates the number of statistics that are not less than the observed statistic of the gene in the joint pooling empirical null distribution:

[0064] in, Indicates the first Under the condition of tag replacement, the first The permutation observation statistics for each target gene Indicates the first Tail enrichment statistics for the optimal tail enrichment direction of each target gene. 1 This is an indicator function; it takes the value 1 if the condition inside the parentheses is true, and 0 if it is false.

[0065] For example, before performing this step, the tag vector is permuted multiple times while keeping the gene expression data matrix unchanged: Without changing the expression values ​​of each gene sample in the gene expression data matrix, the label vector is randomly permuted multiple times to generate multiple sets of permuted label vectors; For each set of permutation label vectors, the calculation process of the bidirectional tail enrichment statistics is repeated to obtain the permutation statistics of each gene in the low-end and high-end directions. Compare the bidirectional permutation statistics and take the larger one as the permutation observation statistics of the corresponding gene under that permutation condition; The permutation statistics of each gene under multiple permutation conditions are summarized to form the set of permutation statistics required for the subsequent construction of the empirical zero distribution.

[0066] It should be noted that after obtaining the significance p-value, the significance p-value can be corrected by Benjamini-Hochberg multiple test to obtain the corrected significance p-value for each gene.

[0067] Step 15: Based on all significant p-values, multiple pathogenic genes are identified from all target genes. Pathogenicity analysis is then performed based on the degree of tail boundary shift and the optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results.

[0068] The above pathogenic gene identification results are used to describe the significance determination of pathogenic genes, the main concentration direction of tail enrichment, and the offset characteristics relative to the tail boundary of non-pathogenic signal vectors.

[0069] Specifically, for each target gene, the significance p-value of the target gene is determined to be less than a preset significance threshold. If so, the target gene is identified as a pathogenic gene. Otherwise, the target gene is not labeled. Then, based on the optimal tail enrichment direction and tail boundary offset of all pathogenic genes, the results of pathogenic gene identification are interpreted, auxiliary labels are added, and ranking references are established to obtain the pathogenic gene identification results.

[0070] For example, the process of interpreting the results, adding auxiliary markers, and providing sorting references for pathogenic genes based on the degree of tail boundary offset and the optimal tail enrichment direction of all pathogenic genes is as follows: For each identified pathogenic gene, the expression abnormality type of the gene is first interpreted based on its optimal tail enrichment direction. When the optimal tail enrichment direction is the low expression end enrichment direction, it indicates that the gene mainly exhibits low expression tail aggregation in the pathogenic signal vector, and it can be marked as a low expression abnormal pathogenic gene. When the optimal tail enrichment direction is the high expression end enrichment direction, it indicates that the gene mainly exhibits high expression tail aggregation in the pathogenic signal vector, and it can be marked as a high expression abnormal pathogenic gene.

[0071] Furthermore, the intensity of abnormal expression is interpreted based on the degree of tail boundary offset corresponding to the pathogenic gene; the greater the degree of tail boundary offset, the more significant the deviation of the abnormal tail of the gene in the pathogenic signal vector from the reference boundary of the non-pathogenic signal vector, and the more prominent its abnormal expression characteristics. Based on the optimal tail enrichment direction and the degree of tail boundary offset, directional markers and offset intensity markers can be added to each pathogenic gene to distinguish and classify different types of pathogenic genes.

[0072] When multiple pathogenic genes simultaneously meet the significance criteria and their significance p-values ​​are close or the same, the degree of tail boundary shift can be further used as a sorting reference. Pathogenic genes with a larger degree of tail boundary shift can obtain a higher sorting priority, so as to make the sorting results more interpretable and better reflect the candidate priority of pathogenic genes.

[0073] It is worth mentioning that the target gene is detected and sorted from two tail directions, rather than from a single direction, which improves the ability to identify tail enrichment information and the interpretability of the results. The degree of deviation and significance of the target gene are expressed by the degree of boundary shift and the p-value, respectively, which improves the information richness of pathogenic gene identification, enables full analysis of pathogenic genes, and effectively improves the accuracy of pathogenic gene identification.

[0074] The method of this application will be illustrated below with a specific example.

[0075] The invention was evaluated on a constructed tail-enrichment simulation dataset, and the traditional differential analysis methods edgeR, DESeq2 and limma were selected for comparison.

[0076] Tail enrichment simulation dataset experimental analysis The proposed method and the comparative method were run on a total of 90 simulated datasets constructed using compcodeR. The area under the receiver operating characteristic curve (AUROC) and the area under the precision-recall curve (AUPRC) for each method on each simulated dataset were calculated, and the results for all simulated datasets were averaged. The overall evaluation results are shown in Table 1.

[0077] Table 1

[0078] As shown in Table 1, the average AUROC of this application on the simulated dataset is 0.987 and the average AUPRC is 0.981, which outperforms traditional difference analysis methods such as edgeR, DESeq2 and limma under experimental conditions.

[0079] To further illustrate the advantages of the present invention relative to the control method, the method with the highest AUPRC among the three control methods (edgeR, DESeq2 and limma) was selected as the optimal control method on each simulated dataset. The improvement of the present application relative to the optimal control method was statistically analyzed, and the results are shown in Table 2.

[0080] Table 2

[0081] As shown in Table 2, in 90 simulated datasets, the method of this application has a higher AUROC on 97.8% of the datasets and a higher AUPRC on 96.7% of the datasets. The average improvement ΔAUROC relative to the best control method is 0.343 and the average improvement ΔAUPRC is 0.329, indicating that the method of this application has good stability and universality under different tail direction, tail sample ratio and perturbation intensity conditions.

[0082] In summary, the method of this application can effectively improve the stability and accuracy of pathogenic gene identification.

[0083] The following is an exemplary description of the pathogenic gene identification device based on abnormal expression provided in this application.

[0084] like Figure 2 As shown, this application provides a pathogenic gene identification device 200, which includes: The acquisition module 201 is used to acquire group tag values ​​of multiple signal vectors; each signal vector includes multiple target genes, and the group tag values ​​are used to describe whether the signal vector is pathogenic. The determination module 202 is used to calculate the tail enrichment statistics of the target gene in two tail directions for each target gene based on the grouping tag values ​​of all signal vectors, and to determine the optimal tail enrichment direction from the two tail directions based on the two tail enrichment statistics; the optimal tail aggregation direction is used to characterize the abnormal expression features of the target gene. The first calculation module 203 is used to calculate the degree of tail boundary offset of the target gene according to the optimal tail enrichment direction of each target gene; the degree of tail boundary offset is used to describe the degree of tail boundary offset when the target gene is on pathogenic signal vectors and non-pathogenic signal vectors. The second calculation module 204 is used to calculate the significance p-value of each target gene; the significance p-value is used to describe the degree of significance of the tail enrichment phenomenon of the target gene. The determination module 205 is used to identify multiple pathogenic genes from all target genes based on all significance p-values, and to perform pathogenicity analysis based on the degree of tail boundary offset and the optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results.

[0085] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0086] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0087] like Figure 3 As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.

[0088] Specifically, when the processor D100 executes the computer program D102, it acquires the grouping tag values ​​of multiple signal vectors, and then, for each target gene, calculates the tail enrichment statistics of the target gene in two tail directions based on the grouping tag values ​​of all signal vectors. Based on the two tail enrichment statistics, it determines the optimal tail enrichment direction from the two tail directions. Then, based on the optimal tail enrichment direction of each target gene, it calculates the tail boundary offset degree of the target gene, and then calculates the significance p-value of each target gene. Finally, based on all the significance p-values, it identifies multiple pathogenic genes from all target genes, and performs pathogenicity analysis based on the tail boundary offset degree and the optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results. The method of detecting and sorting target genes from two tail directions, rather than from a single direction, improves the ability to identify tail enrichment information and the interpretability of results. The degree of shift and significance of target genes are expressed by boundary offset and significance p-value, respectively, which improves the information richness of pathogenic gene identification, enables full analysis of pathogenic genes, and effectively improves the accuracy of pathogenic gene identification.

[0089] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0090] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0091] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0092] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0093] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a device / terminal device for identifying pathogenic genes based on abnormal expression, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0094] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0095] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0096] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention.

Claims

1. A method for identifying pathogenic genes based on abnormal expression, characterized in that, include: Obtain grouping tag values ​​for multiple signal vectors; each signal vector includes multiple target genes, and the grouping tag values ​​are used to describe whether the signal vector is pathogenic; For each target gene, the tail enrichment statistics of the target gene in two tail directions are calculated based on the grouping tag values ​​of all signaling vectors. Based on the two tail enrichment statistics, the optimal tail enrichment direction is determined from the two tail directions. The optimal tail aggregation direction is used to characterize the abnormal expression features of the target gene. Based on the optimal tail enrichment direction of each target gene, the degree of tail boundary offset of the target gene is calculated; the degree of tail boundary offset is used to describe the degree of tail boundary offset when the target gene is on pathogenic signal vectors and non-pathogenic signal vectors. Calculate the significance p-value for each target gene; the significance p-value is used to describe the degree of significance of the tail enrichment phenomenon of the target gene; Multiple pathogenic genes were identified from all target genes based on all significant p-values. Pathogenicity analysis was then performed based on the degree of tail boundary shift and the optimal tail enrichment direction of all pathogenic genes to obtain the pathogenic gene identification results.

2. The method for identifying pathogenic genes according to claim 1, characterized in that, The calculation of the tail enrichment statistics of the target gene in both tail directions based on the grouping tag values ​​of all signal vectors includes: For each tail direction, based on the grouping tag values ​​of all signal vectors, the cumulative tag deviation of the target gene at each sorting position in the tail direction is calculated, and all cumulative tag deviations are normalized to obtain the tail enrichment statistics of the target gene in the tail direction.

3. The method for identifying pathogenic genes according to claim 2, characterized in that, The calculation of the cumulative label deviation of the target gene at each sorting position in the tail direction based on the grouping label values ​​of all signal vectors includes: Through the formula: Calculate the first The target gene is located at the tail end. Next sort position Cumulative deviation of labels at the location ; in, Indicates the first The target gene is located at the tail end. Next sort position At that location, the corresponding group tag value of the signal carrier. This represents the mean of the group labels. Indicates the direction of enrichment at the low expression level. Indicates the direction of high expression enrichment. Indicates the number of sorting positions. , This indicates the number of target genes.

4. The method for identifying pathogenic genes according to claim 3, characterized in that, The normalization of the cumulative deviation of all tags to obtain the tail enrichment statistic of the target gene in the tail direction includes: Through the formula: Calculate the first The target gene is located at the tail end. Tail enrichment statistics ; in, This represents the normalization factor.

5. The method for identifying pathogenic genes according to claim 1, characterized in that, The determination of the optimal tail enrichment direction from two tail directions based on two tail enrichment statistics includes: The tail direction corresponding to the one with the largest value among the two tail enrichment statistics is taken as the optimal tail enrichment direction of the target gene.

6. The method for identifying pathogenic genes according to claim 1, characterized in that, The step of calculating the tail boundary offset of the target gene based on the optimal tail enrichment direction of each target gene includes: Through the formula: Calculate the first The degree of tail boundary offset of each target gene ; in, Indicates the first Optimal tail enrichment direction for target genes enrichment direction for low expression At that time, the average expression value of the pathogenic signaling vector in the optimal tail enrichment direction, Indicates the first Optimal tail enrichment direction for target genes For high expression enrichment direction At that time, the average expression value of the pathogenic signaling vector in the optimal tail enrichment direction, Indicates the first The target gene has a low expression threshold value at the non-pathogenic signaling vector. Indicates the first The boundary value of the high expression end of the target gene in the non-pathogenic signaling vector. Indicates the direction of enrichment at the low expression level. The tail sample set of the pathogenic signaling vector. Indicates the direction of high expression enrichment. The tail sample set of the pathogenic signaling vector.

7. The method for identifying pathogenic genes according to claim 1, characterized in that, The calculation of the significance p-value for each target gene includes: Through the formula: Calculate the first The significance p-value of each target gene ; in, Indicates the number of label replacements. Indicates the number of target genes. This indicates the number of statistics that are not less than the observed statistic of the gene in the joint pooling empirical null distribution: in, Indicates the first Under the condition of tag replacement, the first The permutation observation statistics for each target gene Indicates the first Tail enrichment statistics for the optimal tail enrichment direction of a target gene.

8. The method for identifying pathogenic genes according to claim 1, characterized in that, The identification of multiple pathogenic genes from all target genes based on all significant p-values ​​includes: For each target gene, determine whether the significance p-value of the target gene is less than a preset significance threshold. If so, the target gene is regarded as a pathogenic gene.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for identifying pathogenic genes based on abnormal expression as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for identifying pathogenic genes based on abnormal expression as described in any one of claims 1 to 8.