A rapid detection method for pre-disease status based on gene expression rank

Through the method of gene expression rank, key gene expression data are used to screen key genes and set thresholds to identify pre-morbidity status, solving the complexity and batch effect problems of pre-morbidity status recognition in the prior art, and achieving rapid and accurate pre-morbidity status detection.

CN116312787BActive Publication Date: 2025-08-12BEIJING BAOYING NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310241537.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-13
Publication Date
2025-08-12
Estimated Expiration
2043-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to avoid batch effects when identifying pre-disease states and the calculation is complex, so it is impossible to efficiently identify gene expression changes before the disease occurs.

Method used

Using a method based on gene expression rank, aberrant changes in gene expression rank are calculated by personalizing time sequence gene expression data, key genes are screened out and thresholds are set to identify pre-disease status of the disease.

Benefits of technology

It realizes fast and accurate pre-morbidity status detection, simplifies the calculation process, reduces experimental batch effects and errors, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312787B_ABST
    Figure CN116312787B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for rapid detection of pre-disease states based on gene expression ranks, comprising the following steps: (1) obtaining time-series gene expression data of all genes at each time point, screening out normal data, and calculating baseline expression data of a single individual; (2) sorting the baseline expression data to obtain baseline expression ranks; (3) sorting all time-series gene expression data to obtain the expression rank of each gene at each time point, and calculating the change score of the expression ranks of all genes; (4) screening the gene expression ranks and calculating the individual expression rank change score; (5) determining a threshold for the abnormal score; and (6) identifying the pre-disease state based on the individual rank change score and the threshold. The present invention is based on the relationship between changes in gene expression ranking and the occurrence and development of the disease, has the advantages of simple and rapid calculation, and can also eliminate experimental batch effects and errors to a certain extent based on the ranking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics and computational biology, and in particular to a method for rapidly detecting a pre-disease state based on gene expression ranking and applying abnormal changes in gene expression ranking. Background Art

[0002] According to bifurcation theory, the onset and progression of a disease can experience sudden deterioration. Therefore, the onset and progression of a disease can be divided into three stages: normal, premorbid, and disease. The normal stage refers to the period before the disease develops or when the disease is stable; the disease state refers to the period when the disease develops and rapidly deteriorates; and the premorbid state refers to a critical state where the disease is about to develop, but the entire body system is still relatively normal. Therefore, identifying premorbid states will greatly promote the development of modern medicine. Based on the critical slowing theory of dynamic systems, gene expression in the premorbid state undergoes dramatic changes during the onset and progression of a disease. However, because the onset and progression of a disease often result from the interaction of multiple abnormal genes and biomolecules, the molecular causes of the same disease vary greatly from individual to individual. Furthermore, identifying premorbid states requires observing molecular data from an individual's temporal progression, which can be subject to batch effects during testing. Currently, the mainstream approach uses the variance of individual gene expression changes and changes in co-expression relationships between genes to identify premorbid states. However, these methods suffer from two major issues: first, batch effects cannot be avoided using variance, and second, the calculation of variance and co-expression coefficients is relatively complex.

[0003] To address two issues with existing methods, this present invention uses time-series gene expression data to identify pre-disease states based on abnormal changes in gene expression rankings at each time point. Furthermore, because this invention only uses gene expression rankings with abnormal changes related to pre-disease states, the calculation is simple, convenient, and rapid. Summary of the Invention

[0004] In response to the above-mentioned deficiencies in the existing technology, the present invention provides a method for rapid detection of pre-disease states based on gene expression rank, which uses personalized time-series gene expression data to detect pre-disease states. While ensuring the accuracy of pre-disease state detection, it is more in line with the critical slowing-down theory of dynamic system changes, greatly shortening the calculation time.

[0005] The purpose of the present invention is achieved by adopting the following technical solutions:

[0006] A method for rapid detection of pre-disease state based on gene expression rank, comprising the following steps:

[0007] (1) Obtain the temporal gene expression data of all genes at each time point, filter out normal data, and calculate the baseline expression data of a single individual;

[0008] (2) Sort the baseline expression data of a single individual to obtain the baseline expression rank of the single individual;

[0009] (3) Sort the temporal gene expression data of all genes at each time point to obtain the expression rank of each gene at each time point, and calculate the change score of the expression rank of all genes at each time point;

[0010] (4) Screen the gene expression rank at each time point according to the individual regulation change score, and calculate the individual expression rank change score at each time point;

[0011] (5) Determine the threshold of the anomaly score;

[0012] (6) Identify the pre-disease state based on the individual expression rank change score and threshold at each time point.

[0013] Preferably, in step (1), the normal data is the data of the first four time points in the time-series gene expression data, and the benchmark expression data is the average expression value of the expression value at each time point in the normal data.

[0014] Preferably, in step (2), the sorting method is to sort from large to small according to the benchmark expression data.

[0015] Preferably, in step (3), the sorting method is to sort the time-series gene expression data from large to small.

[0016] Preferably, in step (3), the calculation formula for the gene expression rank change score is as shown in formula (1):

[0017]

[0018] Where s(g i ,t) represents the gene g at time point t i The change score relative to the baseline expression rank, r(g i, t) and r(g i , t0) represent the gene g i The expression rank at time t and the benchmark data (t0), N represents the number of genes contained in the expression data.

[0019] Preferably, in step (4), the calculation formula for the individual regulation change score at each time point is as shown in formula (2):

[0020]

[0021] Where S(t) is the individual's regulation change score at time t, s(g i ,t) represents the gene g at time point t i The change score of expression rank relative to the baseline.

[0022] Preferably, in step (4), the screening condition is the 50 genes with the largest individual regulation change scores at each time point, and the individual expression rank change score is the accumulation of the change values of the expression ranks corresponding to the top 50 genes selected at each time point.

[0023] According to critical slowing theory and dynamic network marker theory, when a critical point occurs, only a small number of genes will experience significant fluctuations in expression. Therefore, in order to accurately identify system abnormalities, the gene expression rank change score of each individual at each time point is determined by the 50 genes with the largest changes at that time point.

[0024] Preferably, in step (5), the threshold is the largest score among the remaining rank change scores after removing a maximum value and a minimum value from the expression rank change scores of all individuals at normal time points.

[0025] Preferably, in step (6), the pre-disease state time point is the time point when the individual's expression rank change score exceeds the threshold for the first time.

[0026] Preferably, for all individuals, the gene expression rank change scores at all time points can be calculated, and abnormal score thresholds can be set based on these rank change scores to assist in identifying the time points of the individual's pre-disease state.

[0027] The beneficial effects of the present invention are:

[0028] 1. The solution of the present invention has good accuracy in identifying the pre-disease state.

[0029] 2. The calculation process of the solution of the present invention is simple and the calculation time is short.

[0030] 3. The present invention is based on the relationship between changes in gene expression ranking and the occurrence and development of diseases, and has the advantages of simple and fast calculation. At the same time, based on the ranking, it can eliminate experimental batch effects and errors to a certain extent.

[0031] 4. The present invention is based on gene expression ranking and can be widely used to quickly identify pre-disease states with abnormal changes in gene expression ranking. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a flow chart of a method for rapid detection of pre-disease status based on gene expression rank according to an embodiment of the present invention;

[0033] Figure 2 This is a graph showing the results of expressing rank change scores on personalized influenza data according to an embodiment of the present invention;

[0034] Figure 3 This is a line graph of the change scores of the symptomatic group (Sx) and the asymptomatic group (Asx) at different time points in an embodiment of the present invention. DETAILED DESCRIPTION

[0035] The following will provide a clear and complete description of the concept, specific structure and technical effects of the present invention in conjunction with the embodiments and drawings to fully understand the purpose, scheme and effects of the present invention.

[0036] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are provided for ease of description only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted based on the understanding of those skilled in the art.

[0037] Example

[0038] The following describes it in detail with reference to specific embodiments.

[0039] Example 1

[0040] A method for rapid detection of pre-disease state based on gene expression rank, comprising the following steps:

[0041] (1) Obtain the temporal gene expression data of all genes at each time point in the influenza dataset in the GEO database, select the first four time points (Baseline, 0h, 5h, 12h) in the temporal gene expression data as the data of the time points in the normal state, and take the average expression value of the gene at these four time points as the baseline expression value to obtain the baseline data;

[0042] Taking the influenza dataset GSE30550 in the GEO database as an example, the specific information of the dataset is shown in Table 1.

[0043] Table 1. Sample quantity information of single-sample time series influenza dataset

[0044]

[0045] (2) Sort the baseline expression data of a single individual from large to small to obtain the baseline expression rank of a single individual;

[0046] (3) Sort the temporal gene expression data of all genes at each time point from large to small to obtain the expression rank of each gene at each time point. According to the formula Calculate the change scores of all gene expression ranks at each time point;

[0047] (4) The 50 genes with the largest expression rank changes at each time point were selected for the gene expression rank at each time point. The expression rank change values corresponding to the top 50 genes at each time point of each individual were accumulated to obtain the expression rank change score, which is:

[0048] (5) Determine the maximum score among the remaining rank change scores after removing the maximum and minimum values from the expression rank change scores of all individuals at normal time points as the threshold;

[0049] (6) Based on the rank change score and threshold of each time point of the individual, the expression rank change score based on the first four time points of all symptomatic samples and all time points of all asymptomatic samples was identified, and the threshold of the abnormal score was determined to be 14.2. The time point when the threshold was exceeded for the first time was the time when the pre-morbid state appeared.

[0050] like Figure 1 FIG. 1 is a flow chart of a method for rapid detection of pre-disease state based on gene expression rank according to the present invention, comprising:

[0051] S101. Calculate the rank of gene expression at each time point;

[0052] S102. Calculate the change score of gene expression rank at each time point;

[0053] S103. Calculate the expression rank change score of each individual at each time point;

[0054] S104. Identify the pre-disease status of an individual based on the individual expression rank change score.

[0055] like Figure 2 The figure below shows the expression rank change score results for personalized influenza data. For all individuals in the influenza dataset, the rapid pre-morbidity detection method was used to calculate the expression rank change scores. Based on the expression rank change scores for the first four time points of all symptomatic samples and all time points of all asymptomatic samples, an abnormal score threshold of 14.2 was determined. The time point when the threshold was first exceeded was the time of pre-morbidity. The results show that this method is very effective in identifying pre-morbidity.

[0056] like Figure 3 As shown, it is a line graph of the change scores of the symptomatic group (Sx) and the asymptomatic group (Asx) at different time points. It can be seen that the temporal gene expression data of all individuals at the initial four time points (Baseline, 0h, 5h, 12h) are all normal.

[0057] The present invention is based on gene expression rank and can be widely used to quickly identify pre-disease states with abnormal changes in gene expression ranking.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for rapid detection of pre-disease state based on gene expression rank, characterized in that: The following steps are involved: (1) Obtain the temporal gene expression data of all genes at each time point, filter out normal data, and calculate the baseline expression data of a single individual; (2) Sort the baseline expression data of a single individual to obtain the baseline expression rank of the single individual; (3) Sort the temporal gene expression data of all genes at each time point to obtain the expression rank of each gene at each time point, and calculate the change score of the expression rank of all genes at each time point; (4) Screen the gene expression rank at each time point according to the regulation change score, and calculate the expression rank change score of each individual at each time point; (5) Determine the threshold of the anomaly score; (6) Identify the pre-disease state based on the expression rank change score and threshold at each time point of the individual; In step (4), the calculation formula for the individual regulation change score at each time point is shown in formula (2): (2); Where S(t) is the individual's regulation change score at time t, s(g i ,t) represents the gene g at time point t i change score relative to baseline expression rank; In step (4), the screening condition is the 50 genes with the largest regulation change scores, and the individual expression rank change score is the accumulation of the expression rank change scores corresponding to the top 50 genes selected at each time point; In step (6), the pre-disease state time point is the time point when the individual expression rank change score exceeds the threshold for the first time.

2. The method for rapid detection of pre-disease state based on gene expression rank according to claim 1, characterized in that: In step (1), the normal data are the data of the first four time points in the time series gene expression data, and the benchmark expression data are the average expression value of the expression value at each time point in the normal data.

3. The method for rapid detection of pre-disease state based on gene expression rank according to claim 1, characterized in that: In the step (2), the sorting method is to sort from large to small according to the benchmark expression data.

4. The method for rapid detection of pre-disease state based on gene expression rank according to claim 1, characterized in that: In the step (3), the sorting method is to sort the time-series gene expression data from large to small.

5. The method for rapid detection of pre-disease state based on gene expression rank according to claim 1, characterized in that: In step (3), the calculation formula of the gene expression rank change score is shown in formula (1): (1); Where s(g i ,t) represents the gene g at time point t i The change score relative to the baseline expression rank, r(g i , t) and r(g i , t0) represent the gene g i The expression rank at time t and the benchmark data (t0), N represents the number of genes contained in the expression data.

6. The method for rapid detection of pre-disease state based on gene expression rank according to claim 1, characterized in that: In step (5), the threshold is the largest score among the remaining rank change scores after removing a maximum value and a minimum value from the expression rank change scores of all individuals at normal time points.

7. The method for rapid detection of pre-disease state based on gene expression rank according to claim 1, characterized in that: For all individuals, the gene expression rank change scores at all time points can be calculated, and abnormal score thresholds can be set based on these rank change scores to assist in identifying the time points of individual pre-disease states.

Citation Information

Patent Citations

  • Biliary atresia potential molecular subtype and recognition method of core gene of biliary atresia potential molecular subtype

    CN114783515A

  • Genetic screening for improving treatment of patients diagnosed with depression

    US20060160119A1