Method, device, electronic device and readable storage medium for detecting pre-disease state
By constructing dynamic network biomarkers and applying causal emergence theory, the problems of insufficient stability and accuracy in single-sample detection are solved, achieving stable and accurate identification of pre-disease states and providing technical support for early intervention in complex diseases.
Patent Information
- Application Number
- CN202411579586.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-07
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-07
AI Technical Summary
In existing technologies, single-sample detection methods based on specific states cannot effectively utilize control samples or other state samples, resulting in insufficient stability and accuracy in pre-disease state detection.
By obtaining dynamic network markers related to the target disease state, including the regulatory network between transcription factors and messenger RNA, the effective information value is calculated using causal emergence theory, and a gene regulatory network is constructed to determine whether the disease state is a pre-disease state.
It achieves more stable and accurate pre-disease state identification and can utilize sample information from multiple disease states to improve the robustness and accuracy of detection.
Smart Images

Figure CN119541637B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of analytical testing, and more specifically, to a method, apparatus, electronic device, and readable storage medium for detecting a pre-disease state. Background Art
[0002] Complex diseases result from the combined effects of genetic and environmental factors, leading to imbalances within the body. Numerous clinical studies have found that the development of complex diseases does not follow a gradual pattern, but rather a mutational pattern. That is, a relatively stable state (e.g., the normal state) rapidly transitions to another relatively stable state (the disease state) after passing a critical point (e.g., a pre-disease state). Therefore, the evolution of complex diseases can be divided into three states: a relatively stable normal state, a pre-disease state characterized by decreased recovery ability and increased susceptibility, and a relatively stable disease state. The pre-disease state is reversible; with good early treatment, it will recover to a normal state; otherwise, it will worsen into a disease state.
[0003] Currently, pre-disease states can be detected using a single sample in a specific state via a dynamic network biomarker (DNB).
[0004] However, using only a single sample in a specific state means that when detecting pre-disease states, it is impossible to refer to control samples or samples in other states, which leads to insufficient stability and accuracy in detecting pre-disease states. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the prior art by providing a method, apparatus, electronic device, and readable storage medium for detecting pre-disease states, which can determine whether a disease state is a pre-disease state by valid information values corresponding to dynamic network markers associated with multiple disease states.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] In a first aspect, the present invention provides a method for detecting a pre-disease state, comprising: acquiring dynamic network biomarkers related to the disease state of a target disease, wherein the dynamic network biomarkers include the size of the regulatory network between transcription factors and messenger RNA in each disease state, and the disease state is used to represent the development stage of the target disease; acquiring effective information values of the dynamic network biomarkers corresponding to each disease state, and determining whether the disease state is a pre-disease state of the target disease based on the effective information values.
[0008] In some implementations, acquiring dynamic network biomarkers related to the disease state of the target disease includes: acquiring first sample data related to the disease state of the target disease, the first sample data including expression profile data of transcription factors and expression profile data of messenger RNA; identifying the regulatory relationship between transcription factors and messenger RNA in each disease state based on the first sample data; obtaining a state-specific network between transcription factors and messenger RNA based on the identified multiple regulatory relationships, the state-specific network being used to represent the regulatory relationship between transcription factors and messenger RNA corresponding to each disease state; and acquiring dynamic network biomarkers based on the state-specific network.
[0009] In some implementations, the first sample data includes disease states corresponding to transcription factors and messenger RNAs. Based on the first sample data, identifying the regulatory relationship between transcription factors and messenger RNAs in each disease state includes: obtaining second sample data from the first sample data, where the disease states corresponding to transcription factors and messenger RNAs have been removed; obtaining a first correlation matrix and a second correlation matrix between the first and second sample data, where the first correlation matrix represents the strength of the correlation between transcription factors and messenger RNAs before and after the removal of disease states; obtaining a third correlation matrix between transcription factors and messenger RNAs and each disease state based on the first and second correlation matrices, where the third correlation matrix represents the strength of the correlation between each disease state and transcription factors and messenger RNAs; and determining, based on the third correlation matrix, whether a regulatory relationship exists between transcription factors and messenger RNAs in each disease state.
[0010] In some implementations, determining whether a regulatory relationship exists between transcription factors and messenger RNA in each disease state based on a third correlation matrix includes: normalizing the third correlation matrix to obtain a fourth correlation matrix, which represents the normalized correlation strength between each disease state and transcription factors and messenger RNA; and determining whether a regulatory relationship exists between transcription factors and messenger RNA in the corresponding disease state for the normalized correlation strength that meets preset conditions in the fourth correlation matrix.
[0011] In some implementations, the preset condition includes: the significance value corresponding to the normalized correlation strength is less than a preset threshold.
[0012] In some implementations, obtaining dynamic network biomarkers based on state-specific networks includes: determining whether transcription factors and messenger RNAs in the state-specific network are target transcription factors or target messenger RNAs using the cumulative probability of a Poisson distribution; extracting the regulatory network between transcription factors and messenger RNAs for each disease state from the state-specific network based on the target transcription factors and target messenger RNAs; and obtaining dynamic network biomarkers based on the size of the regulatory network.
[0013] In some implementations, obtaining the effective information values of dynamic network markers corresponding to each disease state and determining whether the disease state is a pre-disease state of the target disease based on the effective information values includes: calculating the effective information values of the regulatory network between transcription factors and messenger RNA corresponding to each disease state using causal emergence theory based on the dynamic network markers; obtaining the normalized efficiency parameters of the regulatory network between transcription factors and messenger RNA based on the effective information values; and determining the disease state corresponding to the maximum value among multiple efficiency parameters as the pre-disease state of the target disease.
[0014] In the first aspect, whether a disease state is a pre-disease state is determined by the effective information values corresponding to dynamic network markers associated with multiple disease states. By utilizing samples from other states, a gene regulatory network is constructed, and the effective information values from the theory of causal emergence are used as the standard for judging whether a disease state is a pre-disease state, thus achieving a more stable and accurate identification of pre-disease states.
[0015] Secondly, the present invention provides a device for detecting a pre-disease state, characterized by: an acquisition module for acquiring dynamic network biomarkers related to the disease state of a target disease, wherein the dynamic network biomarkers include a regulatory network between transcription factors and messenger RNA in each disease state, and the disease state is used to represent the development stage of the target disease; and a determination module for acquiring valid information values of the dynamic network biomarkers corresponding to each disease state, and determining whether the disease state is a pre-disease state of the target disease based on the valid information values.
[0016] In some implementations, the acquisition module is specifically used to acquire first sample data related to the disease state of the target disease, the first sample data including expression profile data of transcription factors and expression profile data of messenger RNA; based on the first sample data, identify the regulatory relationship between transcription factors and messenger RNA in each disease state; based on the identified multiple regulatory relationships, obtain a state-specific network between transcription factors and messenger RNA, the state-specific network being used to represent the regulatory relationship between transcription factors and messenger RNA corresponding to each disease state; and acquire dynamic network markers based on the state-specific network.
[0017] In some implementations, the first sample data includes disease states corresponding to transcription factors and messenger RNAs; the acquisition module is specifically used to acquire second sample data based on the first sample data, wherein the disease states corresponding to transcription factors and messenger RNAs have been removed from the second sample data; acquire a first correlation matrix and a second correlation matrix between the first sample data and the second sample data, wherein the first correlation matrix represents the strength of the correlation between transcription factors and messenger RNAs before the removal of disease states, and the second correlation matrix represents the strength of the correlation between transcription factors and messenger RNAs after the removal of disease states; acquire a third correlation matrix between transcription factors and messenger RNAs and each disease state based on the first and second correlation matrices, wherein the third correlation matrix represents the strength of the correlation between each disease state and transcription factors and messenger RNAs; and determine, based on the third correlation matrix, whether a regulatory relationship exists between transcription factors and messenger RNAs in each disease state.
[0018] In some implementations, the acquisition module is specifically used to normalize the third correlation matrix to obtain a fourth correlation matrix. The fourth correlation matrix is used to represent the normalized correlation strength between each disease state and transcription factors and messenger RNA. It is determined that in the fourth correlation matrix, the transcription factors and messenger RNA corresponding to the normalized correlation strength that meet the preset conditions have a regulatory relationship in the corresponding disease state.
[0019] In some implementations, the preset condition includes: the significance value corresponding to the normalized correlation strength is less than a preset threshold.
[0020] In some implementations, obtaining dynamic network biomarkers based on state-specific networks includes: determining whether transcription factors and messenger RNAs in the state-specific network are target transcription factors or target messenger RNAs using the cumulative probability of a Poisson distribution; extracting the regulatory network between transcription factors and messenger RNAs for each disease state from the state-specific network based on the target transcription factors and target messenger RNAs; and obtaining dynamic network biomarkers based on the size of the regulatory network.
[0021] In some implementations, the determination module is specifically used to calculate the effective information value of the regulatory network between transcription factors and messenger RNA corresponding to each disease state based on dynamic network markers and causal emergence theory; obtain the normalized efficiency parameter of the regulatory network between transcription factors and messenger RNA based on the effective information value; and determine the disease state corresponding to the maximum value among multiple efficiency parameters as the pre-disease state of the target disease.
[0022] Thirdly, the present invention provides an electronic device including a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the method provided in the first aspect.
[0023] Fourthly, the present invention provides a computer-readable storage medium on which a computer program is stored, and the computer program is executed by a processor to perform the method provided in the first aspect.
[0024] The beneficial effects of the second to fourth aspects can be referred to the first aspect, and will not be elaborated here. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A flowchart illustrating a method for detecting a pre-disease state provided in an embodiment of this application is shown;
[0027] Figure 2 This illustration shows a flowchart of a method for detecting a pre-disease state according to an embodiment of this application, which obtains dynamic network markers related to the disease state of a target disease. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0029] Complex diseases result from the combined effects of genetic and environmental factors, leading to imbalances within the body. Numerous clinical studies have found that the development of complex diseases does not follow a gradual pattern, but rather a mutational pattern. That is, a relatively stable state (e.g., the normal state) rapidly transitions to another relatively stable state (the disease state) after passing a critical point (e.g., a pre-disease state). Therefore, the evolution of complex diseases can be divided into three states: a relatively stable normal state, a pre-disease state characterized by decreased recovery ability and increased susceptibility, and a relatively stable disease state. The pre-disease state is reversible; with good early treatment, it will return to the normal state; otherwise, it will worsen into the disease state. Therefore, identifying the critical point (i.e., the pre-disease state) in the evolution of complex diseases is of significant scientific importance for their prevention and treatment.
[0030] However, accurately detecting pre-disease states or critical points in complex diseases is a challenging task. This is because changes at the molecular level (e.g., molecular biomarkers) of different phenotypes are often subtle or even negligible before entering a disease state. With ongoing research, researchers have discovered that dynamic network biomarkers (DNBs) are helpful in detecting pre-disease states and offer significant advantages compared to traditional biomarkers (molecular biomarkers and network biomarkers).
[0031] Previous research methods have primarily relied on specific state samples, failing to fully utilize information from control samples or other state samples to construct gene regulatory networks. Consequently, the robustness and accuracy of the constructed gene regulatory networks are limited. Furthermore, gene regulatory networks are inherently causal regulatory networks. However, at the network level, existing research methods have not considered the causal emergence of dynamic networks. Therefore, there is an urgent need to develop novel detection methods applicable to both single-sample and multi-sample transcriptome data. These methods should be able to simultaneously predict gene regulatory networks at both single-sample and multi-sample levels, and consider the causal emergence of gene regulatory networks, ultimately contributing to the accurate detection of pre-disease states in complex diseases.
[0032] Based on this, embodiments of the present invention provide a method for detecting a pre-disease state, comprising: acquiring dynamic network biomarkers related to the disease state of a target disease, wherein the dynamic network biomarkers include the size of the regulatory network between transcription factors and messenger RNA in each disease state, and the disease state is used to represent the development stage of the target disease; acquiring effective information values of the dynamic network biomarkers corresponding to each disease state, and determining whether the disease state is a pre-disease state of the target disease based on the effective information values.
[0033] Whether a disease state is a pre-disease state is determined by valid information values corresponding to dynamic network markers associated with multiple disease states. By utilizing samples from other states, a gene regulatory network is constructed, and valid information values from causal emergence theory are used as the standard for judging whether a disease state is a pre-disease state, achieving a more stable and accurate identification of pre-disease states.
[0034] The following describes, with reference to the accompanying drawings, a method for detecting the scavenging effect of flavonoids on reactive oxygen species in living cells, as provided in the embodiments of the present invention.
[0035] Figure 1 A schematic flowchart of a method for detecting a pre-disease state provided in an embodiment of this application is shown. Figure 1 As shown, the method may include S101 and S102.
[0036] S101. Obtain dynamic network biomarkers associated with the disease state of the target disease. The dynamic network biomarkers include the size of the regulatory network between transcription factors and messenger RNA in each disease state. The disease state is used to represent the developmental stage of the target disease.
[0037] Figure 2 This illustration shows a flowchart of a method for detecting a pre-disease state according to an embodiment of this application, which obtains dynamic network markers related to the disease state of a target disease.
[0038] In some implementations, reference Figure 2 To obtain dynamic network biomarkers related to the disease state of the target disease, including:
[0039] S1011. Obtain first sample data related to the disease state of the target disease. The first sample data includes expression profile data of transcription factors and expression profile data of messenger RNA.
[0040] In some implementations, the first sample data may be obtained by preprocessing transcriptome data for a given target disease. For example, the given transcriptome data may be subjected to logarithmic transformation, removal of repetitive genes, and preservation of protein coding. Then, according to prior annotation information, the processed transcriptome data is segmented into transcription factor (TF) expression profile data and messenger RNA (mRNA) expression profile data.
[0041] As an example, the expression profile data corresponding to TF can be represented by the following formula:
[0042]
[0043] The expression profile data corresponding to mRNA can be represented by the following formula:
[0044]
[0045] Where s represents the number of samples, q represents the number of TFs in each sample, and p represents the number of mRNAs in each sample. Assume that s samples of TF and mRNA transcriptome data have m disease states k, and these m disease states k can be recorded as k∈{s1,s2,…,s} k ,…s m}
[0046] S1012. Based on the first sample data, identify the regulatory relationship between transcription factors and messenger RNA in each disease state.
[0047] S1013. Based on the identified multiple regulatory relationships, a state-specific network between transcription factors and messenger RNA is obtained. The state-specific network is used to represent the regulatory relationship between transcription factors and messenger RNA corresponding to each disease state.
[0048] In some implementations, the first sample data includes disease states corresponding to transcription factors and messenger RNAs. Based on the first sample data, identifying the regulatory relationship between transcription factors and messenger RNAs in each disease state includes: obtaining second sample data from the first sample data, where the disease states corresponding to transcription factors and messenger RNAs have been removed; obtaining a first correlation matrix and a second correlation matrix between the first and second sample data, where the first correlation matrix represents the strength of the correlation between transcription factors and messenger RNAs before and after the removal of disease states; obtaining a third correlation matrix between transcription factors and messenger RNAs and each disease state based on the first and second correlation matrices, where the third correlation matrix represents the strength of the correlation between each disease state and transcription factors and messenger RNAs; and determining, based on the third correlation matrix, whether a regulatory relationship exists between transcription factors and messenger RNAs in each disease state.
[0049] In some implementations, obtaining the first and second correlation matrices between transcription factors and messenger RNA between the first and second sample data can be achieved using network identification methods. For example, the network identification method may include at least one of the following three proportionality analysis methods: φ(Phit), φ... s (Phis) and ρ p (Rhop).
[0050] In some implementations, ratio analysis is used to calculate a metric between two given variables. For example, φ(Phit) and φ... s (Phis) can use distance metrics as a metric, ρ p (Rhop) can use similarity metrics as a measure.
[0051] In some implementations, it is assumed that the two given variables are V i and V j The three proportion analysis methods are defined as follows:
[0052]
[0053]
[0054] Where var is the variance function. Similar to the Pearson correlation coefficient, ρ p The metric is a similarity metric and is naturally symmetric, with a value range of [-1, 1]. φ and φ s The indicator is a distance indicator, with a value range of [0,∞].
[0055] As an example, the first correlation matrix X (k) This can be expressed by the following formula:
[0056]
[0057] Second correlation matrix Y (k) This can be expressed by the following formula:
[0058]
[0059] in, and The values represent the correlation strength between TF(j) and mRNA(i) before and after removing state k samples, respectively. and The larger the absolute value, the stronger the correlation between TF(j) and mRNA(i).
[0060] In some implementations, the correlation strength U between TF(j) and mRNA(i) for state k is... (k) (That is, the strength of the correlation between each disease state k and transcription factors and messenger RNA) can be represented by a third correlation matrix, for example, the third correlation matrix can be represented by the following formula:
[0061]
[0062] In some implementations, based on the third correlation matrix, to determine whether a regulatory relationship exists between transcription factors and messenger RNA in each disease state, the third correlation matrix can be normalized to obtain a fourth correlation matrix. This fourth correlation matrix represents the normalized correlation strength between each disease state and the transcription factors and messenger RNA. The fourth correlation matrix determines whether a regulatory relationship exists between the transcription factors and messenger RNA in the corresponding disease state, corresponding to normalized correlation strengths that meet preset conditions. As an example, the preset conditions may include: the significance value corresponding to the normalized correlation strength is less than a preset threshold.
[0063] In some implementations... The correlation basically follows a normal distribution. As an example, the normalized correlation strength... This can be expressed by the following formula:
[0064]
[0065] Where, μ (k) and σ (k) Representing the third correlation matrix U (k) The mean and standard deviation of the fourth correlation matrix Z. (k) This can be expressed by the following formula:
[0066]
[0067] Among them, the strength of each normalized correlation Corresponding to a significance p-value, The corresponding significance p-value It can be calculated using the following formula:
[0068]
[0069] in, express The absolute value of a random number from a standard normal distribution is calculated using the pnorm() function. The probability p-value. The smaller the value, the more likely a regulatory relationship exists between TF(j) and mRNA(i) in disease state k. As an example, the preset threshold for p could be 0.05.
[0070] In some implementations, the regulatory relationships between transcription factors and messenger RNA corresponding to each identified disease state can be integrated to obtain a state-specific network between transcription factors and messenger RNA.
[0071] S1014. Obtain dynamic network markers based on state-specific networks.
[0072] In some implementations, obtaining dynamic network biomarkers based on state-specific networks includes: determining whether transcription factors and messenger RNAs in the state-specific network are target transcription factors or target messenger RNAs using the cumulative probability of a Poisson distribution; extracting the regulatory network between transcription factors and messenger RNAs for each disease state from the state-specific network based on the target transcription factors and target messenger RNAs; and obtaining dynamic network biomarkers based on the size of the regulatory network.
[0073] In some implementations, the target transcription factor and target messenger RNA can be hub TFs and hub mRNAs in a state-specific network. Whether a TF or mRNA is a hub can be assessed using the cumulative probability of a Poisson distribution, as shown by the following formula:
[0074]
[0075] Where λ=np, n represents the number of genes (including TF and mRNA), t represents the number of predicted TF-mRNA regulatory relationships, and C represents the number of all possible TF-mRNA regulatory relationships. A smaller p-value for a gene indicates that the gene is more likely to be a hub. As an example, genes with p-values less than 0.05 can be identified as hubs.
[0076] In some implementations, the regulatory network between transcription factors and messenger RNA for each disease state can be extracted from a state-specific network based on the target transcription factor and target messenger RNA, which can be achieved through the following methods:
[0077] For a state-specific network of m disease state indicators, only the pivotal TFs and mRNAs that simultaneously constitute all m state-specific networks are retained. These pivotal TFs and mRNAs are conserved pivotal TFs and mRNAs across the m state-specific networks. Based on these conserved pivotal TFs and mRNAs, the TF-mRNA regulatory network for each state is extracted (the size of each regulatory network is used as a marker of dynamic network activity). For each state's TF-mRNA regulatory network, the number of TFs and mRNAs is the same, but the TF-mRNA regulatory relationships differ.
[0078] S102. Obtain the effective information values of the dynamic network markers corresponding to each disease state, and determine whether the disease state is a pre-disease state of the target disease based on the effective information values.
[0079] In some implementations, effective information values of the regulatory network between transcription factors and messenger RNA corresponding to each disease state can be calculated based on dynamic network markers using causal emergence theory; normalized efficiency parameters of the regulatory network between transcription factors and messenger RNA can be obtained based on the effective information values; and the disease state corresponding to the maximum value among multiple efficiency parameters can be determined as the pre-disease state of the target disease.
[0080] In some implementations, the effective information (EI) value of the regulatory network between transcription factors and messenger RNA corresponding to each disease state is calculated based on dynamic network markers using causal emergence theory. This can be achieved using the following formula:
[0081] EI = H( <W i out >)- <H(W i out (Formula 10)
[0082] Among them, W i out This represents a control network, which includes multiple network nodes. The parameter in <> refers to the average value from node i to all nodes, and H refers to the entropy, i.e., H( <W i out >) represents the entropy of the average value from node i to all nodes in the control network. <H(W i out The value of EI represents the average entropy from node i to all nodes in the control network. EI varies with network size; therefore, the efficiency parameter can be a standardized metric, effectiveness, which can be used to characterize the causal emergence of the network. As an example, effectiveness can be obtained using the following formula:
[0083]
[0084] Effectiveness is the result of EI normalized by the logarithm of the number of network nodes n.
[0085] In some implementations, through network causal emergence analysis, the critical point or pre-disease state of the target disease often exhibits an Effectiveness peak. Therefore, the disease state corresponding to the maximum Effectiveness value is the critical point or pre-disease state of the target disease.
[0086] The method provided in this application will be further described below through two embodiments.
[0087] Example 1:
[0088] This embodiment uses the identification of precancerous liver disease states in humans as an example to illustrate the method provided in this application.
[0089] In this embodiment, the target disease is liver cancer, so the transcriptome data can be collected from the public database GEO (GeneExpression Omnibus, https: / / www.ncbi.nlm.nih.gov / geo / ) for human liver cancer caused by hepatitis C virus (dataset number GSE6764), and human TF prior information obtained from the CollecTRI database (https: / / github.com / saezlab / CollecTRI).
[0090] After preprocessing the transcriptome data (such as logarithmic transformation, removal of duplicate genes, and retention of protein-coding genes), expression profiles of 1138 TFs and 15,140 mRNAs from 75 liver cancer samples were obtained. Therefore, in this embodiment, and These 75 liver cancer samples included 13 control samples and 62 case samples. The 62 case samples were divided into 7 disease states: cirrhotic liver tissue (10 samples), low-grade dysplastic liver tissue (10 samples), high-grade dysplastic liver tissue (7 samples), very early HCC (8 samples), early HCC (10 samples), advanced HCC (7 samples), and very advanced HCC (10 samples).
[0091] According to the method provided in this application, a state-specific network for human liver cancer was obtained based on transcriptome data (R and T) of 1138 TFs and 15,140 mRNAs from 75 liver cancer samples. In this embodiment, three proportionality analysis methods (Phit, Phis, and Rhop) were used to calculate the strength of the relationship between TFs and mRNAs before and after removing human liver cancer state k samples. The significance p-threshold for the strength of the TF-mRNA relationship was set to 0.05 for each human liver cancer state. In each human liver cancer state k, all TF-mRNA regulatory relationships were fused to obtain seven state-specific networks for human liver cancer, and the number of gene regulatory relationships in human liver cancer in Example 1 is shown in Table 1.
[0092] Table 1
[0093]
[0094] In this embodiment, dynamic network biomarkers for seven human liver cancer states (cirrhotic liver tissue, poorly dysplastic liver tissue, highly dysplastic liver tissue, very early-stage liver cancer, early-stage liver cancer, intermediate-stage liver cancer, and advanced-stage liver cancer) were obtained using three proportionality analysis methods (Phit, Phis, and Rhop). For example, for the Phit method, the dynamic network biomarkers for the seven human liver cancer states were 39,292, 37,502, 38,622, 29,928, 30,893, 28,375, and 30,167, respectively. For the Phis method, the dynamic network biomarkers for the seven human liver cancer states were 53,819, 51,946, 61,331, 42,564, 41,720, 46,228, and 49,151, respectively. For the Rhop method, the dynamic network biomarkers for the seven human liver cancer states were 3137, 4198, 3876, 2719, 1810, 2473, and 1847, respectively. The causal emergence efficiency of these dynamic networks (i.e., TF-mRNA regulatory networks) is shown in Table 2.
[0095] Table 2
[0096]
[0097] Table 2 shows that for the Phit and Phis methods, very early-stage liver cancer was predicted as the critical state of human liver cancer; for the Rhop method, early-stage liver cancer was predicted as the critical state of human liver cancer. This result indicates that the dynamic networks constructed by most proportional analysis methods (including the Phit and Phis methods) predict the critical or pre-disease state of human liver cancer in complete agreement with the results of previous experimental observations.
[0098] Example 2:
[0099] This embodiment uses the identification of pre-lymphoma disease states in mice as an example to illustrate the method provided in this application.
[0100] In this embodiment, the target disease is lymphoma, and the transcriptome data can be mouse lymphoma transcriptome data collected from the public database GEO (GeneExpression Omnibus, https: / / www.ncbi.nlm.nih.gov / geo / ) (dataset number GSE6136), as well as mouse TF prior information obtained from the CollecTRI database (https: / / github.com / saezlab / CollecTRI).
[0101] After preprocessing the transcriptome data (such as logarithmic transformation, removal of duplicate genes and retention of protein-coding genes), expression profiles of 934 TFs and 12,545 mRNAs from 26 lymphoma samples were obtained.
[0102] Therefore, in this embodiment, and These 26 lymphoma samples were divided into 5 states: resting (5 samples), activated (3 samples), marginal lymphoma (6 samples), transitional lymphoma (5 samples), and aggressive lymphoma (7 samples).
[0103] According to the method provided in this application, mouse lymphoma state-specific networks were obtained based on transcriptome data (R and T) of 934 TFs and 12,545 mRNAs from 26 lymphoma samples. In this embodiment, three proportionality analysis methods (Phit, Phis, and Rhop) were used to calculate the strength of the TF-mRNA relationship before and after removing the mouse lymphoma state k sample. The significance p-threshold for the TF-mRNA relationship strength was set to 0.05 for each mouse lymphoma state. In each mouse lymphoma state k, all TF-mRNA regulatory relationships were fused to obtain five mouse lymphoma state-specific networks, and the number of mouse lymphoma state-specific gene regulatory relationships in Example 2 is shown in Table 3.
[0104] Table 3
[0105]
[0106] In this embodiment, dynamic network biomarkers for five mouse lymphoma states (resting, activated, marginal, transitional, and invasive) were obtained using three proportionality analysis methods (Phit, Phis, and Rhop). For the Phit method, the dynamic network biomarkers for the five mouse lymphoma states were 180,758, 239,734, 146,729, 205,938, and 189,554, respectively. For the Phis method, the dynamic network biomarkers for the five mouse lymphoma states were 218,673, 270,865, 201,766, 237,132, and 211,230, respectively. For the Rhop method, the dynamic network biomarkers for the five mouse lymphoma states were 26,861, 28,108, 26,534, 24,110, and 27,749, respectively. The causal emergence efficiency of these dynamic networks (i.e., TF-mRNA regulatory networks) is shown in Table 4.
[0107] Table 4
[0108]
[0109]
[0110] Table 4 shows that for the Phit, Phis, and Rhop methods, the lymphoma marginal state was predicted as the critical state of mouse lymphoma. This result indicates that the dynamic networks constructed by all proportional analysis methods (including the Phit, Phis, and Rhop methods) predict the critical or pre-disease state of mouse lymphoma in complete agreement with the results of previous experimental observations.
[0111] The results in Examples 1 and 2 show that the three ratio analysis methods can accurately predict the critical or pre-disease states of human liver cancer and mouse lymphoma. This result demonstrates that the method proposed in this application can effectively predict pre-disease states. In summary, the pre-disease state detection method proposed in this application can accurately detect disease critical or pre-disease states, providing technical support for the early intervention and treatment of complex human diseases, and has significant biological implications.
[0112] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A device for detecting pre-disease states, characterized by: An acquisition module is used to acquire dynamic network biomarkers related to the disease state of a target disease. The dynamic network biomarkers include a regulatory network between transcription factors and messenger RNA for each disease state, and the disease state is used to represent the developmental stage of the target disease. The determination module is used to obtain the effective information value of the dynamic network marker corresponding to each disease state, and determine whether the disease state is a pre-disease state of the target disease based on the effective information value. The determining module is specifically used to calculate the effective information value of the regulatory network between the transcription factors and messenger RNA corresponding to each disease state based on the dynamic network markers and the causal emergence theory; obtain the normalized efficiency parameter of the regulatory network between the transcription factors and messenger RNA based on the effective information value; and determine the disease state corresponding to the maximum value among the multiple efficiency parameters as the pre-disease state of the target disease. The acquisition module is specifically used to acquire first sample data related to the disease state of the target disease, the first sample data including expression profile data of transcription factors and expression profile data of messenger RNA; based on the first sample data, identify the regulatory relationship between the transcription factors and the messenger RNA in each disease state; based on the identified multiple regulatory relationships, obtain a state-specific network between the transcription factors and the messenger RNA, the state-specific network being used to represent the regulatory relationship between the transcription factors and the messenger RNA corresponding to each disease state; and acquire the dynamic network markers based on the state-specific network. The first sample data includes the disease states corresponding to the transcription factors and the messenger RNA; the acquisition module is specifically used to acquire second sample data based on the first sample data, wherein the disease states corresponding to the transcription factors and the messenger RNA have been removed from the second sample data; acquire a first correlation matrix and a second correlation matrix between the first sample data and the second sample data, wherein the first correlation matrix represents the correlation strength between the transcription factors and the messenger RNA before the removal of the disease states, and the second correlation matrix represents the correlation strength between the transcription factors and the messenger RNA after the removal of the disease states; acquire a third correlation matrix between the transcription factors and the messenger RNA and each disease state based on the first and second correlation matrices, wherein the third correlation matrix represents the correlation strength between each disease state and the transcription factors and the messenger RNA; and determine, based on the third correlation matrix, whether a regulatory relationship exists between the transcription factors and the messenger RNA in each disease state. The acquisition module is specifically used to normalize the third correlation matrix to obtain a fourth correlation matrix, which is used to represent the normalized correlation strength between each disease state and the transcription factor and the messenger RNA. In the fourth correlation matrix, it is determined that the transcription factor and the messenger RNA corresponding to the normalized correlation strength that meets the preset conditions have a regulatory relationship in the corresponding disease state. The acquisition module is further configured to determine whether the transcription factor and the messenger RNA in the state-specific network are target transcription factors or target messenger RNAs by using the cumulative probability of the Poisson distribution; extract the regulatory network between the transcription factor and the messenger RNA of each disease state from the state-specific network based on the target transcription factor and the target messenger RNA; and obtain the dynamic network marker based on the size of the regulatory network.
2. The apparatus according to claim 1, characterized in that, The preset conditions include: The significance value corresponding to the normalized correlation strength is less than a preset threshold.
3. An electronic device, characterized in that, The device includes a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus. The processor executes the machine-readable instructions to perform a method for detecting a pre-disease state, the method comprising: Acquire dynamic network biomarkers associated with disease states of a target disease, wherein the dynamic network biomarkers are used to represent the size of the regulatory network between transcription factors and messenger RNA in each disease state, and the disease state is used to represent the developmental stage of the target disease; Obtain the effective information value of the dynamic network marker corresponding to each disease state, and determine whether the disease state is a pre-disease state of the target disease based on the effective information value; The step of obtaining effective information values of the dynamic network markers corresponding to each disease state and determining whether the disease state is a pre-disease state of the target disease based on the effective information values includes: calculating effective information values of the regulatory network between the transcription factors and messenger RNA corresponding to each disease state using causal emergence theory based on the dynamic network markers; obtaining normalized efficiency parameters of the regulatory network between the transcription factors and messenger RNA based on the effective information values; and determining the disease state corresponding to the maximum value among multiple efficiency parameters as the pre-disease state of the target disease. The step of acquiring dynamic network biomarkers related to the disease state of the target disease includes: acquiring first sample data related to the disease state of the target disease, the first sample data including expression profile data of transcription factors and expression profile data of messenger RNA; identifying the regulatory relationship between the transcription factors and the messenger RNA in each disease state based on the first sample data; obtaining a state-specific network between the transcription factors and the messenger RNA based on the identified multiple regulatory relationships, the state-specific network being used to represent the regulatory relationship between the transcription factors and the messenger RNA corresponding to each disease state; and acquiring the dynamic network biomarkers based on the state-specific network. The first sample data includes disease states corresponding to the transcription factors and messenger RNAs. Identifying the regulatory relationship between the transcription factors and messenger RNAs in each disease state based on the first sample data includes: obtaining second sample data from the first sample data, wherein the disease states corresponding to the transcription factors and messenger RNAs have been removed from the second sample data; obtaining a first correlation matrix and a second correlation matrix between the transcription factors and messenger RNAs between the first and second sample data, wherein the first correlation matrix represents the strength of the correlation between the transcription factors and messenger RNAs before the removal of the disease states, and the second correlation matrix represents the strength of the correlation between the transcription factors and messenger RNAs after the removal of the disease states; obtaining a third correlation matrix between the transcription factors and messenger RNAs and each disease state based on the first and second correlation matrices, wherein the third correlation matrix represents the strength of the correlation between each disease state and the transcription factors and messenger RNAs; and determining whether a regulatory relationship exists between the transcription factors and messenger RNAs in each disease state based on the third correlation matrix. The step of determining whether a regulatory relationship exists between the transcription factor and the messenger RNA in each disease state based on the third correlation matrix includes: normalizing the third correlation matrix to obtain a fourth correlation matrix, which represents the normalized correlation strength between each disease state and the transcription factor and the messenger RNA; and determining whether a regulatory relationship exists between the transcription factor and the messenger RNA in the corresponding disease state for the normalized correlation strength that meets preset conditions in the fourth correlation matrix. The step of obtaining the dynamic network biomarker based on the state-specific network includes: determining whether the transcription factor and the messenger RNA in the state-specific network are target transcription factors or target messenger RNAs by using the cumulative probability of a Poisson distribution; extracting the regulatory network between the transcription factor and the messenger RNA for each disease state from the state-specific network based on the target transcription factor and the target messenger RNA; and obtaining the dynamic network biomarker based on the size of the regulatory network.
4. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium, and when the computer program is run by a processor, it performs the steps of the method for detecting pre-disease states provided in claim 3.
Citation Information
Patent Citations
Precancerous disease state detecting device and method
CN104794321A
Method and system for recognizing disease related factor on basis of functional module
CN106874706A